Company Logo
Software Engineer

Netflix - 1d ago

Company Logo
Senior Software Engineer

Reddit - 4d ago

Senior Machine Learning Engineer, LLM Inference Optimization

Requirements

  • Strong Python and PyTorch engineering skills
  • Hands-on experience deploying or optimizing LLM, VLM, or high-throughput transformer inference systems
  • Practical knowledge of at least one modern inference stack such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, or equivalent internal systems
  • Strong understanding of transformer inference bottlenecks, including KV cache, attention, memory bandwidth, batching, parallelism, and long-context serving
  • Ability to reason quantitatively about latency, throughput, quality, utilization, and cost tradeoffs
  • Strong communication skills and ability to collaborate with research, kernel, infrastructure, product, and customer teams

Nice to Haves

  • Experience with quantization-aware training, post-training quantization, FP8, INT8, INT4, NVFP4, MXFP4, AWQ, GPTQ, SmoothQuant, or related techniques
  • Experience with distillation, speculative decoding, EAGLE, Medusa, multi-token prediction, or other inference acceleration methods
  • Experience with agentic workloads, including tool calling, structured outputs, streaming APIs, high concurrency, and multi-step orchestration
  • CUDA or Triton familiarity, even if the role is not primarily a kernel-engineering role
  • Open-source contributions to vLLM, SGLang, TensorRT-LLM, FlashInfer, LMCache, PyTorch, Triton, Ray, KServe, or related projects

What You'll Be Doing

  • Own optimization work for specific model families, customer endpoints, or serving backends
  • Run engine comparisons and recommend practical serving configurations for specific workloads
  • Debug model quality or performance regressions during production rollouts
  • Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token
  • Deploy, configure, benchmark, and extend inference engines
  • Build and productionize model-compression workflows
  • Implement or integrate speculative decoding, draft-model approaches, KV-cache optimization, and more
  • Build reproducible benchmark harnesses
  • Partner with GPU kernel engineers and platform engineers to diagnose bottlenecks
  • Write clear design docs, performance reports, rollout plans, and customer-facing technical explanations

Perks & Benefits

  • Competitive compensation
  • Career growth and learning opportunities
  • Flexibility and ownership
  • Collaborative and innovative culture
  • Opportunity to work on impactful AI projects
  • International environment and talented teams
AI Summary ✨
Nebius logo

Nebius

London, UK

Experience: Senior
Posted: September 23, 2026
Last seen: an hour ago
Python
machinelearning

Why we track Nebius

Nebius is the Amsterdam-headquartered, Nasdaq-listed AI cloud that spun out of Yandex. They build GPU clusters, their own server hardware, and a full cloud stack for AI training and inference, with engineering concentrated in Amsterdam and teams in London, Prague, Berlin, Finland, and Paris plus remote roles across Europe. Pay is strong: levels.fyi shows a €140K median in the Netherlands (€160K at the G17 level) and a £183K median in London.

Similar jobs

  • 3 hours ago
    New
  • 3 hours ago
    New
  • 3 hours ago
    New
  • See all jobs in UK