Company Logo
Software Engineer

Netflix - 1d ago

Company Logo
Senior Software Engineer

Reddit - 4d ago

Senior Machine Learning Engineer, LLM Inference Optimization

Requirements

  • Strong Python and PyTorch engineering skills
  • Hands-on experience deploying or optimizing LLM, VLM, or high-throughput transformer inference systems
  • Practical knowledge of at least one modern inference stack such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, or equivalent internal systems
  • Strong understanding of transformer inference bottlenecks, including KV cache, attention, memory bandwidth, batching, parallelism, and long-context serving
  • Ability to reason quantitatively about latency, throughput, quality, utilization, and cost tradeoffs
  • Strong communication skills and ability to collaborate with research, kernel, infrastructure, product, and customer teams

Nice to Haves

  • Experience with quantization-aware training, post-training quantization, FP8, INT8, INT4, NVFP4, MXFP4, AWQ, GPTQ, SmoothQuant, or related techniques
  • Experience with distillation, speculative decoding, EAGLE, Medusa, multi-token prediction, or other inference acceleration methods
  • Experience with agentic workloads, including tool calling, structured outputs, streaming APIs, high concurrency, and multi-step orchestration
  • CUDA or Triton familiarity, even if the role is not primarily a kernel-engineering role
  • Open-source contributions to vLLM, SGLang, TensorRT-LLM, FlashInfer, LMCache, PyTorch, Triton, Ray, KServe, or related projects

What You'll Be Doing

  • Own optimization work for specific model families, customer endpoints, or serving backends
  • Run engine comparisons and recommend practical serving configurations for specific workloads
  • Debug model quality or performance regressions during production rollouts
  • Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token
  • Deploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or similar systems
  • Build and productionize model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery
  • Implement or integrate speculative decoding, draft-model approaches, KV-cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving
  • Build reproducible benchmark harnesses for TTFT, TPOT, tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token
  • Partner with GPU kernel engineers and platform engineers to diagnose bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers
  • Write clear design docs, performance reports, rollout plans, and customer-facing technical explanations

Perks and Benefits

  • Competitive compensation
  • Career growth and learning opportunities
  • Flexibility and ownership
  • Collaborative and innovative culture
  • Opportunity to work on impactful AI projects
  • International environment and talented teams
AI Summary ✨

Similar jobs

  • 8 days agoRemote EMEA
  • 8 days agoRemote EMEA
  • 8 days agoRemote EMEA
  • See all jobs in undefined →