Senior Applied Scientist, Efficient LLM Inference & Model Optimization
Senior Applied Scientist, Efficient LLM Inference & Model Optimization
Requirements
PhD in computer science, machine learning, ML systems, computer systems, computer architecture, electrical engineering, applied math, or a related field
Strong publication record or equivalent research artifacts in ML, ML systems, efficient inference, model compression, quantization, distillation, serving systems, or related areas
Strong hands-on coding ability in Python and PyTorch
Deep understanding of LLMs, VLMs, transformer inference, model compression, quantization, and production-serving tradeoffs
Strong experimental design skills
Excellent written and verbal communication
Nice to Haves
First-author publications in NeurIPS, ICML, ICLR, MLSys, ACL, EMNLP, ASPLOS, OSDI, SOSP, ISCA, HPCA, or comparable venues
Experience deploying ML models or inference optimizations in production
Experience with vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo, FlashAttention, FlashInfer, Triton, CUDA, or PyTorch internals
Experience with post-training methods like SFT, DPO, RLHF, RLAIF, preference optimization, or synthetic data generation
Open-source research artifacts, widely used benchmarks, technical blogs, or invited talks in efficient AI systems
What You'll Be Doing
Own focused research projects from hypothesis through production handoff
Prepare internal reports, technical blogs, or papers for external credibility
Partner with MLEs to convert research prototypes into production components
Define and execute research programs in efficient inference with measurable production impact
Invent and productionize methods for various optimization techniques
Build prototypes in PyTorch, Triton, CUDA-adjacent tooling, or inference-serving frameworks
Design rigorous evaluation methodology covering various aspects
Collaborate with teams to choose high-leverage research bets
Mentor engineers and scientists on experimental design and tradeoffs