Requirements:
- Bachelor's or Master's degree in Computer Science, Mathematics, Physics, a related quantitative field, or equivalent practical experience.
- 4 years of experience building, scaling, and debugging machine learning models using deep learning frameworks (e.g., JAX, PyTorch, or TensorFlow).
- Experience in at least one core area: Reinforcement Learning (RL), Post-Training, Agentic Tool-Use, or Inference-Time Search.
Nice to haves:
- PhD in Computer Science, Machine Learning, Physics, or a related quantitative field.
- Experience designing asynchronous agent-environment simulation loops or large distributed post-training pipelines.
- Experience prototyping new hypotheses quickly while maintaining clean, robust, and production-grade shared codebases.
What you'll be doing:
- Operate across the full research-and-engineering lifecycle of frontier reasoning and agentic systems.
- Bridge research to production by solving unsolved problems in agentic reasoning.
- Architect and optimize distributed post-training pipelines and agent-environment simulation loops.
- Design experiments and failure analyses to isolate performance bottlenecks.
- Maintain high code quality and architectural health across shared Reinforcement Learning and modeling codebases.
Perks and benefits:
US: $207,000 - $300,000 (USD) + 20% bonus target + equity + benefits
Opportunity to share preferred working locations - London, UK; Mountain View, CA, USA; New York, NY, USA
Learn more about benefits at Google.