Possess real world experience operating, monitoring, and debugging infrastructure for large-scale AI/ML workloads
Have experience working in cloud compute infrastructure design, preferably GCP
Possess strong programming skills
Have significant experience working and deploying in Kubernetes at scale
Familiarity with the Nvidia GPU generations
Proven track record of building production observability and telemetry stacks
Nice to Have
Have a background in either ML SWE or infrastructure SRE work to build on
Conceptual understanding of ML workload efficiency paradigms
Have experience leading and delivering projects to multidisciplinary stakeholders
Familiarity with Google TPU generations
Familiarity with: workload scheduling; machine learning efficiency research; familiarity with ML-driven R&D cycles; familiarity with hardware benchmarking
What You'll Be Doing
Design, deploy, and scale robust observability systems and telemetry pipelines to monitor fleetwide compute efficiency, hardware health, and workload goodput across distributed clusters
Drive hardware efficiency and node reliability across our accelerator fleet, and integrating new hardware to leverage advancements
Identify compute waste and efficiency bottlenecks across the fleet, partnering with ML and platform teams to actively optimize accelerator utilization and improve workload goodput
Collaborate with teams in the AI org, and work closely with the ML Infrastructure team to identify, instrument, and improve canonical efficiency metrics across both training and inference runs
Contribute to the efforts for consistently improving the reliability of our ML runs.
Operate, maintain, and harden research, development, and production cloud infrastructure and cluster deployments
Partner and collaborate with a diverse set of teams incl. science, research, product, business development and operations
Contribute to core technical decisions (e.g. choice of tooling, infrastructure, and architectural design)
Perks and Benefits
Equal employment opportunities regardless of sex, race, religion or belief, ethnic or national origin, disability, age, citizenship, marital, domestic or civil partnership status, sexual orientation, gender identity, pregnancy or related condition (including breastfeeding) or any other basis protected by applicable law
Hybrid working model with the requirement to be able to come into the office 3 days a week