Netflix - 1d ago
Reddit - 4d ago
Experience building or operating large-scale training platforms
Worked with large scale compute clusters (GPUs)
Proven ability to debug performance and reliability issues across large distributed fleets
Strong problem-solving skills and ability to work independently
Strong communication skills and the ability to work effectively with both internal and external partners
Deep knowledge of modern cloud infrastructure including Kubernetes, Infrastructure as Code, AWS, and GCP
Experience with SLURM
Experience building or operating large-scale training platforms
Maintain research infrastructure, ensuring health, and optimizing components to extract peak performance from the system (both on application, and infrastructure side)
Scale infrastructure to meet growing research demands while maintaining reliability and performance
Collaborate with research teams to deeply understand their infrastructure needs, and design solutions that balance performance with cost efficiency.
Identify and resolve performance bottlenecks and capacity hotspots through deep analysis of distributed systems at scale.
Build and evolve telemetry and monitoring systems to provide deep visibility into infrastructure performance, utilization, and costs across our cloud and datacenter fleets.
Participate in on-call rotations and incident response to maintain system reliability
Base Annual Salary:
EU €100,000 - €230,000 + Equity
US $150,000 - $300,000 + Equity
This role is based in our Freiburg / San Francisco office. We operate a hybrid model and cover reasonable travel costs — relocation is encouraged but not required. We do expect a meaningful in-person presence, and we'll discuss what that looks like for your situation during the process.
Note: Our recruitment process uses AI-assisted tools to help manage and organize applications. All hiring decisions are always made by our team.

Senior Infrastructure Engineer (PostgreSQL DBA)
Berlin, Germany
Poland, Portugal, Ireland, Germany