AI Reliability Engineer, HPC

Requirements

  • Bachelor’s Degree in Computer Science, or related technical discipline AND 4+ years technical experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering OR equivalent experience

Nice to Haves

  • Master’s Degree in Computer Science, or related technical discipline AND 2+ years technical experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering OR equivalent experience
  • Experience with Kubernetes, Docker, container orchestration, and CI/CD pipelines for ML training or inference workloads
  • Experience with public cloud platforms such as Azure, AWS, or GCP, including infrastructure-as-code
  • Experience with monitoring and observability tools such as Grafana, Datadog, or OpenTelemetry
  • Programming or scripting experience in Python, Go, or Bash
  • Experience with distributed systems, networking, storage, and high-performance computing (HPC)
  • Experience operating GPU clusters and workload schedulers for ML/AI workloads
  • Experience with ML training or inference pipelines
  • Experience with capacity planning and cost optimization for GPU-based infrastructure

What You'll Be Doing

  • Reliability & Availability: Ensure uptime, resiliency, and fault tolerance of HPC clusters powering MAI model training and inference
  • Observability: Design and maintain monitoring, alerting, and logging systems to provide real-time visibility into all aspects of HPC systems including GPU, clusters, storage and networking
  • Automation & Tooling: Build automation for deployments, incident response, scaling, and failover in CPU+GPU environments
  • Incident Management: Lead on-call rotations, troubleshoot production issues, conduct blameless postmortems, and drive continuous improvements
  • Security & Compliance: Ensure data privacy, compliance, and secure operations across model training and serving environments
  • Collaboration: Partner with ML engineers and platform teams to improve developer experience and accelerate research-to-production workflows
  • Embody our Culture and Values

Perks and Benefits

  • Software Engineering IC5 - The typical base pay range for this role across the United Kingdom is £93,500.00 - £161,800.00 per year
  • Benefits and other compensation eligibility
  • Equal opportunity employer policy
AI Summary ✨

Similar jobs

  • 2 hours agoNew
  • 9 hours agoNew
  • a day agoNew
  • See all jobs in UK →