Company Logo
Software Engineer

Netflix - 1d ago

Company Logo
Senior Software Engineer

Reddit - 4d ago

Site Reliability Engineer

Requirements

  • A Master’s degree in Computer Science, Engineering, or a related field.
  • 7+ years of experience in a DevOps or SRE role, with strong expertise in cloud computing and distributed systems.
  • Hands-on experience with site reliability issues, including root cause analysis, in-production troubleshooting, and on-call rotations.
  • Proficiency in working with reliability KPIs, such as observability, alerting, and SLAs.
  • Experience with CI/CD, containerization, and orchestration tools like Docker and Kubernetes.
  • Knowledge of monitoring, logging, alerting, and observability tools such as Prometheus, Grafana, ELK Stack, or Datadog.
  • Familiarity with infrastructure-as-code tools like Terraform or CloudFormation.
  • Proficiency in scripting languages (Python, Go, Bash) and a strong understanding of software development best practices.
  • Solid grasp of networking, security, and system administration concepts.
  • Excellent problem-solving and communication skills, with the ability to work effectively in a collaborative environment.
  • Experience in an AI/ML environment, high-performance computing (HPC) systems, or modern AI-oriented solutions (e.g., Fluidstack, Coreweave, Vast) is a plus.

What You'll Be Doing

  • Design, build, and maintain scalable, highly available, and fault-tolerant infrastructures to support web services and ML workloads.
  • Ensure our platform, inference, and model training environments are always highly available and enable seamless replication across HPC clusters.
  • Operate systems and troubleshoot issues in production, including interrupts, on-call responses, and infrastructure scaling.
  • Implement and improve monitoring, alerting, and incident response systems to minimize downtime and optimize performance.
  • Develop and maintain workflows and tools for CI/CD, containerization, orchestration, monitoring, and logging.
  • Participate in on-call rotations to respond to incidents and perform root cause analysis.
  • Drive continuous improvement in infrastructure automation, deployment, and orchestration using tools like Kubernetes, Flux, and Terraform.
  • Collaborate with AI/ML researchers to enable safe and reproducible model-training experiments.
  • Build a cloud-agnostic platform that abstracts infrastructure complexities for science and engineering teams.
  • Design and develop new workflows, tooling, and automation to improve system reliability, availability, and performance.
  • Work with the security team to ensure infrastructure adheres to best practices and compliance requirements.
  • Document processes and procedures to ensure consistency and knowledge sharing across the team.

Perks and Benefits

  • We offer a comprehensive benefits package designed to support well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.
AI Summary ✨
Mistral AI logo

Mistral AI

Paris, France

Experience: Senior
Posted: September 12, 2026
Last seen: an hour ago
Docker
Golang
Kubernetes
Python
Terraform
sitereliability

Why we track Mistral AI

Mistral AI is a Paris-based AI company building frontier large language models. Founded by former DeepMind and Meta researchers, they're one of Europe's most important AI companies. If you want to work on cutting-edge AI research and infrastructure in Europe, Mistral is the one to watch.

Similar jobs

  • mistralai logo

    Site Reliability Engineer, Mistral Cloud

    France, Netherlands, Poland, Germany, UK

    8 days ago
  • 12 days ago
    Remote
  • a month ago
  • See all jobs in France