Company Logo
Software Engineer

Netflix - 1d ago

Company Logo
Senior Software Engineer

Reddit - 4d ago

Senior Site Reliability Engineer, DGX Cloud

Requirements:

  • BS in Computer Science or related technical field, or equivalent experience
  • 10+ years of experience operating production services
  • Expert-level knowledge of Kubernetes administration, containerization, and microservices architecture
  • Experience with infrastructure automation tools (e.g., Terraform, Ansible, Chef, Puppet)
  • Proficiency in at least one high-level programming language (e.g., Python, Go)
  • In-depth knowledge of Linux operating systems, networking fundamentals (TCP/IP), and cloud security standards
  • Proficient knowledge of SRE principles, encompassing SLOs, SLIs, error budgets, and incident handling
  • Experience building and operating comprehensive observability stacks using tools like OpenTelemetry, Prometheus, Grafana, ELK Stack, Lightstep, Splunk, etc.

What you'll be doing:

  • Build, implement and support operational and reliability aspects of large-scale Kubernetes clusters
  • Define SLOs/SLIs, monitor error budgets, and streamline reporting
  • Support services before launch and maintain live services
  • Operate and optimize GPU workloads across various cloud platforms
  • Scale systems sustainably through automation and push for reliability improvements
  • Lead triage and root-cause analysis of incidents, practice blameless postmortems
  • Participate in on-call rotation to support production services

Ways to stand out from the crowd:

  • Operating GPU-accelerated clusters with KubeVirt in production
  • Applying generative-AI techniques to reduce operational toil
  • Experience with workflow orchestration platforms such as Temporal, Cadence, Airflow, Argo Workflows, or Step Functions
AI Summary ✨
NVIDIA logo

NVIDIA

Switzerland

Experience: Senior
Posted: August 18, 2026
Last seen: 34 minutes ago
Aws
Azure
Gcp
Golang
Kubernetes
Python
Terraform
sitereliability

Why we track NVIDIA

NVIDIA has become one of the most important companies in tech thanks to AI and GPU computing. They have EU roles across several countries. If you're interested in hardware, CUDA, or ML infrastructure, they're hard to beat.

Similar jobs

  • proton logo

    Site Reliability Engineer

    Switzerland, France

    2 months ago
  • proton logo

    Site Reliability Engineer (Application Edge)

    UK, Czech Republic, Switzerland, France

    7 months ago
    Still looking
  • a year ago
    Still looking
    Remote
  • See all jobs in Switzerland