Company Logo
Software Engineer

Netflix - 1d ago

Company Logo
Senior Software Engineer

Reddit - 4d ago

Senior Production Engineer - DGX Cloud

Requirements:

  • 8+ years of experience building or operating production services and large-scale distributed systems, including hands-on automation.
  • Strong programming skills in Python, Go, or a comparable language, with experience developing tools for production operations.
  • Experience with infrastructure as code, configuration management, or GitOps, and with building automation for repeatable service deployments and changes.
  • Strong knowledge of Linux, Kubernetes, containers, cloud infrastructure, distributed systems, and networking fundamentals; ability to diagnose failures in production.
  • Understanding of Production Engineering principles, including SLIs, SLOs, error budgets, incident response, and reducing operational toil.
  • Experience instrumenting services and using metrics, logs, and traces to understand system behavior and improve reliability.
  • Clear technical communication and ability to work across engineering teams.
  • BS/MS in Computer Science or equivalent practical experience.

Nice to Haves:

  • Familiarity with technologies such as vLLM, SGLang, PyTorch, TensorRT-LLM, NVIDIA Dynamo, CUDA, or NCCL, and with GPU performance analysis.
  • Experience building Kubernetes operators, controllers, workload orchestration services, fleet management systems, or self-healing automation.
  • Experience with Terraform, Argo CD, CI/CD, policy validation, or safe deployment and rollback systems.
  • Background with developing with AI tools and agents.
  • Experience with production AI inference or agentic workloads, including debugging issues across models, runtimes, Kubernetes, and hardware.

What you'll be doing:

  • Build and operate production software, automation, and tooling for control plane services, model deployments, and inference and agentic workloads across DGX Cloud environments.
  • Improve the reliability of inference and agentic platforms and services.
  • Improve endpoint availability, inference routing, capacity management, and service health.
  • Use infrastructure as code and GitOps to deploy, configure, validate, upgrade, and recover services consistently across environments.
  • Build workflows for service enablement, model releases, handoff, deprecation, and ongoing operations.
  • Define and instrument SLIs and SLOs for inference and control plane services.
  • Participate in on-call and incident response, troubleshoot failures, and collaborate with various engineering teams.

Perks and Benefits:

  • Trailblazing work in Artificial Intelligence, High-Performance Computing, and Visualization.
  • Opportunity to work with cutting-edge technologies and brilliant minds.
  • Creative, hard-working, and self-motivated environment.
AI Summary ✨

Similar jobs

  • 11 hours agoNew
  • sonarsource logo

    Cloud Platform Engineer

    Geneva, Switzerland

    13 hours agoNew
  • 2 days ago
  • See all jobs in Switzerland →