Senior Network Production Engineer, AI Supercomputing

Requirements

  • Significant experience operating large-scale datacenter or high-performance computing networks in production.
  • Deep understanding of Ethernet, RDMA, RoCEv2, congestion control, routing, load balancing, and lossless or near-lossless network design.
  • Experience debugging complex failures across switches, NICs, hosts, drivers, firmware, and distributed applications.
  • Strong Linux systems knowledge and proficiency in Python, Go, C++, or another systems-oriented programming language.
  • Experience building monitoring, automation, diagnostic, or remediation systems for production infrastructure.
  • Demonstrated ability to lead high-severity incidents and coordinate resolution across multiple engineering teams.

Nice to Haves

  • Experience operating GPU training clusters or other tightly coupled distributed-computing systems.
  • Knowledge of MRC or other multi-rail, multi-plane, multipath Ethernet transports.
  • Experience with NCCL, collective communications, CUDA, GPUDirect RDMA, and distributed training frameworks.
  • Familiarity with ConnectX-class NICs, modern Ethernet switch ASICs, SONiC, SAI, and switch telemetry.
  • Experience with InfiniBand and the operational differences between InfiniBand and Ethernet-based AI fabrics.

What You'll Be Doing

  • Own Production Network Health
  • Operate the Network at Frontier Scale
  • Improve Large-Job Reliability and Performance
  • Automate Fleet Operations
  • Drive Cross-Company Execution

Perks and Benefits

  • Microsoft is an equal opportunity employer
  • All qualified applicants will receive consideration for employment without regard to various factors
  • Find additional benefits and pay information here
  • Open position for a minimum of 5 days with ongoing applications until filled
AI Summary ✨

Similar jobs

  • 2 hours agoNew
  • 9 hours agoNew
  • AI Reliability Engineer, HPC

    Microsoft·Greater London, UK

    a day agoNew
  • See all jobs in UK →