Requirements
- Cloud & Infrastructure as Code: Hands-on experience building and managing scalable cloud infrastructure on AWS using Infrastructure as Code (Terraform, CloudFormation, etc.)
- Software Development: Fluency in Python, Go, or a similar language for infrastructure automation, tooling, or backend development
- Databases: Deep experience operating and tuning enterprise databases at scale
- Containerization: Experience working with containerized environments and orchestration tools (e.g., Docker, ECS, Kubernetes)
- Systems Reliability & Observability: Solid understanding of Site Reliability Engineering (SRE) principles
Nice to Haves
- Observability & Telemetry with Datadog, Prometheus
- AI/ML Infrastructure support
- Cloud Cost Optimization
What You'll Be Doing
- Architecting, deploying, and managing low-latency, highly available storage solutions
- Ensuring sub-second end-user latency, near-zero-downtime reliability, and high throughput across storage engines
- Automating infrastructure provisioning and compute pipelines
- Right-sizing compute/storage resources and managing AWS infrastructure costs
- Operating and tuning enterprise databases at scale
- Working with containerized environments and orchestration tools
- Applying SRE principles in monitoring, alerting, health checks, and performance troubleshooting
Perks and Benefits
You will be at the ground floor of building the data-storage systems that powers our flagship products. You'll solve complex, high-scale infrastructure challenges at the intersection of distributed storage and compute.
This is a hybrid role requiring in-office presence two days per week.