4+ years of professional experience in Platform Engineering, DevOps, Data Platform Engineering, MLOps, Backend Engineering, ML Engineering, or a similar role
Hands-on experience on Linux system administration, both in cloud and on-premises environments
Strong engineering skills, including version control, coding best practices, debugging, testing, and production readiness
Hands-on experience building or operating LLM-powered applications, agents, or GenAI platforms in production
Strong platform engineering experience with Kubernetes, Helm, Argo CD, and GitOps practices
Experience with observability tools such as Prometheus, Grafana, Kibana, or similar
Proficient in Python and comfortable working with backend services and APIs
Nice to Haves
Experience working with MCPs
Experience with LLM gateways and model abstraction layers such as LiteLLM
Experience with LLM observability and evaluation tooling such as Langfuse
Experience with high-performance LLM serving frameworks such as vLLM
What You'll Be Doing
Develop tooling and infrastructure to improve the developer experience for building, testing, deploying, and operating GenAI systems
Design and operate services handling areas like model access, tool integration, orchestration, evaluation, observability, and runtime infrastructure
Build scalable and reliable infrastructure on Kubernetes
Work closely with engineering teams to understand use cases and create reusable platform capabilities
Support technically complex use cases, moving teams from experimentation to production
Own projects end to end, from architecture and design to implementation, deployment, and operation
Identify areas of the platform that can be simplified, standardized, or made more reliable and scalable
Collaborate with cross-functional teams and help define standards and best practices