5–8+ years of professional DevOps or SRE experience, with a track record of owning production systems at scale
Deep hands-on experience with large-scale logging infrastructure – Splunk, Loki, ClickHouse, or Elastic
Solid working knowledge of OpenTelemetry: collector configuration, pipelines, and instrumentation
Strong programming skills in Go and/or Python; experience building integrations and applications to large-scale Observability environments
Experience developing, deploying and running distributed applications on cloud platforms; experience with container and orchestration technologies (Docker, Kubernetes)
Comfortable owning on-call across a multi-tool observability stack, including leading Sev 1/2 incident response
Nice to Have
Experience evaluating and prototyping alternative storage/processing backends (e.g., ClickHouse, Loki) as part of a cost or stability migration
Experience with other Observability tooling like Grafana, Cortex, and Tempo
Promote the DevOps/SRE approach
What You Will Be Doing
Design and operate Adobe's observability infrastructure at scale – Splunk, Loki, ClickHouse, Cortex, Tempo, and OTel collector pipelines processing billions of events daily
Drive cost optimisation: analyse pipeline data, build guardrails, and engage engineering teams
Build AI-powered tooling that surfaces actionable insights from high-volume log datasets and automates routine platform workflows
Own SLOs and SLIs for a platform thousands of engineers rely on every day
Own on-call, lead incident response for high-severity issues, and make every shift better than the last