Site Reliability Engineer (SRE), Observability, London
Site Reliability Engineer (SRE), Observability, London
Requirements
Strong sense of ownership and integrity demonstrated through clear communication and collaboration
Experience in managing and scaling distributed systems in a public, private, or hybrid cloud environment
The ability to design, author, and release code in languages like (but not limited to) Go or Python
Acute drive to automate manual operations and to improve them through repeated iteration
Understanding of the Linux Operating System, standard networking protocols, and components
Nice to Haves
Hands-on experience managing large numbers of diverse systems with configuration management or software delivery platforms (such as Puppet and Spinnaker)
Experience with deploying, supporting and monitoring new and existing services, platforms, and application stacks
Experience with scale testing, disaster recovery, and capacity planning
Familiarity with microservices architecture and container orchestration with Kubernetes
What You'll Be Doing
Solving problems using data, teamwork, and expertise in a diverse and challenging technical environment
Owning the full infrastructure stack, from performance debugging to traffic management
Working with Linux and Kubernetes to run systems and utilize various tools for system management
Collaborating with development teams to deliver optimal results for Apple services