Company Logo
Software Engineer

Netflix - 1d ago

Company Logo
Senior Software Engineer

Reddit - 4d ago

Staff Site Reliability Engineer

Requirements

  • 8+ years of experience in Site Reliability Engineering, Platform Engineering, or Infrastructure Engineering, with a track record of staff-level technical leadership
  • Deep, hands-on expertise in Kubernetes at production scale — architecture, multi-tenancy, networking, security, and day-2 operations
  • Experience running large-scale internal infrastructure platforms in a public cloud, preferably AWS
  • Strong expertise in cloud-native architectures, infrastructure-as-code (Terraform), and CI/CD pipelines
  • Experience building or operating infrastructure that hosts AI/ML or agentic workflows — model serving, GPU scheduling, or orchestration frameworks
  • Deep experience with observability platforms and monitoring tools (Grafana, Splunk, APM, or equivalent) in a large-scale environment
  • Demonstrated ability to influence technical direction across teams and geographies without formal authority
  • Effective verbal and written communication skills, with experience mentoring engineers and driving technical alignment across distributed teams
  • Computer Science degree or related field, or equivalent experience

Nice to Haves

  • Building internal developer platforms (IDPs) with a "platform as a product" mindset
  • Hands-on experience with AI agent orchestration, vector databases, model routing, or inference optimization infrastructure
  • Multi-cloud experience (AWS plus Azure or GCP)
  • Service mesh technologies (Istio, Linkerd) or GitOps tooling (ArgoCD, Flux) in production Kubernetes environments
  • Kubernetes certifications (CKA, CKS, CKAD) or equivalent cloud certifications
  • Experience managing or administering an enterprise-scale SCM platform (GitHub, GitLab, or equivalent)

What you'll be doing

  • Owning the architecture and evolution of our Kubernetes platform for the Dublin/EMEA team — cluster strategy, multi-tenancy, networking, and security posture — in close coordination with counterparts in the US and India
  • Leading the technical design and rollout of self-service platform capabilities, golden paths, and self-healing patterns that let internal teams ship and operate workflows without deep platform expertise
  • Driving the platform's readiness to host AI-driven internal workflows reliably at scale — model serving, GPU scheduling, and orchestration infrastructure
  • Taking ownership of the hardest, highest-ambiguity reliability problems: complex production incidents, capacity and scaling challenges, and cross-system failure modes
  • Participating in a follow-the-sun on-call rotation shared with SRE peers in the US and India, and driving root-cause resolution and postmortems for the incidents you own
  • Setting technical standards for SLIs/SLOs, error budgets, and on-call practice, and holding the team accountable to them
  • Improving SDLC processes for infrastructure-as-code, including CI/CD pipeline maturity and change/release management
  • Mentoring engineers across the team and raising the technical bar through design reviews, architecture discussions, and hands-on pairing
  • Partnering with architects, security, and product engineering teams to align infrastructure decisions with the business's reliability, security, and delivery needs
  • Representing the Dublin/EMEA team in global platform architecture discussions, working as a peer with Staff/Principal-level engineers in the US and India

Perks and Benefits

  • Annual base salary range in Ireland: €92.000 - €126.500 EUR
  • Comprehensive healthcare coverage
  • Equity (where applicable) and bonus
  • Paid time off and parental leave
AI Summary ✨
Okta logo

Okta

Dublin, Ireland

Experience: Staff
Posted: September 18, 2026
Last seen: 2 hours ago
Aws
Azure
Gcp
Git
Kubernetes
Rest
Terraform
sitereliability

Why we track Okta

Okta is the leading identity and access management platform. They have a major engineering hub in Dublin and presence in London. The work involves security, authentication infrastructure, and identity protocols at enterprise scale.

Similar jobs

  • 2 days ago
    New
  • 9 days ago
    Remote
  • 9 days ago
  • See all jobs in Ireland