Company Logo
Software Engineer

Netflix - 1d ago

Company Logo
Senior Software Engineer

Reddit - 4d ago

Senior Site Reliability Engineer

Requirements

  • 5+ years of experience in Site Reliability Engineering, Platform Engineering, or Infrastructure Engineering
  • Strong hands-on experience with Kubernetes in production — deployment, networking, security, and troubleshooting
  • Experience operating large-scale internal infrastructure platforms in a public cloud, preferably AWS
  • Solid experience with infrastructure-as-code (Terraform) and CI/CD pipelines
  • Exposure to or interest in building infrastructure that hosts AI/ML or agentic workflows
  • Experience with observability platforms and monitoring tools (Grafana, Splunk, APM, or equivalent)
  • Comfortable working in a fast-moving, evolving environment with ambiguity around process and tooling
  • Effective verbal and written communication skills, with the ability to collaborate across time zones
  • Computer Science degree or related field, or equivalent experience

Nice to Haves

  • Exposure to internal developer platform (IDP) concepts or "platform as a product" thinking
  • Hands-on experience with AI agent orchestration, vector databases, model routing, or inference infrastructure
  • Multi-cloud experience (AWS plus Azure or GCP)
  • Service mesh technologies (Istio, Linkerd) or GitOps tooling (ArgoCD, Flux)
  • Kubernetes certifications (CKA, CKS, CKAD) or equivalent cloud certifications
  • Experience administering an enterprise-scale SCM platform (GitHub, GitLab, or equivalent)

What you'll be doing

  • Building, operating, and improving our Kubernetes platform — cluster management, networking, security posture, and day-2 operations
  • Developing self-service capabilities, golden paths, and automation that let internal teams onboard and operate workflows with less platform expertise required
  • Supporting the platform's readiness to host AI-driven internal workflows reliably at scale
  • Troubleshooting and resolving production issues, including participating in a follow-the-sun on-call rotation shared with SRE peers in the US and India, and driving postmortems for incidents you own
  • Contributing to and upholding reliability practices — SLIs/SLOs, error budgets, and operational runbooks
  • Improving CI/CD pipelines and infrastructure-as-code practices, including change and release management
  • Building and maintaining observability and monitoring coverage for the systems you own
  • Mentoring junior engineers and contributing to design and code reviews
  • Collaborating with the Staff SRE and Dublin team to align local implementation with the broader global platform strategy

Perks and Benefits

  • Supporting Your Well-Being
  • Driving Social Impact
  • Developing Talent and Fostering Connection + Community
AI Summary ✨
Okta logo

Okta

Dublin, Ireland

Experience: Senior
Posted: September 18, 2026
Last seen: 2 hours ago
Aws
Azure
Gcp
Git
Kubernetes
Rest
Terraform
sitereliability

Why we track Okta

Okta is the leading identity and access management platform. They have a major engineering hub in Dublin and presence in London. The work involves security, authentication infrastructure, and identity protocols at enterprise scale.

Similar jobs

  • 2 days ago
    New
  • 9 days ago
    Remote
  • 9 days ago
  • See all jobs in Ireland