Proficiency in at least one systems/scripting language for platform tooling and automation - Python, Go, or Bash.
Hands-on Kubernetes experience at the operational level - cluster upgrades, networking, capacity, and multi-tenancy, not just deploying to a managed cluster.
Strong Linux fundamentals (RHEL especially) in a production environment.
Practical experience owning CI/CD tooling at the platform level (Jenkins, GitLab CI, ArgoCD/Flux, or equivalent) - pipeline templates and shared workflows.
Infrastructure-as-Code and GitOps experience (Helm, Ansible, Terraform, ArgoCD).
A working understanding of platform observability (Prometheus, Grafana, or similar).
The mindset that application teams are your customers - building tools people want to use, and communicating clearly with technical and non-technical colleagues.
Nice to haves:
Experience operating enterprise storage at scale, and database-adjacent platform work.
Kubernetes networking and ingress (Cilium, Calico, ingress controllers, service mesh).
Secrets management and workload identity (SPIFFE/SPIRE, Vault, External Secrets).
Security - experience with cyber security standards such as CIS Benchmarks and frameworks such as NIST CSF.
Experience in a regulated or audited environment.
What you'll be doing:
Operate and harden the Kubernetes estate - cluster operations including upgrades, capacity planning, networking, and multi-tenant namespace management across on-premise and hybrid environments.
Own the CI/CD and deployment platform layer - Jenkins, Artifactory, ArgoCD and related tooling - so pipelines are reliable, fast, and extensible without the platform team being a bottleneck on every new service.
Build self-service capability - paved-path templates and tooling that let application teams deploy and operate without manual handoffs.
Implement security and compliance guardrails in the platform - pipeline and supply-chain security (DefectDojo, Keycloak, security scanning), workload identity (SPIFFE/SPIRE), and policy enforcement at admission rather than manual review gates.
Instrument platform observability and SLOs - ensure services shipped through the platform come with baseline dashboards, structured logging, and alerting; contribute to the Quality-of-Service standard the platform is held to.
Contribute to platform networking - ingress/egress, traffic flows, and connectivity, working with the network team where it touches network operations.
Participate in reliability and incident response - root-cause analysis, drift detection, and continuous improvement of platform stability.