A bias toward experimentation, you form hypotheses, prototype quickly, measure honestly, and let data settle the hard architectural questions.
Nice to Have
Cluster Orchestration & Balancing: Hands-on experience building auto-balancing systems, scheduling frameworks, or multi-dimensional resource allocation algorithms for large-scale distributed systems.
Database Internals & Streaming: Experience with distributed database internals LSM-tree architectures. Deep understanding of streaming mechanics during node join/leave/move operations is a major plus.
Workload Scheduling at Scale: Experience running large-scale stateful workloads on modern orchestration platforms (e.g., Kubernetes), with an understanding of container runtime boundaries, resource limits, and scheduling topologies.
Quota & Traffic Control: Experience implementing admission control, token/leaky-bucket algorithms, or distributed rate limiting in high-throughput environments while considering fairness.
Kernel & OS-Level Intuition: Familiarity with Linux cgroups (v1/v2), eBPF, I/O schedulers, and network topologies that govern resource isolation or the appetite to dig into them on demand.
Willingness to ramp on the JVM, since much of our current stack is Java.
Experience transitioning large single-tenant legacy systems into modern shared-resource, auto-scaling architectures.
What You'll Be Doing
You will own the architecture for resource quota management, node-level capacity estimation, driving performance efficiencies and dynamic cluster rebalancing algorithms across our fleet.
You will figure out how to model CPU, memory, IOPS, and data-streaming requirements per node.
Build software that automatically redistributes data and workloads to optimize the total node count.
Lean on modern workload-scheduling and orchestration primitives to enforce these boundaries, turning raw hardware into a highly efficient, self-healing, self-balancing multi-tenant utility.
Communicate with platform, SRE, and product teams across Apple.