Apple seeks a Senior ML/RL Training Infrastructure Engineer in Zürich to build large-scale reinforcement learning systems for its foundation models.
This position sits within the Europe-based applied ML team under Core Foundation Models. The primary objective is to engineer robust, high-performance training frameworks that accelerate experimentation for researchers by supporting next-generation large-scale machine learning and reinforcement learning infrastructure.
Responsibilities
- Architect and scale systems designed for large-scale reinforcement learning applied to Apple’s foundation models.
- Construct high-performance RL pipelines utilizing TPU-based training with JAX, incorporating distributed actor/learner setups and efficient experience replay mechanisms.
- Manage the complete stack of RL training systems, ranging from low-level compiler optimizations and performance tuning to cluster-level orchestration.
- Guarantee that training pipelines are reliable, reproducible, and observable, allowing research teams to iterate quickly.
- Partner with engineering hubs in New York, Seattle, and Cupertino to enhance tooling for large-scale model training.
- Identify and resolve bottlenecks in large-scale ML jobs, including issues related to I/O, input pipelines, kernel performance, memory usage, and compilation.
Requirements
- A PhD or MSc degree in Computer Science, Computer Engineering, or a closely related discipline.
- Proven hands-on experience in designing, building, or maintaining large-scale ML training infrastructure.
- Strong proficiency in PyTorch or JAX, along with experience running training workloads on GPUs or TPUs.
- A solid understanding of distributed systems concepts, including parallelism strategies, fault tolerance, and synchronization.
Nice to have
- Practical experience in developing or optimizing training loops, RL pipelines, or large-scale model-training frameworks.
- Strong software engineering skills in Python, with a focus on reliability, debuggability, and high-performance execution.
- Deep experience with PyTorch/JAX internals, XLA, debugging, and performance profiling on GPU/TPU architectures.
- Expertise in distributed RL training patterns, including actor/learner architectures, experience replay, and parallel environment execution.
- Experience building training services, orchestration tools, or automated pipelines for large-scale experiments.
- Familiarity with RL-specific infrastructure requirements such as actor/learner architectures and large-scale environment execution.
- Experience working with cloud-scale clusters or specialized accelerators like TPU v5/v6, GPU, or custom hardware.
- Contributions to ML frameworks, distributed training libraries, or high-performance computing systems.
About the company
Apple values diversity and believes that differences in who people are, what they've experienced, and how they think are their greatest strength. The company is committed to creating products that serve everyone by including everyone.
- Apple values diversity and believes that differences in who people are, what they've experienced, and how they think are their greatest strength.
- Apple is committed to creating products that serve everyone by including everyone.
Quelle: öffentlich zugängliche Karriereseite des Arbeitgebers. Batchly ist nicht der Arbeitgeber und steht nicht notwendigerweise in einem Vertragsverhältnis mit dem Unternehmen.