Apertus Engineer: Infrastructure

ETH Zurich · Zürich

ETH Zurich seeks an Infrastructure Engineer to ensure the stability and performance of Apertus ML pipelines on the Alps supercomputer, collaborating closely with CSCS.

This position focuses on maintaining the infrastructure that supports the Apertus project's machine learning workflows. The engineer is responsible for ensuring high throughput and system reliability by managing container images and coordinating with the Swiss National Supercomputing Centre (CSCS) to optimize the underlying compute environment.

Responsibilities

  • Develop, maintain, and update container images for pre-training, post-training alignment, and model deployment phases.
  • Manage the full dependency stack, including CUDA, NCCL, and PyTorch, specifically for the ARM-based (aarch64, Grace-Hopper) architecture of the Alps supercomputer.
  • Ensure that image builds are reproducible, properly versioned, and well-documented, including the implementation of CI processes for builds and upgrades.
  • Validate container images against reference workloads in collaboration with Apertus engineers and maintain functional launch examples.
  • Act as the main technical liaison with CSCS engineers and researchers regarding compute resources, system reliability, and efficiency.
  • Partner with CSCS staff to identify and implement improvements in compute infrastructure efficiency and performance, aligning with ML systems engineering goals.
  • Contribute to systemic enhancements in CSCS resources, such as network, storage, and scheduling, that impact large-scale LLM training.
  • Document and share institutional knowledge regarding CSCS infrastructure and best practices for utilizing high-performance systems.
  • Stress test infrastructure using representative workloads to validate stability and throughput following image upgrades, maintenance, or configuration changes.
  • Collaborate with Apertus engineers to debug cluster-level issues affecting stability, such as node failures, networking bottlenecks, storage performance, checkpointing, and scheduling.
  • Provide support for the Apertus serving stack, which utilizes the same container images, while noting that operational ownership lies with a separate engineer.

Requirements

  • MSc or PhD in Computer Science, Data Science, Artificial Intelligence, Machine Learning, or a related discipline.
  • Practical experience with High-Performance Computing (HPC) environments, including job schedulers like Slurm, shared filesystems, and multi-node GPU systems.
  • Strong proficiency in Linux systems and container technologies, such as Docker, Podman, and HPC runtimes like enroot or Apptainer.
  • Excellent collaboration and communication skills, with the ability to work effectively across research, engineering, and operations teams.
  • Prior hands-on experience in the core domains of this role, which may be gained through projects or studies; formal work experience is preferred.
  • High flexibility to adapt to shifting priorities, tools, and tasks driven by training schedules, releases, and the dynamic nature of the field.

Nice to have

  • Familiarity with LLM training and serving frameworks such as Megatron-LM, PyTorch distributed, vLLM, or SGLang.
  • Experience building or adapting containers for ARM64/aarch64 platforms.
  • Knowledge of HPC networking and communication stacks, including Slingshot, libfabric, NCCL, and debugging techniques.
  • Experience with parallel filesystems like Lustre and storage performance tuning.
  • Experience setting up CI/CD pipelines for container image builds.

Auf Firmen-Website bewerben

Ähnliche Stellen

Quelle: öffentlich zugängliche Karriereseite des Arbeitgebers. Batchly ist nicht der Arbeitgeber und steht nicht notwendigerweise in einem Vertragsverhältnis mit dem Unternehmen.