ML Infra Engineer

Flexion Robotics · Zürich

Flexion Robotics in Zürich seeks a Senior ML Infra Engineer to build core compute and data platforms for humanoid robots. The role involves designing training clusters, architecting data pipelines, and optimizing distributed training for large-scale AI models.

This senior role focuses on developing the foundational infrastructure for humanoid robotics, including training clusters and data pipelines. It is part of the Infrastructure team, which draws on expertise from companies like Google, Meta, and Amazon, and requires on-site presence at the Zürich office.

Responsibilities

  • Design, deploy, and maintain GPU compute clusters for large-scale model training across multiple cloud providers, including job scheduling with Slurm and Kubernetes.
  • Build data platforms and pipelines covering storage, processing, and serving layers for the full data lifecycle—from simulator output and robot telemetry to training datasets.
  • Utilize object storage (S3), parallel filesystems (Lustre), and data formats (Parquet, WebDataset, LeRobot) for data infrastructure.
  • Apply distributed processing frameworks (Ray, Spark) to transform and validate data at scale.
  • Optimize distributed training by collaborating with AI engineers to improve throughput, device utilization, and communication efficiency across multi-node GPU clusters.
  • Enhance distributed training for IsaacLab-based sim-to-real workflows.
  • Assess and integrate new platforms, including cloud providers, GPUaaS services, and emerging tools, to support growing compute needs.

Requirements

  • 3+ years of experience building and operating infrastructure for large-scale deep learning systems.
  • Hands-on experience with training or supporting training of large models (billions of parameters) in distributed multi-node GPU environments, with deep knowledge of DDP, FSDP, and NCCL.
  • Strong experience with at least one major cloud platform (AWS or GCP), including compute provisioning and networking.
  • Experience with job scheduling and orchestration tools such as Slurm or Kubernetes.
  • Experience in building data pipelines and managing large-scale storage, including object stores (S3 or equivalent) and familiarity with parallel filesystems like Lustre.
  • Proficiency in Python and working knowledge of PyTorch.
  • Ownership mindset: ability to make architectural decisions, set direction, and deliver independently in a fast-paced environment.

Nice to have

  • Experience with distributed data processing frameworks such as Ray or Spark.
  • Familiarity with data formats including Parquet, WebDataset, and LeRobot.
  • Experience with additional GPU cloud providers like Lambda Labs, CoreWeave, RunPod, or Nebius.
  • Experience managing on-premise compute infrastructure.
  • Familiarity with robotics simulation environments such as IsaacLab, IsaacGym, or MuJoCo.
  • Experience with infrastructure-as-code tools like Terraform or Ansible.
  • Familiarity with experiment tracking platforms such as Weights & Biases or MLflow.
  • Experience with GPU programming and profiling tools like CUDA or Nsight.

What the company offers

  • Competitive compensation package
  • A front-row seat at one of Europe’s most ambitious robotics companies
  • An energetic, collaborative team with a bias for action

About the company

Flexion Robotics is developing the intelligence layer for next-generation humanoid robots, aiming to accelerate the shift from fragile prototypes to real-world deployment. Founded by experts in robot reinforcement learning (former Nvidia and ETH Zürich researchers) and backed by leading international VC firms, the company achieved rapid progress, deploying humanoid capabilities within months of its first line of code.

  • Building the intelligence layer for humanoid robots
  • Accelerating transition from prototypes to real-world deployment
  • Founded by leading robotics scientists with experience at Nvidia and ETH Zürich
  • Backed by top international venture capital firms
  • Rapid development cycle: from first code to real-world deployment in months

Auf Firmen-Website bewerben

Ähnliche Stellen

Quelle: öffentlich zugängliche Karriereseite des Arbeitgebers. Batchly ist nicht der Arbeitgeber und steht nicht notwendigerweise in einem Vertragsverhältnis mit dem Unternehmen.