Apertus Engineer: Deployment

ETH Zurich · Zürich

ETH Zurich seeks a Mid-level Apertus Engineer to ensure seamless deployment of open-source LLMs across server and personal environments, bridging research and community infrastructure.

The company is looking for an engineer to guarantee that Apertus model releases function immediately upon publication within the broader open-source LLM landscape. This position focuses on enabling both high-performance server inference and accessible local deployment, ensuring the technology is ready for widespread use.

Responsibilities

  • Oversee the technical release process for trained Apertus models, including checkpoint conversion and the preparation of release assets like weights, configurations, tokenizers, and model cards in collaboration with the training team.
  • Develop and integrate support for Apertus model architectures into community libraries such as Hugging Face Transformers, vLLM, SGLang, and llama.cpp, guiding these contributions through the review process to ensure robust community adoption.
  • Confirm that new model releases are compatible with major inference engines and standard model formats from day one.
  • Align release schedules and technical documentation with the community manager to ensure smooth public rollout.
  • Create quantized versions of released models (e.g., FP8, INT4/AWQ/GPTQ, GGUF) optimized for both server-grade and personal hardware.
  • Test quantized variants against established evaluation benchmarks to verify that model quality remains intact.
  • Develop example scripts and reference configurations demonstrating how to serve and utilize Apertus models using vLLM, SGLang, and Transformers.
  • Provide support for local and personal deployment ecosystems, including tools like LM Studio, Ollama, and llama.cpp.
  • Maintain and update deployment documentation and troubleshooting guides for the user community.

Requirements

  • MSc or PhD in Computer Science, Data Science, Artificial Intelligence, Machine Learning, or a closely related discipline.
  • Exceptional BSc graduates with substantial engineering experience may also be considered.
  • Proficient in Python and software engineering, with practical experience in open-source workflows such as pull requests, code reviews, and CI pipelines.
  • Hands-on experience with LLM inference stacks, including Hugging Face Transformers, vLLM, or SGLang.
  • Strong collaboration and communication abilities, with the capacity to work effectively across research, engineering, and community-facing teams.
  • Prior hands-on experience in the core areas of this role, gained through projects or studies; formal work experience is preferred.
  • High flexibility to adapt to shifting priorities, tools, and daily tasks driven by training schedules, release cycles, and the rapid pace of the field.
  • A proven track record of merged contributions to ML or inference libraries such as Transformers, vLLM, SGLang, or llama.cpp.

Nice to have

  • Experience converting models between different formats and frameworks, such as Megatron-LM checkpoints, safetensors, and GGUF.
  • Familiarity with personal and local deployment tools like LM Studio, Ollama, or llama.cpp.
  • Background in writing developer-facing documentation and example code.
  • Published research in relevant domains or familiarity with recent academic publications in these areas.
  • Experience in quantizing models without performance loss (FP8, INT4, AWQ, GPTQ) and evaluating the results.
  • Knowledge of LLM evaluation harnesses and benchmark pipelines.
  • Experience with GPU inference performance tuning and large-scale serving.
  • Experience with Apple Silicon/MLX or other consumer-hardware inference targets.

What the company offers

  • A stimulating academic environment at one of the world's leading technical universities.
  • Access to Alps, one of the largest AI-ready supercomputers in Europe.
  • The opportunity to work alongside and intersect with leading researchers in the field.
  • Collaboration with top researchers and engineers from EPFL, ETH Zürich, CSCS, and other Swiss institutions.
  • Attractive employment conditions and comprehensive benefits, including the ETH Zürich/EPFL pension plans.
  • Flexible working arrangements, including options for remote work.
  • Professional development opportunities, including conference attendance and specialised training.
  • The chance to contribute to open-source projects with global impact.
  • Being part of Switzerland's sovereign AI development, working on technology with national significance.

About the company

ETH Zurich is a premier technical university contributing to Switzerland's sovereign AI development. The institution fosters a collaborative environment where engineers and researchers from leading Swiss institutions work together on technology with national significance and global open-source impact.

  • Contributing to open-source projects with global impact.
  • Being part of Switzerland's sovereign AI development, working on technology with national significance.

Auf Firmen-Website bewerben

Ähnliche Stellen

Quelle: öffentlich zugängliche Karriereseite des Arbeitgebers. Batchly ist nicht der Arbeitgeber und steht nicht notwendigerweise in einem Vertragsverhältnis mit dem Unternehmen.