AI Research Engineer (Model Compression & Quantization)

Tether · Zürich

Tether seeks a Senior AI Research Engineer in Zürich to advance model compression and quantization for multimodal AI systems, enabling efficient deployment on edge devices. The role focuses on reducing model size and latency while preserving accuracy, using techniques like quantization, distillation, and pruning.

As part of Tether’s AI research team, this role drives innovation in making advanced multimodal AI systems—such as large language and vision-language models—run efficiently on resource-limited devices like smartphones. The goal is to deliver low-latency, low-memory AI solutions that maintain high performance and real-world utility.

Responsibilities

  • Apply low-bit quantization techniques to shrink generative AI models (LLMs, VLMs, multimodal) and reduce inference latency without compromising output quality.
  • Use knowledge distillation to transfer capabilities from large models to smaller ones, enabling efficient multimodal reasoning across text, image, and audio.
  • Implement pruning methods to eliminate redundant parameters and attention heads, lowering computational load while preserving performance.
  • Evaluate trade-offs between model efficiency (size, latency, memory) and accuracy across quantization, distillation, and pruning, and suggest improvements based on experimental results.
  • Research and deploy mixed-precision quantization and other advanced compression methods, such as adaptive pruning and intermediate feature matching, to optimize accuracy and performance.
  • Stay updated on the latest model compression research, especially for multimodal and generative architectures.
  • Clearly document research methods, experiments, and outcomes to support reproducibility, team collaboration, and communication with stakeholders.
  • Write and publish technical papers in top AI conferences (e.g., NeurIPS, ICML, ICLR, CVPR, ACL, AAAI) to contribute to the field of model compression for multimodal AI.

Requirements

  • A degree in Computer Science or a related field.
  • Experience with PyTorch or equivalent deep learning frameworks.
  • Practical experience with model quantization, including Quantization-Aware Training (QAT) and Post-Training Quantization (PTQ).
  • Research and hands-on experience in knowledge distillation for compressing large models into smaller, efficient versions.
  • Research and hands-on experience in model pruning for compressing large models into smaller, efficient versions.
  • Strong understanding of neural network architectures and training processes, including transformers (LLMs, VLMs), backpropagation, optimization, and fine-tuning.

Nice to have

  • Familiarity with C++ is beneficial, particularly for developing low-level quantization kernels or inference optimizations.

About the company

Tether is leading a global financial transformation by enabling businesses to use reserve-backed tokens across blockchains. The company values transparency, operates with a lean structure, and brings together a global team of top talent working remotely. Tether is positioned at the intersection of technology and human potential, setting new industry standards.

  • Tether pioneers a global financial revolution with reserve-backed tokens across blockchains.
  • Transparency is central to Tether’s operations, ensuring trust in every transaction.
  • Tether’s team is a global talent network, working remotely from around the world.
  • The company has grown rapidly while maintaining a lean structure and industry leadership.
  • Tether combines technology with human potential, pushing boundaries and setting new standards.

Auf Firmen-Website bewerben

Ähnliche Stellen

Quelle: öffentlich zugängliche Karriereseite des Arbeitgebers. Batchly ist nicht der Arbeitgeber und steht nicht notwendigerweise in einem Vertragsverhältnis mit dem Unternehmen.