Tether seeks a Senior AI Research Engineer in Zürich to optimize AI model serving and inference for edge and mobile devices. The role focuses on low-latency, high-throughput, memory-efficient systems, with remote work available.
As part of Tether’s AI model team, this role drives innovation in model serving and inference architectures, aiming to deliver efficient, scalable AI performance across real-world applications—from resource-limited devices to complex multi-modal systems.
Responsibilities
- Design and deploy advanced model serving systems that maximize throughput and minimize latency while reducing memory usage.
- Ensure efficient operation across diverse environments, including edge platforms and devices with limited resources.
- Define performance goals such as lower latency, faster token responses, and reduced memory footprint.
- Conduct controlled inference tests in simulated and live production settings.
- Monitor key metrics including response time, throughput, memory use, and error rates.
- Document and compare results against benchmarks to validate performance across platforms.
- Create high-quality test datasets and simulation scenarios focused on real-world challenges, especially on low-resource devices.
- Establish measurable criteria to evaluate model performance, latency, and memory use under various conditions.
- Analyze computational efficiency and identify bottlenecks in the serving pipeline using processing and memory metrics.
- Resolve issues like inefficient batch processing, network delays, and high memory consumption to improve scalability and reliability.
- Collaborate with cross-functional teams to integrate optimized inference frameworks into production pipelines for edge and on-device use.
- Define success metrics including improved real-world performance, low error rates, scalability, and optimal memory use, with ongoing monitoring and iterative improvements.
Requirements
- A degree in Computer Science or a related field.
- A PhD in NLP, Machine Learning, or a related area, with a strong publication record in top AI conferences, is preferred.
- Must have knowledge of Metal Shading Language (MSL).
- Should be able to write custom compute shaders from scratch.
- Proven experience in low-level kernel optimizations and inference optimization for mobile devices is required.
- Demonstrated contributions that improved inference latency, throughput, and memory footprint for applications on resource-constrained devices and edge platforms.
- Deep understanding of modern model serving architectures and inference optimization techniques, including methods for low-latency, high-throughput, and efficient memory management in constrained environments.
- Strong expertise in writing GPU kernels for mobile devices such as smartphones.
- In-depth knowledge of model serving frameworks and engines is required.
- Practical experience in developing and deploying full inference pipelines—from model optimization to integration on resource-limited devices.
- Proven ability to apply empirical research to solve challenges in model serving, such as latency, bottlenecks, and memory constraints.
- Proficiency in designing robust evaluation frameworks and iterating on optimization strategies to enhance inference performance and system efficiency.
Nice to have
- Experience with distributed inference systems, including optimization techniques like Tensor Parallelism, Pipeline Parallelism, and Expert Parallelism for large models on GPU clusters.
- Deep understanding of the mathematical and structural foundations of Diffusion Models and Vision Transformers.
- Familiarity with techniques such as Pruning, Quantization, Flash attention, KV Cache, and Speculative Decoding (Eagle).
About the company
Tether is pioneering a global financial revolution by enabling seamless integration of reserve-backed tokens across blockchains. Their solutions support businesses including exchanges, wallets, payment processors, and ATMs. The company values transparency, operates remotely with a global team, and has grown rapidly while maintaining a lean structure.
- Pioneering a global financial revolution with blockchain-based solutions.
- Empowering businesses to integrate reserve-backed tokens across blockchains.
- Commitment to transparency in all operations.
- Global, remote-first team with talent from around the world.
- Fast growth, lean operations, and industry leadership.
- Strong emphasis on innovation and excellence in product development.
Quelle: öffentlich zugängliche Karriereseite des Arbeitgebers. Batchly ist nicht der Arbeitgeber und steht nicht notwendigerweise in einem Vertragsverhältnis mit dem Unternehmen.