NVIDIA seeks a Senior Solutions Architect to guide partners and customers in deploying AI inference services, focusing on LLM integration and GPU optimization within the Worldwide Field Operations team.
This customer-facing position serves as the primary technical liaison between NVIDIA and its global partner network. The specialist supports cloud partners in integrating the NVIDIA AI stack, enabling them to build, launch, and maintain complete AI service lifecycles that span from initial training phases through to final inference deployment.
Responsibilities
- Collaborate closely with NVIDIA Cloud Partners and their end-users to assess technical needs and design optimal solutions.
- Create and showcase implementations leveraging NVIDIA’s proprietary and open-source NLP and Large Language Model technologies, ensuring seamless integration into agentic workflows.
- Conduct rigorous performance analysis and tuning to maximize efficiency on GPU-based infrastructure, specifically for inference tasks and agentic pipelines.
- Coordinate with Engineering, Product, and Sales divisions to architect and plan tailored solutions for client requirements.
- Inform product development by gathering customer insights and validating new features through proof-of-concept testing.
- Establish deep domain expertise in embedding NVIDIA technologies into AI Cloud and Enterprise Computing frameworks.
- Lead technical demonstrations and facilitate strategic discussions with developers, product managers, and executive stakeholders.
- Promote the adoption of the NVIDIA AI platform and streamline the process of moving models into production environments.
Requirements
- Strong proficiency in English, including verbal communication, written documentation, and technical presentation abilities.
- A Master’s degree or Ph.D. in Computer Science, Artificial Intelligence, or comparable professional experience.
- At least five years of industry or academic background in machine learning, deep learning, or data science, with a specific focus on Deep Neural Network (DNN) inference.
- Practical knowledge of modern architectures including Large Language Models (LLMs), Vision-Language Models (VLMs), and diffusion models, particularly those utilizing Mixture of Experts (MoE) structures.
- Familiarity with core DNN inference libraries such as TRT-LLM, Dynamo, and RedHat Inference Server, alongside experience in developing agentic pipelines.
Nice to have
- Hands-on experience with inference workloads for massive MoE architectures across domains like NLP, Computer Vision, or Automatic Speech Recognition.
- Proficiency with DevOps tools including Docker, Kubernetes, and Singularity.
- Understanding of High-Performance Computing (HPC) environments, covering data center layout, InfiniBand interconnects, cluster storage, and scheduling management.
About the company
The team operates in a dynamic, fast-paced environment centered on artificial intelligence. They value individuals who are self-starters, driven by continuous learning, and eager to share discoveries with colleagues.
- A culture that prioritizes constant innovation and rapid development in the AI sector.
- Expectation for team members to maintain a growth mindset and actively contribute knowledge to the group.
Quelle: öffentlich zugängliche Karriereseite des Arbeitgebers. Batchly ist nicht der Arbeitgeber und steht nicht notwendigerweise in einem Vertragsverhältnis mit dem Unternehmen.