About this role
NVIDIA is seeking a Senior Solution Architect to design customer-focused generative AI solutions involving large language models (LLMs), training, deployment and retrieval-augmented generation (RAG). The full-time role is listed in Bengaluru and Mumbai, India. It involves working with customers, partners and NVIDIA engineering teams to turn language-related business challenges into solutions.
Responsibilities
- Architect end-to-end generative AI solutions, including LLM training and deployment and RAG workflows.
- Engage customers and partners to understand requirements; lead workshops and design sessions to define and refine solutions.
- Lead LLM training and optimization using NVIDIA hardware and software, and develop strategies for effective model performance.
- Provide technical guidance on LLM training and RAG implementation. Share customer feedback with engineering teams to help evolve NVIDIA’s generative AI software, and collaborate across business development, marketing and engineering.
Required qualifications
- Master’s or Ph.D. in Computer Science or Artificial Intelligence, or equivalent experience, and at least 7 years of hands-on technical AI experience focused on generative AI, particularly LLM training.
- A track record of deploying and optimizing LLMs for production inference; expertise in LLM training and fine-tuning with Megatron-LM, Megatron-Bridge, AutoModel and PyTorch.
- Proficiency in deployment and inference optimization across hardware platforms, particularly GPUs; knowledge of GPU cluster architecture and parallel processing for accelerated training and inference.
- Strong communication and collaboration skills, including experience leading workshops or training sessions and explaining technical solutions to varied audiences.
Preferred qualifications
- Experience deploying LLMs in AWS, Azure or GCP cloud environments and on-premises infrastructure; optimizing inference speed, memory efficiency and resource use; and using Docker and Kubernetes for scalable deployment.
- Deep understanding of GPU clusters, parallel and distributed computing, plus hands-on NVIDIA GPU and cluster-management experience designing scalable LLM training and inference workflows.
Skills for this role
Generative AILarge Language ModelsAgentic AIRetrieval-Augmented GenerationLLM trainingLLM fine-tuningLLM deploymentInference optimizationMegatron-LMMegatron-BridgeAutoModelPyTorchGPU cluster architectureParallel processingDistributed computingAWSAzureGoogle Cloud PlatformDockerKubernetesNVIDIA GPU technologiesCustomer collaborationTechnical leadershipWorkshop facilitationTechnical presentations