About this role
Set the technical direction for synthetic data generation across NVIDIA's frontier model efforts, building open-source NeMo libraries that generate synthetic datasets across text, code, structured, and multimodal data to feed pre- and post-training of LLMs such as Nemotron. The role combines hands-on software engineering and applied research to build scalable data generation pipelines, advance privacy-preserving synthesis, multimodal generation, and publish original research.
Skills for this role
Synthetic data generationGenerative modelingMultimodal machine learningLarge language models (LLMs)vLLMTGINeMoNemotronData pipeline developmentDistributed inferenceScalable data pipelinesReinforcement learningReward modelingDifferential privacyOpen-source library developmentSDK developmentAPI designGitCI/CDData anonymizationDe-identification