About this role
Own and build the end-to-end evaluation and experimentation pipeline for Nuro's agent fleet, including eval data collection, eval loop construction, and automated hill-climbing; run hands-on post-training experiments on open-weight vision-language models using proprietary driving data and collaborate with engineering to deploy and measure impact.
Skills for this role
Agent systemsEvaluation pipelineClosed-loop evaluationEval data collectionEval loop constructionAutomated hill-climbingExperiment designStatistical analysisSupervised fine-tuning (SFT)Reinforcement learningVision-language modelsOpen-weight modelsPost-trainingData curationProduction systemsPythonLLM agent systemsTransformer modelsTest-time compute optimizationVerifier systemsData labeling managementTask suite constructionMining production tracesHuman feedback collectionA/B testingCausal inferenceMultimodal models