About this role
Develop and deploy LLM post-training pipelines (e.g., SFT, DPO, PPO, Reward Modeling), optimize agent and tool-use training, and iterate core scenario models; manage post-training data workflows and build automated evaluation combining human review and LLM-as-Judge.
Skills for this role
SFTDPOPPOReward ModelingRLHFMegatronveRLDistributed Training1000-GPU-scale Distributed TrainingHyperparameter TuningModel EvaluationOnline RegressionAgent Tool-UseTool CallingMulti-turn DialogueFunction CallingData CollectionData CleaningData SynthesisData AnnotationAnnotation Guideline DevelopmentVendor CoordinationAutomated EvaluationLLM-as-JudgeTraining OptimizationRouter ModelsQuery RewritingWeb Search Decision MakingMemory ModelingE-commerce Search Relevance