About this role
Conduct research to develop, scale, and optimize post-training methods—including reinforcement learning, preference-based optimization, fine-tuning, and alignment—for scientific foundation models. Work includes designing and evaluating post-training pipelines and workflows on leadership-class supercomputers and collaborating with computational scientists and domain researchers to apply adaptive learning systems to scientific problems.
Skills for this role
Reinforcement LearningPost-TrainingPreference OptimizationFine-tuningAlignmentPolicy OptimizationBandit AlgorithmsPreference LearningReinforcement Learning from FeedbackDirect Preference OptimizationReward ModelingModel AdaptationLarge-scale Model TrainingDistributed Learning SystemsDistributed TrainingMulti-accelerator ExecutionPythonCC++PyTorchJAXMathematical OptimizationLinear AlgebraNumerical MethodsData MiningStatisticsSoftware Development PracticesWorkflow OptimizationHigh-performance Computing (HPC)Scalability