About this role
Researcher on the Post-Training team developing post-training and alignment methods for multimodal foundation models, designing preference optimization, evaluation frameworks, and feedback-driven learning. The role includes implementing, debugging, and scaling experimental systems and translating research findings into production-ready systems to improve model reasoning and human alignment.
Skills for this role
Preference optimizationAlignment methodsRLHFMultimodal model trainingGenerative modelsModel evaluationMetrics designFeedback-driven learningHuman-in-the-loop evaluationMachine learning systems engineeringModel debuggingScaling ML systemsReproducibility in trainingState Space Models (SSMs)Foundation models