About this role
The Post-Training researcher develops and evaluates techniques to adapt and align foundation models for enterprise use (e.g., supervised fine-tuning, preference optimization DPO/RLHF/RLAIF, LoRA/PEFT, instruction-tuning, and continual adaptation). This hybrid role is based in San Francisco and New York and involves building prototypes, running experiments, and informing safe, scalable use of foundation models across industries.
Skills for this role
Supervised fine-tuningPreference optimizationRLHFDPORLAIFLoRAPEFTInstruction-tuning pipelinesFoundation model adaptationLLMsSLMsData curationReward modelingContinual pretrainingAgentic collaborationEnsemblingReActGraph-of-thoughtsProgrammingData analysisPrototypingChatGPTCursorPerplexity