About this role
Distyl AI seeks an Applied AI Researcher, Post-Training to adapt foundation models to the performance and alignment needs of enterprise AI systems. The research aims to turn raw model capability into trustworthy, contextually aligned behavior in mission-critical workflows across industries including telecom, healthcare, insurance and manufacturing.
Responsibilities
- Develop and evaluate supervised fine-tuning, preference optimization (including DPO, RLHF and RLAIF), and continual adaptation techniques for foundation models.
- Investigate ways to align large models with human and system-level objectives, balancing generalization against specialization, data efficiency against robustness, and capability against controllability.
- Build prototypes and run experiments that demonstrate the effectiveness of research ideas for enterprise applications.
Requirements
- Deep understanding of post-training methods, including supervised fine-tuning, RLHF/DPO, LoRA/PEFT and instruction-tuning pipelines.
- Experience adapting LLMs or SLMs to specialized domains or behaviors through data curation, reward modeling or continual pretraining.
- Expertise in building compound AI systems with models, including agentic collaboration and techniques such as ensembling, ReAct and graph-of-thoughts.
- A demonstrated research track record, strong programming and data analysis skills, and the ability to show practical results. Candidates should use AI tools such as ChatGPT, Cursor and Perplexity in their workflow. No specific degree or minimum number of years of experience is stated.
Location and compensation:
- This full-time role is based in San Francisco or New York and follows a hybrid schedule with 3+ in-office days per week, Tuesday–Thursday.
- Base salary is $150K–$250K, depending on experience, location and level. The role is also eligible for equity and benefits including medical, dental and vision coverage for employees and dependents, flexible time off, retirement and financial-planning resources, wellness and family-building benefits, and in-office lunches and snacks.
Skills for this role
Foundation modelsLarge language modelsSmall language modelsSupervised fine-tuningPreference optimizationDPORLHFRLAIFLoRAPEFTInstruction tuningData curationReward modelingContinual pretrainingCompound AI systemsAgentic collaborationEnsemblingReActGraph-of-thoughtsProgrammingData analysisExperimentationChatGPTCursorPerplexity