About this role
Design, execute, and operationalize fine-tuning workflows for large language models across supervised, preference-based, and reinforcement learning approaches; lead dataset construction, build scalable distributed training pipelines, and implement rigorous evaluation and safety processes.
Skills for this role
LLM fine-tuningTransformer modelsPythonPyTorchRLHFDPOSupervised fine-tuningLoRAQLoRAAdapter-based methodsDistributed trainingFSDPZeROPipeline parallelismHyperparameter tuningOptimizer configurationTraining stability strategiesDataset constructionData curationInstruction tuningHuman evaluationAutomated benchmarkingSafety evaluationGPU cluster operationMixed precision trainingSequence packingEfficient attention implementationsModel artifact managementReproducibilityEvaluation methodology