About this role
You will own the end-to-end post-training process for a multi-agent system, adapting open-weight models across text, vision, and speech modalities to improve performance on complex, domain-specific data like handwritten deeds and satellite imagery.
Skills for this role
PythonPyTorchMachine LearningReinforcement LearningNatural Language ProcessingComputer VisionLLMsTransformersFine-tuningDPOGRPOSFTHugging FaceGISASRTTS