About this role
Own the quality, reliability, and trustworthiness of clinical AI outputs by building evaluation frameworks, calibrated confidence scoring, automated QA, and scalable quality gates. Contribute to model alignment and fine-tuning and implement feedback loops to continuously improve model outputs.
Skills for this role
PyTorchLLM architecturesModel evaluationBenchmarkingQuality metricsPythonMachine LearningStatisticsResearch paper implementationCommunicationUncertainty quantificationConfidence scoringEvaluation frameworksAutomated evaluation pipelinesFeedback loopsModel alignmentFine-tuningRLHFDPOHealthcare AIProduction LLM evaluationClinician-validated test cases