About this role
Design, evaluate, and improve AI capabilities for legal research and workflows, building prototypes and production-ready features using LLMs, RAG, and agentic systems. Partner with engineers, product teams, and customers to define evaluation methodologies, benchmarks, monitoring, and tooling to ensure quality, reliability, and explainability.
Skills for this role
Applied Data ScienceMachine LearningLarge Language Models (LLMs)Agentic SystemsPrompt EngineeringRetrieval-Augmented Generation (RAG)Retrieval and RankingEmbeddingsSemantic SearchStructured GenerationExperimental DesignA/B TestingCausal InferenceConfidence/Uncertainty QuantificationPythonpandasNumPyscikit-learnSQLLangChainLangGraphLlamaIndexOpenAI APIsAnthropic APIsAWSAzureGoogle Cloud PlatformError AnalysisBenchmarkingHuman-in-the-loop Evaluation