About this role
The LLM Reliability & Evaluation Engineer will design and implement evaluation pipelines to ensure large language models and agentic AI systems meet enterprise standards for accuracy, safety, and trustworthiness. This role involves benchmarking models, conducting adversarial stress testing, and building automated frameworks to monitor performance and mitigate hallucinations.
Skills for this role
LLM architecturesPrompt engineeringEval harnessesRagasOpenAI EvalsDeepEvalPythonLangChainLangGraphLlamaIndexHugging FaceVector databasesRAG pipelinesKnowledge graph integrationBias/fairness testingResponsible AIRLHFRLAIF