New opportunity

AI Engineer, Evaluation

Distyl AI · San Francisco; New York, United States

About this role

Design and implement evaluation frameworks and pipelines for customer-facing AI systems, write production Python code, and build golden test cases and regression suites. Define, calibrate, and operate LLM-based graders and use evaluation signals to guide prompt design, agent logic, model selection, and release readiness.

Skills for this role

PythonEvaluation-Driven DevelopmentExperiment-Driven DevelopmentEvaluation frameworksEvaluation pipelinesLLMsLLM-based gradersTest case designRegression testingAI-assisted test generationPrompt designAgent logicModel selectionScoring functionsProduction software engineeringSystems-oriented designExperimentation frameworksCollaboration with domain experts