About this role
Own how Cardboard measures and improves the quality of its AI agent by defining quality standards, building trusted evaluation datasets from real product usage, creating offline and online evaluations (including automated checks and human review), analyzing model and agent failure patterns, and working with product and engineering to ship improvements.
Skills for this role
LLMsAI agentsTypeScriptPythonEvaluation dataset constructionOffline evaluationOnline evaluationAutomated checksHuman reviewModel evaluationDataset buildingExperiment designStatisticsHuman labelingModel gradersFine-tuningModel selectionRegression checksRelease gatesQuality monitoringLatency trackingCost monitoringProduct judgmentSoftware engineeringMultimodal AIVideoMedia pipelines