New opportunity

Applied AI Researcher, Benchmarking

Distyl AI · San Francisco; New York, USA

About this role

Design and construct evaluation frameworks and benchmarks that measure reasoning depth, interaction quality, reliability, and operational impact; explore adversarial robustness, longitudinal tracking, and human-in-the-loop assessment to quantify emergent capability.

Skills for this role

BenchmarkingEvaluation FrameworksTest SuitesExperimental DesignStatistical AnalysisReproducible ExperimentsModel EvaluationAdversarial Robustness TestingLongitudinal Performance TrackingHuman-in-the-loop EvaluationCompound AI SystemsAgentic CollaborationEnsemblingReActGraph-of-ThoughtsPrototypingProgrammingData AnalysisChatGPTCursorPerplexityDesigning MetricsResearch Publication