About this role
This role involves designing and building reasoning benchmarks across science and engineering domains to evaluate frontier AI systems. You will develop evaluation infrastructure, build AI agents, and contribute to the company's core research agenda in a research-first environment.
Skills for this role
Machine LearningStatisticsProbabilityAI AgentsEvaluation HarnessesSWE-benchSoftware EngineeringResearch Pipelines