About this role
Build environments, evals, and training systems that convert recorded enterprise data into replayable environments and eval sets; design graders, mine tasks and trajectories, and train and evaluate agents over long horizons while scaling evaluation and training runs for enterprise deployment.
Skills for this role
Reinforcement LearningLLM post-trainingEvaluation InfrastructureAgent EngineeringEnvironment DesignReward DesignContext EngineeringGrader DesignMining Tasks from Historical WorkflowsCreating Eval SetsModel EvaluationInference Cost OptimizationLong-horizon Agent TrainingObservabilityReproducibilityScalable Training SystemsProduction Software EngineeringProcessing LargeMessy DatasetsResearch-to-Production Implementation