About this role
Build and scale the research infrastructure that enables experiments on LLM agents (experiment harnesses, evaluation pipelines, and tooling), partner with researchers to design and evaluate experiments on real incident data, and productionize improvements to agent behavior for enterprise customers.
Skills for this role
PythonFastAPIBackend developmentSoftware engineeringExperimentationExperiment designResearch infrastructureAWSAmazon ECSKubernetesPostgresS3ObservabilityMonitoring toolsLarge Language Models (LLMs)Agent architecturesPrompting strategiesEvaluation pipelinesScaling research prototypesAgentic AI systemsReasoning workflowsTool-use policies