About this role
Build and own evaluation and experimentation systems for a scientific agent, working with the Principal Scientist to design reproducible experiments, evaluation datasets, replay and trajectory analysis, control policies, and production improvements. Ship end-to-end solutions including data pipelines, dataset and artifact versioning, experiment tracking, and tooling to enable fast, reliable comparisons between models, prompts, and policies.
Skills for this role
PythonTestingTypingPackagingCode reviewContinuous Integration (CI)ML evaluationExperimentationOffline evaluation harnessesDataset designSplit designMetric designData pipelinesSchema designData versioningAirflowPrefectDagsterRayObject storageReproducibilityPinned environmentsSeeded runsArtifact versioningExperiment trackingTrajectory analysisReproducible replayControl policiesLLM agentsTool calling (agents)