About this role
Lead the design, deployment, and operation of a self-hosted LangSmith evaluation platform, build LLM evaluation tooling and observability (tracing, metrics, alerting) for AI systems in production, and manage a small engineering team responsible for platform reliability and on-call operations.
Skills for this role
PythonLangSmithLangFuseBraintrustWeights & BiasesKubernetesCloud InfrastructureRelational DatabasesAnalytical DatastoresSingle Sign-On (SSO)Access ControlData IsolationCapacity ManagementVendor ManagementDeploymentOperationsUpgradesSecurityEvaluation ToolingLLM EvaluationLLM-as-judgeDataset VersioningRegression TestingContinuous Integration (CI)TracingMetricsAlertingLatency MonitoringCost MonitoringModel Versioning