About this role
Architect and build custom AI infrastructure and hardware solutions, optimizing performance, power consumption, cost and scalability, and advise on technology and vendor selection and integration. Build evaluation and quality engineering capabilities for AI and agentic systems, establishing measurable quality gates and automated evaluation suites using tools like LangSmith, Arize Phoenix, Weights & Biases, MLflow, and OpenAI Evals.
Skills for this role
Large Language Models (LLMs)LangSmithBraintrustArize PhoenixWeights & Biases WeaveMLflowOpenAI EvalsCodexClaude CodeCursorPythonTest automationData analysisStatisticsExperimentationLLM evaluationRAG evaluationAgent trajectory analysisBenchmark designError taxonomyTracingObservabilityAdversarial testingSafety testingProduction monitoringRed teamingRoot-cause investigationCI/CD integrationRelease managementIncident management」「Quality engineering""Evaluation engineering""Hardware in