About this role
Lead the architecture and scaling of automated safety evaluation and enforcement platforms for generative AI, transitioning bespoke evaluation scripts to high-throughput, low-latency LLM-as-a-judge autorater infrastructure. Act as a tech lead across high-scrutiny domains (Health, Civics and Elections, AI Ecosystem) to ensure models have authoritative consensus, robust safety guardrails, and sub-100ms enforcement mitigations.
Skills for this role
Generative AILarge Language Models (LLMs)LLM-as-a-judge autorater systemsAutomated safety evaluationAutomated enforcement platformsReal-time enforcement systemsContinuous regression testingData analysisSignal analysisModel developmentCybersecuritySQLPythonGolangFraud investigationsRisk managementThreat analysisEncrypted communicationsAnti-abuse messaging solutionsML reputation modelingDataset developmentBenchmarking and metricsCross-functional stakeholder communicationSafety policy and enforcement model developmentOperational workflow automation