New opportunity

Sr. Staff AI Engineer, AI Agent Platform (Agent & Evaluation Harness)

GEICO · New York City, United States of America

About this role

GEICO seeks a Sr. Staff AI Engineer to own agent behavior and evaluation methodology for its AI Agent Platform. The platform supports reusable capabilities for workflows including claims, underwriting, and internal associate tools. The role works alongside backend platform engineers responsible for execution, scaling, and operations.

Responsibilities

  • Design reusable agent patterns for planning, reasoning, tool selection and calling, memory, and error recovery. Develop prompt and context strategies covering context-window management, summarization, memory, and RAG.
  • Define MCP-based connections to enterprise tools and systems, reusable packaged agent skills, and evidence-based approaches to sub-agent and multi-agent orchestration. Assess new models for capability, latency, and cost.
  • Build an evaluation harness that teams can use to define, run, and compare offline benchmarks, simulated-user evaluations, and online quality monitoring. Establish metrics, rubrics, golden datasets, LLM-as-judge graders, and human review; account for statistical variance, dataset contamination, and grader bias.
  • Analyze traces and evaluation results, run experiments on prompts, models, architectures, and tools, and turn findings into production-ready patterns. Set technical direction, mentor engineers and data scientists, and partner with product and business teams on measurable quality targets.

Minimum qualifications: At least 10 years in software engineering, ML engineering, or applied science, with strong Python proficiency. Requires technical-lead experience delivering significant systems across engineers or teams; hands-on experience taking LLM-based GenAI and agentic applications beyond prototypes; and experience designing experiments and evaluations to improve performance. Requires practical knowledge of prompt and context engineering, tool calling, RAG, agent architectures, and MCP, plus distributed-systems fundamentals and experience with production services at scale.

Preferred qualifications

Experience with LLM or agent evaluation frameworks, human annotation, MCP servers, packaged agent skills, agent frameworks, and evaluation or observability tools. A background in statistics, fine-tuning or reinforcement fine-tuning, cloud AI platforms, public GenAI contributions, and AI security or responsible AI is also preferred.

This is a hybrid, full-time role listed in New York City, Palo Alto, Bethesda, and Seattle. The annual salary guideline is $115,000–$260,000; GEICO will consider sponsoring a qualified applicant for employment authorization. Benefits mentioned include development programs, mentorship, certification assistance, and flexibility.

Skills for this role

PythonGenerative AILLMsAI agentsAgent architecturesPrompt engineeringContext engineeringTool callingModel Context Protocol (MCP)RAGMulti-agent systemsAgent evaluationLLM-as-judgeExperimental designStatistical significanceHuman annotationTrace analysisDistributed systemsConcurrencyFault toleranceAPI designTechnical leadershipGPTClaudeLlamaQwenInspect AIBraintrustDeepEvalLangGraph. Microsoft Agent Framework

Your skill match

Checking your profile…

YOUR NEXT STEP

Get interview-ready for this role

A focused preparation guide, built around this job’s responsibilities and requirements.

✦ AI-generated guide
Preparation suggestions, not the employer’s actual interview questions. Always check the original posting for current requirements.

Loading this role’s preparation guide…

KEEP EXPLORING

Similar AI jobs

Related skills and specializations in United States of America. Matched to this role, not your profile.

Explore more jobs
Finding similar opportunities…