About this role
ServiceNow seeks a Senior Research Scientist, Agentic AI in Hyderabad to develop the evaluation layer used to validate AI agents before they reach customers. The role owns core components of an automated platform that scores large volumes of agent traces, including pipelines that run traces through model-based judges and scoring logic that produces actionable results. The posting labels the work arrangement “Flexible” but does not specify an on-site, hybrid, or remote schedule.
Responsibilities
- Design evaluation methodologies and benchmarks for agent reasoning, planning, tool use, reliability, and safety, using LLM-as-a-Judge, trajectory-based, and human evaluation.
- Take work from research question through prototype to shipped feature. Build and harden customer-facing evaluation pipelines and scoring logic.
- Curate synthetic and real-world datasets, and assess evaluator consistency and agreement with human labels.
Requirements
- At least 5 years in machine learning, applied AI, prompt engineering, or agentic AI, including experience shipping a product or feature to users.
- Strong Python skills and practical depth in agentic AI and context engineering, including planning, reasoning, memory, tool use, retrieval, and long-context approaches.
- Experience designing evaluation methodologies, rather than only running evaluations; hands-on production experience with LLM APIs, prompt engineering, structured output, and cost and latency tradeoffs.
- Ability to communicate clearly with technical and non-technical audiences.
Preferred
Experience with AI-assisted development tools such as Claude Code or Windsurf.