About this role
Design and deliver production-grade AI systems for mission-critical federal environments. Own retrieval-augmented generation (RAG) pipelines, evaluation systems, and performance optimization, building secure, scalable, observable services across the AI application lifecycle.
Responsibilities
- Build enterprise data ingestion, embedding, indexing, and retrieval pipelines, including document chunking, metadata filtering, hybrid search, and retrieval tuning.
- Develop AI APIs and services; manage prompts, model configurations, and structured outputs across development and production.
- Implement logging, tracing, and metrics for latency, throughput, token usage, cache hit rate, and errors. Improve performance through caching, parallelization, and asynchronous workflows.
- Build automated evaluation using benchmark datasets, LLM-as-a-judge methods, regression tests, and human evaluation. Deploy, monitor, debug, and support production AI services, including incident response.
Required qualifications
- 3+ years of experience building production systems, including AI/ML applications, and a bachelor’s degree.
- Experience with LLM systems using RAG frameworks such as LangChain or LlamaIndex, or custom pipelines and vector stores; API development with FastAPI, Flask, or Node.js; cloud-native AWS architectures; and agent orchestration patterns.
- Experience with observability, performance optimization, evaluation pipelines, and production AI operations. Knowledge of AI security, including prompt injection, jailbreak resistance, data leakage prevention, guardrails, secure prompt design, and responsible AI practices.
- Ability to obtain a TS/SCI clearance. Selected applicants are subject to a security investigation and may need to meet classified-information access requirements.
Preferred qualifications
- Experience supporting DoD or other federal missions; high-performance inference with vLLM, Ray Serve, or Triton Inference Server; semantic caching; guardrail frameworks; and enterprise data platforms or knowledge repositories.
- Experience with AWS GovCloud, IL4 or IL5, other secure cloud environments, event-driven messaging, or Model Context Protocol (MCP).
The role is hybrid, based in Washington, DC, with McLean, Virginia, also listed. Hybrid employees are expected to work frequently from a Booz Allen facility and may need to visit customer facilities. The projected annual salary range is $99,000–$225,000 USD. Benefits may include health, life, disability, financial and retirement offerings, paid leave, professional development, and tuition assistance, subject to eligibility.