About this role
Own the Generation half of Document Intelligence, turning multimodal primitives (keyframes, transcripts, OCR) into schema-valid entities like SOPs; design and iterate prompts, define per-entity schemas and validators, build evaluation datasets and rubrics, assemble multimodal context windows, and choose models and token budgets for production.
Skills for this role
Prompt engineeringStructured outputTool useRetrieval-augmented generation (RAG)Retrieval systemsEval dataset creationEvaluation rubricsDeploying LLMs to productionMultimodal AIOCRDocument understandingVideo understandingFine-tuningLLM observabilitySchema validationContext assemblyModel selectionCost and latency optimization