About this role
RapidClaims is a leader in AI-driven revenue cycle management for the US healthcare sector. The company is seeking a Senior Machine Learning Engineer to own the end-to-end applied LLM, retrieval, and evaluation layer of their healthcare AI platform. This is a production-focused engineering role responsible for building scalable, auditable, and cost-efficient systems that automate revenue cycle workflows including medical coding, claim edits, and denials management. Responsibilities include: Self-Hosted LLM Infrastructure: Deploying, fine-tuning, and operating open-source models (e.g., Llama, Qwen, MedGemma) using vLLM/SGLang/TensorRT-LLM, managing GPU economics, and executing fine-tuning workflows (SFT, LoRA, QLoRA, DPO). Knowledge Graphs & Retrieval: Designing knowledge graphs for medical standards (ICD-10, CPT, etc.) and building embedding-based retrieval systems over clinical notes and payer policies. Evaluation & Monitoring: Building continuous evaluation pipelines, monitoring production drift, and tracking business metrics like coding accuracy and latency. LLM Systems & Prompt Engineering: Designing context pipelines, implementing structured outputs, and managing prompts as a versioned layer. Agentic Workflows: Building MCP servers and multi-step agent workflows with human-in-the-loop checkpoints. Required Qualifications: 5+ years of experience in ML/AI engineering, including 6+ months in production LLM systems. Proficiency in Python, PyTorch, and Hugging Face. Hands-on experience with self-hosted LLM deployment and evaluation infrastructure. Strongly Preferred: Experience with fine-tuning on domain-specific corpora, graph databases (Neo4j, ArangoDB), vector databases, hybrid search, and LLM observability tools. Familiarity with healthcare or regulated domains is a plus.