About this role
Morningstar seeks a Senior Machine Learning Engineer to build and scale its Unified AI/ML Data Collection Platform. The role brings existing AI/ML and LLM-driven data systems into a cohesive platform for data pipelines, model lifecycle management, evaluation and production deployment. The engineer will work with ML engineers, product managers, researchers and business stakeholders, mentor engineers and promote reliable, maintainable and cost-efficient engineering practices.
Responsibilities
- Design scalable data collection and enrichment workflows for structured and unstructured sources, including data ingestion, feature management, model training and scalable inference.
- Build LLM capabilities such as retrieval-augmented generation, prompt orchestration, entity extraction, summarization, classification and automated validation. Develop agentic workflows connecting models to internal tools, APIs, knowledge stores, data sources and workflow systems.
- Deploy and optimize ML and LLM models using versioning, CI/CD, experiment tracking, model registries, rollout strategies and rollback mechanisms.
- Develop evaluation frameworks covering extraction quality, hallucination risk, grounding, consistency, latency, coverage and data reliability. Implement logging, tracing, alerting, cost tracking, performance monitoring, drift detection and reliability dashboards.
- Engineer distributed, event-driven, cloud-native systems using asynchronous processing, message queues, containers and orchestration for high-volume workloads. Evaluate emerging AI/ML tools and approaches to improve automation and developer productivity.
Qualifications
- Bachelor’s or Master’s degree in Computer Science, Data Science, Mathematics or a related technical field; 5+ years of experience in machine learning engineering or data science focused on ML systems, ML platforms or distributed systems.
- Experience building production-grade ML systems; hands-on MLOps experience with CI/CD, monitoring and experiment tracking; strong Python and SQL or similar programming skills.
- Experience with cloud platforms and containerization, such as AWS, GCP or Azure, Docker and Kubernetes; production LLM systems involving RAG, embeddings and vector databases. Requires distributed-systems judgment, technical problem-solving, communication and collaboration across global teams.
Work arrangement: Based in Vashi, Navi Mumbai, with four days per week in the office. Office hours are 12:00pm–5:00pm, followed by remote support from 7:30pm–10:00pm. Limited corporate travel may be required.