About this role
This full-time MLOps role, listed for Mumbai and Noida, owns the platform’s operational ML infrastructure and the lifecycle from model handoff through production deployment, monitoring and retraining. The engineer sets MLOps standards for the squad and ensures models run reliably at scale.
Responsibilities
- Establish and maintain the model registry, experiment tracking, feature store and serving infrastructure. Define standards for model versioning, experiment tracking and production readiness, and build CI/CD pipelines for repeatable model releases.
- Deploy trained models to stable, monitored endpoints. Design serving infrastructure for real-time inference with appropriate latency and throughput; provide rule-based fallbacks when model confidence drops below defined thresholds, plus safe versioning and rollback.
- Monitor accuracy, latency, throughput and data drift; set performance thresholds and alerts. Build automated retraining pipelines triggered by drift or degradation, and connect application-layer validator decisions to retraining workflows.
- Own production monitoring, alerting, runbooks and incident response. Coordinate with the Backend Engineer on endpoint reliability and latency, and conduct load testing and capacity planning as transaction volumes grow.
Required qualifications
- B.E., B.Tech or M.Tech in Computer Science, Information Technology or a related field; 5+ years of engineering experience, including at least 3 years in production MLOps or ML infrastructure. Production-scale model deployment and maintenance, model monitoring and automated retraining experience are required.
- Experience with MLflow; model serving frameworks such as TorchServe, TensorFlow Serving or BentoML; Docker, Kubernetes and Terraform or equivalent infrastructure-as-code; GitHub Actions, GitLab CI or equivalent CI/CD; Prometheus, Grafana or equivalent monitoring; and data drift detection frameworks. Python proficiency, familiarity with Bash automation and experience with at least one of AWS, Azure or GCP are required.
Preferred qualifications
The Databricks MLOps stack is strongly preferred, and Databricks Model Serving is a plus. Other preferences include enterprise real-time inference, feature stores, regulated-domain experience such as fintech or compliance where explainability and auditability matter, A/B testing, shadow deployments, and MLOps open-source contributions or community knowledge sharing.