About this role
Role Overview: We are seeking a highly skilled Data Scientist with deep expertise in Databricks, Machine Learning, and AI-driven analytics to join our growing data organization. You will design, build, and deploy scalable models and data workflows that power enterprise insights, automation, and decision-making, working across engineering, analytics, and business teams to transform raw data into intelligent, production-ready solutions. Key Responsibilities: Develop, train, and deploy machine learning models using Databricks notebooks, MLflow, and the Lakehouse architecture. Build scalable ETL/ELT pipelines leveraging Delta Lake, PySpark, and Databricks workflows. Implement AI/LLM-based solutions, including retrieval-augmented generation (RAG), vector search, and enterprise agent workflows. Partner with data engineering to optimize datasets for analytics, modeling, and real-time inference. Conduct exploratory data analysis (EDA), feature engineering, and statistical modeling. Deploy models into production using Databricks Model Serving, serverless compute, or API endpoints. Ensure governance, security, and compliance using Unity Catalog, including implementing Role-Based Access Control (RBAC) and Attribute-Based Access Control (ABAC). Required Qualifications: Bachelor’s or Master’s degree in Data Science, Computer Science, Statistics, or related field. 3–7+ years of experience building machine learning models in Python (Pandas, Scikit-learn, PySpark, TensorFlow, or PyTorch). Hands-on experience with Databricks (notebooks, Delta Lake, MLflow, Databricks SQL). Strong understanding of Lakehouse architecture, distributed computing, and scalable data processing. Experience deploying ML models into production. Proficiency in SQL and Python. Familiarity with LLMs, embeddings, vector databases, or AI agent frameworks. Preferred Qualifications: Experience with Databricks Model Serving, Vector Search, or serverless warehouses. Background in NLP, deep learning, or generative AI. Experience integrating Databricks with SAP, Snowflake, or enterprise BI tools. Knowledge of MLOps best practices, CI/CD pipelines, and cloud platforms (Azure, AWS, or GCP).