About this role
Based in Bengaluru, this role supports LexisNexis Legal & Professional’s APAC shift toward machine learning, predictive analytics and agentic development. The engineer will deliver reliable models and insights for regional stakeholders on modern data platforms, working across business units and geographies.
Responsibilities
- Clean, transform and join datasets with SQL and Python, addressing missing data, outliers, normalization and data leakage. Build business-focused models using regression, classification, time series analysis and statistical testing, with attention to performance and interpretability.
- Engineer features through aggregation, encoding, interaction terms and time windows; assess their importance and stability. Select algorithms, tune hyperparameters, address bias, variance and class imbalance, and use regularization and ensembling where appropriate.
- Design and validate experiments using train/validation/test splits, cross-validation, statistical significance and confidence intervals. Evaluate models with precision, recall, F1, ROC AUC, calibration, confusion matrices and problem-appropriate cost-sensitive metrics.
- Collaborate with data scientists to turn statistical insights into deployable predictive models. Productionize models for speed and reliability, monitor drift and performance, and manage A/B rollouts using Databricks and related tooling.
- Develop autonomous, adaptive workflows and automated pipelines for data processing, deployment and monitoring, including work with Databricks and Microsoft Fabric. Document methods, assumptions, analyses and business implications for technical and non-technical audiences. Collaborate across regions and support junior analysts and team development.
Requirements and preferences:
- A bachelor’s degree in Data Science, Statistics, Computer Science or a related field is required; a master’s degree is preferred. At least five years of experience in data and machine learning or closely related roles is required, including independent end-to-end delivery from development and testing through production.
- Expertise with Databricks, Microsoft Fabric and Power BI is required to help operate, maintain and provide break/fix coverage for core data platforms. Candidates must be able to discuss SQL, Python, data modeling and evaluation, statistical foundations and ML algorithms in a technical interview. Strong communication and cross-functional collaboration are expected.
- Experience with NLP, agentic models, generative AI, workflow automation, ETL, PowerApps, MLflow, pipeline orchestration, version control, CI/CD and model monitoring is advantageous.
The posting mentions flexible working arrangements, wellbeing initiatives, study assistance, sabbaticals and learning resources; the recruiter will provide location-specific benefits details.