About this role
The Data Scientist will analyze large datasets related to CAS products and scientific domains, including chemistry and life sciences. Working with scientists and other stakeholders, they will translate business problems into data-driven solutions, communicate findings, help shape data aggregation strategy, and develop the capabilities, partnerships and tools needed for complex analytics and AI/ML projects.
Responsibilities
- Build generative AI solutions, including fine-tuned large language models, text generation systems and chatbots. Improve LLM performance on question answering, information extraction, long-form generation, translation and summarization, with particular attention to factual accuracy in scientific tasks.
- Mine large data stores; clean and validate data; analyze errors and improve models. Design experiments and select algorithms, including NLP and named entity recognition for text extraction. Apply machine learning, deep learning and statistical techniques to forecasting, prediction, attribution, recommendations, user-experience measurement, journey orchestration and personalization.
- Use exploratory analysis and data-driven optimization, and present model outcomes and insights through reports and visualizations. Implementing MLOps practices for automated deployment, monitoring and management of production models is desirable.
Requirements
- A bachelor's degree or higher in computer science, statistics, mathematics, engineering or a related field. The listing states 5–10 years of experience; the detailed requirements specify 5–8 years of post-degree work with diverse datasets.
- Expertise in advanced analytics, data mining, modeling and AI/ML; skills in generative AI, LLMs, RAG, fine-tuning and prompt engineering. Approaches mentioned include chain-of-thought prompting, self-consistency, chunking, dialogue resolution and reinforcement learning with human feedback. Candidates must be able to develop LLM evaluation frameworks using human experts and quantitative metrics.
- Experience building public-cloud applications, with AWS preferred; proficiency in Java, Scala, JavaScript, TypeScript, Python or R and Linux/Unix; and experience with relational, NoSQL, RDF/triple-store and vector databases. Communication, initiative and a proactive approach are required. Power BI experience is desirable.
Skills for this role
Data ScienceGenerative AILarge Language ModelsRAGLLM fine-tuningPrompt EngineeringChain-of-thought promptingReinforcement Learning from Human FeedbackLLM EvaluationNatural Language ProcessingNamed Entity RecognitionMachine LearningDeep LearningStatistical ModelingData MiningExploratory Data AnalysisForecastingRecommendation SystemsMLOpsCI/CDAWSPythonRJavaScalaJavaScriptTypeScriptLinuxUnixRelational Databases