New opportunity

Machine Learning Operations Engineer II

S&P Global · Cambridge; New York, United States

About this role

Kensho, S&P Global’s AI innovation hub, seeks a Machine Learning Operations Engineer II to build and support its ML platform. The team develops processes, tooling, and infrastructure that help ML engineers move from research to reliable production models, agents, and products. The role works with ML product and research teams as well as core infrastructure, site reliability, and security teams.

Responsibilities

  • Develop tools, services, and frameworks that make ML workflows robust, auditable, and usable. Work with ML engineers to identify process bottlenecks and improve experimentation, prototyping, production readiness, and developer experience.
  • Provide training and resources on ML best practices. Evaluate open-source and third-party solutions, integrate them into the platform, and encourage adoption across teams.
  • Build scalable, efficient, automated processes for model fine-tuning, reinforcement learning, and evaluation of LLMs and agents. Improve observability of production agentic applications to detect performance, decay, and drift issues. Track emerging tools and frameworks.

Required qualifications

At least 2 years of experience in ML infrastructure, MLOps, ML engineering, or a similar area. Experience managing distributed systems with Kubernetes, including its concepts and trade-offs; understanding of AWS; Python proficiency; familiarity with distributed computing and workflow orchestration, such as Ray and Airflow; software engineering best practices in an ML context; and basic understanding of ML, LLMs, and agents. Candidates should be able to debug distributed systems across infrastructure, networking, and application layers and communicate effectively to drive adoption across teams.

Additional experience of interest: Agentic AI systems and workflows, running workflows on Ray, and MCP server patterns. The team’s tools include Python, Bash, LangGraph, PyTorch, Amazon EKS, Airflow, Jsonnet, Terraform, Git, GitHub, AWS, LangFuse, Sentry, Prometheus, and W&B; it also uses Bedrock and SageMaker.

The position is listed in Cambridge, Massachusetts, and New York, New York. Anticipated base salary is $130k–$175k, with eligibility for an annual incentive bonus and equity plans. Listed benefits include medical, dental, and vision insurance with company-paid premiums, unlimited paid time off, 26 weeks of paid parental leave, a 401(k) with 6% employer matching, and education assistance.

Skills for this role

MLOpsML infrastructureML engineeringKubernetesAWSAmazon EKSAmazon BedrockAmazon SageMakerPythonDistributed systemsDistributed computingWorkflow orchestrationRayAirflowMachine learningLarge language modelsAI agentsSoftware engineeringDebuggingCommunicationModel fine-tuningReinforcement learningLLM evaluationAgent evaluationObservabilityBashLangGraphPyTorchJsonnetTerraform

Your skill match

Checking your profile…

YOUR NEXT STEP

Get interview-ready for this role

A focused preparation guide, built around this job’s responsibilities and requirements.

✦ AI-generated guide
Preparation suggestions, not the employer’s actual interview questions. Always check the original posting for current requirements.

Loading this role’s preparation guide…

KEEP EXPLORING

Similar AI jobs

Related skills and specializations in United States. Matched to this role, not your profile.

Explore more jobs
Finding similar opportunities…