About this role
XenonStack seeks a Machine Learning Engineer, Reinforcement Learning to build adaptive AI agents for enterprise decision-making and workflows in Mohali, India. The full-time position lists a salary of 8–15 LPA.
Responsibilities
- Design, implement and train reinforcement learning algorithms, including PPO, A3C, DQN and SAC. Build custom simulations of business processes and design reward functions that balance efficiency, accuracy and long-term value.
- Develop production-ready agents for dynamic decisions and task orchestration. Integrate RL models with LLMs, knowledge bases and external tools, and build multi-agent systems for collaboration, negotiation and coordination.
- Deploy agents on cloud and hybrid infrastructure, optimize distributed training and inference with frameworks such as Ray RLlib and Horovod, and apply quantization, ONNX and TensorRT for scalable deployment.
- Evaluate agent robustness, reliability and interpretability; implement fail-safes, guardrails and observability; and document experiments and lessons learned.
Requirements and preferences:
- The job-information panel lists 1–3 years of work experience, while the role description specifies 2–5 years applying RL to enterprise-grade systems and the technical requirements specify 2–5 years of hands-on experience with RL frameworks such as Ray RLlib, Stable Baselines, PyTorch RL and TensorFlow Agents.
- Strong Python skills and proficiency with PyTorch or TensorFlow are required, along with experience training RL algorithms, familiarity with simulation environments such as Gymnasium, Isaac Gym or Unity ML-Agents, and experience in reward modeling and optimization.
- Exposure to AWS, GCP or Azure, Docker, Kubernetes and ML CI/CD is sought. Multi-agent and collaborative RL knowledge is a strong plus; familiarity with LLMs and RLHF is desirable.
- The role calls for analytical problem-solving, collaboration across AI, data and platform teams, production-focused engineering, and attention to responsible AI, including bias mitigation, fairness and transparency.
Skills for this role
Reinforcement LearningPythonPyTorchTensorFlowPPOA3CDQNSACActor-Critic methodsRay RLlibStable BaselinesPyTorch RLTensorFlow AgentsGymnasiumIsaac GymUnity ML-AgentsReward modelingSimulation environmentsMulti-agent systemsLLMsRLHFAWSGCPAzureDockerKubernetesCI/CDHorovodQuantizationONNXTensorRT