About this role
Role Overview: As an Engineer – Perception & Learning, you will work at the intersection of advanced sensor hardware, machine vision, and real-time robot state estimation. You will collaborate with motion planning and application teams to build spatial and geometric perception systems, enabling humanoid robots to perceive, navigate, and interact with the physical world. Your work involves building visual intelligence pipelines to translate raw camera inputs into actionable motion planning commands for mission-critical tasks in manufacturing and other sectors. You will also contribute to visual pipelines used to train and condition large VLM/VLA and World models. Responsibilities: Develop traditional machine vision algorithms for object detection, image segmentation, and visual servoing using robotic arms. Research and implement algorithms for various camera types, including event-based, ToF, and stereo cameras. Collaborate with controls, motion planning, and deployment teams to develop robust autonomic stacks for loco-manipulation and grasping. Work on cutting-edge technology involving VLAs and World Models to advance the autonomous stack for Svaya Robots. Collaborate across disciplines to evolve a reliable robot platform. Required Qualifications: MS or MTech (or higher) with a specialization in Robotics, Computer Vision, and/or ML. Proven experience in developing fundamental computer vision and machine learning algorithms. Strong understanding of the Point Cloud Library, filtering algorithms, and modern vision approaches including diffusion and attention models. Background in advanced deep learning and LLMs with demonstrable projects. Familiarity with C++, multi-threaded programming, and high-performance optimization for GPU deployment. Ability to work in a fast-paced startup environment with a strong learning mindset and good communication skills. Preferred Qualifications: Experience in Reinforcement Learning, Vision Language Action (VLA) models, development of simulation environments, building World Models, and working with robotic arms or humanoids.