About this role
AuxoAI is hiring a Senior Applied AI Engineer to design and deploy production-grade computer vision systems that perform reliably in real-world environments. The role owns end-to-end visual intelligence systems—from system design through deployment and real-world performance—and combines deep learning, classical computer vision, geometric methods and multimodal reasoning. Work includes integrating visual perception and understanding into larger AI platforms and agent-based workflows.
Responsibilities
- Develop systems for object detection, segmentation and tracking; scene understanding and structured perception; and video understanding and temporal reasoning.
- Build and optimize CNNs, vision transformers, and detection and segmentation models, including ResNet, EfficientNet, ViT, Swin, DeiT, YOLO, DETR and Mask R-CNN. Develop vision-language capabilities such as CLIP-style models, visual grounding and captioning.
- Implement multi-object tracking with approaches such as SORT, DeepSORT and ByteTrack; feature matching and representation learning; and temporal modeling using RNNs or video transformers. Apply camera calibration, epipolar geometry, pose estimation, and 3D reconstruction or depth estimation where relevant.
- Optimize latency, throughput and scalability for real-time, edge and distributed deployment. Build annotation, dataset-curation and synthetic-data pipelines, and integrate vision systems into multimodal AI, agent-based and decision-making workflows.
Requirements
- At least 5 years building computer vision systems in production environments; the listing displays a 5–10 year experience range. Strong PyTorch or TensorFlow experience and hands-on work with detection, segmentation or tracking, as well as model training, fine-tuning and evaluation.
- Strong understanding of representation learning, contrastive and focal losses, and evaluation metrics including mAP, IoU, precision and recall. End-to-end deployment experience is important; experience limited primarily to academic projects or model experimentation may not be a fit. The structured posting lists a bachelor's degree credential.
Preferred
- Multimodal vision-language work and familiarity with CLIP, BLIP or Flamingo; 3D vision using NeRFs, SLAM or point clouds; action recognition or event detection; active learning, hard negative mining, large-scale datasets or distributed training pipelines.
Location: Mumbai, Bangalore, Hyderabad or Gurgaon, India. Hybrid, with three days per week in the office.