About this role
Zoom’s Audio team is seeking an Audio AI Engineer in Singapore to develop AI-driven features for real-time audio and video communication. The role focuses on low-latency streaming speech models that improve intelligibility and naturalness while preserving speaker identity and expressiveness.
Responsibilities
- Research, design and develop algorithms for accent conversion, voice conversion, speech synthesis and automatic speech recognition.
- Prototype end-to-end audio models and work with product and platform teams to integrate them into real-time communication systems.
- Evaluate and optimize speech quality, latency, robustness and scalability. Keep current with speech-processing research and contribute through patents and internal knowledge sharing.
Qualifications
- A PhD or equivalent experience in a relevant field involving streaming, accent conversion, voice conversion, text-to-speech or automatic speech recognition.
- Proficiency with deep learning frameworks such as PyTorch or TensorFlow, and programming skills in Python, C/C++ or similar languages.
- Understanding of sequence-modeling architectures, including Transformers, RNNs, diffusion models or conformers.
- Experience developing and deploying low-latency, real-time speech or audio models using streaming architectures and optimized pipelines; familiarity with compression and acceleration methods including quantization, pruning and distillation; and experience with real-time audio systems in networked communication environments.
- The posting calls for publication in top-tier conferences such as ICASSP, Interspeech, NeurIPS and ICLR. More than 2 years of relevant industry experience is considered a plus.
The position is full-time in Singapore. Zoom describes an office-and-remote hybrid approach generally but does not specify this role’s work style. Its benefits program offers options supporting health, work-life balance and community involvement.