About this role
Research and develop algorithms for accent conversion, voice conversion, speech synthesis, and ASR on low-latency streaming architectures; prototype end-to-end audio models and integrate them into real-time communication systems while evaluating and optimizing quality, latency, robustness, and scalability.
Skills for this role
Accent conversionVoice conversionSpeech synthesis (TTS)Automatic Speech Recognition (ASR)Low-latency streaming architecturesEnd-to-end audio modelingPyTorchTensorFlowPythonC/C++TransformersRNNsDiffusion modelsConformersModel compressionQuantizationPruningDistillationReal-time audio systemsOptimized inference pipelinesModel evaluation and optimizationPublishing in top-tier conferences (ICASSPInterspeechNeurIPSICLR)