About this role
Plaud seeks an on-site Machine Learning Engineer (Speech/Audio) in Singapore to support speech recognition model development through large-scale data pipelines, model fine-tuning and evaluation. The role is part of its Global Product R&D Center.
Responsibilities
- Own speech and audio data pipelines for collection, cleaning, filtering, labeling, augmentation and quality control at model-training scale. Work with senior speech engineers on mining terms and hotwords from ASR output.
- Contribute to training, fine-tuning and evaluating speech/language models alongside domain specialists. Improve recognition of code-switching, names and product terms, and adapt models with scenario-specific data to recognize industry terminology across languages and business verticals.
- Build keyword- and domain-lexicon-based test sets and evaluation frameworks; benchmark internal models against open-source and commercial alternatives.
Required qualifications
- At least 1 year of hands-on experience in speech, machine learning or large-scale data engineering.
- Experience in at least one of: ASR or SpeechLLM training, fine-tuning or evaluation; LLM or general machine learning model training or fine-tuning; or large-scale audio, video or text data pipelines, such as tens of thousands of hours of audio or TB-scale multimodal data.
- Solid Python and PyTorch fundamentals, plus experience with distributed data processing, such as Spark or Ray.
Preferred qualifications
- Familiarity with SpeechLLM or speech self-supervised learning, or exposure to StepAudio or Qwen3-Omni; experience improving ASR through hotwords, contextual biasing or code-switching; speech-related patents or publications at Interspeech, ICASSP or other leading AI venues; or ownership of a speech-data workstream spanning tens or hundreds of thousands of hours.
The full-time role offers market-competitive compensation, an employee stock ownership plan, AI productivity tools, equipment, company events, and medical insurance and WICA coverage. No salary amount is specified.
Skills for this role
PythonPyTorchSpeech recognitionASRSpeechLLMLarge language modelsMachine learningSpeech model fine-tuningModel evaluationAudio data pipelinesDistributed data processingSparkRayData labelingData augmentationData quality controlDomain adaptationCode-switchingHotword miningContextual biasingSpeech self-supervised learningStepAudioQwen3-Omni