About this role
This researcher will develop reinforcement learning algorithms for multimodal models within Tencent’s Technology Engineering Group. The position is full-time and on-site at Singapore-CapitaSky.
Responsibilities
- Research reinforcement learning for diffusion models used in image and video generation, autoregressive models for multimodal understanding, and unified multimodal frameworks.
- Design reinforcement learning training frameworks and reward-modeling strategies for efficient, large-scale training. Improve training stability and address reward hacking.
- Explore new reinforcement learning approaches that learn more directly and efficiently from environmental feedback.
Required qualifications
- Bachelor’s degree or higher in Computer Science or a related field.
- Strong research capabilities, demonstrated by publications at top conferences such as ICML, NeurIPS, ICLR, CVPR, ICCV, ECCV, or SIGGRAPH.
- Strong engineering and programming skills, with experience implementing deep learning systems, optimizing model training and inference, using CPU/GPU acceleration, and conducting distributed training and inference.
Preferred qualifications
- Experience with diffusion models, autoregressive models, or text-to-image or text-to-video generation.
- Participation in ACM/NOIP Informatics Olympiad competitions.
Skills for this role
Reinforcement LearningMultimodal ModelsDiffusion ModelsAutoregressive ModelsImage GenerationVideo GenerationMultimodal UnderstandingReward ModelingDeep Learning SystemsModel TrainingInference OptimizationCPU AccelerationGPU AccelerationDistributed TrainingDistributed InferenceProgrammingResearch