About this role
Research role focused on video generation and understanding, including parsing semantics, state tracking, consistency modeling, unified generation-understanding architectures, and causal consistency control; collaborate with Agent/RL teams to create a closed loop of generation → understanding → control → feedback and validate controllable interactive capabilities in game scenarios.
Skills for this role
Video understandingVideo predictionReinforcement learningMultimodal generationMultimodal understandingVideo generationVision-language models (VLMs)Diffusion modelsInteractive video generationControllable video generationModel training pipelinesPythonPyTorchGame AISimulation environmentsClosed-loop systemsState trackingConsistency modelingCausal consistency controlLow-latency feedbackAgent collaborationEvaluation systems