I am a final-year PhD student in Computer Science at Queen Mary University of London (QMUL), working with Prof. Ioannis Patras and supported by the Queen Mary Principal's Scholarship. I also collaborate closely with Dr. Ziquan Liu and Prof. Shaogang Gong. Previously, I earned an MSc in Control Science and Engineering from Nanjing University of Information Science & Technology (NUIST), where I worked with Prof. Qingshan Liu.
My research background spans human facial behaviour understanding, vision-language models, and generative models. My current work focuses on generative video models, particularly: (1) long-video generation, with an emphasis on preserving temporal consistency and mitigating error accumulation; (2) real-time video generation using self-rollout autoregressive diffusion models that synthesize each video segment faster than its playback duration; and (3) controllable video generation conditioned on text, audio, and action signals, including interactive world models and audio-driven avatars.
Outside of research, I am passionate about rock music, especially progressive rock, post-rock, Britpop, gothic rock, funk rock, and shoegaze. I enjoy playing guitar and singing, as well as staying active through hiking, cycling, and fitness.