profile photo

Wentao Jiang 蒋文涛

I am a second-year master student at Visual & Data Intelligence Center (VDI) leded by Prof. Jingyi Yu in ShanghaiTech University, advised by Prof. Jingya Wang and Prof. Ye Shi.

I am studying agents and multimodal models particularly focusing on physical real-world applications. My research interests lie in world models, generative models (video & 3D), and multimodal agents.

I believe the ultimate goal of AGI is embodied agents, which requires the integration of many technologies including multimodal understanding, action sequence generation, memory mechanisms, and continual learning. More importantly, it's about how to stably deploy and deliver to users in real world.

My current research agenda encompasses:
1. Multimodal real-time interaction agentic systems (ARFlow, Super Star);
2. harnessing generative models with physical constraints (ARFlow, DSG);
3. Continuously evolving agent systems (Super Star, AgentgenV2).

I am looking for job opportunities for the 27th session. Or potential future Ph.D. opportunities. Feel free to contact me for discussion or collaboration!

Github  /  📧

News
  • [Aug. 2026] We have released the complete source code for Super-Star!
  • [July 2026] Two paper has been accepted to ACM Multimedia 2026 (CCF-A)! Big thanks to my collaborators!
  • [Mar. 2025] Our paper ARFlow is now available on arXiv!
  • [Jun. 2024] Graduated from China University of Mining and Technology as an outstanding graduate !
  • [May 2024] DSG has been accepted to ICML 2024!
  • [Oct. 2023] Join YesAI Lab, ShanghaiTech University to research the guidance and acceleration of diffusion models.
Selected Publications
Super Star : Towards Streaming Real-time Interactive Agents for Digital Humans
Wentao Jiang, Youchen Xie, Haidi Fan, Yajing Chen, Xin Wang, Ye Shi, Jingya Wang
ACM International Conference on Multimedia (ACM MM), 2026
pdf / project page / code

A system framework enables digital humans to interact with users in online real-time based on user multimodal input and generate full-body co-speech gestures, which is trained through our closed-loop self-evolution data flywheel to support continual adaptation to user preferences.

ARFlow: Human Action-Reaction Flow Matching with Physical Guidance
Wentao Jiang, Jingya Wang, Kaiyang Ji, Baoxiong Jia, Siyuan Huang, Ye Shi
ACM International Conference on Multimedia (ACM MM), 2026
pdf / project page / code

A novel paradigm that directly establishes a mapping between the action and reaction distribution, requiring only minor code modifications to significantly enhance the modal fusion capability and computational efficiency of diffusion-based methods. Moreover, we propose a general efficient reprojection guidance method which can incorporate various physical constraints.

Consistency Trajectory Models with Test-time Guidance
Wentao Jiang, Lingxiao Yang, Jingya Wang, Ye Shi
Preprint, 2024

A unified architecture that reduces the cumulative error of the consistency models and prevents the test-time guidance from deviating from the data manifold under Gaussian spherical constraints, enabling image editing, restoration, and deblurring with a small number of steps.

Joint Facial Action Unit Recognition and Landmark Detection via Transformer
Zhiwen Shao*, Wentao Jiang*, Yingjie Xia*
Preprint, 2023

A novel local-global transformer to extract local-global feature, which is good at capturing subtle AU details like vanilla convolution while maintaining the global relational modeling capacity of transformer. Besides, we jointly train facial AU recognition and facial landmark detection, in which the two correlated tasks contribute to each other and further facilitate the learning of local-global feature.



Education
ShanghaiTech University
Master of Science: Computer Science and Technology
September 2024 - Present
YesAI Lab in Visual & Data Intelligence Center (VDI)
Advisor: Prof. Jingya Wang and Prof. Ye Shi
China University of Mining and Technology
Bachelor of Science: Computer Science and Technology
September 2021 - June 2024
Bachelor of Science: Materials and Physics
September 2020 - June 2021
GPA: 90.91/100, Ranking: 1/299
Advisor: Prof. Zhiwen Shao and Prof. Qianqian Xu
Experience
Tencent, Shenzhen, China
Mainly focus on Multimodal Real-time Interactive Agent,
Interactive Video World Models
RedNote, Shanghai, China
Mainly focus on Multimodal Search and Coding Agent,
Agentic RL and Self-evolution
KURN, Shanghai, China
Mainly focus on Robotics
Awards & Honors
  • First Class Scholarship, 2025
  • Outstanding Graduate, 2024
  • National Scholarship, 2021
Skills

Programming: C++, Python, PyTorch, Blender, ROS, Vibe Coding

Languages: Chinese (Native), English (CET-4: 626, CET-6: 564)

Interests: Chinese Chess (Regional Master), Erhu (Grade 10), Basketball, League of Legends


Design inspired by Jon Barron's website.   Since Aug. 2026