Yue Zhang

Hello! I am a postdoctoral research associate in the MURGe-Lab led by Prof. Mohit Bansal, at the University of North Carolina, Chapel Hill. I earned my Ph.D. at Michigan State University, where I was advised by Prof. Parisa Kordjamshidi. I was a visiting scholar at Virginia Tech, collaborating with Prof. Lifu Huang. Prior to Ph.D. study, I obtained my master’s degree from Peking University. My research interests include multimodal learning, embodied agents, and spatial reasoning.

Email  /  CV  /  Google Scholar  /  Github  /  LinkedIn

profile photo

News

  • [2026.6] EgoMemReason is accepted to COLM 2026!
  • [2026.6] Five papers (DEER3D, AnchorWeave, VisionCoach, V-Co, and PQSG) were accepted to ECCV 2026, and MedForget was accepted to TMLR.
  • [2026.5] EPiC is accepted to ICML 2026!
  • [2026.4] Two papers were accepted to the ACL main, and two papers will appear in CVPR Findings!
  • [2025.8] Two papers got accepted by EMNLP 2025!
  • [2025.4] Officially Joined MURGe-Lab!

Research Directions

  • Embodied & Vision-and-Language Navigation: building agents that perceive, reason, and navigate in 3D environments by following natural-language instructions, including VLN Survey, SkillNav, Prune-Then-Plan, VLN-Analogy, NavHint, Dual-Action-VLN-CE, VLN-Trans and LOViS.
  • Spatial & 3D Reasoning: equipping (multimodal) LLMs with situated spatial understanding and 3D grounding, and studying when models should abstain, including SpatialUncertain, AViC, DEER-3D, and SPARTUN3D.
  • Video Generation & World Models: controllable, physics- and world-consistent video generation guided by structured priors, verification, and evaluation, including PhyMotion, SketchVerify, AnchorWeave, EPiC, and PQSG.
  • Video Understanding & Reasoning: long-horizon, memory-driven, and spatial-temporal grounded reasoning over videos, including EgoMemReason and VisionCoach.
  • Multimodal Reasoning & Agents: dynamic expert aggregation, tool recruitment, and routing for general multimodal reasoning, including MEXA, DART, and RGD.
  • Trustworthy & Forensic Multimodal AI: interpretable forgery and deepfake detection and multimodal unlearning for safe deployment, including Deepfake-Agent, M2F2-Det, Deepfake-VQA, and MedForget.

Selected Preprints

Seeing Isn’t Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
Yue Zhang, Zun Wang, Han Lin, Yonatan Bitton, Idan Szpektor, Mohit Bansal
Preprint, 2026
ArXiv/ Code
PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation
Yidong Huang*, Zun Wang*, Han Lin, Dong-Ki Kim, Shayegan Omidshafiei, Jaehong Yoon, Jaemin Cho, Yue Zhang, Mohit Bansal
Preprint, 2026
ArXiv/ Code
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
Shoubin Yu*, Yue Zhang*, Zun Wang, Jaehong Yoon, Huaxiu Yao, Mingyu Ding, Mohit Bansal
Preprint, 2026
ArXiv/ Code
Planning with Sketch-Guided Verification for Physics-Aware Video Generation
Yidong Huang, Zun Wang, Han Lin, Dong-Ki Kim, Shayegan Omidshafiei, Jaehong Yoon, Yue Zhang, Mohit Bansal
Preprint, 2025
ArXiv/ Code

Selected Publications

EgoMemReason: A Memory-driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding
Ziyang Wang*, Yue Zhang* Shoubin Yu, Ce Zhang, Zenqi Zhao, Jaehong Yoon, Hyunji Lee, Gedas Bertasius, Mohit Bansal
COLM, 2026
ArXiv/ Code
AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
Zun Wang, Han Lin, Jaehong Yoon, Jaemin Cho, Yue Zhang, Mohit Bansal
ECCV, 2026
ArXiv/ Code
Error-Driven Scene Editing for 3D Grounding in Large Language Models
Yue Zhang, Zun Wang, Han Lin, Jialu Li, Jianing Yang, Yonatan Bitton, Idan Szpektor, Mohit Bansal
ECCV, 2026
ArXiv/ Code
VisionCoach: Reinforcing Grounded Video Reasoning via Visual-Perception Prompting
Daeun Lee, Shoubin Yu, Yue Zhang, Mohit Bansal
ECCV, 2026
ArXiv/ Code
Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation
Atin Pothiraj, Jaemin Cho, Yue Zhang, Elias Stengel-Eskin, Mohit Bansal
ECCV, 2026
ArXiv/ Code
V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising
Han Lin, Xichen Pan, Zun Wang, Yue Zhang, Chu Wang, Jaemin Cho, Mohit Bansal
ECCV, 2026
ArXiv/ Code
Hierarchy-Aware Multimodal Unlearning for Medical AI
Fengli Wu*, Vaidehi Patil*, Jaehong Yoon, Yue Zhang, Mohit Bansal
TMLR, 2026
ArXiv/ Code
EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance
Zun Wang, Jaemin Cho, Jialu Li, Han Lin, Jaehong Yoon, Yue Zhang, Mohit Bansal
ICML, 2026
ArXiv/ Code
Breaking Down and Building Up: Mixture of Skill-Based Vision-and-Language Navigation Agents
Tianyi Ma, Yue Zhang, Zehao Wang, Parisa Kordjamshidi
ACL (Oral), 2026
ArXiv/ Code
Routing with Generated Data: Annotation-Free LLM Skill Estimation and Expert Selection
Tianyi Niu, Justin Chih-Yao Chen, Genta Indra Winata, Shi-Xiong Zhang, Supriyo Chakraborty, Sambit Sahu, Yue Zhang, Elias Stengel-Eskin, Mohit Bansal
ACL (Oral), 2026
ArXiv/ Code
Deepfake-Agent: Aggregating Semantic Forgery Clues for Generalizable Detection
Xiao Guo, Yue Zhang, Mohit Bansal, Xiaoming Liu
CVPR Findings, 2026
ArXiv/ Code
Prune-Then-Plan: Step-Level Calibration for Stable Frontier Exploration in Embodied Question Answering
Noah Frahm, Prakrut Patel, Yue Zhang, Shoubin Yu, Mohit Bansal, Roni Sengupta
CVPR Findings, 2026
ArXiv/ Code
DART: Leveraging Multi-Agent Disagreement for Tool Recruitment in Multimodal Reasoning
Nithin Sivakumaran, Justin Chih-Yao Chen, David Wan, Yue Zhang, Jaehong Yoon, Elias Stengel-Eskin, Mohit Bansal
EACL, 2026
ArXiv/ Code
MEXA: Towards General Multimodal Reasoning with Dynamic Multi-Expert Aggregation
Shoubin Yu*, Yue Zhang*, Ziyang Wang, Jaehong Yoon, Mohit Bansal
EMNLP Findings, 2025
ArXiv/ Code
Vision-and-Language Navigation with Analogical Textual Descriptions in LLMs
Yue Zhang, Tianyi Ma, Zun Wang, Yanyuan Qiao, Parisa Kordjamshidi
EMNLP, 2025
ArXiv/ Code
Rethinking Vision Language Model in Face Forensic: Multi-modal Interpretable Forged Face Detector
Xiao Guo, Xiufeng Song, Yue Zhang, Xiaohong Liu, Xiaoming Liu
CVPR (Oral), 2025
ArXiv/ Code
SPARTUN3D: Situated Spatial Understanding of 3D World in Large Language Models
Yue Zhang, Zhiyang Xu, Ying Shen, Parisa Kordjamshidi, Lifu Huang
ICLR, 2025
ArXiv/ Code
Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models
Yue Zhang*, Ziqiao Ma*, Jialu Li*, Yanyuan Qiao*, Zun Wang*, Joyce Chai, Qi Wu, Mohit Bansal, Parisa Kordjamshidi
TMLR, 2024
ArXiv/ Code
Narrowing the Gap between Vision and Action in Navigation
Yue Zhang, Parisa Kordjamshidi
ACM MM, 2024
ArXiv/ Code
Common Sense Reasoning for Deepfake Detection
Yue Zhang, Ben Colman, Xiao Guo, Ali Shahriyari, Gaurav Bharaj
ECCV, 2024
ArXiv/ Code
NavHint: Vision and Language Navigation Agent with a Hint Generator
Yue Zhang, Quan Guo, Parisa Kordjamshidi
EACL Findings, 2024
ArXiv/ Code
VLN-Trans: Translator for the Vision and Language Navigation Agent
Yue Zhang, Parisa Kordjamshidi
ACL (Oral), 2023
ArXiv Code
LOViS: Learning Orientation and Visual Signals for Vision and Language Navigation
Yue Zhang, Parisa Kordjamshidi
COLING (Oral), 2022
ArXiv/ Code
Explicit Object Relation Alignment for Vision and Language Navigation
Yue Zhang, Parisa Kordjamshidi
ACL SRW, 2022
ArXiv/ Code
Towards Navigation by Reasoning over Spatial Configurations
Yue Zhang, Quan Guo, Parisa Kordjamshidi
ACL workshop on SpLU-RoboNLP, 2021
ArXiv/ Code

Professional Service

  • [2024.7] Co-organizer of SpLU-RoboNLP @ ACL 2024
  • [2022.11] Invited talk at Sichuan University
  • Reviewer for ACL, EMNLP, NAACL