Zanyi Wang

I am a MS student in Electrical and Computer Engineering (ECE) at the University of California, San Diego (UCSD).

Previously, I received my B.Eng. in Automation Engineering from Xi’an Jiaotong University (XJTU). I also spent a wonderful time as an exchange student at CUHK.

My research interests lie in Video Understanding and Generative Modeling.

Email  /  HF  /  Github

profile photo
News
  • [2026-07] One paper accepted to ACM MM 2026!
  • [2026-06] One paper accepted to ECCV 2026!
  • [2026-02] One paper accepted to CVPR 2026!
  • [2026-01] One paper accepted to ICLR 2026!
  • [2025-07] One paper accepted to ACM MM 2025!
Selected Publications (* denotes equal contribution)
ReChannel teaser
From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models
Zanyi Wang, Xin Lin, Haodong Li, Dengyang Jiang, Yijiang Li
Preprint, 2026
code / arXiv

Readout, not generation.

AffordanceSAM teaser
AffordanceSAM: Segment Anything Once More in Affordance Grounding
Dengyang Jiang, Zanyi Wang*, Hengzhuang Li, Sizhe Dang, Teli Ma, Wei Wei, Guang Dai, Lei Zhang, Mengmeng Wang
ACM International Conference on Multimedia (ACM MM), 2026
arXiv

Transferring SAM to affordance grounding task and showing robust performance for both seen and unseen actions.

DMDR teaser
Distribution Matching Distillation Meets Reinforcement Learning
Dengyang Jiang, Dongyang Liu, Zanyi Wang, Qilong Wu, Liuzhuozheng Li, Hengzhuang Li, Xin Jin, David Liu, Zhen Li, Bo Zhang, Mengmeng Wang, Steven Hoi, Peng Gao, Harry Yang
European Conference on Computer Vision (ECCV), 2026
code / arXiv

Showing that DMD and RL can be trained simultaneously, with RL enabling the student model to surpass the teacher and DMD loss regularizing RL to prevent reward hacking.

Grounding DINO video adaptation teaser
Unlocking the Potential of Grounding DINO in Videos: Parameter-Efficient Adaptation for Limited-Data Spatial-Temporal Localization
Zanyi Wang, Fan Li, Dengyang Jiang, Liuzhuozheng Li, Yunhua Zhong, Guang Dai, Mengmeng Wang
Conference on Computer Vision and Pattern Recognition (CVPR) @ CV4Smalls WS, 2026
arXiv

A parameter-efficient adaptation approach that successfully tackles complex spatio-temporal localization under strict small-data constraints.

RefTON teaser
RefTON: Person-to-Person Virtual Try-On with Unpaired Visual References
Liuzhuozheng Li, Yue Gong, Shanyuan Liu, Zanyi Wang, Dengyang Jiang, Liebucha Wu, Bo Cheng, Yuhang Ma, Dawei Leng, Yuhui Yin
Conference on Computer Vision and Pattern Recognition (CVPR), 2026
code / arXiv

An End-to-End Virtual Try-on model that directly fits the target garment onto the person image.

FlowRVS teaser
Deforming Videos to Masks: Flow Matching for Referring Video Segmentation
Zanyi Wang, Dengyang Jiang, Liuzhuozheng Li, Sizhe Dang, Chengzu Li, Harry Yang, Guang Dai, Mengmeng Wang, Jingdong Wang
International Conference on Learning Representations (ICLR), 2026
code / arXiv / 机器之心

Reformulated RVOS as learning a continuous, text-conditioned flow that deforms a video's content into its target mask.

TriCLIP-3D teaser
TriCLIP-3D: A Unified Parameter-Efficient Framework for Tri-Modal 3D Visual Grounding based on CLIP
Fan Li, Zanyi Wang*, Zeyi Huang, Guang Dai, Jingdong Wang, Mengmeng Wang
ACM International Conference on Multimedia (ACM MM), 2025
arXiv

Developed a unified framework leveraging CLIP’s ViT encoder for efficient tri-modal 3D visual grounding.

d-opsd
D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models
Dengyang Jiang, Xin Jin, Dongyang Liu, Zanyi Wang, Mingzhe Zheng, Ruoyi Du, Xiangpeng Yang, Qilong Wu, Zhen Li, Peng Gao, Harry Yang, Steven Hoi
Preprint, 2026
code / arXiv

On-policy self-distillation for continuously tuning step-distilled diffusion models without sacrificing few-step capacity.

Time conditioning in diffusion models teaser
Exploring Time Conditioning in Diffusion Generative Models from Disjoint Noisy Data Manifolds
Liuzhuozheng Li, Zhiyuan Zhan, Shuhong Liu, Dengyang Jiang, Zanyi Wang, Guang Dai, Jingdong Wang, Mengmeng Wang
Preprint, 2026
code / arXiv

Diffusion models can generate high-quality samples without timestep or class embeddings when noisy data manifolds are kept geometrically disjoint across timesteps.

Thinking in Frames teaser
Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning
Chengzu Li, Zanyi Wang*, Jiaang Li, Yi Xu, Han Zhou, Huanyu Zhang, Ruichuan An, Dengyang Jiang, Zhaochong An, Ivan Vulić, Serge Belongie, Anna Korhonen
Preprint, 2026
arXiv

Exploring how visual context enhances video reasoning capabilities.


Design and source code from Jon Barron's website.