Selected Projects

Open-source research systems from our group and collaborators — spanning embodied AI & robotics, autonomous driving, generative models, and efficient multimodal foundation models.

# equal contribution* corresponding author

  1. AdaRoboVLG: real-world robotic grasping across different hands
    arXiv 2026Embodied AI

    AdaRoboVLG: Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp Synthesis

    Sixu Yan, Shikang Wang, Binhua Huang, Xuanlai Tang, Guohua Fan, Fan Huang, Haoxuan Li, Yongkang Li, Yuhan Li, Bencheng Liao, Zeyu Zhang, Wenyu Liu, Hangxin Liu, Xinggang Wang

    A task-adaptive Vision-Language-Grasp framework that decouples physical grasp synthesis from task-dependent understanding, supporting generalizable grasping across different robotic hands via composable foundation-model priors.

  2. arXiv 2026World Action Model

    DreamWAM: Beyond RGB Future Prediction for World Action Models

    Shanglin Yuan#, Weiheng Zhao#, Xin Shi, Haoyi Jiang, Xianda Guo, Liu Liu, Wenyu Liu, Wei Sui, Xinggang Wang*

    A world action model that reformulates future prediction as structured world modeling beyond RGB, representing future states through complementary views of appearance, motion, geometry, and semantics for robust manipulation.

  3. Faster-WAM framework: inference-time future conditioning for world action models
    arXiv 2026World Action Model

    Faster-WAM: Efficient Inference-Time Future Conditioning for Robust World Action Models

    Weiheng Zhao, Haoyi Jiang, Xin Shi, Liu Liu, Zhizhong Su, Wei Sui, Fan Huang, Xinggang Wang

    A world action model that computes future representations once and selectively reuses them during action denoising, achieving robust out-of-distribution generalization with substantially reduced inference latency.

  4. ReWorld: representation learning framework for world action models
    arXiv 2026Autonomous Driving

    ReWorld: Learning Better Representations for World Action Models

    Tianze Xia#, Lijun Zhou#, Kaixin Xiong, Jingfeng Yao, Yu Zhu, Zhenxin Zhu, Haiyang Sun, Bing Wang, Guang Chen, Hangjun Ye, Wenyu Liu, Xinggang Wang*

    The first representation learning framework for autonomous-driving world action models, explicitly optimizing the latent world-to-action pathway without external encoders or teacher models at only +0.3% training cost.

  5. Moebius: lightweight image inpainting pipeline
    ECCV 2026Generative Model

    Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance

    Kangsheng Duan, Ziyang Xu, Wenyu Liu, Xiaohu Ruan, Xiaoxin Chen, Xinggang Wang

    A 0.2B lightweight image inpainting framework that delivers 10B-level performance.

  6. UniDriveVLA: Mixture-of-Transformers architecture for autonomous driving
    arXiv 2026Autonomous Driving

    UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving

    Yongkang Li, Lijun Zhou, Sixu Yan, Bencheng Liao, Tianyi Yan, Kaixin Xiong, Long Chen, Hongwei Xie, Bing Wang, Guang Chen, Hangjun Ye, Haiyang Sun, Wenyu Liu, Xinggang Wang*

    A unified driving vision-language-action model based on Mixture-of-Transformers, decoupling driving understanding, scene perception, and action planning into specialized experts coordinated through masked joint attention.

  7. RAD-2: generator-discriminator framework for closed-loop planning
    arXiv 2026Autonomous Driving

    RAD-2: Scaling Reinforcement Learning in a Generator-Discriminator Framework

    Hao Gao, Shaoyu Chen, Yifan Zhu, Yuehao Song, Wenyu Liu, Qian Zhang, Xinggang Wang*

    A generator-discriminator framework for closed-loop driving: a diffusion generator proposes diverse trajectory candidates while an RL-optimized discriminator reranks them by long-term driving quality, reducing the collision rate by 56% over strong diffusion planners.

  8. Senna-2: aligning VLM and end-to-end driving policy
    arXiv 2026Autonomous Driving

    Senna-2: Aligning VLM and End-to-End Driving Policy for Consistent Decision Making and Planning

    Yuehao Song, Shaoyu Chen, Hao Gao, Yifan Zhu, Weixiang Yue, Jialv Zou, Bo Jiang, Zihao Lu, Yu Wang, Qian Zhang, Xinggang Wang*

    A VLM-E2E driving policy that explicitly aligns the VLM's high-level decisions with the E2E policy's low-level planning, via consistency-oriented three-stage training and hierarchical reinforcement learning in 3DGS environments.

  9. DriveLaW: unifying planning and video generation in a latent driving world
    CVPR 2026Autonomous Driving

    DriveLaW: Unifying Planning and Video Generation in a Latent Driving World

    Tianze Xia, Yongkang Li#, Lijun Zhou#, Jingfeng Yao, Kaixin Xiong, Haiyang Sun, Bing Wang, Kun Ma, Guang Chen, Hangjun Ye, Wenyu Liu, Xinggang Wang*

    A driving world model that unifies video generation and motion planning by injecting the video generator's latent representation directly into a diffusion planner, achieving state-of-the-art results on both nuScenes video prediction and NAVSIM planning.

  10. DiffusionVL: translating autoregressive models into diffusion vision-language models
    ECCV 2026Multimodal LLM

    DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models

    Lunbin Zeng, Jingfeng Yao, Bencheng Liao, Hongyuan Tao, Wenyu Liu, Xinggang Wang

    A framework that translates any autoregressive model into a diffusion vision-language model, combining the strengths of autoregressive pre-training with parallel diffusion decoding.

  11. VTP: scalable pre-training of visual tokenizers for generation
    ECCV 2026Generative Model

    VTP: Towards Scalable Pre-training of Visual Tokenizers for Generation

    Jingfeng Yao, Yuda Song, Yucong Zhou, Xinggang Wang

    A unified visual tokenizer pre-training framework that jointly optimizes contrastive, self-supervised, and reconstruction objectives, unlocking a new scaling law where generation performance scales with tokenizer pre-training compute, parameters, and data.

  12. InfiniteVL: linear and sparse attention for unlimited-input vision-language models
    arXiv 2025Efficient VLM

    InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models

    Hongyuan Tao, Bencheng Liao, Shaoyu Chen, Haoran Yin, Qian Zhang, Wenyu Liu, Xinggang Wang*

    An efficient vision-language model that combines linear attention for compact long-term memory with sparse attention for precise visual perception, supporting unlimited-input streaming with a constant memory footprint.

  13. 4DLangVGGT: 4D language-visual geometry grounded transformer
    arXiv 20254D Scene Understanding

    4DLangVGGT: 4D Language-Visual Geometry Grounded Transformer

    Xianfeng Wu, Yajing Bai#, Minghan Li, Xianzu Wu, Xueqi Zhao, Zhongyuan Lai, Wenyu Liu, Xinggang Wang*

    The first Transformer-based feed-forward framework for 4D language grounding, jointly integrating geometric perception (StreamVGGT) and language alignment (Semantic Bridging Decoder) — no per-scene optimization, generalizing across dynamic scenes.

  14. TransLight: image-guided customized lighting control with generative decoupling
    ICML 2026Generative Model

    TransLight: Image-Guided Customized Lighting Control with Generative Decoupling

    Zongming Li, Lianghui Zhu, Haocheng Shen, Longjin Ran, Wenyu Liu, Xinggang Wang

    Image-guided customized lighting control: generative decoupling transfers the lighting effects of any reference image onto user content, enabling text-free, precise relighting.

  15. LENS: learning to segment anything with unified reinforced reasoning
    AAAI 2026Segmentation

    LENS: Learning to Segment Anything with Unified Reinforced Reasoning

    Lianghui Zhu, Bin Ouyang, Yuxuan Zhang, Tianheng Cheng, Rui Hu, Haocheng Shen, Longjin Ran, Xiaoxin Chen, Li Yu, Wenyu Liu, Xinggang Wang

    Segment anything with unified reinforced reasoning: a multimodal model that reasons over complex referring expressions and produces precise segmentation masks, trained end-to-end with reinforcement learning.

  16. ReCogDrive: reinforced cognitive framework for end-to-end autonomous driving
    ICLR 2026Autonomous Driving

    ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving

    Yongkang Li#, Kaixin Xiong#, Xiangyu Guo, Fang Li, Sixu Yan, Gangwei Xu, Lijun Zhou, Long Chen, Haiyang Sun, Bing Wang, Kun Ma, Guang Chen, Hangjun Ye, Wenyu Liu, Xinggang Wang*

    A reinforced cognitive driving framework that unifies driving understanding and planning by combining a cognition-enhanced VLM with a diffusion planner, further reinforced by DiffGRPO for safer and more comfortable trajectories.

  17. GaussTR: foundation model-aligned Gaussian transformer for self-supervised 3D spatial understanding
    CVPR 20253D Perception

    GaussTR: Foundation Model-Aligned Gaussian Transformer for Self-Supervised 3D Spatial Understanding

    Haoyi Jiang, Liu Liu, Tianheng Cheng, Xinjie Wang, Tianwei Lin, Zhizhong Su, Wenyu Liu, Xinggang Wang

    Self-supervised 3D spatial understanding: aligning Gaussian-based 3D representations with foundation-model features, enabling open-vocabulary 3D perception without manual 3D annotations.

  18. DiffusionDrive: truncated diffusion model for end-to-end autonomous driving
    CVPR 2025Autonomous Driving

    DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving

    Bencheng Liao, Shaoyu Chen, Haoran Yin, Bo Jiang, Cheng Wang, Sixu Yan, Xinbang Zhang, Xiangyu Li, Ying Zhang, Qian Zhang, Xinggang Wang

    A truncated diffusion policy for real-time end-to-end driving, denoising from an anchored Gaussian prior to generate diverse, high-quality trajectories in just a few steps.

  19. Senna: bridging large vision-language models and end-to-end autonomous driving
    IJCV 2026Autonomous Driving

    Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving

    Bo Jiang, Shaoyu Chen, Bencheng Liao, Xingyu Zhang, Wei Yin, Qian Zhang, Chang Huang, Wenyu Liu, Xinggang Wang

    A driving VLM that bridges large vision-language models and end-to-end autonomous driving, combining high-level scene understanding and reasoning with precise trajectory planning.

  20. VADv2: end-to-end autonomous driving via probabilistic planning
    ICLR 2026Autonomous Driving

    VADv2: End-to-End Autonomous Driving via Probabilistic Planning

    Bo Jiang, Shaoyu Chen, Hao Gao, Bencheng Liao, Qian Zhang, Wenyu Liu, Xinggang Wang

    End-to-end vectorized autonomous driving that outputs a probabilistic distribution over actions learned from large-scale driving demonstrations and samples one action to control the vehicle.

  21. YOLO-World: real-time open-vocabulary object detection
    CVPR 2024Open-Vocabulary Detection

    YOLO-World: Real-Time Open-Vocabulary Object Detection

    Tianheng Cheng, Lin Song, Yixiao Ge, Wenyu Liu, Xinggang Wang, Ying Shan

    Real-time open-vocabulary detection: YOLO enhanced with vision-language modeling via the re-parameterizable RepVL-PAN and region-text contrastive pre-training, reaching 35.4 AP at 52 FPS on LVIS zero-shot.

  22. Vision Mamba: efficient visual representation learning with bidirectional state space model
    ICML 2024Visual Representation

    Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model

    Lianghui Zhu#, Bencheng Liao#, Qian Zhang, Xinlong Wang, Wenyu Liu, Xinggang Wang*

    A generic vision backbone built on bidirectional state space models, delivering Transformer-level representation power with linear complexity and no attention.

  23. 4DGaussians: real-time dynamic scene rendering
    CVPR 20244D Reconstruction

    4D Gaussian Splatting for Real-Time Dynamic Scene Rendering

    Guanjun Wu#, Taoran Yi#, Jiemin Fang, Lingxi Xie, Xiaopeng Zhang, Wei Wei, Wenyu Liu, Qi Tian, Xinggang Wang*

    Real-time rendering of dynamic scenes by modeling 4D scenes with Gaussian splatting and a compact spatio-temporal representation, achieving high-quality novel-view synthesis at interactive frame rates.

  24. MapTR: online vectorized HD map construction
    IJCV 2025Autonomous Driving

    MapTRv2: An End-to-End Framework for Online Vectorized HD Map Construction

    Bencheng Liao, Shaoyu Chen, Yunchi Zhang, Bo Jiang, Qian Zhang, Wenyu Liu, Chang Huang, Xinggang Wang

    An end-to-end framework for online vectorized HD map construction, modeling map elements as point sets with a hierarchical query design for real-time structured map learning.

  25. EVA: exploring the limits of masked visual representation learning at scale
    CVPR 2023Visual Representation

    EVA: Exploring the Limits of Masked Visual Representation Learning at Scale

    Yuxin Fang#, Wen Wang, Binhui Xie, Quan Sun, Ledell Wu, Xinggang Wang*, Tiejun Huang, Xinlong Wang*, Yue Cao*

    Scaling masked image modeling to billion-parameter vision transformers: a vanilla ViT pre-trained with masked image modeling achieves state-of-the-art results across a wide range of downstream vision tasks.

  26. ByteTrack: multi-object tracking by associating every detection box
    ECCV 2022Multi-Object Tracking

    ByteTrack: Multi-Object Tracking by Associating Every Detection Box

    Yifu Zhang, Peize Sun, Yi Jiang, Dongdong Yu, Fucheng Weng, Zehuan Yuan, Ping Luo, Wenyu Liu, Xinggang Wang

    Simple yet effective multi-object tracking: by associating every detection box instead of only high-scoring ones, ByteTrack recovers true objects from low-score detections and filters out background, setting the standard tracker integrated in YOLOv8 and beyond.

  27. FairMOT: on the fairness of detection and re-identification in multiple object tracking
    IJCV 2021Multi-Object Tracking

    FairMOT: On the Fairness of Detection and Re-Identification in Multiple Object Tracking

    Yifu Zhang#, Chunyu Wang, Xinggang Wang*, Wenjun Zeng, Wenyu Liu

    A one-shot multi-object tracker that fairly balances detection and re-identification in a shared homogeneous network, addressing the inherent unfairness of treating re-ID as a secondary task.

  28. CCNet: criss-cross attention for semantic segmentation
    TPAMI 2023Semantic Segmentation

    CCNet: Criss-Cross Attention for Semantic Segmentation

    Zilong Huang#, Xinggang Wang#*, Yunchao Wei, Lichao Huang, Humphrey Shi, Wenyu Liu, Thomas Huang

    Criss-cross attention captures full-image context with sparse criss-cross paths instead of dense attention, an efficient long-range dependency module widely used in segmentation — notably used in AlphaFold.