Selected Projects

Open-source research systems from our group and collaborators — spanning embodied AI & robotics, autonomous driving, generative models, and efficient multimodal foundation models.

# equal contribution* corresponding author

  1. AdaRoboVLG: real-world robotic grasping across different hands
    arXiv 2026Embodied AI

    AdaRoboVLG: Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp Synthesis

    Sixu Yan, Shikang Wang, Binhua Huang, Xuanlai Tang, Guohua Fan, Fan Huang, Haoxuan Li, Yongkang Li, Yuhan Li, Bencheng Liao, Zeyu Zhang, Wenyu Liu, Hangxin Liu, Xinggang Wang

    A task-adaptive Vision-Language-Grasp framework that decouples physical grasp synthesis from task-dependent understanding, supporting generalizable grasping across different robotic hands via composable foundation-model priors.

  2. arXiv 2026World ModelEmbodied AI

    DreamWAM: Beyond RGB Future Prediction for World Action Models

    Shanglin Yuan#, Weiheng Zhao#, Xin Shi, Haoyi Jiang, Xianda Guo, Liu Liu, Wenyu Liu, Wei Sui, Xinggang Wang*

    A world action model that reformulates future prediction as structured world modeling beyond RGB, representing future states through complementary views of appearance, motion, geometry, and semantics for robust manipulation.

  3. Faster-WAM framework: inference-time future conditioning for world action models
    arXiv 2026World ModelEmbodied AI

    Faster-WAM: Efficient Inference-Time Future Conditioning for Robust World Action Models

    Weiheng Zhao, Haoyi Jiang, Xin Shi, Liu Liu, Zhizhong Su, Wei Sui, Fan Huang, Xinggang Wang

    A world action model that computes future representations once and selectively reuses them during action denoising, achieving robust out-of-distribution generalization with substantially reduced inference latency.

  4. ReWorld: representation learning framework for world action models
    arXiv 2026World ModelAutonomous Driving

    ReWorld: Learning Better Representations for World Action Models

    Tianze Xia#, Lijun Zhou#, Kaixin Xiong, Jingfeng Yao, Yu Zhu, Zhenxin Zhu, Haiyang Sun, Bing Wang, Guang Chen, Hangjun Ye, Wenyu Liu, Xinggang Wang*

    The first representation learning framework for autonomous-driving world action models, explicitly optimizing the latent world-to-action pathway without external encoders or teacher models at only +0.3% training cost.

  5. Moebius: lightweight image inpainting pipeline
    ECCV 2026Generative Model

    Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance

    Kangsheng Duan, Ziyang Xu, Wenyu Liu, Xiaohu Ruan, Xiaoxin Chen, Xinggang Wang

    A 0.2B lightweight image inpainting framework that delivers 10B-level performance.

  6. CoRL 2026Embodied AI

    MotionVLA: Injecting Geometric Motion into Vision-Language-Action Model

    Shanglin Yuan, Weiheng Zhao, Xianda Guo, Wei Sui, Li Yu, Wenyu Liu, Xinggang Wang

    Injects geometric motion into a VLA as compact, queryable trajectory-field tokens from strictly past frames, giving the policy a time-continuous motion history for long-horizon manipulation.

  7. RAD-2: generator-discriminator framework for closed-loop planning
    arXiv 2026Autonomous Driving

    RAD-2: Scaling Reinforcement Learning in a Generator-Discriminator Framework

    Hao Gao, Shaoyu Chen, Yifan Zhu, Yuehao Song, Wenyu Liu, Qian Zhang, Xinggang Wang*

    A generator-discriminator framework for closed-loop driving: a diffusion generator proposes diverse trajectory candidates while an RL-optimized discriminator reranks them by long-term driving quality, reducing the collision rate by 56% over strong diffusion planners.

  8. UniDriveVLA: Mixture-of-Transformers architecture for autonomous driving
    arXiv 2026Autonomous Driving

    UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving

    Yongkang Li, Lijun Zhou, Sixu Yan, Bencheng Liao, Tianyi Yan, Kaixin Xiong, Long Chen, Hongwei Xie, Bing Wang, Guang Chen, Hangjun Ye, Haiyang Sun, Wenyu Liu, Xinggang Wang*

    A unified driving vision-language-action model based on Mixture-of-Transformers, decoupling driving understanding, scene perception, and action planning into specialized experts coordinated through masked joint attention.

  9. MoDA: mixture-of-depths attention architecture and visible relationships
    arXiv 2026Efficient LLM

    MoDA: Mixture-of-Depths Attention

    Lianghui Zhu, Yuxin Fang, Bencheng Liao, Shijie Wang, Tianheng Cheng, Zilong Huang, Chen Chen, Lai Wei, Yutao Zeng, Ya Wang, Yi Lin, Yu Li, Xinggang Wang*

    A depth-scaling attention primitive that lets each head attend to both sequence KV and depth KV from preceding layers, with a hardware-efficient Triton kernel reaching 97.3% of FlashAttention-2 efficiency at 64K sequence length.

  10. Senna-2: aligning VLM and end-to-end driving policy
    arXiv 2026Autonomous DrivingMultimodal LLM

    Senna-2: Aligning VLM and End-to-End Driving Policy for Consistent Decision Making and Planning

    Yuehao Song, Shaoyu Chen, Hao Gao, Yifan Zhu, Weixiang Yue, Jialv Zou, Bo Jiang, Zihao Lu, Yu Wang, Qian Zhang, Xinggang Wang*

    A VLM-E2E driving policy that explicitly aligns the VLM's high-level decisions with the E2E policy's low-level planning, via consistency-oriented three-stage training and hierarchical reinforcement learning in 3DGS environments.

  11. DriveLaW: unifying planning and video generation in a latent driving world
    CVPR 2026Autonomous DrivingWorld Model

    DriveLaW: Unifying Planning and Video Generation in a Latent Driving World

    Tianze Xia, Yongkang Li#, Lijun Zhou#, Jingfeng Yao, Kaixin Xiong, Haiyang Sun, Bing Wang, Kun Ma, Guang Chen, Hangjun Ye, Wenyu Liu, Xinggang Wang*

    A driving world model that unifies video generation and motion planning by injecting the video generator's latent representation directly into a diffusion planner, achieving state-of-the-art results on both nuScenes video prediction and NAVSIM planning.

  12. DiffusionVL: translating autoregressive models into diffusion vision-language models
    ECCV 2026Multimodal LLM

    DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models

    Lunbin Zeng, Jingfeng Yao, Bencheng Liao, Hongyuan Tao, Wenyu Liu, Xinggang Wang

    A framework that translates any autoregressive model into a diffusion vision-language model, combining the strengths of autoregressive pre-training with parallel diffusion decoding.

  13. VTP: scalable pre-training of visual tokenizers for generation
    ECCV 2026Generative Model

    VTP: Towards Scalable Pre-training of Visual Tokenizers for Generation

    Jingfeng Yao, Yuda Song, Yucong Zhou, Xinggang Wang

    A unified visual tokenizer pre-training framework that jointly optimizes contrastive, self-supervised, and reconstruction objectives, unlocking a new scaling law where generation performance scales with tokenizer pre-training compute, parameters, and data.

  14. InfiniteVL: linear and sparse attention for unlimited-input vision-language models
    arXiv 2025Multimodal LLM

    InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models

    Hongyuan Tao, Bencheng Liao, Shaoyu Chen, Haoran Yin, Qian Zhang, Wenyu Liu, Xinggang Wang*

    An efficient vision-language model that combines linear attention for compact long-term memory with sparse attention for precise visual perception, supporting unlimited-input streaming with a constant memory footprint.

  15. 4DLangVGGT: 4D language-visual geometry grounded transformer
    arXiv 20253D / 4D Vision

    4DLangVGGT: 4D Language-Visual Geometry Grounded Transformer

    Xianfeng Wu, Yajing Bai#, Minghan Li, Xianzu Wu, Xueqi Zhao, Zhongyuan Lai, Wenyu Liu, Xinggang Wang*

    The first Transformer-based feed-forward framework for 4D language grounding, jointly integrating geometric perception (StreamVGGT) and language alignment (Semantic Bridging Decoder) — no per-scene optimization, generalizing across dynamic scenes.

  16. VGT: visual generation tuning pipeline that turns a pretrained VLM into an image generator
    arXiv 2025Generative ModelMultimodal LLM

    VGT: Visual Generation Tuning

    Jiahao Guo#, Sinan Du#, Jingfeng Yao, Wenyu Liu, Bo Li, Haoxiang Cao, Kun Gai, Chun Yuan, Kai Wu, Xinggang Wang*

    Unleashes visual generation from any pretrained VLM by aligning its semantic encoder with a pixel decoder (VGT-AE) and tuning a lightweight flow-matching head (VGT-AR) — GenEval 0.82, DPG-Bench 81.28, 20× faster convergence.

  17. MobileI2V: fast high-resolution image-to-video generation on mobile devices
    arXiv 2025Generative Model

    MobileI2V: Fast and High-Resolution Image-to-Video on Mobile Devices

    Shuai Zhang#, Bao Tang#, Siyuan Yu#, Yueting Zhu, Jingfeng Yao, Ya Zou, Shanglin Yuan, Li Yu, Wenyu Liu, Xinggang Wang*

    A 0.27B image-to-video model for mobile devices — 5.55× smaller than SVD-XT with comparable quality, generating 720p video in 2.24s on mobile and running 199× faster on an A100.

  18. TransLight: image-guided customized lighting control with generative decoupling
    ICML 2026Generative Model

    TransLight: Image-Guided Customized Lighting Control with Generative Decoupling

    Zongming Li, Lianghui Zhu, Haocheng Shen, Longjin Ran, Wenyu Liu, Xinggang Wang

    Image-guided customized lighting control: generative decoupling transfers the lighting effects of any reference image onto user content, enabling text-free, precise relighting.

  19. LENS: learning to segment anything with unified reinforced reasoning
    AAAI 2026SegmentationMultimodal LLM

    LENS: Learning to Segment Anything with Unified Reinforced Reasoning

    Lianghui Zhu, Bin Ouyang, Yuxuan Zhang, Tianheng Cheng, Rui Hu, Haocheng Shen, Longjin Ran, Xiaoxin Chen, Li Yu, Wenyu Liu, Xinggang Wang

    Segment anything with unified reinforced reasoning: a multimodal model that reasons over complex referring expressions and produces precise segmentation masks, trained end-to-end with reinforcement learning.

  20. ReCogDrive: reinforced cognitive framework for end-to-end autonomous driving
    ICLR 2026Autonomous DrivingMultimodal LLM

    ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving

    Yongkang Li#, Kaixin Xiong#, Xiangyu Guo, Fang Li, Sixu Yan, Gangwei Xu, Lijun Zhou, Long Chen, Haiyang Sun, Bing Wang, Kun Ma, Guang Chen, Hangjun Ye, Wenyu Liu, Xinggang Wang*

    A reinforced cognitive driving framework that unifies driving understanding and planning by combining a cognition-enhanced VLM with a diffusion planner, further reinforced by DiffGRPO for safer and more comfortable trajectories.

  21. GroundingSuite: multi-granular pixel grounding tasks versus RefCOCO, gRefCOCO and GranD
    ICCV 2025SegmentationMultimodal LLM

    GroundingSuite: Measuring Complex Multi-Granular Pixel Grounding

    Rui Hu#, Lianghui Zhu#, Yuxuan Zhang, Tianheng Cheng, Lei Liu, Heng Liu, Longjin Ran, Xiaoxin Chen, Wenyu Liu, Xinggang Wang*

    A comprehensive pixel grounding framework: an automated VLM-based annotation pipeline (4.5× faster than GLaMM), a 9.56M-sample training set with diverse referring expressions, and a 3,800-instance evaluation benchmark.

  22. OmniMamba: unified multimodal understanding and generation architecture based on Mamba-2
    ECCV 2026Multimodal LLMGenerative Model

    OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models

    Jialv Zou, Bencheng Liao, Qian Zhang, Wenyu Liu, Xinggang Wang*

    The first linear-model-based unified multimodal understanding and visual generation model — competitive with only 2M training data, up to 119.2× speedup and 63% GPU memory reduction on long-sequence generation vs Transformer counterparts.

  23. LightningDiT: ImageNet-256 samples generated with the VA-VAE aligned latent diffusion system
    CVPR 2025Generative Model

    LightningDiT & VA-VAE: Taming the Reconstruction-Generation Dilemma in Latent Diffusion Models

    Jingfeng Yao, Bin Yang, Xinggang Wang

    Aligns the VAE latent space with vision foundation models (VA-VAE) and pairs it with an enhanced DiT baseline (LightningDiT) — FID 1.35 with 0.28 rFID on ImageNet-256, converging 21.8× faster than the original DiT (FID 2.11 in 64 epochs). CVPR 2025 Oral.

  24. GaussTR: foundation model-aligned Gaussian transformer for self-supervised 3D spatial understanding
    CVPR 20253D / 4D Vision

    GaussTR: Foundation Model-Aligned Gaussian Transformer for Self-Supervised 3D Spatial Understanding

    Haoyi Jiang, Liu Liu, Tianheng Cheng, Xinjie Wang, Tianwei Lin, Zhizhong Su, Wenyu Liu, Xinggang Wang

    Self-supervised 3D spatial understanding: aligning Gaussian-based 3D representations with foundation-model features, enabling open-vocabulary 3D perception without manual 3D annotations.

  25. DiffusionDrive: truncated diffusion model for end-to-end autonomous driving
    CVPR 2025Autonomous Driving

    DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving

    Bencheng Liao, Shaoyu Chen, Haoran Yin, Bo Jiang, Cheng Wang, Sixu Yan, Xinbang Zhang, Xiangyu Li, Ying Zhang, Qian Zhang, Xinggang Wang

    A truncated diffusion policy for real-time end-to-end driving, denoising from an anchored Gaussian prior to generate diverse, high-quality trajectories in just a few steps.

  26. Senna: bridging large vision-language models and end-to-end autonomous driving
    IJCV 2026Autonomous DrivingMultimodal LLM

    Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving

    Bo Jiang, Shaoyu Chen, Bencheng Liao, Xingyu Zhang, Wei Yin, Qian Zhang, Chang Huang, Wenyu Liu, Xinggang Wang

    A driving VLM that bridges large vision-language models and end-to-end autonomous driving, combining high-level scene understanding and reasoning with precise trajectory planning.

  27. ControlAR: controllable image generation with autoregressive models
    ICLR 2025Generative Model

    ControlAR: Controllable Image Generation with Autoregressive Models

    Zongming Li#, Tianheng Cheng#, Shoufa Chen, Peize Sun, Haocheng Shen, Longjin Ran, Xiaoxin Chen, Wenyu Liu, Xinggang Wang*

    Adds spatial controls (edges, depth, segmentation maps) to autoregressive image generators via conditional decoding on the control sequence — no hand-crafted special tokens, supporting arbitrary-resolution generation.

  28. DiG: high-resolution samples from diffusion models with gated linear attention
    CVPR 2025Generative Model

    DiG: Scalable and Efficient Diffusion Models with Gated Linear Attention

    Lianghui Zhu, Zilong Huang*, Bencheng Liao, Jun Hao Liew, Hanshu Yan, Jiashi Feng, Xinggang Wang*

    Brings gated linear attention into the diffusion backbone with minimal parameter overhead — beating DiT at 256² while scaling far better at high resolutions: DiG-S/2 is 2.5× faster with 75.7% less GPU memory than DiT-S/2 at 1792², and DiG-XL/2 is 1.8× faster than DiT with FlashAttention-2 at 2048².

  29. VADv2: end-to-end autonomous driving via probabilistic planning
    ICLR 2026Autonomous Driving

    VADv2: End-to-End Autonomous Driving via Probabilistic Planning

    Bo Jiang, Shaoyu Chen, Hao Gao, Bencheng Liao, Qian Zhang, Wenyu Liu, Xinggang Wang

    End-to-end vectorized autonomous driving that outputs a probabilistic distribution over actions learned from large-scale driving demonstrations and samples one action to control the vehicle.

  30. EVA-X: self-supervised pre-training pipeline for chest X-ray foundation models
    npj Digital Medicine 2025Medical Imaging

    EVA-X: A Foundation Model for General Chest X-ray Analysis with Self-Supervised Learning

    Jingfeng Yao, Xinggang Wang*, Yuehao Song, Huangxuan Zhao, Jun Ma, Yajie Chen, Wenyu Liu, Bo Wang*

    The first X-ray self-supervised ViT foundation model capturing both semantic and geometric information, spanning 20+ chest diseases with leading results on 11+ detection tasks and strong few-shot capability.

  31. MIM4D: dual masked image modeling on multi-view video with 3D volumetric rendering supervision
    IJCV 2025Autonomous Driving3D / 4D Vision

    MIM4D: Masked Modeling with Multi-View Video for Autonomous Driving Representation Learning

    Jialv Zou#, Bencheng Liao#, Qian Zhang, Wenyu Liu, Xinggang Wang*

    A dual masked-image-modeling pre-training paradigm for multi-view driving videos, supervising pseudo-3D features via volumetric differentiable rendering — improving end-to-end planning (-9% collision), BEV segmentation (+8.7% IoU), 3D detection (+3.5% mAP) and HD map construction (+1.4% mAP) on nuScenes.

  32. YOLO-World: real-time open-vocabulary object detection
    CVPR 2024Detection & Tracking

    YOLO-World: Real-Time Open-Vocabulary Object Detection

    Tianheng Cheng, Lin Song, Yixiao Ge, Wenyu Liu, Xinggang Wang, Ying Shan

    Real-time open-vocabulary detection: YOLO enhanced with vision-language modeling via the re-parameterizable RepVL-PAN and region-text contrastive pre-training, reaching 35.4 AP at 52 FPS on LVIS zero-shot.

  33. Vision Mamba: efficient visual representation learning with bidirectional state space model
    ICML 2024Visual Representation

    Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model

    Lianghui Zhu#, Bencheng Liao#, Qian Zhang, Xinlong Wang, Wenyu Liu, Xinggang Wang*

    A generic vision backbone built on bidirectional state space models, delivering Transformer-level representation power with linear complexity and no attention.

  34. 4DGaussians: real-time dynamic scene rendering
    CVPR 20243D / 4D Vision

    4D Gaussian Splatting for Real-Time Dynamic Scene Rendering

    Guanjun Wu#, Taoran Yi#, Jiemin Fang, Lingxi Xie, Xiaopeng Zhang, Wei Wei, Wenyu Liu, Qi Tian, Xinggang Wang*

    Real-time rendering of dynamic scenes by modeling 4D scenes with Gaussian splatting and a compact spatio-temporal representation, achieving high-quality novel-view synthesis at interactive frame rates.

  35. WeakTr: plain ViT framework for weakly-supervised semantic segmentation with adaptive attention fusion
    IEEE TIP 2026Segmentation

    WeakTr: Exploring Plain Vision Transformer for Weakly-Supervised Semantic Segmentation

    Lianghui Zhu#, Yingyue Li#, Jiemin Fang, Yan Liu, Xin Hao, Wenyu Liu, Xinggang Wang*

    Explores plain ViT for weakly-supervised semantic segmentation: end-to-end CAM generation by adaptively fusing self-attention maps with learned head weights, plus online retraining with a gradient-clipping decoder — 78.5% mIoU on VOC12, 51.1% on COCO14.

  36. MapTR: online vectorized HD map construction
    IJCV 2025Autonomous Driving3D / 4D Vision

    MapTRv2: An End-to-End Framework for Online Vectorized HD Map Construction

    Bencheng Liao, Shaoyu Chen, Yunchi Zhang, Bo Jiang, Qian Zhang, Wenyu Liu, Chang Huang, Xinggang Wang

    An end-to-end framework for online vectorized HD map construction, modeling map elements as point sets with a hierarchical query design for real-time structured map learning.

  37. EVA: exploring the limits of masked visual representation learning at scale
    CVPR 2023Visual Representation

    EVA: Exploring the Limits of Masked Visual Representation Learning at Scale

    Yuxin Fang#, Wen Wang, Binhui Xie, Quan Sun, Ledell Wu, Xinggang Wang*, Tiejun Huang, Xinlong Wang*, Yue Cao*

    Scaling masked image modeling to billion-parameter vision transformers: a vanilla ViT pre-trained with masked image modeling achieves state-of-the-art results across a wide range of downstream vision tasks.

  38. ByteTrack: multi-object tracking by associating every detection box
    ECCV 2022Detection & Tracking

    ByteTrack: Multi-Object Tracking by Associating Every Detection Box

    Yifu Zhang, Peize Sun, Yi Jiang, Dongdong Yu, Fucheng Weng, Zehuan Yuan, Ping Luo, Wenyu Liu, Xinggang Wang

    Simple yet effective multi-object tracking: by associating every detection box instead of only high-scoring ones, ByteTrack recovers true objects from low-score detections and filters out background, setting the standard tracker integrated in YOLOv8 and beyond.

  39. FairMOT: on the fairness of detection and re-identification in multiple object tracking
    IJCV 2021Detection & Tracking

    FairMOT: On the Fairness of Detection and Re-Identification in Multiple Object Tracking

    Yifu Zhang#, Chunyu Wang, Xinggang Wang*, Wenjun Zeng, Wenyu Liu

    A one-shot multi-object tracker that fairly balances detection and re-identification in a shared homogeneous network, addressing the inherent unfairness of treating re-ID as a secondary task.

  40. CCNet: criss-cross attention for semantic segmentation
    TPAMI 2023Segmentation

    CCNet: Criss-Cross Attention for Semantic Segmentation

    Zilong Huang#, Xinggang Wang#*, Yunchao Wei, Lichao Huang, Humphrey Shi, Wenyu Liu, Thomas Huang

    Criss-cross attention captures full-image context with sparse criss-cross paths instead of dense attention, an efficient long-range dependency module widely used in segmentation — notably used in AlphaFold.