Professor Xinggang Wang
Google Scholar ↗ Citations61,854 h-index91

Research interests: computer vision and deep learning, especially, visual representation learning, LLM architectures, physical AI, and agentic models.

Selected Projects

StableVQ: stable vector-quantized tokenizer training StableVQNeurIPS 2026Generative Model Bridge3D: enabling VLA models to see and act in 3D Bridge3DarXiv 2026Embodied AI AdaRoboVLG: adaptive vision-language grasping across robotic hands AdaRoboVLGarXiv 2026Embodied AI DreamWAMarXiv 2026World Model Faster-WAM: efficient inference-time future conditioning Faster-WAMarXiv 2026World Model EOVSAM: one-pass open-vocabulary segmentation with SAM 3 EOVSAMarXiv 2026Segmentation ReWorld: representation learning for world action models ReWorldarXiv 2026World Model Moebius: lightweight image inpainting framework MoebiusECCV 2026Generative Model MotionVLACoRL 2026Embodied AI RAD-2: generator-discriminator framework for closed-loop planning RAD-2arXiv 2026Autonomous Driving UniDriveVLA: unified driving vision-language-action model UniDriveVLAarXiv 2026Autonomous Driving MoDA: mixture-of-depths attention architecture MoDAarXiv 2026Efficient LLM Senna-2: aligning VLM and end-to-end driving policy Senna-2arXiv 2026Autonomous Driving OmniTrackarXiv 2026Embodied AI Spa3R: predictive spatial field modeling for 3D visual reasoning Spa3RarXiv 20263D / 4D Vision DriveLaW: unifying planning and video generation DriveLaWCVPR 2026Autonomous Driving DiffusionVL: diffusion vision-language models DiffusionVLECCV 2026Multimodal LLM VTP: scalable pre-training of visual tokenizers VTPECCV 2026Generative Model InfiniteVL: unlimited-input vision-language model InfiniteVLNeurIPS 2026Multimodal LLM 4DLangVGGT: 4D language-visual geometry grounded transformer 4DLangVGGTarXiv 20253D / 4D Vision VGT: visual generation tuning for pretrained VLMs VGTarXiv 2025Generative Model MobileI2V: fast high-resolution image-to-video on mobile devices MobileI2VarXiv 2025Generative Model TransLight: image-guided customized lighting control TransLightICML 2026Generative Model LENS: segment anything with unified reinforced reasoning LENSAAAI 2026Segmentation ReCogDrive: reinforced cognitive framework for autonomous driving ReCogDriveICLR 2026Autonomous Driving GroundingSuite: multi-granular pixel grounding benchmark GroundingSuiteICCV 2025Segmentation OmniMamba: unified multimodal understanding and generation via state space models OmniMambaECCV 2026Multimodal LLM LightningDiT: fast-converging latent diffusion with VA-VAE LightningDiTCVPR 2025Generative Model GaussTR: foundation model-aligned Gaussian transformer GaussTRCVPR 20253D / 4D Vision DiffusionDrive: truncated diffusion model for end-to-end driving DiffusionDriveCVPR 2025Autonomous Driving Senna: bridging large VLMs and end-to-end driving SennaIJCV 2026Autonomous Driving ControlAR: controllable autoregressive image generation ControlARICLR 2025Generative Model DiG: efficient diffusion models with gated linear attention DiGCVPR 2025Generative Model VADv2: end-to-end autonomous driving via probabilistic planning VADv2ICLR 2026Autonomous Driving EVA-X: self-supervised chest X-ray foundation model EVA-Xnpj Digital Medicine 2025Medical Imaging MIM4D: masked modeling with multi-view video for driving representation learning MIM4DIJCV 2025Autonomous Driving YOLO-World: real-time open-vocabulary object detection YOLO-WorldCVPR 2024Detection & Tracking Vision Mamba: bidirectional state space model for vision Vision MambaICML 2024Visual Representation 4DGaussians: real-time dynamic scene rendering 4DGaussiansCVPR 20243D / 4D Vision WeakTr: plain ViT for weakly-supervised semantic segmentation WeakTrIEEE TIP 2026Segmentation MapTR: online vectorized HD map construction MapTRIJCV 2025Autonomous Driving EVA: masked visual representation learning at scale EVACVPR 2023Visual Representation ByteTrack: multi-object tracking by associating every detection box ByteTrackECCV 2022Detection & Tracking FairMOT: fairness of detection and re-identification in MOT FairMOTIJCV 2021Detection & Tracking CCNet: criss-cross attention for semantic segmentation CCNetTPAMI 2023Segmentation

All projects →

Some Highly Influential Papers

# first author is my student, * corresponding author

Media Coverage

Academic Activities

Awards

Research Funding

Short Bio

He completed his Ph.D. and B.E. in Huazhong University of Science and Technology in 2014 and 2009 respectively. His Ph.D. supervisors were Prof. Wenyu Liu and Prof. Xiang Bai. During his Ph.D. period, he visited UCLA and Temple University where he was supervised by Prof. Alan Yuille and Prof. Longin Latecki, respectively. He also worked in the visual computing group in Microsoft Research Asia as an intern, supervised by Prof. Zhuowen Tu and collaborated with Prof. Yi Ma.