đź“– About Me

I am currently the CTO of UniUbi (宇泛), where I lead R&D at the intersection of Physical AI, embodied intelligence, and legged robotics. Our team has built a high-performance quadruped platform capable of a 720° aerial flip—to our knowledge, a world first—and developed multi-robot swarm control for coordinated real-world missions. This work has been featured in CCTV (央视) interviews.

Previously, I was a Senior Algorithm Expert in the Multimodal Group at 01.AI, focusing on computer vision, vision-language modeling, and controllable multimodal generation (image, video, and speech). Before that, I headed the AI department at Xinhua Zhiyun (an Alibaba-affiliated company), and earlier co-founded UniUbi as CTO. I received my master’s degree in 2015 from the Institute of Automation, Chinese Academy of Sciences, under the supervision of Professor Stan Z. Li.

I am broadly interested in Physical AI, including embodied agents, robot learning, and multimodal foundation models for real-world interaction. Feel free to reach me at (wangtaomarvel at gmail dot com).

🔥 News

  • 2026:  🎉🎉 Interviewed by CCTV on UniUbi’s quadruped robotics R&D and real-world deployment.
  • 2026:  🎉🎉 Built a quadruped robot dog featuring 720° aerial flips and multi-robot swarm / collaborative control.

đź’» Projects

Featured
quadruped robot dog

High-Performance Quadruped Robot Dog @ UniUbi

  • Project Duration: 2026–Present
  • Developed a quadruped platform capable of a 720° aerial flip, among the first of its kind worldwide.
  • Enabled multi-robot swarm control with strong multi-agent collaboration for coordinated tasks.
  • Featured in CCTV media coverage; demonstrated at major robotics events such as the World Robot Conference.
SIDP visual navigation

Self-Imitated Diffusion Policy for Efficient and Robust Visual Navigation

  • Project Duration: 2026
  • Propose SIDP, a self-imitated diffusion policy that selectively imitates high-reward trajectories sampled from itself, reducing reliance on the costly “generate-then-filter” inference pipeline.
  • Introduce a reward-guided self-imitation mechanism that encourages consistently high-quality trajectory generation for efficient and robust visual navigation.
Goal-oriented Navigation Instruction Generation

Goal-oriented Navigation Instruction Generation with Tour Video Priors

  • Project Duration: 2026
  • Introduce VideoNIG, a goal-oriented video-grounded navigation instruction generation task that produces step-by-step instructions from ego-centric tour videos, an initial observation, and a multimodal goal—without relying on maps or graphs.
  • Propose a two-stage curriculum learning framework (Action Warmup + Complexity Progression) that improves long-horizon spatial reasoning and yields instructions executable by downstream VLN agents.
sym

Adaptive-Length Tokenizer for Autoregressive Mask Generation

  • Project Duration: 2025
  • Propose ALTo, an adaptive-length mask tokenizer that, for the first time, enables the model to autonomously determine the number of mask tokens based on the complexity of the input mask.
  • Develop ALToLLM, which integrates ALTo into a multimodal large language model (MLLM), enabling adaptive mask token generation for object segmentation tasks.
sym

Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal Model

  • Project Duration: 2025
  • Compresses a full mask into ≤32 discrete tokens, so inference reduces to the LLM’s native next-token prediction paradigm.
  • Recovers masks without the original image, removing the extra segmentation decoder and simplifying the pipeline.
  • Hierarchical Mask Tokens support progressive detail refinement for high-quality masks.
sym

Text-to-Image Generation Based on Multimodal Large Models

  • Project Duration: 2024
  • This is an Any-to-Any Multimodal LLM project that supports using both images and text as conditions simultaneously.
  • By leveraging the powerful MLLM, it achieves superior text encoding and offers better prompt following compared to open-source models.
sym

Intelligent Image Editing Based on Multimodal Large Models

  • Project Duration: 2024
  • Enables various intelligent image editing tasks.
  • Automatically generates Edit Type, Mask Prompt, and Output Image Prompt based on MLLM.
sym

Multimodal-Driven Digital Human Model for Thousands of Users

  • Project Duration: 2021-2023
  • Supports customization of digital humans for multiple users (>1000) within a single model.
  • Supports multiple input types, including voice, singing, and images.
  • Extremely fast inference speed, utilizing RTX 4090 for 10x video synthesis speed.
speech-driven video synthesis
music-driven video synthesis
sym

Self-Supervised Multi-Speaker TTS Model

  • Project Duration: 2022-2023
  • Seamlessly integrates phoneme prediction, phoneme alignment, and vocoder to achieve true self-supervised training, leveraging the value of big data.
  • Supports multi-speaker training and zero-shot voice cloning.
  • Non-autoregressive design, enabling fast inference speeds.

All Speakers in One Model

Zero Shot Original Voice

Zero Shot Synthesized

📝 Publications

  • Runhua Zhang, Junyi Hou, Changxu Cheng, Qiyi Chen, Tao Wang, Wuyue Zhao. “Self-Imitated Diffusion Policy for Efficient and Robust Visual Navigation”. IEEE Robotics and Automation Letters (RA-L), 2026.
  • Tao Wang, Jianwei Yang, Zhen Lei, Shengcai Liao, Stan Z. Li. “Face Liveness Detection Using 3D Structure Recovered from a Single Camera”. ICB2013. Madrid, Spain, June 4-7, 2013. Citations:162
  • Tao Wang, Changxu Cheng, Lingfeng Wang, Senda Chen, Wuyue Zhao. “HiMTok: Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal Model”. arXiv:2503.13026. Accepted by ICCV 2025. the code is at GitHub
  • Lingfeng Wang*, Hualing Lin*, Senda Chen*, Tao Wang*, Changxu Cheng, Yangyang Zhong, Dong Zheng, Wuyue Zhao. “ALTo: Adaptive-Length Tokenizer for Autoregressive Mask Generation”. arXiv:2505.16495. Accepted by NeurIPS 2025. the code is at GitHub

🎖 Honors and Awards

  • National Scholarship three times
  • Champion of the first Alibaba Tianchi Big Data Competition, with a prize of 200,000 RMB
  • Champion of the 2014 Double 11 Tmall Recommendation Algorithm Challenge, with a prize of 1,000,000 RMB