đź“– About Me

I am currently the CTO of UniUbi (宇泛), leading R&D on quadruped robots (robot dogs). Our team developed a high-performance quadruped capable of a 720° aerial flip—to our knowledge, a world first—and supports multi-robot swarm control for coordinated missions. This work has also been featured in CCTV (央视) interviews.

Previously, I was a Senior Algorithm Expert in the Multimodal Group at 01.AI, working on computer vision, vision-language modeling, and controllable multimodal generation (image / video / speech). Before that, I headed the AI department at Xinhua Zhiyun (an Alibaba-affiliated company), and earlier co-founded UniUbi as CTO. I received my master’s degree in 2015 from the Institute of Automation, Chinese Academy of Sciences, under the supervision of Professor Stan Z. Li.

If you are interested in quadruped robotics, embodied AI, or multimodal systems, feel free to contact me at (wangtaomarvel at gmail dot com).

🔥 News

  • 2026:  🎉🎉 Interviewed by CCTV on UniUbi’s quadruped robotics R&D and real-world deployment.
  • 2026:  🎉🎉 Built a quadruped robot dog featuring 720° aerial flips and multi-robot swarm / collaborative control.

đź’» Projects

Featured
quadruped robot dog

High-Performance Quadruped Robot Dog @ UniUbi

  • Project Duration: 2026–Present
  • Developed a quadruped platform capable of a 720° aerial flip, among the first of its kind worldwide.
  • Enabled multi-robot swarm control with strong multi-agent collaboration for coordinated tasks.
  • Featured in CCTV media coverage; demonstrated at major robotics events such as the World Robot Conference.
SIDP visual navigation

Self-Imitated Diffusion Policy for Efficient and Robust Visual Navigation

  • Project Duration: 2026
  • Propose SIDP, a self-imitated diffusion policy that selectively imitates high-reward trajectories sampled from itself, reducing reliance on the costly “generate-then-filter” inference pipeline.
  • Introduce a reward-guided self-imitation mechanism that encourages consistently high-quality trajectory generation for efficient and robust visual navigation.
Goal-oriented Navigation Instruction Generation

Goal-oriented Navigation Instruction Generation with Tour Video Priors

  • Project Duration: 2026
  • Introduce VideoNIG, a goal-oriented video-grounded navigation instruction generation task that produces step-by-step instructions from ego-centric tour videos, an initial observation, and a multimodal goal—without relying on maps or graphs.
  • Propose a two-stage curriculum learning framework (Action Warmup + Complexity Progression) that improves long-horizon spatial reasoning and yields instructions executable by downstream VLN agents.
sym

Adaptive-Length Tokenizer for Autoregressive Mask Generation

  • Project Duration: 2025
  • Propose ALTo, an adaptive-length mask tokenizer that, for the first time, enables the model to autonomously determine the number of mask tokens based on the complexity of the input mask.
  • Develop ALToLLM, which integrates ALTo into a multimodal large language model (MLLM), enabling adaptive mask token generation for object segmentation tasks.
sym

Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal Model

  • Project Duration: 2025
  • Compresses a full mask into ≤32 discrete tokens, so inference reduces to the LLM’s native next-token prediction paradigm.
  • Recovers masks without the original image, removing the extra segmentation decoder and simplifying the pipeline.
  • Hierarchical Mask Tokens support progressive detail refinement for high-quality masks.
sym

Text-to-Image Generation Based on Multimodal Large Models

  • Project Duration: 2024
  • This is an Any-to-Any Multimodal LLM project that supports using both images and text as conditions simultaneously.
  • By leveraging the powerful MLLM, it achieves superior text encoding and offers better prompt following compared to open-source models.
sym

Intelligent Image Editing Based on Multimodal Large Models

  • Project Duration: 2024
  • Enables various intelligent image editing tasks.
  • Automatically generates Edit Type, Mask Prompt, and Output Image Prompt based on MLLM.
sym

Multimodal-Driven Digital Human Model for Thousands of Users

  • Project Duration: 2021-2023
  • Supports customization of digital humans for multiple users (>1000) within a single model.
  • Supports multiple input types, including voice, singing, and images.
  • Extremely fast inference speed, utilizing RTX 4090 for 10x video synthesis speed.
speech-driven video synthesis
music-driven video synthesis
sym

Self-Supervised Multi-Speaker TTS Model

  • Project Duration: 2022-2023
  • Seamlessly integrates phoneme prediction, phoneme alignment, and vocoder to achieve true self-supervised training, leveraging the value of big data.
  • Supports multi-speaker training and zero-shot voice cloning.
  • Non-autoregressive design, enabling fast inference speeds.

All Speakers in One Model

Zero Shot Original Voice

Zero Shot Synthesized

📝 Publications

  • Runhua Zhang, Junyi Hou, Changxu Cheng, Qiyi Chen, Tao Wang, Wuyue Zhao. “Self-Imitated Diffusion Policy for Efficient and Robust Visual Navigation”. IEEE Robotics and Automation Letters (RA-L), 2026.
  • Tao Wang, Jianwei Yang, Zhen Lei, Shengcai Liao, Stan Z. Li. “Face Liveness Detection Using 3D Structure Recovered from a Single Camera”. ICB2013. Madrid, Spain, June 4-7, 2013. Citations:162
  • Tao Wang, Changxu Cheng, Lingfeng Wang, Senda Chen, Wuyue Zhao. “HiMTok: Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal Model”. arXiv:2503.13026. Accepted by ICCV 2025. the code is at GitHub
  • Lingfeng Wang*, Hualing Lin*, Senda Chen*, Tao Wang*, Changxu Cheng, Yangyang Zhong, Dong Zheng, Wuyue Zhao. “ALTo: Adaptive-Length Tokenizer for Autoregressive Mask Generation”. arXiv:2505.16495. Accepted by NeurIPS 2025. the code is at GitHub

🎖 Honors and Awards

  • National Scholarship three times
  • Champion of the first Alibaba Tianchi Big Data Competition, with a prize of 200,000 RMB
  • Champion of the 2014 Double 11 Tmall Recommendation Algorithm Challenge, with a prize of 1,000,000 RMB