đź“– About Me
I am currently the CTO of UniUbi (宇泛), where I lead R&D at the intersection of Physical AI, embodied intelligence, and legged robotics. Our team has built a high-performance quadruped platform capable of a 720° aerial flip—to our knowledge, a world first—and developed multi-robot swarm control for coordinated real-world missions. This work has been featured in CCTV (央视) interviews.
Previously, I was a Senior Algorithm Expert in the Multimodal Group at 01.AI, focusing on computer vision, vision-language modeling, and controllable multimodal generation (image, video, and speech). Before that, I headed the AI department at Xinhua Zhiyun (an Alibaba-affiliated company), and earlier co-founded UniUbi as CTO. I received my master’s degree in 2015 from the Institute of Automation, Chinese Academy of Sciences, under the supervision of Professor Stan Z. Li.
I am broadly interested in Physical AI, including embodied agents, robot learning, and multimodal foundation models for real-world interaction. Feel free to reach me at (wangtaomarvel at gmail dot com).
🔥 News
- 2026:  🎉🎉 Interviewed by CCTV on UniUbi’s quadruped robotics R&D and real-world deployment.
- 2026:  🎉🎉 Built a quadruped robot dog featuring 720° aerial flips and multi-robot swarm / collaborative control.
đź’» Projects
High-Performance Quadruped Robot Dog @ UniUbi
- Project Duration: 2026–Present
- Developed a quadruped platform capable of a 720° aerial flip, among the first of its kind worldwide.
- Enabled multi-robot swarm control with strong multi-agent collaboration for coordinated tasks.
- Featured in CCTV media coverage; demonstrated at major robotics events such as the World Robot Conference.
Self-Imitated Diffusion Policy for Efficient and Robust Visual Navigation
- Project Duration: 2026
- Propose SIDP, a self-imitated diffusion policy that selectively imitates high-reward trajectories sampled from itself, reducing reliance on the costly “generate-then-filter” inference pipeline.
- Introduce a reward-guided self-imitation mechanism that encourages consistently high-quality trajectory generation for efficient and robust visual navigation.
Goal-oriented Navigation Instruction Generation with Tour Video Priors
- Project Duration: 2026
- Introduce VideoNIG, a goal-oriented video-grounded navigation instruction generation task that produces step-by-step instructions from ego-centric tour videos, an initial observation, and a multimodal goal—without relying on maps or graphs.
- Propose a two-stage curriculum learning framework (Action Warmup + Complexity Progression) that improves long-horizon spatial reasoning and yields instructions executable by downstream VLN agents.
Adaptive-Length Tokenizer for Autoregressive Mask Generation
- Project Duration: 2025
- Propose ALTo, an adaptive-length mask tokenizer that, for the first time, enables the model to autonomously determine the number of mask tokens based on the complexity of the input mask.
- Develop ALToLLM, which integrates ALTo into a multimodal large language model (MLLM), enabling adaptive mask token generation for object segmentation tasks.
Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal Model
- Project Duration: 2025
- Compresses a full mask into ≤32 discrete tokens, so inference reduces to the LLM’s native next-token prediction paradigm.
- Recovers masks without the original image, removing the extra segmentation decoder and simplifying the pipeline.
- Hierarchical Mask Tokens support progressive detail refinement for high-quality masks.
Text-to-Image Generation Based on Multimodal Large Models
- Project Duration: 2024
- This is an Any-to-Any Multimodal LLM project that supports using both images and text as conditions simultaneously.
- By leveraging the powerful MLLM, it achieves superior text encoding and offers better prompt following compared to open-source models.
Intelligent Image Editing Based on Multimodal Large Models
- Project Duration: 2024
- Enables various intelligent image editing tasks.
- Automatically generates Edit Type, Mask Prompt, and Output Image Prompt based on MLLM.
Multimodal-Driven Digital Human Model for Thousands of Users
- Project Duration: 2021-2023
- Supports customization of digital humans for multiple users (>1000) within a single model.
- Supports multiple input types, including voice, singing, and images.
- Extremely fast inference speed, utilizing RTX 4090 for 10x video synthesis speed.
Self-Supervised Multi-Speaker TTS Model
- Project Duration: 2022-2023
- Seamlessly integrates phoneme prediction, phoneme alignment, and vocoder to achieve true self-supervised training, leveraging the value of big data.
- Supports multi-speaker training and zero-shot voice cloning.
- Non-autoregressive design, enabling fast inference speeds.
All Speakers in One Model
Zero Shot Original Voice
Zero Shot Synthesized
📝 Publications
- Runhua Zhang, Junyi Hou, Changxu Cheng, Qiyi Chen, Tao Wang, Wuyue Zhao. “Self-Imitated Diffusion Policy for Efficient and Robust Visual Navigation”. IEEE Robotics and Automation Letters (RA-L), 2026.
- Tao Wang, Jianwei Yang, Zhen Lei, Shengcai Liao, Stan Z. Li. “Face Liveness Detection Using 3D Structure Recovered from a Single Camera”. ICB2013. Madrid, Spain, June 4-7, 2013. Citations:162
- Tao Wang, Changxu Cheng, Lingfeng Wang, Senda Chen, Wuyue Zhao. “HiMTok: Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal Model”. arXiv:2503.13026. Accepted by ICCV 2025. the code is at GitHub
- Lingfeng Wang*, Hualing Lin*, Senda Chen*, Tao Wang*, Changxu Cheng, Yangyang Zhong, Dong Zheng, Wuyue Zhao. “ALTo: Adaptive-Length Tokenizer for Autoregressive Mask Generation”. arXiv:2505.16495. Accepted by NeurIPS 2025. the code is at GitHub
🎖 Honors and Awards
- National Scholarship three times
- Champion of the first Alibaba Tianchi Big Data Competition, with a prize of 200,000 RMB
- Champion of the 2014 Double 11 Tmall Recommendation Algorithm Challenge, with a prize of 1,000,000 RMB