← Back to News List
Industry News2026-08-24

Beijing Humanoid Pelican-VL 2.0 Model Now Commercially Available; New Tiankang Omni Launched | Spotlight on World Robot Conference in 2026

Zhongzhi Chuangxin Beijing's Pelican-VL 2.0 model for humanoid robots goes commercial; new Tiankong Omni launched | Focused on World Robot Conference 2026

On 8 20, Beijing Humanoid unveiled the unified Pelican-Unify model and the open-source foundation TianGong Omni, signing agreements covering 6 countries including China, Germany, Japan, and South Korea. End-to-end shared representations replace modular stitching; industry competition is shifting from "speed" to "standards and ecosystems."

Beijing Humanoid Unveils Unified Model and TianGong Omni: Embodied Intelligence Shifts from "Modular" to "End-to-End"

On 8 20, the second day of the World Robot Conference, during the Embodied Intelligence Application Innovation Theme Event titled "Embodied Awakening · Wisdom in All Forms,"Beijing Humanoid Robot Innovation CenterLaunched the unified Pelican-Unify model and the lightweight humanoid robot TianGong Omni, completed multiple strategic signings, and partnered with organizations across 6 countries including China, Germany, South Korea, Japan, and Spain.

This launch deserves its own spotlight: beyond the "national team" label, it has put a key industry debate on the table—should an embodied brain be built from multiple modules, or should it follow an end-to-end unified approach?Beijing HumanoidChose the latter, paired with an open hardware base, and handed the choice to the ecosystem.

End-to-end is not concatenation; it's shared representation.

Pelican-Unify's core vision is to advance the technology roadmap from modular to end-to-end.

Comparison of two technical approaches: Modular assembly (fragmented capabilities, accumulated errors) vs. End-to-end unified architecture (shared representations, co-evolution).

As officially defined, "end-to-end" does not mean simply concatenating outputs from separate models for vision-language understanding, task planning, action execution, and world prediction. Instead, it involves building a shared representation space where multi-dimensional inputs—such as text, images, videos, and actions—are unified, encoded, aligned, reasoned over, and planned together, ultimately generating future video frames and low-level actions in an end-to-end manner.

Ju Xiaozhu, head of the Beijing Humanoid Large Model, breaks this design down into three unifications: unified understanding, mapping scenarios, instructions, visual context, and action history to a shared semantic space; unified reasoning, transforming task intent, action selection, and future consequences into supervisable evolutionary reasoning; and unified generation, jointly outputting future video and underlying actions within a single diffusion decoding process.

The industry implication is clear: under a modular architecture, frequent model calls lead to error accumulation, making it difficult to iterate collaboratively with real-world data and hitting a performance ceiling. In contrast, shared representations enable capabilities to reinforce each other, closing the loop from "understand → reason → plan → act" — moving from "seeing the world" to "simulating the future and executing precisely."

Brain: From Understanding to Doing

Complementing the unified model is the embodied brain model, Pelican-VL 2.0. Building on spatiotemporal understanding and physical perception, it leverages post-training infrastructure and reinforcement learning to significantly enhance agentic capabilities. Key abilities such as long-horizon task execution and tool invocation have been markedly improved. As officially stated: from "understanding" to "doing."

Public data also shows that Pelican-VL 2.0 significantly outperforms GPT5.5 in embodied task evaluations, with general agent capabilities matching or surpassing leading domestic commercial large models. It is also the first embodied large model in China to obtain regulatory filing approval from the Cyberspace Administration of China (CAC), and it leads the industry in token usage volume.

Positioning "Brain" against general large models signals a shift: Embodied brains are evolving from "robot-specific small models" into composite entities combining "general agent capabilities + physical interfaces." Evaluation metrics have expanded beyond single-operation success rates to include long-horizon tasks and tool-calling as key agent indicators.

Body: Open base of 1.35 meters

If the unified model is the "brain," Tiangong Omni is its matching "body with a cerebellum." Standing at 1.35 meters tall and weighing 39 kilograms, it is officially positioned as the world's first open foundation for Humanoid Native systems that combine perception, mobility, and full-body general-purpose control.

The product line offers three tiers: the flagship version with 31 degrees of freedom, featuring the Thor T5000 compute domain controller and 2070 TFLOPS AI performance, supporting dexterous hands, VR, and homogeneous arm expansion; the standard version with 25 degrees of freedom, focused on autonomous navigation and human-robot interaction; and the base version with 23 degrees of freedom, supporting quick-swap batteries.

Tiangong Omni Open Foundation: Fully open four-layer APIs (joint, sensor, system, motion control) with flagship, standard, and basic product tiers.

The demo didn't showcase high-dynamic feats like sprinting or flips. Instead, it demonstrated walking on plank stilts, navigating stairs, interacting with clones, and office printing—aligning with its positioning: agile maneuvering, complex terrain adaptation, and multi-task execution. It empowers developers by unifying the robot's body, perception, motion, and data collection into a single platform.

Openness is what makes this machine truly remarkable. TianGong Omni exposes four layers of interfaces: joints, sensors, system, and motion control APIs. It comes pre-integrated with a mature robotic cerebellum for motion planning and high-compute multimodal perception, supports edge deployment of large models, and includes a 500-hour dataset of dynamically filtered action sequences.

Xiong Youjun's Strategy: Don't Steal Customers, Build Local Foundations

Beijing Humanoid CEO Xiong Youjun was straightforward: Omni targets industrial scenarios and avoids duplicating what other companies already do. We aim to provide a standard platform for partners to build on, rather than competing directly with vendors like Songyan Dynamics or Accelerate Evolution. I won't chase customers myself; instead, I'll lay the foundation beneath you.

This logic is scarce in the embodied AI industry. The biggest bottleneck for humanoid robots today isn't the hardware itself, but the lack of large-scale application data. By opening our platform to more developers, we get real-world deployment and data accumulation that fuels LLM iteration. Our full-stack product suite, including the "Tiangong + Huishi Kaiwu" dual platforms, is accelerating this closed-loop ecosystem.

Xiong Youjun doesn't shy away from the debate over ROI: while the ROI of deploying humanoid robots in factories currently lags behind that of industrial robots or human labor, the gap is narrowing rapidly. He's betting not on today's single deal, but on the compounding returns of "standardization + ecosystem" over the next two to three years.

Industry Benchmark: Everyone is moving toward the same goal

Zooming out, the information released alongside the conference almost all points toward "unification":Tencent Robotics X LabCollaborated with Tsinghua University and others to launch GeniWorld, an interactive world model that converts numerical actions into visual action representations and predicts future video for synthetic training data.Daimeng RobotIntroducing Daimon-TWM, a tactile anchoring world model that fuses predictive and real-time tactile feedback for motion correction.Ant LingboLingBot-VLA 2.0 integrates 6 hours of real-world physical data, covering 17 manufacturers and 20 configurations, with inference latency reduced to 130 milliseconds.

InfoQ summarizes this trend as "route mixing": the industry has moved beyond the debate of choosing between VLA and world models toward integration.Star MapXpeng MotorsWait until top vendors adopt a unified approach; Gartner predicts that VLA will remain dominant in the next one to two years, with world models gradually integrated. Pelican-Unify effectively implements this "hybrid" strategy by unifying shared representations in one step.

There is also a comparison on the performance side.Unitree RoboticsRelease the "Superman" prototype before the conference, capable of jumping approximately 2 meters in place and achieving a top running speed of 12.66m/s.UBTECH Robotics, Agibot RobotGalaxy UniversalDeployment progress varies across vendors, with clear divergence in strategies: some have pushed "running more" to its limit, while others are focusing on "working smarter" and "ecosystem openness"—Beijing Humanoid Robot falls into the latter category.

The supply chain side is also coordinating.OrbbecProvides 3D visual perceptionLeader Harmonic DriveComponent suppliers will showcase joint modules and reducer solutions at the conference exhibition area.DobotImplement the "One Brain, Multiple Bodies" concept in the bionic dinosaur exhibits at the Shenzhen Natural History Museum.YoubitechLaunched the industrial embodied AI model "Zhihe," enabling one brain to orchestrate multiple forms across wheeled and humanoid robots. Open foundation + shared intelligence is becoming vendors' new choice.

Unanswered questions

The "open foundation" approach also carries uncertainty: platformization means thinner gross margins and reliance on ecosystem growth, yet current use cases remain unproven. It's unclear whether a developer ecosystem will emerge, and the battle over standards is far from settled.

Another highlight on the commercial front: Tiangong 3.0 has entered into a deep strategic partnership with Mercedes-Benz, serving as a "Silicon-Based Recommender" in showrooms, auto shows, racing events, and live streams. It handles end-to-end services including greeting, vehicle explanations, and interactive consultations. This marks the first international validation of the "Standard + Scenario" model.

Embodied AI is still waiting for its "Android moment." This time, what will decide the winner isn't a cooler robot, but an open standard that's easier to adopt and cheaper to deploy—along with the app ecosystem built around it.

Zhongzhi Chuangxin LogoZhongzhi Chuangxin

Zhongzhi Chuangxin (Guangzhou) Artificial Intelligence Technology Co., Ltd.

AI-powered decision-making, embodied intelligence enabling smart manufacturing

Contact Us

  • 18319080137
  • jason@zhyunsoft.com
  • Nansha District, Guangzhou

Business Hours

Monday to Friday

9:00 - 18:00

© 2026Zhongzhi Chuangxin (Guangzhou) Artificial Intelligence Technology Co., Ltd.

All implemented technologies and equipment solutions are self-developed real-scene AI projects by the enterprise.