

Edited by Wang Lingfang
On 2026/8/27, XPeng told the world through a launch event:The era of physical AI has begun.
Today, XPeng held a Physical AI sharing session themed "TIME".Introducing the Next-Gen VLA Large Model: XOS 6.3.0It will be globally debuted on the Xpeng G9L. More importantly, the launch of the second-generation VLA reaffirms that Xpeng is not just an automaker capable of autonomous driving, nor merely an autonomous driving company that builds cars.XPeng takes another solid step toward becoming a global physical AI company focused on embodied intelligence.
This strategic shift is built on XPeng's 12 years of full-stack R&D investment and sustained commitment to physical AI, marking the emergence of a new species.A clear strategic framework of "one AI foundation and two embodied intelligence products" has emerged: atop the AI infrastructure, automobiles andRobot(14.460, 0.04, 0.28%) The two core entities are resonating.The news that XPeng's robotics business secured over $2 billion in its first round of funding and achieved a valuation exceeding $3 billion on 8/24 represents the most direct market validation of this strategy.
The race in physical AI has just begun. Xpeng is among the leaders alongside Tesla.
01
From Seeing the World to Understanding Time
"Past AI was limited by various constraints, focusing primarily on spatial understanding... but more often than not,"To build physical AI, you must understand time. This is the first major shift we've made: introducing the concept of time.”
Liu Xianming, Head of the General Intelligence Center at XPeng Group, unveiled the underlying logic behind the second-generation VLA model upgrade during the launch event. The real challenge in physical AI is not just enabling vehicles to see what's happening, but helping them understand what is occurring now, what happened before, and what could happen next.
Traditional ADAS models can detect vehicles, pedestrians, roads, and traffic signals, but to them, the world is just a series of static images.To make complex decisions in the physical world, simply "seeing" is not enough—vehicles need to understand the past, present, and future.

Xpeng G9L will be the first to feature the new second-generation VLA version.
The second-generation VLA incorporates "time" into the model for the first time, advancing AI's understanding of the world from 3D space to 4D spacetime.
This breakthrough was achieved through several key technologies:
First is the Infini-VLA long-horizon architecture, which enables the model to maintain an effective temporal memory of up to 30 seconds.
The system no longer "sees and forgets" when a vehicle drifts across lanes, the car ahead slows down repeatedly, or pedestrians approach the intersection. Instead, it remembers these targets' prior continuous actions to assess the current situation and predict their next intentions.
Now, let's look at the "present." Previously, Tesla FSD's standout feature in Silicon Valley was its exceptional speed, with no perceptible lag.
How does XPeng's second-generation VLA solve this?
Liu Xianming highlighted the streaming autoregressive inference mode, which boosts end-to-end response speed by 300%.
Traditional models follow a discrete "see → compute → output → see again" pattern, similar to large language models: they generate token by token, with each new token conditioned on all preceding tokens.
But the second-generation VLA enables parallel "see-think-act" processing. When facing sudden events like a lead vehicle braking abruptly, pedestrians turning back, or neighboring cars cutting in, the car responds with the smoothness and agility of an experienced driver.
Certainly, the faster response speed is also linked to XPeng's use of cameras as input signals. Cameras have a higher frame rate than LiDAR, giving them a natural advantage.
Solving the "past" and "present" isn't enough; we must also address the "future" by equipping models with predictive capabilities.
Liu Xianming introduced X-Foresight, a predictive world model that forecasts the possible movements of surrounding traffic participants over the next 6 seconds.He stated that while the model can predict up to 21 seconds, 6 seconds is sufficient. Based on these predictions, the system can evaluate multiple candidate trajectories in advance, adjusting speed, following distance, and route proactively to select safer, more natural driving behaviors.
Will edge-side compute power support model operation after the "join time"? Moreover, Liu Xianming has consistently advocated for the "scaling law," emphasizing continuous increases in model parameter scale.
How to achieve both?
Liu Xianming also introduced the MoT hybrid architecture, which assigns different road scenarios and tasks to specialized "experts"—such as urban autonomous driving, campus roaming, and parking specialists—to prevent cross-scenario interference. During scenario transitions, the system maintains seamless and continuous operation by leveraging visual memory for cross-scenario reuse.
The MoT hybrid architecture is similar to the MoE architecture popularized by DeepSeek-V3, both of which invoke only a subset of sub-networks on demand to reduce inference compute costs.
Of course, "scaling laws" require models to grow. Liu Xianming also explained on-site that only a larger model can achieve sufficient generalization capability, making rapid entry into the global market possible.
Finally,The second-generation VLA edge model features a parameter count 3.5 times higher than the first generation, exceeding mainstream VLA models by more than 15 times.
The prescription is set; what about the results?
The XPeng test vehicle equipped with the second-generation VLA has already demonstrated unprecedented capabilities in intelligent driving.
In the video played at the press event,Vehicles equipped with the 2nd-gen VLA system can anticipate a U-turn intention from the car ahead. Even when there is space to pass, they wait for the vehicle to complete its turn before proceeding, ensuring greater safety.

In other videos, the new version of the intelligent driving system in equipped vehicles demonstrates significantly faster response times. It quickly reacts to sudden obstacles like tires on the road, emerging two-wheeled vehicles, or toys tossed into the path. On construction zones where a dual-lane road narrows to a single lane, if an oncoming vehicle is traveling in the wrong direction, the ego vehicle correctly identifies the situation, stops to yield, and then accelerates promptly once the oncoming vehicle has cleared the way.
More impressively, a video showcases the vehicle's automatic ferry boarding and disembarking. The New Zhiji version can autonomously navigate in unfamiliar environments, follow the lead vehicle to queue and park, and then automatically exit the ferry upon arrival to merge onto public roads.
One sentence summary: It remembers, thinks, and predicts.
XPeng's second-generation VLA has completed localization acceptance testing in Munich, Germany. Trained on Chinese data, this model delivers driving experiences in German urban roads that closely match those in China with minimal additional local training data—demonstroring true global generalization capability.
02
The real moat isn't the model—it's the ability to build models.
With the next-gen VLA large model, XPeng now holds the "Dragon-Slaying Blade"—who can compete?
However, Liu Xianming believes thatThe model will keep evolving, but the true long-term moat is the ability to iterate continuously.
Liu Xianming emphasized that behind the model lies a system dedicated to building a comprehensive Physical AI technology framework spanning data, training, inference, prediction, and action.

The term mentioned by He Xiaopeng and Liu Xianming is AI Infra (AI Infrastructure).
Over the past few years, XPeng has continuously invested in building a complete AI infrastructure covering data, training, simulation, evaluation, deployment, and compute.
Regarding data,Liu Xianming stated that the model's single-training data throughput has reached 1.1 billion clips, representing a 10-fold increase compared to six months ago.
"Data always has two dimensions: scale and quality," said Liu Xianming. "We need not only large-scale data, but also data that is sufficiently diverse and high-quality."
XPeng has established a vector-based data analysis system capable ofMassive Data(13.060, 0.31, 2.43%) for localization, and also to uncover corner cases. When issues arise, this system provides similar data for reinforcement training.
Liu Xianming also highlighted XPeng's AI Infra simulation capabilities. "When real-world vehicle validation can't keep up with model training speed, simulation is the best tool for the job."Over the past two months, XPeng Simulation's daily new model count increased by 290%, nearly tripling."Simulation is also another core moat of ours."
Intelligent driving is now a key factor in consumers' purchasing decisions, yet many automakers still treat it as just another feature or spec—outsourcing it to suppliers and simply selecting the best option.
From XPeng's perspective, intelligent driving is just the tip of the iceberg in its AI ecosystem—a key application scenario for physical AI.
The difference in these two mindsets is stark: when cars evolve into embodied AI agents, the former has no chance. Companies with a vision similar to XPeng may not have moved as early or built as solid a foundation.
XPeng built a systematic capability starting from AI Infra, enabling models to evolve continuously and rapidly.
This is XPeng's true long-term moat.
03
One AI foundation, two intelligent products
"Automobiles are merely the first major platform for scaling physical AI."
Another carrier is, naturally, the robot.
Liu Xianming displayed a slide during the press conference:Control two distinct embodiments using the same VLA foundation model: a XPeng vehicle on the left and an indoor robot navigation system on the right.The model achieves strong generalization across different embodiments through more generalized training and larger scale.

This embodies XPeng's strategy of "one AI foundation and two embodied intelligence products."
As the current physical platform for large-scale deployment, vehicles enable users to experience faster response times, smoother path planning, more natural driving trajectories, and forward-looking decision-making. The long-term value lies in reusing this core capability of understanding, predicting, and acting in the physical world—whether for Robotaxis, robots, or flying cars. These different "bodies" face a shared underlying challenge: how to perceive the real world, anticipate changes, and take effective action within it.

Xiaopeng's second-generation VLA, built on a unified technical foundation, now delivers end-to-end capabilities from L2 to L4 and continues to advance toward higher levels of autonomous driving. Xiaopeng Robotaxis equipped with the second-generation VLA have recently obtained Guangzhou's qualification for remote testing of intelligent connected vehicles, enabling driverless road tests without a safety driver on designated Class 1, 2, and 3 test routes in Guangzhou. This marks the beginning of Xiaopeng's key phase for validating fully driverless operation on public roads.
On the robotics front, on 8/24, XPeng Robotics completed its Series A equity financing round, raising over $2 billion with a post-money valuation exceeding $3 billion (approximately RMB 4 billion), setting a new record for single-round private equity financing in China's embodied AI industry.

He Xiaopeng has consistently emphasized that XPeng is a mobility explorer in the physical AI world and an embodied intelligence company with a global vision.
This clearly reflects the capital market's recognition of XPeng's leadership in physical AI, its technology roadmap, mass production capabilities, and long-term commercial value.
Regarding this news, He Xiaopeng posted on Weibo saying"In the near future, we will bring the first advanced humanoid robot into real-world use. She will be unlike any robot you've seen before: more human-like, safer, no remote control needed, natural interaction, personal memory, and true value creation—not just demos, but mass production that brings her to life."
"Advanced," "remote-free," and "memory-enabled"... XPeng's humanoid robot isn't mass-produced yet, but it starts from a higher ground than today's hottest robotics companies.
Not only autonomous vehicles and Robotaxis, but also flying cars and robots. In this sense, as Liu Xianming says, XPeng has always positioned itself as a "physical AI company."
What does it mean to be a physical AI company? What capabilities must it possess? There is no blueprint. Instead of predicting the future, we are building it—that is what XPeng is doing.
In 2025 12, after experiencing Tesla FSD V14.2, He Xiaopeng made a bet with Liu Xianming: If by 2026 8 30, XPeng's VLA achieves the same overall performance in China as Tesla FSD V14.2 does in Silicon Valley, He Xiaopeng will build a Chinese-style cafeteria in Silicon Valley. Otherwise, Liu Xianming will run naked across the Golden Gate Bridge.
Shortly after the Physical AI sharing session on 8/27, He Xiaopeng posted on Weibo saying,On standard roads, XPeng's 2nd Gen VLA matches the best in the world. In complex scenarios like narrow streets, campuses, and parking lots, it delivers even better performance—enabling direct navigation to your parking spot for a safer, more efficient experience.

He Xiaopeng expressed his endorsement of the new second-generation VLA version on Weibo.
Replacing "world-class products" with Tesla FSD means He Xiaopeng's bet has been settled, and Liu Xianming has won.
Beyond the wager, what matters most is that Xiaopeng's models now understand time itself—finding better solutions in complex physical systems and forging their own path.
