XPENG unveiled a major milestone on the evening of August 27, hosting its Physical AI sharing session and the new-version experience day for its second-generation VLA model. The upgraded XOS 6.3.0 system will make its global debut on the XPENG G9L.
The core breakthrough of this second-generation VLA upgrade lies in the model's newfound grasp of the time dimension. For the first time, XPENG's physical world foundation model incorporates "time" into its framework, advancing AI comprehension from "3D space" to "4D spacetime."
This upgrade also delivers a threefold improvement in end-to-end response speed, while the new XPENG X-Foresight prediction world model debuts in production vehicles, capable of anticipating events within six seconds and forecasting potential behaviors of surrounding traffic participants.
The new version introduces a MoT hybrid architecture that dynamically allocates model capabilities across tasks, minimizing interference between scenarios such as urban roads, campus areas, and parking. With a 3.5-fold increase in on-device model parameters, reaching over 15 times the size of mainstream VLA models, combined with advanced trajectory prediction, extended effective temporal sequences, and faster decision-making, the model achieves a 20-fold enhancement in multi-dimensional comprehensive safety capabilities.
Beyond intelligent assisted driving, the second-generation VLA revamps the vehicle's digital brain with the first Master Agent, integrating VLA and VLM for a unified cockpit experience. Built on XPENG's proprietary Omni full-modality model, Master Agent understands natural language and ambiguous semantics, automatically breaking down user intent into actionable tasks and coordinating vertical agents across driving, chassis, cockpit, and body control to execute them, completing the loop from "comprehension" to "action."
The new version introduces "voice-activated roadside stopping" and "voice-commanded nearby parking," achieving a seamless chain from voice commands to autonomous execution. Leveraging a unified technology foundation, the second-generation VLA bridges L2 through L4 capabilities, paving the way for more advanced autonomous driving.
XPENG's Robotaxi, equipped with the second-generation VLA, recently obtained remote testing qualification for intelligent connected vehicles in Guangzhou, allowing driverless road tests with no safety officer in the main seat across Level 1, 2, and 3 test roads in the city.
The technology is also expanding beyond vehicles to robotics. XPENG's general-purpose humanoid robot IRON, powered by three Turing AI chips with an effective computing power of 2250 TOPS, deploys the physical world foundation model on-device, enabling autonomous completion of complex tasks without remote teleoperation, while ensuring low-latency inference and data security.
Through learned token compression and distillation training, XPENG preserves input information and model capacity while achieving high-quality visual understanding with fewer effective tokens. This allows deployment of the second-generation VLA foundation model on lower-computing platforms. As the foundation model strengthens, the distilled Turing VLA 2.0 Lite also sees significant capability gains, with initial rollout scheduled for September on the XPENG G9L Max trim.
XPENG is also advancing global deployment of the second-generation VLA. Recent localization acceptance tests in Germany showed that the model, trained on Chinese data, delivers a driving experience closely matching domestic performance on German urban roads with minimal additional local training data. The company targets obtaining regulatory approval in Europe by the first half of next year, with sequential deliveries to overseas customers, promoting the global adoption of high-level intelligent assisted driving.