According to an incomplete tally by the STAR Market Daily, world-model-related funding in China exceeded 66.6 billion yuan in the first seven months of 2026, with more than 50 companies receiving investment. The lane is crowded with money: Fei-Fei Li's World Labs closed a $1 billion round in February; Yann LeCun's AMI Labs set a European seed record at $1.03 billion after leaving Meta; and Jijia Shijie raised about 3.5 billion yuan in three rounds within three months. Wang Zhongyuan, president of the Beijing Academy of Artificial Intelligence (BAAI), offered a cooler read in an interview with the STAR Market Daily: 'Honestly, it is still early to talk about stable cash flow from world models.'
[1]The heat and the cold water coexist. One head VC partner said bluntly that under strict pretraining standards, only three to five startups are genuinely building world models; most projects simply swapped the lane name on their pitch decks from simulation, 3D or video generation. In June, Li called world models one of the most important and most abused terms in AI.
Wang's core judgment: VLA is the present, world models are the future. He walks through the likely payers. Autonomous driving and industrial digital twins demand the most physical fidelity and have the strongest willingness to pay, but current models cannot yet meet their core needs. Games and film tolerate hallucination better and may see the first paid cases, but mostly as auxiliary tools. Robotics is stuck on the loop of hardware, data and models - limited embodied capability constrains real-robot data collection, scarce data weakens the model, and a weak model cannot land, which caps capability again. His conclusion: the first stable cash flow will most likely appear where industrial simulation and games-and-film intersect, using world models for synthetic data and pre-evaluation.
How to tell a contract-ready world model from a paper-ready one? Wang offers what he calls the most reliable indicator - closed-loop success rate: plug the world model into a policy as an evaluation scorer, and check whether its success probability inside the model matches its success probability on a real machine. 'If these numbers do not match, nothing else matters.'
He is also doing conceptual deflation. 'Can pigs fly alongside planes in the sky? Not in the real world, but a video-generation model can.' Video generation is not a world model, he insists: video models train on plenty of science-fiction footage, where objects vanish and gravity breaks - imagination in content creation, but an error in the physical world. He lists four routes currently labeled world models - language-centric, pixel-centric, 3D-structure-centric and visual-representation-centric - and says all remain far from a true physical-world foundation model.
On background, BAAI-incubated Niju Zhen (Physis), founded by Peking University researchers Chen Boyuan and Ji Jiaming, released Wujie Physis-v0.1 in June, positioned as the first general world foundation model with next-physical-state prediction at its core, with a flagship model and open-sourced slices planned by year-end. BAAI, a research institute rather than a company, does not raise funding. Asked about the US-China gap, Wang's answer was short: 'On world models, there is no gap.'
Hot money has pushed the term to a peak; frontline researchers are trying to pull the same word back from fundraising narrative into technical definition. Demos are strong and revenue is weak - that is the common disease of this year's world-model companies. The ruler that separates the two kinds is not a launch event; it is closed-loop success rate.
[1]