
The dividend period of large language models has entered the second half, and the physical world has become the next direction that the giants are focusing on. Capital entered the market earlier than the concept. In the first quarter of this year alone, global financing in the field of physical AI exceeded US$6.4 billion, and funds were obviously concentrated in world models and basic models.
Letting AI rehearse the physical world “in the mind” and then execute it in reality is regarded as a path to機器人, autonomous driving and the only way forward for next-generation computing. The earliest and loudest preacher on this road was Li Feifei. Her phrase “The universe is not made of words, it is made of real things”, almost became the opening statement of the entire track.
As soon as this route started, Li Feifei, who founded it, took the company to join AMD.
The next step of AI, starting from the preview world
The transaction is on an all-stock basis for approximately $8.2 billion and upon completion,She will serve as AMD’s executive vice president and chief scientist, reporting directly to Su Zifeng. This company was established just over two years ago. In February this year, it just completed a round of financing of US$1 billion, with a post-investment valuation of approximately US$5.4 billion. Just seven months later, AMD’s price was more than 50% higher.
This is AMD’s second-largest acquisition in history, second only to Xilinx, which will be worth about $50 billion in 2022. But unlike Xilinx, World Labs does not have a mature product line and has not disclosed revenue figures. Its first commercial product, Marble, was only officially launched in November last year, and its new generation model, Atlas, was unveiled less than a month before the deal was announced.
Less than a month before the deal was announced,英偉達Having just signed a contract to acquire the open source model community Hugging Face for approximately US$12.9 billion, the two chip giants successively reached out to the model layer in the same month.
In 2012, a neural network called AlexNet took a significant lead in the ImageNet image recognition competition. It was trained on Nvidia GPUs. That game is widely regarded as the starting point of the deep learning era, and it also indirectly contributed to Nvidia’s AI dominance for more than ten years. The creator of ImageNet is Li Feifei.Today, she chooses AMD。
The news came out after the U.S. stock market closed. AMD official data showed that the stock price rose only slightly by 0.40% after the market closed, and the cumulative increase in the following three trading days was less than 2%. A star company with high hopes, a record-breaking acquisition, but the attitude of the capital market is quite calm. The world model track cannot support a definite valuation at the moment.
However, part of the reason why the valuation cannot be supported lies precisely in the words “world model”. What the big language model is good at is language. It can explain physics clearly, but it may not really “understand” the physical world we live in. The core idea of the world model is to let the AI preview the world “in its mind” before deciding how to act.
The difference between world models, game engines, and physics engines is also very intuitive: the latter are rules “written” line by line by engineers, while the former are rules “learned” by AI from massive amounts of video and sensor data. The rules are precise and verifiable, but they cannot cover the complexity of the real world. The learned rules can theoretically be generalized to unseen scenes, but the cost is that they are not precise enough and may cause “hallucinations” from time to time.
Weather forecasting provides a ready reference. Traditional numerical weather forecasting cuts the atmosphere into grids and solves fluid dynamics equations grid by grid. In the past few years, Google DeepMind’s GraphCast, Huawei’s Pangu Weather, etc.The AI model no longer solves equations, but directly learns how the weather evolves from decades of meteorological data. It is already comparable to traditional methods in many indicators., the speed is much faster. The AI weather model is essentially a “weather-only” world model.
What World Model Company wants to do is to extend this from weather to the entire physical world.
A different world model
Li Feifei issued a document in June this year admitting that the planComputer vision, robotics, reinforcement learning and generative AI each say they are making world models, but they are not talking about the same thing at all.. This article gives a classification that can roughly divide the world models on the market into three levels.
The renderer generates realistic pictures or 3D scenes based on conditions, which is essentially not far from video generation and 3D generation. Given an action, the simulator predicts how the world will change, such as “If the robot hand moves 5 centimeters to the left, will the cup fall over?” The planner uses its internal understanding of the world to directly decide what to do next.This is the original academic meaning of the world model。
It falls on the specific player,騰訊混元HY-World and World Labs’ Marble mainly stay on the first layer, generating a 3D world that can be walked into and exported.Google Genie 3 and昆侖萬維Matrix-GameGoing a little further, the screen will change in real time according to the user’s operations, between the first layer and the second layer.
Israeli company Decart is hedging its bets.Lucy rewrites the video footage in real time, favoring rendering. Oasis is geared toward robotics and autonomous driving, with a bias towards simulation. NVIDIA’s Cosmos is a complete product line, covering everything from picture generation to physical deduction, with the focus falling on robotics and autonomous driving.The driving models of Waymo and Wayve are the most typical representatives of the second layer.。AMI Labs founded by Yang LikunGo straight to the third level, don’t generate pictures at all, and just learn “how the world works” in the abstract space.
However, most of the current commercial “world models” mainly focus on the first layer. Kunlun Wanwei’s Matrix-Game team once split the world model into three steps: understanding the current state, predicting the next state, and rendering.
Most models actually skip the first two steps and go directly to rendering. Fang Han, chairman of Kunlun Wanwei, announced as early as July at the World Artificial Intelligence Conference that 2026 is the “first year of the world model.” Even AMI Labs CEO Le Brun told TechCrunch in March this year that “world model” will be the next hot word.Within six months, every company will call themselves World Model Corporation and go raise money.。The market often uses the third-level vision to price first-level products.。
There are practical ways to use them, and they are very simple.
It is unfair to say that the world model is still in the PPT stage. It already has real users, but its usage is much simpler than advertised.
In February this year, Waymo released the Waymo World Model, which is based on Google Genie 3. It can simulate tornadoes, flooded streets, and even an elephant on the road, which the team has never really encountered. GAIA-4, released by British autonomous driving company Wayve in August, can already put driving AI into a closed loop:The AI makes different decisions, and the road conditions it “sees” change accordingly.. The reason why autonomous driving takes the lead is that the data is the richest and the needs are clearest, and extreme scenarios cannot be tested repeatedly in reality.
The field of robotics regards it as a “data factory”. Companies such as Agility Robotics, Figure AI, and Skild AI are all early users of NVIDIA Cosmos, using it to batch generate training data that is difficult to collect in the real world. On the content creation side, Decart said that more than 100,000 developers have made products based on its model, mainly focusing on e-commerce and live broadcasting. Tencent’s open source HY-World 2.0 is directly aimed at the production of game maps and level prototypes.
“Training robots with imperfect data” is the most common question from the outside world: the model itself doesn’t know much about physics, so using the data it generates to train robots is tantamount to teaching the wrong thing. Industry insiders believe exactly the opposite.The value of synthetic data never lies in being “true”, but in being cheap, diverse, and controllable. Pilots repeatedly practice engine failure in the simulator. The feel of the simulator is not exactly the same as that of a real airplane, but they dare not practice these situations on a real airplane. The same goes for robots. What they lack most is data that cannot be collected in reality.
This approach has proven evidence, as shown in this year’s Ctrl-World study at ICLR.Using only the synthetic data generated by the world model for post-training, the robot’s success rate on unfamiliar objects and new instructions increased from 38.7% to 83.4%.. Enterprises are also very clear about the use of synthetic data. Nvidia’s Cosmos Transfer makes the traditional simulation engine responsible for ensuring the accuracy of physics and geometry, and the world model is only responsible for making the picture realistic and diverse.
Large language models have the entire Internet’s text to learn.The world model is facing a blank. There is no ready-made “physical Internet” in the world. There is no real data of force feedback, collision and object interaction in ordinary videos, and these are exactly what robots need most.。
Zhou Zhihua, an academician of the Chinese Academy of Sciences, pointed out the second trouble at the World Artificial Intelligence Conference in July this year:The errors in the multi-step derivation of the world model are not linearly accumulated, but amplified by square levels. The further the derivation is, the more complete the collapse will be.. MemoBench, jointly released by Harvard, MIT and other institutions, tests this kind of basic skill, “object permanence”. When something leaves the field of vision and comes back, it is still in the same place. This is almost common sense to humans, but according to reports, ten mainstream world models collectively failed this test, with a maximum score of 0.58 out of 1.
The calculation of computing power is even more difficult. According to industry estimates, the GPU computing power required for a world model is several times or even dozens of times that of a large language model of the same scale. Real-time generation of high-definition images often requires multiple cards in parallel.
The problem of closed-loop business is ranked last. Leading companies such as World Labs and Decart have not disclosed revenue data. If the technology can run smoothly, it does not mean that customers are willing to continue to pay. Therefore, today’s world model is more like a capability embedded in autonomous driving, robots and content production processes, rather than a product that can support a company alone. This is also the key to understanding AMD’s acquisition.
Su Zifeng took action and Li Feifei nodded.
World Labs has raised a total of approximately US$1.2 billion in more than two years. This is an astronomical figure in any software track, but it is not generous in the cutting-edge model track. The world model needs to handle video, multiple perspectives, three-dimensional geometry and real-time interaction, and training and inference are more computationally intensive than large language models.
Commercialization is also realistic. Marble is a very good tool that can turn several photos into 3D scenes that can be roamed, edited, and exported to Unreal and Unity. Its users are mainly game, film, television, and VR teams. But it belongs to the “renderer” layer and is a tool business. And what Li Feifei really wants to do is “spatial intelligence”.What is needed is long-term investment regardless of short-term returns.。
Li Feifei said in the announcement that to promote the next generation of AI, “Requires close collaboration between model research, systems and computation“. The world model is inseparable from computing power. Instead of queuing up to buy cards, it is better to “join” directly. The open letter she published on the day of the official announcement was titled “to find a newer world“, taken from the old sailor in Tennyson’s poem “Ulysses” who refused to dock and insisted on sailing again.
“As the company develops, our ambitions continue to expand,” and the expansion of ambitions brings about an expansion in the demand for computing power, Li Feifei said, “If we can work together with AMD,It can really create an opportunity for ourselves, AMD, and the entire ecosystem to accelerate the flywheel effect between software and hardware development.”。
AMD has always been a pursuer when it comes to physics AI. NVIDIA’s Omniverse simulation platform and Isaac robot tool chain have been in operation for many years, and the Cosmos world model launched in early 2025 has now been iterated to its third generation. AMD, on the other hand, has previously publicly released mainly text and video models. It has only made sporadic attempts at world models, and is more of an investor in companies such as World Labs and Odyssey.
AMD has made four acquisitions this year:In June, it bought MEXT, which specializes in memory optimization. In July, it acquired the device-side inference software team FastFlowLM. In August, it acquired the inference chip company Taalas. In September, it acquired World Labs.. The first three deals were all about making up lessons at the inference efficiency and system level. This is the first time World Labs has bet on a direction that even Nvidia has not yet fully occupied.
Taalas’ approach is to “engrave” the AI model directly into the chip. According to Taalas, it only takes about two months from getting the model to making the chip. Now that AMD has a cutting-edge model team, it theoretically has a new possibility. It no longer just allows the hardware to adapt to the model, but the model defines the hardware. It’s just that the world model is iterated for a few months, and once the chip is finalized, it is difficult to change. Whether this road can be followed is still a question mark.
AMD buys more than just “the ability to see needs clearly”
Many analysts believe that AMD spent $8.2 billion to buy the ability to see the demand for AI computing power a few years in advance. It takes three to five years for a chip to go from definition to mass production, so it is certainly important to understand future workloads. But if you just want to see the demand clearly, cooperation is enough: the two parties will jointly optimize model training and inference on AMD GPUs from 2025, and Li Feifei also appeared on AMD’s CES stage earlier this year.
Spending 8.2 billion US dollars to fully buy it, AMD wants itFirst is exclusivity. NVIDIA is also an investor in World Labs. After the acquisition, this team only belongs to AMD. The second step is to complete the puzzle of benchmarking NVIDIA Cosmos. In the longer term, it is necessary to gain the say in defining needs.
Su Zifeng said in the interview: “The deeper you understand the entire end-to-end process, the better you can build a better system. This is why we acquired World Labs.” Talking about the goal, she said bluntly: “You will have both open source models and proprietary models”, and AMD’s ambition is to “define the future of AI computing.”
NVIDIA’s moat back then was not predicted, but “raised” in the hands of AlexNet, CUDA and generations of researchers. It is this process that AMD wants to copy:Let the next generation of AI workloads grow on your own chips and software from the beginning。
This transaction was all paid for with stocks, which was a latecomer’s strategy of playing small and broad, and also made AMD’s expansion highly dependent on its own stock price.
After the acquisition was completed, almost all the major forces in the world model circuit were tied to a computing power giant. Nvidia develops Cosmos on its own and has invested in a number of startups such as World Labs, AMI Labs, Decart and Odyssey. Google has Genie, which has been implemented in Waymo.亞馬遜It used self-developed Trainium chips to sign Decart and Odyssey as AWS customers.
Now, AMD owns World Labs.
Most of the leading companies that are still independent, such as AMI, Decart, Odyssey and General Intuition, have become associated with a certain computing power camp through financing or cooperation. “Neutral” world model companies are becoming a scarce commodity.
Regarding the research itself, this transaction may bring some changes. Su Zifeng and Li Feifei both said in a joint interview on the day of the transaction that AI is still in a very early stage, and future competition lies in the co-evolution of software and hardware. The co-design of models and chips will be accelerated. The computing characteristics of the world model are very different from those of large language models. It must simultaneously process large-scale videos, multi-view spatial information, three-dimensional geometry and low-latency interactions. When the world model team can directly participate in defining the chip for the first time, the computing power barrier may be alleviated from the hardware side.
The 3D route also received a “long-term meal ticket.” Li Feifei promised in an open letter,To build a cutting-edge research organization within AMD that can last for decades and insist on providing a widely accessible open model without having to rush to meet the funding cycle, World Labs can do longer-term research.。
Related Reading
- The pig-killing plate uses AI, and the scam begins batch production2026-10-04
- Original Entrepreneur Wu Yongming “cuts out” the new Alibaba with three swords2026-10-04
- US road rage killer’s sentence quashed because AI video of victim was shown in court2026-10-03
- What we’re learning about the FlyDubai co-pilot who allegedly stabbed the captain and tried to crash the plane2026-10-03
- Robin Li bet on big models, what did he think clearly?2026-10-03