
Produced by Huxiu Technology Group
Author | Liang Karl
Editor | Miao Zhengqing
Header image | Huawei official website
For a long time, NVIDIA has profoundly influenced the global AI process with its advanced manufacturing processes and CUDA ecosystem. But when large model parameters soar to 10 trillion or even 40 trillion, it becomes increasingly difficult for a single chip to support it independently, and the competition for AI computing power relies more on system architecture and interconnection capabilities. In the context of limited manufacturing processes, Huawei is trying to provide system-level solutions to make millions of processors work like one computer through architecture and interconnection innovation.
In September, during the Huawei Full Connection Conference, Huawei Rotating Chairman Xu Zhijun and HiSilicon Chief Scientist Liao Heng disclosed Huawei’s trump card, Peerium computing architecture and Lingqu bus (UnifiedBus, referred to as UB). They tried to break the serial restrictions of the Turing paradigm and von Neumann’s master-slave architecture, and revealed that in statistical data, Ascend’s market share in China has surpassed NVIDIA.
In the past, a common pain point encountered by the industry was that a single chip had extremely powerful computing power, but when thousands of chips were combined into a cluster, the overall computing power efficiency would drop off a cliff due to communication delays and protocol conversions. The usual solution in the industry is to patch. For example, NVIDIA has built NVLink inside the cabinet and InfiniBand network outside the cabinet, but this “multi-protocol conversion” will bring microsecond-level delay overhead when crossing cabinets and regions.
In this regard, Xu Zhijun said that Huawei has no problems with single-chip design. Although it is still limited by domestic manufacturing processes, Huawei is already a global leader in turning the system into a computer through multi-chip interconnection. Behind this are breakthroughs in two core technologies: one is the Peerium computing architecture, which changes the master-slave architecture through nested parallelism and unified memory addressing, achieving strong scalability for million-level processors; the other is the Lingqu bus.
The change in Lingqu is that it uses a unified open protocol to “treat the network as a bus.” It eliminates the cumbersome protocol conversion in traditional solutions. UB adopts a unified protocol. It takes about 150 nanoseconds to transmit data packets from the scale-up network to the scale-out network, which saves an order of magnitude compared to the protocol conversion overhead. Millions of chips physically distributed in different cabinets work logically like one computer.
In addition, Huawei released an NPO-based (Near-chip optical transmission)’s Hi-ONE is also worthy of attention. Huawei chose a built-in light source solution, which reduces the cost by nearly 40% compared to the CPO solution, and once there is a problem with the light engine, the entire AI chip will not be scrapped. This is a compromise between engineering pragmatic choices and cost. Xu Zhijun predicts that it will be difficult for other manufacturers to do it in the next two to three years.
Behind this, the “CUDA barriers” of the software ecosystem are also being eliminated. Dr. Liao Heng pointed out that in the past 18 months, models have become very fine-grained, and the emergence of new compiler languages is gradually replacing and flattening the software gap between Ascend and NVIDIA. In the past, domestic manufacturers did not use domestic chips. The biggest concern was that the software ecosystem was not easy to use. This shift in the underlying programming paradigm can create a software breakthrough window period for domestic chips.
In the face of market demand, Xu Zhijun said, “The key to the future is not how hard we promote it, but how big our supply capacity is.” At present, super clusters with around 200,000 cards have become an urgent need for China’s cutting-edge AI laboratories to train ultra-large-scale basic models. Ascend chips are in a state of shortage of “how much is produced, how much is supplied”. Chinese head model manufacturers are also weighing computing power efficiency, original engineering services, and market-oriented options after large models are urgently needed. With the release of new chips at the end of the year, the bottleneck of domestic computing power will be further loosened.
The following is a summary of the Q&A from the media roundtable exchange during the conference:
Q: The 256,000 cards written in Liao Bo’s paper form a super node. Is this the upper limit of this technology? What is the current overall market demand situation?
廖恆:The figure of 256,000 or approximately 200,000 is not determined arbitrarily.
First of all, this type of large cluster is used for training, and the hardware infrastructure needs to support basic model training, and the current model parameter size has reached 5 trillion levels.
In the next two years, it is expected that about 6 to 7 cutting-edge AI laboratories will appear in China, all with the goal of training basic language models with parameter sizes between 10 trillion and 40 trillion.These numbers roughly match the memory capacity and size of supernodes, and therefore require an infrastructure of this scale to do training.
Second, in China, most data centers are connected to the national grid. Considering the power transmission system design of a single region and a single data center campus, 200,000 is a relatively conservative number.
As I said, everyone must be very familiar with the AI model competition that is happening around the world. In China, at least 3 to 4 first-tier players have gained widespread recognition and achieved first-tier status.
Q: How will Huawei push China’s model manufacturers to use Ascend series chips for training?
徐直軍:The PR version of 950 was first launched, mainly for reasoning, and the current market volume is not large. Supernodes based on 950DT are mainly used for training and are still under testing. Large-scale supply should be at the end of this year or early next year. We have had a lot of communication with various model manufacturers. Judging from the current test results, the training performance of 950DT is very good. I believe that starting from next year, a large number of model training will be built on the 950DT super node. The key in the future may not be how hard we push, but how big our supply capacity is. Now we can supply as much as we can produce. We hope that the industrial chain will accelerate the expansion of production, faster and more productive.
Q: Some leading model manufacturers in the United States are now calling for slowing down the development of the AI industry. What does Huawei think of this?
徐直軍:Judging from the level and pace of AI development in China and the United States, China’s models are basically open source, and relatively speaking, the stage of development is also transparent. However, because several mainstream model manufacturers in the United States have more powerful computing power, only they may currently know the stage of development. The AI risks they feel may not be felt strongly in China. China still has to speed up its pace now. When it can truly feel this risk, it may feel the same as the American head model manufacturers. But I always thought,AI must both develop and manage risks, and make AI “for good” rather than “for evil.”
Q: Does the 950 super node currently have any plans to expand overseas? If so, in which markets will it be launched? What is the domestic market share? How do you view the issue of chip autonomy?
徐直軍:At present, Shengteng is far from enough to meet the domestic market demand, so there are no plans to comprehensively expand overseas markets. But there are indeed a few countries in great need. We have tested and supplied them, but the quantity is very small.
It is difficult to calculate NVIDIA’s market share in the Chinese market, but if we look at the data we can collect, Ascend should have surpassed NVIDIA.
It is an inevitable path for China to promote chip independence.China has a population of 1.4 billion and is a large industrialized country. All life and work are related to chips. The Chinese nation is also a nation with a strong sense of worry. We cannot always be controlled by others. We cannot be told to sell you one today and told not to sell it to you tomorrow. Although the level is a little lower and the advanced nature is a little lower, it should be the inevitable way to solve the “have or not” problem and not worry about it every day. I believe that whether it is the Chinese government or the Chinese industry, including Huawei, promoting the full autonomy of chips and the entire chain of the semiconductor industry is definitely the only way to go, and I also believe that it will become a reality one day.
Q: Huawei launched Lingqu Internet technology, which is based on its long-term accumulation in the computing and communications fields. How much of it is the contribution of Huawei’s network department?
徐直軍:UB is not just the contribution of Huawei’s network department. Based on various interconnection protocols, we finally made it into a unified protocol. In the network field, IP protocols are generally used, and the latency is very high. Computing interconnection is different from traditional networks. We are talking about buses, which require low latency and ultra-high speed. When we decided to make this product, we placed it in the computing department instead of the traditional network department, so that we could truly create a high-speed, low-latency, and unified protocol UB.
Liao Heng: If you think of UB as the crystallization of the wisdom of a group of people, it was conceived by computer architects but raised by network technology.
Q: UB is a technology that takes into account both Scale-out and Scale-up. Nvidia has made NVLink and InfiniBand respectively. Why does Huawei combine Scale-up and Scale-out into one technology? What advantages does this have over other current solutions?
徐直軍:First of all, it’s not a question of whether they can do it, but whether they can do it. The latest NVL72 has a very wide bandwidth in the cabinet, reaching 1.8TB/s, but it drops to 0.2TB/s as soon as it comes out of the cabinet. To turn millions of processors into one computer, it is impossible to achieve such low bandwidth in the cabinet. Everyone hopes that the speed should be the same within the cabinet, between cabinets, between data centers, and even across zones. Only then can a huge physical space become a computer. This is the value of UB, which is based on a unified protocol and achieves high-speed interconnection throughout the entire process. Therefore, what UB’s technical route is doing is a bus, just like the bus inside a computer, not just Link. Based on UB, our inter-cabinet bandwidth can reach up to 800GB/s.
After we opened the UB protocol last year, the largest number of downloads actually came from Silicon Valley. NVIDIA’s earliest goal may not be to make NV72, but NV256, but in the end its 256 became 72; while Huawei’s 910C super node achieved 384, and more than a thousand sets were delivered. The main difference is that we used UB on the 910C super node. At that time, NVIDIA’s NVLink was mainly limited to cabinets. The network speed between cabinets was too slow and it was impossible to expand too much.
Liao Heng: Let me add why a single protocol is adopted. NVIDIA and most other manufacturers use at least two protocols to deal with different scenarios. Just like a trip, you often need to choose multiple means of transportation such as airplanes, high-speed rails, and cars at the same time. The conversion process is quite cumbersome and time-consuming. In these super nodes, a large amount of traffic needs to be transferred. If switching from NVLink to InfiniBand, it will involve at least a microsecond-level conversion time, that is, the transfer time from one protocol to another, and the delay is difficult to meet.
UB uses a unified protocol to transmit data packets from the scale-up network to the scale-out network, which may only take about 150 nanoseconds. In comparison, the conversion overhead with these protocols can be saved by at least an order of magnitude.
Q: In the Atlas 950 using UB, what proportion of optical interconnection and electrical interconnection are occupied in the scale-up scenario? Is it possible to realize optical interconnection in all cabinets in the future? What are the main obstacles?
徐直軍:Huawei has just released the NPO-based Hi-ONE module, which supports 7.2T high bandwidth. The optical engine is only about 5 centimeters away from the main chip. Only this part requires copper wires for electrical connection. It can basically be said to be an all-optical super node. At the same time, the Hi-ONE module encapsulates the light source directly into the module to achieve built-in light source. We predict that it will be difficult for other manufacturers to do it in the next two to three years. This is achieved through Huawei’s years of accumulation in optoelectronics-related fields and mass production.
廖恆:Early in my career, more than 20 years ago, we were already talking about making chip connections optical. Because the transmission distance and attenuation of light are very small, signals can be transmitted farther at high speeds.
From a circuit perspective, the interior of Hi-ONE is actually very simple, with not many components. But we have a very difficult road to mass production: optical technology involves a lot of physical principles. A slight change in temperature will affect the physical characteristics of the diode, change the refractive index, affect the performance of the waveguide, and cause it to fail to work properly.
I think the biggest obstacle to all-optical interconnects is concerns about reliability because devices can fail. At present, most manufacturers choose the external light source route. If a single light source is used to feed 32 signals, the power must be amplified 32 times. Such huge power can easily cause many problems, including problems at the coupling interface and other aspects. We chose the built-in light source solution out of pragmatic engineering choices. Hi-ONE is an excellent compromise solution reached after many trade-offs and will soon enter the mass production stage.
Q: Regarding the process and production capacity issues of Ascend chips, what are the main bottlenecks currently limiting the further release of production capacity? How long does it take for Huawei to achieve large-scale and stable supply of computing chips?
徐直軍:China’s bottleneck problem is basically the same as the global bottleneck problem. The global industry is not prepared for such a surge in development of AI. The reason why all optical module companies and storage companies can sell high prices now is not because of how good their products are, but because of a serious shortage of supply. Therefore, only when the world and China reach a balance between supply and demand at the same time can the bottleneck problem be completely ended.I expect that the global supply and demand balance will be resolved almost in 2029, and China may be a little slower.The balance of supply and demand here also needs to consider the technological changes that continue to occur in the development process of AI, including CPO, NPO, etc.
From Huawei’s perspective, compared to the past communications and IT industries, our supply now has a considerable scale; but compared with the current industry demand for AI infrastructure, the scale is indeed far from enough. To achieve the goal of providing customers with as much as they want, Huawei’s ability to achieve this balance depends on the balance of the entire Chinese industry.
Q: There are reports that Malaysia is considering deploying sovereign AI based on Huawei Ascend 910C. How does Huawei balance the needs of domestic customers and overseas customers?
徐直軍:Judging from the customer’s choice to deploy Huawei solutions, it is definitely based on geopolitical and multi-vendor considerations. Every enterprise must eliminate security risks for its long-term development, so Shengteng must be an option for all customers, although it still has a gap compared to its competitors. As for the balance between domestic and overseas, our principle is simple: domestically, we give priority to Chinese customers who are in urgent need of computing power; overseas, we serve the few customers who cannot buy it, or those who must choose another choice.
Q: What were Huawei’s internal choices or games on the two technical routes of NPO and CPO?
徐直軍:This issue seems to have never been debated at Huawei. Because Huawei makes optical components, optical modules, and Ascend chips, we know which situation is the best choice, so we reached a consensus early on to be an NPO instead of a CPO. And I think that for a long time, NPO will be the best choice from the perspectives of project realization, cost, maintenance, yield, etc.
After several years of debate on NPO and CPO in the industry, judging from the two OIF discussions, the industry’s understanding is becoming clearer. When the OIF was first discussed, some manufacturers did not support NPO and wanted to engage in CPO; but when the NPO project was recently established, there were very few opponents. Compared with the CPO solution, the NPO copper wire connection is about 5 centimeters, which is a bit longer than the CPO solution of about 5 millimeters, but it brings about a nearly 40% reduction in cost. It also brings another benefit, that is, once there is a problem with the optical engine, the entire AI chip will not be scrapped together. Of course, CPO and NPO are not a matter of who eliminates whom, but a balanced choice in various directions such as engineering, cost, industrial chain, and maintenance. There is no right or wrong.
Therefore, we will unswervingly follow the path of NPO, promote the formation of standards in the industry, and the entire industry chain will work together to develop NPO well, but this does not mean that CPO cannot do it. At the OIF, NPO received support from 40 industry chain manufacturers. It should be said that everyone has also seen the benefits of NPO.
Q: Regarding the use of 950DT in model training, what areas and issues did the model manufacturer give Huawei feedback on that need improvement? What chip and underlying adjustments and optimizations has Huawei made for pre-training large models?
廖恆:The first obstacle that Ascend chips encountered before was actually not the hardware itself, but mainly ecological barriers. Just two years ago, there were still huge obstacles at the Ascend software level, and CUDA clearly had an overwhelming advantage in this regard, because the entire academic community grew up with CUDA, and all algorithm developers have become accustomed to the convenience of PyTorch plus CUDA.
Big changes have taken place in the past 18 months.First, the mathematical complexity of the model has changed dramatically, and the programming complexity has also increased. Second, models are becoming very fine-grained, which means it is no longer possible to write programs in PyTorch alone and remain efficient. Therefore, even algorithm developers must shift to the programming style of ultra-large-scale kernels and require new programming languages. This is already the development direction of cutting-edge laboratories.
Today, as new compiler languages emerge, they are gradually replacing and balancing CUDA’s barriers,Cutting edge labs no longer take CUDA as seriously as they did two years ago.This is the removal of barriers at the software level. I think we are approaching or even reaching equilibrium and the gap has narrowed significantly.
Secondly, regarding the chip itself, we have our own advantages, as I mentioned about the UB and Nested BSP models. It’s not just programming a single chip, it’s programming large numbers of chips in a convenient way, and we’re providing those capabilities. Inside the chip, we have also learned from previous lessons and made a lot of improvements.
Overall, we’re closing the gap between NVIDIA and Ascend. I think the main differences have narrowed so much these days that it no longer feels unfamiliar to people when developing with these processors.
Q: Regarding UnifiedBus technology, is this a route with Chinese characteristics that China has taken when advanced processes are limited, or is it a possible direction for the entire industry in the future? Is it possible for NVIDIA to go in this direction?
徐直軍:UB is a path of innovation, and I think it is also the only path for future AI computing.
Because only interconnection technology like UB can realize huge computers with million-level processors. Only when millions of processors work like a computer can we truly achieve better and more efficient training and inference, so it is the only way to go. I believe that in a few years, everyone will be heading this way.
Why open the UB protocol? It is because I believe that the entire industry will follow this route. In this way, moving AI towards AGI requires the joint efforts of the entire industry. Why Liao Bo published the Peerium computing architecture in the form of a paper is because this innovative architecture is the computing architecture of the AI era. The 960 only has a faster pace, and the specifications are the same as what I talked about last year; HBM is in the 960, there is no progress of HBM, and there is no progress of the 960.
Today, UB interconnection technology is the basis for the implementation of the entire Peerium computing architecture. Whether it is UB interconnection technology or Peerium computing architecture, we believe it is the future of the AI era.
I have said many times that there is no problem with our design on a single chip, but it is still subject to domestic technology. But we are already a global leader in interconnecting multiple chips into a computer. This is based on architectural innovation and interconnected technology innovation, and these research efforts did not start after the sanctions, but before the sanctions. The first AI chip and Huawei’s entire AI strategy were released by me at the 2018 HC conference; in 2019, we released super nodes and clusters. Since then, we have believed that in order for AI to support large-scale training and inference, new architectures and new interconnection technologies are needed. Since then, we have continued to invest in research and innovation, and we have the innovative computing architecture and UB interconnection technology we are talking about today.
Related Reading
- Fuel industry warns truck stops against selling ‘red dye’ diesel after Trump lifts restrictions2026-10-08
- Original: Ternus took office to focus on design, and Apple launched intensive new products in October. Is this a strategic shift to use hardware to save itself?2026-10-07
- China’s largest dumpling IPO is coming2026-10-07
- Trillions of dollars are hitting high-speed charging piles, why is there still no solution to the queue?2026-10-07
- Disney increases park prices every year. Here’s how it avoids pricing out families2026-10-06