

Huawei Cloud has provided its answer.

Within 5 days of the launch of the Lovart Beta version, over 100,000 users registered; Genspark achieved over $10 million ARR in just 9 days; the “first-generation top influencer” Manus has repeatedly broken global attention records…
In 2025, the global AI Agent market is set to explode again, and the AI computing power market will usher in a new “craze”.
On one hand, as the complexity of global models and the demand for large-scale real-time interactions progress simultaneously, both domestic and foreign “AI computing power” concept stocks are surging. The global demand for AI computing power has not diminished due to the gradual cooling of the “hundred model war”; instead, the demand is growing.
On the other hand, in the face of the global AI Agent craze, the severe shortage of AI computing power is at the forefront, with cost control and elastic scaling becoming significant challenges for enterprises. Additionally, the configuration and management of a vast AI toolchain are extremely cumbersome, and there is a lack of a comprehensive technical foundation.
If the “hundred model war” is Level 1 of this competition, then after clearing it, a more challenging Level 2 is presented to everyone.
—— In the era of “Agents reigning supreme”, how can we achieve a breakthrough in computing power efficiency in high concurrency and high throughput inference and training scenarios?
At the Huawei Connect 2025 conference, Huawei announced that its “star product” CloudMatrix’s cloud super node specifications will be upgraded from 384 cards to a future 8192 cards; at the same time, it was announced that the CloudMatrix384 AI Token inference service is fully online, and the enterprise-level Agent platform Versatile was launched to help industry clients quickly develop various AI Agents.
In other words, currently, in the face of the increasingly explosive wave of large-scale implementation of AI Agents, Huawei Cloud has built a comprehensive technical foundation covering hardware, computing power, large models, and application open platforms.
In the era of “Agents reigning supreme”, Huawei Cloud’s answer is “CloudMatrix384 x MaaS platform x AI Token service x Versatile”, which may also be the current top configuration.

Powerful computing, globally strong
At the Huawei Connect 2025 conference, Huawei launched the latest super node products Atlas 950 SuperPoD and Atlas 960 SuperPoD, which support 8192 and 15488 Ascend cards respectively, leading the industry in key indicators such as card scale, total computing power, memory capacity, and interconnect bandwidth, and will be the global top computing power super nodes for many years to come.
Why do we need such powerful computing power?
First, let’s look at a set of news. Last year, two AI news unexpectedly went viral in both the biological research and AI communities — Google DeepMind’s latest AI model AlphaFold 3 was published in the top journal Nature, and the Nobel Prize in Chemistry was unprecedentedly awarded to DeepMind founder Demis Hassabis.
According to the data disclosed in the paper, AlphaFold3 used 256 A100 GPUs for about 20 days of training, with a training computation volume of approximately 4E22FLOP — ten times that of AlphaFold 2.
Undoubtedly, modern cutting-edge scientific research is increasingly dependent on high-performance AI computing, from protein folding schemes and interaction methods to simulating neural connection pathways in the human brain, the computational demands of cutting-edge technology have far exceeded general human cognition.
Previously, the domestic research field heavily relied on foreign high-performance AI computing platforms, and independent innovation was often long constrained by external factors. For many years, China has been conducting research on high-performance computing, resulting in excellent achievements such as the “Sunway TaihuLight” supercomputer.
In July of this year, a joint team from the Chinese Academy of Sciences officially released the “Rock Solid Scientific Foundation Model” supported by Cloudmatrix384 Ascend AI cloud services.

Rock Solid Scientific Foundation Model
The Rock Solid model will cover multiple research scenarios of the Chinese Academy of Sciences, trained with professional scientific knowledge and data to serve scientific tasks, achieving in-depth understanding of various scientific modal data such as waves, spectra, and fields; the model connects to 170 million scientific literature and real-time open scientific information, reducing the literature research time from 3 to 5 days to just 20 minutes, and some drug target discovery research efficiency can even be accelerated by more than 10 times.
Science knows no borders, but unfortunately, research does. Only by running our own large models on a super strong computing power foundation of independent innovation can scientific research and development be free from external constraints.
In addition to scientific research, another application scenario with even greater computing power requirements is intelligent vehicles.
As we all know, currently, with the explosive growth of model computing power demands on intelligent driving platforms, traditional computing architectures can no longer support the generational leap of AI technology. The degree of intelligence in vehicles is experiencing explosive growth, increasingly becoming “supercomputing centers on four wheels”, with the complexity of training intelligent driving algorithms and models skyrocketing, leading to a surge in demand for computing power utilization efficiency.

Recently, at the Intelligent Vehicle Conference 2025, Changan became the first state-owned enterprise to apply Huawei Cloud’s CloudMatrix384 super node for intelligent assisted driving research and development.
In response to the common problems faced by current vehicle-side computing power, Huawei Cloud’s CloudVeo has integrated the CloudMatrix384 super node to provide powerful support for intelligent assisted driving model training. Actual test results show that on E2E and VLA models, the performance of the CloudMatrix384 super node exceeds that of H100, making it an extremely suitable computing power platform for intelligent assisted driving model training. According to IDC data, Huawei Cloud has ranked first in China’s automotive cloud market share for several consecutive years.
For any cloud service provider or AI computing power platform, the demand for sufficiently powerful computing power is always the top priority for customers.

Explosive Growth in Token Consumption
From news to automobiles, from scientific research to search, from autonomous driving to companion robots… In the past 18 months, the application of AI Agents in China has shown exponential growth.
Tokens are the “basic units” used to measure and process text in the world of large models, akin to today’s water and electricity; tokens are the water and electricity of the intelligent world, even air.
According to data from the National Bureau of Statistics, at the beginning of 2024, China’s daily token consumption was 100 billion; by the end of June this year, the daily token consumption in China had exceeded 30 trillion, growing more than 300 times in just a year and a half, reflecting the rapid growth of AI application scale.
At the same time, the demand for MaaS (Model as a Service) is also growing rapidly.
By encapsulating AI models into standardized services through cloud platforms, Huawei Cloud’s MaaS service allows users to call powerful AI capabilities through simple API interfaces without complex and expensive model training and operations. For example, in the first half of this year, a 17-year-old high school girl built an AI application on the Huawei Cloud MaaS platform specifically to help children with cleft lip and palate undergo language rehabilitation training, truly achieving “AI for all, technology for good”.
Currently, Huawei Cloud’s MaaS service supports mainstream large models such as DeepSeek, Kimi, Qwen, Pangu, SDXL, and Wan.
Additionally, to further lower the development threshold for AI Agents while improving their performance, model adaptation, and effect tuning, Huawei Cloud’s CloudMatrix384 AI Token inference service was also fully launched at this year’s Huawei Connect conference.
For many enterprises with AI Agent development needs, the AI Token service can effectively shield complex underlying technical implementations, allowing companies to develop Agents more simply and efficiently. The xDeepServe distributed inference framework based on CloudMatrix384, with its extreme separation architecture, allows the super node to release more efficient computing power.
Breaking down the traditional “Transformer”! Attention computation and FFN computation are decoupled, sliced into finer “work islands”, creating a “super high-speed assembly line” for tokens. Single card throughput has reached 2.5-4 times that of H20, with a maximum of 2400 TPS.
For instance, 360 Nano AI, relying on CloudMatrix384’s AI Token inference service, successfully processes tens of millions of content generation requests daily.
Although Nano AI is known for its AI search capabilities, it is not just an AI search engine; it is also an AI Agent with autonomous thinking and task planning capabilities, able to automatically call various tools to complete complex tasks.
Backed by CloudMatrix384’s AI Token inference service, multiple expert intelligent agents of Nano AI can flexibly form groups, nest, and collaborate to complete complex tasks, while also running asynchronously in parallel, significantly reducing the execution time of super tasks.
360 Group founder Zhou Hongyi stated: “The ability of an intelligent agent is measured by the computing power it uses. We have developed an L4-level intelligent agent — swarm intelligent agents, where dozens of intelligent agents collaborate like a team to perform super complex tasks, but a 5-10 minute video can consume tens of millions of tokens, leading to enormous computing power consumption. Huawei Cloud’s computing power architecture can perfectly support the collaborative work of multiple foundational large models.”

Lowering the Development Threshold for Agents:
Faster, Better, More Efficient
In addition to challenges in computing power and applications, the large-scale implementation of AI Agents has always faced a significant barrier — high development thresholds.
Enterprise Agent development processes involve many nodes and complex business requirements, often leading to the “last mile” problem in the large-scale implementation of Agents — those who understand the business may not understand Agent development, while those with development capabilities may not accurately address the needs.
To address this, at the Huawei Connect 2025 conference, Huawei Executive Director and CEO of Huawei Cloud Computing Zhang Pingan officially launched the enterprise-level intelligent agent platform Versatile, which enables enterprise-level Agent generation through a simplified process. Users only need to input business logic descriptions and flowcharts, completing development in two steps, reducing a task that originally required 30 person-days to just 3 person-days, achieving a tenfold increase in efficiency.

Huawei Executive Director and CEO of Huawei Cloud Computing Zhang Pingan
For example, Huitong Travel relied on Huawei Cloud’s enterprise-level intelligent agent platform Versatile to create a dedicated Agent called “Tongbao”.
During employee use, “Tongbao” can integrate travel industry data, enterprise management knowledge, and employee historical travel records to provide real-time reminders for departure, transfer, flight changes, and corresponding solutions. After the trip, it will automatically check reimbursement compliance, making the process fast and worry-free.
With the capabilities of Versatile, “Tongbao” combines industry-leading travel domain large models and various specialized small models to meet various task requirements such as deep intent understanding, real-time data acquisition, complex computation solving, and human-like summarization; it can also connect three layers of data flywheels, timely solidifying real-time business data generated by employees and enterprises, becoming “smarter and smarter”.
Currently, AI Agents are becoming a new application form in the AI era.
According to data from Guosen Securities, currently, 30% of large enterprises with annual revenues exceeding 500 million yuan have established dedicated AI Agent teams, and 63% of B-end enterprises have listed AI Agents as a key layout for the next 12 months. Market research firm CB Insights predicts that by 2032, the AI Agent market size will exceed 100 billion.
Today, the number of global customers for Huawei Cloud AI cloud services has grown from 321 last year to 1805 this year; achieving remarkable results across various industries.
Today, the application scenarios and value of artificial intelligence are transitioning from quantitative changes to qualitative changes, the “hundred model war” is gradually cooling, and the global AI competition is moving towards the “large-scale implementation of Agents”. What we face is the “Level 2” world of large models — the era of AI Agents.
In the era of AI Agents, what kind of computing power configuration do we need?
Huawei Cloud has provided its answer.
CloudMatrix384 x MaaS platform x AI Token service x Versatile = Top configuration for the Agent era.



