The artificial intelligence landscape, long dominated by a singular focus on raw computational power, is undergoing a profound transformation. As Agentic AI applications move from conceptualization to large-scale deployment, the industry’s attention has pivoted from merely accumulating more Graphics Processing Units (GPUs) to a far more intricate challenge: maximizing the efficiency with which these vast compute resources are utilized. This paradigm shift was the central theme of H3C’s major announcements at the highly anticipated 2026 Apsara Conference, where the company articulated a vision for AI infrastructure built not on hardware stacking, but on holistic system-level optimization.
The Evolving Demands of Agentic AI and the Apsara Conference Backdrop
For years, the mantra in AI infrastructure was straightforward: more compute. This drive fueled a massive investment in powerful GPUs, with the belief that increasing the number of processors would proportionally accelerate AI development and deployment. However, the emergence of Agentic AI, characterized by its continuous interaction with foundation models and dynamic task execution, has exposed the limitations of this singular focus. These sophisticated AI agents demand not just brute force, but seamless, low-latency, and highly efficient processing of information at an unprecedented scale, often measured in trillions of tokens.
The Apsara Conference, Alibaba Cloud’s flagship annual technology event, served as the ideal platform for H3C to unveil its refined strategy. Renowned as one of Asia’s most influential technology gatherings, the conference consistently attracts industry leaders, developers, and researchers keen to explore the cutting edge of cloud computing, AI, and big data. H3C’s presence underscored the critical juncture the AI industry has reached, signaling that the future of AI innovation hinges less on incremental hardware gains and more on intelligent system design. The company’s pronouncements at Apsara 2026 positioned it at the forefront of this evolving narrative, advocating for a shift in competitive metrics from raw GPU count to "useful Tokens per GPU."
Beyond Raw Horsepower: The Bottlenecks of Unoptimized Clusters
Zhu Shiyin, General Manager of H3C’s Advanced Technology Research Department, highlighted the inherent complexities of modern AI clusters. He emphasized that as these clusters expand from hundreds to thousands, and potentially tens of thousands of GPUs, the individual processing power of each GPU becomes just one variable in a much larger equation. The efficiency of the entire system can be severely eroded by a multitude of factors: idle compute cycles, network congestion, insufficient data supply, and the intricate operational and failure challenges inherent in managing such colossal distributed systems. Simply put, adding more GPUs does not guarantee a proportional increase in performance if the surrounding infrastructure cannot keep pace. This often leads to significant capital expenditure without corresponding gains in AI model training or inference throughput, directly impacting the return on investment for enterprises.
Industry analysts echo this sentiment, pointing out that the total cost of ownership (TCO) for AI infrastructure is increasingly influenced by operational efficiency rather than just upfront hardware costs. Energy consumption, for instance, has become a major concern. Unoptimized clusters lead to higher Power Usage Effectiveness (PUE) ratios, meaning a larger proportion of electricity is consumed by cooling and other overheads rather than by actual compute, escalating operational expenditures and environmental impact.

H3C’s Holistic Vision: A System-Level Approach to AI Infrastructure
In response to these burgeoning challenges, H3C presented a comprehensive product lineup at the Apsara Conference that underscored a clear, integrated logic: AI infrastructure must transcend hardware stacking and embrace system-level optimization across compute, networking, storage, software, and operations. This strategy is designed to ensure that every component works in concert to maximize the generation of useful tokens, which are the fundamental units of information processed by large language models and agentic AI.
The company’s approach is meticulously structured to address each potential bottleneck, transforming AI infrastructure from a collection of powerful components into a finely tuned, integrated data processing powerhouse. This vision aligns with the growing industry consensus that true AI scale requires a symbiotic relationship between hardware and software, where performance is measured by the end-to-end efficiency of the AI pipeline.
Supercharging Compute: The UniPoD S80000 Series SuperPod
Central to H3C’s compute strategy is the UniPoD S80000 Series SuperPod. This flagship offering embodies the concept of heterogeneous computing, moving beyond a sole reliance on GPUs. The SuperPod supports configurations ranging from 32 to an impressive 1,024 GPUs, with the capacity to scale further to 16,384 GPUs in massive deployments. Critically, it integrates a diverse array of computing resources, including traditional Central Processing Units (CPUs), Graphics Processing Units (GPUs), Neural Processing Units (NPUs), and Data Processing Units (DPUs).
This heterogeneous architecture is vital for Agentic AI, which often involves a mix of tasks requiring different types of processing power—from the general-purpose computation of CPUs to the specialized parallel processing of GPUs, the optimized inference of NPUs, and the data handling and security functions offloaded to DPUs. By managing these varied resources as a unified system, the UniPoD S80000 aims to eliminate resource silos and ensure that the right compute resource is applied to the right task at the right time. Zhu Shiyin emphasized that the real competition in AI clusters is no longer about the sheer number of GPUs a system contains, but about "how many useful Tokens each GPU can actually produce," highlighting a shift from raw capacity to effective output.
Bridging the Gaps: Advanced Networking for AI Interconnects
As GPU performance continues its rapid ascent, the network invariably becomes the next critical bottleneck. The sheer volume of data exchanged between high-performance GPUs in large AI models can quickly overwhelm traditional network infrastructures, forcing GPUs to wait for data, thereby wasting precious compute cycles. H3C addressed this challenge with a multi-layered interconnect strategy designed for three distinct scenarios: Scale-Up (within a single node), Scale-Out (between multiple nodes), and Scale-Across (between geographically dispersed data centers).

For Scale-Up within a node, H3C unveiled the S9828-128EO, a 102.4T NPO (Network Processing Unit) silicon-photonics intelligent computing switch. Silicon photonics, a cutting-edge technology that integrates optical components with electronic circuits on a single silicon chip, significantly reduces power consumption and latency. H3C claims this switch can cut end-to-end latency by 15%, a crucial improvement for intra-node communication where microseconds matter for tightly coupled GPU operations.
For Scale-Out between nodes, the company showcased its 1.6T intelligent computing switch, the S9828-64FP. This switch incorporates advanced 224G SerDes (Serializer/Deserializer) technologies, which are essential for transmitting high-speed data across backplanes and cables with minimal signal degradation. Its design reduces power consumption by over 20% in high-density environments, supporting the large-scale cluster expansion necessary for truly massive AI models.
Addressing the challenge of Scale-Across data centers, H3C introduced the 800G intelligent computing DCI (Data Center Interconnect) switch, the S12500R-64EP. This switch is engineered to facilitate seamless intra-city, cross-data-center compute scheduling and gradient synchronization. As AI training jobs become too large for a single data center, the ability to distribute workloads and synchronize model parameters across multiple sites without significant performance degradation becomes paramount.
H3C’s networking philosophy extends beyond raw bandwidth, acknowledging that protocols, congestion control, link reliability, and fault diagnosis are equally vital. The company stressed the importance of open protocols and industry standards to prevent the creation of proprietary "technology silos" that could hinder innovation and interoperability within the broader AI ecosystem. This commitment to openness is a strategic move to foster a more collaborative and efficient development environment for AI.
Feeding the Beast: High-Performance AI Storage
GPUs, no matter how powerful, are useless without a continuous, high-speed supply of data. In the context of large-model training, data preparation, model parameter exchange, and checkpointing all involve massive data movements. During inference, the management of KV Cache (Key-Value Cache) further intensifies the pressure on storage and caching resources, as models need to quickly access and store past conversation states.
Recognizing storage as an integral part of the "Token production pipeline," H3C introduced the UniStor X20000 series X20836. This high-performance storage solution boasts impressive specifications, delivering up to 200GB/s of bandwidth and 3 million IOPS (Input/Output Operations Per Second) per node. Crucially, it supports interoperability across various protocols, including block, file, object, and HDFS, ensuring flexibility and compatibility with diverse AI workloads.
H3C reported significant efficiency gains from its storage solutions: its "full-speed engine" can reduce GPU waiting time by up to 30%, while XCache inference acceleration can cut time-to-first-token latency by as much as 90% in inference scenarios. These figures highlight that every moment a GPU spends waiting for data translates directly into increased costs and reduced throughput. The takeaway is clear: AI infrastructure is evolving from a mere compute center into a comprehensive data processing system where optimized data flow is as critical as computational power.

The Orchestration Layer: Software as the AI Conductor
Once hardware components reach a certain scale and sophistication, software optimization becomes the ultimate differentiator. Zhu Shiyin underscored that the entire software stack—from the operating system and communication libraries to resource management and scheduling platforms—must work in close concert with the underlying hardware.
Sophisticated software is essential for:
- Communication Optimization: Minimizing latency and maximizing throughput for data exchange.
- Overlapping Computation with Communication: Intelligently scheduling tasks so that GPUs are performing computations while data is being transferred, reducing idle time.
- Task Scheduling and Dynamic Load Balancing: Ensuring that workloads are distributed efficiently across all available resources, adapting to changing demands in real-time.
- Unified Resource Management: Providing a single pane of glass for managing heterogeneous compute, network, and storage resources, simplifying operations for AI developers and infrastructure teams.
This comprehensive software layer is what transforms disparate hardware components into a cohesive, functioning AI super-system. Without intelligent orchestration, even the most powerful hardware risks operating below its potential. The impact of a software bottleneck is not limited to a single metric; it ultimately affects the system’s ability to generate useful tokens per second and, crucially, the cost associated with producing each token. The competition in AI infrastructure is increasingly moving towards the efficacy of the MLOps (Machine Learning Operations) and orchestration platforms that bind these complex systems together.
Industry Reactions and Broader Implications
H3C’s presentations at the Apsara Conference resonated deeply within the industry, underscoring a growing consensus among technology leaders and analysts. Industry experts widely acknowledge that the era of simply throwing more hardware at AI problems is drawing to a close. The shift towards system-level optimization and efficiency metrics like "cost per Token" is not merely a technical adjustment but a strategic imperative.
For AI developers and enterprises, this means a future where advanced AI capabilities, particularly agentic AI, become more accessible and economically viable. Lower latency and higher throughput translate into faster model training, quicker iteration cycles, and more responsive AI services. This efficiency will be crucial for the widespread adoption of AI agents in various industries, from customer service and data analysis to scientific research and autonomous systems.
For infrastructure providers like H3C, the competitive landscape is shifting. Success will depend less on who can produce the fastest chip and more on who can deliver the most integrated, optimized, and energy-efficient end-to-end AI infrastructure solutions. This necessitates deeper collaboration between hardware and software teams and a commitment to open standards to foster a robust ecosystem.

The broader AI ecosystem stands to benefit from this efficiency drive. As the cost per token decreases, the economic barriers to developing and deploying complex AI models will be lowered, potentially democratizing access to cutting-edge AI technologies. Furthermore, the focus on energy efficiency contributes to the sustainability goals of the technology sector, addressing growing concerns about the environmental footprint of large-scale AI deployments.
Conclusion: The Future is Efficient
From H3C’s exhibition booth to its forum sessions at the Apsara Conference, the message was consistent and clear: the next phase of AI infrastructure competition will be defined by compute utilization. Supernodes provide organized compute, high-speed networking enables efficient communication, high-performance storage ensures timely data delivery, and intelligent software orchestrates these elements into continuously and efficiently usable AI capabilities.
As AI models scale to trillions of parameters and agentic AI becomes pervasive, the critical metrics have expanded beyond raw FLOPs (Floating Point Operations Per Second) to encompass compute utilization, Token throughput, latency, energy consumption, and crucially, the cost per Token. H3C’s pursuit of extreme Token cost efficiency is not about simply aggregating more GPUs; it is about ensuring that every GPU, every network link, and every unit of storage within the system spends less time waiting and more time actively contributing to the generation of intelligence.
As AI transitions into a stage of large-scale production and commercialization, the challenge of compute is no longer solely a chip problem. It has evolved into a sophisticated competition over the efficiency, integration, and intelligent management of the entire AI infrastructure system, marking a pivotal moment in the industry’s trajectory.







