How XPUs Meet a World-Class AI Factory
To generate intelligence at scale, AI factories run continuously, and their economics are defined by delivered output: tokens per second, tokens per watt, cost per token, utilization and uptime. That requires AI infrastructure designed and built as a full factory, not a collection of individual accelerators. Hyperscalers and AI-native companies building custom XPUs must consider not just XPU design, but the design and development of the entire AI platform, including scale-up and scale-out networking, rack-scale architecture, production factory software and a robust supplier ecosystem. At AI factory scale, this path is complex and costly, and represents a fundamental obstacle to getting XPUs to market quickly. Breaking the constraint means combining custom XPUs with proven, mature infrastructure — allowing builders to focus innovation where it matters most while harnessing established technology for the rest. NVLink Fusion delivers on that need, connecting XPUs to NVIDIA’s world-leading AI infrastructure to increase performance, accelerate time to market and mitigate risk for semi-custom AI factories.
Unlock XPU Performance With Fast Scale-Up For modern workloads such as running trillion-parameter models, mixture-of-experts architectures and agentic AI, if the scale-up fabric cannot keep up, utilization drops and cost per token rises. A scale-up networking solution must excel on three dimensions: Delivered performance: End-to-end network performance, in-network compute and mature software integration. Factory resiliency: Uptime, continuous health monitoring and telemetry, and component-level serviceability while the factory keeps running. Platform maturity: Reduced operational risk by using a mature technology stack with a demonstrated track record of large-scale deployments and realized return on investment. As an example, NVLink Fusion brings XPUs into the NVIDIA NVLink scale-up domain. Sixth-generation NVLink provides leading high-bandwidth, low-latency networking across a 72-XPU domain. The end-to-end latency for XPU-to-XPU transfers is 3x lower than alternative solutions based on off-the-shelf Ethernet, and the packet rate is 10x higher. For end-to-end performance, NVIDIA GB300 NVL72 systems help deliver significantly higher throughput and better interactivity compared with configurations that don’t use NVL72, and future NVLink roadmap configurations include domains of up to 1,152 accelerators and co-packaged optics. The 72-GPU NVLink scale-up domain enables GB300 NVL72 to deliver higher per-GPU throughput and interactivity compared with NVIDIA B300. Results from NVIDIA’s AI Inference Performance Benchmarks page. NVLink Fusion also includes NVIDIA NVLink-C2C for connecting XPUs to NVIDIA Vera CPUs or other ecosystem CPUs, delivering up to 6x the energy efficiency of a PCIe interface — helping remove barriers between control and compute for agentic systems. A Proven Stack and Ecosystem for Development and Deployment Teams developing custom XPUs often underestimate the effort and complexity of turning XPU innovation into data center deployment. This includes: Integrating high-speed CPU and scale-up interfaces Sourcing and validating a scale-up network solution Designing compute and switch trays Designing and validating a rack architecture, including cooling and power Integrating security and storage Managing a complex supplier ecosystem The ideal platform provides all of this, allowing teams to focus on targeted innovation while using proven solutions for the rest. NVLink Fusion is supported by an ecosystem designed for rapid development, integration and deployment, spanning ASIC design, CPU, and IP and optical interconnect partners. “NVLink Fusion gives customers the ability to choose the CPU architecture, the performance level, the software capabilities that best meet their needs for the workloads that they care about,” said Tim Wilson, vice president and general manager of data center silicon engineering at Intel. NVLink Fusion adopters can also use the NVIDIA MGX rack-scale architecture and the same supply chain used for MGX-based systems such as NVIDIA Vera Rubin NVL72.