SUMMARYNVIDIA said its Vera Rubin NVL72 platform is in full production and is being extended with Groq 3 LPX for faster token generation in long-context agentic AI workloads. The company also introduced Spectrum-X Multiplane for larger AI networks, Scale-In infrastructure powered by BlueField-4 and DOCA, and NVLink Fusion for connecting custom chips to its AI stack. Nebius, CoreWeave and SpaceXAI are among the partners adopting the new infrastructure.

With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agentsblogs.nvidia.com

The next era of AI inference won’t be defined by a single breakthrough chip, network or system. It’ll be defined by how every layer of the AI factory works together. That’s why NVIDIA is extending Vera Rubin NVL72 with fast token generation for agentic systems.

Announced today, the NVIDIA Vera Rubin rack-scale system NVIDIA Groq 3 LPX is in full production. In an Artificial Analysis benchmark running Gemma 4 31B, an open source agentic model, it delivered 3,400 output tokens per second for 100,000-token long-context use cases critical to agentic systems, 4x faster than the nearest alternative platform.

Industry partners worldwide are adopting Vera Rubin platform solutions. SpaceXAI announced that NVIDIA Vera CPUs will power its next generation of agentic AI. CoreWeave has deployed into production Spectrum-X Multiplane, which connects NVIDIA Vera Rubin racks using multiple parallel switches to provide high-bandwidth, flat and lossless AI networks. Nebius is the first AI cloud to adopt NVIDIA Groq 3 LPX.

As AI shifts from training to reasoning and agentic, inference has become the new frontier. Agentic AI systems are generating more tokens, processing dramatically larger context windows and increasingly collaborating with other AI systems to solve complex problems.

These workloads demand a new class of infrastructure optimized not just for performance but for throughput, responsiveness and economics at unprecedented scale.

At the Hot Chips conference this week in Palo Alto, California, NVIDIA is showcasing how extreme codesign is reshaping the AI factory from end to end. By architecting compute, networking and inference acceleration as a unified system, NVIDIA is helping customers build infrastructure purpose-built for the emerging demands of long-context inference and multi-agent systems.

Extreme Codesign Optimizes for Performance

Extreme codesign is the guiding principle behind NVIDIA platforms. Vera Rubin is engineered to accelerate inference as agents reason over increasingly long sequences.

NVIDIA Spectrum-X Ethernet moves those massive data flows efficiently across AI factories, and NVIDIA Groq 3 LPX is built to generate tokens at ultrafast speeds. Together, they show how NVIDIA is optimizing every stage of the AI pipeline, from context and communication to generation, as part of a single, integrated AI factory architecture.

NVIDIA Groq 3 LPX brings a new low-latency inference architecture designed to work alongside Vera Rubin NLV72, the most versatile AI factory platform, helping enterprises and cloud providers deliver the low latency, extreme throughput and scalable economics required for agentic applications.

Breakthrough performance comes not from optimizing individual components in isolation, but from codesigning every layer of the stack. From networking and context processing to large-scale inference, NVIDIA’s full-stack platform turns AI factories into integrated engines for intelligence, built to turn ever-growing volumes of tokens into revenue.

Tuesday, Aug. 24, 8:00 a.m. PT 🔗

NVIDIA Partners Adopt Vera Rubin for Lowest Token Costs

Nebius, a leading AI cloud, is first to adopt NVIDIA Groq 3 LPX, giving developers access to leading token generation speeds for highly responsive agentic AI applications.

Adding NVIDIA Groq 3 LPX to NVIDIA Vera Rubin NVL72 in Nebius Token Factory will boost inference performance so developers can build highly interactive agents, coding systems and other real-time AI experiences at scale.

Connecting NVIDIA Vera Rubin racks, CoreWeave is deploying Spectrum-X Multiplane in production, unlocking advances for its AI cloud infrastructure.

Tuesday, Aug. 24, 8:00 a.m. PT 🔗

SpaceXAI Adopts NVIDIA Vera CPUs for Agentic AI

SpaceXAI plans to build and scale its future AI architecture around NVIDIA Vera Rubin, from data centers on Earth to orbital satellites. The company plans to deploy NVIDIA Vera CPUs to accelerate the CPU-intensive work behind agentic AI, including orchestration, tool use, code execution, data processing and simulation.

The SpaceXAI partnership extends NVIDIA’s full-stack AI platform to SpaceXAI, bringing together Vera CPUs, NVIDIA accelerated computing, networking and software to advance AI at unprecedented scale.

Designed for the agentic era, Vera Rubin provides leading per-core performance, exceptional memory bandwidth and predictable performance under load, helping agents complete tasks faster and keeping valuable GPU infrastructure fully utilized.

Tuesday, Aug. 24, 8:00 a.m. PT 🔗

NVIDIA Groq 3 LPX: The Interactive AI Inference Accelerator

Codesigned with the Vera Rubin NVL72 platform, NVIDIA Groq 3 LPX is helping AI factories deliver tokens at the lowest latency for agentic workloads.

Agentic AI is creating a new performance challenge: decode latency. As AI agents reason, use tools and interact with other systems, they generate responses one token at a time, causing even tiny delays to multiply across complex chains of work. To keep agents operating at the pace users expect, NVIDIA Groq 3 LPX extends the Vera Rubin NVL72 platform with specialized acceleration for token generation.

NVIDIA Rubin GPUs handle large-scale context processing while LPX accelerates latency-sensitive decode workloads. The result is faster, more predictable token generation that helps AI factories deliver responsive reasoning, smoother agent interactions and greater infrastructure efficiency.

Together, Rubin GPUs and LPUs are designed to eliminate the traditional tradeoff between speed and throughput, helping AI providers deliver responsive, large-scale inference for the next generation of agentic AI applications.

Building the Token Factory

As the industry shifts from model training to serving intelligence at scale, infrastructure must evolve into what NVIDIA describes as a “token factory” capable of delivering performance, throughput, intelligence integrity and economic efficiency simultaneously. Agentic AI systems increasingly communicate with other AI systems, access multiple data sources and maintain large amounts of context, creating unprecedented demand for fast inference.

NVIDIA Groq 3 LPX was designed for exactly these workloads. As an extension of the Vera Rubin NVL72, it enables ultrafast responsiveness even across massive context windows while helping service providers maximize throughput and infrastructure utilization.

Extreme Codesign for Inference

Unlike standalone accelerators, NVIDIA Groq 3 LPX combines the strengths of GPUs and LPUs through extreme codesign. Rubin GPUs and LPUs jointly compute every layer of an AI model, enabling new levels of inference performance for agentic workloads.

At scale, fleets of LPUs operate as a giant processor optimized for deterministic inference. A rack-scale NVIDIA Groq 3 LPX deployment can include 256 LP30 accelerators connected through direct chip-to-chip links, creating a highly efficient inference engine built for modern AI factories.

Designed for the Agentic AI Era

As reasoning models grow and agentic workflows generate ever more tokens, the infrastructure required to serve them must evolve. NVIDIA Groq 3 LPX extends the Vera Rubin NVL72 platform with a purpose-built inference architecture designed to maximize responsiveness, throughput and efficiency, helping power the next generation of AI factories.

And this is only the beginning, more optimizations, more models, more performance when paired with Vera Rubin NVL72 — new levels of throughput and interactivity are coming. Stay tuned.

Tuesday, Aug. 24, 8:00 a.m. PT 🔗

NVIDIA Spectrum-X Multiplane Enables Massive AI Factory Scale on a Flatter, More Resilient Network​

As AI factories grow massive, the network has become a critical engine of performance. At Hot Chips, NVIDIA is spotlighting Spectrum-X Multiplane — the latest in the hardware-accelerated Spectrum-X Ethernet architecture that lets Ethernet scale to unprecedented size while avoiding the latency, jitter and cost of adding another network tier.

NVIDIA Spectrum-X Ethernet is designed as an end-to-end, AI-optimized Ethernet platform, combining NVIDIA Spectrum-X Ethernet switches, SuperNICs and software to improve the performance and efficiency of Ethernet-based AI infrastructure for AI factories and clouds. The platform is designed to deliver 1.6x better AI networking performance compared with off-the-shelf Ethernet, while providing consistent, predictable performance in multi-tenant environments.

Multiplane Unlocks Scale Without the Tradeoffs of a New Tier

Scaling an AI factory beyond today’s largest clusters traditionally means adding a third network tier, which adds latency, slows things down unpredictably and drives up the cost of cabling, optics and power. Spectrum-X Multiplane takes a simpler approach: It splits each server’s network connection into several independent paths, or “planes,” each running its own lightweight two-tier network. The result is a flat, simple network that scales to 512,000 GPUs, without the added cost and complexity of a third tier.

This all happens automatically. A dedicated hardware engine inside the NVIDIA ConnectX SuperNIC manages traffic across the planes and instantly reroutes around any failure, so applications and software simply see one fast, reliable connection. In an eight-plane topology, if one plane fails, the network still maintains about 90% of its total bandwidth, with hardware recovery that’s 11x faster than software-based multiplane load balancing. This translates to 1.6x higher AI factory output.

Built Through Extreme Codesign

That reliability comes from extreme codesign of Vera Rubin NVL72, spanning switch silicon, SuperNICs and software. Spectrum-X SN6000 series switches, based on the 102.4Tb/s Spectrum-6 Ethernet ASIC and ConnectX-9 SuperNICs, supporting up to 1,600Gb/s per GPU, are purpose-built for Vera Rubin NVL72 AI factories. Spectrum-XGS Ethernet extends that same codesign across data centers, letting multiple facilities function as a single AI super-factory and accelerating multi-site NCCL collectives by 1.9x.

Tuesday, Aug. 24, 8:00 a.m. PT 🔗

NVIDIA Introduces Scale-In Infrastructure for Agentic AI Factories, Powered by BlueField-4, DOCA

NVIDIA is introducing NVIDIA Scale-In, the fifth pillar of NVIDIA AI networking and a new class of accelerated network infrastructure for agentic AI factories. Scale-In extends purpose-built acceleration to the infrastructure services that secure, manage and operate the AI factory.

Powered by the NVIDIA BlueField-4 processor and NVIDIA DOCA software platform and connected over NVIDIA Spectrum-X Ethernet, NVIDIA Scale-In transforms the traditional north-south access network into a unified, accelerated infrastructure domain.

Cloud computing brought software-defined networking, composability and elasticity to the data center, enabling users, applications, data and services to scale dynamically.

Agentic AI represents the next platform shift. AI factories bring together massive accelerated compute with growing numbers of users, applications and autonomous agents, all continuously interacting with data, storage and services. This transforms the demands on the infrastructure that brings AI to life. Networking, storage, cybersecurity and operations must now be accelerated alongside AI compute, combining software-defined flexibility with purpose-built hardware acceleration and full-stack codesign.

NVIDIA Scale-In delivers multi-tenant networking, high-performance storage access, in-silicon security, elastic provisioning and real-time observability, while keeping infrastructure processing independent of host compute resources. By accelerating and codesigning these services as part of the AI factory, Scale-In helps security, data access and operations scale alongside AI compute. The result is secure, efficient and manageable shared infrastructure for deploying and operating agentic AI at massive scale.

Tuesday, Aug. 24, 8:00 a.m. PT 🔗

NVIDIA NVLink Fusion Connects XPUs to NVIDIA’s Leading AI Platform

NVIDIA NVLink Fusion brings custom silicon into NVIDIA’s world-leading AI infrastructure platform, enabling hyperscalers and AI-native companies to build semi-custom AI factories with greater performance, flexibility and speed.

As AI models grow in size and complexity, raw compute alone is not enough. AI factories require high-bandwidth, low-latency scale-up networking, proven rack-scale architectures and a full ecosystem spanning power, cooling, management software and supply chain. NVLink Fusion addresses these challenges by connecting custom XPUs and CPUs to NVIDIA’s scale-up and scale-out technology stack.

The platform includes sixth-generation NVIDIA NVLink and NVLink Switch purpose-built scale-up networking, as well as NVLink-C2C for energy-efficient connectivity between XPUs and CPUs. Through the NVIDIA MGX ecosystem, adopters can also use production-proven rack designs, components, manufacturing partner solutions and open, extensible software for distributed computing, disaggregated workloads and cluster management.

By standardizing GPU- and XPU-based systems on a unified architecture, NVLink Fusion helps decouple data center buildout from silicon readiness. Operators can share rack footprints, networking, cooling, power delivery and management systems, then adjust the mix of GPUs and XPUs as supply and workload requirements evolve.

NVLink Fusion extends the NVIDIA AI platform’s vertically integrated, horizontally open approach to custom silicon. It gives partners the freedom to innovate where they differentiate while drawing on NVIDIA technologies across compute, networking, infrastructure and software — creating a single, flexible AI factory that no one company could build alone.