Skip to main content
Back to Glossary
ai-infrastructure

GB200 NVL72

NVIDIA's Blackwell-generation rack-scale AI system — 36 Grace CPUs and 72 Blackwell GPUs interconnected at 130 TB/s via NVLink — delivering 1.4 exaFLOPS of AI performance at roughly 120 kW per rack.

Unit: exaFLOPS / ~120 kW per rack

Full Definition

The NVIDIA GB200 NVL72 is a rack-scale computing system introduced as part of the Blackwell architecture in 2024. Each rack contains 18 GB200 Superchip trays, with each tray holding one Grace CPU and two Blackwell GPUs. The 72 GPUs are fully interconnected via fifth-generation NVLink and NVSwitch chips at 130 TB/s all-to-all bandwidth, enabling the entire rack to operate as a unified computing entity for large model training and inference.

The NVL72 requires approximately 120 kW of power per rack and demands direct liquid cooling — air cooling is insufficient. Memory across the rack totals over 13 TB of HBM3e. For inference, the NVL72 delivers 30× the performance of the previous H100 generation on large language model tasks. These requirements place the GB200 NVL72 at the centre of AI factory infrastructure planning, particularly for power delivery, liquid cooling design, and structural floor loading.

Also Known As

Blackwell rackNVL72GB200

Source Reference

NVIDIA GB200 NVL72 Architecture Whitepaper; NVIDIA GTC 2024 Technical Sessions

Related Terms