Skip to main content
Back to Glossary
ai-infrastructure

Graphics Processing Unit

GPU

A massively parallel processor originally designed for graphics rendering, now the primary compute engine for AI training and inference at scale.

Unit: FLOPS / TFLOPS

Full Definition

A Graphics Processing Unit (GPU) contains thousands of smaller, specialised cores designed to process many operations simultaneously. Where a CPU optimises for sequential task performance, a GPU excels at parallel workloads — making it the default accelerator for deep learning model training and inference.

Modern datacentre GPUs such as the NVIDIA H100 and B200 are purpose-built for AI workloads, incorporating Tensor Cores for matrix multiplication, high-bandwidth memory (HBM), and dedicated interconnects (NVLink). A single H100 SXM5 delivers 3.9 petaFLOPS of FP8 performance; a GB200 NVL72 rack delivers 1.4 exaFLOPS. GPU power density ranges from 300 W to 1,000 W per card, requiring purpose-designed power distribution and liquid cooling infrastructure.

Also Known As

graphics cardacceleratorAI accelerator

Source Reference

NVIDIA H100 Tensor Core GPU Architecture Whitepaper; JEDEC HBM3 specification

Related Terms