Floating Point Operations Per Second
FLOPSGraphics Processing Unit
A massively parallel processor originally designed for graphics rendering, now the primary compute engine for AI training and inference at scale.
Full Definition
A Graphics Processing Unit (GPU) contains thousands of smaller, specialised cores designed to process many operations simultaneously. Where a CPU optimises for sequential task performance, a GPU excels at parallel workloads — making it the default accelerator for deep learning model training and inference.
Modern datacentre GPUs such as the NVIDIA H100 and B200 are purpose-built for AI workloads, incorporating Tensor Cores for matrix multiplication, high-bandwidth memory (HBM), and dedicated interconnects (NVLink). A single H100 SXM5 delivers 3.9 petaFLOPS of FP8 performance; a GB200 NVL72 rack delivers 1.4 exaFLOPS. GPU power density ranges from 300 W to 1,000 W per card, requiring purpose-designed power distribution and liquid cooling infrastructure.
Also Known As
Source Reference
NVIDIA H100 Tensor Core GPU Architecture Whitepaper; JEDEC HBM3 specification