CUDA
CUDAMulti-Instance GPU
An NVIDIA GPU feature that partitions a single physical GPU into up to seven isolated instances, each with dedicated compute, memory and bandwidth resources.
Full Definition
<p>Multi-Instance GPU (MIG) is an NVIDIA feature introduced with the A100 that allows a single physical GPU to be partitioned into up to seven independent GPU instances. Each MIG instance has its own dedicated streaming multiprocessors, HBM memory partition and memory bandwidth — providing performance isolation and QoS guarantees that time-sliced vGPU cannot offer.</p><p>MIG is primarily used in inference and multi-tenant environments where smaller GPU slices are more appropriate than a full GPU per workload. A single H100 80GB GPU can be partitioned into 1g.10gb instances (1/7th of the GPU, 10GB HBM) up to 7g.80gb (the entire GPU). MIG instances can be combined in a rack to create pools of inference capacity. MIG is not supported on all GPU generations — it requires Ampere (A100) or later architectures.</p>
Also Known As
Source Reference
NVIDIA Multi-Instance GPU User Guide; NVIDIA A100/H100 Architecture Whitepapers