Public preview · figures are illustrative fixture data, attributed to their source · data to May 2025 · read-only
Hardware · Accelerator · NVIDIA

NVIDIA A100 80GB

80 GB of HBM2e memory at 2,039 GB/s — enough for models up to ~120B parameters at 4-bit on one card.

Memory
80 GBHBM2e
Memory speed
2,039 GB/s#1 of 14 tracked
Power
400 W
Launch price
16 Nov 2020

What it runs

Largest models without offloading on the A100 80GB server (512 GB DDR4), 8K context.

Systems using it

Measured performance

Throughput on systems containing this device.

Model · quantizationSystemRuntimeContextMeasurementsSource
Qwen2.5 32B Instruct AWQ 4-bitA100 80GB server (512 GB DDR4)vLLMcuda · 0.7.38K
48 tok/s · Generation throughput
72 GB · Peak memory
5,200 tok/s · Prompt throughput
95 ms · Time to first token
Mutinai illustrative fixtures

Community results

Contributions are not open yet, so there is nothing here from members.

Reviews

Contributions are not open yet, so there is nothing here from members.

Specifications

Vendor
NVIDIA
Type
Accelerator
Memory
80 GB
Memory type
HBM2e
Bandwidth
2039 GB/s
GPU-usable share
Backends
cuda (NVIDIA GPUs)
TDP
400 W
Released
16 Nov 2020
Launch price

Sources & history

Sources

No source records linked.