Hardware · Accelerator · NVIDIA
NVIDIA A100 80GB
80 GB of HBM2e memory at 2,039 GB/s — enough for models up to ~120B parameters at 4-bit on one card.
- Memory
- 80 GBHBM2e
- Memory speed
- 2,039 GB/s#1 of 14 tracked
- Power
- 400 W
- Launch price
- —16 Nov 2020
What it runs
Largest models without offloading on the A100 80GB server (512 GB DDR4), 8K context.
- VariantLlama 3.1 70B Instruct70.6B · Q4_K_M via llama.cppRuns well~30 tok/s est.
- VariantLlama 3.3 70B Instruct70.6B · Q4_K_M via llama.cppRuns well~30 tok/s est.
- VariantMixtral 8x7B Instruct v0.146.7B · Q4_K_M via llama.cppRuns well~157 tok/s est.
- VariantQwen2.5 32B Instruct32.8B · BF16 via SGLangRuns well~20 tok/s est.
- VariantQwen2.5-Coder 32B Instruct32.8B · Q8_0 via llama.cppRuns well~38 tok/s est.
- VariantDeepSeek-R1-Distill-Qwen-32B32.8B · BF16 via SGLangRuns well~20 tok/s est.
- VariantQwen3 30B-A3B30.5B · Q8_0 via llama.cppRuns well~328 tok/s est.
- VariantGemma 3 27B IT27.4B · BF16 via SGLangRuns well~24 tok/s est.
- VariantGemma 2 27B IT27.2B · BF16 via SGLangRuns well~24 tok/s est.
- VariantMistral Small 24B Instruct 250123.6B · BF16 via SGLangRuns well~28 tok/s est.
Systems using it
- System A100 80GB server (512 GB DDR4)Runs models up to ~120B parameters at 4-bit entirely on the GPU. What runs
Measured performance
Throughput on systems containing this device.
| Model · quantization | System | Runtime | Context | Measurements | Source |
|---|---|---|---|---|---|
| Qwen2.5 32B Instruct AWQ 4-bit | A100 80GB server (512 GB DDR4) | vLLMcuda · 0.7.3 | 8K | 48 tok/s · Generation throughput 72 GB · Peak memory 5,200 tok/s · Prompt throughput 95 ms · Time to first token | Mutinai illustrative fixtures |
Community results
Contributions are not open yet, so there is nothing here from members.
Reviews
Contributions are not open yet, so there is nothing here from members.
Specifications
- Vendor
- NVIDIA
- Type
- Accelerator
- Memory
- 80 GB
- Memory type
- HBM2e
- Bandwidth
- 2039 GB/s
- GPU-usable share
- —
- Backends
- cuda (NVIDIA GPUs)
- TDP
- 400 W
- Released
- 16 Nov 2020
- Launch price
- —
Sources & history
Sources
No source records linked.