Public preview · figures are illustrative fixture data, attributed to their source · data to May 2025 · read-only
System · server

A100 80GB server (512 GB DDR4)

Runs models up to ~120B parameters at 4-bit entirely on the GPU. Single datacenter GPU with large host memory.

GPU memory
80 GB
System RAM
512 GB
Runs
18 of 19model variants, 8K
Estimated cost

What it runs

Largest models that fit without spilling into system memory, at an 8K contextHow much text the model can consider at once, counted in tokens — roughly ¾ of a word each..

Components

Measured performance

Model · quantizationRuntimeContextMeasurementsSource
Qwen2.5 32B Instruct AWQ 4-bitvLLMcuda · 0.7.38K
48 tok/s · Generation throughput
72 GB · Peak memory
5,200 tok/s · Prompt throughput
95 ms · Time to first token
Mutinai illustrative fixtures

Community results

Contributions are not open yet, so there is nothing here from members.

Reviews

Contributions are not open yet, so there is nothing here from members.