Public preview · figures are illustrative fixture data, attributed to their source · data to May 2025 · read-only

A100 80GB server (512 GB DDR4)

What can I run? · 8K context
Change hardwareCurrently: A100 80GB server (512 GB DDR4)
Build your own setupChoose GPUs or a chip and how much memory you have
18run well
1with offload
0too large

Speeds marked measured come from reference results and verified community runs on this exact system; others are estimates.

Runs in accelerator memory 18

Fits entirely in GPU or unified memory — the fast path.

ModelRecommended downloadMemoryGeneration speed
Variant Llama 3.1 70B InstructMeta · 70.6BQ4_K_M via llama.cpp40.2 GiB · LM Studio Community
44.8 of 76.0 GiB
~30 tok/s est.
Variant Llama 3.3 70B InstructMeta · 70.6B
2 other options
Q3_K_M · Unsloth AIllama.cppRuns well36.4 GiB~38 tok/s est.
AWQ 4-bit · Unsloth AISGLangRuns well42.3 GiB~32 tok/s est.
Q4_K_M via llama.cpp40.2 GiB · LM Studio Community
44.8 of 76.0 GiB
~30 tok/s est.
Variant Mixtral 8x7B Instruct v0.1Mistral AI · 46.7B (12.9B active)Q4_K_M via llama.cpp26.6 GiB · LM Studio Community
29.1 of 76.0 GiB
~157 tok/s est.
Variant Qwen2.5 32B InstructQwen Team (Alibaba Cloud) · 32.8B
3 other options
AWQ 4-bit · Qwen Team (Alibaba Cloud)SGLangRuns well20.8 GiB~68 tok/s est.
Q4_K_M · Qwen Team (Alibaba Cloud)llama.cppRuns well21.9 GiB~65 tok/s est.
Q5_K_M · LM Studio Communityllama.cppRuns well25.1 GiB~56 tok/s est.
BF16 via SGLang61.0 GiB · Qwen Team (Alibaba Cloud)
66.0 of 76.0 GiB
~20 tok/s est.
Variant Qwen2.5-Coder 32B InstructQwen Team (Alibaba Cloud) · 32.8B
2 other options
Q4_K_M · Qwen Team (Alibaba Cloud)llama.cppRuns well21.9 GiB~65 tok/s est.
BF16 · Qwen Team (Alibaba Cloud)SGLangRuns well66.0 GiB~20 tok/s est.
Q8_0 via llama.cpp32.4 GiB · LM Studio Community
36.2 of 76.0 GiB
~38 tok/s est.
Variant DeepSeek-R1-Distill-Qwen-32BQwen Team (Alibaba Cloud) · 32.8B
1 other option
Q4_K_M · LM Studio Communityllama.cppRuns well21.9 GiB~65 tok/s est.
BF16 via SGLang61.0 GiB · DeepSeek
66.0 of 76.0 GiB
~20 tok/s est.
Variant Qwen3 30B-A3BQwen Team (Alibaba Cloud) · 30.5B (3.3B active)
5 other options
Q4_K_M · Qwen Team (Alibaba Cloud)llama.cppRuns well19.2 GiB~521 tok/s est.
Q4_K_M · Unsloth AIllama.cppRuns well19.3 GiB~519 tok/s est.
FP8 · Qwen Team (Alibaba Cloud)SGLangRuns well31.2 GiB~342 tok/s est.
Q8_0 · Qwen Team (Alibaba Cloud)llama.cppRuns well32.7 GiB~327 tok/s est.
BF16 · Qwen Team (Alibaba Cloud)SGLangRuns well60.4 GiB~186 tok/s est.
Q8_0 via llama.cpp30.2 GiB · LM Studio Community
32.7 of 76.0 GiB
~328 tok/s est.
Variant Gemma 3 27B ITGoogle DeepMind · 27.4B
1 other option
Q4_K_M · LM Studio Communityllama.cppRuns well20.6 GiB~77 tok/s est.
BF16 via SGLang51.1 GiB · Google DeepMind
57.5 of 76.0 GiB
~24 tok/s est.
Variant Gemma 2 27B ITGoogle DeepMind · 27.2B
1 other option
Q4_K_M · LM Studio Communityllama.cppRuns well19.5 GiB~77 tok/s est.
BF16 via SGLang50.7 GiB · Google DeepMind
56.1 of 76.0 GiB
~24 tok/s est.
Variant Mistral Small 24B Instruct 2501Mistral AI · 23.6B
2 other options
Q4_K_M · LM Studio Communityllama.cppRuns well15.7 GiB~89 tok/s est.
Q6_K · LM Studio Communityllama.cppRuns well20.5 GiB~67 tok/s est.
BF16 via SGLang43.9 GiB · Mistral AI
47.4 of 76.0 GiB
~28 tok/s est.
Variant Qwen2.5 14B InstructQwen Team (Alibaba Cloud) · 14.8B
3 other options
AWQ 4-bit · Qwen Team (Alibaba Cloud)SGLangRuns well10.2 GiB~147 tok/s est.
Q4_K_M · Qwen Team (Alibaba Cloud)llama.cppRuns well10.7 GiB~139 tok/s est.
Q6_K · LM Studio Communityllama.cppRuns well13.7 GiB~105 tok/s est.
BF16 via SGLang27.5 GiB · Qwen Team (Alibaba Cloud)
30.6 of 76.0 GiB
~44 tok/s est.
Variant Phi-4Microsoft · 14.7B
2 other options
Q4_K_M · LM Studio Communityllama.cppRuns well10.7 GiB~140 tok/s est.
BF16 · MicrosoftSGLangRuns well30.5 GiB~44 tok/s est.
Q8_0 via llama.cpp14.5 GiB · LM Studio Community
17.1 of 76.0 GiB
~82 tok/s est.
Variant Gemma 2 9B ITGoogle DeepMind · 9.24B
1 other option
Q4_K_M · LM Studio Communityllama.cppRuns well8.6 GiB~214 tok/s est.
BF16 via SGLang17.2 GiB · Google DeepMind
21.0 of 76.0 GiB
~70 tok/s est.
Variant Hermes 3 Llama 3.1 8BMeta · 8.03B
2 other options
Q4_K_M · NousResearchllama.cppRuns well6.3 GiB~243 tok/s est.
Q6_K · NousResearchllama.cppRuns well7.9 GiB~186 tok/s est.
Q8_0 via llama.cpp8.0 GiB · NousResearch
9.8 of 76.0 GiB
~146 tok/s est.
Variant Llama 3.1 8B InstructMeta · 8.03B
3 other options
Q4_K_M · LM Studio Communityllama.cppRuns well6.3 GiB~243 tok/s est.
FP8 · RedHatAISGLangRuns well10.3 GiB~138 tok/s est.
BF16 · MetaSGLangRuns well17.1 GiB~80 tok/s est.
Q8_0 via llama.cpp7.9 GiB · LM Studio Community
9.8 of 76.0 GiB
~146 tok/s est.
Variant Qwen2.5 7B InstructQwen Team (Alibaba Cloud) · 7.62B
3 other options
AWQ 4-bit · Qwen Team (Alibaba Cloud)SGLangRuns well5.2 GiB~270 tok/s est.
Q4_K_M · Qwen Team (Alibaba Cloud)llama.cppRuns well5.5 GiB~255 tok/s est.
BF16 · Qwen Team (Alibaba Cloud)SGLangRuns well15.7 GiB~84 tok/s est.
Q8_0 via llama.cpp7.5 GiB · Qwen Team (Alibaba Cloud)
8.8 of 76.0 GiB
~154 tok/s est.
Variant DeepSeek-R1-Distill-Qwen-7BQwen Team (Alibaba Cloud) · 7.62B
1 other option
Q4_K_M · LM Studio Communityllama.cppRuns well5.2 GiB~255 tok/s est.
BF16 via SGLang14.2 GiB · DeepSeek
15.5 of 76.0 GiB
~84 tok/s est.
Variant Mistral 7B Instruct v0.3Mistral AI · 7.25B
2 other options
Q4_K_M · LM Studio Communityllama.cppRuns well5.8 GiB~267 tok/s est.
BF16 · Mistral AISGLangRuns well15.5 GiB~88 tok/s est.
Q8_0 via llama.cpp7.2 GiB · LM Studio Community
9.0 of 76.0 GiB
~161 tok/s est.

Runs with partial offload to system RAM 1

Too big for accelerator memory alone; part of it runs from system RAM, which is much slower.

ModelRecommended downloadMemoryGeneration speed
Variant DeepSeek-R1DeepSeek · 671B (37B active)
1 other option
Q3_K_M · Unsloth AIllama.cppWith offload318.7 GiB~9 tok/s est.
IQ2_XXS via llama.cpp161 GiB · Unsloth AI
168.4 of 76.0 GiB · 92.4 in RAM
~21 tok/s est.
How this is calculatedMemory, fit and speed estimates

Memory = download size + fp16 KV cache for 8K tokens + runtime overhead. Usable memory is 95% of dedicated VRAM, a device-specific share of unified memory (75% by default), and 80% of system RAM for runtimes that can offload. Runtimes must load the file format and support a backend present on the hardware.

Estimated speed is bounded by memory bandwidth ÷ bytes read per token (active parameters for mixture-of-experts), shown in italics with est. and a ±35% range. Measured speeds are medians of reference results and verified public community runs on the same system, download and runtime. For each model we recommend the highest-precision download that fits, preferring not to offload.