RTX 5090 workstation (96 GB DDR5)
What can I run? · 8K contextChange hardware
Desktops
Servers
Build your own setup
Speeds marked measured come from reference results and verified community runs on this exact system; others are estimates.
Runs in accelerator memory 16
Fits entirely in GPU or unified memory — the fast path.
| Model | Recommended download | Memory | Generation speed | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
Variant Qwen2.5 32B InstructQwen Team (Alibaba Cloud) · 32.8B2 other options
| Q5_K_M via llama.cpp21.7 GiB · LM Studio Community | 25.1 of 30.4 GiB | ~49 tok/s est. | |||||||||||||||
Variant Qwen2.5-Coder 32B InstructQwen Team (Alibaba Cloud) · 32.8B1 other option
| Q4_K_M via llama.cpp18.7 GiB · Qwen Team (Alibaba Cloud) | 21.9 of 30.4 GiB | ~57 tok/s est. | |||||||||||||||
| Variant DeepSeek-R1-Distill-Qwen-32BQwen Team (Alibaba Cloud) · 32.8B | Q4_K_M via llama.cpp18.7 GiB · LM Studio Community | 21.9 of 30.4 GiB | ~57 tok/s est. | |||||||||||||||
Variant Qwen3 30B-A3BQwen Team (Alibaba Cloud) · 30.5B (3.3B active)3 other options
| Q4_K_M via llama.cpp17.3 GiB · Qwen Team (Alibaba Cloud) | 19.2 of 30.4 GiB | ~458 tok/s est. | |||||||||||||||
| Variant Gemma 3 27B ITGoogle DeepMind · 27.4B | Q4_K_M via llama.cpp15.6 GiB · LM Studio Community | 20.6 of 30.4 GiB | ~67 tok/s est. | |||||||||||||||
| Variant Gemma 2 27B ITGoogle DeepMind · 27.2B | Q4_K_M via llama.cpp15.5 GiB · LM Studio Community | 19.5 of 30.4 GiB | ~68 tok/s est. | |||||||||||||||
Variant Mistral Small 24B Instruct 2501Mistral AI · 23.6B1 other option
| Q6_K via llama.cpp18.0 GiB · LM Studio Community | 20.5 of 30.4 GiB | ~59 tok/s est. | |||||||||||||||
Variant Qwen2.5 14B InstructQwen Team (Alibaba Cloud) · 14.8B2 other options
| Q6_K via llama.cpp11.3 GiB · LM Studio Community | 13.7 of 30.4 GiB | ~92 tok/s est. | |||||||||||||||
Variant Phi-4Microsoft · 14.7B1 other option
| Q8_0 via llama.cpp14.5 GiB · LM Studio Community | 17.1 of 30.4 GiB | ~72 tok/s est. | |||||||||||||||
Variant Gemma 2 9B ITGoogle DeepMind · 9.24B1 other option
| BF16 via SGLang17.2 GiB · Google DeepMind | 21.0 of 30.4 GiB | ~61 tok/s est. | |||||||||||||||
Variant Hermes 3 Llama 3.1 8BMeta · 8.03B2 other options
| Q8_0 via llama.cpp8.0 GiB · NousResearch | 9.8 of 30.4 GiB | ~128 tok/s est. | |||||||||||||||
Variant Llama 3.1 8B InstructMeta · 8.03B3 other options
| Q8_0 via llama.cpp7.9 GiB · LM Studio Community | 9.8 of 30.4 GiB | ~128 tok/s est. | |||||||||||||||
Variant Qwen2.5 7B InstructQwen Team (Alibaba Cloud) · 7.62B3 other options
| Q8_0 via llama.cpp7.5 GiB · Qwen Team (Alibaba Cloud) | 8.8 of 30.4 GiB | ~135 tok/s est. | |||||||||||||||
Variant DeepSeek-R1-Distill-Qwen-7BQwen Team (Alibaba Cloud) · 7.62B1 other option
| BF16 via SGLang14.2 GiB · DeepSeek | 15.5 of 30.4 GiB | ~74 tok/s est. | |||||||||||||||
Variant Mistral 7B Instruct v0.3Mistral AI · 7.25B2 other options
| Q8_0 via llama.cpp7.2 GiB · LM Studio Community | 9.0 of 30.4 GiB | ~141 tok/s est. | |||||||||||||||
| Variant Mixtral 8x7B Instruct v0.1Mistral AI · 46.7B (12.9B active) | Q4_K_M via llama.cpp26.6 GiB · LM Studio CommunityTight fit | 29.1 of 30.4 GiB | ~138 tok/s est. |
Runs with partial offload to system RAM 2
Too big for accelerator memory alone; part of it runs from system RAM, which is much slower.
| Model | Recommended download | Memory | Generation speed | |||||
|---|---|---|---|---|---|---|---|---|
| Variant Llama 3.1 70B InstructMeta · 70.6B | Q4_K_M via llama.cpp40.2 GiB · LM Studio Community | 44.8 of 30.4 GiB · 14.4 in RAM | ~4 tok/s est. | |||||
Variant Llama 3.3 70B InstructMeta · 70.6B1 other option
| Q3_K_M via llama.cpp32.1 GiB · Unsloth AI | 36.4 of 30.4 GiB · 6.0 in RAM | ~8 tok/s est. |
Too large for this hardware
- DeepSeek-R1 671B · 3 downloads checked
How this is calculated
Memory = download size + fp16 KV cache for 8K tokens + runtime overhead. Usable memory is 95% of dedicated VRAM, a device-specific share of unified memory (75% by default), and 80% of system RAM for runtimes that can offload. Runtimes must load the file format and support a backend present on the hardware.
Estimated speed is bounded by memory bandwidth ÷ bytes read per token (active parameters for mixture-of-experts), shown in italics with est. and a ±35% range. Measured speeds are medians of reference results and verified public community runs on the same system, download and runtime. For each model we recommend the highest-precision download that fits, preferring not to offload.