System · desktop
RX 7900 XTX desktop (64 GB DDR5)
Runs models up to ~35B parameters at 4-bit entirely on the GPU. AMD 24 GB GPU desktop on ROCm or Vulkan.
- GPU memory
- 24 GB
- System RAM
- 64 GB
- Runs
- 15 of 19model variants, 8K
- Estimated cost
- ~$2,400
What it runs
Largest models that fit without spilling into system memory, at an 8K contextHow much text the model can consider at once, counted in tokens — roughly ¾ of a word each..
- VariantQwen2.5 32B Instruct32.8B · Q4_K_M via llama.cppTight fit~30 tok/s est.
- VariantQwen2.5-Coder 32B Instruct32.8B · Q4_K_M via llama.cppTight fit~30 tok/s est.
- VariantDeepSeek-R1-Distill-Qwen-32B32.8B · Q4_K_M via llama.cppTight fit~30 tok/s est.
- VariantQwen3 30B-A3B30.5B · Q4_K_M via llama.cppRuns well~245 tok/s est.
- VariantGemma 3 27B IT27.4B · Q4_K_M via llama.cppTight fit~36 tok/s est.
- VariantGemma 2 27B IT27.2B · Q4_K_M via llama.cppRuns well~36 tok/s est.
- VariantMistral Small 24B Instruct 250123.6B · Q6_K via llama.cppRuns well~31 tok/s est.
- VariantQwen2.5 14B Instruct14.8B · Q6_K via llama.cppRuns well~49 tok/s est.
- VariantPhi-414.7B · Q8_0 via llama.cppRuns well~39 tok/s est.
- VariantGemma 2 9B IT9.24B · BF16 via SGLangTight fit~33 tok/s est.
Components
- 1× Hardware AMD Radeon RX 7900 XTX gpu · 24 GB
- 1× Hardware AMD Ryzen 9 7950X cpu
- System RAM bandwidth 96 GB/s
Measured performance
| Model · quantization | Runtime | Context | Measurements | Source |
|---|---|---|---|---|
| Llama 3.1 8B Instruct Q4_K_M (LM Studio Community) | llama.cpprocm · b4600 | 4K | 3,050 tok/s · Prompt processing (512) 94 tok/s · Generation (128) | Mutinai illustrative fixtures |
Community results
Contributions are not open yet, so there is nothing here from members.
Reviews
Contributions are not open yet, so there is nothing here from members.