System · mini pc
Ryzen AI Max+ 395 mini PC (128 GB)
Runs models up to ~160B parameters at 4-bit in unified memory. Strix Halo system with 128 GB unified memory.
- Unified memory
- 128 GB
- System RAM
- Shared
- Runs
- 18 of 19model variants, 8K
- Estimated cost
- ~$2,000
What it runs
Largest models that fit without spilling into system memory, at an 8K contextHow much text the model can consider at once, counted in tokens — roughly ¾ of a word each..
- VariantLlama 3.1 70B Instruct70.6B · Q4_K_M via llama.cppRuns well~4 tok/s est.
- VariantLlama 3.3 70B Instruct70.6B · Q4_K_M via llama.cppRuns well~4 tok/s est.
- VariantMixtral 8x7B Instruct v0.146.7B · BF16 via SGLangTight fit~6 tok/s est.
- VariantQwen2.5 32B Instruct32.8B · BF16 via SGLangRuns well~3 tok/s est.
- VariantQwen2.5-Coder 32B Instruct32.8B · Q8_0 via llama.cppRuns well~5 tok/s est.
- VariantDeepSeek-R1-Distill-Qwen-32B32.8B · BF16 via SGLangRuns well~3 tok/s est.
- VariantQwen3 30B-A3B30.5B · Q8_0 via llama.cppRuns well~41 tok/s est.
- VariantGemma 3 27B IT27.4B · BF16 via SGLangRuns well~3 tok/s est.
- VariantGemma 2 27B IT27.2B · BF16 via SGLangRuns well~3 tok/s est.
- VariantMistral Small 24B Instruct 250123.6B · BF16 via SGLangRuns well~4 tok/s est.
Components
Measured performance
| Model · quantization | Runtime | Context | Measurements | Source |
|---|---|---|---|---|
| Qwen3 30B-A3B Q4_K_M (Unsloth AI) | llama.cppvulkan · b5300 | 4K | 420 tok/s · Prompt processing (512) 51 tok/s · Generation (128) | Mutinai illustrative fixtures |
Community results
Contributions are not open yet, so there is nothing here from members.
Reviews
Contributions are not open yet, so there is nothing here from members.