Ryzen AI Max+ 395 mini PC (128 GB)
What can I run? · 8K contextChange hardware
Desktops
Servers
Build your own setup
Speeds marked measured come from reference results and verified community runs on this exact system; others are estimates.
Runs in accelerator memory 18
Fits entirely in GPU or unified memory — the fast path.
| Model | Recommended download | Memory | Generation speed | |||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Variant Llama 3.1 70B InstructMeta · 70.6B | Q4_K_M via llama.cpp40.2 GiB · LM Studio Community | 44.8 of 96.0 GiB | ~4 tok/s est. | |||||||||||||||||||||||||
Variant Llama 3.3 70B InstructMeta · 70.6B2 other options
| Q4_K_M via llama.cpp40.2 GiB · LM Studio Community | 44.8 of 96.0 GiB | ~4 tok/s est. | |||||||||||||||||||||||||
Variant Qwen2.5 32B InstructQwen Team (Alibaba Cloud) · 32.8B3 other options
| BF16 via SGLang61.0 GiB · Qwen Team (Alibaba Cloud) | 66.0 of 96.0 GiB | ~3 tok/s est. | |||||||||||||||||||||||||
Variant Qwen2.5-Coder 32B InstructQwen Team (Alibaba Cloud) · 32.8B2 other options
| Q8_0 via llama.cpp32.4 GiB · LM Studio Community | 36.2 of 96.0 GiB | ~5 tok/s est. | |||||||||||||||||||||||||
Variant DeepSeek-R1-Distill-Qwen-32BQwen Team (Alibaba Cloud) · 32.8B1 other option
| BF16 via SGLang61.0 GiB · DeepSeek | 66.0 of 96.0 GiB | ~3 tok/s est. | |||||||||||||||||||||||||
Variant Qwen3 30B-A3BQwen Team (Alibaba Cloud) · 30.5B (3.3B active)5 other options
| Q8_0 via llama.cpp30.2 GiB · LM Studio Community | 32.7 of 96.0 GiB | ~41 tok/s est. | |||||||||||||||||||||||||
Variant Gemma 3 27B ITGoogle DeepMind · 27.4B1 other option
| BF16 via SGLang51.1 GiB · Google DeepMind | 57.5 of 96.0 GiB | ~3 tok/s est. | |||||||||||||||||||||||||
Variant Gemma 2 27B ITGoogle DeepMind · 27.2B1 other option
| BF16 via SGLang50.7 GiB · Google DeepMind | 56.1 of 96.0 GiB | ~3 tok/s est. | |||||||||||||||||||||||||
Variant Mistral Small 24B Instruct 2501Mistral AI · 23.6B2 other options
| BF16 via SGLang43.9 GiB · Mistral AI | 47.4 of 96.0 GiB | ~4 tok/s est. | |||||||||||||||||||||||||
Variant Qwen2.5 14B InstructQwen Team (Alibaba Cloud) · 14.8B3 other options
| BF16 via SGLang27.5 GiB · Qwen Team (Alibaba Cloud) | 30.6 of 96.0 GiB | ~6 tok/s est. | |||||||||||||||||||||||||
Variant Phi-4Microsoft · 14.7B2 other options
| Q8_0 via llama.cpp14.5 GiB · LM Studio Community | 17.1 of 96.0 GiB | ~10 tok/s est. | |||||||||||||||||||||||||
Variant Gemma 2 9B ITGoogle DeepMind · 9.24B1 other option
| BF16 via SGLang17.2 GiB · Google DeepMind | 21.0 of 96.0 GiB | ~9 tok/s est. | |||||||||||||||||||||||||
Variant Hermes 3 Llama 3.1 8BMeta · 8.03B2 other options
| Q8_0 via llama.cpp8.0 GiB · NousResearch | 9.8 of 96.0 GiB | ~18 tok/s est. | |||||||||||||||||||||||||
Variant Llama 3.1 8B InstructMeta · 8.03B3 other options
| Q8_0 via llama.cpp7.9 GiB · LM Studio Community | 9.8 of 96.0 GiB | ~18 tok/s est. | |||||||||||||||||||||||||
Variant Qwen2.5 7B InstructQwen Team (Alibaba Cloud) · 7.62B3 other options
| Q8_0 via llama.cpp7.5 GiB · Qwen Team (Alibaba Cloud) | 8.8 of 96.0 GiB | ~19 tok/s est. | |||||||||||||||||||||||||
Variant DeepSeek-R1-Distill-Qwen-7BQwen Team (Alibaba Cloud) · 7.62B1 other option
| BF16 via SGLang14.2 GiB · DeepSeek | 15.5 of 96.0 GiB | ~11 tok/s est. | |||||||||||||||||||||||||
Variant Mistral 7B Instruct v0.3Mistral AI · 7.25B2 other options
| Q8_0 via llama.cpp7.2 GiB · LM Studio Community | 9.0 of 96.0 GiB | ~20 tok/s est. | |||||||||||||||||||||||||
Variant Mixtral 8x7B Instruct v0.1Mistral AI · 46.7B (12.9B active)1 other option
| BF16 via SGLang87.0 GiB · Mistral AITight fit | 92.0 of 96.0 GiB | ~6 tok/s est. |
Too large for this hardware
- DeepSeek-R1 671B · 3 downloads checked
How this is calculated
Memory = download size + fp16 KV cache for 8K tokens + runtime overhead. Usable memory is 95% of dedicated VRAM, a device-specific share of unified memory (75% by default), and 80% of system RAM for runtimes that can offload. Runtimes must load the file format and support a backend present on the hardware.
Estimated speed is bounded by memory bandwidth ÷ bytes read per token (active parameters for mixture-of-experts), shown in italics with est. and a ±35% range. Measured speeds are medians of reference results and verified public community runs on the same system, download and runtime. For each model we recommend the highest-precision download that fits, preferring not to offload.