System · mini pc
GB10 mini workstation (128 GB)
Runs models up to ~160B parameters at 4-bit in unified memory. CUDA system with 128 GB unified memory.
- Unified memory
- 128 GB
- System RAM
- Shared
- Runs
- 18 of 19model variants, 8K
- Estimated cost
- ~$4,000
What it runs
Largest models that fit without spilling into system memory, at an 8K contextHow much text the model can consider at once, counted in tokens — roughly ¾ of a word each..
- VariantLlama 3.1 70B Instruct70.6B · Q4_K_M via llama.cppRuns well~4 tok/s est.
- VariantLlama 3.3 70B Instruct70.6B · Q4_K_M via llama.cppRuns well~4 tok/s est.
- VariantMixtral 8x7B Instruct v0.146.7B · BF16 via SGLangRuns well~7 tok/s est.
- VariantQwen2.5 32B Instruct32.8B · BF16 via SGLangRuns well~3 tok/s est.
- VariantQwen2.5-Coder 32B Instruct32.8B · Q8_0 via llama.cppRuns well~5 tok/s est.
- VariantDeepSeek-R1-Distill-Qwen-32B32.8B · BF16 via SGLangRuns well~3 tok/s est.
- VariantQwen3 30B-A3B30.5B · Q8_0 via llama.cppRuns well~44 tok/s est.
- VariantGemma 3 27B IT27.4B · BF16 via SGLangRuns well~3 tok/s est.
- VariantGemma 2 27B IT27.2B · BF16 via SGLangRuns well~3 tok/s est.
- VariantMistral Small 24B Instruct 250123.6B · BF16 via SGLangRuns well~4 tok/s est.
Components
Measured performance
No performance measurements recorded yet.
Community results
Contributions are not open yet, so there is nothing here from members.
Reviews
Contributions are not open yet, so there is nothing here from members.