Full environment and measurements
- Benchmark
- llama-bench
- Prompt processing (512)
- 485 tok/s
- Generation (128)
- 61.5 tok/s
- Artifact
- Qwen3 30B-A3B Q4_K_M (Unsloth AI)
- Context
- 4K
- Flash attention
- on
- OS
- macOS 15.4
Runs record exactly what was run and on what, so results can be reproduced. Reviews rate specific qualities, not a single score.
“Thinking disabled.”
“Both cards power limited to 280W.”
Q4_K_M with 16K context fits on the 4090 with room to spare. Excellent at targeted edits in Aider; weaker at long agentic loops where it sometimes loses the plan.
Fastest path to new model support and runs on everything. The flags surface is huge; defaults change between builds, so pin versions when benchmarking.
Bought three used. Power limits at 280W lose very little generation speed. Budget time for riser cables and airflow.
Great for getting teammates started. I switch to llama.cpp directly when I need specific KV cache or offload settings.
Runs at interactive speeds in MLX 4-bit. Thinking mode helps on multi-step problems; turn it off for quick chat.
New runs are unverified. Moderators verify runs whose numbers are consistent with the hardware and settings. Only verified, public runs feed the measured speeds in What can I run?.
Everything you contribute is public, unlisted (reachable by link, excluded from listings and averages) or private (only you). Public runs on your own systems show the hardware components, not your system’s name. Contributions are never used for AI training unless you opt in. Read the principles.