61.5 tok/s generation (128)
Qwen3 30B-A3B Q4_K_M
Full environment and measurements
- Benchmark
- llama-bench
- Prompt processing (512)
- 485 tok/s
- Generation (128)
- 61.5 tok/s
- Artifact
- Qwen3 30B-A3B Q4_K_M (Unsloth AI)
- Context
- 4K
- Flash attention
- on
- OS
- macOS 15.4