54.2 tok/s generation throughput
Qwen3 30B-A3B MLX 4-bit
“Thinking disabled.”
Full environment and measurements
- Benchmark
- Interactive throughput
- Generation throughput
- 54.2 tok/s
- Peak memory
- 17.8 GB
- Prompt throughput
- 520 tok/s
- Time to first token
- 610 ms
- Artifact
- Qwen3 30B-A3B MLX 4-bit (MLX Community)
- Context
- 8K
- OS
- macOS 15.4