How open models are compared
A benchmarkA fixed set of questions every model is given, so their scores can be compared on the same task. is a fixed set of questions or a repeatable measurement, so different models can be put on the same scale. Mutinai records who reported each result and when.
What models can do
Fixed question sets a model answers. Each score says who produced it: the model’s developer, or the benchmark’s own maintainers running every model the same way.
GPQA Diamond · counts towards Reasoning
Graduate-level, search-resistant science questions.
Accuracy (%, higher is better) · 10 recorded results
developer-reported
- DeepSeek-R1 671B71.520 Jan 2025
- Qwen3 30B-A3B65.829 Apr 2025
- Qwen2.5 32B62.119 Sept 2024
- Phi-4 14B56.112 Dec 2024
- Llama 3.3 70B50.56 Dec 2024
Best open score over time
Each step is a release that beat the previous best recorded open result.
HumanEval · counts towards Coding
Python function synthesis from docstrings.
pass@1 (%, higher is better) · 9 recorded results
developer-reported
- Qwen2.5-Coder 32B92.712 Nov 2024
- Llama 3.3 70B88.46 Dec 2024
- Qwen2.5 32B88.419 Sept 2024
- Mistral Small 24B84.830 Jan 2025
- Qwen2.5 7B84.819 Sept 2024
Best open score over time
Each step is a release that beat the previous best recorded open result.
IFEval · counts towards Instruction following
Verifiable instruction-following prompts.
Strict prompt accuracy (%, higher is better) · 3 recorded results
developer-reported
- Llama 3.3 70B92.16 Dec 2024
- Gemma 3 27B90.412 Mar 2025
- Llama 3.1 8B80.423 Jul 2024
LiveCodeBench · counts towards Coding
Contamination-aware coding problems collected over time.
pass@1 (%, higher is better) · 4 recorded results
developer-reported
- DeepSeek-R1 671B65.920 Jan 2025
- Qwen3 30B-A3B62.629 Apr 2025
- Qwen2.5 32B57.219 Sept 2024
- Qwen2.5-Coder 32B31.412 Nov 2024
MATH-500 · counts towards Reasoning
500-problem subset of the MATH competition benchmark.
Accuracy (%, higher is better) · 3 recorded results
developer-reported
- DeepSeek-R1 671B97.320 Jan 2025
- Qwen2.5 32B94.319 Sept 2024
- Qwen2.5-Math 7B92.819 Sept 2024
MMLU-Pro · counts towards Knowledge
Harder, 10-option successor to MMLU across 14 domains.
Accuracy (%, higher is better) · 10 recorded results
developer-reported
- DeepSeek-R1 671B84.020 Jan 2025
- Phi-4 14B70.412 Dec 2024
- Qwen2.5 32B69.019 Sept 2024
- Llama 3.3 70B68.96 Dec 2024
- Gemma 3 27B67.512 Mar 2025
Best open score over time
Each step is a release that beat the previous best recorded open result.
How fast they run
Throughput measured on real hardware, by projects and by members. Speed depends on the machine, the runtime and the settings.
Interactive throughput
Runtime-agnostic single-user throughput measurement.
Generation throughput (tok/s, higher is better) · Peak memory (GB, lower is better) · Prompt throughput (tok/s, higher is better) · Time to first token (ms, lower is better) · 2 measured runs
2 runs — every one is in the table below, by model, system and runtime.
llama-bench
Standard llama.cpp throughput benchmark.
Prompt processing (512) (tok/s, higher is better) · Generation (128) (tok/s, higher is better) · 11 measured runs
11 runs — every one is in the table below, by model, system and runtime.
Measured speed, every run
Tokens per second for one model download, on one system, with one runtime. Generation is how fast it writes; prompt processing is how fast it reads what you give it.
| Model | System | Runtime | Context | Generation | Prompt | Test · who measured |
|---|---|---|---|---|---|---|
| Llama 3.1 8BQ4_K_M | RTX 4090 workstation | llama.cppcuda · b4600 | 4K | 128 tok/s | 12,100 tok/s | llama-bencheditorial · Mutinai illustrative fixtures |
| Llama 3.1 8BQ4_K_M | RX 7900 XTX desktop | llama.cpprocm · b4600 | 4K | 94 tok/s | 3,050 tok/s | llama-bencheditorial · Mutinai illustrative fixtures |
| Llama 3.1 8BQ4_K_M | MacBook Pro M3 Max 64 GB | llama.cppmetal · b4600 | 4K | 54 tok/s | 760 tok/s | llama-bencheditorial · Mutinai illustrative fixtures |
| Llama 3.1 8BQ4_K_M | CPU-only Ryzen 9 7950X | llama.cppcpu · b4600 | 4K | 12.4 tok/s | 92 tok/s | llama-bencheditorial · Mutinai illustrative fixtures |
| Llama 3.3 70BQ4_K_M | Dual RTX 3090 | llama.cppcuda · b4600 | 4K | 17.6 tok/s | 390 tok/s | llama-bencheditorial · Mutinai illustrative fixtures |
| Llama 3.3 70BQ4_K_M | Mac Studio M2 Ultra 192 GB | llama.cppmetal · b4600 | 4K | 12.1 tok/s | 135 tok/s | llama-bencheditorial · Mutinai illustrative fixtures |
| Llama 3.3 70BMLX 4-bit | MacBook Pro M4 Max 128 GB | MLX-LMmetal · 0.21.0 | 4K | 11.2 tok/s | 110 tok/s | Interactive throughputeditorial · Mutinai illustrative fixtures |
| Llama 3.3 70BQ4_K_M | MacBook Pro M3 Max 64 GB | llama.cppmetal · b4600 | 4K | 7.4 tok/s | 72 tok/s | llama-bencheditorial · Mutinai illustrative fixtures |
| Qwen2.5 14BQ4_K_M | RTX 4060 Ti 16GB budget build | llama.cppcuda · b4600 | 4K | 25.5 tok/s | 1,250 tok/s | llama-bencheditorial · Mutinai illustrative fixtures |
| Qwen2.5 32BQ4_K_M | RTX 5090 workstation | llama.cppcuda · b4600 | 4K | 61 tok/s | 2,900 tok/s | llama-bencheditorial · Mutinai illustrative fixtures |
| Qwen2.5 32BAWQ 4-bit | A100 80GB server | vLLMcuda · 0.7.3 | 8K | 48 tok/s | 5,200 tok/s | Interactive throughputeditorial · Mutinai illustrative fixtures |
| Qwen3 30B-A3BQ4_K_M | RTX 4090 workstation | llama.cppcuda · b5300 | 4K | 152 tok/s | 3,400 tok/s | llama-bencheditorial · Mutinai illustrative fixtures |
| Qwen3 30B-A3BQ4_K_M | Ryzen AI Max+ 395 mini PC | llama.cppvulkan · b5300 | 4K | 51 tok/s | 420 tok/s | llama-bencheditorial · Mutinai illustrative fixtures |
Developer-reported scores use each lab’s own prompts and settings, so numbers from different labs are not exactly comparable; a benchmark that runs every model itself uses one setup for all of them. Use them to shortlist, then look at what members measured on hardware like yours.