Benchmarks

How open models are compared

A benchmarkA fixed set of questions every model is given, so their scores can be compared on the same task. is a fixed set of questions or a repeatable measurement, so different models can be put on the same scale. Mutinai records who reported each result and when.

What models can do

Fixed question sets a model answers. Each score says who produced it: the model’s developer, or the benchmark’s own maintainers running every model the same way.

  • GPQA Diamond · counts towards Reasoning

    Graduate-level, search-resistant science questions.

    Accuracy (%, higher is better) · 10 recorded results

    developer-reported

    1. DeepSeek-R1 671B71.520 Jan 2025
    2. Qwen3 30B-A3B65.829 Apr 2025
    3. Qwen2.5 32B62.119 Sept 2024
    4. Phi-4 14B56.112 Dec 2024
    5. Llama 3.3 70B50.56 Dec 2024

    Best open score over time

    Each step is a release that beat the previous best recorded open result.

    3080Llama 3.1 8B Instruct: 30.4%Qwen2.5 32B Instruct: 49.5%Llama 3.3 70B Instruct: 50.5%Phi-4: 56.1%DeepSeek-R1: 71.5%DeepSeek-R1 · 71.5Jul 24Jan 25
  • HumanEval · counts towards Coding

    Python function synthesis from docstrings.

    pass@1 (%, higher is better) · 9 recorded results

    developer-reported

    1. Qwen2.5-Coder 32B92.712 Nov 2024
    2. Llama 3.3 70B88.46 Dec 2024
    3. Qwen2.5 32B88.419 Sept 2024
    4. Mistral Small 24B84.830 Jan 2025
    5. Qwen2.5 7B84.819 Sept 2024

    Best open score over time

    Each step is a release that beat the previous best recorded open result.

    80100Llama 3.1 70B Instruct: 80.5%Qwen2.5 32B Instruct: 88.4%Qwen2.5-Coder 32B Instruct: 92.7%Qwen2.5-Coder 32B Instruct · 92.7Jul 24Nov 24
  • IFEval · counts towards Instruction following

    Verifiable instruction-following prompts.

    Strict prompt accuracy (%, higher is better) · 3 recorded results

    developer-reported

    1. Llama 3.3 70B92.16 Dec 2024
    2. Gemma 3 27B90.412 Mar 2025
    3. Llama 3.1 8B80.423 Jul 2024
  • LiveCodeBench · counts towards Coding

    Contamination-aware coding problems collected over time.

    pass@1 (%, higher is better) · 4 recorded results

    developer-reported

    1. DeepSeek-R1 671B65.920 Jan 2025
    2. Qwen3 30B-A3B62.629 Apr 2025
    3. Qwen2.5 32B57.219 Sept 2024
    4. Qwen2.5-Coder 32B31.412 Nov 2024
  • MATH-500 · counts towards Reasoning

    500-problem subset of the MATH competition benchmark.

    Accuracy (%, higher is better) · 3 recorded results

    developer-reported

    1. DeepSeek-R1 671B97.320 Jan 2025
    2. Qwen2.5 32B94.319 Sept 2024
    3. Qwen2.5-Math 7B92.819 Sept 2024
  • MMLU-Pro · counts towards Knowledge

    Harder, 10-option successor to MMLU across 14 domains.

    Accuracy (%, higher is better) · 10 recorded results

    developer-reported

    1. DeepSeek-R1 671B84.020 Jan 2025
    2. Phi-4 14B70.412 Dec 2024
    3. Qwen2.5 32B69.019 Sept 2024
    4. Llama 3.3 70B68.96 Dec 2024
    5. Gemma 3 27B67.512 Mar 2025

    Best open score over time

    Each step is a release that beat the previous best recorded open result.

    6090Llama 3.1 70B Instruct: 66.4%Qwen2.5 32B Instruct: 69.0%Phi-4: 70.4%DeepSeek-R1: 84.0%DeepSeek-R1 · 84.0Jul 24Jan 25

How fast they run

Throughput measured on real hardware, by projects and by members. Speed depends on the machine, the runtime and the settings.

  • Interactive throughput

    Runtime-agnostic single-user throughput measurement.

    Generation throughput (tok/s, higher is better) · Peak memory (GB, lower is better) · Prompt throughput (tok/s, higher is better) · Time to first token (ms, lower is better) · 2 measured runs

    2 runs — every one is in the table below, by model, system and runtime.

  • llama-bench

    Standard llama.cpp throughput benchmark.

    Prompt processing (512) (tok/s, higher is better) · Generation (128) (tok/s, higher is better) · 11 measured runs

    11 runs — every one is in the table below, by model, system and runtime.

Measured speed, every run

Tokens per second for one model download, on one system, with one runtime. Generation is how fast it writes; prompt processing is how fast it reads what you give it.

ModelSystemRuntimeContextGenerationPromptTest · who measured
Llama 3.1 8BQ4_K_MRTX 4090 workstationllama.cppcuda · b46004K128 tok/s12,100 tok/sllama-bencheditorial · Mutinai illustrative fixtures
Llama 3.1 8BQ4_K_MRX 7900 XTX desktopllama.cpprocm · b46004K94 tok/s3,050 tok/sllama-bencheditorial · Mutinai illustrative fixtures
Llama 3.1 8BQ4_K_MMacBook Pro M3 Max 64 GBllama.cppmetal · b46004K54 tok/s760 tok/sllama-bencheditorial · Mutinai illustrative fixtures
Llama 3.1 8BQ4_K_MCPU-only Ryzen 9 7950Xllama.cppcpu · b46004K12.4 tok/s92 tok/sllama-bencheditorial · Mutinai illustrative fixtures
Llama 3.3 70BQ4_K_MDual RTX 3090llama.cppcuda · b46004K17.6 tok/s390 tok/sllama-bencheditorial · Mutinai illustrative fixtures
Llama 3.3 70BQ4_K_MMac Studio M2 Ultra 192 GBllama.cppmetal · b46004K12.1 tok/s135 tok/sllama-bencheditorial · Mutinai illustrative fixtures
Llama 3.3 70BMLX 4-bitMacBook Pro M4 Max 128 GBMLX-LMmetal · 0.21.04K11.2 tok/s110 tok/sInteractive throughputeditorial · Mutinai illustrative fixtures
Llama 3.3 70BQ4_K_MMacBook Pro M3 Max 64 GBllama.cppmetal · b46004K7.4 tok/s72 tok/sllama-bencheditorial · Mutinai illustrative fixtures
Qwen2.5 14BQ4_K_MRTX 4060 Ti 16GB budget buildllama.cppcuda · b46004K25.5 tok/s1,250 tok/sllama-bencheditorial · Mutinai illustrative fixtures
Qwen2.5 32BQ4_K_MRTX 5090 workstationllama.cppcuda · b46004K61 tok/s2,900 tok/sllama-bencheditorial · Mutinai illustrative fixtures
Qwen2.5 32BAWQ 4-bitA100 80GB servervLLMcuda · 0.7.38K48 tok/s5,200 tok/sInteractive throughputeditorial · Mutinai illustrative fixtures
Qwen3 30B-A3BQ4_K_MRTX 4090 workstationllama.cppcuda · b53004K152 tok/s3,400 tok/sllama-bencheditorial · Mutinai illustrative fixtures
Qwen3 30B-A3BQ4_K_MRyzen AI Max+ 395 mini PCllama.cppvulkan · b53004K51 tok/s420 tok/sllama-bencheditorial · Mutinai illustrative fixtures

Developer-reported scores use each lab’s own prompts and settings, so numbers from different labs are not exactly comparable; a benchmark that runs every model itself uses one setup for all of them. Use them to shortlist, then look at what members measured on hardware like yours.