Public preview · figures are illustrative fixture data, attributed to their source · data to May 2025 · read-only
Tool · Inference runtime · ggml.org

llama.cpp

LLM inference in C/C++

License
PermissiveMIT License
Language
C++
Latest release
14 Sept 2026llama.cpp v0.4.1
Measurements
11with this runtime

Relationships

  • Benchmarkllama-benchImplements
  • ToolOllamaFoundation for

Releases & news

  1. llama.cpp v0.4.1
    llama.cpp 0.4.1 adds Maple 20B-A1B, Tencent Hy 4, and Spark2.5 support, improves JSON schema handling, chat parsing, logging, and server child-process management, and updates ggml to v0.24.0.
  2. llama.cpp b5450
    Automated build release.
  3. Official Qwen3 GGUF quantizations published
    The Qwen team published first-party GGUF files for Qwen3 models.

Measured performance with this runtime

Throughput members and projects recorded, in tokens per secondHow fast the model writes its answer; about 10 tokens per second reads like comfortable typing..

Model · quantizationSystemRuntimeContextMeasurementsSource
Llama 3.1 8B Instruct Q4_K_M (LM Studio Community)CPU-only Ryzen 9 7950X (128 GB DDR5)llama.cppcpu · b46004K
92 tok/s · Prompt processing (512)
12.4 tok/s · Generation (128)
Mutinai illustrative fixtures
Llama 3.3 70B Instruct Q4_K_M (LM Studio Community)Dual RTX 3090 (128 GB DDR5)llama.cppcuda · b46004K
390 tok/s · Prompt processing (512)
17.6 tok/s · Generation (128)
Mutinai illustrative fixtures
Llama 3.1 8B Instruct Q4_K_M (LM Studio Community)MacBook Pro M3 Max 64 GBllama.cppmetal · b46004K
760 tok/s · Prompt processing (512)
54 tok/s · Generation (128)
Mutinai illustrative fixtures
Llama 3.3 70B Instruct Q4_K_M (LM Studio Community)MacBook Pro M3 Max 64 GBllama.cppmetal · b46004K
72 tok/s · Prompt processing (512)
7.4 tok/s · Generation (128)
Mutinai illustrative fixtures
Llama 3.3 70B Instruct Q4_K_M (LM Studio Community)Mac Studio M2 Ultra 192 GBllama.cppmetal · b46004K
135 tok/s · Prompt processing (512)
12.1 tok/s · Generation (128)
Mutinai illustrative fixtures
Qwen2.5 14B Instruct Q4_K_MRTX 4060 Ti 16GB budget build (32 GB DDR5)llama.cppcuda · b46004K
1,250 tok/s · Prompt processing (512)
25.5 tok/s · Generation (128)
Mutinai illustrative fixtures
Llama 3.1 8B Instruct Q4_K_M (LM Studio Community)RTX 4090 workstation (64 GB DDR5)llama.cppcuda · b46004K
12,100 tok/s · Prompt processing (512)
128 tok/s · Generation (128)
Mutinai illustrative fixtures
Qwen3 30B-A3B Q4_K_M (Unsloth AI)RTX 4090 workstation (64 GB DDR5)llama.cppcuda · b53004K
3,400 tok/s · Prompt processing (512)
152 tok/s · Generation (128)
Mutinai illustrative fixtures
Qwen2.5 32B Instruct Q4_K_MRTX 5090 workstation (96 GB DDR5)llama.cppcuda · b46004K
2,900 tok/s · Prompt processing (512)
61 tok/s · Generation (128)
Mutinai illustrative fixtures
Llama 3.1 8B Instruct Q4_K_M (LM Studio Community)RX 7900 XTX desktop (64 GB DDR5)llama.cpprocm · b46004K
3,050 tok/s · Prompt processing (512)
94 tok/s · Generation (128)
Mutinai illustrative fixtures
Qwen3 30B-A3B Q4_K_M (Unsloth AI)Ryzen AI Max+ 395 mini PC (128 GB)llama.cppvulkan · b53004K
420 tok/s · Prompt processing (512)
51 tok/s · Generation (128)
Mutinai illustrative fixtures

Reviews

Contributions are not open yet, so there is nothing here from members.

Sources & history

Sources

  • GitHub (live source, 31 records, 22 Sept 2026)
  • GitHub (fixture) (illustrative fixture, 1 record, 13 Sept 2026)
  • Mutinai illustrative fixtures (illustrative fixture, 1 record, 13 Sept 2026)

External identifiers

Field history

  • homepageUrl = "https://llama.app" GitHub (current)
  • homepageUrl = "https://llama.app" GitHub
  • homepageUrl = "https://llama.app" GitHub
  • homepageUrl = "https://llama.app" GitHub
  • homepageUrl = "https://llama.app" GitHub
  • homepageUrl = "https://llama.app" GitHub
  • homepageUrl = "https://llama.app" GitHub
  • homepageUrl = "https://llama.app" GitHub
  • homepageUrl = "https://llama.app" GitHub
  • homepageUrl = "https://llama.app" GitHub
  • homepageUrl = "https://llama.app" GitHub
  • homepageUrl = "https://llama.app" GitHub