Tool · Inference runtime · ggml.org
llama.cpp
LLM inference in C/C++
- License
- PermissiveMIT License
- Language
- C++
- Latest release
- 14 Sept 2026llama.cpp v0.4.1
- Measurements
- 11with this runtime
Relationships
- Benchmarkllama-benchImplements
- ToolOllamaFoundation for
Releases & news
- llama.cpp v0.4.1llama.cpp 0.4.1 adds Maple 20B-A1B, Tencent Hy 4, and Spark2.5 support, improves JSON schema handling, chat parsing, logging, and server child-process management, and updates ggml to v0.24.0.
- llama.cpp b5450Automated build release.
- Official Qwen3 GGUF quantizations publishedThe Qwen team published first-party GGUF files for Qwen3 models.
Measured performance with this runtime
Throughput members and projects recorded, in tokens per secondHow fast the model writes its answer; about 10 tokens per second reads like comfortable typing..
| Model · quantization | System | Runtime | Context | Measurements | Source |
|---|---|---|---|---|---|
| Llama 3.1 8B Instruct Q4_K_M (LM Studio Community) | CPU-only Ryzen 9 7950X (128 GB DDR5) | llama.cppcpu · b4600 | 4K | 92 tok/s · Prompt processing (512) 12.4 tok/s · Generation (128) | Mutinai illustrative fixtures |
| Llama 3.3 70B Instruct Q4_K_M (LM Studio Community) | Dual RTX 3090 (128 GB DDR5) | llama.cppcuda · b4600 | 4K | 390 tok/s · Prompt processing (512) 17.6 tok/s · Generation (128) | Mutinai illustrative fixtures |
| Llama 3.1 8B Instruct Q4_K_M (LM Studio Community) | MacBook Pro M3 Max 64 GB | llama.cppmetal · b4600 | 4K | 760 tok/s · Prompt processing (512) 54 tok/s · Generation (128) | Mutinai illustrative fixtures |
| Llama 3.3 70B Instruct Q4_K_M (LM Studio Community) | MacBook Pro M3 Max 64 GB | llama.cppmetal · b4600 | 4K | 72 tok/s · Prompt processing (512) 7.4 tok/s · Generation (128) | Mutinai illustrative fixtures |
| Llama 3.3 70B Instruct Q4_K_M (LM Studio Community) | Mac Studio M2 Ultra 192 GB | llama.cppmetal · b4600 | 4K | 135 tok/s · Prompt processing (512) 12.1 tok/s · Generation (128) | Mutinai illustrative fixtures |
| Qwen2.5 14B Instruct Q4_K_M | RTX 4060 Ti 16GB budget build (32 GB DDR5) | llama.cppcuda · b4600 | 4K | 1,250 tok/s · Prompt processing (512) 25.5 tok/s · Generation (128) | Mutinai illustrative fixtures |
| Llama 3.1 8B Instruct Q4_K_M (LM Studio Community) | RTX 4090 workstation (64 GB DDR5) | llama.cppcuda · b4600 | 4K | 12,100 tok/s · Prompt processing (512) 128 tok/s · Generation (128) | Mutinai illustrative fixtures |
| Qwen3 30B-A3B Q4_K_M (Unsloth AI) | RTX 4090 workstation (64 GB DDR5) | llama.cppcuda · b5300 | 4K | 3,400 tok/s · Prompt processing (512) 152 tok/s · Generation (128) | Mutinai illustrative fixtures |
| Qwen2.5 32B Instruct Q4_K_M | RTX 5090 workstation (96 GB DDR5) | llama.cppcuda · b4600 | 4K | 2,900 tok/s · Prompt processing (512) 61 tok/s · Generation (128) | Mutinai illustrative fixtures |
| Llama 3.1 8B Instruct Q4_K_M (LM Studio Community) | RX 7900 XTX desktop (64 GB DDR5) | llama.cpprocm · b4600 | 4K | 3,050 tok/s · Prompt processing (512) 94 tok/s · Generation (128) | Mutinai illustrative fixtures |
| Qwen3 30B-A3B Q4_K_M (Unsloth AI) | Ryzen AI Max+ 395 mini PC (128 GB) | llama.cppvulkan · b5300 | 4K | 420 tok/s · Prompt processing (512) 51 tok/s · Generation (128) | Mutinai illustrative fixtures |
Reviews
Contributions are not open yet, so there is nothing here from members.
Sources & history
Sources
- GitHub (live source, 31 records, 22 Sept 2026)
- GitHub (fixture) (illustrative fixture, 1 record, 13 Sept 2026)
- Mutinai illustrative fixtures (illustrative fixture, 1 record, 13 Sept 2026)
External identifiers
- github ggml-org/llama.cpp
Field history
- homepageUrl = "https://llama.app" — GitHub (current)
- homepageUrl = "https://llama.app" — GitHub
- homepageUrl = "https://llama.app" — GitHub
- homepageUrl = "https://llama.app" — GitHub
- homepageUrl = "https://llama.app" — GitHub
- homepageUrl = "https://llama.app" — GitHub
- homepageUrl = "https://llama.app" — GitHub
- homepageUrl = "https://llama.app" — GitHub
- homepageUrl = "https://llama.app" — GitHub
- homepageUrl = "https://llama.app" — GitHub
- homepageUrl = "https://llama.app" — GitHub
- homepageUrl = "https://llama.app" — GitHub