Llama 3.1 8B
A 8 billion-parameterHow many numbers the model learned during training — the usual rough measure of its size. dense model from Meta. 8B, 70B and 405B models with 128K context.
- Size
- 8.03Bdense
- Memory to run
- ~6 GBsmallest, 8K ctx
- Context
- 128Ktokens
- License
- PermissiveLlama 3.1 Community License; Apache License 2.0
- Updated
- 18 Sept 2026
Which version to use
- Variant Llama 3.1 8B Instruct · MetaTuned to follow instructions and hold a conversation. The usual choice.
- Variant Hermes · dedsecisback2026A community or third-party fine-tune of another variant.
- Variant Hermes 3 Llama 3.1 8B · NousResearchA community or third-party fine-tune of another variant.
- Variant neurologist 7B instruct · mainbrainsA community or third-party fine-tune of another variant.
- Variant solana nvidia trading factory 8B · solanaclawdA community or third-party fine-tune of another variant.
Other sizes in Llama 3.1: Llama 3.1 70B
Where it runs
Best-fitting version on each reference system, at an 8K contextHow much text the model can consider at once, counted in tokens — roughly ¾ of a word each..
Runs well 12
- A100 80GB server~146 tok/s est.
- Dual RTX 3090~64 tok/s est.
- GB10 mini workstation~20 tok/s est.
- MacBook Pro M3 Max 64 GB~29 tok/s est.
- MacBook Pro M4 Max 128 GB~39 tok/s est.
- Mac mini M4 Pro 48 GB~20 tok/s est.
- Mac Studio M2 Ultra 192 GB~57 tok/s est.
- RTX 4060 Ti 16GB budget build~21 tok/s est.
- RTX 4090 workstation~72 tok/s est.
- RTX 5090 workstation~128 tok/s est.
- RX 7900 XTX desktop~69 tok/s est.
- Ryzen AI Max+ 395 mini PC~18 tok/s est.
Slowly (CPU or offload) 1
- CPU-only Ryzen 9 7950X~6 tok/s est.
Too large 0
None of the reference systems.
Benchmarks
Each benchmarkA fixed set of questions every model is given, so their scores can be compared on the same task. is developer-reported; prompts and settings differ between labs. Bars are relative to the best open result.
Variants & downloads
Each variant is a separate set of weights. Expand one for its downloads: quantizationStoring each of the model’s numbers with fewer bits, so the file is smaller and needs less memory. trades a little quality for a much smaller file, and the memory column adds the working memoryExtra space the model needs while it answers. It grows with the length of the conversation, on top of the file itself. a conversation needs on top.
Variant Llama 3.1 8BRestricted
- What it is
- Pretrained foundation weights — a starting point for fine-tuning, not for chat.
- Publisher
- Meta
- License
- Llama 3.1 Community License · commercial use restricted
- Released
- 23 Jul 2024
- huggingface
- meta-llama/Llama-3.1-8B
Contributions are not open yet, so there is nothing here from members.
| Quantization | Format | Bits | Download | Memory @ 8K | Publisher |
|---|---|---|---|---|---|
| Download BF16native | safetensors | 16 | 15.0 GiB | Metameta-llama/Llama-3.1-8B |
Variant Llama 3.1 8B InstructRestricted
- What it is
- Tuned to follow instructions and hold a conversation. The usual choice.
- Publisher
- Meta
- License
- Llama 3.1 Community License · commercial use restricted
- Released
- 23 Jul 2024
- huggingface
- meta-llama/Llama-3.1-8B-Instruct
Contributions are not open yet, so there is nothing here from members.
| Quantization | Format | Bits | Download | Memory @ 8K | Publisher |
|---|---|---|---|---|---|
| Download BF16native | safetensors | 16 | 15.0 GiB | Metameta-llama/Llama-3.1-8B-Instruct | |
| Download Q8_0legacy gguf | gguf | 8.5 | 7.9 GiB | LM Studio Communitylmstudio-community/Llama-3.1-8B-Instruct-GGUF | |
| Download FP8fp8 | safetensors | 8.1 | 8.5 GiB | RedHatAIRedHatAI/Meta-Llama-3.1-8B-Instruct-FP8 | |
| Download Q4_K_Mk quant | gguf | 4.89 | 4.6 GiB | LM Studio Communitylmstudio-community/Llama-3.1-8B-Instruct-GGUF | |
| Download MLX 4-bitmlx | mlx | 4.5 | 4.2 GiB | MLX Communitymlx-community/Llama-3.1-8B-Instruct-4bit |
Variant HermesLicense unknown
- What it is
- A community or third-party fine-tune of another variant.
- Publisher
- dedsecisback2026
- License
- Unknown
- Released
- 21 Aug 2026
- huggingface
- dedsecisback2026/Hermes
Contributions are not open yet, so there is nothing here from members.
No downloadable artifacts recorded.
Variant Hermes 3 Llama 3.1 8BRestricted
- What it is
- A community or third-party fine-tune of another variant.
- Publisher
- NousResearch
- License
- Llama 3.1 Community License · commercial use restricted
- Released
- 28 Jul 2024
- huggingface
- NousResearch/Hermes-3-Llama-3.1-8B
Contributions are not open yet, so there is nothing here from members.
| Quantization | Format | Bits | Download | Memory @ 8K | Publisher |
|---|---|---|---|---|---|
| Download Q8_0legacy gguf | gguf | 8.5 | 8.0 GiB | NousResearchNousResearch/Hermes-3-Llama-3.1-8B-GGUF | |
| Download Q6_Kk quant | gguf | 6.56 | 6.1 GiB | NousResearchNousResearch/Hermes-3-Llama-3.1-8B-GGUF | |
| Download Q4_K_Mk quant | gguf | 4.89 | 4.6 GiB | NousResearchNousResearch/Hermes-3-Llama-3.1-8B-GGUF |
Variant neurologist 7B instructPermissive
- What it is
- A community or third-party fine-tune of another variant.
- Publisher
- mainbrains
- License
- Apache License 2.0 · commercial use allowed
- Released
- 20 Aug 2026
- huggingface
- mainbrains/neurologist-7b-instruct
Contributions are not open yet, so there is nothing here from members.
No downloadable artifacts recorded.
Variant solana nvidia trading factory 8BPermissive
- What it is
- A community or third-party fine-tune of another variant.
- Publisher
- solanaclawd
- License
- Apache License 2.0 · commercial use allowed
- Released
- 6 Sept 2026
- huggingface
- solanaclawd/solana-nvidia-trading-factory-8b
Contributions are not open yet, so there is nothing here from members.
No downloadable artifacts recorded.
Architecture details
- Parameters
- 8.03B
- Active / token
- All (dense)
- Architecture
- Dense
- Layers
- 32
- Attention heads
- 32
- KV heads
- 8
- Head dim
- 128
- KV cache @ 8K (fp16)
- 1.00 GiB
- Max context
- 131,072 tokens
Measured performance
Throughput on specific systems and runtimes, with the source of each measurement.
| Quantization | System | Runtime | Context | Measurements | Source |
|---|---|---|---|---|---|
| Llama 3.1 8B Instruct Q4_K_M (LM Studio Community) | CPU-only Ryzen 9 7950X (128 GB DDR5) | llama.cppcpu · b4600 | 4K | 92 tok/s · Prompt processing (512) 12.4 tok/s · Generation (128) | Mutinai illustrative fixtures |
| Llama 3.1 8B Instruct Q4_K_M (LM Studio Community) | MacBook Pro M3 Max 64 GB | llama.cppmetal · b4600 | 4K | 760 tok/s · Prompt processing (512) 54 tok/s · Generation (128) | Mutinai illustrative fixtures |
| Llama 3.1 8B Instruct Q4_K_M (LM Studio Community) | RTX 4090 workstation (64 GB DDR5) | llama.cppcuda · b4600 | 4K | 12,100 tok/s · Prompt processing (512) 128 tok/s · Generation (128) | Mutinai illustrative fixtures |
| Llama 3.1 8B Instruct Q4_K_M (LM Studio Community) | RX 7900 XTX desktop (64 GB DDR5) | llama.cpprocm · b4600 | 4K | 3,050 tok/s · Prompt processing (512) 94 tok/s · Generation (128) | Mutinai illustrative fixtures |
Community results
Reviews
Lineage
How this model’s variants relate to each other and to other models.
Sources & history
Sources
- Hugging Face (fixture) (illustrative fixture, 1 record, 13 Sept 2026)
- Hugging Face Hub (live source, 1 record, 13 Sept 2026)
- Mutinai illustrative fixtures (illustrative fixture, 1 record, 13 Sept 2026)
External identifiers
- huggingface meta-llama/Llama-3.1-8B
- huggingface meta-llama/Llama-3.1-8B-Instruct
- huggingface dedsecisback2026/Hermes
- huggingface NousResearch/Hermes-3-Llama-3.1-8B
- huggingface mainbrains/neurologist-7b-instruct
- huggingface solanaclawd/solana-nvidia-trading-factory-8b
Field history
- license = "llama3.1" — Hugging Face Hub (current)
- license = "llama3.1" — Hugging Face (fixture)
- base = "NousResearch/Hermes-3-Llama-3.1-8B" — Hugging Face Hub (current)
- derivation = "fine_tune" — Hugging Face Hub (current)
- license = null — Hugging Face Hub (current)
- name = "Hermes" — Hugging Face Hub (current)
- base = "meta-llama/Llama-3.1-8B" — Hugging Face (fixture) (current)
- derivation = "fine_tune" — Hugging Face (fixture) (current)
- license = "llama3" — Hugging Face (fixture) (current)
- name = "Hermes 3 Llama 3.1 8B" — Hugging Face (fixture) (current)
- base = "NousResearch/Hermes-3-Llama-3.1-8B" — Hugging Face Hub (current)
- derivation = "fine_tune" — Hugging Face Hub (current)
Timeline
- Hermes 3 fine-tunes of Llama 3.1 released
- Meta releases Llama 3.1 (8B, 70B, 405B)