Mistral 7B
A 7.2 billion-parameterHow many numbers the model learned during training — the usual rough measure of its size. dense model from Mistral AI. Extended vocabulary and function calling support.
- Size
- 7.25Bdense
- Memory to run
- ~6 GBsmallest, 8K ctx
- Context
- 32Ktokens
- License
- PermissiveApache License 2.0
- Updated
- 13 Sept 2026
Which version to use
- Variant Mistral 7B Instruct v0.3 · Mistral AITuned to follow instructions and hold a conversation. The usual choice.
- Variant mistral 7B instruct v0.3 bf16 · VikramPalA community or third-party fine-tune of another variant.
- Variant Mistral 7B Instruct v0.3 Jbliterated · ApolloRainesA community or third-party fine-tune of another variant.
- Variant Mistral BioMed Tool Caller 7B · RumiiiA community or third-party fine-tune of another variant.
Where it runs
Best-fitting version on each reference system, at an 8K contextHow much text the model can consider at once, counted in tokens — roughly ¾ of a word each..
Runs well 12
- A100 80GB server~161 tok/s est.
- Dual RTX 3090~70 tok/s est.
- GB10 mini workstation~22 tok/s est.
- MacBook Pro M3 Max 64 GB~32 tok/s est.
- MacBook Pro M4 Max 128 GB~43 tok/s est.
- Mac mini M4 Pro 48 GB~22 tok/s est.
- Mac Studio M2 Ultra 192 GB~63 tok/s est.
- RTX 4060 Ti 16GB budget build~23 tok/s est.
- RTX 4090 workstation~80 tok/s est.
- RTX 5090 workstation~141 tok/s est.
- RX 7900 XTX desktop~76 tok/s est.
- Ryzen AI Max+ 395 mini PC~20 tok/s est.
Slowly (CPU or offload) 1
- CPU-only Ryzen 9 7950X~7 tok/s est.
Too large 0
None of the reference systems.
Benchmarks
Each benchmarkA fixed set of questions every model is given, so their scores can be compared on the same task. is developer-reported; prompts and settings differ between labs. Bars are relative to the best open result.
Variants & downloads
Each variant is a separate set of weights. Expand one for its downloads: quantizationStoring each of the model’s numbers with fewer bits, so the file is smaller and needs less memory. trades a little quality for a much smaller file, and the memory column adds the working memoryExtra space the model needs while it answers. It grows with the length of the conversation, on top of the file itself. a conversation needs on top.
Variant Mistral 7B Instruct v0.3Permissive
- What it is
- Tuned to follow instructions and hold a conversation. The usual choice.
- Publisher
- Mistral AI
- License
- Apache License 2.0 · commercial use allowed
- Released
- 22 May 2024
- huggingface
- mistralai/Mistral-7B-Instruct-v0.3
Contributions are not open yet, so there is nothing here from members.
| Quantization | Format | Bits | Download | Memory @ 8K | Publisher |
|---|---|---|---|---|---|
| Download BF16native | safetensors | 16 | 13.5 GiB | Mistral AImistralai/Mistral-7B-Instruct-v0.3 | |
| Download Q8_0legacy gguf | gguf | 8.5 | 7.2 GiB | LM Studio Communitylmstudio-community/Mistral-7B-Instruct-v0.3-GGUF | |
| Download Q4_K_Mk quant | gguf | 4.89 | 4.1 GiB | LM Studio Communitylmstudio-community/Mistral-7B-Instruct-v0.3-GGUF |
Variant mistral 7B instruct v0.3 bf16License unknown
- What it is
- A community or third-party fine-tune of another variant.
- Publisher
- VikramPal
- License
- Unknown
- Released
- 13 Aug 2026
- huggingface
- VikramPal/mistral-7b-instruct-v0.3-bf16
Contributions are not open yet, so there is nothing here from members.
No downloadable artifacts recorded.
Variant Mistral 7B Instruct v0.3 JbliteratedPermissive
- What it is
- A community or third-party fine-tune of another variant.
- Publisher
- ApolloRaines
- License
- Apache License 2.0 · commercial use allowed
- Released
- 14 Jul 2026
Contributions are not open yet, so there is nothing here from members.
No downloadable artifacts recorded.
Variant Mistral BioMed Tool Caller 7BPermissive
- What it is
- A community or third-party fine-tune of another variant.
- Publisher
- Rumiii
- License
- Apache License 2.0 · commercial use allowed
- Released
- 24 Aug 2026
- huggingface
- Rumiii/Mistral-BioMed-Tool-Caller-7B
Contributions are not open yet, so there is nothing here from members.
No downloadable artifacts recorded.
Architecture details
- Parameters
- 7.25B
- Active / token
- All (dense)
- Architecture
- Dense
- Layers
- 32
- Attention heads
- 32
- KV heads
- 8
- Head dim
- 128
- KV cache @ 8K (fp16)
- 1.00 GiB
- Max context
- 32,768 tokens
Measured performance
Throughput on specific systems and runtimes, with the source of each measurement.
Community results
Reviews
Lineage
How this model’s variants relate to each other and to other models.
Sources & history
Sources
- Hugging Face Hub (live source, 1 record, 13 Sept 2026)
External identifiers
- huggingface mistralai/Mistral-7B-Instruct-v0.3
- huggingface VikramPal/mistral-7b-instruct-v0.3-bf16
- huggingface ApolloRaines/Mistral-7B-Instruct-v0.3-Jbliterated
- huggingface Rumiii/Mistral-BioMed-Tool-Caller-7B
Field history
- base = "mistralai/Mistral-7B-Instruct-v0.3" — Hugging Face Hub (current)
- derivation = "fine_tune" — Hugging Face Hub (current)
- license = "other" — Hugging Face Hub (current)
- name = "mistral 7B instruct v0.3 bf16" — Hugging Face Hub (current)
- base = "mistralai/Mistral-7B-Instruct-v0.3" — Hugging Face Hub (current)
- derivation = "fine_tune" — Hugging Face Hub (current)
- license = "apache-2.0" — Hugging Face Hub (current)
- name = "Mistral 7B Instruct v0.3 Jbliterated" — Hugging Face Hub (current)
- base = "mistralai/Mistral-7B-Instruct-v0.3" — Hugging Face Hub (current)
- derivation = "fine_tune" — Hugging Face Hub (current)
- license = "apache-2.0" — Hugging Face Hub (current)
- name = "Mistral BioMed Tool Caller 7B" — Hugging Face Hub (current)