Gemma 2 9B
A 9.2 billion-parameterHow many numbers the model learned during training — the usual rough measure of its size. dense model from Google DeepMind. 9B and 27B models with interleaved local/global attention.
- Size
- 9.24Bdense
- Memory to run
- ~9 GB estimatesmallest, 8K ctx
- Context
- 8Ktokens
- License
- RestrictedGemma Terms of Use
- Updated
- —
Where it runs
Best-fitting version on each reference system, at an 8K contextHow much text the model can consider at once, counted in tokens — roughly ¾ of a word each..
Runs well 16
- A100 80GB server~70 tok/s estimate
- Dual RTX 3090~30 tok/s estimate
- GB10 mini workstation~9 tok/s estimate
- MacBook Air M2 16 GB~11 tok/s estimate
- MacBook Air M3 16 GB~11 tok/s estimate
- MacBook Air M4 16 GB~13 tok/s estimate
- MacBook Pro M3 Max 64 GB~42 tok/s estimate
- MacBook Pro M4 Max 128 GB~57 tok/s estimate
- Mac mini M4 16 GB~13 tok/s estimate
- Mac mini M4 Pro 48 GB~29 tok/s estimate
- Mac Studio M2 Ultra 192 GB~84 tok/s estimate
- RTX 4060 Ti 16GB budget build~30 tok/s estimate
- RTX 4090 workstation~34 tok/s estimate
- RTX 5090 workstation~61 tok/s estimate
- RX 7900 XTX desktop~33 tok/s estimate
- Ryzen AI Max+ 395 mini PC~9 tok/s estimate
Slowly (CPU or offload) 2
- CPU-only Ryzen 9 7950X~9 tok/s estimate
- Windows laptop 16 GB~9 tok/s estimate
Too large 1
Measured throughput
Versions & downloads
Each version is a separate set of weights. Expand one for its downloads: quantizationStoring each of the model’s numbers with fewer bits, so the file is smaller and needs less memory. trades a little quality for a much smaller file, and the memory column adds the working memoryExtra space the model needs while it answers. It grows with the length of the conversation, on top of the file itself. a conversation needs on top.
Variant Gemma 2 9B ITRestricted
- Good for
- Publisher
- Google DeepMind
- License
- Gemma Terms of Use · commercial use restricted
- Released
- 27 Jun 2024
- huggingface
- google/gemma-2-9b-it
Contributions are not open yet, so there is nothing here from members.
| Quantization | Format | Bits | Download | Memory @ 8K | Publisher |
|---|---|---|---|---|---|
| Download BF16native | safetensors | 16 | 17.2 GiB | Google DeepMindgoogle/gemma-2-9b-it | |
| Download Q4_K_Mk quant | gguf | 4.89 | 5.3 GiB | LM Studio Communitylmstudio-community/gemma-2-9b-it-GGUF |
Other sizes in Gemma 2: Gemma 2 27B
Benchmarks
Each benchmarkA fixed set of questions every model is given, so their scores can be compared on the same task. score is labelled with who produced it. Developer-reported scores use each lab’s own prompts and settings; a benchmark that runs every model itself uses one setup for all. Bars are relative to the best open result.
Community
Benchmark runs and reviews from people running it on their own hardware.
Benchmark runs
Reviews
Details & sources
Architecture
- Parameters
- 9.24B
- Active / token
- All (dense)
- Architecture
- Dense
- Layers
- 42
- Attention heads
- 16
- KV heads
- 8
- Head dim
- 256
- KV cache @ 8K (fp16)
- 2.63 GiB
- Max context
- 8,192 tokens
Lineage
How this model’s versions relate to each other and to other models.