Model · Gemma 2 · June 2024

Gemma 2 9B

A 9.2 billion-parameterHow many numbers the model learned during training — the usual rough measure of its size. dense model from Google DeepMind. 9B and 27B models with interleaved local/global attention.

Size
9.24Bdense
Memory to run
~9 GB estimatesmallest, 8K ctx
Context
8Ktokens
License
RestrictedGemma Terms of Use
Updated

Where it runs

Best-fitting version on each reference system, at an 8K contextHow much text the model can consider at once, counted in tokens — roughly ¾ of a word each..

Runs well 16

Slowly (CPU or offload) 2

Too large 1

Measured throughput

No performance measurements recorded yet.

Versions & downloads

Each version is a separate set of weights. Expand one for its downloads: quantizationStoring each of the model’s numbers with fewer bits, so the file is smaller and needs less memory. trades a little quality for a much smaller file, and the memory column adds the working memoryExtra space the model needs while it answers. It grows with the length of the conversation, on top of the file itself. a conversation needs on top.

Variant Gemma 2 9B ITTuned to follow instructions and hold a conversation. The usual choice. · Google DeepMind · 2 downloadsRestricted
Good for
Chat
Publisher
Google DeepMind
License
Gemma Terms of Use · commercial use restricted
Released
27 Jun 2024

Contributions are not open yet, so there is nothing here from members.

QuantizationFormatBitsDownloadMemory @ 8KPublisher
Download BF16nativesafetensors1617.2 GiB~21.0 GiB estimateGoogle DeepMindgoogle/gemma-2-9b-it
Download Q4_K_Mk quantgguf4.895.3 GiB~8.6 GiB estimateLM Studio Communitylmstudio-community/gemma-2-9b-it-GGUF

Other sizes in Gemma 2: Gemma 2 27B

Benchmarks

Each benchmarkA fixed set of questions every model is given, so their scores can be compared on the same task. score is labelled with who produced it. Developer-reported scores use each lab’s own prompts and settings; a benchmark that runs every model itself uses one setup for all. Bars are relative to the best open result.

No benchmark results recorded for this model.

Community

Benchmark runs and reviews from people running it on their own hardware.

Benchmark runs

Contributions are not open yet, so there is nothing here from members.

Reviews

Contributions are not open yet, so there is nothing here from members.

Details & sources

Architecture

Parameters
9.24B
Active / token
All (dense)
Architecture
Dense
Layers
42
Attention heads
16
KV heads
8
Head dim
256
KV cache @ 8K (fp16)
2.63 GiB
Max context
8,192 tokens

Lineage

How this model’s versions relate to each other and to other models.

VariantGemma 2 9B IT instruct

Sources

No source records linked.

External identifiers