Qwen3 32B
A 33 billion-parameterHow many numbers the model learned during training — the usual rough measure of its size. dense model from Qwen Team (Alibaba Cloud). Dense and MoE models with hybrid thinking modes.
- Size
- 32.8Bdense
- Memory to run
- —smallest, 8K ctx
- Context
- 40Ktokens
- License
- License unknown
- Updated
- 20 May 2025
Which version to use
Other sizes in Qwen3: Qwen3 0.6B, Qwen3 1.7B, Qwen3 4B, Qwen3 8B, Qwen3 14B, Qwen3 30B-A3B, Qwen3 235B-A22B
Where it runs
Best-fitting version on each reference system, at an 8K contextHow much text the model can consider at once, counted in tokens — roughly ¾ of a word each..
Benchmarks
Each benchmarkA fixed set of questions every model is given, so their scores can be compared on the same task. is developer-reported; prompts and settings differ between labs. Bars are relative to the best open result.
Variants & downloads
Each variant is a separate set of weights. Expand one for its downloads: quantizationStoring each of the model’s numbers with fewer bits, so the file is smaller and needs less memory. trades a little quality for a much smaller file, and the memory column adds the working memoryExtra space the model needs while it answers. It grows with the length of the conversation, on top of the file itself. a conversation needs on top.
Architecture details
- Parameters
- 32.76B
- Active / token
- All (dense)
- Architecture
- Dense
- Layers
- 64
- Attention heads
- 64
- KV heads
- 8
- Head dim
- 128
- KV cache @ 8K (fp16)
- 2.00 GiB
- Max context
- 40,960 tokens
Measured performance
Throughput on specific systems and runtimes, with the source of each measurement.
Community results
Reviews
Lineage
How this model’s variants relate to each other and to other models.
Sources & history
Sources
No source records linked.
Timeline
- Official Qwen3 GGUF quantizations published
- Qwen3 released with MoE and hybrid thinking