Public preview · figures are illustrative fixture data, attributed to their source · data to May 2025 · read-only
Model · Qwen2.5 · September 2024 Live · Hugging Face

Qwen2.5 32B

A 33 billion-parameterHow many numbers the model learned during training — the usual rough measure of its size. dense model from Qwen Team (Alibaba Cloud). Dense models from 0.5B to 72B trained on 18T tokens.

Size
32.8Bdense
Memory to run
~21 GBsmallest, 8K ctx
Context
128Ktokens
License
PermissiveApache License 2.0; MIT License
Updated
13 Sept 2026

Which version to use

Other sizes in Qwen2.5: Qwen2.5 7B, Qwen2.5 14B

Where it runs

Best-fitting version on each reference system, at an 8K contextHow much text the model can consider at once, counted in tokens — roughly ¾ of a word each..

Runs well 11

Slowly (CPU or offload) 2

Too large 0

None of the reference systems.

Benchmarks

Each benchmarkA fixed set of questions every model is given, so their scores can be compared on the same task. is developer-reported; prompts and settings differ between labs. Bars are relative to the best open result.

BenchmarkQwen2.5 32B InstructDeepSeek-R1-Distill-Qwen-32B
Benchmark GPQA DiamondAccuracy
49.5src
62.1src
Benchmark HumanEvalpass@1
88.4src
Benchmark LiveCodeBenchpass@1
57.2src
Benchmark MATH-500Accuracy
94.3src
Benchmark MMLU-ProAccuracy
69.0src

developer reported

Variants & downloads

Each variant is a separate set of weights. Expand one for its downloads: quantizationStoring each of the model’s numbers with fewer bits, so the file is smaller and needs less memory. trades a little quality for a much smaller file, and the memory column adds the working memoryExtra space the model needs while it answers. It grows with the length of the conversation, on top of the file itself. a conversation needs on top.

Variant Qwen2.5 32Bbase · Qwen Team (Alibaba Cloud) · 1 download · Apache License 2.0Permissive
What it is
Pretrained foundation weights — a starting point for fine-tuning, not for chat.
Publisher
Qwen Team (Alibaba Cloud)
License
Apache License 2.0 · commercial use allowed
Released
19 Sept 2024
huggingface
Qwen/Qwen2.5-32B

Contributions are not open yet, so there is nothing here from members.

QuantizationFormatBitsDownloadMemory @ 8KPublisher
Download BF16nativesafetensors1661.0 GiB~66.0 GiBQwen Team (Alibaba Cloud)Qwen/Qwen2.5-32B
Variant Qwen2.5 32B Instructinstruct · Qwen Team (Alibaba Cloud) · 5 downloads · Apache License 2.0Permissive
What it is
Tuned to follow instructions and hold a conversation. The usual choice.
Publisher
Qwen Team (Alibaba Cloud)
License
Apache License 2.0 · commercial use allowed
Released
19 Sept 2024

Contributions are not open yet, so there is nothing here from members.

QuantizationFormatBitsDownloadMemory @ 8KPublisher
Download BF16nativesafetensors1661.0 GiB~66.0 GiBQwen Team (Alibaba Cloud)Qwen/Qwen2.5-32B-Instruct
Download Q5_K_Mk quantgguf5.6921.7 GiB~25.1 GiBLM Studio Communitylmstudio-community/Qwen2.5-32B-Instruct-GGUF
Download Q4_K_Mk quantgguf4.8918.7 GiB~21.9 GiBQwen Team (Alibaba Cloud)Qwen/Qwen2.5-32B-Instruct-GGUF
Download AWQ 4-bitawqsafetensors4.617.5 GiB~20.7 GiBQwen Team (Alibaba Cloud)Qwen/Qwen2.5-32B-Instruct-AWQ
Download MLX 4-bitmlxmlx4.517.2 GiB~20.4 GiBMLX Communitymlx-community/Qwen2.5-32B-Instruct-4bit
Variant DeepSeek-R1-Distill-Qwen-32Bdistill · DeepSeek (third-party) · 3 downloads · MIT LicensePermissive
What it is
A smaller model trained to imitate a larger one.
Publisher
DeepSeek
License
MIT License · commercial use allowed
Released
20 Jan 2025

Contributions are not open yet, so there is nothing here from members.

QuantizationFormatBitsDownloadMemory @ 8KPublisher
Download BF16nativesafetensors1661.0 GiB~66.0 GiBDeepSeekdeepseek-ai/DeepSeek-R1-Distill-Qwen-32B
Download Q4_K_Mk quantgguf4.8918.7 GiB~21.9 GiBLM Studio Communitylmstudio-community/DeepSeek-R1-Distill-Qwen-32B-GGUF
Download MLX 4-bitmlxmlx4.517.2 GiB~20.4 GiBMLX Communitymlx-community/DeepSeek-R1-Distill-Qwen-32B-4bit
Variant RaDaR 32Bfine tune · sczzz (third-party) · 0 downloads · Apache License 2.0Permissive
What it is
A community or third-party fine-tune of another variant.
Publisher
sczzz
License
Apache License 2.0 · commercial use allowed
Released
19 May 2026
huggingface
sczzz/RaDaR-32B

Contributions are not open yet, so there is nothing here from members.

No downloadable artifacts recorded.

Architecture detailsLayers, attention and KV cache geometry
Parameters
32.76B
Active / token
All (dense)
Architecture
Dense
Layers
64
Attention heads
40
KV heads
8
Head dim
128
KV cache @ 8K (fp16)
2.00 GiB
Max context
131,072 tokens

Measured performance

Throughput on specific systems and runtimes, with the source of each measurement.

QuantizationSystemRuntimeContextMeasurementsSource
Qwen2.5 32B Instruct AWQ 4-bitA100 80GB server (512 GB DDR4)vLLMcuda · 0.7.38K
48 tok/s · Generation throughput
72 GB · Peak memory
5,200 tok/s · Prompt throughput
95 ms · Time to first token
Mutinai illustrative fixtures
Qwen2.5 32B Instruct Q4_K_MRTX 5090 workstation (96 GB DDR5)llama.cppcuda · b46004K
2,900 tok/s · Prompt processing (512)
61 tok/s · Generation (128)
Mutinai illustrative fixtures

Community results

Contributions are not open yet, so there is nothing here from members.

Reviews

Contributions are not open yet, so there is nothing here from members.

Lineage

How this model’s variants relate to each other and to other models.

VariantQwen2.5 32B base
VariantDeepSeek-R1-Distill-Qwen-32B fine-tuned from Qwen2.5 32B· distilled from DeepSeek-R1
VariantQwen2.5 32B Instruct instruct

Sources & history

Sources

  • Mutinai illustrative fixtures (illustrative fixture, 1 record, 13 Sept 2026)
  • Hugging Face Hub (live source, 1 record, 13 Sept 2026)

External identifiers

Field history

  • base = "deepseek-ai/DeepSeek-R1-Distill-Qwen-32B" Hugging Face Hub (current)
  • derivation = "fine_tune" Hugging Face Hub (current)
  • license = "apache-2.0" Hugging Face Hub (current)
  • name = "RaDaR 32B" Hugging Face Hub (current)

Timeline

  1. Qwen2.5 family released