Qwen3 30B-A3B
A 31 billion-parameterHow many numbers the model learned during training — the usual rough measure of its size. mixture-of-expertsA model split into parts where only a few run for each word, so it answers faster than its size suggests. model using ~3.3 billion parameters per token from Qwen Team (Alibaba Cloud). Dense and MoE models with hybrid thinking modes.
- Size
- 30.5B3.3B active
- Memory to run
- ~18 GBsmallest, 8K ctx
- Context
- 32Ktokens
- License
- PermissiveApache License 2.0
- Updated
- 18 Sept 2026
Which version to use
- Variant Qwen3 30B-A3B · Qwen Team (Alibaba Cloud)Tuned to follow instructions and hold a conversation. The usual choice.
- Variant Qwen3 30B A3B Instruct 2507 · Qwen Team (Alibaba Cloud)Tuned to follow instructions and hold a conversation. The usual choice.
- Variant CantoneseLLM v2.0 30B A3B Thinking · hon9kon9izeA community or third-party fine-tune of another variant.
- Variant Qwen3 30B A3B ShapleyMCG K34 Validation Reconstruction · brandonmusicA community or third-party fine-tune of another variant.
Other sizes in Qwen3: Qwen3 0.6B, Qwen3 1.7B, Qwen3 4B, Qwen3 8B, Qwen3 14B, Qwen3 32B, Qwen3 235B-A22B
Where it runs
Best-fitting version on each reference system, at an 8K contextHow much text the model can consider at once, counted in tokens — roughly ¾ of a word each..
Runs well 11
- A100 80GB server~328 tok/s est.
- Dual RTX 3090~143 tok/s est.
- GB10 mini workstation~44 tok/s est.
- MacBook Pro M3 Max 64 GB~64 tok/s est.
- MacBook Pro M4 Max 128 GB~88 tok/s est.
- Mac mini M4 Pro 48 GB~44 tok/s est.
- Mac Studio M2 Ultra 192 GB~129 tok/s est.
- RTX 4090 workstation~258 tok/s est.
- RTX 5090 workstation~458 tok/s est.
- RX 7900 XTX desktop~245 tok/s est.
- Ryzen AI Max+ 395 mini PC~41 tok/s est.
Slowly (CPU or offload) 2
- CPU-only Ryzen 9 7950X~13 tok/s est.
- RTX 4060 Ti 16GB budget build~50 tok/s est.
Too large 0
None of the reference systems.
Benchmarks
Each benchmarkA fixed set of questions every model is given, so their scores can be compared on the same task. is developer-reported; prompts and settings differ between labs. Bars are relative to the best open result.
Variants & downloads
Each variant is a separate set of weights. Expand one for its downloads: quantizationStoring each of the model’s numbers with fewer bits, so the file is smaller and needs less memory. trades a little quality for a much smaller file, and the memory column adds the working memoryExtra space the model needs while it answers. It grows with the length of the conversation, on top of the file itself. a conversation needs on top.
Variant Qwen3 30B-A3B BasePermissive
- What it is
- Pretrained foundation weights — a starting point for fine-tuning, not for chat.
- Publisher
- Qwen Team (Alibaba Cloud)
- License
- Apache License 2.0 · commercial use allowed
- Released
- 29 Apr 2025
- huggingface
- Qwen/Qwen3-30B-A3B-Base
Contributions are not open yet, so there is nothing here from members.
| Quantization | Format | Bits | Download | Memory @ 8K | Publisher |
|---|---|---|---|---|---|
| Download BF16native | safetensors | 16 | 56.9 GiB | Qwen Team (Alibaba Cloud)Qwen/Qwen3-30B-A3B-Base |
Variant Qwen3 30B-A3BPermissive
- What it is
- Tuned to follow instructions and hold a conversation. The usual choice.
- Publisher
- Qwen Team (Alibaba Cloud)
- License
- Apache License 2.0 · commercial use allowed
- Released
- 29 Apr 2025
- huggingface
- Qwen/Qwen3-30B-A3B
Contributions are not open yet, so there is nothing here from members.
| Quantization | Format | Bits | Download | Memory @ 8K | Publisher |
|---|---|---|---|---|---|
| Download BF16native | safetensors | 16 | 56.9 GiB | Qwen Team (Alibaba Cloud)Qwen/Qwen3-30B-A3B | |
| Download Q8_0legacy gguf | gguf | 8.5 | 30.2 GiB | LM Studio Communitylmstudio-community/Qwen3-30B-A3B-GGUF | |
| Download MLX 8-bitmlx | mlx | 8.5 | 30.2 GiB | MLX Communitymlx-community/Qwen3-30B-A3B-8bit | |
| Download Q8_0legacy gguf | gguf | 8.5 | 30.3 GiB | Qwen Team (Alibaba Cloud)Qwen/Qwen3-30B-A3B-GGUF | |
| Download FP8fp8 | safetensors | 8.1 | 28.8 GiB | Qwen Team (Alibaba Cloud)Qwen/Qwen3-30B-A3B-FP8 | |
| Download Q4_K_Mk quant | gguf | 4.89 | 17.3 GiB | Qwen Team (Alibaba Cloud)Qwen/Qwen3-30B-A3B-GGUF | |
| Download Q4_K_Mk quant | gguf | 4.89 | 17.4 GiB | Unsloth AIunsloth/Qwen3-30B-A3B-GGUF | |
| Download MLX 4-bitmlx | mlx | 4.5 | 16.0 GiB | MLX Communitymlx-community/Qwen3-30B-A3B-4bit |
Variant Qwen3 30B A3B Instruct 2507Permissive
- What it is
- Tuned to follow instructions and hold a conversation. The usual choice.
- Publisher
- Qwen Team (Alibaba Cloud)
- License
- Apache License 2.0 · commercial use allowed
- Released
- 28 Jul 2025
- huggingface
- Qwen/Qwen3-30B-A3B-Instruct-2507
Contributions are not open yet, so there is nothing here from members.
No downloadable artifacts recorded.
Variant CantoneseLLM v2.0 30B A3B ThinkingPermissive
- What it is
- A community or third-party fine-tune of another variant.
- Publisher
- hon9kon9ize
- License
- Apache License 2.0 · commercial use allowed
- Released
- 1 Aug 2026
Contributions are not open yet, so there is nothing here from members.
No downloadable artifacts recorded.
Variant Qwen3 30B A3B ShapleyMCG K34 Validation ReconstructionPermissive
- What it is
- A community or third-party fine-tune of another variant.
- Publisher
- brandonmusic
- License
- Apache License 2.0 · commercial use allowed
- Released
- 24 Aug 2026
Contributions are not open yet, so there is nothing here from members.
No downloadable artifacts recorded.
Architecture details
- Parameters
- 30.53B
- Active / token
- 3.3B
- Architecture
- Mixture of experts
- Layers
- 48
- Attention heads
- 32
- KV heads
- 4
- Head dim
- 128
- KV cache @ 8K (fp16)
- 0.75 GiB
- Max context
- 32,768 tokens
Measured performance
Throughput on specific systems and runtimes, with the source of each measurement.
| Quantization | System | Runtime | Context | Measurements | Source |
|---|---|---|---|---|---|
| Qwen3 30B-A3B Q4_K_M (Unsloth AI) | RTX 4090 workstation (64 GB DDR5) | llama.cppcuda · b5300 | 4K | 3,400 tok/s · Prompt processing (512) 152 tok/s · Generation (128) | Mutinai illustrative fixtures |
| Qwen3 30B-A3B Q4_K_M (Unsloth AI) | Ryzen AI Max+ 395 mini PC (128 GB) | llama.cppvulkan · b5300 | 4K | 420 tok/s · Prompt processing (512) 51 tok/s · Generation (128) | Mutinai illustrative fixtures |
Community results
Reviews
Lineage
How this model’s variants relate to each other and to other models.
Sources & history
Sources
- Hugging Face Hub (live source, 1 record, 13 Sept 2026)
- Mutinai illustrative fixtures (illustrative fixture, 1 record, 13 Sept 2026)
External identifiers
- huggingface Qwen/Qwen3-30B-A3B-Base
- huggingface Qwen/Qwen3-30B-A3B
- huggingface Qwen/Qwen3-30B-A3B-Instruct-2507
- huggingface hon9kon9ize/CantoneseLLM-v2.0-30B-A3B-Thinking
- huggingface brandonmusic/Qwen3-30B-A3B-ShapleyMCG-K34-Validation-Reconstruction
Field history
- license = "apache-2.0" — Hugging Face Hub (current)
- auto_promotion = {"repo":"Qwen/Qwen3-30B-A3B-Instruct-2507","rule":"first_party_release","evidence":["publisher Qwen is Qwen Team (Alibaba Cloud), developer of Qwen, Qwen Coder, Qwen Math","repo name Qwen3-30B-A3B-Instruct-2507 belongs to family Qwen","config.json: 48 layers, 32 attention heads, 4 KV heads, head dim 128, context 262144","safetensors metadata: 30,532,122,624 parameters, consistent with 30B","128 experts (8 per token); A3B active parameters as named by the publisher","published 2025-07-28 (Hugging Face repository creation date)","variant kind instruct stated by the name"]} — Hugging Face Hub (current)
- base = "Qwen/Qwen3-30B-A3B-Base" — Hugging Face Hub (current)
- derivation = "fine_tune" — Hugging Face Hub (current)
- license = "apache-2.0" — Hugging Face Hub (current)
- name = "CantoneseLLM v2.0 30B A3B Thinking" — Hugging Face Hub (current)
- base = "Qwen/Qwen3-30B-A3B-Base" — Hugging Face Hub (current)
- derivation = "fine_tune" — Hugging Face Hub (current)
- license = "apache-2.0" — Hugging Face Hub (current)
- name = "Qwen3 30B A3B ShapleyMCG K34 Validation Reconstruction" — Hugging Face Hub (current)
Timeline
- Official Qwen3 GGUF quantizations published
- Qwen3 released with MoE and hybrid thinking