Hardware · GPU · NVIDIA
NVIDIA GeForce RTX 4090
24 GB of GDDR6X memory at 1,008 GB/s — enough for models up to ~35B parameters at 4-bit on one card.
- Memory
- 24 GBGDDR6X
- Memory speed
- 1,008 GB/s#3 of 14 tracked
- Power
- 450 W
- Launch price
- $1,599$67 per GB
What it runs
Largest models without offloading on the RTX 4090 workstation (64 GB DDR5), 8K context.
- VariantQwen2.5 32B Instruct32.8B · Q4_K_M via llama.cppTight fit~32 tok/s est.
- VariantQwen2.5-Coder 32B Instruct32.8B · Q4_K_M via llama.cppTight fit~32 tok/s est.
- VariantDeepSeek-R1-Distill-Qwen-32B32.8B · Q4_K_M via llama.cppTight fit~32 tok/s est.
- VariantQwen3 30B-A3B30.5B · Q4_K_M via llama.cppRuns well~258 tok/s est.
- VariantGemma 3 27B IT27.4B · Q4_K_M via llama.cppTight fit~38 tok/s est.
- VariantGemma 2 27B IT27.2B · Q4_K_M via llama.cppRuns well~38 tok/s est.
- VariantMistral Small 24B Instruct 250123.6B · Q6_K via llama.cppRuns well~33 tok/s est.
- VariantQwen2.5 14B Instruct14.8B · Q6_K via llama.cppRuns well~52 tok/s est.
- VariantPhi-414.7B · Q8_0 via llama.cppRuns well~41 tok/s est.
- VariantGemma 2 9B IT9.24B · BF16 via SGLangTight fit~34 tok/s est.
Systems using it
- System RTX 4090 workstation (64 GB DDR5)Runs models up to ~35B parameters at 4-bit entirely on the GPU. What runs
Measured performance
Throughput on systems containing this device.
| Model · quantization | System | Runtime | Context | Measurements | Source |
|---|---|---|---|---|---|
| Llama 3.1 8B Instruct Q4_K_M (LM Studio Community) | RTX 4090 workstation (64 GB DDR5) | llama.cppcuda · b4600 | 4K | 12,100 tok/s · Prompt processing (512) 128 tok/s · Generation (128) | Mutinai illustrative fixtures |
| Qwen3 30B-A3B Q4_K_M (Unsloth AI) | RTX 4090 workstation (64 GB DDR5) | llama.cppcuda · b5300 | 4K | 3,400 tok/s · Prompt processing (512) 152 tok/s · Generation (128) | Mutinai illustrative fixtures |
Community results
Contributions are not open yet, so there is nothing here from members.
Reviews
Contributions are not open yet, so there is nothing here from members.
Specifications
- Vendor
- NVIDIA
- Type
- GPU
- Memory
- 24 GB
- Memory type
- GDDR6X
- Bandwidth
- 1008 GB/s
- GPU-usable share
- —
- Backends
- cuda (NVIDIA GPUs), vulkan (most GPUs (Vulkan))
- TDP
- 450 W
- Released
- 12 Oct 2022
- Launch price
- $1,599
Successor to NVIDIA GeForce RTX 3090 · Succeeded by NVIDIA GeForce RTX 5090
Sources & history
Sources
- Mutinai illustrative fixtures (illustrative fixture, 1 record, 13 Sept 2026)