Hardware · GPU · NVIDIA
NVIDIA GeForce RTX 4060 Ti 16GB
16 GB of GDDR6 memory at 288 GB/s — enough for models up to ~22B parameters at 4-bit on one card.
- Memory
- 16 GBGDDR6
- Memory speed
- 288 GB/s#9 of 14 tracked
- Power
- 165 W
- Launch price
- $499$31 per GB
What it runs
Largest models without offloading on the RTX 4060 Ti 16GB budget build (32 GB DDR5), 8K context.
- VariantQwen2.5 14B Instruct14.8B · Q6_K via llama.cppTight fit~15 tok/s est.
- VariantPhi-414.7B · Q4_K_M via llama.cppRuns well~20 tok/s est.
- VariantGemma 2 9B IT9.24B · Q4_K_M via llama.cppRuns well~30 tok/s est.
- VariantHermes 3 Llama 3.1 8B8.03B · Q8_0 via llama.cppRuns well~21 tok/s est.
- VariantLlama 3.1 8B Instruct8.03B · Q8_0 via llama.cppRuns well~21 tok/s est.
- VariantQwen2.5 7B Instruct7.62B · Q8_0 via llama.cppRuns well~22 tok/s est.
- VariantDeepSeek-R1-Distill-Qwen-7B7.62B · Q4_K_M via llama.cppRuns well~36 tok/s est.
- VariantMistral 7B Instruct v0.37.25B · Q8_0 via llama.cppRuns well~23 tok/s est.
Systems using it
- System RTX 4060 Ti 16GB budget build (32 GB DDR5)Runs models up to ~22B parameters at 4-bit entirely on the GPU. What runs
Measured performance
Throughput on systems containing this device.
| Model · quantization | System | Runtime | Context | Measurements | Source |
|---|---|---|---|---|---|
| Qwen2.5 14B Instruct Q4_K_M | RTX 4060 Ti 16GB budget build (32 GB DDR5) | llama.cppcuda · b4600 | 4K | 1,250 tok/s · Prompt processing (512) 25.5 tok/s · Generation (128) | Mutinai illustrative fixtures |
Community results
Contributions are not open yet, so there is nothing here from members.
Reviews
Contributions are not open yet, so there is nothing here from members.
Specifications
- Vendor
- NVIDIA
- Type
- GPU
- Memory
- 16 GB
- Memory type
- GDDR6
- Bandwidth
- 288 GB/s
- GPU-usable share
- —
- Backends
- cuda (NVIDIA GPUs), vulkan (most GPUs (Vulkan))
- TDP
- 165 W
- Released
- 18 Jul 2023
- Launch price
- $499
Sources & history
Sources
No source records linked.