Public preview · figures are illustrative fixture data, attributed to their source · data to May 2025 · read-only
System · desktop

RTX 4060 Ti 16GB budget build (32 GB DDR5)

Runs models up to ~22B parameters at 4-bit entirely on the GPU. Entry-level 16 GB GPU system.

GPU memory
16 GB
System RAM
32 GB
Runs
8 of 19model variants, 8K
Estimated cost
~$1,400

What it runs

Largest models that fit without spilling into system memory, at an 8K contextHow much text the model can consider at once, counted in tokens — roughly ¾ of a word each..

Components

Measured performance

Model · quantizationRuntimeContextMeasurementsSource
Qwen2.5 14B Instruct Q4_K_Mllama.cppcuda · b46004K
1,250 tok/s · Prompt processing (512)
25.5 tok/s · Generation (128)
Mutinai illustrative fixtures

Community results

Contributions are not open yet, so there is nothing here from members.

Reviews

Contributions are not open yet, so there is nothing here from members.