Public preview · figures are illustrative fixture data, attributed to their source · data to May 2025 · read-only
System · desktop

RTX 4090 workstation (64 GB DDR5)

Runs models up to ~35B parameters at 4-bit entirely on the GPU. Single 24 GB GPU desktop.

GPU memory
24 GB
System RAM
64 GB
Runs
15 of 19model variants, 8K
Estimated cost
~$3,200

What it runs

Largest models that fit without spilling into system memory, at an 8K contextHow much text the model can consider at once, counted in tokens — roughly ¾ of a word each..

Components

Measured performance

Model · quantizationRuntimeContextMeasurementsSource
Llama 3.1 8B Instruct Q4_K_M (LM Studio Community)llama.cppcuda · b46004K
12,100 tok/s · Prompt processing (512)
128 tok/s · Generation (128)
Mutinai illustrative fixtures
Qwen3 30B-A3B Q4_K_M (Unsloth AI)llama.cppcuda · b53004K
3,400 tok/s · Prompt processing (512)
152 tok/s · Generation (128)
Mutinai illustrative fixtures

Community results

Contributions are not open yet, so there is nothing here from members.

Reviews

Contributions are not open yet, so there is nothing here from members.