Public preview · figures are illustrative fixture data, attributed to their source · data to May 2025 · read-only
Member

Basement Cluster

Used 3090s and loud fans.

@basement-cluster · joined 13 Sept 2026 · 2 public runs · 1 public reviews
Sample content. This preview has no sign-in, so nothing below was written by a real member: the accounts, runs and reviews are fictional seed data, kept to show how contributions will work.

Benchmark runs

36.5 tok/s generation throughput
on AMD Ryzen 9 7950X + NVIDIA GeForce RTX 3090 (Triple 3090 rack, 128 GB RAM) · vLLM 0.7.2 · cuda
@basement-cluster · 20 Feb 2025 unverified
Full environment and measurements
Benchmark
Interactive throughput
Generation throughput
36.5 tok/s
Peak memory
44 GB
Prompt throughput
1,780 tok/s
Time to first token
240 ms
Artifact
Qwen2.5 32B Instruct AWQ 4-bit
Context
16K
OS
Debian 12
Driver
560.35
Parameters
tensor_parallel_size=2
Helpful 0
17.1 tok/s generation (128)
@basement-cluster · 5 Jan 2025 verified

Both cards power limited to 280W.

Full environment and measurements
Benchmark
llama-bench
Prompt processing (512)
372 tok/s
Generation (128)
17.1 tok/s
Artifact
Llama 3.3 70B Instruct Q4_K_M (LM Studio Community)
Context
4K
KV cache
f16
Flash attention
on
OS
Debian 12
Driver
560.35
Parameters
split_mode=layer power_limit_w=280
Helpful 2

Reviews

Still the value pick for VRAM per dollar

Bought three used. Power limits at 280W lose very little generation speed. Budget time for riser cables and airflow.

Speed4/5
Reliability4/5
Ease of setup3/5
Value5/5
Helpful 0