Public preview · figures are illustrative fixture data, attributed to their source · data to May 2025 · read-only
Community

What people are running, measuring and recommending.

Runs record exactly what was run and on what, so results can be reproduced. Reviews rate specific qualities, not a single score.

Sample content. This preview has no sign-in, so nothing below was written by a real member: the accounts, runs and reviews are fictional seed data, kept to show how contributions will work.
4members
5public runs
3verified
6reviews

Recent benchmark runs

61.5 tok/s generation (128)
on MacBook Pro M3 Max 64 GB · llama.cpp b5310 · metal
@kestrel · 4 May 2025 verified
Full environment and measurements
Benchmark
llama-bench
Prompt processing (512)
485 tok/s
Generation (128)
61.5 tok/s
Artifact
Qwen3 30B-A3B Q4_K_M (Unsloth AI)
Context
4K
Flash attention
on
OS
macOS 15.4
Helpful 0
54.2 tok/s generation throughput
Qwen3 30B-A3B MLX 4-bit
on Apple M4 Pro (20-core GPU) (member system, 48 GB unified) · MLX-LM 0.24.0 · metal
@quietmodel · 3 May 2025 unverified

Thinking disabled.

Full environment and measurements
Benchmark
Interactive throughput
Generation throughput
54.2 tok/s
Peak memory
17.8 GB
Prompt throughput
520 tok/s
Time to first token
610 ms
Artifact
Qwen3 30B-A3B MLX 4-bit (MLX Community)
Context
8K
OS
macOS 15.4
Helpful 1
36.5 tok/s generation throughput
on AMD Ryzen 9 7950X + NVIDIA GeForce RTX 3090 (Triple 3090 rack, 128 GB RAM) · vLLM 0.7.2 · cuda
@basement-cluster · 20 Feb 2025 unverified
Full environment and measurements
Benchmark
Interactive throughput
Generation throughput
36.5 tok/s
Peak memory
44 GB
Prompt throughput
1,780 tok/s
Time to first token
240 ms
Artifact
Qwen2.5 32B Instruct AWQ 4-bit
Context
16K
OS
Debian 12
Driver
560.35
Parameters
tensor_parallel_size=2
Helpful 0
131 tok/s generation (128)
@tokenwright · 10 Feb 2025 verified
Full environment and measurements
Benchmark
llama-bench
Prompt processing (512)
11,850 tok/s
Generation (128)
131 tok/s
Artifact
Llama 3.1 8B Instruct Q4_K_M (LM Studio Community)
Context
4K
Batch size
512
KV cache
f16
Flash attention
on
OS
Ubuntu 24.04
Driver
565.77
Helpful 0
17.1 tok/s generation (128)
@basement-cluster · 5 Jan 2025 verified

Both cards power limited to 280W.

Full environment and measurements
Benchmark
llama-bench
Prompt processing (512)
372 tok/s
Generation (128)
17.1 tok/s
Artifact
Llama 3.3 70B Instruct Q4_K_M (LM Studio Community)
Context
4K
KV cache
f16
Flash attention
on
OS
Debian 12
Driver
560.35
Parameters
split_mode=layer power_limit_w=280
Helpful 2

Recent reviews

Best local coding model I have used at 24 GB

Q4_K_M with 16K context fits on the 4090 with room to spare. Excellent at targeted edits in Aider; weaker at long agentic loops where it sometimes loses the plan.

Output quality4/5
Coding5/5
Reasoning3/5
Agents & tool use3/5
Hardware efficiency4/5
Helpful 2

The foundation everything else runs on

@kestrel on llama.cpp ·

Fastest path to new model support and runs on everything. The flags surface is huge; defaults change between builds, so pin versions when benchmarking.

Output quality5/5
Speed5/5
Reliability4/5
Ease of setup3/5
Helpful 1

Still the value pick for VRAM per dollar

Bought three used. Power limits at 280W lose very little generation speed. Budget time for riser cables and airflow.

Speed4/5
Reliability4/5
Ease of setup3/5
Value5/5
Helpful 0

Easy setup, fewer knobs

@tokenwright on Ollama ·

Great for getting teammates started. I switch to llama.cpp directly when I need specific KV cache or offload settings.

Speed3/5
Reliability4/5
Ease of setup5/5
Helpful 0

MoE makes a 48 GB Mac feel fast

@quietmodel on Qwen3 30B-A3B ·

Runs at interactive speeds in MLX 4-bit. Thinking mode helps on multi-step problems; turn it off for quick chat.

Coding4/5
Reasoning4/5
Agents & tool use4/5
Hardware efficiency5/5
Helpful 0

How contributions work

VerificationWhy some runs are marked verified

New runs are unverified. Moderators verify runs whose numbers are consistent with the hardware and settings. Only verified, public runs feed the measured speeds in What can I run?.

Visibility and privacyPublic, unlisted and private

Everything you contribute is public, unlisted (reachable by link, excluded from listings and averages) or private (only you). Public runs on your own systems show the hardware components, not your system’s name. Contributions are never used for AI training unless you opt in. Read the principles.