Public preview · figures are illustrative fixture data, attributed to their source · data to May 2025 · read-only
Member

quietmodel

@quietmodel · joined 13 Sept 2026 · 1 public runs · 1 public reviews
Sample content. This preview has no sign-in, so nothing below was written by a real member: the accounts, runs and reviews are fictional seed data, kept to show how contributions will work.

Benchmark runs

54.2 tok/s generation throughput
Qwen3 30B-A3B MLX 4-bit
on Apple M4 Pro (20-core GPU) (member system, 48 GB unified) · MLX-LM 0.24.0 · metal
@quietmodel · 3 May 2025 unverified

Thinking disabled.

Full environment and measurements
Benchmark
Interactive throughput
Generation throughput
54.2 tok/s
Peak memory
17.8 GB
Prompt throughput
520 tok/s
Time to first token
610 ms
Artifact
Qwen3 30B-A3B MLX 4-bit (MLX Community)
Context
8K
OS
macOS 15.4
Helpful 1

Reviews

MoE makes a 48 GB Mac feel fast

@quietmodel on Qwen3 30B-A3B ·

Runs at interactive speeds in MLX 4-bit. Thinking mode helps on multi-step problems; turn it off for quick chat.

Coding4/5
Reasoning4/5
Agents & tool use4/5
Hardware efficiency5/5
Helpful 0