Public preview · figures are illustrative fixture data, attributed to their source · data to May 2025 · read-only
Member

Kestrel

Community moderator. Mostly Apple silicon benchmarks.

@kestrel · joined 13 Sept 2026 · 1 public runs · 2 public reviews
Sample content. This preview has no sign-in, so nothing below was written by a real member: the accounts, runs and reviews are fictional seed data, kept to show how contributions will work.

Benchmark runs

61.5 tok/s generation (128)
on MacBook Pro M3 Max 64 GB · llama.cpp b5310 · metal
@kestrel · 4 May 2025 verified
Full environment and measurements
Benchmark
llama-bench
Prompt processing (512)
485 tok/s
Generation (128)
61.5 tok/s
Artifact
Qwen3 30B-A3B Q4_K_M (Unsloth AI)
Context
4K
Flash attention
on
OS
macOS 15.4
Helpful 0

Reviews

The foundation everything else runs on

@kestrel on llama.cpp ·

Fastest path to new model support and runs on everything. The flags surface is huge; defaults change between builds, so pin versions when benchmarking.

Output quality5/5
Speed5/5
Reliability4/5
Ease of setup3/5
Helpful 1

Strong reasoning, verbose

Good on math and planning tasks. Produces long reasoning traces, so budget context and generation time accordingly.

Coding3/5
Reasoning4/5
Hardware efficiency3/5
Helpful 0