Public preview · figures are illustrative fixture data, attributed to their source · data to May 2025 · read-only
Member

Tokenwright

Local coding agents on a single 4090.

@tokenwright · joined 13 Sept 2026 · 1 public runs · 2 public reviews
Sample content. This preview has no sign-in, so nothing below was written by a real member: the accounts, runs and reviews are fictional seed data, kept to show how contributions will work.

Benchmark runs

131 tok/s generation (128)
@tokenwright · 10 Feb 2025 verified
Full environment and measurements
Benchmark
llama-bench
Prompt processing (512)
11,850 tok/s
Generation (128)
131 tok/s
Artifact
Llama 3.1 8B Instruct Q4_K_M (LM Studio Community)
Context
4K
Batch size
512
KV cache
f16
Flash attention
on
OS
Ubuntu 24.04
Driver
565.77
Helpful 0

Reviews

Best local coding model I have used at 24 GB

Q4_K_M with 16K context fits on the 4090 with room to spare. Excellent at targeted edits in Aider; weaker at long agentic loops where it sometimes loses the plan.

Output quality4/5
Coding5/5
Reasoning3/5
Agents & tool use3/5
Hardware efficiency4/5
Helpful 2

Easy setup, fewer knobs

@tokenwright on Ollama ·

Great for getting teammates started. I switch to llama.cpp directly when I need specific KV cache or offload settings.

Speed3/5
Reliability4/5
Ease of setup5/5
Helpful 0