Public preview · figures are illustrative fixture data, attributed to their source · data to May 2025 · read-only
System · laptop

MacBook Pro M4 Max 128 GB

Runs models up to ~160B parameters at 4-bit in unified memory. Laptop with 128 GB unified memory.

Unified memory
128 GB
System RAM
Shared
Runs
18 of 19model variants, 8K
Estimated cost
~$5,400

What it runs

Largest models that fit without spilling into system memory, at an 8K contextHow much text the model can consider at once, counted in tokens — roughly ¾ of a word each..

Components

Measured performance

Model · quantizationRuntimeContextMeasurementsSource
Llama 3.3 70B Instruct MLX 4-bit (MLX Community)MLX-LMmetal · 0.21.04K
11.2 tok/s · Generation throughput
41 GB · Peak memory
110 tok/s · Prompt throughput
4,200 ms · Time to first token
Mutinai illustrative fixtures

Community results

Contributions are not open yet, so there is nothing here from members.

Reviews

Contributions are not open yet, so there is nothing here from members.