Hardware · SoC / APU · Apple
Apple M4 Max (40-core GPU)
Shares one pool of LPDDR5X memory between CPU and GPU at 546 GB/s. What fits depends on how much memory the system has.
- Memory
- UnifiedLPDDR5X
- Memory speed
- 546 GB/s#7 of 14 tracked
- Power
- —
- Launch price
- —8 Nov 2024
What it runs
Largest models without offloading on the MacBook Pro M4 Max 128 GB, 8K context.
- VariantLlama 3.1 70B Instruct70.6B · Q4_K_M via llama.cppRuns well~8 tok/s est.
- VariantLlama 3.3 70B Instruct70.6B · Q4_K_M via llama.cppRuns well~8 tok/s est.
- VariantMixtral 8x7B Instruct v0.146.7B · Q4_K_M via llama.cppRuns well~42 tok/s est.
- VariantQwen2.5 32B Instruct32.8B · Q5_K_M via llama.cppRuns well~15 tok/s est.
- VariantQwen2.5-Coder 32B Instruct32.8B · Q8_0 via llama.cppRuns well~10 tok/s est.
- VariantDeepSeek-R1-Distill-Qwen-32B32.8B · Q4_K_M via llama.cppRuns well~17 tok/s est.
- VariantQwen3 30B-A3B30.5B · Q8_0 via llama.cppRuns well~88 tok/s est.
- VariantGemma 3 27B IT27.4B · Q4_K_M via llama.cppRuns well~21 tok/s est.
- VariantGemma 2 27B IT27.2B · Q4_K_M via llama.cppRuns well~21 tok/s est.
- VariantMistral Small 24B Instruct 250123.6B · Q6_K via llama.cppRuns well~18 tok/s est.
Systems using it
- System MacBook Pro M4 Max 128 GBRuns models up to ~160B parameters at 4-bit in unified memory. What runs
Measured performance
Throughput on systems containing this device.
| Model · quantization | System | Runtime | Context | Measurements | Source |
|---|---|---|---|---|---|
| Llama 3.3 70B Instruct MLX 4-bit (MLX Community) | MacBook Pro M4 Max 128 GB | MLX-LMmetal · 0.21.0 | 4K | 11.2 tok/s · Generation throughput 41 GB · Peak memory 110 tok/s · Prompt throughput 4,200 ms · Time to first token | Mutinai illustrative fixtures |
Community results
Contributions are not open yet, so there is nothing here from members.
Reviews
Contributions are not open yet, so there is nothing here from members.
Specifications
- Vendor
- Apple
- Type
- SoC / APU
- Memory
- Unified (per system)
- Memory type
- LPDDR5X
- Bandwidth
- 546 GB/s
- GPU-usable share
- 75% by default
- Backends
- metal (Apple silicon), cpu (CPUs)
- TDP
- —
- Released
- 8 Nov 2024
- Launch price
- —
Sources & history
Sources
No source records linked.