MacBook Air M1 8 GB
Change hardware
Not sure what you have? Help me check
On a Mac
Click the Apple logo in the top-left corner of the screen, then About This Mac. Note the Chip (M1, M2, M3, M4…) and the Memory (8 GB, 16 GB…).
On Windows
Press Ctrl + Shift + Esc to open Task Manager, then Performance. Memory shows your RAM. If a GPU entry lists Dedicated GPU memory of 4 GB or more, you have a graphics card; otherwise you don’t.
Then pick the closest match — or the Build your own setup option below for anything else:
Still unsure? Start with the 16 GB laptop: what runs there runs on almost any recent computer.
Laptops
Desktops
Servers
Build your own setup
Speeds marked measured come from reference results and verified community runs on this exact system; others are estimates.
Qwen2.5 7B Instruct
Compact general-purpose model for coding and tool use. From Qwen Team (Alibaba Cloud).
- Speed
- Types faster than you can read estimate~9 tok/s estimate
- Download
- 4.0 GiBMLX 4-bit
- Memory
- 5.1 of 5.4 GB estimatea tight fit: close other apps first
- License
- PermissiveApache License 2.0
Why this one: a general chat model that fits this machine without spilling out of memory, answers faster than you read, and has the strongest benchmark results among those that do.
Now run it
Easiest: LM Studio free desktop app, no terminal
- Download LM Studio from lmstudio.ai, install it and open it.
- Open the search tab (the magnifying glass) and search for
mlx-community/Qwen2.5-7B-Instruct-4bit. - Choose the
MLX 4-bitdownload (4.0 GiB) and click Download. - Go to the chat tab, pick Qwen2.5 7B Instruct at the top, and type your first message.
Everything that runs here grouped by how it runs
Runs in accelerator memory 2
Fits entirely in GPU or unified memory — the fast path.
| Model | Recommended download | Memory | Generation speed |
|---|---|---|---|
| Variant Qwen2.5 7B InstructQwen Team (Alibaba Cloud) · 7.62B | MLX 4-bit via MLX-LM4.0 GiB · MLX CommunityTight fit | 5.1 of 5.4 GiB estimate | ~9 tok/s estimate |
| Variant DeepSeek-R1-Distill-Qwen-7BQwen Team (Alibaba Cloud) · 7.62B | Q4_K_M via llama.cpp4.3 GiB · LM Studio CommunityTight fit | 5.2 of 5.4 GiB estimate | ~9 tok/s estimate |
Too large for this hardware
- DeepSeek-R1 671B · 3 downloads checked
- Llama 3.1 70B Instruct 70.6B · 3 downloads checked
- Llama 3.3 70B Instruct 70.6B · 5 downloads checked
- Mixtral 8x7B Instruct v0.1 46.7B · 2 downloads checked
- Qwen2.5 32B Instruct 32.8B · 5 downloads checked
- Qwen2.5-Coder 32B Instruct 32.8B · 4 downloads checked
- DeepSeek-R1-Distill-Qwen-32B 32.8B · 3 downloads checked
- Qwen3 30B-A3B 30.5B · 8 downloads checked
- Gemma 3 27B IT 27.4B · 3 downloads checked
- Gemma 2 27B IT 27.2B · 2 downloads checked
- Mistral Small 24B Instruct 2501 23.6B · 4 downloads checked
- Qwen2.5 14B Instruct 14.8B · 5 downloads checked
- Phi-4 14.7B · 4 downloads checked
- Gemma 2 9B IT 9.24B · 2 downloads checked
- Hermes 3 Llama 3.1 8B 8.03B · 3 downloads checked
- Llama 3.1 8B Instruct 8.03B · 5 downloads checked
- Mistral 7B Instruct v0.3 7.25B · 3 downloads checked
How this is calculated
Memory = download size + fp16 KV cache for 8K tokens + runtime overhead. Usable memory is 95% of dedicated VRAM, a device-specific share of unified memory (75% by default), and 80% of system RAM for runtimes that can offload. Runtimes must load the file format and support a backend present on the hardware.
Estimated speed is bounded by memory bandwidth ÷ bytes read per token (active parameters for mixture-of-experts), shown in italics with an estimate label and a ±35% range (hover for it). Measured speeds are medians of reference results and verified public community runs on the same system, download and runtime. For each model we recommend the highest-precision download that fits, preferring not to offload.