What can I run? · 8K context

MacBook Air M1 8 GB

Change hardwareCurrently: MacBook Air M1 8 GB
Not sure what you have? Help me checkTwo numbers decide it: memory, and whether there is a graphics card

On a Mac

Click the Apple logo in the top-left corner of the screen, then About This Mac. Note the Chip (M1, M2, M3, M4…) and the Memory (8 GB, 16 GB…).

On Windows

Press Ctrl + Shift + Esc to open Task Manager, then Performance. Memory shows your RAM. If a GPU entry lists Dedicated GPU memory of 4 GB or more, you have a graphics card; otherwise you don’t.

Then pick the closest match — or the Build your own setup option below for anything else:

Still unsure? Start with the 16 GB laptop: what runs there runs on almost any recent computer.

Build your own setupChoose GPUs or a chip and how much memory you have
2run well
0with offload
17too large

Speeds marked measured come from reference results and verified community runs on this exact system; others are estimates.

Start with this one

Qwen2.5 7B Instruct

Compact general-purpose model for coding and tool use. From Qwen Team (Alibaba Cloud).

Speed
Types faster than you can read estimate~9 tok/s estimate
Download
4.0 GiBMLX 4-bit
Memory
5.1 of 5.4 GB estimatea tight fit: close other apps first
License
PermissiveApache License 2.0

Why this one: a general chat model that fits this machine without spilling out of memory, answers faster than you read, and has the strongest benchmark results among those that do.

Now run it

Easiest: LM Studio free desktop app, no terminal

  1. Download LM Studio from lmstudio.ai, install it and open it.
  2. Open the search tab (the magnifying glass) and search for mlx-community/Qwen2.5-7B-Instruct-4bit.
  3. Choose the MLX 4-bit download (4.0 GiB) and click Download.
  4. Go to the chat tab, pick Qwen2.5 7B Instruct at the top, and type your first message.

Everything that runs here grouped by how it runs

Runs in accelerator memory 2

Fits entirely in GPU or unified memory — the fast path.

ModelRecommended downloadMemoryGeneration speed
Variant Qwen2.5 7B InstructQwen Team (Alibaba Cloud) · 7.62BMLX 4-bit via MLX-LM4.0 GiB · MLX CommunityTight fit
5.1 of 5.4 GiB estimate
~9 tok/s estimate
Variant DeepSeek-R1-Distill-Qwen-7BQwen Team (Alibaba Cloud) · 7.62BQ4_K_M via llama.cpp4.3 GiB · LM Studio CommunityTight fit
5.2 of 5.4 GiB estimate
~9 tok/s estimate
Too large for this hardware17 model variants
How this is calculatedMemory, fit and speed estimates

Memory = download size + fp16 KV cache for 8K tokens + runtime overhead. Usable memory is 95% of dedicated VRAM, a device-specific share of unified memory (75% by default), and 80% of system RAM for runtimes that can offload. Runtimes must load the file format and support a backend present on the hardware.

Estimated speed is bounded by memory bandwidth ÷ bytes read per token (active parameters for mixture-of-experts), shown in italics with an estimate label and a ±35% range (hover for it). Measured speeds are medians of reference results and verified public community runs on the same system, download and runtime. For each model we recommend the highest-precision download that fits, preferring not to offload.