Models · 25 tracked

Pick a model by what you have and what you want it for.

Every model here can be downloaded and run on your own hardware. The number that matters most is the memory it needs, so the list is grouped by it.

Every model, by what it needs to run18 of 25

Memory at 4-bit, the usual way people run models locally, with room for a working conversation.

Filters1
Reset

Most laptops

up to 6 GB

Runs on most laptops, even without a graphics card.

  • Compact general-purpose model for coding and tool use.Qwen Team (Alibaba Cloud) · Runs well on most reference systems
    ~6 GB estimateto runPermissive
  • Compact general-purpose model for coding and tool use.Meta · Runs well on most reference systems
    ~6 GB estimateto runPermissive
  • Compact chat model that can call tools.Mistral AI · Runs well on most reference systems
    ~6 GB estimateto runPermissive

A 16 GB graphics card

6–15 GB

Or a Mac with 24 GB of memory.

  • Chat model that can call tools.Mistral AI · Runs well on many reference systems
    ~15 GB estimateto runPermissive
  • General-purpose model for coding and reasoning.Microsoft · Runs well on most reference systems
    ~10 GB estimateto runPermissive
  • General-purpose model for coding and tool use.Qwen Team (Alibaba Cloud) · Runs well on most reference systems
    ~10 GB estimateto runPermissive
  • Compact model for everyday chat.Google DeepMind · Runs well on most reference systems
    ~9 GB estimateto runRestricted

A 24 GB graphics card

15–22 GB

Or a Mac with 32 GB of memory.

  • Efficient general-purpose model for coding, reasoning and tool use.Qwen Team (Alibaba Cloud) · Runs well on many reference systems
    ~18 GB estimateto runPermissive
  • Chat model that also understands images. Among the strongest open models at following instructions.Google DeepMind · Runs well on many reference systems
    ~20 GB estimateto runRestricted
  • Model for coding. Among the strongest open models at coding.Qwen Team (Alibaba Cloud) · Runs well on many reference systems
    ~21 GB estimateto runPermissive
  • General-purpose model for coding, reasoning and tool use. Among the strongest open models at reasoning.Qwen Team (Alibaba Cloud) · Runs well on many reference systems
    ~21 GB estimateto runPermissive
  • Model for everyday chat.Google DeepMind · Runs well on many reference systems
    ~20 GB estimateto runRestricted

A 32 GB graphics card

22–30 GB

Or a Mac with 48 GB of memory.

  • Efficient model for everyday chat.Mistral AI · Runs well on many reference systems
    ~30 GB estimateto runPermissive

A 64 GB Mac or two graphics cards

30–46 GB

Where 70B-class models start to fit.

  • Chat model that can call tools. Among the strongest open models at following instructions.Meta · Runs well only on high-memory systems
    ~37 GB estimateto runRestricted
  • Chat model that can call tools.Meta · Runs well only on high-memory systems
    ~42 GB estimateto runRestricted

Server-class memory

over 92 GB

Datacenter accelerators and multi-GPU servers.

  • Very large general-purpose model for coding and reasoning. Among the strongest open models at coding.DeepSeek · Runs only slowly on reference systems
    ~169 GB estimateto runPermissive

Memory not recorded yet

No download with a known size is on file, so what these need is not estimated.

  • Compact model for everyday chat.Qwen Team (Alibaba Cloud)
    Permissive
  • Very large model for everyday chat.Qwen Team (Alibaba Cloud)
    Permissive

Parameters, context length, capability profiles and side-by-side comparison are in the Technical view.