Public preview · figures are illustrative fixture data, attributed to their source · data to May 2025 · read-only
Model · Llama 3.3 · December 2024 Live · Hugging Face

Llama 3.3 70B

A 71 billion-parameterHow many numbers the model learned during training — the usual rough measure of its size. dense model from Meta. A 70B instruct model with quality close to Llama 3.1 405B.

Size
70.6Bdense
Memory to run
~37 GBsmallest, 8K ctx
Context
128Ktokens
License
RestrictedLlama 3.3 Community License
Updated
18 Sept 2026

Which version to use

Where it runs

Best-fitting version on each reference system, at an 8K contextHow much text the model can consider at once, counted in tokens — roughly ¾ of a word each..

Runs well 7

Slowly (CPU or offload) 5

Too large 1

Benchmarks

Each benchmarkA fixed set of questions every model is given, so their scores can be compared on the same task. is developer-reported; prompts and settings differ between labs. Bars are relative to the best open result.

BenchmarkLlama 3.3 70B Instruct
Benchmark GPQA DiamondAccuracy
50.5src
Benchmark HumanEvalpass@1
88.4src
Benchmark IFEvalStrict prompt accuracy
92.1src
Benchmark MMLU-ProAccuracy
68.9src

developer reported

Variants & downloads

Each variant is a separate set of weights. Expand one for its downloads: quantizationStoring each of the model’s numbers with fewer bits, so the file is smaller and needs less memory. trades a little quality for a much smaller file, and the memory column adds the working memoryExtra space the model needs while it answers. It grows with the length of the conversation, on top of the file itself. a conversation needs on top.

Variant Llama 3.3 70B Instructinstruct · Meta · 5 downloads · Llama 3.3 Community LicenseRestricted
What it is
Tuned to follow instructions and hold a conversation. The usual choice.
Publisher
Meta
License
Llama 3.3 Community License · commercial use restricted
Released
6 Dec 2024

Contributions are not open yet, so there is nothing here from members.

QuantizationFormatBitsDownloadMemory @ 8KPublisher
Download BF16nativesafetensors16131 GiB~139.7 GiBMetameta-llama/Llama-3.3-70B-Instruct
Download Q4_K_Mk quantgguf4.8940.2 GiB~44.8 GiBLM Studio Communitylmstudio-community/Llama-3.3-70B-Instruct-GGUF
Download AWQ 4-bitawqsafetensors4.637.8 GiB~42.3 GiBUnsloth AIunsloth/Llama-3.3-70B-Instruct-AWQ
Download MLX 4-bitmlxmlx4.537.0 GiB~41.4 GiBMLX Communitymlx-community/Llama-3.3-70B-Instruct-4bit
Download Q3_K_Mk quantgguf3.9132.1 GiB~36.4 GiBUnsloth AIunsloth/Llama-3.3-70B-Instruct-GGUF
Architecture detailsLayers, attention and KV cache geometry
Parameters
70.55B
Active / token
All (dense)
Architecture
Dense
Layers
80
Attention heads
64
KV heads
8
Head dim
128
KV cache @ 8K (fp16)
2.50 GiB
Max context
131,072 tokens

Measured performance

Throughput on specific systems and runtimes, with the source of each measurement.

QuantizationSystemRuntimeContextMeasurementsSource
Llama 3.3 70B Instruct Q4_K_M (LM Studio Community)Dual RTX 3090 (128 GB DDR5)llama.cppcuda · b46004K
390 tok/s · Prompt processing (512)
17.6 tok/s · Generation (128)
Mutinai illustrative fixtures
Llama 3.3 70B Instruct Q4_K_M (LM Studio Community)MacBook Pro M3 Max 64 GBllama.cppmetal · b46004K
72 tok/s · Prompt processing (512)
7.4 tok/s · Generation (128)
Mutinai illustrative fixtures
Llama 3.3 70B Instruct MLX 4-bit (MLX Community)MacBook Pro M4 Max 128 GBMLX-LMmetal · 0.21.04K
11.2 tok/s · Generation throughput
41 GB · Peak memory
110 tok/s · Prompt throughput
4,200 ms · Time to first token
Mutinai illustrative fixtures
Llama 3.3 70B Instruct Q4_K_M (LM Studio Community)Mac Studio M2 Ultra 192 GBllama.cppmetal · b46004K
135 tok/s · Prompt processing (512)
12.1 tok/s · Generation (128)
Mutinai illustrative fixtures

Community results

Contributions are not open yet, so there is nothing here from members.

Reviews

Contributions are not open yet, so there is nothing here from members.

Lineage

How this model’s variants relate to each other and to other models.

VariantLlama 3.3 70B Instruct instruct

Sources & history

Sources

  • Hugging Face Hub (live source, 1 record, 18 Sept 2026)
  • Mutinai illustrative fixtures (illustrative fixture, 1 record, 13 Sept 2026)

External identifiers

Field history

  • license = "llama3.3" Hugging Face Hub (current)

Timeline

  1. Llama 3.3 70B released