Skip to content
mlx.app

Llama 3.3 70B Instruct

Meta · 70.6B · up to 128K context · llama-3.3

OVER BY 4.0 GB

Meta tuned this to match Llama 3.1 405B on many benchmarks at a sixth of the size, and a 64GB Mac is the realistic floor for the 4-bit build. It's a strong generalist but the 80-layer depth makes it noticeably slower token-per-second than the 32B tier on the same hardware.

On a M4 Pro with 48 GB

macOS9.6 GBweights39.7 GBKV cache2.7 GBover4.0 GB
Weights
39.7 GB
KV cache
2.7 GB
Usable RAM
38.4 GB
Headroom
0.0 GB
Generation
5.4 tok/sest.
Prompt processing
12.8 tok/sest.
First token (1K prompt)
80.1 sest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q439.7 GB42.4 GBOVER BY 4.0 GB5.4 tok/sLlama-3.3-70B-Instruct-4bit
q875.0 GB77.7 GBOVER BY 39.3 GB2.8 tok/sLlama-3.3-70B-Instruct-8bit

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/Llama-3.3-70B-Instruct-4bit
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Llama-3.3-70B-Instruct-4bit --port 8080

How far can you push the context

2K tokensOVER BY 2.0 GB

cache 0.7 GB

8K tokensOVER BY 4.0 GB

cache 2.7 GB

32K tokensOVER BY 12.0 GB

cache 10.7 GB

128K tokensOVER BY 44.3 GB

cache 42.9 GB

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M1no configuration
M1 Prono configuration
M1 Max64 GB8.0 tok/s
M1 Ultra64 GB16.3 tok/s
M2no configuration
M2 Prono configuration
M2 Max64 GB8.1 tok/s
M2 Ultra64 GB16.5 tok/s
M3no configuration
M3 Prono configuration
M3 Max (14-core CPU)no configuration
M3 Max (16-core CPU)64 GB8.1 tok/s
M3 Ultra96 GB17.1 tok/s
M4no configuration
M4 Pro64 GB5.4 tok/s
M4 Max (14-core CPU)no configuration
M4 Max (16-core CPU)64 GB11.3 tok/s
M5no configuration
M5 Prounverified64 GB6.1 tok/s
M5 Maxunverified64 GB12.8 tok/s