Skip to content
mlx.app

Qwen2.5 72B Instruct

Alibaba · 72.7B · up to 32K context · qwen

OVER BY 5.2 GB

This needs a 64GB Mac to run 4-bit with any breathing room, and a 48GB machine will be tight once the OS takes its share. It's a genuine step up in world knowledge and reasoning over the 32B tier, which is the only reason to pay for that much unified memory.

On a M4 Pro with 48 GB

macOS9.6 GBweights40.9 GBKV cache2.7 GBover5.2 GB
Weights
40.9 GB
KV cache
2.7 GB
Usable RAM
38.4 GB
Headroom
0.0 GB
Generation
5.2 tok/sest.
Prompt processing
12.4 tok/sest.
First token (1K prompt)
82.4 sest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q440.9 GB43.6 GBOVER BY 5.2 GB5.2 tok/sQwen2.5-72B-Instruct-4bit
q877.2 GB79.9 GBOVER BY 41.5 GB2.8 tok/sQwen2.5-72B-Instruct-8bit

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/Qwen2.5-72B-Instruct-4bit
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Qwen2.5-72B-Instruct-4bit --port 8080

How far can you push the context

2K tokensOVER BY 3.2 GB

cache 0.7 GB

8K tokensOVER BY 5.2 GB

cache 2.7 GB

32K tokensOVER BY 13.2 GB

cache 10.7 GB

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M1no configuration
M1 Prono configuration
M1 Max64 GB7.7 tok/s
M1 Ultra64 GB15.8 tok/s
M2no configuration
M2 Prono configuration
M2 Max64 GB7.8 tok/s
M2 Ultra64 GB16.0 tok/s
M3no configuration
M3 Prono configuration
M3 Max (14-core CPU)no configuration
M3 Max (16-core CPU)64 GB7.8 tok/s
M3 Ultra96 GB16.6 tok/s
M4no configuration
M4 Pro64 GB5.2 tok/s
M4 Max (14-core CPU)no configuration
M4 Max (16-core CPU)64 GB10.9 tok/s
M5no configuration
M5 Prounverified64 GB5.9 tok/s
M5 Maxunverified64 GB12.5 tok/s