Skip to content
mlx.app

Qwen2.5 7B Instruct

Alibaba · 7.6B · up to 32K context · apache-2.0

FITS · 22.7 GB FREE

The default general-purpose pick for a 16GB Mac running 4-bit, with enough room left for a browser tab or two. It's a known quantity: broad training data, reliable instruction following, nothing flashy.

On a M4 Pro with 48 GB

macOS9.6 GBweights15.2 GBKV cache0.5 GBheadroom22.7 GB
Weights
15.2 GB
KV cache
0.5 GB
Usable RAM
38.4 GB
Headroom
22.7 GB
Generation
14.0 tok/sest.
Prompt processing
119 tok/sest.
First token (1K prompt)
8.6 sest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q44.3 GB4.7 GBFITS · 33.7 GB FREE49.8 tok/sQwen2.5-7B-Instruct-4bit
q88.1 GB8.5 GBFITS · 29.9 GB FREE26.4 tok/sQwen2.5-7B-Instruct-8bit
bf1615.2 GB15.7 GBFITS · 22.7 GB FREE14.0 tok/sQwen2.5-7B-Instruct-bf16

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/Qwen2.5-7B-Instruct-bf16
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Qwen2.5-7B-Instruct-bf16 --port 8080

How far can you push the context

2K tokensFITS · 23.1 GB FREE

cache 0.1 GB

8K tokensFITS · 22.7 GB FREE

cache 0.5 GB

32K tokensFITS · 21.3 GB FREE

cache 1.9 GB

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M1no configuration
M1 Pro32 GB10.0 tok/s
M1 Max32 GB20.8 tok/s
M1 Ultra64 GB42.6 tok/s
M224 GB4.8 tok/s
M2 Pro32 GB10.1 tok/s
M2 Max32 GB21.1 tok/s
M2 Ultra64 GB43.2 tok/s
M324 GB4.8 tok/s
M3 Pro36 GB7.4 tok/s
M3 Max (14-core CPU)36 GB15.4 tok/s
M3 Max (16-core CPU)48 GB21.1 tok/s
M3 Ultra96 GB44.7 tok/s
M424 GB5.8 tok/s
M4 Pro24 GB14.0 tok/s
M4 Max (14-core CPU)36 GB21.8 tok/s
M4 Max (16-core CPU)48 GB29.5 tok/s
M524 GB7.5 tok/s
M5 Prounverified24 GB16.0 tok/s
M5 Maxunverified36 GB33.5 tok/s