Skip to content
mlx.app

Qwen2.5 14B Instruct

Alibaba · 14.7B · up to 32K context · apache-2.0

FITS · 21.2 GB FREE

Needs a 16GB Mac at minimum for the 4-bit build, and a 24GB or 32GB machine is where it actually feels comfortable with other apps open. It's noticeably more capable than the 7B on reasoning chains without yet needing a 32GB-class machine.

On a M4 Pro with 48 GB

macOS9.6 GBweights15.6 GBKV cache1.6 GBheadroom21.2 GB
Weights
15.6 GB
KV cache
1.6 GB
Usable RAM
38.4 GB
Headroom
21.2 GB
Generation
13.6 tok/sest.
Prompt processing
61.4 tok/sest.
First token (1K prompt)
16.7 sest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q48.3 GB9.9 GBFITS · 28.5 GB FREE25.8 tok/sQwen2.5-14B-Instruct-4bit
q815.6 GB17.2 GBFITS · 21.2 GB FREE13.6 tok/sQwen2.5-14B-Instruct-8bit

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/Qwen2.5-14B-Instruct-8bit
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Qwen2.5-14B-Instruct-8bit --port 8080

How far can you push the context

2K tokensFITS · 22.4 GB FREE

cache 0.4 GB

8K tokensFITS · 21.2 GB FREE

cache 1.6 GB

32K tokensFITS · 16.3 GB FREE

cache 6.4 GB

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M1no configuration
M1 Pro32 GB9.7 tok/s
M1 Max32 GB20.2 tok/s
M1 Ultra64 GB41.5 tok/s
M224 GB4.7 tok/s
M2 Pro32 GB9.9 tok/s
M2 Max32 GB20.5 tok/s
M2 Ultra64 GB42.0 tok/s
M324 GB4.7 tok/s
M3 Pro36 GB7.2 tok/s
M3 Max (14-core CPU)36 GB15.0 tok/s
M3 Max (16-core CPU)48 GB20.5 tok/s
M3 Ultra96 GB43.5 tok/s
M424 GB5.7 tok/s
M4 Pro24 GB13.6 tok/s
M4 Max (14-core CPU)36 GB21.3 tok/s
M4 Max (16-core CPU)48 GB28.7 tok/s
M524 GB7.3 tok/s
M5 Prounverified24 GB15.5 tok/s
M5 Maxunverified36 GB32.6 tok/s