Skip to content
mlx.app

Qwen2.5 Coder 32B Instruct

Alibaba · 32.5B · up to 32K context · apache-2.0

FITS · 18.0 GB FREE

For a while this was the best local coding model available, and on a 32GB Mac it still holds up well against newer general models. Qwen3-Coder-30B-A3B now beats it on speed for similar quality, but this dense model tends to be more consistent on tricky edge cases.

On a M4 Pro with 48 GB

macOS9.6 GBweights18.3 GBKV cache2.1 GBheadroom18.0 GB
Weights
18.3 GB
KV cache
2.1 GB
Usable RAM
38.4 GB
Headroom
18.0 GB
Generation
11.6 tok/sest.
Prompt processing
27.8 tok/sest.
First token (1K prompt)
36.9 sest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q418.3 GB20.4 GBFITS · 18.0 GB FREE11.6 tok/sQwen2.5-Coder-32B-Instruct-4bit
q834.5 GB36.7 GBTIGHT · 1.7 GB FREE6.2 tok/sQwen2.5-Coder-32B-Instruct-8bit

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/Qwen2.5-Coder-32B-Instruct-4bit
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Qwen2.5-Coder-32B-Instruct-4bit --port 8080

How far can you push the context

2K tokensFITS · 19.6 GB FREE

cache 0.5 GB

8K tokensFITS · 18.0 GB FREE

cache 2.1 GB

32K tokensFITS · 11.5 GB FREE

cache 8.6 GB

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M1no configuration
M1 Pro32 GB8.3 tok/s
M1 Max32 GB17.3 tok/s
M1 Ultra64 GB35.4 tok/s
M2no configuration
M2 Pro32 GB8.4 tok/s
M2 Max32 GB17.5 tok/s
M2 Ultra64 GB35.9 tok/s
M3no configuration
M3 Pro36 GB6.2 tok/s
M3 Max (14-core CPU)36 GB12.8 tok/s
M3 Max (16-core CPU)48 GB17.5 tok/s
M3 Ultra96 GB37.2 tok/s
M432 GB4.9 tok/s
M4 Pro48 GB11.6 tok/s
M4 Max (14-core CPU)36 GB18.2 tok/s
M4 Max (16-core CPU)48 GB24.5 tok/s
M532 GB6.3 tok/s
M5 Prounverified48 GB13.3 tok/s
M5 Maxunverified36 GB27.9 tok/s