Skip to content
mlx.app

Qwen2.5-VL 32B Instruct

Alibaba · 33B · up to 125K context · apache-2.0

FITS · 17.7 GB FREE

The 32GB-Mac step up from the 7B vision model, with meaningfully better reasoning about multi-image and multi-step visual tasks. It's slower per image than the 7B, so batch document processing is where the extra memory cost pays off most.

On a M4 Pro with 48 GB

macOS9.6 GBweights18.6 GBKV cache2.1 GBheadroom17.7 GB
Weights
18.6 GB
KV cache
2.1 GB
Usable RAM
38.4 GB
Headroom
17.7 GB
Generation
11.5 tok/sest.
Prompt processing
27.4 tok/sest.
First token (1K prompt)
37.4 sest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q418.6 GB20.7 GBFITS · 17.7 GB FREE11.5 tok/sQwen2.5-VL-32B-Instruct-4bit
q835.1 GB37.2 GBTIGHT · 1.2 GB FREE6.1 tok/sQwen2.5-VL-32B-Instruct-8bit

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/Qwen2.5-VL-32B-Instruct-4bit
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Qwen2.5-VL-32B-Instruct-4bit --port 8080

How far can you push the context

2K tokensFITS · 19.3 GB FREE

cache 0.5 GB

8K tokensFITS · 17.7 GB FREE

cache 2.1 GB

32K tokensFITS · 11.2 GB FREE

cache 8.6 GB

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M1no configuration
M1 Pro32 GB8.2 tok/s
M1 Max32 GB17.0 tok/s
M1 Ultra64 GB34.9 tok/s
M2no configuration
M2 Pro32 GB8.3 tok/s
M2 Max32 GB17.2 tok/s
M2 Ultra64 GB35.3 tok/s
M3no configuration
M3 Pro36 GB6.1 tok/s
M3 Max (14-core CPU)36 GB12.6 tok/s
M3 Max (16-core CPU)48 GB17.2 tok/s
M3 Ultra96 GB36.6 tok/s
M432 GB4.8 tok/s
M4 Pro48 GB11.5 tok/s
M4 Max (14-core CPU)36 GB17.9 tok/s
M4 Max (16-core CPU)48 GB24.1 tok/s
M532 GB6.2 tok/s
M5 Prounverified48 GB13.1 tok/s
M5 Maxunverified36 GB27.5 tok/s