Skip to content
mlx.app

Qwen2.5-VL 7B Instruct

Alibaba · 8.3B · up to 125K context · apache-2.0

FITS · 29.1 GB FREE

A capable everyday vision model that fits a 16GB Mac, able to read screenshots, charts, and dense documents rather than just describing photos. It also does basic video understanding and can point to bounding boxes, which most local vision models skip.

On a M4 Pro with 48 GB

macOS9.6 GBweights8.8 GBKV cache0.5 GBheadroom29.1 GB
Weights
8.8 GB
KV cache
0.5 GB
Usable RAM
38.4 GB
Headroom
29.1 GB
Generation
24.1 tok/sest.
Prompt processing
109 tok/sest.
First token (1K prompt)
9.4 sest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q44.7 GB5.1 GBFITS · 33.3 GB FREE45.6 tok/sQwen2.5-VL-7B-Instruct-4bit
q88.8 GB9.3 GBFITS · 29.1 GB FREE24.1 tok/sQwen2.5-VL-7B-Instruct-8bit

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/Qwen2.5-VL-7B-Instruct-8bit
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Qwen2.5-VL-7B-Instruct-8bit --port 8080

How far can you push the context

2K tokensFITS · 29.5 GB FREE

cache 0.1 GB

8K tokensFITS · 29.1 GB FREE

cache 0.5 GB

32K tokensFITS · 27.7 GB FREE

cache 1.9 GB

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M116 GB5.6 tok/s
M1 Pro16 GB17.2 tok/s
M1 Max32 GB35.8 tok/s
M1 Ultra64 GB73.5 tok/s
M216 GB8.3 tok/s
M2 Pro16 GB17.5 tok/s
M2 Max32 GB36.3 tok/s
M2 Ultra64 GB74.4 tok/s
M316 GB8.3 tok/s
M3 Pro18 GB12.8 tok/s
M3 Max (14-core CPU)36 GB26.5 tok/s
M3 Max (16-core CPU)48 GB36.3 tok/s
M3 Ultra96 GB77.1 tok/s
M416 GB10.1 tok/s
M4 Pro24 GB24.1 tok/s
M4 Max (14-core CPU)36 GB37.7 tok/s
M4 Max (16-core CPU)48 GB50.8 tok/s
M516 GB13.0 tok/s
M5 Prounverified24 GB27.5 tok/s
M5 Maxunverified36 GB57.8 tok/s