Skip to content
mlx.app

Llama 3.2 11B Vision Instruct

Meta · 10.7B · up to 128K context · llama-3.2

FITS · 17.0 GB FREE

Meta's cross-attention vision adapter bolted onto a Llama 3.1 8B backbone, fitting a 16GB Mac comfortably at 4-bit. It's solid for general photo description but noticeably weaker than Qwen2.5-VL on dense documents and charts.

On a M4 Pro with 48 GB

macOS9.6 GBweights21.4 GBKV cache0.0 GBheadroom17.0 GB
Weights
21.4 GB
KV cache
0.0 GBestimated shape
Usable RAM
38.4 GB
Headroom
17.0 GB
Generation
10.0 tok/sest.
Prompt processing
84.4 tok/sest.
First token (1K prompt)
12.1 sest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q46.0 GB6.0 GBFITS · 32.4 GB FREE35.4 tok/sLlama-3.2-11B-Vision-Instruct-4bit
bf1621.4 GB21.4 GBFITS · 17.0 GB FREE10.0 tok/sLlama-3.2-11B-Vision-Instruct

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/Llama-3.2-11B-Vision-Instruct
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Llama-3.2-11B-Vision-Instruct --port 8080

How far can you push the context

2K tokensFITS · 17.0 GB FREE

cache 0.0 GB

8K tokensFITS · 17.0 GB FREE

cache 0.0 GB

32K tokensFITS · 17.0 GB FREE

cache 0.0 GB

128K tokensFITS · 17.0 GB FREE

cache 0.0 GB

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M1no configuration
M1 Pro32 GB7.1 tok/s
M1 Max32 GB14.8 tok/s
M1 Ultra64 GB30.3 tok/s
M2no configuration
M2 Pro32 GB7.2 tok/s
M2 Max32 GB15.0 tok/s
M2 Ultra64 GB30.7 tok/s
M3no configuration
M3 Pro36 GB5.3 tok/s
M3 Max (14-core CPU)36 GB10.9 tok/s
M3 Max (16-core CPU)48 GB15.0 tok/s
M3 Ultra96 GB31.8 tok/s
M432 GB4.1 tok/s
M4 Pro48 GB10.0 tok/s
M4 Max (14-core CPU)36 GB15.5 tok/s
M4 Max (16-core CPU)48 GB20.9 tok/s
M532 GB5.4 tok/s
M5 Prounverified48 GB11.3 tok/s
M5 Maxunverified36 GB23.8 tok/s