Llama 3.2 11B Vision Instruct
Meta · 10.7B · up to 128K context · llama-3.2
Meta's cross-attention vision adapter bolted onto a Llama 3.1 8B backbone, fitting a 16GB Mac comfortably at 4-bit. It's solid for general photo description but noticeably weaker than Qwen2.5-VL on dense documents and charts.
On a M4 Pro with 48 GB
macOS9.6 GBweights21.4 GBKV cache0.0 GBheadroom17.0 GB
- Weights
- 21.4 GB
- KV cache
- 0.0 GBestimated shape
- Usable RAM
- 38.4 GB
- Headroom
- 17.0 GB
- Generation
- 10.0 tok/sest.
- Prompt processing
- 84.4 tok/sest.
- First token (1K prompt)
- 12.1 sest.
Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.
Every quantization, on your Mac
| Quant | Weights | + KV | Verdict | Speed est. | Repo |
|---|---|---|---|---|---|
| q4 | 6.0 GB | 6.0 GB | FITS · 32.4 GB FREE | 35.4 tok/s | Llama-3.2-11B-Vision-Instruct-4bit |
| bf16 | 21.4 GB | 21.4 GB | FITS · 17.0 GB FREE | 10.0 tok/s | Llama-3.2-11B-Vision-Instruct |
We never host weights. Every link goes to Hugging Face.
Run it
Chat
mlx_lm.chat --model mlx-community/Llama-3.2-11B-Vision-InstructServe an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Llama-3.2-11B-Vision-Instruct --port 8080How far can you push the context
2K tokensFITS · 17.0 GB FREE
cache 0.0 GB
8K tokensFITS · 17.0 GB FREE
cache 0.0 GB
32K tokensFITS · 17.0 GB FREE
cache 0.0 GB
128K tokensFITS · 17.0 GB FREE
cache 0.0 GB
Which Macs run this
| Chip | Smallest RAM that fits | Speed est. |
|---|---|---|
| M1 | no configuration | — |
| M1 Pro | 32 GB | 7.1 tok/s |
| M1 Max | 32 GB | 14.8 tok/s |
| M1 Ultra | 64 GB | 30.3 tok/s |
| M2 | no configuration | — |
| M2 Pro | 32 GB | 7.2 tok/s |
| M2 Max | 32 GB | 15.0 tok/s |
| M2 Ultra | 64 GB | 30.7 tok/s |
| M3 | no configuration | — |
| M3 Pro | 36 GB | 5.3 tok/s |
| M3 Max (14-core CPU) | 36 GB | 10.9 tok/s |
| M3 Max (16-core CPU) | 48 GB | 15.0 tok/s |
| M3 Ultra | 96 GB | 31.8 tok/s |
| M4 | 32 GB | 4.1 tok/s |
| M4 Pro | 48 GB | 10.0 tok/s |
| M4 Max (14-core CPU) | 36 GB | 15.5 tok/s |
| M4 Max (16-core CPU) | 48 GB | 20.9 tok/s |
| M5 | 32 GB | 5.4 tok/s |
| M5 Prounverified | 48 GB | 11.3 tok/s |
| M5 Maxunverified | 36 GB | 23.8 tok/s |