Llama 3.2 90B Vision Instruct
Meta · 88B · up to 128K context · llama-3.2
Needs a 64GB Mac for the 4-bit build, and it's the rare vision model at this scale that also handles long text context well. It's a meaningful jump over the 11B version on complex visual reasoning, though most users won't need to pay this much memory for it.
On a M4 Pro with 48 GB
macOS9.6 GBweights49.5 GBKV cache0.0 GBover11.1 GB
- Weights
- 49.5 GB
- KV cache
- 0.0 GBestimated shape
- Usable RAM
- 38.4 GB
- Headroom
- 0.0 GB
- Generation
- 4.3 tok/sest.
- Prompt processing
- 10.3 tok/sest.
- First token (1K prompt)
- 99.8 sest.
Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.
Every quantization, on your Mac
| Quant | Weights | + KV | Verdict | Speed est. | Repo |
|---|---|---|---|---|---|
| q4 | 49.5 GB | 49.5 GB | OVER BY 11.1 GB | 4.3 tok/s | Llama-3.2-90B-Vision-Instruct-4bit |
We never host weights. Every link goes to Hugging Face.
Run it
Chat
mlx_lm.chat --model mlx-community/Llama-3.2-90B-Vision-Instruct-4bitServe an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Llama-3.2-90B-Vision-Instruct-4bit --port 8080How far can you push the context
2K tokensOVER BY 11.1 GB
cache 0.0 GB
8K tokensOVER BY 11.1 GB
cache 0.0 GB
32K tokensOVER BY 11.1 GB
cache 0.0 GB
128K tokensOVER BY 11.1 GB
cache 0.0 GB
Which Macs run this
| Chip | Smallest RAM that fits | Speed est. |
|---|---|---|
| M1 | no configuration | — |
| M1 Pro | no configuration | — |
| M1 Max | 64 GB | 6.4 tok/s |
| M1 Ultra | 64 GB | 13.1 tok/s |
| M2 | no configuration | — |
| M2 Pro | no configuration | — |
| M2 Max | 64 GB | 6.5 tok/s |
| M2 Ultra | 64 GB | 13.3 tok/s |
| M3 | no configuration | — |
| M3 Pro | no configuration | — |
| M3 Max (14-core CPU) | no configuration | — |
| M3 Max (16-core CPU) | 64 GB | 6.5 tok/s |
| M3 Ultra | 96 GB | 13.7 tok/s |
| M4 | no configuration | — |
| M4 Pro | 64 GB | 4.3 tok/s |
| M4 Max (14-core CPU) | no configuration | — |
| M4 Max (16-core CPU) | 64 GB | 9.0 tok/s |
| M5 | no configuration | — |
| M5 Prounverified | 64 GB | 4.9 tok/s |
| M5 Maxunverified | 64 GB | 10.3 tok/s |