Skip to content
mlx.app

Llama 3.2 90B Vision Instruct

Meta · 88B · up to 128K context · llama-3.2

OVER BY 11.1 GB

Needs a 64GB Mac for the 4-bit build, and it's the rare vision model at this scale that also handles long text context well. It's a meaningful jump over the 11B version on complex visual reasoning, though most users won't need to pay this much memory for it.

On a M4 Pro with 48 GB

macOS9.6 GBweights49.5 GBKV cache0.0 GBover11.1 GB
Weights
49.5 GB
KV cache
0.0 GBestimated shape
Usable RAM
38.4 GB
Headroom
0.0 GB
Generation
4.3 tok/sest.
Prompt processing
10.3 tok/sest.
First token (1K prompt)
99.8 sest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q449.5 GB49.5 GBOVER BY 11.1 GB4.3 tok/sLlama-3.2-90B-Vision-Instruct-4bit

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/Llama-3.2-90B-Vision-Instruct-4bit
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Llama-3.2-90B-Vision-Instruct-4bit --port 8080

How far can you push the context

2K tokensOVER BY 11.1 GB

cache 0.0 GB

8K tokensOVER BY 11.1 GB

cache 0.0 GB

32K tokensOVER BY 11.1 GB

cache 0.0 GB

128K tokensOVER BY 11.1 GB

cache 0.0 GB

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M1no configuration
M1 Prono configuration
M1 Max64 GB6.4 tok/s
M1 Ultra64 GB13.1 tok/s
M2no configuration
M2 Prono configuration
M2 Max64 GB6.5 tok/s
M2 Ultra64 GB13.3 tok/s
M3no configuration
M3 Prono configuration
M3 Max (14-core CPU)no configuration
M3 Max (16-core CPU)64 GB6.5 tok/s
M3 Ultra96 GB13.7 tok/s
M4no configuration
M4 Pro64 GB4.3 tok/s
M4 Max (14-core CPU)no configuration
M4 Max (16-core CPU)64 GB9.0 tok/s
M5no configuration
M5 Prounverified64 GB4.9 tok/s
M5 Maxunverified64 GB10.3 tok/s