Skip to content
mlx.app

Llama 3.2 3B Instruct

Meta · 3.2B · up to 128K context · llama-3.2

FITS · 31.5 GB FREE

A step up from the 1B that still fits comfortably on 8GB machines at 4-bit. It's the model most people should try first before assuming they need something bigger.

On a M4 Pro with 48 GB

macOS9.6 GBweights6.4 GBKV cache0.5 GBheadroom31.5 GB
Weights
6.4 GB
KV cache
0.5 GB
Usable RAM
38.4 GB
Headroom
31.5 GB
Generation
33.2 tok/sest.
Prompt processing
281 tok/sest.
First token (1K prompt)
3.6 sest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q41.8 GB2.3 GBFITS · 36.1 GB FREE118 tok/sLlama-3.2-3B-Instruct-4bit
q83.4 GB3.9 GBFITS · 34.5 GB FREE62.4 tok/sLlama-3.2-3B-Instruct-8bit
bf166.4 GB6.9 GBFITS · 31.5 GB FREE33.2 tok/sLlama-3.2-3B-Instruct-bf16

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/Llama-3.2-3B-Instruct-bf16
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Llama-3.2-3B-Instruct-bf16 --port 8080

How far can you push the context

2K tokensFITS · 31.9 GB FREE

cache 0.1 GB

8K tokensFITS · 31.5 GB FREE

cache 0.5 GB

32K tokensFITS · 30.1 GB FREE

cache 1.9 GB

128K tokensFITS · 24.5 GB FREE

cache 7.5 GB

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M116 GB7.6 tok/s
M1 Pro16 GB23.7 tok/s
M1 Max32 GB49.2 tok/s
M1 Ultra64 GB101 tok/s
M216 GB11.4 tok/s
M2 Pro16 GB24.0 tok/s
M2 Max32 GB49.8 tok/s
M2 Ultra64 GB102 tok/s
M316 GB11.4 tok/s
M3 Pro18 GB17.5 tok/s
M3 Max (14-core CPU)36 GB36.4 tok/s
M3 Max (16-core CPU)48 GB49.8 tok/s
M3 Ultra96 GB106 tok/s
M416 GB13.8 tok/s
M4 Pro24 GB33.2 tok/s
M4 Max (14-core CPU)36 GB51.7 tok/s
M4 Max (16-core CPU)48 GB69.7 tok/s
M516 GB17.9 tok/s
M5 Prounverified24 GB37.8 tok/s
M5 Maxunverified36 GB79.4 tok/s