Llama 3.2 3B Instruct
Meta · 3.2B · up to 128K context · llama-3.2
A step up from the 1B that still fits comfortably on 8GB machines at 4-bit. It's the model most people should try first before assuming they need something bigger.
On a M4 Pro with 48 GB
macOS9.6 GBweights6.4 GBKV cache0.5 GBheadroom31.5 GB
- Weights
- 6.4 GB
- KV cache
- 0.5 GB
- Usable RAM
- 38.4 GB
- Headroom
- 31.5 GB
- Generation
- 33.2 tok/sest.
- Prompt processing
- 281 tok/sest.
- First token (1K prompt)
- 3.6 sest.
Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.
Every quantization, on your Mac
| Quant | Weights | + KV | Verdict | Speed est. | Repo |
|---|---|---|---|---|---|
| q4 | 1.8 GB | 2.3 GB | FITS · 36.1 GB FREE | 118 tok/s | Llama-3.2-3B-Instruct-4bit |
| q8 | 3.4 GB | 3.9 GB | FITS · 34.5 GB FREE | 62.4 tok/s | Llama-3.2-3B-Instruct-8bit |
| bf16 | 6.4 GB | 6.9 GB | FITS · 31.5 GB FREE | 33.2 tok/s | Llama-3.2-3B-Instruct-bf16 |
We never host weights. Every link goes to Hugging Face.
Run it
Chat
mlx_lm.chat --model mlx-community/Llama-3.2-3B-Instruct-bf16Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Llama-3.2-3B-Instruct-bf16 --port 8080How far can you push the context
2K tokensFITS · 31.9 GB FREE
cache 0.1 GB
8K tokensFITS · 31.5 GB FREE
cache 0.5 GB
32K tokensFITS · 30.1 GB FREE
cache 1.9 GB
128K tokensFITS · 24.5 GB FREE
cache 7.5 GB
Which Macs run this
| Chip | Smallest RAM that fits | Speed est. |
|---|---|---|
| M1 | 16 GB | 7.6 tok/s |
| M1 Pro | 16 GB | 23.7 tok/s |
| M1 Max | 32 GB | 49.2 tok/s |
| M1 Ultra | 64 GB | 101 tok/s |
| M2 | 16 GB | 11.4 tok/s |
| M2 Pro | 16 GB | 24.0 tok/s |
| M2 Max | 32 GB | 49.8 tok/s |
| M2 Ultra | 64 GB | 102 tok/s |
| M3 | 16 GB | 11.4 tok/s |
| M3 Pro | 18 GB | 17.5 tok/s |
| M3 Max (14-core CPU) | 36 GB | 36.4 tok/s |
| M3 Max (16-core CPU) | 48 GB | 49.8 tok/s |
| M3 Ultra | 96 GB | 106 tok/s |
| M4 | 16 GB | 13.8 tok/s |
| M4 Pro | 24 GB | 33.2 tok/s |
| M4 Max (14-core CPU) | 36 GB | 51.7 tok/s |
| M4 Max (16-core CPU) | 48 GB | 69.7 tok/s |
| M5 | 16 GB | 17.9 tok/s |
| M5 Prounverified | 24 GB | 37.8 tok/s |
| M5 Maxunverified | 36 GB | 79.4 tok/s |