Llama 3.2 1B Instruct
Meta · 1.2B · up to 128K context · llama-3.2
This one runs on almost anything with unified memory, including an 8GB Mac mini with room to spare. It answers fast but loses the thread on anything that needs more than a paragraph of reasoning.
On a M4 Pro with 48 GB
macOS9.6 GBweights2.5 GBKV cache0.3 GBheadroom35.7 GB
- Weights
- 2.5 GB
- KV cache
- 0.3 GB
- Usable RAM
- 38.4 GB
- Headroom
- 35.7 GB
- Generation
- 85.9 tok/sest.
- Prompt processing
- 728 tok/sest.
- First token (1K prompt)
- 1.4 sest.
Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.
Every quantization, on your Mac
| Quant | Weights | + KV | Verdict | Speed est. | Repo |
|---|---|---|---|---|---|
| q4 | 0.7 GB | 1.0 GB | FITS · 37.4 GB FREE | 305 tok/s | Llama-3.2-1B-Instruct-4bit |
| q8 | 1.3 GB | 1.6 GB | FITS · 36.8 GB FREE | 162 tok/s | Llama-3.2-1B-Instruct-8bit |
| bf16 | 2.5 GB | 2.7 GB | FITS · 35.7 GB FREE | 85.9 tok/s | Llama-3.2-1B-Instruct |
We never host weights. Every link goes to Hugging Face.
Run it
Chat
mlx_lm.chat --model mlx-community/Llama-3.2-1B-InstructServe an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Llama-3.2-1B-Instruct --port 8080How far can you push the context
2K tokensFITS · 35.9 GB FREE
cache 0.1 GB
8K tokensFITS · 35.7 GB FREE
cache 0.3 GB
32K tokensFITS · 34.8 GB FREE
cache 1.1 GB
128K tokensFITS · 31.6 GB FREE
cache 4.3 GB
Which Macs run this
| Chip | Smallest RAM that fits | Speed est. |
|---|---|---|
| M1 | 8 GB | 19.7 tok/s |
| M1 Pro | 16 GB | 61.3 tok/s |
| M1 Max | 32 GB | 127 tok/s |
| M1 Ultra | 64 GB | 261 tok/s |
| M2 | 8 GB | 29.4 tok/s |
| M2 Pro | 16 GB | 62.1 tok/s |
| M2 Max | 32 GB | 129 tok/s |
| M2 Ultra | 64 GB | 265 tok/s |
| M3 | 8 GB | 29.4 tok/s |
| M3 Pro | 18 GB | 45.4 tok/s |
| M3 Max (14-core CPU) | 36 GB | 94.4 tok/s |
| M3 Max (16-core CPU) | 48 GB | 129 tok/s |
| M3 Ultra | 96 GB | 274 tok/s |
| M4 | 16 GB | 35.8 tok/s |
| M4 Pro | 24 GB | 85.9 tok/s |
| M4 Max (14-core CPU) | 36 GB | 134 tok/s |
| M4 Max (16-core CPU) | 48 GB | 181 tok/s |
| M5 | 16 GB | 46.3 tok/s |
| M5 Prounverified | 24 GB | 97.8 tok/s |
| M5 Maxunverified | 36 GB | 205 tok/s |