Skip to content
mlx.app

Llama 3.2 1B Instruct

Meta · 1.2B · up to 128K context · llama-3.2

FITS · 35.7 GB FREE

This one runs on almost anything with unified memory, including an 8GB Mac mini with room to spare. It answers fast but loses the thread on anything that needs more than a paragraph of reasoning.

On a M4 Pro with 48 GB

macOS9.6 GBweights2.5 GBKV cache0.3 GBheadroom35.7 GB
Weights
2.5 GB
KV cache
0.3 GB
Usable RAM
38.4 GB
Headroom
35.7 GB
Generation
85.9 tok/sest.
Prompt processing
728 tok/sest.
First token (1K prompt)
1.4 sest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q40.7 GB1.0 GBFITS · 37.4 GB FREE305 tok/sLlama-3.2-1B-Instruct-4bit
q81.3 GB1.6 GBFITS · 36.8 GB FREE162 tok/sLlama-3.2-1B-Instruct-8bit
bf162.5 GB2.7 GBFITS · 35.7 GB FREE85.9 tok/sLlama-3.2-1B-Instruct

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/Llama-3.2-1B-Instruct
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Llama-3.2-1B-Instruct --port 8080

How far can you push the context

2K tokensFITS · 35.9 GB FREE

cache 0.1 GB

8K tokensFITS · 35.7 GB FREE

cache 0.3 GB

32K tokensFITS · 34.8 GB FREE

cache 1.1 GB

128K tokensFITS · 31.6 GB FREE

cache 4.3 GB

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M18 GB19.7 tok/s
M1 Pro16 GB61.3 tok/s
M1 Max32 GB127 tok/s
M1 Ultra64 GB261 tok/s
M28 GB29.4 tok/s
M2 Pro16 GB62.1 tok/s
M2 Max32 GB129 tok/s
M2 Ultra64 GB265 tok/s
M38 GB29.4 tok/s
M3 Pro18 GB45.4 tok/s
M3 Max (14-core CPU)36 GB94.4 tok/s
M3 Max (16-core CPU)48 GB129 tok/s
M3 Ultra96 GB274 tok/s
M416 GB35.8 tok/s
M4 Pro24 GB85.9 tok/s
M4 Max (14-core CPU)36 GB134 tok/s
M4 Max (16-core CPU)48 GB181 tok/s
M516 GB46.3 tok/s
M5 Prounverified24 GB97.8 tok/s
M5 Maxunverified36 GB205 tok/s