Skip to content
mlx.app

Llama 3.1 8B Instruct

Meta · 8.0B · up to 128K context · llama-3.1

FITS · 21.3 GB FREE

A true 128K context window makes this a good choice for a 16GB Mac that needs to chew on long documents, not just chat. It's a generation behind on raw reasoning benchmarks but the long context still earns it a place.

On a M4 Pro with 48 GB

macOS9.6 GBweights16.1 GBKV cache1.1 GBheadroom21.3 GB
Weights
16.1 GB
KV cache
1.1 GB
Usable RAM
38.4 GB
Headroom
21.3 GB
Generation
13.3 tok/sest.
Prompt processing
112 tok/sest.
First token (1K prompt)
9.1 sest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q44.5 GB5.6 GBFITS · 32.8 GB FREE47.1 tok/sMeta-Llama-3.1-8B-Instruct-4bit
q88.5 GB9.6 GBFITS · 28.8 GB FREE25.0 tok/sMeta-Llama-3.1-8B-Instruct-8bit
bf1616.1 GB17.1 GBFITS · 21.3 GB FREE13.3 tok/sMeta-Llama-3.1-8B-Instruct-bf16

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/Meta-Llama-3.1-8B-Instruct-bf16
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Meta-Llama-3.1-8B-Instruct-bf16 --port 8080

How far can you push the context

2K tokensFITS · 22.1 GB FREE

cache 0.3 GB

8K tokensFITS · 21.3 GB FREE

cache 1.1 GB

32K tokensFITS · 18.0 GB FREE

cache 4.3 GB

128K tokensTIGHT · 5.2 GB FREE

cache 17.2 GB

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M1no configuration
M1 Pro32 GB9.5 tok/s
M1 Max32 GB19.7 tok/s
M1 Ultra64 GB40.3 tok/s
M224 GB4.5 tok/s
M2 Pro32 GB9.6 tok/s
M2 Max32 GB19.9 tok/s
M2 Ultra64 GB40.8 tok/s
M324 GB4.5 tok/s
M3 Pro36 GB7.0 tok/s
M3 Max (14-core CPU)36 GB14.6 tok/s
M3 Max (16-core CPU)48 GB19.9 tok/s
M3 Ultra96 GB42.3 tok/s
M424 GB5.5 tok/s
M4 Pro24 GB13.3 tok/s
M4 Max (14-core CPU)36 GB20.7 tok/s
M4 Max (16-core CPU)48 GB27.9 tok/s
M524 GB7.1 tok/s
M5 Prounverified24 GB15.1 tok/s
M5 Maxunverified36 GB31.7 tok/s