Skip to content
mlx.app

Qwen3 8B

Alibaba · 8.2B · up to 40K context · apache-2.0

FITS · 28.5 GB FREE

This is currently the best all-round 8B for a 16GB Mac, with the thinking-mode toggle giving you a manual quality/speed dial. At 4-bit it leaves enough memory free to keep a coding editor and a few Chrome tabs open.

On a M4 Pro with 48 GB

macOS9.6 GBweights8.7 GBKV cache1.2 GBheadroom28.5 GB
Weights
8.7 GB
KV cache
1.2 GB
Usable RAM
38.4 GB
Headroom
28.5 GB
Generation
24.4 tok/sest.
Prompt processing
110 tok/sest.
First token (1K prompt)
9.3 sest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q44.6 GB5.8 GBFITS · 32.6 GB FREE46.2 tok/sQwen3-8B-4bit
q66.7 GB7.9 GBFITS · 30.5 GB FREE32.0 tok/sQwen3-8B-6bit
q88.7 GB9.9 GBFITS · 28.5 GB FREE24.4 tok/sQwen3-8B-8bit

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/Qwen3-8B-8bit
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Qwen3-8B-8bit --port 8080

How far can you push the context

2K tokensFITS · 29.4 GB FREE

cache 0.3 GB

8K tokensFITS · 28.5 GB FREE

cache 1.2 GB

32K tokensFITS · 24.9 GB FREE

cache 4.8 GB

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M116 GB5.6 tok/s
M1 Pro16 GB17.4 tok/s
M1 Max32 GB36.3 tok/s
M1 Ultra64 GB74.4 tok/s
M216 GB8.4 tok/s
M2 Pro16 GB17.7 tok/s
M2 Max32 GB36.7 tok/s
M2 Ultra64 GB75.3 tok/s
M316 GB8.4 tok/s
M3 Pro18 GB12.9 tok/s
M3 Max (14-core CPU)36 GB26.9 tok/s
M3 Max (16-core CPU)48 GB36.7 tok/s
M3 Ultra96 GB78.0 tok/s
M416 GB10.2 tok/s
M4 Pro24 GB24.4 tok/s
M4 Max (14-core CPU)36 GB38.1 tok/s
M4 Max (16-core CPU)48 GB51.4 tok/s
M516 GB13.2 tok/s
M5 Prounverified24 GB27.8 tok/s
M5 Maxunverified36 GB58.5 tok/s