Skip to content
mlx.app

Qwen3 32B

Alibaba · 32.8B · up to 40K context · apache-2.0

FITS · 9.6 GB FREE

Currently the strongest dense text model that a 32GB Mac can run at a usable quant, and its thinking mode holds up on genuinely hard reasoning benchmarks. If you have the memory budget, this outperforms Qwen3-30B-A3B on quality per token even though it's slower.

On a M4 Pro with 48 GB

macOS9.6 GBweights26.6 GBKV cache2.1 GBheadroom9.6 GB
Weights
26.6 GB
KV cache
2.1 GB
Usable RAM
38.4 GB
Headroom
9.6 GB
Generation
8.0 tok/sest.
Prompt processing
27.5 tok/sest.
First token (1K prompt)
37.2 sest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q418.4 GB20.6 GBFITS · 17.8 GB FREE11.5 tok/sQwen3-32B-4bit
q626.6 GB28.8 GBFITS · 9.6 GB FREE8.0 tok/sQwen3-32B-6bit
q834.8 GB37.0 GBTIGHT · 1.4 GB FREE6.1 tok/sQwen3-32B-8bit

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/Qwen3-32B-6bit
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Qwen3-32B-6bit --port 8080

How far can you push the context

2K tokensFITS · 11.2 GB FREE

cache 0.5 GB

8K tokensFITS · 9.6 GB FREE

cache 2.1 GB

32K tokensTIGHT · 3.2 GB FREE

cache 8.6 GB

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M1no configuration
M1 Prono configuration
M1 Max64 GB11.9 tok/s
M1 Ultra64 GB24.3 tok/s
M2no configuration
M2 Prono configuration
M2 Max64 GB12.0 tok/s
M2 Ultra64 GB24.6 tok/s
M3no configuration
M3 Prono configuration
M3 Max (14-core CPU)no configuration
M3 Max (16-core CPU)48 GB12.0 tok/s
M3 Ultra96 GB25.5 tok/s
M4no configuration
M4 Pro48 GB8.0 tok/s
M4 Max (14-core CPU)48 GB12.5 tok/s
M4 Max (16-core CPU)48 GB16.8 tok/s
M5no configuration
M5 Prounverified48 GB9.1 tok/s
M5 Maxunverified48 GB19.1 tok/s