Skip to content
mlx.app

Qwen3 14B

Alibaba · 14.8B · up to 40K context · apache-2.0

FITS · 21.3 GB FREE

Fits 4-bit on a 16GB Mac but is happier on 24GB where the OS isn't fighting it for memory. The reasoning gains over the 8B model are real on harder logic and math prompts, at roughly double the load size.

On a M4 Pro with 48 GB

macOS9.6 GBweights15.7 GBKV cache1.3 GBheadroom21.3 GB
Weights
15.7 GB
KV cache
1.3 GB
Usable RAM
38.4 GB
Headroom
21.3 GB
Generation
13.5 tok/sest.
Prompt processing
61.0 tok/sest.
First token (1K prompt)
16.8 sest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q48.3 GB9.7 GBFITS · 28.7 GB FREE25.6 tok/sQwen3-14B-4bit
q612.0 GB13.4 GBFITS · 25.0 GB FREE17.7 tok/sQwen3-14B-6bit
q815.7 GB17.1 GBFITS · 21.3 GB FREE13.5 tok/sQwen3-14B-8bit

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/Qwen3-14B-8bit
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Qwen3-14B-8bit --port 8080

How far can you push the context

2K tokensFITS · 22.3 GB FREE

cache 0.3 GB

8K tokensFITS · 21.3 GB FREE

cache 1.3 GB

32K tokensFITS · 17.3 GB FREE

cache 5.4 GB

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M1no configuration
M1 Pro32 GB9.7 tok/s
M1 Max32 GB20.1 tok/s
M1 Ultra64 GB41.2 tok/s
M224 GB4.6 tok/s
M2 Pro32 GB9.8 tok/s
M2 Max32 GB20.3 tok/s
M2 Ultra64 GB41.7 tok/s
M324 GB4.6 tok/s
M3 Pro36 GB7.2 tok/s
M3 Max (14-core CPU)36 GB14.9 tok/s
M3 Max (16-core CPU)48 GB20.3 tok/s
M3 Ultra96 GB43.2 tok/s
M424 GB5.6 tok/s
M4 Pro24 GB13.5 tok/s
M4 Max (14-core CPU)36 GB21.1 tok/s
M4 Max (16-core CPU)48 GB28.5 tok/s
M524 GB7.3 tok/s
M5 Prounverified24 GB15.4 tok/s
M5 Maxunverified36 GB32.4 tok/s