Qwen3 14B
Alibaba · 14.8B · up to 40K context · apache-2.0
Fits 4-bit on a 16GB Mac but is happier on 24GB where the OS isn't fighting it for memory. The reasoning gains over the 8B model are real on harder logic and math prompts, at roughly double the load size.
On a M4 Pro with 48 GB
macOS9.6 GBweights15.7 GBKV cache1.3 GBheadroom21.3 GB
- Weights
- 15.7 GB
- KV cache
- 1.3 GB
- Usable RAM
- 38.4 GB
- Headroom
- 21.3 GB
- Generation
- 13.5 tok/sest.
- Prompt processing
- 61.0 tok/sest.
- First token (1K prompt)
- 16.8 sest.
Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.
Every quantization, on your Mac
| Quant | Weights | + KV | Verdict | Speed est. | Repo |
|---|---|---|---|---|---|
| q4 | 8.3 GB | 9.7 GB | FITS · 28.7 GB FREE | 25.6 tok/s | Qwen3-14B-4bit |
| q6 | 12.0 GB | 13.4 GB | FITS · 25.0 GB FREE | 17.7 tok/s | Qwen3-14B-6bit |
| q8 | 15.7 GB | 17.1 GB | FITS · 21.3 GB FREE | 13.5 tok/s | Qwen3-14B-8bit |
We never host weights. Every link goes to Hugging Face.
Run it
Chat
mlx_lm.chat --model mlx-community/Qwen3-14B-8bitServe an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Qwen3-14B-8bit --port 8080How far can you push the context
2K tokensFITS · 22.3 GB FREE
cache 0.3 GB
8K tokensFITS · 21.3 GB FREE
cache 1.3 GB
32K tokensFITS · 17.3 GB FREE
cache 5.4 GB
Which Macs run this
| Chip | Smallest RAM that fits | Speed est. |
|---|---|---|
| M1 | no configuration | — |
| M1 Pro | 32 GB | 9.7 tok/s |
| M1 Max | 32 GB | 20.1 tok/s |
| M1 Ultra | 64 GB | 41.2 tok/s |
| M2 | 24 GB | 4.6 tok/s |
| M2 Pro | 32 GB | 9.8 tok/s |
| M2 Max | 32 GB | 20.3 tok/s |
| M2 Ultra | 64 GB | 41.7 tok/s |
| M3 | 24 GB | 4.6 tok/s |
| M3 Pro | 36 GB | 7.2 tok/s |
| M3 Max (14-core CPU) | 36 GB | 14.9 tok/s |
| M3 Max (16-core CPU) | 48 GB | 20.3 tok/s |
| M3 Ultra | 96 GB | 43.2 tok/s |
| M4 | 24 GB | 5.6 tok/s |
| M4 Pro | 24 GB | 13.5 tok/s |
| M4 Max (14-core CPU) | 36 GB | 21.1 tok/s |
| M4 Max (16-core CPU) | 48 GB | 28.5 tok/s |
| M5 | 24 GB | 7.3 tok/s |
| M5 Prounverified | 24 GB | 15.4 tok/s |
| M5 Maxunverified | 36 GB | 32.4 tok/s |