GPT-OSS 20B
OpenAI · 21B (3.6B active, MoE) · up to 128K context · apache-2.0
OpenAI's first open-weight release since GPT-2, shipped natively in 4-bit MXFP4 so it already fits a 16GB Mac without further quantization. It uses OpenAI's harmony chat format and a configurable reasoning effort setting rather than a simple thinking toggle.
On a M4 Pro with 48 GB
macOS9.6 GBweights22.3 GBKV cache0.4 GBheadroom15.7 GB
- Weights
- 22.3 GB
- KV cache
- 0.4 GB
- Usable RAM
- 38.4 GB
- Headroom
- 15.7 GB
- Generation
- 55.7 tok/sest.
- Prompt processing
- 251 tok/sest.
- First token (1K prompt)
- 4.1 sest.
Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.
Every quantization, on your Mac
| Quant | Weights | + KV | Verdict | Speed est. | Repo |
|---|---|---|---|---|---|
| q4 | 11.8 GB | 12.2 GB | FITS · 26.2 GB FREE | 105 tok/s | gpt-oss-20b-MXFP4-Q4 |
| q8 | 22.3 GB | 22.7 GB | FITS · 15.7 GB FREE | 55.7 tok/s | gpt-oss-20b-MXFP4-Q8 |
We never host weights. Every link goes to Hugging Face.
Run it
Chat
mlx_lm.chat --model mlx-community/gpt-oss-20b-MXFP4-Q8Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/gpt-oss-20b-MXFP4-Q8 --port 8080How far can you push the context
2K tokensFITS · 16.0 GB FREE
cache 0.1 GB
8K tokensFITS · 15.7 GB FREE
cache 0.4 GB
32K tokensFITS · 14.5 GB FREE
cache 1.6 GB
128K tokensFITS · 9.6 GB FREE
cache 6.4 GB
Which Macs run this
| Chip | Smallest RAM that fits | Speed est. |
|---|---|---|
| M1 | no configuration | — |
| M1 Pro | 32 GB | 39.7 tok/s |
| M1 Max | 32 GB | 82.6 tok/s |
| M1 Ultra | 64 GB | 169 tok/s |
| M2 | no configuration | — |
| M2 Pro | 32 GB | 40.3 tok/s |
| M2 Max | 32 GB | 83.7 tok/s |
| M2 Ultra | 64 GB | 172 tok/s |
| M3 | no configuration | — |
| M3 Pro | 36 GB | 29.4 tok/s |
| M3 Max (14-core CPU) | 36 GB | 61.2 tok/s |
| M3 Max (16-core CPU) | 48 GB | 83.7 tok/s |
| M3 Ultra | 96 GB | 178 tok/s |
| M4 | 32 GB | 23.2 tok/s |
| M4 Pro | 48 GB | 55.7 tok/s |
| M4 Max (14-core CPU) | 36 GB | 86.8 tok/s |
| M4 Max (16-core CPU) | 48 GB | 117 tok/s |
| M5 | 32 GB | 30.0 tok/s |
| M5 Prounverified | 48 GB | 63.4 tok/s |
| M5 Maxunverified | 36 GB | 133 tok/s |