Skip to content
mlx.app

GPT-OSS 20B

OpenAI · 21B (3.6B active, MoE) · up to 128K context · apache-2.0

FITS · 15.7 GB FREE

OpenAI's first open-weight release since GPT-2, shipped natively in 4-bit MXFP4 so it already fits a 16GB Mac without further quantization. It uses OpenAI's harmony chat format and a configurable reasoning effort setting rather than a simple thinking toggle.

On a M4 Pro with 48 GB

macOS9.6 GBweights22.3 GBKV cache0.4 GBheadroom15.7 GB
Weights
22.3 GB
KV cache
0.4 GB
Usable RAM
38.4 GB
Headroom
15.7 GB
Generation
55.7 tok/sest.
Prompt processing
251 tok/sest.
First token (1K prompt)
4.1 sest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q411.8 GB12.2 GBFITS · 26.2 GB FREE105 tok/sgpt-oss-20b-MXFP4-Q4
q822.3 GB22.7 GBFITS · 15.7 GB FREE55.7 tok/sgpt-oss-20b-MXFP4-Q8

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/gpt-oss-20b-MXFP4-Q8
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/gpt-oss-20b-MXFP4-Q8 --port 8080

How far can you push the context

2K tokensFITS · 16.0 GB FREE

cache 0.1 GB

8K tokensFITS · 15.7 GB FREE

cache 0.4 GB

32K tokensFITS · 14.5 GB FREE

cache 1.6 GB

128K tokensFITS · 9.6 GB FREE

cache 6.4 GB

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M1no configuration
M1 Pro32 GB39.7 tok/s
M1 Max32 GB82.6 tok/s
M1 Ultra64 GB169 tok/s
M2no configuration
M2 Pro32 GB40.3 tok/s
M2 Max32 GB83.7 tok/s
M2 Ultra64 GB172 tok/s
M3no configuration
M3 Pro36 GB29.4 tok/s
M3 Max (14-core CPU)36 GB61.2 tok/s
M3 Max (16-core CPU)48 GB83.7 tok/s
M3 Ultra96 GB178 tok/s
M432 GB23.2 tok/s
M4 Pro48 GB55.7 tok/s
M4 Max (14-core CPU)36 GB86.8 tok/s
M4 Max (16-core CPU)48 GB117 tok/s
M532 GB30.0 tok/s
M5 Prounverified48 GB63.4 tok/s
M5 Maxunverified36 GB133 tok/s