Skip to content
mlx.app

Qwen3 235B-A22B

Alibaba · 235B (22B active, MoE) · up to 40K context · apache-2.0

OVER BY 212.9 GB

A genuinely usable frontier-scale MoE on a 192GB or 256GB Mac Studio, with 22B active parameters keeping generation speed reasonable despite the huge total. It sits between GPT-OSS-120B and DeepSeek-V3 in both memory demand and benchmark quality.

On a M4 Pro with 48 GB

macOS9.6 GBweights249.7 GBKV cache1.6 GBover212.9 GB
Weights
249.7 GB
KV cache
1.6 GB
Usable RAM
38.4 GB
Headroom
0.0 GB
Generation
9.1 tok/sest.
Prompt processing
41.0 tok/sest.
First token (1K prompt)
24.9 sest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q4132.2 GB133.8 GBOVER BY 95.4 GB17.2 tok/sQwen3-235B-A22B-4bit
q8249.7 GB251.3 GBOVER BY 212.9 GB9.1 tok/sQwen3-235B-A22B-4bit-DWQ

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/Qwen3-235B-A22B-4bit-DWQ
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Qwen3-235B-A22B-4bit-DWQ --port 8080

How far can you push the context

2K tokensOVER BY 211.7 GB

cache 0.4 GB

8K tokensOVER BY 212.9 GB

cache 1.6 GB

32K tokensOVER BY 217.6 GB

cache 6.3 GB

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M1no configuration
M1 Prono configuration
M1 Maxno configuration
M1 Ultrano configuration
M2no configuration
M2 Prono configuration
M2 Maxno configuration
M2 Ultrano configuration
M3no configuration
M3 Prono configuration
M3 Max (14-core CPU)no configuration
M3 Max (16-core CPU)no configuration
M3 Ultra512 GB29.1 tok/s
M4no configuration
M4 Prono configuration
M4 Max (14-core CPU)no configuration
M4 Max (16-core CPU)no configuration
M5no configuration
M5 Prounverifiedno configuration
M5 Maxunverifiedno configuration