Skip to content
mlx.app

Qwen3 Coder 30B-A3B Instruct

Alibaba · 30.5B (3.3B active, MoE) · up to 256K context · apache-2.0

FITS · 12.8 GB FREE

This is the current go-to local coding model for a 32GB Mac, trading some raw code-quality against Qwen2.5-Coder-32B for much faster generation. Its 262K context window means it can hold a real codebase's worth of files in one session, which the dense coder can't do as cheaply.

On a M4 Pro with 48 GB

macOS9.6 GBweights24.8 GBKV cache0.8 GBheadroom12.8 GB
Weights
24.8 GB
KV cache
0.8 GB
Usable RAM
38.4 GB
Headroom
12.8 GB
Generation
79.4 tok/sest.
Prompt processing
274 tok/sest.
First token (1K prompt)
3.7 sest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q417.2 GB18.0 GBFITS · 20.4 GB FREE115 tok/sQwen3-Coder-30B-A3B-Instruct-4bit
q624.8 GB25.6 GBFITS · 12.8 GB FREE79.4 tok/sQwen3-Coder-30B-A3B-Instruct-6bit
q832.4 GB33.2 GBTIGHT · 5.2 GB FREE60.7 tok/sQwen3-Coder-30B-A3B-Instruct-8bit

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/Qwen3-Coder-30B-A3B-Instruct-6bit
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Qwen3-Coder-30B-A3B-Instruct-6bit --port 8080

How far can you push the context

2K tokensFITS · 13.4 GB FREE

cache 0.2 GB

8K tokensFITS · 12.8 GB FREE

cache 0.8 GB

32K tokensFITS · 10.4 GB FREE

cache 3.2 GB

128K tokensTIGHT · 0.7 GB FREE

cache 12.9 GB

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M1no configuration
M1 Prono configuration
M1 Max64 GB118 tok/s
M1 Ultra64 GB242 tok/s
M2no configuration
M2 Prono configuration
M2 Max64 GB119 tok/s
M2 Ultra64 GB245 tok/s
M3no configuration
M3 Pro36 GB42.0 tok/s
M3 Max (14-core CPU)36 GB87.3 tok/s
M3 Max (16-core CPU)48 GB119 tok/s
M3 Ultra96 GB254 tok/s
M4no configuration
M4 Pro48 GB79.4 tok/s
M4 Max (14-core CPU)36 GB124 tok/s
M4 Max (16-core CPU)48 GB167 tok/s
M5no configuration
M5 Prounverified48 GB90.5 tok/s
M5 Maxunverified36 GB190 tok/s