Skip to content
mlx.app

Qwen3 0.6B

Alibaba · 600M · up to 40K context · apache-2.0

FITS · 36.3 GB FREE

Qwen3's smallest model adds an optional thinking mode that the 0.5-class Qwen2.5 didn't have. It still can't hold much context, so keep prompts short and single-purpose.

On a M4 Pro with 48 GB

macOS9.6 GBweights1.2 GBKV cache0.9 GBheadroom36.3 GB
Weights
1.2 GB
KV cache
0.9 GB
Usable RAM
38.4 GB
Headroom
36.3 GB
Generation
177 tok/sest.
Prompt processing
1505 tok/sest.
First token (1K prompt)
680 msest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q40.3 GB1.3 GBFITS · 37.1 GB FREE631 tok/sQwen3-0.6B-4bit
q80.6 GB1.6 GBFITS · 36.8 GB FREE334 tok/sQwen3-0.6B-8bit
bf161.2 GB2.1 GBFITS · 36.3 GB FREE177 tok/sQwen3-0.6B-bf16

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/Qwen3-0.6B-bf16
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Qwen3-0.6B-bf16 --port 8080

How far can you push the context

2K tokensFITS · 37.0 GB FREE

cache 0.2 GB

8K tokensFITS · 36.3 GB FREE

cache 0.9 GB

32K tokensFITS · 33.4 GB FREE

cache 3.8 GB

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M18 GB40.8 tok/s
M1 Pro16 GB127 tok/s
M1 Max32 GB263 tok/s
M1 Ultra64 GB540 tok/s
M28 GB60.8 tok/s
M2 Pro16 GB128 tok/s
M2 Max32 GB267 tok/s
M2 Ultra64 GB547 tok/s
M38 GB60.8 tok/s
M3 Pro18 GB93.8 tok/s
M3 Max (14-core CPU)36 GB195 tok/s
M3 Max (16-core CPU)48 GB267 tok/s
M3 Ultra96 GB566 tok/s
M416 GB74.0 tok/s
M4 Pro24 GB177 tok/s
M4 Max (14-core CPU)36 GB277 tok/s
M4 Max (16-core CPU)48 GB373 tok/s
M516 GB95.6 tok/s
M5 Prounverified24 GB202 tok/s
M5 Maxunverified36 GB425 tok/s