Skip to content
mlx.app

Qwen3 1.7B

Alibaba · 1.7B · up to 40K context · apache-2.0

FITS · 34.1 GB FREE

A comfortable fit on any 8GB Mac and effectively free to run on 16GB. The thinking mode helps on logic puzzles but roughly doubles response time, so it's a tradeoff you feel.

On a M4 Pro with 48 GB

macOS9.6 GBweights3.4 GBKV cache0.9 GBheadroom34.1 GB
Weights
3.4 GB
KV cache
0.9 GB
Usable RAM
38.4 GB
Headroom
34.1 GB
Generation
62.6 tok/sest.
Prompt processing
531 tok/sest.
First token (1K prompt)
1.9 sest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q41.0 GB1.9 GBFITS · 36.5 GB FREE223 tok/sQwen3-1.7B-4bit
q81.8 GB2.7 GBFITS · 35.7 GB FREE118 tok/sQwen3-1.7B-8bit
bf163.4 GB4.3 GBFITS · 34.1 GB FREE62.6 tok/sQwen3-1.7B-bf16

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/Qwen3-1.7B-bf16
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Qwen3-1.7B-bf16 --port 8080

How far can you push the context

2K tokensFITS · 34.8 GB FREE

cache 0.2 GB

8K tokensFITS · 34.1 GB FREE

cache 0.9 GB

32K tokensFITS · 31.2 GB FREE

cache 3.8 GB

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M18 GB14.4 tok/s
M1 Pro16 GB44.7 tok/s
M1 Max32 GB92.9 tok/s
M1 Ultra64 GB191 tok/s
M28 GB21.5 tok/s
M2 Pro16 GB45.3 tok/s
M2 Max32 GB94.1 tok/s
M2 Ultra64 GB193 tok/s
M38 GB21.5 tok/s
M3 Pro18 GB33.1 tok/s
M3 Max (14-core CPU)36 GB68.8 tok/s
M3 Max (16-core CPU)48 GB94.1 tok/s
M3 Ultra96 GB200 tok/s
M416 GB26.1 tok/s
M4 Pro24 GB62.6 tok/s
M4 Max (14-core CPU)36 GB97.7 tok/s
M4 Max (16-core CPU)48 GB132 tok/s
M516 GB33.8 tok/s
M5 Prounverified24 GB71.3 tok/s
M5 Maxunverified36 GB150 tok/s