Skip to content
mlx.app

Qwen2.5 0.5B Instruct

Alibaba · 490M · up to 32K context · apache-2.0

FITS · 37.3 GB FREE

Small enough to load on a base 8GB MacBook Air alongside a browser and an IDE. Treat it as a utility model for routing and formatting, not as a reasoning partner.

On a M4 Pro with 48 GB

macOS9.6 GBweights1.0 GBKV cache0.1 GBheadroom37.3 GB
Weights
1.0 GB
KV cache
0.1 GB
Usable RAM
38.4 GB
Headroom
37.3 GB
Generation
217 tok/sest.
Prompt processing
1843 tok/sest.
First token (1K prompt)
556 msest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q40.3 GB0.4 GBFITS · 38.0 GB FREE773 tok/sQwen2.5-0.5B-Instruct-4bit
q80.5 GB0.6 GBFITS · 37.8 GB FREE409 tok/sQwen2.5-0.5B-Instruct-8bit
bf161.0 GB1.1 GBFITS · 37.3 GB FREE217 tok/sQwen2.5-0.5B-Instruct-bf16

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/Qwen2.5-0.5B-Instruct-bf16
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Qwen2.5-0.5B-Instruct-bf16 --port 8080

How far can you push the context

2K tokensFITS · 37.4 GB FREE

cache 0.0 GB

8K tokensFITS · 37.3 GB FREE

cache 0.1 GB

32K tokensFITS · 37.0 GB FREE

cache 0.4 GB

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M18 GB50.0 tok/s
M1 Pro16 GB155 tok/s
M1 Max32 GB322 tok/s
M1 Ultra64 GB661 tok/s
M28 GB74.5 tok/s
M2 Pro16 GB157 tok/s
M2 Max32 GB327 tok/s
M2 Ultra64 GB669 tok/s
M38 GB74.5 tok/s
M3 Pro18 GB115 tok/s
M3 Max (14-core CPU)36 GB239 tok/s
M3 Max (16-core CPU)48 GB327 tok/s
M3 Ultra96 GB694 tok/s
M416 GB90.6 tok/s
M4 Pro24 GB217 tok/s
M4 Max (14-core CPU)36 GB339 tok/s
M4 Max (16-core CPU)48 GB457 tok/s
M516 GB117 tok/s
M5 Prounverified24 GB247 tok/s
M5 Maxunverified36 GB520 tok/s