Qwen2.5 0.5B Instruct
Alibaba · 490M · up to 32K context · apache-2.0
Small enough to load on a base 8GB MacBook Air alongside a browser and an IDE. Treat it as a utility model for routing and formatting, not as a reasoning partner.
On a M4 Pro with 48 GB
macOS9.6 GBweights1.0 GBKV cache0.1 GBheadroom37.3 GB
- Weights
- 1.0 GB
- KV cache
- 0.1 GB
- Usable RAM
- 38.4 GB
- Headroom
- 37.3 GB
- Generation
- 217 tok/sest.
- Prompt processing
- 1843 tok/sest.
- First token (1K prompt)
- 556 msest.
Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.
Every quantization, on your Mac
| Quant | Weights | + KV | Verdict | Speed est. | Repo |
|---|---|---|---|---|---|
| q4 | 0.3 GB | 0.4 GB | FITS · 38.0 GB FREE | 773 tok/s | Qwen2.5-0.5B-Instruct-4bit |
| q8 | 0.5 GB | 0.6 GB | FITS · 37.8 GB FREE | 409 tok/s | Qwen2.5-0.5B-Instruct-8bit |
| bf16 | 1.0 GB | 1.1 GB | FITS · 37.3 GB FREE | 217 tok/s | Qwen2.5-0.5B-Instruct-bf16 |
We never host weights. Every link goes to Hugging Face.
Run it
Chat
mlx_lm.chat --model mlx-community/Qwen2.5-0.5B-Instruct-bf16Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Qwen2.5-0.5B-Instruct-bf16 --port 8080How far can you push the context
2K tokensFITS · 37.4 GB FREE
cache 0.0 GB
8K tokensFITS · 37.3 GB FREE
cache 0.1 GB
32K tokensFITS · 37.0 GB FREE
cache 0.4 GB
Which Macs run this
| Chip | Smallest RAM that fits | Speed est. |
|---|---|---|
| M1 | 8 GB | 50.0 tok/s |
| M1 Pro | 16 GB | 155 tok/s |
| M1 Max | 32 GB | 322 tok/s |
| M1 Ultra | 64 GB | 661 tok/s |
| M2 | 8 GB | 74.5 tok/s |
| M2 Pro | 16 GB | 157 tok/s |
| M2 Max | 32 GB | 327 tok/s |
| M2 Ultra | 64 GB | 669 tok/s |
| M3 | 8 GB | 74.5 tok/s |
| M3 Pro | 18 GB | 115 tok/s |
| M3 Max (14-core CPU) | 36 GB | 239 tok/s |
| M3 Max (16-core CPU) | 48 GB | 327 tok/s |
| M3 Ultra | 96 GB | 694 tok/s |
| M4 | 16 GB | 90.6 tok/s |
| M4 Pro | 24 GB | 217 tok/s |
| M4 Max (14-core CPU) | 36 GB | 339 tok/s |
| M4 Max (16-core CPU) | 48 GB | 457 tok/s |
| M5 | 16 GB | 117 tok/s |
| M5 Prounverified | 24 GB | 247 tok/s |
| M5 Maxunverified | 36 GB | 520 tok/s |