Codestral 22B v0.1
Mistral AI · 22.2B · up to 32K context · mnpl
A code-specialist that fits a 32GB Mac comfortably at 4-bit and supports fill-in-the-middle completion, which general chat models handle poorly. The license is non-commercial though, so it's a personal-project tool rather than something to ship in a product.
On a M4 Pro with 48 GB
macOS9.6 GBweights23.6 GBKV cache1.9 GBheadroom12.9 GB
- Weights
- 23.6 GB
- KV cache
- 1.9 GB
- Usable RAM
- 38.4 GB
- Headroom
- 12.9 GB
- Generation
- 9.0 tok/sest.
- Prompt processing
- 40.7 tok/sest.
- First token (1K prompt)
- 25.2 sest.
Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.
Every quantization, on your Mac
| Quant | Weights | + KV | Verdict | Speed est. | Repo |
|---|---|---|---|---|---|
| q4 | 12.5 GB | 14.4 GB | FITS · 24.0 GB FREE | 17.1 tok/s | Codestral-22B-v0.1-4bit |
| q8 | 23.6 GB | 25.5 GB | FITS · 12.9 GB FREE | 9.0 tok/s | Codestral-22B-v0.1-8bit |
We never host weights. Every link goes to Hugging Face.
Run it
Chat
mlx_lm.chat --model mlx-community/Codestral-22B-v0.1-8bitServe an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Codestral-22B-v0.1-8bit --port 8080How far can you push the context
2K tokensFITS · 14.3 GB FREE
cache 0.5 GB
8K tokensFITS · 12.9 GB FREE
cache 1.9 GB
32K tokensFITS · 7.3 GB FREE
cache 7.5 GB
Which Macs run this
| Chip | Smallest RAM that fits | Speed est. |
|---|---|---|
| M1 | no configuration | — |
| M1 Pro | no configuration | — |
| M1 Max | 64 GB | 13.4 tok/s |
| M1 Ultra | 64 GB | 27.5 tok/s |
| M2 | no configuration | — |
| M2 Pro | no configuration | — |
| M2 Max | 64 GB | 13.6 tok/s |
| M2 Ultra | 64 GB | 27.8 tok/s |
| M3 | no configuration | — |
| M3 Pro | 36 GB | 4.8 tok/s |
| M3 Max (14-core CPU) | 36 GB | 9.9 tok/s |
| M3 Max (16-core CPU) | 48 GB | 13.6 tok/s |
| M3 Ultra | 96 GB | 28.8 tok/s |
| M4 | no configuration | — |
| M4 Pro | 48 GB | 9.0 tok/s |
| M4 Max (14-core CPU) | 36 GB | 14.1 tok/s |
| M4 Max (16-core CPU) | 48 GB | 19.0 tok/s |
| M5 | no configuration | — |
| M5 Prounverified | 48 GB | 10.3 tok/s |
| M5 Maxunverified | 36 GB | 21.6 tok/s |