Skip to content
mlx.app

Codestral 22B v0.1

Mistral AI · 22.2B · up to 32K context · mnpl

FITS · 12.9 GB FREE

A code-specialist that fits a 32GB Mac comfortably at 4-bit and supports fill-in-the-middle completion, which general chat models handle poorly. The license is non-commercial though, so it's a personal-project tool rather than something to ship in a product.

On a M4 Pro with 48 GB

macOS9.6 GBweights23.6 GBKV cache1.9 GBheadroom12.9 GB
Weights
23.6 GB
KV cache
1.9 GB
Usable RAM
38.4 GB
Headroom
12.9 GB
Generation
9.0 tok/sest.
Prompt processing
40.7 tok/sest.
First token (1K prompt)
25.2 sest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q412.5 GB14.4 GBFITS · 24.0 GB FREE17.1 tok/sCodestral-22B-v0.1-4bit
q823.6 GB25.5 GBFITS · 12.9 GB FREE9.0 tok/sCodestral-22B-v0.1-8bit

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/Codestral-22B-v0.1-8bit
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Codestral-22B-v0.1-8bit --port 8080

How far can you push the context

2K tokensFITS · 14.3 GB FREE

cache 0.5 GB

8K tokensFITS · 12.9 GB FREE

cache 1.9 GB

32K tokensFITS · 7.3 GB FREE

cache 7.5 GB

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M1no configuration
M1 Prono configuration
M1 Max64 GB13.4 tok/s
M1 Ultra64 GB27.5 tok/s
M2no configuration
M2 Prono configuration
M2 Max64 GB13.6 tok/s
M2 Ultra64 GB27.8 tok/s
M3no configuration
M3 Pro36 GB4.8 tok/s
M3 Max (14-core CPU)36 GB9.9 tok/s
M3 Max (16-core CPU)48 GB13.6 tok/s
M3 Ultra96 GB28.8 tok/s
M4no configuration
M4 Pro48 GB9.0 tok/s
M4 Max (14-core CPU)36 GB14.1 tok/s
M4 Max (16-core CPU)48 GB19.0 tok/s
M5no configuration
M5 Prounverified48 GB10.3 tok/s
M5 Maxunverified36 GB21.6 tok/s