Skip to content
mlx.app

Mistral Small 24B Instruct 2501

Mistral AI · 24B · up to 32K context · apache-2.0

FITS · 11.6 GB FREE

Aimed squarely at the 32GB Mac tier and it earns the memory, with quality that tracks close to models twice its size on everyday tasks. It's fully Apache-licensed, which matters if you're shipping a commercial product on top of it.

On a M4 Pro with 48 GB

macOS9.6 GBweights25.5 GBKV cache1.3 GBheadroom11.6 GB
Weights
25.5 GB
KV cache
1.3 GB
Usable RAM
38.4 GB
Headroom
11.6 GB
Generation
8.4 tok/sest.
Prompt processing
37.6 tok/sest.
First token (1K prompt)
27.2 sest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q413.5 GB14.8 GBFITS · 23.6 GB FREE15.8 tok/sMistral-Small-24B-Instruct-2501-4bit
q825.5 GB26.8 GBFITS · 11.6 GB FREE8.4 tok/sMistral-Small-24B-Instruct-2501-8bit

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/Mistral-Small-24B-Instruct-2501-8bit
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Mistral-Small-24B-Instruct-2501-8bit --port 8080

How far can you push the context

2K tokensFITS · 12.6 GB FREE

cache 0.3 GB

8K tokensFITS · 11.6 GB FREE

cache 1.3 GB

32K tokensFITS · 7.5 GB FREE

cache 5.4 GB

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M1no configuration
M1 Prono configuration
M1 Max64 GB12.4 tok/s
M1 Ultra64 GB25.4 tok/s
M2no configuration
M2 Prono configuration
M2 Max64 GB12.5 tok/s
M2 Ultra64 GB25.7 tok/s
M3no configuration
M3 Prono configuration
M3 Max (14-core CPU)no configuration
M3 Max (16-core CPU)48 GB12.5 tok/s
M3 Ultra96 GB26.7 tok/s
M4no configuration
M4 Pro48 GB8.4 tok/s
M4 Max (14-core CPU)48 GB13.0 tok/s
M4 Max (16-core CPU)48 GB17.6 tok/s
M5no configuration
M5 Prounverified48 GB9.5 tok/s
M5 Maxunverified48 GB20.0 tok/s