Skip to content
mlx.app

Mistral 7B Instruct v0.3

Mistral AI · 7.3B · up to 32K context · apache-2.0

FITS · 29.6 GB FREE

Still a dependable, no-drama 7B for a 16GB machine, though newer Qwen3 models have mostly caught up on quality. Its main selling point at this point is a permissive license and function-calling support baked in.

On a M4 Pro with 48 GB

macOS9.6 GBweights7.8 GBKV cache1.1 GBheadroom29.6 GB
Weights
7.8 GB
KV cache
1.1 GB
Usable RAM
38.4 GB
Headroom
29.6 GB
Generation
27.5 tok/sest.
Prompt processing
124 tok/sest.
First token (1K prompt)
8.3 sest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q44.1 GB5.2 GBFITS · 33.2 GB FREE51.9 tok/sMistral-7B-Instruct-v0.3-4bit
q87.8 GB8.8 GBFITS · 29.6 GB FREE27.5 tok/sMistral-7B-Instruct-v0.3-8bit

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/Mistral-7B-Instruct-v0.3-8bit
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Mistral-7B-Instruct-v0.3-8bit --port 8080

How far can you push the context

2K tokensFITS · 30.4 GB FREE

cache 0.3 GB

8K tokensFITS · 29.6 GB FREE

cache 1.1 GB

32K tokensFITS · 26.3 GB FREE

cache 4.3 GB

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M116 GB6.3 tok/s
M1 Pro16 GB19.6 tok/s
M1 Max32 GB40.7 tok/s
M1 Ultra64 GB83.5 tok/s
M216 GB9.4 tok/s
M2 Pro16 GB19.9 tok/s
M2 Max32 GB41.3 tok/s
M2 Ultra64 GB84.6 tok/s
M316 GB9.4 tok/s
M3 Pro18 GB14.5 tok/s
M3 Max (14-core CPU)36 GB30.2 tok/s
M3 Max (16-core CPU)48 GB41.3 tok/s
M3 Ultra96 GB87.6 tok/s
M416 GB11.4 tok/s
M4 Pro24 GB27.5 tok/s
M4 Max (14-core CPU)36 GB42.8 tok/s
M4 Max (16-core CPU)48 GB57.7 tok/s
M516 GB14.8 tok/s
M5 Prounverified24 GB31.3 tok/s
M5 Maxunverified36 GB65.7 tok/s