Mistral 7B Instruct v0.3
Mistral AI · 7.3B · up to 32K context · apache-2.0
Still a dependable, no-drama 7B for a 16GB machine, though newer Qwen3 models have mostly caught up on quality. Its main selling point at this point is a permissive license and function-calling support baked in.
On a M4 Pro with 48 GB
macOS9.6 GBweights7.8 GBKV cache1.1 GBheadroom29.6 GB
- Weights
- 7.8 GB
- KV cache
- 1.1 GB
- Usable RAM
- 38.4 GB
- Headroom
- 29.6 GB
- Generation
- 27.5 tok/sest.
- Prompt processing
- 124 tok/sest.
- First token (1K prompt)
- 8.3 sest.
Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.
Every quantization, on your Mac
| Quant | Weights | + KV | Verdict | Speed est. | Repo |
|---|---|---|---|---|---|
| q4 | 4.1 GB | 5.2 GB | FITS · 33.2 GB FREE | 51.9 tok/s | Mistral-7B-Instruct-v0.3-4bit |
| q8 | 7.8 GB | 8.8 GB | FITS · 29.6 GB FREE | 27.5 tok/s | Mistral-7B-Instruct-v0.3-8bit |
We never host weights. Every link goes to Hugging Face.
Run it
Chat
mlx_lm.chat --model mlx-community/Mistral-7B-Instruct-v0.3-8bitServe an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Mistral-7B-Instruct-v0.3-8bit --port 8080How far can you push the context
2K tokensFITS · 30.4 GB FREE
cache 0.3 GB
8K tokensFITS · 29.6 GB FREE
cache 1.1 GB
32K tokensFITS · 26.3 GB FREE
cache 4.3 GB
Which Macs run this
| Chip | Smallest RAM that fits | Speed est. |
|---|---|---|
| M1 | 16 GB | 6.3 tok/s |
| M1 Pro | 16 GB | 19.6 tok/s |
| M1 Max | 32 GB | 40.7 tok/s |
| M1 Ultra | 64 GB | 83.5 tok/s |
| M2 | 16 GB | 9.4 tok/s |
| M2 Pro | 16 GB | 19.9 tok/s |
| M2 Max | 32 GB | 41.3 tok/s |
| M2 Ultra | 64 GB | 84.6 tok/s |
| M3 | 16 GB | 9.4 tok/s |
| M3 Pro | 18 GB | 14.5 tok/s |
| M3 Max (14-core CPU) | 36 GB | 30.2 tok/s |
| M3 Max (16-core CPU) | 48 GB | 41.3 tok/s |
| M3 Ultra | 96 GB | 87.6 tok/s |
| M4 | 16 GB | 11.4 tok/s |
| M4 Pro | 24 GB | 27.5 tok/s |
| M4 Max (14-core CPU) | 36 GB | 42.8 tok/s |
| M4 Max (16-core CPU) | 48 GB | 57.7 tok/s |
| M5 | 16 GB | 14.8 tok/s |
| M5 Prounverified | 24 GB | 31.3 tok/s |
| M5 Maxunverified | 36 GB | 65.7 tok/s |