Skip to content
mlx.app

M4 · 32 GB

120 GB/s memory bandwidth · 23.0 GB usable for models · 9.0 GB left to macOS

The short answer

27 models run comfortably and 5 more will run tight at 8K context. The largest that fits is Qwen3 30B-A3B Instruct 2507 at q4, generating around 47.8 tokens per second.

macOS9.0 GBweights17.2 GBKV cache0.8 GBheadroom5.1 GB

Every model on this Mac

ModelSizeBest quantVerdictSpeed est.
Qwen2.5 0.5B Instruct490Mbf16FITS · 22.0 GB FREE90.6 tok/s
BGE-M3567Mbf16FITS · 21.9 GB FREE78.3 tok/s
Qwen3 0.6B600Mbf16FITS · 20.9 GB FREE74.0 tok/s
Parakeet TDT 0.6B v2600Mbf16FITS · 21.8 GB FREE74.0 tok/s
Llama 3.2 1B Instruct1.2Bbf16FITS · 20.3 GB FREE35.8 tok/s
Whisper Large v31.6Bbf16FITS · 19.9 GB FREE28.6 tok/s
Qwen3 1.7B1.7Bbf16FITS · 18.7 GB FREE26.1 tok/s
Llama 3.2 3B Instruct3.2Bbf16FITS · 16.2 GB FREE13.8 tok/s
Qwen3 4B4Bq8FITS · 17.6 GB FREE20.9 tok/s
Qwen3 Embedding 4B4Bq8FITS · 17.6 GB FREE20.9 tok/s
Mistral 7B Instruct v0.37.3Bq8FITS · 14.2 GB FREE11.4 tok/s
Qwen2.5 7B Instruct7.6Bbf16FITS · 7.4 GB FREE5.8 tok/s
InternVL3 8B8Bq8FITS · 14.5 GB FREE10.4 tok/s
Llama 3.1 8B Instruct8.0Bbf16FITS · 5.9 GB FREE5.5 tok/s
Qwen3 8B8.2Bq8FITS · 13.1 GB FREE10.2 tok/s
Qwen2.5-VL 7B Instruct8.3Bq8FITS · 13.8 GB FREE10.1 tok/s
Gemma 2 9B IT9.2Bq8FITS · 13.2 GB FREE9.0 tok/s
Llama 3.2 11B Vision Instruct10.7Bq4FITS · 17.0 GB FREE14.8 tok/s
Qwen2.5 14B Instruct14.7Bq8FITS · 5.8 GB FREE5.7 tok/s
Qwen3 14B14.8Bq8FITS · 6.0 GB FREE5.6 tok/s
GPT-OSS 20B21Bq4FITS · 10.8 GB FREE43.9 tok/s
Codestral 22B v0.122.2Bq4FITS · 8.7 GB FREE7.1 tok/s
Mistral Small 24B Instruct 250124Bq4FITS · 8.2 GB FREE6.6 tok/s
Gemma 3 27B IT27Bq4FITS · 7.9 GB FREE5.8 tok/s
Gemma 2 27B IT27.2Bq4FITS · 7.7 GB FREE5.8 tok/s
Qwen3 30B-A3B Instruct 250730.5Bq4FITS · 5.1 GB FREE47.8 tok/s
Qwen3 Coder 30B-A3B Instruct30.5Bq4FITS · 5.1 GB FREE47.8 tok/s
Qwen2.5 32B Instruct32.5Bq4TIGHT · 2.6 GB FREE4.9 tok/s
Qwen2.5 Coder 32B Instruct32.5Bq4TIGHT · 2.6 GB FREE4.9 tok/s
DeepSeek R1 Distill Qwen 32B32.5Bq4TIGHT · 2.6 GB FREE4.9 tok/s
Qwen3 32B32.8Bq4TIGHT · 2.4 GB FREE4.8 tok/s
Qwen2.5-VL 32B Instruct33Bq4TIGHT · 2.3 GB FREE4.8 tok/s
Mixtral 8x7B Instruct v0.146.7BOVER
Llama 3.3 70B Instruct70.6BOVER
Qwen2.5 72B Instruct72.7BOVER
Llama 3.2 90B Vision Instruct88BOVER
GPT-OSS 120B116.8BOVER
Qwen3 235B-A22B235BOVER
DeepSeek V3671BOVER