Skip to content
mlx.app

M3 Max (16-core CPU) · 96 GB

400 GB/s memory bandwidth · 76.8 GB usable for models · 19.2 GB left to macOS

The short answer

36 models run comfortably and 1 more will run tight at 8K context. The largest that fits is Llama 3.2 90B Vision Instruct at q4, generating around 6.5 tokens per second.

macOS19.2 GBweights49.5 GBKV cache0.0 GBheadroom27.3 GB

Every model on this Mac

ModelSizeBest quantVerdictSpeed est.
Qwen2.5 0.5B Instruct490Mbf16FITS · 75.7 GB FREE327 tok/s
BGE-M3567Mbf16FITS · 75.7 GB FREE282 tok/s
Qwen3 0.6B600Mbf16FITS · 74.7 GB FREE267 tok/s
Parakeet TDT 0.6B v2600Mbf16FITS · 75.6 GB FREE267 tok/s
Llama 3.2 1B Instruct1.2Bbf16FITS · 74.1 GB FREE129 tok/s
Whisper Large v31.6Bbf16FITS · 73.7 GB FREE103 tok/s
Qwen3 1.7B1.7Bbf16FITS · 72.5 GB FREE94.1 tok/s
Llama 3.2 3B Instruct3.2Bbf16FITS · 69.9 GB FREE49.8 tok/s
Qwen3 4B4Bq8FITS · 71.3 GB FREE75.3 tok/s
Qwen3 Embedding 4B4Bq8FITS · 71.3 GB FREE75.3 tok/s
Mistral 7B Instruct v0.37.3Bq8FITS · 68.0 GB FREE41.3 tok/s
Qwen2.5 7B Instruct7.6Bbf16FITS · 61.1 GB FREE21.1 tok/s
InternVL3 8B8Bq8FITS · 68.3 GB FREE37.6 tok/s
Llama 3.1 8B Instruct8.0Bbf16FITS · 59.7 GB FREE19.9 tok/s
Qwen3 8B8.2Bq8FITS · 66.9 GB FREE36.7 tok/s
Qwen2.5-VL 7B Instruct8.3Bq8FITS · 67.5 GB FREE36.3 tok/s
Gemma 2 9B IT9.2Bq8FITS · 67.0 GB FREE32.6 tok/s
Llama 3.2 11B Vision Instruct10.7Bbf16FITS · 55.4 GB FREE15.0 tok/s
Qwen2.5 14B Instruct14.7Bq8FITS · 59.6 GB FREE20.5 tok/s
Qwen3 14B14.8Bq8FITS · 59.7 GB FREE20.3 tok/s
GPT-OSS 20B21Bq8FITS · 54.1 GB FREE83.7 tok/s
Codestral 22B v0.122.2Bq8FITS · 51.3 GB FREE13.6 tok/s
Mistral Small 24B Instruct 250124Bq8FITS · 50.0 GB FREE12.5 tok/s
Gemma 3 27B IT27Bbf16FITS · 22.8 GB FREE5.9 tok/s
Gemma 2 27B IT27.2Bq8FITS · 47.9 GB FREE11.1 tok/s
Qwen3 30B-A3B Instruct 250730.5Bq8FITS · 43.6 GB FREE91.3 tok/s
Qwen3 Coder 30B-A3B Instruct30.5Bq8FITS · 43.6 GB FREE91.3 tok/s
Qwen2.5 32B Instruct32.5Bq8FITS · 40.1 GB FREE9.3 tok/s
Qwen2.5 Coder 32B Instruct32.5Bq8FITS · 40.1 GB FREE9.3 tok/s
DeepSeek R1 Distill Qwen 32B32.5Bq8FITS · 40.1 GB FREE9.3 tok/s
Qwen3 32B32.8Bq8FITS · 39.8 GB FREE9.2 tok/s
Qwen2.5-VL 32B Instruct33Bq8FITS · 39.6 GB FREE9.1 tok/s
Mixtral 8x7B Instruct v0.146.7Bq8FITS · 26.1 GB FREE23.3 tok/s
Llama 3.3 70B Instruct70.6Bq4FITS · 34.4 GB FREE8.1 tok/s
Qwen2.5 72B Instruct72.7Bq4FITS · 33.2 GB FREE7.8 tok/s
Llama 3.2 90B Vision Instruct88Bq4FITS · 27.3 GB FREE6.5 tok/s
GPT-OSS 120B116.8Bq4TIGHT · 10.5 GB FREE112 tok/s
Qwen3 235B-A22B235BOVER
DeepSeek V3671BOVER