Skip to content
mlx.app

M3 Max (16-core CPU) · 128 GB

400 GB/s memory bandwidth · 108.8 GB usable for models · 19.2 GB left to macOS

The short answer

37 models run comfortably and 0 more will run tight at 8K context. The largest that fits is GPT-OSS 120B at q4, generating around 112 tokens per second.

macOS19.2 GBweights65.7 GBKV cache0.6 GBheadroom42.5 GB

Every model on this Mac

ModelSizeBest quantVerdictSpeed est.
Qwen2.5 0.5B Instruct490Mbf16FITS · 107.7 GB FREE327 tok/s
BGE-M3567Mbf16FITS · 107.7 GB FREE282 tok/s
Qwen3 0.6B600Mbf16FITS · 106.7 GB FREE267 tok/s
Parakeet TDT 0.6B v2600Mbf16FITS · 107.6 GB FREE267 tok/s
Llama 3.2 1B Instruct1.2Bbf16FITS · 106.1 GB FREE129 tok/s
Whisper Large v31.6Bbf16FITS · 105.7 GB FREE103 tok/s
Qwen3 1.7B1.7Bbf16FITS · 104.5 GB FREE94.1 tok/s
Llama 3.2 3B Instruct3.2Bbf16FITS · 101.9 GB FREE49.8 tok/s
Qwen3 4B4Bq8FITS · 103.3 GB FREE75.3 tok/s
Qwen3 Embedding 4B4Bq8FITS · 103.3 GB FREE75.3 tok/s
Mistral 7B Instruct v0.37.3Bq8FITS · 100.0 GB FREE41.3 tok/s
Qwen2.5 7B Instruct7.6Bbf16FITS · 93.1 GB FREE21.1 tok/s
InternVL3 8B8Bq8FITS · 100.3 GB FREE37.6 tok/s
Llama 3.1 8B Instruct8.0Bbf16FITS · 91.7 GB FREE19.9 tok/s
Qwen3 8B8.2Bq8FITS · 98.9 GB FREE36.7 tok/s
Qwen2.5-VL 7B Instruct8.3Bq8FITS · 99.5 GB FREE36.3 tok/s
Gemma 2 9B IT9.2Bq8FITS · 99.0 GB FREE32.6 tok/s
Llama 3.2 11B Vision Instruct10.7Bbf16FITS · 87.4 GB FREE15.0 tok/s
Qwen2.5 14B Instruct14.7Bq8FITS · 91.6 GB FREE20.5 tok/s
Qwen3 14B14.8Bq8FITS · 91.7 GB FREE20.3 tok/s
GPT-OSS 20B21Bq8FITS · 86.1 GB FREE83.7 tok/s
Codestral 22B v0.122.2Bq8FITS · 83.3 GB FREE13.6 tok/s
Mistral Small 24B Instruct 250124Bq8FITS · 82.0 GB FREE12.5 tok/s
Gemma 3 27B IT27Bbf16FITS · 54.8 GB FREE5.9 tok/s
Gemma 2 27B IT27.2Bq8FITS · 79.9 GB FREE11.1 tok/s
Qwen3 30B-A3B Instruct 250730.5Bq8FITS · 75.6 GB FREE91.3 tok/s
Qwen3 Coder 30B-A3B Instruct30.5Bq8FITS · 75.6 GB FREE91.3 tok/s
Qwen2.5 32B Instruct32.5Bq8FITS · 72.1 GB FREE9.3 tok/s
Qwen2.5 Coder 32B Instruct32.5Bq8FITS · 72.1 GB FREE9.3 tok/s
DeepSeek R1 Distill Qwen 32B32.5Bq8FITS · 72.1 GB FREE9.3 tok/s
Qwen3 32B32.8Bq8FITS · 71.8 GB FREE9.2 tok/s
Qwen2.5-VL 32B Instruct33Bq8FITS · 71.6 GB FREE9.1 tok/s
Mixtral 8x7B Instruct v0.146.7Bq8FITS · 58.1 GB FREE23.3 tok/s
Llama 3.3 70B Instruct70.6Bq8FITS · 31.1 GB FREE4.3 tok/s
Qwen2.5 72B Instruct72.7Bq8FITS · 28.9 GB FREE4.1 tok/s
Llama 3.2 90B Vision Instruct88Bq4FITS · 59.3 GB FREE6.5 tok/s
GPT-OSS 120B116.8Bq4FITS · 42.5 GB FREE112 tok/s
Qwen3 235B-A22B235BOVER
DeepSeek V3671BOVER