Skip to content
mlx.app

M3 Ultra · 512 GB

819 GB/s memory bandwidth · 435.2 GB usable for models · 76.8 GB left to macOS

The short answer

38 models run comfortably and 1 more will run tight at 8K context. The largest that fits is Qwen3 235B-A22B at q8, generating around 29.1 tokens per second.

macOS76.8 GBweights249.7 GBKV cache1.6 GBheadroom183.9 GB

Every model on this Mac

ModelSizeBest quantVerdictSpeed est.
Qwen2.5 0.5B Instruct490Mbf16FITS · 434.1 GB FREE694 tok/s
BGE-M3567Mbf16FITS · 434.1 GB FREE599 tok/s
Qwen3 0.6B600Mbf16FITS · 433.1 GB FREE566 tok/s
Parakeet TDT 0.6B v2600Mbf16FITS · 434.0 GB FREE566 tok/s
Llama 3.2 1B Instruct1.2Bbf16FITS · 432.5 GB FREE274 tok/s
Whisper Large v31.6Bbf16FITS · 432.1 GB FREE219 tok/s
Qwen3 1.7B1.7Bbf16FITS · 430.9 GB FREE200 tok/s
Llama 3.2 3B Instruct3.2Bbf16FITS · 428.3 GB FREE106 tok/s
Qwen3 4B4Bq8FITS · 429.7 GB FREE160 tok/s
Qwen3 Embedding 4B4Bq8FITS · 429.7 GB FREE160 tok/s
Mistral 7B Instruct v0.37.3Bq8FITS · 426.4 GB FREE87.6 tok/s
Qwen2.5 7B Instruct7.6Bbf16FITS · 419.5 GB FREE44.7 tok/s
InternVL3 8B8Bq8FITS · 426.7 GB FREE80.0 tok/s
Llama 3.1 8B Instruct8.0Bbf16FITS · 418.1 GB FREE42.3 tok/s
Qwen3 8B8.2Bq8FITS · 425.3 GB FREE78.0 tok/s
Qwen2.5-VL 7B Instruct8.3Bq8FITS · 425.9 GB FREE77.1 tok/s
Gemma 2 9B IT9.2Bq8FITS · 425.4 GB FREE69.2 tok/s
Llama 3.2 11B Vision Instruct10.7Bbf16FITS · 413.8 GB FREE31.8 tok/s
Qwen2.5 14B Instruct14.7Bq8FITS · 418.0 GB FREE43.5 tok/s
Qwen3 14B14.8Bq8FITS · 418.1 GB FREE43.2 tok/s
GPT-OSS 20B21Bq8FITS · 412.5 GB FREE178 tok/s
Codestral 22B v0.122.2Bq8FITS · 409.7 GB FREE28.8 tok/s
Mistral Small 24B Instruct 250124Bq8FITS · 408.4 GB FREE26.7 tok/s
Gemma 3 27B IT27Bbf16FITS · 381.2 GB FREE12.6 tok/s
Gemma 2 27B IT27.2Bq8FITS · 406.3 GB FREE23.5 tok/s
Qwen3 30B-A3B Instruct 250730.5Bq8FITS · 402.0 GB FREE194 tok/s
Qwen3 Coder 30B-A3B Instruct30.5Bq8FITS · 402.0 GB FREE194 tok/s
Qwen2.5 32B Instruct32.5Bq8FITS · 398.5 GB FREE19.7 tok/s
Qwen2.5 Coder 32B Instruct32.5Bq8FITS · 398.5 GB FREE19.7 tok/s
DeepSeek R1 Distill Qwen 32B32.5Bq8FITS · 398.5 GB FREE19.7 tok/s
Qwen3 32B32.8Bq8FITS · 398.2 GB FREE19.5 tok/s
Qwen2.5-VL 32B Instruct33Bq8FITS · 398.0 GB FREE19.4 tok/s
Mixtral 8x7B Instruct v0.146.7Bq8FITS · 384.5 GB FREE49.6 tok/s
Llama 3.3 70B Instruct70.6Bq8FITS · 357.5 GB FREE9.1 tok/s
Qwen2.5 72B Instruct72.7Bq8FITS · 355.3 GB FREE8.8 tok/s
Llama 3.2 90B Vision Instruct88Bq4FITS · 385.7 GB FREE13.7 tok/s
GPT-OSS 120B116.8Bq8FITS · 310.5 GB FREE125 tok/s
Qwen3 235B-A22B235Bq8FITS · 183.9 GB FREE29.1 tok/s
DeepSeek V3671Bq4TIGHT · 57.8 GB FREE32.7 tok/s