M2 Max · 96 GB
400 GB/s memory bandwidth · 76.8 GB usable for models · 19.2 GB left to macOS
The short answer
36 models run comfortably and 1 more will run tight at 8K context. The largest that fits is Llama 3.2 90B Vision Instruct at q4, generating around 6.5 tokens per second.
macOS19.2 GBweights49.5 GBKV cache0.0 GBheadroom27.3 GB
Every model on this Mac
| Model | Size | Best quant | Verdict | Speed est. |
|---|---|---|---|---|
| Qwen2.5 0.5B Instruct | 490M | bf16 | FITS · 75.7 GB FREE | 327 tok/s |
| BGE-M3 | 567M | bf16 | FITS · 75.7 GB FREE | 282 tok/s |
| Qwen3 0.6B | 600M | bf16 | FITS · 74.7 GB FREE | 267 tok/s |
| Parakeet TDT 0.6B v2 | 600M | bf16 | FITS · 75.6 GB FREE | 267 tok/s |
| Llama 3.2 1B Instruct | 1.2B | bf16 | FITS · 74.1 GB FREE | 129 tok/s |
| Whisper Large v3 | 1.6B | bf16 | FITS · 73.7 GB FREE | 103 tok/s |
| Qwen3 1.7B | 1.7B | bf16 | FITS · 72.5 GB FREE | 94.1 tok/s |
| Llama 3.2 3B Instruct | 3.2B | bf16 | FITS · 69.9 GB FREE | 49.8 tok/s |
| Qwen3 4B | 4B | q8 | FITS · 71.3 GB FREE | 75.3 tok/s |
| Qwen3 Embedding 4B | 4B | q8 | FITS · 71.3 GB FREE | 75.3 tok/s |
| Mistral 7B Instruct v0.3 | 7.3B | q8 | FITS · 68.0 GB FREE | 41.3 tok/s |
| Qwen2.5 7B Instruct | 7.6B | bf16 | FITS · 61.1 GB FREE | 21.1 tok/s |
| InternVL3 8B | 8B | q8 | FITS · 68.3 GB FREE | 37.6 tok/s |
| Llama 3.1 8B Instruct | 8.0B | bf16 | FITS · 59.7 GB FREE | 19.9 tok/s |
| Qwen3 8B | 8.2B | q8 | FITS · 66.9 GB FREE | 36.7 tok/s |
| Qwen2.5-VL 7B Instruct | 8.3B | q8 | FITS · 67.5 GB FREE | 36.3 tok/s |
| Gemma 2 9B IT | 9.2B | q8 | FITS · 67.0 GB FREE | 32.6 tok/s |
| Llama 3.2 11B Vision Instruct | 10.7B | bf16 | FITS · 55.4 GB FREE | 15.0 tok/s |
| Qwen2.5 14B Instruct | 14.7B | q8 | FITS · 59.6 GB FREE | 20.5 tok/s |
| Qwen3 14B | 14.8B | q8 | FITS · 59.7 GB FREE | 20.3 tok/s |
| GPT-OSS 20B | 21B | q8 | FITS · 54.1 GB FREE | 83.7 tok/s |
| Codestral 22B v0.1 | 22.2B | q8 | FITS · 51.3 GB FREE | 13.6 tok/s |
| Mistral Small 24B Instruct 2501 | 24B | q8 | FITS · 50.0 GB FREE | 12.5 tok/s |
| Gemma 3 27B IT | 27B | bf16 | FITS · 22.8 GB FREE | 5.9 tok/s |
| Gemma 2 27B IT | 27.2B | q8 | FITS · 47.9 GB FREE | 11.1 tok/s |
| Qwen3 30B-A3B Instruct 2507 | 30.5B | q8 | FITS · 43.6 GB FREE | 91.3 tok/s |
| Qwen3 Coder 30B-A3B Instruct | 30.5B | q8 | FITS · 43.6 GB FREE | 91.3 tok/s |
| Qwen2.5 32B Instruct | 32.5B | q8 | FITS · 40.1 GB FREE | 9.3 tok/s |
| Qwen2.5 Coder 32B Instruct | 32.5B | q8 | FITS · 40.1 GB FREE | 9.3 tok/s |
| DeepSeek R1 Distill Qwen 32B | 32.5B | q8 | FITS · 40.1 GB FREE | 9.3 tok/s |
| Qwen3 32B | 32.8B | q8 | FITS · 39.8 GB FREE | 9.2 tok/s |
| Qwen2.5-VL 32B Instruct | 33B | q8 | FITS · 39.6 GB FREE | 9.1 tok/s |
| Mixtral 8x7B Instruct v0.1 | 46.7B | q8 | FITS · 26.1 GB FREE | 23.3 tok/s |
| Llama 3.3 70B Instruct | 70.6B | q4 | FITS · 34.4 GB FREE | 8.1 tok/s |
| Qwen2.5 72B Instruct | 72.7B | q4 | FITS · 33.2 GB FREE | 7.8 tok/s |
| Llama 3.2 90B Vision Instruct | 88B | q4 | FITS · 27.3 GB FREE | 6.5 tok/s |
| GPT-OSS 120B | 116.8B | q4 | TIGHT · 10.5 GB FREE | 112 tok/s |
| Qwen3 235B-A22B | 235B | — | OVER | — |
| DeepSeek V3 | 671B | — | OVER | — |