Skip to content
mlx.app

Models

33 of 39 run on a M4 Pro with 48 GB at 8K context. Sorted by fit, then speed.

33 shown

Qwen3 30B-A3B Instruct 2507

Alibaba · 30.5B (3.3B active) · q6

FITS · 12.8 GB FREE

The best speed-to-quality tradeoff currently available on a 32GB Mac, because only 3.3B parameters activate per token even though the full 30B has to sit in memory. It feels closer to a dense 14B model in responsiveness while answering more like a 30B one.

~79.4 tok/s est.8K ctxcommercial ok

Qwen3 Coder 30B-A3B Instruct

Alibaba · 30.5B (3.3B active) · q6

FITS · 12.8 GB FREE

This is the current go-to local coding model for a 32GB Mac, trading some raw code-quality against Qwen2.5-Coder-32B for much faster generation. Its 262K context window means it can hold a real codebase's worth of files in one session, which the dense coder can't do as cheaply.

~79.4 tok/s est.8K ctxcommercial ok

Qwen2.5-VL 32B Instruct

Alibaba · 33B · q4

FITS · 17.7 GB FREE

The 32GB-Mac step up from the 7B vision model, with meaningfully better reasoning about multi-image and multi-step visual tasks. It's slower per image than the 7B, so batch document processing is where the extra memory cost pays off most.

~11.5 tok/s est.8K ctxcommercial ok

DeepSeek R1 Distill Qwen 32B

DeepSeek · 32.5B · q4

FITS · 18.0 GB FREE

Distilled from DeepSeek R1's reasoning traces onto a Qwen2.5 32B base, so on a 32GB Mac you get long chain-of-thought behavior without the 671B parent's memory bill. Expect verbose answers full of visible reasoning steps, which is either useful or annoying depending on the task.

~11.6 tok/s est.8K ctxcommercial ok

Mixtral 8x7B Instruct v0.1

Mistral AI · 46.7B (12.9B active) · q4

FITS · 11.1 GB FREE

One of the earliest MoE models to run well on Apple Silicon, needing a 32GB Mac for 4-bit despite having only 12.9B active parameters per token. It still runs faster than a dense 46B model would, but newer MoE designs like Qwen3-30B-A3B have mostly superseded it on quality.

~29.3 tok/s est.8K ctxcommercial ok

Qwen3 32B

Alibaba · 32.8B · q6

FITS · 9.6 GB FREE

Currently the strongest dense text model that a 32GB Mac can run at a usable quant, and its thinking mode holds up on genuinely hard reasoning benchmarks. If you have the memory budget, this outperforms Qwen3-30B-A3B on quality per token even though it's slower.

~8.0 tok/s est.8K ctxcommercial ok

GPT-OSS 20B

OpenAI · 21B (3.6B active) · q8

FITS · 15.7 GB FREE

OpenAI's first open-weight release since GPT-2, shipped natively in 4-bit MXFP4 so it already fits a 16GB Mac without further quantization. It uses OpenAI's harmony chat format and a configurable reasoning effort setting rather than a simple thinking toggle.

~55.7 tok/s est.8K ctxcommercial ok

Gemma 3 27B IT

Google · 27B · q8

FITS · 9.7 GB FREE

Fixes Gemma 2's biggest weakness by pushing context out to 128K and adding native image input. On a 32GB Mac at 4-bit it's one of the more well-rounded large models available for both text and pictures.

~7.4 tok/s est.8K ctxcommercial ok

Qwen2.5 32B Instruct

Alibaba · 32.5B · q4

FITS · 18.0 GB FREE

This is where a 32GB Mac starts to feel necessary rather than optional, and even then 4-bit leaves a thinner margin than the 24B-class models. Quality is a solid step above 14B on reasoning-heavy prompts, which is the whole reason to pay the memory cost.

~11.6 tok/s est.8K ctxcommercial ok

Qwen2.5 Coder 32B Instruct

Alibaba · 32.5B · q4

FITS · 18.0 GB FREE

For a while this was the best local coding model available, and on a 32GB Mac it still holds up well against newer general models. Qwen3-Coder-30B-A3B now beats it on speed for similar quality, but this dense model tends to be more consistent on tricky edge cases.

~11.6 tok/s est.8K ctxcommercial ok

Mistral Small 24B Instruct 2501

Mistral AI · 24B · q8

FITS · 11.6 GB FREE

Aimed squarely at the 32GB Mac tier and it earns the memory, with quality that tracks close to models twice its size on everyday tasks. It's fully Apache-licensed, which matters if you're shipping a commercial product on top of it.

~8.4 tok/s est.8K ctxcommercial ok

Gemma 2 27B IT

Google · 27.2B · q8

FITS · 9.5 GB FREE

Wants a 32GB Mac to run 4-bit without squeezing everything else out of memory. It's a strong writer for its size class, but the short 8K context means it's a poor fit for document-heavy work.

~7.4 tok/s est.8K ctxcommercial ok

Codestral 22B v0.1

Mistral AI · 22.2B · q8

FITS · 12.9 GB FREE

A code-specialist that fits a 32GB Mac comfortably at 4-bit and supports fill-in-the-middle completion, which general chat models handle poorly. The license is non-commercial though, so it's a personal-project tool rather than something to ship in a product.

~9.0 tok/s est.8K ctxrestricted

Qwen3 14B

Alibaba · 14.8B · q8

FITS · 21.3 GB FREE

Fits 4-bit on a 16GB Mac but is happier on 24GB where the OS isn't fighting it for memory. The reasoning gains over the 8B model are real on harder logic and math prompts, at roughly double the load size.

~13.5 tok/s est.8K ctxcommercial ok

Qwen2.5 14B Instruct

Alibaba · 14.7B · q8

FITS · 21.2 GB FREE

Needs a 16GB Mac at minimum for the 4-bit build, and a 24GB or 32GB machine is where it actually feels comfortable with other apps open. It's noticeably more capable than the 7B on reasoning chains without yet needing a 32GB-class machine.

~13.6 tok/s est.8K ctxcommercial ok

Qwen2.5-VL 7B Instruct

Alibaba · 8.3B · q8

FITS · 29.1 GB FREE

A capable everyday vision model that fits a 16GB Mac, able to read screenshots, charts, and dense documents rather than just describing photos. It also does basic video understanding and can point to bounding boxes, which most local vision models skip.

~24.1 tok/s est.8K ctxcommercial ok

Qwen3 8B

Alibaba · 8.2B · q8

FITS · 28.5 GB FREE

This is currently the best all-round 8B for a 16GB Mac, with the thinking-mode toggle giving you a manual quality/speed dial. At 4-bit it leaves enough memory free to keep a coding editor and a few Chrome tabs open.

~24.4 tok/s est.8K ctxcommercial ok

InternVL3 8B

OpenGVLab · 8B · q8

FITS · 29.9 GB FREE

A well-regarded open vision-language alternative to Qwen2.5-VL that fits the same 16GB tier. It was trained with native multi-image and video input in mind, which shows up in slightly better consistency across a sequence of frames.

~25.1 tok/s est.8K ctxcommercial ok

Llama 3.2 11B Vision Instruct

Meta · 10.7B · bf16

FITS · 17.0 GB FREE

Meta's cross-attention vision adapter bolted onto a Llama 3.1 8B backbone, fitting a 16GB Mac comfortably at 4-bit. It's solid for general photo description but noticeably weaker than Qwen2.5-VL on dense documents and charts.

~10.0 tok/s est.8K ctxcommercial ok

Gemma 2 9B IT

Google · 9.2B · q8

FITS · 28.6 GB FREE

A capable writer for a 16GB Mac, but the 8K context ceiling is dated next to Qwen3's 40K. Google's safety tuning also makes it noticeably more cautious in refusals than comparable models.

~21.7 tok/s est.8K ctxcommercial ok

Llama 3.1 8B Instruct

Meta · 8.0B · bf16

FITS · 21.3 GB FREE

A true 128K context window makes this a good choice for a 16GB Mac that needs to chew on long documents, not just chat. It's a generation behind on raw reasoning benchmarks but the long context still earns it a place.

~13.3 tok/s est.8K ctxcommercial ok

Mistral 7B Instruct v0.3

Mistral AI · 7.3B · q8

FITS · 29.6 GB FREE

Still a dependable, no-drama 7B for a 16GB machine, though newer Qwen3 models have mostly caught up on quality. Its main selling point at this point is a permissive license and function-calling support baked in.

~27.5 tok/s est.8K ctxcommercial ok

Qwen2.5 7B Instruct

Alibaba · 7.6B · bf16

FITS · 22.7 GB FREE

The default general-purpose pick for a 16GB Mac running 4-bit, with enough room left for a browser tab or two. It's a known quantity: broad training data, reliable instruction following, nothing flashy.

~14.0 tok/s est.8K ctxcommercial ok

Qwen3 4B

Alibaba · 4B · q8

FITS · 32.9 GB FREE

This is the sweet spot for a 16GB Mac that still wants headroom for apps. Punches noticeably above its size on coding and math thanks to the Qwen3 training recipe.

~50.1 tok/s est.8K ctxcommercial ok

Qwen3 Embedding 4B

Alibaba · 4B · q8

FITS · 32.9 GB FREE

Built on the Qwen3 4B backbone, this trades embedding model minimalism for MTEB-leaderboard-level retrieval quality, at the cost of needing a 16GB Mac rather than fitting anywhere. It also supports instruction-aware embeddings, so you can bias retrieval toward a task description at query time.

~50.1 tok/s est.8K ctxcommercial ok

Llama 3.2 3B Instruct

Meta · 3.2B · bf16

FITS · 31.5 GB FREE

A step up from the 1B that still fits comfortably on 8GB machines at 4-bit. It's the model most people should try first before assuming they need something bigger.

~33.2 tok/s est.8K ctxcommercial ok

Qwen3 1.7B

Alibaba · 1.7B · bf16

FITS · 34.1 GB FREE

A comfortable fit on any 8GB Mac and effectively free to run on 16GB. The thinking mode helps on logic puzzles but roughly doubles response time, so it's a tradeoff you feel.

~62.6 tok/s est.8K ctxcommercial ok

Llama 3.2 1B Instruct

Meta · 1.2B · bf16

FITS · 35.7 GB FREE

This one runs on almost anything with unified memory, including an 8GB Mac mini with room to spare. It answers fast but loses the thread on anything that needs more than a paragraph of reasoning.

~85.9 tok/s est.8K ctxcommercial ok

Whisper Large v3

OpenAI · 1.6B · bf16

FITS · 35.3 GB FREE

Runs comfortably on any Apple Silicon Mac and remains the most accurate open transcription model across accents and background noise. Real-time factor on an M-series chip is fast enough for near-live captioning, though the turbo variant trades a little accuracy for more speed.

~68.7 tok/s est.8K ctxcommercial ok

Qwen3 0.6B

Alibaba · 600M · bf16

FITS · 36.3 GB FREE

Qwen3's smallest model adds an optional thinking mode that the 0.5-class Qwen2.5 didn't have. It still can't hold much context, so keep prompts short and single-purpose.

~177 tok/s est.8K ctxcommercial ok

Parakeet TDT 0.6B v2

NVIDIA · 600M · bf16

FITS · 37.2 GB FREE

Tiny compared to Whisper Large but tuned for English speed, and it runs many times faster than real time even on an entry-level Mac. It's the right pick when you're transcribing your own English audio locally and don't need Whisper's language breadth.

~177 tok/s est.8K ctxcommercial ok

BGE-M3

BAAI · 567M · bf16

FITS · 37.3 GB FREE

A multilingual embedding model small enough to run on any Mac without noticing it in memory. It supports dense, sparse, and multi-vector retrieval in one model, which covers most local RAG setups without needing a second embedding pass.

~188 tok/s est.8K ctxcommercial ok

Qwen2.5 0.5B Instruct

Alibaba · 490M · bf16

FITS · 37.3 GB FREE

Small enough to load on a base 8GB MacBook Air alongside a browser and an IDE. Treat it as a utility model for routing and formatting, not as a reasoning partner.

~217 tok/s est.8K ctxcommercial ok