Skip to content
mlx.app

Buying a Mac for local AI

6 minute read

Which spec actually changes your experience, and which upgrades are money you will not notice.

Memory decides what you can run

This is the only decision that closes doors permanently. Apple silicon memory is soldered; you cannot add more later. Everything else on the spec sheet affects how pleasant a model is to use, but memory decides whether it runs at all.

Rough guide, allowing for macOS overhead and a realistic context:

  • 16 GB: up to about 8B at 4-bit. Workable, and everything above it is out.
  • 24 GB: 14B at 4-bit comfortably. A good floor if you are serious.
  • 32 GB: 14B at 8-bit, or 32B at 4-bit if you keep context short.
  • 48 GB: 32B at 4-bit comfortably, with room for a long conversation.
  • 64 GB: 70B at 4-bit, tightly. Large MoE models become genuinely good here.
  • 128 GB+: 70B with room to breathe, or very large MoE models.

If you are choosing between more memory and a faster chip at the same price, take the memory.

Bandwidth decides how fast

Decode speed is bandwidth divided by model size. Within a generation, Pro is roughly double base, Max roughly double Pro, and Ultra roughly double Max. Those multiples show up almost directly in tokens per second.

Check the specific SKU rather than the family name: within some generations, the base chip's bandwidth varies between configurations, and a lower-bandwidth part is proportionally slower at generation regardless of how many GPU cores it has.

GPU cores decide prefill

More cores mainly speed up prompt processing. If you paste long documents or work over a codebase, that is the difference between waiting two seconds and waiting eight. If you type short questions, you will barely notice. It is a real upgrade for a specific workload, not a general one.

The Ultra question

An Ultra gives you the highest bandwidth Apple sells and the largest memory configurations, which together make 70B-class dense models comfortable. If that is your requirement, nothing else in the lineup substitutes. If your ceiling is 32B, an Ultra is a large amount of money for headroom you will not use.

Laptop or desktop

Desktops sustain performance indefinitely. Laptops throttle: a long generation on battery will be slower than the same job plugged in on a desk, and thermals matter more than most benchmark posts admit. If inference is a daily job rather than an occasional one, a Studio or a Mac mini is better value per token.

Buying used

Older generations are competitive because bandwidth and memory age well. A previous-generation Max with 64 GB will often outperform a current-generation Pro with 36 GB on anything large, for less money. Compare the actual bandwidth figure and the actual memory, not the generation number.

The short version

Buy the most memory you can afford, then the highest bandwidth tier you can afford, then stop. Use the hardware pages on this site to see exactly which models each configuration unlocks before you spend anything.