Ollama
The `ollama run` command that made local models a one-liner.
When to use it
Use this when you want the easiest pull-and-serve workflow with a huge model library and an OpenAI-compatible API, and don't need MLX-level Apple silicon speed.
Alternatives in Servers
MLX Omni Server
An OpenAI-compatible server that's actually running MLX underneath.
LM Studio MLX Engine
LM Studio's own MLX backend, open-sourced separately from the app.
Open WebUI
A ChatGPT-shaped front end you point at any local backend.
exo
Turns a pile of Macs (and phones) into one inference cluster.