Whisper Large v3
OpenAI · 1.6B · up to 448 context · mit
Runs comfortably on any Apple Silicon Mac and remains the most accurate open transcription model across accents and background noise. Real-time factor on an M-series chip is fast enough for near-live captioning, though the turbo variant trades a little accuracy for more speed.
On a M4 Pro with 48 GB
macOS9.6 GBweights3.1 GBKV cache0.0 GBheadroom35.3 GB
- Weights
- 3.1 GB
- KV cache
- 0.0 GBestimated shape
- Usable RAM
- 38.4 GB
- Headroom
- 35.3 GB
- Generation
- 68.7 tok/sest.
- Prompt processing
- 583 tok/sest.
- First token (1K prompt)
- 1.8 sest.
Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.
Every quantization, on your Mac
| Quant | Weights | + KV | Verdict | Speed est. | Repo |
|---|---|---|---|---|---|
| q4 | 0.9 GB | 0.9 GB | FITS · 37.5 GB FREE | 244 tok/s | whisper-large-v3-mlx-4bit |
| q8 | 1.6 GB | 1.6 GB | FITS · 36.8 GB FREE | 129 tok/s | whisper-large-v3-mlx-8bit |
| bf16 | 3.1 GB | 3.1 GB | FITS · 35.3 GB FREE | 68.7 tok/s | whisper-large-v3-mlx |
We never host weights. Every link goes to Hugging Face.
Run it
Chat
mlx_lm.chat --model mlx-community/whisper-large-v3-mlxServe an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/whisper-large-v3-mlx --port 8080How far can you push the context
Which Macs run this
| Chip | Smallest RAM that fits | Speed est. |
|---|---|---|
| M1 | 8 GB | 15.8 tok/s |
| M1 Pro | 16 GB | 49.0 tok/s |
| M1 Max | 32 GB | 102 tok/s |
| M1 Ultra | 64 GB | 209 tok/s |
| M2 | 8 GB | 23.5 tok/s |
| M2 Pro | 16 GB | 49.7 tok/s |
| M2 Max | 32 GB | 103 tok/s |
| M2 Ultra | 64 GB | 212 tok/s |
| M3 | 8 GB | 23.5 tok/s |
| M3 Pro | 18 GB | 36.3 tok/s |
| M3 Max (14-core CPU) | 36 GB | 75.5 tok/s |
| M3 Max (16-core CPU) | 48 GB | 103 tok/s |
| M3 Ultra | 96 GB | 219 tok/s |
| M4 | 16 GB | 28.6 tok/s |
| M4 Pro | 24 GB | 68.7 tok/s |
| M4 Max (14-core CPU) | 36 GB | 107 tok/s |
| M4 Max (16-core CPU) | 48 GB | 144 tok/s |
| M5 | 16 GB | 37.0 tok/s |
| M5 Prounverified | 24 GB | 78.2 tok/s |
| M5 Maxunverified | 36 GB | 164 tok/s |