Skip to content
mlx.app

Whisper Large v3

OpenAI · 1.6B · up to 448 context · mit

FITS · 35.3 GB FREE

Runs comfortably on any Apple Silicon Mac and remains the most accurate open transcription model across accents and background noise. Real-time factor on an M-series chip is fast enough for near-live captioning, though the turbo variant trades a little accuracy for more speed.

On a M4 Pro with 48 GB

macOS9.6 GBweights3.1 GBKV cache0.0 GBheadroom35.3 GB
Weights
3.1 GB
KV cache
0.0 GBestimated shape
Usable RAM
38.4 GB
Headroom
35.3 GB
Generation
68.7 tok/sest.
Prompt processing
583 tok/sest.
First token (1K prompt)
1.8 sest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q40.9 GB0.9 GBFITS · 37.5 GB FREE244 tok/swhisper-large-v3-mlx-4bit
q81.6 GB1.6 GBFITS · 36.8 GB FREE129 tok/swhisper-large-v3-mlx-8bit
bf163.1 GB3.1 GBFITS · 35.3 GB FREE68.7 tok/swhisper-large-v3-mlx

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/whisper-large-v3-mlx
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/whisper-large-v3-mlx --port 8080

How far can you push the context

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M18 GB15.8 tok/s
M1 Pro16 GB49.0 tok/s
M1 Max32 GB102 tok/s
M1 Ultra64 GB209 tok/s
M28 GB23.5 tok/s
M2 Pro16 GB49.7 tok/s
M2 Max32 GB103 tok/s
M2 Ultra64 GB212 tok/s
M38 GB23.5 tok/s
M3 Pro18 GB36.3 tok/s
M3 Max (14-core CPU)36 GB75.5 tok/s
M3 Max (16-core CPU)48 GB103 tok/s
M3 Ultra96 GB219 tok/s
M416 GB28.6 tok/s
M4 Pro24 GB68.7 tok/s
M4 Max (14-core CPU)36 GB107 tok/s
M4 Max (16-core CPU)48 GB144 tok/s
M516 GB37.0 tok/s
M5 Prounverified24 GB78.2 tok/s
M5 Maxunverified36 GB164 tok/s