Skip to content
mlx.app

Parakeet TDT 0.6B v2

NVIDIA · 600M · up to 0 context · cc-by-4.0

FITS · 37.2 GB FREE

Tiny compared to Whisper Large but tuned for English speed, and it runs many times faster than real time even on an entry-level Mac. It's the right pick when you're transcribing your own English audio locally and don't need Whisper's language breadth.

On a M4 Pro with 48 GB

macOS9.6 GBweights1.2 GBKV cache0.0 GBheadroom37.2 GB
Weights
1.2 GB
KV cache
0.0 GBestimated shape
Usable RAM
38.4 GB
Headroom
37.2 GB
Generation
177 tok/sest.
Prompt processing
1505 tok/sest.
First token (1K prompt)
680 msest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
bf161.2 GB1.2 GBFITS · 37.2 GB FREE177 tok/sparakeet-tdt-0.6b-v2

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/parakeet-tdt-0.6b-v2
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/parakeet-tdt-0.6b-v2 --port 8080

How far can you push the context

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M18 GB40.8 tok/s
M1 Pro16 GB127 tok/s
M1 Max32 GB263 tok/s
M1 Ultra64 GB540 tok/s
M28 GB60.8 tok/s
M2 Pro16 GB128 tok/s
M2 Max32 GB267 tok/s
M2 Ultra64 GB547 tok/s
M38 GB60.8 tok/s
M3 Pro18 GB93.8 tok/s
M3 Max (14-core CPU)36 GB195 tok/s
M3 Max (16-core CPU)48 GB267 tok/s
M3 Ultra96 GB566 tok/s
M416 GB74.0 tok/s
M4 Pro24 GB177 tok/s
M4 Max (14-core CPU)36 GB277 tok/s
M4 Max (16-core CPU)48 GB373 tok/s
M516 GB95.6 tok/s
M5 Prounverified24 GB202 tok/s
M5 Maxunverified36 GB425 tok/s