Parakeet TDT 0.6B v2
NVIDIA · 600M · up to 0 context · cc-by-4.0
Tiny compared to Whisper Large but tuned for English speed, and it runs many times faster than real time even on an entry-level Mac. It's the right pick when you're transcribing your own English audio locally and don't need Whisper's language breadth.
On a M4 Pro with 48 GB
macOS9.6 GBweights1.2 GBKV cache0.0 GBheadroom37.2 GB
- Weights
- 1.2 GB
- KV cache
- 0.0 GBestimated shape
- Usable RAM
- 38.4 GB
- Headroom
- 37.2 GB
- Generation
- 177 tok/sest.
- Prompt processing
- 1505 tok/sest.
- First token (1K prompt)
- 680 msest.
Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.
Every quantization, on your Mac
| Quant | Weights | + KV | Verdict | Speed est. | Repo |
|---|---|---|---|---|---|
| bf16 | 1.2 GB | 1.2 GB | FITS · 37.2 GB FREE | 177 tok/s | parakeet-tdt-0.6b-v2 |
We never host weights. Every link goes to Hugging Face.
Run it
Chat
mlx_lm.chat --model mlx-community/parakeet-tdt-0.6b-v2Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/parakeet-tdt-0.6b-v2 --port 8080How far can you push the context
Which Macs run this
| Chip | Smallest RAM that fits | Speed est. |
|---|---|---|
| M1 | 8 GB | 40.8 tok/s |
| M1 Pro | 16 GB | 127 tok/s |
| M1 Max | 32 GB | 263 tok/s |
| M1 Ultra | 64 GB | 540 tok/s |
| M2 | 8 GB | 60.8 tok/s |
| M2 Pro | 16 GB | 128 tok/s |
| M2 Max | 32 GB | 267 tok/s |
| M2 Ultra | 64 GB | 547 tok/s |
| M3 | 8 GB | 60.8 tok/s |
| M3 Pro | 18 GB | 93.8 tok/s |
| M3 Max (14-core CPU) | 36 GB | 195 tok/s |
| M3 Max (16-core CPU) | 48 GB | 267 tok/s |
| M3 Ultra | 96 GB | 566 tok/s |
| M4 | 16 GB | 74.0 tok/s |
| M4 Pro | 24 GB | 177 tok/s |
| M4 Max (14-core CPU) | 36 GB | 277 tok/s |
| M4 Max (16-core CPU) | 48 GB | 373 tok/s |
| M5 | 16 GB | 95.6 tok/s |
| M5 Prounverified | 24 GB | 202 tok/s |
| M5 Maxunverified | 36 GB | 425 tok/s |