Skip to content
mlx.app

Gemma 2 9B IT

Google · 9.2B · up to 8K context · gemma

FITS · 28.6 GB FREE

A capable writer for a 16GB Mac, but the 8K context ceiling is dated next to Qwen3's 40K. Google's safety tuning also makes it noticeably more cautious in refusals than comparable models.

On a M4 Pro with 48 GB

macOS9.6 GBweights9.8 GBKV cache0.0 GBheadroom28.6 GB
Weights
9.8 GB
KV cache
0.0 GBestimated shape
Usable RAM
38.4 GB
Headroom
28.6 GB
Generation
21.7 tok/sest.
Prompt processing
97.7 tok/sest.
First token (1K prompt)
10.5 sest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q45.2 GB5.2 GBFITS · 33.2 GB FREE41.0 tok/sgemma-2-9b-it-4bit
q89.8 GB9.8 GBFITS · 28.6 GB FREE21.7 tok/sgemma-2-9b-it-8bit

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/gemma-2-9b-it-8bit
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/gemma-2-9b-it-8bit --port 8080

How far can you push the context

2K tokensFITS · 28.6 GB FREE

cache 0.0 GB

8K tokensFITS · 28.6 GB FREE

cache 0.0 GB

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M116 GB5.0 tok/s
M1 Pro16 GB15.5 tok/s
M1 Max32 GB32.2 tok/s
M1 Ultra64 GB66.0 tok/s
M216 GB7.4 tok/s
M2 Pro16 GB15.7 tok/s
M2 Max32 GB32.6 tok/s
M2 Ultra64 GB66.8 tok/s
M316 GB7.4 tok/s
M3 Pro18 GB11.5 tok/s
M3 Max (14-core CPU)36 GB23.8 tok/s
M3 Max (16-core CPU)48 GB32.6 tok/s
M3 Ultra96 GB69.2 tok/s
M416 GB9.0 tok/s
M4 Pro24 GB21.7 tok/s
M4 Max (14-core CPU)36 GB33.8 tok/s
M4 Max (16-core CPU)48 GB45.6 tok/s
M516 GB11.7 tok/s
M5 Prounverified24 GB24.7 tok/s
M5 Maxunverified36 GB51.9 tok/s