Skip to content
mlx.app

Gemma 2 27B IT

Google · 27.2B · up to 8K context · gemma

FITS · 9.5 GB FREE

Wants a 32GB Mac to run 4-bit without squeezing everything else out of memory. It's a strong writer for its size class, but the short 8K context means it's a poor fit for document-heavy work.

On a M4 Pro with 48 GB

macOS9.6 GBweights28.9 GBKV cache0.0 GBheadroom9.5 GB
Weights
28.9 GB
KV cache
0.0 GBestimated shape
Usable RAM
38.4 GB
Headroom
9.5 GB
Generation
7.4 tok/sest.
Prompt processing
33.2 tok/sest.
First token (1K prompt)
30.8 sest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q415.3 GB15.3 GBFITS · 23.1 GB FREE13.9 tok/sgemma-2-27b-it-4bit
q828.9 GB28.9 GBFITS · 9.5 GB FREE7.4 tok/sgemma-2-27b-it-8bit

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/gemma-2-27b-it-8bit
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/gemma-2-27b-it-8bit --port 8080

How far can you push the context

2K tokensFITS · 9.5 GB FREE

cache 0.0 GB

8K tokensFITS · 9.5 GB FREE

cache 0.0 GB

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M1no configuration
M1 Prono configuration
M1 Max64 GB10.9 tok/s
M1 Ultra64 GB22.4 tok/s
M2no configuration
M2 Prono configuration
M2 Max64 GB11.1 tok/s
M2 Ultra64 GB22.7 tok/s
M3no configuration
M3 Prono configuration
M3 Max (14-core CPU)no configuration
M3 Max (16-core CPU)48 GB11.1 tok/s
M3 Ultra96 GB23.5 tok/s
M4no configuration
M4 Pro48 GB7.4 tok/s
M4 Max (14-core CPU)48 GB11.5 tok/s
M4 Max (16-core CPU)48 GB15.5 tok/s
M5no configuration
M5 Prounverified48 GB8.4 tok/s
M5 Maxunverified48 GB17.6 tok/s