Skip to content
mlx.app

Gemma 3 27B IT

Google · 27B · up to 128K context · gemma

FITS · 9.7 GB FREE

Fixes Gemma 2's biggest weakness by pushing context out to 128K and adding native image input. On a 32GB Mac at 4-bit it's one of the more well-rounded large models available for both text and pictures.

On a M4 Pro with 48 GB

macOS9.6 GBweights28.7 GBKV cache0.0 GBheadroom9.7 GB
Weights
28.7 GB
KV cache
0.0 GBestimated shape
Usable RAM
38.4 GB
Headroom
9.7 GB
Generation
7.4 tok/sest.
Prompt processing
33.4 tok/sest.
First token (1K prompt)
30.6 sest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q415.2 GB15.2 GBFITS · 23.2 GB FREE14.0 tok/sgemma-3-27b-it-4bit
q828.7 GB28.7 GBFITS · 9.7 GB FREE7.4 tok/sgemma-3-27b-it-8bit
bf1654.0 GB54.0 GBOVER BY 15.6 GB3.9 tok/sgemma-3-27b-it-bf16

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/gemma-3-27b-it-8bit
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/gemma-3-27b-it-8bit --port 8080

How far can you push the context

2K tokensFITS · 9.7 GB FREE

cache 0.0 GB

8K tokensFITS · 9.7 GB FREE

cache 0.0 GB

32K tokensFITS · 9.7 GB FREE

cache 0.0 GB

128K tokensFITS · 9.7 GB FREE

cache 0.0 GB

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M1no configuration
M1 Prono configuration
M1 Max64 GB11.0 tok/s
M1 Ultra64 GB22.6 tok/s
M2no configuration
M2 Prono configuration
M2 Max64 GB11.2 tok/s
M2 Ultra64 GB22.9 tok/s
M3no configuration
M3 Prono configuration
M3 Max (14-core CPU)no configuration
M3 Max (16-core CPU)48 GB11.2 tok/s
M3 Ultra96 GB23.7 tok/s
M4no configuration
M4 Pro48 GB7.4 tok/s
M4 Max (14-core CPU)48 GB11.6 tok/s
M4 Max (16-core CPU)48 GB15.6 tok/s
M5no configuration
M5 Prounverified48 GB8.5 tok/s
M5 Maxunverified48 GB17.8 tok/s