Skip to content
mlx.app

BGE-M3

BAAI · 567M · up to 8K context · mit

FITS · 37.3 GB FREE

A multilingual embedding model small enough to run on any Mac without noticing it in memory. It supports dense, sparse, and multi-vector retrieval in one model, which covers most local RAG setups without needing a second embedding pass.

On a M4 Pro with 48 GB

macOS9.6 GBweights1.1 GBKV cache0.0 GBheadroom37.3 GB
Weights
1.1 GB
KV cache
0.0 GBestimated shape
Usable RAM
38.4 GB
Headroom
37.3 GB
Generation
188 tok/sest.
Prompt processing
1593 tok/sest.
First token (1K prompt)
643 msest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q40.3 GB0.3 GBFITS · 38.1 GB FREE668 tok/sbge-m3-mlx-4bit
q60.5 GB0.5 GBFITS · 37.9 GB FREE462 tok/sbge-m3-mlx-6bit
q80.6 GB0.6 GBFITS · 37.8 GB FREE353 tok/sbge-m3-mlx-8bit
bf161.1 GB1.1 GBFITS · 37.3 GB FREE188 tok/sbge-m3-mlx-fp16

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/bge-m3-mlx-fp16
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/bge-m3-mlx-fp16 --port 8080

How far can you push the context

2K tokensFITS · 37.3 GB FREE

cache 0.0 GB

8K tokensFITS · 37.3 GB FREE

cache 0.0 GB

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M18 GB43.2 tok/s
M1 Pro16 GB134 tok/s
M1 Max32 GB279 tok/s
M1 Ultra64 GB571 tok/s
M28 GB64.4 tok/s
M2 Pro16 GB136 tok/s
M2 Max32 GB282 tok/s
M2 Ultra64 GB578 tok/s
M38 GB64.4 tok/s
M3 Pro18 GB99.2 tok/s
M3 Max (14-core CPU)36 GB206 tok/s
M3 Max (16-core CPU)48 GB282 tok/s
M3 Ultra96 GB599 tok/s
M416 GB78.3 tok/s
M4 Pro24 GB188 tok/s
M4 Max (14-core CPU)36 GB293 tok/s
M4 Max (16-core CPU)48 GB395 tok/s
M516 GB101 tok/s
M5 Prounverified24 GB214 tok/s
M5 Maxunverified36 GB449 tok/s