Skip to content
mlx.app

Qwen3 Embedding 4B

Alibaba · 4B · up to 32K context · apache-2.0

FITS · 32.9 GB FREE

Built on the Qwen3 4B backbone, this trades embedding model minimalism for MTEB-leaderboard-level retrieval quality, at the cost of needing a 16GB Mac rather than fitting anywhere. It also supports instruction-aware embeddings, so you can bias retrieval toward a task description at query time.

On a M4 Pro with 48 GB

macOS9.6 GBweights4.3 GBKV cache1.2 GBheadroom32.9 GB
Weights
4.3 GB
KV cache
1.2 GB
Usable RAM
38.4 GB
Headroom
32.9 GB
Generation
50.1 tok/sest.
Prompt processing
226 tok/sest.
First token (1K prompt)
4.5 sest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q42.3 GB3.5 GBFITS · 34.9 GB FREE94.6 tok/sQwen3-Embedding-4B-4bit-DWQ
q84.3 GB5.5 GBFITS · 32.9 GB FREE50.1 tok/sQwen3-Embedding-4B-mxfp8

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/Qwen3-Embedding-4B-mxfp8
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/Qwen3-Embedding-4B-mxfp8 --port 8080

How far can you push the context

2K tokensFITS · 33.8 GB FREE

cache 0.3 GB

8K tokensFITS · 32.9 GB FREE

cache 1.2 GB

32K tokensFITS · 29.3 GB FREE

cache 4.8 GB

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M18 GB11.5 tok/s
M1 Pro16 GB35.8 tok/s
M1 Max32 GB74.4 tok/s
M1 Ultra64 GB152 tok/s
M28 GB17.2 tok/s
M2 Pro16 GB36.2 tok/s
M2 Max32 GB75.3 tok/s
M2 Ultra64 GB154 tok/s
M38 GB17.2 tok/s
M3 Pro18 GB26.5 tok/s
M3 Max (14-core CPU)36 GB55.1 tok/s
M3 Max (16-core CPU)48 GB75.3 tok/s
M3 Ultra96 GB160 tok/s
M416 GB20.9 tok/s
M4 Pro24 GB50.1 tok/s
M4 Max (14-core CPU)36 GB78.1 tok/s
M4 Max (16-core CPU)48 GB105 tok/s
M516 GB27.0 tok/s
M5 Prounverified24 GB57.1 tok/s
M5 Maxunverified36 GB120 tok/s