Skip to content
mlx.app

DeepSeek V3

DeepSeek · 671B (37B active, MoE) · up to 128K context · mit

OVER BY 339.0 GB

This is the model that exists mainly to show what's possible: a 512GB Mac Studio is realistically the only consumer machine that can hold the 4-bit weights at all. Unless you own that machine, treat this entry as a ceiling reference rather than something to actually download.

On a M4 Pro with 48 GB

macOS9.6 GBweights377.4 GBKV cache0.0 GBover339.0 GB
Weights
377.4 GB
KV cache
0.0 GBestimated shape
Usable RAM
38.4 GB
Headroom
0.0 GB
Generation
10.2 tok/sest.
Prompt processing
24.4 tok/sest.
First token (1K prompt)
42.0 sest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q4377.4 GB377.4 GBOVER BY 339.0 GB10.2 tok/sDeepSeek-V3-4bit
q8712.9 GB712.9 GBOVER BY 674.5 GB5.4 tok/sDeepSeek-V3-0324-8bit

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/DeepSeek-V3-4bit
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/DeepSeek-V3-4bit --port 8080

How far can you push the context

2K tokensOVER BY 339.0 GB

cache 0.0 GB

8K tokensOVER BY 339.0 GB

cache 0.0 GB

32K tokensOVER BY 339.0 GB

cache 0.0 GB

128K tokensOVER BY 339.0 GB

cache 0.0 GB

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M1no configuration
M1 Prono configuration
M1 Maxno configuration
M1 Ultrano configuration
M2no configuration
M2 Prono configuration
M2 Maxno configuration
M2 Ultrano configuration
M3no configuration
M3 Prono configuration
M3 Max (14-core CPU)no configuration
M3 Max (16-core CPU)no configuration
M3 Ultra512 GB32.7 tok/s
M4no configuration
M4 Prono configuration
M4 Max (14-core CPU)no configuration
M4 Max (16-core CPU)no configuration
M5no configuration
M5 Prounverifiedno configuration
M5 Maxunverifiedno configuration