Skip to content
mlx.app

Benchmarks

Every number below is a modelled estimate from memory bandwidth and quantized weight size — we will not print a measured figure we did not receive. Real runs from real Macs replace estimates as they arrive.

Ran a model on your Mac and timed it? Submit the run — chip, memory, model, quantization, runtime version and tokens per second. Submissions are published as unverified until several independent runs agree.
Estimated generation speed, tokens per second, best quantization that fits at your 48 GB and current context
ChipGB/sLlama 3.2 3B InstructQwen3 4BMistral 7B Instruct v0.3Qwen2.5 7B InstructInternVL3 8B
M1687.611.56.33.25.8
M1 Pro20023.735.819.610.017.9
M1 Max40049.274.440.720.837.2
M1 Ultra80010115283.542.676.2
M210011.417.29.44.88.6
M2 Pro20024.036.219.910.118.1
M2 Max40049.875.341.321.137.6
M2 Ultra80010215484.643.277.2
M310011.417.29.44.88.6
M3 Pro15017.526.514.57.413.2
M3 Max (14-core CPU)30036.455.130.215.427.5
M3 Max (16-core CPU)40049.875.341.321.137.6
M3 Ultra81910616087.644.780.0
M412013.820.911.45.810.4
M4 Pro27333.250.127.514.025.1
M4 Max (14-core CPU)41051.778.142.821.839.1
M4 Max (16-core CPU)54669.710557.729.552.7
M515317.927.014.87.513.5
M5 Prounverified30737.857.131.316.028.5
M5 Maxunverified61479.412065.733.560.0

All figures marked est. — estimates derived from published bandwidth, assuming a warm model, short prompt and no thermal throttling. Expect real Macs to land within roughly ±25%.

Measured runs from real Macs

Loading…

How we would verify a submission

  1. Same model repo, same quantization, same runtime version.
  2. Warm start: the model already resident, so the first load is not counted.
  3. At least 256 generated tokens, reported as the median of three runs.
  4. No other heavy applications open, and the machine on mains power.
  5. Three independent submissions within 15% of each other before a number stops being labelled unverified.