Submit a run
Every speed on this site is currently an estimate. A measured run from your Mac is worth more than any model we can write. Fill this in and copy the result — we publish it as unverified until three independent runs agree.
Your submission
chip: M1 (m1) memory: 8 GB model: Qwen2.5 0.5B Instruct (qwen2.5-0.5b-instruct) quant: q4 runtime: mlx-lm context: 8192 tokens decode: tok/s prefill: tok/s
Read how a run gets verified before timing: warm start, at least 256 generated tokens, median of three runs, mains power.