InternVL3 8B
OpenGVLab · 8B · up to 32K context · mit
A well-regarded open vision-language alternative to Qwen2.5-VL that fits the same 16GB tier. It was trained with native multi-image and video input in mind, which shows up in slightly better consistency across a sequence of frames.
On a M4 Pro with 48 GB
macOS9.6 GBweights8.5 GBKV cache0.0 GBheadroom29.9 GB
- Weights
- 8.5 GB
- KV cache
- 0.0 GBestimated shape
- Usable RAM
- 38.4 GB
- Headroom
- 29.9 GB
- Generation
- 25.1 tok/sest.
- Prompt processing
- 113 tok/sest.
- First token (1K prompt)
- 9.1 sest.
Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.
Every quantization, on your Mac
| Quant | Weights | + KV | Verdict | Speed est. | Repo |
|---|---|---|---|---|---|
| q3 | 3.5 GB | 3.5 GB | FITS · 34.9 GB FREE | 60.8 tok/s | InternVL3-8B-3bit |
| q4 | 4.5 GB | 4.5 GB | FITS · 33.9 GB FREE | 47.3 tok/s | InternVL3-8B-4bit |
| q6 | 6.5 GB | 6.5 GB | FITS · 31.9 GB FREE | 32.8 tok/s | InternVL3-8B-6bit |
| q8 | 8.5 GB | 8.5 GB | FITS · 29.9 GB FREE | 25.1 tok/s | InternVL3-8B-8bit |
We never host weights. Every link goes to Hugging Face.
Run it
Chat
mlx_lm.chat --model mlx-community/InternVL3-8B-8bitServe an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/InternVL3-8B-8bit --port 8080How far can you push the context
2K tokensFITS · 29.9 GB FREE
cache 0.0 GB
8K tokensFITS · 29.9 GB FREE
cache 0.0 GB
32K tokensFITS · 29.9 GB FREE
cache 0.0 GB
Which Macs run this
| Chip | Smallest RAM that fits | Speed est. |
|---|---|---|
| M1 | 16 GB | 5.8 tok/s |
| M1 Pro | 16 GB | 17.9 tok/s |
| M1 Max | 32 GB | 37.2 tok/s |
| M1 Ultra | 64 GB | 76.2 tok/s |
| M2 | 16 GB | 8.6 tok/s |
| M2 Pro | 16 GB | 18.1 tok/s |
| M2 Max | 32 GB | 37.6 tok/s |
| M2 Ultra | 64 GB | 77.2 tok/s |
| M3 | 16 GB | 8.6 tok/s |
| M3 Pro | 18 GB | 13.2 tok/s |
| M3 Max (14-core CPU) | 36 GB | 27.5 tok/s |
| M3 Max (16-core CPU) | 48 GB | 37.6 tok/s |
| M3 Ultra | 96 GB | 80.0 tok/s |
| M4 | 16 GB | 10.4 tok/s |
| M4 Pro | 24 GB | 25.1 tok/s |
| M4 Max (14-core CPU) | 36 GB | 39.1 tok/s |
| M4 Max (16-core CPU) | 48 GB | 52.7 tok/s |
| M5 | 16 GB | 13.5 tok/s |
| M5 Prounverified | 24 GB | 28.5 tok/s |
| M5 Maxunverified | 36 GB | 60.0 tok/s |