Skip to content
mlx.app

InternVL3 8B

OpenGVLab · 8B · up to 32K context · mit

FITS · 29.9 GB FREE

A well-regarded open vision-language alternative to Qwen2.5-VL that fits the same 16GB tier. It was trained with native multi-image and video input in mind, which shows up in slightly better consistency across a sequence of frames.

On a M4 Pro with 48 GB

macOS9.6 GBweights8.5 GBKV cache0.0 GBheadroom29.9 GB
Weights
8.5 GB
KV cache
0.0 GBestimated shape
Usable RAM
38.4 GB
Headroom
29.9 GB
Generation
25.1 tok/sest.
Prompt processing
113 tok/sest.
First token (1K prompt)
9.1 sest.

Decode speed is bandwidth ÷ active weight bytes, at 273 GB/s and 78% efficiency. These are modelled figures, not measurements.

Every quantization, on your Mac

QuantWeights+ KVVerdictSpeed est.Repo
q33.5 GB3.5 GBFITS · 34.9 GB FREE60.8 tok/sInternVL3-8B-3bit
q44.5 GB4.5 GBFITS · 33.9 GB FREE47.3 tok/sInternVL3-8B-4bit
q66.5 GB6.5 GBFITS · 31.9 GB FREE32.8 tok/sInternVL3-8B-6bit
q88.5 GB8.5 GBFITS · 29.9 GB FREE25.1 tok/sInternVL3-8B-8bit

We never host weights. Every link goes to Hugging Face.

Run it

Chat
mlx_lm.chat --model mlx-community/InternVL3-8B-8bit
Serve an OpenAI-compatible endpoint
mlx_lm.server --model mlx-community/InternVL3-8B-8bit --port 8080

How far can you push the context

2K tokensFITS · 29.9 GB FREE

cache 0.0 GB

8K tokensFITS · 29.9 GB FREE

cache 0.0 GB

32K tokensFITS · 29.9 GB FREE

cache 0.0 GB

Which Macs run this

ChipSmallest RAM that fitsSpeed est.
M116 GB5.8 tok/s
M1 Pro16 GB17.9 tok/s
M1 Max32 GB37.2 tok/s
M1 Ultra64 GB76.2 tok/s
M216 GB8.6 tok/s
M2 Pro16 GB18.1 tok/s
M2 Max32 GB37.6 tok/s
M2 Ultra64 GB77.2 tok/s
M316 GB8.6 tok/s
M3 Pro18 GB13.2 tok/s
M3 Max (14-core CPU)36 GB27.5 tok/s
M3 Max (16-core CPU)48 GB37.6 tok/s
M3 Ultra96 GB80.0 tok/s
M416 GB10.4 tok/s
M4 Pro24 GB25.1 tok/s
M4 Max (14-core CPU)36 GB39.1 tok/s
M4 Max (16-core CPU)48 GB52.7 tok/s
M516 GB13.5 tok/s
M5 Prounverified24 GB28.5 tok/s
M5 Maxunverified36 GB60.0 tok/s