Skip to content
mlx.app

Ferret

Apple's research VLM for referring and grounding inside images.

When to use it

Use this when you need region-level grounding — 'what is inside this box' — rather than whole-image captioning or chat.

Alternatives in Vision