AI Fundamentals

Vision-Language Model (VLM)

A multimodal model that accepts both images and text as input and typically produces text output. Frontier VLMs (GPT-5, Claude Opus 4.7, Gemini 2.5 Pro) handle document understanding, UI comprehension, chart reading, and visual reasoning as first-class tasks alongside pure language work.

Related terms

Next steps

Where this fits

Vision-Language Model (VLM) is part of the AI Fundamentals vocabulary used in the Generative AI Maturity Framework. See the full glossary for the complete set of 149 defined terms, or take the free maturity assessment to see where your organisation stands.