Vision-Language Model (VLM)
A multimodal model that accepts both images and text as input and typically produces text output. Frontier VLMs (GPT-5, Claude Opus 4.7, Gemini 2.5 Pro) handle document understanding, UI comprehension, chart reading, and visual reasoning as first-class tasks alongside pure language work.
Related terms
- Large Language Model (LLM)
A type of AI model trained on vast amounts of text data to understand and generate human-like text. Examples include GPT-4, Claude, and Llama. LLMs power many GenAI applications.
- Foundation Model
A large pre-trained model (LLM, vision, or multimodal) intended as a general-purpose substrate that downstream apps consume via prompting, fine-tuning, or agent scaffolding. Under the EU AI Act these are General-Purpose AI (GPAI) models and their providers carry specific obligations.
Related on this site
Where this fits
Vision-Language Model (VLM) is part of the AI Fundamentals vocabulary used in the Generative AI Maturity Framework. See the full glossary for the complete set of 149 defined terms, or take the free maturity assessment to see where your organisation stands.