AI Fundamentals

Multimodal Embedding

A vector representation of an input (text, image, audio, video) in a shared embedding space, enabling cross-modal search — e.g., finding images relevant to a text query, or vice versa. Powers modern multimodal RAG and semantic search across mixed content types.

Related terms

Next steps

Where this fits

Multimodal Embedding is part of the AI Fundamentals vocabulary used in the Generative AI Maturity Framework. See the full glossary for the complete set of 149 defined terms, or take the free maturity assessment to see where your organisation stands.