Retrieval-Augmented Generation (RAG)
A pattern where an LLM answers a query by first retrieving relevant documents from an external corpus (typically via a vector database or hybrid search) and injecting them into the prompt as grounding context. RAG addresses freshness and factual grounding without model retraining.
Related terms
- Vector Database
A database optimized for storing and querying high-dimensional vector embeddings via approximate nearest-neighbor search. Common choices include Pinecone, Weaviate, Qdrant, Milvus, and pgvector. Central to most RAG architectures.
- Embedding
A numeric vector representation of text (or images, audio, etc.) produced by a model such that semantically similar inputs map to nearby vectors. Used to power search, clustering, and retrieval in modern GenAI stacks.
- GraphRAG
A variant of RAG that first extracts entities and relationships from source documents into a knowledge graph, then queries the graph to assemble context. Better for questions requiring multi-hop reasoning across documents than vector-similarity RAG.
Related on this site
Where this fits
Retrieval-Augmented Generation (RAG) is part of the AI Fundamentals vocabulary used in the Generative AI Maturity Framework. See the full glossary for the complete set of 149 defined terms, or take the free maturity assessment to see where your organisation stands.