Glossary

What is RAG (Retrieval-Augmented Generation)?

RAG (Retrieval-Augmented Generation) is an AI architecture that grounds large language model output in external knowledge by retrieving relevant documents at query time and injecting them into the prompt. It reduces hallucinations, keeps answers current, and lets teams ship LLM features over their own private data.

By Antoni Elkenbracht·Updated January 15, 2026

How RAG works in three stages

  1. Ingestion. Documents are chunked (commonly 200-800 tokens with overlap), embedded with a model like text-embedding-3-large, voyage-large-2 or gemini-embedding-001, and stored in a vector database with metadata.
  2. Retrieval. At query time the user's question is embedded and the top-k semantically closest chunks are returned. Hybrid retrieval (vector + BM25) typically improves recall by 10-20% on knowledge-base workloads.
  3. Generation. Retrieved chunks are inserted into the prompt template ("Use the following sources to answer..."). The LLM produces a grounded answer with citations.

When to use RAG

RAG is the dominant pattern when answers must reflect private, dynamic, or post-training-cutoff knowledge: internal documentation, customer records, product catalogs, regulations, or news. The 2024 Stanford HELM-RAG benchmark found grounded responses reduce hallucination rates by roughly 30-50% on closed-book QA tasks compared to vanilla LLM prompts.

Typical applications include support copilots, document analysis (legal, financial, medical), code search, sales enablement, and compliance Q&A. RAG is also the backbone of most enterprise "chat with your data" products.

Common pitfalls

  • Chunking too small or too large. Small chunks lose context; large chunks dilute relevance. Test with your eval set.
  • No reranking. Top-k vector search alone is noisy. A cross-encoder reranker (e.g. Cohere Rerank, BGE-reranker) reorders by relevance and is usually worth the latency cost.
  • Missing evals. Without retrieval recall, faithfulness, and answer-quality metrics, regressions go unnoticed. Open-source tools: Ragas, TruLens, DeepEval.
  • Stale indexes. Embeddings drift as models update; rebuild periodically or pin embedding versions.

RAG, agentic RAG, and ontology-aware RAG

Beyond vanilla RAG, two extensions matter for production: agentic RAG (the LLM iteratively decides what to retrieve, often used with tool-calling models) and ontology-aware retrieval (retrieval is scoped by a typed knowledge graph so the LLM reasons over entities and relationships, not just text snippets). Ontology-aware retrieval is the emerging standard for regulated domains where typed entities, hierarchies, and provenance matter.

How Enqualia handles RAG (Retrieval-Augmented Generation)

Enqualia exposes this capability as a versioned, metered API on isolated Cloud Run. Read the deployment guide to see schemas, latency targets, and pricing.

Read: RAG Platform

Frequently asked questions

How is RAG different from fine-tuning?
Fine-tuning adjusts a model's weights to specialize behavior. RAG keeps weights frozen and supplies fresh, domain-specific context at inference time. RAG is faster to iterate on, doesn't require retraining when the source data changes, and lets you cite sources. Many production systems combine both, using a fine-tuned model for tone and a RAG pipeline for facts.
What are the main components of a RAG pipeline?
An ingestion pipeline (chunking, embeddings, indexing into a vector database), a retriever (semantic search plus optional keyword/hybrid scoring), a reranker (often a cross-encoder), a prompt assembler that injects retrieved chunks, and a generation model that synthesizes the answer. Production systems add caching, query rewriting, and evaluation harnesses.
Which vector databases are commonly used for RAG?
Pinecone, Weaviate, Qdrant, Milvus, Chroma, pgvector (Postgres), and Vespa are the most common. Choice depends on scale, hybrid-search needs, hosting model, and metadata-filtering performance. For mid-sized workloads, pgvector is often enough; multi-billion-vector workloads typically use Pinecone or Vespa.
What's the difference between RAG and agentic RAG?
Vanilla RAG retrieves once per query. Agentic RAG lets an LLM decide whether to retrieve, what to retrieve, and whether to re-query. This is sometimes called iterative, self-correcting, or agentic retrieval. It costs more tokens but handles complex multi-hop questions better.

Related concepts