What is RAG (Retrieval-Augmented Generation)?
RAG (Retrieval-Augmented Generation) is an AI architecture that grounds large language model output in external knowledge by retrieving relevant documents at query time and injecting them into the prompt. It reduces hallucinations, keeps answers current, and lets teams ship LLM features over their own private data.
How RAG works in three stages
- Ingestion. Documents are chunked (commonly 200-800 tokens with overlap), embedded with a model like text-embedding-3-large, voyage-large-2 or gemini-embedding-001, and stored in a vector database with metadata.
- Retrieval. At query time the user's question is embedded and the top-k semantically closest chunks are returned. Hybrid retrieval (vector + BM25) typically improves recall by 10-20% on knowledge-base workloads.
- Generation. Retrieved chunks are inserted into the prompt template ("Use the following sources to answer..."). The LLM produces a grounded answer with citations.
When to use RAG
RAG is the dominant pattern when answers must reflect private, dynamic, or post-training-cutoff knowledge: internal documentation, customer records, product catalogs, regulations, or news. The 2024 Stanford HELM-RAG benchmark found grounded responses reduce hallucination rates by roughly 30-50% on closed-book QA tasks compared to vanilla LLM prompts.
Typical applications include support copilots, document analysis (legal, financial, medical), code search, sales enablement, and compliance Q&A. RAG is also the backbone of most enterprise "chat with your data" products.
Common pitfalls
- Chunking too small or too large. Small chunks lose context; large chunks dilute relevance. Test with your eval set.
- No reranking. Top-k vector search alone is noisy. A cross-encoder reranker (e.g. Cohere Rerank, BGE-reranker) reorders by relevance and is usually worth the latency cost.
- Missing evals. Without retrieval recall, faithfulness, and answer-quality metrics, regressions go unnoticed. Open-source tools: Ragas, TruLens, DeepEval.
- Stale indexes. Embeddings drift as models update; rebuild periodically or pin embedding versions.
RAG, agentic RAG, and ontology-aware RAG
Beyond vanilla RAG, two extensions matter for production: agentic RAG (the LLM iteratively decides what to retrieve, often used with tool-calling models) and ontology-aware retrieval (retrieval is scoped by a typed knowledge graph so the LLM reasons over entities and relationships, not just text snippets). Ontology-aware retrieval is the emerging standard for regulated domains where typed entities, hierarchies, and provenance matter.
How Enqualia handles RAG (Retrieval-Augmented Generation)
Enqualia exposes this capability as a versioned, metered API on isolated Cloud Run. Read the deployment guide to see schemas, latency targets, and pricing.
Read: RAG Platform