Skip to content

RAG architecture for enterprise AI: what to build and buy

A practical guide to RAG architecture for enterprise AI: when you need it, hybrid search, reranking, GraphRAG, and when to buy a hosted option instead.

Guides6 min read
By the AI App Hunters editors · Updated 1 Oct 2026

Retrieval-augmented generation (RAG) is still the standard way to make an AI model answer from your company's own data. The architecture that works in 2026 is not exotic: clean document parsing, hybrid keyword plus vector search, a reranker, permission filtering, and citations. Most teams should buy or use hosted retrieval first and only build custom pipelines where the data or the stakes demand it.

Key takeaways

  • Check whether you need RAG at all. Small, stable knowledge bases can go straight into the prompt.
  • Hybrid search plus reranking beats pure vector search for most business documents.
  • Your existing database (Postgres, MongoDB, Redis) can probably hold the vectors.
  • Hosted file search from OpenAI and Google now covers simple use cases with very little code.
  • Permissions and document quality cause more failures than model choice.

What is RAG and why do businesses still use it?

A language model only knows what it was trained on. RAG adds a lookup step. When someone asks a question, the system searches your documents, picks the most relevant passages, and gives them to the model with the question. The model answers from those passages and can point to the source.

That solves three problems businesses care about. Answers can use private data the model never saw. They stay current, because you update the index rather than retrain a model. And they can be checked, because each answer can cite a document.

Long context windows did not kill RAG. They changed when you need it. Anthropic's own guidance is that if your knowledge base is under about 200,000 tokens, you can often include all of it in the prompt and use prompt caching, which it says can cut latency by more than 2x and cost by up to 90%. Past that size, or when data changes daily, or when different people may see different documents, retrieval is still the right call.

What does a modern enterprise RAG architecture look like?

Think of it as two pipelines.

Ingestion runs ahead of time:

  1. Connect sources. SharePoint, Google Drive, Confluence, ticketing, CRM, databases.
  2. Parse. Turn PDFs, slides, scans and spreadsheets into clean text with structure kept. Tables and charts are where cheap parsers fail. Tools like LlamaIndex's LlamaParse exist for exactly this.
  3. Chunk. Split documents into passages a few hundred tokens long, along headings where possible.
  4. Add context. Prefix each chunk with a line saying what document and section it came from.
  5. Index. Store embeddings for vector search and text for keyword search, plus metadata such as owner, date and access rights.

Query runs live:

  1. Optionally rewrite the user's question into one or more search queries.
  2. Run keyword and vector search together and merge results.
  3. Filter by the user's permissions.
  4. Rerank the top candidates and keep the best handful.
  5. Generate the answer with citations, and log everything.

In 2026 many teams wrap this in an agent: the model decides whether to search, what to search for, and whether to search again. That works well, but the retrieval underneath still has to be good. An agent with weak search just fails more slowly.

Which retrieval techniques actually improve accuracy?

Three techniques have the clearest evidence behind them.

Hybrid search. Vector search finds passages that mean the same thing. Keyword search finds exact strings like "SKU-4471" or a contract clause number. Business documents are full of exact strings, so you want both.

Contextual chunks. A chunk that says "revenue grew 3% over the previous quarter" is useless without knowing which company and which quarter. Anthropic's Contextual Retrieval write-up prepends a short generated context to every chunk before indexing. In its tests, that cut failed retrievals by 35% with contextual embeddings alone, by 49% when combined with contextual BM25 keyword search, and by 67% when reranking was added.

Reranking. Retrieve broadly (say 50 to 150 candidates), then let a reranking model score each one against the question and keep the top few. Cohere Rerank 4, released in December 2025, supports 100+ languages and can run in Cohere's cloud, in your VPC or on-premises. Fewer, better passages also means fewer tokens per answer.

GraphRAG is the fourth option, and it is narrower. Microsoft's GraphRAG builds a knowledge graph from your text, groups related entities into communities and summarizes them. It helps with questions that span a whole collection ("what are the recurring complaints across all 2026 tickets?") or that require connecting facts across documents. Baseline RAG is weak at both. For a typical policy or product Q&A bot, it is overkill.

Do you need a dedicated vector database?

Usually not at first. The database you already run probably supports vectors:

  • Postgres with pgvector supports HNSW and IVFFlat indexes and can combine vector search with Postgres full-text search. Supabase builds on this and supports semantic, keyword and hybrid search.
  • MongoDB Atlas Vector Search keeps vectors next to operational data, supports hybrid search in a single query, and offers automated embeddings using Voyage AI models.
  • Redis offers vector search for low-latency use cases.

Keeping vectors beside the source data removes a sync job and a second security model. Dedicated vector databases such as Pinecone make sense at large scale or with strict latency needs. Pinecone has a free Starter tier, and its Standard plan has a $50 per month minimum (as of October 2026).

Should you build RAG or buy it?

Option Examples Best for Watch out for
Hosted file search in a model API OpenAI file search, Gemini File Search Prototypes and single-app use cases Less control over chunking and ranking
Enterprise search platform Glean Company-wide Q&A across many SaaS tools Price on request, less custom logic
End-user notebooks Gemini Notebook, projects in Claude or ChatGPT Teams working with a fixed set of documents Not a system other apps can call
Custom pipeline LangChain, LlamaIndex, your database Products, regulated data, complex permissions You own parsing, evals and upkeep

The hosted API tools are better than many teams realize. OpenAI's file search in the Responses API runs semantic and keyword search over vector stores you upload and returns file citations. Google's Gemini File Search chunks and indexes files for you, and storage and query embeddings are free; you pay for indexing at embedding rates and for retrieved tokens.

Glean is the main buy option for company-wide search. It lists 250+ connectors and is permission-aware, so people only see answers from documents they can already open. It does not publish prices.

Build custom when retrieval is part of your product, when permissions are complex (per-client data rooms, for example), or when documents need special parsing. If your goal is to show up in public AI answers rather than answer internal questions, that is a different job: see our GEO guide. For customer-facing support bots, compare dedicated vendors in AI agents for customer support.

Where do enterprise RAG projects go wrong?

  • Messy source documents. Duplicate policies, outdated versions and scanned PDFs. Retrieval finds the wrong version confidently. Clean and date your sources first.
  • Permissions as an afterthought. If access rights are not stored with each chunk and filtered at query time, the bot can leak salary bands or board notes. Design this on day one.
  • No evaluation set. Write 50 to 100 real questions with known correct sources before you tune anything. Measure whether the right passage is retrieved, separately from whether the answer reads well.
  • Treating the model as the fix. If the right passage was never retrieved, a better model will not find it.
  • No freshness plan. Decide how fast changes must show up, then schedule re-indexing to match.

How to choose

Start with the smallest thing that works. If the documents fit in a prompt, skip RAG. If one team needs answers from one set of files, use a notebook tool or hosted file search. If the whole company needs search across many systems with permissions, evaluate Glean against what your existing suite offers. Build custom only when retrieval is core to your product or your data rules demand it, and when you do, use hybrid search, contextual chunks and a reranker from the start.

Bottom line: good enterprise RAG is mostly good data plumbing. Get parsing, permissions and evaluation right and almost any current model will give useful answers.

Frequently asked questions

What is RAG in enterprise AI?+

Retrieval-augmented generation (RAG) finds the most relevant passages in your own documents or databases and passes them to a language model with the question. The model answers from that material, so answers can reflect private, current company data and cite where they came from.

Do I always need RAG?+

No. Anthropic suggests that for a knowledge base under about 200,000 tokens you can often put the whole thing in the prompt and use prompt caching instead. RAG earns its place when the corpus is large, changes often, or has per-user permissions.

What is hybrid search in RAG?+

Hybrid search combines keyword search (such as BM25 or full-text) with vector similarity search and merges the results. It catches exact terms like product codes and names that pure semantic search tends to miss.

Do I need a dedicated vector database?+

Often not. Postgres with pgvector, MongoDB Atlas Vector Search and Redis all store vectors next to your existing data, which avoids syncing a second system. Dedicated vector databases make sense at very large scale or with strict latency targets.

When does GraphRAG make sense?+

GraphRAG helps with broad questions across a whole dataset, like recurring themes in thousands of support tickets, and with questions that require connecting facts across documents. It adds indexing work, so start with hybrid search plus reranking and add a graph only if those questions matter.

Sources, checked 1 Oct 2026

  1. Anthropic: Introducing Contextual Retrieval
  2. Microsoft GraphRAG documentation
  3. Cohere Rerank
  4. OpenAI file search guide
  5. Gemini API File Search
  6. pgvector on GitHub
  7. Supabase AI and vectors
  8. MongoDB Atlas Vector Search
  9. Pinecone pricing
  10. Glean
  11. LlamaIndex

Keep reading

App of the Week6 min read

App of the Week: Attio, the AI-native CRM

Attio is an AI-native CRM with agents, workflows and an MCP server. What it does, who it suits, October 2026 pricing, limits and the best alternatives.

Attio logo
App of the Week5 min read

App of the Week: HeyGen, AI Avatar Video

HeyGen makes avatar videos from text and translates video into 175+ languages. What it does, October 2026 pricing, limits and the best alternatives.

HeyGen logo
App of the Week4 min read

App of the Week: Profound, AI Search Visibility

Profound tracks how your brand shows up in ChatGPT, Claude, Perplexity and Gemini, then runs agents to improve it. Features, pricing, limits, alternatives.

Profound logo