Back to Blog
RAG

Why Naive RAG Fails: Hybrid Search, Reranking and GraphRAG for Enterprise

By Keyved Engineering Team··6 min read

Short answer

Naive RAG — fixed-size chunks, vector search only, top results straight into the prompt — fails on enterprise data because it splits meaning, misses exact terms, ignores document structure and permissions, and can't answer questions that span many documents. Fix it with structure-aware parsing and chunking, hybrid keyword-plus-vector search, a reranker, metadata and permission filters, query rewriting, GraphRAG for relationship questions, and an evaluation set that measures retrieval and answers separately.

Key takeaways

  • Most RAG quality problems are retrieval problems, not model problems.
  • Parse documents by structure, not by character count.
  • Hybrid search plus a reranker is the highest-value upgrade for most systems.
  • Measure retrieval quality separately from answer quality.

Building a basic RAG system takes an afternoon. Split documents into chunks, embed them, store them in a vector database, retrieve the top five matches for each question and pass them to an LLM.

The demo works. Then real users arrive with real questions about real documents, and the answers are vague, incomplete or wrong.

We see this pattern constantly — including on projects we've inherited from other teams. The good news is that the causes are well understood. Here's why naive RAG fails on enterprise data and what fixes each problem.

What is naive RAG?

"Naive RAG" describes the simplest pipeline:

  1. Split documents into fixed-size chunks (say, 500 tokens)
  2. Embed each chunk as a vector
  3. Embed the user's question and find the most similar chunks
  4. Put the top few chunks into the prompt and generate an answer

It's a good starting point and a poor endpoint. If you're still choosing between RAG and other approaches, see RAG vs fine-tuning vs long context.

Why does naive RAG fail on enterprise data?

FailureWhat happensFix
Bad parsingTables, headings and layouts are lost or scrambledStructure-aware parsing
Arbitrary chunkingRelated information split across chunksStructure-based chunking with context
Vector-only searchMisses exact terms: codes, names, numbersHybrid search
Noisy top-kIrrelevant chunks crowd out the right oneReranking
No metadataOld versions, wrong region, wrong product retrievedMetadata filters
No permissionsUsers see content they shouldn'tPermission-aware retrieval
Vague questionsThe query doesn't match how documents are writtenQuery rewriting
Cross-document questionsAnswer requires connecting many sourcesGraphRAG or agentic retrieval
No evaluationNobody knows which step is failingRetrieval and answer evaluation

Let's go through them.

Fix 1: Parse documents by structure

Enterprise documents aren't clean text. They're PDFs with multi-column layouts, scanned pages, tables, headers, footnotes, slides and spreadsheets. Basic text extraction flattens all of that — tables become a jumble of numbers, headings vanish and reading order breaks.

What works:

  • Layout-aware parsers (such as Unstructured, LlamaParse, Azure Document Intelligence or Google Document AI) that keep headings, lists and tables
  • Converting tables to a structured format (Markdown or JSON) the model can read
  • OCR with quality checks for scanned documents (intelligent document processing)
  • Keeping document-level metadata: title, date, version, owner, section path

Parsing quality sets the ceiling for everything that follows.

Fix 2: Chunk by meaning, not by character count

Fixed-size chunks cut sentences, tables and clauses in half. A policy exception lands in one chunk and the rule it modifies in another.

What works:

  • Chunk along the document's own structure: sections, sub-sections, clauses, table rows
  • Add context to each chunk — the document title and section heading path, or a short generated summary of where the chunk sits
  • Use parent–child retrieval: search small, precise chunks but pass the larger parent section to the model
  • Keep tables whole where possible

We hit this directly in legal work, where a clause in section 14 can change a payment term in section 4. Our contract review case study describes how we handled it.

Vector search is good at meaning: "how do I get my money back" matches "refund policy." It's weak at exact terms: product SKUs, error codes, invoice numbers, people's names and acronyms.

Hybrid search runs keyword search (BM25) and vector search together and merges the results, commonly with reciprocal rank fusion. Most mature search and vector databases — including Elasticsearch, OpenSearch, Weaviate, Qdrant and PostgreSQL with pgvector plus full-text search — support it.

For enterprise content full of identifiers and jargon, hybrid search is one of the most reliable improvements you can make.

Fix 4: Add a reranker

Initial retrieval is optimised for speed across millions of chunks, so its ranking is approximate. A reranker (a cross-encoder model, such as those from Cohere, Voyage AI, Jina or open-source alternatives) reads the question and each candidate passage together and scores relevance much more accurately.

The pattern: retrieve 30–100 candidates with hybrid search, rerank, and send the best 5–10 to the model. It adds a little latency and cost, and usually a clear jump in answer quality.

Fix 5: Filter with metadata and permissions

Many wrong answers aren't retrieval misses — they're retrieval of the wrong version: last year's price list, another region's policy, a superseded procedure.

What works:

  • Store metadata with every chunk: date, version, status, region, product, department
  • Filter before or during search ("current policies for the UK only")
  • Let the model or a router infer filters from the question when users don't state them

Permissions belong here too. If a user can't open a document in your systems, retrieval must not return it. Enforce access control in the retrieval layer, using the user's identity — never rely on the prompt to hide content.

Fix 6: Rewrite and expand queries

Users ask short, vague or conversational questions ("what about for contractors?") that don't match how documents are written.

What works:

  • Rewrite follow-up questions into standalone queries using the conversation history
  • Generate several query variants and merge results
  • Break complex questions into sub-questions, retrieve for each and combine

Fix 7: Use GraphRAG for relationship questions

Some questions can't be answered from any single passage: "Which suppliers are involved in contracts that renew next quarter and have unusual liability caps?" or "What are the main themes across all customer complaints this year?"

GraphRAG extracts entities (people, companies, products, clauses) and relationships from documents into a knowledge graph, and uses the graph during retrieval. It's strong for connected and summary-level questions.

The trade-off: building and maintaining the graph costs more, both in processing and engineering. Use it when relationship questions are a core need, not by default. An alternative is agentic retrieval, where an agent runs several searches and queries structured data before answering.

Fix 8: Make the model use the context well

Even with perfect retrieval, generation can go wrong. Good practices:

  • Instruct the model to answer only from provided sources and to say when it doesn't know
  • Require citations to specific passages, and display them to users
  • Use structured output when answers feed other systems
  • Check groundedness: is each claim supported by the retrieved text?

Fix 9: Evaluate retrieval and answers separately

Without evaluation you're guessing which fix to apply. Build a test set of real questions with expected answers and the source passages that support them. Then measure:

StageMetrics
RetrievalRecall@k (did the right passage come back?), precision, MRR
GenerationCorrectness, completeness, groundedness (faithfulness), citation accuracy
End-to-endUser-rated helpfulness, escalation rate

Open-source tools such as Ragas and DeepEval, plus LLM-as-judge scoring calibrated against human ratings, make this practical. Our guide to evaluating LLM applications goes deeper.

What does a production RAG architecture look like?

  1. Ingestion pipeline: connectors, structure-aware parsing, OCR, chunking with context, metadata and permissions, embeddings — running continuously as documents change
  2. Indexes: vector plus keyword (and a graph, if needed)
  3. Query pipeline: query rewriting → hybrid search with filters → reranking → context assembly
  4. Generation: grounded answers with citations and refusal when unsupported
  5. Evaluation and observability: offline test sets plus production tracing (LLM observability)

The ingestion pipeline is usually the biggest piece of work, and the most neglected. It's a data engineering problem as much as an AI one.

How we build enterprise RAG at Keyved

Our RAG and knowledge systems are built on the pattern above: structure-aware parsing, hybrid search, reranking, metadata and permission filters, citations and evaluation from day one. We add GraphRAG or agentic retrieval only when the questions require it.

If you have a RAG system that "mostly works," we can usually find the failing stage within a week by building an evaluation set and tracing real questions. See examples on our projects page or talk to our engineers.

Frequently asked questions

Why is my RAG system giving wrong answers?

The most common causes are poor document parsing (tables and layouts lost), chunks that split related information, vector-only search missing exact terms like product codes, irrelevant chunks crowding out good ones, missing metadata filters, and no evaluation to detect which step is failing.

What is hybrid search in RAG?

Hybrid search combines keyword search (such as BM25), which is good at exact terms and names, with vector search, which is good at meaning and paraphrases. Results from both are merged, usually with reciprocal rank fusion, to improve recall.

What is a reranker in RAG?

A reranker is a model that scores how relevant each retrieved passage is to the question, more accurately than the initial search. Retrieving a larger candidate set and reranking it to the best few passages typically improves answer quality significantly.

What is GraphRAG?

GraphRAG builds a knowledge graph of entities and relationships from your documents and uses it during retrieval. It helps answer questions about connections and themes across many documents, which plain passage retrieval handles poorly. It costs more to build and maintain.

How do you evaluate a RAG system?

Use a test set of real questions with known answers and source documents. Measure retrieval (did the right passages come back?) and generation (is the answer correct, complete and grounded in the sources?) separately, so you know which part to fix.

Want to see how we build these systems for clients?

Let's Talk

Keep reading