Skip to content AI
  • Overview

    • AI news
    • Technical blog
    • Live AI events
    • Inference explained
    • See our approach
  • Products

    • Red Hat AI Enterprise
    • Red Hat AI Inference
    • Red Hat Enterprise Linux AI
    • Red Hat OpenShift AI
    • Explore Red Hat AI
  • Engage & learn

    • Learning hub
    • AI topics
    • AI partners
    • Services for AI
Hybrid cloud
  • Platform solutions

    • Artificial intelligence

      Build, deploy, and monitor AI models and apps.

    • Linux standardization

      Get consistency across operating environments.

    • Application development

      Simplify the way you build, deploy, and manage apps.

    • Automation

      Scale automation and unite tech, teams, and environments.

  • Use cases

    • Virtualization

      Modernize operations for virtualized and containerized workloads.

    • Digital sovereignty

      Control and protect critical infrastructure.

    • Security

      Code, build, deploy, and monitor security-focused software.

    • Edge computing

      Deploy workloads closer to the source with edge technology.

  • Explore solutions
  • Solutions by industry

    • Automotive
    • Financial services
    • Healthcare
    • Industrial sector
    • Media and entertainment
    • Public sector (Global)
    • Public sector (U.S.)
    • Telecommunications

Discover cloud technologies

Learn how to use our cloud products and solutions at your own pace in the Red Hat® Hybrid Cloud Console.

Products
  • Platforms

    • Red Hat AI iconartificial intelligence, Red Hat Enterprise Linux AI, Red Hat OpenShift AI, RHEL AI, machine learning38382025-03-12T19:43:40.963Zimage/svg+xmlRed Hat AI iconartificial intelligence, Red Hat Enterprise Linux AI, Red Hat OpenShift AI, RHEL AI, machine learningIconno2025-03-12T19:39:59.817ZTechnology iconStandardRed Hat AI

      Develop and deploy AI solutions across the hybrid cloud.

    • Red Hat Enterprise Linux iconRHEL, Linux platforms, CentOS2024-03-01T15:26:42.958ZpendingTRA3b65dd25-844d-49bb-93c1-30f5b34684f1Icon2024-03-01T15:26:42.958Ztruepending2024-03-21T00:40:29.326Zrhcc-audience:internalnoTechnology iconDER3b65dd25-844d-49bb-93c1-30f5b34684f1Standardyesrhcc-product:red-hat-enterprise-linuxTechnology iconimage/svg+xml2024-05-10T14:11:29.114ZRed Hat Enterprise Linux iconRHEL, Linux platforms, CentOSActivateActivate2024-05-10T14:11:29.836Zworkflow-process-serviceActivateworkflow-process-servicefalse2024-05-10T14:11:29.836Zworkflow-process-service2024-05-10T14:11:29.836ZUse technology icons to represent Red Hat products and components. Do not remove the icon from the bounding shape.Red Hat Enterprise Linux

      Support hybrid cloud innovation on a flexible operating system.

    • Red Hat OpenShift iconCloud, Containers, Kubernetes2024-03-01T15:26:53.684ZpendingTRA9ec76aa9-ef09-4c49-8816-01dd13970ca7Icon2024-03-01T15:26:53.684Ztruepending2024-03-21T00:39:44.126Zrhcc-audience:internalnoTechnology iconDER9ec76aa9-ef09-4c49-8816-01dd13970ca7Standardyesrhcc-product:red-hat-openshiftrhcc-product:red-hat-openshift-on-ibm-cloudrhcc-product:microsoft-azure-red-hat-openshiftrhcc-product:red-hat-openshift-service-on-awsrhcc-product:red-hat-openshift-container-platformrhcc-product:red-hat-openshift-platform-plusTechnology iconimage/svg+xml2024-05-10T14:18:23.703ZRed Hat OpenShift iconCloud, Containers, KubernetesActivateActivate2024-05-10T14:18:25.221Zworkflow-process-serviceActivateworkflow-process-servicefalse2024-05-10T14:18:25.221Zworkflow-process-service2024-05-10T14:18:25.221ZUse technology icons to represent Red Hat products and components. Do not remove the icon from the bounding shape.Red Hat OpenShift

      Build, modernize, and deploy apps at scale.

    • Red Hat Ansible Automation Platform iconManagement, edge2024-03-01T15:26:35.068ZpendingTRA759b57c4-760b-45a0-a939-821f47181964Icon2024-03-01T15:26:35.068Ztruepending2024-03-21T00:39:55.923Zrhcc-audience:internalnoTechnology iconDER759b57c4-760b-45a0-a939-821f47181964Standardyesrhcc-product:red-hat-ansible-automation-platformTechnology iconimage/svg+xml2024-05-10T14:04:00.014ZRed Hat Ansible Automation Platform iconManagement, edgeActivateActivate2024-05-10T14:04:01.784Zworkflow-process-serviceActivateworkflow-process-servicefalse2024-05-10T14:04:01.784Zworkflow-process-service2024-05-10T14:04:01.784ZUse technology icons to represent Red Hat products and components. Do not remove the icon from the bounding shape.Red Hat Ansible Automation Platform

      Implement enterprise-wide automation.

      New version
  • Featured

    • Lightwell
    • Red Hat AI Enterprise
    • Red Hat OpenShift Virtualization Engine
    • Red Hat Desktop
    • See all products
  • Try & buy

    • Start a trial
    • Buy online
    • Integrate with major cloud providers
  • Services & support

    • Consulting
    • Product support
    • Services for AI
    • Technical Account Management
    • Explore services
Training
  • Training & certification

    • Courses and exams
    • Certifications
    • Skills assessments
    • Red Hat Academy
    • Learning subscription
    • Explore training
  • Featured

    • Red Hat Certified System Administrator exam
    • Red Hat System Administration I
    • Red Hat Learning Subscription trial (No cost)
    • Red Hat Certified Engineer exam
    • Red Hat Certified OpenShift Administrator exam
  • Services

    • Consulting
    • Partner training
    • Product support
    • Services for AI
    • Technical Account Management
Learn
  • Build your skills

    • Documentation
    • Hands-on labs
    • Hybrid cloud learning hub
    • Interactive demos
    • Training and certification
  • More ways to learn

    • Blog
    • Events and webinars
    • Podcasts and video series
    • Red Hat TV
    • Resource library

For developers

Discover resources and tools to help you build, deliver, and manage cloud-native applications and services.

Partners
  • For customers

    • Our partners
    • Red Hat Ecosystem Catalog
    • Find a partner
  • For partners

    • Partner Connect
    • Become a partner
    • Training
    • Support
    • Access the partner portal

Build solutions powered by trusted partners

Find solutions from our collaborative community of experts and technologies in the Red Hat® Ecosystem Catalog.

ConsoleDocsSupport Search

I'd like to:

  • Start a trial
  • Buy a learning subscription
  • Manage subscriptions
  • Contact sales
  • Contact customer service
  • See Red Hat jobs

Help me find:

  • Documentation
  • Developer resources
  • Tech topics
  • Architecture center
  • Security updates
  • Customer support

I want to learn more about:

  • AI
  • Application modernization
  • Automation
  • Cloud-native applications
  • Linux
  • Virtualization
New For you

Recommended

We'll recommend resources you may like as you browse. Try these suggestions for now.

  • Product trial center
  • Courses and exams
  • All products
  • Tech topics
  • Resource library
Log in

Get more with a Red Hat account

  • Console access
  • Event registration
  • Training & trials
  • World-class support

A subscription may be required for some services.

Log in or register
Contact us
Red Hat logo
  • Home
  • Resources
  • Engineering RAG for the enterprise: Index, retrieve, and generate

Engineering RAG for the enterprise: Index, retrieve, and generate

July 20, 2026•
Resource type: E-book
Download PDF

Introduction

You've probably heard someone say it by now: "RAG is dead."

It's a provocative claim, and it captures a real shift in the AI conversation. Two years ago, retrieval-augmented generation (RAG) was a trending topic in AI. Today, the hype has moved on to AI agents, autonomous workflows, and multistep reasoning. RAG, the argument goes, was a stepping stone. We've outgrown it.

Except we haven't. If anything, RAG has become more important, because the systems replacing simple chatbots are the ones that need it most.

Why RAG matters more than ever

Consider what's changed. In 2024, a typical RAG system retrieved a few paragraphs from a knowledge base and fed them to a model that generated an answer. The stakes were low. If the answer was slightly off, a human reviewed it and moved on.

Now, AI agents are making decisions. For example, a claims-processing agent at an insurance company evaluates coverage. A compliance agent at a bank flags suspicious transactions. A procurement agent approves vendor contracts. These systems take action—they execute workflows that affect real outcomes and real money.

An agent that isn't grounded in what the enterprise actually knows will make confident, well-formatted, but ultimately wrong decisions. RAG is the mechanism that connects autonomous AI systems to the organization's own data, giving agents the context they need to act accurately and be auditable.

This isn't just about text anymore. Enterprise data is multimodal: documents with embedded charts, infographics designed for human readers, scanned forms, audio recordings, and spreadsheets used as databases are just a few examples. The context an agent needs might be spread across a PDF, a Confluence page, and a quarterly earnings slide deck, all at once. Modern RAG systems must be able to handle this reality. They have to find, correlate, and deliver the right information from across data types, sources, and formats.

Addressing the gap

The challenge is that RAG is deceptively simple in concept and extremely difficult to get right in production. It involves dozens of interdependent decisions, from how documents are parsed and chunked, to which embedding model is selected, to how retrieval results are ranked. There are no universal defaults. The right combination depends entirely on your data, your domain, and your use case.

This guide walks through the full practitioner journey: building a solid baseline, rigorous evaluation, model fine-tuning, and Auto RAG, where many of these decisions can be optimized automatically. Start wherever makes sense for your situation.

3D letters: R A G

Chapter 1: Why enterprise RAG is harder than it looks

Enterprise data falls into 2 broad categories. The 1st is structured sources, such as databases and APIs, which are relatively straightforward for machines to process. The data has a schema, fields have types, and queries return predictable results.

The 2nd category is unstructured sources, where RAG gets difficult. In most organizations, unstructured data is where knowledge actually lives. The following are examples of the most common challenges from unstructured data. 

Documents designed for humans

Enterprise documents are increasingly rich and visual. Annual reports have infographics. Technical specs have multicolumn layouts. Training manuals mix diagrams with body text. Tables span multiple pages. These documents make perfect sense to a human reader who can follow the visual flow. But to a machine attempting to index and extract that content, they present a significant challenge.

A table that wraps across 2 pages loses its structure when parsed as plain text. A pie chart embedded in a PDF carries meaning that disappears entirely during text extraction. A scanned document may not have an embedded text layer at all, or may become garbled by the physical transfer process. Each of these requires a different parsing strategy, and choosing incorrectly means the data that enters your pipeline is already corrupted before chunking and embedding even begin.

A modern two-story building rendered in white with red architectural accents

Context fragmentation

The answer to a question rarely lives in a single place. A query about contractor data access policies in the European Union (EU) might require information from a human resources (HR) policy document, a regional legal framework, and an internal compliance memo. Those 3 sources might be in different formats, stored in different systems, and written in different styles.

This is among the biggest challenges with enterprise RAG—assembling complete context from fragments scattered across documents and methods. It gets more complex when you factor in multimodal data, where the relevant context might include text, images, structured data, and audio. A naive RAG pipeline that retrieves a single text chunk from a single source will miss this entirely.

Rapidly evolving data

Enterprise knowledge changes. Product features shift between releases. Financial data updates in near real time. Policies get revised quarterly or more often. A RAG pipeline that indexed last month's documentation may be returning outdated answers today.

This is part of why RAG exists at all. Fine-tuning a model on enterprise data creates a snapshot. RAG, done well, can pull from current sources. But "done well" means the indexing pipeline has to keep pace with the rate of change in the underlying data, and that's an infrastructure challenge most demos don't account for.

Security and prompt injection

Every time your RAG system retrieves content and places it in a model's context window, that content becomes a potential attack vector. If a retrieved document contains text crafted to look like system instructions, it can override the model's intended behavior. This isn't a theoretical risk—it's a known class of vulnerability that any production RAG system needs to address.

Good RAG design means treating retrieved content as untrusted input: clearly delineating augmented context from system instructions so the model can distinguish between the 2. Essentially, good RAG design has guardrails that exist between data and instructions.

The demo-to-production gap

These challenges compound. A demo pipeline using a handful of clean, well-formatted documents will perform well because none of these problems are present. The questions are simple, the data is tidy, and everything fits within the model's context window.

Production is different. Real users ask compound, ambiguous questions. Documents are messy, and context is fragmented across sources. The retrieval system returns results that are plausible but wrong, or right but buried under less relevant hits. The model fills its context window with noise and starts hallucinating.

The result is a system that stakeholders saw work perfectly in a meeting and that now fails unpredictably in the field. The frustration this creates is real, and it's often what pushes teams to start treating RAG as an engineering discipline rather than a quick integration.

The next chapter examines the practical decisions involved in building a production RAG pipeline, starting with a distinction that most tutorials skip entirely: your indexing pipeline and your retrieval pipeline are not the same thing.

Chapter 2: Building your first enterprise RAG pipeline

Building your first enterprise RAG pipeline

Most tutorials treat a RAG pipeline as a single flow: ingest, chunk, embed, store, retrieve, generate. In production, that framing breaks down on first contact with enterprise data. A real enterprise RAG system is 3 distinct pipelines, each with a different cadence, a different owner, and a different failure mode. Conflating them is one of the most common architectural mistakes teams make early on.

The 3-pipeline architecture

  • The indexing pipeline prepares your data. It parses raw enterprise documents, normalizes the content, enriches it with metadata, chunks it, generates embeddings, and stores everything in your vector store. This is the 1st pipeline a team builds, and the one most tutorials cover.
  • The retrieval pipeline is what your application calls. When a user or agent submits a query, the retrieval pipeline searches the prepared data, ranks and filters results, constructs a prompt with retrieved context, and sends it to the model for generation.
  • The maintenance pipeline keeps the indexed corpus up to date over time. Source documents change, get withdrawn, or have their access rules updated. Embedding models get upgraded. Chunking strategies get refined. The maintenance pipeline handles all of this without taking retrieval offline. This pipeline is often the most critical, and the one that needs to be highly customized by an organization beyond the tooling a platform provides.

Each pipeline has a different operating profile. Indexing runs in bulk, often as a one-off or scheduled job. Retrieval runs continuously at query latency, where milliseconds count. Maintenance runs incrementally, triggered by upstream events such as document edits, access control changes, or model upgrades.

A problem in retrieval quality might originate in any of the 3. Debugging a RAG system requires knowing which pipeline to look at.

Pipeline 1: Indexing

If the parser can't handle your source documents, nothing else in the pipeline will compensate. Indexing is where the cost of getting data wrong is highest, because every downstream step inherits the result.

  • Parser selection. A single parser does not fit every use case. A multicolumn PDF needs different treatment than a scanned document. A programmatically generated PDF parses differently from a slide deck exported to PDF. The parser needs to match the document types you're working with, and in most enterprises that means handling several formats at once. Red Hat® AI integrates Docling, an open source document-parsing tool that handles unstructured enterprise data at scale across PDFs, HTML, PPTX, and other formats, ensuring model outputs are grounded in your organization's knowledge base.
  • Metadata extraction. When you extract content, you also need to preserve its provenance: source file, date, version, owner, and access control tags. This metadata enables filtering at retrieval time and allows the model to cite sources, thereby supporting auditability. It is also what the maintenance pipeline uses to detect what needs updating. If your indexing pipeline strips metadata during parsing, you permanently lose all 3 capabilities.
  • Normalization. Raw documents carry noise such as headers, footers, encoding artifacts, and layout markup. Normalization cleans this up and extracts structured elements such as tables and charts in a way that preserves their meaning.
  • Chunking strategy. Content is split into chunks for embedding. Fixed-size chunking is fast but context-unaware. Semantic and recursive chunking provides better recall on documents with mixed formatting. Document-aware chunking respects explicit structure such as headings, sections, and tables. There is no universally right chunk size: the right combination depends on the embedding model, the document characteristics, and the queries your users will ask. Evaluation (which we will get into in Chapter 3) tells you when to change it.
  • Embedding model selection. The embedding model converts chunks into vectors, and the quality of those vectors determines how well retrieval matches queries to content. A general-purpose model is a reasonable starting point for text-heavy use cases. Domain-specific limitations show up quickly in specialized verticals. Multilingual deployments often require either multilingual models or fine-tuning on the language mix in use. Like chunking, this is an evaluation question.

Pipeline 2: Retrieval

The retrieval pipeline is where customer-facing quality is felt, and where small configuration changes produce outsized effects.

  • Vector store selection in an enterprise is often an infrastructure decision. If your organization is standardized on PostgreSQL, pgvector may be the path of least resistance. Purpose-built vector databases such as Milvus offer optimized performance for vector-specific workloads, but extensions on general-purpose databases are often sufficient when the alternative is introducing entirely new infrastructure. Red Hat AI supports multiple backends.
  • Retrieval strategy directly affects answer quality. The 3 main approaches are dense retrieval (vector similarity), sparse retrieval (keyword-based, such as BM25), and hybrid retrieval combining both. In most production systems, hybrid outperforms either alone. Reranking adds a 2nd pass to reorder results using a cross-encoder model, often yielding substantial improvement with minimal effort.
  • Access control filtering is enforced at retrieval time using the metadata captured during indexing. A user's permissions, group membership, or regional jurisdiction can all narrow the candidate set before similarity ranking runs. 
  • Prompt construction is where the retrieved context meets the model. How you delineate retrieved content from system instructions matters both for answer quality and for prompt-injection defense. Treat retrieved content as untrusted input.

Switching from dense to hybrid retrieval, or adding a reranking step, can move faithfulness scores by double-digit percentages. This is one of several areas where Auto RAG's ability to test configurations systematically pays off.

Pipeline 3: Maintenance

The maintenance pipeline is what most tutorials skip and what every production team eventually builds, usually under pressure, after an incident. Building it deliberately is cheaper.

  • Update and delete propagation. Source documents change. Policies get revised, contracts get superseded, products get discontinued. The maintenance pipeline detects these changes through webhooks, change feeds, or scheduled comparison and updates the vector store accordingly.
  • Right-to-be-forgotten and data residency. Regulated environments require the ability to remove specific records on request, often with a defined service-level agreement (SLA). The maintenance pipeline is where these requests are executed, and where the audit trail proving compliance is generated. Without a dedicated pipeline, teams run ad-hoc queries against production, which is exactly the kind of unaudited access that triggers compliance findings.
  • Reembedding on model change. When you upgrade the embedding model, every vector in the store is now in a different semantic space than incoming queries. Retrieval silently degrades. The maintenance pipeline handles this either as a full reindex or as a phased migration, with both old and new embeddings live until the switch is complete.
  • Rechunking on strategy change. If evaluation tells you a different chunking approach would improve retrieval, the maintenance pipeline applies the change to the existing corpus rather than requiring a full rebuild.
  • Access control synchronization. When permissions change in source systems (someone leaves a team, a project becomes restricted, a region adopts new data handling rules), the metadata on the corresponding vectors needs to follow.
  • Corpus drift detection. This is distinct from model drift. As source content evolves, the distribution of what's in your vector store shifts. Your evaluation dataset, built against an earlier snapshot, may no longer represent the corpus the system is actually searching. The maintenance pipeline is where corpus drift is detected and where reevaluation cycles are triggered.
  • Provenance and versioning. For regulated customers, knowing which version of which document was indexed at which point, and which answers it contributed to, is a compliance requirement. The maintenance pipeline owns this audit trail. Red Hat AI integrates MLflow for experiment tracking across pipeline configurations, and the same discipline applies to the corpus itself.
A red 3D square with a white Automation icon in a line of clear 3D squares

How the 3 pipelines interact

The pipelines are decoupled by design, but they share state through the vector store and its metadata. Indexing writes the initial state. Retrieval reads it. Maintenance updates it continuously, in response to events from source systems and from internal triggers such as model upgrades or evaluation findings.

A practical operating pattern: Indexing runs once per newly onboarded data source. Retrieval runs at query time, every time. Maintenance runs continuously in the background. This is the architecture that survives contact with enterprise data over years rather than months.

At this point, you have a pipeline. The natural next question: How do you know if it's working well? That's covered in Chapter 3.

diagram showing A practical example of how the 3 pipelines interact in a RAG setup

Figure 1. A practical example of how the 3 pipelines interact in a RAG setup

Chapter 3: Evaluating RAG

Most enterprise RAG projects skip evaluation until something goes wrong in production. A stakeholder reports a bad answer. A user screenshots a hallucination. The team scrambles to figure out whether the problem is retrieval, generation, or both, and has no baseline to compare against.

This is expensive. Fixing a RAG pipeline without evaluation data is guesswork. You change the chunk size, rerun a few queries by hand, and hope the answers look better. With evaluation, you can pinpoint exactly where the pipeline is breaking and measure whether your fix actually worked.

The 4 metrics that matter

RAG evaluation separates into 2 concerns: First, is the retrieval good? Second, is the generation good? These 4 metrics cover both.

  1. Context recall measures whether the right documents are being retrieved. If the answer to a question exists in your knowledge base but the retrieval system isn't finding it, context recall will be low. This points to problems in the indexing pipeline: parsing failures, poor chunking, or an embedding model that doesn't represent your domain well.
  2. Context precision measures whether the retrieved documents are actually relevant. High recall with low precision means the system is finding the right documents but also returning a lot of noise alongside them. That noise fills the context window, which degrades generation quality even when the relevant content is technically present.
  1. Answer faithfulness measures whether the generated response is grounded in the retrieved context. A faithfulness score tells you how much of the answer the model actually derived from the documents it was given, versus how much it fabricated. Low faithfulness with high recall and precision usually means the generation model is the bottleneck, either because it's too small for the task or because the prompt construction isn't directing it to stick to the provided context.
  2. Answer relevance measures whether the response actually addresses the question. A model can produce a perfectly faithful answer that's grounded in retrieved documents and still miss the point of what the user asked.

Together, these 4 metrics give you a diagnostic map. Low recall? Look at your indexing pipeline. Low precision? Look at your retrieval configuration and reranking. Low faithfulness? Look at your generation model and prompt design. Low relevance? Look at how queries are being interpreted.

Building an evaluation harness

The practical barrier to evaluation is usually data. You need a set of questions with known, good answers to test against, and most enterprise teams don't have them available.

There are 2 paths here. The 1st is to build a golden dataset manually by gathering representative questions from real users, having domain experts write reference answers, and using that as your benchmark. This produces high-quality evaluation data but takes time.

The 2nd is synthetic data generation. SDG Hub, which is part of Red Hat AI, can generate question-answer pairs from your enterprise documents, giving you an evaluation dataset without the cost of manual annotation. The synthetic data won't be perfect, but it gives you a working evaluation loop fast, which is better than having no evaluation at all.

Once you have evaluation data, the workflow is straightforward. Run your RAG pipeline on the dataset, score each response across the 4 metrics using LM-Eval, and review the results. A context recall of 0.62 with a faithfulness of 0.91 tells you the model is doing its job well when it gets the right context, but the retrieval system is only finding the right documents about 60% of the time. That's a retrieval problem, and the fix is in your indexing pipeline or retrieval configuration, not in the generation model.

Evaluation as a continuous practice

The most important shift this chapter asks for is treating evaluation as ongoing, not as a one-time gate before deployment.

Define your accuracy targets up front. What context recall do you need? What faithfulness score is acceptable for your use case? Once those targets exist, every decision in the pipeline can be measured against them. Should you change the chunking strategy? Should you switch to a different embedding model? Should you try a larger generation model and absorb the additional inference cost? In any case, the best course of action is to run an evaluation.

Red Hat AI evaluation capabilities offer granular experiment tracking of benchmarks, AI safety and performance profiles for AI artifacts such as prompts, retrievals, datasets, RAG configurations, and the ability to export these results as verifiable and demonstrable immutable Open Container Initiative (OCI) artifacts. This supports regulatory reporting requirements such as the EU AI Act and promotes the reproducibility of results.

Red Hat tooling for evaluation

Red Hat AI offers a range of tools to help you build and test your RAG pipeline. Four of the most important are: 

  • EvalHub, a unified evaluation control plane (preview) that replaces ad-hoc manual testing. It provides a comprehensive interface to scientifically benchmark, score, and audit models, RAG pipelines, and AI agents before they reach production.
  • LM-Eval, which handles benchmarking across the 4 RAG metrics as well as broader model capabilities such as logical reasoning and domain-specific accuracy. It's the primary tool for measuring whether a change to your pipeline actually improved results.
  • GuideLLM, which focuses on a different dimension: throughput and capacity planning. Once your pipeline is accurate, GuideLLM helps you understand how it performs under production load.
  • AI Guardrails, which provide runtime safety checks on model inputs and outputs. While LM-Eval measures quality in a test environment, guardrails monitor quality in production and flag responses that fall outside acceptable bounds.

As we continue, Chapter 4 uses these metrics to diagnose common failure modes. Chapter 5 uses them to determine when fine-tuning is worth the investment. And Chapter 7's Auto RAG uses them as the objective function for automated pipeline optimization.

Chapter 4: Common pitfalls and how to fix them

If your RAG pipeline is underperforming and you're not sure why, the problem almost always traces back to one of 2 root causes. This chapter maps these root causes to the evaluation metrics from Chapter 3, so you can diagnose what's going wrong and focus your effort in the right place.

Messy data in, bad answers out

Enterprise data is unstructured by nature. When the parser can't properly extract content from your source documents, every downstream step inherits the problem. For example, multicolumn layouts get flattened into nonsense text, tables that span pages lose their structure, and scanned documents produce garbled text or nothing at all. 

Chunking and embedding will produce poor results regardless of how well they're configured if the input data is already corrupted.

  • What poor metrics indicate: Low context recall, even when you know the right document exists in the knowledge base. This results in answers that miss key details or hallucinations that appear to come from nowhere but actually stem from malformed context.
  • Where to look: The indexing pipeline. Validate extraction quality on a representative sample of your documents before you index everything. Spot check the chunks your pipeline produces against the original source. If the chunks don't make sense to a human reader, they won't make sense to an embedding model either. This is where to invest in parsing quality and where Docling comes into use.

Wrong configuration

RAG pipelines have many interdependent parameters, such as chunking strategy, chunk size, overlap, embedding model, retrieval method, and reranking. The right combination depends entirely on your data and use case. Accepting defaults without evaluating them, or tuning 1 parameter in isolation without measuring the downstream effect, leads to pipelines that produce inconsistent results.

  • What poor metrics indicate: Variable retrieval quality across query types. Low faithfulness scores. Answers that are thematically related to the question but technically wrong. Or the inverse problem: high accuracy at unnecessarily high compute cost because the pipeline is over-provisioned.
  • Where to look: This is where systematic evaluation becomes essential. Test different combinations of chunking strategy, embedding model, and retrieval configuration against your evaluation dataset. Measure the results. Change 1 variable, reevaluate. This is tedious work when done manually, and it's precisely what Auto RAG (Chapter 7) automates: it runs the permutations, evaluates each configuration, and surfaces the combination that best fits your accuracy targets and cost constraints.

The cost-accuracy tradeoff

A common instinct when quality is low is to throw a bigger model at the problem. Sometimes that helps, but sometimes it doesn't because the bottleneck was in retrieval rather than generation. And sometimes a smaller, well-configured pipeline outperforms a larger one that's poorly tuned.

The ability to map cost against accuracy across different pipeline configurations is one of the most practical benefits of building evaluation into your development process. It keeps teams from overspending on inference when a configuration change would have solved the problem, and from underinvesting when the model genuinely needs more capacity. Auto RAG provides this tradeoff analysis directly, and Red Hat AI offers llm-d—an open, Kubernetes-native serving stack for scaling and optimising distributed LLM inference in production—which uses resource optimization for predictable, scalable performance without unnecessarily overprovisioning.

The next chapter covers what to do when your retrieval quality is good but the generation model still isn't interpreting domain content accurately. That's where fine-tuning comes in.

Chapter 5: Improving RAG through model customization

There's a common misconception that RAG and fine-tuning are competing approaches, but in practice, enterprise-grade RAG pipelines typically need both. They solve different problems, and combining them produces better results than either one alone.

RAG versus fine tuning: When to use which

  • RAG is best for data that changes frequently, such as standard operating procedures, product documentation, policies, or pricing. The model doesn't need to memorize this information—it simply needs access to the latest version at query time. This is what your retrieval pipeline provides.
  • Fine-tuning is best for knowledge that's stable, such as domain-specific reasoning patterns, terminology, or response formats. Fine-tuning teaches the model common terminology within the domain context to help it interpret retrieved content more accurately. An embedding model may know to retrieve documents about financial hedging, but if the generation model doesn't understand the difference between a hedge fund strategy and a garden hedge, the answer will still be wrong. Fine-tuning closes that gap.
  • Retrieval-augmented fine-tuning (RAFT) combines a fine-tuned model and RAG to deliver deep alignment to enterprise data for greater accuracy and relevance than either technique alone. It trains the model specifically on the task of reasoning over retrieved documents, teaching it to identify which parts of retrieved context are relevant and which are distractors. For domain-specific enterprise use cases, RAFT consistently outperforms either technique in isolation.

How evaluation tells you when to fine-tune

This chapter comes after evaluation because most teams don't realize they need fine-tuning until they've run a proper evaluation pass and seen the results.

The signal is specific: high context recall and precision, but low faithfulness or answer relevance. That pattern means the retrieval pipeline is doing its job and the right documents are being found and delivered. The problem is that the generation model can't interpret them well enough to produce an accurate answer.

When you see that pattern, prompt engineering is the 1st approach to try. Adjusting how you instruct the model to use the retrieved context can sometimes close the gap. If it doesn't, fine-tuning is the next step. It's an extension of the pipeline you've already built, not a restart.

The labeled data problem

The most common objection to fine-tuning in enterprise is, "We don't have enough labeled data." Building a training dataset by hand is expensive and slow, especially in specialized domains where the annotators need subject matter expertise.

When private data is unavailable or insufficient, Red Hat AI provides modular workflows to generate synthetic data. SDG Hub (part of Red Hat AI) addresses this by generating synthetic training data from your existing enterprise documents, allowing teams to expand and refine their datasets. You can use SDG Hub to feed in your PDFs, text files, and structured data to produce a curated dataset of question-answer pairs aligned to your domain. The result is training data that would have taken weeks to create manually, produced in a fraction of the time.

Synthetic data isn't a complete replacement for expert-labeled data in every case, but it gets teams past the cold-start problem and into a working fine-tuning loop. You can always augment with manually labeled examples later.

Fine-tuning with Training Hub

Fine-tuning a model for enterprise use involves several moving parts: selecting a base model appropriate for your domain, choosing between parameter-efficient methods such as low-rank adaptation (LoRA) and full fine-tuning, managing the compute infrastructure to run the job, and then verifying that the fine-tuned model actually improved things before you put it in production. Coordinating all of this across different tools and environments adds friction, especially for teams running their 1st fine-tuning job.

With Red Hat AI, you can use Training Hub to manage the full workflow in 1 place.

  1. Select a base model from the validated model repository.
  2. Configure the job, choosing between LoRA and full fine-tuning, depending on the scale of customization you need.
  3. Run it on Red Hat AI distributed training infrastructure.
  4. Use LM-Eval to evaluate the fine-tuned model before promoting it to your RAG pipeline.

That last step matters. A fine-tuned model should measurably improve your evaluation scores on the metrics that triggered the fine-tuning decision in the 1st place. If faithfulness was the problem, faithfulness should go up. If it doesn't, the training data or configuration needs adjustment, and you have the evaluation framework from Chapter 3 to guide that iteration.

Agentic RAG

Fine-tuning isn’t the only tool that can complement RAG.

While standard RAG follows a rigid, linear pipeline, agentic RAG integrates agentic workflows to create an active, decision-making ecosystem. These workflows help the LLM to build step-by-step plans to tackle complex problems systematically, treating retrieval as a dynamic tool rather than a hardcoded sequence.

Autonomous reasoning

Central to this is the ReAct (Think, Act, Observe) pattern, shifting the system from one-off lookups to autonomous reasoning where the agent decides what, when, and how often to retrieve:

  1. Think. The agent analyzes what data is missing or required and plans its next step.
  2. Act. It invokes the appropriate tool, such as calling a Model Context Protocol (MCP) tool to retrieve contracts or claims.
  3. Observe. It evaluates the results. If context is sufficient, it generates the final response. Otherwise, it loops back to step 1.
  4. Repeat. The pattern repeats until the agent has enough verified context to make a decision.

Key pillars of agentic RAG

  • Autonomous tool selection: The LLM independently decides which RAG tools to call and in what order on the fly.
  • Multisource retrieval: The system chains queries across disparate databases, APIs, and document stores in a single session.
  • Human-in-the-loop escalation: If confidence falls below a safe threshold, the agent routes the case to a human reviewer with the full reasoning trail.
Shiny red letters A, R, and G with two silver star-shaped sparkles

The next chapter shows how all of these components, from parsing to evaluation to fine-tuning, connect in Red Hat AI and how different roles divide the work.

Chapter 6:The Red Hat RAG workflow in practice

The previous chapters covered each piece of the RAG system individually: the 3 pipelines (indexing, retrieval, maintenance), evaluation, and model customization. This chapter shows how those pieces fit together on Red Hat AI, and how different roles in your organization divide the work.

The end-to-end workflow

A production RAG workflow on Red Hat AI runs as a continuous loop across 5 activities.

  1. Build the indexing pipeline. Raw enterprise documents are parsed and ingested using Docling. Content is normalized, chunked, embedded, and stored in the vector store of your choice, with metadata preserved for downstream filtering, citation, and audit.
  2. Stand up the retrieval pipeline. The retrieval configuration goes live: dense, sparse, or hybrid retrieval, with optional reranking and prompt construction discipline. Inference runs on the vLLM runtime. To handle scale, Red Hat AI uses llm-d, a distributed inference framework that provides predictable, scalable performance by disaggregating the inference pipeline into modular services. AI Guardrails monitors model inputs and outputs in real time.
  3. Evaluate. LM-Eval runs the pipeline against your evaluation dataset, scoring context recall, context precision, answer faithfulness, and answer relevance. If scores fall below your accuracy targets, the metrics point you toward the bottleneck: retrieval, generation, or both. EvalHub provides a unified control plane for these runs and an audit trail of how the system performed before and after each change.
  4. Improve when needed. If the gap is in retrieval, the response is configuration: reevaluate chunking, embedding model, retrieval strategy, or reranking. If the gap is in generation, SDG Hub generates synthetic training data from your enterprise documents, and Training Hub runs the fine-tuning job. The fine-tuned model is reevaluated with LM-Eval before it replaces the baseline. MLflow tracks every configuration and result, so changes are reproducible rather than tribal knowledge.
  5. Maintain continuously. The maintenance pipeline keeps the corpus correct as the underlying data changes. It propagates updates and deletes from source systems, executes right-to-be-forgotten requests with an audit trail, reembeds when models or chunking strategies change, synchronizes access control updates, and detects corpus drift. When drift is detected, it triggers a fresh evaluation cycle, closing the loop back to step 3.

This isn't a setup you only do a single time. It is a cycle. As enterprise data changes and user needs evolve, the system continuously reevaluates and improves, rather than silently degrading in production.

Roles across the workflow

Red Hat AI provides role-based interfaces that map to how enterprise teams actually organize around AI projects. Each role touches all 3 pipelines, but the center of gravity is different for each.

  • Platform engineers work through AI Hub. They deploy and manage the foundational infrastructure: models, MCP servers, API endpoints, graphics processing unit (GPU) resource allocation, and access controls. With llm-d, platform engineers have the tools to predictably manage the resources of a large cluster. This role also owns the compliance and governance layer, which means role-based access control (RBAC) configuration, data privacy controls, and ensuring that sensitive internal data isn't exposed to public models or third-party libraries. On the maintenance pipeline, platform engineers own the policies for retention, deletion, and audit, along with the infrastructure that executes them.
  • AI engineers work through Gen AI Studio. They discover available models, experiment with retrieval and generation configurations, prototype RAG workflows, and promote validated configurations to production. The day-to-day challenge for this role is getting usable data out of the systems where it actually lives: customer relationship management (CRM) systems, data warehouses, PDFs, emails, and traditional systems. Gen AI Studio provides a workspace to iterate on those integrations without waiting for infrastructure changes. AI engineers also define the triggers that drive the maintenance pipeline, deciding when source-system changes warrant reindexing.
  • Machine learning (ML) engineers are the connective tissue between experimentation and production. They take the workflows that AI engineers prototype and industrialize them into production-ready pipelines with proper error handling, logging, and automation. They own observability and monitoring across all 3 pipelines, making sure that applications in production don't become gaps where performance degrades unnoticed. Corpus drift detection, model drift detection, and the evaluation cycles they trigger sit with ML engineers.

The next chapter looks at where this workflow is headed: Auto RAG, and the shift from static pipelines to systems that optimize themselves.

Chapter 7: Auto RAG and the road ahead

Everything covered in the previous chapters assumes a manually configured pipeline. A user asks a question, the system retrieves chunks, and the model generates a response. The retrieval strategy, chunking approach, and embedding model are chosen upfront and stay fixed unless someone manually reconfigures them.

Auto RAG removes the guesswork from that configuration process. Instead of relying on intuition or trial and error to select the right pipeline settings, Auto RAG systematically evaluates different combinations of parameters and recommends the configuration that performs best for your data and use case.

From manual tuning to automated evaluation

The shift from traditional RAG to Auto RAG is, in essence, the shift from hand-tuned pipelines to empirically optimized ones. An Auto RAG system tests different chunking strategies, embedding models, retrieval methods, and pipeline settings against your actual data. It measures each combination's performance and surfaces the configuration that delivers the best results.

This matters because real enterprise data varies widely. A pipeline optimized for short FAQ-style documents may underperform on long regulatory filings. A chunking strategy that works for code documentation may miss critical context in legal contracts. Without systematic evaluation, teams are left guessing, and those guesses compound across every stage of the pipeline.

What AutoRAG optimizes

A RAG pipeline has many moving parts, and each choice affects accuracy, latency, and cost. AutoRAG treats the pipeline as a modular optimization problem and systematically evaluates configurations across these stages:

  • Document parsing and chunking. Different document types respond to different splitting strategies. AutoRAG evaluates recursive, hierarchical, and hybrid chunking methods with varying chunk sizes and overlap settings to find configurations that preserve the most useful signal for retrieval.  
  • Embedding model selection. The choice of embedding model determines how well queries match relevant documents. AutoRAG benchmarks multiple embedding models against your evaluation dataset to identify which produce the best retrieval results for your specific content.
  • Retrieval strategy. AutoRAG tests simple vector similarity search, hybrid search (combining dense vector and sparse keyword-based retrieval), and different reranking approaches. The optimizer explores which retrieval method yields the highest precision and recall for your use case.
  • Generation parameters. AutoRAG evaluates different foundation models and generation settings to find the combination that produces the most faithful and correct answers from retrieved context.

AutoRAG on Red Hat AI

Building a production-grade RAG pipeline manually means weeks of trial and error: testing chunking strategies, comparing embedding models, tuning retrieval, and evaluating answer quality. The optimal configuration for dense technical documentation differs from that for short product descriptions or legal filings. AutoRAG replaces this manual tuning with a structured, automated process.

Available as a Technology Preview in Red Hat OpenShift® AI 3.4, AutoRAG provides a guided experience through Gen AI Studio. AI engineers upload their documents and an evaluation dataset, select candidate foundation models and embedding models, and launch an optimization run. AutoRAG then systematically tests RAG pipeline configurations, evaluating each combination on the evaluation dataset using metrics such as answer faithfulness, answer correctness, and context correctness.

Two glossy, three-dimensional red gears

Next steps and resources

Where to start

If you're building your 1st RAG pipeline, begin with the pipeline architecture in Chapter 2 and set up LM-Eval before you do anything else. Having evaluation in place from day 1 will save you weeks of guesswork later.

If you already have a pipeline and it's underperforming, start with the evaluation framework in Chapter 3. Run a pass, read the metrics, and use the diagnostic guidance in Chapter 4 to identify your specific bottleneck.

If evaluation is pointing you toward fine-tuning, explore SDG Hub and Training Hub to get started without waiting for manually labeled data.

If you want to skip the manual configuration tuning entirely, explore Auto RAG to automate pipeline optimization from the start.

  • Learn more about the tools to build your RAG pipeline that are part of Red Hat AI.
  • Learn more about Red Hat AI.
  • Learn about AI/ML services from Red Hat.
  • Schedule a complimentary discovery session.
  • Hear more about large-scale data processing for RAG. 

Tags:Artificial intelligence

Red Hat logo

About Red Hat

Red Hat is the open hybrid cloud technology leader, delivering a trusted, consistent and comprehensive foundation for transformative IT innovation and AI applications. Its portfolio of cloud, developer, AI, Linux, automation and application platform technologies enables any application, anywhere—from the datacenter to the edge. As the world's leading provider of enterprise open source software solutions, Red Hat invests in open ecosystems and communities to solve tomorrow's IT challenges. Collaborating with partners and customers, Red Hat helps them build, connect, automate, secure, and manage their IT environments, supported by consulting services and award-winning training and certification offerings.

  • North America
  • Asia Pacific
  • Latin America
  • Europe, Middle East, and Africa
  • 888-REDHAT1
  • +6564904200
  • +5443297300
  • +0080073342835
  • www.redhat.com
  • apace@redhat.com
  • info-latam@redhat.com
  • europe@redhat.com
  • @red-hat
  • @redhat
  • @redhat
  • @red_hat

Copyright © 2026 Red Hat. Red Hat, the Red Hat logo, Ansible, and OpenShift are trademarks or registered trademarks of Red Hat, LLC or its subsidiaries in the United States and other countries. Linux® is the registered trademark of Linus Torvalds in the U.S. and other countries. The OPENSTACK logo and word mark are trademarks or registered trademarks of OpenInfra Foundation, used under license. All other trademarks are the property of their respective owners.

Red Hat logoLinkedInYouTubeFacebookXInstagram

Platforms

  • Red Hat AI
  • Red Hat Enterprise Linux
  • Red Hat OpenShift
  • Red Hat Ansible Automation Platform
  • See all products

Tools

  • Training and certification
  • My account
  • Customer support
  • Developer resources
  • Find a partner
  • Red Hat Ecosystem Catalog
  • Documentation

Try, buy, & sell

  • Product trial center
  • Red Hat Store
  • Buy online (Japan)
  • Console

Communicate

  • Contact sales
  • Contact customer service
  • Contact training
  • Social

About Red Hat

Red Hat is an open hybrid cloud technology leader, delivering a consistent, comprehensive foundation for transformative IT and artificial intelligence (AI) applications in the enterprise. As a trusted adviser to the Fortune 500, Red Hat offers cloud, developer, Linux, automation, and application platform technologies, as well as award-winning services.

  • Our company
  • How we work
  • Customer success stories
  • Analyst relations
  • Newsroom
  • Open source commitments
  • Our social impact
  • Jobs

Change page language

Red Hat legal and privacy links

  • About Red Hat
  • Jobs
  • Events
  • Locations
  • Contact Red Hat
  • Red Hat Blog
  • Inclusion at Red Hat
  • Cool Stuff Store
  • Red Hat Summit
© 2026 Red Hat

Red Hat legal and privacy links

  • Privacy statement
  • Terms of use
  • All policies and guidelines
  • Digital accessibility