Skip to content

AI & automation

RAG vs fine-tuning: which approach does your AI product need?

  • 11 min read
  • By ELIXIR Creative
Diagram in two lanes. Above, RAG: a query is used to retrieve from sources into a context, outlined in gold, which an unchanged model reads to give an answer. Below, fine-tuning: examples train the model itself, outlined in gold, so its outputs take the same shape every time.

RAG changes what a model can read; fine-tuning changes how it behaves. How to tell which problem you have, when to combine them, and when to use neither.

Teams building an AI feature often reach a point where the model alone isn't enough: it doesn't know the company's products, it answers in the wrong format, or it gets policy questions wrong. Two techniques come up straight away — retrieval-augmented generation (RAG) and fine-tuning — and they are often presented as alternatives. They aren't, quite. They change different things, and choosing between them starts with working out which thing is wrong.

The short answer

  • RAG changes what the model can read. While a request runs, the system finds relevant information — in your documents, databases or knowledge base — and places it in the model's context.
  • Fine-tuning changes how the model behaves. Additional training on examples adjusts the model itself, so it follows a format, a style or a task pattern more consistently.
  • Missing, private or changing knowledge points to retrieval. Inconsistent behaviour that instructions and examples can't fix points to fine-tuning.
  • Neither makes answers correct on its own, and both need evaluation.
  • Often the right first step is neither: better instructions, a few examples in the prompt, a database query or ordinary software.

RAG and fine-tuning solve different problems

A language model answers from two places: what it learned in training, and what is in its context — the text supplied with the request. RAG works on the second. Fine-tuning works on the first.

RAG changes what the model can access

RAG leaves the model unchanged. It adds a retrieval step that selects relevant information and supplies it alongside the question, so the knowledge behind a given answer is whatever was retrieved for it.

Fine-tuning changes how the model behaves

Fine-tuning continues training a model on examples — inputs paired with the outputs you want — so that the adjusted model produces that kind of output more consistently. The information supplied at run time stays the same; the model's tendencies change.

OpenAI's guide to optimizing LLM accuracy draws the line in the same place. It recommends working on context when "the model lacks contextual knowledge because it wasn't in its training set", when its knowledge "is out of date", or when it "requires knowledge of proprietary information" — and working on the model itself when it "is producing inconsistent results with incorrect formatting", when "the tone or style of speech is not correct", or when "the reasoning is not being followed consistently". The guide is written for OpenAI's models, but the distinction it draws is architectural and applies more widely.

What RAG actually does

A conceptual flow — real implementations vary a great deal:

  1. A query arrives — a user's question, or a task.
  2. Retrieval. The system searches a prepared collection of sources and selects what is most relevant. Often that means documents split into passages and indexed for semantic search; it can also be a keyword search, a database query, or a mix.
  3. Context construction. The selected passages are assembled with the instructions and the question into what the model will read — usually with each passage's source, so the answer can cite it.
  4. Generation. The model writes its answer from that context.

The term comes from a 2020 research paper that combined a pre-trained model's "parametric memory" with a "non-parametric memory": "a dense vector index of Wikipedia, accessed with a pre-trained neural retriever". It has since widened to mean almost any system that retrieves information and supplies it to a model while a request runs.

RAG's strengths follow from where the knowledge lives: outside the model. Update a document and the next answer can reflect it. Restrict which documents a user may retrieve and you restrict what the model can tell them. Show the retrieved sources and a reader can check the answer.

Its weaknesses are just as structural. An answer can only be as good as what was retrieved, and retrieval can fail: the right passage isn't indexed, it was split in the wrong place, the search ranks it too low, or the source itself is wrong or out of date. OpenAI's guide names the two classic failures: "You can supply the wrong context, so the model can't possibly answer, or you can supply too much irrelevant context, which drowns out the real information and causes hallucinations." Retrieval grounds an answer in sources. It doesn't make the answer correct.

What fine-tuning actually does

Fine-tuning starts from an existing model and trains it further on your examples, each showing an input and the output you want. The result is a new version of the model whose behaviour has shifted towards those examples. Which models can be fine-tuned, which methods are available and what it costs all depend on the provider or the open model you use.

The examples are the work. OpenAI's guide calls preparing them "the most critical step, as your fine-tuning examples must exactly represent what the model will see in the real world." Inconsistent examples teach inconsistency.

Fine-tuning suits behaviour: a consistent output format, a house style, a specific classification or extraction task, a particular way of handling a kind of request. It is a poor substitute for a knowledge base that changes. What a model absorbs in training can't be updated without training again, can't be restricted per user, and can't be traced back to a source document. Using a fine-tuned model as a knowledge base means every change to the facts needs a new training run and a new round of evaluation.

It also adds work for the life of the product: a dataset to maintain, a training process to repeat, a model version to track, and a fresh comparison whenever the underlying base model changes.

RAG vs fine-tuning side by side

RAGFine-tuning
Primary purposeSupply relevant information while a request runsAdjust the model's behaviour through additional training
Knowledge freshnessAs current as the indexed sourcesFixed when training happened
Changing informationUpdate the source and re-index itRetrain, then evaluate again
Behaviour and styleShaped only through instructionsCan be made more consistent
Data requirementsA curated, maintained collection of sourcesA set of high-quality input and output examples
MaintenanceKeep sources current; maintain the index and retrieval qualityMaintain the dataset; retrain when needs or base models change
EvaluationRetrieval quality and answer quality, measured separatelyOutput quality against the target behaviour, compared with the base model
Typical failure modesWrong, missing or outdated context, and plausible answers built on itBehaviour learned too narrowly or too broadly; stale facts stated confidently
TraceabilityAnswers can cite the passages they usedNo link from an output back to a training example
Access controlRetrieval can be filtered by what each user may seeWhat the model learned is available to everyone who uses it

Choose RAG when...

  • The knowledge is private — internal documentation, contracts, product specifications: information no public model was trained on.
  • The information changes — prices, stock, policies, release notes: anything that would be out of date in a model soon after training.
  • Answers need sources, because a reader, an auditor or a colleague must be able to check where an answer came from.
  • Access differs by user, because one customer, team or role may see information that another may not.

Hypothetical examples:

  • An internal assistant answers employees' questions from the HR handbook, IT procedures and expense policy, and links to the paragraph it used. When a policy changes, the document is updated and re-indexed; nothing is retrained.
  • An online shop's product assistant answers questions from the current catalogue — specifications, sizes, availability — which changes every day.
  • A customer-service tool drafts replies using the customer's own order history and the current returns policy, retrieving that customer's records and no one else's.

Choose fine-tuning when...

Consider fine-tuning when the problem is how the model responds, and instructions and examples in the prompt haven't solved it:

  • The output format must be consistent across many variations of input.
  • A specific style or tone must hold across every response.
  • A narrow task pattern — classifying, extracting, transforming — must be performed the same way at high volume, where long instructions in every request become slow or costly.
  • A domain-specific way of responding — what to include, what to leave out, how to handle a kind of request — is easier to show by example than to describe.

Hypothetical examples:

  • A claims system must turn free-text incident descriptions into a fixed structured record using the organisation's own categories. Prompting handles most cases but drifts on unusual ones; a model fine-tuned on reviewed examples follows the schema more consistently.
  • A retailer wants thousands of product descriptions written in a precise house style, and long style instructions in every request are slow and still inconsistent.

Notice what these examples don't use fine-tuning for: teaching the model facts. Domain knowledge doesn't, by itself, call for fine-tuning. It usually belongs in retrieval, where it can be updated and cited.

When you need both

The two can serve different layers of one system. Retrieval supplies the facts for each request; fine-tuning shapes how the model uses them — the structure of the answer, how it cites its sources, what it says when the retrieved context doesn't contain an answer. OpenAI's guide puts it plainly: "These techniques stack on top of each other."

Combining them isn't automatically better. In the same guide's worked example — correcting grammar in Icelandic text, a behaviour problem rather than a knowledge problem — adding retrieval to the fine-tuned model "actually decreased accuracy". Each layer is something more to build, maintain and evaluate, and the combination has to be measured against each part on its own.

Hypothetical: a support tool retrieves the relevant help articles for each ticket, and a fine-tuned model writes the reply in the company's required structure — a summary, the steps, links to the articles used — and says plainly when the articles don't cover the question.

When you need neither

Before building a retrieval pipeline or a training dataset, check whether something simpler does the job:

  • Better instructions and context. Clear instructions, a few well-chosen examples in the prompt and the relevant information supplied directly solve many problems. OpenAI's guide says prompt engineering "is typically the best place to start" and is "often the only method needed for use cases like summarization, translation, and code generation."
  • Information that fits in the request. If the relevant material is small and stable — a short policy, one product sheet — supplying it directly is simpler than building retrieval.
  • A database query. "How many of this item are in stock?" is a query, not a language problem. A model can help turn a question into a query, but the answer should come from the database.
  • Deterministic software. Calculations, validations, eligibility rules and anything else that must be exactly right every time belong in ordinary code.
  • An ordinary application workflow. A good form, a search box and a well-organised help centre sometimes serve people better than a generated answer.

Starting simple also produces something you need anyway: a baseline. OpenAI's guide recommends having "a solid evaluation set from prompt engineering which you can use as a baseline" — the standard that any added complexity then has to beat.

A practical decision framework

  1. Is the problem primarily knowledge or behaviour? Missing or wrong facts point to retrieval. Inconsistent format, tone or task handling point to better instructions first, then possibly fine-tuning.
  2. Does the information change frequently? If it does, it belongs in a source you can update, not in a model's weights.
  3. Where does the source of truth live? In a database: query it. In documents: retrieval may fit. In people's heads: neither technique helps until it is written down.
  4. Does the output need a particular behaviour or format? Try instructions and examples first, and structured-output features where your provider offers them. Fine-tune only if those don't hold.
  5. How will correctness be evaluated? Build a test set with known good answers before choosing. For RAG, measure retrieval and answers separately; for fine-tuning, compare against the base model.
  6. What needs to be updated after launch? Documents may change daily, behaviour requirements occasionally, base models on the provider's schedule. Choose the approach whose upkeep you can sustain.
  7. What are the privacy and security constraints? Who may see which information? May the data be used for training, and where is it processed? Retrieval can enforce access on every request; a fine-tuned model can't.
  8. What is the simplest architecture that meets the requirement? Start there, measure, and add complexity only where the measurements show a gap.

What this means for an AI product

For a product, choosing between RAG and fine-tuning is an architecture decision with long consequences. RAG turns the knowledge layer into a data problem: sources, indexing, permissions, freshness. Fine-tuning turns part of the model into a versioned artefact you own: datasets, training runs, regression tests. Either way, evaluation becomes part of the product rather than a launch task — a test set that runs whenever the sources, the prompts, the model or the training data change.

Design so that you can change your mind. Keep retrieval behind a clear interface, keep prompts and model versions under version control, and keep the evaluation set independent of either approach. Moving from prompting to retrieval, or adding fine-tuning later, then becomes a change you can measure rather than a rewrite. Integration, retrieval and evaluation designed together are what AI development for products and workflows consists of in practice.

Conclusion

RAG and fine-tuning answer different questions. Retrieval decides what the model can read; fine-tuning shapes how it responds. Start by diagnosing the problem: knowledge points to retrieval, behaviour to better instructions and then, perhaps, fine-tuning. Some products need both, in different layers, and many need neither — a better prompt, a database query or ordinary software will do. Whichever you choose, the deciding factor is evaluation: a clear measure of a correct answer, applied before and after every change.

Sources

Checked on 2 October 2026.

Related articles

  • AI & automation11 min read

    AI agents for business: where they actually fit — and where they don't

    Where AI agents can earn a place in a business, where they're a poor fit, and what a process, its data and its owners need before an agent is worth building.

    ELIXIR Creative

  • AI & automation10 min read

    AI agents vs automation: what's the difference — and when to use each

    Automation follows rules you define; an AI agent decides its next step within limits you set. How they differ, when each fits, and how they work together.

    ELIXIR Creative

  • AI & automation11 min read

    What is an AI agent? How agents use tools, data and workflows

    An AI agent is a language model in a loop: it reads its context, chooses a tool, acts, checks the result and decides what comes next. Its parts, and its limits.

    ELIXIR Creative

THE AUDIT

Have a system that isn't working?

Describe it in the audit. ELIXIR looks at it before recommending anything.

Start an audit