Skip to content

RAG vs Fine-Tuning: Which One Does Your Business Chatbot Actually Need?

RAG gives a model your knowledge; fine-tuning changes its behaviour. A plain-English guide to choosing, with a comparison table and a decision checklist.

AI & Automation · 6 Oct 2026 · 4 min read

RAG vs Fine-Tuning: Which One Does Your Business Chatbot Actually Need?

The letters AI in white on a dark teal circuit board

ByAhsan NadeemLead Full-Stack Engineer

"Can you train ChatGPT on our documents?" is one of the most common requests we get. It's a fair question, but it usually mixes up two different techniques. In most cases what the business needs isn't training at all. It's retrieval-augmented generation, or RAG.

This guide explains both in plain English, shows when each one wins, and lists what a chatbot needs before you can put it in front of customers.

The one-line difference

A large language model already knows a lot about the world, but nothing about your refund policy, your price list or last week's product update. There are two ways to close that gap:

  • RAG looks up the relevant passages from your own content at the moment a question is asked, and hands them to the model along with the question. The model answers from what it was just shown.
  • Fine-tuning continues training a model on hundreds or thousands of example conversations, so it learns a pattern: a tone, a format, a way of classifying things.
Use RAG when the model is missing knowledge. Use fine-tuning when it's missing a behaviour.

How RAG works in practice

  1. Prepare the content. Documents, help-centre articles, PDFs and database records are split into small passages ("chunks").
  2. Index it. Each chunk is turned into an embedding, a list of numbers that captures its meaning, and stored in a vector index. If you already run PostgreSQL, the pgvector extension is often enough.
  3. Retrieve. When a user asks a question, the system finds the most relevant chunks. Combining keyword search with vector search (hybrid search) usually beats either one alone.
  4. Generate. The question and the retrieved chunks go to the model with instructions such as "answer only from these sources and cite them".
  5. Cite. The answer links back to the passages it used, so users and your team can check it.

When your documents change, you re-index the changed files. Nothing has to be retrained, which is why RAG suits content that changes weekly or daily.

How fine-tuning works

Fine-tuning takes an existing model and continues its training on your examples: an input and the ideal output, many times over. The model gets better at producing that kind of output, consistently and with shorter prompts.

What it doesn't do well is store facts you can rely on. A fine-tuned model may still state an old price with full confidence, and every time your facts change you'd have to rebuild the training data and train again. That's why vendors such as Anthropic point most customers towards good prompting and retrieval before fine-tuning, and why OpenAI's own guidance recommends trying prompt engineering first. Provider options and prices change often, so check their current documentation before you commit.

RAG vs fine-tuning, side by side

RAG

Fine-tuning

Best for

Answering from your own, changing content

Consistent tone, format or narrow tasks

Keeping facts current

Re-index the changed documents

Retrain with new examples

Citations

Natural: answers point to source passages

Not built in

Data you need

Your existing documents

Hundreds to thousands of curated examples

Time to a first version

Days to a few weeks

Weeks, mostly spent preparing data

Main ongoing cost

Retrieval plus longer prompts per answer

Training runs, and evaluation after each

Control over wrong answers

Strong: restrict answers to retrieved sources

Weaker: facts live inside the model

Access control

Filter what each user can retrieve

Hard: everything trained is in the model

A five-question decision checklist

  1. Does the answer depend on facts in your documents or systems? → RAG.
  2. Do those facts change more than once a quarter? → RAG.
  3. Do users need to see where an answer came from? → RAG.
  4. Have you tried a clear system prompt with good examples, and measured that it still fails? → only then consider fine-tuning.
  5. Is the task narrow and repetitive, like routing tickets into twenty categories at high volume? → fine-tuning (or a smaller fine-tuned model) can pay off.

When you need both

A customer-support assistant might use RAG for answers and a fine-tuned model to produce replies in your exact house style. In our experience a strong system prompt with a few examples gets you most of the way on style, so start there and add fine-tuning only when you can show the gap on real test questions.

What a production chatbot needs (that demos skip)

A RAG demo takes an afternoon. A chatbot you can put in front of customers needs more:

  • An evaluation set. Fifty to two hundred real questions with known good answers, run after every change, so you can tell whether a tweak helped or hurt.
  • Permission-aware retrieval. If an employee can't open a document, the assistant mustn't quote it to them either.
  • A freshness pipeline. New and edited documents are re-indexed automatically, and deleted ones are removed.
  • "I don't know" and a handoff. When retrieval finds nothing relevant, the bot should say so and offer a human, not guess.
  • Logging and feedback. Thumbs up or down on answers, and a way to review the bad ones.
  • Privacy rules. Decide what personal data may be sent to the model provider, and check the provider's data-retention terms.

What drives the cost

Exact prices depend on the provider and change frequently, but the cost drivers are stable:

Cost driver

RAG

Fine-tuning

One-off

Embedding your content once

Preparing data and running training

Per answer

Tokens for the question plus the retrieved passages

Tokens for the question (often shorter prompts)

When content changes

Embedding only the changed documents

Rebuilding examples and retraining

Engineering

Ingestion, retrieval quality, evaluation

Data curation, training, evaluation

For most businesses the engineering time, not the API bill, is the biggest line item. That's another reason to start simple and measure.

Frequently asked questions

Does RAG train the model on my data?

No. With RAG your documents are stored in your own index and only the relevant passages are sent with each question. The underlying model is not retrained.

Will the AI provider use my data to train its models?

Major API providers, including OpenAI and Anthropic, state that they don't train on business API data by default. Terms differ between consumer apps and APIs, and they change, so read the current data-usage policy of the provider you choose.

Do I need a separate vector database?

Not always. If you already run PostgreSQL, the pgvector extension handles many workloads. Dedicated vector databases become worthwhile at very large scale or for specialised search features.

Can a RAG chatbot read PDFs, Google Drive or Notion?

Yes. Each source needs a connector that extracts the text and keeps the index up to date. Scanned PDFs need text recognition (OCR) first.

How do I stop the chatbot from making things up?

Instruct it to answer only from the retrieved sources, show citations, return "I don't know" when nothing relevant is found, and test it against an evaluation set before every release.

We build RAG assistants and LLM features in Python and TypeScript, on top of the systems you already run. See our LLM integration and AI chatbot service, or tell us what you'd like your assistant to answer.

Back to all insights

Keep reading.

Have a project in mind?

Tell us what you're building, what's slowing your team down or what you'd like to automate. We'll come back with honest next steps and a clear estimate.

No obligation and no sales script.

Popular searches

Change theme