- Retrieval is the default for company knowledge: it stays current, cites its sources, and is cheap to change.
- Fine-tune to change how a model behaves, not what it knows.
- Three questions — freshness, traceability, behavior — settle most RAG vs fine-tuning decisions in five minutes.
- Hybrids are normal: retrieval for the facts, a small fine-tuned model for one repetitive step.
- Fine-tuning doesn't cure hallucinations; grounding, citations, and permission to decline reduce them, and an evaluation set tells you by how much.
It comes up in almost every first technical call, usually from the engineer the owner brought along: should we train the model on our documents? It is the architecture question we hear most often, and the answer is almost never "train."
Default to retrieval — RAG — for company knowledge: it stays current, it cites its sources, and it costs little to change. Fine-tune when you need to change how the model behaves — a strict format, a house tone, a narrow classification — not what it knows. Many production systems use both: retrieval for the facts, a small fine-tuned model for one repetitive step.
What each actually does
Retrieval hands the model the right pages at the moment of the question. Fine-tuning changes the model itself. Prompt engineering changes the instructions sent with each request. They solve different problems, and most of the confusion comes from treating all three as rival ways to teach a model facts.
Retrieval-augmented generation (RAG) — at question time, the system searches your documents, pulls the few passages most relevant to the question, and hands them to the model, which answers from them. The model itself never changes. Under the hood: documents split into chunks and stored as embeddings (numeric fingerprints of meaning), usually beside keyword search, because part numbers and names don't embed as gracefully as sentences.
Fine-tuning — continuing to train an existing model on example inputs paired with the outputs you want, so the new behavior lives in its weights. It is very good at teaching a pattern — a format, a tone, a set of labels — and a poor way to teach facts. The model absorbs them unevenly, can't say where it learned them, and repeats last quarter's version until someone trains it again.
Prompt engineering — changing the instructions, examples, and context that travel with every request. It is the cheapest lever and the one to exhaust first: say who is asking, show three good answers, and give the model explicit permission to say it doesn't know.
The mental model: retrieval is an open-book exam, and you can check which page the model used. Fine-tuning is studying — practice changes how the student writes, but on exam day they answer from memory, which is where confident mistakes come from. Prompt engineering is writing a clearer exam question.
The three questions that decide it
Three questions settle most cases in about five minutes. Does the knowledge change? Must every answer point to its source? Are you changing what the model knows, or how it behaves? We call this the freshness, traceability, behavior test. A yes to either of the first two means retrieval; only a behavior gap points to fine-tuning.
1. Freshness — does the knowledge change? Prices, policies, procedures, who handles what. With retrieval, an update is a re-index, and the next answer reflects it. With fine-tuning, it's a new dataset, a training run, and more testing — and between runs the model is confidently out of date. Even "rarely" is a retrieval answer: rarely is exactly when nobody remembers to retrain.
2. Traceability — must every answer point to its source? If anyone downstream has to verify an answer — an attorney, an auditor, a customer asking "says who?" — it needs a citation. Retrieval gives you that almost for free, because the system knows which passages it handed the model. A fine-tuned model can't show its work: there is no page to point at, only weights.
3. Behavior — are you changing what the model knows, or how it acts? "It doesn't know our return policy" is a knowledge gap: retrieval. "It won't hold our output structure across ten thousand calls" is a behavior gap. So is "it writes like a brochure, and we need it to sound like our adjusters." Those earn fine-tuning a hearing, once prompting has had a fair try.
| Question | If yes | If no |
|---|---|---|
| Freshness: does the knowledge change? | Retrieval | Keep going |
| Traceability: must answers cite a source? | Retrieval | Keep going |
| Behavior: is the gap how it acts? | Fine-tuning, after prompting fails | Retrieval and a good prompt |
Rule of thumb: If the fix for a wrong answer is "update the document," you want retrieval. If the fix is "show it fifty more examples of the right output," you may want fine-tuning.
What they cost — build, run, change
Prompting is the cheapest to build and to change. Retrieval costs more up front — ingestion, search, evaluation — and stays cheap to keep current. Fine-tuning puts its cost into a labeled dataset and training, then charges again whenever the knowledge moves. The price of launch matters less than the price of change.
| Prompt engineering | Retrieval (RAG) | Fine-tuning | |
|---|---|---|---|
| Build effort | Hours to days | Days to weeks | Weeks, dataset first |
| Per-query cost | Grows with prompt length | Model call plus retrieved text | Lowest when a small model suffices |
| Cost to update knowledge | Edit the prompt | Re-index the changed document | New data, retrain, retest |
| Traceability | Only what's in the prompt | Every answer can cite its source | None; it's in the weights |
| Typical failure mode | Instructions drift as prompts grow | Wrong passage, confident answer | Stale or blended facts, stated fluently |
| Data you need | A few good examples | Your documents, clean and permissioned | Many labeled input–output pairs |
The update row settles most company-knowledge projects by itself: if your documents move monthly, fine-tuning turns every edit into a training cycle. The per-query row is where fine-tuning wins honestly: at high volume, a small fine-tuned model can be faster and far cheaper than a large model reading retrieved text. That's a reason to fine-tune a step, not the knowledge.
The table also can't show model churn. We test new models the week they ship: a retrieval system moves to a better one with an evaluation run, while a fine-tune is pinned to its base model and must be retrained when that model is retired.
No column shows the costliest part of every option: evaluation, guardrails, monitoring, and the people who review output. We itemized those in the real cost of an AI assistant in production.
The default answer, and when it flips
Our default for company knowledge is retrieval plus a well-built prompt. It flips toward fine-tuning in four situations: a strict output format at high volume, domain language or tone prompting can't hold, latency or cost that calls for a small model, and classification with many labeled examples.
Strict output format at scale. When a downstream system needs the exact structure on every call, small deviations become incidents. Structured-output modes in most model APIs now handle the syntax — valid JSON, required fields — so start there. Fine-tune when the structure is semantic: which field a phrase belongs in, what counts as "resolved."
Domain language or tone. Adjusters, clinicians, and engineers write in a shorthand a general model reads well but doesn't write naturally. Examples in the prompt get you far; when thousands of drafts must sound like your people, fine-tuning holds the voice without a two-page style guide in every request.
Latency and cost on a small model. Some steps run constantly and need no brilliance: routing, tagging, pulling the same six fields. A small model fine-tuned for that one step is faster and far cheaper per call than a frontier-class model, and an open-weight one can run inside your environment when the data can't leave.
Classification with many labeled examples. If your help desk holds years of tickets labeled by the people who handled them, you already own a fine-tuning dataset. On a narrow classification, a small fine-tuned model often beats a large prompted one — on accuracy, not just cost.
Field notes
Both knowledge systems below run on retrieval, and the framework shows why. In one, the documents arrive new with every case; in the other, they are policies that get revised. In both, a person must be able to check an answer against its source. Freshness and traceability settle it.
Field note: For a US injury-law practice, medical records arrive as hundred-page scans, and a summary an attorney can't verify line by line is worse than no summary. The pipeline digitizes each record set, builds embedding-based retrieval over it, and drafts a chronological summary in which every sentence cites its source page; a paralegal verifies it before use. An evaluation harness measures citation accuracy before any prompt or model change ships. First case evaluation now takes hours instead of days.
In framework terms, fine-tuning never enters the room: every case brings a new record set, so there is nothing to train on in advance, and traceability is the job itself. The stack behind the injury-law records pipeline is in our case notes.
Field note: A children's-services nonprofit kept its program knowledge in binders and in the heads of long-tenured staff. The assistant retrieves from a curated store of policies, program manuals, and reporting templates — deliberately no case records and no personal data about families. Every answer cites its source document, or the assistant declines; it has nothing to guess from. New staff now self-serve the routine answers.
The decline is the detail worth copying: a retrieval system knows what it was handed, so it can tell when nothing relevant came back. A fine-tuned model can't, and answers anyway.
Where the documents live. In the injury-law pipeline, records never leave the firm's own cloud tenancy: models run through Amazon Bedrock in the firm's AWS account under a business associate agreement, with zero data retention and access mirroring existing case permissions. The nonprofit's assistant runs inside its existing cloud workspace, behind single sign-on. And an index can forget a document; a fine-tune can't, because what it learned lives in its weights.
Hybrid patterns that work
Production systems rarely pick a side. Retrieval supplies the facts, the prompt sets the rules, and a small fine-tuned model — where one appears at all — does a single repetitive step. Four patterns cover most of what we recommend, including the one where the honest answer is neither.
Retrieval plus a small fine-tuned extractor. A capable general model answers from retrieved passages; beside it, a small fine-tuned model does one narrow, high-volume job, such as pulling the same fields from every incoming document into a fixed schema, or labeling documents by type before indexing. Extraction is its own discipline, covered in our note on turning forms, PDFs, and voicemail into structured data.
Fine-tuned embeddings for domain vocabulary. When retrieval misses, the cause is often vocabulary. Teach the retriever that "MVA" in a medical record and "car accident" in a question mean the same thing, and answers improve without touching the model that writes them. Try hybrid keyword-and-vector search and a reranker first; when the gap persists, the embedding model is often the most useful thing to fine-tune.
Prompt caching for long, stable context. If the knowledge is modest and rarely changes — one handbook, one catalog — skip retrieval and put the whole document in the prompt. Most major model APIs now cache a repeated prompt prefix, making repeat calls faster and cheaper. Ask for quoted sections to keep a form of traceability.
When neither: good search and good documents. Sometimes people don't need a synthesized answer; they need the right document, quickly. A well-configured search over clean, well-titled documents costs less than either technique and can't invent anything.
Rule of thumb: Fine-tune the smallest thing that fixes the problem — the retriever before the writer, one step before the whole pipeline, a small model before a large one.
When not to do either
Most requests for fine-tuning aren't fine-tuning problems, and some aren't model problems at all. Before building anything, check whether the prompt ever got a fair chance, whether your documents are actually right, and whether the thing you want fixed is hallucination — which fine-tuning won't cure and retrieval only reduces.
Most "we need fine-tuning" requests are retrieval and prompt problems. The model "doesn't know our products" because nobody showed it the products. The answers feel generic because the prompt never said who is asking or what good looks like. Fix both before anyone books a training run.
Some are "your documents are wrong." Three versions of the returns policy in three folders, the newest unlabeled. Retrieval answers from whichever ranks highest; fine-tuning blends them into a confident average. The real project is a document cleanup, which costs less than either build.
Don't fine-tune to fix hallucinations. Training on your facts doesn't teach a model to stop inventing; it teaches it to invent in your house style. What reduces hallucination is grounding — answers built from retrieved passages, a citation per claim, permission to decline — plus an evaluation set that catches misses before users do.
And sometimes, don't build at all — not with us, not with anyone. If your documents already live in a platform with a built-in assistant, and nobody needs answers cited or permissioned beyond what it offers, use that. Custom retrieval is worth building when the documents, the access rules, or the cost of a wrong answer are yours alone.
Run it yourself
You can run the freshness, traceability, behavior test on your own case this week with one folder and an hour. Write down three answers, count one number, and name one person. The result usually tells you which architecture you need before anyone writes code.
- Write down the three answers for your case, one sentence each: does the knowledge change, must answers cite a source, is the gap knowledge or behavior?
- Count how often the knowledge changed last quarter. Open the folder the assistant would answer from and count documents edited in the last ninety days. If the number isn't zero, freshness has already voted.
- Check whether anyone will need to verify an answer. Name the person who would be blamed for a wrong one. If they'd want to see the source, traceability has voted too.
- If behavior is the gap, collect twenty real examples of the output you want and put them in the prompt first. If the behavior still won't hold, you have the start of a fine-tuning dataset — and the evaluation set you need either way.
Retrieval versus fine-tuning gets argued like a question of faith. In practice it's three questions and a cost table. The craft is in the boundary — what the model is handed, what it may say, how you find out when it's wrong — which is most of what AI and LLM engineering means, and the part of the work we enjoy most.
Questions we get asked
Is RAG better than fine-tuning?
For company knowledge, usually yes. Retrieval-augmented generation (RAG) answers from your current documents, can cite a source for every answer, and updates as soon as a changed document is re-indexed. Fine-tuning is better at changing how a model behaves — a strict output format, a house tone, a narrow classification — but it is a poor way to add facts. Many production systems use both: retrieval for knowledge, a small fine-tuned model for one repetitive step.
When should you fine-tune an LLM?
Fine-tune when the problem is behavior rather than knowledge, and good prompting has already failed. Typical cases are output that must follow a strict format across thousands of calls, domain language or tone that instructions can't hold, a narrow step you want to move to a smaller, faster, cheaper model, and classification where you already have many labeled examples. Don't fine-tune to add facts that change, or to stop hallucinations.
Can you combine RAG and fine-tuning?
Yes, and many production systems do. The most common pattern is retrieval for the facts and a small fine-tuned model for one repetitive step, such as extracting fields into a fixed schema or classifying documents before they are indexed. Another is fine-tuning the embedding model so search understands your domain vocabulary. Add each part only when it fixes a problem you have measured; otherwise you pay twice for complexity.
Does RAG stop hallucinations?
No, but it reduces them and makes them checkable. A retrieval system hands the model the passages it should answer from, so you can require a citation for every claim, let the model decline when nothing relevant is found, and measure citation accuracy before each change ships. It can still retrieve the wrong passage or misread the right one, so an evaluation set and human review of high-stakes answers stay in the design.
How much data do you need to fine-tune a model?
Less than most people expect for a narrow behavior, and more than most have for anything broad. A consistent output format or a single classification task usually needs hundreds of clean, representative input-and-output examples, sometimes a few thousand, plus a separate held-out set to test against. Coverage of the awkward cases matters more than raw volume. If you can't assemble that set, you aren't ready to fine-tune.
Is prompt engineering enough instead of RAG or fine-tuning?
Often, and it should always be tried first. Clear instructions, a few worked examples, and explicit permission to say 'I don't know' fix many problems blamed on the model. If the knowledge fits comfortably in the prompt and rarely changes, you may not need retrieval at all. Once documents outgrow the prompt, change often, or need citations, add retrieval. Fine-tune only when a behavior still won't hold after that.