The problem it solves
A language model knows nothing about your business. Asked directly, it will produce a plausible, confident, invented answer. Retrieval-augmented generation changes the question: search your own material first, find the passages that bear on the question, and hand those to the model with an instruction to answer using only this.
How it works
- Preparation. Documents are split into passages and stored in a searchable index of meaning.
- A question arrives and is converted the same way.
- Retrieval. The index returns the closest passages — search by meaning, not just keywords.
- Generation. The model answers from the supplied passages, and should say so when the answer is not there.
- Citation. Because you know which passages were used, the answer can link to its sources.
RAG or fine-tuning?
Fine-tuning adjusts behaviour — tone, format, specialised style. It is poor at teaching facts, and those facts cannot be updated without training again. For business knowledge the answer is almost always RAG. If your policy changes on Monday, you want the system correct on Monday.
The four things that determine whether it works
- What you index. Feed it three superseded versions of a policy and it will cite the wrong one.
- How you split documents. Splitting along sections and headings beats cutting every thousand characters.
- How you retrieve. Combine meaning-based search with keyword search, then re-rank.
- What happens when nothing is found. The system must say it does not know.
A question, worked through
Someone types: “What is the notice period in the current staff handbook?” The system does not ask the model to remember. It searches the index, finds the 2025 handbook section on leaving, ignores the 2019 draft that is still on the drive because you never indexed drafts, and hands the model three paragraphs plus the instruction: answer only from this, quote the clause, say if it is missing. The screen shows the answer and a link to page 14. The colleague can open the PDF. That is why people keep using it after week two.
Ask instead: “How many contracts contain a late-delivery clause?” RAG will retrieve a few contracts that mention it and guess. That question is a report. Write a query. Do not pretend a language model is a database.
What to put in the index — and what to leave out
Current policies, current product specs, current rate cards, closed job notes you are allowed to reuse. Not: every email since 2016, three copies of the same SOP, personal folders, or anything a given role must not see. Access control belongs in retrieval. If the warehouse role cannot open the salary file in the ordinary system, the assistant must not retrieve it either.
This curation is the unglamorous majority of a six-to-ten-week build. The model call is the last ten percent. Teams that skip curation get a demo that looks clever and a production system that cites the wrong year.
Demo versus something staff will trust
A demo indexes a clean PDF and answers three planted questions. Production has scans, headings that lie, tables the splitter cuts in half, and people who type “that thing we told the Bahrain client”. Hybrid search — meaning plus keywords — plus a re-ranker is usually the largest quality jump after “stop indexing junk”. Showing sources is the largest trust jump. “I don’t know” is the largest safety jump.
How we usually assemble it on .NET
An indexing job that splits on headings and stores passages with the source URL and the roles allowed to see them. A search service that does meaning plus keywords, then re-ranks. A controller action that retrieves, builds a prompt with “answer only from this”, calls the model, and returns the text plus citations. If retrieval is empty, the action returns “I don’t know” and does not call the model. That last branch is the one demos skip.
None of this requires you to run your own model. It requires you to own the index and the permissions. The model is a service you call. Choosing which one is an afternoon. Cleaning the drive is the month.
A policy change on Monday
Notice period moves from thirty days to sixty. Someone drops the new handbook in the folder and, if you designed it, unpublishes the old one from the index the same morning. The next question cites page 14 of the 2026 file. That is the point of RAG. If the old file is still marked current, the system will be correctly wrong. Retrieval quality is a publishing problem as much as a search problem.
Give one person the job of “what is in the index”. It is not glamorous. It is why staff still use the assistant in month four. A shared drive with write-access for everyone and no owner is how you recreate the filing cabinet, with citations.
What a six-to-ten-week RAG build actually spends time on
Week 1–2: what is current, who may see it, kill the drafts. Week 3–4: split on headings, hybrid search, a prompt that refuses empty retrieval. Week 5–6: a real team asking real questions, citations on the screen, a log of “I don’t know”. Week 7–10: the ugly files — scans, tables, the query that means “that Bahrain job”. That last block is quality. Teams that stop at week 4 have a demo.
Fine-tuning does not appear on that calendar. Neither does “train on all email”. If a proposal leads with those, they are solving a different problem than “answer from our current policies”. Send them this page and ask which week they are pricing.
Cost sits with AI integration, not with a model licence. The model bill on a policy assistant for a fifty-person firm is usually small next to the week you spend killing drafts. Budget the owner of the index as a named hour each week after go-live, or the citations will rot and staff will stop asking.
If the question is a count or a total, stop. Write a query. RAG will retrieve three examples and guess. That guess is how you lose trust in week two. Keep RAG for “what does the current handbook say?”, not “how many contracts contain X?”.
When RAG is the wrong tool
Counts, totals, “how many contracts contain X”, anything that should be a query or a report. A model that retrieved three examples will guess the rest. Live balances belong in the database, not in a passage. Fine-tuning belongs when you need a tone or a format, not a fact. If the set of documents is three pages you could paste, just paste them. RAG is for when paste does not fit.
On ASP.NET Core this is ordinary software. See adding AI to an existing application and the wider guide to AI for business.
If staff will not open a citation, the assistant is already failing. Show the page. Make “I don’t know” louder than a fluent guess. That is RAG working, not a model upgrade.