The short version
A generative model is a very large statistical model of language. Given some text, it predicts what should plausibly follow, one piece at a time, and repeats until the answer is complete. Everything it produces is generated in the moment rather than retrieved from a store of facts.
The model is producing the most plausible continuation, not the true one. When those coincide, the results are excellent. When they diverge, the model states something wrong in exactly the same confident register as everything else.
Why it is good at language and bad at facts
Fluency and accuracy are separate problems. Grammar, tone, summarising a long document, rewriting something aimed at engineers so a customer can follow it — these are pattern problems, and the models are genuinely better at them than most people.
Specific facts are a different matter. Your customer’s account balance, the clause in your standard contract, last month’s delivery figures — none of that was in the training data. So it produces something shaped like the right answer. This is usually called hallucination. It is not a malfunction: it is the system working exactly as designed, applied to a question it should never have been asked unaided.
The engineering answer is to stop asking the model to remember things and instead hand it the relevant material at the moment of the question. That approach is retrieval-augmented generation.
What it is actually useful for
| Task | Why it suits generative AI |
|---|---|
| Summarising documents and threads | The source material is supplied, so accuracy is checkable |
| Drafting routine correspondence | A person reviews before sending; drafting is the slow part |
| Extracting structured data from documents | Output has a rigid shape, so errors are easy to validate |
| Classifying and routing inbound work | Accuracy is measurable against a sample |
| Answering questions from your own documents | Works well only when grounded in retrieved material |
Three failure modes to design around
- Confident invention. Ground every factual answer in supplied material, and show the user which source it came from.
- Inconsistency. The same question can produce differently worded answers. Where structure matters, constrain the output format and validate it.
- No sense of authority. The model cannot tell your approved policy from a draft somebody abandoned in 2019. Curating what it can see is part of the build.
A worked example
A supplier sends a two-page PDF. A person used to open it, type the supplier name, invoice number, date, GST and totals into accounts, then file the PDF. With generative AI in front of that job, the system reads the PDF, fills a form the accounts package already understands, and shows the human a side-by-side of the extract and the page it came from. The human corrects the one field the model misread — usually the GST treatment — and posts.
Notice what the model did not do. It did not decide whether to pay. It did not invent a supplier that was not in the book. It did not write to the bank. It drafted structured fields from a document you already had. That is the shape of work these models are actually good at.
What you should never ask it, unaided
- Last month’s margin, unless you just handed it the report.
- What a named customer is owed, unless you just retrieved the ledger line.
- Whether a clause in “our standard contract” is current, unless the current file is in the prompt.
- Anything that will be sent to a customer or a court without a person reading it.
Those are not edge cases. They are the questions people type first, because they sound like the questions ChatGPT handles on the open internet. Inside a business they are the questions that get you sued or simply not trusted.
How this shows up in software you already run
In an ASP.NET application the model is a service you call, not a product you log into. A controller action gathers the record, optionally retrieves related passages with RAG, calls the model with a tight prompt and a JSON schema, validates the result, and writes a draft the existing permission system can accept. The user never sees a blank chat box unless that is actually the job.
That wrapping — auth, schema, validation, logging — is most of the build. Choosing “which model” is an afternoon. See adding AI to an existing application.
Cost, honestly
A drafting or extraction feature on one document type, writing into one system, is typically three to six weeks if the files are already digital and the API exists. An assistant over a curated set of policies is six to ten weeks, mostly because of retrieval quality. Running cost is a few rupees to a few tens of rupees per call, which is nothing on a pilot and visible if the whole company asks all day.
What “grounded” means in a meeting
It means the model was handed the invoice, the policy, or the last three notes, and instructed to use only that. It does not mean “we connected SharePoint”. A dump of every file since 2016 is how you cite the wrong year. Curation — current documents, current permissions — is the work. RAG is the name for that pattern when the set is too large to paste.
If a vendor says the model “knows your business” after a week of training, ask what happens when the rate card changes on Monday. If the answer is “we retrain”, you bought the wrong thing. If the answer is “we re-index the file”, you are in the right conversation.
The same test applies to “we fine-tuned on your tickets”. Fine-tuning can teach a tone. It will not teach Tuesday’s balance. If the demo answers a question you did not put in the prompt or the retrieved pile, ask where that sentence came from. Silence is the answer you needed.
Bring one ugly PDF to the next meeting — the scan, the stamp over the total, the second page that is a stamp. Ask whether the extract still produces a schema you can validate. Demos use clean files. Your mailbox does not. That file is the project more than the model card on the slide.
Where staff will actually use it
On the screen they already open. A draft reply on the enquiry. Extracted fields on the invoice. A summary on the case file. A separate chat site is a demo. People will try it, then go back to the inbox. The wrapping — auth, schema, validation, logging — is why this is AI integration and not a prompt pack.
If your input is a table of numbers and you want a prediction, classical machine learning is cheaper. If the task requires several steps and taking action, what you need is an AI agent. For the wider picture, start with AI for business.