Do not start with a platform
If the application already runs the business, the job is to attach a capability, not to rewrite the capability in a new stack. The sequence we use is boring on purpose.
- Pick the task. One document type, one mailbox, one “draft this reply” button. Write the definition of done in a sentence.
- See the data. Can the application already reach the files or records? If not, that integration is the first sprint.
- Ground the model. Use RAG or a structured extract. Do not ask the model to remember your customers.
- Read-only first. Show a draft or a suggested field. A person accepts. Only then write back, using the same roles the screen already has.
- Measure. Hours saved, corrections made, cases escalated. Without that, you cannot decide whether to widen.
When the existing application is the problem
If there is no API, no audit trail, and no way to add a button without a vendor, you are in legacy modernisation first. Bolting AI onto a system you cannot change produces a second system beside it — and staff will use neither reliably.
A concrete first feature on a .NET system
You already have an enquiry or case screen. Add “Draft a reply from this record and the last three notes.” The action gathers those notes, optionally retrieves a policy passage with RAG, calls the model with a JSON schema, and puts the draft in the box the user already edits. They send it. You log the draft, the final text, and whether they changed it. After a hundred sends you know if the feature is saving time or creating cleanup.
That is a feature. It is not a chatbot product. Staff do not leave the screen they already trust. The same pattern works for “extract fields from this attached PDF” and “summarise this file for the next person”.
Risks that are easy to skip
Sending customer data to a public model without a contract. Letting the model see files the current user could not open. Writing back without a schema, so a bad parse corrupts a record. Measuring nothing, then arguing about feelings in month three. None of these require a research team to avoid. They require the same discipline as any other integration.
What must already be true in week one
The files or records exist somewhere a program can open. The user already has a screen where the result should appear. Someone can say, in one sentence, what “correct” looks like — a field filled, a draft sitting in a box, a ticket tagged. If those three are missing, the first sprint is not a model. It is finding the files, adding an API, or writing the definition of done.
A shared drive with three copies of every policy is not “data”. It is a curation job. Budget it. Teams that skip it get a demo that answers planted questions and a production assistant that cites 2019.
How to measure the first hundred results
Log the draft, the final text, and whether a person changed it. After a hundred uses you will know the change rate. If people rewrite every draft, the prompt or the retrieval is wrong — or the job was never a drafting job. If they send half unedited, you have something. Hours saved is the number that matters; “people like the button” is not.
Pick the number before you start. Invoice keying time. Time to first reply on the routine ten questions. Number of “where is that file” pings to the same colleague. Move one number. Then talk about a second feature.
When to stop, or not to start
Stop if the application cannot accept a write you can undo. Stop if nobody will own the exceptions. Do not start if the request is “add ChatGPT to the company” with no task. A chat skin on the intranet is a login that goes quiet in week three.
If the sequence never varies and the input is already a form, skip the model. Use workflow automation. The model is for reading and drafting. The workflow is for moving the structured result.
A statement of work you can reuse
“Add a button on the enquiry screen: draft a reply from this record and the last three notes. Optionally retrieve the current policy passage. Return JSON that matches this schema. The user edits and sends. Log draft, final, and whether they changed it. The model must not see files this user could not open. No write to the customer without the click. After a hundred sends we look at change rate and hours.”
That paragraph is a project. “Integrate ChatGPT with our CRM” is not. Send the paragraph. If the reply is a platform diagram, keep looking. If the reply is a list of assumptions about your API and your roles, you have a partner.
Where this sits on a legacy screen
If you cannot add a button without a vendor change request, the first sprint is access, not the model. If there is no API, you are in legacy modernisation or a sidecar that staff will not open. Be honest about that in the brief. Bolting a chat widget onto a desktop app from 2009 produces a screenshot for the board and a process that still lives in Excel.
If you can add a controller action on ASP.NET, you are in the lucky case. Gather the record, retrieve if needed, call the model with a schema, validate, show the draft. That is a feature. It should be estimated like a feature — weeks — not like a transformation. The wider menu of what to bolt on is in AI for business.
Measure from day one. A hundred drafts with a change-rate is a result. A demo that answers three planted questions is not. If the change-rate stays high, fix retrieval or the task — do not add a second button. Widening a feature nobody trusts is how “AI projects” get a name in the office that you will not enjoy.
The wider choice of what to build is in AI for business. If you need a partner for the integration itself, that is AI integration.
Roles, logs, and the file the clerk must not see
The model call inherits the signed-in user. If that user cannot open the salary PDF in the ordinary system, retrieval must not return it. If they can open the case file, the draft may use it. Write that as a test, not as a slide: two accounts, two folders, one button. We implement it with the same authorisation the screen already has — not a second permission model the model “kind of understands”.
Log the user, the record id, the passages retrieved, the raw model output, and the text they sent. That log is how you answer “why did it say that?” on Thursday. Without it you will argue about feelings. With it you can fix a prompt or a folder in an afternoon.