The real cost of a custom AI assistant (the demo is 20% of it)

A demo proves the model can answer; the other 80% of the budget pays for knowing it keeps answering correctly — and catching it when it doesn't.

// key takeaways
  • A demo proves the model can answer; production proves it keeps answering correctly — that's the 80% you're paying for.
  • Evaluation is a test set of real questions scored before every change, not a feeling.
  • Running costs are usually modest; engineering and review time are the real budget.
  • A quote with a line for the model and no line for evaluation is a quote for a demo.
  • If nobody owns the review queue, don't build.

The demo works. Everyone in the meeting nods. Three weeks after launch, the assistant tells a customer the office is open on Sunday — and nobody can say whether that was the first wrong answer or the fortieth, because nobody bought the part that measures.

In our experience, a custom AI assistant costs roughly five times its demo. The demo is a model plus a prompt. Production adds evaluation (how often is it right), guardrails (what it must never do), data plumbing (retrieval, freshness, permissions), monitoring, and maintenance as models and documents change. The model bill is usually modest. The real budget — what we call the invisible 80% — is engineering and review time.

Here is what's inside that 80%, what it typically costs, and how to spot it in a quote.

The demo is 20% of the work

A demo proves one thing: that a model, given your documents and a good prompt, can answer the questions someone chose to ask it. It cannot prove the assistant will keep answering correctly — on questions nobody rehearsed, after the documents change, after the model is updated, at the volume real customers bring.

Demos are easy now, and that is the trap: a frontier-class model with a few pages of context answers most reasonable questions well. But the questions in the room come from people who know the documents. Customers misspell things, ask two questions at once, and ask about whatever the documents never covered.

Failures like the Sunday answer usually have a mundane cause — an old holiday-hours page still in the index, retrieved and trusted — and nothing is measuring, so the error surfaces as a complaint or not at all. That is how cheap builds die: quietly. They deliver the demo, priced as a product, and nobody misses the other 80% until months later, when staff double-check every answer and usage has faded. The 20/80 split is a rule of thumb, not a measurement; the direction is what matters.

The invisible 80%, itemized

The invisible 80% is five line items a demo doesn't need and a production assistant can't live without: evaluation, guardrails, data plumbing, monitoring and observability, and maintenance and ops. Cheap builds skip most of them, not out of malice but because none of them shows up in a demo — so none shows up in the quote.

Line itemWhat it isWhat skipping it looks like in month three
EvaluationReal questions, expected answers, scored before every changeNobody knows how often it's wrong
GuardrailsScope, refusals, personal-data handling, toneIt answers what it should decline, confidently
Data plumbingRetrieval, freshness, permissions mirroring existing accessStale answers; files reaching the wrong people
MonitoringLogs, drift, cost per query, escalation rateProblems arrive as complaints, weeks late
Maintenance and opsModel updates, document changes, the review queueBehavior shifts after an update; flags go unread

Evaluation: knowing how often it's right

Evaluation is a fixed set of real questions with known answers, scored before every change. It turns "it seems fine" into a number you can compare over time.

What cheap builds skip: the set itself; testing means someone types ten questions and decides it looks fine. What breaks: nobody can tell whether a change helped or hurt, so every change is a gamble and every complaint is an anecdote.

Guardrails: what it must never do

Guardrails are the rules around the model: what's in scope, when to decline or hand over to a person, how to treat personal information, what tone to hold. The reliable ones are checks in code, not sentences in the prompt.

What cheap builds skip: everything but one prompt line saying only answer questions about our business. What breaks: it answers what it should decline — a price it can't promise, a legal question, a detail from someone else's account — in the voice it uses for office hours.

Data plumbing: what it knows, and who may see it

Data plumbing is how the right documents reach the model: ingestion, retrieval, an index kept fresh, and permissions that mirror who can already open what. An assistant is only as current as its index and only as safe as its access model.

What cheap builds skip: freshness and permissions — documents uploaded once, by hand, visible to everyone. What breaks: old holiday hours keep getting quoted, a retired policy outlives its replacement, and eventually someone gets an answer from a file they were never meant to open.

Monitoring and observability: what it's doing right now

Monitoring means seeing the assistant at work: questions and answers logged under sensible retention, sources cited, hand-overs to a person, cost per query, and whether any of it is drifting.

What cheap builds skip: the logs, or the habit of reading them. What breaks: problems surface as complaints, weeks late. A rising escalation rate — often the earliest sign that something changed — goes unnoticed, and cost surprises arrive with the invoice.

Maintenance and ops: keeping it right as things change

Maintenance is the recurring work: re-running evaluation when the provider ships or retires a model version, re-indexing when documents change, growing the test set, and staffing the review queue where flagged answers land.

What cheap builds skip: the budget line and the owner; the build is priced as a project with an end date. What breaks: behavior shifts after an upstream model update and nobody connects the dates. And an assistant whose flagged answers nobody reads is an unreviewed assistant with extra steps.

What "evaluation" concretely means

Concretely, evaluation is a test set of typically 100 to 300 real questions, each with the answer a knowledgeable person would give, scored for correctness, citation, and appropriate refusal. It runs on every prompt, model, or document change, and someone reviews what flipped from right to wrong before anything ships.

The questions come from real traffic — the support inbox, call notes, what new staff ask in their first month — not from the team building the assistant. Include the awkward ones: misspelled, two-in-one, out of scope, recently changed. Each answer is then scored on three things:

  • Correctness. Does it match the expected answer? Dates and prices can be checked mechanically; a model grades the rest, and a person reviews a sample of every run.
  • Citation. Does it cite the right document, and does that document actually say it? A right answer with the wrong source is a near-miss worth knowing about.
  • Refusal. Does it decline what it should, and answer what it should? Over-refusal is a failure too: an assistant that declines everything is safe and useless.

Field note: For a US injury-law practice, we built a pipeline that turns hundred-page medical-record scans into a chronological timeline and a draft case summary, every statement linked to its source page. Not a chat assistant, but the same risk in sharper form: a summary an attorney can't verify line by line is worse than none. So an evaluation harness measures citation accuracy before any prompt or model change ships; a change that reads better but cites worse is caught before it reaches a case file.

Evaluation tells you whether the assistant is right, not whether it was worth building. That takes a baseline from before the build, which we cover in how to measure an AI project before and after.

What it costs — honest bands

In our experience, a proof of concept takes days, a production assistant takes weeks to a few months of engineering, and model usage at small-business volumes runs from tens to low hundreds of dollars a month. These are bands we typically see, not quotes; what moves a project is scope, sources, and the cost of a wrong answer.

Line itemWhat we typically seeWhat moves it
Proof of conceptDaysHow clean and reachable the documents are
Production assistantWeeks to a few monthsSources, permissions, integrations, risk
Model usage at SMB volumesTens to low hundreds of dollars a monthQuestion volume, context length, model tier
MonitoringMostly build time; a modest running lineRetention rules, alerting
Human reviewRegular hours from someone who knows the answersLaunch weeks, new question types
Maintenance, yearlyA recurring share of the build, never zeroModel retirements, document churn

Vendors like to discuss the model bill because it's small and easy to estimate. It climbs when every question goes to the largest model, or when untuned retrieval drags in far more context than each question needs.

The real budget is people's time: engineering builds the invisible 80%, and review keeps it honest. Review is heaviest in the first weeks after launch, then settles into a routine.

Three things move an estimate most. Architecture: a proposal to train a model on your documents changes both the build and the maintenance line; we compare the options in RAG vs fine-tuning for company knowledge. Actions: an assistant that books, updates, or sends things has become an agent, and every action needs its own checkpoints — see what AI agents for business actually are. And the price of a wrong answer, which sets the size of the test set and the hours of review.

Rule of thumb: If a quote has a line for the model and no line for evaluation, it's a quote for a demo.

Field notes

Two assistants we've built show what the invisible 80% looks like when it's designed in rather than bolted on. In both, the engineering that mattered most was deciding what the assistant would never see and never say — and in both, those limits are features, not compromises.

Field note: For a children's-services nonprofit, we built a private assistant that answers staff questions from the organization's own program documentation and policies. Deciding what it would never see came first: the knowledge base holds policies, program manuals, and reporting templates — deliberately no case records and no personal data about families. Access rides the existing single sign-on, scoped per role. Every answer cites its source document or the assistant declines; it has nothing to guess from. New staff self-serve routine answers, and program leads get hours back.

Field note: For a senior-living organization, the highest-value case in an AI-readiness evaluation was a family-facing assistant answering from the organization's own published materials. The phase-one prototype was deliberately grounded only in public and marketing materials — zero resident data — so it could ship while data governance matured on its own timeline. It shipped as phase one, with feedback capture wired in to inform the phase-two decision.

Notice what these limits do to the budget. Cite or decline is a guardrail and an evaluation criterion at once: every answer can be checked against its source. And an assistant that sees only public materials needs lighter plumbing and governance, so it can go live while the harder questions take the time they need.

When an assistant does touch client or customer data, decide where that data goes first. In our work that has meant four design decisions: personal identifiers masked inside the client's environment before any external model call; those calls made under zero-data-retention terms; for the most sensitive records, models consumed inside the client's own cloud tenancy; and access that mirrors existing permissions, so nobody can open a file through the assistant that they couldn't open already. All four are cheap on day one and expensive to retrofit.

Build, buy, or don't

Buy an off-the-shelf tool when the questions are generic, the answers are public, and a wrong answer costs little. Build a custom assistant when it must answer from private documents, respect permissions, or connect to your systems. And don't build at all when there's no owner for the review queue, no documents worth answering from, or no volume.

When to buy. If you need a website assistant that answers hours, services, and directions from your public pages, a $30-a-month off-the-shelf tool will do it well enough, and we'd recommend it over anything we could build. Buying moves the software half of the invisible 80% to the vendor. The other half stays with you: someone still checks what it says and updates the pages it reads. The Sunday answer can happen on a subscription tool too.

When to build. Private documents, per-role permissions, integration with the systems where work happens, and a real cost to a wrong answer. That is where the invisible 80% earns its keep.

When not to build — from anyone, including us. Three conditions stop the project:

  • No owner. If nobody will own the review queue — reading flagged answers, fixing sources, adding test questions — don't build. An unowned assistant degrades quietly until people stop using it.
  • No documents worth answering from. If the knowledge lives in two people's heads, or the documents contradict each other, the assistant will faithfully reproduce the confusion. Writing the documents down is the project.
  • No volume. If these questions come up a handful of times a week, a well-organized shared page and the person who knows the answers will beat any assistant.

We would rather say this before a build than explain it after one.

Run it yourself

Before you sign anything, do four things this week: send every vendor the thirty questions, ask to see their evaluation set, ask who owns the review queue, and start your own test set. A vendor who comes back blank on the first three is selling a demo priced as a product.

Ask the 30 questions. Our checklist, 30 questions to ask before you buy an AI assistant, covers evaluation, guardrails, data handling, running cost, and what happens when you want to leave — the questions we'd want asked of us. Send it to every vendor on your shortlist, us included, and compare the answers.

Ask to see the evaluation set. Not a description — the set itself, or a representative slice: questions, expected answers, the latest scores, and what happened to them the last time the model changed. A vendor with an evaluation practice can show you this quickly. A vendor without one will describe a rigorous testing process.

Ask who owns the review queue. On their side and yours: who reads flagged answers, how fast, and who fixes the source when the source was the problem. If your side's answer is "we'll figure it out," figure it out first.

Start your own test set. Pull real questions from your inbox or support log and write down the answer your best person would give. A few dozen is enough to start, and nothing is more useful to whoever builds this.

A demo is a promise that the model can answer. Production is keeping that promise on questions nobody rehearsed, after changes nobody announced, with a named person accountable when it slips. That work is most of the budget, and every quote you read should show it. If you'd like the invisible 80% priced line by line for your case, that is what our AI and LLM engineering practice does.

Questions we get asked

How much does it cost to build a custom AI assistant?

In our experience, a proof of concept takes days of engineering and a production assistant takes weeks to a few months. The demo is roughly 20% of that work; the rest is evaluation, guardrails, data plumbing, monitoring, and maintenance. The estimate moves with the number of sources, permissions, integrations, and how expensive a wrong answer is. Treat any quote without an evaluation line as a quote for a demo.

What does an AI assistant cost per month to run?

What we typically see at small-business volumes is model usage of tens to low hundreds of dollars a month, depending on question volume, how much context each question carries, and which model tier answers. That is usually the smallest line. The larger running costs are people's time: someone reviewing flagged answers and keeping documents current, plus periodic engineering when models or sources change.

What is LLM evaluation and why does it matter?

LLM evaluation is a fixed test set of real questions with expected answers, scored for correctness, correct citation, and appropriate refusal, and re-run before every prompt, model, or document change. It matters because language models fail quietly and confidently. Without evaluation, nobody can say how often the assistant is wrong or whether the last change made it better or worse; every complaint stays an anecdote.

What are guardrails in an AI assistant?

Guardrails are the rules around the model: which topics are in scope, when it must decline or hand over to a person, how it treats personal information, and what tone it keeps. The dependable ones are checks in code, not just instructions in the prompt. For an assistant that answers from documents, the strongest guardrail is simple: cite a source or decline. A good refusal is a feature, not a failure.

Why do AI assistants fail after launch?

In our experience, AI assistants rarely fail dramatically after launch; they fail quietly. The demo proved the model could answer, but nobody built the parts that keep it right: no evaluation set to catch regressions, stale documents in the index, no monitoring of escalations, and no owner for the review queue. Answers drift, people learn not to trust it, usage fades, and nobody can say why.

Should we build a custom AI assistant or buy one?

Buy an off-the-shelf tool when the questions are generic, the answers come from public pages, and a wrong answer costs little. Build custom when the assistant must answer from private documents, respect existing permissions, or connect to your systems. Don't build at all if nobody will own the review queue, the documents aren't worth answering from, or the volume is too low. Either way, you still own whether the answers are right.

// written by

Bogdan Jovanovic · Founder & Partner · Lead & Innovation

Engineer first, consultant second. Bogdan leads Tensorika's client work and its engineering direction — the systems, the standards, and the new ideas that become services.

the people behind Tensorika →

// talk to an engineer

Building an assistant that has to work in production?

We design and ship LLM systems with evaluation and guardrails built in — and we're honest about the 80% of the work the demo doesn't show.