- If you can't say what "worked" means on day zero, it won't.
- Capture five numbers before the build and the same five after, over comparable weeks: hours, turnaround, error rate, coverage, and cost per unit.
- The success criterion belongs in the scope, in one sentence with one number and one time window.
- Whether people use the output or quietly re-check it decides adoption more than any accuracy score.
- Permission to stop at day 30 is part of a well-run AI project, not evidence of a failed one.
The question arrives at month three, usually in a budget meeting: did it work? Nobody can answer, because nobody wrote down what the work looked like at month zero. Every AI project we've been asked to look at after it "didn't work" had the same missing document: the baseline.
Measure an AI project by capturing five numbers before the build starts — hours spent, turnaround time, error or rework rate, coverage, and cost per unit — then reading the same five after, over the same kind of weeks. Write the success criterion into the scope. If you can't say what "worked" means on day zero, it won't.
Why "did it work?" is unanswerable without day zero
Without a day-zero number, "did it work?" turns into a debate about feelings. The people who built it remember the wins, the people who use it remember the bad week, and finance remembers the invoice. A baseline turns the question into arithmetic: the same five numbers, before and after, over comparable weeks.
The business consequence cuts both ways. Without a baseline, AI project ROI is a guess, and an unmeasured project fails twice: good ones get cut because nobody can prove they worked, and bad ones keep running because nobody can prove they didn't. The second mistake bills you every month.
Vendors rarely ask for a baseline, and not only because it means a week of unglamorous counting. A baseline is a commitment: a number the project can fail against. A demo is compared against nothing, which is why demos never fail. A baseline is compared against an ordinary Tuesday in your business — the only comparison that matters.
The five numbers
The day-zero baseline is five numbers about one workflow: hours spent, turnaround time, error or rework rate, coverage, and cost per unit. They are business KPIs, not model metrics, and each can be captured in one ordinary week with a shared spreadsheet. None needs a data team, and none depends on the AI existing yet.
Hours spent. Person-hours per week across everyone who touches the workflow, not only the person whose job title matches. Each person logs minutes per item, or a rough end-of-day estimate, in one shared sheet.
Turnaround time. Elapsed time from arrival to the moment the next step can start — what customers feel. Ten minutes of work can sit inside three days of waiting. Most systems already record both timestamps.
Error or rework rate. The share of items that come back after they were "done": a correction, a follow-up call, a re-keyed field. Mark each one that returns, and note where the problem was found — at the door, or three steps downstream.
Coverage. The share of items handled end to end without a person. Before the build it's usually zero, so capture the split instead: tag each item standard or exception. The standard share is the ceiling for what any system can cover.
Cost per unit. Everything the workflow costs — hours at a loaded rate, tools, rework — divided by the items it processed. After launch, add model usage, hosting, maintenance, and review time; leave those out and every project looks profitable.
| Number | Capture before (one week) | Read after | The trap |
|---|---|---|---|
| Hours spent | Minutes per item, per person | Same, plus review time | Hours saved that nobody redeploys |
| Turnaround | Arrival and completion timestamps | Same, logged by the system | Busy season vs. quiet season |
| Error or rework | Items that came back, and where | Same tally, plus escalations | Trusting demo accuracy |
| Coverage | Standard or exception, per item | Share finished without a person | Counting re-checked items as covered |
| Cost per unit | Hours × loaded rate + tools, ÷ items | Add model, hosting, upkeep, review | Leaving review and upkeep out |
Rule of thumb: If a number can't be captured in a week with a spreadsheet, it's the wrong number for day zero. A rough baseline taken this month beats a precise one that never gets taken.
The qualitative signals that actually decide adoption
The five numbers tell you whether the workflow got faster or cheaper. Four qualitative signals tell you whether it will survive: whether people use the output or check it, how often they escalate, whether a shadow process persists, and whether someone got their evening back.
Trust: do people use it, or check it? Ask what people do when the system hands them an answer. "I send it" and "I compare it with the original first" describe two different projects. Checking is fine when it's designed in, as a review step with an owner; undeclared checking adds a step the hours number never sees.
Escalation rate. The share of items handed to a person, read next to the error rate. Low escalation with rising errors means the system is confidently wrong; high escalation with few errors means it's cautious and can be tuned. Early on, a high rate is often honesty, not failure.
The shadow process test. Does the old spreadsheet still exist? If the team kept the previous process running "just in case," adoption hasn't happened, and you are paying for two processes. A shadow process in week two is normal. In week ten, it's a verdict.
Someone got their evening back. The least measurable signal, and often the most telling. When the person who stayed late every reporting day stops staying late, the project changed something no spreadsheet shows: whether good people stay. Ask about it at every review.
Field notes — what before-and-after looked like
Across four anonymized, representative cases from our case notes, the clearest change was turnaround: days became same-day or hours, and a monthly scramble became a weekly rhythm. In one, errors started getting caught at the door; in another, the saved hours went into decisions. We report outcomes qualitatively, by policy — the exact figures belong to the clients.
Document intake, a professional-services firm in Ontario. Before: staff re-keyed client forms, scans, and attachments, and errors surfaced downstream. After: no re-keying for standard submissions, turnaround from days to same-day, errors caught at the door. The lesson: log where each error is found, not just how many — caught at intake, a mistake is a quick fix; three steps later, it's a correction and a phone call.
Medical-records review, a US injury-law practice. Before: days of senior time reading hundred-page scans before anyone could evaluate a case. After: a first evaluation in hours, from a summary a paralegal verifies, with every statement linked to its source page. The lesson: speed only counts if the output is checkable. An unverifiable summary gets re-read from scratch, and the hours come right back.
Lead intake, a residential roofing company. Before: leads from phone, web form, and marketplaces waited, in storm season, until someone got to them. After: responses in minutes and quotes out the same day. The lessons: "when someone got to it" is a common day-zero answer and a finding in itself, and storm season has to be compared with storm season.
Client reporting, a digital agency. Before: a monthly scramble of hand-built reports, delivered too late to act on. After: a weekly rhythm, with account leads spending the saved time on decisions. Copy that last part: hours saved only count when they land somewhere the business values.
Write the success criterion into the scope
A success criterion is one sentence with one number and one time window, written into the scope before anyone builds. If "worked" isn't defined in the document everyone signs, it gets defined later by whoever is most disappointed or most invested — and neither is a fair judge.
Here is the shape, with an illustrative number: "By day 60, at least 80% of standard intake forms reach the case system on the business day they arrive." Your number comes from your baseline week. Compare that with the criteria we usually find in proposals: "improve efficiency," "save staff time," "modernize intake with AI." None has a number or a window, and all of them will be declared met.
Writing it down changes what gets built. That criterion requires three things a demo never needs:
- An evaluation set — real items from the baseline week with known correct answers, scored before every change. The week you spent counting doubles as test data.
- A review queue with an owner — the other 20% have to go somewhere, and someone is accountable for clearing it.
- Logging — arrival and completion timestamps recorded by the system, so the "after" numbers don't depend on anyone's tally.
None of that is exotic, and none of it is free: it's a large part of the cost side of the ROI math, which we itemized in the real cost of an AI assistant in production.
Field note: For a US injury-law practice, the records pipeline ships with an evaluation harness that measures citation accuracy before any prompt or model change goes live. Traceability to the source page stopped being a launch-day hope and became a test that runs on every change.
A note on privacy: if the workflow touches client data, so does the evaluation set, and it stays where production data lives — same environment, same access rules, identifiers masked before any external model call. The baseline sheet holds counts and timestamps, never names or file contents.
The 30/60/90 review — and permission to kill it
Put three review dates in the scope the day you sign it. Day 30 checks direction, day 60 checks the numbers against the baseline, and day 90 decides: scale, fix, or stop. Day 30 carries one extra job — giving everyone explicit permission to stop, because some projects should, and early is when stopping is cheap.
At day 30, the question isn't whether the criterion is met but whether the project is heading toward it. A "no" usually looks like coverage stuck because real inputs are messier than the baseline week suggested, review that takes longer than the old way of doing the work, or rising errors that a customer noticed before you did. Each is a reason to stop, re-scope, or swap the model step for plain software. None gets cheaper at month six.
The most expensive AI project isn't the one that stopped at day 30. It's the one that should have, and ran to month nine because stopping felt like admitting something. Sunk cost is the line item nobody writes down.
The same honesty applies to pilots. A pilot that can't reach production is a proof of concept, not a failure — if the criterion said so. A scope reading "in three weeks, find out whether the model reads our real forms well enough to justify a pilot" can end in a no and still be money well spent. The failure is a proof of concept sold as a pilot, or a pilot sold as production.
When not to do this: If nobody on your side can spend a week counting before the build, or nobody will own the day-30 review, don't start — not with us, not with anyone. The project will run and cost money, and at month three nobody will know whether it should have.
Common traps
Most bad AI measurement isn't dishonest; it's careless in five predictable ways. Teams compare the wrong weeks, count hours nobody redeploys, forget review time, trust demo accuracy, and move the goalposts once the project is underway. Each one can make a mediocre project look good, or a good one look bad.
Measuring the wrong week. A quiet-season baseline against a busy-season "after," or the reverse. Launch week doesn't count either: everyone is watching, so everyone is careful. Compare like with like, or compare per unit.
Hours saved that nobody redeploys. If saved hours dissolve into the rest of the day, the saving is real on the timesheet and zero on the P&L. Decide in the scope where the time goes.
Ignoring review time. A draft that takes seconds to generate and longer to check than it took to write by hand is a net loss with good marketing. Count reviewer minutes from day one.
Demo accuracy read as production accuracy. The demo ran on clean, hand-picked samples; production gets the blurry scan at 4:55 on a Friday. Measure accuracy on your real items — another reason the baseline week doubles as the evaluation set.
Moving the goalposts. When the criterion isn't met, the temptation is to redefine success: "the numbers are flat, but the team likes it." That may matter, but it had to be named on day zero to count. Change a criterion only in writing, with a reason, before the next review.
Run it yourself
You can take a day-zero baseline this week without buying anything. Pick one workflow, set up one sheet, have the people who do the work fill it in for five working days, and write the one-sentence criterion before anyone talks to a vendor, including us.
- Pick the workflow. One, not three: real volume, a clear owner, mistakes that cost something. Choosing between several? Our note on what to automate first has the triage score we use to rank them.
- Set up the sheet. One row per item, with columns for arrived, completed, minutes of hands-on work, standard or exception, came back, and where the problem was found. Counts and timestamps only — no client names, no file contents.
- Capture an ordinary week. Not a holiday week, not a launch week — and note which kind of week it was, so the "after" is read against the same kind.
- Do the arithmetic. Total hours, median turnaround (so one stuck item doesn't skew it), share that came back, share that was standard, cost per unit.
- Write the criterion. One sentence, one number, one time window — with the three review dates next to it.
If the counting was harder than it sounds — no timestamps, a process that changes with whoever is in, no agreement on what counts as an error — that's a readiness finding, not a measurement failure. Our honest AI-readiness checklist is the scorecard we use to sort that out before anyone builds.
A baseline is the least glamorous document in an AI project and the only one that answers the month-three question with a number instead of a mood. Skipping it saves a week and costs you the ability to know whether to keep paying. If you want an outside read of your criterion before the build starts, that's the kind of work our advisory practice does.
Questions we get asked
How do you measure the ROI of an AI project?
Capture five numbers for the workflow before the build starts: hours spent, turnaround time, error or rework rate, coverage (the share handled without a person), and cost per unit. After launch, read the same five over comparable weeks, and add the new costs: model usage, maintenance, and review time. The return is the difference, judged against a success criterion written into the scope before any work began.
What KPIs should an AI project track?
Track the workflow, not the model. The five that matter to a business are hours spent, turnaround time, error or rework rate, coverage (the share of items handled without a person), and cost per unit including review time. After launch, add two signals: the escalation rate, and whether staff use the output or quietly re-check it. Model accuracy matters only when it is measured on your real inputs.
Why do AI projects fail?
The pattern we see most often is not a bad model but a missing definition of success. Without a baseline and a written criterion, nobody can tell a slow start from a dead end, so weak projects run on and good ones get cut. Close behind: real-world inputs messier than the demo's, review time nobody budgeted, and a process that was never stable enough to automate.
What is a good ROI for AI?
There is no universal number, and anyone quoting one hasn't seen your workflow. A good return meets the success criterion written into the project scope and pays back the full cost — build, running, maintenance, and the time people spend reviewing output — within a period agreed before starting. If the hours saved aren't redeployed to something the business values, the return exists only on paper.
How long before an AI project shows results?
A well-scoped first workflow should show its direction within 30 days of real use and support a clear decision by day 90. At day 30, check whether people use it and whether the numbers are moving the right way. At day 60, compare them with the numbers captured before the build. At day 90, decide to scale, fix, or stop. No movement by day 30 is a result too.
What is the difference between an AI pilot and production?
A pilot runs the system on a limited slice of real work to test it against a written success criterion. Production means it runs daily on all eligible work, with monitoring, an owned review queue, and a maintenance budget. A proof of concept answers one technical question. A pilot that can't reach production is a proof of concept rather than a failure, if its criterion defined what a no looks like.