"Agentic" is a workflow, not magic: AI agents for business, explained

What an AI agent actually is — a workflow with judgment calls and checkpoints — and the three questions that puncture any agent pitch.

// key takeaways
  • An AI agent is a workflow with judgment calls and checkpoints — not a new kind of employee.
  • Checkpoints belong wherever an action is irreversible, moves money, reaches a customer, or touches sensitive data.
  • Three questions puncture any agent pitch: show me the steps and where it stops, what happens when it's wrong, and who reviews what.
  • Autonomy is a dial you turn up one step at a time as evaluation earns it, not a feature you buy on day one.
  • Most "agents" are automations with a language-model step — design them, price them, and trust them like one.

An executive has heard the word agents eleven times this month, twice from vendors with polished demos. She asked both what their agent does when it's wrong. Both answered with another demo.

An AI agent for business is software that carries a task through several steps, makes judgment calls with a language model where a rule won't do, and stops at checkpoints for a person when the next action is irreversible, expensive, or customer-facing. It is a workflow with judgment in it — not a new kind of employee. That's good news: a workflow is something you can draw, test, and price.

The demo versus the job

A demo proves an agent can finish a task once, on clean input, with its builder watching. A job needs it to finish the task correctly every day, on messy input, with nobody watching — and to stop when it can't. Projects fail in that gap, and "fully autonomous" is how the gap gets sold.

Demos show autonomy: the model plans, calls tools, recovers from an error, and narrates. That's the easiest version of the problem — one run, curated inputs, no consequences. A job runs hundreds of times on inputs nobody curated, like the form with the address in the wrong field. What matters is how often it succeeds, what it does when it doesn't, and whether anyone finds out.

Multi-step work also compounds its errors. String together ten steps that are each right nineteen times in twenty, and the whole run comes out right only about three times in five. That's not a statistic about anyone's product; it's multiplication. It's why the agents we trust in production are short, fenced, and interrupted on purpose.

So "fully autonomous" is a warning sign for most businesses, especially a small business with nobody free to watch a dashboard. The actions worth handing to software — quotes, customer messages, orders — are exactly where an unreviewed mistake costs money or trust. A pitch that leads with autonomy usually means the stopping points haven't been designed. Autonomy is a dial, not a feature: you turn it up one step at a time, as evaluation shows that step has earned it.

What an agent actually is

An agent is a workflow that makes some of its own decisions: steps someone designed, judgment calls where a language model reads the situation and chooses what happens next, and checkpoints where it stops for a person or a rule. As a formula: agent = workflow + judgment calls + checkpoints.

That's all an agentic workflow is: a workflow in which some steps are decided rather than scripted.

The workflow is the backbone: the steps, their order, the tools the agent may touch — the CRM, the calendar — and the states a case moves through. It's mostly deterministic, and that's what makes it auditable.

The judgment calls are where no rule reaches: what this email is asking for, whether this exception is routine. In our builds, each one returns a structured answer from a fixed menu and is validated before anything acts on it. A model that can return anything can do anything.

The checkpoints are where it stops: for a person to approve, for a validation rule to pass, or because the model isn't confident enough to continue. An agent without checkpoints isn't more advanced. It's just unsupervised.

The formula also settles two common comparisons. AI agent vs automation: the difference is the judgment calls. An automation follows the same path every time a trigger fires; an agent reads messy input and chooses between paths. AI agent vs chatbot: a chatbot talks, an agent acts. A chatbot's output is a message to a person; an agent's output is an action in one of your systems, or a proposal to take one.

ChatbotAutomationAgent
InputA person's messageA trigger: form, schedule, webhookA task or event, often messy
Decides?What to say nextNo — it follows fixed rulesWhich step comes next, within limits
Acts?No — it answersYes, the same way every timeYes, through the tools it's given
Stops?When the person doesWhen the steps run outAt checkpoints, or when unsure
Where it failsConfident wrong answersInputs nobody planned forSmall errors compounding across steps

Where the checkpoints belong

Checkpoints belong wherever the next action is irreversible, moves money, reaches someone outside the business, or touches compliance-sensitive data. Before those steps, the agent proposes and a person decides. Everywhere else — reading, sorting, looking things up, drafting — let it run, and log everything it does.

  • Irreversible actions. Sent emails, deleted records, submitted filings. If it can't be undone, a person sees it first.
  • Money. Quotes, discounts, credits, orders — anything with a price attached.
  • External communication. Anything a customer reads with your name on it. Routine answers can graduate once evaluation supports it; promises about prices or dates keep their person.
  • Compliance-sensitive data. Health information, client files, personal identifiers.

Human-in-the-loop means a person approves, edits, or rejects the agent's proposed action before it takes effect. That differs from a person on the loop, who watches the log, audits a sample, and can step in, but doesn't approve each action. Both are legitimate: in-the-loop for the four categories above, on-the-loop for everything around them.

A checkpoint also has to be cheap for the human: proposal, inputs, and reason on one screen, one click to approve, edit, or reject, and every decision recorded as evaluation data. If approving takes as long as the work, people rubber-stamp, and the checkpoint becomes theater.

Field note: For a residential roofing company, leads sat unanswered during storm season while every crew was on a roof. The system we built answers routine questions, collects job details, and drafts estimates for the owner's approval; follow-up continues until the job is scheduled or closed. The checkpoint sits on the price, not on every message. Response time went from "when someone got to it" to minutes.

Field note: A regional e-commerce retailer handled delivery exceptions — a wrong address, a weather hold — by phone and spreadsheet. We built a deterministic backbone of queues, retries, and an audit trail. A model triages exceptions and drafts customer notifications, each sent by rule or approved by a person, and cases that need judgment get a clean handoff to someone who has it. Exceptions now get handled before the truck leaves instead of after.

Data gets a checkpoint too, and it comes first: decide what the agent may see before deciding what it may do. Depending on the data, our builds mask personal identifiers before any external model call, use zero-data-retention terms on model calls, keep processing inside the client's own cloud tenancy, and log every action against the input that caused it.

The patterns behind the personas

Behind the persona names — vendors' and ours — agents are built from a handful of well-understood reasoning patterns: plan-then-execute, explore-and-act loops, self-critique, and explicit human handoff. We compose them like parts, choosing per step. Our own AI department runs on the same patterns, which is the most honest demo we can offer.

  • Plan-execute. The model writes a plan; an executor carries it out, checking each result against it. It suits long jobs with a known shape, and a readable plan can double as a checkpoint.
  • Explore-and-act (ReAct-style). Reason, take one action — a search, a lookup, an API call — read the result, choose the next move. It suits paths nobody can map in advance, and it needs a step limit and a tool list, or it wanders.
  • Self-critique (reflexion). Draft, check the draft against explicit criteria — cites its sources, follows the policy, answers the question asked — then revise. It doesn't replace the person; it makes their review faster.
  • Human-in-the-loop. A designed handoff: what the person sees, what they can change, and how the agent resumes afterward.

Most production agents are compositions: plan-execute on the outside, explore-and-act inside the steps that need it, a critique pass before anything leaves, and a human handoff at every checkpoint. That's how the formula's judgment calls and checkpoints get built.

Our AI department works the same way. The persona names are a briefing convenience; each role underneath is one of these patterns on a model tier matched to the work, from frontier-class models for building and high-stakes reasoning to fast lightweight ones for screening and routing. One rule doesn't bend: every AI role reports to a human. Judgment is shared; control isn't.

How to evaluate an agent before you buy one

Ask what we call the three questions, and listen for boring, specific answers: show me the steps and where it can stop; what happens when it's wrong; who reviews what, and how often. A real system answers all three with diagrams, logs, and names. A demo answers with a better demo.

Think of them as AI agent evaluation from the buyer's side of the table. They work on any pitch, ours included.

1. Show me the steps — and where it can stop. A good answer is a one-page list: each step, marked rule or judgment, the tools it can touch, and every point where it halts. A worrying answer is "it figures out the steps itself." If that's partly true, ask to see the fence: which tools it may call, how many steps it may take, what it may never do.

2. What happens when it's wrong? A good answer names failure modes and the response to each: a misread input lands in a review queue, a failed tool call retries and then escalates, a step that dies halfway leaves no half-made order behind. "It's very accurate" is not a failure plan. An agent also inherits every reliability problem of the automation around it, which is why it needs the four properties of automations that don't break before it needs a smarter model.

3. Who reviews what, and how often? A good answer names a role at each checkpoint, a schedule for auditing a sample of the actions that ran unreviewed, and an evaluation set re-run before every prompt or model change. "The AI checks itself" is half an answer: self-critique helps, but it isn't oversight.

Then ask for four artifacts by name: the step list, the failure handling, the review cadence, and the evaluation set — real past cases with known right answers, scored before anything changes. Cheap builds skip the evaluation set. It's the first line item of the invisible 80%, the part of an AI project no demo shows, and an agent needs more of it than an assistant does, because its mistakes are actions, not sentences.

Rule of thumb: If a vendor can't draw the steps on one page, including where it stops, you're looking at a demo. If they can, ask to see a failed run. Successful runs all look alike.

Most "agents" are automations with an LLM step — and that's fine

Most systems sold as agents are automations with one or two language-model steps: a fixed path with a judgment call inside. That isn't a con. It's usually the right design. Price it, build it, and trust it like an automation, and be wary of anyone pricing it like an employee.

A fixed path with a model reading the email and filling in the fields is easier to test, cheaper to run, and fails in fewer ways than an open loop that picks its own tools. Much of what gets marketed as agentic automation has this shape, and so does most of what we ship in our AI and LLM engineering work: a deterministic backbone, a few well-fenced judgment calls, checkpoints where the stakes are.

A true agent, one that picks its own path, earns its place when the path varies from case to case, the steps can't be listed in advance, and every action it can take is reversible or behind a checkpoint. Investigating a stalled order across three systems qualifies. Answering the same four questions from a policy manual doesn't. And if the task has no judgment call at all, you may not need a model, let alone an agent — see when plain software beats AI.

Pricing follows the design: you're paying for a workflow build, an evaluation set, and review time, not a salary. If a pitch compares the agent's price to a headcount, ask what the agent does on the day it's wrong. The employee has an answer to that.

When not to do this: Don't let an agent act unreviewed where a mistake can't be undone and nobody would catch it before a customer does: sending quotes or invoices, changing records other systems depend on, speaking for you on anything contractual. Don't buy one if nobody will own the review queue. And don't buy one — from anyone, us included — if the vendor can't answer the three questions in plain language.

Run it yourself

Pick one task someone wants an agent for and draw its steps on one page. Mark each step rule, judgment, or human. The marks tell you whether you're looking at an automation, an agent worth scoping, or a risk — and they hand you the exact questions to put to any vendor.

  1. Pick a real task — the one a vendor pitched, or the one your team complains about most. One task, not a department.
  2. List every step in order, including the ones people do without noticing: checking the calendar, looking up the customer's history.
  3. Mark each step. Rule if it can be written as if/then, judgment if it needs reading or deciding, human if it must be a person.
  4. Read the marks. No judgment steps: it's an automation. Judgment steps with a human step before every irreversible, money, or customer-facing action: an agent worth scoping. Judgment steps, no human step, and money moves: a risk, not a system.
  5. Collect twenty or thirty real past cases with the right outcome for each. That's the seed of an evaluation set.
  6. Circle the step you'd least like to get wrong. Your first checkpoint goes there, and it's your first question for any vendor.

Agents are real and useful, and less mysterious than the pitch decks suggest. Under the persona and the narration sits the engineering it always was: steps, decisions, stopping points, and a person accountable for each. The builders worth hiring are proud of their step lists. Ask to see one.

Questions we get asked

What is an AI agent, in plain terms?

An AI agent is software that carries a task through several steps, uses a language model to make judgment calls where no rule fits, and stops for a person before actions that are irreversible, expensive, or customer-facing. In practice it is a workflow with judgment in it: designed steps, a few decisions made by a model, and checkpoints. It is not a new kind of employee.

What is the difference between an AI agent and a chatbot?

A chatbot talks; an agent acts. A chatbot answers a person's message, and the person decides what happens next. An agent takes a task or an event, decides which step comes next, and carries it out through the tools it has been given, such as updating a record or drafting a quote, stopping at checkpoints for a person wherever the stakes are high.

What is the difference between an AI agent and automation?

The difference is judgment calls. An automation follows the same fixed path every time a trigger fires. An agent reads messy input and chooses between paths: which tool to call, what an email is asking for, whether an exception is routine. Remove the judgment calls and an agent becomes an automation. Most systems sold as agents are automations with one or two model steps, which is often the right design.

What does human-in-the-loop mean for AI agents?

Human-in-the-loop means a person approves, edits, or rejects an agent's proposed action before it takes effect. It belongs before anything irreversible, anything that moves money, anything a customer will read, and anything touching sensitive data. It differs from human-on-the-loop, where a person monitors the agent's log, audits samples, and can intervene, but does not approve each action in advance.

How do you evaluate an AI agent before buying?

Ask three questions: show me the steps and where it can stop; what happens when it's wrong; who reviews what, and how often. Then ask for four artifacts: the step list with its checkpoints, the failure handling, the review cadence, and the evaluation set, meaning real cases with known right answers that are re-run before every change. A vendor without concrete answers is selling a demo.

Are AI agents safe for a small business to use?

They can be, if the agent is built as a workflow with checkpoints. Let it read, sort, look things up, and draft on its own; require a person's approval before it sends anything that commits you to a customer, changes a price, or deletes records. Keep sensitive data out of what it can see unless a step needs it, log every action, and make sure someone owns the review queue.

// written by

Stefan Jovanovic · Partner · Frontier & Agents

Stefan lives on the frontier: testing new models the week they ship, and building multi-agent systems — swarms of specialized agents that carry real workflows end-to-end.

the people behind Tensorika →

// talk to an engineer

Not sure AI is the answer? Ask us — we'll tell you if it isn't.

A short email describing the workflow is enough. We reply with a verdict, not a pitch: where AI helps, where plain software wins, where neither is worth it.