Skip to content
Anil Thapa
Agentic & AI systems

When the assistant costs more than the queue

A launched assistant has a sponsor. A retired one has nobody, so retirement happens by neglect or not at all. Four numbers that say whether to widen, narrow or switch it off, the meeting that reads them, and what a good retirement keeps.

7 min read

Every assistant I have seen launched had a sponsor, a date and a demonstration. I have not seen one retired that way. Retirement, when it happens, happens by neglect: the usage drops, the person who curated the knowledge base moves on, the answers get worse, and one day someone asks whether the thing in the channel still works. Nobody decided. Nobody was asked to.

That is a gap in how teams run these systems, and it is the opposite of how they are launched. The launch decision gets a scope, a reviewer and a failure that would pause it. The decision to continue, to narrow, or to stop needs the same, and it needs to be scheduled, because nothing in the normal course of operating an assistant will raise it.

The queue it replaced

The assistant exists to replace something: usually a queue of routine questions that waited on an analyst. That queue had a cost, measured in analyst hours and in the days a decision waited for a number. The assistant is worth running while its full cost is below that, and its full cost is larger than the bill.

The model bill is the smallest line. The larger ones are people. In the deployment I wrote up in What the assistant was allowed to know, the curation ran to a standing half-day a week of a domain analyst’s time: reviewing flagged answers, writing the pages, settling the definitions the assistant exposed. Add the review of sampled answers, the corrections users make after accepting a wrong one, and the hardest line to see, the decisions a wrong answer reached. The last one does not appear on any invoice, and it is the cost that ends assistants.

So the comparison is the queue’s cost against all of that, and it is a comparison that changes month by month, which is why it needs a meeting.

Four numbers, one meeting

I would put four numbers in front of the sponsor, the reviewer and the analyst who carries the curation, once a month, and ask one question: widen, hold, narrow, or stop. Each number has a direction that argues for each answer.

Number Where it comes from What it argues
Routine requests replaced, net of corrections The request queue against its baseline, less the time users and analysts spend checking and correcting answers Rising net replacement says hold or widen. Replacement that is eaten by correction says the scope is wrong
Open failures that could mislead a decision The failure log, counting only the failures in supported scope that reached or could reach a consequential decision Any open one says hold. A pattern of them in one area says narrow that area until its definitions are settled
Curation hours against their budget The analyst’s time on pages, reviews and definitions, against what was agreed at launch Steady or falling says the knowledge base is maturing. Rising for two quarters says the assistant is generating governance work faster than the team can absorb
Questions outside scope, by subject The declined questions, grouped This is the only number that argues for widening, and only where the subject’s definitions have an owner

The first number is the one most teams track, and it is misleading alone: “requests replaced” is flattering until the corrections are subtracted. The case study lists what to watch; this is the meeting that reads it. The second number is the one that should be able to stop the assistant on its own, and it is the one I would read first. The third is the one that creeps.

RAND researchers interviewed 65 experienced data scientists and engineers about why AI projects fail, in a 2024 report that notes estimates of failure above 80 percent, and identified five root causes: misunderstanding or miscommunicating what problem the AI is meant to solve, lacking the data to train a model, choosing the newest technology over the problem in front of the users, lacking the infrastructure to manage data and deploy a model, and applying AI to a problem too hard for it. Those are interview findings rather than a measured failure rate, and they concern AI projects broadly rather than analytics assistants. My reading is that three of the five are decisions about what to attempt, made or skipped before anything was built, and that a monthly review with four numbers is the cheapest way to keep making them after launch.

Narrowing is a result, not a retreat

The option teams skip is the middle one. An assistant that answers questions about twelve metrics well and about forty badly is not a failed assistant. It is an assistant with the wrong scope, and the fix is to decline the forty by name until their definitions have owners, which is the same move that made the narrow launch work in the first place. An analyst in the chat window argued for starting smaller than feels reasonable. The monthly review is where the scope stays honest afterward, in both directions.

Narrowing has a cost in goodwill with the people who had the forty, and the line to hold is the one from the launch: the narrowness is why the answers can be believed. A scope that only ever widens is a scope nobody is reviewing.

What a good retirement keeps

If the numbers say stop, most of what the assistant built is not the assistant.

The definitions it forced into the open stay in the platform. The knowledge base of context, the caveats that lived in analysts’ heads, the pages written for every failure, stay readable by the next system and by people. The evaluation set, the verified questions and answers, is the most expensive artifact in the whole deployment and the one most worth keeping, because it is the acceptance test for whatever comes next. And the habit the assistant forced, that a number has an owner and a definition before it is answered, outlives the channel it ran in.

That is the case for building the assistant on the platform rather than on the prompt, made from the other end. The case study’s rule is that a definition reaching the agent as a tool is followed by default, and The dashboard stopped one step short argues that every briefing built on such a tool then computes it the same way. Retirement is where that pays: an assistant built on governed tools can be switched off without switching off the governance, and one built on a prompt and a knowledge base nobody else reads takes its only asset with it.

What a bad retirement looks like is the other thing: the channel goes quiet, the pages stop being written, the definitions drift back into heads, and the next assistant starts from a demo.

What it costs

The sponsor’s standing. A retirement is a public reversal of a launch someone announced. The monthly review makes it smaller, because the decision arrives as the fourth reading of the same numbers rather than as a verdict, and because narrowing was on the table every month before stopping was.

The users who liked it. An assistant that answered a third of their questions quickly is missed, and the queue it replaced comes back. The review has to say what returns and who carries it.

The team’s own pride. The people who built and curated it will read the numbers as a judgment of their work. The honest reading is that the queue was cheaper than the curation for this scope at this time, which is a fact about the scope, and the evaluation set they built is the proof of what they learned.

Stopping something that was about to work. The risk of any review with a trigger. The guard is a horizon: a narrowing or a stop is proposed with the condition that would reverse it, which is the same shape every recommendation on this site takes.

The meeting itself. An hour a month for three people, and the discipline to hold it when nothing seems wrong, which is exactly when the third number is creeping.

The queue the assistant replaced was a cost until it changed a decision. So is the assistant. The difference is that the queue was never going to be mistaken for progress.