Skip to content
Anil Thapa
Agentic & AI systems

An AI assistant for self-service, on a governed data layer

What the assistant was allowed to know

An assistant in the team chat tool, launched against a dozen governed metrics and a knowledge base of what they mean, with every failure fed back as context. The narrow start was the decision. The request queue fell by roughly forty percent. The curation work it created is permanent.

WarehouseCatalogRetrievalLLM agentChat platformEvaluation

What this is about

  • Give the agent metrics as tools and context as documents, never raw tables and a prompt
  • Launch narrower than feels reasonable, and let the unanswered questions set the roadmap
  • Every failure gets answered twice: once to the person, once to the knowledge base

One deployment, described by shape. Figures are rounded and measured as the outcome states; no source system, metric name or organization is identified.

Situation

By the time a language model could write a competent query, the platform underneath it had spent a year getting its house in order: named owners for data, shared metric definitions in code, lineage and a catalog, and observability on the pipelines. That work had been done for people. This case study records what happened when the thing reading it stopped being a person.

The pull was obvious. Questions arrived in the team chat tool, addressed to whoever on the data team had answered last time, and waited in a queue measured in days. A model connected to the warehouse could answer them in seconds. Several vendors were offering to do exactly that, and a demo takes an afternoon.

The demo works because the schema is small and the question is one the builder knows how to answer. The warehouse was neither. It carried a decade of history, tables that had been superseded and never dropped, and several defensible definitions of revenue depending on refunds and what partners were owed. A person querying that warehouse pauses at the ambiguity, some of the time. A model does not pause. It picks one and speaks fluently.

So the decision was never whether to build an assistant. It was what the assistant would be allowed to know, who would be allowed to ask it, and how the team would find out it was wrong before an executive did.

Constraint

  • A channel is an audience. An answer posted in a shared channel reaches everyone in it, including people who could not query the source themselves. Access had to hold at delivery, not only at retrieval.
  • No verified answers existed. Nobody had a set of questions with known-correct results to test against. The first evaluation set would have to be built before anything could be measured.
  • No headcount. The team of ten-plus was already running the platform. The build had to fit in the margins of one engineer’s week, and the running of it inside analyst time that was supposed to be freed, not consumed.
  • One plausible wrong number would end it. Trust in a data system falls faster than it recovers, and the first audience included people whose numbers appear in leadership reviews.
  • Definitions were governed in the center and thin at the edges. The metrics the company was run on had owners and code. The long tail did not, and the assistant would be asked about the long tail on day two.

The first and the last of these shaped the design more than the model did. Everything else was engineering.

Decision

Make the platform, not the prompt, responsible for what the assistant knows. A prompt is a guideline, and a guideline is followed by whoever read it. A definition that lives in the platform and reaches the agent as a tool is followed by default.

Metrics as tools, context as documents

The assistant was given no raw tables. It was given three things.

A metric lookup that returns, for a governed metric, its formula, grain, scope, exclusions and freshness, from the same definitions the dashboards used. A table lookup that returns a business description, the key columns, the join keys and the known caveats. And a read-only query path with a row cap and a timeout, for the questions that still needed SQL after the first two had constrained it.

Around those sat the knowledge base: a few hundred pages at launch, retrieved by meaning rather than by keyword. Metric definitions in plain language. Table notes. The caveats that had lived in analysts’ heads, such as which region reports a week late and which measure finance takes to the board. And, from the second week, the questions already answered, so that a question asked twice was answered the same way.

The hard part was not retrieval. It was writing down what analysts had been quietly working around for years, and finding out in the process how much of the platform’s apparent consistency had been people.

Launch narrower than feels reasonable

The first version answered questions about twelve metrics in two decision areas, and declined the rest by name. It ran in one channel whose members all held the same data access, so the audience problem was solved by the guest list before it was solved by the system. About a dozen people from three departments were invited, and the invitation was explicit: break it. Ask it what you already know. Ask it the same thing two ways. Ask it something it should refuse. Say when it is right for the wrong reason.

The narrow phase ran six weeks. Widening was tied to evidence rather than to a date: the evaluation results below, and a failure log with nothing open that could mislead a decision.

The feedback loop is the product

Every answer was logged with the question, the query it ran, the sources it cited and the reaction it got in the thread. Every flagged answer, and a sample of the unflagged ones, was reviewed weekly by an analyst who knew the domain. Failures were sorted into a short taxonomy: period and date handling, definition ambiguity, a wrong or superseded table, an ambiguous question answered instead of clarified, and an out-of-scope question answered instead of declined.

Each failure was answered twice. Once to the person, in the thread. Once to the knowledge base, as a new page, a corrected definition or a scope note, so the same failure could not recur in the same form. The verified question-and-answer pairs became the evaluation set: sixty at launch, built by hand, growing to about two hundred and fifty within two quarters, a third of them from real failures. The set was re-run after every change to the definitions, the retrieval or the model.

Stress tests before widening

Before the second channel opened, two structured sessions tried to break it on purpose: requests for data outside the askers’ access, questions designed to produce contradictory answers, instructions pasted into a question to see whether the assistant would follow them, and requests for a citation to see whether it would invent one.

Three controls came out of those sessions. Access is enforced in the query path and checked against the channel at delivery, never in the prompt. Citations are derived from the query that actually ran, not generated. And there is an explicit path for “not here”: a sensitive question in a public channel gets a pointer to a private one, not an answer.

Graduate the glue

The first version was built on a workflow-automation tool, because that was the fastest way to change it every week, and weekly change was the whole method. It was also glue: no tests, a thin runbook, and one person who understood it end to end.

Once the definitions and the context had settled, they moved out of the workflow and into tools on an MCP server, so that any agent can call them, including the ones people were already bringing to their desks. The assistant in the chat window became one client of that layer rather than the layer itself, which is where the work that mattered had been all along. What that layer makes possible once business users start commissioning their own reports is a separate piece.

The tradeoff I accepted

I traded coverage for trust, and paid for it in goodwill and in curation.

The narrow launch disappointed people who had seen the demo. About a third of the questions asked in the first month were outside the twelve metrics, and the assistant said so, politely, to people who had expected magic. Holding that line took explaining, more than once, that the narrowness was the reason the answers could be believed.

The data team’s work moved rather than shrank. Less answering, more curating: writing the pages, reviewing the failures, settling definitions the assistant had exposed. That is a standing half-day a week of a domain analyst’s time, and it appears on no roadmap as a deliverable. It is also the most effective governance forcing function I have encountered. Every inconsistency a human analyst had quietly worked around was now surfaced in a public channel, in front of the person who asked. Nothing about the mess was new. The audience was.

Building on glue first was the right call, and it had a bill. By the time the layer graduated into proper tools, the workflow had become load-bearing in a way its runbook did not deserve, and the graduation came later than it should have.

The model bill was the smallest cost of all. An answer costs cents. The people who make the answers trustworthy cost what people cost.

Outcome

Over the first quarter after widening, the routine analyst request queue fell by roughly forty percent against the previous quarter’s count, and a supported question was answered in under a minute instead of waiting a day or more. Analyst capacity moved to work that still needed investigation and judgment, and to the curation described above, so part of the reduction is work that moved rather than work that disappeared.

On the evaluation set, the share of supported questions answered correctly rose from roughly seventy percent at launch to above ninety after two quarters of the feedback loop. The failures that remained were mostly ambiguity the definitions had not yet settled, which is a governance backlog with an owner rather than a model problem. Out-of-scope and restricted questions were declined correctly throughout, which was the criterion that mattered most and the one the stress tests had been for.

What this does not establish: that any particular decision improved because the number arrived faster. For that I would record which decisions cited an answer, and what happened next.

For a new rollout I would watch the same things, in this order: whether restricted questions are declined, whether supported questions match verified answers, how many open failures could mislead a decision, and how much analyst time the curation is taking. Message counts measure curiosity, which is free.

What I’d do differently

Build the first sixty verified answers before launch, not from it. The set arrived in the first two weeks anyway. Having it on day one would have turned the first fortnight from discovery into measurement.

Settle period and time-zone rules in the definitions first. Date handling was the largest failure category in the first month, and every instance was the same failure wearing a different question. It belonged in the metric definitions, not in the knowledge base as a series of corrections.

Set the graduation criterion in advance. “When the definitions stop changing weekly, the layer moves into tools” is easy to say before the workflow becomes load-bearing and hard to schedule afterward.

Publish the failure log to the users, from the first week. The instinct is to fix the embarrassing answers privately. A log that starts ugly and visibly shortens builds more trust than a clean launch, for the same reason a variance report does.