After the wrong number has been used
A corrected dashboard does not update the staffing plan made from it. Following a data error into the decisions it reached, returning the choice to its owner, and making the cost of recovery visible.
Consider an illustrative incident. A weekly capacity report shows a support queue growing. The manager moves people from another queue to cover it. Later, an analyst finds that a join to case history counted some open cases more than once. The team repairs the model, checks the corrected totals against the source, and refreshes the dashboard.
The staffing plan is still in effect. Nobody has taken the corrected backlog back to the manager who moved the people.
That gap is easy to leave out of an incident process. The engineering work has a visible finish: the repair is deployed, the affected history is rebuilt, the checks pass. The decision made from the old result lives elsewhere, in a schedule, a meeting note, or somebody’s understanding of what needs attention. Refreshing the dashboard reaches none of those automatically.
In Earning trust in the numbers, I describe the reliability work around critical metrics: ownership, detection, visible limitations, and a response people can inspect. Here is where I would extend that response. Once a wrong number has been used, someone needs to follow the correction into the decision it informed.
For the organization paying for the platform, that is part of recovering from the error. For the person whose analysis carried it into the meeting, it is part of being trusted with the next decision. Both need evidence that the correction reached the place where the number mattered.
Follow the number past the dashboard
Lineage can identify the models and reports downstream of a faulty source. Usage records can help identify who opened them. Neither tells you, by itself, what someone decided after reading.
Start with the affected metric, period, and population. Then identify the reports and regular decision forums that used them. A report owner may know that the weekly review led to a staffing change even when the platform has no record of the meeting. Someone who downloaded an extract may have carried it into a planning document the data team cannot see.
The question I would ask each owner is specific: what did you change, approve, or defer after using this report? Asking whether the error had any impact is too easy to answer with a glance at the corrected chart.
Keep a short record of the decisions found, their owners, and when a change would still be useful. Include what remains unknown. A report being opened is evidence of possible exposure. A manager confirming that it informed the staffing plan is evidence of use. Those deserve different entries.
This does not require finding every screenshot before anyone can act. Begin with the decisions that carry the largest consequences or are about to become difficult to reverse. Expand the search when evidence points to further use. Document the limits of that search instead of turning an incomplete usage log into a claim that nobody else was affected.
A dashboard can be corrected in one place. The decisions made from it have to be found one by one.
Send the warning while the answer is still being checked
Waiting for a complete root cause can leave the wrong figure in use through another planning cycle. Once there is credible evidence that a consequential number is unreliable, the people using it need to know what they can safely do while the team investigates.
For the illustrative queue incident, an initial message might read:
The open-case totals in Monday’s capacity report include duplicate cases. The size of the overstatement is still being checked. Please pause further staffing changes based on that report and flag any already agreed. The report owner will provide a checked replacement or a progress update before tomorrow’s scheduling meeting.
That message identifies the affected result, the uncertainty, the immediate request, and the next update. It gives the manager a decision to make while the repair continues. The report itself should carry the warning too, and any scheduled distribution of the affected result needs attention.
There may be a usable fallback: a previously validated extract, a narrower slice that has been checked, or an operational source with understood limits. Explain the limits alongside the fallback. If no reliable answer is available, say so. Supplying another plausible total to keep the meeting moving would repeat the failure.
Someone needs explicit responsibility for this communication. In a small incident, the person coordinating the repair may carry it. When the work expands, assign it separately so the engineer investigating the join is not also trying to reconstruct every conversation the report entered.
Google’s SRE incident-response guidance separates operational work from a communications role that updates stakeholders and handles inquiries. I would apply that division to a data incident and extend the follow-up to the decisions already made. That extension is my recommendation; the guidance does not establish that sending an update repairs a business decision.
Give the decision back to its owner
A corrected number can change the case for a decision. It does not give the data team authority to reverse it.
In the queue example, the manager might still keep the staffing change. Absences, the age of unresolved cases, or a coming product release could make it sensible even with a smaller backlog. Those facts need to enter the review. Automatically undoing the move would substitute one incomplete account for another.
The useful handoff contains the original result, the checked replacement, and what changes in the interpretation. If the original recommendation depended on a threshold the corrected number no longer crosses, say that plainly. If the recommendation still holds, explain which evidence supports it now.
Then ask the owner to decide whether to keep, change, or stop the action. Record the reason and any follow-up date. An acknowledgment that the corrected report was received is a different event from reviewing the staffing plan.
This is also why I want an analyst’s recommendation to name a decision and an owner when it is first made. That record gives a later correction somewhere to go. Without it, the team can find the report and still be unable to find the person who can act on the repair.
Some decisions will be too late to reverse. The review then concerns the remaining consequences: what can still be recovered, what needs to change next time, and who owns that work. Calling the data accurate again cannot recover the hours already spent on the wrong queue.
Keep the corrected history visible
A backfill changes what the warehouse says about the past. The record of what people knew when they acted still needs to be explainable.
Keep enough of the affected report’s context to show its period, filters, definition, and the version used in the decision. Mark the old output as superseded and link it to the correction. Use the access and retention rules that already apply; an incident is no reason to scatter copies of sensitive extracts into another workspace.
This matters when somebody returns to the decision months later. If the only surviving report contains the repaired figure, the original staffing choice can look inexplicable. The correction record should let a reader reconstruct both the evidence available then and what was learned afterward.
The notice also needs to travel through the channels that carried the error. A banner on the dashboard helps the next visitor. The people reading an exported deck need the correction attached to that deck or sent to its audience. A meeting owner may need to amend the decision record. Name who will do that, rather than assuming that a message in the data team’s channel will eventually reach them.
The same issue appears when a scheduled AI briefing supplies an interpretation. Repairing its source does not withdraw the paragraph already delivered. The owner needs to check whether the explanation and proposed action still hold, and send the correction to the original audience. A freshly generated answer can be correct while yesterday’s recommendation remains in use.
Record the repair and the decision review separately
I would record technical recovery and follow-up on affected decisions separately. The first can finish while the second still has work assigned.
Technical recovery needs evidence that the affected data has been repaired, checked, and made usable again. The prevention work may have its own owner and deadline. For each consequential decision identified, the follow-up record should show one of these states:
| State | Evidence to keep |
|---|---|
| Awaiting review | Named decision owner, next contact, and deadline before further action |
| Changed | What the owner changed after reviewing the correction, and who will carry it out |
| Kept | The owner’s reason for retaining the decision on the corrected evidence |
| No longer reversible | Consequences reviewed, with any remaining action assigned or explicitly declined |
These are working states for the incident record. They do not need to become a new platform or an approval chain for every report. A few lines in the existing tracker can be enough.
If the owner has not responded and the consequences are material, escalate through the person accountable for the decision. Keep the item visibly open or hand it over explicitly. Silence is weak evidence on which to declare the follow-up complete.
This also makes the operational reporting more honest. The team can report when reliable data became available again while showing that an affected staffing decision is still awaiting review. There is no need to keep engineers in incident mode until every meeting has happened. There is a need to make the remaining responsibility visible.
“Data restored” should tell you what is safe to use now. It should not hide what still needs a decision.
The correction takes longer than the backfill
Following decisions costs time that a pipeline repair alone would not. An analyst has to return to a meeting, a manager has to reconsider a plan, and some other work waits. That cost belongs in the incident record. Otherwise the platform appears cheaper to operate than the people using it experience.
It also creates exposure. A quiet correction is easier to make than a message to the people who acted on the original result. The team may have to explain how an error passed its checks while the recipient explains a choice made from it. Keeping the account focused on the evidence available at each point helps make that discussion useful. It does not make it comfortable.
The response must be proportionate. A display rounding error that could not change a decision does not warrant the same search as duplicated cases in a capacity report. Record why a wider review is unnecessary when that is the judgment. Spend the effort where the correction could change an action or explain a material consequence.
Nor should the number of reversed decisions become a success measure. A careful review may leave the original choice intact. The evidence to inspect is whether the owner received the correction, considered its effect, and recorded what follows. Reversal is one possible result.
For the support manager in the opening example, the useful finish is a staffing plan reconsidered with the corrected backlog and the other evidence that matters. The manager may move the people back or keep them where they are. Either way, the reason belongs beside the decision, where the wrong number was used.