Changing the metric without erasing the decision
A better definition can make the same history tell a different story. Who approves the change, when it applies, how to compare periods, and what to preserve about the decisions made under the old one.
A metric can become less useful without becoming wrong.
Consider an illustrative product measure. An activated workspace used to mean one that created a project within its first fourteen days. The product now depends on collaboration, and the proposed definition requires completing a shared project in that window. The new definition asks a more demanding question. It may be better suited to deciding where onboarding needs work.
It also produces a lower activation rate for the same workspaces. If the new calculation replaces the old one behind an unchanged chart, the next review can begin with a conversation about declining performance that the chart has manufactured.
No pipeline has failed. The code may do exactly what its author intended. The failure is allowing a change in meaning to arrive as though it were a change in the business.
In One number, two histories, I describe reconciling definitions across platforms. A definition change within one platform carries a related problem across time. People need a useful measure now, a defensible way to compare it with the past, and a record of what the earlier decisions were based on. Those needs require separate choices.
Name the decision the new definition will serve
Before changing the calculation, write down what the current definition fails to distinguish and which decision would improve if it did. In the workspace example, creating an empty project might be enough to assess initial setup. It tells less about whether people have completed useful work together.
That does not automatically make the old measure obsolete. If setup remains a decision worth tracking, give the two measures names that distinguish their uses. Replacing one broad label with another broad label can leave the same argument waiting for the next product change.
The person who maintains the model can explain what is feasible and what will break. The person accountable for the use of the metric needs to approve its meaning. Consult the affected consumers, then name who can settle a disagreement. Ownership that exists only in the catalog will not resolve an argument about which team now appears to be missing its target.
Also name the kind of change. If the agreed definition already excluded test workspaces and the query accidentally included them, that is an error to correct, and the decisions made from it need their own follow-up. If the agreed definition is being replaced because the business needs a different measure, explain that choice. Calling an error a definition update would conceal something the reader needs to know.
The UK Office for National Statistics makes this distinction in its revisions and correction policies: revisions incorporate improved methods or data unavailable at first publication, while corrections amend mistakes found afterward, and it flags both to readers. I would carry that distinction into an internal metric change. The approval and migration process here is my recommendation, rather than a requirement that policy places on a company.
Put the effective date beside the approval
A code deployment date leaves important questions unanswered. Does the new definition apply to workspaces created from that day onward? To the next reporting period? To every historical cohort the platform can reconstruct? Those choices produce different series even with identical calculation code.
I would record when the change was approved, which periods or cohorts it applies to, and when consumers will first receive it. They can be different dates. A team might approve the definition before the next quarter, publish a comparison while consumers prepare, and change the official scorecard when the new quarter begins.
Alongside those dates, keep the eligibility rules, the event that qualifies a workspace, the observation window, and the exclusions. Include the reason for the change and examples that sit on either side of the boundary. A workspace that created a project but never completed one is more useful in the review than a sentence saying the new measure reflects engagement better.
The examples should become part of validation. Schema checks can pass while the meaning of a column changes completely: the name and numeric type can remain identical. A release therefore needs a comparison of results and a review of the intended interpretation, as well as a successful build.
This is the responsibility behind owning a model. Shipping the calculation also ships an answer to a business question. Its owner needs a way to show which question that version answers.
Compare two periods under one definition
Before switching the scorecard, calculate both definitions over the same inputs for a useful overlap period. Hold the eligible population and observation window fixed so the difference can be attributed to the rule being changed. If source corrections are also required, explain their effect separately rather than hiding everything inside one adjustment.
Here is an illustrative comparison. Both cohorts have completed the full fourteen-day observation window, and each column uses the same eligible workspaces within a cohort. The percentages are invented for this example.
| Cohort | Created a project | Completed a shared project |
|---|---|---|
| Earlier cohort | 60% | 40% |
| Later cohort | 64% | 44% |
Under either definition, the later cohort is four percentage points higher. Joining the earlier cohort’s old result to the later cohort’s new result would instead show a sixteen-point fall. The individual values are correctly calculated; their combination tells a misleading story.
The overlap needs interpretation too. Identify which workspaces move out of the measure and whether the effect differs by product plan, acquisition route, or another relevant segment. An aggregate adjustment can conceal a definition change that affects one part of the business much more than another. There is rarely a sound basis for applying one conversion factor to every old result and calling the history comparable.
A line chart can turn a change in definition into a story about performance.
If the required events exist reliably in older data, publishing a restated series can make comparison easier. Give that series a clear definition and state how far back it has been reconstructed. It is history calculated using the new rule.
If shared-project completion was only recently instrumented, the older history may be impossible to reconstruct. A missing event record does not prove that the event never happened. Keep the earlier series under its old definition and show the break. A shorter comparable history is more useful than a continuous line built from an unsupported assumption.
Keep the view that the earlier decision used
A restated series answers how past activity looks under the current definition. A report preserved from the time answers what people saw when they made a decision. Both can be useful, and they need labels that keep them apart.
Suppose the earlier activation rate supported a decision to reduce onboarding work. A later review should be able to find the measure used, its definition, the period, the relevant filters, and the recommendation it informed. It can then consider whether that definition was appropriate and whether the newer evidence warrants changing the plan.
Keeping that record does not establish that the earlier choice was sound. Someone may have used a limited measure to support a claim it could not bear. Preserving the evidence makes that judgment possible; overwriting it makes the discussion depend on memory.
The retained record might be a versioned report with its definition reference and decision note. Where reproducibility is required, it also needs the appropriate historical inputs. Old SQL alone cannot recreate an old result if the source records have since changed. The depth of retention should fit the consequence of the decision and the rules already applying to the data.
Make the current route obvious. A reader asking about today’s activation should reach the current definition. A reader investigating the earlier decision should be able to reach the version used then. Preserving a record does not require keeping every obsolete calculation in the default report picker.
Recomputing last quarter can improve the comparison. It cannot change what was available when last quarter's decision was made.
Move the targets and consumers with the calculation
A target chosen against the old definition does not become meaningful under the new one merely because its label still matches.
In the example, keeping the same activation target after requiring a completed shared project could turn a measurement change into an unannounced increase in the expected performance. Replacing the target with an easier number would also be a decision that needs an owner. Neither should happen as a side effect of a model release.
Review the baseline, the intended improvement, and the date from which the revised target applies. Preserve the original target beside earlier reporting. A restated trend can support planning without silently rewriting how past commitments were assessed. If an earlier commitment itself is being revised, record that agreement explicitly.
Then follow the metric into its consumers. A dashboard may be only one of them. An alert threshold, a recurring deck, an exported analysis, and an assistant’s saved description can all retain the old interpretation after the warehouse starts returning the new value.
Lineage helps identify technical dependencies. The named owners of the decisions help identify uses outside the graph. For material consumers, agree who updates the calculation or reference, who updates the explanation, and how to verify that the intended version is being used. The governance essay describes why notification and agreement cover different parts of that work.
The change note should explain the reason, the effective date, the effect on comparison, and where to find the old basis. Put it where people read the metric. A release announcement in a technical channel is insufficient evidence that the scorecard’s readers understand why the number moved.
Give the old calculation an end to routine use
Running both versions temporarily is useful for explaining the difference and moving consumers. Running them indefinitely because nobody can decide which is current recreates the disagreement the governance work was meant to settle.
Agree the migration period and the conditions for retiring the old version from routine use. The important consumers need to have moved, the material differences need explanations, and any remaining use needs an explicit reason and owner. A genuine setup measure can remain as a separately named metric. An abandoned activation calculation should not remain simply because one export has no owner.
There is a real cost to doing this carefully. Someone has to build the comparison, investigate differences, revisit targets, and help consumers interpret the change. Keeping evidence of past reporting consumes attention and sometimes storage. A bounded overlap can delay a cleaner model or a requested report. Those costs should appear in the change plan.
Keeping the old definition forever has a cost as well. It can reward behavior that no longer represents the intended outcome and keep decisions attached to an increasingly weak measure. The purpose of the process is to make a necessary change possible with its consequences understood.
For the workspace example, the useful finish is a current measure whose purpose is clear, comparisons made on a stated basis, and earlier decisions that remain inspectable. The next review can then discuss whether onboarding produces useful collaboration. It need not spend its time reconstructing why the activation line suddenly fell.