When the agent writes the pipeline, review is the platform
For twenty years a data team's constraint was writing. When a model drafts the pipeline, the model and the query, the constraint moves to review: what a reviewer checks when the author is a model, which standards have to run rather than be remembered, and what the team stops practicing.
For as long as I have run data teams, the queue has had the same shape. A request arrives, someone with the skill and the time writes the change, and the writing is the slow part. Review was the faster half: one engineer reading another’s work, usually in less time than it took to write, catching what the author’s own tests had not.
That queue has inverted. A model drafts the pipeline, the dbt model, the query, the test, and the migration, and the draft arrives as a pull request before anyone has decided whether the change should exist. Writing is now the fast part. The constraint is review, and most of what a team knows about review was built for a world where the author could explain the change.
I have watched this happen on a team I ran, and it did not announce itself. The first months looked like a productivity gain. The queue of open changes grew, the batches got bigger, and the senior people spent their days reading.
The first draft is free. The review is the platform.
What changed in the queue
Three things, and the third is the one that matters.
Volume. A change that took a day to write takes an hour to generate, so more changes get proposed. What the platform costs already notes the version of this that reaches the warehouse bill: more queries proposed, and whether the bill rises depends on which ship. The same logic applies to the review queue, which has no auto-suspend.
Batch size. A model asked to add a column will add the column, update the three models downstream, write tests for all four and adjust the documentation, in one change. Each piece is reasonable. The batch is larger than a person would have made, and larger batches are riskier, which is a finding that predates AI and holds under it.
DORA’s 2024 Accelerate State of DevOps report, from a survey of more than 39,000 professionals, associated a 25 percent increase in AI adoption with a 1.5 percent decrease in delivery throughput and a 7.2 percent decrease in delivery stability, while most respondents reported feeling more productive. The report’s own hypothesis for the mechanism is larger changes: it says changelists are likely growing in size, and that larger changes are slower and less stable. It is a survey correlation across many organizations, not an experiment, so it does not say what any one team will see. My reading is that the gain shows up where the batches are kept small and the review keeps pace, and the loss shows up where neither is managed.
Authorless changes. A pull request used to come with a person who had thought about it. Now it may come with a person who prompted it, read the summary, and submitted it. When the reviewer asks why the join is left rather than inner, the answer is sometimes “I’ll ask it.” That is the change that breaks the old review, because the old review was a conversation between two people who both understood the work.
What the reviewer checks when the author is a model
The checks move. Syntax, style and the obvious bug are the model’s strengths, and a reviewer who spends the review on them is reviewing the wrong thing. What is left is what the model cannot know and what it is worst at admitting it does not know.
| When a person wrote it | When a model drafted it |
|---|---|
| Does the code do what the author intended? | Does the author know what they intended? Can the person who submitted it explain the change without the tool? |
| Is the logic right? | Is the definition right? A model will compute revenue three plausible ways and pick one without saying which |
| Are the tests passing? | Did the model write the tests to pass? A test generated beside the code tends to assert what the code does, not what the business means |
| Is the change scoped? | What else did it touch? The downstream models, the documentation and the three small fixes nobody asked for |
| What does this cost to run? | The model does not see the bill. A full refresh of a large table reads as the simplest correct answer |
| Will the author maintain it? | Who owns this now? The person who prompted it, or nobody |
The last row is the one I would start with. A change with no owner is a change that will be re-generated the next time it breaks, by someone who understands it less than the last person did.
Standards have to run, not be remembered
A review culture used to live in people. Senior engineers knew the naming conventions, the layering rules, which models were contracts and which were scratch, and they enforced them by reading. That worked at the volume a person could write. It does not work at the volume a model can produce.
A standard that lives in a reviewer's head does not survive the volume a model can produce.
So the standards move into things that run. Tests on every model, with the business examples that a key constraint cannot express. CI that builds what changed and what depends on it, and blocks the merge. Linters for naming and layering. Contracts on the shared models, so a breaking change fails before a person has to notice it. None of that is new; it is the guardrail set that made giving the models away survivable, and the argument for it is now stronger because the author may not be a person.
What cannot be automated is the review of meaning and consequence: whether the definition is the one the business uses, whether the grain matches the decision, whether the change is worth its cost. That review is scarce, and it is the one to spend the senior people on.
Plan the review like the work
If review is the constraint, it is planned like one. The shape is the one in The review queue has one person’s name on it: estimate the recurring load from actual changes, put it on the plan, and decide what moves to make room. One rule that piece already had, now with a reason it cannot be skipped, and one addition.
Route by consequence, not by size. A model-drafted change is often large and mostly mechanical. The review goes to the part that touches a shared definition or a contract, and the rest gets the automated checks and a lighter read. The routing rule has to be written down, because a reviewer faced with a 400-line change will either read all of it or none of it.
Make the submitter the first reviewer. Before a change enters the queue, the person who prompted it writes, in their own words, what it changes, why, and what they checked. If they cannot, the change is not ready, whatever the tests say. This is the cheapest control available and it also keeps the skill alive, which is the next problem.
What the team stops practicing
The thing nobody plans for is atrophy. A team whose first drafts are all generated stops practicing the first draft, and the first draft is where intuition about the data was formed: the join that doubled rows, the timestamp in the wrong zone, the table that looked right and was deprecated. A reviewer who never writes loses, over a year or two, the instinct that made them a good reviewer.
METR ran a randomized trial in early 2025, published that July, in which sixteen experienced open-source developers completed 246 tasks on their own repositories with and without AI tools. With the tools they took 19 percent longer. Before the study they had expected to be 24 percent faster, and afterward they still believed the tools had sped them up by about 20 percent. It is a small study on one kind of work with the tools of early 2025, so it does not say what a data team will see with later tools on pipeline work. What transfers is the gap between felt speed and measured speed, which is the gap a team has to measure for itself rather than report from how the work feels.
What I would do about it is unglamorous: keep some writing unassisted on purpose, rotate people through it, and treat debugging a generated change as a skill to develop rather than a chore to delegate back to the tool. DORA’s 2025 report describes AI as an amplifier of what a team already does well or badly, and the thing a data team most needs to keep doing well is understanding its own data.
What it costs
Review becomes the bottleneck, and the reviewers are the senior people. The capacity that used to write the hard changes now reads the easy ones. That is the right use of it only if the easy ones are routed away first.
The platform bill moves. More proposed queries, bigger rebuilds, more materializations that looked simplest to the model. The cost review in What the platform costs has to apply to generated changes with the same rigor as to human ones, which means someone has to look.
Ownership gets thin. The person who prompted a change owns it less than the person who wrote it would have. Who owns the model? argued that ownership became authorship once systems started answering. The same argument now applies to the changes themselves: whoever submits a generated change is its author, with the responsibility that carries, or the change has no author at all.
The feeling of speed is not evidence of it. The team will feel faster. The measurement may disagree, and the only way to know is to measure throughput and stability the way they were measured before, rather than counting changes proposed.
What to measure
Not lines written or changes generated. Those go up by construction.
Time from a change being proposed to a reviewed change in production, and whether that time is spent waiting for a reviewer. Change failure rate and the time to repair, which is where larger batches show up first. Review load per senior engineer, in hours, against the plan. And the measure the rest of this site keeps asking for: of the changes that shipped, which ones changed a decision, and which were motion.