Plans are hypotheses about how to build something, so we should analyze and draw conclusions from them.
If you read the planning loop post, you know the bulk of my time spent on a feature is in the planning phase. By the time the agent starts building, the plan has been poked at from multiple angles. And it’s still wrong somewhere. Not wildly wrong. A name changes. A library doesn’t behave the way the plan said it would. A reviewer spots a gap nobody saw.
I love data, and I love any attempt at improvement, so of course I wanted some measurements. I already have a retro skill I can run at the end of a planning session, but it’s really aimed at the session itself. What I wanted to know was different. Based on the plan we created, did the code come out according to it? How close did we hit the mark? Where we deviated, why? Then I wanted to take the feedback left on my PR and fold that in too. If we made fixes during review, why didn’t we anticipate them in the first place? What went off the rails, and what missed the mark entirely?
That’s the last step of the scientific method, really. You analyze the data to see whether it supports or refutes the hypothesis. Supported, and it feeds into the bigger theory. Refuted, and you reject it or revise it to fit what you found.
Doing that for every piece of work means that over time I collect patterns, and then I can amend my systems to account for them. That’s the loop. Plan, build, review, write down the gap, learn from the pile, plan better.
Where it plugs in
Two skills, internal-reconciliation and external-reconciliation, slot into steps I was already running. Neither one blocks anything. Each one records what it found and hands off.
Your workflow isn’t mine, and it doesn’t need to be. What matters is the timing. One runs before anybody else has seen the work, and the other runs after review is over…
Plan
Build and open the PR
internal-reconciliationcompares the plan with the branch diff, before the PR description gets written.
Review and merge
external-reconciliationruns once feedback is addressed and the PR has merged. It compares the plan with what shipped, and uses the review to work out why.
LearnEvery piece of work leaves a record. Over time, the records get mined.
mining the records
- Gather. Pull every reconciliation record across many plans.
- Find the patterns. Look for the misses that keep repeating, and where they cluster.
- Amend the system. Change the skill, rule, or check that could have caught it. Those changes shape the next plan.
Two skills, one comparison
Both skills run the same comparison underneath. They differ in when they run, what they compare against, and whether they can say why something moved.
internal-reconciliation runs right before the PR description gets written. It finds this branch’s plan, compares it with the branch diff, and records what diverged. What it can’t do is look at review. Nobody has seen the code yet.
external-reconciliation I run after the feedback has been addressed and the PR has merged. By then the whole story is in. It can see what feedback was given, whether we fixed it, and how that compares to the plan. So it can get at the question I really care about… how, and why, did we miss what came up in review?
It also records how it matched each piece of feedback to a divergence. A match on the exact file and name is strong. A match on “this comment is kind of near that line” is weak, and I want the weak ones to look weak when I read them later instead of sitting there with the same confidence as the strong ones.
Every divergence gets a class
If you want to count records across plans, they have to use the same words. So there’s a fixed vocabulary, and anything that doesn’t fit gets kept as unclassified. Never dropped. And if enough of those start to look alike, that bubbles up as a suggestion for a new class.
What diverged, which both skills record:
-
approach pivot: the plan’s mechanism got replaced with a different one. -
discovered constraint: the build hit something the plan didn’t know about. -
falsified by measurement: the plan asserted something, and running it proved otherwise. -
renamed after planning: a name the plan coined shipped under a different one. -
named but never built: a file the plan said it would create doesn’t exist.
Why it diverged, which only external can record:
-
reviewer caught what the plan missed: a review thread found a real gap. -
reviewer overturned a considered decision: the plan thought about this and chose otherwise, and review went the other way. -
internal cause: the change came from the author, even if it happened to land during review. -
predates review: the change was already on the branch before the first review comment, so no thread could have caused it.
Built to report honestly
A tool like this is only useful if you can believe it when it says nothing happened. So most of the design is about not lying, in either direction.
- Absence isn’t divergence. A plan step with no code yet isn’t wrong. It might be the next PR. Only a contradiction counts.
- Not run isn’t clean. If a check couldn’t run, say a fetch failed, the record says “not run,” never “no divergence.”
- It never blocks. It records what it found, and the PR flow keeps going.
- It’s safe to re-run. Entries update in place. Fixed ones get marked resolved, not deleted.
“Nothing went wrong” and “nothing got checked” look identical in a summary, so make sure they look different.
Reading the pile
One record says almost nothing. The value shows up when you read a lot of them together, and the questions you ask will depend on your own workflow. Here are the kinds I find useful…
- What shows up before review, and what shows up after? That’s the difference between what the author finds while building and what a fresh reader finds.
- What do reviewers keep catching? If the reviewer-caught entries cluster around a theme, that’s something planning isn’t looking for.
- What does the first pass miss? Entries that were already true before review show where the early check has a blind spot.
- Does the same area keep churning? One file or one piece of copy diverging across several PRs usually means a decision nobody has actually settled.
-
Does the vocabulary fit? Lots of
unclassifiedentries means the classes are wrong, not the entries.
Whatever you find, aim the fix at the step that could have prevented the gap, not the PR that happened to show it. If plans keep asserting how a library behaves and getting disproved, check those claims against the docs at plan time. If reviewers keep catching the same kind of gap, move that check into plan review, before any code is written. If the same decision keeps flipping across PRs, settle it once and write it down where the next plan will find it.
Balancing the book
The word reconciliation comes from bookkeeping. You sit down with two records of the same money, your register and the bank statement, and go line by line. The point isn’t to decide which one is right. Sometimes the bank charged a fee you never wrote down. Sometimes the mistake was yours. Sometimes a check just hasn’t cleared yet. Every difference gets explained.
That’s all this is. The plan and the code are two records of the same feature, and the differences are where the learning is. If you already write plans before you build, you’re most of the way there. Keep the plan after it ships and explain where the two disagree. That’s your first reconciliation.