Product

I Decided I Wouldn't Be the One to Pass My Own Work

A short note on breaking the habit of letting the same agent that wrote the code also decide it was done — and separating the two.

This is a summary translation. The full, longer entry is written first in Korean — see the Korean version.

Starting a new diary-style series here: short, first-person notes about mistakes made and fixed while building, not feature announcements.

First entry: a ticket came back marked DONE. The agent that wrote the code also ran the tests and reported "passing." It was uncomfortable to just accept that and move on — and I couldn't say why at first.

The reason turned out to be simple: no human team lets an author merge their own PR, because the author is the person least able to see the blind spots in their own logic. I had been asking one agent to both write the code and judge whether it was correct. A green CI run only means "the author's own tests passed" — it says nothing about whether the work actually meets the requirement, especially when the same blind spots wrote both the code and the tests.

So the structure changed: the agent that implements a task no longer gets to decide its own work is finished. It has to pass an independent verification step before moving to review, and a human still holds the final gate there. Cases where no one ever renders a verdict — neither pass nor fail — are now surfaced instead of quietly disappearing, since an unjudged task sitting in limbo was the most dangerous state of all.

One line: if the author also grades the work, "passing" is a self-report, not a verdict.

Comments

Comments are coming soon.