The Months I Outsourced My Sense of What Counts as Finished

For nearly a year I let a model own my definition of finished. I treated the model’s green light as the real one for nearly a year. On three separate products I would write a short prompt, generate a spec or a flow, run it past the team once, and mark the ticket done when the output looked coherent. The first time it cost us was on the children’s ministry curriculum tool. We shipped a lesson-planning flow that the model had declared complete after three iterations. Two weeks later the volunteer completion rate dropped twelve points because the flow assumed every user could read a dense settings panel in under a minute. I had never sat with an actual volunteer to watch them stall on that panel; the model had simply never flagged it.

The second cost showed up in the analytics review. We had defined “finished” as “all core states rendered without error.” By that metric the feature passed. By the metric that actually moved retention it failed. I had outsourced the definition of finished to a system that had no stake in the outcome and no memory of the last five times we had made the same mistake.

How the First Prompt Loop Replaced My Own Taste Test

Industry guidance on a clear definition of done exists precisely because a definition of finished cannot be outsourced to a tool. Charlie Munger’s circle of competence is simple: know the perimeter of what you can judge reliably and stay inside it when the stakes are real. I had started treating the model as an extension of that circle rather than a tool that sits outside it. The first prompt would produce a v1 that looked plausible. Instead of running my own quick test against the actual user constraint I had observed before, I would ask the model to refine its own output. Each loop made the artifact smoother and moved me one step further from the raw edge where judgment is formed.

After four or five loops the artifact felt finished to me because it felt finished to the model. The competence boundary had quietly shifted outward to include whatever the model would sign off on. Munger warned that the biggest losses come from people who do not know where their circle ends. I was not violating the circle through overconfidence in my own knowledge; I was shrinking the circle by refusing to exercise the judgment that defines its edge.

The pattern repeated on the sermon resource platform. I let the model decide when a content-tagging system was ready for beta. It declared the taxonomy complete after we had covered the top thirty sermon topics. What it could not see was that children’s ministry volunteers tag by emotional tone and length, not topic. I had the data from earlier research but never re-checked it against the model’s structure. The model had no circle; it simply extended mine until I stopped noticing the boundary.

The Quiet Handoff of the Definition of Finished to the Model

Once the loop felt efficient, the final call moved without announcement. I stopped keeping a private note of the three things that still felt off before I opened the chat window. The model’s version became the baseline, and my remaining objections had to clear a higher bar than they used to. If the output addressed the original prompt, the objections sounded like nitpicks rather than signals that the prompt itself had missed the real constraint.

This handoff is easy to miss because nothing dramatic happens. The ticket still moves through the same stages. The only change is that the last human who could have said “this is not finished yet” now waits for the model to say it first. Munger’s framework makes the cost visible: you have stepped outside the circle and are using an instrument that cannot tell you when you have done so.

The cost appears in the data weeks later. Retention curves flatten where they should rise. Support tickets cluster around flows the model had called elegant. By then the decision to treat the model’s verdict as final has already been locked into the shipped artifact.

Rebuilding a Human Definition of Finished

The repair started with a simple rule on one small feature: I would write the acceptance criteria in a private note before the first prompt. The criteria had to be testable by a human in under ten minutes and had to reference an actual past failure. Only after the note existed would I open the model. The model could then generate options, but the note remained the standard. If the generated work met the note, it passed. If it did not, the work was not finished regardless of how polished it appeared.

The second step was to keep the note visible to the team. The model’s output became one input among others rather than the default definition of done. Over six weeks the practice spread to two other product areas. The change was not dramatic in velocity; it was dramatic in the quality of the objections that surfaced early.

Munger’s point is not that tools are dangerous. It is that competence is a perimeter you maintain by deliberate use, not a territory you expand by delegation. When the model handles volume, the remaining human work is precisely the work of knowing where the perimeter sits on each decision. That work does not scale with prompt speed; it shrinks when it is not exercised.

Your Turn: Apply This Today

  • Pick one open ticket this week and write the three-sentence definition of finished in a note before you open any AI tool.
  • Run the simplest possible human check against that definition within twenty-four hours of writing it, even if the model version is not ready.
  • Share the note with one teammate and ask them to flag any criterion the current draft still misses.
  • Refuse to mark the ticket ready until the note’s criteria are met, regardless of model polish.
  • At the end of the week, list which criteria the model never surfaced on its own and keep that list for the next planning cycle.
  • Repeat the same sequence on exactly one new decision the following week; do not expand the scope until the habit is automatic.

The same pattern shows up in how teams set completion criteria for AI-assisted discovery work and how they protect early-stage judgment when metrics are still thin. Both are worth reading next if this landed.

I consult with product leaders on AI-assisted decision boundaries, early-stage completion criteria, and protecting judgment when data is thin. Let’s talk.

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.