The cursor blinked after the final paste. She had dropped every lesson outline from the past three years into the same thread, added the volunteer notes from last quarter, and hit send. The printer in the corner started its warm-up whine while her eight-year-old tugged at the edge of the table asking about snack plans. She clicked generate, waited twelve seconds, then hit print without reading the full output.
The pages stacked warm in her hands. She folded the top sheet to fit the folder she would carry into the classroom the next morning. Nothing in the text told her whether the restless boys in the back row would actually sit through the activity or whether the story hook would survive once the first paper airplane flew. She closed the laptop anyway.
That moment showed the exact point where the inner discipline gets handed off. The model produced something usable on the surface, yet the coordinator still carried the real weight of deciding whether it would serve the actual children who would show up. Most product teams treat that handoff as solved once the output appears. It is not. The fix is a habit of agent outcome checks: deciding what good looks like before the model is ever called.
The Rule That Forces Agent Outcome Checks Before Model Calls
John Wesley boiled his movement’s daily practice down to three rules that had to be answered before any action: do no harm, do good, stay in love with God. The order mattered. Harm had to be ruled out first. Only then could the next two rules apply. The same sequence exposes what current agent workflows skip.
When a children’s ministry lead feeds three years of outlines into a model, the first rule requires naming the possible harm before generation starts. Harm here looks like a lesson that assumes every child can read at grade level or that ignores the one volunteer who always arrives ten minutes late. Running the check after the output lands wastes the moment the model could have been steered away from those edges. Teams that insert the harm question before the call produce fewer pages that still require heavy rewriting on Saturday night.
The second rule then narrows what counts as good. Good is not more content. It is the single activity the lead can finish in the seven minutes between unloading the car and unlocking the classroom door. The model rarely surfaces that constraint unless the prompt names the real validator and the real time budget first.
How Volunteer Handoffs Expose Missing Guardrails
This is the same instinct behind writing the test before the implementation: define the acceptable result first, then let the system try to meet it.
I once watched a coordinator print a generated lesson on a Thursday, then spend Saturday night cutting the activity into two smaller pieces because the original required six separate supplies. The volunteer who actually taught it never saw the original version. She only saw the version that fit the tote bag she could carry on the bus. The model had optimized for completeness, not for the handoff that actually happened.
Wesley’s rules surface here as a routing decision. The harm check asks whether the generated steps can survive being edited by someone who never met the model. The do-good check asks whether the final printed page still moves the needle on the one outcome the church set for that age group. When those questions sit inside the workflow instead of after it, the printed sheet carries fewer surprises into the room.
The same pattern appears in larger platforms. A feature that auto-suggests small group questions often ships without asking which volunteer will read the suggestion aloud and which child might push back on the wording. The missing guardrail is not more model calls. It is the named human who must sign off on the outcome before the file reaches the printer.
Building Agent Outcome Checks Into a Daily Practice That Survives Model Changes
Models update. The interface changes. The prompt that worked last quarter fails when the safety layer shifts. Wesley’s third rule, stay in love with God, points to the habit that outlasts any single tool. It is the repeated act of naming the actual person who will use the output and the actual Sunday moment that will test it.
Product teams that build this habit keep a running log of every model-assisted task. Each entry records the named outcome for that Sunday and the named person who will validate it before the lesson reaches the classroom. When the model changes, the log still tells the team which tasks moved the needle and which ones only moved pixels. The practice does not require new tooling. It requires the same three questions asked at the start of every session rather than the end.
The discipline compounds. After four weeks the log shows which prompts consistently produce lessons that survive the seven-minute volunteer window. After eight weeks the team can retire prompts that only look impressive in the chat window. The model can be swapped without losing the thread that connects the output to the children who will sit on the carpet.
Your Turn: Apply This Today
- Pick one recurring model-assisted task this week and write the exact Sunday outcome it must serve before you open the chat window.
- Name the single human who will validate that outcome on paper or screen before the file leaves your desk.
- Log the task, the outcome, and the validator in a simple note each time you use the model; keep the note next to the printed result.
- At the end of the week, mark which entries still match what actually happened on Sunday and which ones required changes after the model output.
- Delete or rewrite any prompt that produced pages the validator had to rewrite for more than five minutes.
- Run the same log for a second week with the revised prompts and compare the validator time required.
The Goal Loop I Set That Ignored Sunday Validation shows what happens when the log never reaches the person who actually teaches. The 8x Claim That Still Leaves the Print Step Alone tracks how generation speed hides the real cost that appears only after the pages reach the volunteer.
I consult with product leaders and ministry teams on AI routing decisions, volunteer handoff constraints, and outcome logging practices that survive model changes. Let’s talk.

