The question I keep getting is this: with agents collapsing the build step, what does a product manager actually own now? The answer is narrower and more decisive than most teams expect. The PM stops shipping features and starts choosing the exact places where human taste and judgment still touch real people on Sunday morning.
This shift exposes a common misread. Teams treat the model output as the finished product and assume the remaining work is review or polish. In practice the model removes volume but not consequence. The places where output meets a volunteer printing curriculum or a pastor checking a sermon illustration still require someone to decide what stays and what gets cut.
John Wesley’s three simple rules give the right lens. Do no harm. Do good. Attend to the ordinances of God. Applied to product work, these rules move from personal piety to operational discipline. They force the PM to treat every handoff as a point of potential harm or grace rather than a simple quality gate.
The Orchestration Layer That Replaced the Spec Doc
Spec documents used to carry the weight of decisions. Now the model generates options so quickly that the real artifact is the set of rules that decide which options reach a human. I watched one children’s ministry tool replace a 40-page spec with three checkpoints: does this output require a volunteer to explain it out loud, does it assume a printer is available, and does it leave the user with a next action that takes under seven minutes.
Those checkpoints are not features. They are filters the PM owns. Without them the model produces polished but unusable lessons that volunteers abandon after the first page. The orchestration layer is the living document that updates after each real Sunday failure rather than before each sprint.
Wesley would recognize this pattern. His societies did not rely on new rules written in advance. They relied on class leaders who knew exactly where harm could enter the weekly meeting. The PM now fills that role for the product.
Taste Decisions That Survive Agent Output
Most taste decisions used to happen during build. Now they happen at the moment of selection. In one case a model produced 12 possible illustrations for a kids’ lesson on generosity. Eleven were theologically fine and visually strong. One used an image that worked in English but translated as mockery in two of the target languages. The model had no way to flag it.
The PM’s job is not to improve the prompt until the bad option disappears. The job is to keep the final choice inside the three Wesley rules. Does this option risk harm to a child reading it in another culture? Does it do obvious good for the volunteer who has to teach it? Does it leave room for the local church to adapt it as a means of grace rather than a finished script?
These questions cannot be fully automated because they depend on knowledge of actual Sunday contexts. The model can surface options. Only the PM can apply the filter that protects the user who never asked for more options.
The Handoff Metric No Dashboard Tracks Yet
Completion rate still matters, but it arrives too late. The metric worth tracking is the percentage of agent output that reaches a human without requiring rework on Monday morning. In the SermonCentral workflow we began logging every instance where a generated resource had to be edited after a volunteer tried to use it in real time. The number dropped only when the PM owned the three taste checkpoints instead of delegating them.
This metric is invisible to standard analytics because the rework happens outside the product. It shows up in email threads, printed pages with handwritten corrections, and volunteer churn. Wesley’s first rule, do no harm, translates directly into keeping that rework number near zero at the point of human contact.
The teams that still measure only velocity or output volume miss the actual constraint. The constraint is no longer production. It is the quality of the last human decision before the resource leaves the building.
Your Turn: Apply This Today
- Map the three specific handoff points in your current feature where model output first touches a real user and write them down before the next sprint planning.
- Define the exact taste checkpoint at each handoff using one of Wesley’s three rules and log it in the same place your team keeps acceptance criteria.
- Review the last agent-generated artifact that reached a Sunday morning user and mark every element that would have failed one of the three checkpoints.
- Assign yourself ownership of the final selection step for one upcoming release instead of routing it through an automated review.
- Track the Monday morning rework count for the next two weeks and note which checkpoint would have prevented each instance.
- Write a one-paragraph description of what “attend to the ordinances” means for your product’s actual users and share it with the two engineers closest to the agent pipeline.
The same pattern showed up in The Inner Skills I Let the Model Train Out of Me and The PM Who Inherited the 8x Target. Both pieces trace what happens when the human filter disappears.
I consult with product leaders building ministry tools on AI orchestration layers, taste decisions at human handoffs, and defining non-dashboard metrics. Let’s talk.

The volunteer coordinator’s laptop sat open on the scarred wooden table, its screen reflecting the overhead bulb. Her youngest reached across for another piece of bread while the older one asked about homework. She ignored both long enough to drag three phone photos into the chat window and type the phrases she had heard at the door that morning. What followed was a small lesson in screenshot-driven discovery: raw images and real words, not a written brief, shaped the next build.
A product team tracking 2,400 feature requests last quarter reported that 68 percent came from power users who logged in at least three times a week. The obvious reading was that the backlog now reflected the clearest priorities. The data said the opposite once the team examined who never appeared in the logs at all. Continuous discovery exists to close exactly this gap between what the data records and what the work actually requires.
Most product teams treat verification layers as the thing that slows an AI rollout, but ministry contexts show the opposite pattern: tools without explicit trust mechanisms reach a usage ceiling within weeks and never recover. The missing piece is almost always an AI audit trail: a record, inside the workflow, of what the model produced and who verified it.
The children’s director shoved her laptop across the kitchen table. The screen showed three Figma frames she had built after midnight. One had a big green button for “daily verse push.” Another split the screen between parent notes and a child progress bar. The third tried to guess what a family needed based on last week’s attendance. This is the moment every AI ministry tool faces: the dataset points one way and the person who knows the work points another.