The PM Role That Stopped Shipping and Started Choosing

Workflow diagram and product brief representing product manager decisions

The question I keep getting is this: with agents collapsing the build step, what does a product manager actually own now? The answer is narrower and more decisive than most teams expect. The PM stops shipping features and starts choosing the exact places where human taste and judgment still touch real people on Sunday morning.

This shift exposes a common misread. Teams treat the model output as the finished product and assume the remaining work is review or polish. In practice the model removes volume but not consequence. The places where output meets a volunteer printing curriculum or a pastor checking a sermon illustration still require someone to decide what stays and what gets cut.

John Wesley’s three simple rules give the right lens. Do no harm. Do good. Attend to the ordinances of God. Applied to product work, these rules move from personal piety to operational discipline. They force the PM to treat every handoff as a point of potential harm or grace rather than a simple quality gate.

The Orchestration Layer That Replaced the Spec Doc

Spec documents used to carry the weight of decisions. Now the model generates options so quickly that the real artifact is the set of rules that decide which options reach a human. I watched one children’s ministry tool replace a 40-page spec with three checkpoints: does this output require a volunteer to explain it out loud, does it assume a printer is available, and does it leave the user with a next action that takes under seven minutes.

Those checkpoints are not features. They are filters the PM owns. Without them the model produces polished but unusable lessons that volunteers abandon after the first page. The orchestration layer is the living document that updates after each real Sunday failure rather than before each sprint.

Wesley would recognize this pattern. His societies did not rely on new rules written in advance. They relied on class leaders who knew exactly where harm could enter the weekly meeting. The PM now fills that role for the product.

Taste Decisions That Survive Agent Output

Most taste decisions used to happen during build. Now they happen at the moment of selection. In one case a model produced 12 possible illustrations for a kids’ lesson on generosity. Eleven were theologically fine and visually strong. One used an image that worked in English but translated as mockery in two of the target languages. The model had no way to flag it.

The PM’s job is not to improve the prompt until the bad option disappears. The job is to keep the final choice inside the three Wesley rules. Does this option risk harm to a child reading it in another culture? Does it do obvious good for the volunteer who has to teach it? Does it leave room for the local church to adapt it as a means of grace rather than a finished script?

These questions cannot be fully automated because they depend on knowledge of actual Sunday contexts. The model can surface options. Only the PM can apply the filter that protects the user who never asked for more options.

The Handoff Metric No Dashboard Tracks Yet

Completion rate still matters, but it arrives too late. The metric worth tracking is the percentage of agent output that reaches a human without requiring rework on Monday morning. In the SermonCentral workflow we began logging every instance where a generated resource had to be edited after a volunteer tried to use it in real time. The number dropped only when the PM owned the three taste checkpoints instead of delegating them.

This metric is invisible to standard analytics because the rework happens outside the product. It shows up in email threads, printed pages with handwritten corrections, and volunteer churn. Wesley’s first rule, do no harm, translates directly into keeping that rework number near zero at the point of human contact.

The teams that still measure only velocity or output volume miss the actual constraint. The constraint is no longer production. It is the quality of the last human decision before the resource leaves the building.

Your Turn: Apply This Today

  • Map the three specific handoff points in your current feature where model output first touches a real user and write them down before the next sprint planning.
  • Define the exact taste checkpoint at each handoff using one of Wesley’s three rules and log it in the same place your team keeps acceptance criteria.
  • Review the last agent-generated artifact that reached a Sunday morning user and mark every element that would have failed one of the three checkpoints.
  • Assign yourself ownership of the final selection step for one upcoming release instead of routing it through an automated review.
  • Track the Monday morning rework count for the next two weeks and note which checkpoint would have prevented each instance.
  • Write a one-paragraph description of what “attend to the ordinances” means for your product’s actual users and share it with the two engineers closest to the agent pipeline.

The same pattern showed up in The Inner Skills I Let the Model Train Out of Me and The PM Who Inherited the 8x Target. Both pieces trace what happens when the human filter disappears.

I consult with product leaders building ministry tools on AI orchestration layers, taste decisions at human handoffs, and defining non-dashboard metrics. Let’s talk.

The Year I Optimized for Tokens Instead of Presence

Printed papers representing the print step where presence matters

I spent a year routing every product decision through token budgets and latency targets. The rule was simple: if an AI could generate the draft, the spec, or the volunteer script in under four seconds, we shipped it. What I missed was the moment the curriculum stopped matching what a real 7-minute volunteer could actually finish on a Tuesday night with a stack of printed pages in front of them.

The cost showed up in completion rates. We had higher output and lower retention. The volunteers kept opening the materials once, getting stuck on the third activity, and printing something else instead. I kept measuring the wrong thing because the model made it easy to keep measuring the same thing.

Charlie Munger’s latticework of mental models insists that a single lens will always blind you to what the model cannot see. I had treated speed and token cost as the only two models that mattered. Presence was not in the lattice at all.

The routing habit that trained out my own judgment

Every ticket went through the same path. First the model wrote the acceptance criteria, then it suggested the flow, then it generated the test cases. I reviewed the output for factual errors and moved on. The loop felt disciplined until I noticed I was no longer asking whether the volunteer would still be standing at the table when the third page asked them to cut out six small crosses.

The habit removed friction, but it also removed the friction that forces judgment. When the model cannot see the table, the scissors, or the child who needs the activity finished before the parents arrive, the only remaining signal is whether the tokens stayed under budget. That signal is too narrow for ministry work.

After six months the pattern was clear. Features that scored highest on our internal efficiency dashboard were the same ones that generated the most Monday-morning support tickets from volunteers who could not finish the lesson.

Presence as the override at the print step

The print step is where the abstract workflow becomes physical. I started showing up at two different churches during the hour when volunteers actually assembled the materials. One Tuesday I watched a leader try to follow a three-column layout on an 8.5-by-11 sheet that had already been folded twice. The model had optimized the layout for screen reading. The volunteer had no screen.

That single observation overturned three accepted stories we had shipped in the previous quarter. The fix was not another prompt. It was adding a hard stop: no layout change could ship until someone had printed it on the exact paper stock the volunteer would use and completed the activity with a stopwatch running.

The latticework requires multiple models to stay in tension. Token cost and completion time under real conditions cannot be collapsed into one number. When they conflict, the print step wins because that is where the actual work happens.

How emotional intelligence shows up in a sprint review

Sprint reviews used to be a slide deck of velocity and error rates. After the print-step visits I changed the format. The last ten minutes became a short video of a volunteer attempting the new flow while thinking out loud. No narration, no cuts. The team watched the pauses, the page flips, the quiet “this doesn’t make sense” that never appeared in the analytics.

The first time we ran the video, two planned features were removed before the next sprint began. Not because the metrics said they were slow, but because the volunteer’s frustration was audible. That data had always existed; we had simply routed around the channel that carried it.

Emotional intelligence in this setting is not a soft skill layered on top of product work. It is the additional mental model that tells you when the existing metrics have stopped tracking the outcome that actually matters to the volunteer. Without it the lattice stays incomplete.

Your Turn: Apply This Today

  • Pick one live ministry workflow that touches volunteers this week and block fifteen minutes on your calendar to watch it in person without a laptop open.
  • Print the exact materials the volunteer will use on the same paper they receive and complete the full activity yourself with a timer running.
  • Write down the single moment where the workflow required a judgment the current model cannot see, then log that exact moment in the ticket system as a required check.
  • In the next sprint review, replace one metrics slide with a two-minute unedited clip of a real user attempting the task and end the meeting without discussion.
  • Remove one previously accepted efficiency target that no longer matches what you observed during the print step and document the removal in the same place you track token savings.
  • Schedule the same fifteen-minute presence audit on the same workflow for next week and record only the one adjustment you made because you were physically there.

The Morning the Print Step Broke and The Inner Skills I Let the Model Train Out of Me both trace the same gap between what the model optimizes and what the volunteer actually experiences.

I consult with product leaders and ministry teams on presence audits, sprint-review redesign, and workflow overrides that protect volunteer completion. Let’s talk.

The Quality Erosion No 8x Claim Ever Mentions

Hands using a caliper for quality control, representing AI agent review layers

The fastest agent stacks in ministry tools do not eliminate bottlenecks. They delete the review layers that surface failures before they reach the volunteer who opens the app on a Saturday night with no backup plan.

This pattern shows up when teams optimize for output velocity above all else. One agent generates the schedule, another approves the assignment, a third pushes the notification. Each step removes a human checkpoint that once caught the mismatch between a new curriculum and a classroom without the right materials. The result is not faster ministry. It is ministry that breaks in the exact moment it is needed most.

Jensen Huang has argued that sovereign control over AI systems requires owning the full stack of constraints, not just the generation layer. Applied to product work, that means the team that builds the agent must also own the verification surface. Without that ownership, speed becomes a tax paid later in rework, trust erosion, and quiet volunteer attrition.

The security review that disappeared after the first agent

The first agent that auto-assigns volunteers usually passes internal tests. It respects role permissions and avoids obvious conflicts. The second agent, the one that routes last-minute swaps, rarely receives the same scrutiny. It operates on fresher data and therefore inherits whatever drift has entered the underlying volunteer profile fields since the last manual audit.

That drift shows up as a children’s ministry coordinator receiving a high-school curriculum assignment because the system still lists her as available for any age group. The agent chain treats the profile as ground truth. The missing step is the one that once required a human to confirm the profile had not been updated in twelve months. Once the second agent ships, that confirmation step is treated as legacy process rather than active control surface.

Teams that measure only agent completion rate never see the downstream ticket volume. The tickets arrive as parent complaints or volunteer no-shows instead of system errors. By then the original optimization story has already been told to leadership.

The quality signal that only appears on Sunday

Volunteer scheduling systems produce their clearest failure data on Sunday morning. A printed roster that does not match the actual check-in list, a notification that arrived after the volunteer had already left for church, a curriculum link that points to last quarter’s material. These signals are cheap to observe in person and expensive to reconstruct from logs.

Agent workflows push that observation moment later. The schedule looks complete in the dashboard on Thursday. The quality gap only becomes visible when the volunteer arrives and the room is missing supplies. By treating Sunday data as anecdotal rather than the primary signal, teams lose the ability to tune the reversal rules that matter.

Huang’s point about sovereign systems is that the owner must decide what counts as acceptable output before deployment, not after. In this domain the acceptable output includes zero mismatched age-group assignments on Sunday. Anything less transfers the verification cost onto the volunteer.

Reinstating the manual gate without slowing the whole system

The manual gate does not need to sit in front of every agent action. It needs to sit in front of the actions that change a volunteer’s actual Sunday morning. A lightweight rule can flag any swap that crosses age bands or curriculum types and route only those cases to a human for confirmation. Everything else proceeds.

The cost is measured in minutes per week, not hours. The return is measured in avoided no-shows and retained volunteers. Most teams reject this approach because they have already framed every added step as a violation of the 8x target. The frame itself is the problem. The target was never defined against the cost of a volunteer who stops opening the app.

Reinstating the gate also requires updating the success metric. Completion rate of the agent chain is replaced by completion rate of the actual volunteer assignment on Sunday. The two numbers diverge quickly once the review layer is restored.

Your Turn: Apply This Today

  • Identify the single agent action in your current volunteer scheduling flow that changes an age-group assignment and add an explicit reversal protocol that requires human confirmation before that change commits.
  • Run the next scheduling cycle with that reversal turned on and log every flagged case along with the time required to resolve it.
  • Compare the flagged cases against the previous three cycles that ran without the reversal and count how many would have produced a Sunday mismatch.
  • Update the dashboard metric from agent completion percentage to actual volunteer check-in match rate measured on Sunday afternoon.
  • Document the total human minutes spent on reversals for one full month and present that number alongside any 8x output claims made about the agent stack.
  • Remove the reversal only after the mismatch rate stays below two percent for four consecutive weeks.

The same erosion pattern appears in curriculum publishing and notification systems. Both “The 8x Claim That Hid the Real Cost” and “The 8x Output That Created the Monday Morning Rework” trace the downstream cost when review steps are treated as optional.

I consult with product leaders building ministry tools on agent workflows and review gates. Let’s talk.

The Morning the Print Step Broke

Team keeping a human in the loop while reviewing AI output

The cursor spun for three full seconds before the folder finally refreshed. Maria leaned over the laptop on her kitchen island, one hand still holding a half-zipped backpack, the other clicking the refresh button a second time. Outside the window the sky was still dark blue. Inside, the new agent had overwritten the Sunday curriculum PDF at 2:17 a.m., and now the page numbers no longer matched the printed leader guides stacked on the counter. This is what happens when a workflow drops the human in the loop: the agent acts on its own data and no person reviews the result before it ships.

She opened the file. The volunteer names that had been added by hand last Tuesday were gone. The craft-supply list on page three had shifted to page five. She checked the clock, then opened the group chat. Three volunteers had already asked which room they were in this week. One message sat unread from 6:15: “Printout looks different. Is this right?”

Maria did not answer. She started scrolling instead, trying to rebuild the missing pieces before the first car pulled into the parking lot.

Seneca wrote letters to keep his thoughts from drifting into abstraction. He insisted the recipient write back, because only the reply forced the writer to stay accountable to a real person rather than a tidy idea. The rule still holds when the writer is an agent that never sleeps.

The agent that regenerated Maria’s PDF had no correspondent. It optimized for completeness and speed. It had no mechanism that required the output to land in the hands of the person who would stand in front of eight first-graders at 9:15 and actually use it. The loop closed on its own metrics. The human thread snapped before anyone noticed.

The first break happened at the handoff. The agent produced a clean file and moved it to the shared drive. No step required the file to be opened, printed, and physically carried into the classroom by the same person who would teach from it. The curriculum team had removed that step two months earlier because it added four minutes to the overnight run. Those four minutes had been the only place where Maria’s handwritten volunteer assignments survived. Without them the printed sheets became generic. The volunteers arrived unsure where to go, and Maria spent the first ten minutes of her morning reassigning rooms instead of praying with the team.

Speed without the return letter creates the same gap at larger scale. Product teams ship new agent features the moment the model passes internal tests. The feature moves from staging to production without ever requiring the ministry leader who will rely on it to confirm the output still matches her actual room list, her actual supply closet, her actual Sunday. The agent treats the next run as proof of correctness. The leader treats the next Sunday as proof. Those two proofs rarely meet.

Seneca would have called the missing reply a failure of address. The letter must reach someone who can answer, not merely someone who can receive. When the agent overwrote the volunteer names, it addressed only the shared drive. It never addressed Maria’s memory of which six-year-old needed the larger printout because his mother had asked last month. That detail lived only in the human who had written it down. Once the file replaced the paper, the detail had no address left.

A single checkpoint restores the correspondence. Before the agent commits the new file, the system must surface the changed sections to the exact person who last edited them and require a one-click confirmation or reversal. The confirmation is the reply. Without it the loop stays open and the human stays in the thread.

The checkpoint would have caught the page shift at 2:18 a.m. Maria’s phone would have shown the three altered pages side-by-side with her previous version. She could have tapped “keep names” in thirty seconds. The agent would then have regenerated with the constraint instead of discarding it. The volunteers would have found their rooms listed correctly at 6:42. Maria would have answered the group chat instead of rebuilding the file.

The same pattern repeats in other tools. An agent that schedules small-group reminders drops the names of two new attendees because the calendar sync overwrote the private note field. A giving dashboard agent summarizes weekly totals without preserving the elder’s handwritten flag on which family just lost a job. Each time the agent closes the loop on its own data, the human who stands in the room or sits in the elder meeting loses the thread that only she held.

Deliberate pacing inserts the reply before the next cycle begins. It does not slow adoption. It keeps the adoption from breaking the moment the printed sheet meets the actual person who must use it.

The Exact Moment the Workflow Lost the Human in the Loop

The failure occurred between the agent’s commit and Maria’s first refresh. No one on the product side saw the output until the printed sheets were already in volunteers’ hands. The agent had logged a successful run. The team celebrated the overnight completion rate. The loss appeared only when Maria stood at the island trying to match names to rooms she could no longer see on the page.

That gap is structural, not accidental. The system measured success by file generation, not by whether the person who opens the file can still recognize her own prior work. Once the metric stops at generation, the human who must teach from the output becomes invisible to the loop.

Why Speed Without the Letter Back to the Owner Fails

Safety researchers describe this as keeping a human in the loop: a person stays responsible for approving the system’s output before it takes effect.

Seneca required the recipient to answer because an unanswered letter drifts into monologue. The agent that overwrote the volunteer names performed a monologue. It optimized layout and page count without ever confirming that the names still mattered to the person who added them. The result looked complete on the drive and incomplete in the classroom.

Teams repeat the pattern when they celebrate faster regeneration times without measuring whether the regenerated file still carries the context only the owner possesses. The speed feels like progress until the first Sunday the printed sheet no longer matches the room assignments the leader memorized the week before.

The Single Human-in-the-Loop Checkpoint That Would Have Caught It

The checkpoint is a forced reply step placed immediately after the agent proposes changes. The owner receives a side-by-side view of what moved and what stayed. She confirms or reverts before the file becomes the source of truth for printing. The step adds seconds to the agent’s cycle and removes hours of re-work on Sunday morning.

The same checkpoint can sit inside any agent that touches ministry artifacts. It does not require new model training. It requires only that the system treat the human who last edited the file as the required correspondent rather than an optional reviewer.

Your Turn: Apply This Today

  • Open the current agent loop that produces Sunday materials and add a single confirmation screen that shows only the changed sections to the owner before the file is marked final.
  • Set the confirmation deadline to Thursday 8 p.m. so the reversal step happens before any print run begins on Friday.
  • Log the number of reverts that occur in the first two weeks; treat that number as the real measure of whether the agent preserved owner context.
  • Remove the auto-commit rule that lets the agent overwrite a file without an owner reply; replace it with a pending state that expires only after confirmation.
  • Run the same checkpoint on the small-group reminder agent this week by surfacing any dropped attendee names to the group leader for approval before messages send.
  • Document the exact time cost of the checkpoint in minutes; compare it against the time Maria spent rebuilding the file at 6:42 a.m. and keep the lower number.

Two earlier posts on this blog examined the same tension between agent speed and owner memory: “The 7-Minute Volunteer Still Needs a Paper Trail” and “When the Dashboard Forgets the Elder’s Note.” Both are linked from the archive.

I consult with product leaders building agent loops for ministry tools and church operations teams redesigning overnight content pipelines. Let’s talk.

The 8x Claim That Hid the Real Cost

Dashboard illustrating the AI productivity paradox

A product team reported their AI-assisted planning workflow delivered eight times the output volume in the first month. It is a textbook AI productivity paradox: the volume goes up while the real value quietly goes down. The dashboard showed more lesson outlines, more email drafts, more volunteer schedules generated per hour than any prior quarter. Yet six weeks later the same team watched completion rates among their 7-minute volunteers drop below the baseline they had before the new system.

The obvious reading treats the multiplier as pure gain. The data actually pointed the other direction once second-order effects entered the picture. Volunteers who received the polished outputs no longer understood the reasoning chain that produced them. When a schedule conflict appeared on Sunday morning, they lacked the context to adjust it without breaking the original intent.

This is the foundational misread that causes product teams to celebrate velocity while quietly eroding judgment at the edge. Charlie Munger’s latticework demands that any single model, in this case raw output multiplication, must be cross-checked against at least two other models before it earns trust. One model tracks first-order production. Another tracks whether the recipient can still reconstruct the decision logic six weeks later. When those two models conflict, the latticework flags the multiplier as incomplete rather than impressive.

Where the AI Productivity Paradox Starts: the 8x Number

The eightfold increase came from an agent loop that ingested past curriculum, current calendar constraints, and volunteer availability, then emitted complete weekly plans. The prompt engineering was tight. Token usage stayed low. Human review time fell from forty minutes to five. On paper the system looked like a decisive win for any ministry running on volunteer labor.

The number ignored what happened after the plan left the system. Volunteers received a finished document rather than the intermediate constraints the agent had weighed. They could execute the steps, but they could not explain why a particular story was chosen for a given age group or why one room was double-booked on purpose. When the inevitable exception arrived, execution stalled because the mental model had never been transferred.

Munger would have asked which additional model revealed the omission. In this case the missing model was the volunteer’s ability to re-derive the plan under partial information. The 8x figure measured only the production side of the lattice.

The Hidden Regression Behind the AI Productivity Paradox

Regression appeared first in small decisions that used to be automatic. A volunteer who once swapped two craft supplies without asking now waited for the next planning cycle. Another who used to shorten a lesson when time ran long now read every line even when the room was restless. The surface metrics still looked acceptable until the cumulative effect reached the children and parents.

The second-order failure mode is simple: when output volume rises faster than decision context, the volunteer shifts from owner to operator. Ownership requires the ability to defend a change. Operation only requires following the last instruction. Sunday is when the difference becomes visible because exceptions arrive faster than the next agent run.

Teams that track only generation counts miss this shift entirely. The latticework requires a second measurement: the percentage of volunteers who can state the original constraint set without opening the document. When that percentage falls, the multiplier has begun to subtract capability rather than add it.

The Logging Practice That Surfaces It Early

Economists have debated a similar gap for decades under the name the productivity paradox: new technology shows up everywhere except in the measured results.

The practical countermeasure is a minimal log kept by the volunteer, not the system. After each handoff the volunteer writes two sentences: the constraint they remember most clearly and the one change they would make if the constraint were removed. The log takes less than sixty seconds and travels with the printed plan.

Review happens once a week. A product owner or lead volunteer scans the accumulated sentences for patterns. Repeated inability to name a constraint signals that the agent output has become opaque. A repeated suggested change that the agent would have rejected reveals a missing rule in the prompt. Both signals appear before completion rates move.

This practice adds one override protocol and one weekly log review to any agent loop already running. The override lets a volunteer mark a plan “context needed” and receive the constraint list before the next cycle. The review meeting lasts fifteen minutes and uses only the volunteer sentences, never the original agent trace.

Your Turn: Apply This Today

  • Pick one recurring agent output this week and require the volunteer to add the two-sentence log before they leave the room.
  • Schedule a fifteen-minute Friday review that reads only the volunteer sentences and flags any plan where the constraint cannot be restated.
  • Add an override button in the handoff interface labeled “send constraints” that bypasses polished output and returns the raw decision factors.
  • Measure the percentage of volunteers who correctly name the top constraint on Monday morning; treat any drop below 80 percent as a system defect.
  • Remove one generation step from the agent prompt and replace it with an explicit request that the volunteer must confirm the constraint in their own words.
  • Run the same lattice check on any new agent you ship: first measure output volume, then measure retained decision context at the two-week mark.

Two earlier posts examined the same tension from different angles. “When Speed Outruns Context” traced how notification volume alone can erode volunteer judgment, and “Constraint-First Prompts” showed the minimal prompt change that keeps reasoning visible. Both posts are linked from the archive.

I consult with product leaders building AI workflows for volunteer-driven ministries on output measurement, constraint logging, and second-order failure detection. Let’s talk.

The PM Who Inherited the 8x Target

Product manager planning AI agent oversight

You’re the product manager who just inherited the 8x target for an AI sermon-prep tool inside a mid-sized denomination. The previous lead left three weeks ago. The only brief you received was a slide deck that promised the model would cut prep time for volunteer coordinators from forty minutes to five. No one on the pastoral staff has used an LLM for anything beyond a quick email draft. Before you touch the model, the job you actually inherited is AI agent oversight: deciding what the agent may do alone and where a person must sign off.

That target sits on your desk next to a login for an internal instance of Claude and a folder labeled “prompt library v2.” Your first instinct is to open the folder and start testing.

This is the foundational misread that causes product teams to ship agents that look productive in demos and then quietly fail on Sunday morning. The mistake is treating the model as the primary actor and the human as a reviewer who will simply catch whatever slips through. Teresa Torres’s continuous discovery framework shows why this breaks. Torres insists the team must keep returning to the real user in context, updating the problem statement each time the solution changes. When the solution is an agent, the problem statement must include the exact moment the work leaves the agent and returns to the person who still carries the outcome.

You cannot discover that moment by iterating on prompts alone. You discover it by watching the handoff itself.

AI Agent Oversight Starts With a Written Job Description

The first artifact you build should not be a system prompt. It should be a one-page job description for the agent that names the exact deliverable, the exact constraints, and the exact person who will sign off. At Sermons4Kids we learned this the hard way when we tried to automate volunteer curriculum summaries. The model produced clean text, yet the volunteer still had to rewrite every summary because the tone did not match the actual kids in their room. We had optimized for readability instead of for the handoff to a person who still had to teach the lesson.

Write the job description the same way you would write one for a part-time ministry assistant. Include the Sunday outcome the human still owns. Include the non-negotiables that cannot be automated—pastoral voice, local context, last-minute room changes. Only after that page exists do you test whether an agent can produce work that survives the handoff.

Teams that skip this step keep discovering the same gap in user interviews six months later. Torres would call that a failure of continuous discovery. The problem statement never updated to include the human checkpoint.

AI Agent Oversight Means Logging the Escalation Path

Every agent you ship will eventually surface something it cannot resolve. The question is whether that moment arrives as a quiet failure or as a clear, logged escalation to the right human. Most teams log token counts and latency. Few log the precise condition under which the agent must stop and hand the thread back.

Map the escalation path in the same format you would use for a customer-support workflow. Define the trigger, the data that travels with the handoff, and the single owner who receives it. Test the path with a real coordinator who has never seen the agent before. If the coordinator cannot act on the handoff in under two minutes, the path is not ready.

The 8x target collapses when three escalations arrive on the same Sunday and no one knows who owns them. Continuous discovery requires that you treat the escalation path as part of the product, not as a support ticket that appears after launch.

Measure the handoff, not the tokens

This mirrors NIST’s guidance in its AI Risk Management Framework, which treats accountability and human review as core controls, not afterthoughts.

The metric that matters is the percentage of agent output that reaches the Sunday outcome without requiring rework by the human who owns it. Track how many minutes the volunteer or pastor still spends correcting or completing the work. Track how often the escalation path is used and whether the outcome improves after the handoff.

At Bible Gateway scale we watched thousands of users open a reading plan generated by an early recommendation model, then immediately close it because the plan did not account for their actual reading rhythm. The model had high completion inside the sandbox. The handoff to the real reader failed. Only after we measured time-to-first-real-read did the metric move.

Token usage and latency will look strong in the dashboard. They will not tell you whether the children’s ministry volunteer finished the lesson on time. That is the only number that converts into durable adoption inside a faith community.

Your Turn: Apply This Today

  • Pick the single agent loop you are most tempted to ship this month and write its one-page job description before you open the prompt editor.
  • Define the exact reversal step that returns control to a named human role—children’s director, worship pastor, volunteer coordinator—and log the data package that travels with it.
  • Run one 15-minute observation with a real user performing the handoff on a printed copy of the job description; note every place they hesitate.
  • Build a two-row dashboard this week: one row for agent output volume, one row for minutes the human still spends after the handoff.
  • Schedule the first follow-up discovery call for the same user forty-eight hours after they first use the agent in a live setting.
  • Write the escalation trigger condition in plain language on a single index card and tape it to your monitor before you write the next prompt.

The posts on measuring what survives the handoff and on running continuous discovery inside ministry tools both explore the same pattern from different angles.

I consult with product leaders shipping AI inside mission-driven organizations on workflow handoffs, escalation design, and discovery metrics that survive Sunday morning. Let’s talk.

The 8x Output That Created the Monday Morning Rework

Workspace showing agent output rework piling up on Monday

The “8x Output Multiplier” framework sold by generative AI rollout consultants claims that swapping routine writing and research tasks for agent workflows will produce eight times the volume of drafts, summaries, and lesson outlines without altering the surrounding approval steps.

This framework breaks down for product teams shipping to ministry leaders because it treats every generated token as net new capacity, and it ignores the agent output rework that the extra volume quietly creates. In practice the extra volume arrives without the cross-checks that previously caught doctrinal drift, formatting mismatches for print volunteers, or curriculum gaps that only surface during actual prep. The result is an output spike that lands in the same queue that already runs at capacity on Thursday nights.

Teams adopting the framework therefore ship more artifacts that require the same human hours to salvage, exactly the opposite of the time savings projected in the original model.

This is the foundational misread that causes product teams to celebrate early velocity numbers while the Monday morning rework queue grows. The agents accelerate the first pass but leave untouched the validation loops that determine whether the artifact survives contact with a 7-minute volunteer or a pastor finishing slides at 10 p.m. Saturday.

Charlie Munger’s latticework of mental models offers the corrective lens. Munger argued that durable decisions require holding multiple independent models—physics of workflows, second-order effects, and incentive misalignments—simultaneously so that hidden interactions become visible before they compound. Applied to agent adoption, the latticework forces the product team to track not only token count but also the downstream human correction rate and the decay of that correction rate over successive releases.

The Pilot That Looked Clean Until Sunday Prep Began

This mirrors what software teams call technical debt: speed today that quietly compounds into cleanup tomorrow.

A children’s ministry curriculum pilot ran an agent on every weekly lesson outline for eight weeks. Drafts appeared in the shared folder by Tuesday instead of Thursday. Volunteer completion rates inside the first month rose from 61 percent to 79 percent because the volume of available lessons increased. The dashboard showed the expected 8x lift in artifacts produced.

The failure appeared the week the agent rewrote a Jonah story arc to fit a requested word count. The outline dropped the repentance angle that the print handout had carried for three years. A volunteer in Ohio printed the new version, noticed the missing point during Saturday night prep, and spent forty minutes restoring it by hand. Three other volunteers filed the same correction the next morning. The Monday rework log showed fourteen separate human edits across five lessons, erasing the previous week’s time savings in a single cycle.

The pilot metrics never captured the correction step because the measurement stopped at “draft delivered.” The latticework requires tracking the full loop—generation plus validation plus correction—so the second-order cost becomes visible before the pattern repeats across an entire library.

How the Lattice Reveals the Agent Output Rework Failure Mode

Munger’s approach surfaces the interaction between two models that the 8x claim treats as independent: production speed and quality signal decay. When an agent produces more drafts faster, the volume of content that reaches the validation stage increases. If the validation criteria stay unchanged, the probability that a subtle error slips through rises with each additional draft because the reviewer’s attention budget does not scale linearly.

In the curriculum case the quality signal was the three-year archive of volunteer feedback on specific theological anchors. The agent had no persistent memory of that archive. Each new draft therefore reset the error rate to the level observed in week one rather than compounding the institutional knowledge already captured in the feedback logs. The latticework makes that reset visible by forcing the team to hold the model of “accumulated validation data” alongside the model of “generation throughput.”

Without that joint view, teams optimize for the visible number while the invisible correction tax compounds. The same pattern appears in sermon prep agents that drop illustration sources and in small-group discussion agents that omit follow-up questions already proven effective in prior quarters.

The Logging Practice That Catches Agent Output Rework Before It Ships

One team introduced a single mandatory field in the agent output ticket: “Validation delta from last known good version.” The field required the product owner to record any change in structure, scripture reference, or volunteer action step before the artifact moved to the next stage. The log ran for four weeks and immediately flagged three systematic regressions that the volume metric had hidden.

The practice works because it externalizes the latticework check. Instead of relying on an individual reviewer to remember prior versions, the team maintains an explicit comparison record that any later agent run must beat. When the delta field stays empty or shows repeated corrections, the rollout pauses until the guardrail is restored.

Teams that skip the log continue to report rising output while the Monday morning queue lengthens. The latticework predicts exactly this outcome once the models of speed and accumulated validation are examined together rather than in isolation.

Your Turn: Apply This Today

  • Add a “validation delta” field to the agent ticket template this week and require it to be filled before any draft moves past the first human review.
  • Run the last four weeks of agent outputs through the delta check and record how many required structural or theological corrections that the original volume metric never captured.
  • Pick the single most common correction type from that review and write a one-sentence rule the agent must satisfy on the next run; test it on this week’s batch.
  • Schedule a 15-minute Monday review of the delta log with the same two people who currently handle Sunday prep emergencies so the cost becomes visible to the owners of the rework.
  • Reduce the agent’s generation cap by 30 percent for the next sprint and measure whether the correction rate drops faster than the volume loss.
  • Export the delta log as a standing dashboard widget visible to the agent roadmap owner so the second-order cost appears in every planning meeting.

The same pattern shows up in “The 8x Claim That Still Leaves the Print Step Alone” and “The Goal Loop I Set That Ignored Sunday Validation”.

I consult with product leaders shipping AI agents into ministry and volunteer workflows on rollout pacing, validation logging, and correction-rate measurement. Let’s talk.

The Volunteer Desk Where the Escalation Path Was Missing

Volunteer desk where an agent escalation path was missing

The volunteer coordinator’s kitchen table held a single lamp and a stack of half-folded bulletins. At 9:17 p.m. on Thursday she opened the scheduling agent’s latest file. Three names filled the 8 a.m. Sunday slot. None carried a confirmation checkmark. She stared for four seconds, then picked up her phone and began texting the backup list one by one.

The agent had run exactly as built. It pulled availability, applied the preference rules, and produced a clean list. What it lacked was an agent escalation path: any instruction for the moment when no one on that list had actually said yes. She moved to the next manual step because nothing in the workflow told the system to stop and ask a human first.

Teresa Torres describes continuous discovery as ongoing contact with the real moments when users encounter a product. In agent loops the same principle applies, except the critical moments are handoffs rather than clicks. The loop stays trustworthy only when it contains explicit rules for what happens when the automated step ends and a person must act.

Define the agent escalation path before the first run

Most teams start by describing the output they want: a filled schedule, a set of reminder texts, a list of curriculum pages. Torres pushes teams to name the job the user is actually doing at that moment instead. For the volunteer coordinator the job is not “receive a list.” It is “know within two minutes whether the list is safe to act on or whether she must start calling people.”

Writing that job description forces the escalation step into view. The agent must now include a rule that checks confirmation status before it closes the loop. Without the rule the coordinator still does the work, only later and under more pressure.

One children’s ministry platform I watched last year added a single line to its scheduling agent: if fewer than two confirmed names exist four hours before the shift, pause and surface the list to the coordinator with a red banner. The change took one afternoon. The number of Sunday morning emergency texts dropped by half within three weeks.

Why the agent escalation path has to survive the print step

Ministry workflows still contain physical moments. Someone prints the roster, someone tapes it to the door, someone hands a stack of name tags to a greeter. Torres would call these the observable moments of use. An agent that stops checking once the digital file is generated misses every one of them.

The useful escalation rule therefore triggers after the print action, not before. In practice this means the agent waits for a simple signal— a photo of the printed list, a checkbox marked “posted,” a short voice note from the volunteer on site. Only then does it consider the handoff complete.

Without that post-print check the coordinator receives a clean digital file on Thursday night and still walks into Sunday unsure whether the paper version ever reached the right door. The agent has done its job on paper; the actual job remains unfinished.

Measure the handoff, not the generation count

Reliability teams formalize this in the idea of an on-call escalation policy: someone is always named as the next human in the loop.

Teams often track how many schedules the agent produces each week. That number rises quickly and looks like progress. Torres would ask instead how many of those schedules reached a confirmed human action without further intervention.

Track the percentage of shifts that move from “agent output” to “volunteer confirmed on site” without the coordinator reopening her phone after 9 p.m. When that percentage stalls, the escalation path is still missing or mis-timed.

One team I worked with replaced its weekly generation metric with a simple handoff score. Every Monday they counted how many of the prior week’s agent outputs required zero follow-up texts. The number started at 41 percent. After they added the post-print confirmation rule it reached 78 percent in six weeks. The generation count stayed roughly flat; the usable schedules increased.

Your Turn: Apply This Today

  • Pick one agent loop already running in your ministry tools and write the exact sentence that tells it when to stop and surface a human instead.
  • Add a single confirmation field that must be marked after the physical print or door-post step before the loop can close.
  • Replace the “schedules generated” number on your dashboard with a handoff completion rate for the past seven days.
  • Schedule a ten-minute review each Monday where you look only at the shifts that still needed manual rescue and note the exact missing escalation rule.
  • Write the escalation rule in plain language a volunteer coordinator could read out loud, then paste it into the agent prompt or workflow description.
  • Test the new rule on next week’s schedule and record whether the coordinator sent any after-9 p.m. texts for that shift.

Two earlier posts on this blog explored the same pattern from different angles. “The Niche Tool That Still Needs a Human Print Step” showed how digital outputs still require a physical checkpoint, and “The Goal Loop I Set That Ignored Sunday Validation” described what happens when the measurement stops at generation instead of handoff.

I consult with product leaders building agentic workflows for ministry teams on continuous discovery loops, escalation rule design, and handoff metrics. Let’s talk.

The Inner Skills I Let the Model Train Out of Me

Builder reclaiming judgment after delegating judgment to the model

I spent two years delegating judgment to the model, treating its output as the default first pass on every feature spec. The pattern has a name: automation bias, the habit of trusting a confident machine over your own checks. The assumption was simple: if the language model could surface the obvious paths and edge cases faster than my own notes, then my real work started after the first draft. That habit cost me the ability to see when a proposed flow would break for the exact users who had no one else to call.

The first time it showed up was on a children’s ministry scheduling tool. The model produced a clean volunteer assignment flow that looked efficient on paper. I shipped the review version without walking through the Sunday morning handoff myself. Two weeks later three small churches reported that the print step still required a second person to stand at the copier because the model had optimized for digital completion rates that did not exist in their rooms.

This is the foundational misread that causes product teams to treat judgment as an automatic output of scale rather than a muscle that must be exercised against real constraints. John Wesley’s three rules—do no harm, do good, attend to the means of grace—were never meant as slogans. They were weekly disciplines that forced leaders to examine whether their daily habits actually produced the character the movement needed when pressure hit.

The Rule That Reversed My Automation Bias

Wesley required class leaders to ask the same three questions of every participant each week. The questions were not abstract. They required concrete answers about money, speech, and actions that week. The discipline worked because it refused to let good intentions stand in for observable behavior.

Most agent loops in product work skip this step. The model generates a revised workflow, the PM accepts or tweaks it, and the next sprint begins. There is no recurring checkpoint that asks whether the change still does no harm to the volunteer who prints at 9:15 on Saturday night. Over months the absence of that checkpoint trains the team to accept model fluency as evidence of soundness.

The practical result shows up in error logs rather than user interviews. When the model hallucinates a step that never existed in the original paper process, the team has already lost the habit of catching it before release because nothing in the weekly rhythm forced them to rehearse the failure case.

Where Automation Bias Hides the Real Decision

Agent loops are built to optimize for task completion. They surface the next action, execute it, and report success. The hidden decision is almost never inside the loop; it sits one layer before the loop is allowed to start. That layer is the judgment about whether the task itself is worth automating for this specific user group.

On the curriculum side we once let an agent propose new lesson outlines based on engagement data. The loop ran cleanly and produced outlines that scored higher on predicted completion. Only after launch did we notice the outlines removed every hands-on activity that required a table or floor space—precisely the activities small churches without dedicated rooms relied on. The model had optimized for the metric we fed it; we had lost the weekly discipline of checking whether that metric still served the users who could not change their room.

The decision point was never inside the agent. It was the prior choice to let the agent define which activities counted as valuable. Wesley’s first rule would have required us to ask whether the change risked harm before the loop was permitted to run at all.

The Second-Order Cost When Small Churches Copy Enterprise Defaults

Psychologists describe a related effect as automation bias: people over-trust a confident machine and stop checking its work.

Enterprise teams can absorb model errors because they have staff who can override the system the next day. Small churches have one volunteer who prints on Thursday and no second person to catch the mistake on Sunday. When those churches adopt the same agent-generated defaults, the cost is not lower efficiency; it is visible failure in front of families who already feel marginal.

The pattern repeats across tools. A scheduling agent suggests automatic reminder timing based on large-church data. The small-church volunteer receives the reminder on the only evening they have child care and deletes the app. The model records lower engagement and suggests even more reminders. The loop compounds the original misread.

Wesley’s third rule—attend to the means of grace—translated into product terms means building recurring contact with the actual environment where the tool is used. Without that contact the inner skill of noticing when an optimization has become harmful atrophies, and the model is left to train on its own outputs.

Your Turn: Apply This Today

  • Pick one agent-generated recommendation from this week’s work and trace it backward to the single assumption the model was allowed to make without your override.
  • Log the date, the assumption, and the one-sentence reason you accepted it; keep the log in the same place you keep sprint notes so it cannot be ignored.
  • Once a week, before the next agent loop runs, state out loud or in writing the exact failure that would constitute harm for the smallest user segment you support.
  • Require that statement to be reviewed by one person who has used the tool in a room with no dedicated tech support.
  • At the end of the month, count how many logged overrides changed the shipped version; treat a count below two as evidence the discipline has weakened.
  • Schedule the next review on the calendar now rather than when the next agent output arrives.

The same pattern appears in the niche tools that still need a human print step and in the guardrail conversations that skip the question of whose judgment is actually being replaced.

I consult with product leaders and ministry tool teams on agent loop design, weekly override disciplines, and the second-order costs small-church users absorb when enterprise defaults ship unchanged. Let’s talk.

The Niche Tool That Still Needs a Human Print Step

Volunteer reviewing a printed schedule, the human print step

POST TITLE: The Niche Tool That Still Needs a Human Print Step

CURRENT HTML (full post):
Niche AI tools for ministry planning reduce overhead only when they force a visible handoff to a printed artifact instead of treating the digital output as complete. Most teams assume that smaller, focused models will simply slot into existing volunteer flows without extra steps. The result is the opposite: hidden friction surfaces exactly where the model stops and the physical page begins.

This assumption treats the tool as a closed system. In practice the tool produces a plan or curriculum outline that still requires a human to verify layout, count copies, and confirm timing against room capacity. The gap appears because product decisions rarely model the second loop—the physical verification that happens after the agent finishes.

Charlie Munger’s latticework of mental models requires holding several frames at once rather than optimizing inside one. Applied here, the relevant models are second-order effects, the map-territory distinction, and the value of redundant checkpoints. When teams keep only the software model active, they miss how the physical print step functions as the actual constraint on Sunday morning reliability.

Why low-overhead tools still collide with the human print step

Small-agent workflows promise faster iteration for children’s ministry planners. A focused model can generate a four-week series outline in minutes instead of hours. Yet that workflow still depends on a human print step: the output arrives as formatted text that must still be turned into paper copies for volunteers who will not open a second screen on-site.

The collision occurs because the model optimizes for content completeness, not for physical production time. A volunteer desk that receives the file at 8 p.m. on Thursday still needs twenty minutes to adjust margins, add page numbers, and run the correct number of copies. That interval is invisible inside the agent’s success metric but becomes the dominant variable on Saturday night.

Teams that measure only digital completion rate discover the mismatch only after the first service where the printed sheets are missing or misordered. The low-overhead product did exactly what it was asked to do; the missing model was the physical handoff duration.

The second-order cost when niche tools skip the volunteer desk check

Research on how people process information shows that readers skim screens but commit differently to paper, which is exactly why the desk check matters.

Skipping the explicit check at the volunteer desk creates downstream rework that compounds across multiple services. One missed layout adjustment means every small group receives an incomplete set, which then requires phone calls and last-minute reprints. The cost is not the extra paper; it is the eroded trust that the next digital plan will actually be usable without additional human labor.

Munger’s latticework highlights the second-order effect: the time saved by the agent is transferred to the least-resourced person in the chain—the volunteer who arrives thirty minutes before doors open. When that transfer is invisible in the product metrics, the tool quietly increases the cognitive load on the very users it claims to serve.

Observed cases show the pattern repeats across curriculum tools and sermon-prep agents alike. The model produces clean JSON or markdown; the printed packet still requires a human to confirm that the memory verse fits on one page and that the craft supply list matches the actual bin contents in the storage room.

Latticework moves that protect the human print step in the decision path

Teams that apply multiple models simultaneously insert one redundant checkpoint before the agent hands off. The checkpoint is a single printed sample page generated by the workflow itself, reviewed by the same volunteer who will run the room. This step uses the map-territory model to force the digital output to confront physical constraints early.

They also track a combined metric: digital completion plus physical verification time. The second number surfaces whether the agent’s speed gain is real or merely shifted. When verification time drops below a threshold, the product team knows the handoff model is working; when it rises, they adjust the agent prompt or output format rather than asking volunteers to absorb the difference.

These moves do not add bureaucracy. They add one visible decision node that the agent cannot bypass. The node forces the product to carry the physical constraint as an input rather than treating it as an after-the-fact cleanup task.

Your Turn: Apply This Today

  • Identify the single agent workflow your team runs most often for weekend materials and add a forced “print sample page” step that must be acknowledged before the final file is marked complete.
  • Assign one volunteer desk owner to record the actual minutes between receiving the file and finishing the printed packet for the next two weekends; share the raw numbers with the product owner without interpretation.
  • Modify the agent prompt to include a one-sentence physical constraint field (room capacity, copier tray size, volunteer arrival time) that the model must restate before generating the outline.
  • Run a two-week test where every digital completion is blocked until the printed sample receives a one-word approval (“usable”) from the desk owner; count the number of Saturday-night fixes before and after.
  • Create a shared log that records only the physical verification time and the number of copies actually needed; review the log in the next planning meeting instead of the agent’s internal success score.
  • Remove any auto-approval rule that lets the agent close the loop without the physical checkpoint; replace it with a simple status that stays open until the printed artifact is confirmed.

The Goal Loop I Set That Ignored Sunday Validation and The 8x Claim That Still Leaves the Print Step Alone both trace the same pattern of hidden handoffs that surface only after the model finishes.

I consult with product leaders building agentic tools for ministry workflows on keeping physical handoffs explicit in decision paths and measuring verification time alongside digital completion. Let’s talk.

The Question the Guardrail Conversation Always Skips

Printed card describing a guardrail reversal protocol

A ministry leader asked me last month how to keep an AI scheduling agent from assigning the same exhausted volunteer to three consecutive events without anyone noticing until the calendar went live. The answer is a guardrail reversal protocol: guardrails survive only when they are written as explicit reversal triggers a volunteer can execute without asking permission from the model.

Most teams treat guardrails as static rules the model must obey. That framing collapses the moment real ministry conditions appear—missing state data, a volunteer who cannot be reached, or an event that moved rooms. The guardrail becomes decoration rather than a working part of the loop.

The Meta regression that started with missing state checks

The pattern is familiar to anyone who has read postmortems on blameless incident response: the failure is rarely the model, it is the missing state check.

Teams often begin by listing constraints: no more than two events per weekend, no back-to-back children’s rooms, notify the coordinator if a name repeats. These lists look complete on a whiteboard. They fail in production because the agent never receives an up-to-date view of who actually showed up last week or who called in sick on Saturday night.

Teresa Torres’s continuous discovery work shows the problem clearly. Discovery is not a quarterly workshop; it is repeated contact with the people who will live with the output. When the only contact happens through the prompt, the model fills gaps with assumptions. The regression begins the first time the agent acts on stale volunteer availability data.

I watched one children’s ministry platform lose three reliable volunteers in a single month because the agent kept reassigning them based on an outdated signup sheet. The guardrail existed in the prompt. It never existed in the actual workflow the volunteer saw on her phone.

Why timer-based loops hide the moment your guardrail reversal protocol should fire

Many agent designs poll on a schedule—every four hours, every morning at 6 a.m.—and then push updates. The interval feels safe because it limits cost and noise. It also creates a blind window where a reversal trigger should have fired but cannot.

Consider a volunteer who texts at 9:17 p.m. that she cannot lead the 8 a.m. class. If the next loop runs at midnight, the system has already printed name tags and the children’s director has already adjusted the room plan. The timer did not protect the workflow; it delayed the moment anyone could act.

Torres emphasizes that discovery must happen at the frequency of the decision, not the frequency of the sprint. In volunteer scheduling that frequency is measured in hours, not days. Timer loops that ignore this create the exact condition they claim to prevent: silent failure until a human finally opens the dashboard.

Writing the guardrail reversal protocol that fits on a single printed card

The workable guardrail is not another rule inside the model. It is a short, printed sequence any volunteer can follow when the output looks wrong. The card states three things: the exact condition that triggers reversal, the single action required, and the person who receives the reversal notice. Nothing else.

One church printed the card on the back of the Sunday schedule. The condition read: “If your name appears twice in one day or three times in one week, text REVERSE to the coordinator number before you accept.” The action was one text. The notice went to the same coordinator who already handles last-minute swaps. The model never needed to interpret the reversal; it simply received an updated state from the coordinator’s reply.

This approach matches Torres’s insistence that opportunities and solutions must be tested with the people who will use them. The card is the test. If a volunteer cannot execute it in under thirty seconds while holding a stack of Bibles, the protocol is still too long.

Your Turn: Apply This Today

  • Pick the single agent loop that currently touches volunteer scheduling and list every state field it reads before making an assignment.
  • Write the reversal condition, the one action, and the recipient on one index card; keep the total under forty words.
  • Print five copies and give them to the actual volunteers who would receive the next automated schedule.
  • Observe what happens the first time one of them uses the card; note the exact time between spotting the problem and the state being corrected.
  • Update the agent’s input data source to include the reversal notice as a first-class event rather than an exception report.
  • Run the same loop again the following week and measure whether the reversal rate drops because the model now receives fresher state.

The reversal protocol is the missing piece in both “The Question Every Agent Mandate Still Leaves Unanswered” and “Letter to the PM Who Just Inherited the Agent Mandate.”

I consult with faith-tech product leaders and ministry tool builders on continuous discovery loops, reversal protocol design, and volunteer workflow state management. Let’s talk.

The Inner Game I Outsourced Before the First Volunteer Shift

Coordinator defining agent outcome checks before a model call

The cursor blinked after the final paste. She had dropped every lesson outline from the past three years into the same thread, added the volunteer notes from last quarter, and hit send. The printer in the corner started its warm-up whine while her eight-year-old tugged at the edge of the table asking about snack plans. She clicked generate, waited twelve seconds, then hit print without reading the full output.

The pages stacked warm in her hands. She folded the top sheet to fit the folder she would carry into the classroom the next morning. Nothing in the text told her whether the restless boys in the back row would actually sit through the activity or whether the story hook would survive once the first paper airplane flew. She closed the laptop anyway.

That moment showed the exact point where the inner discipline gets handed off. The model produced something usable on the surface, yet the coordinator still carried the real weight of deciding whether it would serve the actual children who would show up. Most product teams treat that handoff as solved once the output appears. It is not. The fix is a habit of agent outcome checks: deciding what good looks like before the model is ever called.

The Rule That Forces Agent Outcome Checks Before Model Calls

John Wesley boiled his movement’s daily practice down to three rules that had to be answered before any action: do no harm, do good, stay in love with God. The order mattered. Harm had to be ruled out first. Only then could the next two rules apply. The same sequence exposes what current agent workflows skip.

When a children’s ministry lead feeds three years of outlines into a model, the first rule requires naming the possible harm before generation starts. Harm here looks like a lesson that assumes every child can read at grade level or that ignores the one volunteer who always arrives ten minutes late. Running the check after the output lands wastes the moment the model could have been steered away from those edges. Teams that insert the harm question before the call produce fewer pages that still require heavy rewriting on Saturday night.

The second rule then narrows what counts as good. Good is not more content. It is the single activity the lead can finish in the seven minutes between unloading the car and unlocking the classroom door. The model rarely surfaces that constraint unless the prompt names the real validator and the real time budget first.

How Volunteer Handoffs Expose Missing Guardrails

This is the same instinct behind writing the test before the implementation: define the acceptable result first, then let the system try to meet it.

I once watched a coordinator print a generated lesson on a Thursday, then spend Saturday night cutting the activity into two smaller pieces because the original required six separate supplies. The volunteer who actually taught it never saw the original version. She only saw the version that fit the tote bag she could carry on the bus. The model had optimized for completeness, not for the handoff that actually happened.

Wesley’s rules surface here as a routing decision. The harm check asks whether the generated steps can survive being edited by someone who never met the model. The do-good check asks whether the final printed page still moves the needle on the one outcome the church set for that age group. When those questions sit inside the workflow instead of after it, the printed sheet carries fewer surprises into the room.

The same pattern appears in larger platforms. A feature that auto-suggests small group questions often ships without asking which volunteer will read the suggestion aloud and which child might push back on the wording. The missing guardrail is not more model calls. It is the named human who must sign off on the outcome before the file reaches the printer.

Building Agent Outcome Checks Into a Daily Practice That Survives Model Changes

Models update. The interface changes. The prompt that worked last quarter fails when the safety layer shifts. Wesley’s third rule, stay in love with God, points to the habit that outlasts any single tool. It is the repeated act of naming the actual person who will use the output and the actual Sunday moment that will test it.

Product teams that build this habit keep a running log of every model-assisted task. Each entry records the named outcome for that Sunday and the named person who will validate it before the lesson reaches the classroom. When the model changes, the log still tells the team which tasks moved the needle and which ones only moved pixels. The practice does not require new tooling. It requires the same three questions asked at the start of every session rather than the end.

The discipline compounds. After four weeks the log shows which prompts consistently produce lessons that survive the seven-minute volunteer window. After eight weeks the team can retire prompts that only look impressive in the chat window. The model can be swapped without losing the thread that connects the output to the children who will sit on the carpet.

Your Turn: Apply This Today

  • Pick one recurring model-assisted task this week and write the exact Sunday outcome it must serve before you open the chat window.
  • Name the single human who will validate that outcome on paper or screen before the file leaves your desk.
  • Log the task, the outcome, and the validator in a simple note each time you use the model; keep the note next to the printed result.
  • At the end of the week, mark which entries still match what actually happened on Sunday and which ones required changes after the model output.
  • Delete or rewrite any prompt that produced pages the validator had to rewrite for more than five minutes.
  • Run the same log for a second week with the revised prompts and compare the validator time required.

The Goal Loop I Set That Ignored Sunday Validation shows what happens when the log never reaches the person who actually teaches. The 8x Claim That Still Leaves the Print Step Alone tracks how generation speed hides the real cost that appears only after the pages reach the volunteer.

I consult with product leaders and ministry teams on AI routing decisions, volunteer handoff constraints, and outcome logging practices that survive model changes. Let’s talk.

The Goal Loop I Set That Ignored Sunday Validation

Sunday validation

The goal loop I built ignored Sunday validation, and that omission hid the real failure for months. I spent months building a goal loop for a children’s ministry curriculum tool that measured success by how cleanly a volunteer completed the lesson prep workflow inside the app. The loop closed when the PDF downloaded and the checklist hit 100 percent. I treated that as the finish line because the data looked strong and the retention curves held steady in the first weeks.

What I missed was that the real test happened after the download. The loop never checked whether the printed materials made it to the classroom table or whether the 7-minute volunteer could adapt the content on Sunday morning without a second device. The cost showed up in quiet churn: teams kept the account active but stopped using the new features because the Sunday handoff kept failing in ways the metrics never surfaced.

This is the foundational misread that causes product teams to optimize agent loops for task closure instead of mission outcome. Teresa Torres’s continuous discovery framework exposes the flaw directly. Her method insists that teams interview users in the context of the actual work, not after the digital step is done, so the success criteria stay tied to the observable result the user needs.

The Loop That Shipped Clean but Never Reached the Volunteer Desk

The discipline of validating against real outcomes echoes Nielsen Norman Group guidance: a clean internal result means nothing until it is tested where the work actually happens.

The curriculum tool’s agent watched for three signals: content selected, notes added, and file exported. Once those three steps occurred inside the same session, the loop scored the interaction as successful and moved the user into a “completed” cohort. We celebrated the numbers because they looked better than the old manual process.

Yet the exported file often sat in an email inbox or downloads folder until the volunteer had five minutes on Saturday night. At that point the layout no longer matched what the classroom needed, and the agent had already marked the job finished. The loop had no way to register that the materials never made it to the table.

Continuous discovery would have caught this earlier. Torres’s practice of weekly interviews in the actual work environment forces the team to watch the printout travel from screen to printer to classroom. Without that observation, the success criteria stayed inside the app.

How Continuous Discovery Forces Sunday Validation into the Criteria

Torres teaches that opportunity solution trees must be updated against real customer moments, not against internal task definitions. Applied to ministry agents, that means every goal loop needs an explicit Sunday handoff node. The node records whether the volunteer opened the printed page at the right time and could find the next activity without searching.

In practice this changes the data collected. Instead of only tracking export completion, the loop now logs whether the volunteer marked the lesson “used” on Sunday morning and whether any adaptation notes were added afterward. Those two signals sit downstream of the digital steps and become the real stopping condition.

The shift also changes what the agent is allowed to optimize. It can no longer declare victory on a clean PDF. It must wait for evidence that the content survived the handoff and served the actual teaching moment. That requirement surfaces new questions about print layout, timing reminders, and simple fallback instructions that the original loop never considered.

Why Timer Loops Hid the Missing Sunday Validation for Months

Many agent loops default to time-based success because elapsed time is easy to measure. Our loop rewarded any prep session finished inside forty-eight hours of the scheduled class. The timer created the appearance of reliability even when the materials arrived too late for meaningful review.

The hidden cost was delayed feedback. Volunteers who struggled on Sunday rarely reported the problem inside the app, so the timer kept scoring the prior week as successful. Months passed before usage data showed that the high completion rate was not translating into repeat classroom use.

Continuous discovery interrupts this pattern by requiring teams to watch the outcome moment repeatedly. When the interview or observation happens at the Sunday table, the timer loses its authority. The loop must now track the observable sign that the handoff succeeded, which a simple elapsed-time rule cannot provide.

Your Turn: Apply This Today

  • Pick one existing agent loop in your current product and add a single Sunday handoff signal that only fires when the volunteer confirms the materials reached the classroom table.
  • Schedule two 15-minute calls this week with volunteers who used the tool last Sunday and ask them to show you exactly what they had in their hands at 9:45 a.m.
  • Rewrite the loop’s success definition so it requires both the digital completion and the Sunday confirmation before the agent records the interaction as done.
  • Remove any time-based reward that currently closes the loop and replace it with the observable handoff event.
  • Test the revised loop on three active ministry accounts and log whether the new signal changes which users are counted as successful.
  • Share the before-and-after success rates with your team in a single paragraph that names the exact Sunday moment being measured.

The Workflow Handoff Where Models Finally Earned Their Keep and The Question Every Agent Mandate Still Leaves Unanswered both explore what happens when outcome criteria stay anchored to the final human moment rather than the digital step.

I consult with ministry product teams on agent goal loops and Sunday outcome criteria. Let’s talk.

The 8x Claim That Still Leaves the Print Step Alone

Print step handoff

The print step handoff is where the 8x claim quietly breaks down. Teams tracking AI model deployment in content pipelines report an average 8x lift in output volume within the first sixty days. The number shows up in dashboards that count generated outlines, rewritten sections, and draft emails. What the dashboards omit is the unchanged Saturday-night print-and-fold step that still consumes the same volunteer hours it did before any model was installed.

The obvious reading treats the 8x figure as proof that the models scaled the entire workflow. The opposite reading is the accurate one. The measured gains stop at the digital handoff. Everything after that handoff continues to run on the same physical constraints that existed in 2019.

Where the 8x number was actually captured

Analysis from McKinsey on AI adoption shows productivity gains concentrate upstream, leaving the final handoff steps untouched unless they are designed in.

The 8x claim surfaces most often inside the content-creation layer. A team ships a prompt chain that turns sermon notes into children’s ministry scripts, small-group questions, and social posts in one pass. The internal metric tallies those artifacts and divides by the hours the prompt engineer spent. The result looks dramatic because the old baseline counted every manual rewrite.

That baseline never included the next step: exporting the final PDF, loading the copier, and collating sets for volunteers who arrive thirty minutes before the first service. Those tasks remain outside the model loop by design. The 8x number therefore records only the portion of work that already lived inside a computer.

When the same teams later audit total cycle time from idea to printed sheet in a volunteer’s hand, the multiplier drops below 2x. The compression happened entirely upstream of the printer tray.

The print step handoff that stayed outside the model loop

Ministry resource sites have long optimized for the seven-minute volunteer. The person who prints the lesson at home or at the church office still needs a single-sided, correctly collated packet that fits in a folder without extra staples. Models that generate beautiful digital layouts rarely touch the printer driver settings or the paper-size defaults that break that packet.

One children’s curriculum platform tracked volunteer completion rates for three years. The rate held steady at 71 percent even after the organization introduced an AI-assisted script generator. The only variable that moved the needle was the addition of a one-click “print packet exactly as last week” button that bypassed the new digital workflow entirely.

The print step functions as an unmeasured dependency. Because it sits after the model output, teams can celebrate velocity gains while the actual delivery bottleneck stays fixed. The metric hides in plain sight because it lives on a physical machine rather than in the SaaS dashboard.

Latticework that connects engineering output to the print step handoff

Charlie Munger described a latticework of mental models as the habit of pulling ideas from multiple disciplines to see where they intersect. In this case the relevant models are queueing theory from operations research and the doctrine of total depravity from theology. Both insist that friction does not disappear just because one segment of the chain improves.

Queueing theory shows that the slowest station determines throughput. If the print-and-collate station still requires the same labor minutes, upstream acceleration simply creates a larger backlog in front of that station. The model output piles up as unprinted PDFs.

Total depravity supplies the corresponding human observation: people default to the path of least resistance. When the new digital pipeline adds even one extra click or file-format conversion, the volunteer reverts to the old printed master that has worked for years. The latticework therefore predicts that isolated model gains will be absorbed rather than multiplied unless the physical handoff itself is instrumented.

Teams that applied both models together began logging the timestamp when the final PDF reached the copier and when the last packet left the building. Those two timestamps revealed that the 8x digital gain translated into a 1.3x end-to-end gain. The difference was not a failure of the model; it was the predictable result of measuring only the segment the model touched.

Your Turn: Apply This Today

  • Add a single timestamp field in your workflow tool that records when the final approved file is sent to the printer or copier queue.
  • Run a seven-day audit on one recurring resource: note the exact minute the print job starts and the minute the last collated set is placed in the volunteer bin.
  • Compare that elapsed time against the digital generation time you already track; surface the ratio in the same dashboard that currently shows the 8x claim.
  • Identify the one volunteer who performs the print-and-fold step and ask them to log any extra clicks or reformatting required by the new AI output.
  • Set a target to reduce the print-to-bin interval by fifteen percent this month without changing the digital generation speed.
  • Share the updated ratio with the engineering team in the next sprint review so the next model improvement is scoped against the physical constraint rather than the digital one.

The Workflow Handoff Where Models Finally Earned Their Keep and The 10/10 Rule That Still Gets Skipped for the Novel Agent both trace similar gaps between model output and actual delivery.

I consult with product leaders shipping AI workflows for ministry teams on measuring end-to-end handoffs and volunteer completion rates. Let’s talk.

Letter to the PM Who Just Inherited the Agent Mandate

Inherited agent mandate
Dear PM who just inherited the agent mandate at a denomination you still can’t name without checking the org chart, You opened the first planning doc and the words “autonomous agents” and “workflow automation” sat there like they belonged. The team expects you to turn that mandate into something that ships before Q4. You know the real users are volunteers who already give up their only free hour on Sunday, but the spec keeps drifting toward model capabilities instead of what actually leaves their shoulders. This is the same pattern that turns every new AI initiative into another dashboard nobody opens after month two. The foundational misread is treating the agent as a set of capabilities rather than a single, named emotional relief the user feels the moment the task disappears. Shreyas Doshi’s discipline of naming the one-word core offering forces that clarity before any routing logic gets written. When you lock the word first, every later decision stops floating.

The One-Word Relief the Volunteer Actually Buys

Continuous discovery work, as Teresa Torres describes it, starts by naming the outcome a user actually buys—not the task the system completes.

If you just inherited the agent mandate, start here. Most ministry agents still optimize for task completion. The volunteer who prints curriculum at 9:15 p.m. on Saturday does not buy task completion. She buys relief from the low-grade dread that she might stand in front of eight-year-olds unprepared.

That relief has one word: ready. When the agent returns the lesson plan, the craft list, and the three backup activities already formatted for the printer, the word she feels is ready. Everything else is noise.

You see the difference in the usage data from tools like Sermons4Kids. Completion rates jump when the final artifact lands in a single printable packet. They flatline when the agent asks her to choose between three versions or reformat anything herself. The word ready explains why.

How an Inherited Agent Mandate Changes Every Routing Decision

Once ready is fixed, the model no longer needs to demonstrate creativity. It needs to demonstrate absence of friction. That single constraint kills most of the clever multi-step agent loops that look impressive in demos.

Routing now favors the shortest path that produces a print-ready file over the path that gathers extra context. If the volunteer has already told the system her age group and room size in a previous week, the agent should not ask again. Ready means no new questions on Sunday morning.

The same word also dictates what the agent must refuse. It must refuse to generate extra options, extra illustrations, or extra discussion questions. Those extras feel like generosity until the volunteer is standing at the copier with six minutes left. Ready removes them by design.

The Prioritization Rule for an Inherited Agent Mandate

Any roadmap item that lengthens the time between request and printed artifact loses priority. Any item that reduces that time gains it, even if it looks less intelligent.

Apply the rule to the three agent features still sitting in your backlog. The one that adds voice input for last-minute changes stays only if it still produces the same printable packet in under four minutes. The one that lets the volunteer edit the lesson in a browser tab gets cut, because editing breaks the feeling of ready.

The rule also surfaces the real metric. Track the percentage of weeks where the volunteer downloads and prints without opening any other screen. That number tells you whether the core offering is landing, not model accuracy scores.

Your Turn: Owning the Inherited Agent Mandate Today

  • Open the current agent spec and write the single word that describes the volunteer’s relief; keep it to one word only.
  • Take the top three open roadmap items and score each one against whether it shortens or lengthens the path to that relief.
  • Delete or defer the item that adds any new screen, choice, or confirmation step before the printable output appears.
  • Find the last three weeks of volunteer usage data and count how many sessions ended with a single download and print action.
  • Write the new acceptance criteria for the next agent build using only that one-word relief as the success condition.
  • Share the revised one-page spec with the two engineers who will actually implement the routing changes this sprint.

The Single Feeling Most Ministry Tools Still Refuse to Name and The Question Every Agent Mandate Still Leaves Unanswered both trace the same pattern back to the moment the core offering gets named or skipped.

I consult with product leaders on naming one-word core offerings for agentic ministry tools, routing decisions that protect volunteer time, and prioritization rules that survive handoff to volunteer users. Let’s talk.

The Kitchen Counter Where Boldness Beat Process

Person hesitating over a bold decision and the cost of being wrong

Lowering the cost of being wrong is how boldness beats process. She stood at the counter with peanut butter on the knife and the notification still unsent. One thumb hovered over the screen while the other hand reached for the bread. The families were already in the system. The event details were ready. Only the approval chain stood between her and the send button.

She sent it anyway.

Ten minutes later the replies started coming in, and the director was still refreshing an empty inbox.

The senior product manager who later reviewed the same flow would have opened a three-week discovery track. She would have requested usage data from three prior campaigns, scheduled interviews with six parents across two time zones, and produced a one-page brief for legal. The document would have listed risks, success metrics, and rollback criteria. By the time the brief reached the volunteer’s team, the event date would have passed.

The volunteer’s test produced three concrete signals in under an hour. The message needed a clearer time stamp. Families wanted the option to confirm rather than just receive. The open rate on the first send predicted the second send almost exactly. None of those signals required a brief. They required only the ability to change one variable and watch what happened next.

AI tools now collapse the cost of that first change. A notification can be rewritten, segmented, and scheduled inside the same interface the volunteer already uses for attendance. The model suggests variants based on past open rates without requiring a data analyst to build a query. The experiment stays small enough that a single person can own both the change and the result. When the cost of being wrong drops that low, the circle of competence stops being a fence and starts being a starting line.

Teams still default to process because process feels like protection. A documented approval chain limits personal exposure when the test fails. The volunteer who sent the message carried the full risk herself. If open rates had collapsed, she would have been the one explaining the choice at the next meeting. Experienced operators learn to avoid that exposure. They route every decision through layers that distribute blame if the outcome is poor.

The cost is not only time. It is the quiet removal of people whose competence has not yet been recognized by the chart. The twenty-three-year-old volunteer never asked permission because she did not yet know the permission existed. Once she learns the process, the quick test becomes someone else’s job. The signal disappears with it.

The Experiment That Never Reached the Roadmap

Research summarized by Harvard Business Review shows that organizations which lower the cost of intelligent failure learn faster than those that punish it.

The senior PM later described the notification flow as “high risk” because it touched families who had already opted out of two previous messages. Her plan required a controlled rollout to ten percent of the list, a two-week observation window, and a decision gate. The volunteer had already run the equivalent test on twelve families and seen the replies in real time. The difference was not data quality. It was ownership of the outcome before any committee met.

Lowering the Cost of Being Wrong Early

AI makes the first version cheap enough that the volunteer can afford to be imprecise. She can generate three subject lines, send them to staggered groups, and delete the weakest performer before most recipients notice. The same model can surface the exact phrasing that produced the three confirmation replies. None of this replaces judgment. It simply moves the first judgment earlier, before the circle of competence has time to close around it.

Why a Low Cost of Being Wrong Protects the Nerve

The director eventually approved the final message, but only after the volunteer showed the open-rate numbers from her own tests. The process remained on the books. What changed was the willingness to let one person stand outside it long enough to gather evidence the process itself could not produce. That willingness does not scale through new policy. It scales through deliberate assignment of small, unfiltered tests to people whose competence is still forming.

Your Turn: Apply This Today

  • Pick one junior team member or volunteer who has never owned a live notification or feature flag and give them a 48-hour window to change one message or timing variable on a real audience segment of at least fifty users.
  • Remove the requirement for pre-approval on that single change; require only a short log of what was sent and the first three replies or open-rate numbers within 24 hours of send.
  • Set a hard stop at the end of the window where the test either rolls back automatically or moves to the next scheduled send with no further review unless metrics drop below a pre-agreed threshold.
  • Ask the same person to write a two-sentence summary of what they would change next time, then schedule a ten-minute review with only the direct manager present.
  • Repeat the assignment with a different junior person the following week so the pattern becomes visible to the rest of the team before any new process document appears.
  • Track how many of these micro-tests produce a measurable lift that later appears in the official roadmap; note the elapsed time between the first test and the roadmap entry.

The Trust Layer the Roadmap Still Treats as Optional and The Volunteer Desk Where the Trust Toggle Appeared both trace the same pattern of early experiments that later shaped larger decisions.

I consult with product leaders and ministry technology teams on running low-cost AI experiments, protecting early judgment outside formal process, and measuring what actually moves volunteer completion. Let’s talk.

The Question Every Agent Mandate Still Leaves Unanswered

Diagram of agent trust constraints in an AI workflow

Agent trust constraints are the question every agent mandate still leaves unanswered. You’re at your desk when the Slack thread updates itself. The agent has already moved the meeting, pinged the volunteers, and attached a revised agenda. Nothing looks wrong. But nothing asked you either.

That pause before you type “looks good” is the real problem. The decision to trust happened after the action, not inside it.

Most mandates treat that first move as neutral ground. It isn’t. The moment the agent steps forward is where the relationship either holds or starts to thin.

The Follow-Up That Booked Itself

Guidance from the NIST AI Risk Management Framework treats trust boundaries as a design-time requirement, not a setting bolted on after launch.

A children’s ministry director described an agent that scanned the calendar and offered a make-up session to a family that had missed small group. The agent sent the message, chose a time that fit the family’s pattern, and added it to the leader’s schedule. The leader found out when the confirmation landed in her inbox. The family was grateful. The leader felt sidelined.

The agent had done nothing technically wrong. It had optimized for attendance and convenience. What it lacked was any test against the first of Wesley’s rules at the moment it acted. No check existed to confirm whether the family wanted the outreach or whether the leader had already decided this was a case where presence mattered more than completion. The harm was not in the booking; it was in the removal of the leader’s judgment from the loop.

The same pattern appears when an agent reschedules a pastoral visit or flags a giving anomaly without first surfacing the context the pastor already holds. The cost is not lost efficiency. It is lost confidence that the tool still serves the person who carries the relationship.

Wesley’s Rules Translated to Agent Trust Constraints

Do no harm becomes a requirement that the agent must surface any action that could affect a person’s standing or privacy before the action commits. In practice this means the handoff packet must include the exact data the agent used and the exact person who would feel the effect. The leader either approves or supplies the missing context in one step. No silent execution.

Do good becomes a narrower test: the agent must state what measurable good it believes will result and how that good will be verified within seven days. If the verification step cannot be named at handoff, the action does not proceed. This rule stops agents from optimizing for activity metrics that never connect back to the actual health of a volunteer or a small group.

Stay in love with God is the rule most often ignored in product work. It requires the agent to preserve the conditions under which the leader can still exercise care. That means the agent cannot consume so much of the leader’s attention that the leader loses capacity for the people the agent was meant to serve. The constraint shows up as a hard limit on the number of proactive messages the agent may initiate in a week before it must pause and ask whether the leader still wants the volume.

These three constraints are not added later as policy. They are encoded in the handoff object the agent must construct before it takes any step that touches another person.

The Real Cost of Treating Agent Trust Constraints as a Later Toggle

When teams push the trust decision downstream, they create two problems that compound. First, the leader develops a habit of checking the agent’s work after the fact rather than shaping it at the point of delegation. Second, the agent learns that its initiative will usually stand, so it optimizes for speed and volume. The trust layer becomes a review step instead of a design constraint.

Over time the leader either turns the agent off for anything that matters or begins to treat every agent action as provisional. Neither outcome delivers the autonomy the mandate promised. The agent mandate flattens the very judgment it was meant to extend.

The pattern is visible in teams that moved from simple reminder agents to proactive scheduling agents without rebuilding the handoff. The reminder agent succeeded because its output was always visible and reversible. The scheduling agent failed because its output was visible only after it had already altered someone else’s calendar. The difference was not the model. It was the absence of a Wesley-style test at the moment the agent decided to act.

Your Turn: Apply This Today

  • Pick one proactive workflow your agent already runs and write the exact three-line handoff message the agent must send before it acts; include the data used, the person affected, and the verification step that will close the loop within seven days.
  • Build that message into the agent’s execution path so the workflow pauses until the leader replies or the seven-day window expires.
  • Set the agent’s weekly cap at the number of proactive handoffs the leader can realistically review without losing capacity for direct ministry; start with five and adjust after two weeks of real usage.
  • Log every declined or revised handoff for thirty days and review the reasons with the team that owns the agent; look for patterns where the constraint itself needs tightening.
  • Remove any dashboard that shows completed agent actions without also showing the handoff messages that preceded them; the leader should never have to hunt for context after the fact.
  • Run the same three Wesley tests against the next new agent feature before any engineering work begins; if the feature cannot satisfy all three at the handoff point, do not build it.

The Trust Layer the Roadmap Still Treats as Optional and The Mandate That Flattened the Autonomy They Needed both trace the same pattern from different angles.

I consult with product leaders and ministry technology teams on agent handoff design, constraint definition, and workflow ownership. Let’s talk.

The 10/10 Rule That Still Gets Skipped for the Novel Agent

Product manager running a novel agent test against a discovery log

The novel agent test is the 10/10 rule teams still skip. I sat with the support log at 11:40 p.m. and watched the same three users hit the first screen, stall, then close the tab. We had just shipped the agent that would “handle everything after intake.” No one had checked whether intake itself still worked.

The demo video looked clean because we tested it on our own devices with our own data. Real people on real phones never reached the clever part.

That gap keeps showing up whenever the new thing feels more interesting than the thing we already shipped.

The discovery log that never showed the real handoff

Continuous discovery practice, as described by Teresa Torres, depends on testing the riskiest assumption first rather than the most novel feature.

Most discovery logs I see record interviews with ministry volunteers who already know the product. The notes list feature requests and pain points, yet they rarely capture the moment the volunteer steps away from the screen to finish the actual task. In one children’s ministry platform the handoff happened after the lesson plan was printed: the volunteer then spent seven minutes locating physical supplies that the digital workflow assumed were already organized. That step never made it into the opportunity solution tree because the interview ended at the print button.

The gap matters once an agent is introduced to “optimize” the plan. The agent can generate supply lists faster, but the underlying organization problem remains untouched. Continuous discovery requires watching the full loop, not stopping at the digital artifact. Without that observation, the agent multiplies output that still requires the same manual rescue work.

Teams that close the loop discover the handoff is often where trust erodes. A pastor may accept AI-generated talking points yet still retype them because the source material does not match the translation or reading level used on Sunday. The log shows acceptance; the real workflow shows revision.

Where the novel agent test exposes the gap

The 10/10 rule is simple to state and difficult to satisfy: run the current workflow with ten consecutive target users and require every one of them to complete it without hesitation, confusion, or workarounds. Most teams test for usability issues inside the interface. They rarely test for whether the interface itself is the right place to solve the job. When the test is applied to sermon or lesson preparation tools, the pass rate drops once the user must move content into a different format for delivery.

An agent that auto-generates slides or small-group questions looks impressive in a demo. The same agent fails the 10/10 test when the tenth volunteer prints the material, discovers mismatched scripture references, and opens a separate Bible app to correct them. The novel capability did not address the translation or formatting layer that sits between the product and the actual Sunday use.

Teams that reach 10/10 on the base workflow before adding agents report fewer downstream revisions. The agent then operates on a verified foundation rather than amplifying an incomplete one. The data on agent adoption therefore misleads when it counts features shipped instead of workflows cleared.

Why the novel agent test arrives too early

Novelty bias inside product organizations rewards the visible addition over the invisible cleanup. Roadmaps list “AI agent for volunteer matching” or “AI sermon illustration generator” because those items read as innovation to stakeholders. The prior steps—standardizing supply lists, aligning translations, confirming delivery formats—stay invisible because they do not produce new screenshots.

Continuous discovery counters this bias by requiring the team to keep interviewing until the opportunity is solved for the current users, not the imagined next set. When the 10/10 bar is met first, the agent can be evaluated against a stable baseline. When the bar is skipped, the agent becomes another variable that must be debugged by the same volunteers who were already patching the original flow.

The pattern repeats across curriculum platforms and resource sites. The teams that later remove the agent do so after realizing the core workflow still required manual reconciliation. The teams that keep the agent are the ones that first drove the manual version to consistent completion. The difference is not technical sophistication; it is the order of validation.

Your Turn: Apply This Today

  • Pick one ministry workflow that currently ends with a handoff to physical materials or another app; schedule three 20-minute observation sessions this week with users completing the full loop, not just the screen portion.
  • Run the 10/10 test on that workflow before any agent scoping begins; document every point where the tenth user still improvises.
  • Update the opportunity solution tree to include only problems that appeared in at least seven of the ten sessions rather than feature requests gathered in meetings.
  • Delay the next agent-related ticket on the roadmap until the base workflow passes 10/10; move the agent work to a separate discovery track that starts only after the baseline is cleared.
  • Share the revised tree and the 10/10 results in the next roadmap review so stakeholders see the validation order rather than the novelty list.
  • Repeat the 10/10 test one month after any agent ships to confirm the addition did not reintroduce friction at the handoff points previously cleared.

The same pattern of validating the base before layering new capability appears in The Trust Layer the Roadmap Still Treats as Optional and The Mandate That Arrived Before the Workflow Existed. I consult with product leaders on applying continuous discovery to ministry workflows and sequencing AI agents only after core tasks reach unqualified user completion. Let’s talk.

The Workflow Handoff Where Models Finally Earned Their Keep

Team mapping a workflow handoff between people and AI models
The children’s director’s cursor blinked on the second form as she retyped the same three family names from memory, the spreadsheet still open on her left monitor. Steam rose from the coffee she’d poured twenty minutes earlier. Outside the office door, the first small-group leaders were already unlocking the kids’ wing at 6:52 a.m. She had tried the new model the night before. It produced clean rosters in seconds, then promptly invented two children who did not exist and dropped the one who needed an allergy note. She reverted to manual copy-paste because the output could not be trusted in the six minutes she actually had. This is the exact moment most AI tooling for ministry still fails. The model generates content faster than a human can type, yet it never arrives inside the handoff that determines whether the volunteer finishes before the first child walks in. The gap is not speed. It is the absence of a daily rule that decides when the model is allowed to touch the work at all. John Wesley’s three rules were never meant as inspirational slogans. They were operating constraints for people running class meetings with limited time and high stakes: do no harm, do good, attend to the ordinances. When ministry teams adopt models without equivalent constraints, they optimize for visible output while the actual handoff—the place where responsibility transfers to a real person—remains brittle.

The Monday Routing Decision

Work by Harvard Business Review on operationalizing AI underscores that the value shows up at the handoff, not in the model itself.

Every workflow that touches Sunday morning contains one decision point where the model either receives a fixed assignment or it does not. The assignment is not a prompt. It is a rule that states exactly which data the model may touch and which data it may never see. Last month the same children’s director tried letting the model draft the entire small-group list. It produced names faster than she could review. Two families received the wrong meeting time because the model had merged two similar last names from the previous quarter. The correction took eleven minutes she did not have. The Monday routing decision fixes the model to one narrow task—extracting first names and grade levels from the master spreadsheet—and nothing else. The rest of the roster stays in human hands. The rule is written down before any prompt is run. If the output violates the rule, the model is not used that week.

The Workflow Handoff That Actually Mattered

The real value appears only after the model finishes. The children’s director now receives a single clean column of names in the attendance app. She checks it once, adds the allergy note from memory, and hits save. The entire step takes ninety seconds instead of seven minutes. What changed was not the model’s capability. What changed was the handoff rule. The model is permitted to write the first draft only when the human has already confirmed the source data in the spreadsheet. The model is forbidden from inventing fields that do not exist in that source. The director’s job is no longer typing; it is verification against a known standard. Without that explicit handoff, the model drifts into background assistance. It offers suggestions that look plausible and still require the same manual cleanup as before. The volunteer experiences no reduction in cognitive load because the verification step was never removed.

What Breaks in the Workflow Handoff When the Model Stays Hidden

When the model is treated as an optional helper rather than a constrained worker, two failures compound. First, the volunteer never learns the boundary of what the model is allowed to touch, so every week carries the same verification burden. Second, the team loses any shared definition of acceptable output. One week the roster is clean, the next week it contains invented names, and no one can say why. Wesley’s rules prevent both problems by making the constraint visible and repeatable. The model either receives its fixed assignment on Monday or it does not run. The handoff either meets the standard or the human redoes the step. There is no middle ground where the model lingers in the background offering unaccountable help. Teams that skip this discipline discover the same pattern the director described: faster generation followed by identical or greater correction time. The model has not earned its keep until the handoff itself changes.

Your Turn: Apply This Today

  • Pick one observable ministry workflow that repeats every week and write the exact Monday routing rule on a single index card before any model runs.
  • Assign the model one narrow output only—names and grades from the master list, for example—and block it from adding any field that does not already exist in the source data.
  • Define the reversal trigger in the same rule: if the output contains any invented entry, the model is turned off for that workflow until the source data is re-verified by a human.
  • Run the fixed assignment once this week and time only the verification step after the model finishes; record whether the handoff takes less than two minutes.
  • If the verification still exceeds two minutes, adjust the routing rule before the next cycle rather than adding more prompt instructions.
  • Share the written rule with the one person who owns the final handoff so the constraint is visible to both the model and the human.
The same pattern appears in the volunteer desk where responsibility actually transfers and in the token budgets that never connect back to Sunday outcomes. Both posts trace how unconstrained model use creates invisible work that only surfaces when the first child arrives.

The workflow handoff is where models finally earned their keep. I consult with ministry product leaders and church technology teams on workflow handoffs, model routing rules, and verification steps that survive real Sunday morning constraints. Let’s talk.

The Single Feeling Most Ministry Tools Still Refuse to Name

Church volunteer holding a printed list with volunteer certainty

Volunteer certainty is the feeling most ministry tools still refuse to name. Ministry tools keep adding features that promise efficiency while ignoring the one feeling volunteers need most on Sunday morning: the quiet certainty that everything is already in place. This is not a minor oversight. It is the reason so many platforms see high signups and low weekly return.

Feature checklists reward visible progress. They do not reward the moment a volunteer opens an app and knows the printed list is already correct. Charlie Munger’s latticework of mental models shows why this gap persists. When a product team draws only from engineering and growth models, it misses the simple incentive structure that governs real use: people return to tools that remove a specific dread, not tools that offer more options.

Volunteer certainty starts with the list in her hand

A children’s ministry volunteer arrives at 8:15 with two kids in tow and a printed curriculum packet she needs to teach at 9:00. She does not need another dashboard. She needs the packet to match the one her coordinator emailed on Thursday. When the app instead surfaces suggested activities or AI-generated discussion prompts, it adds decision load at the exact moment she has none to spare.

Teams often describe this as a content problem. It is not. It is an outcome problem. The volunteer does not measure success by how many resources the platform contains. She measures it by whether she can hand the coordinator a completed checklist without hunting through menus. Munger’s latticework reminds us to pull the model from operations management here: the last responsible moment for error detection is the moment before the volunteer leaves the house, not the moment she opens the lesson.

Most roadmaps still optimize for the first model. They track page views, feature adoption, and session length. None of those numbers capture whether the single required artifact arrived intact and on time. The result is a product that looks impressive in a demo and fails on Saturday night.

Why outcome statements collapse under agent defaults

Nielsen Norman Group research on user confidence shows that perceived certainty, not raw efficiency, drives whether people trust and keep using a tool.

When teams write outcome statements, they usually begin with the user. They rarely test whether the system still protects that outcome once automated agents or recommendation engines are added. Munger’s approach requires holding multiple models at once: the user model, the agent model, and the error model. Most faith-tech roadmaps drop the error model.

An agent default might reorder lessons by engagement score or insert an extra activity that looks helpful. To the system this counts as personalization. To the volunteer it breaks the one guarantee she needed: that the printed list matches the coordinator’s version exactly. The outcome statement still exists on the product brief, but the live system no longer serves it.

The collapse happens because no one added a constraint that says “do not alter the artifact after the coordinator signs off.” Without that constraint, the agent optimizes for a different objective. Munger would call this a failure of inversion: the team asked what the volunteer wants instead of asking what would make the volunteer discard the tool.

The override that protects volunteer certainty

The fix is not more features. It is a deliberate override that treats the agreed artifact as non-negotiable once the coordinator has approved it. This override can be as simple as a locked export state that agents cannot touch and that surfaces a clear “this version is final” indicator to the volunteer.

Teams resist this because it feels like limiting the product. In reality it is the only way to make the product reliable for the person whose time is scarcest. The volunteer does not need optionality on Sunday morning. She needs the system to honor the prior agreement without introducing new variables.

Munger’s latticework makes the cost visible. When engineering, growth, and user models are held together, the override is not a restriction. It is the mechanism that keeps the lattice from pulling in contradictory directions. Without it, the product keeps optimizing for metrics that the actual user never sees.

Your Turn: Apply This Today

  • Pick one existing workflow that ends with a volunteer receiving a printed or digital artifact and write the single-sentence outcome that volunteer needs to feel certain the artifact is correct.
  • Add a logging step that records whether the artifact delivered to the volunteer matches the version the coordinator last approved, using a simple timestamp comparison.
  • Identify the first agent or recommendation rule that could change that artifact after approval and write an explicit constraint that blocks the change.
  • Run the constraint against the last thirty days of logged deliveries and count how many times the artifact would have been altered.
  • Share the single-sentence outcome and the mismatch count with the two people who own the coordinator and volunteer views of the product.
  • Schedule a thirty-minute review next week to decide whether the override needs to become a permanent rule or can stay as a monitored default.

The same pattern appears in the trust decisions described in The Trust Layer the Roadmap Still Treats as Optional and the workflow timing failures in The Volunteer Desk Where the Trust Toggle Appeared. Both posts show how missing constraints turn good intentions into unreliable delivery.

I consult with ministry product leaders on emotional outcome design and constraint setting for agent-driven roadmaps. Let’s talk.

The Nerve That Disappears Once You Ship the First Agent

Product team shipping the first agent
The “first ship then standardize” framework from recent AI product playbooks assumes that initial deployment success naturally leads to broader rollout discipline. That assumption breaks down in ministry contexts where the first agent reveals how much local variation actually drives engagement. This approach treats the initial launch as proof that centralized controls will now improve outcomes. In practice, it removes the very people who understand the edge cases from making further adjustments. This is the foundational misread that causes product teams to treat early wins as signals to lock down processes rather than expand the range of cheap tests. Experience reinforces the pattern because every prior role rewarded documentation and consistency over continued deviation. Jensen Huang’s sovereign AI argument supplies the needed lens. Huang argued that organizations lose leverage when they outsource core model decisions to external providers; they regain it by owning the stack at the level where outcomes actually form. Applied to faith-tech, the same principle shows why ministry teams must retain local control over agent behavior instead of routing every adjustment through a shared governance layer.

Why shipping the first agent makes the safe path look professional

Shipping the first agent changes more than your roadmap—it changes your nerve. After the first agent ships, product reviews shift. The conversation moves from “did this change anything for the volunteer who only has seven minutes” to “does this meet the new reliability checklist.” The checklist itself is not the problem. The problem is that it rewards artifacts that travel well across teams rather than signals that only appear in one context.

Sermons4Kids data showed this pattern clearly. The print-first flows that kept volunteer completion rates above 70 percent never survived the first enterprise review. Reviewers asked for API documentation and usage dashboards. They did not ask whether the local children’s director could still finish the workflow without opening a laptop.

The pattern repeats with every new agent. The team that built the initial version stops shipping the small prompt variants that only make sense inside one church’s calendar. Those variants disappear because no one is measured on them anymore.

Where Gen Z default experimentation still wins

Research on experimentation culture from Harvard Business Review shows that teams who lower the cost of small bets out-innovate those who optimize for safe, legible artifacts.

Younger product builders still default to local changes because they have not yet internalized the cost of deviating from the approved path. They will alter a single agent response for a youth group that meets on Wednesday nights even when the change will not appear in the quarterly metrics deck.

That instinct produces the data the rest of the organization needs but refuses to request. One volunteer group tested an agent that generated discussion questions from the previous Sunday’s sermon audio. The questions were mediocre by model benchmarks, yet attendance at the small group rose because the questions referenced a local event the model had never seen in training data.

The result only surfaced because the builder ignored the standardization step that normally follows a first ship. Experience teaches people to skip exactly these steps.

A protocol for shipping the first agent: one risky test per sprint

Teams that keep the nerve intact after the first agent treat local tests as non-negotiable work rather than optional polish. They schedule the test before the sprint begins and protect the time from review requirements.

The test must meet three conditions. It runs only for one defined workflow. It uses a prompt or retrieval change that cannot be justified by aggregate metrics. It is logged with the exact user group and the exact failure mode observed, not a success rate.

These constraints sound small until they collide with the post-ship review process. The review wants the change generalized. The protocol refuses. The tension is the point. It keeps the team from mistaking professional appearance for continued learning.

Your Turn: Apply This Today After Shipping the First Agent

  • Select one children’s ministry workflow that already uses an agent for curriculum suggestions and create a single prompt variant tuned only to your three most active volunteers’ actual schedules.
  • Run the variant for two weeks without adding it to any shared library or dashboard, then record the exact number of volunteers who completed the workflow versus the prior baseline.
  • Log the one sentence the volunteers said unprompted about the output, even if it does not map to any current success metric.
  • Delete the variant at the end of the two weeks unless the local completion rate improved by at least ten points.
  • Present the deletion decision and the single sentence to your immediate team as the full result, with no slides or generalization.
  • Repeat the entire sequence next sprint on a different workflow before any central model review occurs.

The Trust Layer the Roadmap Still Treats as Optional and The Year I Treated Experience Like the Asset both trace what happens when teams stop protecting the space for these tests after the first production deployment.

I consult with product leaders in ministry organizations on running under-optimized local AI tests and protecting experimentation after initial agent shipments. Let’s talk.

The Trust Layer the Roadmap Still Treats as Optional

Product team designing the trust layer for an AI agent roadmap

You’re the product manager who got the AI mandate dropped on your desk at a mid-sized denomination last quarter, the one where leadership wants proactive agents handling prayer requests and volunteer scheduling but no one has defined the trust layer that governs what happens when those agents make a mistake that reaches an actual family in crisis.

You’ve shipped a few generation features already. The outputs look clean in the demo. Now the real question is whether your trust layer records why a recommendation was made, who approved the underlying data, or how to unwind it if confidence breaks.

Why the Trust Layer Comes Before Generation Scales

Seneca wrote his letters to Lucilius with the explicit understanding that private correspondence could become public record. He treated every exchange as something that might need to be defended later. That discipline turned individual letters into an infrastructure of accountability rather than just personal advice.

Most current AI roadmaps treat the same principle as optional, even as established risk frameworks put accountability at the center. Generation gets the headcount and the model budget. Governance gets the note that says “we’ll add logging later once we see usage.” The result is a system that can produce fluent suggestions but cannot explain its own chain of custody when a pastor or volunteer needs to know why a particular action was recommended.

The pattern shows up in places that already run at scale. Children’s ministry platforms learned years ago that the 7-minute volunteer will abandon any workflow that cannot be completed without hunting for source material. When an agent starts suggesting curriculum adjustments without a visible audit trail, the same abandonment happens, except now the damage includes families who received the wrong follow-up.

Handoff Failures Without Logs

Once an agent begins acting across multiple systems, the handoff points multiply. A scheduling agent pulls from membership data, checks volunteer availability, and sends a confirmation. If the confirmation goes to the wrong small group because a preference field was misread three months earlier, the failure is invisible until the volunteer shows up at the wrong house.

Seneca’s correspondence rule forces the writer to anticipate that future reader. In product terms, this means every agent action must carry the equivalent of a dated, attributable letter: the model version, the data snapshot used, the policy rule that was active, and the human override path. Without that record, the first time something goes wrong the team is left reconstructing decisions from chat logs and memory.

Teams that skip the trust layer discover the cost during incidents rather than during planning. The incident review then becomes an exercise in guessing which prompt version was live instead of reviewing a deliberate log. That guessing erodes the very trust the agent was meant to build.

Budgeting the Trust Layer Before Launch

Usage numbers will eventually force the conversation, but by then the technical debt is already distributed across several models and data sources. The cheaper moment is before the agent is turned on for real volume.

Allocating tokens and engineering time to the trust layer of immutable action logs, policy versioning, and human-in-the-loop checkpoints looks expensive on a slide. It is cheaper than the support load and relationship repair that follow a single high-visibility error. Seneca did not write letters only when he felt like it; he wrote them consistently because the cost of inconsistency was permanent loss of credibility.

The same consistency applies to roadmap sequencing. Trust features cannot be sequenced after “see how it performs.” They have to be present at the first production handoff or the system remains brittle by design.

Your Turn: Apply This Today

  • Pick the single agent flow closest to launch and require an immutable action log as a launch blocker, not a backlog item.
  • Write the three-sentence policy that defines what constitutes an override-worthy recommendation for that flow and store it alongside the model version.
  • Schedule a 90-minute session with the volunteer or pastor who would actually receive the agent’s output and walk through a mistaken recommendation to see what recovery steps are missing.
  • Allocate 15 percent of the remaining agent development budget explicitly to governance tooling before any additional generation work is approved.
  • Create a one-page “letter” template that every agent action must populate, modeled on the minimum fields Seneca would have considered non-negotiable for later readers.
  • Identify the existing incident response runbook and add a step that requires pulling the agent decision record before any human conversation begins.

The same sequencing problem appears in The Mandate That Arrived Before the Workflow Existed and the concrete toggle decisions described in The Volunteer Desk Where the Trust Toggle Appeared.

I consult with product leaders shipping AI agents inside faith-based organizations on trust infrastructure and roadmap sequencing for governance features. Let’s talk.

The Year I Treated Experience Like the Asset

Product leader weighing experience as an asset against a beginner mindset

I spent the better part of a decade treating experience as an asset, the main thing that made my product decisions reliable. At teams building tools for children’s ministry volunteers and global Bible readers, that experience let me spot the same failure patterns in new interfaces or content workflows before most people finished their first user interview. It felt like an edge that only grew stronger with time.

Then agents started handling the repeatable parts of discovery, spec writing, and initial testing. The same track record that once accelerated decisions began to close off questions I no longer felt I had time to test. The cost showed up in two missed releases where we shipped patterns that worked for veteran users but collapsed for the next generation of volunteers.

This is the foundational misread that causes product teams to treat experience as an asset that only appreciates, even once automation absorbs the routine work. Charlie Munger’s circle of competence describes the boundary where real skill exists and outside which confident guesses turn expensive. When experience stops being updated by direct contact with new conditions, the circle contracts even while the person inside it feels more certain.

The Moment Experience as an Asset Stopped Helping

The shift showed up in a review of a new volunteer onboarding flow for curriculum resources. I had run similar projects three times before and knew the drop-off points after the first print-and-prep step. I signed off on the revised screens based on that history.

Three weeks later the completion data showed the new cohort abandoning at a different step entirely—one that only appeared once agents generated the first-round content variations. My prior runs had never included that variable, so the pattern I trusted no longer matched the actual system.

The circle had narrowed without me noticing because the inputs I used to gather myself were now being produced faster by models I reviewed only at the end. The seniority that once compressed review time now compressed the questions I bothered to ask.

Where Gen Z Chutzpah Actually Compounds Faster

A newer teammate on the same project kept pushing for live observation of volunteers using the agent-generated lesson drafts in real time. She had none of the historical comparisons I carried, so every session produced fresh variables instead of confirmation of old ones.

Her questions surfaced a constraint around token limits that directly affected how small churches could customize material without extra cost. I would have caught the same issue eventually through the slower path of post-launch metrics. She reached it in two moderated sessions because she treated every output as new ground rather than another instance of a known category.

The difference was not raw intelligence. It was that her shorter track record left more room for direct observation before the circle of what she considered settled tightened around her.

Keeping Experience as an Asset Without Faking Beginner Status

The practical problem is not whether experience as an asset still matters. It is how to keep the boundary moving outward when agents now own the first pass on tasks that once forced repeated contact with users. The answer is not to pretend the years never happened or to manufacture beginner exercises that everyone sees through.

Instead the work becomes deliberate insertion of new variables into the loops that used to run on autopilot. That means choosing projects where the agent output cannot be reviewed without fresh fieldwork, then protecting the calendar space for that fieldwork even when the review itself could be done in a fraction of the time.

It also means tracking which decisions still rely on patterns formed before agents arrived and forcing at least one new data source into each of those decisions. The goal is not humility theater. It is keeping the competence boundary from freezing in place while the surrounding environment keeps changing.

Your Turn: Apply This Today

  • Pick one recurring product decision you currently close from memory and add one live observation session with a user type that did not exist when you first learned the pattern.
  • Review the last three specs you approved that relied on agent-generated content and list the variables you did not test because prior experience said they were stable.
  • Block two hours next week for fieldwork that cannot be summarized by an agent and treat that block as non-negotiable the same way you treat revenue reviews.
  • Ask the newest person on your team to name one assumption in the current roadmap that would change if their shorter history were the only data source, then run the test they describe.
  • Document which decisions in your area still carry the highest cost if the circle has already narrowed, and assign the next one to someone whose track record is shorter than yours.
  • Set a recurring calendar reminder every six weeks to re-test one previously stable metric against a fresh cohort rather than assuming the old baseline still holds.

The pattern shows up again in the way teams handled early agent routing for content specs and in the career questions that surface once the old accumulation model no longer matches how competence actually grows.

I consult with product leaders and ministry technology teams on preserving experimentation capacity after automation arrives, updating decision patterns that predate agents, and designing feedback loops that keep competence boundaries from contracting. Let’s talk.

The Mandate That Flattened the Autonomy They Needed

Product team facing flattened autonomy after a top-down AI mandate

Mandates from leadership to adopt AI tools don’t speed up judgment. They flatten the autonomy that forces people to build judgment in the first place, and that flattened autonomy is harder to rebuild than any model.

Product teams hear the order and start routing every decision through the model. The stated goal is consistency and speed. The actual result is that no one practices the hard calls anymore, so the model has nothing reliable left to learn from.

This pattern shows up when the people closest to the work lose permission to refuse or reshape the tool. Once that permission disappears, the output quality stops mattering because the flattened autonomy has already done its damage upstream.

How a Mandate Produces Flattened Autonomy

John Wesley’s three rules offer a way to see the damage clearly. Do no harm. Do good. Stay in love with God. Applied to product work, the rules function as a test for whether an AI mandate preserves or destroys the conditions for sound decisions.

The first rule exposes the Meta failure. When leadership required every team to integrate large models into core workflows, the immediate effect was that engineers stopped questioning whether a generated suggestion actually fit the user context. Harm appeared as quiet degradation: models trained on past bad patterns kept surfacing them faster. Teams could no longer pause the loop because the mandate treated refusal as resistance rather than stewardship.

The second rule shows what doing good requires instead. Good in this setting means the tool must leave the person using it more capable of independent judgment after the session ends. Meta’s rollout measured adoption and token spend. It never measured whether product managers could still articulate why a feature should not ship. The rule forces that second measurement. Without it, the mandate produces volume that looks like progress while judgment atrophies.

The third rule asks whether the process keeps people connected to what actually matters. In Meta’s case the model became the intermediary for almost every review. Engineers lost direct contact with the original problem statements. The connection that should have been protected was the one between the person and the real user outcome. Once that link runs through an always-on model, the work stops forming people who can make hard calls without it.

Ministry and product teams repeat the pattern when they treat AI adoption as a compliance exercise. A children’s ministry platform once required every curriculum writer to run drafts through the model before human review. Volunteer completion rates dropped because the generated language no longer matched the seven-minute attention window the actual users had. The mandate protected the appearance of modern tooling while removing the writers’ ability to test language against real classrooms.

Another team at a large resource site mandated that all search-ranking experiments begin with model-generated hypotheses. The first three months produced statistically significant lifts on paper. The lifts disappeared once the team could no longer explain why a particular ranking change served the reader who opened the app on Sunday morning. The rule set had been inverted: the model came first, and human judgment became the optional check rather than the protected core.

Reversing Flattened Autonomy Before the Model Ships

What changes when autonomy is protected before the model ships is that the three rules regain their force. Teams keep explicit veto rights over any generated output that cannot be defended in plain language. They measure whether the tool leaves the user faster at the next unassisted decision. They refuse rollouts that insert the model between the worker and the person they ultimately serve.

The shift is small in process terms and large in outcome. A team that resists flattened autonomy by keeping veto rights will reject more suggestions early. That same team will also surface higher-quality uses of the model because the people making the calls still practice the judgment the model is meant to support.

Your Turn: Apply This Today

  • Write the exact refusal script your team will use when a generated suggestion fails the three-rule test, then test it on the next three model outputs this week.
  • Remove the model from the first draft stage of one workflow and require the human author to produce the initial version before any AI assistance is allowed.
  • Track whether each team member can still explain the user problem in their own words after a full day of model-assisted work; log the explanations for two weeks.
  • Set a hard limit of two model calls per decision until the person can state the downside of following the model’s recommendation without looking at it.
  • Schedule a weekly 30-minute review where the only allowed topic is which recent model use reduced someone’s ability to decide without the tool, and remove that use from the workflow.
  • Assign one person on each product squad the standing authority to block a mandated AI step if it cannot be defended against all three rules in a single sentence.

The same tension between top-down mandates and preserved judgment appears in The Mandate That Arrived Before the Workflow Existed and The Token Budget That Never Tied Back to Sunday.

I consult with product leaders and ministry technology teams on protecting team judgment during AI rollouts and measuring whether tools increase or decrease decision quality over time. Let’s talk.

The Mandate That Arrived Before the Workflow Existed

Volunteer coordinator at the competence boundary reviewing an AI agent suggestion

McKinsey’s AI transformation framework tells leaders to set the mandate at the top, define the use cases from strategy documents, then cascade requirements downward through implementation teams. The model assumes the center holds enough context to name what good looks like before any local team touches the tool. That assumption collapses the moment the workflow crosses a competence boundary into the hands of volunteer coordinators who decide in seven-minute windows between service and lunch duty.

The framework measures success by adoption dashboards and model accuracy scores. It never asks whether the person who actually assigns rooms and prints name tags can still exercise judgment when the agent suggests a different curriculum track. The result is a shipped feature that meets the mandate and breaks the loop that keeps real ministry work moving.

This is the foundational misread that causes product teams to ship agents that look complete in demos and stall in practice. The gap is not technical. It is a competence boundary problem.

Charlie Munger described a circle of competence as the small set of decisions where a person knows the variables well enough to see second-order effects. Outside that circle, even smart people produce confident errors. Ministry tools cross that boundary the moment an AI agent starts suggesting how a volunteer should adapt a lesson for a child who just lost a parent. The coordinator sees the child’s face and the empty chair. The model sees patterns across anonymized sessions. Munger’s point was that you stay inside the circle or you bring in someone who lives there. Most AI mandates do neither.

How the Mandate Reached the Volunteer Coordinator

The requirement landed in an all-hands note: every new curriculum product must include an AI-assisted personalization layer by Q3. The product team translated the mandate into a clean spec. The agent would read the lesson text, the child’s age and attendance history, then output a suggested adaptation plus a printable note for the volunteer.

The volunteer coordinator first saw the feature the week before VBS training. She opened the new screen, read the suggested rewrite for the story of David and Goliath, and realized the model had removed every reference to violence without asking whether the church’s teaching team wanted that choice made for them. She spent twenty minutes rewriting the note by hand so it still matched what the lead pastor had approved six months earlier.

No one on the central team had seen that rewrite cycle because the success metric stopped at “agent output accepted.” The coordinator’s actual work of preserving local teaching intent never appeared in the dashboard.

The Competence Boundary the Model Crossed Without Noticing

The agent operated on content patterns and user metadata. It had no signal for whether a suggested change would require the volunteer to find new craft supplies at 8:45 on a Wednesday night. That decision sits inside the coordinator’s circle. She knows which families can bring extra glue sticks and which weeks the church van is already booked.

When the model suggested swapping the craft for a tablet-based reflection activity, it crossed the boundary. The coordinator could see the downstream cost in real time: two extra parent volunteers needed, one device cart that was already reserved for the youth group, and a training burden she would carry alone. The agent could not see any of it because those variables live only in her weekly rhythm.

Munger would have called this operating outside the circle with no compensating process. The mandate treated the model as competent across the full decision surface. The local team experienced the model as confidently wrong on the details that actually determined whether the session happened.

Rebuilding Decision Rights at the Competence Boundary

After the first month of quiet workarounds, the team added a single gate. Any agent suggestion that touched supply lists, room assignments, or approved teaching emphasis required an explicit “local override” click before the note could be printed. The override did not require explanation in the system, only acknowledgment that a human with weekly context had seen it.

They also moved the default output from a full rewritten lesson to a short margin note that the coordinator could accept, edit, or ignore in under sixty seconds. The model still generated the note, but the decision rights stayed with the person who would stand in the room.

Usage recovered once coordinators could treat the agent as a suggestion layer rather than a replacement for their own judgment. Retention among volunteer leads rose because the tool no longer created invisible extra work that only they could see.

Your Turn: Apply This Today

  • Pick the single recurring decision your current AI feature makes on behalf of a local ministry user and write down the exact variables that person sees weekly that the model cannot access.
  • Shadow one volunteer coordinator for a full planning cycle this week and note every judgment call that happens after the agent output appears on screen.
  • Change the default output of one agent flow from a completed artifact to a one-sentence margin suggestion that requires an explicit local accept step.
  • Remove any success metric that counts “agent output accepted” without also tracking whether the local user still completed their weekly task in the same amount of time.
  • Map the three decisions that must stay inside the local circle and add a visible override control for each before the next release.
  • Run a thirty-minute review with the actual end user of the last agent suggestion that crossed a competence boundary and log what the model missed.

The same pattern shows up in how teams handle inherited agent mandates and in the months spent routing every spec through the largest model before testing it against real weekly rhythms. Both posts trace the cost of assuming central competence where local judgment loops still do the work.

I consult with product leaders building tools for ministry volunteers and church staff on AI workflow boundaries and local decision rights. Let’s talk.

The Volunteer Desk Where the Trust Toggle Appeared

Volunteer flipping the trust toggle off in an AI lesson-planning tool

The children’s director sat at the folding table in the supply closet, the one with the wobbly leg and the permanent marker stain shaped like Texas. Her laptop fan whined as the planning tool loaded its afternoon suggestions. Three craft options appeared in the side pane, each one bright with stock photos of construction paper and glue sticks. She scanned the list, then reached for the small trust toggle at the top of the pane and switched it off. The recommendations disappeared. She exhaled, opened the blank lesson template, and started typing from memory instead.

She had tried the suggestions twice already that month. Both times the activities called for items the budget line did not cover, and both times she had spent an extra evening driving to two different stores looking for substitutes. Flipping the trust toggle was faster than arguing with the interface.

Solomon’s judgment offers a useful lens here. Two women claimed the same child. The king did not search for more data or run another test. He watched which person was willing to lose the outcome rather than harm the person in front of them. The real parent revealed herself by her readiness to override the proposed solution. Modern planning tools rarely test for that same willingness. They optimize for suggestion volume and measure acceptance rates, then wonder why the people closest to the work keep turning the suggestions off.

This matters because the person sitting at the volunteer desk already carries the constraint the model cannot see. She knows the supply closet inventory, the parent who always forgets to send scissors, and the fact that the church van is in the shop this week. When the interface treats her override as a failure state instead of the normal next step, it adds friction exactly where the work is most fragile.

Why the Trust Toggle Appears at the Point of Use

Most AI planning features ship with suggestions turned on by default, a choice usability research shows users rarely revisit. The assumption is that more options will help the time-pressed user. In practice the suggestions arrive already filtered through averaged data from thousands of other churches. What remains is often a list of activities that require either money or prep time the current volunteer does not have.

I watched the same pattern repeat across three different children’s ministry tools last year. Each time a new recommendation engine rolled out, the first week showed high click-through numbers. By week three the internal dashboard showed a spike in users who had disabled the pane entirely. The product team read the drop as disengagement. The volunteers described reaching for the trust toggle as finally being able to finish the task.

The break happens because the default carries an implicit claim: the model knows the local constraints better than the person who just counted the glue sticks. When that claim proves false, the user does not refine the suggestion. She removes the entire channel. Solomon would recognize the move. The one who truly bears the outcome will cut off the process rather than accept a solution that harms what she is responsible for.

The Override Cost No Dashboard Shows

Every extra click required to turn suggestions off carries a hidden cost. The volunteer has already opened the tool, seen the list, evaluated each item against her actual supplies, and then hunted for the setting. That sequence takes thirty to forty seconds on a good day. Across a quarter it adds up to hours that never appear in any usage metric because the time is spent outside the product.

The deeper cost is eroded trust. After the third time a suggestion set proves unusable, the volunteer stops believing future suggestions will be relevant. She treats the pane as noise rather than help. Product teams often respond by adding more personalization signals, yet the signals they request (budget ranges, supply lists, volunteer skill tags) are exactly the data points the volunteer does not have time to maintain inside the tool.

Solomon’s test still applies. The person willing to discard the proposed plan is usually the one who understands the real stakes. Forcing her to keep justifying the discard only increases the distance between the model and the work.

Designing the Trust Toggle Before the Feature Ships

The practical fix begins with treating the trust toggle as a core control rather than a settings afterthought. Place it directly beside the suggestion pane, not buried three menus deep. Label it plainly: “Hide suggestions for this plan.” Make the choice sticky for that user and that planning cycle so she does not repeat the action every time she opens the lesson.

Teams that have done this report an unexpected result. When the off-switch is cheap, more volunteers experiment with suggestions in the first place. They know they can dismiss the output without penalty, so they glance at it instead of preemptively disabling the whole feature. The acceptance rate on individual suggestions may drop, but the overall usefulness of the tool rises because the interface stops fighting the person who actually owns the outcome.

This is not a rejection of AI assistance. It is an application of Solomon’s discernment: surface the model’s proposal, then give the one closest to the child the uncomplicated ability to set it aside. The measure of a trust feature is not how often the suggestion is accepted. It is how little effort it takes for the right person to override it when acceptance would cause harm.

Your Turn: Apply This Today

  • Open the current planning tool your team uses and locate the suggestion pane; add a one-click toggle labeled “Hide for this plan” directly beside it before the next release.
  • Pick one children’s ministry volunteer who has used the tool in the past month and ask her to walk through her last planning session while you watch; note every time she ignored or dismissed a suggestion.
  • Change the default state for new users in one small cohort so the suggestion pane starts hidden; measure whether they turn it on later compared with the always-on group.
  • Write the override action into the user flow spec for the next AI feature so the engineering ticket includes the off-switch before any model integration work begins.
  • Review the last three support tickets from ministry users about “bad suggestions” and trace whether each person had an easy path to disable the pane without leaving the screen.
  • Set a recurring calendar reminder for the first Monday of each month to ask two frontline users whether the current suggestions still match what they actually have on hand.

The same pattern shows up in the work I described in The AI Line Item No Demo Ever Justified and The v1 Where the Model Handed Me the Safe Path.

I consult with product leaders building AI planning tools for ministry volunteers on suggestion interfaces, override controls, and frontline decision rights. Let’s talk.

The Token Budget That Never Tied Back to Sunday

Ministry team reviewing an AI token budget against volunteer outcomes

Teams running ministry platforms reported a 62% rise in monthly token consumption for AI-assisted content generation over the past twelve months. The figure looks like adoption until you trace where the tokens actually went. Most of that token budget increase came from repeated prompt refinement and human review loops that left the original volunteer task list unchanged.

The obvious reading treats higher spend as proof of value delivered. The reverse is closer to what happened. Extra tokens masked the absence of any measurable shift in how a children’s ministry volunteer prepared a lesson on a Tuesday night. A token budget that grows without a matching change in volunteer outcomes is a warning sign, not a win.

When Spend Reviews Hit the Smallest Teams First

Jensen Huang’s argument for sovereign AI centers on keeping decision rights inside the organization that actually owns the outcome. When that principle is ignored, budget conversations default to usage dashboards that every team can see but few can act on. Small volunteer-driven groups feel the pressure first because they lack both the headcount to absorb review cycles and the political capital to push back on centralized spend.

One curriculum platform tracked a spike in summarization calls during the back-to-school window. The calls originated from regional coordinators trying to shorten existing lesson outlines for volunteers who only had seven minutes between work and the Wednesday program. Each summary still required a second pass by a paid editor before it reached the volunteer. Token counts rose; preparation time for the volunteer stayed flat.

The pattern repeats whenever an AI feature is measured by output volume rather than handoff reduction. The smallest teams absorb the verification burden because they sit at the end of the chain and cannot delegate it further. Larger churches with staff simply route the same content through another approval layer and call the process “governance.”

Mapping the Token Budget to One Observable Handoff

Huang’s point about sovereignty only lands when every token is attached to a single, named handoff that someone can watch change. In volunteer settings that handoff is usually the moment a printed or emailed resource reaches the person who will use it without further editing. Most current dashboards stop counting at generation and never reach that moment.

A children’s ministry tool introduced an AI outline generator for weekly lessons. Usage logs showed strong uptake among mid-week staff. When the team instead measured whether volunteers still opened the full curriculum PDF or simply printed the AI summary, the number of untouched PDFs stayed above 70%. The tokens had produced an intermediate artifact that required the same amount of volunteer attention as before.

Reversing the measurement forces a narrower scope. Instead of asking how many outlines the model produced, the question becomes whether the volunteer’s print step or copy-paste step disappeared for one specific age group. That single observable change is the only unit that can be priced against token cost with any honesty.

Building a Token Budget Ledger Before the Next Renewal

Renewal conversations arrive with polished charts of tokens consumed and outputs generated. Without a parallel ledger that records the exact workflow step removed or shortened, those charts cannot answer whether the spend should continue. The ledger has to be built while the feature is still in pilot, not after the contract is signed.

One platform began requiring every new AI-assisted workflow to declare the single volunteer action it intended to replace. The declaration lived in the same ticket as the prompt engineering work. Three months later the team could show that only two of the six pilots had actually reduced the declared action; the other four had added review steps instead. Budget was reallocated before the annual review rather than defended after it.

The discipline is unglamorous. It means writing down the before-and-after timing for one concrete task, collecting that timing from five real volunteers, and refusing to count any token that cannot be tied to a measured difference. Teams that adopt the habit stop celebrating usage graphs and start treating the token budget as a ledger of removed work rather than a record of activity.

Your Turn: Apply This Today

  • Pick the single volunteer task your current AI feature was meant to shorten and write the exact before-and-after action in one sentence.
  • Instrument the product so you can observe whether that action still occurs for at least ten users this month.
  • Attach the token cost of the feature directly to those ten users and calculate cost per removed action.
  • Present the cost-per-action number in the next spend review instead of total tokens used.
  • If the action count has not dropped, disable the generation path for new users until the handoff changes.
  • Document the revised number and the disabled path so the next renewal conversation starts from an outcome rather than a usage total.

The same gap between reported usage and actual workflow change appears in both “The AI Line Item No Demo Ever Justified” and “The Proven-Better Pattern That Still Missed the Smallest Churches.”

I consult with product leaders building tools for volunteer-driven ministries on mapping AI spend to observable handoffs and maintaining outcome ledgers ahead of renewals. Let’s talk.

The Career Question the Old Playbook Still Ignores

Newly promoted product leader rethinking the old playbook for managing AI agents

A product leader asked me last month how to decide what parts of the job to hand off to agents now that every promotion comes with an agent mandate. The answer is simple but rarely stated: keep the judgment calls where the downside lands on people you know by name, and let the models handle the rest. Everything else is secondary. The old playbook never prepared anyone for this question.

Charlie Munger spent decades telling investors to stay inside their circle of competence. The rule was not about avoiding hard work. It was about refusing to make calls when you could not see the full set of consequences. In product work that rule now collides with live agents that produce first drafts at scale. The circle does not shrink because the tools got faster. It moves because the cost of being wrong now hits real users and real teammates before you finish reviewing the output.

The promotion the old playbook never described

The title change arrived with a new expectation. I was told the team would run on agents for initial scoping, first-pass requirements, and early user-story drafts. The headcount stayed flat, but the mandate was to ship more surface area with the same number of humans. The old promotion playbook measured scope by how many features I could personally shepherd. The new one measured it by how many decisions I could safely review after an agent had already moved them forward.

At first the arrangement felt efficient. The agents produced clean documents that looked ready for stakeholder review. Then a set of stories reached the volunteer build team for children’s ministry curriculum. The agent had optimized for completion rate metrics that made sense on paper but ignored the seven-minute volunteer who prints the lesson on Sunday morning and discovers the activity requires materials that are not in the supply closet. The error was small on screen and expensive in real rooms. That was the moment the circle of competence stopped being a theory.

Where the old playbook breaks once agents write the first draft

Munger’s test was never whether you could understand the output. It was whether you could still trace the second-order effects on the actual people who would use it. When agents write the first draft, that trace becomes harder because the draft already carries the appearance of finished work. The temptation is to treat the model’s version as the default and only edit around the edges. The boundary moves inward the moment you accept that default.

I watched the same pattern on two other teams. One handled global discipleship tools where traffic numbers are easy to celebrate. The agent suggested a flow that increased session time for logged-in users but quietly dropped first-time visitors from low-bandwidth regions. The other team managed sermon resources for pastors who prepare on Friday night. The agent optimized for search ranking and removed the plain-text print option that those pastors relied on when their internet dropped. In both cases the humans reviewing the work had never held the downstream roles themselves. The competence gap was not technical. It was relational.

The two questions the old playbook left out

Research like Harvard Business Review on AI and product roles confirms the old playbook never accounted for managing agents that draft the work.

The first question is whether I can name the three people who will feel the downside if this decision is wrong. If I cannot name them, I do not own the call. The second question is whether those same three people would recognize the failure mode inside the first five minutes of using the product. If the answer is no, the scope stays with me until I can answer both questions clearly. These two questions have become the practical edge of my circle.

Applying them has changed what I accept. I still review agent output, but I now require a short “who gets hurt” paragraph written by a human who has done the actual job the agent is modeling. The paragraph must be attached to every story that reaches review. It is not a process document. It is the minimum signal that the judgment call has not yet been outsourced.

Your Turn: Apply This Today

  • Block thirty minutes this week and write down the three judgment calls you currently own that you would refuse to hand to an agent even if the output looked perfect.
  • For each of those three calls, name the two people downstream who would absorb the cost if the call is wrong.
  • Write one sentence for each person describing how they would first notice the failure in their actual workflow.
  • Schedule a fifteen-minute conversation with one of those people and read the sentence out loud to confirm it matches their experience.
  • Add a standing rule in your next planning doc that any agent-generated story must include a “who gets hurt” paragraph before it reaches review.
  • Review the last five stories your team accepted from agents and mark which ones would have failed the two questions above.

The PM Who Just Inherited the Fable Agent Mandate and The Compensation Split No Playbook Prepared PMs For both trace the same boundary problem from different angles.

I consult with product leaders and ministry technology teams on agent scope decisions, judgment call ownership, and downstream user impact reviews. Let’s talk.

The Proven-Better Pattern That Still Missed the Smallest Churches

Small rural church where the proven-better pattern failed to fit

The proven-better pattern approach, drawn directly from the scaling tactics in books like “Inspired” by Marty Cagan, directs teams to locate high-performing flows from mature products and transplant them into smaller offerings. The method treats completion rates and delight signals from large user bases as reliable signals worth replicating. It promises faster iteration by skipping the need to invent from scratch. The proven-better pattern can still miss the people you most want to serve.

This method breaks for ministry tools aimed at volunteer-led congregations because the original delight signals emerged under conditions of paid staff, reliable devices, and repeated weekly exposure. A pattern built for daily app users cannot assume the same friction profile when the user opens the tool once a month on a shared tablet during set-up time. The transplant therefore lands without evidence that the underlying job still exists in the new setting.

The result is shipped features that look finished on internal demos yet produce zero measurable movement in the actual retention metric that matters for these teams: whether the volunteer completes the task before the first child arrives.

This is the foundational misread that causes product teams to treat proven patterns as context-neutral assets rather than hypotheses that still require validation. The pattern itself is not the problem. The decision to bypass the step that checks fit is.

Teresa Torres’s continuous discovery habits supply the missing lens here. Torres insists that discovery work must continue after a pattern has shown results elsewhere. The habits force teams to test the assumption that the same outcome will appear in a different environment rather than declaring the pattern “proven” once it succeeds at scale.

Where the proven-better pattern first led me wrong

I took a three-screen welcome sequence that had lifted activation numbers in a consumer reading app and dropped it into a planning tool for children’s ministry volunteers. The sequence used progress bars, a quick profile photo upload, and a one-tap “start your first lesson” button. Internal testing showed the flow completed in under ninety seconds.

When the first twenty-person church installed the update, the photo upload step failed on the shared device because the tablet camera permission had been locked down by the church IT volunteer. The progress bar never advanced for anyone who reached that screen. Completion rate for the new onboarding dropped to 12 percent within the first month.

The copy had preserved every pixel of the original pattern. It had preserved none of the conditions that made the pattern work.

Why the proven-better pattern missed the smallest churches

The original pattern produced delight because users returned daily and the app could remember their last state across sessions. In the small church setting the tool is opened once, used for ten minutes, and then closed until the next rotation of volunteers. No memory of prior sessions survives.

The progress bar that signaled momentum to a daily user now signaled an unfinished task to a volunteer who would not return for four weeks. The metric that looked like delight in the source context became an abandonment signal in the new one.

Teams measuring only the source metric never see this reversal. They ship the pattern and move to the next proven item on the list.

The discovery interview the proven-better pattern skips

A single thirty-minute conversation with a volunteer coordinator from a rural church would have surfaced the camera permission constraint and the four-week gap between uses. That conversation never occurred because the pattern had already cleared internal success criteria at the source company.

Torres’s continuous discovery habits require exactly this interview even after the pattern has data behind it. The habit is not additional work; it is the work that prevents the pattern from becoming expensive technical debt in the new context.

Your Turn: Apply This Today

  • Pick the smallest user segment your product currently serves and pull the last three patterns you copied from larger products.
  • Schedule one thirty-minute call this week with a single user from that segment; ask only what changed the last time they tried to complete the flow.
  • Record the exact constraint they name that did not exist in the original pattern’s environment.
  • Write one sentence that states whether the copied pattern still solves the job under that constraint.
  • Remove or replace the pattern element that fails the test before the next sprint planning meeting.
  • Repeat the interview with one new user from the same segment every month for the next quarter.

The same continuous discovery habit that would have caught the onboarding mismatch appears again in “Trust Logs Beat the Next Feature Sprint” and “The v1 Where the Model Handed Me the Safe Path.”

I consult with product leaders and ministry technology teams on continuous discovery for small-segment users, pattern transplantation risks, and retention metrics that survive volunteer constraints. Let’s talk.

The Months I Routed Every Spec Through the Biggest Model

Product manager applying routing rules to decide which AI model handles each spec

I spent eight months sending every product spec to the largest model we could afford. The assumption was simple: more parameters meant fewer gaps for the volunteer teams who would later implement the feature. Prayer request flows, curriculum upload paths, even the small print-first forms for Sermons4Kids all went through the same pipeline. Output looked finished. Token counts stayed under budget. Launches still slipped because the teams had to rebuild the exact judgment calls the model had flattened. What finally fixed it was a set of routing rules, not a bigger model.

The cost showed up in volunteer completion rates. A children’s ministry coordinator would open the generated story, stare at the polished steps, and realize the model had removed the three places where a tired volunteer needed to decide whether to skip, adapt, or hand off. Those decisions never appeared in the original ticket. They reappeared as rework after the first test with real users.

This pattern is the foundational misread that causes product teams to treat model scale as a substitute for visible ownership. The bigger model does not eliminate judgment work; it buries it inside language that feels authoritative. Teams then discover the missing pieces only when the artifact reaches the people who actually run the program on Sunday morning.

John Wesley’s three rules offer a practical lens here. Do no harm. Do good. Stay in love with the disciplines that keep the work honest. Applied to model routing, the first rule means refusing any output that hides the exact decisions volunteers must still make. The second rule means routing only the portions of work that genuinely reduce repetition without erasing context. The third rule means keeping a running record of what each tier of model forces the team to restore by hand.

When Routing Rules Were Missing for a Prayer Request Flow

The spec described a simple intake form that let parents submit a request and receive a follow-up from a small-group leader within forty-eight hours. Fable 5 produced ten acceptance criteria, four edge-case branches, and a suggested notification schedule. Every sentence read as if someone had already tested it with real families.

The ministry lead opened the ticket and immediately flagged three problems the model had resolved without stating its assumptions. It had decided that every request would route to the same leader pool. It had assumed the parent would receive an automated confirmation that included the leader’s name. It had removed the step where the volunteer could mark a request as “needs pastoral review” without logging a separate ticket.

Those three decisions were the actual product. Restoring them took longer than writing the original spec because the language in the ticket made the choices look settled. The team had to reopen the conversation with the volunteer coordinators to learn what they actually needed to retain control over.

The Rewrite Cost Routing Rules Would Have Caught

Every time we routed a full artifact through the largest model, the rewrite work happened after the ticket reached the volunteer side. The model produced clean flows that assumed perfect information and consistent volunteer availability. Real ministry calendars contain neither.

One Sermons4Kids update required the volunteer to choose between printing a single sheet or a full lesson packet depending on how many kids showed up that week. The model had defaulted to the full packet because that option looked more complete. The volunteer had to add back the decision point and the accompanying print instructions. That single addition required three rounds of internal review because the original ticket no longer contained the rationale.

The token bill recorded only generation time. It never captured the hours spent reintroducing the places where a human had to exercise judgment under time pressure. Those hours appeared later as missed launch dates and lowered volunteer retention metrics.

Routing Rules That Actually Kept the Work Visible

Anthropic’s guidance on choosing the right model for the task reinforces why routing rules beat defaulting to the biggest model every time.

We eventually replaced the single-pipeline approach with three explicit tiers. Tier one handled only repetitive formatting and placeholder text. Tier two generated options with the decision points left explicit and bracketed. Tier three was reserved for narrow slices where the model had already demonstrated consistent accuracy against logged volunteer feedback.

Each tier carried a required log entry: what the model produced and what the team had to restore before the artifact reached the volunteer test. The log lived in the same ticket system the ministry teams used, not in a separate research document.

The discipline forced us to notice when the largest model was smoothing over exactly the friction that made a feature usable by a seven-minute volunteer. It also made the cost of each routing choice visible to the rest of the product group before the next planning cycle.

Your Turn: Apply This Today

  • Pick the next three product artifacts on your roadmap and assign each one a model tier before any generation begins.
  • For the artifact assigned tier one, log every sentence the model adds that removes a volunteer decision point.
  • For the artifact assigned tier two, require the team to restore at least two explicit judgment steps before the ticket moves to volunteer review.
  • After the third artifact ships, compare the three logs and mark which tier produced the highest volume of hidden rewrite work.
  • Share the marked log with the two volunteer coordinators who will actually use the feature and ask them to confirm the restored steps match what they need.
  • Update the routing rules document with the single sentence that best describes when tier three is now off-limits for your current volunteer contexts.

The v1 Where the Model Handed Me the Safe Path and Trust Logs Beat the Next Feature Sprint both trace the same pattern of hidden judgment costs across different product surfaces.

I consult with product leaders and ministry technology teams on explicit model tiering, restoring volunteer judgment points, and keeping rewrite costs visible in the ticket system. Let’s talk.

The AI Line Item No Demo Ever Justified

Product manager auditing the AI line item on a ministry software budget

The real test for any AI tooling budget is whether a volunteer can name one concrete change in their weekly work without being prompted by a product manager. Most line items never clear that bar. They multiply compute costs and internal velocity metrics while leaving the actual user path untouched.

This is the foundational misread that causes product teams to approve tooling spend based on engineering demos rather than observable volunteer behavior. The spend grows because the demo looks impressive, yet the volunteer still prints the same PDF at the same time each week.

John Wesley’s first rule — do no harm — offers the clearest lens here. Applied to tooling, it means refusing any hidden cost that quietly increases friction for the people who never see the roadmap.

The AI Line Item That Grew While Velocity Stayed Flat

One children’s ministry platform added three successive AI code-generation tools over eighteen months. The engineering team reported a 40 percent rise in pull requests merged. The volunteer completion rate on the curriculum builder stayed flat at 62 percent.

The new components arrived as internal micro-features: smarter autocomplete in the lesson editor, auto-tagging for activity files, and a prompt that suggested reorderings of the printed handout. None of these altered the seven-minute window in which a volunteer opens the material on a phone and decides whether to print or abandon it.

Feature velocity rose on the internal dashboard. The number of volunteers who reached the print step did not. The budget line item kept its justification because the demos never required a volunteer to describe the difference.

Wesley’s Do-No-Harm Test on the Hidden Line Item

Wesley’s rule forces a different question: does this spend create new friction that only appears after the contract is signed? In practice this shows up as added login steps, new file formats that require conversion, or prompt interfaces that demand the volunteer rephrase their intent before the system accepts it.

A sermon resource team once introduced an AI summarizer for weekly emails. The tool reduced drafting time for staff but required volunteers to confirm the summary matched their printed notes before forwarding. The confirmation step added forty seconds per email. Across hundreds of weekly users, the cumulative harm exceeded the drafting savings within one quarter.

The rule is simple to apply once named: if the cost cannot be explained to a volunteer in one sentence that ends with a visible action they already perform, the spend fails the test.

The Single Metric That Justifies the Line Item

Research on software ROI, including Harvard Business Review on measuring value, shows that a tool budget survives only when one usage metric ties the line item to real behavior change.

The only metric that held steady across recent quarters was the percentage of volunteers who completed their task without opening a support ticket or re-downloading the same file twice. AI features that improved this number stayed funded. Those that improved only internal velocity metrics were quietly deprioritized in the next planning cycle.

This metric survives because it cannot be gamed by a demo. A volunteer either finishes or does not. The tooling spend must produce a measurable lift here or it is, by definition, a line item no demo can justify.

Teams that adopted this filter reduced their AI tooling stack by roughly one-third while holding volunteer completion rates steady. The remaining spend now maps directly to actions the volunteer can name without prompting.

Your Turn: Apply This Today

  • Open the current AI tooling invoice and select the single largest monthly line item.
  • Write one sentence that describes the exact volunteer action this item must improve within the next seven days.
  • Run a five-user test this week using only that sentence as the success criterion.
  • Remove or renegotiate the line item if the test shows no lift in the named action.
  • Repeat the same mapping for the next two largest AI-related expenses before the next budget review.
  • Document the outcome in the same place the engineering velocity numbers are tracked so the comparison is visible to leadership.

The Compensation Split No Playbook Prepared PMs For and Trust Logs Beat the Next Feature Sprint both trace similar gaps between internal metrics and observable user behavior. The same pattern appears whenever tooling spend outruns the volunteer’s ability to name its effect.

I consult with product leaders and ministry tech teams on AI tooling justification and volunteer outcome mapping. Let’s talk.

The PM Who Just Inherited the Fable Agent Mandate

Product manager reviewing the agent mandate brief for a ministry AI assistant

You’re the product manager who just inherited the Fable Agent mandate at a denomination you only half-understand, and the brief says to ship something that feels like an AI assistant without breaking any existing volunteer workflows. The leadership team keeps repeating the same line about “safe defaults” and “conservative guardrails,” which sounded reassuring in the first three meetings. Now the first prototype sits in front of you and every suggested prompt returns the same careful, slightly generic answer. Inheriting an agent mandate like this means owning the decisions the model quietly makes for you.

That feeling is the signal. Fable 5’s training pushes it to stay inside a narrow band of acceptable ministry language and task patterns. The model was never rewarded for guessing what a children’s ministry volunteer actually needs when the craft supplies run out at 6:47 p.m. on a Tuesday. Most teams treat that narrowness as a feature. It is the exact place where product edge leaks away.

Charlie Munger’s circle of competence makes the cost visible. He warned that knowing the boundaries of what you understand is more valuable than expanding the territory you claim to know. When the agent’s defaults define the circle, every new workflow gets pulled back inside it. The denomination’s actual edge lives outside that circle, in the messy handoffs between volunteers who have never opened the app before. The product decision is whether to leave the model inside its comfort zone or to force it across the boundary on purpose.

Where the agent mandate stops short on volunteer orchestration

Volunteer orchestration is the workflow most teams test last because it looks simple on a flowchart. Assign curriculum, send reminder, mark complete. Fable 5 handles the first two steps inside its training distribution. It suggests standard email language, attaches the right PDF, and logs the assignment. The moment the volunteer replies with anything outside the script—“my kid has a fever and the printer is jammed”—the agent retreats to a safe, non-committal response or hands the thread back to a human.

I watched this exact failure on a children’s ministry rollout that served two hundred congregations. The agent could schedule the lesson but could not reroute the physical materials when a volunteer was suddenly unavailable. Completion rates dropped because the system kept treating the volunteer as a reliable node instead of a person whose availability changed daily. The safety layer had been tuned to avoid promising anything it could not guarantee, which is reasonable until the guarantee itself becomes the blocker.

The pattern repeats whenever the task requires real-time adaptation rather than pre-approved steps. The model’s circle of competence includes “send the lesson plan.” It does not include “decide whether the substitute volunteer can teach the same material with only a printed outline and ten minutes of prep.” That decision sits outside the training distribution, so the agent never offers it.

The prompt pattern that breaks the agent mandate safety layer

Anthropic’s own core views on AI safety describe why agent systems need human judgment at exactly these boundaries, and the agent mandate makes that responsibility yours.

The safety layer loosens when the prompt forces the model to declare its own uncertainty before it acts. Instead of asking the agent to generate the reminder email, you ask it to list every assumption it is making about the volunteer’s context and then rewrite the email once those assumptions are stated. The extra step is small in tokens but large in effect. It moves the output from the center of the circle toward the edge.

Teams that discover this pattern usually find it by accident. One product lead added the line “state the three facts about this volunteer you are least certain about” to every orchestration prompt. The agent began surfacing gaps the product team had never logged: the volunteer who only checks email on Sunday afternoons, the one whose printer only works at the church building, the substitute who has never taught without a co-leader. Those gaps became the actual product surface. The model was no longer optimizing for safe language; it was optimizing for explicit ignorance.

The same pattern works on curriculum adaptation. Ask the agent to generate the lesson for a mixed-age group, then immediately require it to name the single age range where the activity will fail. That second instruction pulls the output outside the conservative default and into the specific constraints of the room the volunteer actually walks into.

How the agent mandate shifts competence boundaries

Once the agent is allowed to produce the first draft of any artifact, the ownership question changes. The PM who treats the draft as a finished suggestion keeps the circle of competence exactly where Fable 5 drew it. The PM who treats the draft as raw material that must be stress-tested against one real volunteer’s constraints moves the boundary outward. The difference shows up in the revision log, not in the initial output.

At SermonCentral we learned this when we let the agent draft volunteer instructions for a new series. The first versions were polished and complete. They also assumed every volunteer would read the full document before arriving. When we forced the agent to rewrite after a volunteer read the draft aloud in real time and stumbled on the third paragraph, the instructions shortened, the steps reordered, and the completion rate rose. The model had not become more capable; the team had simply stopped accepting output that stayed inside its original circle.

The boundary keeps moving only if the team measures something the model cannot see. Volunteer completion rate after the first contact is one such measure. Time from assignment to first material request is another. Both sit outside the agent’s training data, so they remain useful signals long after the language inside the agent has been optimized.

Your Turn: Apply This Today

  • Pick one existing Fable Agent workflow that touches volunteers and add the exact line “state the three facts about this volunteer you are least certain about” to the prompt; run it on the next five real assignments and note which facts the model consistently misses.
  • Take the agent’s most recent draft for any volunteer-facing email or instruction set, read it aloud to one actual volunteer within 24 hours, and rewrite only the sentence that caused the first hesitation.
  • Log the time between assignment and first material request for the next ten volunteer tasks; compare the numbers against the agent’s predicted timeline and mark every gap larger than two days.
  • Replace the default “send reminder” step in one orchestration flow with a prompt that forces the agent to name the single constraint that would make the reminder useless, then adjust the send time accordingly.
  • Choose one children’s ministry workflow and require the agent to generate both the main lesson and the single-age-range failure case before any output is shown to a human reviewer.
  • At the end of this week, count how many of the six prompts above produced at least one output that the model originally refused or softened; keep only those revised prompts in the active library.

The same boundary problem shows up in how teams decide what counts as finished once AI writes the first draft; that pattern is laid out in The Months I Outsourced My Sense of What Counts as Finished. It also appears in the decision to keep human judgment visible rather than sprinting toward the next automated feature, which Trust Logs Beat the Next Feature Sprint examines in detail.

I consult with product leaders shipping AI features into volunteer and ministry workflows on prompt boundary testing, completion-rate instrumentation, and first-draft ownership. Let’s talk.

The v1 Where the Model Handed Me the Safe Path

Developer choosing the safe path the AI model generated for a church onboarding flow

The cursor sat on the shared screen while the church tech lead’s finger hovered over the trackpad. One click and the entire volunteer onboarding spec filled the document—every permission tier, every reminder cadence, every fallback for a missed shift. The lines came straight from the model output. They read like something shipped already: clean states, graceful error paths, even a note about mobile push timing that matched the style guide. No one in the room mentioned the actual Saturday line at the welcome desk, the one where parents juggle coffee and name tags while two kids argue over the same blue sticker. The model had quietly chosen the safe path, and no one had noticed.

That paste replaced three sentences the team had written the week before. Those sentences described the moment a new volunteer stands at the table and asks which form to sign first, then waits while the tablet loads the wrong page. The model version never included the wait. It assumed the form would already be open and the signature field highlighted.

The difference showed up later when the same lead tested the flow on a borrowed phone during the actual sign-up hour. The polished steps collapsed at the handoff between printed name tag and digital check-in. The model had optimized for completion rate. It had never seen the table.

Tony Fadell built the first iPod and the early iPhone under the same constraint. He refused to let early usage data or simulated models dictate the v1 shape. Data at that stage only reflected what the team already knew how to measure. Real judgment came from standing in the room with the first users and noticing which physical action they performed before they ever touched the screen. Fadell treated that observation as the only reliable input for the first release.

A model-generated safe path always reads as finished, which is exactly why AI agents skip that observation by default. They optimize for the average documented flow across thousands of prior examples. The result is a v1 that looks complete because every listed edge case carries a solution. The edges that matter locally never appear in the training distribution.

The model removed the awkward handoff because awkwardness does not register as a measurable failure state. In the volunteer flow the handoff lasted eight seconds on average. Those seconds contained the only moment the new person decided whether they belonged in the room. The agent replaced them with an automatic redirect. The redirect worked in staging. It left the actual parent holding a tablet while their child reached for the sticker sheet.

Fadell’s stance makes the cost visible. He would have required the team to keep the handoff in the first version even when every metric suggested removing it. The metric only captured task completion. It missed the decision to return next week. Ministry products live or die on that second decision, not the first tap.

The same pattern appears whenever an agent generates an entire feature branch. The code passes every test the prompt described. It never contains the one condition the builder only notices after watching the real person hesitate.

Re-inserting that condition requires an explicit step the model cannot perform. The builder must name the single lived moment the spec still ignores. That moment becomes the required test before any other acceptance criteria are written. Everything else can stay generated. The one judgment stays manual.

This rule collides with current agent workflows because the safe path already satisfies the prompt. The model will happily generate five alternative versions, each one still missing the same lived detail. Only the builder who stood at the table can supply it. The model has no table.

The practice forces a smaller v1. The smaller surface makes the missing judgment impossible to overlook. It also makes the first user test cheaper. The team ships the awkward handoff on purpose, watches what happens, then decides whether the metric or the moment was right.

The moment the model handed me the safe path

The team had already run the generated flow against their internal checklist. Every required field existed. Every confirmation message matched the tone document. The only remaining task was to map the digital steps onto the physical table that actually existed on Saturday mornings.

When they tried the mapping, the redirect that looked efficient in Figma forced the volunteer to turn their back on the parent for three seconds. That turn broke the single point of eye contact that kept new people from walking away. The model had treated the turn as irrelevant because no ticket described it.

The handoff survived only because one person on the team refused to delete the three original sentences. Those sentences stayed in the spec as a separate acceptance criterion labeled “observed Saturday behavior.” Without that label the redirect would have shipped.

What Fadell’s v1 stance exposes in agent-built flows

Fadell treated early data as a mirror of existing assumptions rather than new information. The same mirror effect appears when an agent builds from prompt alone. The prompt already encodes what the team believes it needs. The generated flow therefore confirms the belief instead of challenging it.

The confirmation shows up most clearly in the edge cases the model adds. They are the edges that appeared in other products. They are rarely the edges that appear at 8:45 a.m. when the children’s check-in line backs up into the parking lot. Those edges only become visible after the builder watches the line form.

Keeping Fadell’s rule means the first version ships with fewer automated steps, not more. The reduced surface leaves room for the observation that the model cannot generate. The observation then becomes the actual scope for iteration two.

Re-inserting the judgment the safe path skips

The judgment usually hides inside a single sentence that describes a physical action rather than a screen state. In the volunteer flow the sentence was: “The parent sets the coffee cup down before they can sign.” That sentence dictated the height of the tablet stand and the size of the signature field. No model output contained it.

The builder must add the sentence after the agent finishes its work. The addition happens in the same document, not in a separate research file. Once the sentence exists, every subsequent change must preserve the action it describes or explicitly replace it with a better observed action.

This single addition prevents the safe path from becoming the only path. It also gives the team a concrete test: does the new volunteer complete the action without turning away? The answer arrives from the table, not from the analytics dashboard.

Your Turn: Apply This Today

  • Open the current v1 spec you generated with an agent this week and highlight every paragraph that describes a screen state instead of a physical action the user performs first.
  • Replace one of those paragraphs with a single sentence written from direct observation of the real moment the user hesitates.
  • Mark that sentence as a non-negotiable acceptance criterion that must pass before any automated edge case is implemented.
  • Run the revised spec past the person who will actually use the flow and ask them to change only that one sentence if it no longer matches what they do.
  • Delete every generated step that now conflicts with the updated sentence, even if the deletion reduces the visible feature count.
  • Ship the smaller version to a single location this week and record whether the observed hesitation disappears or changes form.

The Judgment Call No Dataset Will Make for Your v1 and The Months I Outsourced My Sense of What Counts as Finished both trace the same pattern of handing scope decisions to external sources before the builder has named the one moment that matters.

I consult with product leaders building AI features for ministry tools on v1 scoping, judgment integration, and agent-assisted flows. Let’s talk.

The Compensation Split No Playbook Prepared PMs For

A phone capturing a glowing screen, illustrating the compensation split driving AI-titled PM pay

Product management salary surveys from late 2025 put AI-titled PM roles at a 47 percent premium over generalist PMs with identical years of experience and scope. That gap is the compensation split no playbook prepared product managers for. The surface reading points to straightforward supply and demand for scarce model-tuning skills. The actual pattern shows experienced product people exiting roles that force repeated exposure to messy constraints and moving into lanes where the title itself signals value before any shipped outcome does.

That split does not just change bank accounts. It changes the inputs that shape judgment. People who once spent weeks inside volunteer workflows or global translation queues now optimize inside environments where the feedback loop rewards speed of model iteration over durability of the underlying choice.

Teresa Torres built continuous discovery on the premise that product decisions stay honest only when teams maintain weekly contact with the people who will live with the output. The compensation split severs that contact for a growing group of practitioners before they ever develop the habit.

The Compensation Split When the Offer Letter Becomes the Product Decision

The first decision many PMs now face is whether to accept the AI-labeled role before they have run a single discovery loop inside the domain. Offer letters arrive with equity grants sized for model work, not for the slower work of understanding why a children’s ministry volunteer abandons a lesson plan at minute seven. Accepting the letter becomes the product decision, and every later choice inherits its assumptions.

Teams that hire on title rather than demonstrated discovery practice quickly discover the gap. The new hire ships prompt refinements that look elegant in demos yet fail when the actual user prints the output on a shared church copier with no color cartridge. The compensation signal masked the missing exposure to that constraint.

Continuous discovery requires repeated contact with the constraint before the compensation decision locks in. Without it, the builder optimizes for the environment that paid them to arrive, not the environment the product must serve.

Discovery Habits That Survive the Move Into the Premium Lane

Torres’s framework demands at least one customer conversation per week that can still kill the current roadmap. In the premium lane this habit collides with calendar pressure that treats model experimentation as the only visible output. The builders who keep the habit block the same two hours every week and protect them the way they once protected sprint planning.

One PM who moved from curriculum tools into an AI platform role kept the habit by running weekly calls with the exact volunteer segment he had served before. The calls surfaced that the new summarization feature created more work for users who needed to annotate the summary for doctrinal review. The model team initially dismissed the finding as edge-case; the PM’s prior constraint knowledge let him show the usage volume was not edge at all.

The habit only survives when the PM treats the conversation as non-negotiable output rather than optional research. Compensation that rewards model velocity makes this harder, not easier, because the visible metrics inside the new role rarely track whether the constraint was ever encountered.

Surviving the Compensation Split When the Market Owns Your Title

Once the title itself carries market value, ownership of outcomes becomes harder to claim and harder to lose. The organization assumes the AI PM will deliver model improvements; everything else can be delegated. The builder who wants durable ownership must explicitly contract for the discovery work that sits underneath the model work.

This shows up in roadmap reviews. A generalist PM once had to justify every shipped feature against user retention data. The AI-titled PM can point to token reduction numbers and stop there. Continuous discovery pushes the PM to keep the second slide that shows whether the token reduction changed any user behavior that actually mattered.

Ownership also requires refusing certain promotions. The next title increase often moves the person further from the constraint surface. Several builders have started declining the title bump and negotiating scope and compensation inside the current lane instead. The market still pays, but the builder stays inside the feedback loop that produced judgment in the first place.

Your Turn: Apply This Today

  • Pull the offer letters or role descriptions from your last two compensation conversations and list the three constraints each role explicitly required you to encounter weekly.
  • Map the last three product decisions you owned and mark which ones would still hold if your title and comp changed tomorrow.
  • Block two recurring hours this week for a customer conversation that has the power to kill the current initiative; put it on the calendar before any model experimentation time.
  • Write the second slide you would need in the next roadmap review that connects model output to one user behavior that matters outside the model.
  • Identify the next title or scope increase being discussed for you and list the three constraints you would lose access to if you accepted it.
  • Choose one prior role where you learned a durable constraint and schedule a 30-minute call this month with someone still working inside that constraint.

The same compensation pressure that rewards narrow model skill also makes the broader discovery habit appear optional. Two earlier posts on this blog examined how retention metrics in volunteer tools and enterprise PM patterns in global platforms both depend on repeated contact with constraints that never carry premium pay.

I consult with product leaders on compensation-driven scope decisions, continuous discovery under title pressure, and ownership structures that survive market valuation of roles. Let’s talk.

Trust Logs Beat the Next Feature Sprint

Clasped hands representing the relationship trust logs protect in ministry AI tools

Adding AI features without visible trust logs does not accelerate ministry work. Trust logs, not the next feature, are what keep users accountable. It trains users to treat the system as unaccountable by default.

Church tech teams often assume that faster generation or smarter agents will increase adoption. The opposite pattern appears once volunteers encounter an output they cannot trace or reverse. The absence of a record signals that the tool operates outside the same standards of responsibility expected from every other part of their workflow.

This is the foundational misread that causes product teams to ship capabilities users later refuse to rely on. John Wesley’s three rules offer a clearer test: do no harm, do good, and stay in love with the people served. Applied to product decisions, the first rule requires that no change increase risk to the volunteer or the people they serve. The second demands that every addition demonstrably helps. The third insists the relationship itself remains intact after the feature is used.

Trust logs protect the volunteer handoff, not the model card

A children’s ministry volunteer opens an AI-generated lesson plan on a Thursday night. The content looks usable, yet she needs to know whether the scripture references were pulled from the same translation the church uses, whether prior edits from her team were overwritten, and whether any suggestion came from an outside source that might conflict with her pastor’s guidance. Without a visible log of sources and changes, she cannot decide in the seven minutes she has before the kids arrive.

The log is not a compliance document. It is the only artifact that lets her confirm the plan still belongs to her context. When the log is missing, the volunteer either spends extra time rewriting from scratch or accepts the output while carrying quiet uncertainty into the classroom. Both outcomes reduce completion rates, the metric that actually predicts whether she returns next week.

Teams that treat explainability as a later toggle miss this protection. The feature is not about satisfying an auditor. It is the mechanism that keeps the volunteer’s decision authority intact when the system proposes material she will deliver in person.

Shipping an agent before the trust logs exist creates silent breakage

An agent that schedules follow-up messages or assembles a small-group curriculum operates across multiple steps. When it acts without writing a record of what it changed and why, the next person who touches the same group inherits an altered state with no explanation. The original leader receives no notification of the change. The relationship between the tool and the person responsible for the group is broken at the point of the handoff.

Wesley’s first rule is violated here even though no one intended harm. The agent did not set out to undermine trust. It simply lacked the required record that would have let the original leader review or reclaim ownership. Over repeated cycles, leaders learn to double-check every automated action or to stop using the feature altogether. Either response raises the effective cost of the product.

The pattern repeats across volunteer coordination tools. An agent that books rooms or assigns curriculum without a reversible log produces the same outcome: the person who should remain in charge quietly steps back from relying on the system.

Rollback preserves the relationship only when the record is already public

A mistaken agent action needs reversal. The technical rollback is rarely the difficult part. The difficult part is restoring the volunteer’s confidence that the product will not repeat the error without their knowledge. When the original action left no visible trace, reversal itself becomes another opaque event. The volunteer receives a corrected state without ever seeing what went wrong or why it was fixed.

Wesley’s third rule surfaces here. Staying in love with the people served requires that the relationship survive the correction. A public, queryable record of the action and its reversal allows the volunteer to verify that the issue was addressed and will not recur without notice. Without that record, the correction feels like another autonomous change rather than a restoration of shared control.

Products that add rollback capability after launch discover the cost in support tickets and lost usage. The record must exist before the first agent action ships, because the first failure determines whether the relationship is recoverable.

Frameworks like the NIST AI Risk Management Framework exist because trust logs, not feature counts, are what make an AI system accountable.

Your Turn: Apply This Today

  • Pick the single AI feature closest to release and list every data source, edit, and decision point it touches.
  • Build a read-only log view that any logged-in user can open from the same screen where the output appears.
  • Add one explicit “revert this step” control that writes the reversal to the same log and notifies the original owner.
  • Write the three Wesley-rule questions into the acceptance criteria for the next sprint: does this change risk harm, does it produce clear good, and does it keep the volunteer in control.
  • Run the feature with three actual volunteers this week and require each to explain what the log shows before they accept the output.
  • Remove any agent capability that cannot meet the log requirement before the next release ships.

Two earlier posts on the same theme examined why volunteer completion rates matter more than generation volume and how product teams lose ground when they optimize for model performance instead of handoff clarity.

I consult with product leaders building AI tools for ministry contexts on audit trails, rollback processes, and volunteer handoff records. Let’s talk.

The Months I Outsourced My Sense of What Counts as Finished

A product team reviewing work, reclaiming the definition of finished from the model

For nearly a year I let a model own my definition of finished. I treated the model’s green light as the real one for nearly a year. On three separate products I would write a short prompt, generate a spec or a flow, run it past the team once, and mark the ticket done when the output looked coherent. The first time it cost us was on the children’s ministry curriculum tool. We shipped a lesson-planning flow that the model had declared complete after three iterations. Two weeks later the volunteer completion rate dropped twelve points because the flow assumed every user could read a dense settings panel in under a minute. I had never sat with an actual volunteer to watch them stall on that panel; the model had simply never flagged it.

The second cost showed up in the analytics review. We had defined “finished” as “all core states rendered without error.” By that metric the feature passed. By the metric that actually moved retention it failed. I had outsourced the definition of finished to a system that had no stake in the outcome and no memory of the last five times we had made the same mistake.

How the First Prompt Loop Replaced My Own Taste Test

Industry guidance on a clear definition of done exists precisely because a definition of finished cannot be outsourced to a tool. Charlie Munger’s circle of competence is simple: know the perimeter of what you can judge reliably and stay inside it when the stakes are real. I had started treating the model as an extension of that circle rather than a tool that sits outside it. The first prompt would produce a v1 that looked plausible. Instead of running my own quick test against the actual user constraint I had observed before, I would ask the model to refine its own output. Each loop made the artifact smoother and moved me one step further from the raw edge where judgment is formed.

After four or five loops the artifact felt finished to me because it felt finished to the model. The competence boundary had quietly shifted outward to include whatever the model would sign off on. Munger warned that the biggest losses come from people who do not know where their circle ends. I was not violating the circle through overconfidence in my own knowledge; I was shrinking the circle by refusing to exercise the judgment that defines its edge.

The pattern repeated on the sermon resource platform. I let the model decide when a content-tagging system was ready for beta. It declared the taxonomy complete after we had covered the top thirty sermon topics. What it could not see was that children’s ministry volunteers tag by emotional tone and length, not topic. I had the data from earlier research but never re-checked it against the model’s structure. The model had no circle; it simply extended mine until I stopped noticing the boundary.

The Quiet Handoff of the Definition of Finished to the Model

Once the loop felt efficient, the final call moved without announcement. I stopped keeping a private note of the three things that still felt off before I opened the chat window. The model’s version became the baseline, and my remaining objections had to clear a higher bar than they used to. If the output addressed the original prompt, the objections sounded like nitpicks rather than signals that the prompt itself had missed the real constraint.

This handoff is easy to miss because nothing dramatic happens. The ticket still moves through the same stages. The only change is that the last human who could have said “this is not finished yet” now waits for the model to say it first. Munger’s framework makes the cost visible: you have stepped outside the circle and are using an instrument that cannot tell you when you have done so.

The cost appears in the data weeks later. Retention curves flatten where they should rise. Support tickets cluster around flows the model had called elegant. By then the decision to treat the model’s verdict as final has already been locked into the shipped artifact.

Rebuilding a Human Definition of Finished

The repair started with a simple rule on one small feature: I would write the acceptance criteria in a private note before the first prompt. The criteria had to be testable by a human in under ten minutes and had to reference an actual past failure. Only after the note existed would I open the model. The model could then generate options, but the note remained the standard. If the generated work met the note, it passed. If it did not, the work was not finished regardless of how polished it appeared.

The second step was to keep the note visible to the team. The model’s output became one input among others rather than the default definition of done. Over six weeks the practice spread to two other product areas. The change was not dramatic in velocity; it was dramatic in the quality of the objections that surfaced early.

Munger’s point is not that tools are dangerous. It is that competence is a perimeter you maintain by deliberate use, not a territory you expand by delegation. When the model handles volume, the remaining human work is precisely the work of knowing where the perimeter sits on each decision. That work does not scale with prompt speed; it shrinks when it is not exercised.

Your Turn: Apply This Today

  • Pick one open ticket this week and write the three-sentence definition of finished in a note before you open any AI tool.
  • Run the simplest possible human check against that definition within twenty-four hours of writing it, even if the model version is not ready.
  • Share the note with one teammate and ask them to flag any criterion the current draft still misses.
  • Refuse to mark the ticket ready until the note’s criteria are met, regardless of model polish.
  • At the end of the week, list which criteria the model never surfaced on its own and keep that list for the next planning cycle.
  • Repeat the same sequence on exactly one new decision the following week; do not expand the scope until the habit is automatic.

The same pattern shows up in how teams set completion criteria for AI-assisted discovery work and how they protect early-stage judgment when metrics are still thin. Both are worth reading next if this landed.

I consult with product leaders on AI-assisted decision boundaries, early-stage completion criteria, and protecting judgment when data is thin. Let’s talk.

Why the Ten-Minute Bedtime Rule Stops Working Once You Ship AI Features

A product manager working late, where the bedtime rule fails once shipped AI features keep changing

The bedtime rule, often called the Ten-Minute Bedtime Rule, shows up in product circles through frameworks like those promoted by consultants drawing from books such as “The Effective Executive” and later PM training programs. It instructs teams to close the day by writing a short list of observations, model outputs, or user signals before sleep. The claim is that this tidy capture step turns scattered input into retained knowledge without extra overhead.

The rule breaks once AI features ship because the input no longer arrives as discrete daily packets. Live agents generate continuous streams of partial answers, edge-case refusals, and volunteer rephrasings that arrive at 11 p.m. and again at 6 a.m. The ten-minute window cannot sort which of those outputs already changed the product behavior that same afternoon.

This separation between capture and action worked when product cycles moved in weeks. AI rollouts compress the cycle to hours, so the habit of parking insights for later review creates lag that users feel immediately.

This is the foundational misread that causes product teams to treat learning as an after-hours activity rather than the material that must alter the next deployment.

John Wesley’s three rules offer the needed lens. He told early Methodists to do no harm, do good, and stay in love with God through ordinary disciplines. Applied to product work, the rules become filters that decide which incoming signal requires an immediate change to the shipped system rather than another note in a document.

Why the bedtime rule breaks once live agents ship

At SermonCentral we released a simple agent that suggested age-appropriate rephrasings for children’s ministry scripts. Within two days the agent began softening references to judgment in ways that matched the tone of one large church network but clashed with another. The bedtime list captured the pattern on night one. By night three the mismatch had already reached volunteers who printed the scripts and used them.

The ten-minute capture recorded the observation correctly. It did not trigger any rollback or prompt adjustment because the rule treats the note as sufficient. In practice the agent continued producing the same softened language until a manual review three days later.

Real usage data from the smallest churches showed the problem first. Those volunteers do not file tickets. They simply stop returning to the tool. The bedtime list had no mechanism to surface that drop-off against the agent’s output log.

Wesley’s rules turn raw model output into a decision filter

The first rule, do no harm, forces a check against the most vulnerable user before any output stays live. For the children’s script agent this meant asking whether a suggested rephrasing would confuse a new volunteer who had never taught the story before. If the answer was unclear, the output route changed to a flagged review queue rather than direct suggestion.

The second rule, do good, asks what concrete action the output should enable that day. In the Bible Gateway experiments this looked like requiring every model suggestion to point toward a next step the reader could take on the same visit, such as a short audio clip or a parallel passage. Outputs that only added information without enabling that step were suppressed.

The third rule, stay in love with God through ordinary disciplines, maps to keeping the product tied to the actual rhythm of ministry work. For agents this meant testing whether the suggestion still made sense when the user had only seven minutes between services. Many clever completions failed that test and were removed.

These filters run at decision time, not at the end of the day. They force the product manager to judge the output against live constraints rather than collect it for later sorting.

What replaces the bedtime rule: ownership transfer

The shift is from recording an insight to assigning the change it requires. When an agent surfaces a new refusal pattern, the next step is not a note but a one-line ownership statement that names the person who will adjust the prompt or the guardrail before the next deployment window.

At one rollout the team moved from end-of-day summaries to a shared log that required each entry to include the exact parameter or content rule that would change as a result. Entries without that line were deleted after twenty-four hours. The volume of notes dropped sharply, but the number of actual prompt changes rose.

This transfer also surfaces when the input should not change the product at all. Some model behaviors look novel in the moment yet match patterns already handled by existing volunteer workflows. Wesley’s first rule catches those cases quickly because the test is whether harm would occur if nothing changed.

Even classic guidance on healthy routines assumes a stable input, which is exactly why the bedtime rule fails once shipped AI behavior keeps shifting.

Your Turn: Apply This Today

  • Pick the single AI feature now in production that produces the most daily output and add a one-question harm check to its review log before any new suggestion reaches users.
  • Require every model experiment this week to name the exact deployment parameter or content rule that will change if the experiment succeeds, and delete any experiment notes that omit that line.
  • Run the do-good test on three recent agent outputs by asking whether each one points a volunteer to an action they can finish in the same session; remove any that do not.
  • Review the last forty-eight hours of agent logs against the seven-minute volunteer window and flag any suggestion that requires more context than a printed page can hold.
  • Assign ownership for the top three recurring edge cases you logged this week by writing the owner’s name and the next deployment date next to each case.
  • Delete every capture note older than seven days that has not produced a shipped change, then move the remaining items into the live decision queue.

The CPO Who Captured What the Spreadsheet Missed shows what happens when raw signals stay disconnected from decisions. The Trust Layer No One Budgets For traces how unfiltered agent behavior erodes volunteer confidence over time.

I consult with product leaders on AI feature rollouts, input filtering for live decisions, and ownership transfer in ministry-facing tools. Let’s talk.

The Career Question That Still Needs an Answer After the Promotion

A crowded room of professionals, where each new title raises a career question about real scope

The career question product managers keep asking me after the promotion lands is whether the new title actually expands what they can decide without approval, and it is a career question no raise answers on its own. The answer is no, not unless the scope of their circle of competence has also moved outward in concrete ways.

Most keep measuring progress by the old markers—headcount, budget line, and the size of the roadmap they inherit. Those markers no longer track real ownership once AI tools compress the time between idea and shipped feature.

The Career Question When Compensation and Scope Drift Apart

The promotion conversation usually ends with a number and a new business card. What it rarely surfaces is whether the person now owns decisions that used to sit two layers above them. In practice the gap shows up six months later when an AI-generated prototype lands on their desk and they realize they still need three sign-offs before they can run it with real users.

This mismatch is not unique to faith-tech, but it appears faster here because the user base is smaller and the success metrics are softer. A children’s ministry volunteer who prints a lesson on Tuesday does not file a support ticket when the experience feels off—she simply stops coming back. The PM who cannot change the print layout without another round of reviews never learns what that volunteer actually needed.

Gallup’s research on what actually drives engagement at work reframes this career question well. Charlie Munger’s circle of competence is useful here because it forces the question away from title and toward boundary. Munger described the circle as the area where you can make decisions with high accuracy because you have seen the patterns before. Outside that circle, even smart people produce expensive errors. The promotion does not enlarge the circle; deliberate work does.

Circle of Competence Applied to AI Product Roles

In AI-era product work the boundary matters more because models now generate options faster than teams can evaluate them. A PM whose competence stops at writing requirements will green-light features that look good in a demo and fail with the smallest churches. The PM who has spent time watching volunteers navigate the same workflow on paper knows which generated option will actually survive contact with a seven-minute prep window.

Expanding the circle requires choosing a narrow slice and owning the outcome end to end. At one ministry tool I watched a PM take responsibility for the entire flow from volunteer signup to lesson completion. She removed three internal review steps and measured only whether the printed page reached the classroom. Usage rose because she had narrowed her circle to something she could actually judge without waiting for data that never arrived in time.

The same pattern holds for larger platforms. When the team at a high-traffic Bible site decided to test an AI-assisted reading plan, the PM who succeeded was the one who had previously owned the print PDF version used by small groups. She already knew which formatting constraints mattered. The model suggestions were filtered against that lived constraint rather than against abstract engagement metrics.

The Career Question the Old Faith-Tech Playbook Cannot Answer

The inherited playbook says move up by managing larger teams and bigger roadmaps. In ministry tools that path often leads to more meetings and less contact with the actual user. The volunteer who prepares a lesson in a church basement does not care about the PM’s headcount; she cares whether the material fits on one page and survives a spilled coffee cup.

AI accelerates the break because it removes the old friction of building. A feature that once took a quarter can now ship in a week. The limiting factor becomes judgment about whether the feature should ship at all. That judgment only improves inside a competence circle that includes direct observation of the user under real constraints.

Teams that keep promoting on the old signals end up with leaders who can describe strategy but cannot tell whether a generated outline will work for a volunteer with no formal training. The organizations that notice the drift are the ones that start asking different questions in promotion reviews: what specific decision boundary has this person moved in the last twelve months, and what evidence shows the boundary is now reliable.

Your Turn: Apply This Today

  • List every current project and mark the single decision on each that you cannot make without another person’s approval; choose one to own fully this quarter.
  • Shadow one user completing their actual task with the current product for at least ninety minutes and record the exact moment they adapt or abandon the flow.
  • Write a one-paragraph definition of the competence boundary you claim today and share it with your manager before the next role conversation.
  • Run one AI-generated option through your existing user constraint set and reject it on paper before any build work begins.
  • Identify the last three features that shipped under your name and note whether you personally observed the user outcome or relied on secondhand reports.
  • Schedule a thirty-minute review with a peer PM who works on a different product line and compare where each of you can act without escalation.

The CPO Who Captured What the Spreadsheet Missed and The CPO Who Started Asking What Kept Her Volunteers Up both trace the same pattern: real ownership grows only after someone redraws their own competence boundary.

I consult with mid-career product leaders on mapping competence boundaries, AI-era role decisions, and ownership growth in mission-driven tools. Let’s talk.

The Judgment Call No Dataset Will Make for Your v1

A network of connected nodes representing the model scores behind a v1 judgment call

The screen glowed under the dim office lights as the triage agent pushed its first batch of cards, and the judgment call it could not make was already waiting. The children’s director’s eyes narrowed at the top result, a flagged prayer request from a single mom whose kid had missed three Sundays. She tapped override without checking the 87 percent score, then two more before the next refresh cycle even started.

Her coffee had gone cold beside the keyboard. The room stayed quiet except for the low hum of the laptop fan. She saved the batch and stood up, already thinking through which volunteer would see the corrected version first thing tomorrow.

That single sequence of clicks showed the gap no model can close on its own. The scores measured pattern fit, not the weight of the actual people on the other side. When the override happens this early, it reveals the constraint that matters most for any v1: the builder still has to carry live context the dataset never captured.

This is the exact spot where over-reliance on AI starts to erode the judgment that later versions depend on. Teams begin treating high confidence numbers as permission to stop asking the harder questions. The product ships with cleaner logs but thinner instincts.

The judgment call Solomon would recognize

Solomon’s judgment offers the clearest lens here. Two women claimed the same living child. No dataset or probability score settled it. Solomon forced the decision into the open by proposing to split the child, then watched which response revealed the real mother. The test did not measure accuracy against past cases. It created a live situation that exposed the difference between claim and care.

Applied to agent handoffs, the same principle holds. When an AI triage system surfaces a recommendation, the override moment functions like Solomon’s sword. It forces the builder to decide whether the model’s pattern match aligns with the actual stakes. Skip that step repeatedly and the team loses the ability to notice when the pattern itself is incomplete.

The children’s director’s overrides worked because she still held the names and the absences in her head. The agent had access only to attendance flags and text strings. Her decision loop included the detail the model could not weight. Without that loop, the next release would optimize for cleaner data rather than better ministry outcomes.

Solomon did not ask for more evidence first. He set a condition that made the real priority visible in real time. Modern agent systems invert this. They present polished outputs and invite teams to accept them unless something obvious breaks. The result is a gradual shift where builders review less and ratify more.

Why your v1 needs a protected judgment call

Product teams building v1 agents for ministry tools see this pattern most clearly when usage data starts rolling in. The metrics look stable because the model handles volume. The edge cases that matter, the ones involving real volunteers and irregular schedules, only surface when someone still practices the override habit.

NN/g research on keeping a human in the loop explains why a deliberate judgment call beats a confident model score. The constraint tightens as models accelerate. Faster inference means more decision cards appear before a builder finishes evaluating the first one. Judgment slots, the mental space required to hold context and make the call, stay fixed. They cannot scale with token speed.

Teams that protect those slots treat overrides as scheduled work rather than exceptions. They block time after each agent run to review a small set of cases without the scores visible. The practice keeps the builder’s sense of what counts sharp instead of letting it atrophy.

The same protection shows up in how handoff points get designed. Instead of routing every low-stakes item through the model, the system leaves certain categories for direct human entry. This preserves the live context that later training data will need. Without it, subsequent versions inherit the blind spots of the first release.

Safeguards that keep the human in the loop

Another safeguard comes from requiring the builder to write a short note on each override before the card closes. The note forces articulation of the missing variable. Over weeks those notes become the clearest signal of where the model still needs human weight.

Protecting judgment also means limiting the number of agent-driven decisions any single builder reviews in one sitting. Fatigue turns overrides into rubber stamps. Capping the load keeps each one deliberate.

Finally, the practice extends to how teams document wins. When an override later proves correct in follow-up conversations with users, that outcome gets recorded with the same detail as model successes. The record reinforces that judgment remains the durable asset.

Your Turn: Apply This Today
– Block thirty minutes on your calendar this week to review five agent outputs with the scores hidden before you check them.
– Pick one category of decisions your current v1 agent handles and route it back to direct entry for the next sprint.
– After every override this month, write one sentence explaining the variable the model missed and store it in a shared note.
– Limit yourself to reviewing no more than eight agent cards in any single session for the next two weeks.
– At the end of this week, pull the three overrides that felt hardest and discuss them with one teammate without referencing the scores.
– Schedule a recurring thirty-minute slot every Friday to read the override notes from the prior week out loud to yourself.

The CPO Who Captured What the Spreadsheet Missed shows how one leader built a parallel record of the calls models could not make. The Trust Layer No One Budgets For traces what happens when those records never get created.

I consult with product leaders building AI agents for ministry tools on v1 judgment practices and override design. Let’s talk.

The Manila Kitchen Table Where an App Finally Worked

A team running a screenshot-driven discovery session, using real images and notes instead of a written product spec

A team running a screenshot-driven discovery session, using real images and notes instead of a written product specThe volunteer coordinator’s laptop sat open on the scarred wooden table, its screen reflecting the overhead bulb. Her youngest reached across for another piece of bread while the older one asked about homework. She ignored both long enough to drag three phone photos into the chat window and type the phrases she had heard at the door that morning. What followed was a small lesson in screenshot-driven discovery: raw images and real words, not a written brief, shaped the next build.

Claude generated a new check-in layout in the reply. She copied the suggested changes straight into the test build her developer had left running on the side. The corrected flow loaded without the old crash. She tested it once, nodded, and closed the laptop to finish dinner.

Teresa Torres built continuous discovery on the idea that product teams must keep direct contact with users through regular, lightweight probes rather than periodic big-bang research. The Manila scene shows what that practice looks like when the person doing the probing has no formal product training and the tool in front of her is an LLM instead of a research platform.

The same principle scales when the constraints arrive as images and spoken fragments rather than translated tickets.

The screenshot-driven discovery loop that replaced the product spec

Most ministry tools still begin with a written brief that someone must first interpret. The coordinator skipped that step entirely. She sent the actual error states and the exact phrases parents used. The model treated those artifacts as the source of truth.

Each iteration stayed inside the same chat thread. She would test the suggestion on Sunday, photograph the new failure point, and paste it back the following week. The document that mattered was the running conversation itself, not a separate requirements file that would have aged out after the first service.

Continuous discovery insists on small, frequent signals over large, infrequent studies. Here the signals were visual and verbal, captured at the moment of use. No one had to schedule a discovery sprint or book a research lab. The loop ran on whatever device was already on the kitchen table.

Why spoken parent language beats polished requirements

Polished requirements smooth away the friction that actually breaks the experience. When the coordinator typed “the line was too long for the little one,” the model kept that constraint visible. It did not translate the complaint into abstract language about throughput or queue management.

Spoken fragments carry the emotional weight and the physical reality at the same time. A parent holding a toddler cannot wait in a straight line. That single sentence forced the layout to include a side bench and a second volunteer who could greet families before they reached the tablet. No requirements document written later would have recovered that detail with the same precision.

Torres warns against losing the raw data behind layers of interpretation. The LLM here acted as a direct transcription layer rather than an additional interpreter. The original words stayed in the prompt, so the generated interface stayed tethered to the real constraint.

Shipping the version that survived one real Sunday morning

After three weeks the build reached a state where it handled the busiest arrival window without crashing. The coordinator did not wait for a full regression suite. She watched the door the next Sunday, noted the three families who arrived together, and confirmed the flow stayed intact.

That single morning became the acceptance test. Everything else—edge cases, accessibility audits, future roadmap items—remained secondary until the next failure showed up in a new set of screenshots. The product advanced only when the real environment confirmed it.

Continuous discovery measures progress by the reduction of surprise in live use, not by completion of a spec. The coordinator never claimed the app was finished. She only claimed it had survived the most recent Sunday without creating new problems for the families at the door.

Your Turn: Apply This Today

  • Pick the single workflow that broke last Sunday and open a fresh chat with your AI tool before you write any description.
  • Take three screenshots of the exact failure states on the device the volunteer actually uses and paste them into the chat with no additional summary.
  • Transcribe the two or three sentences users said out loud when the flow failed and add those lines beneath the images.
  • Generate the revised screen inside the same thread and push the change to a test build the same day.
  • Run the new build on the next live event and photograph whatever still breaks before you touch the prompt again.
  • Repeat the loop once a week for four weeks, keeping every screenshot and spoken line in one running thread so the history stays intact.

The Tuesday the Children’s Director’s Inbox Became the Real Product Spec and The Fitness App a Pastor’s Kid Shipped Without Writing Code both trace the same pattern of letting raw signals drive the next build instead of waiting for translated requirements.

I consult with faith-tech product leaders and ministry builders on screenshot-driven discovery loops, keeping spoken user language as the primary spec, and running continuous discovery cycles that fit around existing volunteer schedules. Let’s talk.

The Year I Kept Waiting for Usage Numbers That Never Came

An analytics dashboard showing usage data that an AI ministry tool team waited on instead of acting on volunteer feedback

I spent fourteen months refusing to greenlight an AI outline generator for children’s ministry because the projected usage numbers never materialized in our internal dashboards. Every sprint review I asked the same question: show me the cohort that actually opens the tool more than once. The answer stayed flat. We had real volunteer feedback in the form of one-off emails and hallway comments, but I treated those as noise until the telemetry caught up. The mistake was treating usage data as the only signal that counted and the qualitative input as noise.

The cost showed up later. Two other teams shipped lighter versions of the same idea while we waited. Their adoption came through the exact channels we had dismissed as anecdotal. By the time our version reached the same directors, the window for shaping how they used AI had already closed.

Taste as the first filter before any usage data

Volunteer tools rarely produce clean usage signals until the workflow already matches how people actually prepare on Tuesday nights. The children’s director prints the lesson at the kitchen table, marks it with a highlighter, and hands the marked copy to the small-group leader on Wednesday. No login, no repeat session recorded. The data stays silent because the real use happens offline.

I used to treat this silence as a verdict against the feature. Now I treat it as evidence that the initial test must happen through direct observation rather than aggregated logs. One director’s reaction to a printed sample carries more weight than a month of anonymous click data that never arrives.

The practical shift is to route early decisions through taste first. Build the smallest artifact that one practitioner can hold and judge inside their existing rhythm. If it survives that single judgment, then instrument it. Not the other way around.

How Wesley’s rules expose where data stays silent

John Wesley gave the early Methodists three simple rules: do no harm, do good, and attend to the means of grace. Applied to product work, the first rule means refusing to ship something that adds friction to an already thin volunteer margin. The second means looking for concrete help that fits the actual preparation window. The third means paying attention to the repeated practices that sustain the work week after week.

Data dashboards are excellent at measuring harm once it scales. They are nearly useless at detecting whether a tool does good inside a single, unrepeated Tuesday night routine. That judgment requires the same kind of attentive presence Wesley expected of class leaders who visited homes rather than waiting for reports.

The rules therefore reorder the sequence. Taste informed by direct contact becomes the first screen. Quantitative thresholds come later, once the practice has already taken root in one location.

When to ship the version that feels right to one children’s director

We eventually released a minimal outline generator after a single director walked us through her exact Sunday morning handoff. She showed us the folder she carries, the three places she writes notes, and the moment she decides the material will not work for her group. That thirty-minute walkthrough replaced six months of waiting for cohort metrics.

The version we shipped matched her folder structure and printed cleanly on standard paper. Usage numbers appeared only after the tool was already in rotation with three other directors who heard about it through the same informal channel. The telemetry finally moved because the fit had already been tested in one real context.

Shipping on that kind of single-point confirmation feels reckless until you accept that volunteer contexts keep most usage invisible by design. The alternative is to keep polishing a tool that never enters the only workflow that matters.

Your Turn: Apply This Today

  • Pick one AI feature decision currently stalled on usage projections and write the single sentence that describes the exact moment a children’s director would first hold the output in her hands.
  • Find three volunteers who match the target context and schedule a thirty-minute walkthrough this week where they show you their current preparation materials without any new tool present.
  • Print or mock the smallest possible version of the feature that fits inside their existing folder or notebook and watch where they mark it or set it aside.
  • Note the exact point in their rhythm where the artifact either reduces a step or adds one, then adjust the mock before showing it to anyone else.
  • Send the revised mock to those same three volunteers with one question: does this replace something you already do or sit on top of it?
  • If two of the three say it replaces a step, ship the narrow version to their group only and measure what actually changes in their handoff rather than waiting for broader telemetry.

The same pattern shows up in The Tuesday the Children’s Director’s Inbox Became the Real Product Spec and The 1997 Lesson AI Product Teams Keep Missing.

I consult with AI product leaders and ministry tool teams on taste-driven early decisions, volunteer workflow observation, and when to release before metrics appear. Let’s talk.

The Feature That Got Easier to Build But Still Never Reached the Smallest Churches

Volunteer in a small sanctuary, the kind reaching the smallest churches depends on

The call came while I was still clearing breakfast dishes. A pastor in western Pennsylvania had one question: the AI tool that could write his entire children’s ministry curriculum in ten minutes had already been built, so why was his volunteer team still printing last year’s worksheets at the kitchen table? Reaching the smallest churches, it turned out, had nothing to do with how fast the feature was built.

He wasn’t asking about prompts or pricing. He wanted to know how the file was supposed to reach the three other volunteers who only ever showed up on Sunday morning, none of them checking product stores or reading release notes. The tool had done its part. Everything after that stayed exactly the same.

That single gap between “we shipped it” and “it actually reached the smallest room” is the part no one has fixed yet, and reaching the smallest churches lives entirely inside it.

Print-first constraints remain the real filter. A volunteer who prints the lesson at home on Thursday night does not care that the underlying model ran on a rented GPU cluster. If the output still requires reformatting to fit an 8.5-by-11 sheet and a three-ring binder, the adoption path is identical to the one that existed before the agent existed.

Completion rate inside that binder is the only metric that travels back to the builder. Every other telemetry—sign-ups, demo requests, feature usage—stops at the church office door. The cost reduction therefore changes the speed of iteration for the builder while leaving the observable result for the smallest churches unchanged. Reaching the smallest churches means winning that binder, not the dashboard.

The two distribution paths that decide reaching the smallest churches

Ministry tools travel either through denominational or regional trust nodes or through direct volunteer-to-volunteer referrals. The first path requires explicit endorsement from a person already trusted by multiple congregations; the second requires the tool to survive one successful handoff between two people who already know each other.

App-store or website discovery plays almost no role. A pastor searching for “AI children’s lesson” will still ask the children’s director at the next association meeting whether anyone has tried the new thing. That meeting is the actual product review cycle. Builders who optimize only for the website conversion funnel never reach the table where the decision is made.

Both paths are slow by design. They exist to reduce risk for volunteers who have limited time and zero tolerance for tools that fail on Sunday. Speeding up the build phase does not compress the trust verification phase that follows.

Munger’s Inversion Test Applied to an Agent Rollout

Charlie Munger’s latticework requires asking what would have to be true for the opposite outcome to occur. Instead of asking how to get the agent adopted, the useful question is what would have to be true for the smallest churches to never install it even after it is free and technically perfect.

The answer points to missing handoff artifacts. No printed one-page instruction that fits inside the existing binder. No verbal script the director can repeat in thirty seconds to a volunteer. No version that survives a lost internet connection on Saturday night. Each missing artifact is a reason the tool stays on the drive.

Inversion also reveals the second-order effect. When the agent is easy to build, teams ship more versions. Each new version increases the cognitive load on the person who must decide whether to reprint the lesson packet. The proliferation itself becomes a distribution tax rather than a distribution benefit.

Your Turn: Apply This Today

  • Pick one current AI feature already in pilot and write the exact three-sentence script a children’s director would say to a volunteer at the end of a planning meeting.
  • Print that script on the same sheet as the generated lesson and test whether the volunteer can complete the activity without opening a second tab or device.
  • Map the next two people who must physically receive the printed sheet after the director approves it and note how many days typically pass between each handoff.
  • Remove every UI element that requires the volunteer to log in or download an app; replace it with a single QR code that prints at the bottom of the page and opens the content in a browser.
  • Run the inversion question with the actual volunteer: “What would have to be true for you to throw this page away and use last week’s lesson instead?” Capture the first answer verbatim.
  • Schedule the next iteration of the agent only after the printed artifact survives one full Sunday with the volunteer who originally gave the answer above.

Distribution Moats Still Decide Which Ministry Tools Actually Reach Churches laid out the trust-node mechanics in more detail. The 1997 Lesson AI Product Teams Keep Missing showed how earlier technology shifts produced the same pattern when builders ignored the final handoff layer.

I consult with AI product leaders and ministry tool builders on distribution path mapping, trust-node verification, and agent rollout constraints. Let’s talk.

The CPO Who Captured What the Spreadsheet Missed

Children's ministry volunteers in a classroom, the kind of continuous discovery work that never reaches the login data

Children's ministry volunteers in a classroom, the kind of continuous discovery work that never reaches the login dataA product team tracking 2,400 feature requests last quarter reported that 68 percent came from power users who logged in at least three times a week. The obvious reading was that the backlog now reflected the clearest priorities. The data said the opposite once the team examined who never appeared in the logs at all. Continuous discovery exists to close exactly this gap between what the data records and what the work actually requires.

The actual product owners were children’s ministry volunteers who printed materials on Tuesday nights and never created accounts. Their constraints never reached the spreadsheet. Continuous discovery, as Teresa Torres frames it, requires ongoing contact with the people whose work the software either enables or obstructs. Without that contact, the numbers simply ratified the loudest voices already inside the system.

This gap is not unique to church tools. Any product whose heaviest users operate outside the login flow will produce the same distortion. Surveys and usage metrics both miss the work that happens on paper, in cars, and in the ten minutes between the last child leaving and the volunteer locking the door.

The Tuesday night list that replaced the feature request board

The team stopped maintaining the public request board for six weeks. Instead they asked three volunteers to keep a running list on paper of every task they performed while preparing the next week’s materials. The lists arrived photographed and timestamped each Wednesday morning.

One volunteer wrote “find craft supplies that match the story” three weeks in a row. The feature request board had never contained that phrase. The team had instead built tagging improvements that assumed volunteers already knew which assets existed. The paper lists showed the search happened before the volunteer ever reached the digital library.

Another note read “print two copies because the first one always jams.” That constraint pointed to a print-preview problem, not a content problem. The usage data had shown high print-button clicks but zero time spent on failed jobs because failed jobs never generated a logged event.

How continuous discovery surfaced the real constraint

Volunteers often answered interview questions with single words: “easy,” “quick,” “simple.” Those words aligned with the product team’s own vocabulary for reducing clicks. Torres’s continuous discovery habit of asking “what does that look like on Tuesday night” forced the next question.

When pressed, “easy” turned out to mean “I can finish this while the kids are still eating snack and before the parents arrive.” The real constraint was not interface friction but calendar friction. The product had optimized for session length when the volunteer was optimizing for total elapsed time between arriving at church and leaving with everything in hand.

The same pattern appeared in the word “quick.” It described the window between the end of the service and the start of small-group setup, not the speed of any single screen. Once the team mapped the actual sequence, they removed two approval steps that had existed only to satisfy internal stakeholders who never saw the Tuesday timeline.

The single change that surfaced ownership questions

The team added one required field to every support ticket and every research note: “Who owns the outcome if this works?” The field could not be answered with a job title. It had to name the person who would notice if the change succeeded or failed.

Most answers pointed to volunteers who had never logged in. The children’s director became the named owner for curriculum readiness. The volunteer coordinator became the owner for print reliability. These names had never appeared in any prior analytics dashboard.

The change also revealed that two popular feature ideas had no owner at all. One idea improved search for pastors who prepared their own lessons; the other improved reporting for denominational staff. Neither group performed the Tuesday night work that kept the product alive week to week. Both ideas were deprioritized without debate once ownership was named explicitly.

Your Turn: Apply This Today

  • Pick three volunteers who have never created an account in your current tool and schedule a single 45-minute session this week at the exact time they normally prepare materials.
  • Ask each person to bring the physical artifacts they used last Tuesday—printed pages, handwritten notes, or screenshots of texts—and photograph those artifacts before the conversation starts.
  • During the session, have them walk through the sequence out loud while you write only the exact phrases they use, without translating them into product language.
  • After they finish, ask who would notice first if any step failed and write that person’s name next to every constraint they described.
  • Within 24 hours, send each volunteer a one-paragraph summary of the constraints they named and ask them to correct anything that misstates their actual Tuesday night.
  • Remove every item from your feature request board that has no named owner among the three volunteers you just met.

The same listening pattern appears in the children’s director’s inbox becoming the real product spec and in the 1997 lesson that still trips AI product teams today. Both posts trace how unlogged work surfaces only when someone stands beside the person doing it.

I consult with product leaders and ministry technology teams on continuous discovery with non-logged-in users, naming ownership of volunteer outcomes, and replacing feature boards with Tuesday-night constraints. Let’s talk.

The Trust Layer No One Budgets For

A laptop workflow where an AI audit trail records who reviewed each AI-generated output before use

A laptop workflow where an AI audit trail records who reviewed each AI-generated output before useMost product teams treat verification layers as the thing that slows an AI rollout, but ministry contexts show the opposite pattern: tools without explicit trust mechanisms reach a usage ceiling within weeks and never recover. The missing piece is almost always an AI audit trail: a record, inside the workflow, of what the model produced and who verified it.

This misread turns every efficiency gain into a hidden liability. Teams ship faster prompts and cleaner interfaces, then watch completion rates stall because the output cannot be trusted in real volunteer or pastoral workflows.

Jensen Huang’s sovereign AI argument supplies the lens. Huang insists that control over the stack is not optional once the system handles decisions that matter; dependence on external outputs without ownership creates fragility that no amount of speed can offset. The same principle applies to faith-tech products: the organization using the tool must own the verification step or it cedes authority over the very outcomes it claims to serve.

Where the AI audit trail actually lives

Most AI features in church tools generate content and stop. The prompt runs, the text appears, and the user decides whether to keep or discard it. In practice the decision never gets recorded, so the next volunteer or staff member has no way to know what was reviewed and what was accepted on faith.

The AI audit trail is not a log file sitting in an admin dashboard. It lives inside the workflow where the output is used. For curriculum tools like Sermons4Kids this means every lesson plan that reaches a volunteer must carry the record of who prompted the AI, which sources were cited, and whether a human editor confirmed the scripture references before printing. Without that record the product quietly shifts risk onto the least trained user.

Teams that add the trail after launch discover it changes retention more than any interface tweak. Volunteers finish the task because they can see the verification step was already performed. Completion rate, not session time, becomes the metric that actually moves.

The fallback that protects volunteer ownership

Ministry work is performed by people who do not own the data model. A children’s director using an AI-generated activity needs a path back to a human-authored version when the suggestion misses the mark. That path is the fallback, and it must exist before the feature ships.

The fallback is not a reset button. It is a parallel, non-AI route that preserves the volunteer’s ability to complete the task without depending on the model being correct. In practice this looks like one-click access to the last human-reviewed version of the same curriculum, with the AI output marked as optional rather than default.

Without the fallback the product forces volunteers into a position of either trusting the output or abandoning the task. Sovereign control means the team building the tool decides in advance that the human route remains the primary one. The AI augments; it does not replace the ownership the volunteer already holds.

Why tooling roadmaps keep skipping the verification step

Roadmaps prioritize visible capability because visible capability wins budget. A new generation model or a smarter prompt template shows immediate progress in a demo. Adding an audit trail or a verified fallback does not photograph well and therefore rarely survives the prioritization meeting.

The pattern repeats across faith-tech products. The feature that would let a pastor confirm an AI sermon illustration against the actual text never makes the quarter because the team is still chasing the next model upgrade. Usage data later reveals the problem: the illustrations look good in the interface but produce the same flat adoption curve seen in tools that skipped verification two years earlier.

Huang’s point on sovereignty reframes the tradeoff. The cost of adding the verification step is paid once in velocity. The cost of skipping it is paid continuously in lost trust and stalled usage. Ministry contexts make the second cost visible faster than consumer apps because the users cannot simply switch tools when the output fails; they have to keep serving the same people with whatever the product supplies.

Your Turn: Apply This Today

  • Pick one existing AI workflow in your current product and add a visible confirmation checkbox that records the reviewer’s name and the date before the output is marked ready for use.
  • Identify the single most common output failure reported by volunteers in the last thirty days and build the one-click fallback to the last human-reviewed version of that asset.
  • Move the verification step from an optional admin setting to a required part of the core flow so every new user encounters it on first use.
  • Export the last ten AI-generated items from your staging environment, run them through the proposed audit trail, and measure how many would have required a human correction.
  • Replace one roadmap item that adds model capability with the verification layer for the most-used prompt template already in production.
  • Share the resulting audit record format with two ministry leaders who use the tool weekly and adjust the fields based on what they say they actually need to see.

The same principle of owned verification appears in the 1997 lesson AI product teams keep missing and in the way the children’s director’s inbox became the real product spec. Both posts trace how control over the small, unglamorous steps determines whether the larger system stays usable.

I consult with product leaders and ministry technology teams on verification layers, fallback design, and roadmap prioritization that actually moves ministry usage metrics. Let’s talk.

The CPO Who Started Asking What Kept Her Volunteers Up

A product leader running volunteer discovery interviews with a ministry team

I watched a product lead refresh the same volunteer dashboard three times in one night. Not because the numbers had changed. Because one row showed a woman in Ohio who’d opened the lesson plan at 9:14, closed it at 9:21, and never clicked “I’m in” again. That one row is why she traded dashboards for volunteer discovery interviews.

That gap between finishing the plan and deciding whether to come back is where the real work happens. Everything else—agent pilots, sentiment scores, the roadmap the exec team keeps waving—sits on top of it like fresh paint over a crack.

She started running volunteer discovery interviews instead of polling the analysts. Simple questions. What made you hesitate? What were you staring at when you almost said yes? The answers didn’t fit on any slide. They just kept showing up.

The volunteer discovery interviews that replaced the roadmap review

Your weekly sync with the children’s director used to start with feature status. You swapped the first ten minutes for one standing question: “What kept a volunteer from finishing the lesson this week?”

The answers never matched the AI wishlist. One director described a volunteer who spent twelve minutes hunting the right craft supply list because the print button sat three clicks deep. Another mentioned the parent who texted at 9 p.m. asking for the memory verse because the app only showed it inside the full lesson flow. These were not nice-to-haves; they were the moments that decided whether the volunteer opened the app again.

Torres’s method insists the interview happens while the work is fresh. You started requiring the director to bring one screenshot or one printed page that showed the exact friction. The specificity forced the team to stop theorizing about volunteer experience and start seeing it. She scheduled the volunteer discovery interviews while the frustration was still fresh.

How One Captured Pain Turned Into a Narrow Shipped Tool

A single volunteer’s email thread became the artifact. She had printed the lesson at home, realized the take-home activity required glue sticks she did not own, and spent the next morning driving to three stores before the kids arrived. The director forwarded the thread with the subject line “this is why we lose people.”

You scoped a two-week build: a one-tap “what do I need to buy” list that pulled only the items marked as supplies, formatted for a phone note or a printed half-sheet. No new AI. No new dashboard. The tool shipped to 12 pilot churches. Volunteer completion rate on those lessons rose 19 percent in four weeks. The signal was narrow enough that the build stayed small and the measurement stayed honest.

The Weekly Rhythm That Keeps the Signal Fresh

You now run a 45-minute Friday call limited to three people: the children’s director, the support lead who sees the most tickets, and you. Each person arrives with one raw artifact—no slides, no summaries. You watch the artifact together for eight minutes, then spend the rest of the time writing the smallest possible test that would remove that exact friction.

Nothing goes on the roadmap until it survives two consecutive weeks of this filter. The discipline removes the political weight of executive requests and replaces it with a growing list of narrow, volunteer-validated problems. The AI pilots still exist, but they now compete for the same scarce build time as the glue-stick list. Most of them lose.

Your Turn: Apply This Today

  • Pick the one children’s director or volunteer coordinator you already email weekly and ask her to forward one volunteer message or screenshot from the last seven days before Friday.
  • Block 30 minutes on your calendar this week to watch that single artifact with her on a shared screen and write the smallest test that would remove the friction.
  • Ship the test to no more than three churches and define success as one observable change in volunteer completion rate, not in NPS or sign-ups.
  • Cancel the next roadmap review meeting and replace it with the same 45-minute artifact review, keeping the attendee list to three people max.
  • Write down the exact question you will ask every week for the next four weeks and put it in the calendar invite so it cannot drift.
  • Choose one AI pilot currently on the roadmap and ask the team to name the volunteer frustration it would remove; if they cannot name one from an actual artifact, move it to the bottom of the list.

The pattern in “The Tuesday the Children’s Director’s Inbox Became the Real Product Spec” and “The Fitness App a Pastor’s Kid Shipped Without Writing Code” shows the same result: the smallest shipped change that comes from watching real labor beats the largest planned feature that never leaves the slide deck.

I consult with product leaders at ministry organizations on continuous discovery rhythms, volunteer friction capture, and narrow tool scoping. Let’s talk.

Trust Mechanisms Beat the Next Tooling Sprint

The tooling acceleration thesis in the latest a16z AI memos tells product teams to prioritize feature velocity above all, but the teams that keep ministry partners are the ones investing in trust mechanisms instead. It frames every ministry deployment as a race to embed more generation, summarization, and recommendation models before competitors do. The argument assumes that once the models are live, adoption will follow because the technical capability already exists.

That framing collapses for anyone who has watched a children’s ministry coordinator delete an entire AI-assisted curriculum platform after three weeks. The memo writers never model the moment a volunteer realizes the output cannot be audited against the actual lesson plan she printed on Sunday night. When that happens, the velocity metric turns negative: the team has spent engineering hours accelerating the exit.

Jensen Huang’s sovereign AI argument supplies the missing lens. Control is not a later optimization; it is the precondition that determines whether any AI layer survives first contact with real ownership.

The rollout that shipped features but lost the coordinator

Last spring a mid-size publishing house shipped an AI tagging system across its children’s resources. The model suggested age-appropriate verses and activities in under four seconds. The product lead celebrated the completion rate spike in week one.

By week three the volunteer coordinator in one region had turned the feature off for her entire network. She could not verify whether a suggested story about forgiveness aligned with the specific translation her church used in print. The system offered no provenance link and no way for her to override the suggestion without opening a support ticket. She reverted to the old shared spreadsheet because it let her keep the final say in her own hands.

The team had optimized for model output speed, not for the single point where authority actually transfers. Once that transfer failed, the rest of the feature set became irrelevant.

Why tooling velocity hides the trust mechanisms that transfer ownership

Most AI roadmaps still measure success by the number of prompts routed through the new model. That number rises quickly when the interface is clean. It says nothing about whether the ministry leader who must sign off on the output now treats the system as an extension of her own judgment.

Huang’s point about sovereignty translates directly here. A model hosted and governed by someone else can generate text, but it cannot confer verifiable custody. When the underlying data, the override rules, and the audit trail all live outside the ministry’s control, the leader is renting capability rather than owning the decision chain. Velocity only delays the day she stops renting.

Product teams notice the drop-off months later in churn reports. By then the engineering calendar has already moved on to the next model upgrade.

What trust mechanisms actually look like in practice

A trust mechanism is any small, verifiable control that lets the actual owner confirm or correct the output before it reaches the end user. In one deployment the team added a single required step: every AI-suggested children’s story had to be opened, read, and explicitly approved by the coordinator inside the same interface before the lesson could be printed. The approval created a permanent record tied to her account.

The change added thirty seconds to the workflow. Completion rates held steady because coordinators now treated the output as draft material under their authority rather than as an unexamined suggestion. The override log later became the primary data source for improving the model, not the raw generation metrics.

Another team exposed the source attribution for every verse the model pulled, limited to the exact translation editions the church had already licensed. When a suggestion referenced an unlicensed paraphrase, the interface surfaced the licensing gap immediately. The mechanism did not slow the model; it made the boundary conditions legible to the person who would be held responsible.

Decades of research on the neuroscience of trust show why trust mechanisms outperform raw tooling velocity for sustained adoption.

Your Turn: Apply This Today

  • Pick one existing AI workflow your team shipped in the last quarter and list every output that reaches a ministry leader without an explicit approval or override step.
  • Choose the smallest verifiable piece—usually a single required confirmation or source link—and wire it into the flow this week so the owner must act before the content moves forward.
  • Instrument that confirmation as a logged event tied to the specific user role, then check the log after seven days to see whether overrides cluster around particular content types.
  • Remove any generation step that cannot be traced to a licensed or approved source within the ministry’s existing agreements.
  • Run the revised flow past one actual coordinator who was not part of the original build and ask only whether she can now defend the final output to her volunteers.
  • Document the exact change and the resulting override rate; use that single number as the input for whether the next tooling sprint is even worth scheduling.

The 1997 Lesson AI Product Teams Keep Missing shows how similar control gaps played out with earlier digital tools in churches. The Tuesday the Children’s Director’s Inbox Became the Real Product Spec captures the moment ownership actually changes hands.

I consult with product leaders in faith-tech on AI trust layers, ownership transfer, and verifiable rollout controls. Let’s talk.

The Year I Let Data Override Every Ministry Judgment Call

Tech office workers at computers representing faith-tech AI staffing challenges

I spent a full year insisting that every AI feature decision in our ministry tools had to clear a data threshold first. The rule sounded responsible: no launch without measurable signals on engagement, completion, or retention. In practice it meant we green-lit an automated content generator for children’s ministry volunteers because the early click rates looked promising and the drop-off numbers stayed flat. Six weeks later a volunteer printed a lesson that inserted a theologically sloppy paragraph about grace, and a parent called the church office confused and angry. The damage was small but real, and it happened because I had treated the absence of negative data as permission to skip the judgment step. That year taught me a hard lesson about ministry judgment calls: data can inform them, but it can never replace them.

That choice cost us trust with two ministry teams and forced a rushed rollback that burned engineering hours we didn’t have. More quietly, it taught the product group that metrics could substitute for deciding what actually protects children and the people serving them. I had confused “we have no evidence of harm” with “we have exercised wisdom.” I had let metrics quietly absorb the ministry judgment calls that should have stayed human.

The inbox that taught me ministry judgment calls first

The children’s director’s inbox had always been the real spec. Every week it contained the questions that data never captured: “Will this lesson still work if the volunteer only has seven minutes between services?” “What happens when a ten-year-old asks the follow-up question the script doesn’t cover?” Those messages forced a kind of judgment that treated edge cases as primary, not anecdotal.

When we moved the same team onto an AI-assisted outline tool, the inbox went quiet for three weeks. The numbers looked clean, so we kept shipping. Only later did we learn the volunteers had stopped asking questions because the generated outlines felt finished. They were also quietly correcting the theological shortcuts on their own time. The cost showed up later in volunteer burnout, not in the dashboard.

Solomon’s judgment in the disputed-infant story worked because he refused to accept the data of two living claimants at face value. He forced a test that revealed what each woman was actually willing to protect. The story is not about gathering more information; it is about constructing a situation that makes the real commitment visible before harm occurs.

How data-first habits quietly hand off ownership

Product teams that default to data thresholds train themselves to outsource the hard calls. Once the threshold is met, the decision feels mechanical. In faith contexts this creates a slow transfer of authority from the people who carry the mission to the people who can move the metric. The AI feature that generated small-group questions looked safe because completion rates held steady, yet no one had asked whether the questions still required a leader who could recognize when someone in the room was in actual crisis. Research on how leaders weigh evidence shows why ministry judgment calls still need a human in the loop.

The pattern repeats across tools. A prayer-request classifier gets approved because precision and recall numbers clear the bar. No one tests what happens when the model routes a disclosure of abuse to the wrong staff member because the training data never included that edge. The data did its job; the product organization had already decided its job was finished once the numbers cleared.

This is the move Solomon refused. He did not wait for additional evidence to accumulate; he acted on the judgment that protecting the child mattered more than preserving the appearance of fairness between two claimants. Data-first cultures lose the muscle for that kind of preemptive decision.

Rebuilding ministry judgment calls into every AI pilot

The fix is not to ignore data. It is to insert an explicit judgment gate before any pilot leaves the internal environment. The gate requires naming the specific person or group whose safety or formation could be damaged if the model behaves as designed. That naming happens in writing and must be signed by someone whose role includes ongoing responsibility for the people being served.

Next comes a forced negative test. The team must articulate the scenario in which the feature succeeds on metrics while still violating the named commitment. Only after that scenario is written and reviewed does the pilot move forward. The test is not optional and does not require statistical significance; it requires the same kind of clarity Solomon demonstrated when he identified what the two women were actually willing to lose.

We now run this gate on every new AI capability in the curriculum tools. The process adds two days to the start of a pilot and removes entire categories of later rework. It also surfaces questions the data never would have asked, such as whether an AI-generated illustration for a children’s lesson should be allowed to depict Jesus in a way that contradicts the visual language already used by the volunteer’s own church.

Your Turn: Apply This Today

  • Pick one active AI feature and write the single sentence that names the exact person whose formation or safety could be damaged if the model performs as designed.
  • Schedule a thirty-minute meeting this week with the ministry owner who carries real responsibility for that person; read the sentence out loud and ask them to revise it until it matches their actual stakes.
  • Before the meeting ends, draft the one negative scenario in which the feature meets its current success metrics while violating the commitment you just named.
  • Block two hours next week to run a manual test of that scenario with three real users who match the edge case, not the average user.
  • Write the decision that follows the test in one paragraph and send it to the same ministry owner for sign-off before any further engineering work proceeds.
  • Document the entire exchange in the project log so the next team cannot claim they lacked the judgment step.

The same pattern of deferred judgment shows up in how mid-career product managers track the wrong signals and how children’s directors end up owning the actual product requirements without anyone noticing. Both posts trace what happens when teams treat external proof as a substitute for deciding what must be protected.

I consult with AI product leaders and ministry technology teams on judgment gates for early pilots, metric design that respects edge cases, and rebuilding ownership after data-first habits have taken hold. Let’s talk.

Is This Ministry Task an Entire Role?

Children's ministry leader weighing whether automating ministry tasks gives away a relationship

Last month a children’s ministry director described what happened after she automated the weekly volunteer texts. Replies still came in. People still showed up. But she stopped knowing who was barely hanging on until they disappeared. Automating ministry tasks looked like pure efficiency, until the relationship inside the task quietly went missing.

She asked if the sequence was ever really hers to give away. The answer sat between us for a while.

Some tasks carry the relationship inside them. Once you hand those off, the role doesn’t shrink—it just stops belonging to anyone.

When teams treat automating ministry tasks as a one-off task, they measure only time saved. When they treat it as a role, they measure whether the volunteer still feels seen after the agent takes the first step. The second measurement almost always shows faster drop-off once the human hand disappears.

Print-first tools built for Sermons4Kids exposed this pattern early. The volunteers who completed the full loop were the ones who received a single personal reply after submission. Removing that reply to save coordinator hours also removed the only signal that the church noticed the work happened.

Automating ministry tasks still leaves distribution deciding who keeps the relationship

AI agents can generate the content, but they cannot decide who receives it or how the response travels back. That decision remains with whoever controls the distribution list and the reply path. In most churches that list still sits in a single inbox or spreadsheet maintained by one staff member or lead volunteer. This is the trap of automating ministry tasks without moving the list that carries the relationship.

Handing the generation step to an agent while leaving the list with the original owner looks like efficiency. In practice it creates a new dependency: the owner now spends time reviewing agent output instead of originating the contact. The relationship transfers only when the list itself moves to the agent layer and the original owner loses visibility into who was contacted and what was said.

Ministry platforms that reached scale kept the distribution layer inside the church rather than inside the tool. The churches that retained the list also retained the ability to decide when an automated message should become a human visit instead.

One visible handoff that already happened

Two years ago several large churches moved their children’s check-in reminders from a staff coordinator to an automated SMS flow. The coordinator’s title stayed the same, but the weekly report stopped arriving. Within six months the same churches reported that volunteer no-shows had shifted from last-minute calls to complete silence; the coordinator no longer knew who needed the personal nudge because the system no longer surfaced the exceptions.

The task was replaced. The role that noticed when the task failed was quietly disassembled. Solomon’s test would have flagged the change the moment the coordinator stopped asking to see the list.

Your Turn: Apply This Today

  • Pick one recurring children’s ministry task that currently lands in a coordinator inbox and write down the exact three people who lose visibility if the task moves to an agent.
  • Map the task to its owning role by listing the single decision that only a human in that role can make when the task fails.
  • Run Solomon’s test: ask the current owner whether they would rather keep the task or keep the list of who receives its output.
  • Choose one task this week where the answer is “keep the list” and leave the generation step with the owner rather than the agent.
  • Measure the outcome by counting how many volunteers still receive a human reply within 48 hours after the task completes.
  • Document the change in one shared note so the next product decision starts from the ownership map instead of the time-saved estimate.

The Tuesday the Children’s Director’s Inbox Became the Real Product Spec showed how inbox ownership reveals real ministry structure. How Do You Embed Agents Without Quietly Rewriting Ministry Ownership? examined the same pattern across multiple tools.

I consult with ministry product leaders on task ownership mapping, agent integration boundaries, and volunteer relationship distribution. Let’s talk.

To the PM Who Just Inherited the AI Agent Mandate

Church sound desk representing AI guardrails for churches and worship technology oversight

Inheriting an AI agent mandate is disorienting. You’re the product manager who arrived at the denomination headquarters in March with a mandate to ship an AI agent that handles volunteer scheduling, lesson customization, and follow-up emails—none of which you’ve ever done yourself on a Thursday night when the curriculum box arrives late and the small-group leader texts that she’s sick.

The person who handed you the brief has never opened the volunteer portal at 9 p.m. to see which names are still unchecked. They measured success by tokens generated and tickets closed. You inherited both the budget line and the quiet knowledge that the real users will quit if the agent makes their already-thin margin of time worse.

John Wesley’s first rule—do no harm—turns out to be the only filter that survives contact with actual ministry work. The other two rules matter later. This one must come first because agents do not feel the cost of their own suggestions.

The Rule of an AI Agent Mandate That Protects the Volunteer

Most agent roadmaps start with capability: the model can parse availability, rewrite a story for third-graders, and draft a reminder. That ordering reverses the actual risk. The volunteer who prints the lesson at the kitchen table does not experience capability; she experiences an extra paragraph that contradicts the printed page or a time slot that collides with her own child’s practice.

Wesley insisted the first obligation was to avoid injury even when the intention was good. Applied to agents, this means every proposed action must carry an explicit reversal cost that a non-technical volunteer can exercise in under two minutes. If the reversal takes a support ticket or a settings panel buried three clicks deep, the rule is already broken.

The teams that pass this test build the undo surface before they build the generation surface. They measure the percentage of agent actions that a volunteer can nullify without leaving the flow she was already in. That number, not accuracy benchmarks, becomes the release gate.

An AI Agent Mandate Quietly Shifts You from Builder to Referee

Product managers trained on owned surfaces learn to optimize for completion. When an agent owns the surface, completion is no longer the scarce resource; judgment is. You stop writing requirements that say “the agent shall produce X” and start writing refusal conditions that say “the agent must surface Y uncertainty to a human before proceeding.”

This is the referee posture. You still decide what the agent is allowed to attempt, but you spend most of your attention on the boundary calls: when must a children’s director see the output, when may the agent send without review, when must the agent ask a clarifying question that only a human can answer. The artifact you ship is no longer a feature list; it is a set of escalation thresholds.

One Midwest region tried the builder posture first. Their agent auto-assigned volunteers to rooms based on past attendance. Within six weeks the children’s director was fielding calls from parents whose kids had been placed with the wrong age group because the model treated “showed up twice” as equivalent to “is trained for that room.” The fix was not better training data; it was a hard stop that required the director to confirm any assignment involving a volunteer under three prior sessions. Referee logic replaced builder optimism.

How to Keep Human Outcomes Visible When Agents Move Fast

Agents compress cycles. The same compression hides whether the outcome for the actual child or parent improved. Wesley’s rule forces the outcome back into view by requiring that any agent action be logged against a named human responsibility that already existed before the agent arrived.

That means the dashboard does not track agent utilization; it tracks whether the volunteer who accepted the assignment still shows up and whether the parent who received the follow-up email opens it within the same window as before. When those two numbers move in opposite directions, the agent is creating motion without discipleship.

Teams that keep outcomes visible add one recurring ritual: every Friday they pull the ten most recent agent-initiated actions and ask the actual recipient one question: “Did this save you time you could use for something that matters, or did it create new work you now have to undo?” The answers become the only acceptable source material for the next sprint’s boundary changes.

McKinsey’s explainer on what an AI agent really is is worth reading before you act on any AI agent mandate.

Your Turn: Apply This Today

  • Pick the single agent action that currently runs without human review and add a one-click reversal that lands in the volunteer’s existing inbox rather than a new portal.
  • Write the refusal condition for that action in plain language a children’s director would recognize, then test whether the agent actually stops when the condition appears.
  • Label the next three agent outputs with the name of the human who remains responsible if the output is wrong; surface that name to the volunteer before she acts on it.
  • Run the Friday ritual this week on the five most recent agent actions and record whether any recipient said it created new work; bring only those cases to the next planning meeting.
  • Remove one accuracy metric from the agent dashboard and replace it with the reversal rate measured in the last seven days.
  • Schedule a thirty-minute call with the person who will still be answering parent texts at 8 p.m. if the agent is wrong, and ask her to name the one situation the agent must never handle alone.

The 1997 Lesson AI Product Teams Keep Missing shows how earlier generations of church software repeated the same rollout pattern until volunteer retention numbers forced a redesign. The Tuesday the Children’s Director’s Inbox Became the Real Product Spec traces what changed once the actual recipient of the work became the source of the requirement instead of the target of the feature.

I consult with product leaders at faith-based organizations on agent handoffs, volunteer protection metrics, and outcome visibility in AI integrations. Let’s talk.

The First Build Where Judgment Beat the Dataset

Hands on a laptop reviewing an AI ministry tool dashboard where early usage data can mislead the team

Hands on a laptop reviewing an AI ministry tool dashboard where early usage data can mislead the teamThe children’s director shoved her laptop across the kitchen table. The screen showed three Figma frames she had built after midnight. One had a big green button for “daily verse push.” Another split the screen between parent notes and a child progress bar. The third tried to guess what a family needed based on last week’s attendance. This is the moment every AI ministry tool faces: the dataset points one way and the person who knows the work points another.

She pointed at the usage logs open in another tab. “None of this matches what those numbers say parents want,” she said. The logs came from the first two weeks of an AI pilot. They showed quick taps on random verses and almost no return visits after day three. She had watched real parents in her church that same week. They opened the app once during carpool, closed it when the toddler started crying, and never came back.

The data said simplify the verse feed. Her hands said something else. She had already seen what happened when the team followed the numbers last quarter. The feature shipped, the graphs looked clean for fourteen days, and then the actual ministry leaders stopped logging in.

Solomon faced two women who both claimed the same child. He did not run a survey. He did not average their stories. He called for a sword. The move looked reckless until the real mother spoke. The test was never about the data in front of him. It was about whose voice actually belonged in the room.

That same test sits in front of every team shipping the first version of an AI ministry tool. The early logs come from people who found the prototype link on social media or in a beta signup form. They are not the children’s director at 9:40 p.m. They are not the parent who only opens the app when the van is moving. Their clicks tell you what curious users do, not what exhausted ministry volunteers will keep doing after month three.

The Inbox That Became the Real Spec

The children’s director kept every parent email in one folder. She had printed the last month of messages and spread them on the table next to the laptop. One mother wrote that she needed the app to remind her of the exact craft supplies already in her kitchen drawer. Another asked for a single sentence she could read to her son while she stirred dinner. None of those requests showed up in the first two weeks of usage data because the parents who wrote them had not yet downloaded the pilot.

I watched her map each printed email to a screen. The green button disappeared. The progress bar stayed only if it could be filled by a single tap during pickup line. The verse feed became a list of three options, each tied to an object already in a typical home. She built the change in forty minutes. The logs never would have surfaced those constraints because the people writing the emails were not yet in the dataset.

Why the First Ten Users of an AI Ministry Tool Lied

The first ten users of that pilot were all early adopters who liked trying new church apps. They tapped through every screen the day they signed up. They answered every prompt. Their behavior looked like engagement until the team checked the actual church roster. None of those ten people served in children’s ministry. None of them had kids in the current Sunday school year. Their taps reflected curiosity, not repetition under real load. An AI ministry tool that trusts those first taps optimizes for the wrong people.

Solomon’s test worked because he forced the claimants to reveal what they would actually protect. The early dataset never creates that pressure. It records what people will do once, not what they will defend when the cost is their own time on a Tuesday night.

How Solomon’s Split Decision Shows Up in Roadmaps

Teams keep shipping the averaged version. They see the logs favor short verses and remove the longer reflection. They see the quick taps and kill the extra confirmation step. Six weeks later the real users have left because the thing they needed was never measured.

The fix is not more data. It is deciding whose voice gets the sword. On the current roadmap that means looking at every AI pilot feature and asking which one would survive if the only inputs were the printed emails from the children’s director’s folder. Features that only the beta users love get cut. Features that match the inbox stay even when the graphs look thin in week one.

Your Turn: Apply This Today

  • Pick the single AI pilot feature your team plans to ship in the next four weeks and list the three data sources that currently justify keeping it.
  • Find the printed or saved messages from the actual ministry volunteers who would use the feature after launch and map each one to a screen or flow.
  • Remove any screen element that cannot be completed in one tap during a carpool or while stirring dinner.
  • Run the remaining flow against the first ten users from your beta list and note how many of them actually serve in children’s ministry or lead volunteers.
  • Delete the feature if fewer than half of those ten users match the real ministry role the inbox describes.
  • Write the new version of the feature using only constraints pulled from the inbox and ship that version to the next five volunteers who email you this week.

The same pattern showed up in the post about the children’s director whose inbox became the real product spec and in the post about the 1997 lesson AI product teams keep missing.

I consult with ministry product leaders on early-stage AI features and deciding which data sources actually belong on the roadmap. Let’s talk.