A growing “Agent-First Product Management” playbook circulating among product leaders tells teams to hand discovery, prioritization, and even outcome measurement to autonomous agents so the human PM can focus on “higher-order bets.” The model assumes agents will surface the right problems, run the experiments, and flag only the decisions that truly require taste or authority. In faith-tech work, that assumption breaks immediately because the cost of a wrong call lands on real people whose spiritual formation, not just their engagement metrics, is at stake.
The framework treats verification as a residual task that gets smaller as models improve. It does not account for contexts where the signal is relational and the downside is not a lost quarter but a broken trust that takes years to repair. When the product shapes how volunteers teach children Scripture or how users encounter the Bible in crisis, an agent that optimizes for completion rate can ratify the wrong outcome without ever knowing it missed.
This is the foundational misread that causes product teams to strip the verification layer out of their operating model. They optimize for speed of iteration and treat human review as an expensive add-on rather than the non-delegable core. The result is clean dashboards that hide the moments when an agent quietly chose the expedient over the faithful.
Solomon’s judgment offers the clearest lens here. Two women claimed the same child. Solomon did not ask for more data or run another test. He created a forced choice that revealed which claimant bore the actual cost of the outcome. The test worked because only the true mother would refuse the split. Agents can simulate preferences and optimize for stated goals, but they carry none of the relational stake that makes the refusal meaningful. In ministry products, that stake is the difference between a feature that serves formation and one that merely moves a number.
The Waymo Parallel That Doesn’t Travel to Ministry
Waymo’s autonomous vehicles succeeded because the environment is bounded, the failure modes are physical and observable, and the cost of error falls on insurance models that can price the risk. The same logic does not transfer to products that shape how a children’s ministry volunteer prepares a lesson at 9:30 p.m. on a Tuesday. The variables are not lane markings and traffic signals; they are the volunteer’s attention span, the child’s home situation, and the theological weight of the story being told.
Teams that import the Waymo pattern assume that once the agent handles content selection and personalization, human oversight can drop to quarterly reviews. In practice, the agent begins to favor lessons that score high on the proxy metrics it was given—completion time, visual appeal, repeat usage—while quietly dropping stories that require more context or that surface harder questions. The volunteer still prints the material, but the formation goal has already drifted.
The verification step that remains is not another prompt. It is a human who carries responsibility for the actual child in the room and who will notice when the content no longer fits. That person cannot be replaced by higher model accuracy because the signal lives outside the data the agent sees.
What Breaks When Discovery Gets Automated
Automated discovery promises to replace the slow work of talking to users with continuous agent-driven interviews and synthesis. The pitch is that the agent can run hundreds of lightweight conversations, cluster themes, and surface opportunities faster than any human PM. In faith contexts the method produces plausible themes that still miss the constraint that matters most: the user is often serving under time pressure and emotional weight that no interview transcript captures.
A children’s ministry coordinator may say in a recorded session that she wants shorter lessons. The agent logs the preference. What it cannot register is that the same coordinator will accept a longer lesson if it reduces her Sunday morning panic, because the real job is not efficiency but survival until the service ends. The agent optimizes for the stated preference. The PM who still talks to coordinators in person learns the unstated one.
When discovery is fully automated, the team loses the ability to notice which problems users will tolerate and which ones they will quietly route around. That distinction only appears when someone with skin in the outcome is in the loop. Without it, the roadmap fills with features that solve clean problems and ignore the messy ones that actually determine whether the product stays in use.
The Decision Point No Agent Can Hold
Every week, a product decision arrives that cannot be reduced to trade-offs the agent has been trained to weigh. A new agent-generated curriculum track scores well on engagement but quietly under-emphasizes a core doctrine the ministry has taught for decades. The data does not flag the gap because the doctrine was never encoded as a metric. The human PM must decide whether to ship the track, rewrite it, or kill it.
Solomon’s test applies directly. Only the person who will answer for the outcome in front of the actual ministry leaders will pause to ask whether the doctrine matters more than the engagement lift. An agent can surface the trade-off, but it will not carry the cost of choosing poorly. That cost sits with the PM whose name is attached to the product in the eyes of the churches using it.
Rebuilt operating models that remove this layer do not become faster; they become brittle. The first time an agent makes a call that contradicts the ministry’s actual convictions, the entire trust relationship must be rebuilt from the beginning. The verification step that looked like overhead turns out to be the only part of the system that could have prevented the breach.
Your Turn: Apply This Today
- Pick one agent workflow currently running in your product and add a single human verification checkpoint this week that requires the PM to review the output against the ministry’s stated convictions before it reaches users.
- Write down the exact decision the agent is making today that carries relational risk, then schedule 30 minutes on Friday to test whether the current prompt would have caught the mismatch in a past incident you still remember.
- Run the verification test on the oldest active agent flow in your stack by Friday and log whether the human reviewer changed the output or accepted it unchanged.
- Identify the one proxy metric the agent is optimizing that has no direct counterpart in your ministry’s actual success criteria, then replace it with a qualitative check that only a person carrying outcome responsibility can perform.
- Document the last three times an agent surfaced a “ready to ship” recommendation and note which one required a Solomon-style judgment call that the model could not have resolved.
- Schedule a 45-minute session with the ministry leader who owns the outcome for your highest-volume feature and ask them to name the single failure mode an agent would never flag on its own.
The same pattern appears in “The Verification Layer No One Budgeted For” and “The PM Role That Broke When Agents Took the Wheel,” where the missing human step turned clean agent output into downstream trust problems that took quarters to repair.
I consult with product leaders building AI tools for ministry on agent verification workflows, discovery automation limits, and the decision points that still require human ownership. Let’s talk.
Discover more from Dr. Joshua Read
Subscribe to get the latest posts sent to your email.

