The novel agent test is the 10/10 rule teams still skip. I sat with the support log at 11:40 p.m. and watched the same three users hit the first screen, stall, then close the tab. We had just shipped the agent that would “handle everything after intake.” No one had checked whether intake itself still worked.
The demo video looked clean because we tested it on our own devices with our own data. Real people on real phones never reached the clever part.
That gap keeps showing up whenever the new thing feels more interesting than the thing we already shipped.
The discovery log that never showed the real handoff
Continuous discovery practice, as described by Teresa Torres, depends on testing the riskiest assumption first rather than the most novel feature.
Most discovery logs I see record interviews with ministry volunteers who already know the product. The notes list feature requests and pain points, yet they rarely capture the moment the volunteer steps away from the screen to finish the actual task. In one children’s ministry platform the handoff happened after the lesson plan was printed: the volunteer then spent seven minutes locating physical supplies that the digital workflow assumed were already organized. That step never made it into the opportunity solution tree because the interview ended at the print button.
The gap matters once an agent is introduced to “optimize” the plan. The agent can generate supply lists faster, but the underlying organization problem remains untouched. Continuous discovery requires watching the full loop, not stopping at the digital artifact. Without that observation, the agent multiplies output that still requires the same manual rescue work.
Teams that close the loop discover the handoff is often where trust erodes. A pastor may accept AI-generated talking points yet still retype them because the source material does not match the translation or reading level used on Sunday. The log shows acceptance; the real workflow shows revision.
Where the novel agent test exposes the gap
The 10/10 rule is simple to state and difficult to satisfy: run the current workflow with ten consecutive target users and require every one of them to complete it without hesitation, confusion, or workarounds. Most teams test for usability issues inside the interface. They rarely test for whether the interface itself is the right place to solve the job. When the test is applied to sermon or lesson preparation tools, the pass rate drops once the user must move content into a different format for delivery.
An agent that auto-generates slides or small-group questions looks impressive in a demo. The same agent fails the 10/10 test when the tenth volunteer prints the material, discovers mismatched scripture references, and opens a separate Bible app to correct them. The novel capability did not address the translation or formatting layer that sits between the product and the actual Sunday use.
Teams that reach 10/10 on the base workflow before adding agents report fewer downstream revisions. The agent then operates on a verified foundation rather than amplifying an incomplete one. The data on agent adoption therefore misleads when it counts features shipped instead of workflows cleared.
Why the novel agent test arrives too early
Novelty bias inside product organizations rewards the visible addition over the invisible cleanup. Roadmaps list “AI agent for volunteer matching” or “AI sermon illustration generator” because those items read as innovation to stakeholders. The prior steps—standardizing supply lists, aligning translations, confirming delivery formats—stay invisible because they do not produce new screenshots.
Continuous discovery counters this bias by requiring the team to keep interviewing until the opportunity is solved for the current users, not the imagined next set. When the 10/10 bar is met first, the agent can be evaluated against a stable baseline. When the bar is skipped, the agent becomes another variable that must be debugged by the same volunteers who were already patching the original flow.
The pattern repeats across curriculum platforms and resource sites. The teams that later remove the agent do so after realizing the core workflow still required manual reconciliation. The teams that keep the agent are the ones that first drove the manual version to consistent completion. The difference is not technical sophistication; it is the order of validation.
Your Turn: Apply This Today
- Pick one ministry workflow that currently ends with a handoff to physical materials or another app; schedule three 20-minute observation sessions this week with users completing the full loop, not just the screen portion.
- Run the 10/10 test on that workflow before any agent scoping begins; document every point where the tenth user still improvises.
- Update the opportunity solution tree to include only problems that appeared in at least seven of the ten sessions rather than feature requests gathered in meetings.
- Delay the next agent-related ticket on the roadmap until the base workflow passes 10/10; move the agent work to a separate discovery track that starts only after the baseline is cleared.
- Share the revised tree and the 10/10 results in the next roadmap review so stakeholders see the validation order rather than the novelty list.
- Repeat the 10/10 test one month after any agent ships to confirm the addition did not reintroduce friction at the handoff points previously cleared.
The same pattern of validating the base before layering new capability appears in The Trust Layer the Roadmap Still Treats as Optional and The Mandate That Arrived Before the Workflow Existed. I consult with product leaders on applying continuous discovery to ministry workflows and sequencing AI agents only after core tasks reach unqualified user completion. Let’s talk.

