Why Confidence Theater Still Passes Internal Reviews

The “metrics over theater” mandate from Marty Cagan’s Inspired framework still collapses when product teams move from enterprise software into ministry tools. Cagan calls for replacing executive confidence with observable experiments and clear success criteria. That prescription assumes every handoff can be instrumented inside a single codebase or dashboard. Ministry products break the assumption at the point where an AI draft leaves the screen and enters a volunteer workflow that ends in a printed page or a live children’s lesson.

The framework treats the experiment as complete once token counts, latency, and click-through numbers move. It never accounts for the physical or relational transfer that determines whether the output actually gets used. In practice, teams celebrate a 40 percent drop in generation time while the curriculum still sits unopened because the volunteer cannot finish the customization in seven minutes.

This is the foundational misread that causes product teams to ship pilots that look rigorous in review but never survive the first real gate. The misread treats measurement as a software problem rather than a sequence of human transfers that each carry their own failure modes.

Charlie Munger’s latticework of mental models demands that any new tool be checked against multiple independent lenses at once. One model is the incentive model. Another is the handoff model. A third is the completion model. Most AI pilots in faith settings run only the incentive model and stop. They reward speed and cost reduction while ignoring whether the next human in line can still do their job without new friction. The lattice reveals the gap because each model must be satisfied before the pilot earns the label “working.”

The Experiment That Never Reached the Print Step

A recent Sermons4Kids pilot generated full lesson outlines in under two minutes. The internal review celebrated the output quality and the 60 percent reduction in writer hours. No one logged what happened after the file left the platform. Volunteers still had to reformat the material for their specific room size, age mix, and available supplies. The print step added twenty-three minutes of manual work that the experiment never measured.

The same pattern appears when teams test AI-generated discussion questions for small groups. The model produces theologically sound text. The handoff to the group leader who must decide whether the questions fit the actual people in the room remains invisible. When the leader skips the questions, the metric dashboard still shows success because the generation event completed.

Munger would require a second model that tracks whether the printed or displayed artifact survives the next human filter. Without it, the experiment only proves the model can produce text, not that the text can travel the full distance to use.

Metrics That Reward Speed Over Gatekeeping

Internal reviews often track generation cost, model version, and time-to-first-draft. These numbers move quickly and look decisive on a slide. They do not record whether a ministry director rejected the draft because it lacked local application points. The rejection happens outside the tracked system, so the metric stays clean.

When the same team later adds a “human review” checkbox, the checkbox becomes another speed metric. Reviewers learn to clear items fast to keep the dashboard green. The real gatekeeping judgment gets compressed into a binary field that carries no weight in the next sprint planning meeting.

The latticework exposes the missing model here: the gatekeeper model. It asks whether the person who must say yes or no still has enough context and time to exercise judgment. When that model is absent, speed metrics quietly train teams to bypass the very humans the product claims to serve.

The Lattice Model Missing from Most AI Pilots

Most pilots apply only the engineering model of iteration. They ship, measure, adjust, repeat. Munger’s approach adds the second-order effect model and the reputation model in parallel. The second-order effect model asks what behavior the new speed incentive creates among volunteers who now receive faster but less contextual material. The reputation model asks whether pastors will continue to trust the platform once they notice the gap between polished output and usable output.

A pilot that only watches token spend cannot surface either effect. The first sign of trouble appears weeks later when volunteer completion rates drop or when a pastor quietly switches to another source. By then the internal review has already declared the pilot a win.

Teams that run all three models together produce different artifacts. They log the moment the AI file reaches the volunteer inbox, the time until the volunteer either completes or abandons customization, and the explicit reason given for abandonment. Those logs become the actual experiment data rather than the generation metrics that stop at the screen.

Your Turn: Apply This Today

  • Pick one current AI pilot and add a single handoff log field that records the exact minute the output left the platform and reached its first human recipient.
  • This week, require every pilot review to include the completion timestamp from the volunteer or leader side before any success claim is accepted.
  • Create a two-column template that forces the team to list both the generation metric and the handoff metric for the same item before the item can move to the next sprint.
  • Schedule one thirty-minute call with a children’s ministry volunteer who used last week’s generated material and ask only what broke during the print or prep step.
  • Remove any speed or cost metric from the next internal review slide deck and replace it with the handoff completion rate for the same period.
  • Write down the second model you will run alongside the current experiment and commit to sharing its result in the same review that shows the token numbers.

The Trust Step Kent Beck Made Non-Negotiable and The Open-Weight Number That Still Required a Human Gate both trace how unmeasured handoffs quietly undo technical wins. I consult with ministry product leaders on AI pilot design and handoff measurement systems. Let’s talk.

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.