The real test for any AI tooling budget is whether a volunteer can name one concrete change in their weekly work without being prompted by a product manager. Most line items never clear that bar. They multiply compute costs and internal velocity metrics while leaving the actual user path untouched.
This is the foundational misread that causes product teams to approve tooling spend based on engineering demos rather than observable volunteer behavior. The spend grows because the demo looks impressive, yet the volunteer still prints the same PDF at the same time each week.
John Wesley’s first rule — do no harm — offers the clearest lens here. Applied to tooling, it means refusing any hidden cost that quietly increases friction for the people who never see the roadmap.
The AI Line Item That Grew While Velocity Stayed Flat
One children’s ministry platform added three successive AI code-generation tools over eighteen months. The engineering team reported a 40 percent rise in pull requests merged. The volunteer completion rate on the curriculum builder stayed flat at 62 percent.
The new components arrived as internal micro-features: smarter autocomplete in the lesson editor, auto-tagging for activity files, and a prompt that suggested reorderings of the printed handout. None of these altered the seven-minute window in which a volunteer opens the material on a phone and decides whether to print or abandon it.
Feature velocity rose on the internal dashboard. The number of volunteers who reached the print step did not. The budget line item kept its justification because the demos never required a volunteer to describe the difference.
Wesley’s Do-No-Harm Test on the Hidden Line Item
Wesley’s rule forces a different question: does this spend create new friction that only appears after the contract is signed? In practice this shows up as added login steps, new file formats that require conversion, or prompt interfaces that demand the volunteer rephrase their intent before the system accepts it.
A sermon resource team once introduced an AI summarizer for weekly emails. The tool reduced drafting time for staff but required volunteers to confirm the summary matched their printed notes before forwarding. The confirmation step added forty seconds per email. Across hundreds of weekly users, the cumulative harm exceeded the drafting savings within one quarter.
The rule is simple to apply once named: if the cost cannot be explained to a volunteer in one sentence that ends with a visible action they already perform, the spend fails the test.
The Single Metric That Justifies the Line Item
Research on software ROI, including Harvard Business Review on measuring value, shows that a tool budget survives only when one usage metric ties the line item to real behavior change.
The only metric that held steady across recent quarters was the percentage of volunteers who completed their task without opening a support ticket or re-downloading the same file twice. AI features that improved this number stayed funded. Those that improved only internal velocity metrics were quietly deprioritized in the next planning cycle.
This metric survives because it cannot be gamed by a demo. A volunteer either finishes or does not. The tooling spend must produce a measurable lift here or it is, by definition, a line item no demo can justify.
Teams that adopted this filter reduced their AI tooling stack by roughly one-third while holding volunteer completion rates steady. The remaining spend now maps directly to actions the volunteer can name without prompting.
Your Turn: Apply This Today
- Open the current AI tooling invoice and select the single largest monthly line item.
- Write one sentence that describes the exact volunteer action this item must improve within the next seven days.
- Run a five-user test this week using only that sentence as the success criterion.
- Remove or renegotiate the line item if the test shows no lift in the named action.
- Repeat the same mapping for the next two largest AI-related expenses before the next budget review.
- Document the outcome in the same place the engineering velocity numbers are tracked so the comparison is visible to leadership.
The Compensation Split No Playbook Prepared PMs For and Trust Logs Beat the Next Feature Sprint both trace similar gaps between internal metrics and observable user behavior. The same pattern appears whenever tooling spend outruns the volunteer’s ability to name its effect.
I consult with product leaders and ministry tech teams on AI tooling justification and volunteer outcome mapping. Let’s talk.

