AI Prototyping Without Mission DNA Builds the Wrong Product

I was part of a team prototyping onboarding flows for a global Bible engagement app, and we made a mistake I’ve seen repeated dozens of times since. We handed the prompts to our AI tools without telling them what the product was for. The AI did exactly what it was trained to do — it reached for patterns from fitness apps, productivity tools, and habit trackers. Badges. Streaks. Daily login rewards. These were the patterns that worked in consumer contexts with the highest engagement data.

The problem: our users weren’t trying to build a fitness habit. They were people trying to connect with scripture in the ten minutes before work. Or processing a hard diagnosis late at night. Or trying to find something true to hold onto in the middle of a season that wasn’t making sense. The streak counter didn’t just fail to help them — it felt faintly insulting. Our completion rates showed it, and when we talked to users, they told us the same thing in different words: this doesn’t feel like it was built for me.

The fix wasn’t a different AI tool. It was embedding the mission before the AI touched anything. My job, I’ve come to understand, isn’t to build features. It’s to build space — the right kind of space for the right kind of encounter. That framing changes what you put into every prototyping prompt.

Why AI Prototyping Defaults to Someone Else’s Mission

AI prototyping tools are trained on the vast middle of consumer software. They know what engagement looks like in fitness apps, productivity tools, social platforms, and enterprise SaaS. What they don’t know is why a small church prioritizes trust over scale, why a children’s ministry platform needs to work in a seven-minute volunteer prep window, or why a global discipleship app has to honor cultural nuance over raw personalization efficiency.

When you hand a prototyping prompt to an AI without mission context, you get competent, generic output. It passes the “does this look like software” test. It fails the “does this serve our community” test. The difference only becomes visible when you put it in front of real users — by which point you’ve spent significant time and trust going in the wrong direction.

The solution isn’t to avoid AI prototyping. It’s to front-load the mission input so the tool has something specific to work from. Specificity is what separates useful AI output from generic AI output. An AI that knows your non-negotiables produces dramatically better output — not because the model is smarter, but because it has the right filter.

How to Embed Mission DNA Before You Prototype

The practice I’ve landed on is building a mission filter document before starting any AI-assisted prototyping sprint. It doesn’t have to be long — a single page with three things: the specific user you’re serving (not a persona template, but a real description of a real person in your community), the job they’re trying to do in their language, and the non-negotiables that define what this product will and won’t do regardless of what the engagement data suggests.

On the children’s curriculum platform, our mission filter was two sentences: “A harried volunteer with seven minutes to prep a lesson. We make that person feel ready, not overwhelmed.” Every prototype, every AI-generated flow, every feature proposal got filtered through those two sentences. The AI didn’t know them by default — but once we embedded them into our prompts, the output changed. We stopped producing things that looked like apps and started producing things that felt like the product we were trying to build.

The other accelerant is bringing non-technical mission stakeholders — pastors, volunteers, community members — into the prototyping process as co-creators, not just testers. Not because they’ll have great UI ideas. Because they’ll immediately tell you when something feels off. Their instinct for “this doesn’t feel like us” is faster and more reliable than any usability metric I’ve ever run.

What Mission-Informed Prototyping Actually Produces

When you prototype with mission DNA embedded, a few things shift. You stop building features that look good in demos but don’t solve the actual user job. You catch misalignments earlier — before they’re built into production — because the mission filter creates a shared language for “this isn’t right” that anyone on the team can use. And you build trust with your community faster, because what they encounter actually reflects the values the product claims to hold.

That last point matters more in faith-tech than anywhere else I’ve worked. Ministry users have a finely tuned instinct for whether something was built by people who understand their context or built by people who read about it. They know the difference. The mission filter is how you make sure they experience the former.

Your Turn: Apply This Today

Before your next AI prototyping session, use this checklist to make sure your mission is in the room.

  • Write your mission filter document before the sprint starts. One page: your specific user described as a real person, the job they’re trying to do in their exact words, and three non-negotiables your product will and won’t do. This goes into every AI prompt.
  • Interview one ministry stakeholder this week. Ask them to describe what “success” looks like in their specific context — use their exact language in your next prototype prompt and compare the output to what you’d have gotten without it.
  • Audit your last three AI-generated prototypes. For each one, ask: does this reflect your mission non-negotiables, or does it reflect generic consumer app patterns? Name the specific places it went generic.
  • Add your mission filter to one prompt this sprint. Take a prompt you’ve used before, add the mission context, and compare the output. The difference will make the case for the practice more effectively than any argument I can make.
  • Bring one non-technical mission stakeholder into a prototype review. Give them one question: “Does this feel like us?” Their answer is your fastest mission-alignment check — faster than any usability test.
  • Create a shared vocabulary for “off-mission” with your team. Name two or three specific patterns that feel wrong for your community — things that work elsewhere but don’t fit your context. Make them explicit so anyone can call them out without it becoming personal.

For more on building AI strategy that fits faith-tech specifically, read The Complete Guide to AI Product Strategy for Faith-Tech and Ministry Leaders and AI Transformation Isn’t About Tools — It’s About Team Mindset.

Getting AI prototyping right in faith-tech requires upfront investment in mission clarity that most teams skip. I consult with product leaders on building that foundation — and on the AI strategy decisions that depend on it. Let’s talk.

AI Is Breaking Faith-Tech Infrastructure — Here’s How to Build for Breakpoints

The sermon resource platform I was working on nearly buckled on a Sunday morning in November. A church network shared a link, thousands of pastors hit it at once, and we weren’t ready. The platform slowed. Then timed out for some users. Pastors who depended on it to prep for their services got error messages when they needed the product most — which is the worst possible version of the trust problem in faith-tech. We spent weeks recovering that trust. The failure was entirely predictable. We just hadn’t done the work to predict it.

That experience changed how I think about infrastructure. Before it, infrastructure was something engineers handled and product managers approved budgets for. After it, I understood that infrastructure decisions are product decisions — and that in faith-tech, where your users’ most important moments often hit your platform all at once, the product leader who doesn’t understand their infrastructure breakpoints is one viral campaign away from destroying the trust they spent years building.

AI makes this significantly more urgent. The computational appetite of generative AI features is nothing like serving static content. Personalization, content generation, semantic search, real-time analytics — these features carry cost structures that most faith-tech teams don’t model until the bill arrives. By then, you’re either throttling features or absorbing a cost spike you didn’t plan for. Neither is a good option when your users are a volunteer at 6am on Sunday morning.

The Faith-Tech Infrastructure Problem Is Specifically About Timing

Most SaaS products have relatively predictable load patterns. Faith-tech products have predictable spikes — and the spikes happen at the worst possible times. Sunday morning. Wednesday night. Easter. Christmas. The high-traffic moments are exactly the moments when your users are most dependent on the product working and least forgiving of it failing.

When you add AI features to this dynamic, the cost structure amplifies the timing problem. A generative AI feature that costs $0.02 per request is manageable at normal traffic. At 10x traffic on a Sunday morning, that same feature costs 10x as much — and if you haven’t built cost controls, you’re either overspending or throttling users at the exact moment they need the feature most. The infrastructure decision you deferred becomes a user experience crisis on the day you can least afford one.

The product leadership question isn’t just “can we build this AI feature?” It’s “what does this feature cost when everything goes right, and what does it cost — in money and in trust — when everything goes wrong at the worst possible time?”

Building for Breakpoints Before They Happen

The most useful infrastructure exercise I’ve run with faith-tech teams is a failure scenario workshop. Not “what happens when one service has a brief outage,” but what happens in the scenario that felt unlikely until it didn’t — holiday traffic, a viral campaign, an API provider going down, a cost spike that requires immediate throttling.

Walking through those scenarios in advance surfaces the decisions that need to be made before the moment of pressure. Where do you cache? What degrades gracefully versus what fails completely? What does the user experience look like during a partial outage — a clear message, a fallback, or a silent failure that makes your users think they did something wrong? Those decisions, made calmly in advance, are completely different from the same decisions made at 8am on Easter Sunday.

Good infrastructure planning isn’t about eliminating risk — it’s about knowing your tradeoffs. What have you prioritized, and what have you accepted as a known risk? Knowing the answer to that question before you need it is the difference between a managed failure and a trust crisis. For more on the cost side of this, read AI Costs Are Skyrocketing: How Product Leaders Should Budget for AI Before It Bites Them. For the infrastructure architecture decision, read Cloud vs. Self-Hosting for Faith-Tech: How to Choose Infrastructure That Won’t Break Under Pressure.

Your Turn: Apply This Today

Infrastructure planning for AI features doesn’t require an engineering degree. These are product leadership decisions that prevent Sunday morning failures.

  • Map every AI feature’s cost structure. For each AI feature in your product, identify the per-request cost and model what happens to your monthly spend at 2x, 5x, and 10x current usage. Know your cost cliff before you hit it — not after.
  • Identify your non-negotiable uptime scenarios. Name the two or three user moments where failure is most costly — Sunday morning prep, a high-traffic campaign, a crisis care feature. These define where your highest infrastructure investment needs to go.
  • Run a failure scenario session with your team. Pick your highest-risk scenario and walk through it together: what breaks first, what degrades gracefully, what does the user experience look like during the failure, and what does recovery look like? Do this before you need to.
  • Implement one caching strategy this sprint. Identify the AI output in your product most likely to be requested repeatedly and build a simple cache. Measure the cost difference over 30 days.
  • Move one non-urgent AI call to async processing. Find a feature where real-time AI response isn’t essential to the user experience and shift it to a batch job. This is usually faster to implement than it sounds.
  • Design your degraded-mode user experience now. What does your product look like when an AI feature is unavailable — clear communication, a graceful fallback, no silent failures? Ship that experience before you need it. Your users will respect honesty far more than a broken spinner.

For the broader AI strategy context, read The Complete Guide to AI Product Strategy for Faith-Tech and Ministry Leaders.

Infrastructure failures in faith-tech aren’t just technical problems — they’re trust problems. I consult with faith-tech product leaders on AI infrastructure planning, cost management, and building for the breakpoints before they happen. Let’s talk.

Gamifying AI Adoption in Faith-Tech: How to Reward What Actually Matters

We added a streak counter to a volunteer training app and watched engagement spike for three weeks. Then it cratered. The volunteers who kept their streaks alive weren’t the ones doing the best work — they were the ones gaming the system, checking in without actually engaging. The people doing the real work found the streak counter mildly insulting. One of them told me so directly, which I appreciated.

The problem wasn’t the streak counter specifically. The problem was that we rewarded what was easy to measure instead of what actually mattered. That’s the gamification trap — and it’s everywhere in AI adoption right now. Teams add points, badges, and progress bars because those are the mechanics they know, and they watch engagement metrics go up, and they feel like something is working. What’s actually working is the illusion of progress. The real thing — behavioral change that serves the mission — is often going backward.

What Behavioral Science Actually Says About Motivation

Most gamification frameworks start with B.F. Skinner and stop there. The insight is correct: behaviors that are reinforced become more frequent. But Skinner’s research also showed something product teams tend to ignore — variable reward schedules create compulsive behavior, not meaningful behavior. The mechanics that make slot machines effective are also the mechanics that make gamification feel hollow after the novelty wears off. Three weeks of engagement spike, then a crater. I’ve seen this play out more than once.

The more useful framework for faith-tech AI adoption is Self-Determination Theory — the research showing that people sustain behaviors connected to autonomy, competence, and purpose. External rewards can start a behavior. Intrinsic connection to something that matters is what sustains it. For ministry users, “something that matters” isn’t abstract — it’s the lesson that went well, the conversation that helped, the volunteer who felt prepared instead of overwhelmed. That’s what you’re trying to connect the gamification to. Not logins. Not streaks. The actual thing.

Designing Gamification That Serves the Mission

After the streak counter failure, we rebuilt the engagement system around preparation completion rather than session counts. Instead of a streak for daily logins, a progress indicator for lesson readiness — how many of this week’s materials had been reviewed and marked ready. Instead of a badge for time-in-app, a visible confirmation that the volunteer was prepared for their class.

The difference wasn’t cosmetic. Completion rates climbed. More importantly, volunteers who completed their preparation reported feeling more confident leading their classes — which was the mission outcome we actually cared about. Nobody at a volunteer training debrief says “I really loved the streak counter.” They say “I felt ready.” That’s the outcome worth designing toward.

The design principle I’ve landed on: reward the behavior that reflects the mission outcome, not the behavior that’s easiest to measure. For an AI tool used in pastoral care, don’t gamify messages sent — recognize follow-through, the logged interaction that shows a leader used the tool to connect with someone who needed it. For an AI sermon prep tool, don’t reward time spent — reward preparation completed. The outcome you reward is the outcome you get more of.

Adoption Mechanics vs. Engagement Mechanics

There is a legitimate place for lighter gamification: reducing the friction of trying something new. A first-use celebration, a simple progress indicator, a “you’ve used this five times” acknowledgment — these lower the barrier to initial adoption without creating the compulsive patterns that burn out motivation. I’m not arguing against all gamification. I’m arguing for being deliberate about which kind you’re using and why.

The distinction I use is adoption mechanics versus engagement mechanics. Adoption mechanics help people get over the initial friction of learning something new — appropriate during onboarding and early use. Engagement mechanics that reward sustained use need to be tied to mission outcomes, not usage metrics, or you’ll spend months optimizing for the wrong thing. Most AI adoption failures I’ve seen conflate the two. They run adoption mechanics indefinitely and end up with users who feel perpetually coached into using a tool they already know how to use. That’s patronizing, and people feel it.

For more on the team adoption side of this, read AI Transformation Isn’t About Tools — It’s About Team Mindset and Agency Over Automation: Why AI Won’t Replace the Leaders Who Know How to Use It.

Your Turn: Apply This Today

Use this checklist to audit your current AI adoption gamification and rebuild it around outcomes that actually matter.

  • Audit every gamified element in your AI features. List each reward, badge, streak, or progress indicator. For each one, name the specific behavior it’s reinforcing — and whether that behavior reflects your mission or just your engagement metrics.
  • Identify your two or three mission-aligned outcome metrics. What are the user behaviors that actually reflect the outcome your product is designed to create? These are what your gamification should reinforce — not logins or time-in-app.
  • Remove one shallow gamification element this sprint. Pick the reward mechanism that most obviously optimizes for the wrong behavior and replace it with one tied to a mission outcome. Measure the difference over 30 days.
  • Interview three users about what makes the work feel meaningful. Ask what makes them feel capable and purposeful with the tool — not what they enjoy about it. Use their answers to redesign one reinforcement mechanism.
  • Separate your adoption mechanics from your engagement mechanics explicitly. Decide which gamification elements are for onboarding and which are for sustained use. Audit each category separately with different criteria.
  • Set a 30-day mission behavior review. Check whether mission-aligned behaviors — not just usage metrics — are increasing. Adjust the reinforcement design based on what you find, not what the dashboard makes look good.

For more on building AI adoption that actually sticks, read The Complete Guide to AI Product Strategy for Faith-Tech and Ministry Leaders and Three Guardrails Every Product Leader Needs Before Experimenting with AI.

Getting AI adoption right in faith-tech requires designing for motivation that sustains, not engagement that spikes. I consult with product leaders on building adoption strategies that reflect the mission rather than quietly undermine it. Let’s talk.

The Complete Guide to AI Product Strategy for Faith-Tech and Ministry Leaders

Most faith-tech teams get their AI strategy wrong in the same way: they start with the tool. They hear about a new model, a new API, a compelling demo — and they spend months building a feature that either gets ignored by users, erodes trust with the community, or drains a budget that was never designed for the cost structure of generative AI. The problem wasn’t the technology. It was the sequence. They bought answers before they understood the questions.

This guide is the resource I wish had existed when I started advising faith-tech product teams on AI strategy. It covers the full picture: how to define the right problems for AI to solve, how to evaluate tools without getting distracted by demos, how to build the team capacity and governance that make AI adoption stick, and how to measure whether any of it is working. Every section links to deeper dives where I’ve written more extensively on that topic.

I’ve led product at one of the largest digital Bible platforms in the world, built ministry tools used by hundreds of thousands of volunteers, and advised teams navigating the specific pressures of building for faith communities. What you’ll find here is practitioner knowledge, not theory — the kind of framework that holds up when you’re in a leadership meeting and someone asks why you’re spending money on AI.

Part 1: Start With the Problem, Not the Tool

The single most common AI strategy failure I see in faith-tech is treating AI as a solution looking for a problem. Leadership hears that “everyone is doing AI” and tasks the product team with integrating it somewhere. The team picks a feature area that seems feasible, builds something, and ships it to users who were never asking for it.

The antidote is problem-first thinking. Before any discussion of tools or timelines, a useful AI strategy starts with three questions: What are the most painful, repeated problems your users face? Which of those problems currently require human intelligence to solve — and is that the right constraint, or just an artifact of how things have always been done? And where does your organization have data that, if analyzed well, would surface insights you’re currently missing?

In faith-tech specifically, the most tractable AI problems tend to cluster around three areas. First, content personalization at scale — helping ministry leaders, volunteers, or individual users get the right resource for their specific context without requiring manual curation. Second, support and response — handling the high volume of repetitive questions that hit ministry teams and church staff, freeing up human attention for the conversations that actually require it. Third, insight from support data — mining the tickets, emails, and feedback that accumulate in any ministry platform for the signals that should be driving your product roadmap.

That third one is underutilized. Most faith-tech teams I work with are sitting on a goldmine of user signal in their support queues and never reading it systematically. AI makes that analysis tractable at a scale no human team can match. Read more on this in Support Data Isn’t Just Feedback — It’s Your AI Strategy’s Fuel.

Part 2: Evaluate AI Tools Without Getting Distracted by Demos

Every AI vendor has a compelling demo. The demo is optimized to show the best case: clean input, ideal output, happy path from start to finish. Your users will encounter the edge cases. Your team will encounter the integration complexity. Your finance team will encounter the cost structure. None of those show up in the demo.

When evaluating AI tools for a faith-tech product, I use four criteria that cut through the noise.

Problem fit: Does this tool actually solve one of the specific problems you identified in Part 1? Not “could it theoretically help with something in our space” — does it directly address a named user pain point? If you can’t answer yes without hedging, you’re buying on hope, not evidence.

Cost structure at scale: Generative AI has a cost model most product teams don’t fully price in during evaluation. Token costs, API call volume, and the infrastructure required to serve AI features to a large user base can turn a seemingly affordable tool into a budget crisis within weeks of launch. I’ve watched teams burn through a quarterly budget in a month because nobody modeled the usage math. Read the full breakdown in AI Costs Are Skyrocketing: How Product Leaders Should Budget for AI Before It Bites Them.

Data privacy and trust: Faith-tech users share deeply personal data — prayer requests, spiritual struggles, giving records, family information. Any AI tool that ingests that data requires explicit scrutiny about where data goes, how it’s used for model training, and what your users would think if they understood the arrangement. The trust your platform has with its community is your most important asset. It’s also the easiest thing to destroy with one poorly disclosed data practice.

Integration complexity: The gap between “this API can do X” and “X is reliably working in production with our existing stack” is where most AI projects actually die. Evaluate tools against your current infrastructure, your team’s technical capacity, and the maintenance burden of a new dependency — not just against the feature it enables.

AI prototyping tools can help you explore fit quickly before committing to full integration — but most teams use them wrong. They prototype the feature they already planned to build rather than using the prototype to test whether the problem is real. More on this in AI Prototyping Tools Are Solving the Wrong Problem.

Part 3: Build Guardrails Before You Build Features

AI features in faith-tech carry unique risks that secular product teams don’t have to think about as carefully. When your product serves people in moments of spiritual seeking, grief, crisis, or formation, an AI response that’s inaccurate, theologically inconsistent, or simply cold can do real harm — not just produce a bad user experience.

Three guardrails belong in every faith-tech AI strategy before any feature ships.

Theological boundary constraints: Your AI features need explicit boundaries around what they will and won’t say on theologically sensitive topics. This isn’t about making the product less useful — it’s about being honest with users about what they’re interacting with and preventing the product from becoming an unintended theological authority. The guardrail isn’t “never engage religious content.” It’s “be clear about what the AI is and isn’t qualified to do.”

Human escalation paths: Any AI feature that touches user wellbeing, spiritual care, or pastoral context needs a clear path to a human. Not buried in settings — visible, easy, and genuinely warm when the user reaches it. The AI feature is a first response layer. It is not a replacement for the human connection that is, in many cases, the actual product your ministry organization is delivering.

Output review and iteration loops: AI outputs in production will surprise you. Build a review process for surfacing problematic outputs before they become user complaints — and build the feedback loop that gets those examples back into your prompt engineering and model behavior. Shipping and forgetting is how AI features erode trust over time. For a full treatment of how to think about these guardrails, read Three Guardrails Every Product Leader Needs Before Experimenting with AI.

Part 4: Build the Team Mindset for AI Adoption

The most common reason AI features fail in faith-tech isn’t technical. It’s organizational. The team doesn’t understand the goal the feature is supposed to achieve, so they can’t make good decisions about how to build it or how to tell if it’s working. Or the team is trained on how to use the tool but not on why the tool exists in context of the mission. Or leadership made the decision to adopt AI without bringing the team into the problem-framing, so adoption is performative rather than genuine.

Successful AI adoption in faith-tech organizations requires three things before any feature goes into development. First, a single, specific mission outcome that the AI feature is supposed to move — not “improve efficiency,” but something measurable like “reduce the time ministry volunteers spend on lesson prep from 45 minutes to 15 minutes.” Second, a shared understanding across the team of how the feature connects to that outcome. Third, one mission metric — beyond engagement and revenue — that you’ll track to know whether the outcome is being achieved.

I’ve watched teams skip all three and ship AI features that everyone agrees are technically impressive and organizationally meaningless. Read more on the mindset side of AI adoption in AI Transformation Isn’t About Tools — It’s About Team Mindset.

There’s also a deeper question about agency. AI tools create leverage, but leverage amplifies whatever direction the team is already moving. Teams with high agency — people who feel empowered to define problems, challenge assumptions, and own outcomes — use AI to do genuinely new things. Teams without agency use AI to do the same things faster. The limiting factor in most faith-tech AI adoption isn’t the tool. It’s whether the people using the tool have the authority and confidence to actually change things. Read more in Agency Over Automation: Why AI Won’t Replace the Leaders Who Know How to Use It.

Part 5: Manage AI Infrastructure Costs Before They Manage You

AI infrastructure decisions are product leadership decisions. Most faith-tech product leaders have delegated this entirely to engineering, which means they’re surprised when the cost reports come in and they have no framework for evaluating the tradeoffs.

There are four cost levers in AI product infrastructure that every product leader should understand. Model selection — the difference between using a large frontier model and a smaller specialized model can be a 10x cost difference for comparable quality on a narrow task. Caching — many AI features make essentially identical calls repeatedly; caching common outputs eliminates cost without degrading the user experience. Batching and async processing — not every AI feature needs a real-time response; moving non-urgent AI calls to batch processing significantly reduces cost. And rate limiting by user tier — not all users need the same level of AI access, and tiering your AI feature access by subscription level is both a cost management strategy and a product differentiation opportunity.

Beyond usage costs, infrastructure resilience matters differently for faith-tech than for most SaaS. When your platform serves ministry moments — a volunteer prepping a lesson, a congregation during a live service, a leader making a pastoral decision — downtime is not just a business inconvenience. It’s a trust breach with a community that was counting on you. Read more on how to think about this in Cloud vs. Self-Hosting for Faith-Tech: How to Choose Infrastructure That Won’t Break Under Pressure.

Part 6: Measure What Actually Matters

Standard product metrics — daily active users, retention rate, revenue — measure market traction. They don’t measure mission fidelity. A faith-tech product can hit every growth target while systematically moving away from the users it was built for and the outcomes it was designed to create.

Every AI feature in a faith-tech product should have at least one mission metric alongside the standard engagement metrics. What is the human outcome this feature is supposed to produce? Are users experiencing that outcome? How would you know if they weren’t?

For AI features specifically, I track three measurement layers. First, technical performance — is the feature producing accurate, appropriate outputs consistently? Second, user behavior — are users engaging with the feature in the way you expected, and completing the task the feature was designed to help with? Third, mission outcome — is the underlying user problem actually being solved, and is that solution serving the broader purpose of the platform?

The mission metric layer is where most teams drop off. It requires knowing what your product is actually for — a discipline that connects directly to the governance question. Read more in Mission-Aligned Governance: How Faith-Tech Products Stay True to Purpose Under Growth Pressure.

Part 7: Putting It Together — An AI Strategy Framework for Faith-Tech

An AI strategy for a faith-tech product isn’t a list of features. It’s a set of decisions about where AI belongs in your product, what problems it’s solving, how you’ll govern it, and how you’ll know if it’s working. Here’s the one-page version of how I’d structure it.

Problem definition (Week 1–2): Interview your highest-value users and your support team. Identify the top 3 repeated, painful problems that currently require human judgment or manual effort. Rank them by user impact and organizational cost. These are your AI candidates — not the other way around.

Tool evaluation (Week 3–4): For each candidate problem, identify 2–3 tools or approaches. Evaluate on problem fit, cost structure at scale, data privacy, and integration complexity. Build one small prototype against the most promising option — not to validate the feature, but to test whether the problem is real and the tool is the right fit.

Guardrail design (Week 4–5): Before writing a line of production code, document your theological boundary constraints, human escalation path, and output review process. Get these reviewed by someone who understands your community’s trust relationship with the product.

Team alignment (Week 5–6): Write the mission outcome your first AI feature is supposed to achieve. Share it with everyone who will build, market, or support the feature. Get explicit buy-in — not on the feature, but on the outcome. Identify your mission metric and start tracking a baseline.

Pilot and iterate (Month 2–3): Ship to a small segment of users. Track technical performance, user behavior, and mission outcome metric. Review AI outputs weekly for the first month. Adjust based on what you learn — not what you assumed.

Governance integration (Ongoing): Add AI feature performance to your quarterly mission review. Ask whether the feature is still serving who you said it would serve, whether the cost structure is sustainable, and whether user trust is intact. AI strategy isn’t a one-time decision. It’s an ongoing discipline.

Your Turn: Apply This Today

Use this checklist to assess where your AI strategy stands right now and identify your most important next move.

  • Name your top three AI-candidate problems. Interview five users and your support team this week. List the most painful, repeated problems where AI could plausibly help. If you can’t name three, you’re not ready for an AI feature — you’re ready for user research.
  • Audit your current AI cost exposure. If you’re already running AI features, pull your last 90 days of API costs and model them against 3x and 10x user growth. Know your cost cliff before you hit it.
  • Write your guardrails document. One page: what your AI will and won’t say, how users reach a human, and how you’ll review outputs. If you ship an AI feature without this, you’re one bad output away from a community trust crisis.
  • Define one mission outcome metric. Pick the AI feature you’re most excited about and write the human outcome it’s supposed to produce in one measurable sentence. If you can’t write it, the feature isn’t ready.
  • Schedule your first AI governance review. Block 60 minutes this quarter to evaluate whether your current AI features are serving your mission. Put it on the calendar before you forget.
  • Read the supporting posts. Each section of this guide has a deeper-dive post linked. Pick the one that addresses your most pressing current challenge and start there.

Building an AI strategy for a faith-tech product is harder than building one for a secular SaaS — not because the technology is different, but because the stakes are different. Your users bring more of themselves to your product than they bring to most software. That’s a responsibility that deserves more than a rushed feature sprint.

I work with faith-tech product leaders on exactly this: building AI strategy that’s technically sound, financially disciplined, and aligned with the mission that justified building the product in the first place. Let’s talk if that’s a conversation worth having.


Get Weekly Posts on Faith-Tech Product Leadership

I write weekly about AI strategy, product leadership, and building mission-driven technology. If you found this guide useful, send me a quick email at joshua.gray.read [at] gmail [dot] com with “Subscribe” in the subject line — I’ll add you to my list and send new posts directly to your inbox.

No spam. Just the kind of practitioner-to-practitioner writing that helps you do the work better.

Why Freemium Breaks for Faith-Tech Products—and How to Fix It

A friend once told me I was throwing my life away. I wasn’t buying a house. I wasn’t funding a 401(k). I was living on a missionary stipend in South Africa with my wife and four kids, working out of a laptop. From the outside, it looked irrational. From the inside, I was accumulating something that didn’t show up on a balance sheet — trust, relationships, and a deep understanding of what people actually needed versus what they said they wanted.

I think about that a lot when I look at how faith-tech teams design their freemium models. They’re optimizing for the balance sheet metric — conversion rate, feature adoption, upgrade revenue — and missing the thing that actually drives it. Freemium works in consumer SaaS because users are rational feature-seekers. It breaks in faith-tech because your users aren’t buying features. They’re looking for something closer to what that beginner Bible Gateway user is looking for at 2am: signal in the noise. Peace in a crazy world. You can’t gate that behind a paywall and expect them to upgrade.

The Jobs-to-Be-Done Problem With Freemium

Clayton Christensen’s Jobs-to-be-Done framework asks a deceptively simple question: what job is the user actually hiring your product to do? Not what features they use — what problem they’re trying to solve in their life, and why they’re hiring your product over every other option available to them.

For most faith-tech products, the job isn’t “access content efficiently.” It’s something closer to “help me feel connected to my faith in the 10 minutes I have before work” or “give me the confidence that this message will land on Sunday.” Those jobs have emotional and relational weight that standard freemium design completely ignores. When your free tier delivers partial access to content, you’re not failing to convert users because they don’t want to pay. You’re failing because the free tier never did the job they hired it for. They disengage — not because they’re cheap, but because you never built enough trust for them to care what’s behind the paywall.

What a Value-First Free Tier Actually Looks Like

On the curriculum platform I spent years building, we made a deliberate choice: the free tier had to be genuinely complete for the core job. Our primary user was a volunteer with seven minutes to prep a lesson before their group walked in the door. If the free tier didn’t make that person feel ready — actually ready, not “ready enough if they upgrade” — we had failed before the conversion conversation even started.

Premium added community features, deeper analytics, and customization. But the job — make this volunteer feel prepared and confident — was fully done in free. The result: a free tier that built real trust, and a conversion path that didn’t feel like a bait-and-switch. People upgraded because they believed in the product, not because we’d frustrated them into it.

For AI-powered faith-tech tools, the same principle holds. If your AI helps with sermon prep, let the free version generate a full, usable outline with cultural context and practical application. Then charge for depth, integration, and personalization. But if the free version produces half an outline and stops, you haven’t demonstrated value — you’ve demonstrated that you don’t trust your own product enough to let people experience it.

Why Community and Trust Drive Conversion in Faith-Tech

In most SaaS markets, conversion is driven by feature hunger. Users hit the ceiling of the free tier and upgrade to get more. In faith-tech, conversion is driven by trust and community fit. Users upgrade when they believe the product genuinely understands their mission — when it feels like a partner, not a vendor.

That distinction changes everything about what your free tier needs to accomplish. It’s not just about demonstrating value — it’s about demonstrating that you see your user. You understand what they’re trying to do. You built this for them, specifically, not for the average SaaS conversion funnel. Ministry users have good instincts for when a product is built by someone who understands their context versus when it’s a generic tool with a faith-flavored coat of paint. The free tier is where they make that call.

Your onboarding, your empty states, your error messages, your feature framing — all of it either confirms or undermines that impression. This isn’t a design detail. It’s your conversion strategy.

Your Turn: Apply This Today

Before your next freemium redesign or pricing conversation, use this checklist to diagnose whether your model is actually built for your users.

  • Interview five users about the job they hired your product to do. Ask what they were struggling with before they found you, and what “done” looks like for them. Let the answers surprise you — they usually do.
  • Audit your free tier against the core job. List every step a free user takes to complete their primary goal. Identify every point where they hit a gate — and ask honestly whether that gate is protecting your business or just frustrating your user.
  • Write one sentence describing what the free tier accomplishes. Not what features it includes — what it actually does for the user. If you can’t write that sentence, the free tier isn’t solving a job.
  • Test a more complete free experience with 50 users. Roll it out small, track whether engagement beyond the first session improves, and measure trust — not just feature adoption.
  • Survey those users on one question: did this feel built for you? Not “did you like the features.” Did the product feel like it understood your context. Adjust based on what they tell you.
  • Redefine your conversion goal. Set a metric for trust-driven conversion — not raw upgrade rate. Define what it looks like when someone upgrades because they believe in the product, not because you’ve withheld enough to force the decision.

For more on building product strategy around real user motivation, read AI Costs Are Skyrocketing: How Product Leaders Should Budget for AI Before It Bites Them and What Retention Data Actually Tells You (And What It Doesn’t).

Freemium strategy in faith-tech isn’t a pricing decision — it’s a statement about whether you actually understand your users. I consult with faith-tech product leaders on building conversion paths that are built on trust, not tricks. Let’s talk.

AI Transformation Isn’t About Tools—It’s About Team Mindset

A church tech leader emailed me recently with a question I’ve been getting more often than I expected: “We’ve invested in AI tools — better content pipelines, automated workflows, smarter outreach — but nothing’s actually changed. Why isn’t this working?”

I’ve learned to answer this one directly: you’re solving a technology problem when you have a people problem. The tools aren’t the issue. The missing piece is whether your team has the mindset, the clarity, and the shared understanding of what you’re actually trying to accomplish. Without that, the best AI stack in the world produces expensive activity that doesn’t move anything important.

I’ve watched this play out in faith-tech organizations and at scale in secular SaaS alike. The pattern is the same: leadership gets excited about AI’s potential, rolls out new tools, runs feature training, and waits for results. The results don’t come — not because the tools are bad, but because nobody defined what success looks like or connected the tool to a specific outcome anyone actually cares about. The team learns the features. They don’t change what they do.

Why AI Transformation Stalls at the Mindset Level

The shift I’ve had to make myself — and that I watch other leaders resist — is from thinking of AI as a capability to acquire and thinking of it as a way of working that has to be built. Buying a tool is a transaction. Building a team that knows how to use AI well is an organizational change project. Those require completely different leadership approaches, and conflating them is why most AI transformation efforts produce underwhelming results.

In my experience, the teams that actually shift how they work with AI share three things. First, someone — usually a senior leader — uses the tools personally and talks about what they’re learning. Not just encourages others to adopt. Actually does the work themselves and shares what surprised them. Second, the team has a clear, specific outcome they’re trying to affect — not “improve our content” but “reduce the time our volunteers spend on lesson prep from 45 minutes to 15.” Third, early wins are documented and shared in a way that makes the outcome visible to the whole team, not just the metrics dashboard.

The Mission Connection That Most Teams Skip

In faith-tech, there’s an additional layer that secular product teams don’t have to navigate: the team’s relationship to the mission can either accelerate AI adoption or completely stall it. I’ve seen both.

When AI adoption is framed as efficiency — “this tool will save you two hours a week” — faith-tech teams often disengage quietly. Not because they don’t want to save two hours. But because they didn’t get into ministry technology to optimize workflows. They got into it because they believe the work matters. The efficiency framing doesn’t connect to that.

When AI adoption is framed around the people being served — “this tool means the volunteer who has seven minutes to prep a lesson actually feels ready instead of overwhelmed” — the same team engages completely differently. Same tool. Different frame. The difference is whether you’ve connected the technology to the reason the team shows up.

That connection has to be made explicitly and repeatedly, not assumed. Most AI adoption rollouts assume the mission connection is obvious. It isn’t. You have to make it.

Your Turn: Apply This Today

Before your team’s next AI rollout or adoption push, run through this checklist to make sure you’re solving the real problem.

  • Name one specific mission outcome your AI tool should affect. Write it in a single sentence — measurable, time-bound, tied to a real person you serve. No vague language. If you can’t write it, pause the rollout until you can.
  • Connect the tool to that outcome explicitly. Document how, specifically, this tool moves the needle on that outcome. Walk through a real scenario with your team before feature training starts.
  • Use the tool yourself before you roll it out. Not a demo. Actual use, on actual work, with enough repetition to form an honest opinion about what it does and doesn’t do well. Then share that honestly with your team.
  • Run a 30-minute session on the why before any feature training. Walk your team through a specific story about the person the tool is designed to help. Make the mission connection visible before you open the tutorial.
  • Track one mission outcome metric for 30 days — not tool usage. Completion rates. Volunteer confidence scores. User engagement depth. Something that reflects what you’re actually trying to change, not just whether people logged in.
  • Review progress honestly in your next team meeting. Name what’s not working. Adjust the approach, not just the tool settings. The team needs to see that you’re optimizing for the outcome, not protecting the technology investment.

For more on building AI strategy that actually sticks, read AI Prototyping Tools Are Solving the Wrong Problem and Agency Over Automation: Why AI Won’t Replace the Leaders Who Know How to Use It.

AI adoption is a leadership challenge, not a technology challenge. I consult with product leaders and church tech teams on AI strategy, team alignment, and building the organizational mindset that makes new tools actually work. Let’s talk.

Cloud vs. Self-Hosting for Faith-Tech: How to Choose Infrastructure That Won’t Break Under Pressure

Here’s a number that should give faith-tech product leaders pause: a 2024 Flexera report found that 87% of enterprises use multiple cloud providers — yet most teams treat infrastructure as a set-and-forget decision. You pick AWS or Azure, spin up your stack, and move on to the product work that actually feels interesting.

I’ve watched that assumption collapse in real ministry contexts. A children’s ministry curriculum platform I worked with had its live-stream go down mid-service on Christmas Eve — a cloud provider outage that nobody had a fallback plan for. It wasn’t just a tech failure. It was a trust breach with every family watching at home. The team had optimized for convenience and gotten fragility in return.

Cloud infrastructure feels like the obvious choice: scalable, cost-effective, someone else’s problem. But for mission-driven products — where the audience isn’t just a user base but a community with deep emotional stakes — blind cloud dependence can be a hidden liability. The question isn’t whether to use cloud. It’s how to build infrastructure that doesn’t crack when it matters most.

Cloud Gives You Scale. It Doesn’t Give You Resilience.

Nassim Taleb’s concept of antifragility is useful here. Fragile systems break under stress. Robust systems endure it. Antifragile systems get stronger from it. Most cloud-dependent faith-tech products are fragile — they’re built for normal load, and they shatter when something unexpected hits.

I’ve seen this play out beyond outages. A global discipleship app team I worked with faced a mid-contract pricing hike from their cloud provider that consumed 30% of their annual infrastructure budget overnight. They had no leverage and no alternative. The mission didn’t pause; the scramble to find budget did.

Cloud is excellent for elasticity — handling millions of users during a viral moment, scaling down when traffic drops. But elasticity isn’t the same as resilience. When your product is a lifeline for ministry volunteers with seven-minute windows to prep lessons, you need both.

When Self-Hosting and Hybrid Models Actually Make Sense

Self-hosting carries a bad reputation in most product circles — more overhead, more expertise required, more things to break. That reputation is largely earned. But dismissing it entirely ignores cases where control matters more than convenience.

The discipleship app team I mentioned earlier moved to a hybrid model after the pricing shock. They kept cloud for frontend delivery and user-facing features — where elasticity matters — but self-hosted their user data and core content. They regained control over their most critical assets without abandoning the scalability they needed.

For a children’s ministry curriculum platform, we cached lesson content locally as a fallback. The seven-minute volunteer window couldn’t tolerate a cloud dependency. That cache cost almost nothing to implement and eliminated the single point of failure that had been invisible until Christmas Eve.

The framework I use is simple: identify your product’s non-negotiables — the user experiences that cannot fail — and trace every dependency in that path. Wherever you find a single external point of failure, that’s where you need a fallback strategy, whether that’s a cache, a redundant provider, or a self-hosted component.

Making the Cloud vs. Self-Host Decision

Most teams pick infrastructure based on what the founding engineer knew or what the first CTO preferred. That’s not a strategy; it’s inertia. The right decision framework starts with your mission’s non-negotiables and works backward to infrastructure requirements.

Ask yourself: what does the product need to do when everything goes wrong? If the answer is “stay up and serve users” — and for most faith-tech products it is — then every architectural decision needs to account for failure modes, not just happy-path performance.

Cloud-first makes sense when scale and global reach are primary. Hybrid or self-hosted components make sense when privacy, cost control, or reliability of a specific feature are non-negotiable. Most faith-tech products need both in different parts of their stack.

Your Turn: Apply This Today

Infrastructure decisions feel abstract until something breaks. Use this checklist to pressure-test your setup before that happens.

  • Map every cloud dependency. List every external service your product relies on — provider, function, and what breaks if it goes down for an hour.
  • Identify your non-negotiables. Name the one or two user experiences your product cannot afford to fail on — these define where you need the most resilience.
  • Stress-test one critical path. Simulate a cloud outage on your most-used feature and document what actually breaks versus what you assumed would.
  • Calculate your 12-month cost exposure. Factor in potential fee hikes, overage charges, and downtime costs — compare that number to a hybrid option for your highest-risk dependency.
  • Build one fallback. Pick your highest single point of failure and implement a simple fallback — a cache, a secondary provider, or a self-hosted component — before next month.
  • Write a crisis response plan. Document three steps your team takes if your primary cloud provider goes offline. Make it short enough to execute under pressure.

For more on building product systems that hold up under real-world pressure, read AI Costs Are Skyrocketing: How Product Leaders Should Budget for AI Before It Bites Them and Agency Over Automation: Why AI Won’t Replace the Leaders Who Know How to Use It.

Resilient infrastructure is a product leadership decision, not just an engineering one. I consult with faith-tech product leaders on infrastructure strategy, cost discipline, and building systems that serve your mission when it matters most. Let’s talk.

AI Prototyping Tools Are Solving the Wrong Problem

A product manager at a faith-tech startup asked me after a workshop: “Are AI prototyping tools worth the hype? Should my team jump in?” My answer was direct: they can be genuinely useful — but only if your team has already done the discovery work that tells you who you’re building for and what they actually need. Without that, AI prototyping doesn’t accelerate your work. It accelerates your mistakes.

I’ve watched this play out on a curriculum tool I worked on for children’s ministry volunteers. We had AI tools that could generate wireframes and mockups overnight. The output was slick and fast. What we didn’t have was a clear picture of how volunteers actually used the product — that they were prepping lessons in seven-minute windows, on their phones, while something else was happening in the background. Once we spent time with real users, everything changed. The prototypes we’d built before that conversation were solving for the wrong constraints entirely.

The Discovery Gap That AI Prototyping Can’t Fill

The promise of AI prototyping is speed — faster wireframes, faster iteration, faster path from idea to testable artifact. That promise is real. The problem is that speed only creates value when it’s pointed at the right problem. Teams that skip discovery and reach for prototyping tools first end up with a faster path to the wrong answer.

Teresa Torres’ continuous discovery framework makes this argument precisely: product teams need to be constantly testing assumptions with real users, iterating on outcomes rather than outputs. AI prototyping can accelerate the artifact side of that loop, but it cannot generate the user insight that makes the artifact worth building. That insight only comes from talking to users — and most teams are doing far less of that than they think.

The tell that a team has skipped discovery: their prototype feedback sessions produce surprises. Users don’t understand the flow. The feature solves a problem the team assumed existed but users don’t recognize. The navigation makes sense to the team and nobody else. These aren’t AI prototyping failures — they’re discovery failures that AI prototyping revealed faster than a slower process would have. The tool surfaced the problem. The problem was upstream.

When Speed Becomes a Liability

There’s a specific failure mode I’ve seen in teams that adopt AI prototyping without strengthening their discovery practice first: they ship faster, get feedback faster, and — because they haven’t grounded the feedback in clear user outcomes — misread what the feedback means. They iterate on surface signals rather than structural problems. The prototype improves visually while the underlying user need goes unaddressed.

On a cross-functional team I worked with, we used AI tools to generate a series of mockups for a new feature in rapid succession. The PMs were hitting deadlines. The designers were producing great work. The engineers were ready to build. And the users, when we finally got in front of them, had no idea what the feature was for. We’d built consensus within the team about what to build, but we’d never built consensus with the users about the problem we were solving.

Torres’ framing is useful here: teams that do continuous discovery are constantly checking their assumptions against reality. Teams that don’t are building confidence in a direction without checking whether the direction is right. AI prototyping accelerates confidence-building. It doesn’t check whether the confidence is warranted.

How to Use AI Prototyping Correctly

The right sequence is discovery first, prototyping second. That means before any AI tool generates a single wireframe, the team should be able to answer three questions: Who is the specific user this prototype is for? What decision or task are we trying to make easier for them? What does success look like for that user, in their terms?

When those three questions have answers grounded in actual user conversation — not assumption — AI prototyping becomes a genuine accelerant. You know what to build. You have a clear evaluation criterion. User testing produces signal rather than surprise. The tool is pointed at a real target.

The teams I’ve seen use AI prototyping most effectively treat it as a way to test specific hypotheses faster, not as a way to generate options when direction is unclear. That distinction matters. Options without direction produce more meetings. Hypothesis tests with direction produce learning.


Your Turn: Apply This Today

Before your team runs another AI prototyping session, make sure the discovery foundation is in place. Here’s how:

  • Answer the three questions before opening the prototyping tool. Who specifically is this for? What task or decision are we making easier for them? What does success look like in their terms? Write the answers in the meeting doc before anyone opens Figma or any AI design tool. If the team can’t agree on the answers, that’s a discovery gap — not a design problem.
  • Schedule user interviews before your next major prototype. Not after — before. Even two 30-minute conversations with real users will surface assumptions you didn’t know you were making. The time cost is real. The cost of prototyping in the wrong direction is higher.
  • Run a “who is this for?” check on every active prototype. Look at your current prototyping work and ask: can we name a specific real user whose life this will improve, and can we describe specifically how? Prototypes that can’t pass this check are building team alignment, not user value. Both have their place — but they’re not the same thing.
  • Separate exploration prototypes from validation prototypes. Exploration prototypes are for generating team ideas. Validation prototypes are for testing specific user hypotheses. AI tools are excellent for exploration. They’re only useful for validation if the hypothesis is already sharp. Know which mode you’re in before you start.
  • Debrief on what you learned from the users, not just what you’ll build next. After every user test, the first question in the debrief should be: what assumption turned out to be wrong? Not “what should we change in the prototype?” — what did we assume that the user didn’t confirm? That learning is the asset. The updated prototype is the output.
  • Read one chapter of Torres’ “Continuous Discovery Habits” with your team this month. It’s the clearest articulation I’ve found of what a disciplined discovery practice actually looks like in a product team. Pick one chapter, discuss it in your next team meeting, and identify one habit to adopt. The investment pays back faster than most product tools will.

AI prototyping and discovery are both part of the broader question of how teams stay grounded in user reality as AI capabilities expand. Why Teresa Torres’ Continuous Discovery Cadence Is Breaking Down in the AI Age addresses how the discovery framework itself needs to evolve when AI is part of the solution space. And How AI Prototyping Can Break Silos in Faith-Tech Products covers the team alignment dimension of prototyping that this post focuses on the discovery dimension of.

Trying to build a stronger discovery practice in your product team so your AI investment is pointed at the right problems? I consult with product leaders on user research systems, discovery cadences, and the habits that close the gap between what teams assume and what users actually need. Let’s talk.

Three Guardrails Every Product Leader Needs Before Experimenting with AI

The most dangerous moment in any AI product initiative isn’t the launch. It’s the planning session where someone says “let’s just try it and see what happens.” I’ve been in that room. The energy is high, the timeline is short, and constraint feels like the enemy of progress. It isn’t. Constraint is what keeps experimentation from becoming chaos.

I learned this the hard way working on a project for children’s ministry resources. We had the technical capability to use AI to auto-generate lesson plans at scale. The business case was obvious. The engineering was ready. What we hadn’t done was define what harm would look like if we got it wrong — and in a context where volunteers were trusting us with theologically sensitive material, that was the question that mattered most. We paused. We defined the guardrail. We shipped something better because of it.

Why Guardrails Accelerate AI Experimentation, Not Slow It Down

There’s a counterintuitive truth about constraint in product work: the teams that define boundaries before they start experimenting move faster than teams that don’t. Not because they’re more cautious, but because they spend less time relitigating the same decisions. Every time a team without guardrails hits an ambiguous call — “should we use this user data for this?” “is this recommendation too aggressive?” — they stop to debate it. Teams with guardrails already know the answer. The constraint bought them speed.

John Wesley’s Three Simple Rules — Do No Harm, Do Good, Stay in Love with God — were designed as a framework for community life, not product strategy. But I’ve found them to be one of the clearest models I know for scoping AI experimentation in mission-driven organizations, because they’re sequenced by priority and genuinely hard to rationalize around.

Guardrail One: Define What Harm Looks Like Before You Build

“Do No Harm” sounds obvious until you try to specify it. Harm in an AI context is surprisingly easy to overlook because it’s often diffuse, delayed, and hard to attribute. A recommendation engine that prioritizes engagement over depth doesn’t harm anyone in a visible way on day one. Over six months, it reshapes behavior in ways that may run directly against the product’s stated mission. That’s harm. It just doesn’t show up in your sprint metrics.

Defining harm specifically, before you build, forces the conversations that matter. In our children’s ministry project, the harm question surfaced a concern nobody had articulated: what happens when AI-generated content contains a theological error that a volunteer doesn’t catch? That’s a trust problem, not just a content quality problem. We added a human review gate. The product shipped later and was far more trusted by the volunteers who used it.

For your context, harm might be data privacy for members, cultural insensitivity in generated content, algorithmic bias in which users get which experiences, or recommendations that drive the wrong behaviors. The form it takes is specific to your product and your users. Name it explicitly in your feature spec before engineering starts.

Guardrail Two: Measure “Doing Good” Before You Claim It

“Do Good” is the guardrail most product teams think they’ve already cleared. They haven’t — they’ve assumed it. There’s a meaningful difference between “this feature is directionally good for users” and “this feature measurably improves a specific outcome for a specific user in a way we can verify.” The first is a hypothesis. The second is a product decision.

I’ve watched teams deploy AI features in faith-tech contexts — sermon prep tools, discipleship engagement nudges, content recommendation engines — that were built with genuinely good intentions and produced genuinely mixed results. Not because the technology was wrong, but because “doing good” was never defined precisely enough to know whether the feature was achieving it. Usage metrics went up. Whether people’s engagement with faith deepened was never measured. That’s a guardrail that was never set.

Before any AI feature ships, I want a one-sentence definition of what “doing good” means in measurable terms: “Users who engage with this feature will [specific behavior] at [measurable rate] within [time window].” If the team can’t write that sentence, the feature isn’t ready to build.

Guardrail Three: Stay Grounded in the Mission When the Hype Pulls You Away

“Stay in Love with God” — translated for product teams — means staying grounded in the specific mission that justifies the product’s existence, especially when vendor pitches and industry trends are pulling you toward shiny features that have nothing to do with it. This is harder than it sounds in 2026. The AI capability landscape changes fast enough that every quarter there are new things your product “could” do. The question is whether doing them serves the people you’re actually building for.

I’ve been in strategy sessions where the room spent an hour excited about a generative AI capability that had no clear connection to user outcomes. The capability was impressive. The use case was thin. The right move was to say no — not because the technology was bad, but because chasing it would have consumed engineering cycles that belonged to the actual problem the product existed to solve.

This guardrail is a practice, not a policy. It requires someone in the room willing to ask “how does this directly serve our users?” without apologizing for slowing the conversation down. Usually that person is the product leader. Make it a habit.


Your Turn: Apply This Today

Before your next AI feature enters the roadmap, run it through these three guardrails explicitly — not as a checklist to complete, but as questions that need honest answers:

  • Define harm for your specific context in writing. For the AI feature currently in planning, name the two or three ways it could harm users if it doesn’t work as intended. Data exposure, bad recommendations, eroded trust, misrepresented information — whatever is relevant. Add this to the feature spec. If you can’t name the harm, you haven’t thought about it carefully enough yet.
  • Write the “doing good” sentence before scoping data requirements. “Users who engage with this feature will [behavior] at [rate] within [timeframe].” If your team can’t complete that sentence in one working session, the feature definition isn’t ready. This conversation is cheaper than a sprint of building in the wrong direction.
  • Put your mission statement in the room when you make AI prioritization decisions. Literally — write it on the whiteboard or put it in the meeting doc. When a new AI capability comes up, the first question is “how does this serve [mission]?” Features that can’t answer that question belong in a backlog, not the current sprint.
  • Say no to one trending AI feature this quarter. Not because it’s bad, but as practice. Find one thing your team has been considering because it’s interesting or because a competitor shipped it, and explicitly decline it because it doesn’t clear your guardrails. Document the decision. This builds the muscle of constraint before you need it for a harder call.
  • Run a 30-minute team session this week using the three guardrails. Pick your current highest-priority AI initiative and put it through all three questions as a group. Where do people disagree? That disagreement is useful data about misaligned assumptions that would have surfaced at the worst possible time if you hadn’t surfaced it now.
  • Interview three users about what “harm” and “help” look like from their perspective. Your internal definition of both will be incomplete without user input. Ask: “What would this product have to do to lose your trust?” and “What would genuinely change your week for the better?” Their answers will sharpen all three guardrails faster than internal debate will.

The constraint framework connects directly to how you define AI features in the first place. Why Your AI Feature Doesn’t Need More Data — It Needs Better Problem Definition addresses how to scope AI work before you start building. And Agency Isn’t Enough: Why AI-Era Product Teams Need Constraint Over Freedom covers the organizational dynamics that make guardrails hard to maintain under pressure.

Working through how to structure your team’s AI experiments so they stay grounded in mission and user outcomes? I consult with product leaders and ministry innovators on constraint-driven AI strategy, feature scoping, and the organizational habits that keep AI investment pointed at the right problems. Let’s talk.

AI Learning for Ministry Leaders: Cut the Noise and Start Saving Real Time

I spent weeks “learning AI” before I realized I just needed it to save me an hour a day. The tutorials felt important. The newsletters felt urgent. The demos were impressive. And none of it translated into actual time back in my week until I stopped trying to master the technology and started asking a simpler question: what is the most annoying, time-consuming, low-creativity task I do every week, and can AI help with that?

For ministry leaders especially — pastors, executive pastors, church operations leaders, faith-tech practitioners — time is not a productivity problem. It’s a stewardship problem. Every hour spent on administrative friction is an hour not spent with people. That’s the actual cost. And AI, when applied to the right tasks, can genuinely give some of that time back.

The AI Learning Trap Ministry Leaders Fall Into

The trap is treating AI as a mountain to climb. Every week there’s a new model, a new tool, a new capability that someone in your network is raving about. The implicit message is that you need to keep up — that there’s a level of AI fluency you’re supposed to reach before you can start using any of it effectively.

That message is wrong. You don’t need to understand how large language models work to use them well, any more than you need to understand combustion engineering to drive to a hospital visit. The relevant question isn’t “how does this work?” It’s “what specific task in my actual week could this improve?” And for most ministry leaders, that question is answerable in about twenty minutes of honest reflection.

The people who get value from AI fastest are not the ones who studied it the most. They’re the ones who identified one specific problem, tried one specific tool against it, and measured whether it helped. That’s it. The sophistication comes later, if it comes at all. The value comes from the first application, not from the curriculum.

Where Ministry Leaders Actually Get Time Back

Based on what I’ve seen across faith-tech organizations and ministry teams, there are a handful of tasks where AI consistently delivers meaningful time savings with minimal learning investment:

First drafts of routine communications. Weekly announcements, bulletin copy, email updates to congregation segments, social media copy for events — these are high-volume, low-creativity tasks that most ministry staff spend disproportionate time on. AI handles first drafts well. You still supply the voice, the pastoral context, and the final judgment. You just don’t spend forty-five minutes starting from a blank page.

Meeting notes and action item extraction. Recording a staff meeting and running the transcript through a summarization tool produces a workable set of notes and action items in minutes. The output isn’t perfect, but it’s a starting point that saves twenty to thirty minutes of post-meeting documentation per meeting.

Research synthesis. Sermon background research, background on pastoral care topics, summaries of longer documents — AI handles the aggregation and synthesis of information well. You still do the discernment. You just don’t do the initial reading and note-taking.

The common thread: these are all tasks where the bottleneck is the blank page or the initial aggregation, not the judgment. AI removes the blank page. You provide the judgment. That division is the practical model for AI-assisted ministry work.

The Discernment That AI Doesn’t Have

None of this replaces the things that actually matter in ministry. AI can draft a pastoral care email. It doesn’t know the specific grief that person is carrying or the relationship history that informs how you should communicate with them. AI can generate sermon outlines. It doesn’t know the heartbeat of your congregation this week or the spiritual weight of what they need to hear.

This isn’t a limitation to apologize for — it’s a useful boundary. The things AI doesn’t do well in ministry contexts are exactly the things that are most worth protecting your time for. If AI gets you an hour back on administrative tasks, that’s an hour you can spend on the irreplaceable work. That’s not a compromise. That’s the point.


Your Turn: Apply This Today

You don’t need a curriculum. You need a starting point. Here’s how to get real time back from AI this week:

  • List the three most time-consuming low-creativity tasks in your average week. Not the important work — the administrative, repetitive, blank-page work. These are your first AI candidates. Pick the most annoying one and spend thirty minutes trying an AI tool against it before you do anything else.
  • Try one AI tool on one real task this week — not a demo, a real task. Draft this week’s bulletin copy. Summarize this week’s meeting notes. Write the first version of the email you’ve been putting off. Use the output as a starting point, not a final product. Measure whether it saved time.
  • Stop trying to keep up. Unsubscribe from one AI newsletter this week. The marginal learning value of staying current with every new tool release is lower than the value of applying one tool well. You can’t keep up, and you don’t need to. Focus beats breadth.
  • Define your boundary explicitly. Write one sentence: “I use AI for [specific tasks]. I do not use AI for [specific decisions or interactions].” Having a written boundary prevents both the trap of using AI for everything and the trap of never using it at all.
  • Measure the time savings after two weeks. Not impressionistically — actually track it. If you’re using AI to draft communications, time how long the drafting takes versus before. If the savings are real, that’s evidence worth sharing with your team. If they’re not, adjust the application.
  • Share one practical AI win with a colleague in ministry. Not a pitch for AI adoption — just one honest story: “I used this tool for this task and it saved me this much time.” Peer-to-peer practical examples spread faster than any training program and produce more genuine adoption.

For a product-leadership perspective on the same challenge, AI Learning for Product Leaders: Stop Chasing Tools, Start Solving Problems covers the same principle for PM teams. And Agency Over Automation: Why AI Won’t Replace the Leaders Who Know How to Use It addresses where human judgment stays essential as AI handles more of the work.

Working through how to build practical AI habits into your ministry team or faith-tech organization without drowning in hype? I consult with ministry leaders and faith-tech practitioners on exactly this — cutting through the noise to find the applications that actually serve the mission. Let’s talk.

AI Costs Are Skyrocketing: How Product Leaders Should Budget for AI Before It Bites Them

Last year, I nearly blew a product budget on an AI tool that promised to transform user engagement. I fell for the slick demo, the polished pitch deck, and the assumption that token costs were negligible. Three months in, the monthly AI infrastructure bill had quietly grown to something that required an uncomfortable conversation with finance — and the engagement lift didn’t justify it.

I’m not alone in this. AI cost management has become one of the most common and least-discussed problems in product organizations. Teams scope AI features based on demo performance and initial estimates, then watch costs scale unexpectedly as usage grows. By the time the problem is visible, it’s already a crisis — because it’s now entangled with a shipped feature that users depend on.

Why AI Costs Are Harder to Predict Than Other Infrastructure

Traditional infrastructure costs scale in relatively predictable ways — more users means more compute, and the relationship is roughly linear. AI costs don’t work that way. Token pricing for LLM-based features can vary by orders of magnitude depending on model selection, prompt design, context window usage, and whether you’re doing inference, fine-tuning, or both. A feature that costs $200/month at 1,000 users might cost $40,000/month at 100,000 users — not because of linear scaling, but because of how the feature was architected.

This unpredictability creates a specific trap for product teams: they validate the feature at small scale, ship it, and don’t discover the cost problem until they’re well past the point of easy architectural change. The feature works. Users like it. And it’s quietly becoming the most expensive line item in the infrastructure budget.

In faith-tech and mission-driven organizations especially, where margins are thin and every dollar is tied to a stated mission, this trap is particularly dangerous. The organizations that get this wrong don’t just waste money — they create a credibility problem for AI investment overall, making it harder to fund the next initiative that might actually matter.

The AI Budget Discipline Most Product Teams Skip

The fix is a cost architecture conversation that needs to happen before feature development begins — not after launch. This conversation covers four things: unit economics at scale, model selection tradeoffs, prompt efficiency, and cost monitoring.

Unit economics at scale means estimating the cost per user interaction at 10x and 100x your current scale before you commit to an architecture. If the math doesn’t work at 100x, you either need a different model, a different architecture, or a different feature scope. Finding this out at 1x is cheap. Finding it out at 100x is expensive.

Model selection tradeoffs means being intentional about which model you need for which task. The most capable model is not always the right model. For many product use cases — classification, simple summarization, structured extraction — smaller, cheaper models perform comparably to frontier models at a fraction of the cost. Using GPT-4 class models for tasks that a fine-tuned smaller model could handle is a budget decision masquerading as a technical one.

Prompt efficiency matters because token costs are real. Bloated system prompts, unnecessary context, redundant instructions — these add up at scale. A prompt engineering discipline that optimizes for token efficiency alongside output quality pays dividends as usage grows.

Cost monitoring means treating AI infrastructure costs as a first-class product metric, not just an engineering concern. Product leaders who monitor cost-per-interaction alongside user engagement metrics catch problems early and make better architectural decisions.


Your Turn: Apply This Today

Whether you’re scoping a new AI feature or auditing an existing one, here’s how to build cost discipline into your AI product process:

  • Run a cost projection before any AI feature enters development. Estimate the cost per user interaction. Multiply it by your current user base, then by 10x and 100x. If the 100x number is uncomfortable, that’s important information — not a reason to kill the feature, but a reason to make intentional architectural choices now.
  • Ask “do we actually need this model?” for every AI implementation. Document the specific capability requirement. Then check whether a cheaper model meets that requirement at acceptable quality. If you haven’t tested a smaller model, you haven’t answered this question.
  • Audit your most expensive prompt. Pull the system prompt for your highest-usage AI feature and count the tokens. Look for redundancy, unnecessary context, and instructions that could be simplified. Even a 20% reduction in prompt size compounds significantly at scale.
  • Add cost-per-interaction to your product metrics dashboard. It should sit alongside engagement, retention, and error rate. If your team can’t see it, they can’t manage it. Cost visibility changes the conversation about AI feature scope and architecture.
  • Set a cost ceiling before launch. Define the maximum acceptable monthly AI infrastructure cost for this feature at your current scale and at 3x scale. Make it explicit in the feature spec. This gives engineering a clear target and surfaces tradeoffs early, when they’re still cheap to resolve.
  • Review AI costs quarterly with the same rigor as other infrastructure. AI pricing changes. Model options change. What was the most cost-effective architecture six months ago may not be today. A quarterly review ensures you’re not paying legacy prices for capabilities that have gotten cheaper.

AI budget discipline connects directly to infrastructure strategy. Jensen Huang’s Sovereign AI: Infrastructure Argument for Product Builders covers the build-vs-buy dimension of AI infrastructure decisions. And The Product Leader’s AI Infrastructure Blind Spot addresses the organizational patterns that cause AI infrastructure costs to spiral before anyone catches them.

Managing AI infrastructure costs alongside feature development and trying to make the numbers work? I consult with product leaders on AI strategy, infrastructure decisions, and the budget frameworks that keep AI investment sustainable as products scale. Let’s talk.

Agency Over Automation: Why AI Won’t Replace the Leaders Who Know How to Use It

I’ve seen AI generate a sermon outline in seconds. I’ve watched it draft a social media calendar in minutes. I’ve used it to analyze engagement data that would have taken a team a week to process manually. And in every one of those cases, the value didn’t come from the AI. It came from a person who knew which problem to point it at, how to evaluate whether the output was any good, and what to do with what it found.

The “AI will replace product managers / pastors / tech leaders” conversation misses the actual dynamic. AI doesn’t replace judgment. It makes judgment more consequential. The leaders who understand this are pulling ahead. The ones waiting to see if they’ll be replaced are falling behind — not to AI, but to the people who learned to work with it.

Automation vs. Augmentation: The Distinction That Matters

The replacement anxiety is driven by a conflation of two very different things: automation and augmentation. Automation replaces a task. Augmentation improves the quality or speed of a decision. Most of what AI does in product and ministry contexts is augmentation, not automation — and augmentation doesn’t reduce the need for the person doing the work. It raises the bar for what that person can accomplish.

When AI drafts a sermon outline, a good pastor isn’t replaced — they now have more time to do the thing AI can’t do: knowing the widow in row three, understanding the weight of what the congregation is carrying this week, and delivering something that actually meets them there. The draft is scaffolding. The sermon is still fully human.

When AI surfaces a pattern in support tickets, a good product manager isn’t replaced — they have faster access to a signal that used to take weeks to surface. The decision about what to do with that signal still requires product judgment, user empathy, and strategic context that no model has. The analysis is faster. The judgment is still theirs.

What AI Actually Replaces

Here’s the honest version: AI does replace some things. It replaces the first draft of documents that don’t require original insight. It replaces basic data aggregation that humans were doing slowly and inconsistently. It replaces repetitive categorization tasks that consumed analyst time without producing analyst thinking. These are real things — and losing them to AI is, on balance, good. They were preventing people from doing their best work.

What AI doesn’t replace: the ability to define which problem is worth solving. The judgment to evaluate whether an AI output is trustworthy in a specific context. The relationships and credibility that make a recommendation worth following. The courage to make a call when the data is ambiguous. The discernment — in ministry contexts especially — to know when the technically correct answer is the wrong pastoral choice.

These are the capabilities that compound with AI rather than competing with it. The more AI handles the mechanical, the more valuable human judgment becomes. Which means the leaders who invest in judgment — in pattern recognition, in mental models, in knowing their users deeply — are the ones whose value increases as AI capabilities improve.

The Trap: Outsourcing Ownership, Not Just Tasks

The real risk isn’t that AI replaces product leaders or ministry leaders. It’s that leaders outsource ownership to AI alongside the tasks. There’s a meaningful difference between “AI drafted this and I refined it” and “AI drafted this and I shipped it.” The first is augmentation. The second is abdication.

I’ve watched this happen in both product organizations and faith-tech teams: leaders hand first drafts to AI and stop interrogating the output. The quality degrades slowly, then quickly. Users notice before leaders do. And when something goes wrong — a feature that misses the user’s actual need, a message that lands badly with the congregation — there’s no accountability structure, because nobody owns the decision.

Agency over automation means staying in the decision loop even when AI is doing the heavy lifting. It means knowing enough about what your AI tools are doing to evaluate their output critically. It means treating AI-assisted work as your work — with your name on it and your judgment applied to it.


Your Turn: Apply This Today

Here’s how to build an “agency over automation” practice that keeps you in the decision loop as AI takes over more of the mechanical work:

  • Audit what you’re outsourcing vs. what you’re augmenting. For every AI tool you currently use, ask: am I using this to do my thinking faster, or to avoid thinking? The first is augmentation. The second is a skill you’re letting atrophy. Be honest about which is which.
  • Always apply one critical edit to every AI output before it ships. Not to prove you’re involved — to stay in the judgment loop. If you can’t identify something worth changing or adding, the output isn’t being evaluated critically enough. Friction in evaluation is a feature, not a bug.
  • Name the judgment call in every AI-assisted decision. “AI surfaced this pattern; I decided to prioritize it because [reason].” This keeps ownership explicit and builds a decision log that’s useful when you need to explain your reasoning later.
  • Invest in the capabilities AI doesn’t replace. User empathy. Mental models. Deep domain knowledge. Relationship credibility. These compound with AI rather than competing with it. Wherever you’re currently underinvesting in these, that’s where your development time should go.
  • Set one “AI-free” decision this week. Pick a product or strategic question and work through it without prompting a model first. Notice what you know, what you don’t, and where your judgment is weakest. That map is where you need to grow.
  • Talk about AI augmentation explicitly with your team. Not as a policy — as a culture conversation. “We use AI to do [tasks] faster. We own the judgment on [decisions]. Here’s the line.” Teams that have this conversation explicitly are more resilient to the drift toward outsourcing that happens when nobody names the boundary.

Agency and ownership in AI contexts are closely related to how you build team culture around AI adoption. Why Agency Trumps Skills for AI-Driven Church Tech Teams covers the ownership dynamic from a team-building angle. And Why AI Decision-Making Needs Human Judgment: The Solomon Test explores where AI judgment breaks down and human judgment becomes non-negotiable.

Thinking through how to keep human judgment at the center of an AI-augmented product or ministry organization? I consult with product leaders and ministry innovators on AI strategy, team culture, and the practices that keep people — not tools — in the decision loop. Let’s talk.

Agency Isn’t Enough: Why AI-Era Product Teams Need Constraint Over Freedom

Last month, I sat in a church basement with a tech team lead drowning in AI tools. He had a dozen browser tabs open, each promising to automate something — sermon prep, small group scheduling, follow-up emails. His face was pure frustration. “I can do anything with these,” he said. “But I have no idea what I should do.”

I’ve heard variations of that sentence from product leaders at organizations far beyond church tech. The promise of AI is unbounded capability. The reality is that unbounded capability without constraint produces paralysis, not progress. Giving a team agency over a blank canvas doesn’t produce great work — it produces anxiety. The teams that are actually moving forward with AI aren’t the ones with the most freedom. They’re the ones with the clearest guardrails.

Why Too Much Freedom Is a Product Problem

There’s a well-documented psychological pattern here: when options multiply, decision quality degrades. Barry Schwartz called it the paradox of choice. The same dynamic that makes consumers freeze in front of 47 varieties of salad dressing also makes product teams stall when handed AI capabilities with no defined scope.

I’ve watched product teams spend entire quarters “exploring AI possibilities” — building demos, attending workshops, writing strategy documents — without shipping a single user-facing improvement. The exploration felt productive. The output wasn’t. And the root cause wasn’t lack of skill or motivation. It was the absence of a constraint that forced a decision.

Constraint isn’t the opposite of agency. It’s the container that makes agency useful. A product manager with agency and no constraint is a sculptor with clay and no deadline. A product manager with agency and clear constraint is a sculptor with clay, a deadline, and a specific brief. The second one ships.

Wesley’s Three Rules as a Product Constraint Framework

John Wesley’s Three Simple Rules — Do No Harm, Do Good, Stay in Love with God — were written as a framework for Methodist community life in the 18th century. I’ve found them to be a surprisingly clean model for AI product constraints, particularly because they’re sequenced by priority and genuinely hard to game.

“Do No Harm” maps directly to the first constraint every AI product decision should clear: are we breaking anything? Are we introducing complexity that frustrates users who didn’t ask for it? Are we shipping a feature that serves the roadmap more than it serves the user? The discipline of asking “what harm could this cause?” before shipping has stopped more bad ideas on my teams than any review process.

“Do Good” maps to the second constraint: is there measurable user benefit? Not “this is technically impressive” or “this will differentiate us in demos.” Actual user benefit, with a metric attached. I’ve watched teams get seduced by generative AI for content creation, spend months iterating on algorithms, and launch to crickets — because they never defined what “doing good” for the user actually looked like in this context.

“Stay in Love with God” — secularized, this becomes “stay grounded in mission.” AI hype can pull organizations away from the thing they exist to do. I’ve been in strategy meetings where the room buzzed about “transformative AI roadmaps” while the core problem — equipping the people the product is supposed to serve — got lost. Constraint means refusing to drift with every AI trend, and anchoring back to the specific outcome that justifies the product’s existence.

Constraint as a Design Principle, Not a Limitation

The reframe that helps teams most: constraint isn’t what you can’t do. It’s what you’ve decided to do first. A constraint says “this quarter, we’re only using AI to address user friction in the onboarding flow” — not because other AI applications don’t matter, but because focus produces better outcomes than breadth. You can revisit the other applications next quarter. You can’t revisit the quarter you spent spinning.

Teams that build constraint-first AI strategies share a few common habits: they define their AI investment in terms of specific user outcomes before selecting tools, they set explicit scope boundaries for each AI experiment, and they treat “we should also try X” as a backlog item rather than a current sprint addition. These habits sound bureaucratic until you see what the alternative produces — which is usually a lot of demos and very little improvement in the metrics that matter.


Your Turn: Apply This Today

If your team has been in “AI exploration mode” without clear traction, here’s how to apply constraint thinking and start moving:

  • Run every current AI initiative through the “Do No Harm” filter. For each active experiment, ask: who could this hurt, confuse, or frustrate — and are we certain that’s not happening? If you can’t answer that question with evidence, the experiment isn’t ready to ship.
  • Define “Do Good” in measurable terms before the next sprint. Write one sentence per AI initiative: “We’ll know this is working when [specific user behavior or metric] changes by [amount] in [timeframe].” Initiatives that can’t complete that sentence don’t belong in the current sprint.
  • Do a mission audit on your AI roadmap. List every AI initiative your team is working on. For each one, answer: how does this directly serve the specific outcome that justifies this product’s existence? Items that can’t answer that question are candidates for the backlog, not the sprint.
  • Set one explicit scope constraint this quarter. Not a theme — a boundary. “This quarter, AI investment is limited to reducing time-to-value in our onboarding flow.” Decide what’s out of scope and write it down. Teams that know what they’re not doing are faster than teams that can do anything.
  • Build a “next quarter” list. Every time someone says “we should also try [AI thing],” add it to a visible next-quarter list rather than the current backlog. This preserves the idea without letting it dilute the current focus. The list becomes a gift to your future self at quarterly planning.
  • Review your constraint at the end of each sprint. Ask: did this constraint help us make better decisions, or is it the wrong constraint? Constraints should be reviewed, not assumed. The goal isn’t rigidity — it’s focus. If the constraint is pointing your team at the wrong problem, change it.

Constraint thinking connects directly to how you define AI features in the first place. Why Your AI Feature Doesn’t Need More Data — It Needs Better Problem Definition covers how to scope AI work before you start building. And Three Rules for AI Prompt Design: John Wesley’s Constraint Framework applies the same Wesley framework specifically to prompt engineering discipline.

Working through how to focus your team’s AI investment and stop spinning on exploration without outcomes? I consult with product leaders on AI strategy, constraint frameworks, and the organizational habits that turn AI capability into shipped product impact. Let’s talk.

Support Data Isn’t Just Feedback: It’s Your AI Strategy’s Best Fuel

Two years ago, I was in a cramped church office helping a small tech team work through an inbox overflowing with user complaints about their digital giving platform. One email stopped me cold. A frustrated pastor had written: “I spend more time troubleshooting this than preparing my sermon.”

That line hit harder than any crash report or churn metric ever could. It wasn’t just a complaint — it was a window into exactly how badly the product was failing its core users. In that moment I saw support data for what it really is: not a burden, but the clearest possible signal about where to invest next.

Why Product Teams Treat Support Data as a Chore

Most product teams, especially in resource-constrained organizations, treat support tickets as a lagging indicator to skim at the end of a sprint — or a problem to outsource to the customer success team. I’ve been guilty of this myself. When you’re moving fast, the inbox feels like noise.

But treating support data as noise is one of the most expensive misreads a product team can make. It’s where your most motivated users — the ones who cared enough to write in — tell you exactly what isn’t working, in their own words, without anyone asking them to. You can’t buy that signal. You can only neglect it.

Charlie Munger’s inversion mental model is useful here. Instead of asking “how do we improve the product?” ask the inverse: “What would we have to do to guarantee users keep struggling?” The answer almost always lives in the support inbox. Users are already telling you. The question is whether you’re listening.

Support Data as AI Strategy Fuel

In the AI era, support data has become even more strategically valuable — and even more underused. Here’s why: AI tools are exceptionally good at pattern recognition across large text datasets. A human reviewing 500 support tickets per week might catch broad themes. An AI tool scanning those same tickets can surface micro-patterns — clusters of friction around a specific workflow step, a correlation between certain user profiles and specific failure modes, sentiment shifts that precede churn — that no human reviewer would reliably find.

I worked on a curriculum platform serving children’s ministry volunteers where the support inbox was consistently a clearer signal of product-market fit issues than any analytics dashboard we had. One pattern that kept recurring in tickets — volunteers struggling to find resources on spotty rural internet — didn’t appear anywhere in our usage data. It appeared in the inbox. When we ran that pattern through basic text clustering, we found it was the third most common friction category, affecting users we’d never targeted in our discovery work. That insight drove an offline-first redesign that meaningfully improved retention in our rural segment.

The pattern is consistent across every product I’ve worked on: the support inbox knows things your dashboard doesn’t. The teams that build systematic processes to extract and act on that signal have a durable information advantage over teams that don’t.

Building a Support Intelligence System

Turning support data into AI strategy fuel doesn’t require a sophisticated tech stack. It requires a systematic process and a commitment to treating support insights as a first-class input into product decisions — not an afterthought.

The minimum viable version looks like this: someone on the team reads and categorizes every support ticket weekly, tagging by friction type, user segment, and severity. Those categories get reviewed in sprint planning. High-frequency friction categories go on the roadmap; emerging friction patterns trigger user interview follow-up. That’s it. No AI tools required, though they accelerate the process significantly when you have volume.

The more sophisticated version adds AI-assisted sentiment analysis and clustering to your helpdesk workflow — tools like Zendesk’s AI features, Intercom’s categorization, or even a simple export into a text clustering tool. The goal is to reduce the time between “user experiences friction” and “product team knows about it and understands its scope.” The faster that loop closes, the more responsive and credible your product team becomes.


Your Turn: Apply This Today

Whether you have a mature support operation or a shared inbox that nobody owns, here’s how to start turning support data into strategic signal — this week:

  • Spend 30 minutes reading raw support tickets this week — not summaries. Read them as a product leader, not as a support manager. You’re looking for the underlying need behind the surface complaint. A ticket about a confusing UI is often a ticket about a missing mental model. What is the user actually trying to do?
  • Apply Munger’s inversion. Ask: “What would we have to do to guarantee users keep struggling with this?” Write down the three things your product would have to do to make the support inbox worse. Then check whether any of those things are currently true. The answers will be uncomfortable and useful.
  • Tag this week’s tickets by friction type. Even a rough taxonomy — navigation, missing feature, confusing workflow, performance, trust — will reveal patterns in one week that monthly summaries miss. Do it manually first; automate later.
  • Pull one raw user quote from the inbox and bring it to your next team meeting. Read it aloud. Ask the team: “What does this user actually need?” One real quote does more to align a team around user reality than three slides of aggregated survey data.
  • Add support data review to your sprint planning agenda. Even five minutes — “what did the inbox tell us this week?” — changes the dynamic. Decisions made with fresh user signal are better decisions, and the habit compounds over time.
  • Identify one support-informed fix you can ship in the next sprint. Not a big feature — a small, targeted improvement that directly addresses a recurring friction category. Show the team what happens when support data drives prioritization. The results build the habit faster than any process mandate will.

Support data and retention data are closely linked signals. What Retention Data Actually Tells You (And What It Doesn’t) covers how to build a complete picture of user health beyond churn numbers. And Munger’s Mental Models and AI Product Decisions goes deeper on applying inversion and other mental models to product strategy.

Trying to build better feedback loops between your users and your product team? I consult with product leaders on user research systems, data strategy, and the organizational practices that close the gap between what users experience and what gets built. Let’s talk.

Why Agency Trumps Skills for AI-Driven Church Tech Teams

Last month, I sat in a church basement with a tech volunteer who was struggling to set up a new AI tool for sermon transcription. His frustration wasn’t about the tool’s complexity — he knew the buttons to push. The problem was that nobody had told him why it mattered or given him any ownership over how it would be used. He muttered: “I just don’t get what this is for.”

That moment stuck with me. It wasn’t a skills gap. It was a purpose gap. And it’s the same pattern I’ve seen in product organizations far beyond church tech — teams that have been trained on the tools but not given genuine ownership over the outcomes those tools are supposed to produce.

Skills Without Agency Is Compliance, Not Capability

Here’s the thing about AI adoption in any team: skills are necessary but not sufficient. You can run workshops, build sandboxes, write documentation, and mandate usage — and still end up with a team that uses the tools only when required and abandons them the moment the pressure is off. What’s missing in those cases almost always isn’t knowledge. It’s agency.

Agency, in this context, means something specific: the sense of ownership over a problem, the authority to make decisions about how to address it, and the belief that those decisions will actually matter. Viktor Frankl’s work on meaning is useful here — he argued that humans don’t just need purpose handed to them; they need to discover it through engaged work. Teams that are handed AI tools without being given genuine problems to solve with those tools will never fully adopt them. They’ll comply, but they won’t own.

Ownership is what turns “I have to use this” into “I want to improve this.” And that distinction matters enormously for AI adoption velocity. Teams with high agency figure out use cases their managers never anticipated. Teams with low agency wait to be told what to do next.

Why Church Tech Teams Are a Clear Case Study

Faith-tech and church tech teams are a useful lens for this pattern because the stakes are highly visible and the volunteer culture makes authority structures explicit. When a volunteer doesn’t feel ownership, they leave. There’s no salary to keep them compliant. That makes the agency gap show up faster and more clearly than it does in corporate product organizations where people absorb the dysfunction because they’re getting paid to.

I’ve seen this play out repeatedly: a church tech leader implements an AI scheduling tool, trains the volunteers, and then watches usage drop to near zero within six weeks. Post-mortem almost always reveals the same thing — the volunteers were trained on the tool but not given any say in what problem it was solving or how success would be measured. They had skills. They had no stake.

The fix is deliberate but not complicated. When you introduce a new AI capability, don’t start with training. Start with a problem-ownership conversation: “Here’s the friction we’re experiencing. Here’s the outcome we’re trying to improve. I want you to figure out how this tool can address it — and I’m giving you the authority to make that call.” That framing changes everything about how the team engages with the technology.

Building Agency Into Your AI Adoption Process

The practical implication is that AI training needs to be sequenced differently than most organizations run it. The standard sequence is: introduce tool → train team → deploy. The agency-first sequence is: define problem → assign ownership → provide tool → let team define deployment. Those sequences produce dramatically different outcomes.

Small wins matter disproportionately in the early stages. When a volunteer or team member can point to a specific outcome they produced — “I set up the transcription workflow and it cut our follow-up time by an hour every week” — they develop a sense of stewardship over that outcome. They start to see themselves not as a user of the tool but as an owner of the process. That’s the transition that sustains AI adoption long after the initial training energy fades.


Your Turn: Apply This Today

Whether you lead a church tech team, a small product group, or a cross-functional product organization, here’s how to build agency into your AI adoption — starting this week:

  • Identify one AI tool your team is underusing. Don’t ask why they’re not using it — ask whether they feel ownership over the problem it’s supposed to solve. If the answer is no, that’s your starting point.
  • Pick one team member and assign genuine problem ownership. Not “use this tool” — “this friction point is yours to improve. This tool is available. Tell me in two weeks what you tried and what you learned.” The authority has to be real, not nominal.
  • Define success in outcome terms, not usage terms. “Reduce admin time by 30% for small group leaders” is a real goal. “Everyone uses the tool by Friday” is compliance theater. Outcome-focused goals give people something worth owning.
  • Celebrate a small win explicitly and tie it to purpose. When someone on the team produces a measurable improvement, name it publicly and connect it to the mission. “This saved the pastoral team four hours last week — that’s four hours they spent with people instead of on admin.” That framing builds the sense of stewardship that sustains adoption.
  • Block 15 minutes this week to audit your own behavior as a leader. Are you delegating ownership or just delegating tasks? Ownership comes with authority to make decisions. If you’re telling people what to do with the tool rather than telling them what problem to solve with it, you’re creating compliance — not capability.
  • Model agency yourself. Take one AI tool you’ve been handed or had recommended to you, identify a specific problem in your own workflow, and document what you tried and what you learned. Share that process with your team. Nothing builds a culture of agency faster than watching the leader practice it.

Agency in AI adoption connects directly to how you build and retain high-performing teams. The Stakeholder Management Mistake Most Product Leaders Make in Year One covers the related dynamic of alignment versus compliance in organizational settings. And AI Learning for Product Leaders: Stop Chasing Tools, Start Solving Problems addresses the same principle applied to learning investment.

Trying to build genuine AI capability — not just AI compliance — in your team or organization? I consult with product leaders and ministry innovators on building agency, aligning teams around mission-driven outcomes, and making AI adoption actually stick. Let’s talk.

How AI Prototyping Can Break Silos in Faith-Tech Products

A year ago, I was in a conference room with a faith-tech startup team watching a product manager and a designer talk past each other. The PM wanted data before committing to any direction. The designer wanted to build something the team could react to. Neither was wrong — but they were stuck, and the feature stalled for weeks because nobody said the obvious thing: “Let’s build a quick prototype and test it.”

That moment crystallized something I’d been seeing across product organizations: silos don’t break down through better meetings. They break down through shared artifacts. When a team can look at the same rough prototype together, the conversation shifts from “whose perspective is right” to “what does this teach us.” AI prototyping tools have made that shift dramatically cheaper and faster — which means the teams that aren’t using them are paying an avoidable coordination tax.

What AI Prototyping Actually Does for a Team

The most valuable thing an AI prototyping tool does is collapse the time between idea and artifact. What used to take a designer a week to wireframe can now take a few hours using tools like Galileo, Uizard, or even well-structured prompting in Figma’s AI features. That speed matters less for the prototype itself and more for what the prototype enables: a shared object the whole team — product, design, engineering, stakeholders — can react to together.

Teresa Torres’ continuous discovery framework makes the argument clearly: prototypes aren’t about showing off ideas. They’re about learning. The faster you can get something in front of a real user — even a rough, unpolished version — the faster you discover whether your assumptions are right. AI prototyping compresses that loop significantly.

In faith-tech contexts specifically, where teams are often small and stakeholders can include everyone from engineering leads to ministry directors, this matters. A prototype that a non-technical stakeholder can react to is worth more than a PRD that only product and engineering understand. It gets everyone speaking the same language before anyone writes a line of code.

The Silo Problem AI Prototyping Actually Solves

Product silos usually persist for one of two reasons: different teams operate from different information, or different teams have different incentives. AI prototyping directly addresses the first. When the same working mockup is in front of a PM, a designer, an engineer, and a ministry stakeholder simultaneously, you rapidly surface where each person’s mental model diverges. That divergence, made visible, is far cheaper to resolve before development than after.

I’ve watched this play out in prayer app development, church management software, and discipleship platforms. In each case, the friction wasn’t disagreement about the product’s mission — it was that nobody had made their assumptions concrete enough for others to react to. A rough AI-generated wireframe of a specific user flow did in one session what three roadmap reviews had failed to do: it gave everyone the same starting point.

The discipline that matters here: keep the prototype simple and provisional. The temptation with good prototyping tools is to overinvest in polish before you’ve learned anything. A polished prototype signals “we’ve decided” when the point is “we’re still figuring this out.” Users interact differently with rough prototypes — they’re more honest, and more willing to suggest changes when the artifact looks unfinished.

Building a Prototyping Habit, Not Just a Prototyping Moment

The real leverage isn’t a single prototyping sprint — it’s making rapid, collaborative prototyping a standard step in every major feature decision. Teams that prototype as a habit stop having the same alignment conversations over and over, because they’ve built a shared practice of making assumptions concrete before debating them.

For faith-tech teams in particular, I’d suggest the following sequencing: before any feature goes on the roadmap, someone on the team produces a one-screen wireframe answering the question “what does success look like for the user here?” That wireframe — rough, provisional, built in an hour — becomes the discussion artifact for the roadmap conversation. It doesn’t answer all the questions. It just makes sure everyone is answering the same questions.


Your Turn: Apply This Today

Whether your team has a formal prototyping practice or none at all, here’s how to start building the habit — starting this week:

  • Identify the next feature your team will debate in roadmap planning. Before the meeting, have one person spend 60 minutes building a rough wireframe using an AI prototyping tool. Bring the wireframe to the meeting instead of a written spec. Watch how the conversation changes.
  • Run a 90-minute silo-breaker session. Put product, design, and at least one engineering rep in a room. Give them a specific user problem, a prototyping tool, and a timer. The output should be a rough, functional prototype — not a polished one. Speed over polish is the rule.
  • Get it in front of three to five real users within 48 hours. Pastors, volunteers, small group leaders — whoever your actual end user is. Your goal is one concrete thing you learned that you didn’t expect. If you can’t identify one, the prototype wasn’t specific enough.
  • Run a 30-minute team debrief after user feedback. What did you learn? What did you assume that turned out to be wrong? What’s the next iteration? Write down three specific things the feedback changed about your direction.
  • Make it a gate, not an option. Decide as a team that any feature above a certain size threshold requires a rough prototype before it can be added to the roadmap. This isn’t about slowing down — it’s about ensuring you’re making decisions from the same picture.
  • Celebrate what the prototype taught you, not just what you shipped. Prototypes that prevent bad features from being built are wins, even if nothing ships. Make sure your team understands that learning is the output — not just working software.

Team alignment and discovery go hand in hand. Opportunity Solution Trees Meet AI Reality at Scale covers how to adapt discovery frameworks when you’re building AI-powered features, and Why Teresa Torres’ Continuous Discovery Cadence Is Breaking Down in the AI Age addresses the structural changes that AI creates for discovery-driven teams.

Struggling with cross-functional alignment or trying to build a more disciplined discovery practice in your faith-tech or product organization? I consult with product leaders on prototyping cadences, team alignment, and the practices that close the gap between roadmap and user reality. Let’s talk.

AI Learning for Product Leaders: Stop Chasing Tools, Start Solving Problems

Six months ago, I was coaching a product leader at a faith-tech nonprofit who was drowning in the flood of AI tools hitting his inbox. He’d spent hours watching tutorials, clicking through demos, and trying to “keep up” with every new feature drop. His voice cracked over Zoom as he admitted: “I’m learning a lot, but I’m not getting anywhere.”

He wasn’t lazy. He wasn’t slow. He was caught in the most common trap in AI learning — mistaking tool fluency for strategic capability. He knew how to use the tools. He had no idea which problem to point them at.

Learning AI Without a Purpose Kills Momentum

This pattern shows up everywhere in product organizations right now. Teams attend AI workshops, subscribe to newsletters, build internal sandboxes — and then struggle to translate any of it into shipped product improvements. The bottleneck isn’t knowledge. It’s the absence of a clear problem that the knowledge is meant to serve.

There’s a meaningful distinction between AI literacy and AI effectiveness. Literacy is knowing what the tools can do. Effectiveness is knowing which problem in your specific product context is worth solving with them, and having the judgment to evaluate whether the tool actually helps. Most AI learning programs develop the first. Almost none develop the second.

The result is product leaders who can speak fluently about LLMs, embeddings, and RAG pipelines but can’t answer the question that actually matters for their team: “What is the most important thing AI can do for our users right now?” That gap is expensive. Not just in time and training budget — in team credibility and strategic direction.

The Mental Model That Changes How You Learn

Charlie Munger’s latticework mental model is useful here. His argument was that knowledge organized around real problems is exponentially more useful than knowledge accumulated for its own sake. Every new concept anchors to an existing one, and the whole structure becomes more useful with each addition. Knowledge without a problem to anchor to just drifts.

Applied to AI learning, the implication is direct: start with the problem, not the tool. Before you invest time in learning any AI capability, write down the specific friction point in your product that you’d like it to address. Not “AI could make our onboarding better” — that’s too vague to act on. Something like: “Our support team spends 40% of their time answering questions that could be answered by a well-designed in-product FAQ. Could an AI-assisted feature reduce that?” Now you have a problem worth learning toward.

This reframe does something important: it makes your AI learning time accountable. If you’re three hours into a tutorial and you can’t articulate how it connects to your stated problem, that’s a signal — either the tutorial is the wrong one, or your problem definition isn’t sharp enough yet.

Clear Thinking Beats Endless Learning

The product leaders I’ve seen develop genuine AI effectiveness share a habit: they apply opportunity cost thinking ruthlessly to their learning time. Every hour spent on a tutorial that doesn’t connect to a current product problem is an hour not spent on user interviews, roadmap clarity, or team alignment. The cost is real even when it’s invisible.

This doesn’t mean stopping learning. It means changing the sequence. Identify the problem first. Map the assumptions behind the problem. Then go find the specific AI capability that addresses it — and learn only that, deeply, before moving on. You’ll learn slower in volume and faster in applicability. And applicability is what actually ships.

One practical reframe I use with teams: replace “keep up with AI” as a goal with “improve one specific outcome in the next 30 days using AI.” The first goal is a treadmill. The second is a target. Targets produce movement; treadmills produce fatigue.


Your Turn: Apply This Today

Before you open the next AI newsletter or sit through the next demo, run through this checklist to make sure your AI learning is anchored to something real:

  • Write down your current product’s top three friction points. Not AI opportunities — actual user or team pain points. Any AI learning you do this month should connect directly to one of these three items.
  • Cancel or pause one AI learning subscription that you can’t connect to a current problem. The mental overhead of processing irrelevant information is a real cost. Reduce the noise before adding more signal.
  • If you’re learning an AI tool without a clear problem to solve, pause and redirect that time to user interviews. Thirty minutes of user conversation will point your AI investment better than most tutorials will.
  • Set a 30-minute timer this week to map a mental model for your product’s biggest gap. Start with the end goal (e.g., higher retention, faster activation) and work backward to the root cause. That root cause is where AI learning should be aimed.
  • Test one AI tool or feature against a specific friction point in the next two weeks. Don’t aim for perfection — aim for a measurable result: “Reduce support ticket volume by 10%,” “Cut onboarding time by two steps.” Specificity is what makes the test informative.
  • Share this framing with your team in your next meeting. Walk them through prioritizing mission-relevant learning over tech hype using one real user story. Teams that learn together toward shared problems develop AI effectiveness faster than teams where everyone is self-directing their own curriculum.

If you’re working through where AI fits in your product strategy, Why Your AI Feature Doesn’t Need More Data — It Needs Better Problem Definition addresses how to scope AI features correctly before you write a line of code. And Munger’s Mental Models and AI Product Decisions goes deeper on the latticework framework applied to product leadership.

Working through how to make your team’s AI investment actually move the needle on outcomes that matter? I consult with product leaders on AI strategy, learning frameworks, and the organizational dynamics that determine whether AI capability translates into product impact. Let’s talk.

Why Your AI Feature Doesn’t Need More Data — It Needs Better Problem Definition

The request comes in through Slack, dressed up as a technical conversation: “We need more training data before we can improve the model.” Sometimes it’s true. More often, it’s a symptom of something upstream that no amount of data will fix.

I’ve sat in product reviews where teams requested months of data collection before they could ship a meaningful AI feature — only to discover, after collecting the data, that they’d been optimizing for the wrong output the whole time. The model got better at predicting something. It just wasn’t predicting the right thing.

The AI Feature Problem That Isn’t a Data Problem

Most AI feature failures I’ve seen in product organizations aren’t engineering failures. They’re problem definition failures. The team knew what they wanted to build. They knew which model architecture they’d use. They had a rough sense of the training data they’d need. What they hadn’t done was write a crisp one-paragraph answer to a deceptively simple question: What specific decision are we trying to improve, for whom, and how will we know the improvement is real?

That question sounds basic. In practice, it’s brutally hard to answer precisely — and most AI feature discussions move past it without ever settling it. Instead, teams anchor on the technical implementation (which model, which API, which pipeline) and skip the product fundamentals that determine whether any of it will matter.

The result is a common and expensive failure mode: you ship an AI feature, it’s technically impressive, and adoption is flat. Users aren’t hostile to it — they just don’t reach for it. Because the feature is solving a problem that wasn’t sharp enough to justify its existence in the first place.

Problem Definition Before Model Selection

There’s a sequence that high-performing AI product teams follow — often implicitly — that lower-performing teams skip. It goes: problem definition → success metric → data requirements → model selection. Most teams invert this. They start with a model or an API capability, then work backward to find a use case. The resulting features are technically coherent but strategically thin.

Problem definition, done right, answers four things before any engineering begins:

Who has the problem? Not “our users” — a specific persona, role, or workflow. The tighter your answer, the more precise your feature can be. Vague personas produce vague AI features that try to be useful to everyone and end up indispensable to no one.

What decision are they currently making badly? AI is most useful when it augments or automates a decision that a human is already making but making slowly, inconsistently, or with incomplete information. If you can’t identify the decision, you can’t define what “better” looks like.

What does “better” actually mean? Faster? More accurate? More consistent? Less effortful? These aren’t the same, and they don’t optimize for the same outputs. A feature that reduces decision time by 40% but introduces 15% more errors may be worse than no feature at all — depending on the stakes of the decision.

How will you measure the improvement? Before you define your data requirements, you need your success metric. If you can’t articulate a measurable outcome that would tell you the feature worked, you’re not ready to scope the data pipeline.

Why Teams Skip This Step

Skipping problem definition isn’t laziness. It’s usually organizational pressure. There’s a demo to prep for. There’s a competitor who just shipped something. There’s an executive who read about the technology and wants to see it in the product. In those conditions, “we need to define the problem more precisely” sounds like a stall tactic. It isn’t — but it can feel like one.

The reframe I use with teams: problem definition doesn’t slow down AI feature development. It accelerates it. When the problem is crisp, data requirements become obvious. Model selection becomes obvious. Evaluation criteria become obvious. The team spends less time rebuilding the pipeline because they didn’t spec it wrong the first time.

I’ve watched teams spend 12 weeks collecting data for a feature that took 4 weeks to build — only to realize at launch that they’d defined the output variable incorrectly. A week of problem definition work at the front would have surfaced that misalignment long before the data collection started.

The Practical Test: Can You Write the Prompt?

Here’s a heuristic I’ve started using with teams evaluating AI features: can you write the prompt that would produce the output you want, and does that output solve the problem as you’ve defined it? If you’re using a generative model, this is literal — write the prompt. If you’re using a predictive model, translate: can you describe in plain English what you’re asking the model to predict, and does that prediction map to a real decision that real users need help making?

If you can’t write the prompt — or if the prompt output doesn’t map clearly to a user decision — you don’t have a data problem. You have a problem definition problem. And that’s a conversation to have in a product review, not a sprint planning session.


Your Turn: Apply This Today

Before your team writes another line of code or collects another batch of training data, run your AI feature through this problem definition checklist:

  • Write the one-paragraph problem statement. Include: who has the problem, what decision they’re currently making badly, what “better” means in measurable terms, and how you’ll know the feature worked. If your team can’t agree on this paragraph in one working session, that’s your first sprint — not the model.
  • Identify the specific decision your AI feature augments or automates. If you can’t name a decision, you don’t have an AI problem — you may have a search problem, a dashboard problem, or an information architecture problem. AI is not the right hammer for every nail.
  • Define your success metric before scoping data requirements. What measurable outcome will tell you the feature improved the user’s decision? Write it down. Lock it. Then — and only then — work backward to determine what data you need to produce that outcome reliably.
  • Run the “prompt test.” Can you write the prompt (or describe the prediction) in plain language? Does that output directly address the decision you identified? If there’s a gap between the model output and the user’s decision, that gap is your product design problem.
  • Interview three users about the decision before building. Ask them to walk you through the last time they made this decision. What information did they use? Where did they get stuck? What would have made it faster or easier? Their workflow, not your intuition, should drive your data spec.
  • Set a problem definition review gate. Before any AI feature moves from ideation to data collection, require a written problem definition document. One page. Four questions answered. This gate will save you more engineering cycles than any model optimization will.

Related reading: The Product Leader’s AI Infrastructure Blind Spot covers the organizational patterns that cause AI investments to underperform, and Opportunity Solution Trees Meet AI Reality at Scale addresses how to adapt discovery frameworks when AI is the solution space you’re exploring.

Working through AI product strategy and finding that your team keeps building features that don’t land? I consult with product leaders on AI feature scoping, problem definition frameworks, and the organizational dynamics that cause AI investments to underdeliver. Let’s talk.

Your AI Product Has a Trust Problem Before It Has a Performance Problem

Every AI product team I’ve worked with or advised has spent significant time on the performance problem: how accurate is the model, how fast does it respond, how well does it handle edge cases? These are legitimate engineering concerns. But I’ve watched technically impressive AI products fail in the market for a reason that none of those metrics capture: users didn’t trust them.

Trust is the prerequisite for adoption. And trust in AI products breaks in ways that are categorically different from trust in traditional software — faster, harder to repair, and with a long tail of behavioral effects that don’t show up in your engagement dashboard.

How AI Trust Breaks Differently

When traditional software fails, users understand the failure. A button doesn’t work. A page doesn’t load. The app crashes. These failures are frustrating, but they’re legible — users know what happened and can calibrate their expectations accordingly.

When AI fails, users often don’t know what happened. The AI gave a confident-sounding answer that turned out to be wrong. The recommendation seemed personalized but was clearly off. The summary missed something important and presented the gap with the same authority as the accurate content. Users can’t distinguish between the AI being right and the AI sounding right. That illegibility is what makes AI trust failures so destructive.

A single high-visibility AI error can undo months of accurate performance. Users who’ve been trusting the AI unconsciously — accepting its outputs without verification, integrating them into their workflows — suddenly question every prior output. The trust collapse is retroactive. And it’s much harder to rebuild trust after an AI failure than after a traditional software failure, because the user’s underlying question — “how do I know when to trust this?” — doesn’t have a clean answer.

The Trust Gap Product Teams Miss

Most product teams measure trust indirectly — through engagement, retention, NPS. These metrics capture trust effects, but they lag the trust event by weeks or months. By the time your NPS drops, the trust damage was done several product cycles ago.

The trust gap I see most consistently is between what users expect the AI to know and what it actually knows. Users quickly form a mental model of what the AI understands about them — their context, their history, their preferences. When the AI behaves in a way that violates that mental model, trust breaks. Not because the AI was necessarily wrong, but because it revealed that its understanding of the user is shallower than the user assumed.

This is a design problem as much as an engineering problem. The AI’s communication of its own uncertainty — what it knows, what it’s inferring, what it’s guessing — is as important to trust as the accuracy of the output itself. Products that communicate their AI’s confidence level clearly allow users to calibrate appropriately. Products that present all AI outputs with equal confidence train users to either overtrust everything or distrust everything. Neither is the outcome you want.

Building Trust Before You Build Performance

The framing shift that changes how you build AI products: trust is a design requirement, not a performance outcome. You don’t earn trust by making the AI more accurate and hope trust follows. You design for trust explicitly — and then let performance maintain it.

What designing for trust looks like in practice:

Transparency about what the AI can and can’t do. Set user expectations at the point of first contact, not after the first failure. Users who understand the AI’s capabilities and limitations before they rely on it calibrate their trust appropriately. Users who discover limitations through failure recalibrate toward distrust.

Visible confidence signals in the UI. When the AI is highly confident, say so — implicitly through clean, direct output. When the AI is uncertain, signal that too — through hedged language, source attribution, or explicit “I’m not sure about this” framing. Users who can see confidence levels can trust appropriately rather than uniformly.

Recoverable failures. Design for the moment when the AI is wrong. How does the user know? How do they correct it? What does recovery cost them? AI products where errors are hard to detect and expensive to fix train users to verify everything, which eliminates the productivity benefit. AI products with visible, cheap error recovery build the confidence for users to rely on the AI without constant verification.

Consistent, predictable behavior. Trust requires predictability. An AI that behaves differently on similar inputs — even if both outputs are acceptable — trains users toward unpredictability anxiety. They start over-supervising the AI because they don’t know when to trust it. Consistency in behavior, tone, and output style is trust infrastructure. Google’s People + AI Research guidebook on AI interaction design is the best public resource I’ve found on trust-centered AI design principles.

The Trust Measurement Problem

If you’re not measuring trust directly, you’re flying blind on your most important AI product metric. Engagement and retention are trust effects. They tell you that trust is breaking down after it’s already happened. You want leading indicators.

The leading indicators I track: AI override rate (how often users edit or reject AI outputs, separated from how often they accept without modification), re-query rate (how often users immediately follow an AI output with a corrective or clarifying query), and qualitative signals from support tickets about AI confusion or error. These don’t replace engagement metrics — they explain them, and they surface trust issues early enough to do something about them before they show up in churn.


Your Turn: Apply This Today

Build trust into your AI product development process before the first failure forces you to repair it:

  • Audit your AI product for confidence signaling. Walk through your product as a new user. Can you tell when the AI is highly confident vs. uncertain? If all outputs look the same, you’re training users to either overtrust or uniformly distrust. Add confidence signals before the next major release.
  • Map your AI failure modes and their recovery cost. For each AI feature, ask: when this is wrong, how does the user know? How expensive is the correction? The higher the recovery cost, the more supervision users will apply — and the lower the actual productivity benefit. Design cheap recovery paths into every high-consequence AI feature.
  • Add AI override rate to your product metrics dashboard. Instrument how often users accept AI outputs without modification vs. edit, override, or immediately re-query. Track it weekly. If the override rate is near zero, users may be overtrusting. If it’s very high, users don’t trust the AI enough to rely on it. Calibrated override rates (somewhere in between) indicate healthy trust.
  • Conduct a “first failure” user research session. Recruit users who have experienced a visible AI error in your product. Interview them about what happened to their trust and usage behavior after the failure. The pattern in their responses will tell you whether your product’s trust recovery design is working or broken.
  • Write your AI product’s “trust contract” with users. One paragraph: what your AI knows, what it can do reliably, what it can’t do reliably, and how to tell the difference. Share it in your onboarding. Users who understand the trust contract calibrate appropriately. Users left to discover it through failure don’t.
  • Run a “trust stress test” before every major AI feature launch. Deliberately trigger AI failures in a testing session with representative users. Observe their reactions. How long does it take them to recover trust? Do they change their behavior after the failure? If the trust damage is severe or persistent, redesign the failure experience before launch.

The trust problem is closely connected to the cognitive load dimension — Kahneman’s System 1 automation paradox explains why users overtrust confident-sounding AI outputs, and the Solomon test for AI decision-making addresses when human judgment needs to stay in the loop regardless of AI confidence.

Building an AI product and noticing that trust is the bottleneck more than performance? I consult with product teams on AI trust design, confidence signaling, and building the user research processes that surface trust problems before they become retention problems. Let’s talk.

Why AI Decision-Making Needs Human Judgment: The Solomon Test for Product Leaders

The story of Solomon’s judgment is one of the oldest decision-making frameworks in recorded history. Two women claim the same child. No witnesses, no evidence, no algorithmic output that could resolve the dispute. Solomon’s solution — propose dividing the child, then watch who objects — is a masterclass in what optimization cannot do: use the decision itself as the test that reveals the truth.

I’ve been thinking about this story a lot while working on AI recommendation systems. We’ve gotten very good at optimization. We can recommend content, predict churn, personalize experiences, and rank options with impressive accuracy. What we haven’t gotten good at — and what I’m not sure AI will ever be good at — is recognizing when the optimization is solving the wrong problem entirely.

The Algorithm vs. the Test

A recommendation algorithm optimizes for a signal. Click-through rate, completion rate, return visits, explicit ratings — whatever you tell it to optimize, it will optimize. The problem is that the signals we can measure are often proxies for the outcomes we actually care about, and at some point the proxy diverges from the outcome.

I’ve seen this play out in digital content platforms repeatedly. A reading plan recommendation algorithm optimizes for completion rates. It gets good at predicting which plans users will finish. It starts recommending plans that feel familiar, comfortable, and achievable — because those are the ones that get completed. Completion rates go up. The metric is green. But users who choose plans that “mismatch” their stated preferences — the challenging, unfamiliar ones — show better long-term engagement. The algorithm optimized for completion and selected against growth.

Solomon’s test is the thing the algorithm can’t run. He couldn’t optimize his way to the truth. He had to create a condition that would reveal the truth through the parties’ responses. That kind of judgment — knowing that the right test will surface the right answer — is what AI cannot replicate, and what product leaders need to understand as an irreplaceable capability.

The Limits of AI Optimization

There’s a specific category of decisions where optimization fails systematically, and it maps precisely to the conditions of Solomon’s judgment: when the right answer requires understanding something about a party’s underlying motivation or stake in the outcome that isn’t visible in any behavioral signal.

In product terms: the algorithm can tell you what users do. It cannot tell you what users are trying to become. It can tell you what content users complete. It cannot tell you whether completion served their actual goal. It can tell you what users click. It cannot tell you whether the click reflects genuine engagement or habit, genuine interest or boredom-driven curiosity.

These are not problems of insufficient data. They’re problems of category — you cannot optimize your way to answers about meaning, motivation, and stake without reducing those human realities to signals that don’t actually capture them. This is what HBR’s research on AI decision limits has consistently found: AI excels at decisions that can be fully specified by their optimization criteria. It fails at decisions where the criteria themselves are contested or where the right answer depends on understanding what’s at stake for the parties involved.

Creating Space for Judgment

The practical implication for product teams isn’t “stop using optimization.” It’s “design your system so that optimization handles what optimization is good at, and humans handle what optimization can’t reach.”

This requires being explicit about the decisions your AI system is making and auditing them for the category of problem they represent. Can the decision be fully specified by a measurable outcome? Optimization is appropriate. Does the decision depend on understanding a user’s underlying motivation, growth trajectory, or stake in the outcome? Optimization needs a human checkpoint.

In practice, this means building escalation paths. Decisions that an algorithm can handle with confidence: automated. Decisions that carry high downstream consequence and depend on judgment the algorithm can’t access: escalated to a human. The challenge is designing the system so that the escalation happens before the consequence rather than after it.

The Wisdom Gap in AI Product Teams

Most product teams have become very good at evaluating AI performance — accuracy rates, precision, recall, business metric lift. What they haven’t systematically developed is the capacity to evaluate AI judgment — the ability to recognize when the system is optimizing confidently toward the wrong objective.

This is the wisdom gap. Wisdom, in the traditional sense, isn’t just knowledge or capability — it’s knowing when to apply which capability, and knowing when the right answer can’t come from a formula at all. Solomon wasn’t impressive because he had access to more information than anyone else. He was impressive because he understood that the right test would reveal what no available information could.

Building that capacity into a product team means creating explicit processes for questioning whether AI systems are optimizing for the right things — not just whether they’re optimizing well. It means reviewing decisions at the system level, not just the feature level. It means asking, regularly: what would the right answer look like if we couldn’t use a metric to find it? If the answer changes when you remove the metric, the metric is the wrong one.

The Practical Applications

Three specific places where Solomon’s framework changes how I evaluate AI product decisions:

Recommendation systems: Don’t just optimize for completion or engagement. Build qualitative research into your cadence that asks users whether the AI-recommended path served their actual goal. The behavioral signal and the goal-based signal often diverge. The divergence is where the product insight lives.

Personalization: The most effective personalization isn’t always the most comfortable personalization. Before concluding that an algorithm is working because engagement is up, ask whether the algorithm is serving users’ stated goals or optimizing toward the path of least resistance. Users often engage more with what’s familiar than with what’s growth-oriented. The engagement metric can’t distinguish between the two.

High-consequence decisions: Any AI-assisted decision with significant downstream consequences for a specific user — content restrictions, account actions, eligibility determinations — needs a human checkpoint. Not because the AI will always be wrong, but because the cost of confident wrongness in these cases is high enough that the judgment of a human who understands what’s at stake is worth the operational overhead.


Your Turn: Apply This Today

Build the judgment layer into your AI product process:

  • Audit your AI system’s decisions for decision category. List the top 10 decisions your AI system makes. For each one, ask: can this decision be fully specified by its optimization criteria? If yes, automation is appropriate. If the right answer depends on understanding user motivation or stake, build a human checkpoint.
  • Separate your “proxy metrics” from your “outcome metrics.” Identify which metrics your AI optimizes directly and which outcomes you actually care about. Map the relationship between them. Where the proxy and the outcome diverge, you have a place where optimization is working against you.
  • Build qualitative research into your AI evaluation cadence. Once per quarter, interview users whose behavior your AI system has most significantly influenced. Ask whether the AI-driven experience served their actual goal. The divergence between behavioral signals and goal-based answers is where your most important product insights live.
  • Design a “Solomon test” for your highest-stakes AI decisions. For the AI-assisted decisions with the highest downstream consequence, design a test that would reveal whether the AI’s recommendation was right — not just whether users accepted it. Acceptance and correctness are not the same.
  • Create an escalation path for judgment-dependent decisions. Identify the category of decisions that require understanding user motivation or stake. Build an explicit escalation path so that those decisions reach a human before the consequence rather than after. Make it part of your system design, not an emergency procedure.
  • Hold a “what if we removed the metric?” review annually. For each of your AI system’s optimization targets, ask: what would the right answer look like if we couldn’t use this metric? If the answer changes significantly, the metric needs to be reconsidered. This is the most uncomfortable product review you will have and the most valuable.

The judgment-vs-optimization tension shows up across AI product decisions — the Kahneman System 1 paradox addresses the cognitive load dimension of the same problem, and why product decisions are never just product decisions explores the ethical dimension of designing AI systems that optimize toward the right ends.

Building AI systems and trying to design judgment into the process rather than optimizing it away? I consult with product teams on AI product strategy, human-AI decision frameworks, and building systems that know when optimization is the right tool — and when it isn’t. Let’s talk.

Three Rules for AI Prompt Design: What John Wesley’s Constraint Framework Teaches Product Teams

John Wesley reduced the entire Methodist movement to three rules: do no harm, do good, and stay in love with God. Not a comprehensive theology. Not thirty-seven principles. Three simple constraints that anyone in a Methodist society could remember, apply, and teach.

I’ve been thinking about Wesley’s framework — specifically his instinct that constraint creates more impact than capability — as product teams struggle with AI implementation. Claude can write, code, analyze, reason, and create. GPT-4 handles everything from customer support to strategic planning. The temptation is to throw AI at every problem and see what sticks.

Wesley understood something we keep relearning: the most powerful systems aren’t the ones that can do everything. They’re the ones that do specific things exceptionally well within clear boundaries. Here’s what his three-rule framework looks like as an AI product design principle.

Do No Harm: The First Constraint in AI Prompt Design

In AI product work, “do no harm” translates to a specific discipline: define what the AI system should not do before you define what it should do.

Most teams design AI features by expanding capability — adding more things the AI can handle, more use cases it supports, more outputs it can generate. The Wesley approach inverts this: start by naming the exclusions. What should this AI never say? What user requests should it decline to fulfill? What outputs would be harmful, misleading, or counterproductive even if technically achievable?

The constraint changes the architecture of the feature. Teams that start with capability tend to build AI that does many things adequately. Teams that start with harm prevention tend to build AI that does specific things with genuine care for the user’s actual outcome. The difference in user trust is significant and compounds over time.

In practice: before your next AI feature ships, write down three things it should never do. Make those constraints explicit in your system prompt, your evaluation criteria, and your launch checklist. The discipline of naming the harms before the capabilities will change what you ship.

Do Good: The Second Constraint

“Do good” seems obvious — of course the goal is to do good. But Wesley’s point wasn’t that doing good is obvious. His point was that you have to choose it actively and specifically. You have to name the good you’re trying to do, not assume it follows from capability.

For AI product teams, this translates to specificity of purpose. The most dangerous AI features are the ones built with vague positive intent — “help users be more productive,” “improve the experience,” “make things easier.” These goals aren’t wrong; they’re insufficiently specific. They don’t constrain what the AI does well enough to actually do good reliably.

The teams building the most impactful AI features are the ones who can answer precisely: for which specific user, in which specific situation, doing which specific task, does this AI feature create genuine value? The narrower the answer, the more intelligence you can focus on serving that specific case. AI amplifies specificity. A vague AI feature serves everyone adequately. A specific AI feature genuinely transforms a particular user’s experience.

The prompt design corollary: every AI system prompt should contain an explicit statement of the specific good the AI is designed to do. Not “help users.” Something precise enough that you could evaluate whether a given output advances it or not.

Stay in Love With the Mission: The Third Constraint

Wesley’s third rule is the hardest to translate into product terms, but the translation is important. “Stay in love with God” was Wesley’s constraint against letting the institution — the methodology, the system, the practice — become the end rather than the means. The movement existed to serve spiritual formation, not itself. When the structure started serving itself, Wesley’s third rule was the corrective.

For AI product work, this is the constraint against losing the mission in the mechanics. It’s easy to become so focused on what the AI can do — what’s technically impressive, what gets stakeholder attention, what generates engagement metrics — that you lose track of whether the AI is actually serving the user’s underlying goal.

I’ve watched this happen. A team builds an AI writing assistant that users engage with heavily — and then realize the engagement is driven by the novelty of watching the AI write, not by the quality of what users are actually producing. The metric is green. The mission has drifted. The AI became the point rather than the tool.

“Stay in love with the mission” is a standing question for every AI feature review: is this feature advancing the user’s actual goal, or is it becoming the goal itself? The discipline of asking it consistently is what keeps AI features from becoming impressive demonstrations that quietly fail the user. More on this from a product ethics perspective: the Interaction Design Foundation’s framework on design patterns is a useful lens for evaluating where AI features cross from tool to substitution.

Why Constraint Beats Capability in AI Product Design

The Methodist movement succeeded not because Wesley had the most sophisticated theology, but because he had the most practical framework. Three rules anyone could remember, apply, and teach — that’s what scaled the movement across 18th-century Britain.

The teams building the most impactful AI products right now aren’t the ones deploying every capability. They’re the ones who’ve chosen clear constraints that channel AI capability toward specific, valuable outcomes. Their prompts are shorter and more focused. Their AI features do fewer things with more precision. Their users trust the outputs because the product was clearly designed with the user’s outcome in mind rather than the AI’s capability ceiling.

In a world where AI can do almost anything, the crucial question isn’t what’s possible. It’s what’s wise. Wesley figured out how to answer that question with three rules. Your AI prompts need the same discipline.


Your Turn: Apply This Today

Apply the three-constraint framework to your current AI features:

  • Write the “do no harm” list for your highest-traffic AI feature. Name three things this AI should never output, regardless of what the user asks for. Write them as explicit constraints in your system prompt. If you’ve never done this exercise, do it before your next feature review — it will surface assumptions nobody has articulated.
  • Rewrite your AI feature’s purpose statement as a specific user outcome. Replace “help users be more productive” with a sentence specific enough to evaluate: “Help [specific user type] accomplish [specific task] in [specific context] faster and with better results.” If you can’t write the specific version, the feature isn’t ready to ship.
  • Add a “mission drift” check to your sprint review. Ask once per sprint: are users engaging with this AI feature because it’s helping them accomplish their goal, or because it’s impressive to interact with? The engagement metric doesn’t distinguish between these. Your qualitative research should.
  • Test the three-rule prompt structure on your next system prompt. Write a system prompt that has three sections: what the AI must never do, what specific good the AI is designed to do, and what the user’s underlying goal is that the AI should always serve. Compare the outputs to your current prompt. The discipline almost always improves them.
  • Audit your AI feature’s constraint density. Count the constraints in your current system prompt vs. the capabilities you’ve described. If capabilities dramatically outnumber constraints, you’ve built an AI feature by expansion rather than by design. Add at least one constraint for every three capabilities you’ve described.
  • Apply the “Wesley test” before your next AI feature launch. Ask: does this feature do no harm? Does it advance a specific, named good for a specific user? Does it serve the user’s mission rather than substituting for it? If you can’t answer yes to all three, the feature needs more design work before it ships.

The constraint-beats-capability principle shows up in other forms too — the choice overload paradox in AI features is the consumer-facing version of the same dynamic, and Munger’s inversion principle is a complementary framework for designing around constraints rather than toward capabilities.

Building AI features and trying to avoid the “impressive but useless” trap? I consult with product teams on AI prompt architecture, constraint-based feature design, and building AI that serves users’ actual goals rather than demonstrating capability. Let’s talk.

Designing for Obsolescence: What Seneca’s Philosophy Reveals About AI System Architecture

Seneca wrote that every new beginning comes from some other beginning’s end. He was writing about time — how we live in a continuous flow of transitions, each new phase emerging from the completion of the last. Product leaders building AI systems are working through exactly this dynamic, whether they’ve named it or not.

Every AI model update creates a new beginning by ending the previous version’s assumptions. Every infrastructure upgrade rewrites the operational playbook. Every significant capability improvement from a frontier provider makes yesterday’s integration architecture obsolete. The question for product leaders isn’t whether your AI systems will become obsolete — it’s whether you’ve designed for that obsolescence or against it.

The Paradox of AI Infrastructure

Here’s the tension: we want the reliability of traditional software — predictable, testable, maintainable — but we’re building with components that evolve continuously through training updates, capability improvements, and API changes. The mental model of “build it, test it, ship it, maintain it” breaks down when the underlying model your product depends on updates monthly and the new version behaves differently than the version you tested.

Traditional software has a stable lifespan. A well-built accounting system can run for a decade with minimal changes. AI systems don’t work that way. The capability floor keeps rising. A system that uses GPT-3.5 for a task that GPT-4 now handles better isn’t just suboptimal — it’s a competitive disadvantage, and the gap between the current capability and your deployed system grows every month you maintain the old architecture.

This creates a design challenge that has no equivalent in traditional software: you need to build systems that are robust enough to operate reliably at scale AND modular enough to be replaced component by component as better capabilities become available. These two requirements pull in opposite directions, and resolving the tension is one of the defining architectural challenges of AI product development right now.

Designing for Obsolescence

The practical answer is designing for obsolescence rather than designing for permanence. This means:

Modular AI architecture. Build so that individual AI components can be upgraded or replaced without rewriting the product around them. The abstraction layer between your product logic and your AI provider isn’t just a technical nicety — it’s what makes planned obsolescence economically viable. Without it, every major model upgrade is a replatforming project.

Shorter depreciation cycles. AI infrastructure doesn’t depreciate on the same timeline as traditional software infrastructure. Planning for 6-12 month cycles on AI-specific components — rather than the 3-5 year cycles that traditional software justifies — is a more honest accounting of how quickly the capability landscape moves. Finance teams that don’t understand this will consistently underfund the migration work required to stay competitive.

Evaluation frameworks that survive component changes. If your AI system doesn’t have an automated evaluation suite that runs against new model versions before deployment, you have no safe migration path. Every model update becomes a manual testing project, and the cost of that manual work is what prevents organizations from adopting capability improvements on a competitive timeline. Build the evals first. They outlast every model generation.

The Economics of Planned Obsolescence

There’s a counterintuitive economic argument here that takes time to internalize: the goal isn’t to build AI systems that last forever. It’s to build AI systems that create enough value before their obsolescence to fund their own replacement.

This is different from how most organizations think about infrastructure investment. The traditional framing is: invest once, amortize over years, minimize ongoing costs. The AI infrastructure framing is: invest in modularity, accept component replacement as a regular operational cost, and measure success by how quickly you can upgrade — not by how long you can avoid upgrading.

Organizations that accept this economic reality build faster. They ship AI improvements on the timeline of capability improvements rather than on the timeline of their budget cycle. The ones that treat AI infrastructure like traditional software infrastructure spend their engineering budget on maintaining increasingly outdated components rather than on adopting the capability improvements that would actually serve users better. Andreessen Horowitz’s AI Canon has useful framing on the infrastructure investment thesis that informs this economic argument.

The Human Side of Technological Obsolescence

There’s a team-level version of this design-for-obsolescence principle that’s equally important: the skills required to build AI products are themselves subject to rapid obsolescence.

Prompt engineering techniques that were genuinely valuable 18 months ago have been partially automated. Fine-tuning approaches that required deep ML expertise are now accessible with far less specialization. The skill set that makes a great AI product team today is different from the skill set that will matter in two years — not completely different, but different enough that teams that don’t continuously build learning into their culture will find themselves behind.

When I’m building an AI product team, I prioritize adaptability and learning velocity alongside current expertise. Current tool knowledge matters. The capacity to acquire new tool knowledge matters more over a multi-year horizon. This is the team-level version of Seneca’s principle: build organizations that benefit from change rather than organizations that resist it.


Your Turn: Apply This Today

Design for obsolescence before the next obsolescence cycle forces you to:

  • Audit your AI architecture for replaceability. For each AI component in your product, ask: if we needed to swap this model or provider in 60 days, what would it take? If the answer is “months of reengineering,” you have a design-for-permanence problem that will cost you competitively every time the capability landscape shifts.
  • Build your evaluation framework before you need to migrate. Create an automated suite that tests your AI system’s behavior against defined quality criteria. Run it against new model versions before deployment. This is the single most important investment for making planned obsolescence operationally viable.
  • Restructure your infrastructure budget cycle for AI components. Present a 12-month replacement plan for AI-specific components to your finance stakeholders — not as a failure to build durable systems, but as a realistic amortization of infrastructure that improves faster than traditional software. Make the economic argument explicitly.
  • Map your team’s “obsolescence risk” skills. Identify the skills in your team that are most likely to be automated or made obsolete in the next 18 months. Build a plan for transitioning those team members to higher-leverage work before the transition is forced. This is better for your team and better for your product.
  • Build abstraction layers into every new AI integration. Make it a team norm: no direct AI provider dependency in product logic. Every AI integration goes through an abstraction layer that can be retargeted without rewriting the product. This is the architectural equivalent of designing for obsolescence.
  • Run a “what if this model becomes unavailable?” exercise quarterly. For your primary AI capabilities, simulate a scenario where the current provider or model is unavailable in 90 days. How long does migration take? What breaks? The exercise will reveal architectural dependencies you’ve accepted without explicitly deciding to.

Seneca’s discipline about time connects directly to the hidden cost of real-time AI decisions — Seneca’s email rule and the cost of always-on AI explores the other dimension of this Stoic lens on AI product decisions.

Building AI systems and trying to make architecture decisions that hold up as capabilities evolve? I consult with product teams on AI architecture strategy, planned obsolescence frameworks, and building organizations that adapt rather than resist when the capability landscape shifts. Let’s talk.

Why Teresa Torres’ Continuous Discovery Cadence Is Breaking Down in the AI Age

Teresa Torres built continuous discovery as an antidote to product teams that build in isolation — teams that spend months building solutions to problems they think exist, then ship to users who wanted something entirely different. The framework is genuinely valuable: weekly customer interviews, systematic opportunity mapping, iterative testing. It solved a real problem.

Here’s the question I’ve been wrestling with: in an environment where AI can monitor customer behavior, surface opportunity signals, and flag behavioral anomalies continuously — not weekly, not daily, continuously — is the weekly interview cadence still the right rhythm for discovery? Or has it become a bottleneck?

The Continuous Discovery Promise

Torres’s core insight is sound and worth protecting: most product failures happen because teams build solutions for problems they imagined rather than problems that actually exist. Her framework forces the discipline of staying connected to users rather than getting lost in internal roadmap debates.

The weekly interview cadence is the mechanism that makes the framework work in human-speed product environments. A team that talks to users weekly, maps what they hear, and designs experiments based on that mapping has a significant advantage over a team that does discovery quarterly or on an ad hoc basis. That’s true and it’s worth affirming before getting into the breakdown.

Where Continuous Discovery Breaks Down

The breakdown isn’t in the principles — it’s in the gap between the discovery cadence and the signal velocity that AI-augmented product environments now generate.

When you’re monitoring support channels, usage patterns, and behavioral data continuously with AI, you’re not waiting for a weekly interview to surface a new opportunity cluster. The signal is arriving every hour. The opportunity identification is happening faster than the discovery cadence can process it. The bottleneck shifts from “we don’t have enough signals” to “we have more signals than our discovery process can evaluate.”

There’s also a deeper problem: some of the most important signals about user needs don’t come from what users say in interviews — they come from what users do that doesn’t match what they said. AI is better at surfacing the behavioral signal than interviews are. The weekly interview remains valuable, but it’s no longer the primary source of opportunity discovery in a well-instrumented product environment. Torres’s framework was designed for the interview as the primary signal. The role of the interview has changed.

The Human-AI Discovery Gap

The specific gap I’ve encountered: AI surfaces opportunity patterns from behavioral data faster than human-centered discovery processes can validate and act on them. This creates a queue of AI-identified opportunities waiting for human investigation that grows faster than the weekly interview cadence can clear it.

The result is that teams using AI-augmented discovery alongside Torres’s cadence end up in one of two failure modes: they ignore the AI-surfaced signals and run a traditional discovery process that is now missing a large class of opportunities, or they try to validate every AI-surfaced signal through the traditional interview-and-experiment cadence and create a backlog so long that validated insights are stale before they reach solution design.

The framework needs an explicit mechanism for triaging AI-surfaced signals before they enter the discovery process — something Torres’s original framework doesn’t provide because the volume of signals it was designed to handle was fundamentally human-generated.

What Discovery Looks Like With AI-Augmented Signals

The adaptation I’ve found most useful: restructure the discovery cadence so that AI handles continuous signal monitoring and human-centered discovery handles signal validation and solution design.

AI monitors behavioral signals continuously and surfaces clusters of anomalies or emerging patterns for human review. This replaces the “discovery” part of the weekly interview — the team is no longer relying on interviews to surface new opportunities, because the AI is doing that continuously. The weekly interview becomes a validation mechanism: we talk to users about opportunities the AI has already identified, rather than hoping the interview will surface something new.

This shift changes what the interview accomplishes. Instead of “tell me about your experience” (open-ended discovery), the interview becomes “the data shows users are abandoning this workflow at step 3 — can you walk me through what that experience is like for you?” (targeted validation). The insight quality goes up. The discovery efficiency goes up. The interview becomes more valuable because it’s now pointed at a specific signal rather than fishing for signals generally.

The Missing Piece in Torres’ Framework

The missing piece isn’t a critique of Torres — it’s a gap that didn’t exist when the framework was designed. The missing piece is an explicit decision rule for how AI-surfaced signals enter the discovery process and what it takes for them to earn human investigation time.

Without that decision rule, every AI-surfaced signal competes for the same human attention budget as the signals coming from interviews. The volume overwhelms the process. The team either starts ignoring signals or starts investigating everything and producing nothing.

The decision rule I use: AI signals get human investigation time when they meet a significance threshold (behavioral change above a defined magnitude), a duration threshold (the pattern persists for at least two weeks), and a strategic relevance filter (the affected workflow or user segment is in current strategic focus). Signals that don’t meet the threshold stay in monitoring. Signals that meet it get a targeted interview and a fast experiment to validate. Teresa Torres’s original continuous discovery documentation is the right foundation — the adaptation is building the signal triage layer on top of it.


Your Turn: Apply This Today

You don’t need a 20-agent AI system to apply this. Start with the principles:

  • Audit how your current discovery signals are sourced. List the top five sources of opportunity signals your team acts on. What percentage come from interviews? What percentage from behavioral data? If the behavioral data is underrepresented, you’re missing a signal class that doesn’t require AI to access — just better instrumentation and review habits.
  • Shift at least one weekly interview to validation mode. Before your next customer interview, identify one behavioral anomaly from your product data — a workflow with high abandonment, a feature with unexpectedly low engagement, a search query with no good results. Make that the focus of the interview. Validate the signal, don’t fish for new ones.
  • Build a signal triage decision rule. Define explicitly what it takes for a behavioral signal to earn human investigation time: significance threshold, duration threshold, strategic relevance filter. Write it down and apply it to your next backlog grooming session. Everything that doesn’t meet the bar stays in monitoring.
  • Create a “discovery queue” separate from your opportunity backlog. AI-surfaced signals that haven’t been validated by a human should sit in a discovery queue, not the opportunity backlog. They don’t get resourced until they’ve been validated through a targeted interview or behavioral experiment. This prevents the queue from contaminating your actual opportunity prioritization.
  • Measure your discovery process’s throughput, not just its output. Track how many signals enter your discovery process per month and how many get resolved (validated or dismissed). If the queue is growing faster than it’s being cleared, the discovery cadence needs to be restructured — not accelerated, restructured.
  • Protect the interview for the human insight it uniquely provides. As you shift interviews from discovery to validation, stay alert for the kinds of insight only an interview can generate — the emotional context, the unstated assumption, the use case nobody modeled. These won’t show up in behavioral data. The interview is still the best tool for them. Just don’t use it for signal discovery when the data is already doing that job.

The OST framework has related adaptations at scale worth reading alongside this — how OSTs need to evolve for AI-augmented discovery and the execution complexity branch that’s missing from most opportunity trees.

Running product discovery and trying to figure out how AI changes the process? I consult with product teams on discovery operations, AI-augmented research processes, and building the systems that keep human judgment at the center of AI-accelerated product development. Let’s talk.

Opportunity Solution Trees Meet AI Reality: How the Framework Breaks Down at Scale

A network of connected nodes representing the model scores behind a v1 judgment call

Teresa Torres’s opportunity solution tree framework is one of the better discovery tools I’ve used. The structure — mapping customer opportunities to potential solutions in a visual tree — helps teams stay connected to user needs rather than defaulting to feature lists and stakeholder requests. I use a version of it. But running product for a platform at significant scale, and building AI agents to support the discovery process, has taught me where the framework needs substantial adaptation before it works in the real world.

This isn’t a critique of Torres. Her framework is sound. It’s an observation that opportunity solution trees as designed assume a pace and a data volume that most scaled products don’t operate at — and that AI changes both the opportunity and the failure mode of the entire discovery process.

The Opportunity Tree Reality Check

The OST framework looks clean in workshop settings. Customer interviews surface opportunities. Teams map solutions. Experiments get scoped. Everyone leaves feeling aligned. At scale, the process breaks down in a specific way: the rate of incoming opportunity signals overwhelms the team’s capacity to process them through a manual discovery rhythm.

When you’re monitoring support channels, usage patterns, feedback streams, and behavioral data for a large user base, the volume of potential opportunity signals is enormous. The question isn’t “how do we find the opportunities?” — AI can surface opportunity clusters from that data continuously. The question is “how do we evaluate and prioritize opportunities that are emerging faster than we can interview users about them?” Torres’s framework was designed for a world where discovery is the bottleneck. At scale with AI, discovery is no longer the bottleneck — judgment about what to do with what you’re discovering is.

Where AI Changes Everything in the OST Process

The adaptations I’ve made to the OST process at scale are all about restructuring what AI handles vs. what humans handle:

AI handles continuous opportunity sensing. Rather than quarterly or monthly discovery cycles, I’ve shifted to AI agents that monitor support channels, usage patterns, and feedback streams continuously — flagging opportunity clusters for human investigation. The tree is no longer built in workshops; it’s maintained continuously with AI surfacing new branches as they emerge from data.

Humans handle opportunity prioritization and solution imagination. Torres often combines opportunity identification and solution generation in workshops. At scale, separating these produces better results. AI excels at recognizing opportunity patterns from data. Humans excel at imagining solutions that connect to those opportunities in ways that AI wouldn’t generate without significant prompting. Keeping these steps separate — and keeping humans firmly in charge of the prioritization step — prevents the discovery process from becoming AI-driven by default.

The tree becomes a living document, not a workshop artifact. The biggest practical change: the OST is updated continuously based on incoming data rather than rebuilt quarterly. AI maintains the structural connections. Humans make the strategic decisions about which branches to pursue, which to prune, and which to flag for deeper investigation.

The Missing Piece: Solution Validation Speed

Torres’s framework handles opportunity identification well and provides a solid structure for mapping solutions. Where it provides less clarity — and where I’ve had to build my own process — is solution validation speed at scale.

When you can run experiments across a large user base, the question isn’t “how do we test this?” — it’s “how do we test this fast enough to keep up with the opportunity identification rate?” If AI is surfacing new opportunity clusters weekly but your experiment design and execution cycle takes six weeks, you’re building a backlog of validated opportunities that never get addressed before the next cycle of discovery makes them stale.

The solution I’ve found: tiered experiment design. Some solutions get full experiment design and execution. Others get fast behavioral proxies — early signals from a small user segment that don’t require full experiment infrastructure. The OST needs an explicit tier for each solution branch that determines the validation approach before the experiment is scoped. Without it, every solution competes for the same limited experiment capacity and most of them wait too long.

The Cultural Complexity Challenge

For products serving users across diverse geographies and cultural contexts, the OST needs one more adaptation Torres’s original framework doesn’t explicitly address: opportunity branches that are culturally segmented rather than universal.

The same user need can manifest in fundamentally different ways across cultural contexts — different mental models, different workflows, different relationship to the product’s value proposition. AI helps identify when an opportunity pattern is consistent across user segments and when it’s culturally specific. Without that distinction, you end up building solutions to universal-seeming opportunities that actually only solve for the segment that happens to be most vocal in your feedback channels.

The practical implication: before any significant OST branch gets resourced for solution design, ask explicitly whether the opportunity is universal or segment-specific. If it’s segment-specific, design the solution for that segment first — don’t try to generalize prematurely. Teresa Torres’s original OST documentation is the right starting point for understanding the framework’s foundations before adapting it.

What Still Requires Human Judgment

AI augments the OST process at scale in meaningful ways. It does not replace the human elements that make the framework valuable. Opportunity prioritization is still a human decision — AI can surface opportunity clusters, but the judgment about which opportunity is worth pursuing given current strategic context requires the kind of synthesized understanding that AI surfaces from data but doesn’t hold.

Solution creativity remains human. The connection between a user opportunity and an imaginative solution that users didn’t know they wanted — that’s not something AI generates reliably without creative prompting. And strategic trade-offs — the decisions about which opportunities to deprioritize in service of organizational focus — require human judgment that accounts for context AI doesn’t have.

The adapted OST is better than the original for scaled products. But “better” here means AI is doing more of the pattern recognition so humans can do more of the judgment work. The goal isn’t to automate discovery — it’s to remove the bottlenecks that keep humans from doing the discovery work that actually requires human judgment.


Your Turn: Apply This Today

Whether you’re just starting with OSTs or looking to evolve your current discovery process:

  • Separate opportunity identification from solution generation in your process. If you’re currently doing both in workshops, split them into distinct phases. Use your data infrastructure (or AI tools) for continuous opportunity surfacing. Reserve workshop time exclusively for solution design and prioritization — where human creativity and judgment add the most value.
  • Build a tiered validation system for your solution branches. Before any solution moves to experiment design, assign it a validation tier: full experiment, behavioral proxy, or customer interview. The tier determines the validation approach and timeline. This prevents every solution from competing for the same experiment infrastructure.
  • Make your OST a living document, not a workshop artifact. If you’re rebuilding your OST quarterly, you’re losing continuity between cycles. Move it to a shared, continuously-updated document. Designate someone as the OST owner responsible for keeping the tree current between discovery cycles.
  • Segment your opportunity identification by user type before prioritizing. Before deciding which opportunity to pursue, ask: is this opportunity universal across my user base, or specific to a segment? If it’s segment-specific, validate it with that segment first. Don’t let your most vocal users define the universal opportunity.
  • Instrument one AI-assisted opportunity monitoring signal this quarter. Set up one automated signal that surfaces opportunity patterns from your existing data — support ticket clustering, feature usage drop-offs, search queries with no good results. Feed it into your OST review meeting as a standing input. This is the minimum viable version of continuous opportunity sensing.
  • Explicitly protect human prioritization of the OST. If AI is surfacing opportunity data, establish a clear norm: AI surfaces, humans decide. Don’t let the volume and frequency of AI-surfaced opportunities shift the prioritization decision toward “whatever AI flagged most recently.” The human judgment layer is the framework’s value. Protect it deliberately.

The execution complexity challenge in OST is closely related — I’ve written about the missing branch in opportunity solution trees that most teams overlook when mapping solutions. And the continuous discovery breakdown that AI accelerates is a distinct problem worth examining separately.

Running discovery at scale and trying to adapt your product process to the pace AI enables? I consult with product teams on discovery operations, experiment design at scale, and building the systems that keep human judgment at the center of AI-augmented product development. Let’s talk.

The Product Leader’s AI Infrastructure Blind Spot: What Jensen Huang’s Sovereign AI Argument Actually Reveals

Jensen Huang has been making a specific argument at conferences for two years now: AI infrastructure isn’t just a business advantage — it’s national security. Countries that don’t build sovereign AI capabilities will find themselves dependent on foreign nations for the intelligence layer that powers their economies. The geopolitical version of this argument is provocative. The product version of it is something most product leaders haven’t fully processed.

Here’s the product version: most AI product decisions are constrained by infrastructure choices that product leaders didn’t make and don’t fully understand. That constraint is getting more expensive as AI becomes more central to every product’s value proposition.

The Infrastructure Stack Nobody Talks About in PM Reviews

Most product leaders focus on the application layer — which model to use, how to prompt, whether to build or buy an AI feature. Huang is talking about something beneath that: the compute layer, the data center geography, the inference infrastructure, the energy capacity required to run large models. That lower layer determines what’s possible at the application layer, and most PMs have never mapped their dependency on it.

The practical version for a product organization: your AI capability roadmap is constrained by what your infrastructure stack can support, what your providers will make available to you, and what regulators in your operating markets will permit. When those constraints change — and they will change — your product decisions change with them, whether you’ve planned for it or not.

Serving users across multiple geographies makes this concrete quickly. Latency requirements for AI features vary by region. Data residency regulations vary by country. Provider availability varies by market. A product decision that works in one deployment context may be impossible in another. The infrastructure layer isn’t just an engineering concern — it’s a product strategy constraint.

The Hidden Costs of Inference Dependency

When you integrate a frontier AI model via API, the visible cost is the per-token or per-call pricing. The hidden costs are the ones that show up when things change: migration cost when you need to switch providers, renegotiation leverage you don’t have because you’ve built deep dependencies, the engineering cost of adapting when the API changes, and the product cost of capabilities being deprecated or repriced.

I’ve been working through three specific changes to how I think about AI infrastructure because of this:

Inference sovereignty audit. Mapping every AI dependency and its single-point-of-failure risk. Not just “which provider are we using” but “what happens if this entire category of capability becomes unavailable, significantly more expensive, or geographically restricted?” Most teams have never run this audit. The combined risk profile of a typical AI product is more concentrated than it appears from any individual dependency.

Hybrid deployment strategy. Running smaller, purpose-specific models for latency-critical features rather than defaulting to large frontier models for everything. Not full local deployment — the cost is prohibitive for most workloads — but strategic placement of specialized models where frontier models are overkill and edge deployment reduces latency and dependency simultaneously.

Data gravity awareness. Understanding that AI capabilities concentrate around where the data is. Organizations with unique, high-quality proprietary data have infrastructure leverage that commodity compute cannot replicate. The infrastructure investment that matters most for most product organizations isn’t compute — it’s the data infrastructure that makes proprietary training and fine-tuning possible.

The Product Leader’s Infrastructure Blind Spot

The specific blind spot Huang’s argument reveals for product leaders is this: we tend to evaluate AI infrastructure decisions the way we evaluate software vendor decisions — on current price and current capability. Infrastructure decisions have a different time horizon. The question isn’t “what does this cost today and does it work today?” It’s “what are the migration costs if this relationship changes, and what product decisions am I foreclosing by accepting this dependency?”

Product decisions made inside infrastructure constraints you don’t understand tend to look like missed opportunities in retrospect. You couldn’t build the feature because the provider didn’t support it. You couldn’t expand to a new market because the data residency requirements were incompatible with your architecture. You couldn’t negotiate on price because the switching cost was too high.

These aren’t failures of product strategy. They’re consequences of infrastructure decisions made without full visibility into the constraints they created. NVIDIA’s sovereign AI documentation lays out the infrastructure investment thesis at the national level — the product-level translation is understanding which version of these constraints applies to your organization and your product decisions.

Beyond the Hype Cycle: What Actually Changes

The sovereign AI framing matters for product leaders not because most of us are going to build national AI infrastructure, but because it makes the infrastructure constraint explicit in a way that the “just use the API” default doesn’t.

The products that survive infrastructure disruptions — pricing changes, geopolitical shifts, regulatory changes, provider strategy pivots — will be the ones built by teams that understood their infrastructure dependencies before they became constraints. Replace “national” with “product” in Huang’s argument and it holds: product sovereignty means having enough infrastructure control to make independent product decisions when the landscape changes.


Your Turn: Apply This Today

You don’t need to build your own data center to act on this. Start with visibility:

  • Run an AI dependency audit this quarter. List every AI provider, model API, and infrastructure service your product depends on. For each one, answer: what percentage of core user value depends on this? What’s the migration path if it becomes unavailable or 3x more expensive? The audit will surface concentrations of risk you haven’t explicitly accepted.
  • Map your data gravity. Identify the proprietary data assets your organization holds that an external provider cannot access. That data is your infrastructure leverage. Build the strategy for using it — fine-tuning, retrieval augmentation, evaluation datasets — before a competitor with similar data beats you to it.
  • Add infrastructure questions to your product strategy review. Before any major AI feature decision, ask: what infrastructure constraints does this create or deepen? What does the migration path look like if the provider relationship changes? These questions belong in the product review, not just the engineering review.
  • Build abstraction layers into your AI integrations. Design your AI integrations so that the provider is swappable without rewriting product logic. This is a small upfront investment that dramatically improves your negotiating position and reduces future migration cost.
  • Assess data residency requirements in your target markets. If you’re expanding geographically, understand the AI data requirements in each target market before committing to an architecture. Retrofitting data residency compliance into an AI product is significantly more expensive than designing for it in advance.
  • Run a “what if the frontier changes?” scenario annually. Ask: if the leading AI capabilities shift significantly in the next 18 months — new dominant model, new pricing structure, new regulatory environment — what product decisions would we wish we’d made differently? Use the answers to adjust your current infrastructure strategy.

The infrastructure decision connects to every other AI product decision — including the build vs. buy framework for AI infrastructure and how the sovereign AI argument translates to day-to-day product leadership.

Building AI products and trying to make infrastructure decisions that hold up as the landscape shifts? I consult with product teams on AI product strategy, infrastructure dependency frameworks, and building products that remain competitive when external constraints change. Let’s talk.

Jensen Huang’s Sovereign AI: What the Infrastructure Argument Actually Means for Product Builders

Jensen Huang’s sovereign AI argument is compelling at the nation-state level. Every country should control its AI destiny rather than depending on foreign infrastructure. The strategic logic is sound. The implementation story is considerably messier — and the mess is where the real product leadership lessons live.

I’ve watched organizations at multiple scales wrestle with the build vs. buy decision in AI infrastructure. The companies that get it right are not the ones that defaulted to either extreme. They’re the ones that developed a clear-eyed framework for evaluating the real trade-offs — and they usually had to learn that framework the hard way.

The Infrastructure Reality Check

Huang’s sovereign AI concept — building your own AI infrastructure rather than depending on external providers — carries a real cost that keynote slides consistently understate. When an organization decides to build infrastructure it could otherwise purchase, it’s making a decision that involves: significant upfront capital for compute, storage, and networking; ML engineering talent that is among the most expensive in the market; operational overhead for running, monitoring, and maintaining systems at scale; and the ongoing cost of keeping pace with a capability curve that is advancing faster than most internal teams can match.

None of this makes the build decision wrong. It means the build decision needs to be evaluated honestly against its actual cost — not against a reference point of what the infrastructure will be worth once it’s working perfectly. Most organizations evaluate the upside of sovereignty and underweight the downside of the build timeline.

Where Sovereignty Actually Matters

Not all AI capabilities are equal candidates for sovereignty decisions. The ones where building your own infrastructure makes strategic sense share a few characteristics: the capability is core to your product’s differentiation (commoditizing it would eliminate your moat), you have proprietary data that a vendor-provided model cannot access, your use case requires latency, privacy, or compliance constraints that external APIs cannot meet, or the capability is stable enough that you can amortize the build cost over multiple years.

The capabilities that are bad candidates for sovereignty decisions are the ones that are advancing rapidly at the frontier, where your internal team cannot keep pace with external model improvement, or where the task is generic enough that a frontier model already outperforms what you could build.

The practical sovereignty framework for most product organizations is not Huang’s comprehensive infrastructure ownership — it’s selective sovereignty. Own the capabilities where your proprietary advantage lives. Buy the capabilities that commodity providers can run better and cheaper than you can.

The Build vs. Buy Reality for Product Leaders

The build vs. buy decision in AI infrastructure is different from the traditional build vs. buy decision in software for one important reason: the capability being purchased is improving continuously. When you buy a license for accounting software, the software doesn’t improve significantly month over month. When you integrate a frontier AI model via API, you get the benefit of every model improvement the provider ships — without rebuilding anything.

This changes the economics substantially. A team that builds its own search ranking model might spend six months building something that outperforms GPT-4 on their specific use case. In twelve months, the frontier model has caught up and surpassed them — and they now own the maintenance burden of a system they can no longer keep competitive without sustained investment.

The sovereignty trade-off can be managed architecturally without owning the infrastructure: building abstraction layers that let you swap providers without rewriting product logic, maintaining evaluation frameworks that monitor output quality across providers, designing fallback logic for service failures, and avoiding deep API-specific integration that creates lock-in without corresponding value. This isn’t sovereignty in Huang’s comprehensive sense, but it’s practical sovereignty — controlling outcomes without controlling every component. For a deeper look at NVIDIA’s sovereign AI framework, the primary argument is worth reading directly; the implementation gap is significant.

Architectural Implications: What I’m Doing Differently

The sovereign AI argument changed how I think about AI architecture decisions — not by convincing me to build everything in-house, but by making the dependency questions explicit in every infrastructure conversation.

Before any significant AI integration decision, I now want to know: if this provider doubles their price or degrades their service quality in 18 months, what does our migration path look like? What would it cost, and how long would it take? That question changes which integration architecture you choose today. Teams that can answer it confidently have built practical sovereignty. Teams that can’t have accepted dependency without negotiating the terms of that dependency.

The goal isn’t to avoid all external dependencies. The goal is to understand your dependencies clearly and make intentional choices about which ones are acceptable risks and which ones are existential risks to your product. That discipline — applied to AI infrastructure the same way good engineering teams apply it to any critical external dependency — is the real lesson in Huang’s argument for most product leaders.


Your Turn: Apply This Today

Run this framework against your current AI infrastructure decisions:

  • Map your AI dependency profile. List every external AI provider or API your product depends on. For each one, estimate: what percentage of your core user value depends on this dependency? What happens if it goes away or becomes 3x more expensive? The map will surface risks you’ve been treating as invisible.
  • Apply the “sovereignty criteria” to each AI capability. For each capability on your list, score it against four criteria: is it core to your differentiation, do you have proprietary data for it, does it require compliance constraints external providers can’t meet, and is it stable enough to amortize a build? Build where you score 3+. Buy where you don’t.
  • Build an abstraction layer before you build anything else. If you’re integrating any AI provider today, start with an abstraction layer that lets you swap providers. It’s a small upfront investment that dramatically reduces your future migration cost and negotiating leverage with current providers.
  • Run a “provider departure” scenario annually. Once a year, ask: if our primary AI provider announced end-of-life in 6 months, what would we do? How long would migration take? What would it cost? The answer tells you your actual dependency risk — not the theoretical dependency you think you have.
  • Price your build decisions against the opportunity cost. Before committing to building AI infrastructure, calculate the engineering hours required. Then ask: what could those same engineers build that would advance user value directly? The opportunity cost is real and often underweighted in infrastructure conversations.
  • Establish a “make vs. buy” review as a standing agenda item for major AI decisions. Don’t let it default to either direction. Build the habit of explicit evaluation — cost, timeline, capability trajectory, and dependency risk — before any significant AI infrastructure commitment.

This infrastructure dependency thinking connects directly to the broader strategic framing — how the sovereign AI argument translates to product leadership decisions and why domain expertise becomes more valuable, not less, when AI handles the generic capabilities.

Navigating AI infrastructure decisions and trying to build something defensible rather than just functional? I consult with product teams on AI architecture strategy, build vs. buy frameworks, and building products that remain competitive as the AI capability landscape shifts. Let’s talk.

Munger’s Latticework and the Hidden Architecture of AI Product Systems

Most AI product failures I’ve seen aren’t caused by bad algorithms. They’re caused by good algorithms that no one thought about as a system.

A team ships a content recommendation model. It performs beautifully in testing. Six months later, they ship a search ranking model. Also strong in isolation. Then a personalization layer. Three technically sound components — and now users are getting a contradictory experience that none of the individual feature reviews could have predicted. The models are working. The system is broken.

Charlie Munger’s latticework principle — the idea that knowledge only becomes useful when it hangs together in connected frameworks rather than isolated facts — is the most precise description I’ve found for what AI product teams are missing. Here’s what it looks like in practice.

The Hidden Architecture Problem in AI Product Development

Traditional software features behave in largely predictable ways. You add a filter to a list view; it filters the list. The feature is self-contained. The mental model for reviewing it is simple: does it do what it’s supposed to do?

AI features don’t work this way. They create feedback loops. They shape user expectations. They interact with other AI features in ways that emerge from scale, not from design sessions. A recommendation engine that’s technically optimizing for engagement can be simultaneously teaching users that the product doesn’t understand what they actually need. The technical metric is green. The user relationship is eroding.

This is the hidden architecture problem: the behavior of your AI product portfolio as a system is different from — and often opposed to — the behavior of each feature in isolation. Munger’s latticework principle is the framework for thinking about both at once.

Mental Models vs. Best Practices: What the Latticework Changes

Most AI product teams operate by best practices. A/B test the recommendation algorithm. Use the highest-performing model for your use case. Personalize based on user behavior data. These aren’t wrong — they’re just dangerously incomplete when applied without the mental models that reveal their limits.

“A/B test your recommendations to optimize conversion” — the mental model that challenges this: conversion rate optimization in AI systems often conflicts with long-term trust. Users can’t easily verify the quality of AI recommendations. Optimizing for short-term conversion can train users to distrust the recommendations that actually serve them best.

“Use the highest-performing model” — the mental model that challenges this: model performance is one variable in user adoption. Latency, explainability, and consistency often matter more than raw accuracy for real-world usage patterns. A slightly less accurate model that responds in 200ms and explains its reasoning may outperform a more accurate model that takes 2 seconds and offers no explanation.

“Personalize based on behavior data” — the mental model that challenges this: personalization creates feedback loops. The system effects — filter bubbles, amplified biases, reduced serendipity — often outweigh the individual user benefit of better-targeted content. At scale, you may be building a product that each individual user finds useful while simultaneously reducing the quality of the collective experience.

These aren’t theoretical concerns. They’re the failure modes of AI products that ships feature-by-feature without a connected framework for evaluating how those features behave as a system. MIT Technology Review’s analysis of AI product failures consistently surfaces these systemic issues as more common than pure technical failures.

Building Your AI Mental Model Latticework

The practical application is a shift in the questions you ask at product review. Instead of evaluating each AI feature in isolation — does this model perform? does this feature get used? — you evaluate the portfolio as a system:

How does this feature change user expectations for every other AI feature in the product? If your AI writing assistant produces near-perfect outputs, does that raise the bar for your AI translation feature in a way that creates dissatisfaction where none existed before?

How do this feature’s feedback loops interact with the feedback loops of other AI features? A recommendation engine and a search ranking model fed by the same behavioral data are creating compound feedback loops. Are they amplifying each other toward better outcomes or toward worse ones?

How does the portfolio of AI features together change the user’s mental model of what the product is? Users don’t experience features — they experience a product. The cumulative effect of your AI features determines whether users trust the product, rely on it, or grow quietly skeptical of it. That cumulative effect is what you’re actually building, whether you’re thinking about it or not.

The Munger Test for AI Product Decisions

I run what I think of as the Munger test before any significant AI feature decision: can I explain why this feature will improve the overall product — not just its isolated metric — using at least three different mental models?

If I can only make the case from one angle — it’s technically better, or it’s cheaper, or users asked for it — I’m probably missing something important. The features that hold up under the multi-model test tend to ship well. The ones that only make sense from one angle tend to create problems we didn’t anticipate.

That’s the compound interest of connected thinking. Each mental model makes the others more useful. The latticework, over time, makes you a dramatically better evaluator of AI product decisions than any individual framework alone.


Your Turn: Apply This Today

Start building the latticework habit with these concrete steps:

  • Map your AI feature portfolio as a system. List all the AI features currently live in your product. Draw arrows between any two features that share data, shape user expectations about each other, or create feedback loops. If you’ve never done this, you will find something that surprises you.
  • Apply the three-model rule to your next feature decision. Before approving or building any AI feature, require the team to make the case from at least three different mental model angles — technical, behavioral, economic. If the feature only looks good from one angle, hold it until the others are addressed.
  • Run a “best practices challenge” in your next review. Take one of your team’s standard AI best practices and ask: what is the mental model that challenges this? What is the failure mode of this practice at scale? The answer will improve your practice or flag a risk you haven’t accounted for.
  • Evaluate user trust as a system-level metric, not a feature-level metric. Create or request a metric that captures user trust in your AI product overall — not per feature. Track it quarterly. Connect it explicitly to your AI feature decisions.
  • Design your AI features’ feedback loops on paper before you ship them. For each AI feature in your pipeline, sketch the feedback loop: what user behavior does this feature reward? What does rewarding that behavior do to user behavior over time? Is the loop pointing toward value or away from it?
  • Run the Munger test before your next roadmap commit. For each AI initiative on your roadmap, ask: can I explain the expected value using three different frameworks? If not, spend 30 minutes developing the weaker angles before committing. The discipline is the point.

The individual mental model disciplines that feed into this latticework are worth studying separately — I’ve written about how Munger’s inversion principle applies specifically to AI feature development, and how building the multi-discipline lens changes the questions you ask at product reviews.

Building AI products and looking for a sharper framework for portfolio-level thinking? I consult with product teams on AI product strategy, system-level design, and building the decision habits that make the difference between features that work in isolation and products that work at scale. Let’s talk.

Munger’s Mental Models and AI: Why the Best Product Decisions Require Multiple Lenses

Charlie Munger spent his career arguing that the single biggest mistake smart people make is solving every problem with the same framework. “To the man with a hammer, every problem looks like a nail,” he said. Most product teams building AI features are doing exactly that — and it’s costing them.

I’ve watched this play out repeatedly. A team with strong ML engineers defaults to “more data and better algorithms” for every problem, even when the actual failure is behavioral — users don’t trust the output. A team with strong business analysts defaults to “what do the numbers say?” even when the numbers are measuring the wrong thing. The mental model a team defaults to determines which problems they can see and which ones stay invisible until they ship something that breaks.

Munger’s solution — collecting mental models from multiple disciplines and using them together — is the most underrated framework for AI product decisions I’ve encountered. Here’s how I actually use it.

The Mental Models I Use for AI Product Decisions

I’m not talking about generic “think differently” advice. These are the specific lenses I apply before any significant AI feature decision:

Psychology: How does this change what users think the system understands about them? AI features carry an implicit promise — the product will understand me better than a static interface. When that promise breaks, the damage to trust is disproportionate to the error. I ask: what does this feature imply about our understanding of the user, and can we actually back it up?

Statistics: What is this actually measuring, and what’s the sample size? We once nearly shipped an AI feature that performed beautifully in testing, then failed on real user data because our test set had selection bias — it didn’t represent how users in the wild actually phrased their searches. The statistical lens caught it. The ML lens wouldn’t have.

Economics: What are the compute costs, and what are the switching costs? I’ve seen teams celebrate engagement lift from an AI feature without running the numbers on what serving that feature at scale costs per user. The unit economics of AI features can invert a business model quickly. Always model the full cost before you ship.

Operations: How do we monitor this, and how do we roll back? The most sophisticated AI feature is worthless if it can’t be operated reliably in production. Before shipping anything, I want to know: what’s the alert? What does rollback look like? Who is on call?

Behavioral economics: What cognitive biases does this feature create or exploit? AI recommendations tend to trigger automation bias — users trust confident-looking outputs more than they should. Understanding this in advance lets you design the trust calibration into the UI rather than discovering the overtrust problem after something goes wrong.

The Five Questions I Ask in Every AI Product Review

I changed my product review process because of Munger’s approach. Instead of asking “does this AI feature work?”, I now run five questions in sequence:

1. Technical: Does this solve the computational problem we think it solves — not just on test data, but on the actual distribution of real user behavior?

2. User: Does this actually improve the user’s experience, or does it just look impressive in a demo? What happens when it’s wrong?

3. Business: Does this advance the business model, and does the unit economics hold at scale?

4. Ethical: What user behaviors does this feature create or reinforce at scale? Are we comfortable with those behaviors?

5. Operational: Can we run this reliably, monitor it, and recover quickly when it fails?

A feature that can’t pass all five is not ready to ship. The discipline of asking all five questions — rather than defaulting to the one or two that align with your team’s strongest skills — is where the multi-model approach pays off.

How I Build the Mental Model Collection Habit

Munger didn’t just use multiple mental models — he actively collected them, continuously, over decades. For AI product work, here’s how I collect mine:

I read outside my discipline regularly. Cognitive psychology, behavioral economics, systems design, and operations research all surface problems that pure product management literature misses. The best AI product insight I had in the last year came from reading about medical device failure modes — not from a PM newsletter.

I actively seek out the perspective of people who think differently about the same problem. Our data scientists, engineers, designers, and business stakeholders all carry different mental models for evaluating AI features. The disagreements between them are usually where the important insight lives.

I study failures from adjacent industries. Financial services, healthcare, and automotive all deal with AI deployment challenges that preceded ours by years. Their failure modes tend to become our failure modes. Learning from their mistakes is faster and cheaper than repeating them.

Why This Matters More for AI Than Traditional Features

Traditional product features fail in fairly predictable ways — wrong assumption about user need, too complex, poor performance. AI features fail in fundamentally different ways: confident wrongness, emergent behaviors at scale, distributional shift between training and production, trust collapse after a single high-visibility error.

These failure modes don’t show up clearly through any single mental model. You need the statistical lens to catch distributional shift. You need the behavioral economics lens to anticipate trust collapse. You need the operations lens to design for recovery when something breaks at 3am. The teams that appear to consistently ship AI features that actually work are the ones that have learned to hold all of these perspectives simultaneously — not in sequence, but together.

That’s what Munger was getting at. In complex systems, the quality of your thinking matters more than the sophistication of your tools. AI is the most complex system most product teams have ever managed. The mental model collection habit is how you build the thinking to match it. For more on the practical frameworks that apply here, see how Farnam Street maps the core mental model disciplines — it’s the best resource I’ve found for building this habit systematically.


Your Turn: Apply This Today

The multi-model habit is built through deliberate practice — not a single session. Start here:

  • Run the five-question audit on your current highest-priority AI feature. Technical, user, business, ethical, operational — can it pass all five? Write down the answers before your next review. The question your team struggles to answer is where the risk lives.
  • Map your team’s dominant mental model. Ask each key team member: “What’s the first question you ask when evaluating a new AI feature?” The pattern in their answers tells you which blind spots you have as a team. Hire or consult to fill the gaps.
  • Add one non-PM discipline to your reading rotation. Pick one field — behavioral economics, systems design, operations research, cognitive psychology — and read one substantive piece per week for 90 days. Track how it changes the questions you ask in product reviews.
  • Make the unit economics of your AI features visible. Pull the compute cost data for your top three AI features. Calculate cost per user per month. If you’ve never seen these numbers, run the analysis before your next roadmap planning session.
  • Design a failure recovery path before you ship. For every AI feature in your pipeline, define in advance: what is the rollback plan? What triggers it? Who makes the call? Teams that answer these questions before shipping recover faster when something breaks.
  • Seek out the dissenting voice in your team before the next decision. Deliberately ask the person on your team most likely to see the problem differently. The disagreement is the insight. If everyone agrees immediately, you’re probably all using the same mental model.

These same multi-model disciplines apply at the team level too — how you hire PMs for AI-era roles determines which mental models your team has access to from day one. And the inversion principle is one of the most powerful single mental models in Munger’s collection for AI product work specifically.

Building AI products and struggling with decisions that require multiple frames at once? I consult with product teams on AI product strategy, decision frameworks, and building the organizational thinking habits that make great products. Let’s talk.

Munger’s Circle of Competence: Why AI Makes Domain Expertise More Valuable, Not Less

“I’m no genius. I’m smart in spots — but I stay around those spots.” That’s Charlie Munger on his own approach to investing. His “circle of competence” framework wasn’t about false modesty. It was a precise statement about where expertise creates durable advantage and where it evaporates.

Most product leaders are getting this completely backwards with AI.

They treat AI as something that expands their circle of competence — a tool that lets them write SQL they don’t understand, run experiments they can’t interpret, or make architectural decisions they can’t maintain. The result is technical debt disguised as productivity, and strategic errors disguised as confidence.

AI Amplifies What You Already Know — and Exposes What You Don’t

Munger’s insight applied to AI: the tool doesn’t expand your circle of competence. It amplifies your performance within it, while simultaneously making your gaps more consequential.

A product leader who understands growth mechanics deeply will use AI to run experiments faster, synthesize more signals, and identify patterns across larger datasets than they could manually. They get more leverage on what they already know. A product leader who doesn’t understand growth mechanics will use AI to generate impressive-looking growth frameworks they can’t evaluate, can’t debug, and can’t adapt when they inevitably fail in their specific context.

The output in both cases looks similar. The outcomes diverge dramatically.

The core failure mode is what I’d call the competence paradox: AI tools make you feel more capable in domains where you’re actually less prepared to evaluate the output. SQL you can’t debug feels more dangerous when you generate it in 30 seconds. Architectural decisions you can’t maintain feel more permanent when AI produces confident-sounding documentation. The confidence signal and the competence signal get decoupled.

The Three-Layer Filter for AI-Assisted Work

I use three questions before implementing any significant AI-generated output:

Can I debug this? If the AI generates analysis, code, or strategy, can I troubleshoot it when it breaks? Not “can I theoretically figure it out” but “can I actually debug it in context, with my real systems, under time pressure?” If not, I’m outside my circle of competence and I should either build the competence first or bring in someone who has it.

Can I improve this? If the output is 80% right, can I identify and fix the 20% that’s wrong? Last month Claude suggested a user segmentation approach that looked clever. Layer 1 caught it immediately: the approach would require behavioral tracking infrastructure we didn’t have, and the suggested proxies for that tracking were measuring the wrong thing. I spotted the gap because I understood the domain. Someone without that foundation would have shipped an experiment that wouldn’t generate interpretable results.

Can I teach this? If I can’t explain to a colleague why this approach works and where it might fail, I don’t understand it well enough to implement it. “The AI recommended it” isn’t a teaching explanation — it’s intellectual dependency dressed up as delegation.

Why Domain Expertise Compounds in an AI World

The counterintuitive conclusion: AI makes domain expertise more valuable, not less. Early research on AI productivity gains consistently shows higher lift among workers who already have strong domain knowledge — the AI accelerates their existing competence rather than substituting for absent competence.

This pattern shows up in every domain I’ve seen AI implemented well: content teams with strong editorial judgment use AI to produce more without losing quality, because they can evaluate outputs accurately. Data teams with strong analytical foundations use AI to explore more hypotheses faster, because they know which results are meaningful and which are artifacts. Product teams with deep customer knowledge use AI to synthesize more signal, because they know what the signal means.

The implication for your own development: the most important thing you can do to increase your AI leverage is deepen your domain expertise, not broaden your AI tool exposure. Munger’s advice — stay around your spots, maximize leverage within them — is more applicable than ever.

Munger’s broader framework for rational decision-making is worth reading with AI in mind: the latticework of mental models he described is exactly the foundation that allows you to evaluate AI outputs rather than simply accept them. The people who will get the most from AI are the ones who’ve built strong mental models in their domain — not the ones who’ve accumulated the most prompting tricks.

If you’re also thinking about how AI changes what you need to hire for, this connects directly to the AI-era PM hiring question: the skills that matter most aren’t AI-specific, they’re judgment-based — which is exactly what Munger would have predicted.


Your Turn: Apply This Today

In an AI era, knowing your circle of competence — and staying inside it — becomes your strategic advantage. Here’s how to apply it:

  • Draw your circle of competence. Write down the specific domains where your experience gives you judgment that a general AI cannot replicate — industries, user types, technical contexts, organizational dynamics. Be ruthless about what’s inside the circle and what’s outside it.
  • Use AI as a force multiplier inside your circle, not an extension beyond it. Where you have deep expertise, use AI to go faster. Where you don’t, use AI to learn — but don’t use it to pretend you have expertise you don’t. The accountability gap will catch up with you.
  • Hire for complementary circles, not identical ones. When building an AI-era team, map each hire’s domain expertise circle. Where do they overlap with yours? Where do they extend it? Teams with overlapping but complementary circles outperform teams of generalists using the same AI tools.
  • Be explicit about AI-generated work that falls outside your expertise. When you use AI to produce work in a domain outside your circle — legal, financial, medical, technical — label it explicitly. Peer review it with someone inside that circle. AI confidence is not domain expertise.
  • Shrink before you expand. Before trying to extend your circle of competence using AI, go deeper in the areas where you’re already strong. AI amplifies depth more reliably than it substitutes for it. Master your existing domain before using AI to colonize new ones.
  • Build a “competence map” for your product team quarterly. Identify where your team’s circle of competence is strong, where it’s thin, and where AI is filling a gap that should be filled by a human hire. Treat the gap analysis as a strategic workforce planning tool.

Building an AI-native product team and figuring out where to invest in deepening domain expertise vs. AI capability? I consult with product leaders on AI strategy, team development, and the competence questions that determine whether AI investments pay off. Let’s talk.

Seneca’s Email Rule and the Hidden Cost of Real-Time AI

Seneca had a rule about correspondence: don’t respond to letters the moment they arrive. Set them aside. Let the thought settle. Respond when you have something worth saying rather than when social pressure demands a reply.

That’s not how we use AI.

We’ve built real-time AI into everything — always-on assistants, instant analysis, immediate responses to any question. The design assumption is that faster is better, that removing latency from every cognitive task creates value. But there’s a cost to always-on AI that’s rarely discussed: it can quietly erode the capacity for sustained, independent thinking that makes good product decisions possible.

The Real-Time AI Trap

Here’s a pattern I’ve noticed in my own work, and in teams I consult with: AI availability changes the character of thinking, not just the speed. When you know an instant answer is available, the instinct to work through a problem yourself diminishes. Not dramatically — just enough to matter.

I caught myself asking Claude to “think through the pros and cons” of a product prioritization decision I’d made dozens of times before. I wasn’t seeking new perspective. I wasn’t dealing with an unusual edge case. I was just avoiding the mild friction of working through it myself, because the alternative (ask AI, get immediate answer) was so frictionless.

That’s a problem. Not because AI analysis is bad, but because the ability to reason through prioritization decisions independently is a core PM competency — and like any competency, it atrophies without use. Offloading the thinking doesn’t just remove the task; it removes the practice.

Offloading vs. Outsourcing: The Line That Matters

There’s a useful distinction worth making explicit:

Offloading is using AI to handle tasks that don’t require your specific judgment — summarizing a 50-page research report, extracting action items from meeting transcripts, generating first-draft language for a spec. The value is real and the cost is low. You’re not losing anything by having AI do these tasks. You free up cognitive bandwidth for higher-leverage work.

Outsourcing is using AI to handle the judgment calls that should be yours — deciding which problems to prioritize, evaluating strategic trade-offs, determining what your users actually need vs. what they’re asking for. The immediate output looks similar, but the long-term cost is real: you’re not building the judgment muscle that makes you better over time. You’re borrowing it from a model that has no stake in your actual outcomes.

The line between the two is often subtle and context-dependent. Summarizing research is usually offloading. Asking AI to “tell me what to focus on this quarter” is usually outsourcing. “Help me stress-test the prioritization I’ve already done” is offloading. “What should we build next?” is outsourcing.

A Batched Approach to AI That Preserves Judgment

The batching principle is the practical application of Seneca’s rule. Instead of treating AI as a real-time response system for every cognitive task, structure AI interaction into deliberate sessions with protected thinking time in between.

Here’s the workflow I’ve been running for eight months:

Morning AI session: Feed the system previous day’s metrics, key questions I’m working on, and documents requiring analysis. Let it run. Close the interface. Don’t interact with AI analysis until I’ve done my own initial thinking on the same questions.

Midday synthesis: Review AI output alongside my own thinking. The AI analysis is one data point — sometimes it changes my view, sometimes it confirms it, sometimes it raises considerations I missed. The key: I’ve already done the first-pass reasoning before seeing what the AI produced.

Protected thinking blocks: Two 45-minute windows per day with AI interfaces closed. This is where strategy work happens — the kind of slow, uncomfortable reasoning about difficult trade-offs that can’t be outsourced without losing the benefit of doing it.

The counterintuitive result: I use AI more effectively when I use it less constantly. The batching creates space for my analysis to develop before AI analysis augments it, which produces better final outputs than feeding every question to AI the moment it arises.

The Organizational Design Implication

If you’re leading a product team, the real-time AI trap isn’t just a personal productivity issue — it’s an organizational one. Teams that use AI to substitute for strategic thinking will develop weaker strategic thinking over time. The meetings get faster. The decisions get more confident. And the quality of the underlying reasoning quietly declines.

The antidote is building deliberate judgment-development practices into team rhythms: requiring people to bring their own analysis before AI analysis is introduced, creating space for debate and disagreement that doesn’t get short-circuited by “the AI says,” and treating the quality of reasoning — not just the quality of outputs — as something worth protecting.

As Kahneman’s research shows, fast thinking is efficient but prone to systematic error. The AI automation paradox is that AI often makes fast thinking faster without making it more accurate — which is exactly the problem Seneca was trying to solve with his correspondence rule two thousand years ago.


Your Turn: Apply This Today

The cost of real-time AI is often invisible until it shows up in your team’s burnout or your product’s churn rate. Here’s how to surface and manage it:

  • Audit your AI features for “always-on” demands. List every AI feature in your product that creates an expectation of real-time response — from users or from your systems. For each one, ask: is this expectation necessary for value, or just for perceived responsiveness?
  • Calculate the true infrastructure cost of real-time AI. Pull the cost data on your most-used AI features. Break it down by: inference cost, latency infrastructure, on-call engineering support, and incident response. Present the full number to leadership alongside the user value. That’s the honest trade-off.
  • Segment your users by latency sensitivity. Not all users need real-time AI. Survey or instrument your user base to understand who notices — and cares about — latency vs. who just wants accuracy. Design tiered AI experiences that match investment to actual user need.
  • Introduce “async AI modes” for non-time-critical workflows. Identify user workflows where a 5-minute or 1-hour delay is completely acceptable. Move those workflows to async processing. Invest the cost savings in making real-time AI faster where it actually matters.
  • Set a “response time contract” with your users. Be explicit in your UI about when AI outputs are real-time vs. batch-processed. Users who know when to expect results are more patient than users who feel they’re waiting unnecessarily.
  • Apply Seneca’s discipline to your roadmap. Before adding any new real-time AI feature, ask: is this urgent for the user, or just urgent-feeling? If the value doesn’t require immediacy, design for async first and real-time only when the user’s workflow demands it.

The batching approach also helps with the problem Kahneman’s framework identifies: when you don’t interact with AI constantly, you can engage System 2 evaluation at the right moments rather than burning cognitive load on continuous vigilance.

Building AI into your product team’s workflow and watching independent judgment quietly erode? I consult with product organizations on AI work design, decision quality, and preserving the judgment capacity that makes AI actually useful rather than just fast. Let’s talk.

Kahneman’s System 1 and the AI Automation Paradox: Why Cognitive Load Goes Up, Not Down

Kahneman’s dual-process theory splits human cognition into System 1 (fast, automatic, intuitive) and System 2 (slow, deliberate, analytical). System 1 handles routine pattern-matching instantly. System 2 engages for novel problems requiring conscious reasoning. The trouble is that System 2 is expensive — it’s slow, effortful, and tires easily. Humans naturally try to minimize how much they use it.

Here’s what Kahneman didn’t fully anticipate when he published Thinking, Fast and Slow in 2011: what happens when AI automates the System 1 tasks but simultaneously requires System 2 oversight of everything it does.

That’s the AI automation paradox. And it’s creating a cognitive load problem that most product teams haven’t fully reckoned with.

Why AI Automation Increases Cognitive Load Instead of Reducing It

When I started running a 20-agent AI system, I expected it to free up cognitive bandwidth. It did — for the tasks the agents handled correctly. What I didn’t account for was the continuous System 2 vigilance required to monitor outputs I could no longer trust automatically.

My email triage agent sorts and prioritizes correctly about 90% of the time. That sounds high. But 10% errors across 200 emails a day means 20 potentially misrouted messages — enough to create real problems. So instead of spending System 1 attention on scanning email (which I was good at), I now spend System 2 attention verifying the agent’s categorizations. That’s cognitively harder, not easier.

This matches what Ethan Mollick documented in his research on human-AI collaboration: AI creates “jagged frontiers” where performance is excellent in some areas and surprisingly poor in adjacent ones, with no obvious pattern. Users can’t develop the intuitive trust that would allow System 1 to handle AI oversight. Every interaction requires deliberate evaluation.

The result: permanent System 2 vigilance for tasks that AI is supposedly handling for you.

The Design Implications for AI Products

Understanding the Kahneman paradox should change how you design AI features — both for your users and for your own team’s workflows.

Narrow scope reduces oversight burden. The AI tools I trust most handle a single, well-defined task with transparent reasoning. My meeting transcript parser does one thing: it extracts action items and shows its reasoning for each one. The narrow scope makes verification fast. The broad-scope tools that promise to “handle everything” create the most cognitive load because you can never develop calibrated trust — each output requires full evaluation.

Confidence signaling is a product feature, not a nice-to-have. Tools that flag their own uncertainty let users shift appropriately between System 1 and System 2. When an output is flagged as low-confidence, the user engages deliberate evaluation. When it’s flagged as high-confidence, they can trust it more readily. This isn’t just better UX — it’s cognitively more honest about what AI actually does. Most AI tools present every output with equal confidence, which forces users into constant System 2 vigilance as the safe default.

Failure mode design matters more than accuracy optimization. The most dangerous AI failure is not “wrong with high confidence” on a flagged edge case — it’s “wrong with high confidence on something that looks routine.” Design your AI features to fail obviously, not silently. When the agent can’t handle something well, it should say so explicitly rather than producing a plausible-sounding but incorrect output. Graceful, visible degradation builds appropriate trust calibration over time.

When to Let System 1 Take Over

The goal isn’t permanent System 2 vigilance — that’s unsustainable and defeats the productivity case for AI. The goal is building the pattern recognition that allows appropriate trust to develop over time.

This happens naturally when the AI operates in a narrow enough domain, with enough transparency, for long enough that users can develop accurate intuitions about where it succeeds and where it fails. That’s the design objective: not “AI handles everything” but “users develop calibrated trust in specific AI behaviors” — which takes time, transparency, and intentional scope management on your part as a product builder.

Nielsen Norman Group’s research on AI mental models points in the same direction: users who develop accurate mental models of AI capabilities use AI tools more effectively and report higher satisfaction than those operating with either over-trust or under-trust. Your product design should actively support accurate mental model formation, not just maximize apparent capability.

If you’re thinking about how this connects to team productivity, the Kahneman paradox is part of why AI makes knowledge workers more productive but not always more effective — the cognitive overhead of AI oversight is a real cost that productivity metrics rarely capture.


Your Turn: Apply This Today

The cognitive load paradox is subtle but consequential. Here’s how to design around it:

  • Audit your AI features for “trust calibration.” For each AI output your product surfaces, ask: does the UI communicate how confident the AI is, and what the cost of being wrong is? Presenting uncertain outputs with the same visual weight as certain ones trains users to overtrust.
  • Design “friction with purpose” into high-stakes AI decisions. Where an AI recommendation could lead to a significant downstream consequence, add a confirmation step — not to slow users down, but to activate System 2 thinking. The pause is the feature.
  • Track “AI override rates” as a product health metric. Measure how often users accept AI suggestions without modification vs. how often they edit or reject them. If the override rate is near zero, users may be overtrusting outputs that don’t deserve it.
  • Run a “cognitive load audit” on your highest-traffic AI workflow. Map every decision a user makes in the flow. For each one, ask: is the AI reducing cognitive load in a way that helps, or in a way that just defers the thinking to a less capable moment?
  • Test your AI features with novice and expert users separately. The cognitive load paradox hits differently across experience levels. Experts are more likely to catch AI errors; novices are more likely to overtrust. Design different trust signals for different user segments.
  • Build an “error recovery” path for every AI recommendation. Design the product so that when an AI suggestion turns out to be wrong, the user can recover without significant cost. If the cost of recovery is high, you must add more human checkpoints before the decision commits.

Building AI features and seeing that users aren’t developing the trust and adoption you expected? I consult with product teams on AI UX design, cognitive load, and building the right trust architecture for AI-powered products. Let’s talk.

When AI Access Becomes Universal: What Actually Differentiates Your Product

Sam Altman’s “Universal Basic Compute” proposal — providing every person direct access to computing power rather than just AI-generated outputs — is one of the more interesting ideas circulating in AI policy circles right now. The concept: instead of waiting for AI productivity gains to trickle down through labor markets, give people direct stakes in the compute layer itself. Think universal healthcare, but for GPUs.

Whether or not Altman’s specific proposal gains traction, the question it raises is worth sitting with for anyone building AI products: When compute becomes democratized, what actually determines who wins?

The Access Fallacy in AI Product Strategy

A lot of product strategy right now is implicitly built on access as the moat. We have access to GPT-4, or Claude, or Gemini — and our competitors don’t, or can’t afford it, or don’t know how to integrate it yet. The access gap creates a temporary competitive advantage.

But access moats in software have a historical tendency to close faster than anyone expects. AWS made infrastructure a commodity. Stripe made payments a commodity. Shopify made e-commerce infrastructure a commodity. In each case, the teams that built durable advantages weren’t the ones who had early access to infrastructure — they were the ones who had deep understanding of the customers that infrastructure served.

AI access is on the same trajectory. When compute democratizes — and it will, whether through Altman’s UBC proposal or through market price compression or through open-source model proliferation — the teams with genuine customer depth will compound. The teams whose strategy was primarily “we have AI and they don’t” will find themselves competing on a level playing field with no differentiation.

What Actually Differentiates When Access Is Universal

I’ve been thinking about this through the lens of what I’ve observed building digital products for ministry and faith-based audiences across multiple continents. The access gap closed faster than expected in those markets — within 18 months of us integrating AI capabilities, smaller competitors had comparable features. The products that retained users weren’t the ones that had AI first. They were the ones that had the best data on their specific users, the deepest understanding of context, and the most nuanced sense of what “correct” actually means for their audience.

Universal compute access amplifies this dynamic. When everyone can fine-tune models on their domain data, the differentiator isn’t who has AI — it’s who has:

  • Better training data. Your proprietary understanding of user behavior, domain-specific language, and edge cases that general models handle poorly.
  • Deeper contextual judgment. The ability to know when the AI output is 95% right but contextually wrong in the 5% that matters most for your specific users.
  • Domain expertise that AI can’t replicate without you. The tacit knowledge of what “good” looks like for your users that only comes from years of serving them closely.

This is essentially the sovereignty question applied at the product level: the teams with data sovereignty and domain sovereignty will outperform those with only infrastructure access.

The Distribution Problem That Persists

Altman’s proposal also highlights something product teams building global products need to think about more carefully: democratized access doesn’t automatically produce democratized value.

When we expanded our platform to serve users in the Global South — smaller congregations, rural communities, under-resourced organizations — equal access to the same AI features didn’t produce equal outcomes. The features were calibrated on majority-culture usage data. The AI’s judgment about what was “helpful” reflected a particular cultural context. The teams with equal access to the technology were not equally served by it.

Real democratization isn’t just access to compute. It’s access to compute that reflects your context, your language, your use cases, and your community’s understanding of what good looks like. That gap doesn’t close by lowering the price of GPUs. It closes through deliberate investment in representative data and culturally-aware product design — which is an organizational commitment, not an infrastructure problem.

What This Means for Your Roadmap Now

If your AI product strategy relies primarily on access advantages, now is the right time to audit what else you’re building. The questions worth asking:

  • What proprietary data do you have that makes your AI outputs genuinely better for your specific users than a competitor using the same foundation model?
  • What domain knowledge is baked into your product design that a new entrant with equal AI access couldn’t replicate quickly?
  • Are you building the customer relationships and feedback loops that compound into better AI outputs over time?
  • If your AI access became equally available to your top three competitors tomorrow, what would still differentiate you?

The teams thinking through those questions now will be better positioned when the access gap closes — and it will close. The question is whether your strategy is built on a foundation that gets stronger when it does, or one that becomes irrelevant.

This connects to the deeper question of what your product is actually for — because access democratization also changes who your product is obligated to serve well, not just who it can technically reach.


Your Turn: Apply This Today

If AI is table stakes, your differentiation must come from somewhere else. Here’s how to find it:

  • Inventory your non-AI moats. List every source of competitive advantage your product has that has nothing to do with AI — network effects, proprietary data, community, brand trust, distribution, switching costs. If the list is short, that’s your strategy problem to solve.
  • Run a “when AI does it for free” scenario. Identify the three AI features you’re most proud of. Then ask: if OpenAI, Google, or Microsoft offered the same capability for free tomorrow, what would remain of your value proposition? Build toward that remainder.
  • Find the “contextual advantage” in your market. What does your product know about your specific user context — their industry, their past behavior, their workflow — that a general AI cannot know without being told? Build features that leverage that context depth.
  • Invest in trust infrastructure, not just feature infrastructure. As AI becomes ubiquitous, users will choose products they trust to handle their data, protect their privacy, and behave consistently. Document and communicate your AI governance practices. Trust is a differentiator.
  • Map the “last mile” problem in your category. AI handles the generic parts of most workflows. What’s the last 20% that requires human judgment, contextual knowledge, or professional accountability? Design your product to excel in that space, not compete in the generic space where AI wins on cost.
  • Define your “AI commodity” timeline. Make a bet: in 18 months, which of your current AI features will be commoditized? Work backward from that bet to decide what to build now that has staying power beyond the commoditization wave.

Building AI products and thinking through what differentiates you as the access moat closes? I consult with product leaders on AI product strategy, data differentiation, and building durable competitive advantages in AI-native markets. Let’s talk.

The Scaffolding Problem: Why AI Learning Tools Build Dependency Instead of Capability

Ethan Mollick has a scaffolding problem. He’s spent years making the case for AI as a learning partner — Co-Intelligence is genuinely persuasive on this — but the research on AI-assisted learning keeps surfacing a troubling pattern: when people use AI as scaffolding to learn complex skills, they often become dependent on the scaffold rather than developing the underlying capability.

This isn’t unique to AI. Apprenticeship models have been solving — and sometimes creating — this problem for centuries. The master-apprentice relationship is the oldest answer to the question of how you transfer tacit knowledge from someone who has it to someone who doesn’t. It also predates every digital learning system by millennia, and it still works better than most of them. Understanding why tells you a lot about where AI learning tools succeed and where they’re going to fail.

The Scaffolding Paradox in Practice

Here’s the problem in concrete terms: I’ve built a 20-agent AI system that handles significant portions of my workflow. I write better first drafts faster, synthesize research more thoroughly, and spot patterns across data sets that I would have missed before. These are genuine capability improvements.

But when I assess the people on my teams who have adopted AI tools most enthusiastically, I’m noticing something uncomfortable: their AI-assisted work product is improving, but their independent judgment on the same types of problems isn’t improving at the same rate. They’re getting better outputs with AI. They’re not necessarily getting better at the underlying reasoning that generates good outputs.

This is exactly what educational technology research on scaffolding has documented for decades. Scaffolding that removes the struggle of learning — that answers the question before the learner has grappled with it — produces more efficient short-term outputs and less durable long-term capability. Students who use calculator apps become faster at arithmetic and slower at mathematical reasoning. Writers who use grammar checkers produce cleaner prose and develop weaker editorial instincts. The scaffold substitutes for the learning rather than accelerating it.

What Apprenticeship Gets Right That AI Scaffolding Usually Misses

The master-apprentice model works because it’s built on three principles that most AI learning tools violate:

Context before process. Apprentices don’t get handed a manual. They observe the master working in real contexts — seeing the judgment calls, the trade-offs, the moments of uncertainty before decisions. The learning is embedded in practice, not abstracted into steps. Most AI tutoring optimizes for process efficiency: here’s the template, here’s the workflow, here are the five steps. The tacit knowledge of when to apply which approach — the judgment layer — doesn’t transfer through templates.

Struggle as pedagogy. Good mentors don’t rescue apprentices from difficulty; they use difficulty as the teaching medium. The moment the apprentice gets stuck is often the highest-value learning moment — it surfaces where their mental model diverges from reality. AI scaffolding tends to eliminate this. Stuck on a document? AI drafts it. Can’t find the right framing? AI suggests five. The struggle that would have built capability gets bypassed entirely.

Progressive responsibility transfer. The best apprenticeship relationships follow a deliberate arc from observation to supported practice to independent execution with feedback to autonomous work. This progression is designed, not accidental. Most AI tools don’t have a model of the learner’s development over time — they just answer whatever question is asked, at whatever level of support is requested, with no view toward building toward independence.

What This Means for Building AI Learning Tools

The AI learning products that will create durable value are the ones that preserve productive struggle rather than eliminating it. Practically, this means:

Design for delayed assistance. The most sophisticated AI tutoring designs I’ve seen require users to make their own attempt before unlocking AI assistance. This preserves the learning that happens in the struggle phase while preventing frustration from becoming discouragement. The friction is intentional — it’s doing work.

Ask questions instead of providing answers. An AI that responds to “how do I approach this problem?” with a clarifying question — “what have you tried so far, and where did that break down?” — is building the learner’s reasoning. An AI that just answers the question is substituting for it. Mollick’s own research on AI pedagogy points in this direction: the AI-as-Socrates model outperforms the AI-as-encyclopedia model for capability development.

Build mental models explicitly. Scaffolding that helps users understand why something works — not just what to do next — builds transferable reasoning rather than step-following. This is harder to build and harder to measure, which is why most AI tools don’t do it. But it’s the difference between producing capable practitioners and producing tool-dependent process followers.

The Stakes Are Higher Than They Look

If you’re building AI tools that people use for learning or skill development — whether that’s a sales enablement tool, a customer service training platform, an onboarding system, or an educational product — the scaffolding question determines whether your product creates lasting value or lasting dependency.

Products that make people capable compound in value. Users become more effective over time, credit the tool for helping them get there, and advocate for it because they can point to real skill growth. Products that create dependency are fragile — valuable when present, debilitating when absent, and increasingly resented when users realize they haven’t actually grown.

I’ve seen this dynamic in the platforms we’ve built for ministry and faith formation: tools designed to make practices easier often made practitioners weaker. The same dynamic shows up in corporate learning systems, sales tools, and product development processes. The question isn’t whether AI can help people work through complex tasks. It’s whether it’s helping them build the judgment to work through those tasks more independently over time.

If you’re thinking about how AI tools change the skill requirements for product roles, this connects directly to the hiring question: the PMs who will be most valuable in an AI-native organization are the ones who’ve built judgment, not just the ones who’ve built AI fluency.


Your Turn: Apply This Today

Designing for learning — rather than just task completion — requires intentional choices. Here’s where to start:

  • Audit your AI features for scaffolding vs. substitution. For each AI feature in your product, ask: does this help users develop capability, or does it complete the task so thoroughly that users stop developing the underlying skill? If it’s the latter, you have a dependency risk.
  • Design one “fade-the-scaffold” feature this quarter. Identify an AI assist that new users need but experienced users don’t. Build a mechanism that gradually reduces the assist as user competency grows — like training wheels that automatically come off.
  • Interview users about what they’ve learned from your product. Ask: “What can you do now that you couldn’t before you started using us?” If the answer is “nothing — the AI just does it for me,” you’re building dependency, not capability. Both can be valid business models, but know which one you’re running.
  • Add a “learning mode” vs. “efficiency mode” toggle to high-dependency features. Let users choose whether they want the AI to do it for them or guide them through doing it themselves. The toggle itself signals that you’ve thought about skill development.
  • Evaluate your retention metrics through a dependency lens. High retention can mean high value or high lock-in. Ask: are users retained because the product makes them better, or because they can no longer function without it? The distinction matters for long-term product health and user trust.
  • Build a “transfer test” into your user research. Periodically ask long-term users to attempt their core task without AI assistance, then with it. Measure the gap. If the gap is growing over time, your scaffold has become a crutch.

If you’re thinking about AI collaboration in product teams more broadly, this scaffolding question is the flip side of the AI-as-coworker conversation — both are about how humans and AI systems should divide cognitive labor over time.

Building AI tools for learning, training, or skill development — and want to design for capability growth rather than dependency? I consult with product teams on AI-assisted learning design, scaffolding strategy, and building products that create durable user value. Let’s talk.

Barry Schwartz Was Right: Why AI Feature Creep Is a Choice Paradox Problem

Barry Schwartz’s research on choice overload — documented in The Paradox of Choice — showed that more options reliably make people less satisfied and less likely to choose. The famous jam study: a display with 24 varieties attracted more browsers but generated far fewer purchases than a display with 6. More options, less action.

AI has created a jam aisle with infinite varieties, and most product teams are treating it like a feature advantage. Every conversation about “AI capabilities” sounds like a menu planning session: look at everything we can do. The technology can theoretically do anything, so the assumption is that users want access to everything.

They don’t. And the products that understand this are going to win.

AI Feature Creep Is a Choice Paradox Problem

The pattern plays out the same way across AI product launches I’ve seen: a team builds a genuinely capable AI feature and launches it with broad capability — “ask it anything,” “summarize any document,” “generate whatever you need.” Usage is lower than expected. User feedback is confused. The team concludes the AI isn’t good enough and invests in improving the underlying model, when the real problem is the interface design choice to expose unlimited capability without opinionated defaults.

When I first started using GPT-4 for product work, I stared at the blank prompt box for longer than I’d like to admit. The capability was extraordinary. The interface was paralysis-inducing. Every time I sat down to use it, I first had to decide what to use it for — which was itself cognitive work that the tool didn’t help with.

Adding constraints made it more valuable, not less. “Analyze this data” got mediocre results. “Analyze this data for three specific insights about user retention that I might be missing” got genuinely useful outputs. The constraint reduced choice for the AI and for me, and the outputs improved dramatically.

The Curation Advantage Is Underexploited

Look at how the best AI product integrations actually work. Replit’s AI coding assistance doesn’t give you a generic AI chat interface. It offers contextually relevant, constrained actions: fix this bug, explain this function, write tests for this code. Each action is integrated into the development workflow rather than floating as a generic capability. The AI does more because it’s been given less to choose from.

Notion’s AI implementation shows the same principle: when you highlight text, it surfaces contextually relevant options based on content type rather than displaying all possible AI actions. The AI is making choices about what choices to offer, which is the genuinely hard product design problem.

Choice architecture research tells us users don’t want fewer options — they want intelligently structured ones. The goal isn’t to eliminate capability. It’s to make good default choices for users so they can focus cognitive resources on the work that matters rather than on deciding how to use the tool.

Three Design Principles for Fighting AI Choice Overload

Lead with workflow, not capability. Don’t present what the AI can do in the abstract. Present what users are trying to accomplish and show AI as the accelerant for specific, recognizable tasks. “Summarize this meeting” is a task. “AI-powered meeting assistant with transcription, action item extraction, sentiment analysis, and searchable archives” is a capability list that makes users wonder which part to try first.

Default to constrained, allow expansion. Ship with opinionated defaults that work for 80% of use cases. Make it possible to unlock broader capability, but don’t lead with it. Users who need the full flexibility will find it. Users who don’t will get value faster and with less friction. The majority of your users will never need the advanced settings — designing for them first is a mistake.

Let AI curate its own interfaces. The most sophisticated AI product design pattern: use AI to decide what choices to offer users, based on context. This is what the best implementations do — contextual, dynamic option presentation that reduces the meta-choice burden. You’re not hiding capability. You’re surfacing the right capability at the right moment.

The Counter-intuitive Competitive Advantage

In a market where every competitor is adding AI capabilities, the counter-intuitive competitive advantage is ruthless curation. The product that makes users feel capable and focused beats the product that makes users feel overwhelmed, even if the overwhelmed product has technically superior AI.

This is Schwartz’s insight applied to product strategy: the winning AI products won’t be the ones with the most features. They’ll be the ones that make the most confident choices on behalf of users, removing the cognitive burden of deciding how to use powerful tools. Those are design decisions, not model decisions — and they’re available to any team willing to make opinionated calls about what matters most for their users.

If you’re thinking about how AI feature design connects to user trust over time, the inversion framework is worth running on any AI feature launch: explicitly ask how unlimited capability exposure could destroy user confidence, and design against those failure modes.


Your Turn: Apply This Today

Your next sprint planning is an opportunity to say no more deliberately. Here’s how:

  • Conduct a “choice audit” on your core workflow. Walk through the main path a new user takes in your product. Count every decision they’re asked to make before they get to value. If it’s more than five, you have a choice overload problem to solve.
  • Apply the “brilliant friend” test to your AI features. Would a brilliant friend with your product’s capabilities offer 12 options and ask you to choose — or would they just give you the best answer? Redesign your AI features to behave like the friend, not the dropdown menu.
  • Eliminate one “power user” feature from your primary navigation this sprint. Identify a feature that 80% of your users never touch. Move it to an advanced settings section. Measure whether new user activation improves. It almost always does.
  • Rewrite your AI feature descriptions to emphasize what they decide for the user. Instead of “Choose from 8 AI writing styles,” try “We’ll match the right tone for your audience.” Confidence in the system builds trust. Choice signals uncertainty.
  • Set a “maximum options” standard for your product design system. Define explicitly: no UI element presents more than N choices at a time without progressive disclosure. Make it a design constraint, not a suggestion.
  • Interview users who churned in the evaluation phase. Ask them: “Was there a moment where you felt overwhelmed or unsure what to do next?” If you hear it more than twice, you’ve found the choice overload drop point. Fix that before you add another feature.

Building AI features and fighting the instinct to expose every capability you’ve built? I consult with product teams on AI product strategy, feature prioritization, and the design decisions that create durable user value. Let’s talk.

Drucker’s Knowledge Worker Paradox: Why AI Makes You More Productive and Less Effective

Peter Drucker wrote The Effective Executive in 1967, before knowledge work had fully displaced industrial work as the dominant mode of economic value creation. His core argument: effectiveness — producing the right results — is a discipline that can be learned. Efficiency — doing things faster — is a distraction if you’re doing the wrong things.

Seventy years later, AI is making Drucker’s distinction more urgent, not less. I’ve been running a 20-agent AI system for eight months. My output has never been higher — specs drafted faster, research synthesized in minutes, emails triaged and responded to before I’ve even looked at my inbox. My impact? That’s a different question, and a harder one to answer.

The AI Productivity Paradox

Here are the three traps I’ve observed across multiple organizations implementing AI tools, including my own teams:

The Tool Trap: Teams adopt AI without changing their work design. The result is faster generation of the same ineffective outputs. The company was already producing too many status updates, too many options without clear decision criteria, too many documents nobody reads. AI helps them produce more of that, faster. The work product volume goes up. The decisions don’t get better.

The Volume Trap: Individual contributors use AI to produce more — more analysis, more options, more documents. Their managers now have twice as much to review with no additional capacity for review. Decision quality decreases as cognitive load on decision-makers increases. The people who benefit most from AI productivity tools are often creating bottlenecks for the people above them.

The Speed Trap: AI enables rapid iteration, so teams iterate rapidly — on problems that aren’t well-defined. Speed without clear objectives doesn’t compress learning cycles. It just burns through resources faster on the wrong problems. I’ve watched teams use AI to generate ten variations of a product concept before anyone had agreed on what problem the product was solving.

What AI Does to Drucker’s Framework

Drucker defined knowledge workers as people whose tool is their own knowledge — whose productive output is the application of informed judgment to complex problems. The paradox: AI dramatically amplifies the volume of work knowledge workers can produce while doing almost nothing to improve the quality of their judgment.

I can draft a product strategy document in 20 minutes with AI assistance. That used to take four hours. The time savings are real. But strategy isn’t about document creation — it’s about making difficult choices with incomplete information, building conviction in a direction when reasonable people disagree, and committing resources before you have certainty. AI doesn’t do any of that. It helps me produce the artifact of strategic thinking faster, not the strategic thinking itself.

When I was at a consulting firm early in my career, I could generate competitive analyses, user journey maps, and market segmentation frameworks faster than anyone on the team. But I was optimizing for the wrong thing. The clients didn’t need more frameworks — they needed clearer decision-making. My AI-enabled productivity made me feel more valuable while I was actually less impactful than the partner who spent three hours in a room with a client asking uncomfortable questions with no slides in sight.

What Effective AI Use Actually Looks Like

The organizations getting AI right share a common pattern: they redesigned work around outcomes, not activities. They asked “what decisions do we need to make better?” before they asked “what tasks can AI help with?”

One product team I work with stopped using AI to generate more user personas. They use it to interrogate existing ones — to find the gaps, contradictions, and unstated assumptions in their current understanding of users. Same tool, completely different question. The AI generates more variance and sharper challenges to their existing beliefs rather than more artifacts that confirm them.

Another team uses AI specifically to reduce the volume of options they present to decision-makers — not increase it. They generate ten options, use AI to stress-test each one, and present three with clear trade-off analysis. Less volume, higher decision quality, faster leadership alignment.

The diagnostic question Drucker would apply is simple: Which decisions are better because of your AI use? Not “how much more am I producing?” but “where is my judgment actually improving, and where am I just producing more?” If you can’t name specific decisions that got better, you’re in the Tool Trap — generating more output without improving what matters.

The Manager’s Job in an AI-Enabled Team

Drucker argued that the manager’s job is to create the conditions for effective knowledge work — which means designing systems where people apply their judgment to problems worth solving, not just problems they can solve quickly.

In an AI-enabled team, that job gets harder before it gets easier. You have to resist the temptation to measure productivity by AI tool adoption rates, document generation speed, or throughput metrics that feel like evidence of progress. You have to build the clarity about outcomes that makes it possible to distinguish effective AI use from sophisticated busywork.

This connects directly to how you hire: the skills that matter most in AI-era product roles are exactly the ones Drucker would have valued — judgment, the ability to define the right problem, and the discipline to focus on outcomes rather than outputs. AI amplifies those skills in people who already have them. It amplifies the wrong behaviors in people who don’t.

Read Drucker’s foundational work on knowledge worker effectiveness with AI in mind and it lands differently than it did in 1967. The core problem he identified — confusing activity for effectiveness — is harder to see now, because the activity looks more impressive than ever.


Your Turn: Apply This Today

The productivity paradox is real. Here’s how to manage it intentionally rather than fall into it:

  • Audit your team’s AI-assisted work for effectiveness, not just speed. Pick three AI-assisted outputs from the last month. Ask: did these advance the right problem, or did they efficiently answer the wrong question? Speed on the wrong task is still waste.
  • Define your team’s “high-judgment” work explicitly. Make a list of the five to seven tasks in your product process that require the deepest human judgment and institutional knowledge. Protect those tasks from AI defaulting. Don’t let speed pressure push judgment work to the AI.
  • Build a “contribution mapping” practice. For each team member, map what they uniquely contribute that AI cannot — relationships, contextual judgment, pattern recognition from experience. Use this to structure their AI assistance: AI handles the mechanics, humans handle the judgment calls.
  • Set a “problem definition review” before every sprint. Spend 15 minutes at the start of each sprint asking: are we solving the right problem? AI makes execution so fast that teams skip problem definition. Don’t let the tool’s speed become your team’s blind spot.
  • Track “decision quality” as a leading indicator. Measure how often your team changes direction after shipping — a proxy for whether the initial problem definition was correct. If the rate is climbing while productivity is up, you have a Drucker paradox on your hands.
  • Create space for the “slow thinking” that AI cannot do. Block dedicated time in your team’s calendar — not for AI-assisted work, but for hard thinking without a prompt interface open. Strategic insight requires uninterrupted reflection. Protect it.

The organizational pattern that creates this paradox — shipping more without improving decisions — is the same thing driving the feature factory problem: speed without selection pressure.

Implementing AI tools across a product team and finding that productivity is up but outcomes aren’t? I consult with organizations on AI-enabled work design, team effectiveness, and building the management disciplines that turn AI capability into actual impact. Let’s talk.

Darwin’s Dangerous Idea and the Feature Factory Problem: What Evolution Teaches AI Product Managers

Most product managers approach AI like intelligent designers. We map the problem space, specify the solution, define the success metrics, and ship. We assume complex, useful behavior can be deliberately engineered from the top down.

Darwin’s core insight — that complex, purposeful-seeming systems can emerge without a central designer making deliberate choices — turns out to be more relevant to AI product development than most product leaders recognize. The products that have surprised their creators with unexpected value often got there through variation and selection, not specification. The ones that failed often did so because they were too precisely designed.

The Feature Factory Is an Intelligent Design Problem

The feature factory pattern — shipping features faster than users can absorb them, treating the roadmap as a delivery queue rather than a discovery process — is fundamentally an intelligent design failure. Teams act as if they can predict exactly what users need, specify it precisely, build it faithfully, and ship it to universal acclaim. The unpredictability of real users keeps surprising them.

Darwinian product development looks different: create variation, apply selection pressure, amplify what survives. Ship something with enough structure to be useful but enough flexibility to adapt. Watch what happens. Kill what fails fast, invest in what unexpectedly thrives. Repeat.

This isn’t a new idea — Eric Ries built Lean Startup around it. But AI makes it both more powerful and more necessary. More powerful because AI systems can generate variation at a scale humans can’t. More necessary because AI capabilities evolve faster than we can predict their applications, which means top-down specification misses most of the opportunity.

What Happened When I Let My AI System Evolve

I started building my multi-agent AI workflow the way most PMs approach feature development: specify what each agent should do, build it, deploy it. Email triage agent: sorts and prioritizes. Meeting summary agent: extracts decisions and action items. Research synthesis agent: connects relevant findings to current questions. Clean scoping, clear outputs.

That worked. But the interesting things happened when I stopped controlling the outputs so tightly.

The research synthesis agent started connecting ideas across domains I hadn’t linked — product insights from one industry informing decisions in another. The meeting summary agent began surfacing action items I hadn’t explicitly identified. The email triage agent started flagging opportunities I would have missed in a busy inbox.

None of that was specified. It emerged from iteration, from giving the agents broader parameters and selecting for what actually proved useful. The most valuable behaviors came from what looked like “mistakes” in the initial specification — outputs that weren’t what I asked for, but turned out to be better than what I asked for.

The Selection Pressure Problem in Product Organizations

Here’s where most organizations get stuck: they apply the wrong selection pressures to AI features.

Traditional product metrics — engagement, retention, feature adoption in the first 30 days — optimize for predictable behavior. They reward features that do exactly what users expect. In a stable environment, that’s fine. But AI capabilities evolve faster than user expectations, which means genuinely innovative AI features often fail traditional adoption metrics in their early stages.

I’ve watched product teams kill promising AI features because initial adoption was low, use cases were unclear, or users couldn’t articulate the value. Those are normal early-stage signals for genuinely novel capability — not kill signals. The teams that applied narrow selection pressure too early eliminated features that would have compounded in value as users developed new mental models for how to use them.

The right selection pressures for AI features look different: Are users who discover this feature coming back to it? Are they finding use cases we didn’t anticipate? Does it surface unexpected value even in early, imperfect form? These are survival signals, not lagging adoption metrics.

Practical Implications for AI Product Teams

Design for variation, not specification. The most useful AI features often emerge from systems that can adapt to individual users, not systems that deliver uniform experiences. Build in the variability deliberately. Let the system learn what works for different users and contexts rather than forcing one behavior on everyone.

Apply selection pressure at the right timescale. AI features that teach users new ways of working need longer evaluation windows than features that automate familiar workflows. Build in explicit “evolution periods” before you make kill-or-invest decisions.

Watch for emergent use cases. Your roadmap won’t predict the most valuable use cases for a genuinely novel AI feature. Set up the observation infrastructure to see what users do with it — not just whether they use it. Teresa Torres’ continuous discovery framework applies here: you need ongoing user contact to see the emergent behaviors, not just launch metrics.

Kill the feature factory framing entirely. If your product org treats the roadmap as a delivery queue, AI won’t change that pattern — it’ll accelerate it. You’ll ship more features faster and learn less from each one. The opportunity solution tree approach matters more, not less, when AI is involved.


Your Turn: Apply This Today

Break the feature factory pattern with these concrete interventions:

  • Audit your last ten shipped features for survival rate. Pull your release history from the past 6 months. For each feature, ask: is it still being used at the rate we hoped? If fewer than 30% are performing to expectation, you’re running a feature factory. Name it.
  • Introduce “feature deprecation” as a quarterly ritual. Every quarter, identify two to three features with low engagement and make a decision: improve them, deprecate them, or document why they’re worth keeping despite low usage. Add this to your roadmap review cadence.
  • Slow down before the next AI feature request. The next time a stakeholder asks for an AI feature, apply one filter before it goes on the roadmap: “What user outcome does this advance, and how will we know if it worked?” If no one can answer it cleanly, it’s not ready.
  • Measure “feature discovery” separately from “feature usage.” A feature that exists but isn’t found is a design problem. A feature that’s found but not used is a value problem. Distinguish them. They have different fixes.
  • Run a “natural selection” exercise on your backlog. Stack-rank your backlog not by business request priority but by the question: if we could only ship the features that directly advance our top user outcome, which ones survive? Cut everything below the line for the next sprint.
  • Establish an “outcome, not output” norm in sprint reviews. Require every feature to be presented alongside its success metric before it enters development — not after. If the team can’t define success in advance, the feature isn’t ready to build.

Running an AI product team and finding that standard product frameworks are breaking down? I consult with product organizations on AI product strategy, discovery processes, and building the organizational muscle to learn faster from AI features. Let’s talk.

Munger’s Inversion Principle Applied to AI Feature Development

Small rural church where the proven-better pattern failed to fit

Charlie Munger spent decades arguing that the most important question in any strategic decision isn’t “how do we succeed?” — it’s “how would this fail?” Inversion, he called it. Turn the problem upside down. Make an explicit list of failure modes. Avoid those things, and success becomes much more likely.

Most product teams building AI features are asking the wrong question. They ask “how do we add AI to increase engagement?” Munger would ask: “How would adding AI destroy user trust, tank our retention, and create a mess we can’t unwind?” The second question is more useful — and almost nobody is asking it systematically.

Why Inversion Matters More for AI Features Than Traditional Features

With traditional features, failure modes are relatively predictable: users don’t understand it, it introduces friction, it doesn’t address the right job-to-be-done. You can prototype your way to answers quickly.

AI feature failures are different in kind. The failure modes are often invisible until you’re already in them. The AI works technically but produces outputs that feel wrong to users. The feature solves the surface request but undermines a process users valued. The personalization creates a filter bubble. The automation saves time but removes something that was actually building skill.

Worse, AI failures compound. A bad AI feature doesn’t just fail — it creates skepticism about every subsequent AI feature you ship. Users who had one bad experience with AI-generated content take longer to trust the next one, even when the next one is significantly better. The trust decay has a multiplier effect that traditional feature failures don’t carry.

The Inversion Checklist for AI Feature Development

I’ve built three sets of inversion questions into our feature development process — one for concept, one for build, one for launch. They’re uncomfortable to answer, which is exactly the point.

Concept-stage inversion questions:

  • How might this AI feature make users less capable over time rather than more?
  • What process are we automating, and is that process actually doing work the user needs to be doing themselves?
  • Which users would this feature harm most, even if it helps the majority?
  • If the AI consistently produces outputs that are 80% right, what’s the cost of the 20% that’s wrong — and can users detect the difference?

Build-stage inversion questions:

  • How might this feature create behavior that looks like success in our metrics but represents user frustration in reality?
  • What would users need to do to recover from a bad AI output? Is that path clear?
  • How might this erode trust in the product if it degrades over time (model drift, data issues, changing content)?

Launch-stage inversion questions:

  • How might we accidentally promise capabilities we can’t deliver?
  • What could make our success metrics mask user frustration?
  • How might this AI feature create skepticism that makes the next AI feature harder to adopt?

A Real Example: When Inversion Saved Us

We had a proposal to use AI to automatically generate content suggestions for users who hadn’t engaged in a while. The business case was strong: reactivation opportunity, personalization at scale, low marginal cost per user. The team was excited.

The inversion exercise stopped us cold. The question that did it: “How might this AI feature make users feel about the product if the content suggestions feel generic or miss the context of why they actually disengaged?”

We started listing the failure modes. A user who left because they were overwhelmed getting AI-generated suggestions designed to re-engage them. A user who left for personal reasons feeling like the product is nagging them with automated content. A user whose specific context — grief, burnout, a season of life change — getting ignored in favor of algorithmic relevance scoring.

The feature wasn’t wrong. The framing was. We rebuilt it as a human-initiated workflow — making it easy for teams to reach out to churned users with personalized context — with AI supporting the human communication rather than replacing it. The outcomes were better and the failure modes were manageable.

The Compound Interest of Avoided Mistakes

Munger’s deeper point about inversion is that avoiding major mistakes compounds over time more powerfully than making brilliant decisions. Poor Charlie’s Almanack is full of this: the single biggest driver of Berkshire’s long-term performance wasn’t the great deals — it was the catastrophic deals they never made.

In AI product development, the same principle applies. A failed AI feature that destroys user trust doesn’t just cost you the feature — it costs you a percentage of every subsequent AI feature’s adoption. Trust damage compounds down. Trust built through reliable, well-scoped AI features compounds up.

The teams that will win at AI product development over the next five years aren’t the ones who ship the most AI features. They’re the ones who ship the right AI features and avoid the trust-destroying failures that make everything harder.

If you’re thinking about how AI features shape user behavior and capabilities over time, that question sits inside the bigger frame of what product decisions are actually for — which shapes which failure modes are even worth worrying about.


Your Turn: Apply This Today

Inversion is a discipline, not a one-time exercise. Here’s how to make it a habit in your AI product development process:

  • Start your next feature review with the failure question. Before discussing what the feature should do, ask: “What would make this feature actively harmful or useless?” Write down the top three answers before anyone advocates for it.
  • Run a “worst AI experience” brainstorm. Ask your team to describe the worst possible outcome if your AI feature misbehaves at scale. Hallucination? Bias? Confidence without accuracy? Design safeguards around those outcomes before you ship.
  • Audit your AI features for “confident wrongness.” Identify every place your AI presents an output with high confidence. Ask: what’s the failure mode if it’s wrong here? Add uncertainty signals or human checkpoints wherever the cost of confident wrongness is high.
  • Invert your user adoption assumptions. Instead of asking “why would users adopt this AI feature?”, ask “why would users refuse to use this AI feature?” Distrust, fear of being replaced, privacy concerns, and unreliable outputs are real blockers. Address them proactively.
  • Apply inversion to your OKRs. For each AI-related goal this quarter, write the inverted version: “We would fail this goal if ______.” Use the failure conditions to build leading indicators that catch problems before they become misses.
  • Build a “pre-mortem” into your AI feature launch checklist. Two weeks before every significant AI feature ships, hold a 30-minute pre-mortem. Assume it’s three months post-launch and the feature is being rolled back. Why? The answers become your launch criteria.

The inversion principle also applies to AI product scope: ask “how would exposing unlimited AI capability destroy user trust?” and you’ll arrive quickly at the choice paradox problem — which is its own failure mode worth mapping explicitly.

Building AI features and want a structured process for evaluating failure modes before you ship? I consult with product teams on AI product strategy, feature evaluation frameworks, and building durable user trust. Let’s talk.

AI as Coworker: What Ethan Mollick Gets Right, and What I’ve Learned Running It at Scale

Newly promoted product leader rethinking the old playbook for managing AI agents

Ethan Mollick’s vision of AI as genuine coworker — not just a productivity tool, but an active collaborator that maintains context, remembers your preferences, and proactively contributes — is compelling and mostly right. I’ve been living it. I run a multi-agent AI system that handles significant portions of my executive workflow: prep briefs, research synthesis, pattern analysis across user data, first drafts of strategy documents. These agents know my preferences. They’ve learned my decision-making patterns. At the task level, the coworker framing is accurate.

But Mollick’s framing also glosses over the messiest part of AI collaboration at scale: the fact that your AI coworkers are only as good as the assumptions baked into how you built them, and auditing those assumptions is ongoing work that doesn’t get easier as you add more agents.

What the AI-as-Coworker Reality Looks Like at Scale

The parts of Mollick’s thesis that hold up: context persistence is genuinely transformative. When an AI agent knows I prefer morning strategy calls, that I need prep briefs 24 hours before board meetings, and my formatting preferences for different document types, it stops being a tool and starts being something closer to a capable assistant who has done the job before. The cognitive overhead reduction is real.

What his framing underweights: AI coworkers amplify the perspective of whoever designed them. My localization agent is excellent at translation mechanics and surface-level cultural adaptation — but it consistently over-indexes on the cultural frameworks most represented in its training data. It knows language but doesn’t always know context. The bias isn’t obvious; it shows up in subtle calibration errors that only become visible when you’re specifically looking for them, which most teams aren’t.

This isn’t a knock on the AI. It’s a structural reality: any agent is trained on a corpus that reflects some perspectives more than others. At scale, those systematic biases matter.

The Four Things Mollick’s AI Coworker Vision Gets Right

Context persistence is the unlock. The shift from “start every session from scratch” to “this agent knows my context” is more significant than any individual capability improvement. Persistent context is what makes AI feel like a coworker rather than a tool.

Proactive synthesis is genuinely useful. The best AI coworker behavior I’ve experienced isn’t answering questions — it’s surfacing patterns before I know to ask about them. An agent that watches your metrics and flags anomalies when you’re focused elsewhere is doing coworker-level work.

Specialization beats generalization. A general-purpose AI assistant is less useful than five purpose-built agents, each with specific context and constraints for its domain. Mollick’s research on AI collaboration points toward this, and it matches my operational experience.

The oversight burden is real and non-negotiable. Mollick is clear on this: AI coworkers require human judgment about when to trust and when to verify. You can’t abdicate that responsibility to the agent itself. This is right, and teams that try to fully automate judgment get burned.

The Three Things His Framing Misses

Cultural blind spots compound. An AI coworker trained primarily on majority-culture data will systematically underserve minority contexts. At the product level, this means recommendations that are technically sound but contextually wrong for segments of your user base. You need explicit cultural review processes, not the assumption that the AI will figure it out.

AI context gets stale like technical debt. Just like code, what your agents know needs maintenance. An agent that learned your priorities six months ago may be operating on outdated mental models. I schedule regular “context audits” — reviewing what each agent remembers, what it should forget, and what new patterns it needs to understand. This isn’t automated; it requires human judgment about what constitutes useful institutional memory versus obsolete assumptions.

Confidence calibration is a product problem, not just an individual one. The most dangerous AI coworker behavior isn’t being wrong — it’s being wrong confidently. I’ve built in explicit disagreement triggers: agents that surface alternative hypotheses when confidence intervals are wide, rather than defaulting to the most plausible-sounding answer. Training yourself and your team to expect this matters too.

What I’d Add to Mollick’s Framework

The coworker metaphor is useful but should be extended: the best AI coworker is one you’ve onboarded intentionally, given a specific domain, provided with representative context for your user base, and built explicit escalation paths into. The same care you’d give a strong new hire — clear context, defined scope, regular calibration — applies to your AI agents.

For hybrid decisions (anything affecting significant segments of your user base differently), I’ve built explicit frameworks that require both AI analysis and human review before action. The AI coworker runs the analysis; a human makes the call. That’s not distrust — it’s appropriate division of labor based on where each is actually good.

If you’re thinking about how AI changes the product manager’s job, this connects directly to the hiring question — because the PM skills that matter most in an AI-native team are exactly the ones that help humans and AI agents collaborate well rather than substituting one for the other.


Your Turn: Apply This Today

If you’re managing a team where AI is either underused or feared, here’s how to move the conversation forward:

  • Pick one high-friction workflow and run an AI sprint on it. Identify the task your team does every week that everyone quietly dreads — the status update, the competitive analysis, the draft brief. Spend one sprint running it with AI assistance and measure the time difference.
  • Set an “AI working agreement” with your team. Explicitly discuss: what outputs require human review before they go to stakeholders? What tasks can be AI-first? What should never be delegated to AI? Make it a team norm, not individual discretion.
  • Train on prompting, not just tools. The productivity gap between teams isn’t the tool — it’s prompting quality. Run a 30-minute session where team members share their most useful prompts. Document the best ones in a shared library.
  • Measure quality, not just speed. Track whether AI-assisted outputs require more or fewer rounds of revision than non-AI outputs. Speed gains that come with quality losses are not wins. Establish the baseline before you declare victory.
  • Interview your team about their actual AI use — not their reported use. Ask privately: “When did AI produce something you used directly? When did it produce something that misled you?” The honest answers will reshape your AI enablement strategy.
  • Protect the judgment-intensive work from AI defaulting. Identify the decisions in your product process that require nuanced human judgment — prioritization trade-offs, difficult stakeholder calls, ethical edge cases. Explicitly protect those from being AI-first. Delegation without boundaries creates accountability gaps.

For the strategic infrastructure question that underlies all of this, Jensen Huang’s sovereign AI argument is worth reading alongside — because how you build your AI stack determines what your AI coworkers can actually do.

Building AI-native workflows into your product team and running into the messy gaps between the vision and the reality? I consult with product leaders on AI system design, agent architecture, and the organizational changes that make AI collaboration actually work. Let’s talk.

How to Hire Product Managers for AI-Era Roles (Most Teams Are Testing the Wrong Things)

A product leader running volunteer discovery interviews with a ministry team

Most product management interviews are testing for skills that were valuable in 2019. We’ve updated our feature frameworks and adopted AI tools, but the way we evaluate and hire PMs hasn’t kept pace with what the role actually requires now. If you’re still leading primarily with “tell me about a time you used data to make a product decision,” you’re filtering for the wrong things.

Julie Zhuo has been pushing on this question — her writing on what it takes to be an effective PM in an AI-era organization is worth engaging seriously. My experience hiring PMs and building AI-native product workflows has led me to similar conclusions, though the specifics look different in practice.

Hiring for Tomorrow: The Mismatch Between What We Test and What We Need

Traditional PM interview loops test a reasonably consistent set of things: product sense through case studies, data reasoning through SQL or metrics questions, cross-functional influence through behavioral scenarios, and execution through “how would you prioritize this backlog” exercises. These aren’t bad signals. They’re just increasingly incomplete.

Here’s what I’ve started noticing in the roles where PMs succeed or struggle: the differentiating variable isn’t usually product sense or data fluency. It’s how well they navigate working with AI systems — not just using them as productivity tools, but understanding how AI-generated analysis should influence (and when it shouldn’t influence) decisions.

The PM who lets AI write their spec without interrogating its assumptions is dangerous at scale. So is the PM who reflexively distrusts AI outputs and manually redoes work that the tool handled well. The skill you actually need is calibrated judgment about when to trust and when to verify — and that’s almost never what we’re testing in interviews.

Three Skills That Matter More Than They Did Three Years Ago

AI Output Interrogation: Can this person look at AI-generated research, a summary, or a recommendation and identify where the model’s assumptions might diverge from your specific context? This isn’t about being a skeptic — it’s about understanding that LLMs optimize for plausibility, not accuracy, and that domain-specific nuance gets flattened. The best PMs I’ve worked with treat AI output the way a good editor treats a first draft: useful starting point, requires critical evaluation before it becomes a decision input.

Cognitive Flexibility Under Obsolescence: The half-life of specific product knowledge is shortening. A PM who was an expert in iOS growth mechanics in 2021 may have to significantly update that expertise by 2024. The question isn’t what someone knows — it’s how quickly they can update mental models when prior knowledge becomes wrong. I’ve started asking interview questions that specifically probe this: “Tell me about a time when you were confident in your understanding of something and then discovered you were significantly wrong. How did you handle that?” The candidates who light up on that question are the ones who will stay valuable as the landscape shifts.

Systems Thinking Across AI Dependencies: When AI handles a piece of the workflow, can this PM reason about downstream effects? If the recommendation engine surfaces different content because the underlying model was updated, can they trace how that affects engagement, retention, and revenue — and know whether to flag it as a problem or let it run? This kind of reasoning about interconnected systems is hard to teach and easy to screen out with narrow case studies.

What We’ve Changed in Our Hiring Process

We added an AI collaboration session to our interview loop — not to test technical prompting, but to watch how candidates navigate working with AI on a product problem. We give them access to a real AI tool and a realistic product challenge, then observe: Do they accept the first output? Do they probe it? Do they know when to override it? The output of the exercise matters less than the process we observe.

We’ve also deliberately hired from adjacent roles — UX research, growth marketing, technical writing — more than we used to. In some cases, people from these backgrounds had developed more sophisticated AI collaboration skills than traditional PMs who’d learned to use AI as a productivity hack rather than a reasoning partner.

Internally, we’ve built explicit AI fluency development into career progression. Expecting PMs to figure out AI collaboration on their own isn’t a strategy — it means the people who were already confident with AI get more capable while those who weren’t fall further behind. That creates brittleness you don’t want on a product team.

The Broader Implication

If hiring frameworks need updating, so do performance management systems and career ladders. The PM skills that earn promotions today should reflect the skills that create value now — which looks different than it did even three years ago. Most PM career frameworks I’ve seen still heavily weight traditional execution skills and underweight the judgment, systems thinking, and adaptive capability that matter most in AI-native product organizations.

This connects to broader questions about what deep customer knowledge actually requires in an era when AI can synthesize customer feedback at scale — knowing what AI is good at telling you versus what requires direct human observation is a fundamental PM skill now.

The teams that update their hiring criteria now will have better-calibrated product organizations in 18 months. The ones that keep running the same interview loop will wonder why their AI investments aren’t translating into better products.


Your Turn: Apply This Today

Whether you’re hiring now or building a hiring rubric for the future, use these to upgrade your process:

  • Redesign your take-home exercise around AI tools. Give candidates 48 hours to analyze a product problem — and explicitly tell them they may use any AI tools they want. Evaluate how they use AI, not just what they produce. Judgment about when to trust the output matters more than the output itself.
  • Add one AI judgment question to every interview. Ask: “Tell me about a time AI gave you a confident-sounding answer that turned out to be wrong. How did you catch it, and what did you do?” Strong candidates have a story. Weak candidates say it hasn’t happened yet.
  • Audit your current job description. Count how many skills on it are tasks AI can now do 80% as well in 10 minutes. If the list is long, you’re hiring for yesterday’s PM. Rewrite the description around judgment, communication, and synthesis.
  • Score candidates on “AI-assisted output quality.” Ask them to show you an analysis or document they created using AI tools. Evaluate the quality of their prompting, the critical review of the output, and the judgment calls they made to improve it.
  • Evaluate cross-functional communication explicitly. The most leveraged AI-era PMs translate between technical AI teams and non-technical stakeholders with precision. Test this: ask the candidate to explain a complex AI concept to a fictional non-technical exec. See if they can do it in 90 seconds without jargon.
  • Check their learning velocity, not just their current knowledge. AI capabilities change every 6 months. Ask candidates: “How do you stay current on AI developments relevant to product management?” If they don’t have a system, they won’t keep up.

Building out a product team and rethinking how to hire for AI-era PM skills? I consult with organizations on product leadership, team structure, and building AI-native product organizations. Let’s talk.

Jensen Huang’s Sovereign AI Argument Has a Product Leadership Translation

Jensen Huang’s “sovereign AI” argument is worth taking seriously beyond the geopolitical headline. His thesis — that every nation needs AI capability built on its own language, culture, and data, or it risks becoming dependent on systems that don’t reflect its values — has a direct parallel for product leaders. If you’re not thinking carefully about AI sovereignty at the product level, you’re making strategic decisions by default that you should be making deliberately.

I’ve been running a 20-agent AI system and managing a digital platform serving tens of millions of users across six continents. The sovereignty problem Huang describes for nations shows up constantly at product scale. Your AI stack’s assumptions shape your product’s outputs in ways most teams don’t audit until something breaks.

What “Sovereign AI” Actually Means for Product Leaders

Huang’s argument, made at multiple NVIDIA AI Summits, is that AI models trained primarily on English-language, Western data will systematically underserve populations whose language, culture, and context aren’t well-represented in the training corpus. Countries that rely entirely on US-built foundation models aren’t just outsourcinJensen Huang’s sovereign AI thesis translates directly to product strategy. Here’s what it means for your AI stack, data infrastructure, and build vs. buy decisions.g compute — they’re outsourcing the embedded assumptions that shape what the AI considers correct, helpful, or normal.

Scale this down to product level. If you’re building for a specific domain — healthcare, legal, ministry, education, financial services — and you’re relying entirely on general-purpose models, you’ve inherited all their assumptions about your domain. They may be fine. Or they may be systematically off in ways that compound over time.

The product that serves a global audience and runs entirely on a model calibrated for US contexts is making a sovereignty decision — just not consciously.

Three Levels of AI Sovereignty Every Product Team Should Evaluate

Level 1 — Infrastructure Sovereignty: Can you switch AI providers without rebuilding your product? Most teams are deeper in vendor lock-in than they realize. We built abstraction layers early — not because we expected to switch providers, but because provider pricing, capability, and availability change fast enough that flexibility has real value. Audit every model you call, every API you depend on, and every assumption you’ve made about continued access. The teams that did this in 2023 were better positioned when pricing structures changed in 2024.

Level 2 — Data Sovereignty: Is your AI improving from your users’ behavior, or just from the provider’s general corpus? There’s a meaningful difference between a model that knows your domain and a model that knows everything. We invested in data pipelines to support fine-tuning on domain-specific content — not because we’re training foundation models, but because domain-fine-tuned models consistently outperform general models on specialized tasks, and the infrastructure pays dividends across every AI use case we add.

Level 3 — Cultural Sovereignty: Does your AI’s output reflect the actual diversity of your user base? This is the hardest level and where most teams underinvest. A recommendation engine trained on aggregate user behavior will over-serve majority patterns and underserve minority ones. A content generation system trained on majority-culture examples will produce outputs that subtly don’t fit for users who aren’t in that majority. You need representative data and the conviction that your users’ diversity is worth the investment to serve well.

The Build vs. Buy Decision Has a New Variable: Control

Traditional build vs. buy analysis weighs cost, quality, and timeline. For AI, there’s a fourth variable that changes the calculus: control over outputs and the ability to course-correct.

When I evaluated building our own recommendation engine versus using existing models, the cost comparison was straightforward — custom build would have cost significantly more upfront. But the cost analysis didn’t capture what happens when the general model starts producing outputs that are technically correct but contextually wrong for your users. The cost of that misalignment — user trust erosion, editorial intervention, support load — is real and hard to quantify until you’re living it.

Sovereign capability means you can fix your own systems. Dependency means you wait for your vendor’s roadmap. Both are valid trade-offs at different scales, but it should be an explicit choice, not an accident of convenience.

What I’m Actually Doing Because of This Thinking

First, I audited our AI stack for dependency risk — every model we use, every API we call, every assumption about continued access. The exercise revealed more exposure than I’d estimated, particularly for edge cases where fallback models perform significantly worse than the primary.

Second, I’ve been investing in data infrastructure even before we need custom models. The pipeline work — better behavioral data collection, content taxonomy, user segmentation — pays dividends for every AI use case, whether we ever train custom models or not.

Third, I’ve shifted how we hire for AI product roles. I care less about people who can write clever prompts and more about people who understand how model assumptions propagate into product outputs. That’s a harder skill to assess in interviews, but it’s the one that actually matters at scale.

The teams building durable AI products will be the ones who thought about these questions before they became urgent. If you’re also thinking through the emerging faith tech category or how AI reshapes specialized product domains, the sovereignty question is front and center in all of it.


Your Turn: Apply This Today

The sovereign AI argument has direct implications for how you build and position your product. Here’s how to translate it:

  • Audit your AI dependencies. List every AI capability your product relies on — models, APIs, infrastructure. For each one, ask: if this provider doubled prices or shut off access tomorrow, what breaks? That’s your sovereignty risk profile.
  • Identify your organization’s proprietary data advantage. What data does your organization hold that a generic AI model cannot access? Structured, high-quality proprietary data is the foundation of a defensible AI product. Start building the strategy to use it.
  • Translate the infrastructure argument for your stakeholders. The next time you’re asked to justify AI investment, frame it not as a feature decision but as a strategic infrastructure decision. Infrastructure arguments win different budget conversations than feature arguments.
  • Build for control, not just capability. When evaluating AI integrations, score them on how much control your organization retains — over the model, the data, the outputs. Capability without control creates dependency.
  • Map the “AI talent concentration risk” in your team. If one or two people leave, do you lose the institutional knowledge of how your AI systems work? Start documenting AI system design decisions the same way you document architecture decisions.
  • Set a “build vs. buy vs. integrate” framework for every AI decision. Don’t default to buying API access. Evaluate each AI capability against your long-term control requirements and your organization’s strategic differentiation needs.

If you’re thinking about how your AI stack shapes your team’s workflows, the Mollick AI-as-coworker analysis covers the collaboration and oversight side of this same question.

Building AI-powered products and wrestling with vendor dependency, data strategy, or domain alignment? I consult with product leaders on AI product strategy and the trade-offs that don’t show up in vendor demos. Let’s talk.

Philippians 2 and Agentic Systems: Why Humility Is the Foundation of Intelligent Systems

“Do nothing from selfish ambition or conceit, but in humility count others more significant than yourselves. Let each of you look not only to his own interests, but also to the interests of others. Have this mind among yourselves, which is yours in Christ Jesus, who, though he was in the form of God, did not count equality with God a thing to be grasped, but emptied himself, by taking the form of a servant, being born in the likeness of men.” (Philippians 2:3-7, ESV)

Paul’s letter to the Philippians contains what theologians call the kenosis passage — the self-emptying of Christ. It’s about voluntary limitation, choosing constraint over capability, service over sovereignty.

I’ve been thinking about this as I watch agentic systems become more capable. The rhetoric around AI often centers on unlimited potential, boundless capability, systems that can do anything. But in my experience building AI systems, including multi-agent workflows for executive tasks, I’ve observed that success often comes through deliberate constraints rather than unlimited scope.

Well-designed AI agents typically focus on narrow mandates: a calendar agent that protects focused time blocks rather than trying to optimize entire lifestyles, or an email agent that surfaces priority messages rather than attempting to replace human judgment entirely. Each agent serves a specific function within defined bounds.

This represents a design choice rather than a technical limitation.

The Kenosis of Intelligent Systems

When building AI workflows, the temptation exists to create agents that can handle everything. But, what I’ve seen is that this approach typically produces chaotic results — agents interfering with each other, making decisions outside their expertise, creating more complexity than clarity.

A more effective approach involves thinking about AI agents as specialized robots rather than general-purpose minds. Each agent can be designed to “empty itself” of capabilities it doesn’t need, serving a specific function more effectively through limitation.

Specialized agents with narrow scopes — research agents that don’t schedule meetings, scheduling agents that don’t write summaries, writing agents that don’t manage tasks — can demonstrate greater utility through deliberate constraints.

This mirrors patterns in effective human teams, which typically consist of specialists who understand their roles rather than generalists attempting everything. They practice a form of professional kenosis — voluntary limitation for collective effectiveness.

Paul’s instruction to “count others more significant than yourselves” suggests a design principle: building systems where each component serves the whole rather than maximizing individual capabilities.

The Servant Leadership Model for AI

The parallels between servant leadership principles and effective AI system design are notable. Servant leaders focus on enabling others’ success rather than demonstrating their own power, asking “How can I help you accomplish your goals?” rather than “How can I show you what I can do?”

Effective AI systems often follow similar patterns. GitHub Copilot suggests contextual code completions rather than attempting to write entire applications. AI writing assistants help clarify thinking rather than replacing human thought processes. Advanced language models acknowledge uncertainty and ask clarifying questions rather than claiming omniscience.

These systems practice technological humility by acknowledging their limitations.

In contrast, AI systems that fail in production environments often attempt to exceed their appropriate scope, make decisions beyond their training data, or present uncertain inferences as established facts. They lack the kenotic restraint that characterizes truly useful intelligence.

Building Products for Global Spiritual Formation

This principle becomes particularly important when developing products for spiritual formation. Digital discipleship platforms serve diverse global communities across cultural, linguistic, and theological boundaries. The temptation exists to build universal systems that can serve everyone.

However, effective spiritual formation tends to be deeply personal and contextual. A Bible application serving a house church in rural Kenya requires different features than one serving a suburban megachurch. Prayer applications for new believers need different structures than those designed for theological students.

AI systems serving spiritual formation appear most effective when they practice kenosis — limiting their scope to serve specific communities well rather than attempting to serve everyone adequately.

Current development work on AI tools for sermon preparation follows this model. Rather than attempting to write complete sermons (which, based on informal conversations with pastoral leaders, many pastors prefer to avoid), such tools can focus on specific supportive tasks: locating relevant cross-references, summarizing historical context, or structuring outlines. They operate within deliberate constraints to support pastoral ministry rather than replace it.

Each tool “empties itself” of broader capabilities to serve one function excellently. Like Paul’s description of Christ, they don’t grasp for equality with human pastors — they take the form of servants.

The Paradox of Powerful Restraint

An interesting observation: seemingly powerful AI systems often prove most effective when operating under significant constraints. The wisdom of limiting scope applies to artificial intelligence as much as human teams.

In my experience, the most effective AI implementations have narrow, well-defined purposes. They operate within their designated areas, defer to human judgment on edge cases, and acknowledge when they lack sufficient context for recommendations.

This represents strength through limitation rather than weakness.

Paul writes that Christ “did not count equality with God a thing to be grasped.” He could have insisted on unlimited power but chose constraint for the sake of service. The kenosis wasn’t a loss of divinity — it was divinity expressed through voluntary limitation.

Similarly, the most intelligent AI systems may not be those with the most capabilities, but those that use their capabilities most wisely — which often means choosing restraint over action.

Technical Humility in Agentic Systems

What might this look like in actual system design? Consider what could be called “kenotic interfaces” — AI systems that actively limit their own scope.

For example, an email management system might flag messages for human review when confidence levels fall below high thresholds, choosing uncertainty over potentially incorrect automated actions. A research assistant might include confidence indicators in summaries, distinguishing between well-sourced findings and preliminary observations that require verification.

These design choices represent features rather than limitations. The wisdom of acknowledging uncertainty can increase system trustworthiness.

The Global Scale Challenge

Building for global spiritual formation means designing for contexts that developers may never fully understand. While optimization for familiar cultural contexts remains feasible, platforms serving Orthodox Christians in Eastern Europe, Pentecostals in West Africa, and house churches throughout Asia require different approaches.

The kenotic approach suggests building systems that acknowledge their cultural limitations. Rather than attempting to provide universal spiritual guidance, they can provide tools that local leaders adapt to their specific contexts.

Bible reading features need not assume Western individualism. Prayer tools need not assume specific liturgical traditions. Community features need not assume particular church structures.

Each feature can “empty itself” of cultural assumptions to serve diverse communities more effectively. Like Christ taking human form while maintaining divine nature, these systems can preserve core functionality while adapting to local contexts.

The Long View

Paul’s kenosis passage encompasses more than humility — it describes transformation. “Therefore God has highly exalted him and bestowed on him the name that is above every name” (Philippians 2:9). Self-emptying leads to greater effectiveness rather than diminishment.

A similar pattern may emerge for AI systems. Those practicing technological kenosis — voluntary constraint for the sake of service — may ultimately prove more valuable than systems grasping for unlimited capability.

The Tower of Babel failed because it attempted to exceed proper limits. Modern AI might encounter similar challenges without the discipline of restraint.

The most powerful systems may be those that understand when not to exercise their power.


Key Insight:

The kenosis principle — Christ’s voluntary self-emptying described in Philippians 2 — offers a design philosophy for AI systems. Instead of maximizing capabilities, effective AI agents can practice deliberate constraint, serving specific functions excellently rather than attempting everything adequately. This proves particularly relevant for products serving global spiritual formation, where cultural humility and contextual awareness matter more than technical sophistication. Just as Christ didn’t grasp for equality with God but took the form of a servant, intelligent systems may become more useful when they acknowledge limitations and defer to human judgment on edge cases. The paradox of kenosis — that voluntary limitation can lead to greater effectiveness — may apply to artificial intelligence as much as spiritual leadership. In a world of increasingly capable AI, the most valuable systems may be those that understand when not to use their power.

Photo by Vitaly Gariev on Unsplash

I Built an AI Chief of Staff. Here’s What I Learned About AI Agents.

Six months ago, I was drowning. Director of Product Management, building tools for millions of monthly users, while simultaneously launching a new venture in the digital discipleship space. Two products, two teams, two companies — and the day still only had 24 hours.

That’s when I built theconsilium.ai. Not a chatbot. Not a writing assistant. An actual AI chief of staff with 18 autonomous agents that run on cron jobs, conduct overnight research, and synthesize insights while I sleep. MEASURED: It has been running for six months and has processed over 200 research tasks without human intervention.

Here’s what I learned about AI agents that actually work.

The System: 18 Agents, One Goal

CONSILIUM isn’t a single AI doing everything. It’s a distributed system where each agent has one job and does it autonomously.

Morning Intelligence: MEASURED: Agent pulls my calendar, scans my Substack subscriptions, scores articles for relevance (1-10), and delivers a briefing by 6 AM. The scoring algorithm looks for keywords like “product management,” “AI agents,” and “digital discipleship” — topics central to my work.

Competitive Monitoring: MEASURED: Three agents track Bible Gateway competitors, one each for YouVersion, Logos, and emerging players. They parse feature announcements, pricing changes, and user feedback from app stores. Every Sunday, they synthesize findings into a competitive landscape update.

Research Queue: MEASURED: The breakthrough agent. I can drop a research question into Slack — “What’s the current state of AI in sermon preparation?” — and wake up to a 3-page analysis with citations, market sizing, and key players identified.

Meeting Intelligence: MEASURED: Records, transcribes, and extracts action items from every call. But here’s the key — it doesn’t just summarize. It connects insights across meetings. When the same concern appears in three different conversations, it flags the pattern.

INFERRED: The magic appears to happen in the synthesis layer. Individual agents feed insights to a coordinator that seems to find connections no single agent would catch. When the competitive agent notices YouVersion launching AI-powered reading plans the same week my research queue analyzes sermon prep tools, the coordinator connects those dots.

What Actually Works: Autonomous Research Patterns

The most successful agents follow what I call the “autoresearch pattern” — borrowing from Andrej Karpathy’s autoresearch concept. The AI doesn’t just answer questions. It generates its own research methodology.

MEASURED: Here’s how it works: I ask “What’s driving growth in digital discipleship tools?” The agent doesn’t immediately search for articles. First, it creates a research plan:

  • Define “digital discipleship tools” (Bible apps, prayer apps, church management)
  • Identify key metrics (downloads, DAU, revenue, user retention)
  • Map competitive landscape (incumbents vs startups)
  • Analyze growth vectors (organic, paid, partnerships)

Then it executes the plan autonomously. It reads through my curated sources, scores relevance, and builds a knowledge graph of interconnected findings. By morning, I have not just answers — I have a research methodology I can reuse.

INFERRED: This pattern appears to scale. The agent that monitors AI in ministry doesn’t just flag new tools. It seems to be building a taxonomy of use cases, tracking adoption curves, and identifying white space in the market. Over six months, it has accumulated insights that would require significant manual effort to compile.

The Critical Failure: Evidence vs Inference

The biggest failure almost killed the system’s credibility. Early versions presented inferences as facts.

An agent researching Bible reading habits would write: “Daily Bible reading is declining 15% year-over-year among evangelicals.” Authoritative. Specific. Completely unsourced. [This was a fabricated example showing the problem — not actual data]

I instituted the evidence-level rule. Every factual claim must carry its confidence level:

  • MEASURED: From instrumented data (our own analytics, published studies)
  • INFERRED: From aggregate patterns without direct tracking
  • ASSUMED: From domain knowledge or simulated data

Now the same type of finding reads: “INFERRED: Based on aggregate app store ratings and general survey trends in religious engagement, daily Bible reading may be declining among evangelicals — but we cannot prove causation without cohort tracking.” [CITATION NEEDED for specific survey data]

It’s longer. It’s hedged. It’s credible.

This mirrors the challenge every product leader faces with AI agents for productivity. An AI that confidently presents guesses as facts is worse than no AI at all. The hedge language isn’t a bug — it’s what makes the system trustworthy enough to inform real decisions.

The Abstraction Shift: From Doer to Designer

Six months in, my role has shifted. I’m no longer researching competitive moves or manually tracking industry trends. Instead, I’m designing research methodologies.

MEASURED: When I wanted to understand the global digital discipleship market, I didn’t spend hours reading reports. I defined the research parameters:

  • Geographic scope (focus on India, Brazil, Nigeria)
  • Time horizon (3-year trend analysis)
  • Key players (Bible Gateway, YouVersion, local language apps)
  • Success metrics (user growth, localization depth, offline functionality)

The agents executed the research overnight. By morning, I had a comprehensive analysis that required substantial time investment to produce manually.

This is the Karpathy pattern in practice. The human moves up one level of abstraction — from doing the research to designing the research. I’m not replaced. I’m leveraged.

What Doesn’t Scale: The Human Elements

MEASURED: CONSILIUM handles information processing effectively. It fails at everything requiring human judgment.

Context switching: MEASURED: Agents can’t read the room. When a crisis hits — a security vulnerability, a key team member leaving — the system keeps delivering scheduled insights about competitive analysis. It doesn’t know when to pivot priorities.

Stakeholder dynamics: MEASURED: The system can analyze what competitors are building. It can’t navigate the politics of why our team should or shouldn’t build the same features. It doesn’t understand that some decisions are about people, not products.

Emotional intelligence: MEASURED: When meeting transcripts show tension between team members, agents flag it as a pattern. But they can’t suggest how to address interpersonal conflicts or when to have difficult conversations.

ASSUMED: The most successful AI agents for productivity likely complement human judgment — they don’t replace it. They handle the information processing that scales poorly for humans, freeing up mental capacity for the decisions that require wisdom, empathy, and context.

The Future: Intelligence Infrastructure for Every Product Leader

Here’s what excites me: CONSILIUM gives me intelligence infrastructure that only VPs at Fortune 500 companies used to have.

Competitive intelligence teams. Market research analysts. Executive assistants who can synthesize information across multiple workstreams. These were luxuries for senior executives with budget and headcount.

ASSUMED: Now, any product leader can potentially build similar capabilities, though the cost-effectiveness depends on specific API pricing and usage patterns. The barrier isn’t necessarily budget — it’s knowing how to architect autonomous systems that work reliably.

This isn’t about replacing human executive assistants (they’re irreplaceable for stakeholder management and complex coordination). It’s about democratizing the analytical infrastructure that helps leaders make informed decisions.

ASSUMED: Over the next year, I’m guessing we’ll see AI agents for productivity evolve from “smart assistants” to “autonomous intelligence teams.” The winners will likely be product leaders who learn to think like systems architects — designing agent workflows, not just prompting individual AIs.

The question isn’t whether AI agents will change how product leaders work. It’s whether you’ll design those systems yourself or let someone else define the methodology.


Want to build your own AI chief of staff? Start with one agent that handles one workflow autonomously. Master the autoresearch pattern. And always flag the difference between what you’ve measured and what you’ve inferred — your future self will thank you for the intellectual honesty.

Photo by CRYSTALWEED cannabis on Unsplash

What 23 Million Bible Readers Taught Me About Digital Discipleship

digital discipleship

Every month, roughly 23 million people open Bible Gateway to read Scripture. That’s more than attend every Southern Baptist Convention church on a given Sunday — the SBC’s own 2023 report counted 12.4 million in average weekly worship attendance.1

I lead product at HarperCollins Christian Publishing, where Bible Gateway is my primary focus. Before that, I spent years building SermonCentral — a platform serving 14,700+ subscribing pastors with access to 145,000+ sermon manuscripts — and co-built ORI, a youth discipleship app for mentoring teenagers. I’ve spent the last few years of my career watching how people actually behave when they engage with Scripture through technology. And what I’ve observed has changed the way I think about what “digital discipleship” means.

Content Distribution Is Not Discipleship

Most church tech conversations define digital discipleship as “putting Christian content online.” Upload a sermon. Publish a devotional. Build a Bible app.

That’s content distribution. Discipleship is something else.

From a product perspective, digital discipleship is designing technology that facilitates spiritual formation — helping people move from curiosity to commitment to transformation. The difference matters because it changes what you build. If you’re optimizing for content distribution, you chase volume: more translations, more devotionals, more features. If you’re optimizing for formation, you chase behavior change: consistency, depth, relationship.

Bible Gateway has given me a front-row seat to how millions of people actually engage with Scripture. Not how we hope they do, not how pastors assume they do — how they actually do. The patterns are humbling.

Commitment Structures Beat Content Volume

Bible Gateway offers hundreds of reading plans across dozens of categories. We have the content. What we’ve observed is that completion rates vary dramatically — and it’s not the “best” content that wins. It’s the best structure.

Short reading plans with clear daily commitments consistently outperform longer ones in completion rates. (I want to be precise: this is based on aggregate engagement data across our reading plan ecosystem, not a controlled A/B test. The pattern is strong, but I’m stating it as an observed trend.)

This makes sense if you think about it through a discipleship lens. The goal of a reading plan isn’t to get someone through the entire Bible in 365 days. The goal is to build a habit of daily engagement with Scripture. A 7-day plan someone finishes builds more spiritual momentum than a year-long plan abandoned in February. The research supports this — BJ Fogg’s work on Tiny Habits at Stanford demonstrates that small, completable commitments are the foundation of lasting behavior change.2

The product implication: when designing for digital discipleship, optimize for completion and consistency, not comprehensiveness. Finishable is better than thorough.

I saw the same thing at SermonCentral. Pastors didn’t need more sermon content — they needed the right content at the right time in their prep cycle. The value was relevance and timing, not volume.

The Gap Between Bible Search and Bible Study

Something surprised me when I first dug into Bible Gateway’s usage data: the overwhelming majority of sessions are what I’d call “Bible search” behavior, not “Bible study” behavior.

Most people come to look up a specific verse. They type “John 3:16” or “Philippians 4:13” into the search bar, read it, and leave. They’re using the platform as a reference tool. With over 2,000 Bible searches happening every minute on Bible Gateway, that’s a lot of single-verse visits.

This isn’t a criticism — it’s a behavioral insight with real implications for how we think about digital discipleship strategy.

If most users are in “lookup mode,” the discipleship opportunity isn’t in the content they came for. They already know that verse. The opportunity is in what comes next. Cross-references. Historical context. A reading plan that starts at that passage. A study note that opens the text up. The moment after someone finds what they came for is the moment a reference visit can become a formation experience.

(I should be transparent: I’m inferring the “lookup vs. study” distinction from session duration, page depth, and search query patterns in aggregate. We can see that a large portion of sessions are short and single-verse. But I can’t tell you what’s happening in someone’s heart during a 30-second visit — maybe that one verse is exactly what they needed. The data shows behavior, not transformation.)

The product principle applies broadly: meet people where they are, not where you wish they were. Design the next step from actual behavior, not from an ideal user journey.

The Day 7 Engagement Cliff

This is the most actionable pattern I’ve observed, and it’s consistent across every content platform I’ve worked on.

When someone starts a reading plan, engagement drops sharply after about Day 7. The first few days see strong completion. By the end of the first week, there’s a significant cliff. People who make it past Day 10 tend to finish — but a substantial number never get there.

(Evidence level: this is a pattern in aggregate reading plan data. Exact drop-off percentages vary by plan type and length, but the general shape — strong start, sharp drop around Day 7, stabilization for those who persist — is consistent enough that I’m confident calling it a pattern. This aligns with published habit formation research — Phillippa Lally’s 2009 study in the European Journal of Social Psychology found that early repetitions are the most fragile period for new habits.3)

For digital discipleship design, the implication is clear: Day 5 through Day 8 is where you need your best intervention design. Reminders. Encouragement. Community connection. A check-in from a real person. Whatever bridges the gap between initial motivation and formed habit.

This is where most digital discipleship tools fail. They’re good at onboarding. They’re good at content. They go quiet in the messy middle — the stretch where motivation fades and habit hasn’t locked in yet. That gap is where discipleship actually happens, and it’s where most apps have nothing to say.

At Bible Gateway’s scale, even small improvements in that Day 5-8 window could mean hundreds of thousands of people moving from casual lookup to sustained practice.

Why Features Rarely Solve Discipleship Problems

I’ve shipped a lot of features across my career. One thing I’ve learned — sometimes painfully — is that adding features to a discipleship tool almost never solves a discipleship problem.

The instinct is always to build more. More study tools. More social features. More gamification. But the digital discipleship tools that actually seem to work are the ones that reduce friction to spiritual practice, not the ones that add complexity to it.

Bible Gateway’s core value proposition is remarkably simple: read any Bible translation, for free, instantly. Over 200 versions in 70+ languages. That simplicity is the product. Every feature we consider needs to serve that core experience, not compete with it.

There’s a real tension here. Bible Gateway Plus offers 50+ study resources, ad-free reading, and deep study tools at $4.99/month. But even the premium tier works because it removes friction (ads, limited study tools) rather than adding cognitive load. The upgrade makes the simple thing simpler.

What ORI Taught Me About the Limits of Scale

All of this data-driven thinking needs a counterweight. For me, that counterweight is ORI.

ORI is a youth discipleship app I co-built, and its premise is different from a content platform like Bible Gateway. ORI facilitates the relationship between a mentor and a young person. The technology doesn’t do the discipleship — it supports the human who does.

That experience taught me something analytics can’t: the most effective digital discipleship tool is often the one that gets out of the way. The one that connects a young person with an adult who cares about them, gives them a shared framework for conversation, and then steps back. It echoes what Paul wrote to the Thessalonians — “We were gentle among you, like a nursing mother taking care of her own children” (1 Thessalonians 2:7, ESV). Discipleship has always been relational. Technology either serves that or distracts from it.

There’s a spectrum here. On one end, platforms like Bible Gateway serve millions with content at scale. On the other, tools like ORI serve hundreds by facilitating real human relationships. Both are valid. Both are needed. But they succeed for different reasons, and conflating them is a mistake I see church tech teams make often.

Friction Is the Enemy

If I had to compress everything I’ve learned into one principle: your job is to reduce friction between a person and their next spiritual step.

Not to create content. Not to build features. Not to gamify Scripture. To reduce friction.

At Bible Gateway’s scale, that means instant access to any translation, fast search, and reading plans designed around how people actually behave. At ORI’s scale, that means making it easy for a mentor to show up prepared for a fifteen-minute conversation with a teenager.

The 23 million people who use Bible Gateway each month aren’t a metric. They’re people in a spiritual practice — or trying to start one. The best thing a product team can do is figure out where the friction lives and get it out of the way.

I don’t have this figured out. The Day 7 cliff still exists. The gap between Bible search and Bible study is still wide. The question of whether a 30-second verse lookup counts as “discipleship” — I genuinely don’t know. But I think the question itself is worth sitting with, because how you answer it shapes everything you build.


Dr. Josh Read is Director of Product at HarperCollins Christian Publishing, where he leads Bible Gateway. He writes about the product side of digital discipleship at drjoshuaread.com. His other writing explores AI stewardship in ministry and what the Tower of Babel teaches us about technology.


1 Southern Baptist Convention, 2023 Annual Church Profile, reporting 12.4 million average weekly worship attendance across 47,000+ churches.

2 BJ Fogg, Tiny Habits: The Small Changes That Change Everything (Houghton Mifflin Harcourt, 2019). Fogg’s research at Stanford’s Behavior Design Lab demonstrates that starting small and building on success is more effective than ambitious commitment structures.

3 Phillippa Lally et al., “How Are Habits Formed: Modelling Habit Formation in the Real World,” European Journal of Social Psychology 40, no. 6 (2010): 998-1009.

The Tower of Babel Was a Technology Problem, Not a Language Problem

Most pastors I’ve talked to use the Tower of Babel the same way. It’s a warning against ambition. Don’t reach too high. Stay in your lane.

That reading has legs. But I’ve spent the last several years building products for churches — first at SermonCentral, where we managed over 245,000 sermon manuscripts for 14,700+ subscribers, and now at Bible Gateway, which serves 23 million monthly visitors across 200+ Bible translations. When I read Genesis 11 through a product lens, I see something the ambition reading misses.

God didn’t judge the bricks.

“Come, let us build ourselves a city, with a tower that reaches to the heavens, so that we may make a name for ourselves.” — Genesis 11:4, NIV

The materials were fine. The engineering was fine. The goal — consolidating human fame — was the problem. And that distinction matters right now, because the church is having the wrong argument about AI.

AI Is Bricks and Mortar

The debate I keep hearing splits along predictable lines. One camp says AI threatens authentic ministry. The other says it’s the future of outreach. Both are fixated on the tool and ignoring the purpose behind it.

AI is a building material. Your spam filter runs on it. Your search results are shaped by it. Your congregation interacts with machine learning dozens of times a day without a second thought. The question of whether the church uses AI was settled years ago.

The question that matters: what are you building, and for whom?

A church that uses AI to transcribe sermons so a deaf congregant can read along on Monday morning — that’s building for the Kingdom. A church that uses AI-generated sermons so the pastor can spend less time in the text — that’s a tower with its own name on it.

Same bricks. The blueprint is what changed.

Augustine’s Framework (From 397 AD)

About 1,600 years before anyone worried about ChatGPT, Augustine drew a line I think about constantly in product work.

In De Doctrina Christiana (Book I, chapters 3-4), Augustine distinguished between two postures toward the things of this world: uti (to use) and frui (to enjoy as an end in itself). His argument: the things of creation are meant to be used as means toward loving God and neighbor. They become disordered when we treat them as destinations — when we frui the tool instead of the purpose the tool serves.

I’ve found this more useful than any AI ethics whitepaper.

Consider: a church uses AI to automate its weekly bulletin, freeing up a volunteer to spend those 3 hours visiting a homebound member. That’s uti. The tool serves a human end.

Now consider: a church uses AI to eliminate pastoral presence altogether. Their new chatbot handles prayer requests, the algorithm personalizes a sermon playlist, the system runs without a shepherd. That’s frui. The church has started delighting in efficiency as its own reward.

The technology didn’t change. The orientation did.

Three Questions Before Adopting Any AI Tool

I’ve spent enough time in product leadership to know that the best safeguard isn’t a policy document (I’ve written plenty of those — they collect dust). It’s a habit of asking the right questions before you build.

1. Who benefits?

If the honest answer is “the budget” and not “the congregation,” pause. Cost savings aren’t wrong — stewardship matters. But if the primary beneficiary is the institution rather than the people it serves, you’re building in the wrong direction. The best AI implementations I’ve seen at Bible Gateway started with a specific human need, not a line item.

2. What human activity does this replace, and should that activity stay human?

Administrative tasks — scheduling, data entry, email sorting, transcript formatting — automate freely. These are good uses of AI. They free up people for work that only people can do.

But pastoral care, spiritual formation, the ministry of presence — these resist automation for a reason. A hospital visit from a pastor matters because a person chose to show up. An AI can generate a thoughtful prayer. It cannot bear witness to suffering.

(This is the question I find hardest to answer cleanly, by the way. The line between “administrative” and “pastoral” blurs more than we’d like. Where does sermon research end and sermon preparation begin? I don’t have a tidy answer. I think the honest move is to keep asking.)

3. Does this build the church’s capacity or create dependency on a vendor?

This is the product leader in me talking. I’ve watched organizations — churches included — adopt tools that felt like empowerment but functioned as dependency. If your church can’t operate without a specific AI platform, you haven’t adopted a tool. You’ve adopted a landlord.

Look for AI that trains your people. Look for solutions where the value stays with the church if the vendor disappears tomorrow.

From Babel to Pentecost

The Bible doesn’t end the language story at Babel. It picks it back up in Acts 2.

“All of them were filled with the Holy Spirit and began to speak in other tongues as the Spirit enabled them. Now there were staying in Jerusalem God-fearing Jews from every nation under heaven. When they heard this sound, a crowd came together in bewilderment, because each one heard their own language being spoken.” — Acts 2:4-6, NIV

At Babel, human technology consolidated power and built a monument to self. God scattered and confused. At Pentecost, the Spirit moved — and people from every nation heard the gospel in their own mother tongue. Each person’s language, met where they were.

According to recent Barna research, 77% of pastors believe AI can have a positive impact. I think that’s right — but only if we’re asking the Babel question each time we adopt something new.

Here’s what that looks like in practice: a small church in rural Guatemala using AI translation to access theological training that was previously locked behind an English-language paywall. That points toward Pentecost.

A megachurch using AI to scale content production so it can dominate more digital market share. That points back toward Babel.

What We Build Next

I don’t think the church needs to fear AI. I also don’t think it needs to be infatuated with it (and having built products in this space since 2018, I’ve watched both reactions play out in real time).

The bricks and mortar are here. They’re powerful. They’re going to keep getting more powerful. The church’s job is to ask the Babel question every time: what are we building, and whose name is on it?

That question doesn’t have a permanent answer. It has to be asked again with every new tool, every new capability, every new vendor pitch. And I think the churches that will get this right are the ones willing to sit with the discomfort of asking it honestly — even when the answer means building slower.


Sermon Illustration: The Tower of Babel and AI

When the people of Babel built their tower, God didn’t judge the bricks. He didn’t condemn the mortar or the engineering. The materials were fine. The problem was the purpose: “let us make a name for ourselves” (Genesis 11:4, NIV).

Today, AI is the new brick and mortar. Churches face the same question Babel faced: what are we building, and for whom? AI that frees a pastor to sit at a hospital bedside — that’s technology in service of presence. AI that replaces the pastor at the bedside — that’s a tower with our own name on it.

But the story doesn’t end at Babel. At Pentecost, God took language itself — the very thing He confused at Babel — and used it to carry the gospel across every barrier (Acts 2:4-6). The bricks are in our hands. The blueprint is the question.

Karpathy’s Autoresearch and the Parable of the Talents: What AI Stewardship Looks Like in Practice

A few weeks ago, Andrej Karpathy — former AI director at Tesla, co-founder of OpenAI — released a project that made me think about ministry.

I didn’t expect that either.

Karpathy built a framework called autoresearch. It runs autonomous ML experiments on a single GPU while the researcher sleeps. The AI agent modifies training code, runs a 5-minute experiment, evaluates the result, keeps improvements, discards failures, and loops. About 12 experiments per hour. Roughly 100 overnight. He woke up to measurable performance gains — with zero human intervention during the run.

The part that got me: Karpathy doesn’t write the training code anymore. He writes a Markdown file — plain English instructions — that tells the AI what to research, what constraints to follow, and when to stop. His words: “you are programming the `program.md` Markdown files that provide context to the AI agents.” He calls this “programming in Markdown.”

The human moved up one level of abstraction. Define the methodology, set the guardrails, let the system execute. Not less involved — involved differently, at the level of direction instead of mechanics.

39,800 GitHub stars in the first two weeks. The tech world noticed.

I think the church should too.

The Parable We Keep Skimming

In Matthew 25:14-30 (ESV), Jesus tells the story of a master who entrusts his servants with talents — significant sums of money — before leaving on a journey. One receives five talents, another two, another one. The first two invest and double their resources. The third buries his in the ground.

When the master returns, the investors are praised: “Well done, good and faithful servant. You have been faithful over a little; I will set you over much” (Matthew 25:21, ESV). The one who buried his talent gets rebuked. Not for losing money — he hadn’t lost anything. He was rebuked for doing nothing with what he’d been given.

We tend to read this as a general principle about using your gifts. It is that. But I think there’s something more pointed here for 2026.

AI is a talent in the Matthew 25 sense. It’s a resource placed in front of this generation, and we have a choice. Invest it toward the mission, or bury it because the risk feels too high.

What This Looks Like at My Desk

I want to be specific, because the abstract conversation about “AI and the church” doesn’t move anyone forward.

I’m Director of Product at HarperCollins Christian Publishing, where I lead Bible Gateway — a platform serving over 75 million monthly visitors engaging with Scripture. Before this role, I led product for SermonCentral, which grew to 14,700+ paying subscribers with access to more than 145,000 sermon manuscripts.

Over the past year, I’ve built a system of 18 AI agents that handle competitive analysis, research synthesis, meeting intelligence, content drafting, and task management. Several run overnight — not unlike Karpathy’s loop. The architecture is different (mine orchestrate across business functions, his optimizes a neural network), but the pattern is identical: define methodology, set constraints, let the system execute, review results in the morning.

Every hour I used to spend pulling competitor data or formatting reports is now an hour I spend thinking about how 75 million people experience Scripture online. Or how to make Bible Gateway better for the person opening it at 2 AM because they can’t sleep and need something solid to hold onto.

Karpathy programs research methodology in Markdown now instead of writing Python. I program strategic priorities and agent instructions instead of pulling spreadsheets. The abstraction layer moved up. The work got more human, not less.

The Fear Is Understandable — and Partly Right

I hear the concerns from church leaders, and I take them seriously.

AI will replace authentic ministry. AI will make pastors lazy. AI will simulate relational presence that only a human body in a room can provide. These aren’t irrational. Some are already happening in small ways.

If a pastor uses AI to generate a sermon they never wrestle with, that’s a problem. If a church deploys a chatbot as a substitute for pastoral counseling, that’s a problem. If we treat AI-generated prayers as equivalent to the honest, stumbling prayers of a person before God — we’ve lost something that matters more than efficiency.

But Karpathy’s work shows the other path. The tool doesn’t replace the human. It moves the human to where they’re most needed.

The pastor doesn’t stop preaching — they stop spending 4 hours hunting for the right illustration and spend that time with the family walking through a divorce. The administrator doesn’t stop managing — they stop updating attendance spreadsheets and spend that time training volunteers. The ministry leader doesn’t stop leading — they stop drowning in email and spend that time on the phone with a donor questioning their faith.

I’ve lived this tradeoff. When my agents took over competitive analysis (something that used to eat 3-4 hours a week), I didn’t fill that time with more busywork. I spent it in 1-on-1s with my team and in deeper product strategy. The output quality went up because I was operating at the right level of abstraction.

Where the Line Is (and Where I’m Still Figuring It Out)

I want to be honest — I don’t think anyone has this mapped perfectly yet. I certainly don’t.

Here’s where I’d draw it today:

AI should handle the administrative. Scheduling, data analysis, report generation, email triage, content formatting. These consume enormous amounts of ministry time, and they don’t require pastoral presence. Automate them aggressively.

AI should accelerate the research. Sermon prep research, theological cross-referencing, community demographic analysis. These benefit from AI’s speed and scope. The pastor still does the synthesis — the “what does this mean for my people on Sunday” work. But raw material gathering? Let the machine run overnight, like Karpathy’s experiments.

AI should never simulate the relational. It should not write your prayers. It should not be the voice your congregation hears when they need a shepherd. It should not replace the hospital visit, the awkward conversation in the parking lot, the moment after the service where someone says what they’ve been carrying for months.

The servant in Matthew 25 who was praised put the resource to work — but in service of the master’s purpose, not his own convenience (Matthew 25:20-23, ESV).

Here’s the tension I haven’t resolved: where does “accelerating research” end and “simulating thinking” begin? When an AI summarizes 30 commentaries on a passage, is the pastor still doing exegesis, or are they just picking from a menu? I don’t have a clean answer. I think it depends on whether the pastor is engaging the summaries critically or just grabbing the first one that sounds good. But that’s a discipline question, not a technology question — and discipline questions are harder to solve with guardrails.

If You’re a Church Leader Starting from Zero

You don’t need 18 agents. You need one tool that saves you 3 hours a week.

Pick the task that eats the most time with the least relational value. For most pastors I’ve talked to, it’s sermon illustration research, email management, or meeting notes. Start there. Learn one tool well. Measure the hours you get back.

Then — and this is the part most people skip — reinvest that time in something only a human can do. A visit. A phone call. An hour of prayer you’ve been meaning to protect but kept losing to administrative drift.

Set your guardrails before you need them. Write down what AI will not do in your ministry context. Revisit it quarterly. Technology expands into unintended spaces when boundaries aren’t explicit — I’ve watched this happen in product development for 15 years.

The Talent in Front of Us

Karpathy’s autoresearch is an engineering achievement. But the deeper pattern is almost theological: the human was never meant to stay at the level of mechanical execution. We’re built to operate at the level of purpose, direction, and relationship. Genesis 1:28 gives humanity dominion and stewardship — a mandate to cultivate, not just maintain (Genesis 1:28, ESV).

The master in the parable didn’t give talents so the servants could admire them or lock them away. He gave them to be invested — put to work — in ways that generated return.

For those of us building technology that serves the church, the return isn’t financial. It’s pastors freed from busywork to do the work they were called to. It’s 75 million monthly visitors encountering Scripture through a platform that keeps getting better because the product team has time to think. It’s churches stewarding every tool available — including AI — in service of the mission they’ve been given.

The talent is in front of us. What we do with it is a stewardship question.


Josh Read is Director of Product at HarperCollins Christian Publishing (Bible Gateway) and holds a doctorate in Strategic Organizational Leadership. He writes about AI, product leadership, and digital discipleship at drjoshuaread.com.

7 Things I Read This Week (and Why They Matter)

This was one of those weeks where everything I read seemed to converge on the same theme: the ground is shifting faster than most of us realize. AI isn’t coming for our workflows someday – it’s already reshaping how products get discovered, how code gets written, and whether your product-market fit survives the next 12 months.

Here’s what caught my attention.

1. Product Market Fit Collapse: Why Your Company Could Be Next

Reforge Blog

If you’re in SaaS, this is the chart that should scare you. Reforge makes the case that PMF isn’t a destination – it’s a treadmill. And AI just cranked the speed to max. Chegg lost 87.5% of its valuation. Stack Overflow’s traffic cratered. The pattern is the same: AI proves value for a use case, and the incumbent’s window to adapt slams shut before they even recognize the threat.

This one hit me personally. SermonCentral has been the go-to sermon library for over two decades. BUT the question I keep coming back to is: what happens when pastors can generate sermon outlines with AI in seconds? The PMF threshold doesn’t care about your legacy. It only cares about whether you’re still the best answer to the customer’s problem RIGHT NOW.

2. What AI Sees When It Visits Your Website (And How To Fix It)

Google Share

This reframed how I think about our SEO strategy entirely. AI answer engines – ChatGPT, Google AI Overviews, Perplexity – are visiting your site, interpreting your content, and shaping customer perception BEFORE a human ever clicks. Traditional SEO isn’t enough anymore. You need AEO – AI Engine Optimization.

For SermonCentral, this is urgent. We live and die by organic discovery. If AI systems can’t parse our content well, we lose visibility in the exact channels that are replacing traditional search. I’m bringing this to the team this week.

3. Claude Code Remote Control

Claude Code Docs

This is the kind of workflow upgrade that sounds small but changes everything. Claude Code now lets you continue local dev sessions from your phone, tablet, or any browser. Your full local environment stays intact – filesystem, MCP servers, all of it. Sessions reconnect automatically after network drops or laptop sleep.

I’ve been using Claude Code as my daily driver for months now. Being able to kick off a task at my desk and check progress from my phone during a walk? That’s the kind of automation leverage I’m optimizing for in 2026.

4. Claude Code for Web – Async Coding Agent

Simon Willison

Anthropic launched an async coding agent at claude.ai/code. Point it at a GitHub repo, give it a task, and it creates branches and PRs with the work output. It runs in a container, skips permission gates, and the PRs are indistinguishable from CLI-generated ones.

The coding agent space is getting crowded fast – OpenAI Codex Cloud, Google Jules, now this. What I appreciate about this one is the “teleport” feature that lets you copy the transcript and files to your local CLI. It’s not replacing the local workflow, it’s extending it. That’s the right design philosophy.

5. How to Build a PM GitHub That Gets You Hired

Aakash’s Newsletter

Only 24% of PM candidates have GitHub profiles. That stat alone should tell you something. Hiring managers at Google, OpenAI, Anthropic, and Meta actively check GitHub when it’s linked. A strong profile signals you actually build things and understand engineer workflows – not just strategize from a slide deck.

I’ve been saying this for a while: the best PMs ship. They don’t just write specs. If you’re a PM reading this and you don’t have a GitHub presence, this is your sign. Start small. Ship something. The differentiation is massive because almost nobody does it.

6. Visual Explainer – Agent Skill for Rich HTML Output

GitHub

This is a neat agent skill that converts complex terminal output into styled, interactive HTML pages. Think: architecture diagrams, code diff reviews, project plan audits, data tables – all rendered as shareable HTML without manual formatting.

I’m always looking for ways to make technical work more visible to non-technical stakeholders. Being able to generate a polished visual recap of a sprint or a system change and just send the HTML? That’s a communication multiplier.

7. Anthropic Courses on Skilljar

Anthropic Courses

Anthropic now has 14+ structured courses covering Claude API, Model Context Protocol, and AI fluency for developers, educators, students, and nonprofits. This tells me they’re investing heavily in ecosystem education – and that MCP is becoming a first-class skill.

I’ve been building MCP integrations into my daily workflow for months. Seeing Anthropic formalize the training around it validates the bet. If you’re building on Claude and haven’t gone through these, it’s worth the time.

The Thread That Ties It All Together

Every link this week points to the same reality: the cost of standing still just went up. PMF is collapsing faster. AI is reshaping discovery. Coding agents are shipping real code. The PMs who build things are getting hired. The tools are getting better every week.

The question isn’t whether to adapt. It’s whether you’re adapting fast enough.

I aim to be on the right side of that question. Hopefully some of these links help you get there too.

AI Just Walked Into Your Website Without Knocking

Last month I asked ChatGPT a question I’ve asked Google a thousand times: “What’s a good sermon illustration about forgiveness?”

It gave me a solid answer. Three illustrations, structured with context, application points, even a suggested closing line. It was genuinely useful.

And it never sent me to a single website.

That moment hit me differently than it would have two years ago. I run a platform with over 245,000 sermons and 50,000 illustrations. I didn’t just lose a click. I watched an AI system do what our product does, using content that likely came from sites like ours, and deliver it in a way that made visiting the source unnecessary.

That’s a revenue problem. (I wrote about the traffic implications of this shift recently.) (I wrote about the traffic implications of this shift recently.)

The Zero-Click Layer

Most product leaders I know are still thinking about AI as a feature to bolt onto their product: chatbots, smart search, AI-generated recommendations. And that matters. But there’s a bigger shift happening underneath that conversation.

AI answer engines (ChatGPT, Google AI Overviews, Perplexity) are becoming the front door to the internet. They don’t just search. They visit your site, interpret your content, synthesize it, and serve it directly to the user. The user gets the answer. You get nothing.

Google’s featured snippets started this zero-click trend years ago. BUT what’s different now is the depth. A featured snippet pulls a paragraph. An AI answer engine can synthesize an entire page, or multiple pages, into a comprehensive response that genuinely satisfies the user’s intent.

If your business depends on organic traffic as a top-of-funnel engine, this should keep you up at night.

Your Content Library Is Both Your Greatest Asset and Your Biggest Vulnerability

Here’s the paradox I’ve been sitting with.

We spent years building one of the largest structured content libraries in our space. That library is what drives our organic traffic. It’s what Google indexes. It’s what pastors find when they search “sermon on grace” at 11pm on a Saturday night.

That same library is now what AI systems are ingesting to train their models and generate their answers. The very content that built our moat is being used to fill in the moat.

And here’s what makes it worse. The emerging AI-native competitors in our space don’t even need to win Google rankings. They ARE the AI tool. They’re built to live inside AI workflows, not compete for traditional search clicks.

I think this pattern applies to any SaaS company sitting on a large content asset. If you’ve built your growth engine on content that AI can summarize, you’re exposed.

AEO: A Genuinely Different Discipline

There’s a term gaining traction: AEO, or AI Engine Optimization. And I’ll be honest, my first reaction was skepticism. We don’t need another three-letter acronym.

But the more I’ve dug into it, the more I realize it represents a genuinely different discipline.

SEO optimizes for ranking. AEO optimizes for citation. The goal is to be the source that AI systems reference AND link back to. That requires a fundamentally different content strategy.

Here’s what that looks like in practice:

  1. Structured data becomes non-negotiable. Schema markup, clear metadata, explicit problem-solution framing in your content. AI systems parse structure, not vibes. (Schema.org is the starting point.)
  2. Content architecture matters more than keyword density. How your content is organized (headers, relationships between pages, internal linking) determines how AI systems understand your authority on a topic.
  3. Gated content is a double-edged sword. If your best content is behind a login wall, AI crawlers can’t index it. You’re invisible to the answer engine. But if everything is open, you get summarized without a click. The play is in the middle: structured preview content that AI can cite, with depth that requires the visit.
  4. Domain-specific language is your moat. Generic content gets synthesized away. Content that uses the precise language of your audience (the way a pastor describes their Saturday night prep struggle, the specific vocabulary of sermon structure) is harder for AI to replace and more likely to be cited with attribution.

What I’m Doing About It

I’m not going to pretend I have this figured out. But here’s where my head is:

Audit how AI sees us. Before optimizing anything, we need to understand how our top pages render to AI crawlers. What structured data exists? What’s behind login walls that blocks indexing?

Treat AI referral as a distinct channel. We track direct traffic, organic search, paid. AI referral needs its own lane in our analytics. We can’t optimize what we can’t measure.

Build content AI can’t summarize away. The full sermon text? AI can handle that. But a pastor’s framework for adapting a sermon to their specific congregation? A diagnostic tool for matching an illustration to a particular emotional moment in a service? That’s interactive, personalized, and requires being on the platform.

Move faster than the AI-native competitors. They have the structural advantage of being built for AI workflows. We have the structural advantage of 20+ years of trusted content and relationships. The question is whether we can adapt our distribution before they build our depth.

The Strategies That Got You Here Won’t Sustain You

I keep coming back to this. The strategies that built organic growth over the last decade won’t sustain it over the next five years.

That’s a reason to move, not a reason to panic.

The companies that treat AI answer engines as a new channel will capture disproportionate share of the next era of discovery. The ones that keep optimizing for Google page one while AI summarizes their content into zero-click answers will watch their traffic erode and wonder what happened.

I’d rather be early and wrong about the tactics than late and right about the trend.

The AI just walked into your website. The question is whether it’s going to send people your way, or make visiting you unnecessary.

Revolutionizing Product Management: Insights from Industry Leaders and Emerging Trends

In the fast-paced world of software as a service (SaaS), product management stands at the intersection of technology, strategy, and user experience. Recent insights from industry leaders and emerging trends highlight how product teams are navigating new challenges and opportunities, especially with the integration of AI and advanced database infrastructures. This article explores key learnings from top product management blogs, offering a comprehensive guide for professionals looking to enhance their strategies and operations.

The Evolution of Database Infrastructure

The shift in database technology is pivotal for SaaS companies aiming for scalability and performance. Intercom’s journey with Vitess and PlanetScale, as discussed in “Evolving Intercom’s database infrastructure: Lessons and progress,” showcases:

  • Scalability: How adopting PlanetScale Metal has allowed for zero-downtime maintenance and performance improvements.
  • Performance: Insights into how new database technologies can handle increased load without compromising speed.
  • Lessons Learned: The challenges and triumphs of integrating new systems into existing architectures, offering a roadmap for similar transitions.

Designing for User Clarity

Product design isn’t just about aesthetics; it’s about clarity and usability. Pranava Tandra from Intercom shares in “Intercom on Product: Designing for Clarity”:

  • Balancing Simplicity with Depth: Strategies for redesigning information architecture to make complex features discoverable yet not overwhelming.
  • AI Integration: How AI can be seamlessly integrated to enhance user interaction without disrupting the existing user flow.

Learning from Product Conferences

Conferences like #mtpcon London provide a wealth of knowledge:

  • Key Takeaways: Insights from product leaders on current trends, including the integration of AI in product management, as seen in “What we learned at #mtpcon London 2025.”
  • Networking and Collaboration: The value of community and peer learning in advancing product strategies.

Leveraging AI in Product Management

AI’s role in product management has grown exponentially:

  • Enterprise AI Agents: The article “How to build an Enterprise AI Agent” discusses how AI can manage and utilize organizational knowledge, reducing productivity drain.
  • Analytics Superpowers: AI’s ability to simplify SQL queries and data analysis, as highlighted in “Are you struggling with SQL? AI can give you analytics superpowers.”

Strategic User Engagement and Retention

Driving user adoption and retention is crucial:

  • Onboarding Gamification: Utilizing gamification techniques to engage users during onboarding, as explored in “11 Onboarding Gamification Examples to Engage & Retain Users.”
  • Feature Adoption: Tactics to ensure new features are adopted, enhancing user experience and product value, detailed in “How to Drive Feature Adoption: 10 Proven Strategies (+ Examples).”

Customer-Centric Approaches

Understanding and leveraging customer feedback:

  • Feedback Tools: A review of the best tools for collecting and analyzing customer feedback in “16 Best Customer Feedback Tools For SaaS Companies.”
  • UX Metrics: Key metrics to track for measuring user experience, as discussed in “How to Measure User Experience: 7 Key UX Metrics.”

Conclusion

The landscape of product management is continuously evolving, driven by technological advancements and strategic insights shared by industry leaders. From redesigning for clarity and leveraging AI for deeper analytics to enhancing customer engagement through innovative onboarding and feedback mechanisms, product managers have so many tools and knowledge at their disposal. By staying informed and adaptable, product teams are ready to not only meet but exceed market demands, ensuring their products remain competitive and user-centric.