The question I keep hearing from product teams is how to move from shipping AI features to actually building agents that improve over time. The answer is not better model selection or tighter prompt engineering. It is shifting from feature velocity to systems that accumulate team judgment.
This question surfaces because most roadmaps still treat agents as a new category of deliverable. Teams list capabilities like “handles scheduling conflicts” or “summarizes user feedback” and then measure progress by how many of those items ship. The result is agents that perform isolated tasks but never learn from the context the rest of the team already holds.
Teresa Torres’s continuous discovery framework shows why this pattern repeats. Torres argues that discovery must run in parallel with delivery, not in advance of it. When teams apply that lens to agents, the problem becomes clear: most current agent work still runs on a delivery-first cadence. The agent gets defined, built, and released before anyone checks whether the underlying assumptions about team knowledge have changed.
Why agent roadmaps keep collapsing into feature lists
Feature lists reward visible output. An agent that books rooms looks like progress even if it books the wrong room for the same reasons the last three tools did. The list does not track whether the agent knows why a certain room was rejected last quarter or which stakeholder vetoed it. Over time the agent repeats the same errors because nothing in the roadmap forces the team to record those reasons.
Continuous discovery changes the unit of work. Instead of adding another agent behavior, the team runs weekly interviews focused on recent decisions the agent should have influenced. The output is not a new feature but an updated set of assumptions that get fed back into the agent’s context layer. This keeps the roadmap from drifting into isolated capabilities.
One team I watched tried to build an agent for volunteer scheduling across multiple church sites. Their initial roadmap listed twenty behaviors. After running discovery interviews with the actual coordinators, they dropped twelve of those behaviors and instead built a simple way to capture why a volunteer was declined. The agent improved faster because the team now had a living record of past judgments rather than a growing list of functions.
Discovery loops that actually capture institutional knowledge
Most teams run discovery as a research phase that happens before build. Torres’s approach treats it as an ongoing loop that runs alongside delivery. For agent work this means the discovery questions must target how past decisions get reused, not just what users say they want next.
The practical version looks like this: every week the product manager pulls the three most recent agent failures or overrides. They sit with the people who made the override and ask what information the agent was missing. That single piece of context gets turned into a structured note that the agent can reference. Over months the notes form a lightweight institutional memory that no feature list ever captured.
Without this loop the agent stays frozen at the knowledge level of its first release. With the loop the agent slowly absorbs the pattern of choices the team has already made. The discovery work is not separate from the agent; it is the mechanism that makes the agent accumulate judgment.
The org design move most teams skip
Continuous discovery requires someone to own the loop, not just the output. Most agent teams assign ownership to an engineer or a single product manager who already carries delivery pressure. The discovery work gets deprioritized the moment shipping dates slip.
The teams that make progress create a rotating discovery owner role that sits outside the delivery sprint. That person runs the weekly interviews, maintains the decision notes, and updates the agent’s context rules. Delivery still happens, but the discovery owner protects the accumulation of judgment from getting lost in velocity metrics.
This role does not require a new headcount. It requires a temporary shift in how existing time gets allocated. One team I worked with gave the discovery owner four protected hours each week. Those hours produced more lasting agent improvement than the previous three months of feature additions.
Your Turn: Apply This Today
- Pick the agent your team ships most often and list the last five times a human overrode its output.
- Schedule three thirty-minute conversations this week with the people who made those overrides; ask only what context the agent lacked.
- Turn the clearest missing piece from each conversation into a one-sentence note and add it to whatever context store the agent already uses.
- Assign one person on the team to own the weekly override review for the next four weeks and protect those four hours on their calendar.
- Measure whether the same override pattern appears again after the notes are added; track it for two releases.
- Write a short internal note on what the agent now handles differently and share it with the broader product group.
Two earlier posts on the same thread walk through how judgment shows up in product systems and why most discovery work fails to stick once delivery pressure returns.
I consult with product leaders on AI agent roadmaps, continuous discovery loops, and building systems that accumulate team judgment. Let’s talk.
Discover more from Dr. Joshua Read
Subscribe to get the latest posts sent to your email.

