Product Management
20+ discovery calls to a top-3 roadmap brief
· Updated · 5 min read
S
Get AI Workflow Teardown Template
The template I use to map AI workflows before writing a spec.
No spam, unsubscribe anytime. I only email teardowns.
· Updated · 5 min read
The template I use to map AI workflows before writing a spec.
No spam, unsubscribe anytime. I only email teardowns.
Discovery isn't call stacking; it's a ranked deliberation. At Upcore we ran 20+ prospect interviews across 12 verticals, distilled them into an opportunity brief, and surfaced the top-3 agentic-AI use cases by severity, ROI, and buildability. The scoring frame is the deliverable, not the call logs.
Twenty discovery calls will give you twenty compelling stories, and a roadmap can only hold three. That gap — between what you heard and what you're allowed to build — is where most discovery quietly turns into a highlight reel of the loudest prospect. Teresa Torres' continuous-discovery research makes the same point structurally: interview first, map what you heard into an opportunity space, and never let a solution skip the mapping step just because it's exciting. The fix isn't more calls. It's a decision stack ruthless enough that any executive can trace why use case #1 beat use case #7.
We didn't wing it. Every call followed the same 45-minute structure:
"What does your week look like? Walk me through the last time [problem area] came up."
"If we built exactly what you described, what would make you not buy it?"
That last one surfaced more kill criteria than any feature list. It's the same instinct behind The Mom Test: people will happily bless your idea; they'll only tell the truth about their constraints.
The synthesis wasn't thematic clustering — it was a scoring machine, close in spirit to RICE prioritization but reshaped for discovery instead of a backlog: no Reach factor, because at this stage every opportunity is still unproven, and no single Effort denominator, because feasibility varies too much by vertical to reduce to one number.
Every opportunity got a triple score. The top-3 use cases by this triple are what a roadmap should be allowed to defend. Everything else is a tasteful "not yet."
| Score | Severity | ROI | Buildability |
|---|---|---|---|
| 5 | Stops work daily; regulatory risk | Budget allocated, urgent | Core competency, data ready, GTM clear |
| 4 | Weekly blocker; workarounds expensive | Budget signaled, timeline defined | Adjacent to core, data accessible |
| 3 | Monthly annoyance; workarounds exist | Budget possible next cycle | New domain, data needs work |
| 2 | Quarterly friction; tolerable | No budget, maybe later | Significant R&D needed |
| 1 | Nice to have; no urgency | No budget, no owner | Research project |
We scored independently (PM + eng lead), then reconciled. Disagreements usually meant "we don't understand the build" — which is a signal itself.
After scoring independently, PM and eng lead sit down with the spreadsheet. The rules:
This meeting is where the roadmap actually gets written. The scoring is just the prep.
| Rank | Use case | Verticals | Severity | ROI | Buildability | Composite |
|---|---|---|---|---|---|---|
| 1 | Agentic dispatch (manufacturing/logistics) | Mfg, Logistics, Wholesale | 5 | 5 | 4 | 14 |
| 2 | AI routing & scheduling (logistics/field services) | Logistics, Field Svcs, Utilities | 4 | 5 | 4 | 13 |
| 3 | Agentic onboarding (SaaS/B2B) | SaaS, Fintech, Edtech | 4 | 4 | 5 | 13 |
The tie at 13 was broken by learning velocity: dispatch had the clearest 0→1 path.
The discovery brief wasn't just a roadmap artifact — it became the pricing framework. For each top-3 use case, we modeled:
| Use case | Pricing metric | Willingness to pay (median) | Model |
|---|---|---|---|
| Agentic dispatch | $/truck/month | $180 | Per-asset SaaS |
| AI routing | $/route optimized | $0.45 | Usage-based |
| Agentic onboarding | $/seat/month | $45 | Seat-based |
These numbers came directly from the budget signals in discovery. We didn't guess — we asked "what would you pay?" and the median answer became the anchor.
Discovery's deliverable isn't what you heard. It's a ranking any skeptic can audit.
For the mechanics I keep alive, the deep-dive method covers evidence. The downstream scoring system this fed is in the Upcore lead-scoring case study.