When I joined a $10M GenAI transformation program at Deloitte Digital, the client brief was three words: "AI in the contact center." That's it. No feature spec. No success criteria. Just a category and a budget.
The client — a major US health insurer with 2,000+ agents and roughly a million customer interactions per year — had been briefed on what GenAI could do. They were excited about one feature in particular: a real-time AI assistant that would listen to live customer calls and whisper the right answers to agents in real time. They had seen a demo. They expected it in Sprint 1.
My recommendation was to defer that feature entirely to Phase 2. That conversation was the hardest one I had on the program. And it was the right call — proven six months later by a 78% AI acceptance rate and a program that finished 20% ahead of schedule.
Here's the framework that made the argument, and how to apply it to any GenAI roadmap with too many ideas and not enough guidance on where to start.
The Problem With Most GenAI Roadmaps
Standard product prioritization frameworks — RICE, ICE, value/effort matrices — are built for a world where the cost of being wrong is recoverable. You ship a feature, it underperforms, you iterate. The feedback loop is the mechanism.
GenAI breaks this assumption in regulated environments. The cost of being wrong is not always recoverable. An AI system that's right 90% of the time but wrong 10% in dangerous ways is worse than no AI at all. In healthcare, a wrong AI output during a live customer call isn't just a bad experience — it's a compliance incident and potentially a patient safety risk.
When I mapped 180+ distinct agent tasks across three business units, I noticed a pattern: the features stakeholders were most excited about had the highest failure cost. The features nobody mentioned in the kickoff meeting had the most benign failure modes and the fastest path to measurable ROI.
The Four Questions
Before any GenAI feature gets a priority ranking, I now ask four questions. All four answers have to inform the sequencing decision — not just the first one.
What's the upside if the AI gets it right?
Time saved, cost reduced, quality improved. The standard question. Every roadmap asks this one. Necessary but not sufficient.
What's the downside if the AI gets it wrong?
This is the question most roadmaps skip. Wrong output for a post-call summary = an agent spends 30 seconds correcting it. Wrong output during a live healthcare call = potential compliance incident, patient harm, or agent embarrassment that kills adoption permanently. The failure mode shapes everything.
Is the training data ready?
AI features built on insufficient domain-specific data produce confident but wrong outputs. Data readiness isn't a technical detail — it's a risk multiplier. A high-risk feature with poor data readiness is a compliance incident waiting to happen.
Can you measure whether the AI is actually helping?
Features where success is ambiguous are dangerous to ship first. If you can't measure the outcome, you can't detect early failures or justify the investment to stakeholders who are watching closely.
Only after answering all four do you have enough information to sequence the feature correctly.
Applied to the HCSC Roadmap
Here's how those four questions changed the priority order for this specific program. The left column is what stakeholders wanted. The right column is what the framework produced.
Risk-Adjusted GenAI Feature Prioritization — HCSC Contact Center
| Feature | Upside if Right | Downside if Wrong | Data Ready? | Priority |
|---|---|---|---|---|
| Smart email drafting | 8–12 min saved per email, high volume | Low — agent reviews before send | Yes — email corpus available | P0 |
| Post-call documentation | 12–18 min saved per call, QA coverage | Low — agent confirms before submitting | Yes — call transcripts available | P0 |
| QA sentiment flagging | QA coverage from <10% to 80%+ | Medium — false positives waste supervisor time | Partial — needs calibration | P1 |
| Escalation routing | Faster resolution, right agent every time | Medium — wrong routing = customer frustration | Partial — historical routing data | P1 |
| Real-time agent guidance | Very high — correct answers live on calls | High — wrong guidance during live call = compliance liability, patient harm | No — needs HCSC-specific training | P2 → Phase 2 |
The counterintuitive result: the feature with the highest upside scored last, not first. High upside + high downside + poor data readiness = maximum risk, minimum delivery confidence. Shipping it first — before delivery credibility was established — would have been betting the entire program on the feature most likely to fail catastrophically.
The key insight: In standard prioritization, high upside pushes a feature up. In AI prioritization in regulated environments, high upside paired with high downside risk and poor data readiness pushes it down — because the cost of an early failure is not just one bad sprint. It's trust damage that can kill the program.
The Conversation That Mattered
HCSC leadership had been sold on the real-time guidance vision. It was the demo they'd seen, the feature they mentioned in every planning meeting, the thing they came to the kickoff ready to build. Telling them it would not be in Phase 1 required a specific kind of argument.
I didn't frame it as "this feature is too risky." That sounds like I didn't believe in the vision. I framed it as a sequencing problem: "Let's earn the right to do the high-stakes feature by proving we can nail the low-stakes ones first."
Defer real-time guidance to Phase 2. Ship email drafting first.
"If we ship smart email drafting on time and it performs, you'll trust us with the feature that matters most. If we rush real-time guidance and it fails — even once, in a compliance-sensitive context — the entire program's credibility is at risk. Phase 1 is about building the delivery trust that makes Phase 2 possible."
What I risked: stakeholder frustration in the short term. If Phase 1 underdelivered, deferring the headline feature would look like the wrong call in retrospect. I needed Phase 1 to be undeniable.
Agents accepted the AI-generated draft without significant edits 78% of the time. That number became the internal proof of concept — and the stakeholders who initially wanted real-time guidance in Phase 1 became the biggest advocates for Phase 2 expansion after they saw what the framework produced.
One More Thing: Design for the Skeptic
The risk-adjusted framework handles sequencing. There's a separate UX decision that shapes whether the features you ship actually get adopted.
When we shipped the first AI features, the most technically capable one had the lowest initial adoption. Agents were skeptical — they had experienced automation that got it wrong and felt embarrassed in front of customers. "Technically correct 85% of the time" is not enough for a person whose professional credibility is on the line in front of a customer.
Clean AI output, no friction
Assumes the user trusts the AI. Fast to build. Feels modern. Delivers low adoption when users are skeptical — which most enterprise users are, especially after prior automation projects failed them.
AI confidence score + easy override
Shows the model's confidence. One-click override. Agent stays in control. Slower to build. Adoption jumped ~30% vs. baseline assumption. Override data also feeds model improvement.
The AI confidence score wasn't technically necessary — it was a UX decision. But it was the right one for the actual users, who were not AI enthusiasts. They were healthcare agents under time pressure who needed to trust the tool before they'd let it touch their workflow.
The lesson generalizes: your AI feature adoption will be determined by your most skeptical user, not your most enthusiastic one. Design for the skeptic. Make the override obvious. Show the confidence. Let them stay in control. Adoption comes from trust, not from having the best model.
How to Apply This to Your Roadmap
If you have a backlog of AI feature ideas and need to sequence them, the practical version of this framework is a one-page table. For each candidate feature, fill in four cells: upside if right, downside if wrong, data readiness, measurability. Then look for the pattern.
P0 candidates look like this: moderate to high upside, low downside risk (because a human reviews before anything goes out), existing data to train on, clear metric to track. These are your "earn trust" features. Ship them first, hit the numbers, unlock the harder features.
P2 candidates look like this: very high upside, very high downside risk, insufficient domain-specific data, ambiguous success criteria. These are your "Phase 2" features. Not because they're not valuable — they may be the most valuable thing on the roadmap. But they need delivery credibility to be built before they're deployed, and they need model quality that usually requires training data that only Phase 1 can generate.
Before sequencing any AI feature, ask: "What happens the first time this gets it wrong — and can the program survive that?" If the answer is no, the feature is not Phase 1, regardless of how exciting the upside looks.
Standard value/effort matrices don't work for AI in regulated environments. They optimize for upside and ignore downside risk. Add an explicit "downside if wrong" column to every AI prioritization exercise — it will change the order.
The most-wanted feature is often the most dangerous to ship first. Excitement and risk are correlated in GenAI because the most impressive demos are usually the features where wrong AI output has the highest cost. Separate excitement from sequencing.
Data readiness is a risk multiplier, not a technical detail. A high-risk feature built on insufficient training data is a compliance incident waiting to happen. Data readiness should gate prioritization, not follow it.
Delivery credibility is currency. Phase 1 earns Phase 2. If you can demonstrate that you shipped the low-risk features on time and on spec, you'll have the stakeholder trust to deploy the high-stakes features — and they'll be better for the additional training data Phase 1 generates.
Design your AI UX for the skeptic, not the enthusiast. Show confidence scores. Make overrides easy. Keep humans in control. Adoption is a product problem, not a training problem — the best model with a bad UX for skeptical users will still have low adoption.