$10M program. 2,000+ agents. A client who wanted "AI in the contact center" but couldn't define what that meant. I built a risk-adjusted prioritization framework — and argued against shipping the feature everyone wanted in order to ship the one that would prove we could be trusted.
Deloitte's Salesforce-based platform for enterprise contact center management. Handles ticket routing, agent workflows, post-call documentation, and CRM integration. Used by Fortune 500 clients including Apple Inc. and HCSC across multiple verticals.
Health Care Service Corporation — one of the largest US health insurers. ~1 million customer interactions per year across 2,000+ agents in three business units: individual plans, employer benefits, and provider relations. The program goal: "use AI to improve agent efficiency."
That's not a product brief. That's a category. My first job was to turn "improve agent efficiency with AI" into specific, measurable, shippable features — in the right order. That required understanding the problem before touching a feature list.
HCSC leadership had been briefed on GenAI capabilities and arrived with a vision: an AI assistant that could guide agents in real time during customer calls. It was compelling, impressive, and — by my analysis — the wrong place to start.
The category excitement around GenAI created a specific PM challenge: the most visually impressive features were also the highest risk. Shipping the flashy feature first, before establishing delivery credibility, would have exposed the program to catastrophic failure in a regulated healthcare environment.
"Use AI to improve agent efficiency across 2,000+ agents handling 1M+ customer interactions per year."
In healthcare, wrong AI output during a live call = compliance liability + patient safety risk. The cost of being wrong is asymmetric. A system that's right 90% of the time but wrong 10% in dangerous ways is worse than no AI.
Agents were skeptical — they'd seen automation promises fail before. They'd experienced suggestions that were wrong or tone-deaf. Designing AI features for the enthusiast (who trusts AI by default) produces different products than designing for the skeptic (who needs to feel in control). We designed for the skeptic. That's why adoption worked.
Ran structured discovery sessions across all three business units. Mapped 180+ distinct agent tasks. For each task, scored: automation potential (can AI reliably do this?), business impact (what's the cost if this task is slow or wrong?), and data readiness (how much labeled data exists?).
Duration: 3 weeks · Method: Structured workshops, task taxonomy mappingThe most-requested feature (real-time agent guidance during live calls) scored highest on business impact — but also highest on downside risk. The least-requested features (post-call documentation, email drafting) scored highest on ROI per unit of effort because their failure mode was benign: agent reviews before anything goes out.
Finding: High excitement ≠ high priority when failure cost is asymmetricStandard GenAI prioritization frameworks optimize for impact and effort. They miss the critical third dimension for regulated environments: the asymmetric cost of being wrong.
For each proposed AI feature, I added four explicit questions before scoring:
Each feature scored on: (1) Upside if right, (2) Downside if wrong, (3) Data readiness, (4) Measurability
| Feature | Upside if Right | Downside if Wrong | Data Ready? | Priority |
|---|---|---|---|---|
| Smart email drafting | Saves 8–12 min/email, high volume | Low — agent reviews before send | Yes — email corpus available | P0 |
| Post-call documentation | 12–18 min saved per call, QA coverage | Low — agent confirms before submitting | Yes — call transcripts available | P0 |
| QA sentiment flagging | QA coverage from <10% to 80%+ | Medium — false positives waste supervisor time | Partial — needs calibration | P1 |
| Escalation routing | Faster resolution, right agent every time | Medium — wrong routing = customer frustration | Partial — historical routing data | P1 |
| Real-time agent guidance | Very high — correct answers during live call | High — wrong guidance during live call = compliance liability, patient harm | No — needs extensive HCSC-specific training | P2 → Phase 2 |
The framework reveals the counterintuitive truth: the feature with the highest upside has the highest downside and the lowest data readiness. Shipping it first maximizes the chance of failure in a healthcare environment where wrong = compliance incident.
HCSC leadership had been sold on the real-time agent assistant vision. It was the demo they'd seen, the vision they'd bought into, and the feature they expected in Sprint 1. My recommendation was to defer it to Phase 2.
My argument: "Let's earn the right to do the high-risk feature by proving we can nail the low-risk ones first. Phase 1 is about building delivery credibility. If we ship smart email drafting on time and it performs, you'll trust us with the feature that matters most. If we rush real-time guidance and it fails, the entire program's credibility is at risk in a regulated environment."
What I risked: Stakeholder frustration in the short term. If Phase 1 underdelivered, deferring the headline feature would look like the wrong call in retrospect.
Agents accepted the AI-generated draft without significant edits 78% of the time. That number became the proof point for expanding the program — and gave us the credibility to deploy real-time guidance in Phase 2 with stakeholder trust already established.
Every month, new GPT capabilities were announced. HCSC stakeholders kept adding "can we also do X?" to the backlog. Without discipline, this would have killed the delivery timeline.
HCSC's IT, Legal, Operations, and Business teams all had different success criteria. IT cared about integration stability. Legal cared about data handling. Operations cared about efficiency metrics. Business cared about cost reduction.
| Decision | What I Chose | Why | Alternative Rejected |
|---|---|---|---|
| Feature sequencing | Low-risk, high-visibility features first | Build trust before deploying high-stakes AI; earn the right to Phase 2 | Flagship feature first — stakeholders wanted it, but failure risk was too high |
| Real-time guidance timing | Deferred to Phase 2 | Insufficient training data; downside risk in healthcare context too high; needed Phase 1 to establish credibility | Ship in Phase 1 — would have exposed program to compliance incident before trust was established |
| AI UX design | Confidence scores + easy override | Skeptical agents adopt when they feel in control, not replaced; override data also improves the model | Clean AI output, no override — faster to build, significantly worse adoption |
| Scope control | Formal 3-step change control | Prevents timeline slippage, protects delivery credibility, keeps trust with HCSC leadership | Accommodate all requests — satisfies short-term stakeholder desire, catastrophic for long-term delivery |
The program became an internal Deloitte case study for GenAI transformation methodology. The TruServe product roadmap work was used in pitches to Fortune 500 clients including Apple Inc.
Every stakeholder wanted AI. Nobody could tell me what problem they wanted it to solve. My job was to translate category excitement into specific, measurable tasks that AI could actually do reliably — and sequence them by the asymmetry of their failure mode.
An AI system that's right 90% of the time but wrong 10% in dangerous ways is worse than a system that does nothing. This shaped every feature prioritization decision. Standard value/effort matrices don't capture this — you need an explicit downside risk column.
The most technically impressive feature had the lowest adoption until we redesigned the UX to make agents feel in control. Feature quality ≠ adoption. You have to design for the skeptic: give them an easy override, show them the model's confidence, let them feel like the AI is assisting them rather than replacing them.
The reason we could expand into higher-risk, higher-value features in Phase 2 is because Phase 1 shipped on time and on budget. The stakeholders who initially wanted real-time guidance in Phase 1 became the biggest advocates for Phase 2 expansion after Phase 1 delivered. Trust is earned before it's spent.