Case Study · GenAI Product Management · Enterprise Delivery

Which GenAI feature do you
build first? I had a framework.

$10M program. 2,000+ agents. A client who wanted "AI in the contact center" but couldn't define what that meant. I built a risk-adjusted prioritization framework — and argued against shipping the feature everyone wanted in order to ship the one that would prove we could be trusted.

Deloitte Digital Product Consultant Nov 2021 – Jan 2024 GenAI · Healthcare · Fortune 500 · $10M Program
60%
Operational Efficiency
Gain (HCSC)
35%
Agent Handle Time
Reduction
78%
AI Draft Acceptance
Rate (Sprint 4)
20%
Ahead of Schedule
on $10M Program
10+
GenAI Features
Shipped
01 Discovery
Mapped 180+ agent tasks
Scored each task on automation potential × business impact × downside risk of wrong AI output
02 Framework
Risk-adjusted prioritization
Explicit downside risk weighting — the most-wanted feature had the highest failure cost
03 The Bet
Email drafting before real-time assist
Argued against the headline feature. 78% acceptance rate in Sprint 4 became proof of concept
04 Scale
Phase 1 earned Phase 2
Delivery credibility unlocked the high-risk, high-value features everyone wanted from the start
business_center Context
The Platform, The Client, The Problem

TruServe — The Product

Deloitte's Salesforce-based platform for enterprise contact center management. Handles ticket routing, agent workflows, post-call documentation, and CRM integration. Used by Fortune 500 clients including Apple Inc. and HCSC across multiple verticals.

HCSC — The Client

Health Care Service Corporation — one of the largest US health insurers. ~1 million customer interactions per year across 2,000+ agents in three business units: individual plans, employer benefits, and provider relations. The program goal: "use AI to improve agent efficiency."

That's not a product brief. That's a category. My first job was to turn "improve agent efficiency with AI" into specific, measurable, shippable features — in the right order. That required understanding the problem before touching a feature list.

The Problem
Everyone Wanted AI. Nobody Could Define It.

HCSC leadership had been briefed on GenAI capabilities and arrived with a vision: an AI assistant that could guide agents in real time during customer calls. It was compelling, impressive, and — by my analysis — the wrong place to start.

The category excitement around GenAI created a specific PM challenge: the most visually impressive features were also the highest risk. Shipping the flashy feature first, before establishing delivery credibility, would have exposed the program to catastrophic failure in a regulated healthcare environment.

The Stated Goal

"Use AI to improve agent efficiency across 2,000+ agents handling 1M+ customer interactions per year."

The Real Challenge

In healthcare, wrong AI output during a live call = compliance liability + patient safety risk. The cost of being wrong is asymmetric. A system that's right 90% of the time but wrong 10% in dangerous ways is worse than no AI.

The Key Insight About Users

Agents were skeptical — they'd seen automation promises fail before. They'd experienced suggestions that were wrong or tone-deaf. Designing AI features for the enthusiast (who trusts AI by default) produces different products than designing for the skeptic (who needs to feel in control). We designed for the skeptic. That's why adoption worked.

search Discovery
Mapping 180+ Tasks to Find the Real Opportunity
1

Agent task mapping workshops with HCSC ops leads

Ran structured discovery sessions across all three business units. Mapped 180+ distinct agent tasks. For each task, scored: automation potential (can AI reliably do this?), business impact (what's the cost if this task is slow or wrong?), and data readiness (how much labeled data exists?).

Duration: 3 weeks · Method: Structured workshops, task taxonomy mapping
2

The contrarian finding

The most-requested feature (real-time agent guidance during live calls) scored highest on business impact — but also highest on downside risk. The least-requested features (post-call documentation, email drafting) scored highest on ROI per unit of effort because their failure mode was benign: agent reviews before anything goes out.

Finding: High excitement ≠ high priority when failure cost is asymmetric
filter_list Prioritization
The Risk-Adjusted Value Framework

Standard GenAI prioritization frameworks optimize for impact and effort. They miss the critical third dimension for regulated environments: the asymmetric cost of being wrong.

For each proposed AI feature, I added four explicit questions before scoring:

Risk-Adjusted Value Framework — Applied to HCSC Features

Each feature scored on: (1) Upside if right, (2) Downside if wrong, (3) Data readiness, (4) Measurability

Feature Upside if Right Downside if Wrong Data Ready? Priority
Smart email drafting Saves 8–12 min/email, high volume Low — agent reviews before send Yes — email corpus available P0
Post-call documentation 12–18 min saved per call, QA coverage Low — agent confirms before submitting Yes — call transcripts available P0
QA sentiment flagging QA coverage from <10% to 80%+ Medium — false positives waste supervisor time Partial — needs calibration P1
Escalation routing Faster resolution, right agent every time Medium — wrong routing = customer frustration Partial — historical routing data P1
Real-time agent guidance Very high — correct answers during live call High — wrong guidance during live call = compliance liability, patient harm No — needs extensive HCSC-specific training P2 → Phase 2

The framework reveals the counterintuitive truth: the feature with the highest upside has the highest downside and the lowest data readiness. Shipping it first maximizes the chance of failure in a healthcare environment where wrong = compliance incident.

The Bet
Arguing Against the Feature Everyone Wanted
The Non-Obvious Call

Defer real-time agent guidance to Phase 2. Ship email drafting first.

HCSC leadership had been sold on the real-time agent assistant vision. It was the demo they'd seen, the vision they'd bought into, and the feature they expected in Sprint 1. My recommendation was to defer it to Phase 2.

My argument: "Let's earn the right to do the high-risk feature by proving we can nail the low-risk ones first. Phase 1 is about building delivery credibility. If we ship smart email drafting on time and it performs, you'll trust us with the feature that matters most. If we rush real-time guidance and it fails, the entire program's credibility is at risk in a regulated environment."

What I risked: Stakeholder frustration in the short term. If Phase 1 underdelivered, deferring the headline feature would look like the wrong call in retrospect.

Sprint 4 result: 78% AI email draft acceptance rate.

Agents accepted the AI-generated draft without significant edits 78% of the time. That number became the proof point for expanding the program — and gave us the credibility to deploy real-time guidance in Phase 2 with stakeholder trust already established.

engineering Execution
The Hardest Problems Weren't Technical
1
Scope creep from AI excitement

Every month, new GPT capabilities were announced. HCSC stakeholders kept adding "can we also do X?" to the backlog. Without discipline, this would have killed the delivery timeline.

Solution: Formal change control process — any scope addition required 3-step evaluation (timeline impact, budget impact, integration complexity) before consideration. Prevented 3 significant scope expansions that would have delayed delivery by 6–10 weeks.
2
Adoption was a product problem, not a training problem

HCSC's IT, Legal, Operations, and Business teams all had different success criteria. IT cared about integration stability. Legal cared about data handling. Operations cared about efficiency metrics. Business cared about cost reduction.

Solution: Monthly joint reviews with all four functions. Slow? Yes. Did it prevent a compliance-related launch delay that would have cost 6 weeks? Also yes. Cross-functional alignment is product work, not admin overhead.
Key Decisions
The Tradeoffs, Made Explicit
DecisionWhat I ChoseWhyAlternative Rejected
Feature sequencing Low-risk, high-visibility features first Build trust before deploying high-stakes AI; earn the right to Phase 2 Flagship feature first — stakeholders wanted it, but failure risk was too high
Real-time guidance timing Deferred to Phase 2 Insufficient training data; downside risk in healthcare context too high; needed Phase 1 to establish credibility Ship in Phase 1 — would have exposed program to compliance incident before trust was established
AI UX design Confidence scores + easy override Skeptical agents adopt when they feel in control, not replaced; override data also improves the model Clean AI output, no override — faster to build, significantly worse adoption
Scope control Formal 3-step change control Prevents timeline slippage, protects delivery credibility, keeps trust with HCSC leadership Accommodate all requests — satisfies short-term stakeholder desire, catastrophic for long-term delivery
trending_up Results
What the Program Delivered
60%
↑ Improvement
Operational Efficiency
Gain (HCSC)
35%
↓ Reduction
Average Agent
Handle Time
78%
Acceptance Rate
AI Email Draft
Accepted by Agents
20%
Ahead of Schedule
$10M Program
Delivery
10+
Shipped
GenAI Features
in Production
0
Clean
Compliance Incidents
During Deployment

The program became an internal Deloitte case study for GenAI transformation methodology. The TruServe product roadmap work was used in pitches to Fortune 500 clients including Apple Inc.

Lessons Learned
What I Carry Forward
Lesson 01

"AI" is not a product requirement. A job-to-be-done is.

Every stakeholder wanted AI. Nobody could tell me what problem they wanted it to solve. My job was to translate category excitement into specific, measurable tasks that AI could actually do reliably — and sequence them by the asymmetry of their failure mode.

Lesson 02

In regulated industries, the cost of being wrong is asymmetric.

An AI system that's right 90% of the time but wrong 10% in dangerous ways is worse than a system that does nothing. This shaped every feature prioritization decision. Standard value/effort matrices don't capture this — you need an explicit downside risk column.

Lesson 03

Adoption is a product problem, not a training problem.

The most technically impressive feature had the lowest adoption until we redesigned the UX to make agents feel in control. Feature quality ≠ adoption. You have to design for the skeptic: give them an easy override, show them the model's confidence, let them feel like the AI is assisting them rather than replacing them.

Lesson 04

Delivery credibility is currency. Spend it carefully.

The reason we could expand into higher-risk, higher-value features in Phase 2 is because Phase 1 shipped on time and on budget. The stakeholders who initially wanted real-time guidance in Phase 1 became the biggest advocates for Phase 2 expansion after Phase 1 delivered. Trust is earned before it's spent.