Skip to main content

AI Feedback & Insights Agent

Agentic Workflow DesignAI DesignUX Research

AI Feedback & Insights Agent

Willis Towers Watson · Individual Marketplace · 2025–2026

During Medicare Open Enrollment, researchers were spending entire days manually categorizing participant feedback. I realized the bottleneck wasn't the volume, it was the workflow. So I built a system that automated the process, covered the researchers' product blind spots, and reduced synthesis from hours to minutes.

A product manager querying the Participant Listening Agent in Microsoft Teams, shown on a laptop. The agent reports the top trending issue as Plan Comparison Logic at 38% of 256 comments, with representative participant quotes.

95%

Categorization & synthesis accuracy

8+ hrs → min

Synthesis time

7 days → same day

Insight delivery

My Role

AI Product Design & Systems Design (self-initiated, cross-functional)

Stack

Copilot Studio (GPT-5), Qualtrics API and TextIQ, Azure OpenAI Service, Dataverse, Power Automate

Timeline

2025–2026

Type

Agentic AI workflow, internal tooling

A billion-dollar book of business generating feedback nobody could keep up with

Willis Towers Watson ran Via Benefits, a private marketplace where individuals, mostly retirees, shop for and enroll in health coverage. It carried an annual book of business exceeding $1B across Medicare, Individual and Family Plans, Dental, Vision, and Hearing, and other ancillary benefits.

Feedback arrived constantly from hundreds of thousands of people: website surveys, mobile app feedback, post-call NPS, and CSAT.

The weekly pass was entirely manual. A researcher pulled the raw feedback out of Qualtrics, categorized each comment, and scrubbed PHI and PII line by line before any of it could be shared. The cleaned summary went to a Teams channel every Monday, covering the week prior. Live access to the raw data required a Qualtrics seat, so everyone else asked a researcher and waited, often a full day for a single question.

Three problems that reinforced each other

These three problems didn't happen in isolation. Each one made the other two worse.

The synthesis tax. A researcher spent a full day each week on categorization, cleanup, and redaction. That was 20% of weekly capacity going to work that required domain expertise they didn't have, producing outputs that arrived too late to act on.

The expert gap. Categorization was really a routing decision. Choosing a bucket meant deciding which team owned the problem, and the people making that call weren't subject-matter experts on every product feature. Getting it wrong sent feedback to a team that couldn't act on it while the team that could never saw it. It also corrupted every trend report built on those categories, which is how a real pattern gets buried and a phantom one gets reported.

The distribution lag. Insights sat in a spreadsheet until the Monday post. Feedback arriving on a Monday waited a full week before anyone outside the research team saw it. During Open Enrollment the cycle stayed the same length, but ten times the volume meant ten times as many issues sitting undelivered.

I ran into the expert gap myself. I helped with the feedback process from time to time, and I worked on the Shopping and Quoting team, so I was closer to the product than most researchers were. That proximity still had limits. If a comment touched a known technical limitation in another team's domain, I might file it as a bug when it wasn't one. The person doing the categorization couldn't know what they didn't know. That's structural.

Researcher hours were the smallest part of the cost. Feedback that arrives late and miscategorized is a compliance and retention risk. The signal that something is failing a participant reaches the people who can fix it after the window to fix it has closed.

Interested?

Want to talk through the methodology or the build?

Next Case Study

People-First Enrollment Redesign

Read case study