AI Feedback & Insights Agent
AI Feedback & Insights Agent
Willis Towers Watson · Individual Marketplace · 2025–2026
During Medicare Open Enrollment, researchers were spending entire days manually categorizing participant feedback. I realized the bottleneck wasn't the volume, it was the workflow. So I built a system that automated the process, covered the researchers' product blind spots, and reduced synthesis from hours to minutes.

95%
Categorization & synthesis accuracy
8+ hrs → min
Synthesis time
7 days → same day
Insight delivery
My Role
AI Product Design & Systems Design (self-initiated, cross-functional)
Stack
Copilot Studio (GPT-5), Qualtrics API and TextIQ, Azure OpenAI Service, Dataverse, Power Automate
Timeline
2025–2026
Type
Agentic AI workflow, internal tooling
The Context
A billion-dollar book of business generating feedback nobody could keep up with
Willis Towers Watson ran Via Benefits, a private marketplace where individuals, mostly retirees, shop for and enroll in health coverage. It carried an annual book of business exceeding $1B across Medicare, Individual and Family Plans, Dental, Vision, and Hearing, and other ancillary benefits.
Feedback arrived constantly from hundreds of thousands of people: website surveys, mobile app feedback, post-call NPS, and CSAT.
The weekly pass was entirely manual. A researcher pulled the raw feedback out of Qualtrics, categorized each comment, and scrubbed PHI and PII line by line before any of it could be shared. The cleaned summary went to a Teams channel every Monday, covering the week prior. Live access to the raw data required a Qualtrics seat, so everyone else asked a researcher and waited, often a full day for a single question.
The Problem
Three problems that reinforced each other
These three problems didn't happen in isolation. Each one made the other two worse.
The synthesis tax. A researcher spent a full day each week on categorization, cleanup, and redaction. That was 20% of weekly capacity going to work that required domain expertise they didn't have, producing outputs that arrived too late to act on.
The expert gap. Categorization was really a routing decision. Choosing a bucket meant deciding which team owned the problem, and the people making that call weren't subject-matter experts on every product feature. Getting it wrong sent feedback to a team that couldn't act on it while the team that could never saw it. It also corrupted every trend report built on those categories, which is how a real pattern gets buried and a phantom one gets reported.
The distribution lag. Insights sat in a spreadsheet until the Monday post. Feedback arriving on a Monday waited a full week before anyone outside the research team saw it. During Open Enrollment the cycle stayed the same length, but ten times the volume meant ten times as many issues sitting undelivered.
I ran into the expert gap myself. I helped with the feedback process from time to time, and I worked on the Shopping and Quoting team, so I was closer to the product than most researchers were. That proximity still had limits. If a comment touched a known technical limitation in another team's domain, I might file it as a bug when it wasn't one. The person doing the categorization couldn't know what they didn't know. That's structural.
Researcher hours were the smallest part of the cost. Feedback that arrives late and miscategorized is a compliance and retention risk. The signal that something is failing a participant reaches the people who can fix it after the window to fix it has closed.
Interested?