Cases / Applied AI · 2026
Synthetic AI Focus Groups: Project Avatar
A SaaS platform where AI persona agents walk through a digital product and report where and why they drop off — taken from discovery to a closed-cohort pilot.

Context
Real user research (interviews, usability tests, beta cohorts) is slow and expensive, so teams skip it, run it too late, or run it on too few people to trust. I wanted a fast, cheap layer between “we have an idea” and “we have real users”.
The idea: instead of one big prediction model, run many independent LLM agents. Each agent is assembled from modular persona “skills” — core identity, decision style, domain literacy, emotional triggers, communication style — walks through a product on its own, and logs a structured trace of what it did and why.
My role
I own the product end to end: vision, discovery, scope, the roadmap and every trade-off. Engineering is done with AI coding agents under my specification and review.
What I did
- Discovery. Analysed an existing synthetic-user tool and ran a head-to-head on a live retail website. The comparison surfaced an architectural decision: a lightweight mode (one page load plus static analysis) and a heavy interactive mode (a real multi-step browser session) carry very different anti-bot risk, so they are built as two separate capabilities.
- A falsifiable PoC. Scoped the first version to a single hypothesis: agents built from modular persona-skills behave meaningfully differently on the same product, and the variation comes from the persona, not from model noise. Five grocery-shopping personas, unauthenticated browsing, single sessions, and an explicit list of what is out of scope.
- Maturity ladder. Defined four stages — concept interview, prototype click-through, live-product multi-day simulation, validation against real data — each with its own interaction mode and output.
- Validation as the trust product. Designed validation as a pluggable provider (panel vendors early, the client’s own analytics later) and started a Model Failures Log of concrete, checkable reasoning errors as calibration data. LLM agents are more rational than real users; the product has to say so rather than hide it.
- Release 1 for a closed cohort. Admin-provisioned accounts instead of self-service sign-up, per-user data isolation, a usage quota, an explicit consent step before any URL is tested, protection against browser jobs reaching internal addresses, human-readable error states, and real token-usage logging for every model call.
- Product decisions with cost in mind. For bring-your-own API keys I chose a policy where users without a key run on the platform key under a quota, while an own key lifts token units but not server capacity. Estimates are tracked against actuals per task; work that touches authentication consistently took at least twice the visible scope.
- Release discipline. Separate staging environment with a deploy script, smoke tests and a spend-capped key; production only after staging passes.
Result
- A working MVP with persona and project management, five testing modes (concept brief, image review, lightweight page test, guided and interactive browsing) and a supervisor layer: contradiction check, live-site verification, executive summary, reproducibility confidence and PDF export.
- Measured cost per persona run, from small samples: roughly 2,400 tokens for an image review up to about 35,800 for a full interactive session. This is the basis for the usage quota.
- Release 1 hardening is in progress on staging; the next step is a walkthrough with outside testers. Demo available on request.