Synthetic Respondents Concept Testing: The Fidelity Data
Synthetic respondents concept testing measured against human baselines: where the 70 and 85 percent fidelity numbers come from and what they let you decide.
Eleven packaging routes are taped to a meeting room wall. There is budget for three. Somebody has to say which three, nobody actually knows, so the loudest opinion wins and the other eight ideas are never tested. That meeting is why synthetic respondents concept testing now gets a hearing in consumer goods teams that were skeptical two years ago.
Key Takeaways
- Consumer goods teams still narrow concepts in a meeting, because the research cycle cannot run eleven times on one launch calendar.
- Synthetic personas built from five demographic fields reach roughly 70 percent accuracy against human behavior.
- Built from dense qualitative input with a memory and reflection architecture, they reach roughly 85 percent fidelity against a human test and retest baseline.
- Across five PNAS studies, a simulation replicated four correctly and predicted the failure of the fifth, later confirmed against 1,000 real humans.
- Synthetic respondents are a screening instrument. The finalists still go to a human sample.
What concept pre-testing actually is
Concept pre-testing is the research step that measures reaction to an idea while the idea is still cheap to change, using a defined target audience and a comparable stimulus for each route being considered. In consumer goods that means packaging routes, claim wording, a new flavor, a campaign line. The field calls it screening when the job is narrowing many ideas and validation when the job is confirming the few that survived. The methodological vocabulary, including monadic and sequential monadic designs, is documented by ESOMAR.
The mechanics are unglamorous. Someone writes the concept cards. Someone else decides how many routes a respondent can absorb before fatigue distorts the answers, which caps a monadic test at three or four. That cap is where the problem starts, because the innovation funnel does not produce three ideas. It produces thirty.
How concept screening is done today
PERSONAL EXPERIENCE This is the sequence we watch run inside enterprise insight teams, and it has barely changed in fifteen years.
- Collect the routes from the innovation team, usually in a slide deck with no consistent format.
- Cut the list down in a stakeholder meeting, on judgment, to what the budget carries.
- Write the concept cards and the screener for the surviving routes.
- Program the questionnaire, then walk every path by hand to catch a broken skip before field opens.
- Recruit the sample, one to two weeks for a general population target and longer for a low incidence one.
- Run field for another two weeks.
- Code the open ends, build the tables, write the readout.
Step two is the step nobody documents. It happens before any instrument exists, it leaves no trace in the report, and it is the most consequential decision in the whole study. Eight of the eleven routes die there, in a room, on instinct. The research that follows is rigorous about the three that survived and silent about the eight that did not.
Step four has its own literature, treated separately in [Testing a Questionnaire Before Field](/blog/survey-testing-before-fieldwork), and step seven in [From Transcript to Recommendation Without Losing the Evidence](/blog/qualitative-coding-evidence-chain). Cycle time pressure on concept testing is a standing beat at Quirk's.
What the narrowing costs and who answers for it
UNIQUE INSIGHT The cost of concept screening is not the fieldwork invoice. It is the untested eight, and specifically the possibility that the winner was among them.
Look at the accounting. The invoice for a three cell monadic test is a line item the insights director can defend. The eight cut routes cost nothing on paper, so they never enter the conversation about whether the study was worth it. Then the launch underperforms and the post mortem asks whether the research was wrong. The research was fine. It answered a question about three routes chosen by people who had no data when they chose them.
You only launch once. A wrong bet on a national rollout costs millions and burns a year of calendar, and the person who signed off on the three finalists answers for it. Innovation failure rates and their drivers are tracked in the industry analysis published by NIQ.
Everyone in that meeting knows step two is guesswork. It stays that way because the alternative, eleven monadic cells, does not fit the launch calendar.
What a synthetic respondent is
Synthetic respondents are AI-generated participants built from qualitative material about a defined audience, used to simulate how that audience would react to a concept before it goes to field. They stand in for human respondents during early screening. They do not replace the human sample that validates the finalists.
Two terms travel with them, and both get used loosely in vendor decks.
A digital twin, in consumer research, is one synthetic respondent built to correspond to a specific real person or documented segment, carrying that person's stated attitudes, history and constraints rather than a demographic average.
A memory and reflection architecture is the design that gives a synthetic respondent a stored record of its own past statements, plus a periodic step where it summarizes that record into higher level conclusions about itself. Without it, the model answers each question from scratch and contradicts itself across a session. The version our simulation derives from was published by Stanford and Google researchers in Generative Agents: Interactive Simulacra of Human Behavior.
That is why the gap between a persona prompt and a simulated agent is real rather than marketing. A persona prompt is a paragraph. An agent has three parts:
- A memory stream, which stores what the agent said and did, retrievable by relevance rather than recency alone.
- A reflection engine, which periodically reads that stream and writes conclusions about the agent's own preferences and constraints.
- A planning module, which turns those conclusions into intended behavior, so the agent acts consistently instead of answering each prompt in isolation.
Running several models instead of one has its own consequences, handled in [One Model, One Bias](/blog/single-model-bias-market-research).
What the fidelity numbers say
ORIGINAL DATA Fidelity depends almost entirely on what goes in. Measured internally against a human test and retest baseline, personas built from five demographic fields reach roughly 70 percent accuracy, and personas built from dense qualitative input with a memory and reflection architecture reach roughly 85 percent. Source: Cassi.ai What-IF Machine benchmark, 2026.
| Approach | Fidelity to human behavior | Input required | Best used for |
|---|---|---|---|
| Five demographic fields | ~70% | Minutes | Rough directional reads |
| Dense qualitative input with memory and reflection | ~85% | 2-hour depth interview transcripts | Screening many concepts before field |
| Real human sample | Baseline | Weeks of recruitment and full field cost | Validating the finalists |
Source for the two synthetic rows: Cassi.ai What-IF Machine benchmark, 2026. The human row is the reference the other two are measured against.
Read the 70 percent row carefully, because it is the number most teams have experienced. "Woman, 34, São Paulo, married, two children" produces a stereotype, and a stereotype agrees with whatever the concept says. That is the version of this technology that burned a lot of insight directors, and their skepticism is earned.
The second number needs a harder test than self-report, so we ran one. ORIGINAL DATA Across five published PNAS studies, the simulation replicated four correctly and predicted the failure of the fifth, later confirmed against a sample of 1,000 real humans. Effect size correlation ranged from 0.85 to 0.90. Source: Cassi.ai What-IF Machine benchmark, 2026.
Predicting a failure matters more than replicating a success. A simulation that agrees with everything is useless for screening, because screening is the act of saying no to eight things.
Persona fidelity to human behavior
Persona fidelity to human behavior. horizontal bar data: Qualitative input 85; Demographics only 70.Source: Cassi.ai What-IF Machine benchmark 2026.
Where Cassi.ai comes in
Cassi.ai is a software engineering company specialized in the pains of market research, innovation and insights, working with enterprise research teams and agencies across banking, health, retail, media and consumer goods. The What-IF Machine is the module built for the step two problem above.
PERSONAL EXPERIENCE The sequence we run on a client study:
- Ingest the qualitative material, typically 2-hour depth interview transcripts from the brand's own existing studies.
- Build the digital twins as biopsychosocial personas, carrying attitudes, constraints and history rather than demographic fields.
- Run the scenario against every route, including message pre-tests, packaging reactions and crisis response.
- Rank the full set and narrow it to the finalists, with the reasoning attached to each cut.
- Validate those finalists with a real human sample, in the normal way.
Step four is where the cut becomes documented. When the launch is reviewed a year later, there is a record of why eight routes were dropped, which the meeting room version never produced.
We govern this with a four level ladder of trust: possibilities, qualitative attitudes, quantitative reads, macroeconomic simulation. Most client work sits at the first two, and we say so out loud, because claiming the third on a screening instrument is how this field loses credibility.
When to use which approach
UNIQUE INSIGHT The decision turns on two things only: how many routes you have, and what the output will be used to decide. Belief in the technology has nothing to do with it.
| If your situation is | Then | Because |
|---|---|---|
| Three or four routes, budget for all of them | Test with humans directly | There is nothing to narrow |
| Ten or more routes, budget for three | Screen synthetically, then validate | The gain is in the eight you can now examine |
| Recurring innovation waves | Build the twins once, reuse across waves | The qualitative ingestion cost amortizes |
| Regulated claim or legal submission | Human sample only | Auditability of the respondent base is the requirement |
| No existing qualitative material | Run depth interviews first | Demographic fields give you the 70 percent version |
The last row is the one teams try to skip, and skipping it is how you end up with the stereotype problem. The architecture is only as good as what it read.
FAQ
Do synthetic respondents replace human research?
No. They shorten the first pass and widen how many ideas get examined. The decision on the finalists stays with a human sample. UNIQUE INSIGHT Any vendor claiming replacement is describing a product that will not survive a launch review.
How accurate is synthetic respondents concept testing?
ORIGINAL DATA Roughly 85 percent fidelity against a human test and retest baseline when built from dense qualitative input with a memory and reflection architecture. Roughly 70 percent when built from five demographic fields. Source: Cassi.ai What-IF Machine benchmark, 2026.
How much qualitative input does a reliable persona need?
The 85 percent figure comes from 2-hour depth interview transcripts. Thinner material moves the result toward the demographic baseline. Studies the brand already commissioned are usually enough, per the Cassi.ai portfolio.
Can synthetic respondents be used for regulatory submissions?
No. Regulated claims and legal submissions require a human sample, because auditability of who answered is the requirement, and a simulated respondent has no auditable identity. Guidance on respondent provenance is maintained by ESOMAR.
What makes this different from a persona prompt in ChatGPT?
A persona prompt answers each question independently. A simulated agent carries a memory stream, a reflection engine and a planning module, which is what produces consistency across a session. The architecture is described in Generative Agents: Interactive Simulacra of Human Behavior.
The eight routes that die in a meeting room are the most expensive research your company never ran. Nobody invoices for them, and nobody can tell you afterward whether the winner was among them.
If your last innovation wave started with more than ten routes and went to field with three, ask who made that cut and what they had in front of them when they made it. If you want to see the simulation run on routes from a study you already own, we will build it with your own qualitative material.
Published by Cassi.ai. Read the full article at https://www.cassiai.com/blog/synthetic-respondents-concept-testing.