EVIDENCE INDEX·METHODOLOGY·11 MIN READ

We re-ran two of our own studies with synthetic panels. Here are the numbers.

Two replica studies — a B2B channel survey and a consumer concept test — fielded again with synthetic personas and scored against the original human respondents on distribution distance. What matched, what collapsed, and what calibration actually fixed.

#synthetic#methodology#ai-research

Grade your own homework, in public

In May we published a decision matrix for synthetic research. Its core rule: treat every synthetic output as a hypothesis until a human benchmark says otherwise. Fair question back: had we benchmarked our own synthetic panel that way?

We had. This post is the readout.

The method is a replica study. Take a study we already fielded with real respondents. Run the identical questionnaire against synthetic personas. Compare answer distributions question by question. No demo data, no cherry-picked subset — the same instrument, scored against the humans who already answered it.

We did this twice, in two different worlds:

Replica 1 — B2B. A commercial-refrigeration channel study (the one behind our published refrigeration sample report). Benchmark cohort: the 67 human respondents who passed every trust gate — engagement scoring plus zero picks on trap brands (real companies that don't make this equipment; anyone "familiar" with them is answering on autopilot). Synthetic cohort: 149 personas from our refrigeration-distribution pool.

Replica 2 — consumer. A licensed-toy concept test with 220 usable human respondents (also on this blog). Synthetic cohort: 220 personas, quota-matched on the axes the study screened for.

Same questionnaire. Same options. One panel human, one synthetic. Then measure the gap — per question, as a distribution.

Why we score distributions, not 'accuracy'

Most synthetic vendors publish a single match percentage. We think that number hides more than it shows. A synthetic panel can nail the average and still be useless — the insight lives in the spread, and the spread is what models erase first.

So the primary score here is Total Variation Distance (TVD): 0 means the synthetic and human answer distributions are identical; 1 means they don't overlap at all. It punishes exactly the failure a headline percentage forgives — everyone piling onto one plausible answer. For ranking questions we report mean absolute rank error instead: how far each item lands from where the humans put it.

What matched out of the box

The behavioral and structural layers held up well before any calibration.

SurfaceHuman vs syntheticRead
Customer-mix allocation (B2B, constant-sum)QSR 20 vs 27 · convenience 25 vs 18 · hotels 11 vs 15 · grocery 15 vs 16Within 2–7 points on every segment
Premium-brand ratings matrix (B2B)sales strength 3.99 vs 4.01Near-exact on the strongest axis
Exclusivity posture (B2B)TVD 0.205Same option ordering
Brand familiarity (consumer)near-exact distribution, 8 brandsPinned to persona facts
Purchase frequency (consumer)distributions matchGrounded in stated behavior

Also worth reporting: zero trap-brand picks across all 149 B2B synthetic respondents, zero piped-question violations, zero constant-sum violations. The synthetic panel is, if anything, more internally consistent than the humans — real distributors rate almost everything 3.5–4.2 and trip traps at high rates; only 67 of the study's 144 usable human respondents passed every trust gate.

What broke, specifically

Now the part vendors don't publish. Before calibration, the attitudinal layer failed in three repeatable ways.

Central-tendency collapse. On "how likely are you to evaluate new brands," 51% of human distributors said "very likely." The synthetic panel put 100% on the same hedge option — "not actively evaluating, but open." TVD 0.881, the worst score we measured anywhere. A second attitude question collapsed to 72% "neutral" (TVD 0.584).

Rank lock. Asked to rank 11 purchase criteria, the synthetic panel produced textbook logic — reliability first, nearly unanimously — while humans rank by operational reality (availability first). Mean absolute rank error: ~2.4 positions per item.

Scale floor. In the consumer replica, the 1–10 interest scale came back with a mean of 2.3 — 191 of 220 synthetic respondents answered "1" on purchase likelihood (stdev 0.36). The model rated a parent's personal taste for the category in a vacuum, ignoring that the same persona buys these toys monthly. The synthetic report then contradicted itself: "passionate demand" in the narrative, 1.1/10 in its own tables.

One rule explains all three, plus a gender inversion and an income collapse we also caught: anything we did not explicitly pin or quota drifted to the pool's defaults. That rule now governs how we design every synthetic run.

Anything we did not explicitly pin or quota drifted to the pool's defaults. That sentence cost us a month, and it now governs every run.

What calibration fixed, with numbers

The fix is not a better prompt. It's an engine change: pin every factual answer to the persona's stored attributes, ground attitude questions in the persona's own behavior, and move the final answer draw out of the model into seeded code — so the distribution is targeted, auditable, and reproducible bit-for-bit.

FailureBeforeAfter calibration
Worst B2B attitude question (TVD vs humans)0.881 — 100% on one option0.113 — 60% top share, real spread. Fixed.
Second B2B attitude (TVD)0.584 — 72% neutral0.266 observed at n=20; 0.095 expected. Mechanism verified; the n=20 draw is a P=.039 tail, reproducible by design.
Ranking lock (mean abs rank error, 11 items)~2.4 — unanimous #10.482 — 20/20 distinct orderings, availability back to rank ~2 (humans: 2.9). Fixed.
Untouched control question (TVD)0.2150.160 — improved incidentally.
Consumer scale floor (1–10 scales)interest mean 2.3; purchase likelihood 191 of 220 at “1”mean 7.71, 0 of 220 at floor; recent buyers 8.05 vs lapsed 6.87. Fixed.

Two honesty notes on that table. The 0.266 row is a mechanism that provably targets 0.095 but drew a tail at n=20 — we report the observed number anyway, because that's the point of a readout. And the consumer scales have no human benchmark: the real study never asked those questions. What we claim there is the elimination of an internal incoherence — the floor artifact — verified by the ratings now grading correctly with each persona's actual buying behavior.

What we still don't claim

This readout is internal. No third party has audited it. The acceptance runs behind the calibration table are n=20; the full replica cohorts are 67 (trusted B2B humans) and 220 (consumer). Two domains — refrigeration distribution and licensed toys — not a universal warranty. Distribution match on attitude questions does not make synthetic respondents valid for pricing commitment, novel categories, or the go/no-go — our own decision matrix still routes those to humans.

The next step is structural: a per-study evidence card attached to every report — which dimensions were calibrated against human data, at what distance, with what confidence, and what was not covered. Ships Q3 2026. Until then, the full replica protocol, cohort definitions, and per-question tables are available to any evaluating customer on request.

The receipt

Synthetic panels fail in specific, measurable, fixable ways — and the fixes hold when scored against real respondents. But you only learn that by running the replica and publishing the misses next to the hits. A match percentage without the misses is marketing.

Our position stands: synthetic is the best pre-flight tool this field has ever had, and it earns exactly the decisions its evidence supports. This is what that evidence looks like when someone shows the work.

See the evidence trail in your own data.

A 30-minute working session on a real study of yours — no slides, no demo theater.

Book a working session