A synthetic panel is a group of AI-simulated survey respondents used in place of, or alongside, a traditional research panel. Instead of recruiting real people over weeks, you describe an audience, an AI platform generates simulated panelists, and those panelists answer your questions in minutes.
Every synthetic panel produces the same kind of output: clean percentages, confident rationale, professional charts. What separates a synthetic panel you can act on from one that manufactures plausible fiction is a single question: where does the data behind it come from?
This guide explains how synthetic panels work, why the data source decides whether the results are trustworthy, and what to ask before you bring AI-generated research into a real decision.
Jump to a section
- How does a synthetic panel work?
- Why the data behind a synthetic panel is everything
- Are synthetic panels accurate?
- How StatSocial Digital Twins answers the data question
- Related terms: AI personas and synthetic data
- Side-by-side comparison
- Questions to ask any synthetic panel vendor
- Put the data foundation to the test
- FAQ
How does a synthetic panel work?
Most synthetic panels follow the same recipe. A large language model (LLM) receives a prompt describing a respondent, often a set of demographic traits like age, gender, income, and region. The model then generates the answers it predicts that person would give. Repeat a few hundred times and you have a panel.
The appeal is obvious. No recruiting costs. No six-week fielding timelines. No panel fraud, straight-lining, or professional survey takers. Hard-to-reach audiences become as easy to field as the general population.
But notice what is actually happening in that recipe. The respondents are not sampled from anywhere. They are generated from a description. Which means the panel is only as real as the data that shaped it, and in most cases, that data is nothing more than the model’s general knowledge filtered through a demographic prompt.
Why the data behind a synthetic panel is everything
The AI layer is nearly identical across every platform in this category. Everyone has access to the same frontier language models (Claude, ChatGPT, etc.). The intelligence generating the answers is not the differentiator. The data grounding those answers is.
When a synthetic panel tells you that 42 percent of your audience would try a new product, that number came from somewhere. There are only three possibilities, and they produce very different levels of trust:
- The model’s general training data. The AI answers from what it learned about people in general, shaped by a demographic description. This is the foundation of most synthetic panels. It produces answers that sound like your audience but were never derived from your audience. Demographics tell you who someone is on paper. They do not tell you what that person follows, buys, watches, or believes.
- A seed dataset. Responses are extrapolated from an existing survey or customer file. Better, but the output can never contain more signals than the seed. Gaps and biases in the source become gaps and biases in every generated answer, and the data ages the moment it is collected.
- Observed behavior from real people. Answers are modeled from what a defined set of real individuals actually follows, watches, reads, and buys, refreshed continuously. This is the only foundation where the data describes your specific audience rather than a general approximation of it.
The stakes are not theoretical. Industry publication Quirk’s Media covered academic research on synthetic respondents showing that LLM-simulated respondents built on demographic profiles alone performed poorly at predicting real-world outcomes, and that accuracy depended heavily on whether the underlying inputs included behavioral and attitudinal information rather than demographics alone.
In other words, the research community keeps arriving at the same conclusion: the model matters far less than the data. A synthetic panel without real behavioral data underneath it is a very fast way to generate a stereotype.
Are synthetic panels accurate?
It depends entirely on the foundation, and you should never accept an accuracy claim without a published benchmark.
Ungrounded panels, the kind built on demographic prompts, tend to produce answers that regress to the average. They flatten the exact differences between audiences that research exists to find. They can also exhibit what researchers call hyper-accuracy distortion, where outputs look impossibly precise and consistent in ways real humans never are.
Grounded panels are a different story. When each simulated respondent is anchored to the observed behavior of real people, accuracy becomes measurable, and it can be benchmarked directly against real-world survey results. That benchmark, stated as a published error rate against real human data, is the single clearest signal that a vendor’s data foundation is real.
Which brings us to how StatSocial approached the problem.
How StatSocial Digital Twins answers the data question
StatSocial Digital Twins delivers what synthetic panels promise, speed and reach, on the foundation most of them lack: real people and observed behavior.
Digital Twins are anonymized, AI-modeled representations of real audiences, built from StatSocial’s patented PeopleGraph and KnowledgeGraph, which span 150M+ U.S. adults. The data comes from cross-platform public social signals, connected with household and offline data. Each twin is grounded in hundreds of observed behavioral attributes per person: media consumption, influencer affinities, interests, professional signals, and demographics. Nothing is invented, and nothing is extrapolated from a thin seed. The behavior was observed before the question was ever asked.
That data-first design was the founding premise of the product. As StatSocial CEO David Barker said in the Digital Twins launch announcement: “Most AI-driven market research today relies on synthetic personas that aren’t grounded in real human behavior. We took a different approach. Digital Twins starts with real groups and the behavioral signals tied to what they follow, watch, read, and buy.”
The data foundation shows up in three ways that a demographic-prompt panel cannot replicate:
Measured accuracy. Across more than 40 benchmark studies, Digital Twins has averaged 3.3 points mean absolute error against real-world survey results, compared to 5 to 6 points MAE for typical opt-in online panels. Accuracy is not asserted. It is measured against sources including U.S. Census and Pew Research benchmarks and continuously re-tested.
Reach into audiences that panels cannot recruit. Because twins are built from behavioral signals rather than recruitment, Digital Twins can survey low-incidence and hard-to-reach groups, from niche fan communities to investor segments to a specific creator’s followers. If the behavioral signal exists in PeopleGraph, the audience can be fielded.
Transparent, weighted answers. Digital Twins returns quantitative results alongside a written rationale explaining how the audience is likely to think. And with StatSocial Focus Groups, the same weighted audience becomes conversational, with every voice traceable to the share of real buyers it represents. You can read more in our post on AI focus group platforms built on real audiences.
Related terms you will hear: AI personas and synthetic data
Two adjacent terms get used interchangeably with synthetic panels. They should not be, and the data question separates them cleanly.
AI personas are fictional characters role-played by a language model: “You are Sarah, a 34-year-old suburban mom who cares about sustainability.” Where the data comes from: nowhere. A persona has no data source beyond the prompt and the model’s general knowledge. Personas are useful for brainstorming and rehearsing objections, and nothing more. There is no way to know how many real buyers, if any, a persona represents.
Synthetic data is the umbrella term for data that is generated rather than collected. Instead of asking real people questions, software studies an existing dataset, like a past survey or customer file, and produces new records that follow the same patterns. Synthetic data has real uses. It can fill out a small sample, protect people’s privacy, and add responses for groups a survey did not reach enough of. But it is a copy of a copy. It can stretch the information that already exists in the original data. It cannot know anything the original data never knew.
A synthetic panel built on either of these foundations inherits their limits. A panel built on observed behavior does not.
Synthetic panel approaches compared
| Typical Synthetic Panel | AI Personas | Synthetic Data | StatSocial Digital Twins | |
|---|---|---|---|---|
| Where the data comes from | LLM training data plus demographic prompts | The prompt; no data source | A seed dataset or generative model | Observed cross-platform social, household, and offline data on 150M+ U.S. adults |
| Grounded in real people? | Rarely | No | Sometimes, via seed data | Yes |
| Audience specificity | Broad demographic segments | Whatever you describe | Limited by source data | Any audience with behavioral social signals, including niche and low-incidence groups |
| Accuracy validation | Varies widely; often unpublished | None | Depends on methodology | 3.3 MAE across 40+ benchmark studies vs. 5 to 6 for opt-in panels |
| Transparency | Typically a black box | None | Varies | Written rationale per answer; weighted, traceable responses |
| Best for | Fast directional reads | Brainstorming and ideation | Data augmentation and privacy | Quantitative surveys and qualitative research you can act on |
Questions to ask any synthetic panel vendor
Four questions, in order of importance. The first one settles most evaluations on its own.
- Where does the data come from? A credible vendor should answer in one sentence. If the answer is “a demographic prompt” or “the model,” you are getting a stereotype, not an audience. If the answer is a vague reference to proprietary AI, keep asking. Look for observed behavioral data at the individual level, and ask what specific signals are collected, how they are connected, and how often they are refreshed.
- How is accuracy validated, and is it published? Ask for benchmark studies against real-world survey results and a stated error rate. If they cannot produce one, assume the worst.
- Can it reach your actual audience? General population reads are easy. The test is whether the platform can model your specific buyers: a creator’s followers, a competitor’s customers, a hard-to-reach professional segment.
- Can you inspect the reasoning? A percentage with no explanation behind it gives you nothing to check.Look for platforms that show why an audience responds the way it does and how much of the market each response represents.
Digital Twins was built to pass all four, starting with the first.
Identify any audience you want to survey: your buyers, a competitor’s customers, or a segment no panel has ever been able to recruit. In a live demo, we will build that audience from the observed behavior of 150M+ U.S. adults, field your questions against it, and show you the reasoning behind every answer.
Book a Digital Twins demo
Frequently asked questions
What is a synthetic panel in market research?
A synthetic panel is a group of AI-simulated respondents used in place of human survey panelists. A model generates answers it predicts a described audience would give, delivering results in minutes instead of weeks. Quality depends entirely on whether the respondents are grounded in real behavioral data or only in demographic prompts.
Where does the data behind a synthetic panel come from?
For most synthetic panels, the data comes from the language model’s general training data, shaped by a demographic prompt. No proprietary information about your specific audience is involved. StatSocial Digital Twins takes a different approach: every respondent is modeled from observed cross-platform social, household, and offline data on real people, so answers reflect actual audience behavior rather than a general approximation.
Are synthetic panels accurate?
Accuracy depends on the data foundation. Research covered by Quirk’s found that LLM respondents built from demographics alone performed poorly at predicting real outcomes. By contrast, StatSocial Digital Twins, which model respondents from observed behavior across 150M+ U.S. adults, average 3.3 points MAE against real-world survey results across 40+ benchmark studies, better than the 5 to 6 points typical of opt-in online panels.
Is a synthetic panel the same as synthetic data?
No. Synthetic data is the broader category of any artificially generated data, used for augmentation, privacy protection, and modeling. A synthetic panel is one application of synthetic data: AI simulated respondents answering survey questions.
Do Digital Twins use real people's identities?
No. Digital Twins are anonymized models. They are built from real behavioral signals but no individual identity is exposed, which lets teams survey audiences without privacy risk.



