Back to Blog Digital Twins

Synthetic Audiences Are Fast. But Are They Actually Modeling People?

A version of this article by Michael Hussey, Founder and President of StatSocial, appeared originally on Advertising Week 360.

Synthetic audiences have become one of the hottest ideas in market research, and for understandable reasons. Traditional research can be expensive, slow, and dependent on panels that are not always as reliable as brands need them to be. So when AI-driven platforms promised to simulate audience reactions faster and at lower cost, the industry moved quickly.

The appeal is obvious. Why wait weeks for research when a model can generate answers in seconds?

But speed is not the same as understanding. Many synthetic audience systems are not actually modeling how real people behave. They are modeling how different groups get described online.

The internet is not the audience

Most synthetic systems are still inference engines. They are trained on massive amounts of publicly available internet text, then generate responses based on the language, assumptions, and patterns surrounding certain groups.

That can produce outputs that sound convincing. But sounding convincing is not the same as being grounded in real audience behavior.

For broad categories, the problem can be less obvious. There is a lot of public discourse around groups like dog owners, Gen Z consumers, or sports fans. A model can often approximate those audiences because there is so much language attached to them.

The weakness shows up when the audience gets more specific. An audience of oncologists, luxury travelers, niche investors, regional voters, or fans of a specific creator may not leave behind the same volume of public signal. When that happens, synthetic systems often fill in the gaps with inference. The output can still sound polished, but the representation underneath may be thin, stereotyped, or incomplete.

Some audiences are more visible than others

Online visibility is uneven. Some groups are talked about constantly. Others are not. That imbalance creates a serious problem for synthetic research. If a model is relying heavily on public language, it may overrepresent the groups that are most discussed online and flatten the groups that are less visible.

A 2026 paper from Google DeepMind found that when models are asked to generate diverse personas, the output can collapse around a narrow cluster of stereotypical responses.

That finding points to the central concern: The model may be generating its sense of an audience from how that audience is described, rather than from an independent understanding of who is in it.

Black boxes make the problem harder

Traditional research has its own flaws. Panels can be low quality. Respondents can be unreliable. Incentives can distort participation. But at least researchers can usually examine how a study was built. They can review how participants were selected, how the data was collected, and what factors may have shaped the results.

With many synthetic audience platforms, that visibility disappears. Researchers see the output, but they cannot always inspect how the audience was constructed or what assumptions shaped the response. That becomes a major issue when modeled outputs start guiding real business decisions.

Fluency is not fidelity

Synthetic audience tools can be useful for rapid testing and directional learning. They can help teams move faster, explore ideas, and pressure-test assumptions. The risk comes when companies start treating modeled behavior as interchangeable with direct audience understanding.

AI models are getting better quickly. But speed and fluency do not solve the core problem if the system is still learning from distorted or incomplete signals.

People are inconsistent. They are contradictory. They often make decisions for reasons that do not line up neatly with demographics or broad audience labels. Synthetic systems can smooth out that inconsistency and produce cleaner-looking profiles that feel more dependable than they really are. That is where the danger lives: mistaking coherence for understanding.

The question brands should ask

The useful test is simple: Can you inspect the behavioral inputs underneath the model? And do those inputs come from the audience you care about, or from public language about that audience? If the answer is unclear, the output may be closer to a coherent guess than a reliable research finding.

AI will absolutely improve parts of the research process. Faster iteration has real value. But research exists to help companies understand how real people behave, not just to produce confident-sounding answers.

Once outputs start sounding believable, people become less likely to ask whether the system truly understands the audience it is speaking for. That is the line marketers, researchers, and communicators need to watch closely.

See what real behavioral grounding looks like

The test in this article applies to any AI research tool, including ours: can you inspect the behavioral inputs underneath the model?

With Digital Twins, the answer is yes. Every Twin is grounded in observed behavior from 150M+ U.S. adults, what real people follow, watch, read, and buy, not in how audiences get described online. Because the inputs are inspectable, you can see exactly why an audience responded the way it did.

Request a demo and bring your hardest-to-reach audience. We’ll show you the behavioral data behind every answer.