The stereotype ceiling holding back synthetic research
Editor’s note: Derrick McLean is product scientist, Edge Center of Excellence, at Qualtrics, Chicago, with 13 years of experience in experimental psychology, data science and survey methodology. He holds a B.S. from Seattle University, a master’s degree and Ph.D. in psychology from Claremont Graduate University, as well as an AI competency credential from Mindstone. Find McLean on LinkedIn. Jordan Harper is principal AI thought leader, Edge Center of Excellence, at Qualtrics, Chicago, with over 20 years of experience in software engineering, digital transformation and customer experience strategy. He holds a Master of Physics from the University of Leicester. Find Harper on LinkedIn.
Synthetic research – using AI models to simulate consumer responses and behavior – is fundamentally changing how organizations understand their customers. Faster iteration, lower cost, broader reach. The promise is real.
But the first generation of tools has a design problem that many early adopters aren't asking about. The models can't reproduce the messiness and disagreement that make research valuable in the first place, and that flaw is shaping every output they produce.
To understand why, it helps to think about how we handle uncertainty in other fields.
Every morning, meteorologists run dozens of weather forecasts simultaneously. Each model starts with slightly different assumptions. The forecasters don't average those models into a single confident answer. They study the spread. When the models converge, confidence is high. When they diverge, that divergence is itself the forecast – the atmosphere is in a genuinely uncertain state. The variance isn't noise; it's the critical information that helps shape an accurate forecast. The same principle underlies what makes a research audience worth having. A well-constructed panel draws its value from the independence and diversity of what individuals contribute.
Real customers hold contradictory opinions about the same brand. They make trade-off decisions based on mood, recent experience and factors they can't articulate. They are sometimes wrong about what they want until they encounter it. These aren't flaws in the data. The behavioral variance, outlier sentiment and internal inconsistency are precisely what market research exists to surface.
A model trained to eliminate those qualities simulates a rationalized ideal of a person, which isn’t very useful for consumer research.
But that's a design challenge for synthetic research, not a fatal flaw. The messiness is the signal, and the models that earn trust will be the ones built with the rigor and the data to preserve it.
Synthetic research tools: Why breadth isn’t depth
Most first-generation synthetic research tools or digital twins are built on general-purpose AI models like GPT-4 or Gemini. These foundational models are trained on the open internet and possess remarkable breadth, capable of writing poetry, coding in Python, summarizing legal documents and analyzing data.
But in market research, breadth is not depth.
When a general-purpose model is asked to simulate a consumer weighing price against convenience, it doesn't draw on the full range of how real people resolve that tension. It draws on a training set built from Reddit threads, Wikipedia articles and public forums – the version of that consumer that the internet chose to represent.
Human panel research consistently shows that price sensitivity isn't a single trait. It shifts depending on how a question is framed, what trade-offs are on the table and what emotional context surrounds the decision.
Some travelers accept an inconvenient layover without hesitation. Others will pay significantly more to avoid it, driven by anxiety they may not fully articulate. Others report satisfaction with the cheaper choice in the moment and dissatisfaction weeks later.
These models smooth out that variance by design, preventing researchers from seeing a segment they didn't know existed: the price sensitivity that only surfaces under certain framing, or the emotional driver no one on the team anticipated.
The illusion of personalization
The industry's answer to this limitation has been more data. Enrich the model with purchase histories, behavioral profiles, CRM records, past survey responses. The more you give it, the more accurately it will simulate your customer.
A 2026 study put this directly to the test, and found the answer wasn't so simple.
Researchers in the study, titled “Digital Twins as Funhouse Mirrors,” provided a model with over 500 individual data points on one person and measured whether that enrichment improved the model's ability to predict that person's actual survey responses. Compared to a version using only basic demographic information, the improvement was statistically insignificant (p=0.37). Five hundred data points, but no meaningful gain.
The answers didn't change because the architecture didn't change. The researchers called it an "illusion of personalization."
The LLM architecture problem and the stereotype ceiling
The real explanation of the problem is architectural. As noted, general-purpose large language models are trained on the breadth of public text and are, by design, generalists. Their intelligence is built on aggregation. The models learn what the average person across the internet tends to think, feel and say.
When you feed in specific customer data, it stretches its average persona to accommodate the new information. No matter how much context you layer onto a general model, you cannot override the fundamental orientation of its training. It will give you the internet's best guess at what your customer thinks, not a simulation of what your customer actually does.
Researchers call this the stereotype ceiling: The point at which more personalization data stops making a difference.
What the next generation of synthetic models looks like
The stereotype ceiling is not a verdict on synthetic research. It's a maturation challenge, and one the field is beginning to address.
Rather than building a consumer simulation tool on top of a general-purpose model, the alternative is to train a foundational model not on public text, but on structured survey response data: millions of real respondents, real questions, real behavioral patterns across demographics and product categories.
The goal is to teach a model how different humans behave when they encounter questions for the first time – the uncertainty, the contradictions, the rationalization that happens in real time as people form opinions.
Early work in this space, including research conducted at Qualtrics, suggests that when training data crosses a critical mass of real survey responses, something shifts.
In our study, the model was reproducing the full distribution, including the variance, the outliers and the disagreements that make research data worth having. Compared against general-purpose alternatives, the fine-tuned model showed accuracy improvements that, in internal benchmarking, were 10 to 12 times more accurate at representing real human responses.
This distinction matters for anyone evaluating synthetic tools. What a model was trained on determines what it can and cannot surface. No amount of clever prompting or data enrichment will override that constraint.
Where synthetic adds value, and where it doesn’t
Even the most carefully trained synthetic model is not a replacement for human research. These models perform well when you're asking people what they think, what they prefer and how they feel: opinion questions, trade-offs between product features, concept testing, emotional reactions. These are the kinds of questions real survey data trains a model to understand deeply.
They are significantly weaker when answers depend on personal experience. A model cannot know how your restaurant visit went last Thursday. For these cases, human respondents remain irreplaceable, and vendors who claim otherwise are not running the tests that would surface their own limitations.
The right frame for synthetic research is not to replace the panel. It is to run better research. Simulation built on real human data enables research teams to explore a problem space, develop sharper hypotheses, test survey instruments and identify key segments before investing in full-scale human data collection. It compresses research cycles without sacrificing integrity, but only when it's built on the right foundation.
Questions to ask before trusting a synthetic promise
Synthetic research represents a genuine advance in how organizations can understand their customers … when the underlying models are built for the task.
The question for leaders is whether the tools they're evaluating were designed to preserve the signal that makes research worth commissioning. When a vendor claims its model is 90% accurate, don't take it at face value. Ask for the underlying research:
- What was the study design?
- What was the methodology?
- Is the data replicable?
- What has changed in the model since that validation was conducted?
Then ask the harder question: If this output had come from a human panel, would you find it convincing? If the synthetic model is built on top of an LLM, be wary of the stereotype ceiling and the illusion of personalization.
The organizations that capture the most value from synthetic research will be the ones that ask those questions now, before the stereotype ceiling becomes invisible, baked into insights that look right, feel right and quietly erase the variance where the real answers live.