Measure, generate or simulate?
Editor's note: Alex Mangoff is a senior account consultant at Burke with more than two decades of experience helping organizations navigate complex research and analytics challenges. He specializes in translating data into insights that inform smarter business decisions. As part of Burke's synthetic data team, he contributes to research exploring how generated data can be rigorously evaluated and validated for real-world applications. Find Mangoff on LinkedIn.
Synthetic data is not new. We’ve been doing Monte Carlo simulations since 1949 and predicting answers for incomplete data through conjoint for decades. But it's more ubiquitous now and the discourse around it is intensifying as large language models enter the mix of methods, and academics, vendors and researchers all weigh in. Business teams are inundated with promises of instant samples and lower costs. They wonder whether synthetic data – and AI, more broadly – is the solution that can help them test more ideas, reach hard-to-find audiences or simply move faster than the competition. This interest makes sense. Synthetic data offers a compelling answer for marketers and insights leaders facing tighter deadlines and shrinking budgets.
The promise is real. But so is the risk.
Synthetic output can sound human, look precise and still lead a team to a different conclusion than data from human respondents would. So, the debate shouldn't start with whether AI can "be human." It should start with a more practical question: can it support the decision we're about to make?
That question changes the conversation. Instead of treating synthetic data as a novelty or a golden ticket, we can treat it as part of a larger decision system; one made up not of a single tool but many.
Knowing your tools
Synthetic data is a well-known term, but different methods sit underneath it. Treating them as interchangeable creates confusion.
Broadly, there are two families: large language models (LLMs) and generative data models.
When most people think of synthetic data, they picture LLMs: systems that generate plausible, human-sounding responses without an actual person behind them. Synthetic personas, digital twins and synthetic panels all fall under this umbrella.
Generative data models work differently and are what we’ve seen leveraged in market research for decades. Instead of starting with language, they start with structured data already captured in a study or dataset and they estimate new responses based on what's already known. They're especially useful, then, for filling gaps and creating more stable reads from high-quality source data.
Language models are great for language. Generative data models are great for data.
That distinction matters because many business decisions need more than plausible responses; they need reliable evidence. And the riskier the decision, the more reliance we should place on the evidence. If a decision requires a measured read on what people think, do or intend, then the method behind it needs to be judged on one thing: whether it preserves the conclusion a trusted source of truth would support.
In practice, that judgment usually starts with one number: accuracy.
Accuracy is not decision reliability
A lot of the conversation around LLM-based synthetic methods centers on accuracy. Published research often cites performance around 80% compared to live data or human consistency.1 On the surface, it sounds encouraging. Eighty percent feels close and it's tempting to read that number as proof that synthetic methods are ready to replace or scale back parts of the traditional research process.
But accuracy alone does not answer business questions.
Data can look close overall and still change the recommendation. A synthetic panel might produce responses that feel coherent while still misstating the market underneath them. Even a model that holds up on broad patterns can flip the one number that drives a decision: the top-ranked concept, the majority view, the leading brand.
As an example, many researchers have experience working with typing tools for segmentation, where 80% or greater accuracy in predicting across segments is an acceptable baseline. But, hidden in that overall accuracy could be a single segment where accuracy is no better than the flip of a coin.
This reality has also been reflected in our own validation work. We found that LLM-based synthetic results landed in that same general range often cited by others in the marketplace, but at roughly 80% accuracy, the synthetic data gave us a different answer three out of five times.2 Here's what we can take from that: Eighty percent accuracy isn't meaningless, but it also isn't automatically decision-grade. It might be enough for exploration, early language development or first-pass prioritization. However, it’s a harder case for final validation or an investment-level call.
In other words, accuracy tells us whether the data looks close to the source of truth. Decision reliability tells us whether the recommendation holds. For business leaders, that's a critical distinction.
A decision standard
Synthetic data needs a decision standard because not every decision requires the same level of evidence.
A team narrowing 100 early ideas into a shortlist can tolerate more uncertainty than a team choosing which concept to launch. Pressure-testing objections to a marketing message is a different bar than deciding which campaign gets the budget. The gap shows up in analysis too: a broad hypothesis about which drivers might matter needs far less certainty than a final, ranked model used to guide spend.
The standard itself doesn't need to be complicated. Synthetic data is fit for a decision when its method, source data, validation and level of uncertainty are all appropriate for the business risk of being wrong.
To apply that standard, teams need to answer three questions:
- What decision are we trying to support? Are we exploring, prioritizing, validating or deciding?
- What evidence does that decision require? Does it need a current read from people or can it lean on data that's already been measured?
- What happens if the result is wrong? If a different source of truth changes the recommendation, that's a business risk, not a theoretical one.
Once those questions are answered, the evaluation becomes more practical. Instead of asking only whether the synthetic data looks accurate, we're asking what kind of risk remains and whether it's one we can live with. A broader quality framework is what turns that question into something a team can actually use.
We've already separated accuracy from decision reliability. What decision reliability actually demands is a sharper set of questions: Do the toplines hold? Do individual records look plausible? Do the relationships needed for deeper analysis survive? Accuracy alone can't answer any of those, which is exactly the gap the next framework needs to close.
Reframing the decision
Burke evaluates synthetic data using three dimensions: fidelity, authenticity and resolution. Together, they form the FAR Quality Score. The aim of FAR is to separate different kinds of quality that are not captured in a single accuracy claim.
Fidelity asks whether the topline picture holds. If a trusted dataset shows that 40% of respondents are top-box on purchase intent, does the synthetic data produce a similar read? Fidelity matters because many business decisions start with toplines: who leads, what changed and which option is strongest.
Authenticity asks whether the records look realistic. If you inspect individual synthetic rows, do they look like plausible respondents or do they show unrealistic patterns? Authenticity also helps guard against simply memorizing or copying real records. It gives us a reason to believe the generated data behaves like usable respondent-level data.
Resolution asks whether the important relationships in the data still hold. This is where many important research decisions live. Do we see the same differences across subgroups? Do relationships between attitudes and behaviors remain intact? Do the patterns needed for drivers, segmentation and crosstabs survive? A dataset can preserve the topline and still fail to preserve the relationships needed for deeper analysis.
FAR connects naturally to decision reliability as a result. Weak fidelity puts topline conclusions at risk. Weak resolution threatens the analysis underneath them. And if authenticity is weak, the data may not be plausible or responsible enough to use in the first place. Together, the three dimensions help move the conversation from "How accurate is it?" to "What kinds of decisions can it support?"
A high-level score is still not the whole answer. Some decisions depend mostly on topline stability. Others depend on relationships, rankings or subgroup patterns. But FAR is designed to give teams a more useful way to evaluate synthetic data than a single broad accuracy claim.
Why choosing the right method matters
The validation study referenced earlier not only informed the development of the FAR Quality Score but also provided clear examples of the gap between LLMs and generative data models when simulating quantitative data. In the study, we surveyed 5,000 people, generated 1,000 synthetic personas and built several test datasets to compare against a human baseline.
One comparison paired a fully human dataset against a blended one; 75% human responses extended with 25% generative estimates (Figure 1). The results showed the promise of generative methods at every level of the FAR framework, not just in aggregate. On fidelity, toplines on attitudinal questions like saving-money intent were nearly identical between the two datasets. On resolution, that consistency held at a more granular level too, with behavioral measures like weekly online shopping activity tracking closely, even when broken out by age group.2
That result comes with a catch. It still required collecting three-quarters of the data upfront. That's not a flaw in generative methods; it's the trade-off they're built around. They extend real data but not the need for it.
We pushed further in another comparison and asked what happens with zero measured data, relying entirely on a synthetic panel built with no pretrained knowledge of the study. Results here were markedly different. Strong saving-money sentiment came in well below the human figure and overall weekly shopping activity undershot by a wide margin, with the gap widening consistently for respondents over 35. This indicates a bias in LLM models and public data that skews our understanding of real human behavior (Figure 2).
This example is a strong argument for validating any fully synthetic panel against a trusted source. It also demonstrates why LLMs are better suited for early-stage decisions versus higher stakes and final assessments. 
Measure, generate or simulate?
These results reinforce the need to properly address the practical question of what role AI or synthetic data should play in business decisions.
The results also show that we don’t need to frame the discussion as real data versus synthetic data. Instead, it’s a more nuanced conversation about whether the data needed should be measured, generated or simulated.
When you need data, measure it
What is the temperature outside? This is a question that warrants measurement.
If the decision requires a current, direct read from people or from the world, synthetic data should not be in the consideration set. AI is not a sensor. It can reason over information but it doesn't automatically know what people think, feel or choose at a given moment. In these situations, there's no substitute for measuring directly, especially when the business question depends on measured reality: how consumers respond to a new offer, which message changes intent or how a specific audience experiences a different brand position. Done well, direct measurement remains essential.
Not every decision requires a large traditional study but we cannot treat synthetic output as measured evidence in absence of foundational data and understanding.
When you need better data, generate it
I have temperature data in many Texas cities. Can you estimate it for Fort Worth? This is a question generative data is built for.
Generative data models are ideal when trusted data exists but needs to be made more useful than its current form. They can fill gaps, boost small segments or repair trend breaks in data that's already been collected. In each case, the output is a more complete dataset extended from real respondents, not invented to replace them.
Generated data is still grounded in what was originally measured, which is an important difference. It can create more stability, more usable coverage and more room to examine smaller groups, but it shouldn't be interpreted the same way as additional, independently measured respondents. Generative models make high-quality data work harder; they do not turn weak data into strong evidence.
When you need reasoning, simulate it
Should I bring a rain jacket with me to Houston? This is a great question for a language model, because there's no single right answer, just better and worse ones. Language models are built for exactly that kind of reasoning – helping people think through a problem and how to approach it.
LLM-based methods are strongest when the goal is exploration, not measurement. They can help teams pressure-test early ideas, surface objections and reason across different perspectives before deciding where stronger evidence is needed.
The watchout is treating simulation as measurement. LLMs can help teams think better but they should not automatically be used to decide what a market believes, what share of consumers will behave in a certain way or which concept will win. Used intentionally, they improve the path to a decision. Used carelessly, they can make weak evidence feel more convincing than is deserved.
What this looks like in practice
The value of a decision standard becomes clearer when the method is matched to the job.
In early-stage innovation, LLM-based simulation can help teams narrow a wide field of ideas before investing in human testing, the same kind of decision we described earlier. In one project, synthetic personas weren't reliable enough to pick the winning idea but they were useful for narrowing a field of 92 candidates down to a workable shortlist. This is simulation doing exactly what it's suited for: not the final call but the work of deciding where stronger evidence is worth collecting.
Generative data models play a different role. In brand measurement, for example, they can help extend trusted respondent data when coverage is uneven across brands, segments or attributes. The goal is not to replace respondents or invent a market from scratch. It is to make already-measured data more complete and more useful for decision-making.
Both examples point to the same conclusion: Synthetic data works best when the method fits the decision. Use simulation to explore and prioritize. Use generative models to extend trusted data. Use direct measurement when the decision requires current evidence from people or the market.
What insights leaders should own
For insights leaders, the real opportunity is building better decision systems not just adopting synthetic data faster.
Practically, it comes down to creating a shared language for when to measure, when to generate and when to simulate, plus clear expectations for what each method is actually suited to support. Skip that discipline and teams end up using synthetic data because it's available rather than because it's appropriate. Build it in and synthetic data can expand the reach of research and make data strategy more responsive to the business.
The questions leaders should ask are straightforward. Two of them are the same ones from the decision standard: what decision are we trying to support and what evidence does it require? The rest speak to how the organization uses synthetic data at all, not just this one decision:
- Are we measuring, generating or simulating?
- What source of truth are we validating against?
- What decisions should this output not be used for?
These questions are not barriers to using synthetic data. They’re what make it usable. Synthetic data stops being a shortcut around evidence and becomes one input in an intentional decision system – one that builds confidence without overstating certainty.
Defensible confidence
Synthetic data doesn't replace human data. It stretches what we can do with it. The future of research isn't synthetic versus human; it's an integration of direct measurement, generative models, LLM-based simulation and expert judgment, each used where it fits.
Defensible confidence comes from using human data where current evidence matters, extending trusted data where the evidence holds up and simulating where it helps a team think through a problem. It is not artificial certainty or a shortcut around understanding people.
Synthetic data is not “the decision.” It's one input among several and it's fit for use only when the method, the evidence and uncertainty match what is at risk if the answer is wrong.
References
1 See Joon Sung Park et al., “LLM agents grounded in self-reports enable general-purpose simulation of individuals,” arXiv (2024, revised 2026), arXiv:2411.10109; Tianyi Peng et al., “A mega-study of digital twins reveals strengths, weaknesses and opportunities for further improvement,” arXiv (2025), arXiv:2509.19088; and Sarah Schröder et al., “Large language models do not simulate human psychology,” arXiv (2025), arXiv:2508.06950.
2 Burke, Inc. (March–June 2026). Synthetic Data and Decision Reliability Study.