Brand Analytics · Synthetic Research

Challenges With Synthetic Personas and Their Best Use Cases

A synthetic estimate for a whole market generally combines how each group answers with how large each group is. A grounded language model can supply the answers. However, the size of each group has to come from measurement.

Anton Dudarenko · 12 min read · 19 September 2026
TL;DR A grounded language model can supply the answers in a market estimate. However, the size of each group has to be measured.
  • A synthetic result for a whole population generally combines how each group answers with how large each group is. The mix of people inside a language model does not match the population of interest, so the group sizes have to come from real data.
  • Grounding improves the answers: interview-grounded agents reached 83% of participants' own two-week consistency, against 74% for demographics-only agents.
  • In one study synthetic responses alone were 24 to 86% biased, and under 5% once corrected with real responses.
  • For consumer markets Lift-Off proposes grouping people by psychological profile and by occasion, each with a known share.
  • Synthetic personas fit stated-attitude questions, concept screening and questions that arrive after the survey has closed. Price tests and behaviour forecasts need more care, and the result depends on how the study is set up.

Synthetic personas and how they are built

This article argues that a synthetic persona is reliable to the extent that the simulation is tied to real data. A synthetic persona is a language model instructed to answer as a described person. A synthetic respondent is that persona answering a research question. Many of them together form a synthetic sample.

Li, Chen, Namkoong and Peng write that persona-based simulations hold promise for changing disciplines that rely on feedback from whole populations. These include social science, economic analysis, marketing research and business operations. Maier and colleagues write that consumer research costs companies billions annually. They also write that it suffers from panel biases and limited scale.

A persona can be built from a prompt, from a real person's record, or from data about the population. The simplest way places a description of a subpopulation in the prompt, and the model is steered by that description. A more demanding way builds the agent from a real person's own record, such as a two-hour, semi-structured interview, structured survey answers, or the interview and the survey together. Park and colleagues built their agents from records of this kind. This article calls a persona built from a real person's record grounded.

The persona's attributes can also be drawn from real population distributions. Sun and colleagues test this with demographic distributions, and we are increasingly using this approach.

Li and colleagues add a warning. Today's methods for generating personas are improvised and rest on rules of thumb. They do not guarantee sound method or precise simulation, and the result is systematic bias in the work that depends on them.

Documented challenges with prompted personas

Published research documents these challenges when a model is given a prompted persona.

  1. The model invents context, filling any gap the prompt leaves with its own assumptions (Gui and Toubia).
  2. The model distorts the spread of answers when it is asked directly for a numerical rating (Maier and colleagues).
  3. The model reacts to the order of the answer options, and its answers shift when the order changes (Dominguez-Olmedo and colleagues).
  4. The model gives less accurate results when it writes the persona's details itself (Li and colleagues).

Terms used in this article and its scope

The sections below use a few more terms. Psychological profile, occasion, demand space and weight belong to the approach Lift-Off proposes later in this article. Test-retest consistency comes from the published studies.

A psychological profile describes a person's values and attitudes to life, their patterns of behaviour and the way they make decisions, and it stays the same from one occasion to the next. An occasion is the situation in which a product is chosen or used: the time, the place, the company present and the need.

A demand space is one kind of consumer in one occasion, and it is the unit the market is divided into here. A weight is the share of all occasions that one occasion holds, taken from measurement. Test-retest consistency is how closely the same people repeat their own answers when asked again.

The sections below cover, in order, group answers and group sizes in a population estimate, and the sources of error in each. They then cover Lift-Off's proposal to group consumers by psychological profile and occasion, a worked example of weighting occasions, and the published accuracy results to date. Finally they cover practical remedies for the documented challenges, the best use cases and the cases that need more care, and a test of a synthetic result against an existing study.

Group answers and group sizes in a population estimate

A synthetic result reported for a whole population, a market, an electorate or a customer base generally combines how each group answers with how large each group is. Argyle and colleagues, in the paper that introduced silicon sampling, write a population estimate in exactly that form. Silicon sampling is their name for simulating a survey sample with a language model. In their terms, these are the answer given a person's background, and the distribution of backgrounds in the population.

A language model can supply the group answers well when it is grounded in real material. Park and colleagues built agents from two-hour interviews with 1,052 participants. On survey items that were not used to build the agents, those agents reached 83% of the participants' own two-week consistency, against 74% for agents built from demographics alone.

However, the model is unreliable on the size of each group. Argyle and colleagues point out that the mix of people a language model has learned from is different from the mix of people in the population being studied. If nothing corrects for that, conclusions about overall patterns in the population, which they call marginal patterns, come out skewed.

Their fix is to take the mix of people from a known, nationally representative sample, and then let the model supply the answers inside each group. They state that this works for any chosen population, as long as the model captures the answers inside a group well.

Further results support the same principle. Sun and colleagues found that group-level demographic distributions alone were enough for language models to generate response distributions remarkably similar to actual U.S. public opinion polls, and that the match varied by demographic group and by topic. Krsteski and colleagues found that synthetic responses alone introduced substantial bias, 24 to 86%, and that combining them with a correction based on a set of real human responses reduced bias below 5%.

In each of these papers the structure of the population comes from real data, and the model supplies the answers inside it.

Sources of error in the answers and the group sizes

The approach has been tested against 57 product surveys conducted by a leading personal care corporation. The decision this quarter is whether a synthetic read enters a plan, and the entry test has to cover the answers and the group sizes.

The answers can be confounded. In Gui and Toubia's demand experiment, the curves from human participants followed the classic downward slope, while the GPT-simulated curves followed an inverted-U shape. The simulated world changed with the price. The model filled in details the prompt never specified, and those details rose with the price of the product: the last price paid, the competitor's price and the days to expiry.

Gui and Toubia explain why this happens. In an experiment with real people, only the price changes and everything else about the product stays the same. A simulated person is given only the prompt. When the price in the prompt changes, the model also changes its picture of the details the prompt left out, and those details should have stayed fixed. They tested this against a real experiment with 40 products. The simulated person was answering a different question from the one the humans answered.

Human demand curve against the curve produced by LLM personas Schematic. Human purchase probability falls steadily as price rises. The LLM personas produce an inverted U, with a region at lower prices where purchase probability rises with price. Region where the simulated purchase probability rises as price rises Price Purchase probability Human respondents: purchase probability falls as price rises LLM personas: an inverted U Human respondents LLM personas simulated probability rises with price
Schematic of the pattern Gui and Toubia report, drawn to show the two forms and plotted from no dataset. The solid line is the human panel and the dashed line is the simulated one.

The answers can also lose their spread. Bisbee and colleagues prompted a language model with personas and found that its average scores matched a national survey closely, while the responses varied less than the real ones. Asked directly for numerical ratings, language models produce unrealistic response distributions. A concept screen that reports the right mean and the wrong spread passes any check made on the mean, and the decision risk is in the spread.

The group sizes can be wrong as well. Treating every group as equal in size can make the population-level number wrong, however good the answers inside each group are. No volume of synthetic respondents repairs wrong answers or wrong group sizes. Repeating a simulation with unchanged input gives a more precise estimate of what the model says, and it adds no information about the population.

The size of each group in a real market is a measurement, and a language model has no access to it.

Lift-Off's proposal to group consumers by psychological profile and occasion

Published methods for building simulated groups define those groups by demographics. In Argyle's method the groups are defined by people's backgrounds and taken from a national survey. In the method of Sun and colleagues the groups are defined by demographic distributions.

Published studies show the weakness of grouping people by demographics. Wang and colleagues compared real survey respondents from different socioeconomic groups with Llama-8B agents built from demographic profiles. They used a silhouette score, which shows how sharply groups separate. A score near 0 means the groups overlap. The real people scored -0.02: they varied widely inside each group and overlapped with the other groups. However, the agents scored 0.19: they were more alike inside each group, and further apart from the other groups, than real people are.

In the authors' words, a demographic profile is "only a sparse and static account of an individual".

Liu, Diab and Fried found that language models are 9.7% less steerable towards personas whose views are unusual for their demographic. Less steerable means harder to keep in character. The models sometimes gave the view that is typical of the demographic in place of the view they were asked to hold.

For consumer categories we propose grouping people in two ways at once: by psychological profile and by occasion. The size of each profile starts from real data about the country's population. Our demand-space work states: "The same person is a different buyer in different contexts."

A demand space is one psychological profile crossed with one consumption occasion, and each is sized by weekly occasions, by category volume and by annual value in pounds. Each demand space has a measured share of all occasions in the category. A typical map has 8-12 demand spaces built from 25-40 micro-spaces, which are smaller clusters of occasions.

Our approach works top-down and bottom-up. Top-down, the measured map supplies the groups and their sizes before any simulation begins, so the structure of the sample comes from market data. Bottom-up, a persona built from real population data answers inside one specified occasion. We also write the occasion into the prompt, so that the model has less to invent. Gui and Toubia explain that a simulated subject exists only as the prompt describes it, and they found that telling the simulated subject how the study was designed improved every model they tested.

Grouping consumers by psychological profile and by occasion A grid. The rows are psychological profiles, each with a known share of the population. The columns are occasions, each with a measured share of the market. Inside each cell a persona with that profile answers in that occasion. Each cell's answer times that cell's known share, summed, gives one market answer. TOP-DOWN two sets of known shares, fixed before any simulation OCCASIONS measured share of the market Occasion 1 Occasion 2 Occasion 3 Occ. 4 PSYCHOLOGICAL PROFILES known share of the population Profile A Profile B Profile C One cell: a persona with Profile A answers inside Occasion 2 A Profile A persona answers inside Occasion 2 BOTTOM-UP inside each cell, a persona with that profile answers in that occasion Market answer each cell's answer x that cell's known share, summed
The approach Lift-Off proposes. People are grouped by psychological profile and by occasion, and both sets of shares are known before any simulation. Row heights and column widths are illustrative.

We build the personas from real data about the population. Demographic data for the country and the mix of the 16 personality types in that country set the starting point for who the personas are and in what proportions. Each psychological profile is built from those personality types.

The mix of personas therefore starts from the country's real proportions. We then adjust it for how much each kind of person buys in the category, before any answer is simulated.

Our method is to start top-down from the measured market, divide it by psychological profile and by occasion, simulate each profile inside each occasion, and weight the simulated answers by the share of the market that each combination is known to hold.

Weighting a synthetic sample to known group sizes is established in the published work above. However, using psychological profiles and occasions as the groups is our proposal.

We group by psychological profile and by occasion because each narrows what the model has to guess. The profile fixes how a person tends to decide, and the occasion fixes the situation that person is in.

Published work shows that knowing more about a person than their demographics improves prediction, as the interview study described above found. That study grounded its agents in interviews, and we build our profiles differently, so it supports our reasoning and does not test our method.

We expect profile and occasion together to predict better than profile alone or occasion alone.

Invented persona detail and real persona detail have opposite effects. Invented detail made predictions worse. Li and colleagues compared four persona types in a US presidential election case study, and as the LLM-generated persona attributes increased, the simulated results deviated further from the real outcome, ending with every state voting for the Democratic candidate.

However, detail taken from real interviews raised accuracy from 74% to 83% of the human benchmark. RAIN Group's test for an insight that matters includes that it is contrary to conventional wisdom, and the conventional wisdom here is that a better persona is a more detailed persona. In the election study and the interview study, accuracy depended on the source of the detail.

A worked example of weighting occasions

A market-level answer is each occasion's answer multiplied by that occasion's share, then summed.

An illustration with invented figures makes the cost of skipping the weights concrete. Take a category with two occasions. A weekday breakfast occasion holds 70% of occasions with 20% purchase intent. A weekend treat occasion holds 30% of occasions with 60% purchase intent.

The weighted market-level intent is 0.70 x 20% + 0.30 x 60% = 32%. A simple average of the two occasions gives 40%, eight points too high.

Illustration of a weighted market answer against a simple average Invented figures. A weekday breakfast occasion holds 70 per cent of occasions with 20 per cent purchase intent. A weekend treat occasion holds 30 per cent with 60 per cent intent. Weighted by share the market intent is 32 per cent. A simple average gives 40 per cent. INSIDE EACH OCCASION Weekday breakfast 70% of occasions Weekday breakfast: 20% purchase intent 20% intent Weekend treat 30% of occasions Weekend treat: 60% purchase intent 60% intent MARKET LEVEL Weighted by occasion share 0.70 x 20% + 0.30 x 60% Weighted by occasion share: 32% 32% Simple average of occasions (20% + 60%) / 2 Simple average of the two occasions: 40% 40% eight points too high
An illustration with invented figures. Both occasion-level answers are identical in the two market rows. Only the weights differ, and the market answer differs by eight points.

Every answer inside each occasion can be right, and the number for the whole market can still be wrong. In this example the mistake makes the idea look more popular than it is: 40% against a true 32%. If the shares were the other way round, with the weekend treat as the larger occasion, the same mistake would make the idea look less popular than it is. Without the weights there is no way to know which way the error goes.

Argyle and colleagues compare this to Simpson's paradox. A pattern seen in a whole group can differ from the pattern inside each of the smaller groups that make it up.

This suggests one reason, offered as reasoning, why simulated answers vary less than real ones. A persona that is asked a question with no occasion attached has nothing to tell it which situation matters more. The model gives one averaged answer that covers many situations at once.

Published accuracy results to date

Inside a defined group, the spread of answers can be recovered. Maier and colleagues asked the model to answer in words. They then matched those words to a rating scale, such as 1 to 5, by comparing them with reference statements. Across 57 personal care product surveys and 9,300 human responses, this reached 90% of human test-retest reliability.

The spread of the simulated answers also stayed close to the real spread, with a KS similarity above 0.85. KS similarity measures how closely the shapes of two sets of answers match, and 1 means identical. The way the question was asked preserved the shape of the distribution.

Training the model on real answers closes more of the gap. Describing a group in the prompt has struggled to predict how that group's answers are spread. SubPOP is a dataset of 3,362 survey questions with 70K records of how different groups answered them.

Suh and colleagues trained the model further on those real answers, a step called fine-tuning. This cut the gap between the model's answers and people's answers by up to 46% compared with earlier methods, and the improvement held on surveys and groups the model had not seen.

Those results improve the answers inside a group that someone has already defined. However, the size of each group comes from elsewhere. In Argyle's method it comes from a nationally representative sample. In the method of Sun and colleagues it comes from a demographic distribution. Argyle's method and Sun's method offer no way to group people in a consumer market. Our proposal addresses that gap.

Argyle and colleagues add a caution that applies to any frame. Being able to sample a group's answers does not guarantee that the answers are faithful to that group, so each domain and each group needs its own check.

Accuracy also has a human ceiling. Park's figures are shares of the participants' own two-week test-retest consistency, so 83% means 83% of how consistently the same people answer the same questions twice. The benchmark is human repeatability.

Agent accuracy as a share of human test-retest consistency, by grounding source Park and colleagues, held-out General Social Survey items. Demographics-only agents 74 per cent. Survey-only 82 per cent. Interview-only 83 per cent. Interviews and surveys combined 86 per cent. 100 per cent is the participants' own two-week consistency. 100 = how consistently the same people answer twice Demographics only Demographics-only agents: 74% of human test-retest consistency 74% Survey answers Survey-grounded agents: 82% 82% Two-hour interview Interview-grounded agents: 83% 83% Interview and survey Agents grounded in both sources: 86% 86% Agents grounded in each person's own record (teal) against agents built from demographics alone (grey).
Source: Park et al., arXiv 2411.10109, held-out General Social Survey items, 1,052 participants. Accuracy is a share of the participants' own two-week test-retest consistency.

Weighting by occasion does only part of the job. It deals with context invented by the model, with group sizes the model cannot know, and with persona detail invented by the model. It does nothing, however, for order bias, for statistical significance, or for the gap between what people say and what they do. Each of those has its own remedy in the next section.

To measure the accuracy of the weighting step, a weighted result would be compared with real survey answers on how closely the averages match and how closely the spreads match. The result would then be read against how consistent the real answers are with themselves.

Practical remedies for the documented challenges

A synthetic result should be validated against human data before it is used. Park and colleagues express every accuracy figure as a share of the participants' own two-week consistency, and any synthetic result can be held to the same standard.

Order bias is large. Dominguez-Olmedo and colleagues tested 43 language models on the American Community Survey and found the responses governed by ordering and labelling biases, for example towards answers labelled with the letter "A". Once the answer order was randomised to adjust for those biases, the models trended towards uniformly random responses, irrespective of model size. The remedy is durability testing: randomise the order, reword the question, and keep only the findings that survive.

Attitudes and behaviour diverge. The strongest published accuracy evidence is for survey experiments. Across 70 pre-registered survey experiments with 476 measured effects, predictions from simulated responses correlated with the real effects at r = 0.85, where 1 would be a perfect match. A language model is trained on text, so evidence about stated responses does not extend to behaviour in the field without a separate check. The remedy is to ask attitude questions that are known proxies for the behaviour, and to validate the proxy against real data.

Persona construction needs validation. More LLM-generated detail pushed one set of election simulations further from the real result, so a new construction should be validated against known human ground truth before use.

A correlation metric shows whether the averages are right, and a distribution-similarity metric shows whether the spread is right. Maier and colleagues report 90% of human test-retest reliability, and a KS similarity above 0.85. Where the same people cannot be re-tested, split the human data in two at random, correlate the halves, repeat the split many times and average the result. That average is the accuracy ceiling for the data.

Sample size has no remedy. More synthetic respondents add precision about the model and no statistical significance about the population.

The standard a synthetic answer has to beat depends on the alternative available at that moment. When the question arrives after the survey has closed, the alternative may be no research, or one person's opinion. A synthetic answer has to beat that alternative, and it has to be reported with its ceiling.

In a vendor conversation, this reduces to asking which measured distribution the output was weighted to, and how closely the spread of the simulated answers matched the real spread.

Challenge Remedy
Invented context and confounded answers Specify the situation and reveal the study design in the prompt
Group sizes the model cannot know Take them from real data: a representative sample, a demographic distribution, or known profile and occasion shares
Invented persona detail Ground the persona in real interviews or survey records, then validate it against human data
Order bias Durability tests: randomise the order, reword the question, keep what survives
Stated attitudes diverging from behaviour Ask attitude questions that are validated proxies for the behaviour
An average that hides a wrong spread Report a correlation metric and a distribution-similarity metric together
Unknown accuracy ceiling Split-half consistency of the human data
More synthetic respondents No remedy: they add no statistical significance

Best use cases for synthetic personas and cases that need more care

The strongest use case is a question about stated attitudes and reactions to something shown to the respondent. Across 70 pre-registered survey experiments, researchers compared the change each experiment caused in real people with the change predicted from simulated responses. The real changes and the predicted changes correlated at r = 0.85. For unpublished studies that could not have been in the training data, the correlation was r = 0.90.

Concept and message screening is another good fit. The elicitation method matters here. Purchase intent elicited as text and mapped to the Likert scale by semantic similarity reached 90% of human test-retest reliability across 57 personal care product surveys.

A question that arrives after the survey has closed is also a good fit, when the alternative may be no research, or one person's opinion. A synthetic answer is useful there on one condition, that it is reported with its accuracy ceiling.

A market-level read is a further fit, in a category where each occasion's share of total occasions is already measured. The weights exist before any simulation starts, so each occasion's simulated answers can be weighted by the share that occasion is known to hold. Its first use should include the test of averages and spread described above.

The weakest fit is a thin subgroup, read for direction only. Simulation can produce more answers for an underrepresented group, and those answers add no statistical significance.

Price tests and behaviour forecasts need more care, and the result depends on how the study is set up. A price test fails when the simulated subject is blind to the study design, because the simulated world then changes with the price. Revealing the design to the simulated subject improved every model Gui and Toubia tested.

A forecast of behaviour can work. However, the accuracy evidence above is for survey experiments, so a behaviour forecast should be checked against real behaviour, such as past sales, before it is relied on.

A larger synthetic sample adds no statistical significance, because repeating a simulation with unchanged input adds no information about the population.

Testing a synthetic result against an existing study

Take the most recent segmentation or usage-and-attitude study. Write down each occasion and its share of total occasions. Then ask of any synthetic result already circulating which of those shares it was weighted to. If the answer is none, the market-level number is an unweighted average, and the illustration above shows what that costs. Next, score one past synthetic result against its human equivalent with a correlation metric and a distribution-similarity metric. These steps use existing material and can be done before any purchase decision.

NavigatorLab, Lift-Off's demand-space mapping tool, shows each occasion with its share of the market, so the weights exist before any simulation starts.

This is the kind of question we work through with insight teams. If you hold a measured occasion map and a synthetic result you have to defend, get in touch.

Sources and further reading