Synthetic audiences: where they help, where they stop, and how to run one

© Illustration by Aurélie Garnier

Khalil A. Cassimally is an audience consultant and coach who helps organisations align teams around change, user needs and AI.

Synthetic audiences promise fast, cheap reader feedback, but the claims made for them are easy to misread. Drawing on recent research and his own practice, he shares a practical four-step way to use synthetic audiences to:

  • Recognise where synthetic feedback stops being trustworthy
  • Run a panel grounded in material your newsroom already has
  • Test language, narrow options and surface objections before committing
Key takeaways to copy & share
> Audience consultant Khalil Cassimally's rule is to use synthetic audiences to narrow, never to decide. They rank options and surface objections, but cannot tell you magnitudes like willingness to pay

> Start with the decision you need to make, not the question you want to ask. Most synthetic research fails by producing interesting material nobody was ever going to act on

> Ground the panel in words real readers actually wrote. Old interview transcripts, cancellation reasons, unread survey comments. The test is whether you can name the humans inside the system

> Most teams only ever count responses. The better question is where your draft conflicts with what you have documented about a reader, which is analysis rather than impersonation

> Tell it to answer as a reader rather than an analyst, and that indifference is valid. Then validate whatever survives with real people, which is where the saving comes from

Audience research in most newsrooms comes in two states. Either we talk to real people or we don’t do it at all. In my experience, the second is far more common than the first.

And where research does happen, rather than being systematic, it tends to happen once in a while. Worse still, its insights usually go into a report or a slide deck, which just sits in a shared drive.

What we lose by not doing research

Organisations that research their audiences well outperform the ones that don’t. The gap is not small. A study compared companies with mature research practices against those with weak ones across a dozen business outcomes. On revenue, 47% of high-maturity organisations reported a positive impact against 11% of low-maturity ones. On customer retention, 58% against 24%. Averaged across every outcome measured, the research-mature organisations came out 2.3 times ahead.

audience research is important

None of that is unique to the media but the mechanism is familiar to anyone who has been a part of a redesign that’s gone badly. Knowing what readers actually think changes what gets built. Not knowing means guessing, and guesses compound.

The questions most research doesn’t tackle

Few newsrooms are commissioning representative surveys. Meanwhile interviews cost real money and real weeks. Plenty of newsrooms also treat editorial instinct as the senior evidence, with audience feedback consulted after instinct has already decided. The effect is the same either way. When there is appetite for research, it goes to the bigger decisions while everything else gets settled without evidence.

That leaves a whole category of question permanently unanswered. The recurring ones, where we need to know:

  • How a piece of content will serve the users we are trying to reach
  • Which of several options is weakest, before we invest in any of them
  • What objections a change is likely to run into
  • How different groups of users will react to the same thing
When to consider synthetic research

To take one example: whether a new onboarding flow feels like a data grab. Too small to commission a study for, and it creates uncertainty that festers within the onboarding team. So somebody decides on instinct, usually whoever is most senior, and the instinct becomes the decision. It might well be right. Nobody will ever know though because nothing was tested.

This is the gap synthetic audiences fill. They don’t have to replace the research we already do well. What they change is the dozens of small calls a month that currently run on somebody’s hunch. And for the many newsrooms doing no audience research at all, the comparison isn’t between synthetic feedback and properly-designed research. It’s between synthetic feedback and nothing, and synthetic feedback done well beats nothing.

The value is that a whole class of decisions moves from opinion to evidence, at a cost and speed that makes asking routine.

What a synthetic audience is

“Synthetic audience” is an umbrella term, and the field hasn’t settled on what sits under it. You’ll see people use synthetic personas, synthetic users, digital twins, synthetic panels and silicon samples loosely and sometimes interchangeably. I’m using three here because the differences between them change what we can trust.

  • A synthetic user is an LLM given a description of a type of person and asked to respond as they would. We write the description – the segment’s age range and media habits, whatever we know – put it in the prompt, and ask the question. Nothing sits underneath it. What comes back is the model’s averaged sense of what somebody matching that description would say, assembled from everything in its training data about people like that.
  • A digital twin is built from one real person’s own material: an interview transcript, their survey answers, their support emails, their behaviour on our site. That material goes into the prompt and the model answers new questions as that person. It’s also possible to fine-tune a commercially available LLM on the material instead of putting it in the prompt. Academic studies from Stanford and Columbia tested that route, and neither found it beat simply placing the person’s words in the prompt. Which is good news for small teams: building a twin should not require any substantial engineering lift.
  • A synthetic panel is either of the above, multiplied. We can build twenty or fifty across our segments, put the same question to all of them, and read the pattern rather than any single answer.

What synthetic audiences are good for

A number of use cases have emerged. These four have the clearest evidence behind them:

Testing language before it goes out. Paywall messages, membership appeals, onboarding copy, campaign framing. We put the draft in front of a synthetic panel and ask what they find unclear and what would make them stop reading. This is the highest-value use, partly because language is what these models handle best, and partly because it’s a decision we make most often.

Narrowing a field of options. Six newsletter concepts, or four pricing structures. The panel won’t tell us which one to launch. But it will reliably tell us which ones are weakest. So, the real research budget goes to the survivors. Spindrift Beverage Co., a US drinks brand, used a synthetic panel to choose between new product concepts. It reached the same answer in a week that a 500-person consumer study had taken two months to produce.

Surfacing objections we haven’t thought of. Ask the panel what would stop them subscribing, or what they’d assume about us from our about page, and we get a list of concerns to check. Where synthetic panels have been compared against real research, they have identified the same themes and the same objections. The magnitude and the proportions differ, the list doesn’t.

Comparing groups against the same thing. Put one paywall message to a synthetic lapsed-subscriber panel and a synthetic loyal-reader panel, then read the difference. Relative comparison is where these systems are most useful.

You may have spotted that all of the four use cases narrow or pressure-test rather than decide.

How accurate are synthetic audiences

A figure of around 85% circulates in coverage of this field, usually without much explanation. Read quickly, it suggests a synthetic version of a person gets 85 out of every 100 answers right. That isn’t what was measured.

The figure comes from the Stanford study. Researchers interviewed just over a thousand people at length. They built a digital twin of each person from their interview transcript. Then they tested whether the twin could predict that person’s answers to a survey the person had already completed.

But they didn’t score the twin’s answers against the persons’ answers and leave it at that. Because if we ask a person to fill the same survey again two weeks later, they match their own earlier answers only about eight times in ten. (Memory shifts!) So, the fair question is not whether a twin matches perfectly. It’s how a twin does against that eight-in-ten human benchmark.

Researchers found that the synthetic twins reached roughly 85% of that human benchmark – so, not 85% of the answers.

how synthetic responses match to human responses

The figure tells us that synthetic twins built on rich individual material carry genuine signals about the people they came from.

The evidence in this field is still early, and most of it comes from outside media. For me, that’s another reason to use synthetic audiences for narrowing rather than deciding, which is what the rest of this piece argues for.

Where synthetic audiences stop

They can’t tell us how much. Synthetic responses cluster. Real people spread out more across opinions, while synthetic ones converge towards each other. A panel can tell us which option comes out ahead. It can’t tell us by how much, which rules out adoption rates, willingness to pay, satisfaction scores and anything else where the size of the number is the finding.

They can’t show us the small segments that matter. Clustering also means that the users at the edges disappear: the share of subscribers already on their way to cancelling, or the small core of superfans who account for a disproportionate slice of revenue. Those groups are small because they’re specific, and specificity is something a synthetic panel loses.

They tilt in a consistent direction. Synthetic respondents tend to be WEIRD (Western, Educated, Industrialised, Rich, and Democratic). They come out more comfortable with technology, less worried about privacy, better informed and more agreeable than the people they’re modelled on.

They shouldn’t speak for communities we haven’t spoken to. A team in the US building a WhatsApp media-literacy service for the Indian diaspora tried synthetic personas and got the diaspora as Westernised media imagines it. They caught it because they knew what the actual community sounds like. This is the kind of stereotype that synthetic responses can perpetuate.

How to start experimenting with synthetic audiences

Running a synthetic panel takes four steps, and each exists to make the next one worth doing.

The first fixes what you’re actually trying to settle, so you don’t generate material nobody acts on. The second decides what evidence the panel stands on, which sets how much weight its answers can carry. The third picks the kind of question you’re putting to it, which determines whether the answer is usable at all. The fourth gets an honest answer out, then records it in a form a colleague can judge for themselves rather than take on trust.

Set up once, the marginal cost of the next question is close to nothing. That’s the point. You’re building something you put questions to whenever they come up.

1. Fix the decision before you write the question

Most synthetic research fails before a panel is ever built, by producing interesting material nobody was going to act on in the first place.

Understand what you’re trying to achieve – what decision you’re hoping to make. Then determine the insights you need to inform the decision. Then figure out how to get those insights from a synthetic panel.

This ensures that the work you’re going to do is aligned with what you’re working towards. This is the step that stops the output becoming a slide nobody uses.

How to work towards objectives

2. Decide what data the panel is standing on

What a panel stands on sets the ceiling on everything after it. The same question put to a well-grounded panel and an ungrounded one produces answers that may look similar but instead mean completely different things.

There are three levels:

  • Nothing. A prompt describing a reader, with no evidence underneath it. What comes back is the model’s general impression of people like that.
  • Demographics. Market or audience data about the segment you’re aiming at. Better, and still a description of a category rather than of anybody.
  • Real material from real readers. Words that actual people you can name said or wrote.

The test is whether you can name the humans whose words are in the system. If you can’t, you’re on one of the first two levels.

Reaching the third level may be easier than it looks, depending on what your newsroom has kept. Worth looking for:

  • Past interview transcripts, from any research project, however old
  • Open-text answers from surveys (which usually go unread)
  • Subscriber cancellation reasons
  • Support and complaints emails
  • Comments and community posts
How to consolidate data

Prepare the material by anonymising the person and keeping the detail that makes them specific: their circumstances, and what they came to you for. Spread it across your segments rather than piling it up. Fifteen people covering five reader types beats forty of the same type.

3. Pick the job

There’s more than one kind of question you can put to a synthetic panel. Two are worth setting out here, because most teams use the first but never the second. Choosing deliberately between them is what separates a usable answer from a merely plausible one.

  • Count them. Put the same question to the whole panel and use the aggregate: which of six concepts is weakest, or which objection comes up most often. Read the ranking and ignore the margins. This is the workhorse, and it’s subject to the tilt described earlier.
  • Analyse against them. Instead of asking how users would react to a membership appeal, ask where that appeal conflicts with what you’ve documented about them. The model isn’t being asked to be anyone. It compares your draft against a profile you wrote and can check, then hands back specific points of friction you can verify or throw out. Most teams never try this, and it’s the job best matched to what these systems genuinely do well, which is analyse language.

4. Make it argue back, then record what you did

Left alone, these models are agreeable, and an agreeable panel tends to tell you what you already believe.

Three things to build into the prompt:

  • Tell it to answer as a reader rather than as an analyst. Otherwise the responses arrive sounding like a dull report, with tidy conclusions no human user would share.
  • Tell it that being indifferent is a valid answer. Plenty of readers genuinely don’t care, and a panel where everyone holds a strong view is already wrong.
  • Show it a few examples of how real people answered a similar question, with any polished phrasing taken out.

Then narrow with the panel and validate the surviving one or two options with real people. You’re only testing what has already survived something, which is where the saving comes from.

Finally, write up what you ran: who the panel was defined as, what it was grounded in, what you showed it, what you asked, what came back, and what still needs checking. Synthetic output is convincing whether or not it happens to be right. In one study, experienced researchers reviewing persona material could not reliably say which of it had been written by people. Making the output less convincing doesn’t help either, because results that read as obviously machine-made don’t get acted on. Since you can’t remove the believability, the write-up is what stops a plausible answer becoming a fact by repetition.

Start with the research you already have

The case for synthetic audiences isn’t that they closely match real readers. Academic testing has consistently found that they don’t. The case is that most audience questions in most newsrooms are currently answered by whoever is most senior in the room, and a synthetic panel grounded in things our actual readers said is better than that.

So, use them to narrow, not to decide. Ground them in real material, and know which level you’re on. Ask them to compare and to analyse. Keep them away from magnitudes, from the small segments that may matter most, and from communities you haven’t met. And, finally, validate what survives with people.

To get started: find the research your newsroom already commissioned and hasn’t opened since – the launch interviews, the subscriber survey, last year’s churn study – and work out which of this month’s decisions it could answer, if only you could ask it questions.