Can synthetic AI measurement predict what real users see?
Visibility in AI is usually measured by sending controlled prompts to models. But a person gets the answer in their own account, with their own history, memory and settings. The protocol below tests how far apart those two worlds are.
When a brand appears in 30% of an AI's answers, what does that mean? That 30% of people see it? Not necessarily. It means it appears in 30% of the answers generated under the experiment's conditions. Those are two different statements, and a dashboard rarely says which one it is showing.
This study doesn't assume synthetic measurement is correct. It treats it as a hypothesis and describes the experiment that can confirm it, limit it or reject it. All three outcomes are useful.
- AI-visibility benchmarks sample model outputs under controlled conditions, not answers received by users. How closely the two correspond remains to be shown.
- On the same commercial intent, the study compares the brand distribution from synthetic queries with the one from answers real users receive in their own AI accounts, using quantitative metrics and confidence intervals.
- The hypothesis concerns the aggregate distribution, not the individual answer. The result may be strong agreement, conditional agreement by category, or weak agreement, and all three get published.

Status: protocol published, no real-user data. The real-user panel has not been collected, so this page contains no result of the comparison. The numbers in the examples are illustrative.
RAVI (Reality-Adjusted Visibility Intelligence) is our tool for measuring market share inside AI answers, and this study was conceived as its validation layer. The tool is switched off; the research question remains open.
30% in a benchmark doesn't mean 30% of people
„Brand A appeared in 30% of responses generated under the experimental conditions defined by the methodology.”
„30% of real users see Brand A.”
A real AI answer depends on far more than the current prompt. A synthetic monitor controls only part of that environment; a real user lives inside the rest of it:
- conversation history
- account memory
- personalization
- location
- product settings
- model routing
- active tools
- time of the query
- stochastic variability
The same question, two populations
Controlled, systematic, repeated queries against AI systems.
The synthetic recommendation distribution
Ps(B | I, M, Cs)
- B
- the recommended or mentioned brand
- I
- the intent / prompt
- M
- the AI model
- Cs
- the experiment's standardized context
Real people, in their own AI accounts, receive a controlled research task.
The real-user recommendation distribution
Pr(B | I, M, U, Cr)
- B
- the recommended or mentioned brand
- I
- the intent / prompt
- M
- the AI model
- U
- the user: history, memory, personalization
- Cr
- the real product and conversation context
D(Ps, Pr)
The study estimates this distance. It doesn't assume it is zero, and it doesn't assume it is large: it is the quantity we measure.
The central question and the hypotheses
To what extent does the brand-recommendation distribution obtained through standardized synthetic querying reproduce the distribution received by real users in their own AI environments?
An opinion poll works because it asks people. You don't ask simulated voters and assume they represent the electorate. Here it is the other way round: you send 10,000 prompts to a model and sample model outputs, not users.
- H₁. Across a sufficiently large and representative set of commercial intents, Ps(B | I) ≈ Pr(B | I): the synthetic distribution can work as a proxy for the real one.
- H₀.The synthetic distribution doesn't reproduce the real one well enough to serve as a proxy for real exposure.
The hypothesis does not say an API answer resembles a user's answer. It says something narrower: individual answers may differ a lot while the aggregate brand distribution still converges. These are two different hypotheses: the literature warns against the first, and as far as we found, the second has not been tested for brand recommendations.
Three experiments and one rule
- 01
Controlled prompt: user variance
A large group of participants receives exactly the same prompt, in their own accounts. The prompt stays constant, the users change. If the distributions differ a lot, personalization and product context matter. If they converge, standardized measurement becomes more plausible.
- 02
Intent variation: phrasing variance
People don't phrase things alike. “Best bank for UK–China B2B payments” and “Which bank should my company use for suppliers in China?” express the same intent. Participants and the synthetic system receive the same variants, and phrasing variance is compared with user variance.
- 03
Many intents: the limits of the method
The process repeats across many intents, industries (banking, insurance, technology, automotive, SaaS), recommendation types, AI systems and phrasings. It may turn out to work well for generic product recommendations and less well for personalized financial decisions. That doesn't invalidate the method; it draws its boundaries.
- Note
We look for the cases where it fails
A serious benchmark doesn't stop at where it works. If agreement varies by category, each market can get its own reliability score, instead of the same certainty for all.
Seven metrics, not “looks about the same”
| Metric | What it asks |
|---|---|
| Recommendation-share error | How far is each brand's synthetic share from its observed real-user share? |
| Mean absolute error | What is the average error, in percentage points, across brands? |
| Rank correlation | Does synthetic measurement reproduce the competitive ordering? |
| Top-1 agreement | Does it identify the same market leader? |
| Top-3 overlap | Does it identify the same main set of recommendations? |
| Distribution distance | How similar are the complete probability distributions? |
| Confidence intervals | How much uncertainty surrounds the observed differences? |
The goal isn't “they look roughly similar”. It is “how similar are they, quantitatively”.
Three outcomes, all publishable
- A
Strong agreement
Synthetic and real distributions are consistently similar. Controlled querying can be a scalable proxy for aggregate exposure, and AI recommendation share gains a basis as a market metric.
- B
Conditional agreement
Synthetic measurement predicts real exposure well in some markets and poorly in others. The result is category-level reliability models: we know exactly where to trust it and where not to.
- C
Weak agreement
The differences are large. It would refute the strongest version of the “market share in AI” hypothesis and show that synthetic visibility methodologies don't represent real exposure without additional contextual modelling.
- Principle
A refuted hypothesis is still a result
If the data contradicts the hypothesis, we change the model, not the data. Methodological uncertainty gets published, not turned into false precision.
What we could say, and what we couldn't
| Level | The question | How it is obtained |
|---|---|---|
| 1 · AI recommendation share | What do AI systems recommend under standardized conditions? | Directly, through synthetic experiments. |
| 2 · Validated share | How well does standardized measurement reproduce the aggregate real-user distribution? | Through repeated human-panel validation. |
| 3 · Actual exposure | What recommendations are actually shown across the whole population of AI users? | Would require platform-level telemetry or representative observational data. |
We don't claim to have the level-3 dataset, and the distinction stays explicit. “Market share” needs a denominator: without validation, synthetic measurement only tells us how brands stand inside the experimental universe we built.
What we don't collect
The study doesn't ask participants to hand over their private conversation history: no months of messages, no documents, no full account profile. The unit of research is the assigned prompt and the response received, plus only the metadata needed to interpret the experiment:
- the AI product used and the visible model;
- the country and the time of the query;
- a new or an existing conversation;
- memory / personalization on or off.
The principle: the minimum information needed to answer the question.
What is already known, and what we didn't find
Three bodies of literature converge on the problem. None of them answers our exact question.
- Visibility in AI is a distribution, not a point. GEO-bench measures visibility from sets of queries [1]; measuring once is unreliable, because answers vary across runs, prompts and time [3]; AI services differ from one another in phrasing sensitivity and source diversity [2].
- Personalization changes recommendations. A user's revealed identity significantly influences (p < 0.001) the recommendations of several consumer chatbots [4].
- Synthetic users are not real ones. Simulators favour popular items and correlate little with human preferences [5]; a synthetic-user win rate of 87% became 72% with real humans [6]; on 1,000 real dialogues across 16 domains, simulated users are easier interlocutors: less friction, more positive feedback [7]; and real users found nine kinds of personalization errors that LLM judges had missed [8].
The gap.In the sources we consulted, we didn't find a study that compares the brand distribution obtained through standardized synthetic querying with the one received by real users, in their own accounts, on the same commercial intent. This is not a systematic review and we don't claim to be first. The literature says neither that synthetic equals real nor that it is unrelated; it says the relationship has to be validated empirically and may vary by domain. If you know a study that does exactly this, write to us.
Questions about the study
What is The AI Recommendation Reality Study?
A research protocol that tests whether the brand distribution recommended in controlled synthetic queries reproduces the distribution received by real users in their own AI accounts. The page publishes the problem, the hypotheses, the experimental design and the metrics. It contains no results of the comparison.
Does the study have results?
No. The real-user panel has not been collected, and the numbers in the page's examples are illustrative. We publish the protocol so it can be criticised, replicated or picked up by someone else.
Why isn't querying an API thousands of times enough?
Because that samples model outputs under controlled conditions, not answers received by people. A real user gets the answer inside an environment with history, memory, personalization, location and active tools. Whether those differences wash out in the aggregate is exactly what the study tests.
What would a negative result mean?
That some of the AI-visibility metrics used today describe model behaviour, not real user exposure, and need to be limited, recalibrated or reinterpreted. It would be as publishable as a positive one.
What data would a participant provide?
Only the assigned prompt, the response received and minimal metadata: the AI product, the visible model, the country, the time, whether the conversation was new and whether memory or personalization was on. Not their private conversation history.
Cited literature
The title, authors, year and venue of each paper were checked against the paper's own page. The list is not a systematic literature review.
- 1linkAggarwal et al.2024
GEO: Generative Engine Optimization
KDD 2024. Builds GEO-bench, a benchmark of diverse queries and relevant web sources, to evaluate content visibility in generative-engine answers.
- 2linkChen et al.2025
Generative Engine Optimization: How to Dominate AI Search
arXiv 2509.08919. Controlled experiments across several AI services: a systematic bias towards earned media, and differences between services in domain diversity, freshness, cross-language stability and phrasing sensitivity.
- 3linkSchulte et al.2026
Don't Measure Once: Measuring Visibility in AI Search (GEO)
arXiv 2604.07585. Argues that visibility in AI search must be measured repeatedly and characterised as a distribution, not a fixed point.
- 4linkKantharuban et al.2025
Stereotype or Personalization? User Identity Biases Chatbot Recommendations
Findings of ACL 2025. A user's revealed identity significantly influences the recommendations of several consumer chatbots, without being transparently indicated.
- 5linkYoon et al.2024
Evaluating Large Language Models as Generative User Simulators for Conversational Recommendation
NAACL 2024. A five-task protocol for synthetic users: simulators favour popular items and correlate little with human preferences.
- 6linkSingh et al.2026
FSPO: Few-Shot Optimization of Synthetic Preferences Effectively Personalizes to Real Users
ICLR 2026. Up to 1,500 synthetic users; 87% win rate with synthetic users, 72% with real humans, on open-ended question answering.
- 7linkLiu et al.2026
Synthetic Users, Real Differences: an Evaluation Framework for User Simulation in Multi-Turn Conversations
arXiv 2605.02624 (realsim). 1,000 real dialogues, 16 domains, eight dimensions: simulated users introduce less friction and give more positive feedback and context.
- 8linkBalepur et al.2026
Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users
ACL 2026. Nine kinds of personalization errors, undetectable by LLM judges, found only with real users.
Got a criticism or a missing reference?
The protocol is open. If you see a weak point in the design, a paper we didn't find, or you want to replicate the study on another market, write to us.