← All posts

AEOsim Blog

Bridge The Real2Sim Gap, not Sim2Real: Why Simulated Data Beats Real Data in AI Visibility Tools

Article 2/n on synthetic data in LLM visibility tools. See Article 1: How we Build the Prompt Set.

Table of Contents

Our last post covered how a prompt set gets built: prompt families, randomisation inside them, an audit trail, and a 60/20/20 split across real, task-elicited, and generated prompts. This one answers the objection that follows.

If the prompts are made up, what is the measurement worth?

Often more than the real alternative.

Real data can be worthless

A dataset is worth what it reduces your uncertainty about the thing you are measuring. Origin does not appear in that definition. A prompt whose answer you could have predicted carries no information, however faithfully it was harvested.

Most observed query volume is that. The head of any category is questions with stable answers where your position is known. Log ten thousand and the number moves when the query mix shifts and sits still when your position changes.

So: which prompts carry information about your position? Standardised testing has spent a century on that question.

What standardised testing figured out

Nobody builds the SAT or the JEE from questions students happen to ask. Every item is authored, measured, then kept or cut on evidence. The framework is item response theory, and some of it transfers.

Difficulty and discrimination decide an item, not content. The standard two-parameter model says a candidate of ability θ gets item i right with probability

P(θ) = σ( a·(θ − b) )

where b is difficulty and a is discrimination, the slope of the curve. A good question on an important topic with a flat slope gets cut.

The same form transfers. Read θ as a brand's strength in a category, b as how strong a brand must be to get recommended on that prompt, and a as how sharply mention-probability separates strong brands from weak ones. A prompt every competitor wins and a prompt only the leader wins are both low-information.

Information peaks near difficulty. How much an item tells you about a given candidate is

I(θ) = a²·P(θ)·(1 − P(θ))

which is maximised at P = 0.5. The most informative prompts for a brand are the ones it wins about half the time. Prompts you always win and prompts you never win are both dead spend, and this is client-specific: the prompts that discriminate for a challenger are not the ones that discriminate for the incumbent. Volume cannot fix it. One prompt set sold to every client in a category measures a few of them well and the others badly.

Take a mid-market project management tool. It shows up on "best project management software" in 95% of runs, on "project management tool for a 40-person agency that bills hourly" in about half, and on "enterprise PM platform for regulated industries" in 2%. Plug those into the formula and the middle prompt carries roughly five times the information of the first and twelve times the third, at identical cost per run. Two of the three are on the dashboard telling you nothing.

Items are calibrated before they count. Exams embed unscored trial items in live sittings, and items that fail their statistics never score anybody. The equivalent is a pilot round estimating a prompt's discrimination before it enters the reported number. Most prompt sets have never had this done, so nobody knows which prompts are doing work.

Anchor items keep scores comparable across years. When a form retires, a subset carries into the next one. This answers prompt set drift: a set that never changes goes stale, a set that changes wholesale destroys your time series. Freeze an anchor block, rotate the remainder, handle the boundary explicitly.

A blueprint constrains information maximisation. Adaptive tests maximise information under content-balancing constraints, because optimising for information alone measures precisely and represents nothing. Coverage across intent classes is your blueprint. Discrimination is what you optimise inside it.

Exposure is managed, because known items stop working. Here the pressure is stronger, since clients optimising against the set is intended behaviour. A reported prompt is a target from the moment it is reported, and held-out families are the only way to separate real gains from gains against the instrument.

Where it breaks. IRT assumes one latent trait and independent items. Visibility is not one trait, and prompts from the same category share enough wording that outcomes correlate. Test-takers cannot change their ability mid-exam, whereas moving the measured quantity is the point of your programme. Use it for design discipline, not as a model to fit.

Two adjacent ideas finish the picture. Optimal experimental design places observations where estimates are uncertain rather than in proportion to natural frequency, so a prompt you win every week is a spent design point. Stratified sampling allocates toward strata where responses disagree most, so follow variance across intent classes rather than search volume. Migration questions, objection-shaped questions, compliance-gated evaluations and comparisons against an incumbent are rare and highly variable, and volume weighting starves all four.

One caveat: selection introduces bias. A chosen sample can drift from the population and mislead on absolute levels even while it sharpens comparisons. Keep a small volume-weighted panel alongside as a calibration check.

Three things simulation does that observation cannot

Adversarial probing. What does the engine say when the prompt carries your top three sales objections in a buyer's words, or names a competitor first? Waiting for real users to ask is how you find out after the deal is lost. The same control protects the metric: a clean branded and unbranded split can be audited for poisoning, and an organic mix arrives contaminated by people who already knew your brand.

Forced elicitation. Engines resist picking. Ask openly and you get a shortlist with no recommendation in it, so share-of-recommendation computed from that text is counting mentions under a better name. Ask an engine to commit and it commits. In one category, terminal turns demanding a single choice produced a pick in nearly every conversation, and the shares separated brands far more sharply than the mention grid, which had them clustered within a few points.

Multi-turn control. Buying conversations survey early and commit late, so the decision sits at the end of a trajectory that sampling opening questions never reaches. Simulation lets you fix which constraint enters at which turn and vary one thing at a time. This is a method we developed at AEOsim for Analytika, and you can read more about it in our guest article on Conversation Analytics, published on The GEO Community, or the AEOsim republish.

That's a non-exhaustive list.

Synthetic is a spectrum, not binary

Provenance still matters, on a separate axis from usefulness, and it runs continuously. One overarching buyer intent → five positions:

  1. Transcribed word for word from a sales call
  2. The same question with client name and jargon normalised out
  3. A customer given the task and asked to type what they would send
  4. A template built from constraints observed across calls, filled with plausible values
  5. A model handed the category name and asked to produce prompts

A set cannot be labelled real or simulated, only described by its mix, which is why 60/20/20 works better as a disclosure format than a target ratio. The right mix differs by category.

Position on the ladder does not determine information value. Rung one can be a settled question that separates nobody, rung five can be the probe that finds a real gap. The ladder determines failure mode: top rungs risk being unrepresentative of scale, bottom rungs risk model voice, a tidiness real buyers lack, which biases you toward content written in the same register.

Measurement starts with a baseline

A prompt set is only useful if you know what it is being compared against.

Before optimising, establish a baseline: which engines are being measured, which competitors are in the set, which intent classes are covered, where the brand is currently mentioned, recommended, omitted, or misrepresented, and which results are stable enough to treat as signal.

Treat that baseline as a reference point for later movement, not as the score itself. Without it, an increase in visibility could reflect a better prompt mix, a model change, a shift in competitor coverage, or a real improvement in how the engine represents your brand.

For a practical framework on setting that baseline, see Kojable's guide to AI visibility tracking baselines.

Where simulated sets fail

Saturation. Prompts lose discrimination as a category converges, and the set then reports stability that reflects the instrument.

Drift. A set authored eighteen months ago encodes an old buying conversation. Refreshing without anchors destroys comparability.

Goodhart. A reported set becomes a target. Held-out families are the check.

Unfalsifiability. A simulated set can be tuned until the client looks good. The guard is external validation: referral traffic from chat surfaces, questions logged by sales, shifts in inbound phrasing.

What to ask a vendor

Not whether the prompts are real. Ask which prompts currently separate you from your competitors, and on what evidence. Ask what happens to a prompt when it stops separating anything. Ask whether the set is positioned against your standing or shared across the category. Ask what carries across a refresh, who can add prompts, and whether anything in the set is adversarial to you.

A well-built simulated set answers all of those. "Our data is real" answers none of them.

References

  • Kolen, M.J. & Brennan, R.L. (2014). Test Equating, Scaling, and Linking: Methods and Practices, 3rd ed. Springer.
  • Lord, F.M. (1980). Applications of Item Response Theory to Practical Testing Problems. Lawrence Erlbaum Associates.
  • Martínez-Plumed, F., Prudêncio, R.B.C., Martínez-Usó, A. & Hernández-Orallo, J. (2016). Making sense of item response theory in machine learning. European Conference on Artificial Intelligence (ECAI), pp. 1140-1148.
  • Settles, B. (2009). Active learning literature survey. Computer Sciences Technical Report 1648, University of Wisconsin-Madison.

Arnav Narang is founder of AEOsim, and Piush Vaish is founder of Kojable. This is Article 2/n on synthetic data in LLM visibility tools. Read Article 1. Measure multi-turn AI visibility with Analytika.