AEOsim Blog
Play the Games You Might Lose: Item Information Peaks at P=0.5, Prompt Difficulty, and Why Always-Win Dashboards Waste Compute
Article 4/n on synthetic data in LLM visibility tools. See Article 1: How we Build the Prompt Set and Article 2: Bridge The Real2Sim Gap.
Table of Contents
Imagine a game with 100 levels.
As you read this next line, imagine your cat just spawned into this game's world. You have 9 lives. In front of your eyes, text flashes past: "Your mission, should you choose to accept it: figure out what level your cat is at."
The rules of the game are simple:
- You do not know what level your cat is at. You can assume (A) it is randomly generated and can be any number between 1 and 100.
- The game has 100 levels which you can arbitrarily pick and choose from.
- Every time you play a game-level, you lose one life, irrespective of whether you succeed.
- You are almost sure to win levels that are much below yours, and almost sure to lose levels that are much above yours. However, you may win or lose game-levels near your cat-level. You have a >50% chance of winning a level below you, <50% chance of winning a level above you, and a 50/50 chance of winning the exact level you are on.
- You have 9 lives. You retain the memories of your previous life every time you respawn.
- You have to estimate what level you are at with the best possible accuracy and precision. For example, if your cat-level (unknown to you) is 35 and you guess you are between cat-levels 30 and 40, this is better than guessing that you are between 25 and 45, because the former is more precise.
Strange game.
But this is a simplified version of one of the problems we faced while making our synthetic prompt sets for AI visibility tools. We have previously talked about how synthetic data can be more useful than real data, and today want to reinforce that intuition.
Your problem: maximize the information gained per life. Our problem: maximize the information gained per multi-turn chat unit.
Life 1: play level 50, not 1 or 100
You have to pick a level. What do you play?
Level 1 is tempting. It is safe, you will almost certainly win, and winning feels like progress. It is also a wasted life. Before you played it you were already sure you would win. After you played it you are still sure you would win. You spent a life and your estimate of your own level did not move by a hair.
Level 100 is exactly as bad. You lose, you learned that you are not above 100, which you already knew.
Play level 50 instead. Before you play it you have no idea what will happen. Win and you know you are probably above 50. Lose and you know you are probably below it. Either outcome moves you.
So:
An outcome you can predict teaches you nothing. Only the games you might lose are worth playing.
The rest of this post attaches arithmetic to that.
Rule 4 as a logistic: P(win) = σ(a · (θ − j))
Rule 4 describes a curve. Write your unknown cat-level as θ and the level you choose to play as j. Then
P(win) = σ( a · (θ − j) )
where σ is the logistic function, which turns any number into a probability between 0 and 1. That is rule 4 restated. When you play a level below yours, θ − j is positive and your win chance is above half. When you play a level above yours, it is negative and your win chance is below half. When you play your own level exactly, θ − j is zero and the curve reads 0.5.
a is the steepness. It controls how fast "almost sure to win" turns into "almost sure to lose". A big a means a knife edge: three levels below you is a walkover, three above is hopeless. A small a means a gentle slope where even distant levels are coin flips.
For a cat sitting at level 35:
Ninety of those hundred levels have an outcome you could have called in advance.
Information I = a² · P · (1 − P)
Now put a number on what each level teaches you. That quantity is called information, and for this curve it is
I = a² · P · (1 − P)
where P is your chance of winning the level you chose. That product P(1−P) is at its largest when P is one half, and it collapses toward zero as P approaches either 0 or 1.
So the arithmetic agrees with the intuition. A level you win 99% of the time carries about 0.01 of the quantity. A level you win half the time carries 0.25. The coin-flip level is worth twenty-five of the easy ones.
Outside that band you are spending lives to confirm what you already believed. The same peak at P ≈ 0.5 is the reason Article 2 argued that volume-weighted head prompts are often dead spend, and why misalignment monitoring should favour families near even odds rather than settled ones.
Aim at your current 50/50 estimate, then update
Look at the peak again. It sits exactly at your own level.
So to play the most informative level, you would need to already know your level, which is the thing you spawned here not knowing.
You do not need the true answer, only your current best guess. Play the level you would bet 50/50 on. Watch what happens. Update. Play the new level you would bet 50/50 on. Rule 5 gave you memory across lives precisely so you can do this, and without memory across lives, nine of them would buy you almost nothing.
Life 1, play 50. You win, so you are probably higher. Life 2, play 70. You lose, so you are probably between. Life 3, play 60. And so on, each play landing near the current estimate, each one a genuine coin flip, each one worth twenty-five of the plays a more timid cat would have made.
This is what a modern computerised exam does when it gives you a harder question after you answer one correctly. It is chasing the peak of that curve. Article 1 is how we build the prompt set that makes that chase possible in visibility tools.
Nine peak plays vs nine random plays
Information adds up across plays. Nine plays at the peak give you 9 × 0.25 = 2.25, and the width of your final estimate is one over the square root of that, about ±0.67 levels.
Now suppose you had ignored all of this and picked nine levels at random. Most would land in the flat regions and contribute almost nothing. Average information per random play works out to roughly 0.01, so nine random lives give you 0.09, and your final estimate is about ±3.3 levels wide.
To reach the precision that nine well-chosen lives gave you, the random cat needs around 225.
Same game, same rules, same cost per life, twenty-five times the budget, purely from where the lives were spent.
Map the game onto AI visibility prompts
Everything above was a story about a game. Change five nouns and it is our problem.
| In the game | In AI visibility |
|---|---|
| Your cat-level θ | How strongly an engine associates a brand with a category |
| A game-level j | A prompt, with its own difficulty |
| Playing a level | Running a multi-turn chat and seeing what the engine says |
| Winning | The brand shows up, or gets recommended |
| A life | A unit of compute, which is a real invoice |
| Steepness a | How sharply that prompt separates strong brands from weak ones |
A prompt where every brand in the category appears is level 1. A prompt where only the market leader ever appears is level 100. Both are on somebody's dashboard right now, refreshing weekly and teaching nobody anything.
And because the peak sits at your level, the informative prompts for a challenger are not the informative prompts for the incumbent. A single prompt set sold to every client in a category is, for most of those clients, a list of level 1s and level 100s.
Five places reality is harder than the game
Games have tidy rules. Ours does not. Five places where reality is harder, roughly in order of how much trouble they cause.
(A) Your level was never uniformly random. Rule 1 said any number from 1 to 100 with equal chance, which is why level 50 was the right opening move. Real brands are not drawn from a hat. You usually know roughly where a brand stands before you run anything, from category position, from the last measurement cycle, from what the sales team hears. That prior lets you start life 1 near the informative band instead of burning plays to find it. It also cuts the other way: a wrong prior sends you confidently to the flat part of the curve, and a few unlucky outcomes there look exactly like a correct prior. Kojable's baseline guidance is how you keep that prior honest across cycles.
(B) You were told the ordering, not the probabilities. Rule 4 promised that higher levels are harder. It never told you how much harder, and the honest version of this game does not hand you a either. Worse, a is not one number: every prompt has its own steepness, and a prompt with a shallow slope is nearly worthless no matter how well you aim, because even a perfectly targeted play barely moves your estimate. So you are estimating the properties of the instrument and the thing being measured at the same time, from the same data. That is solvable, but it needs more plays than the tidy version, and some of those plays go entirely to pricing the prompts.
(C) The levels do not come numbered. In the game, level 60 announces itself as level 60. A prompt arrives with no label. Nothing in its text tells you whether it is a walkover or a wall, and the same prompt has different difficulty against a different category, on a different engine, in a different month.
Call that map the Chat Difficulty Function: prompt in, difficulty out, before you have spent anything running it. Without it you cannot aim, and everything above depended on aiming. Two ways at it. Measure difficulty by piloting prompts and watching what happens, which works and costs lives. Or predict it from features of the prompt, how many constraints it carries, how specific they are, how deep in the conversation the decision is asked for, then calibrate that predictor once so it prices new prompts for free. Prediction is largely open. Conversation Analytics is where those constraint and decision turns live in the buying chat.
(D) A chat is not one play. Rule 3 charged you one life per level, and each play was independent of the last. A multi-turn conversation is not five independent plays. The engine takes a position in turn two and carries it forward, so the turns inside one chat agree with each other more than they should. Five correlated turns are worth substantially less than five separate coin flips, and counting them as five will convince you that you have five times the precision you have. Clarification as a Branch Point shows how a single clarifying answer can reshuffle the set far more than repeating the same branch.
(E) Your cat keeps levelling up mid-game. Rule 1 fixed θ at spawn. In our version the whole point of the exercise is that the client is actively trying to raise θ by publishing, fixing pages, and getting cited. A test-taker cannot get smarter halfway through the exam. A brand can, and is paying you to help. So the instrument has to distinguish movement in the brand from drift in the engine, which means holding some prompts fixed as reference points across cycles even when better-aimed ones exist. Pure information maximisation would retire those anchors. Do not.
Play the games you might lose
One line, from life 1.
Play the games you might lose.
A dashboard full of prompts your brand always wins is a cat spending every life on level 1. The number on it will be high, and stable, and it will stay high and stable through a quarter in which you lost ground everywhere that mattered, because nothing on the board was ever capable of moving.
You pay full price for a number that was already determined before you ran anything. Run the adaptive panel in product with Analytika.
Kojable: redundancy, clustering, and representative prompts
There is another efficiency problem alongside prompt difficulty: redundancy. Even if every prompt is individually relevant, a large monitoring library can still spend substantial compute repeatedly measuring questions that are different in wording but very similar in meaning.
Kojable explored this in a study of 180 B2B finance prompts run through Gemini. The prompts deliberately varied by persona, industry, geography, integration requirements, intent and phrasing rather than simply swapping synonyms. Across 16,110 prompt pairs, semantic similarity between prompts was strongly associated with similarity between the overall answers they produced (r = 0.878). In practical terms, prompts that were closer in meaning tended to generate answers that were closer in meaning as well.
That creates an important distinction between prompt breadth and information gain. Kojable's broader AI-visibility measurement guide puts it simply: "More prompts can increase breadth. They do not automatically increase relevance." A prompt library ultimately defines the buyer universe that a visibility system measures. A set dominated by branded questions, for example, measures something quite different from one built around category discovery, comparisons, use cases and recommendation questions.
One practical approach is therefore to build the prompt universe around real buyer questions first, cluster questions that express similar underlying needs, and then select representative prompts from those clusters. Prompt difficulty could add another selection layer: among commercially meaningful questions, favour prompts where the outcome is sufficiently uncertain to reveal something about the brand's current position rather than repeatedly confirming an outcome that is already highly predictable.
That does not mean one representative prompt can safely replace every variation. Kojable's study explicitly cautions that "A representative prompt is useful, not automatically sufficient." Two prompts can produce broadly similar answers while still differing in the details that matter commercially: whether a particular brand appears, which company is recommended first, what sources are cited, or whether the brand is described accurately.
A more efficient measurement design may therefore use clusters rather than an ever-growing flat prompt list: a representative seed prompt for recurring monitoring, selected validation prompts for commercially important variations, and a smaller number of stable anchor prompts preserved over time so that genuine movement can still be distinguished from changes in the measurement panel itself. Kojable's measurement guidance similarly argues that comparable verification requires keeping prompt versions, platforms, markets, repeated-run design and sampling rules visible across measurement cycles.
The objective is not simply to track fewer prompts, or more prompts. It is to construct a measurement panel in which each additional prompt earns its place by contributing either new buyer coverage, new information about competitive position, or validation of an outcome important enough that redundancy is worth the cost.
References
- Lord, F.M. (1980). Applications of Item Response Theory to Practical Testing Problems. Lawrence Erlbaum Associates.
- Kojable (2026). AI Prompt Similarity: 180 Gemini Prompts Tested.
- Kojable. AI visibility monitoring and baselines.
Arnav Narang is founder of AEOsim, and Piush Vaish is founder of Kojable. This is Article 4/n on synthetic data in LLM visibility tools. Read Article 1, Article 2, and Article 3. Measure multi-turn AI visibility with Analytika.