So what?
What can Telltail do for you?
You already paid to run the evaluation. Telltail helps you use its question-by-question results to decide what to fix, test, or fund next.
- Find the hard cases by name. Open the exact questions everyone missed, only one system solved, or that still separate the field.
- Know what to try next. See whether a stronger model, another attempt, or more search reaches new questions.
- Reveal which LLMs are reliable. Find out which tasks truly discriminate between different model choices.
- Spend money and compute where it counts. Re-rank tasks by trainability, instead of repeated random sampling.
First: Two "equal" models.
A tie? Or many unique stories?
Let's say you have a four-question test. ChatGPT gets two questions right. Claude gets two questions right. On a leaderboard, both score 50%. Most people today would call them "equal" and move on.
But a score tells us how many, not which. They might solve the same two questions. They might share one. Or they might solve completely different questions, covering all four between them. Select the three different overlap patterns below. The two scores never change.
Next: four models, the same four questions
The map shows what the average erased
Now let's imagine that four models answer the same four questions. The results below divide the map into four regions.
- Success ceiling: every model answers the question correctly, so the success is shared across the entire group.
- Mixed interior: more than one model answers correctly, but at least one model does not, revealing disagreement within the group.
- Singleton frontier: exactly one model answers correctly, showing a capability that none of the other models demonstrate.
- Failure floor: every model answers incorrectly, so the failure is shared across the entire group.
Select a question to name its pattern.
Telltail gives the two ends short names. is the share of questions everybody missed. is the share everybody solved. Here each is 25%.
Those two regions share a quirk: they cannot rank the models. On a floor question every model got it wrong; on a ceiling question every model got it right. Either way the models look identical, so half of this test says nothing about which model is better. The leaderboard order is decided entirely by the other two regions.
These numbers describe this selected group of models and questions. Because each map is unique to the set of questions and models that were chosen, the results are always relative to at least three decisions: which models were tried, which questions were asked, and how they were scored.
Now: run the same evaluation again
Language models are non-deterministic
Language models do not always give the same answer twice. So we run the same models A–D on the same questions across R = 4 saved runs. R is the total number of runs we saved. Each square shows one dot for every run currently included; the dots are repeated tries, not new models.
Each model hides a different story. A’s score swings from 1/4 to 4/4. B has a fixed blind spot: it always solves Q1 and never solves Q3. C has two strong questions it solves 3 times out of 4 and two weak questions it solves only once. D’s wins move around: it alternates between two pairs, getting every question right twice and wrong twice.
Yet B, C, and D score exactly 2/4 in every run. Press Scramble run order, then set k: how many of those R saved runs you choose to look at. R is what you have; k is what you choose. The figure introduces two common AI-evaluation ideas: pass@k, the chance at least one of k attempts passes, and pass^k, the chance all k attempts pass. The dots use the first k saved runs; the summaries consider every possible k-run subset of the R saved runs. Scrambling changes which outcomes appear first, but not the summaries. No model changed. Only the sampled runs changed.
Each cell shows exactly the first 2 of 4 run outcomes. The order-free summaries below choose k distinct runs uniformly from all four saved runs.
As k grows, pass@k can only rise, while pass^k can only fall.
| Model | Observed any-success shown first k runs | pass@k at least one of k passes | pass^k all k pass |
|---|---|---|---|
| A | 50.0% | 83.3% | 41.7% |
| B | 50.0% | 66.7% | 33.3% |
| C | 75.0% | 75.0% | 25.0% |
| D | 100.0% | 83.3% | 16.7% |
Why do we call it evaluation geometry?
An evaluation has a shape
Now let the same idea grow: 8 made-up models answer 64 made-up questions, with R = 8 saved runs.
We give the models different stories. One is steady. One is just as capable on average, but jumpy. One has a fixed blind spot. Two are good at different subjects. One resembles another model. One gets the same number right each run, but not the same questions. One rarely has a very bad run.
The questions fall into the same four regions from the smaller map: success ceiling, mixed interior, singleton frontier, and failure floor. Some stay in one region across runs; others move.
Each green square is a pass. Each red square is a fail. Put the highest-scoring models at the top, then put the questions passed by the most models on the left. The white success curve traces the question solve rate—the share of models that passed each question—so it runs from high on the left to low on the right.
The white curve shows the share of models that passed each question: 100% at the top, 0% at the bottom. The vertical lines mark region boundaries.
At this scale, the map begins with the success ceiling, runs through the mixed interior and singleton frontier, and ends at the failure floor. That larger pattern was hiding behind eight averages: easy head, disputed middle, hard tail, and model-specific holes.
Here we made the data, so we know why its patterns exist. In a real evaluation, that recipe is hidden. Telltail shows the shape without pretending it knows the cause.
Open up the third dimension
How can we see all attempted runs at once?
The last picture showed the same R = 8 saved runs one at a time. Together, those saved runs are a third dimension of the evaluation. Now keep all R visible at once.
Combine the runs and each cell becomes a c/R bar: how many of the R saved runs passed. Keep the run index and the same cell becomes R stripes in the displayed order. Shuffle that order to see what changes—and what cannot. These runs are declared interchangeable draws, so “Run 1” is just a label: shuffling changes which selection you are looking at, never what was measured.
8 models × 64 questions × 8 runs
Put R inside each cell
The matrix never changes its data. Choose whether each cell combines the runs into one c/R bar or preserves their displayed order as R stripes.
Now we can start to see some "topology" emerging. This is what we mean by the "outcome geometry" of the evaluation.
With one run per cell, Telltail’s four regions told the whole story. With R runs, each question needs two answers: how many models ever reach it, and whether they hold it every run. The familiar names now need qualifiers. The success ceiling splits into an all-run success ceiling (every model, every run) and an intermittent success ceiling (every model at least once, but not every time). The singleton frontier splits the same way. The mixed interior stays one zone, and the failure floor is always all-run: a single success anywhere already leaves it. The columns below are grouped into all six zones, with five white dividers separating them. The per-run stripes show why each column earned its qualifier.
Model rows can be sorted three ways. Each is an observed share over all R saved runs and requires no independence assumption: average per run (the per-attempt rate), at least once (the share of questions the model ever solved), and every run (the share it solved on all R runs, using the same 1[c = R] fact behind the all-run zones). Those rankings need not agree. The group order and the average-per-run curve provably cannot both be monotone at once. That disagreement is real structure, not a rendering bug.
If you remember nothing else from scrolling this far, remember this:
The mean is not the map.
Why Telltail exists
The mean is useful. It is just not the whole evaluation.
A leaderboard average compresses a whole grid into one number. In the process it throws away which questions were shared wins, shared failures, disagreements, or lucky reruns. Telltail keeps that structure visible and keeps the question names behind every count.
What structure existed in the evaluation before we averaged it away?
Don’t settle for the average.
Tune into the tail
Sign up for updates about telltail.