Telltail

Benchmarks report means, and throw away the map that generated them.
Telltail is a statistical analysis tool for exploring that map.
The map is the full evaluation matrix of × .

So what?

What can Telltail do for you?

You already paid to run the evaluation. Telltail helps you use its question-by-question results to decide what to fix, test, or fund next.

  • Find the hard cases by name. Open the exact questions everyone missed, only one system solved, or that still separate the field.
  • Know what to try next. See whether a stronger model, another attempt, or more search reaches new questions.
  • Reveal which LLMs are reliable. Find out which tasks truly discriminate between different model choices.
  • Spend money and compute where it counts. Re-rank tasks by trainability, instead of repeated random sampling.

First: Two "equal" models.

A tie? Or many unique stories?

Let's say you have a four-question test. ChatGPT gets two questions right. Claude gets two questions right. On a leaderboard, both score 50%. Most people today would call them "equal" and move on.

But a score tells us how many, not which. They might solve the same two questions. They might share one. Or they might solve completely different questions, covering all four between them. Select the three different overlap patterns below. The two scores never change.

ChatGPT 2/4 · 50%
Claude 2/4 · 50%
Q1
Q2
Q3
Q4
ChatGPT
Claude
Both right 2 of 4
Both wrong 2 of 4
At least one right 2 of 4
With ChatGPT’s answers fixed, these three patterns represent 1 + 4 + 1 = 6 assignments of the named questions. The two 50% scores never move.
Two identical scores can describe many answer patterns, including the opposite pattern.

Next: four models, the same four questions

The map shows what the average erased

Now let's imagine that four models answer the same four questions. The results below divide the map into four regions.

  • Success ceiling: every model answers the question correctly, so the success is shared across the entire group.
  • Mixed interior: more than one model answers correctly, but at least one model does not, revealing disagreement within the group.
  • Singleton frontier: exactly one model answers correctly, showing a capability that none of the other models demonstrate.
  • Failure floor: every model answers incorrectly, so the failure is shared across the entire group.

Select a question to name its pattern.

Model
A
B
C
D
Pattern
Success ceiling All four models got this question right.
1 of 4 · 25% questions everybody missed
1 of 4 · 25% questions everybody solved
The average is one number. The grid shows where that number came from.

Telltail gives the two ends short names. is the share of questions everybody missed. is the share everybody solved. Here each is 25%.

Those two regions share a quirk: they cannot rank the models. On a floor question every model got it wrong; on a ceiling question every model got it right. Either way the models look identical, so half of this test says nothing about which model is better. The leaderboard order is decided entirely by the other two regions.

These numbers describe this selected group of models and questions. Because each map is unique to the set of questions and models that were chosen, the results are always relative to at least three decisions: which models were tried, which questions were asked, and how they were scored.

Now: run the same evaluation again

Language models are non-deterministic

Language models do not always give the same answer twice. So we run the same models A–D on the same questions across R = 4 saved runs. R is the total number of runs we saved. Each square shows one dot for every run currently included; the dots are repeated tries, not new models.

Each model hides a different story. A’s score swings from 1/4 to 4/4. B has a fixed blind spot: it always solves Q1 and never solves Q3. C has two strong questions it solves 3 times out of 4 and two weak questions it solves only once. D’s wins move around: it alternates between two pairs, getting every question right twice and wrong twice.

Yet B, C, and D score exactly 2/4 in every run. Press Scramble run order, then set k: how many of those R saved runs you choose to look at. R is what you have; k is what you choose. The figure introduces two common AI-evaluation ideas: pass@k, the chance at least one of k attempts passes, and pass^k, the chance all k attempts pass. The dots use the first k saved runs; the summaries consider every possible k-run subset of the R saved runs. Scrambling changes which outcomes appear first, but not the summaries. No model changed. Only the sampled runs changed.

Sampling order Now: 1 → 2 → 3 → 4
Choose k from R = 4 saved runs
Q1
Q2
Q3
Q4
A
B
C
D

Each cell shows exactly the first 2 of 4 run outcomes. The order-free summaries below choose k distinct runs uniformly from all four saved runs.

As k grows, pass@k can only rise, while pass^k can only fall.

ModelObserved any-success
shown first k runs
pass@k
at least one of k passes
pass^k
all k pass
A50.0%83.3%41.7%
B50.0%66.7%33.3%
C75.0%75.0%25.0%
D100.0%83.3%16.7%
Exact probabilities for uniformly choosing k distinct runs from these four saved runs; they do not estimate or guarantee an unsaved future attempt. These finite saved-set summaries ignore display order; the shown first-k result does not.
Repeated question-level results reveal four unique stories that the averages hide.

Why do we call it evaluation geometry?

An evaluation has a shape

Now let the same idea grow: 8 made-up models answer 64 made-up questions, with R = 8 saved runs.

We give the models different stories. One is steady. One is just as capable on average, but jumpy. One has a fixed blind spot. Two are good at different subjects. One resembles another model. One gets the same number right each run, but not the same questions. One rarely has a very bad run.

The questions fall into the same four regions from the smaller map: success ceiling, mixed interior, singleton frontier, and failure floor. Some stay in one region across runs; others move.

Each green square is a pass. Each red square is a fail. Put the highest-scoring models at the top, then put the questions passed by the most models on the left. The white success curve traces the question solve rate—the share of models that passed each question—so it runs from high on the left to low on the right.

8 models × 64 questions
Run 1 of 8
success ceiling 1mixed interior 44singleton frontier 8failure floor 11
B good but jumpy 58%
A steady all-rounder 53%
F similar family 53%
G wins move around 50%
E different specialist 34%
C one fixed blind spot 33%
D specialist 27%
H rare bad run 3%
pass fail success curve · question solve rate

The white curve shows the share of models that passed each question: 100% at the top, 0% at the bottom. The vertical lines mark region boundaries.

Sorted this way, the evaluation has a visible shape: common wins on the left, disagreements in the middle, and common failures on the right.

At this scale, the map begins with the success ceiling, runs through the mixed interior and singleton frontier, and ends at the failure floor. That larger pattern was hiding behind eight averages: easy head, disputed middle, hard tail, and model-specific holes.

Here we made the data, so we know why its patterns exist. In a real evaluation, that recipe is hidden. Telltail shows the shape without pretending it knows the cause.

Open up the third dimension

How can we see all attempted runs at once?

The last picture showed the same R = 8 saved runs one at a time. Together, those saved runs are a third dimension of the evaluation. Now keep all R visible at once.

Combine the runs and each cell becomes a c/R bar: how many of the R saved runs passed. Keep the run index and the same cell becomes R stripes in the displayed order. Shuffle that order to see what changes—and what cannot. These runs are declared interchangeable draws, so “Run 1” is just a label: shuffling changes which selection you are looking at, never what was measured.

8 models × 64 questions × 8 runs

Put R inside each cell

The matrix never changes its data. Choose whether each cell combines the runs into one c/R bar or preserves their displayed order as R stripes.

Cell view
Sort model rows by
All-run success ceiling 1Intermittent success ceiling 9Mixed interior 37All-run singleton frontier 1Intermittent singleton frontier 8All-run failure floor 8
F 56%
B 51%
G 50%
A 47%
E 32%
C 30%
D 29%
H 24%
pass / green share fail / red share full run range middle half mean solve rate
Columns are grouped into the six repeat-aware zones over all R = 8 saved runs. The success ceiling and singleton frontier each split into an all-run variant (succeeded on every saved run) and an intermittent variant (reached at least once, but not every run). The failure floor is always all-run: any success anywhere already leaves it. The overlay is calculated once from all 8 runs, so toggling, shuffling, or re-sorting rows does not change it. Its curves are visually smoothed across nearby columns within each zone; the saved values and statistical summaries do not change.

Now we can start to see some "topology" emerging. This is what we mean by the "outcome geometry" of the evaluation.

With one run per cell, Telltail’s four regions told the whole story. With R runs, each question needs two answers: how many models ever reach it, and whether they hold it every run. The familiar names now need qualifiers. The success ceiling splits into an all-run success ceiling (every model, every run) and an intermittent success ceiling (every model at least once, but not every time). The singleton frontier splits the same way. The mixed interior stays one zone, and the failure floor is always all-run: a single success anywhere already leaves it. The columns below are grouped into all six zones, with five white dividers separating them. The per-run stripes show why each column earned its qualifier.

Model rows can be sorted three ways. Each is an observed share over all R saved runs and requires no independence assumption: average per run (the per-attempt rate), at least once (the share of questions the model ever solved), and every run (the share it solved on all R runs, using the same 1[c = R] fact behind the all-run zones). Those rankings need not agree. The group order and the average-per-run curve provably cannot both be monotone at once. That disagreement is real structure, not a rendering bug.

If you remember nothing else from scrolling this far, remember this:
The mean is not the map.

Why Telltail exists

The mean is useful. It is just not the whole evaluation.

A leaderboard average compresses a whole grid into one number. In the process it throws away which questions were shared wins, shared failures, disagreements, or lucky reruns. Telltail keeps that structure visible and keeps the question names behind every count.

What structure existed in the evaluation before we averaged it away?

Don’t settle for the average.