Part 1 ended with twenty-four runs and a gap of 12 between two averages. Before we can ask whether that gap is real, we need to be careful about the word "average", because there are several, they disagree, and the choice between them is the first place a write-up can mislead without anyone lying. This part is the arithmetic of summarising many numbers into one, and then into two, and it ends with the special case that trips up more papers than any other: a rate, like "16 of 18 runs passed", and the interval that belongs beside it.
Two middles
Here are the twelve scores from the runs with the tool, sorted: 49, 55, 58, 62, 66, 71, 79, 84, 86, 88, 91, 95.
The mean is what most people call the average: add them up and divide by how many. The sum is 884, there are 12, so the mean is 73.7.
The median is the middle value once they are sorted. With twelve numbers there is no single middle, so it is halfway between the sixth and seventh: 71 and 79, giving 75. Half the runs scored below 75 and half above.
Here they nearly agree, 73.7 and 75, and when they agree it does not much matter which you report. The interesting case is when they do not.
The runaway run
Now the costs. Here are twelve token counts, in thousands, for twelve runs of an agent building the same app, again invented to look like the real distributions: 181, 204, 238, 257, 269, 291, 312, 348, 402, 517, 784, 3,120.
Eleven of them sit between 181 and 784. The twelfth is a run that got stuck in a loop, retried the same failing test forty times, and burned 3.1 million tokens before it gave up. This happens. It happened in the first study more than once.
| Summary | Value (thousands of tokens) | What it is telling you |
|---|---|---|
| Mean | 577 | The total spend divided by twelve. Dragged up by the one runaway run, which alone is 45 per cent of the total |
| Median | 302 | The middle run. Nine of the twelve cost less than the mean, which is a sign the mean is not "typical" of anything |
| Trimmed mean (drop the lowest and highest, average the rest) | 362 | The mean with the extremes removed. Closer to the median, but the amount of trimming is a choice you have to declare |
| Geometric mean (average of the logarithms, then undo the log) | 382 | The mean on a multiplicative scale. A run that costs twice as much moves it as far as one that costs half as much |
The mean says "a run costs about 577K tokens". The median says "a typical run costs about 300K, and occasionally one goes wild". Both are true. Only the second is what you want to know when budgeting, and only the first is what you want to know when paying the bill at the end of the month, because the bill is the total and the total is twelve times the mean.
This is why the first study summarised cost with the median and reported ranges rather than means, and said so in one sentence. It is also why the second study compared token counts on a logarithmic scale, where a run costing twice as much sits one fixed step to the right regardless of whether it went from 100K to 200K or from 1M to 2M. On that scale the average gap between the two conditions came out at 0.80 in natural-log units, which is a multiplier of about 2.2, and the medians told the same story in plain numbers: 615K tokens with the verification tool against 262K without, a factor of 2.35. When you see a paper report "log tokens", that is all it means: it is comparing multipliers rather than differences, because for costs that is the honest comparison.
Try it: the middle-finder
Drag any of the twelve costs along the line and watch the three middles. Drag the runaway run back into the pack, or drag an ordinary run out to the far right, and see which summary cares.
Try dragging the runaway run all the way to 3,500 and then all the way down to 200. The mean travels across a quarter of the axis. The median barely notices, because the median only cares that the run is on one side of the middle, not how far. That insensitivity is called robustness, and it is the reason the median is the default for anything with a long tail: costs, latencies, file sizes, and the time an agent takes to finish.
Two other means worth knowing
A weighted mean counts some items more than others. A rubric is a weighted mean in disguise: the first study's rubric had 14 criteria worth up to 42 points, so a criterion worth 5 points has five times the pull of one worth 1. When a paper says "weighted score", it means each item was multiplied by its weight before being added. The trap is that the weights are a judgement, and a different judge would set them differently, which is why the second study wrote its weights down and froze them before any run existed.
The geometric mean is the right average for things that multiply rather than add: growth rates, speedups, price ratios, and token multipliers. If one change makes a run twice as expensive and another makes it half as expensive, the arithmetic mean of 2 and 0.5 is 1.25, which claims a net cost increase that does not exist. The geometric mean is the square root of 2 × 0.5, which is 1, the right answer. Averaging the logarithms and then undoing the log is the same calculation, which is why "log tokens" and "geometric mean" turn up together.
How spread out
Two sets of runs can share a mean and be nothing alike. The twelve scores with the tool have a mean of 73.7 and range from 49 to 95. A different tool might give twelve runs that all land between 70 and 78 with the same mean. The second is far more useful, and a summary that reports only the mean cannot tell them apart. So beside every middle goes a spread.
The range is the lowest to the highest, 49 to 95. Easy to read, but it is set entirely by the two most extreme runs, which are the two least representative.
The standard deviation (SD) is the typical distance of a run from the mean. For the twelve scores it is 15.5, which you can read as "a run usually lands within about 15 points of the average, sometimes 30". For those who want the recipe: take each score's distance from the mean, square it, average the squares (dividing by one less than the count, for a reason that does not matter here), and take the square root. The squaring means the standard deviation, like the mean, is pulled hard by outliers: the twelve token costs have a standard deviation of 818 thousand, larger than eleven of the twelve values, which is the runaway run again.
The interquartile range (IQR) is the median's companion. Sort the values, find the value a quarter of the way up and the value three quarters of the way up, and the IQR is the stretch between them, the middle half of the data. For the token costs it runs from 252 to 431 thousand. The runaway run does not touch it.
A box plot draws all of that in one shape, and it is the picture to look for in any results section: the box is the interquartile range, the line inside it is the median, the whiskers stretch to the rest of the data, and any run too far out is drawn as a lone dot, the outlier. The first study's cost figures are box plots on a logarithmic axis, which is the standard way to show something like token counts.
A rate is an average in disguise
The first study's headline was a rate: at the higher reasoning setting, 16 of 18 runs were perfect on the first try, against 5 of 18 at the lower setting. Rates feel like a different kind of number from scores, but they are not. Write a 1 for each perfect run and a 0 for each imperfect one, and the rate is the mean of those ones and zeros: 16 ones and 2 zeros average to 0.889, which is 88.9 per cent. Everything that applies to means applies to rates, including the fact that a rate from 18 runs is a noisy estimate of the rate you would see from 18,000.
So a rate needs an interval beside it, a range of true rates the data are compatible with, and this is where a textbook formula goes wrong in a way that is worth seeing once, because you will meet it in published work.
Four intervals for "16 out of 18"
Call the observed rate p (here 0.889) and the number of runs n (18). Each method below produces a range that is meant to contain the true rate 95 times out of 100 if you used it on many studies. Part 6 says exactly what that promise means. For now, look at the ranges.
| Method | 5 of 18 (27.8%) | 16 of 18 (88.9%) | 0 of 18 (0%) | 18 of 18 (100%) |
|---|---|---|---|---|
| Wald (the textbook formula) | 7.1% to 48.5% | 74.4% to 103.4% | 0% to 0% | 100% to 100% |
| Wilson score | 12.5% to 50.9% | 67.2% to 96.9% | 0% to 17.6% | 82.4% to 100% |
| Clopper-Pearson (exact) | 9.7% to 53.5% | 65.3% to 98.6% | 0% to 18.5% | 81.5% to 100% |
Wald. The interval is p plus or minus 1.96 times the square root of p(1 − p)/n. It is the one in most introductory courses and most spreadsheets, and it fails exactly when it matters. For 16 of 18 it says the true pass rate could be as high as 103 per cent. For 0 of 18 it says the true rate is exactly zero, with no uncertainty at all, after eighteen runs. The formula assumes the rate is far from 0 and 100 and the sample is large, and small agent studies are usually neither.
Wilson score. Instead of centring the interval on the observed rate, Wilson asks a slightly different question: which true rates would make an observation like 16 of 18 unsurprising? The answer is an interval that is pulled toward 50 per cent, never crosses 0 or 100, and behaves well even for 0 of 18, where it says the true rate is somewhere below 17.6 per cent, which is the honest statement. Its centre is (p + z²/2n) divided by (1 + z²/n), with z = 1.96, and its half-width is z times the square root of p(1 − p)/n + z²/4n², divided by the same (1 + z²/n). Nobody computes it by hand, but the shape of the formula shows what it does: it adds about two imaginary successes and two imaginary failures to the count, which is why it never runs off the end.
Clopper-Pearson. The exact method. Find the lowest true rate at which seeing 16 or more of 18 would still happen at least 2.5 per cent of the time, and the highest true rate at which seeing 16 or fewer would. That pair is the interval. It never lies, but it is a little wider than it needs to be, which is a known trade: it guarantees at least 95 per cent coverage and usually delivers more. Papers that say "exact binomial interval" mean this.
For most work, Wilson is the sensible default, and the second study's survival rate, +13 percentage points with an interval of 10 to 16, is the kind of number these intervals produce. The Wald interval is worth recognising only so that you can distrust a paper that reports a rate of 103 per cent, or an interval of exactly zero width.
Newcombe's method: the gap between two rates
The first study's claim was not about one rate but about the gap between two: 88.9 per cent minus 27.8 per cent, a 61-point difference. An interval for a difference of two rates is a different calculation, and the well-behaved version is due to Robert Newcombe, from a 1998 paper that compared eleven methods and recommended this one.
The idea is simple once you have Wilson. Compute the Wilson interval for each rate on its own. Then, for the lower end of the difference, take how far the first rate could plausibly be below its estimate and how far the second could be above its estimate, and combine those two distances the way you combine the sides of a right-angled triangle: square each, add, take the root. Do the same with the other two distances for the upper end.
For 16 of 18 against 5 of 18 that gives a difference of 61 points with an interval from 29 to 78 points. Read that as: the data are consistent with the higher reasoning setting adding anywhere from about 30 to about 80 percentage points of first-try success, and not consistent with it adding nothing. Compare the twenty-four runs of this series, where 7 of 12 with the tool and 5 of 12 without scored 70 or more: a 17-point gap with a Newcombe interval from −21 to +48, which is to say that twelve runs a side cannot tell you whether the tool changes the pass rate at all. Same arithmetic, and the honest answer is different, because the numbers are smaller and the gap is narrower.
Try it: intervals for a rate
Set two pass rates with the sliders and see the three intervals for each, plus Newcombe's interval for the gap between them. Push a rate to the edge and watch Wald misbehave.
Two things to try. Set condition 1 to 18 of 18 and watch Wald announce a pass rate of exactly 100 per cent with no doubt at all, while Wilson says "somewhere above 82 per cent", which is what eighteen clean runs actually justify. Then keep the rates where they are and push both run counts to 200: the intervals shrink to a few points wide and the three methods converge, because with enough runs the textbook formula stops mattering.
In Python
SciPy has the middles, the spreads and the honest rate intervals built in, and it deliberately leaves out the Wald interval. Newcombe's method is a few lines on top of Wilson.
import numpy as np
from scipy import stats
scores = np.array([49, 55, 58, 62, 66, 71, 79, 84, 86, 88, 91, 95]) # with the tool, sorted
costs = np.array([181, 204, 238, 257, 269, 291, 312, 348, 402, 517, 784, 3120]) # thousand tokens
print(f"scores: mean {scores.mean():.1f}, median {np.median(scores):.1f}, "
f"sd {scores.std(ddof=1):.1f}, IQR {np.percentile(scores, 25):.1f} to {np.percentile(scores, 75):.1f}")
print(f"costs: mean {costs.mean():.0f}, median {np.median(costs):.0f}, "
f"trimmed {stats.trim_mean(costs, 1/12):.0f}, geometric {stats.gmean(costs):.0f}")
# Intervals for a rate: 16 of 18 perfect runs
k, n = 16, 18
for method in ("wilson", "exact"): # scipy has no Wald, because nobody should use it
ci = stats.binomtest(k, n).proportion_ci(confidence_level=0.95, method=method)
print(f"{k}/{n} {method:>6}: {100*ci.low:.1f}% to {100*ci.high:.1f}%")
p = k / n
print(f"{k}/{n} wald: {100*(p - 1.96*np.sqrt(p*(1-p)/n)):.1f}% to {100*(p + 1.96*np.sqrt(p*(1-p)/n)):.1f}%")
# Newcombe's interval for the gap between two rates, built from two Wilson intervals
def wilson(k, n, z=1.959964):
p, d = k / n, 1 + z*z/n
c, h = (p + z*z/(2*n)) / d, z*np.sqrt(p*(1-p)/n + z*z/(4*n*n)) / d
return c - h, c + h
def newcombe(k1, n1, k2, n2):
p1, p2 = k1/n1, k2/n2
l1, u1 = wilson(k1, n1); l2, u2 = wilson(k2, n2)
d = p1 - p2
return d, d - np.sqrt((p1-l1)**2 + (u2-p2)**2), d + np.sqrt((u1-p1)**2 + (p2-l2)**2)
d, lo, hi = newcombe(16, 18, 5, 18)
print(f"gap 16/18 vs 5/18: {100*d:.1f} points, {100*lo:.1f} to {100*hi:.1f}")
which prints:
scores: mean 73.7, median 75.0, sd 15.5, IQR 61.0 to 86.5
costs: mean 577, median 302, trimmed 362, geometric 382
16/18 wilson: 67.2% to 96.9%
16/18 exact: 65.3% to 98.6%
16/18 wald: 74.4% to 103.4%
gap 16/18 vs 5/18: 61.1 points, 29.4 to 78.4
Where this leaves us
There are several averages, and the choice between them is a claim. The mean is the total shared out equally, the median is the typical case, and when they disagree the data have a tail, which is the moment to report both and say why. Every middle needs a spread beside it, and the box plot is the picture that shows both at once. A rate is a mean of ones and zeros, so it needs an interval too, and the honest intervals for small counts are Wilson for one rate and Newcombe for the gap between two. What none of this has answered is the question Part 1 posed: whether a gap of 12 points, or 61, is more than the noise could produce. To ask that properly we first have to state exactly what we are claiming, which is the next part.
Next: The Claim and Its Shadow: what a hypothesis is, why every one comes with a null hypothesis attached, one-sided and two-sided claims, and the six hypotheses from the verification study written out as bets.