Part 5 left two questions the p-value cannot answer. How big is the effect, in terms someone could act on? And how precisely do we know it? Those are the questions that decide whether a result matters, and the numbers that answer them, the effect size and its confidence interval, should sit beside every p-value in a results table, as they do in the verification study's.

How big, in units that mean something

The gap in the twenty-four runs is 11.9 points on a 100-point rubric. That is the effect size, and the first thing to say about it is that it is already in the right units. Twelve points is roughly two rubric criteria's worth of functionality: an app that starts and whose main page loads, versus one that does not. A reader who knows the rubric can picture it. When a paper reports effects in the outcome's own units, points, seconds, percentage points, tokens, keep them there.

Two standardised effect sizes turn up when papers compare results across different outcomes or different studies, and they are worth recognising.

Cohen's d divides the gap by the typical spread of the runs. The two groups have standard deviations of 15.5 and 17.4, roughly 16.5 between them, so d = 11.9 / 16.5 = 0.72. The rule of thumb attached to it, 0.2 small, 0.5 medium, 0.8 large, is crude but useful: 0.72 says the tool's effect is a bit less than one typical run-to-run wobble, which matches the picture in Part 1 where the two rows of dots overlapped substantially.

Cliff's delta asks a question with no units at all: pick one run with the tool and one without at random, and how often does the first score higher? For the twenty-four runs the answer is 71 per cent of pairs, against 29 per cent the other way, and delta is the difference, 0.42. It is the natural effect size for the Mann-Whitney test of Part 4 and for anything where the raw numbers are not on a meaningful scale.

Effect In the outcome's units As Cohen's d As "how often does with beat without"
The twenty-four runs +11.9 points 0.72 71 per cent of pairs
P1, the shell on functional score +12.3 points not reported, and not needed
P2, the boot probe on survival +13 percentage points
P3, the shell on tokens ×2.2

The verification study reported every effect in natural units and left the standardised ones out, which I think is right when the units are meaningful. The standardised versions earn their place when they are not.

How sure: the confidence interval

An effect from a sample is an estimate, and Part 1's demo showed how far an estimate from six runs a side can wander from the truth. A confidence interval is the range of effects that the data cannot rule out, and a 95 per cent confidence interval is one built by a procedure with a specific promise: if you ran the same study many times and built the interval each time, about 95 of every 100 intervals would contain the true effect.

That is the promise, and it is subtly different from the sentence everyone wants to say, "there is a 95 per cent chance the true effect is in this interval". The true effect is a fixed number. It is either in this particular interval or it is not. What is 95 per cent is the procedure's hit rate over many studies, and this interval is one draw from that procedure. In practice the distinction rarely changes a decision, and the working reading is the safe one: the interval is the set of effects that are compatible with what we saw. Values inside it, the data cannot distinguish from the truth. Values outside it, the data argue against.

Two intervals from the verification study show the range of what "compatible" can mean.

Effect 95 per cent interval Width What it says
P2, boot probe on survival +13 points +10 to +16 6 Precise. The benefit is somewhere between a tenth and a sixth of all runs
P5, modification vs fresh build −0.7 points −13.7 to +12.7 26 The data are compatible with a 13-point benefit either way. This interval says almost nothing, and the paper says so

The width of an interval is the honest measure of what a study learned. A narrow interval around zero is a finding. A wide interval around zero is a shrug.

Building one by resampling: the bootstrap

The classical way to get an interval is a formula involving the standard deviation and the square root of the sample size. It works when the noise is bell-shaped and the statistic is a plain mean. For anything else, medians, ratios, differences within strata, the modern way is to let the data build the interval themselves, and the method is called the bootstrap.

The idea is a cousin of the permutation test. There, we asked what the world would look like if the labels did not matter, by shuffling labels. Here, we ask what other samples from the same world might have looked like, by resampling with replacement: draw twelve runs from the twelve "with" runs, putting each one back after drawing it so that some are picked twice and some not at all, and do the same for the twelve "without" runs. Compute the gap. Repeat thousands of times. The spread of those gaps is the spread of the estimate.

Here is one resample of the twelve "with" scores, sorted, with the original beside it.

Scores
Original 49, 55, 58, 62, 66, 71, 79, 84, 86, 88, 91, 95
One resample 49, 55, 55, 58, 58, 58, 62, 71, 79, 88, 88, 91

The resample drew 55 twice, 58 three times, 88 twice, and never drew 66, 84, 86 or 95. Its mean is 67.7 against the original 73.7. That is one imaginary study, drawn from the world our twelve runs describe. Do the same for the "without" side, take the gap, and you have one bootstrap gap. After 5,000 of them, sort the gaps and read off the values that 2.5 per cent fall below and 2.5 per cent fall above. Those two values are the percentile interval.

For the twenty-four runs, resampling each condition as one pool gives an interval from −1.0 to +24.6. It includes zero, which lines up with the pooled p-value of 0.093 from Part 4: the data pooled across models cannot rule out a gap of zero.

Now resample the way the experiment was built, within each of the four cells, six runs from each. That is a stratified bootstrap, the exact counterpart of the stratified shuffle, and it gives an interval from +6.3 to +17.7. Zero is out, which lines up with the stratified p-value of 0.0014. Same data, and once again the design decides how much it tells you.

The verification study built every interval this way: 5,000 resamples, drawn within each model-task-condition cell, with a fixed random seed of 20260703 written into the released code so that anyone can regenerate the exact numbers. P1's +12.26 came with an interval of +7.04 to +17.47. As a cross-check the study also fitted a completely different kind of model, a mixed-effects regression that gives each of the six models its own baseline, and the estimates agreed to two decimal places. When two unrelated methods agree, the agreement is worth more than either method's assumptions.

Try it: the resampler

Step draws one resample of each condition, shows which runs were picked (dimmed means not picked, a number means picked that many times), and drops the gap onto the histogram. Run draws 5,000. The interval is read from the histogram's tails. The checkbox resamples within each model.

Press Step five or six times and watch the chips: each resample leaves some runs grey and stamps others with ×2 or ×3. Then Run 5,000 with the box unticked and read the interval off the green lines, and again with it ticked. The stratified interval is about half as wide and does not touch zero.

Intervals and p-values are the same information, twice

If a 95 per cent interval for a gap excludes zero, the p-value for "the gap is zero" is below 0.05, and the reverse. They are two views of one calculation, and a paper that reports both is not padding: the interval shows the size and the precision, the p-value shows how the result stands against the convention. The twenty-four runs illustrate it twice over. Pooled: interval −1.0 to +24.6, p = 0.093, both say "cannot rule out zero". Stratified: +6.3 to +17.7, p = 0.0014, both say "zero is out".

The one place they appear to disagree is instructive. P4 in the verification study had an interval of +0.8 to +13.4, which excludes zero, and a corrected p-value of 0.0826, which is above the line. The interval was computed for P4 on its own. The p-value was corrected for P4 being one of six tests. Part 7 explains the correction, and the resolution is that an interval corrected the same way would have been wider and would have reached zero.

For completeness, the formula-based interval for the pooled twenty-four runs: the standard error of the gap is the square root of (15.5² / 12 + 17.4² / 12), which is 6.7, and the interval is the gap plus or minus 1.96 standard errors, −1.3 to +25.1. Nearly identical to the pooled bootstrap. When the formula's assumptions hold, the bootstrap reproduces it, and when they do not, only the bootstrap is left standing.

Which interval, when

Situation Use Why
A single pass rate, or the gap between two Wilson and Newcombe from Part 2 Closed-form, instant, and well-behaved at 0 and 100 per cent
A difference in means with bell-shaped noise and no strata The standard-error formula It is what the bootstrap will reproduce anyway
Medians, ratios, log-scale effects, or any design with strata The bootstrap, resampling within the cells No formula to trust, and the resampling follows the design
Several hypotheses tested together Whichever of the above, then be aware the intervals are uncorrected The p-values will be corrected and the intervals usually are not. Read the paper's note on which is which
Two conditions run on the same tasks The paired bootstrap of Part 11 Resampling tasks rather than runs removes the task difficulty from the noise

In Python

Effect sizes, both bootstraps, and the formula interval for comparison. The last call is SciPy's own bootstrap, which uses a slightly better construction of the interval called BCa (bias-corrected and accelerated) that you will see named in papers.

import numpy as np
from scipy import stats

A_with, A_without = np.array([88, 91, 79, 84, 95, 86]), np.array([78, 84, 70, 75, 89, 62])
B_with, B_without = np.array([62, 58, 71, 49, 66, 55]), np.array([46, 53, 40, 58, 49, 37])
with_tool, without_tool = np.concatenate([A_with, B_with]), np.concatenate([A_without, B_without])

# Effect sizes
gap = with_tool.mean() - without_tool.mean()
pooled_sd = np.sqrt((with_tool.var(ddof=1) + without_tool.var(ddof=1)) / 2)
wins = np.mean([a > b for a in with_tool for b in without_tool])
losses = np.mean([a < b for a in with_tool for b in without_tool])
print(f"gap {gap:+.1f} points, Cohen's d {gap / pooled_sd:.2f}, Cliff's delta {wins - losses:.2f}")

# Bootstrap interval, resampling each condition as one pool
rng = np.random.default_rng(20260703)
boot = [rng.choice(with_tool, 12).mean() - rng.choice(without_tool, 12).mean() for _ in range(5000)]
print(f"pooled bootstrap 95% interval: {np.percentile(boot, 2.5):+.1f} to {np.percentile(boot, 97.5):+.1f}")

# Stratified bootstrap: resample within each of the four cells
def one_resample():
    a = rng.choice(A_with, 6).mean() - rng.choice(A_without, 6).mean()
    b = rng.choice(B_with, 6).mean() - rng.choice(B_without, 6).mean()
    return (a + b) / 2
boot_s = [one_resample() for _ in range(5000)]
print(f"stratified bootstrap 95% interval: {np.percentile(boot_s, 2.5):+.1f} to {np.percentile(boot_s, 97.5):+.1f}")

# The formula-based interval, for comparison
se = np.sqrt(with_tool.var(ddof=1) / 12 + without_tool.var(ddof=1) / 12)
print(f"standard error {se:.1f}, formula interval: {gap - 1.96*se:+.1f} to {gap + 1.96*se:+.1f}")

# SciPy's bootstrap does the pooled version in one call, with a better (BCa) interval
res = stats.bootstrap((with_tool, without_tool), lambda a, b, axis: a.mean(axis) - b.mean(axis),
                      n_resamples=5000, random_state=20260703)
print(f"scipy BCa interval: {res.confidence_interval.low:+.1f} to {res.confidence_interval.high:+.1f}")

which prints:

gap +11.9 points, Cohen's d 0.72, Cliff's delta 0.42
pooled bootstrap 95% interval: -1.0 to +24.6
stratified bootstrap 95% interval: +6.2 to +17.7
standard error 6.7, formula interval: -1.3 to +25.1
scipy BCa interval: -1.3 to +24.3

Where this leaves us

The effect size says how much, in the outcome's own units when it has them. The confidence interval says how precisely, and its width is the honest record of what the study learned. The bootstrap builds the interval from the data's own shape by resampling the runs, within the cells the design was built from, and the verification study's intervals and p-values came from the same shuffling machinery, seeded so that anyone can reproduce them. The next part is about the cost of asking six questions at once, which is why P4's interval and P4's p-value seem to disagree.


Next: Six Hypotheses, One Bar: why six tests at 0.05 give chance a one-in-four shot at a false positive, Bonferroni, Holm's step-down worked on the study's own six p-values, what a "family" is, and why the secondary contrasts were reported uncorrected and called exploratory.