Every part so far has been about detecting an effect. This one is about the harder job of showing there is none, which comes up more often than you would think. Does the testing tool improve the score? The first study found that it did not, and that it cost 42 to 68 per cent more tokens. Does a shell help a modification task less than a fresh build? The verification study's P5 expected "about the same". A study I am planning now has, as its most likely honest headline, that a feature makes no difference. Each of those is a claim of absence, and the machinery of Parts 4 to 7 cannot make it.
The asymmetry
Recall the logic. We assume the null, the tool does nothing, and ask how surprising the data would be. If very surprising, we reject the null. If not, we keep it. Keeping it is not confirming it, for the reason Part 3 gave: not guilty is not innocent. The test was built to detect, and a detector that did not go off is a detector that did not go off.
P5 is the cleanest example in the study. The shell's benefit on the modification task minus its benefit on the fresh build came to −0.7 points, with a 95 per cent interval from −13.7 to +12.7 and a corrected p-value of 0.9989. The gap is as close to zero as a gap gets. And the paper, correctly, did not report "the shell helps modification and fresh builds equally". It reported that the hypothesis was not supported, and explained why in one line: support would have needed "a confidence interval close to zero, not merely a non-significant permutation p-value". The interval was 26 points wide. A 13-point benefit either way was still on the table. A shrug at p = 0.9989 is still a shrug.
Absence of evidence, evidence of absence
The phrase is a cliché because it is exactly right. "We found no significant effect" is absence of evidence: the detector did not go off, and the interval tells you how insensitive the detector was. "The effect, if any, is smaller than five points" is evidence of absence: a positive claim, with a number in it, that a study can support or fail to support like any other.
The difference between the two intervals in the figure is not their centre. Both sit on −0.7. It is their width, and width is bought with runs. The lower interval is what P5 would have looked like with roughly four times the data, and it supports the stronger claim because it fits inside a band. Which raises the question of where the band came from.
The smallest effect worth caring about
To say "no meaningful difference" you have to say what "meaningful" means, in the outcome's units, before you see the data. This number is the smallest effect size of interest, and choosing it is a judgement about the world rather than a statistical calculation. On a 100-point functional rubric, five points is less than one criterion, less than the disagreement between two human graders, and less than the run-to-run noise of a single configuration. A tool whose effect is inside ±5 is, for every practical purpose, a tool that does nothing. So five is a defensible bound, and the point is not that it is right but that it is written down first.
Write it down first because otherwise the bound will be chosen to fit the interval. If the interval had come out at −8 to +7, a bound of 10 would make the null look defended, and a reader has no way of knowing whether 10 was the plan or the rescue. This is the single strongest argument for the preregistration of Part 10, and it is why my planned null-result study will deposit its bounds with a third party before the first run.
The equivalence test
With a bound in hand, the claim "the effect is within ±5 points" becomes testable, and the test is called an equivalence test. The standard form is two one-sided tests, usually written TOST, and the name describes it exactly.
Test one: is the effect significantly greater than −5? Test two: is the effect significantly less than +5? If both pass at the 5 per cent level, the effect is inside the band, and the null has been defended: not "we could not find a difference" but "the difference, if any, is smaller than five points".
There is a picture that makes the procedure obvious. Two one-sided tests at 5 per cent each are the same thing as checking whether a 90 per cent confidence interval lies entirely inside the band from −5 to +5. Ninety, not ninety-five, because each one-sided test spends its 5 per cent on one end. So the recipe is: build the 90 per cent interval by the bootstrap of Part 6, draw the band, and look.
The four things that can happen, on one axis:
| The 90 per cent interval | Inside the band? | Includes zero? | Conclusion |
|---|---|---|---|
| −3.5 to +2.1 | yes | yes | Equivalent. The null is defended |
| +1.0 to +4.0 | yes | no | A real effect, but too small to matter. Also a defended null, in practice |
| +6.3 to +17.7 | no | no | A real effect that matters. The twenty-four runs, stratified |
| −13.7 to +12.7 | no | yes | Inconclusive. The study was too small to say either thing. P5 |
The last row is the one that costs the most, because it means the runs were spent and nothing was learned, and it is the row that a well-designed null study is built to avoid. (Part 15 gives the same band a probability instead of a verdict, which handles the middle cases more gracefully.)
Power: the chance of seeing what is there
The reason P5 landed in the last row was not bad luck. It was arithmetic that could have been done before the study, and the concept is statistical power: the probability that a study will detect an effect of a given size, if that effect is really there.
Power depends on three things. The size of the effect you are looking for, the noise between runs, and the number of runs. Bigger effects, less noise and more runs all raise it. The convention is to aim for 80 per cent, meaning that if the effect exists, four studies in five will find it and one in five will miss it, a rate of misses that the field has decided to live with.
Here is what it looks like on the scale of the twenty-four runs, with about 15 points of noise between runs, for the ordinary two-sided test at the 0.05 line.
| Runs per condition | Smallest gap detectable with 80 per cent power |
|---|---|
| 6 | about 27 points |
| 12 | about 18 points |
| 24 | about 12 points |
| 48 | about 9 points |
| 100 | about 6 points |
| 192 | about 4 points |
Six runs a side, the size of one cell in the twenty-four, can reliably detect only a gap of 27 points, more than twice the tool's real effect. The verification study's 192 runs a side in its core conditions could see a 4-point gap, which is why P1's 12 points came through with room to spare. P6, with about 28 runs a side, could reliably see 11 points and was looking at 15, which is a coin toss made slightly generous, and it came out just on the wrong side of the line.
The figures come from the standard power calculation for a two-sample test, and a rule of thumb gets close for anything but the smallest samples: the smallest detectable gap is about 2.8 times the noise divided by the square root of half the runs per condition. Any statistics package will give it, and a power analysis in a methods section is just this calculation run in reverse: decide the gap you need to see, and it tells you how many runs to buy. For a 12-point gap with 15 points of noise the answer is 26 runs a side.
How many runs to defend a null
For an equivalence test the question is how many runs make the 90 per cent interval narrow enough to fit in the band. With 15 points of noise and a band of ±5, the interval's half-width is about 1.645 × 15 × the square root of 2 divided by the runs per condition, and setting that to 5 gives about 50 runs a side. That is the number at which the interval on average just fits, and it is a trap, because the interval is centred on a measured gap that itself wanders a few points either side of zero. Simulate it and only one null study in twenty-five comes out equivalent at 50 runs a side, and one in four at 70. For the usual 80 per cent power the arithmetic says about 155 runs a side, three times the naive number, and that is the figure a power analysis for a null-result study has to produce before the runs are bought.
Now P5 again. The modification task had 144 runs, which sounds like plenty, but P5 was a difference of differences: the shell's gap on the modification task minus the shell's gap on the fresh build. That is four groups and two subtractions, and every subtraction adds noise: the interval for a difference of two gaps is about twice as wide as the interval for one gap on the same number of runs. Twice the width means four times the runs to get back to the same precision, and the study did not have them. A power analysis done in advance would have said so, and the honest report of an inconclusive result is what the paper printed instead. The planned study does the analysis first.
Try it: bounding a null
Set the band, the noise, the true gap and the runs per condition. Draw a study simulates one experiment, builds its 90 per cent interval, and gives the verdict. Draw 20 shows how often a study of this size reaches each verdict, which is its power.
Two things to try. With the true gap at 0, 15 points of noise and 12 runs a side, press Draw 20 and read the tally: nearly every study is inconclusive, which is P5's situation. Slide the runs up to 70 and draw 20 again: only about a quarter come out equivalent, even though the interval "fits on average" at 50. Slide to 160 and most do. Then set the true gap to 12 and runs to 6, and count how many studies find it. That fraction is the power, and it will be poor.
In Python
The equivalence test is the interval rule, four lines long. The simulation is the honest way to find a sample size: run the study you are planning a few thousand times on a true null and count how often it comes out equivalent. The power calculation gives the other numbers in this part.
import numpy as np
from scipy import stats
from statsmodels.stats.power import TTestIndPower
# Equivalence test (TOST) by the interval rule: does the 90% interval fit inside the band?
def tost(a, b, bound):
gap = a.mean() - b.mean()
se = np.sqrt(a.var(ddof=1) / len(a) + b.var(ddof=1) / len(b))
lo, hi = gap + stats.t.ppf([0.05, 0.95], len(a) + len(b) - 2) * se # 90% interval = two one-sided tests
return ("equivalent" if -bound <= lo and hi <= bound else
"inconclusive" if lo <= 0 <= hi else "different")
# Simulate a true null (no effect, 15 points of noise) many times at two sample sizes
rng = np.random.default_rng(20260703)
for n in (12, 70):
verdicts = [tost(rng.normal(62, 15, n), rng.normal(62, 15, n), bound=5) for _ in range(2000)]
eq = verdicts.count("equivalent") / 2000; inc = verdicts.count("inconclusive") / 2000
print(f"{n:>2} runs a side, band ±5: equivalent {eq:.0%}, inconclusive {inc:.0%}, wrong {1 - eq - inc:.0%}")
# Power: how many runs to detect a gap of 12 points with 15 points of noise, 80% of the time
power = TTestIndPower()
n = power.solve_power(effect_size=12 / 15, alpha=0.05, power=0.8)
print(f"runs per condition for 80% power on a 12-point gap: {n:.0f}")
# The other way round: the smallest gap a given number of runs can reliably detect
for n in (6, 12, 24, 48, 100, 192):
d = power.solve_power(nobs1=n, alpha=0.05, power=0.8)
print(f"{n:>3} runs a side detect a gap of about {15 * d:.0f} points")
which prints:
12 runs a side, band ±5: equivalent 0%, inconclusive 90%, wrong 10%
70 runs a side, band ±5: equivalent 25%, inconclusive 65%, wrong 9%
runs per condition for 80% power on a 12-point gap: 26
6 runs a side detect a gap of about 27 points
12 runs a side detect a gap of about 18 points
24 runs a side detect a gap of about 12 points
48 runs a side detect a gap of about 9 points
100 runs a side detect a gap of about 6 points
192 runs a side detect a gap of about 4 points
Where this leaves us
The ordinary test detects. It cannot certify absence, because a detector that stayed silent may have been too weak to hear. To claim "no meaningful effect" you must say in advance what meaningful means, in the outcome's units, and then show that the whole 90 per cent interval fits inside that band: two one-sided tests, or one picture. How many runs that takes is a calculation, power analysis, that can and should be done before a run is bought, and the verification study's one inconclusive result is what happens when it is not. The last two parts leave the arithmetic and turn to the two things that decide whether any of it is believed: how the study was built, and when its choices were made.
Next: Designed to Be Believed: assigned versus randomised conditions, confounds and the one the first study admitted to, blinding the grader, freezing the rubric, the intention-to-treat rule borrowed from medicine, seeds and reproducibility, and how to read a kappa of 0.973.