Part 3 ended with a promise: that the null hypothesis, unlike the claim, says exactly what the data should look like, and that we could check it mechanically. This part keeps the promise. The method is the permutation test, it needs no formulas, and once you have done it by hand you will never again read "p = 0.03" as a mystery.

The whole idea fits in a sentence. If the tool makes no difference, then the labels "with" and "without" on the twenty-four runs are arbitrary, so any way of dealing the twenty-four scores into two hands of twelve is as legitimate as the real one. Deal them again and again, note the gap each time, and see how often a gap of 12 or more turns up by dealing alone.

The observed gap

Here are the twenty-four scores again, and the gap we are trying to explain.

Scores Mean
With the tool 88, 91, 79, 84, 95, 86, 62, 58, 71, 49, 66, 55 73.7
Without 78, 84, 70, 75, 89, 62, 46, 53, 40, 58, 49, 37 61.8
Gap +11.9

The number we compute from the data, here the difference between the two means, is called the test statistic. It could have been the difference between the medians, or the difference divided by the spread. The choice is free, provided it is made before looking, and the paper says which one it used. The verification study used the difference in means for scores and the difference in log tokens for cost.

Five steps

Step 1. Pool. Forget the labels. There are now just twenty-four numbers in a bag: 88, 91, 79, and so on down to 37.

Step 2. Shuffle. Deal twelve of them at random into a hand called "with" and the other twelve into "without".

Step 3. Compute. Take the gap between the two hands' means, exactly as before.

Step 4. Repeat. Do steps 2 and 3 thousands of times, and keep every gap.

Step 5. Count. How many of those shuffled gaps were at least as far from zero as the real gap of 11.9? That count, divided by the number of shuffles, is the p-value.

Here are the first five shuffles I did, with a fixed random seed so that they are reproducible.

Shuffle Mean of the "with" hand Mean of the "without" hand Gap
1 70.5 64.9 +5.6
2 68.6 66.8 +1.8
3 68.0 67.4 +0.6
4 68.2 67.3 +0.9
5 64.0 71.4 −7.4

Nothing about the tool went into those five gaps. They are pure dealing. And already they show the shape of the answer: shuffled gaps wander several points either side of zero, and none of the first five reached 11.9. Keep going to ten thousand and the gaps pile up into a hill centred on zero, the null distribution, which is the picture of what "the tool does nothing" looks like in this experiment.

−25−15−50+5+15+25 gap between the two shuffled hands, in points (10,000 shuffles) observed +11.9 −11.9 9% of shufflesland out here
The null distribution for the twenty-four runs, pooled. The real gap sits inside the hill's shoulder, not out in its tail. About one shuffle in eleven produces a gap this big or bigger without any tool at all.

The count

For twelve against twelve the shuffling can actually be done exhaustively. There are 2,704,156 ways to choose which twelve of the twenty-four scores go in the "with" hand, and a laptop can compute the gap for every one of them in under a minute. Of those, 252,182 give a gap at least as far from zero as 11.9. That is 0.093, and it is the exact two-sided p-value for this data.

Ten thousand random shuffles give 0.097, close enough, and for larger studies random shuffling is the only option: the verification study used 10,000 permutations per hypothesis, because with 192 runs a side the number of possible deals has more digits than this page.

So: if the tool did nothing, a gap of 12 would show up about once in every eleven attempts. That is not rare. By the convention explained in the next part, a result is called surprising when it would show up less than once in twenty, and this one would not. Pooled, the twenty-four runs do not support the claim that the tool helps.

Which is odd, because look at the table at the top. Within each model, every single run with the tool scored higher than the average run without it. Something is being thrown away.

Shuffle within the model

What is being thrown away is the design. The strong model scores about 30 points above the weak one, regardless of the tool. When the pooled shuffle deals cards, it sometimes puts five strong-model runs in one hand and one in the other, and that alone produces gaps of 15 or 20 that have nothing to do with the tool. The null distribution is wide because it is full of model effects, and the tool's 12 points get lost in them.

The fix is to shuffle in a way that respects how the experiment was actually run. The tool was switched on or off within each model. So shuffle the labels within Model A's twelve runs, and separately within Model B's twelve, and never let a strong-model score wander into the weak model's hand. This is a stratified permutation test, and each model is a stratum, a layer that the shuffle is not allowed to cross.

Pooled shuffle Shuffled within each model
What is swapped any of the 24 scores for any other with and without labels inside Model A, and inside Model B
Possible deals 2,704,156 853,776
Deals with a gap ≥ 11.9 252,182 1,160
p-value 0.093 0.0014
Verdict at the 0.05 convention not surprising surprising: about one deal in 740

The observed gap did not change. The question got sharper. "Could dealing alone produce a gap of 12?" became "could dealing alone, among runs of the same model, produce a gap of 12?", and the answer to the second is almost never.

This is what the verification study means by "stratified by model and task". Its six models differ far more than they agree, and its seven tasks differ too, so every shuffle was done inside a model-task cell, and the label swapping never crossed a cell boundary. Without that, the study would have been asking whether the shell helps on average across a random mix of models, which is a question with far more noise in it than the one it wanted to answer.

Try it: the shuffler

Below are the twenty-four runs as chips, coloured by model. Step deals them once, shows the new hands and the gap, and drops one bar onto the histogram. Run deals a thousand times. The checkbox switches to shuffling within each model. Watch the hill narrow.

Press Step a dozen times and watch the chips: a pooled deal will often put four or five red chips in one hand, and the gap follows the red. Then Run 1,000 twice, once with the box unticked and once ticked, and compare the width of the two hills. The real gap never moves. What moves is how surprising it is.

What the test assumes, and what it does not

The permutation test asks one thing of the data: that under the null hypothesis the labels are exchangeable, meaning any run could have carried either label. That is true whenever the condition was assigned to runs independently of how they would score, which is what a controlled experiment arranges. It is not true if, say, the runs with the tool were all done on a Tuesday when the model's provider had a fast day, which is a confound and the subject of Part 9. Stratifying handles the confounds you know about and recorded, like model and task. It cannot handle the ones you did not.

It also assumes the runs are separate draws. Two runs that share a random seed and produce nearly the same output are one run counted twice, and shuffling them as if they were independent overstates the evidence.

What the test does not assume is anything about the shape of the noise. The classical alternative, the t-test (Part 12 has the whole family), assumes the scores are spread in a bell curve and computes the p-value from a formula. When the bell holds, the two tests agree closely. When it does not, as with the token costs of Part 2, the permutation test still gives the right answer because it never assumed a shape: it used the data's own shape, by shuffling the data itself. That is why it is the default for small, oddly distributed samples, and why the verification study used it for all six hypotheses. Two other tests you will meet are relatives: the Mann-Whitney test is a permutation test on ranks rather than raw scores, and Fisher's exact test, which the next part explains, is a permutation test on counts.

In Python

The pooled shuffle, the stratified shuffle, and the exact count over all 2.7 million deals. The exact loop takes about twenty seconds. Change the seed and the first two numbers move in the third decimal place, which is the resolution of 10,000 shuffles.

import numpy as np
from itertools import combinations

A_with, A_without = [88, 91, 79, 84, 95, 86], [78, 84, 70, 75, 89, 62]
B_with, B_without = [62, 58, 71, 49, 66, 55], [46, 53, 40, 58, 49, 37]
with_tool, without_tool = np.array(A_with + B_with), np.array(A_without + B_without)
observed = with_tool.mean() - without_tool.mean()

# 1. Pooled permutation test, 10,000 random shuffles
rng = np.random.default_rng(20260703)
pool = np.concatenate([with_tool, without_tool])
shuffled_gaps = np.empty(10_000)
for i in range(10_000):
    rng.shuffle(pool)
    shuffled_gaps[i] = pool[:12].mean() - pool[12:].mean()
p_pooled = np.mean(np.abs(shuffled_gaps) >= abs(observed))     # two-sided
print(f"observed gap {observed:+.2f}, pooled p = {p_pooled:.3f}")

# 2. Stratified: shuffle within each model, never across
def strat_gap(rng):
    a = rng.permutation(A_with + A_without); b = rng.permutation(B_with + B_without)
    return ((a[:6].mean() - a[6:].mean()) + (b[:6].mean() - b[6:].mean())) / 2
strat = np.array([strat_gap(rng) for _ in range(10_000)])
print(f"stratified p = {np.mean(np.abs(strat) >= abs(observed)):.4f}")

# 3. Exact answer for the pooled test: every one of the 2,704,156 possible deals
count = total = 0
for idx in combinations(range(24), 12):
    hand = pool[list(idx)]
    count += abs(hand.mean() - (pool.sum() - hand.sum()) / 12) >= abs(observed) - 1e-9
    total += 1
print(f"exact pooled p = {count} / {total} = {count / total:.4f}")

which prints:

observed gap +11.92, pooled p = 0.095
stratified p = 0.0010
exact pooled p = 252182 / 2704156 = 0.0933

Where this leaves us

The null hypothesis says the labels do not matter, so the permutation test swaps them and watches. The fraction of swaps that beat the real gap is the p-value, and for the twenty-four runs it is 0.093 pooled and 0.0014 within model, which is the difference between "not surprising" and "surprising", from the same numbers, depending on whether the shuffle respects the design. That p-value is now a count you have watched being made. The next part is about what it does and does not mean, and about the number 0.05.


Next: What 0.0006 Means, and What 0.08 Does Not: the p-value defined and misdefined, where the 0.05 convention came from, why "significant" is not "important", why "not significant" is not "no effect", Fisher's exact test for counts, and the six p-values from the verification study read one at a time.