The first ten parts used two resampling tools for two jobs: shuffling labels to get a p-value (Part 4) and resampling runs to get an interval (Part 6). Papers use the same two words, permutation and bootstrap, in more ways than that, and this part collects the ones you will meet. The bootstrap as a test. The paired bootstrap, which is the right comparison for two conditions that ran on the same tasks. The bootstrap for statistics that have no formula. And the jackknife, the bootstrap's older and simpler relative, which answers a question of its own.

The bootstrap as a test

Part 6 built a confidence interval by resampling each condition's runs with replacement. If that interval excludes zero, the gap is significant at the matching level, so the bootstrap already gives you a test of a sort. There is also a direct version, and it is worth seeing because it makes the difference between a bootstrap test and a permutation test precise.

A permutation test asks: if the two conditions were the same in every way, how often would shuffling produce this gap? It builds the null by mixing the two groups into one pool.

A bootstrap test asks something slightly narrower: if the two conditions had the same mean but kept their own spreads, how often would resampling produce this gap? It builds the null differently. Slide each group so that its mean sits on the common mean of all twenty-four runs, leaving every score's distance from its own group's mean untouched. Now the null is true by construction, each group keeps its own shape, and resampling from the two shifted groups shows what gaps a world with equal means would produce.

Permutation test Bootstrap test
The null it builds the two groups are one population, so labels are exchangeable the two groups have equal means, and may differ in spread
How it builds it pool and shuffle shift each group to the common mean, then resample each with replacement
What it assumes exchangeability under the null that resampling each group mimics sampling from its population, which needs a dozen or so runs a side
On the twenty-four runs, pooled p = 0.093 p = 0.065

The two numbers differ because the groups differ in spread, 15.5 against 17.4, and the bootstrap null keeps that while the permutation null erases it. Neither is wrong. They answer different questions, and for most experiments the permutation test's question, "are these two conditions distinguishable at all", is the one being asked, which is why it is the default and why the verification study used it. The bootstrap test earns its place when the spreads clearly differ and the mean is the only thing you are claiming about, or when the statistic is something the permutation test cannot handle, which is the subject of the next section but one.

The paired bootstrap

Here is a design that comes up constantly in agent research and that the tools so far have handled badly. Seven tasks. Each task built once with the tool and once without. Fourteen scores, not twenty-four, and they come in pairs.

Task With the tool Without Difference
kanban 81 70 +11
calendar 74 61 +13
log explorer 66 57 +9
chat 79 68 +11
invoice 70 55 +15
gallery 62 52 +10
quiz 84 75 +9
Mean 73.7 62.6 +11.1

Look down the difference column: every task gained, by between 9 and 15 points. The tasks themselves vary far more, from the gallery at 52 without the tool to the quiz at 75. Treat the fourteen scores as two independent groups of seven, resample each, and the task-to-task variation floods the interval: −2.9 to +19.1, which includes zero and would make the tool's effect look uncertain. That is the wrong analysis, because the two groups are not independent. Each task's "with" score and "without" score share the task's difficulty, and the difficulty cancels in the difference.

The paired bootstrap resamples the unit that was actually independent, the task. Draw seven tasks with replacement, take the mean of their differences, repeat. The interval is +9.7 to +12.7: five times narrower, from exactly the same numbers, because the pairing removed a source of noise that was never about the tool. This is the resampling version of the paired t-test of Part 12, and the general rule behind both is the one from Part 4: resample or shuffle the unit that the design made independent. The verification study's stratification by model and task is the same rule at a larger scale.

405060708090 kanbancalendarlog explorerchatinvoicegalleryquiz withoutwith −50+10+20 unpaired paired
Seven tasks, each scored under both conditions. The rows sit at very different heights, but every line is about the same length. Pairing measures the lines and ignores the heights, which is why its interval for the mean gap is a fifth the width of the unpaired one.

Anything you can compute, you can bootstrap

The formula intervals of Part 6 exist for means, and with some effort for a few other things. The bootstrap does not care what the statistic is. Compute it on a resample, repeat, read the tails. Three that come up in agent papers:

A ratio of medians. The verification study reported that the shell cost 2.35 times the tokens of no verification, medians of 615K against 262K. A ratio of medians has no textbook interval. Resample each group, take the ratio of the two resampled medians, and 10,000 repeats give an interval directly. On twelve invented token counts a side that resemble the paper's, the multiplier is ×2.38 with an interval of ×2.16 to ×2.69.

A correlation. Does the number of screenshots a model took track how much it gained from them? Part 13 computes the correlation, and an interval for it comes from resampling the pairs (run, screenshots, gain) and recomputing.

A difference of differences. P5's statistic was the shell's gap on one task minus its gap on another. Four groups, resampled within each cell, recombined, repeated. The paper's interval for it, −13.7 to +12.7, was built exactly this way.

The one rule is to resample at the level of the design: pairs stay together, cells stay separate, and a statistic computed on a resample is computed the same way as on the original.

The jackknife

Before the bootstrap there was the jackknife, from the 1950s, and it is still the right tool for one question: which single observation is doing the work? Leave each run out in turn, recompute the statistic without it, and look at how far each omission moves the answer.

On the twelve token costs of Part 2, leaving out the runaway run moves the mean from 577 to 346 thousand. Leaving out any other run moves it by 36 at most. That is a numerical way of saying what the picture said: one run owns the mean. A paper that reports "results were robust to leaving any single run out" has done this, and it is a cheap, convincing check for a small study. The jackknife also gives a standard error, and for smooth statistics like means it agrees with the bootstrap's, but for anything with a step in it, medians and quantiles especially, it can be badly wrong, which is the main reason the bootstrap replaced it.

When resampling lets you down

The bootstrap is not magic, and three of its failures are worth knowing by name.

Too few runs. Resampling six runs with replacement produces only a few hundred distinct resamples, and an interval drawn from them is jagged and too narrow. Below about eight or ten a side, prefer a permutation test, which is exact at any size, and treat any interval as rough.

The extremes. The bootstrap cannot estimate the uncertainty of a maximum or a minimum, because a resample can never contain a value larger than the largest one you have. "The best run scored 95" has no bootstrap interval worth reporting.

Heavy tails. With a distribution like the token costs, where one run in twelve can be ten times the others, the resampled means depend on how many times the runaway run happens to be drawn, and the interval for the mean is unstable. This is the same lesson as Part 2: report the median, bootstrap the median, and the problem goes away.

The verification study stayed inside these limits: forty-eight or more runs per condition, means of bounded rubric scores, and medians for the tokens.

In Python

Each of the four ideas above is a few lines once you have the resampling loop from Part 6. The paired example uses the seven-task table, the ratio example invents twelve token counts a side that resemble the paper's, and the jackknife runs on Part 2's costs.

import numpy as np

with_tool    = np.array([88, 91, 79, 84, 95, 86, 62, 58, 71, 49, 66, 55])
without_tool = np.array([78, 84, 70, 75, 89, 62, 46, 53, 40, 58, 49, 37])
observed = with_tool.mean() - without_tool.mean()
rng = np.random.default_rng(20260703)

# 1. The bootstrap test: move both groups onto a common mean, so the null is true,
#    then resample and ask how often a gap as large as the real one appears
grand = np.concatenate([with_tool, without_tool]).mean()
null_with, null_without = with_tool - with_tool.mean() + grand, without_tool - without_tool.mean() + grand
gaps = np.array([rng.choice(null_with, 12).mean() - rng.choice(null_without, 12).mean() for _ in range(10_000)])
print(f"bootstrap test, pooled: p = {np.mean(np.abs(gaps) >= abs(observed)):.3f}   (permutation gave 0.093)")

# 2. The paired bootstrap: seven tasks, each run under both conditions. Resample tasks, not runs.
tasks = ["kanban", "calendar", "log explorer", "chat", "invoice", "gallery", "quiz"]
with_by_task    = np.array([81, 74, 66, 79, 70, 62, 84])
without_by_task = np.array([70, 61, 57, 68, 55, 52, 75])
diff = with_by_task - without_by_task
paired = np.array([rng.choice(diff, 7).mean() for _ in range(10_000)])
unpaired = np.array([rng.choice(with_by_task, 7).mean() - rng.choice(without_by_task, 7).mean() for _ in range(10_000)])
print(f"gap {diff.mean():+.1f}: paired interval {np.percentile(paired, 2.5):+.1f} to {np.percentile(paired, 97.5):+.1f}, "
      f"unpaired {np.percentile(unpaired, 2.5):+.1f} to {np.percentile(unpaired, 97.5):+.1f}")

# 3. Anything you can compute, you can bootstrap: the token multiplier as a ratio of medians
tok_with    = np.array([612, 540, 720, 655, 590, 810, 575, 640, 700, 530, 615, 690])   # thousand tokens
tok_without = np.array([250, 300, 262, 240, 275, 310, 255, 290, 230, 270, 265, 245])
ratio = np.median(tok_with) / np.median(tok_without)
boots = [np.median(rng.choice(tok_with, 12)) / np.median(rng.choice(tok_without, 12)) for _ in range(10_000)]
print(f"token multiplier x{ratio:.2f}, 95% interval x{np.percentile(boots, 2.5):.2f} to x{np.percentile(boots, 97.5):.2f}")

# 4. The jackknife: leave each run out in turn and see how far the estimate moves
costs = np.array([181, 204, 238, 257, 269, 291, 312, 348, 402, 517, 784, 3120])
loo = np.array([np.delete(costs, i).mean() for i in range(12)])
worst = np.argmax(np.abs(loo - costs.mean()))
print(f"mean cost {costs.mean():.0f}; leaving out run {worst + 1} (cost {costs[worst]}) moves it to {loo[worst]:.0f}; "
      f"no other run moves it by more than {np.delete(np.abs(loo - costs.mean()), worst).max():.0f}")

which prints:

bootstrap test, pooled: p = 0.065   (permutation gave 0.093)
gap +11.1: paired interval +9.7 to +12.7, unpaired +2.9 to +19.1
token multiplier x2.38, 95% interval x2.16 to x2.69
mean cost 577; leaving out run 12 (cost 3120) moves it to 346; no other run moves it by more than 36

Where this leaves us

Permutation and bootstrap are two ways of asking "what else could the data have looked like", and they build different imaginary worlds: one where the labels are interchangeable, one where each group keeps its shape but its mean is pinned. The bootstrap's real power is that it works for any statistic, a ratio, a correlation, a difference of differences, as long as the resampling follows the design, and the paired case is where following the design matters most: resample the tasks, not the runs, and an interval shrinks fivefold from the same numbers. The next part turns to the tests that predate all of this, which still fill most results sections, and shows which resampling test each one is a formula for.


Next: The Classical Toolbox: Student's t and Welch's t, the paired t-test, Mann-Whitney and Wilcoxon, chi-square against Fisher, ANOVA and Kruskal-Wallis, with a table that says which one to reach for and which resampling test it corresponds to.