The permutation test and the bootstrap need a computer. For most of the twentieth century there was no computer, and the tests that were invented instead work by assuming a shape for the noise, usually the bell curve, and looking up the answer in a table. Those tests are still what most papers report, and they are worth knowing for two reasons: you will read them constantly, and when their assumptions hold they give the same answer as the resampling tests in a thousandth of the time. This part goes through the ones you will meet, each with the question it answers, the assumption it makes, and the resampling test it stands in for.

Two independent groups: the t-test

The t-test compares two means. It computes the gap, divides it by the standard error from Part 6, and asks how often a ratio that large would arise if the true gap were zero and the noise were bell-shaped. On the twenty-four runs, pooled: gap 11.9, standard error 6.7, ratio 1.77, p = 0.091. The permutation test of Part 4 gave 0.093 on the same numbers. When the bell holds, they agree, and here it holds well enough.

There are two versions and the choice matters more than textbooks admit. Student's t-test assumes the two groups have the same spread. Welch's t-test does not, and costs almost nothing when they do. Two conditions of an agent study very often have different spreads, a tool that fixes crashes removes the low scores from one group and not the other, and Student's test then gives the wrong p-value in a direction that depends on which group is larger. Use Welch by default, and treat a paper that reports "t-test" without saying which as having probably used Student's, since that is what most software defaults to.

The t-test's assumption is bell-shaped noise, or a sample large enough that the average is bell-shaped regardless, which happens surprisingly quickly, by thirty or so runs a side for anything that is not wildly skewed. Token costs are wildly skewed. Rubric scores usually are not.

Two paired groups: the paired t-test

Part 11 introduced seven tasks each run under both conditions, and showed the paired bootstrap giving an interval five times narrower than the unpaired one. The paired t-test is the formula version. Take the seven differences, one per task, and run a one-sample t-test on them against zero. It never sees the fourteen original scores, only the seven gaps, so the task difficulty that varied from 52 to 84 has already cancelled.

Test on the seven tasks p-value
Welch's t, treating the fourteen scores as two independent groups 0.028
Paired t, on the seven differences 0.00001

Same numbers. The wrong test says "probably real", the right one says "no doubt", because the wrong test is fighting noise that the design already removed. This is the single most common mistake in small agent studies, and the fix is always the same question: were these runs paired by task, seed or prompt? If they were, the differences are the data.

When the shape is wrong: rank tests

For skewed outcomes and small samples the bell assumption fails, and the classical answer is to replace the scores with their ranks, first to twenty-fourth, and test those. Ranks cannot be skewed, so the shape problem disappears, at the cost of throwing away how far apart the scores were.

The Mann-Whitney U test (also called the Wilcoxon rank-sum test) is the rank version of the two-group t-test, and it is, exactly, a permutation test on ranks. Its natural effect size is Cliff's delta from Part 6. On the twenty-four runs it gives p = 0.088, close to the others because the scores were not badly skewed. On the token costs of Part 2 it would be the right choice and the t-test would not.

The Wilcoxon signed-rank test is the rank version of the paired t-test. On the seven tasks it gives p = 0.0156, and that number is worth a pause: with seven pairs that all point the same way, 0.0156 is the smallest p-value the test can produce, because there are only 2⁷ = 128 ways the signs could have fallen and the two most extreme are the observed ones. Rank tests with tiny samples have a floor, and a paper reporting p = 0.0156 from seven pairs has hit it, not measured it.

Counts: chi-square and Fisher

For a two-by-two table of counts, 16 of 18 perfect against 5 of 18, there are two tests. Fisher's exact test, from Part 5, enumerates every possible table and is right at any size. The chi-square test compares the observed counts with the counts you would expect if the two rows were the same, and looks the discrepancy up in a table. It is an approximation, and the standard warning is that it needs an expected count of at least five in every cell. The 18-versus-18 table has a smallest expected count of 7.5, so both are fine, and they give 0.0007 and 0.0005. With three of eighteen against one of eighteen the chi-square approximation breaks and Fisher is the only honest choice. For large tables, hundreds of runs, the two are indistinguishable and chi-square is what you will see.

Three or more groups: ANOVA

The verification study had eight tool conditions, not two. Testing every pair with a t-test would be twenty-eight tests, and Part 7 explained what that does to the false-positive rate. The classical first step is analysis of variance (ANOVA), which asks one question of all the groups at once: is there any difference among these means, or could they all have come from one population? It compares the spread between group means with the spread within groups, and reports an F statistic and a p-value.

Add a third condition to the twenty-four runs, a boot probe that lands between the other two, and the one-way ANOVA across the three gives F = 1.75, p = 0.189. Not significant, because pooled across models the noise is too large, which is the lesson of Part 4 again: the classical version of stratifying is to add the model as a second factor, a two-way ANOVA, and the model's variance stops counting against the tool. The rank version of one-way ANOVA is the Kruskal-Wallis test, p = 0.179 here.

ANOVA says whether anything differs, not what. The follow-up is pairwise comparisons, and they need the correction of Part 7. The classical package is Tukey's HSD, which is built for exactly this, and Holm across the pairs does the same job. On the three conditions, the three pairwise Welch tests give raw p-values of 0.091, 0.219 and 0.581, and Holm turns them into 0.272, 0.438 and 0.581, none significant, in line with the ANOVA. The verification study did not use ANOVA. It pre-specified six contrasts and corrected across them, which is the cleaner design when you know in advance which comparisons matter.

Which test

Outcome Design Classical test Its resampling equivalent Effect size to report
A score or measurement two independent groups Welch's t-test permutation test on means, stratified if there are strata gap in points, Cohen's d, 95 per cent interval
A score, skewed or tiny sample two independent groups Mann-Whitney U permutation test on ranks Cliff's delta, difference of medians
A score two conditions on the same tasks or seeds paired t-test paired bootstrap, or permutation of signs mean difference and its interval
A score, skewed paired Wilcoxon signed-rank sign permutation median difference
A rate or count two groups Fisher's exact test (chi-square if every expected count is 5 or more) it is a permutation test already Newcombe interval for the gap
A score three or more groups one-way ANOVA, then Tukey or Holm pairs permutation of the F statistic per-pair gaps with corrected intervals
A score, skewed three or more groups Kruskal-Wallis, then pairwise Mann-Whitney with Holm permutation on ranks per-pair Cliff's delta
A score with strata two or more groups and a known factor such as model or task two-way ANOVA, or the regression of Part 13 stratified permutation adjusted gap and its interval

Two rules cover most of the table. If the runs came in pairs, use the paired row. If the outcome is skewed or the sample is small, use the rank row or a resampling test. Everything else is Welch.

Try it: which test

Answer three questions and the box names the test, its resampling equivalent, and the effect size to put beside it.

In Python

SciPy has every test in the table. The script runs each on the numbers from this part: the twenty-four runs, the seven paired tasks, the 16-of-18 table, and a third condition invented to sit between the other two.

import numpy as np
from scipy import stats
from statsmodels.stats.multitest import multipletests

with_tool    = np.array([88, 91, 79, 84, 95, 86, 62, 58, 71, 49, 66, 55])
without_tool = np.array([78, 84, 70, 75, 89, 62, 46, 53, 40, 58, 49, 37])

# Two independent groups: Student assumes equal spreads, Welch does not. Welch is the default.
print(f"Student t: p = {stats.ttest_ind(with_tool, without_tool).pvalue:.3f}")
print(f"Welch t:   p = {stats.ttest_ind(with_tool, without_tool, equal_var=False).pvalue:.3f}")
print(f"Mann-Whitney U: p = {stats.mannwhitneyu(with_tool, without_tool).pvalue:.3f}")

# Two paired groups: the seven tasks of Part 11, each run under both conditions
with_by_task    = np.array([81, 74, 66, 79, 70, 62, 84])
without_by_task = np.array([70, 61, 57, 68, 55, 52, 75])
print(f"paired t: p = {stats.ttest_rel(with_by_task, without_by_task).pvalue:.1e}   "
      f"(unpaired Welch on the same numbers: p = {stats.ttest_ind(with_by_task, without_by_task, equal_var=False).pvalue:.3f})")
print(f"Wilcoxon signed-rank: p = {stats.wilcoxon(with_by_task, without_by_task).pvalue:.4f}")

# Counts: chi-square wants expected counts of 5 or more in every cell; Fisher does not care
table = [[16, 2], [5, 13]]
chi2, p_chi, _, expected = stats.chi2_contingency(table)
print(f"chi-square: p = {p_chi:.4f} (smallest expected count {expected.min():.1f}); Fisher exact: p = {stats.fisher_exact(table).pvalue:.4f}")

# Three or more groups: add a third condition, a boot probe, and ask whether the three differ at all
boot_probe = np.array([84, 86, 77, 80, 91, 83, 58, 55, 66, 47, 63, 52])
f, p_anova = stats.f_oneway(with_tool, boot_probe, without_tool)
print(f"one-way ANOVA across three conditions: F = {f:.2f}, p = {p_anova:.3f}")
print(f"Kruskal-Wallis: p = {stats.kruskal(with_tool, boot_probe, without_tool).pvalue:.3f}")

# Then which pairs differ, with Holm across the three comparisons
pairs = {"shell vs none": (with_tool, without_tool), "probe vs none": (boot_probe, without_tool), "shell vs probe": (with_tool, boot_probe)}
raw = [stats.ttest_ind(a, b, equal_var=False).pvalue for a, b in pairs.values()]
_, holm, _, _ = multipletests(raw, method="holm")
for name, p, h in zip(pairs, raw, holm):
    print(f"  {name}: raw p = {p:.3f}, Holm p = {h:.3f}")

which prints:

Student t: p = 0.091
Welch t:   p = 0.091
Mann-Whitney U: p = 0.088
paired t: p = 1.0e-05   (unpaired Welch on the same numbers: p = 0.028)
Wilcoxon signed-rank: p = 0.0156
chi-square: p = 0.0007 (smallest expected count 7.5); Fisher exact: p = 0.0005
one-way ANOVA across three conditions: F = 1.75, p = 0.189
Kruskal-Wallis: p = 0.179
  shell vs none: raw p = 0.091, Holm p = 0.272
  probe vs none: raw p = 0.219, Holm p = 0.438
  shell vs probe: raw p = 0.581, Holm p = 0.581

Where this leaves us

The classical tests are formulas for the resampling tests, valid when the noise has the shape they assume, and they are what you will read in most papers. Welch's t for two groups, the paired t when the runs came in pairs, ranks when the outcome is skewed, Fisher for small counts, ANOVA when there are several groups followed by corrected pairwise comparisons. The paired row is the one to remember, because pairing turned 0.028 into 0.00001 on the same fourteen numbers. Every one of these tests compares groups. The next part is about the other shape a question can take: not "do these two differ" but "how does one thing change with another", which is where lines, correlations, and the regression that cross-checked the verification study come in.


Next: Lines Through the Data: correlation and what it does not mean, fitting a line and reading its slope, regression as the general form of every comparison so far, the mixed-effects model that gave each model its own baseline, and why P5 was an interaction.