Every test so far has compared groups: with the tool against without, higher reasoning against lower. The other shape a question takes is how does one thing change with another: does spending more tokens buy a better score, does taking more screenshots buy a bigger gain, does the tool's effect grow with the size of the task. Those questions are answered with lines, and the machinery of lines, correlation and regression, turns out to contain everything from the earlier parts as special cases. It is also the machinery the verification study used to cross-check its permutation results, so by the end of this part you will be able to read the one line in that paper that has not yet been explained.
Correlation
Give the twenty-four runs a second number each: the tokens they spent, in thousands. Runs with the tool spend more, around 640K, because the tool runs code and reads the output. Runs without spend around 266K. Plot tokens against score and ask the correlation question: when one is high, is the other?
Pearson's r answers it with a number between −1 and +1: +1 means the points lie exactly on a rising line, −1 exactly on a falling one, 0 means no straight-line relationship at all. Across all twenty-four runs, r = 0.34. Tokens and score rise together, moderately. Spearman's rho is the same idea on ranks, so it does not care whether the relationship is a straight line, only whether it is consistently upward, and it is the one to use with skewed variables like tokens. Here rho = 0.30, close to r because nothing is very skewed.
Now look inside each condition. Among the twelve runs with the tool, the correlation between tokens and score is 0.00. Among the twelve without, also 0.00. Spending more tokens buys nothing within a condition. The correlation across all twenty-four is entirely the tool: it raises the score and it raises the token count, and a line through the whole cloud connects the two clusters without saying anything about what happens inside them.
This is what "correlation is not causation" means in practice. It is rarely that the correlation is spurious. It is that a third thing, here the condition, moves both variables, and the line faithfully reports their joint motion. The verification study met the reverse case and reported it: the models that took the most screenshots were not the ones that gained the most from the screenshot tool, "so we cannot credit the pooled benefit to looking". A correlation that is absent where the mechanism says it should be present is evidence against the mechanism, which is a legitimate and underused kind of finding.
Two more cautions. Pearson's r is as sensitive to a single outlier as the mean is, and one runaway run can create or destroy a correlation on its own, which the demo below lets you do by hand. And r says nothing about the size of the effect: a correlation of 0.9 can describe a slope of half a point per million tokens if the noise is tiny. For the size, you need the line.
The line, and how to read it
Linear regression fits the line that comes closest to the points, closest meaning that the squared vertical distances add up to the least, and reports two numbers: the intercept, where the line crosses the vertical axis, and the slope, how much the outcome changes per unit of the input. On the twenty-four runs the line is score = 54.5 + 0.029 × tokens, so every extra 100K tokens goes with 2.9 more points. The slope is the effect size, in the outcome's units per unit of input, and like every effect size it needs an interval, which the bootstrap of Part 11 gives by resampling the pairs.
The square of the correlation, r², is the share of the outcome's variation that the line accounts for. Here it is 0.11: tokens explain 11 per cent of why scores differ, and 89 per cent is other things. In a paper, a slope with a tight interval and an r² of 0.1 means "a real but small influence among many". A slope with an r² of 0.9 means the input nearly determines the outcome.
Regression is every comparison so far
Here is the trick that makes regression the general tool. Give the tool a number, 1 for with and 0 for without, and fit score = a + b × tool. The slope b is the gap between the two means, 11.9, and the intercept is the mean without the tool. The t-test of Part 12 is a regression with one yes-or-no input.
Now add the model as a second input, 1 for the weak model and 0 for the strong: score = a + b × tool + c × model. The fit gives b = +11.9, c = −28.1, and the interval on b is +5.4 to +18.5. That interval is the one from Part 6's stratified bootstrap, +6.2 to +17.7, arrived at by a different road. Adding the model as a term does exactly what shuffling within the model did: it takes the 28-point difference between the models out of the noise, so that the tool's 12 points can be seen against 15 points of run-to-run spread instead of 30 points of model spread. This is called adjusting for or controlling for the model, and it is what a paper means when a results table says "adjusted for model and task".
| Way of asking "does the tool help, given that models differ?" | Answer | Where |
|---|---|---|
| Shuffle labels within each model | p = 0.0014 | Part 4 |
| Resample within each cell | +6.2 to +17.7 | Part 6 |
| Regression with the model as a term | +5.4 to +18.5 | this part |
| Two-way ANOVA with model as a factor | the same as the regression | Part 12 |
Four names, one idea. Each one says: compare like with like, and let the design remove the noise it was built to remove.
Mixed effects: letting each model have its own baseline
With two models, adding the model as a term is fine. The verification study had six, and it treated them differently. Instead of estimating six separate baselines as fixed numbers, a mixed-effects model treats the models as draws from a population of possible models, each with its own baseline pulled from a shared distribution, and estimates the tool's effect as something that holds across that population. The "mixed" is that the tool's effect is a fixed effect, the same for everyone, and each model's baseline is a random effect, one draw per model.
On the twenty-four runs, the mixed model gives the tool +11.9 with a standard error of 3.2, the same as the plain regression, because with two models there is little difference. The estimated spread of model baselines is 19.7 points, which is the model gap seen from the mixed model's side. With six models and seven tasks, both as random effects, the mixed model does the stratifying automatically and gives an honest standard error for a tool effect that is meant to generalise to models and tasks not in the study.
That is the cross-check the verification study reports. Every one of its six hypotheses was tested by stratified permutation, and then re-estimated by "a completely different statistical method, a mixed-effects regression", and the two sets of estimates matched to two decimal places. When a resampling method that assumes nothing and a model-based method that assumes a lot agree, the agreement says the assumptions were not doing the work, which is the strongest kind of robustness check a small study can offer.
Interactions: why P5 was so wide
The last thing regression adds is the ability to ask whether an effect depends on something. Does the tool help more on a modification task than on a fresh build? That is P5, and in regression language it is an interaction: score = a + b × tool + c × modification + d × (tool × modification). The coefficient d is the tool's gap on the modification task minus its gap on the fresh build, the difference of differences that Part 6 mentioned, and P5's hypothesis was that d is about zero.
Interactions are expensive. On simulated data with the same 12-point gap on both tasks and 12 runs per cell, the tool's main effect has a standard error of 7.2 and the interaction has a standard error of 10.2, because d is built from two subtractions and every subtraction adds noise. Halving the standard error of an interaction takes four times the runs. That is the arithmetic behind P5's interval of −13.7 to +12.7 on 144 runs, and it is the general warning: any claim of the form "the effect is bigger for X than for Y" needs several times the data of the claim "there is an effect".
Try it: the line
Twelve draggable points, a fitted line, and the correlation and slope updated live. Drag one point far away and watch r and the slope follow it.
Press Add a runaway run: one run that spent nearly a million tokens and scored 30 turns a flat line into a falling one and r from 0 to −0.5. Then drag it up to a score of 100 and the line rises instead. Twelve honest runs and one strange one, and the one decides the story. Spearman's rho would move much less, because it only sees that the new point is last in one ranking and first in the other.
In Python
The table of runs gets a tokens column, and the rest is SciPy for the correlations and the line, and statsmodels for the regression, the mixed model, and the interaction, using the formula language you will see in papers' supplementary code: score ~ tool + model means "score as a function of tool and model".
import numpy as np, pandas as pd
from scipy import stats
import statsmodels.formula.api as smf
# The twenty-four runs as a table, with the tokens each run spent (thousands)
runs = pd.DataFrame({
"score": [88, 91, 79, 84, 95, 86, 62, 58, 71, 49, 66, 55, 78, 84, 70, 75, 89, 62, 46, 53, 40, 58, 49, 37],
"tokens": [530, 810, 575, 590, 615, 720, 540, 690, 655, 700, 612, 640, 300, 230, 290, 250, 265, 270, 240, 255, 275, 262, 310, 245],
"tool": [1] * 12 + [0] * 12,
"model": (["A"] * 6 + ["B"] * 6) * 2,
})
# 1. Correlation: do runs that spend more tokens score higher?
r, p = stats.pearsonr(runs.tokens, runs.score)
rho, _ = stats.spearmanr(runs.tokens, runs.score)
print(f"all 24 runs: Pearson r = {r:.2f} (p = {p:.3f}), Spearman rho = {rho:.2f}")
for t, grp in runs.groupby("tool"):
print(f" within {'with' if t else 'without'} the tool: r = {stats.pearsonr(grp.tokens, grp.score)[0]:+.2f}")
# 2. A line through the data, and how to read its slope
fit = stats.linregress(runs.tokens, runs.score)
print(f"score = {fit.intercept:.1f} + {fit.slope:.3f} x tokens, so +100K tokens goes with {100 * fit.slope:+.1f} points; r^2 = {fit.rvalue**2:.2f}")
# 3. Regression as the general comparison: the tool's gap, adjusted for model
ols = smf.ols("score ~ tool + model", data=runs).fit()
print(f"OLS: tool {ols.params['tool']:+.1f} (SE {ols.bse['tool']:.1f}, 95% {ols.conf_int().loc['tool', 0]:+.1f} to {ols.conf_int().loc['tool', 1]:+.1f}), "
f"model B {ols.params['model[T.B]']:+.1f}")
# 4. The same thing as a mixed-effects model: model as a random intercept, the paper's cross-check
mixed = smf.mixedlm("score ~ tool", data=runs, groups=runs["model"]).fit()
print(f"mixed: tool {mixed.params['tool']:+.1f} (SE {mixed.bse['tool']:.1f}), spread of model baselines {np.sqrt(mixed.cov_re.iloc[0, 0]):.1f}")
# 5. P5 as an interaction: does the tool's gap differ between a fresh build and a modification task?
rng = np.random.default_rng(20260703)
p5 = pd.DataFrame({"task": np.repeat(["fresh", "modify"], 24), "tool": np.tile(np.repeat([0, 1], 12), 2)})
p5["score"] = 60 + 12 * p5.tool - 8 * (p5.task == "modify") + rng.normal(0, 15, 48) # same 12-point gap on both tasks
inter = smf.ols("score ~ tool * task", data=p5).fit()
print(f"interaction (tool x modify): {inter.params['tool:task[T.modify]']:+.1f}, SE {inter.bse['tool:task[T.modify]']:.1f}; "
f"the main tool effect's SE on the same data: {inter.bse['tool']:.1f}")
which prints:
all 24 runs: Pearson r = 0.34 (p = 0.106), Spearman rho = 0.30
within without the tool: r = -0.00
within with the tool: r = +0.00
score = 54.5 + 0.029 x tokens, so +100K tokens goes with +2.9 points; r^2 = 0.11
OLS: tool +11.9 (SE 3.2, 95% +5.4 to +18.5), model B -28.1
mixed: tool +11.9 (SE 3.2), spread of model baselines 19.7
interaction (tool x modify): +3.5, SE 10.2; the main tool effect's SE on the same data: 7.2
Where this leaves us
A correlation is a number for how consistently two things move together, and the most common reason they do is a third thing moving both, which is why the twenty-four runs show a correlation between tokens and score that vanishes inside each condition. A regression line puts a size on the relationship, and regression with a yes-or-no input is the t-test, with a second input it is the stratified analysis, with the models as random draws it is the mixed model the verification study used as its cross-check, and with a product term it is the interaction that made P5 so hard to pin down. One tool, and every comparison in the series is a special case of it. The next part takes all of this to the place most readers of this series will actually use it: a leaderboard.
Next: Reading a Leaderboard: a benchmark score is a rate with an interval, pass@k and the estimator that gets it right, seeds and why temperature zero is not deterministic, comparing two models on the same tasks with McNemar's test and the paired bootstrap, and what twenty benchmarks do to the word "wins".