The permutation test of Part 4 rests on one assumption: that under the null, any run could have carried either label. The intervals of Part 6 rest on the scores being honest measurements. Neither is guaranteed by the maths. Both are guaranteed, or not, by how the study was built, and no amount of statistics afterwards can repair a design that let something other than the tool decide the scores.
This part is the list of design choices that make a result believable, each attached to the place where one of my two studies got it right or wrong. The first study is the honest starting point again: I did it the way most people do a first study, and the second study's methods section is, in large part, a list of the things I changed.
Assigned or randomised
The first study's methods section contains a sentence I would now write differently: the study was observational, "with conditions assigned rather than randomized". What that means in practice is that I decided which runs got the testing tool, or the higher reasoning setting, as I went, in batches, according to what I wanted to learn next.
That sounds harmless and is not, because anything else that varied between batches now travels with the condition. In an agent study the list is long.
| Something that drifts | How it becomes a confound |
|---|---|
| The model itself | Providers update models under the same name. A batch run in week three may be running a different model from week one |
| Load on the provider | Slow days produce more timeouts and more retries, which cost tokens and sometimes points |
| Your own prompt | Every small rewording between batches is a second variable |
| Task order | If the hard tasks were run last, and the tool was added last, the tool inherits the hard tasks |
| Your attention | The configurations you repeated are the ones you found interesting, which is not a random sample of the ones you ran |
Any of these can produce a gap between conditions that has nothing to do with the condition, and the shuffle of Part 4 cannot tell the difference, because from the data's point of view a batch effect and a tool effect look identical. The first study named one of them, batch timing, as a residual confound for one of its comparisons, which is the right thing to do when it has already happened. The better thing is to prevent it.
Randomisation prevents it. Decide the order of all the runs by lot before starting, so that the tool and no-tool runs are interleaved across time, tasks and provider moods. Then anything that drifts is equally likely to land in either condition, and the only systematic difference between the two groups is the thing you changed. That is what makes the labels exchangeable, and it is what makes the permutation test's assumption true rather than hoped for.
The verification study went further and removed the drift where it could. Its agent was a small program under my own control, and the tool list was the single thing that changed between conditions: the prompts were identical byte for byte across all 1,116 runs, and the paper says so, with the hash that proves it. When only one thing varies, only one thing can explain the result.
Blinding
The grades in both studies were given by a human, me, against a rubric. In the first study I knew, as I graded, which condition each run came from. I do not think that changed my scores. I also cannot prove it, and neither could you, and a reviewer is entitled to assume it did.
The bias is not dishonesty. It is expectation. A grader who knows a run had the testing tool looks a little harder for evidence that the tests helped, gives a borderline criterion the benefit of the doubt, and does the opposite for the run they expect to be worse. Two or three points a run, in the direction of the hypothesis, is enough to manufacture a finding from nothing, and the demo at the end shows exactly how little it takes.
Blinding removes the expectation by removing the information. The verification study graded every run under an opaque identifier, in an order shuffled by a fixed seed, with the model and condition fields stripped from every scorecard before it reached me. The paper's phrase is that blinding was "enforced by the setup, not left to discipline", which is the standard to aim for: not "I tried not to let it influence me" but "I could not have known".
Freezing the rubric
The rubric is the list of criteria and their weights, and the moment it is written matters as much as what it says. Write it after seeing a few runs and it will drift toward the things the favoured condition happens to do well, without anyone intending it. The verification study wrote and froze its scorecard before a single run existed, and applied it unchanged to all 1,116.
Freezing has a companion rule for the things a rubric cannot anticipate. When an app would not start without, say, a port number, the study allowed the grader to supply one, but only from within "the artifact's own declared configuration space": a value the app itself listed as valid. Anything more, a missing file, a guessed environment variable, was a failure the app had earned. The rule sounds fussy and is the difference between a grader who evaluates and a grader who helps.
Counting every run: intention to treat
Two situations come up in every agent study, and both invite a quiet form of cheating.
The first: a model is given the shell and never uses it. Does that run count in the shell condition? The verification study's answer is yes, always, and the paper borrows the name from clinical medicine: intention to treat. A patient who never picks up the prescription still counts in the drug group, because dropping them would leave the drug group with only the patients who were well enough to reach the pharmacy. A model that ignores its tool still counts in the tool condition, because dropping it would leave the tool condition with only the models that use tools well, which is not the question.
The second: some runs produce apps that never start, and score close to zero. The temptation is to set them aside as "not real attempts". Here is why that is not allowed, with numbers.
| Runs | Started | Score of the ones that started | Mean counting every run | Mean counting only starters | |
|---|---|---|---|---|---|
| Condition A | 10 | 6 | 80 each | 48 | 80 |
| Condition B | 10 | 10 | 60 each | 60 | 60 |
Condition A crashes four times in ten and is excellent when it does not. Counting every run, B is better, 60 to 48. Counting only the apps that started, A is better, 80 to 60. Excluding the failures does not clean the data. It rewards the condition that fails most, because failing removes its worst runs from the average. The paper calls this a selection effect and refuses it in one sentence: excluding non-starters "would quietly reward the conditions that produce the most non-starting applications". Every run that was launched is graded, and a crash is a low score, not a missing one.
Can someone else get the same numbers
Two words that sound alike and are not. A result is reproducible if someone with your data and your code gets your numbers. It is replicable if someone with new data and their own code gets your conclusion. The first is a matter of engineering and is entirely within your control. The second is what science is for, and it is out of your hands.
The verification study's engineering is one line: the random seed, 20260703, is written into the released analysis code, so that rerunning it regenerates every permutation, every bootstrap interval, and every number in the paper exactly. The grading order was shuffled with a fixed seed too. When a number in a paper cannot be regenerated, the reader has to take the author's word for it, and the whole apparatus of this series exists so that they do not have to.
Reading a kappa
Even a blinded grader is a single grader, and a reviewer will ask whether a second pass would give the same scores. The verification study answered by regrading a 10 per cent sample, 112 runs, four weeks later, still blind, and reporting three numbers: 98.8 per cent of individual rubric items were scored identically, the linearly weighted kappa was 0.973, and the intraclass correlation on the per-run total was 0.997.
Cohen's kappa is agreement corrected for luck. Two graders who each mark most runs "pass" will agree most of the time by chance alone, so raw agreement flatters them. Suppose they agree on 90 per cent of items, and that from their individual pass rates you would expect them to agree on 60 per cent by chance. Kappa is the agreement beyond chance as a fraction of the agreement that was available beyond chance: (0.90 − 0.60) / (1 − 0.60) = 0.75. A kappa of 1 is perfect, 0 is chance, and the usual reading is that above 0.8 is excellent. Weighted kappa gives partial credit for near misses, a 3 against a 4 counting as closer than a 3 against a 0, which is right for graded criteria. The intraclass correlation does the same job for continuous totals. At 0.973 and 0.997, the two passes were as close as two passes get, and the mean absolute difference per run, 0.52 points out of 100, is smaller than any effect the paper reports. That is what the numbers are for: they say that grader noise cannot be what produced the results.
Four ways a study can be wrong
Methodologists sort the ways a conclusion can fail into four kinds, and a threats to validity section walks through them. Here they are, with the questions to ask of any paper.
| Threat | The question | In the two studies |
|---|---|---|
| Internal | Did the condition cause the gap, or did something that travelled with it? | Confounds, grader expectation, non-starters. The first study admitted batch timing. The second randomised, blinded and counted every run |
| External | Does the result hold beyond these models, tasks and prompts? | Six models and seven tasks say more than one of each. Neither says anything about next year's models, and both papers say so |
| Construct | Does the outcome measure the thing the claim is about? | A functional rubric measures what the rubric lists. "Quality" is a bigger word. The verification study's screenshot tool could not click, so its visual conditions understate what sight could do, and the paper flags it |
| Statistical conclusion | Were the tests, corrections and sample sizes adequate? | Parts 4 to 8 of this series. The second study's P5 is the example of a sample size that was not |
A good limitations section is not an apology. It is this table, filled in without flinching, so that the reader knows which of the four the author has ruled out and which remain.
Try it: the grader who knows
The demo simulates a study where the tool does nothing at all, and a grader who, knowing which runs had the tool, leans by a few points in its favour. See how small a lean it takes to produce a "finding".
Set the lean to 3 points, less than the disagreement between two honest graders, and 48 runs a side, and press Grade 20 studies. About one in five will report a finding, from a tool that does nothing, four times the honest false-alarm rate. Push the runs to 200 a side and it is one in two. Tick the blinding box and press again. Then notice the uncomfortable part: more runs make the problem worse, not better, because a systematic lean is exactly the kind of small, consistent effect that large samples are good at detecting.
In Python
Kappa is one call in scikit-learn. The two small simulations after it are the two tables above: the crash-exclusion selection effect, and the grader who leans.
import numpy as np
from sklearn.metrics import cohen_kappa_score
# Two grading passes over the same 20 rubric items, scored 0 to 3
first = [3, 3, 2, 3, 0, 3, 1, 3, 3, 2, 3, 3, 0, 3, 2, 3, 3, 1, 3, 3]
second = [3, 3, 2, 3, 0, 3, 2, 3, 3, 2, 3, 3, 0, 3, 3, 3, 3, 1, 3, 3]
agree = np.mean(np.array(first) == np.array(second))
print(f"raw agreement {agree:.0%}")
print(f"Cohen's kappa {cohen_kappa_score(first, second):.3f}")
print(f"linearly weighted kappa {cohen_kappa_score(first, second, weights='linear'):.3f}")
# Why excluding the crashes flatters the condition that crashes most
A = [0, 0, 0, 0, 80, 80, 80, 80, 80, 80] # starts 6 times in 10, excellent when it does
B = [60] * 10 # always starts, always adequate
print(f"counting every run: A {np.mean(A):.0f}, B {np.mean(B):.0f}")
print(f"counting only starters: A {np.mean([x for x in A if x]):.0f}, B {np.mean(B):.0f}")
# A grader who leans 3 points toward the tool turns a null into a finding, given enough runs
from scipy import stats
rng = np.random.default_rng(3)
hits = 0
for _ in range(200):
a = rng.normal(62, 15, 48) + 3 # the tool does nothing; the grader adds 3
b = rng.normal(62, 15, 48)
hits += stats.ttest_ind(a, b).pvalue < 0.05
print(f"a 3-point lean with 48 runs a side: {hits / 2:.0f}% of null studies report a finding")
which prints:
raw agreement 90%
Cohen's kappa 0.803
linearly weighted kappa 0.890
counting every run: A 48, B 60
counting only starters: A 80, B 60
a 3-point lean with 48 runs a side: 18% of null studies report a finding
Where this leaves us
The tests assume exchangeable labels and honest scores, and the design is what makes those true. Randomise the order so drift cannot travel with the condition, change one thing at a time, blind the grader so expectation has nowhere to go, freeze the rubric before the runs exist, count every run including the ones that crashed and the ones that ignored their tools, seed everything so that every number can be regenerated, and measure your own consistency so that grader noise can be ruled out by arithmetic rather than by assurance. Then fill in the four-row table without flinching. What remains is the question of when the choices were made, which is the last part.
Next: When You Decide Matters: post-hoc, exploratory and confirmatory analysis, the freeze, preregistration and registered reports, results-blind review, and a demo in which you try to find a significant result in pure noise, then lock your choices first and try again.