The twenty-four runs of Part 1 gave a gap of 12 points. Before asking whether chance could produce it, we have to say what we think produced it, and say it precisely enough that the data could prove us wrong. That statement is a hypothesis, and the discipline of writing one down before looking is most of what separates a study from an anecdote.
This part is about how to write one, what its shadow is, and what it means for a hypothesis to be "supported". It uses the six from the verification study, because they were written before any analysis, and because three of them were not supported, which makes them better teaching material than six that were.
Four ingredients
A hypothesis is a bet, and like any bet it has to be specific enough to be settled. Here is one from the study, and its four parts.
Agents given a shell to run their code will produce applications with a higher functional score than agents given no verification tools, across the six models and the three tasks whose behaviour can be checked by automatic probes.
| Ingredient | In that sentence | What goes wrong without it |
|---|---|---|
| Comparison | shell vs no verification tools | "Verification helps" compared to what? A version with fewer tools, a different model, last month's run? |
| Outcome | functional score | "Better" could mean the score, the pass rate, the tokens, or the look of the thing, and each could point a different way |
| Direction | higher | Without it, any change counts as a win, including a drop |
| Scope | six models, three probe-checkable tasks | A result on one model and one task is a story about that model and that task |
Here are three hypotheses that look fine and are not.
"The testing tool is useful." No comparison, no outcome, no direction. It cannot lose, which means it cannot win either.
"Reasoning effort improves the agent." Improves what? The first study found that a higher reasoning setting lifted first-try perfect runs from 5 of 18 to 16 of 18 and left the final score unchanged, because the lower setting got there after corrections. Whether "improves" is true depends entirely on which outcome you name.
"The screenshot tool will change the interface score." A comparison and an outcome, but no direction, and "change" is not what anyone believes. The study's actual P4 said "higher", and then, as we will see, had to report that it could not confirm it.
The shadow
Every hypothesis casts a shadow: the statement that nothing is going on. For the shell hypothesis, the shadow is "the shell makes no difference to the functional score". This is the null hypothesis, and the odd, important fact about statistics is that it is the shadow that gets tested, not the claim.
Why test the thing you do not believe? Because the shadow is the only one of the two whose consequences can be calculated. If the shell makes no difference, then the label "shell" or "no shell" on each run is just a label, and the scores would have come out the same with the labels swapped. That is a precise, mechanical statement, and Part 4 turns it into a procedure: swap the labels, thousands of times, and see how big a gap swapping alone produces. The claim, by contrast, says only that the shell helps by some amount, and "some amount" does not tell you what the data should look like.
So the logic runs backwards from what intuition expects:
- Assume the shadow. The tool does nothing.
- Work out how surprising the observed gap of 12 would be if that were true.
- If it would be very surprising, abandon the shadow, and by elimination the claim stands. If it would not be surprising, keep the shadow, and the claim is not supported.
Note the wording of the last step. Keeping the shadow is not the same as proving it. A courtroom is the standard analogy and it is a good one.
| In court | In a study |
|---|---|
| The defendant is presumed innocent | The null hypothesis is presumed true |
| The prosecution presents evidence | The data are collected |
| "Guilty beyond reasonable doubt" | The gap would be very unlikely under the null, so the null is rejected |
| "Not guilty" | The null is not rejected. It does not mean innocent. It means the evidence was not enough |
| The standard of proof is set in advance | The threshold for "unlikely" is set in advance, conventionally 5 per cent |
A verdict of not guilty leaves open that the defendant did it and the case was weak. A hypothesis that is not supported leaves open that the tool helps and the study was too small to see it. The verification study's P4 is exactly this: the screenshot tool scored 6.9 points higher on visual tasks, and the study could not rule out chance to the required standard, so P4 was reported as not supported rather than as "screenshots do not help". Part 8 is about what it takes to say the stronger thing.
The alternative, and which way it points
The claim itself, once the shadow has been rejected, is called the alternative hypothesis. It comes in two shapes.
A one-sided alternative names a direction: the shell raises the score. A two-sided alternative says only that the score changes, up or down. The difference sounds like a technicality and is not, because a one-sided test only counts surprise in the named direction, which makes it easier to pass, and that ease is open to abuse: pick the direction after seeing which way the data went and a one-sided test flatters you.
The verification study expected a direction for five of its six hypotheses and still tested all of them two-sided. That is the cautious convention and the one I would recommend. A tool can hurt as well as help, and being surprised by a drop should count as being surprised. When you read a paper, the phrase "two-sided" or "two-tailed" beside a p-value is a small mark of good faith.
Six bets, written out
Here are the six primary hypotheses from the verification study, each reduced to its four ingredients, with what happened. The effect sizes and the corrected p-values will be explained in Parts 5 to 7. For now, read the first four columns as bets placed before the data existed, and the last as how they were settled.
| Comparison | Outcome | Expected | Scope | Result | |
|---|---|---|---|---|---|
| P1 | shell vs no verification | functional score | higher | 6 models, 3 probe-checkable tasks | +12.3 points. Supported |
| P2 | boot probe vs no verification | survival (did the app start) | higher | 6 models, all 7 tasks | +13 percentage points. Supported |
| P3 | shell vs no verification | tokens used (log scale) | higher, a cost | 6 models, all tasks | ×2.2. Supported |
| P4 | screenshots vs shell | human interface score | higher | 2 visually graded tasks | +6.9 points. Not supported |
| P5 | shell's benefit on a modification task vs its benefit on a fresh build | functional score | about the same | the kanban task, both versions | −0.7, interval −13.7 to +12.7. Not supported |
| P6 | shell vs no verification | human score on a performance-critical task | higher | one task, about 28 runs a side | +15.3 points. Not supported |
Three things to notice.
P3 is a hypothesis about a cost, and it was expected to come out "higher". Supported, here, means "yes, the tool is expensive, as we said it would be". A hypothesis does not have to be good news.
P5 is a hypothesis whose expected answer was no difference: the shell should help a modification task about as much as a fresh build. That is a legitimate thing to predict, but it cannot be settled by the machinery above, because the machinery only ever rejects the shadow, and here the shadow is the claim. The study reported it plainly: the gap was near zero, and the interval around it was so wide that it would have been near zero whatever the truth was. Part 8 covers how to make a "no difference" hypothesis testable, and the planned follow-up study is built around doing so.
P6 has the largest effect of the six, 15 points, and was not supported, because it rested on one task with about 28 runs a side. That is the sample-size lesson of Part 1 showing up in a real table: a big gap on few runs is weaker evidence than a modest gap on many.
Primary, secondary, pre-specified
The six above are primary hypotheses, the study's headline bets. Because there were six of them, chance had six tries at producing an impressive-looking gap, and Part 7 explains the correction the study applied for that. The study also asked a further set of questions after the fact, things like "does the boot probe alone get most of the shell's benefit?", and reported them as secondary contrasts: exploratory, uncorrected, and labelled as such, so that no reader mistakes them for bets that were placed in advance.
Pre-specified means the hypotheses were written, with all four ingredients and the test to be used, before the analysis was run. The verification study's methods section says the six were "decided in advance", and then adds a sentence I would encourage anyone to copy: that they were "pre-specified rather than pre-registered in the external-registry sense", because the plan was fixed in a private document rather than deposited with a third party. Part 10 is about that distinction and why it matters more for some studies than for others.
Try it: the hypothesis builder
Pick the four ingredients and the box writes the hypothesis, its shadow, and what each verdict would mean. The presets are the six from the study.
Press P5 and read what the box says about a claim whose expected answer is "the same". Then set the direction to "different, either way" for P1 and notice that the verdict text changes: a two-sided claim is supported by a surprise in either direction, which is why it is the safer default.
In Python
A hypothesis is data, and writing the six down as records makes it obvious when one is missing an ingredient, and which one cannot be settled by the ordinary test.
# A hypothesis is data: write the six down with all four ingredients, then check nothing is missing
hypotheses = [
dict(id="P1", compare=("shell", "no verification"), outcome="functional score", direction="higher", scope="6 models, 3 probe-checkable tasks"),
dict(id="P2", compare=("boot probe", "no verification"), outcome="survival", direction="higher", scope="6 models, 7 tasks"),
dict(id="P3", compare=("shell", "no verification"), outcome="log tokens", direction="higher", scope="6 models, all tasks"),
dict(id="P4", compare=("screenshots", "shell"), outcome="human interface score", direction="higher", scope="2 visually graded tasks"),
dict(id="P5", compare=("shell benefit on modification", "shell benefit on fresh build"), outcome="functional score", direction="same", scope="the kanban task"),
dict(id="P6", compare=("shell", "no verification"), outcome="human interface score", direction="higher", scope="the performance task"),
]
for h in hypotheses:
a, b = h["compare"]
null = f"{a} makes no difference to {h['outcome']} compared with {b}"
testable = "needs an equivalence bound (Part 8)" if h["direction"] == "same" else "two-sided permutation test"
print(f"{h['id']}: {a} vs {b}, {h['outcome']}, expected {h['direction']}. Null: {null}. Test: {testable}")
which prints:
P1: shell vs no verification, functional score, expected higher. Null: shell makes no difference to functional score compared with no verification. Test: two-sided permutation test
P2: boot probe vs no verification, survival, expected higher. Null: boot probe makes no difference to survival compared with no verification. Test: two-sided permutation test
P3: shell vs no verification, log tokens, expected higher. Null: shell makes no difference to log tokens compared with no verification. Test: two-sided permutation test
P4: screenshots vs shell, human interface score, expected higher. Null: screenshots makes no difference to human interface score compared with shell. Test: two-sided permutation test
P5: shell benefit on modification vs shell benefit on fresh build, functional score, expected same. Null: shell benefit on modification makes no difference to functional score compared with shell benefit on fresh build. Test: needs an equivalence bound (Part 8)
P6: shell vs no verification, human interface score, expected higher. Null: shell makes no difference to human interface score compared with no verification. Test: two-sided permutation test
Where this leaves us
A hypothesis is a bet with four parts: what is compared, on which outcome, in which direction, and over what scope. Its shadow, the null hypothesis, is the statement that the comparison makes no difference, and it is the shadow that gets tested, because only the shadow says exactly what the data should look like. "Supported" means the shadow was rejected. "Not supported" means it was not, and that is all it means. The next part does the rejecting, by taking the shadow at its word: if the labels on the twenty-four runs do not matter, swap them and see.
Next: Shuffle the Labels: the permutation test, one shuffle at a time. Why the gap of 12 in the twenty-four runs is not surprising when you shuffle everything, and very surprising when you shuffle within each model, and what that says about how the verification study ran its tests.