Part 5 listed the things a p-value is not, and the first was the one everybody wants: the probability that the hypothesis is true. Nothing since has provided it either. Confidence intervals came with a promise about a procedure's long-run hit rate, not about this interval. The whole apparatus has been about how surprising the data would be under various assumptions, never about how probable the assumptions are given the data.
There is a framework that answers the wanted question directly, and it is older than any of the tests in the series: Bayesian inference, named after Thomas Bayes, an eighteenth-century clergyman whose one theorem does all the work. This part is the shortest possible honest introduction: what it asks for that the other approach does not, what it gives back, and where it lands on the examples we already know.
The price and the product
To say how probable a hypothesis is after seeing the data, you have to say how probable it was before. That is the price, and it is the whole difference between the two approaches. The prior is your belief about the effect before the study: the tool probably helps a little, or effects in this field are usually small, or I have no idea. The likelihood is what the data say: how probable this result would be at each possible value of the effect, which is the same object the p-value was computed from. Bayes' theorem multiplies the two and rescales, and the result is the posterior: your belief about the effect after the study, as a full distribution over every possible value.
In the courtroom of Part 3, the classical juror asks "how likely is this evidence if the defendant is innocent?" and stops. The Bayesian juror also asks "how likely was guilt before the evidence?" and combines the two. Neither is wrong. The second answers the question people actually have, and pays for it by having to write the prior down.
From the posterior come the things the earlier parts could only approximate. A credible interval is the range holding 95 per cent of the posterior, and it means exactly what people wrongly think a confidence interval means: given the data and the prior, there is a 95 per cent probability the effect is in there. The probability that the effect is positive is an area under the curve. The probability that it is inside Part 8's band is another.
A rate, and why Wilson was nearly right
Start with 16 of 18 perfect runs. Say you knew nothing beforehand: every pass rate from 0 to 100 per cent equally likely, a flat prior. The likelihood of 16 successes in 18 is highest near 89 per cent and falls away on either side. Multiply, and the posterior is a distribution called Beta(17, 3): the flat prior contributed one imaginary success and one imaginary failure, the data contributed 16 and 2. Its 95 per cent credible interval is 66.9 to 96.6 per cent.
Wilson's interval from Part 2 was 67.2 to 96.9. The two methods, from different philosophies, land a third of a point apart, and that is not a coincidence: Wilson's formula adds about two imaginary successes and two failures, and a flat prior adds one of each, so they are doing nearly the same thing. When the prior is flat and the sample is not tiny, Bayesian and classical intervals agree, and the argument between the schools is about interpretation rather than numbers.
What the posterior adds is the ability to answer questions the interval cannot. Is the true rate above 80 per cent? The area of the posterior beyond 0.8 is 0.76, so: probably, three to one. No classical test produces that sentence.
A gap, and what a sceptical prior does
Now the twenty-four runs. Stratified by model, the gap is +11.9 with a standard error of 3.2 (Part 13). With a flat prior the posterior is a bell curve centred on 11.9 with that spread, the credible interval is +5.6 to +18.2, and the probability that the gap is positive is, to three decimals, 1.000. Same numbers as the classical interval, same conclusion.
Now suppose you came to the study sceptical: tools in this field usually make a few points' difference, so your prior is centred on zero with a spread of 5 points. Bayes' theorem combines the prior and the data by weighting each by its precision, one over its spread squared:
| Centre | Spread | Weight | |
|---|---|---|---|
| Sceptical prior | 0 | 5.0 | 1 / 25 = 0.04 |
| The data | +11.9 | 3.2 | 1 / 10.2 = 0.10 |
| Posterior | +8.4 | 2.7 | 0.14 |
The posterior centre is the weighted average, (0 × 0.04 + 11.9 × 0.10) / 0.14 = 8.4, and the posterior spread is narrower than either input. The data pulled the sceptic most of the way, but not all the way: the estimate shrinks toward the prior by about a third, and the credible interval is +3.1 to +13.7. The probability that the gap is positive is still 0.999, because 12 points on 24 runs is a lot of evidence even for a sceptic. The probability that it lies inside ±5 is 0.10.
Shrinkage is not a distortion. It is what a reasonable person does with a surprising result from a small study, and it is why Bayesian estimates from small samples are usually closer to the truth than the raw numbers: the raw gap of 11.9 was one draw from a noisy process, and the prior's pull toward the typical is a hedge against the draw having been lucky. The cost is that the hedge must be declared, and a prior chosen after seeing the data is the garden of forking paths with a new name.
The Bayes factor
Sometimes the question is not "how big is the effect" but "which of two stories fits better", and the Bayesian answer is the Bayes factor: how many times more probable the data are under one hypothesis than under the other. It is the ratio of two likelihoods, each averaged over its hypothesis's prior.
Take the first-try rate again, and two stories. Under the first, the reasoning setting makes no difference and the rate is a coin toss, exactly 50 per cent. Under the second, the rate could be anything from 0 to 100. The data, 16 of 18, are about 90 times more probable under the second story than the first. That is a Bayes factor of 90 to 1, and by the conventional scale, where 3 is "some evidence", 10 is "strong" and 30 is "very strong", it is decisive. Had the count been 11 of 18, the factor would have been 0.4 to 1, mild evidence for the coin toss, which is something no p-value can express: a p-value can fail to reject the null, but only a Bayes factor can say the null is the better story.
The catch is the prior on the second story. "Anything from 0 to 100" is a generous alternative, and it spreads its probability thin, which is why 11 of 18 counts against it. Choose a narrower alternative, "somewhere between 50 and 100", and the factor changes. Bayes factors are honest about this in the sense that the prior is on the page, and fragile in the sense that reasonable priors can give factors that differ by a multiple of ten. A paper that reports one should report the sensitivity too.
The region of practical equivalence
Part 8 defended a null with two one-sided tests against a bound chosen in advance. The Bayesian version is simpler to state: draw the same band, ±5 points, call it the region of practical equivalence (ROPE), and read off how much of the posterior falls inside it. If nearly all of it does, the effect is practically nothing. If nearly none does, the effect matters. If it is split, the study did not settle it. On the twenty-four runs with the sceptical prior, 10 per cent of the posterior lies inside ±5, so the tool's effect is probably not negligible, but a one-in-ten chance that it is remains.
The correspondence with Part 8 is close: the classical equivalence test asks whether the 90 per cent interval fits inside the band, and the ROPE asks what fraction of the posterior does, and with a flat prior those are nearly the same question. The Bayesian version gives a number instead of a verdict, which is often more useful and always more honest about the middle cases.
When to reach for it
Bayesian methods earn their place in three situations. When the sample is small and there is real prior knowledge, from earlier studies or from physics, that it would be wasteful to ignore. When the question is a decision, ship or do not ship, that needs a probability rather than a verdict. And when the analysis will be updated as data arrive, because a posterior can be updated indefinitely without the multiple-looks problem that makes peeking at a p-value a sin.
They do not earn a pass on Part 10. The prior is an analytic choice, and a prior chosen after seeing the data is exactly as corrupting as an outcome chosen after seeing the data. Preregister the prior with the rest of the plan, or report the analysis under several priors so the reader can see how much the conclusion depends on it. And be honest about the one thing the framework cannot supply: a prior that everyone agrees on. When two readers hold different priors they will reach different posteriors from the same data, and that is not a flaw in the arithmetic. It is a true statement about the two readers.
Try it: prior, data, posterior
Set what the data say (the gap and its standard error), how sceptical you were beforehand, and the band that counts as no meaningful effect. The curves and the three probabilities update.
Three things to try. Slide the prior spread all the way right to make it flat and watch the posterior sit exactly on the data. Bring it to 2, a very sceptical prior, and watch a 12-point gap shrink to 5 while the probability that it is positive stays high: scepticism about the size is not scepticism about the sign. Then set the standard error to 10, a six-run study, and see the prior take over entirely, which is what should happen when the data are weak.
In Python
The three calculations of this part: the Beta posterior for a rate, the precision-weighted update for a gap under two priors, and a Bayes factor computed from the marginal likelihoods, which for a rate with a flat prior is a Beta function and nothing harder.
import numpy as np
from scipy import stats
from math import lgamma, exp
# 1. A rate: 16 of 18 perfect runs, with a flat prior. The posterior is Beta(17, 3).
post = stats.beta(1 + 16, 1 + 2)
lo, hi = post.ppf([0.025, 0.975])
print(f"16 of 18, flat prior: posterior mean {post.mean():.3f}, 95% credible interval {100*lo:.1f}% to {100*hi:.1f}% (Wilson gave 67.2% to 96.9%)")
print(f"probability the true rate is above 80%: {1 - post.cdf(0.8):.2f}")
# 2. A gap: the stratified estimate +11.9 with standard error 3.2, and two priors
gap, se = 11.9, 3.2
for name, prior_sd in (("flat prior", np.inf), ("sceptical prior, centred on 0 with sd 5", 5.0)):
prec = 1 / se**2 + (0 if np.isinf(prior_sd) else 1 / prior_sd**2)
mean, sd = (gap / se**2) / prec, np.sqrt(1 / prec)
band = stats.norm(mean, sd).cdf(5) - stats.norm(mean, sd).cdf(-5)
print(f"{name}: posterior {mean:+.1f} ± {1.96*sd:.1f}, P(gap > 0) = {1 - stats.norm(mean, sd).cdf(0):.3f}, P(inside ±5) = {band:.2f}")
# 3. A Bayes factor for the rate: "the setting does nothing, p = 0.5" against "p could be anything"
def log_beta(a, b): return lgamma(a) + lgamma(b) - lgamma(a + b)
k, n = 16, 18
log_m1 = log_beta(k + 1, n - k + 1) - log_beta(1, 1) # marginal likelihood under a flat prior on p
log_m0 = k * np.log(0.5) + (n - k) * np.log(0.5) # likelihood at p = 0.5 exactly
print(f"Bayes factor for 'p is not 0.5' vs 'p = 0.5', 16 of 18: {exp(log_m1 - log_m0):.0f} to 1")
k, n = 11, 18
print(f"the same for 11 of 18: {exp(log_beta(k+1, n-k+1) - log_beta(1,1) - n*np.log(0.5)):.1f} to 1")
which prints:
16 of 18, flat prior: posterior mean 0.850, 95% credible interval 66.9% to 96.6% (Wilson gave 67.2% to 96.9%)
probability the true rate is above 80%: 0.76
flat prior: posterior +11.9 ± 6.3, P(gap > 0) = 1.000, P(inside ±5) = 0.02
sceptical prior, centred on 0 with sd 5: posterior +8.4 ± 5.3, P(gap > 0) = 0.999, P(inside ±5) = 0.10
Bayes factor for 'p is not 0.5' vs 'p = 0.5', 16 of 18: 90 to 1
the same for 11 of 18: 0.4 to 1
Where this leaves us
The Bayesian framework answers the question the rest of the series could not, the probability of the hypothesis given the data, and charges for it: a prior, written down, before the data. With a flat prior its intervals agree with the classical ones almost to the decimal, so the two schools differ in what the numbers mean more than in what they are. With an informative prior it shrinks small-study estimates toward what was expected, which is usually wise and always has to be declared. The Bayes factor compares stories and can favour the null, which nothing else in the series could do. And the region of practical equivalence is Part 8's bound with a probability attached.
That is the end of the series, and the end is the same as the beginning. A number from an experiment is a sample from a spread. Every tool here, from the median to the posterior, is a way of being honest about the spread, and every design rule is a way of making sure the spread is the only thing between you and the truth. Whichever school you use, the discipline is the same: decide before you look, count everything, correct for how many questions you asked, and report the interval, not just the verdict.
The glossary: every term the series uses, one line each, with the part that explains it.