Luck or skill 965,470 team-matches measured
The question behind every record on this site

Is that a feat, or a coincidence?

In Moneyball, Bill James works out whether Jack Chesbro's 41–12 season was the pitcher or the team behind him. Chesbro's side won 52% of its games, so James asks how likely 41–12 is from an ordinary pitcher on that team. The number separates what happened from what the circumstances already explain — and almost every figure on this site shows the first half and none of the second.

The longest run 50 Manchester City, no draw in 50 matches
The model expected 9.1 draws across those same 50 fixtures
Chance of that run 1 in 26,400 asked of this run alone — the overstated figure
Runs that long, expected anywhere 0.46 across every run in the database. There is 1.

Those last two tiles are the whole page. One in 26,400 sounds like a miracle. But the run was not picked at random — it was picked for being the longest, out of every run every team has ever had here. Ask instead how many runs of 50 ought to turn up somewhere in a database this size, and the answer is about 0.46. A record is not the same thing as a miracle.

The records this is about are on the all-time no-draw leaderboard, and every team's own figure is on its team page. The run in those tiles is Manchester City's 50, match by match.

Two distributions, one dataset Poisson-binomial against binomial

James had one probability: his pitcher's team won 52% of its games, every game. Under a constant chance, the number of successes in n independent tries is the binomial distribution, and his factorial formula is exactly its probability mass function. Our matches do not share a chance. Liverpool at home to a promoted side and Liverpool away at Arsenal are not the same trial, and the five-feature logistic draw model, recalibrated by its isotonic map gives each of them its own figure. A sum of independent coin flips with different chances is not binomial — it is the Poisson-binomial distribution, and that is what every number on this page uses.

Reading? Poisson-binomial? Binomial at the average chance? The binomial's error?
Spread over a career, summed across 1,595 teams ? 174,080 176,237 +1.24%
Chance of a record run happening at all, averaged over 664 of them ? 1 in 403 1 in 388 1.04× too likely

Measured, not assumed — and the honest reading is that the two rows disagree about how much this matters. Over a whole career the correction is worth about one percent, because draw chances live in a narrow band between roughly 7% and 40% and p(1−p) is nearly flat across it. Anyone claiming this page needs the Poisson-binomial because the binomial's spread is wrong would be overselling it. Over a single run it is worth far more, and always in the same direction: the binomial makes a record look more ordinary than it was.

The real reason for it is neither. The expectation is only available per match. Comparing a team against the sum of its own fixtures' chances — what these matches deserved — is the entire point, and once the average is per-match, the matching distribution is this one. Putting a binomial beside a per-match average is mixing two models.

The longest runs, and what each was worth 664 runs of 20+ measured · the full leaderboard ›

For each run: how many draws the model expected across those exact fixtures, the chance that none of them landed, and what a binomial at the run's own average chance would have said instead. Every figure in this table is asked of a run that was selected for being long, which overstates it. The table underneath is the version that is not.

Team Run? Between Draws expected? Chance of none? A binomial would say?
Manchester City 50 Dec 2020 – Sep 2021 9.1 1 in 26,400 1 in 23,700
Wales 46 Feb 2008 – Mar 2013 10.1 1 in 104,200 1 in 92,600
East Stirlingshire 42 Sep 2003 – Sep 2004 5.3 1 in 315 1 in 295
Borussia Dortmund 39 Apr 2021 – Dec 2021 7.4 1 in 3,900 1 in 3,500
SL Benfica 37 Jul 2010 – Mar 2011 7.3 1 in 3,500 1 in 3,300
Cambodia 37 Jun 2015 – Mar 2018 6.9 1 in 2,400 1 in 2,100
Gaziantep FK 36 Jan 2023 – Dec 2023 7.9 1 in 8,000 1 in 7,300
PSV Eindhoven 36 Dec 2013 – Oct 2014 7.6 1 in 5,300 1 in 5,100
Arsenal 35 Feb 2022 – Oct 2022 8.3 1 in 14,200 1 in 13,600
Hebei 35 Jun 2022 – Dec 2022 7.3 1 in 4,100 1 in 3,600
RKC Waalwijk 35 Aug 2009 – May 2010 7.3 1 in 4,000 1 in 3,700
Celtic 35 Aug 2001 – Jan 2002 6.3 1 in 1,100 1 in 1,000
Spain 35 Jun 2008 – Jul 2010 5.9 1 in 729 1 in 647
Kashima Antlers 34 Jul 2013 – May 2014 8.0 1 in 10,400 1 in 9,700
Manchester United 34 Apr 2023 – Nov 2023 8.1 1 in 10,300 1 in 10,000
How often should a run that long turn up anywhere? the honest question

Asking how unlikely this run was, after picking it for being the longest, is asking the wrong question. The right one is how many runs of at least that length ought to exist somewhere in the database at all. That is worked out by walking every position a run could start from — every match after a draw, and every match after a gap in coverage — and adding up the chance that a run starts there and survives.

A run of? On record? The model expects? One flat draw rate expects? Observed ÷ expected?
5+ 52,415 53,120 55,599 0.99
10+ 11,094 11,506 13,484 0.96
15+ 2,624 2,723 3,414 0.96
20+ 664 676 872 0.98
25+ 183 175 223 1.05
30+ 52 48 57 1.09
35+ 13 14 15 0.96
40+ 3 4.18 3.82 0.72
45+ 2 1.36 0.99 1.47
50+ 1 0.46 0.26 2.16

Expectations add whether or not the matches are independent, so this column does not lean on that assumption at all — only the spreads and the tails above it do.

Watch which way the flat column goes, because it reverses. It expects more short runs than the model and fewer long ones. Two different effects are at work and they point opposite ways. Inside a single run, averaging the chances first always makes the run look more likely — that is the concavity from the table above. Across the whole database, giving every team the same chance makes a long run far less likely, because a long run is overwhelmingly produced by the teams that hardly ever draw, and a flat rate deletes them. The second effect is the bigger one at the top of this table: teams are not alike, and a league in which they were would almost never produce a fifty.

Teams the model did not expect since 2015 · 100+ matches
Name? Matches? Drew? Model expected? League average expected? Standard errors ? Adjusted ? Reading?
Bristol Rovers 592 123 160.8 156.1 −3.50 0.185 far fewer than expected — and still worth checking the model before the team
South Africa 146 56 36.9 37.4 +3.67 0.185 far more than expected — and still worth checking the model before the team
Sheffield United 582 119 155.3 153.1 −3.42 0.185 far fewer than expected — and still worth checking the model before the team
Deportivo La Coruña 345 121 92.5 93.7 +3.49 0.205 far more than expected — and still worth checking the model before the team
Kifisia 103 37 22.7 23.1 +3.41 0.226 far more than expected — and still worth checking the model before the team
Åsane Fotball 356 98 73.3 73.3 +3.26 0.226 far more than expected — and still worth checking the model before the team
PAS Lamia 311 99 74.7 72.1 +3.25 0.226 far more than expected — and still worth checking the model before the team
Fredrikstad FK 309 90 66.4 64.7 +3.29 0.226 far more than expected — and still worth checking the model before the team
Plymouth Argyle 580 125 157.9 154.4 −3.07 0.249 far fewer than expected — and still worth checking the model before the team
Lahti 357 114 88.8 85.5 +3.11 0.275 far more than expected — and still worth checking the model before the team
Cameroon 128 48 32.6 31.6 +3.15 0.289 far more than expected — and still worth checking the model before the team
ES Troyes AC 448 101 128.7 131.2 −2.90 0.307 clearly fewer than expected, though one subject in twenty looks like this by chance

Ordered by how far each sits from what the model expected of its own fixtures. The adjusted column is what makes an ordered table defensible at all: test this many subjects and a handful will clear any threshold on noise alone. Read it as "if I call every row above this one interesting, this is roughly the share that is not". Note that the top of this table has adjusted values nowhere near small — which is the finding, not a disappointment.

Referees the model did not expect all matches · 100+ matches
Name? Matches? Drew? Model expected? League average expected? Standard errors ? Adjusted ? Reading?
C Scott 142 51 35.0 34.7 +3.13 0.326 far more than expected — and still worth checking the model before the team
J Linington 485 107 135.3 132.3 −2.87 0.326 clearly fewer than expected, though one subject in twenty looks like this by chance
T Harrington 365 76 100.0 98.1 −2.83 0.326 clearly fewer than expected, though one subject in twenty looks like this by chance
J Moss 511 105 132.5 134.8 −2.78 0.326 clearly fewer than expected, though one subject in twenty looks like this by chance
Grant Hegley 268 56 75.3 73.0 −2.63 0.326 clearly fewer than expected, though one subject in twenty looks like this by chance
Brian Curson 184 68 51.5 50.2 +2.72 0.326 clearly more than expected, though one subject in twenty looks like this by chance
P Wright 227 81 63.0 60.5 +2.68 0.326 clearly more than expected, though one subject in twenty looks like this by chance
C Hicks 304 64 83.5 81.2 −2.51 0.326 clearly fewer than expected, though one subject in twenty looks like this by chance
Steve Dunn 121 20 31.8 31.5 −2.44 0.326 clearly fewer than expected, though one subject in twenty looks like this by chance
Alan Freeland 123 17 28.7 29.9 −2.51 0.326 clearly fewer than expected, though one subject in twenty looks like this by chance
L Swabey 287 62 79.3 77.3 −2.28 0.554 clearly fewer than expected, though one subject in twenty looks like this by chance
J Busby 327 73 90.7 88.6 −2.19 0.630 clearly fewer than expected, though one subject in twenty looks like this by chance

Ordered by how far each sits from what the model expected of its own fixtures. The adjusted column is what makes an ordered table defensible at all: test this many subjects and a handful will clear any threshold on noise alone. Read it as "if I call every row above this one interesting, this is roughly the share that is not". Note that the top of this table has adjusted values nowhere near small — which is the finding, not a disappointment.

What this does not say read before quoting anything above
Matches are not quite independent.
The distribution assumes they are, given their chances. They are not: the same squad, the same manager, the same injury list, and a side that starts playing for draws keeps doing it. The model's per-match figure absorbs much of that — it is built from recent form — but the leftover correlation makes the true spread a little wider than the one quoted here, so the tails above are, if anything, slightly too small. No fudge factor has been applied to hide that.
A small p-value is not skill.
It says the model did not expect this. For us that is as likely to be a gap in the model as a property of the team, and when the model is wrong that is a finding about the model.
The probabilities are in sample.
Every played match here was part of the five-feature logistic draw model, recalibrated by its isotonic map's own training set, so its chance was fitted already knowing the result. That flatters the model, and every figure built on it inherits the flattery. football:report:draw-holdout is the standing out-of-sample check.
How the tails were computed.
Exactly, by folding one match into the distribution at a time — no factorials, which lose every digit they have long before 700! overflows. Above 1,000 matches a skew-corrected normal approximation is used instead and the row says so; measured against the exact answer it is within 0.006 in the tails. Errors in the exact path accumulate at about 1e-13 over a thousand matches.
Draws are ninety-minute draws.
As everywhere on this site: a shoot-out win stays a draw.
Players and managers are missing.
The striker version of this question — "were those goals him, or the side around him?" — needs to know who was on the pitch, and there are no per-match appearances in this database. Managers have no source at all. Neither is being guessed at.

Probabilities from the five-feature logistic draw model, recalibrated by its isotonic map. Last measured 22 Sep 2026. Descriptive statistics over historical results, not betting advice.