Docteur Bet

Methodology

How these probabilities are computed

A model published in 1997, a few thousand matches, and fifty thousand replayed seasons. Nothing esoteric — and all of it verifiable, including what does not work.

Our approach

We model football as we would any statistical phenomenon: from data, with models whose assumptions are known and checkable — not from an intuition dressed up as a percentage.

Two consequences. Our models rest on established, published methods, not on something invented for the occasion. And when a model has limits, we write them down: a good model is one whose blind spots are known.

The rest of this page walks through the construction of the league model, from the most naive model up to the one we publish. The other two, both proprietary, have their own pages: they say what they set out to do, how we verify them and where they fail — not how they are built.

Starting with the simplest thing

The most naive model looks at nothing: it takes the average number of goals scored at home and away across the league and applies it to every fixture. In Ligue 1 that gives about 1.9 goals for the home side and 1.7 for the visitor, whatever the fixture.

It is a bad model, but it asks the right question: from two expected goal numbers, you can already compute the probability of every scoreline, and therefore of a win, a draw, or over 2.5 goals. All the work that follows consists of getting better numbers.

Giving each team its own strengths

In 1982 Michael Maher proposed describing each team with two numbers rather than one: an attacking strength and a defensive strength. A single parameter would not distinguish a side that wins 3-2 from one that wins 1-0.

λ = attack(home) × defence(away) × home advantage μ = attack(away) × defence(home)

An attack of 1.60 means the team scores 60% more than the league average against an average defence. A defence reads the other way round: it is a multiplier of goals conceded, so the lower the better. Home advantage is common to the whole league, and comes out at about ×1.23 in our fits.

These strengths cannot be read off the table: they are estimated over thousands of matches, and adjusted for the quality of the opponents faced. A side that started well against promoted teams does not earn the same strength as one that did as much against the top of the table.

What the Poisson distribution cannot do

Maher assumes the two sides score independently of one another, following two Poisson distributions. That is convenient, and it is wrong on one specific point, which Dixon and Coles measured in 1997.

Tight scorelines are more frequent than independence predicts. There are more 0-0s, 1-0s, 0-1s and 1-1s in reality than in the model. That is no accident: at 0-0 in the 80th minute, the two sides play differently from how they would at 3-1. The score influences the game, so goals are not independent.

Forgetting the past, but slowly

Dixon and Coles's second idea is that a match from 2019 does not say as much as one from last week. Each match receives a weight that decays over time. A match one half-life old counts for half.

That half-life is not chosen by intuition. It is calibrated league by league: we replay each past season week by week, predict the next matchday under several settings, and keep the one that was least wrong. The result varies enormously, and that variation is itself information — where the half-life is short, hierarchies move fast.

LeagueHalf-lifeρMatchesEffective sampleLast computed
Admiral Bundesliga365 j—955426September 21, 2026
Allsvenskan365 j—1,756675September 27, 2026
Bundesliga (Germany)365 j-0.0862,139866September 27, 2026
Champions League— j—14418October 5, 2026
Championship (England)240 j—3,8471,040September 27, 2026
Ekstraklasa1095 j—2,0261,678September 27, 2026
Eliteserien180 j—1,710337September 27, 2026
Eredivisie (Netherlands)365 j—2,063868September 27, 2026
La Liga (Spain)365 j-0.0122,6841,076September 27, 2026
La Liga 2 (Spain)240 j—3,199824October 5, 2026
Liga Portugal (Portugal)1095 j0.0361,9871,644September 27, 2026
Ligue 1 (France)730 j-0.0502,3331,589September 27, 2026
Premier League (England)180 j-0.1102,652530September 27, 2026
Premiership (Scotland)730 j—1,5041,038September 27, 2026
Pro League (Belgium)365 j—2,093834September 21, 2026
Serie A (Italy)240 j0.0112,674699September 27, 2026
Serie B (Italy)120 j—2,681344September 27, 2026
Super League (Switzerland)240 j—1,404404September 27, 2026
Süper Lig (Turkey)120 j—2,367276September 27, 2026
Superliga— j—1,2451,245September 27, 2026

The effective sample is the number of matches it would take, all of equal weight, to carry as much information as the weighted set. In the Premier League, 2,652 matches are worth only 530: that is the price of a short memory, and it is also why probabilities there are less clear-cut.

The case of promoted sides

A newly promoted team has no history in its new league. Estimating a strength from three matches would produce nonsense: three wins would make it a title contender.

Its estimate is therefore shrunk towards the average level of a promoted side, measured over 149 promotions observed since 2020 across ten European leagues. On average, a promoted side attacks at ×0.82 and concedes at ×1.34 relative to the rest of its new league. The more it plays, the less that anchor weighs. Promoted sides are flagged as such in the tables.

Ten leagues, not six: we estimate this level wherever we hold data, including where we publish nothing. A promoted Norwegian club informs none of our pages, but it informs the Norwegian prior — and that is the only way to obtain one.

Comparing two leagues: our adaptation of Dixon-Coles

Dixon-Coles was written for a single league, and there is a reason for that. At every fit, the attack strengths are recentred: their mean is brought back to zero in logarithm. That is what makes the model identifiable — without such an anchor, multiplying every attack by some number and dividing every defence by the same one would leave the intensities unchanged.

The consequence is that each index reads relative to its own league. Porto's reads “relative to Portugal”, Manchester City's “relative to England”. Setting the two side by side and drawing a conclusion is comparing two temperatures where one is in Celsius and the other in Fahrenheit.

So we added two terms per league to the model: δ for attack, ε for defence. A fixture between team i from league L and team j from league M is then written:

λ_home = α_i · β_j · exp(δ_L + ε_M + g)
λ_away = α_j · β_i · exp(δ_M + ε_L)

g is the home advantage specific to European competition, estimated on these fixtures rather than carried over from the domestic leagues: the travel, often longer, adds to it. The Dixon-Coles τ correction is kept, with its own ρ.

Do these terms earn their place? The question is settled out of sample, on a full season never seen during fitting — 138 fixtures. The model with offsets scores a log loss of 0.933 and an RPS of 0.204; without the offsets, 1.007 and 0.230; betting on historical frequencies alone, 1.023 and 0.237.

The second figure is worth pausing on: comparing two domestic indices with no correction barely beats informed guesswork. That is the measure of what the shortcut is worth, and of what the offset avoids.

These probabilities are published only where both clubs play in one of our six leagues — a minority of European fixtures. Elsewhere we do not measure the opponent's level, and we do not invent it.

Replaying the season

The remaining fixture list is replayed 50,000 times. Each match is drawn at random according to the model's probabilities, points accumulate, and the final table is settled by the league's real rules — points, then goal difference, then goals scored.

A title probability of 60% means exactly this: across fifty thousand simulated seasons, that team finishes first in thirty thousand of them. And is wrong four times in ten.

The uncertainty on the strengths themselves

Strengths are estimated, and therefore uncertain. Treating them as known would make the probabilities too clear-cut: a favourite would cross the season without its advantage ever being questioned.

The model is therefore refitted 60 times on resamples of past matches, and each simulated season draws at random which of those fits it uses. The published probabilities thus carry the uncertainty on the strengths, not only the randomness of the matches.

When the figures move

The model is only refitted if matches have been played since the last computation. Without a new result it would return exactly the same figures under a more recent date — which would suggest a movement that did not happen. In practice: one computation a week, two in midweek-fixture weeks, none during international breaks.

Every page carries the date of its computation. If the nightly chain fails, the site keeps serving the last valid computation — and that date is the only thing that tells you so.

What this model does not know

It knows results and nothing else. Not injuries, not transfers, not a change of manager, not a match played with ten men. A side that loses its striker in January will keep its attacking strength until results say otherwise.

And the Dixon-Coles correction, while it settles the question of tight scorelines, does not lift the independence assumption: it corrects it on four cells, not everywhere. That is an acknowledged limit of the model, and it has been known since 1997.

What it was not built for

Dixon-Coles explains a result well from the overall strength of two sides. It was not built for a different task: finding precisely where an operator has mispriced. An odds market is already largely efficient — beating it does not call for a good result model, but for a model calibrated for that particular contest.

That is why, and not out of any taste for secrecy, we built two further models, distinct from this one.

Beyond Dixon-Coles

Both models stay proprietary. Unlike Dixon-Coles, we do not detail their construction here: that is what carries the value of our products. What does remain verifiable is their performance over time — the history of the results is available, losses included, not only the wins that get highlighted.

Each has its own page, stating what it sets out to do, how we verify it, and where it fails.

  • The player model — what a player is worth on a given date, and what his absence costs his side. It is the one that reads line-ups, where Dixon-Coles cannot.
  • The application's model — the one that looks for where a price is wrong, and why that task calls for a different construction from a result model.

The limits we own

  • A probability is not a prediction. A 61% chance of the title means that, across a hundred similar seasons, you would expect it to happen in about sixty-one of them — not that it is written in advance.
  • The model always starts from a history. An unprecedented event — a major injury, a radical change of manager — takes time to be absorbed, since it is only seen through results.
  • No model replaces expertise. Neither a coaching staff's nor a doctor's. Our tools serve to support a decision, not to take it.

A question about the method?

If you work with us — media, club, researcher — and you are wondering what our models can and cannot cover for your particular case, write to us rather than guess. An honest answer about what is not measurable saves you more time than a promise.

Question about the method

Describe your use case and the question you need settled. We answer on what is measurable — and on what is not.

Going further

We have a model that lifts these limits

The model published here is open, documented and deliberately simple — it has to be reproducible. The one we use internally goes further: it lifts the independence assumption between the two sides instead of correcting it on four scorelines, it takes market odds as input, and it is trained market by market. It is not published here, but it is available.