Docteur Bet

Methodology

The player model

Dixon-Coles does not read line-ups. A second model estimates what each player is worth on a given date, and what his absence actually costs his side.

Why a second reading

The league model describes a side with two numbers, estimated from its results. That is enough to replay a season, and not enough to read a single fixture: it does not know who plays. A side deprived of its best winger keeps its attacking strength until results say otherwise — several weeks later.

The player model answers the question the other one cannot ask: what is this side worth tonight, with these eleven?

What it measures

It assigns each player a level across a set of actions in the game — what he produces going forward, what he breaks up, what he secures — separately from the side he plays in. An excellent midfielder in a weak team and an average midfielder in a strong one produce the same raw statistics; separating the two is precisely the point of the model.

Those levels are then aggregated across the announced eleven and converted into attacking and defensive intensities comparable to those of the league model.

What it allows: removing a player

This is the use that separates this model from the first. The eleven can be rebuilt without a player and he can be replaced by whoever would actually play in his place — not a theoretical deputy, but the stand-in that recent line-ups point to, and who is neither injured nor suspended — then the full score distribution recomputed.

The output is therefore not "this player matters" but a measured gap: so many points of win probability, so many expected goals, a quantified shift in the law of scores.

How we verify it

On fixtures the model never saw while being fitted, comparing its probabilities to what actually happened. It is the same criterion as everywhere else on this site: the likelihood of what occurred, not how convincing a prediction looks after the fact.

The model must also beat a witness. A gap that does not survive a change of test window is not a result: it is noise, and we treat it as such.

What it does not know

It starts from the announced line-ups. An hour before kick-off it has nothing more to say than the league model.

It reads neither tactics, nor the game plan, nor a player's form on the day. A player shifted into a position he never occupies is badly assessed, because he is assessed from what he produced elsewhere.

And its coverage follows our data. In an international fixture, some of the players have no usable history in our leagues: we then publish nothing rather than filling the gap with an average.

What we do not detail

The list of actions used, how individual levels are estimated, and how they recombine into a team strength all stay internal. That is not coyness: it is the work that sets our products apart, and writing it down would make it reproducible.

What remains verifiable is what matters to a reader: the evaluation method, the limits above, and the track record.

The league model, by contrast, is described in full — it was published in 1997 and we invented none of it. Read the full methodology.

Going further

We have a model that lifts these limits

The model published here is open, documented and deliberately simple — it has to be reproducible. The one we use internally goes further: it lifts the independence assumption between the two sides instead of correcting it on four scorelines, it takes market odds as input, and it is trained market by market. It is not published here, but it is available.