Docteur Bet

Methodology

The application's model

Explaining a result well and finding where a price is wrong are two different tasks. The second one needs a different model, and a different measure of success.

The problem is not the obvious one

Dixon-Coles answers one question well: given the strength of two sides, what is the probability of each score? That is what replaying a league requires, and it is not what the application requires.

The application answers a different question: is this displayed price wrong, and by how much? An odds market already aggregates the information of thousands of participants; it is not easy to beat, and a good result model does not suffice. It takes a model built for that particular confrontation.

What that changes in the construction

Three consequences, statable without going into detail.

The target is the full law of scores, not the 1X2. A model tuned to name the winner is blind to most markets — goal totals, margins, half-times. The entire distribution is modelled, and the markets are derived from it.

Calibration matters more than average accuracy. Announcing 70% when the event happens 75 times in 100 is a mistake that costs money, even when the favourite named is the right one. A model that is rarely wrong but states its certainties badly is unusable here.

Dependence between markets counts. One side's goal count and its opponent's are not independent, and a model that treats them as such produces prices that contradict each other — and therefore false opportunities.

How its performance is judged

On fixtures later than its fitting, never on those used to build it, and with the same criterion as the rest of the site: the likelihood of the score that actually occurred.

It is set against a witness — the Dixon-Coles model described in our methodology, fitted on the same data and evaluated on the same fixtures. A proprietary model that fails to beat one published in 1997 does not deserve to exist, and that is the first thing we check.

What we publish

The complete track record of the application's outputs, losses included. It is the only verification a user can carry out unaided, and it is the one that counts: a success shown without its denominator proves nothing.

What we do not detail

Its construction. It carries the value of our products, and writing it down would be giving it away. We publish neither its family of distributions, nor its variables, nor its settings.

That reserve stops where your interest begins: the evaluation criterion, the witness, the track record and the limits below are public.

The limits

A model better calibrated than an operator on one market is not better on all of them, nor permanently. Prices move, operators adjust, and an edge measured over one period has to be measured again over the next.

None of this guarantees a profit. A probability remains a probability, including when it is correct: a losing run is not proof the model was wrong, and a winning run is not proof of the opposite. Only volume settles it.

And our blind spots are the same as everywhere else on this site: what the data does not contain, the model does not know. The player model fills part of that gap by reading line-ups; it does not fill all of it.

Going further

We have a model that lifts these limits

The model published here is open, documented and deliberately simple — it has to be reproducible. The one we use internally goes further: it lifts the independence assumption between the two sides instead of correcting it on four scorelines, it takes market odds as input, and it is trained market by market. It is not published here, but it is available.