WC26 modelDashboard
Final · Spain won · 104 matches
Glossary
Three ideas do most of the work on this site — here they are, drawn out

The review talks in expected goals, an xG-based win expectancy called xWDL, and a per-player quality rating. None is complicated once you see it once. The three cards below each show the idea with a worked example; the statistical terms further down are the standard machinery, defined briefly with a link to read more.

Companion to the tournament review and the method page · the model-specific terms carry a worked example; the statistical terms link out to a fuller reference
Model term · xG
Expected goals: how good a chance was, before anyone shot

Every shot is graded on how often a chance like it — the distance, the angle, the body part, the pressure — gets scored, from long historical data. A tap-in is worth most of a goal; a hopeful long shot, almost nothing. Add a team's chances up and you get the goals it ought to have scored, which is often a truer read of a match than the scoreline the finishing produced.

Penalty
0.79
Close-range header
0.35
Edge-of-box shot
0.09
Speculative long shot
0.03

xG per chance, on a scale where a certain goal = 1.00.

Illustrative single-chance xG values · FotMob supplies this site's xG
Model term · xWDL
xWDL: turning a match's chances into a deserved win, draw or loss

A scoreline is one noisy sample; the chances a match produced are a much bigger one. xWDL feeds each side's expected goals through the same score model the forecasts use and reads off how often that balance of chances ends in a home win, a draw, or an away win. It is the answer to “who deservedit?” — the yardstick the review scores both the model and the market against.

Home 2.4 xGvsAway 0.6 xG
76·16·8
Home winDrawAway win

A side that out-creates its opponent 2.4 xG to 0.6 “deserves” the win roughly three times in four — but still draws or loses one time in four. That gap is why a single result proves little.

Worked example · post-match xG for each side → win/draw/loss probability
Model term · player quality
Player quality: one rating that puts a keeper and a striker on the same ruler

The model rates every player from their club output — shots created and taken, defensive actions, minutes, all age-adjusted — and expresses it as a percentile against everyone else, so a 94th-percentile full-back and a 94th-percentile goalkeeper mean the same thing. Team strength is these ratings rolled up over the expected starting eleven; that is what lets the lineup, not the badge, decide the forecast.

Squad playerPedro Porro (94th)
0th50th — median player100th
Illustrative percentiles within the tournament player pool · see the method page for the full formula
Statistical terms
The standard machinery, briefly
z-scoreHow many standard deviations a value sits above or below the average. A striker at +2 is well above a typical striker; −1 is below. It puts goalkeepers and forwards on one comparable scale. Wikipedia ↗
EloA points rating, born in chess and now standard for national teams: you gain points beating strong opponents and lose them slipping up against weak ones. The gap between two ratings maps directly to a win probability. Wikipedia ↗
Poisson distributionThe textbook model for counting independent, fairly rare events in a fixed window — here, goals in a match. Give it an expected count and it returns the probability of 0, 1, 2, 3 … goals. Wikipedia ↗
λ (lambda)The single input to a Poisson distribution: a team's expected number of goals in the match. A bigger favourite gets a higher λ.
Dixon–ColesA standard correction to the basic two-Poisson football model (Dixon & Coles, 1997). It nudges the probabilities of the lowest scores (0-0, 1-0, 1-1), which independent Poisson gets slightly wrong because level teams play more cautiously.
Monte Carlo simulationEstimating an outcome by repeated random trial: play the tournament thousands of times, each match decided by a draw from its score distribution, then count how often each team advances. Wikipedia ↗
log-lossA scoring rule for probability forecasts. It rewards putting confidence on what actually happened and punishes confident-but-wrong calls. Lower is better; a coin-flip forecaster scores about 1.10 across three outcomes. Wikipedia ↗
Brier scoreA second forecast-scoring rule: the squared distance between the probabilities you assigned and what actually happened. Like log-loss, lower is better — but it punishes a confident miss less harshly, so the two are shown side by side. Wikipedia ↗
t-testA significance test for whether two sets of paired numbers differ by more than chance — here, the model's per-match log-loss against the market's. The p-value is the chance you'd see a gap this large if the two were really equal. Wikipedia ↗
vig / vig-free consensusVig (vigorish, or “juice”) is the bookmaker's built-in margin: turn a book's odds into probabilities and they sum to more than 100% — that excess is the house's cut. Strip it out and rescale to a clean 100%, take the median across books, and you get the vig-free consensus — the market line this site compares the model against. Wikipedia ↗

More on how these fit together is on the method page; the numbers they produce are in the tournament review.

Established methods — follow a link for the full treatment