Guide

Forecast Accuracy Scoring

The four measures Hinsley uses to grade a probabilistic forecast, and the arithmetic behind each one.

Forecast scoring

4 parts, 33 minute read

A forecast of "70% chance" is not right or wrong when the day arrives. It is one claim in a long run of claims, and the only honest way to judge it is to score the whole run. This guide explains the four measures Hinsley uses to do that, why each one exists, and exactly how each is calculated.

Scoring matters for a variety of reasons. It gives forecasters a feedback mechanism to help improve their forecasting skill. It tells an organization which of its people and models to trust on which subjects. And it lets an aggregation weight its inputs by demonstrated skill instead of by seniority or volume.

The four measures build on one another, so they are best read in order. "What is a Brier score?" defines forecasting error measurement. "What is a relative Brier score?" makes error comparable between forecasters who did not answer the same questions on the same days. "What is an ordinal Brier score?" handles questions whose answers sit on an ordered scale, where some wrong answers are closer than others. And "What is calibration?" asks a different question altogether: not how much error there was, but whether the forecasted probabilities matched the real-world occurrence rates.

Which measure answers which question

The four are not alternatives. A scorecard in Hinsley shows several of them at once, because each one answers a question the others cannot.

Measure The question it answers Direction
Mean Brier score How much total error did this forecaster produce? Lower is better
Relative Brier score Did this forecaster beat the crowd, accounting for when they showed up? Higher is better
Ordinal Brier score On a question with ordered answers, how far off was the answer? Lower is better
Calibration When this forecaster says 70%, does it happen 70% of the time? Closer to the diagonal is better

A common mistake is to read one of them alone. A low mean Brier score can mean an easy set of questions. A strong relative score with poor calibration means a forecaster who ranks well but whose stated probabilities should not be taken at face value. Good calibration with no relative skill means a forecaster who is honest about uncertainty but adds nothing the crowd did not already know.

Where these numbers appear in Hinsley

Every forecast in Hinsley is stored with the day it was made, and it stands until it is replaced. When a question resolves, Hinsley walks the scoring window one day at a time, scores whatever forecast stood on each day, and rolls the daily scores up into the per-question numbers described in these four pages. Aggregations are scored the same way as people, so an AI forecaster, a weighted crowd aggregation, and an individual analyst all appear on the same leaderboard on the same scale.

Because scoring is per day rather than per question, the timing of a forecast is part of its score. A forecaster who was accurate early is credited for every day they were accurate. That single rule shapes everything else, and the relative Brier score page shows what it does to a leaderboard.

This guide was written with the help of AI, but was reviewed and edited by a human.

Put the method into practice

Ready to measure forecast accuracy properly?