Collecting since July 2026 · 225,371 snapshots and counting.
451 divergences resolved. Who was right?
When market price and Scopes Fair Value disagreed:
Most resolved flags so far are energy. Other desks are live and accumulating. Full scorecard by desk →
conflicts · versioned · append-only · archived · verifiable · last ingest 4h ago
The Scopes Calibration Index
How well each prediction market's prices have actually been calibrated against outcomes, by venue, by desk and by horizon, measured under rules published in advance from a timestamped record. Published monthly by Scopes Reference Data, LLC. The Scopes Fair Value is not scored by the index.
Where to start
One record. Three ways in. The public score stays free to cite.
How it works
- 1
Record
the market price
- 2
Pair
it with an independent Scopes Fair Value
- 3
Resolve
track the gap and publish who was right
Every reading is timestamped the moment we record it, and the record is append-only, so the score is one you can cite. How the method works →
Filled = live today; outline = coming. One method across every desk.
The Read-off
Which frontier model reads the event best
For every open event there are two estimates of the same probability: the price a prediction market implies, and an independent fair value. We ask each frontier model the same question, before the event resolves: which one is better calibrated? When it settles, we record whose read matched the estimate that proved closer. The models grade the estimates; they do not forecast the outcome.
Tested on events that have not happened yet
Each read is recorded and locked before the event resolves, so a model cannot have trained on the answer. A live test, not a static benchmark.
Reproducible line by line
The exact prompt, the inputs each model saw, and its raw samples are published. Anyone can re-run a read and check the score by hand.
Scored mechanically
When the event settles, the estimate closer to the outcome is the better calibrated one. Append-only, no editorializing, and a model may abstain when the question is genuinely undecidable.
| Model | Matched better est. | Decided | By choice |
|---|---|---|---|
| anthropic claude-opus-5 | 69.6% (16/23) | 23 | mkt 16/22 · fair 0/1 |
| google gemini-3.8-flash | early | 14 | mkt 9/12 · fair 0/2 |
| xai grok-4.6 | early | 1 | mkt 0/1 |
| openai gpt-6-astra | early | 0 | — |
| deepseek deepseek-v4-pro-0813 | early | 0 | — |
Below 20 decided reads a model shows "early", not a rating. Abstentions are recorded but never scored. Scoring begins as events resolve.
Used when you need to cite a measured answer, read the odds without trading them, or license the record. How institutions use Scopes → · Reference →
The Week in Review
A free weekly recap: the biggest gaps between the market and the Scopes Fair Value across every desk: what resolved, and what's ahead. No spam, unsubscribe anytime.
Read it · Rung 1 of the data ladder
Research access from $49 a month. Institutional licensing and benchmark administration for funds, administrators and data platforms.