Collecting since July 2026 · 222,293 snapshots and counting.
451 divergences resolved. Who was right?
When market price and Scopes Fair Value disagreed:
Most resolved flags so far are energy. Other desks are live and accumulating. Full scorecard by desk →
conflicts · versioned · append-only · archived · verifiable · last ingest 4h ago
Where to start
One record. Three ways in. The public score stays free to cite.
How it works
- 1
Record
the market price
- 2
Pair
it with an independent Scopes Fair Value
- 3
Resolve
track the gap and publish who was right
Every reading is timestamped the moment we record it, and the record is append-only, so the score is one you can cite. How the method works →
Filled = live today; outline = coming. One method across every desk.
The Read-off
Which frontier model reads the event best
For every open event there are two estimates of the same probability: the price a prediction market implies, and an independent fair value. We ask each frontier model the same question, before the event resolves: which one is better calibrated? When it settles, we record whose read matched the estimate that proved closer. The models grade the estimates; they do not forecast the outcome.
Tested on events that have not happened yet
Each read is recorded and locked before the event resolves, so a model cannot have trained on the answer. A live test, not a static benchmark.
Reproducible line by line
The exact prompt, the inputs each model saw, and its raw samples are published. Anyone can re-run a read and check the score by hand.
Scored mechanically
When the event settles, the estimate closer to the outcome is the better calibrated one. Append-only, no editorializing, and a model may abstain when the question is genuinely undecidable.
| Model | Matched better est. | Decided | By choice |
|---|---|---|---|
| anthropic claude-opus-5 | 72.7% (16/22) | 22 | mkt 16/22 |
| google gemini-3.8-flash | early | 13 | mkt 9/12 · fair 0/1 |
| xai grok-4.6 | early | 1 | mkt 0/1 |
| openai gpt-6-astra | early | 0 | — |
| deepseek deepseek-v4-pro-0813 | early | 0 | — |
Below 20 decided reads a model shows "early", not a rating. Abstentions are recorded but never scored. Scoring begins as events resolve.
Used when you need to cite a measured answer, read the odds without trading them, or license the record. How institutions use Scopes → · Reference →
The Week in Review
A free weekly recap: the biggest gaps between the market and the Scopes Fair Value across every desk: what resolved, and what's ahead. No spam, unsubscribe anytime.
Read it · Rung 1 of the data ladder
Research access from $49 a month. Institutional licensing and benchmark administration for funds, administrators and data platforms.