Scopes

The scorekeeperfor prediction markets

Everyone quotes the odds.
We keep score.

So you know how much to trust a number before you cite it or trade on it: the odds everyone quotes become something you can calibrate, not something you take on faith.

Today's widest live divergence: Jaguars v Bengals. Kalshi says 56.5%, our independent benchmark says 58.8%. Yours to call:

Every figure is fingerprinted and timestamped to the Bitcoin blockchain. Verify the record →

Publisher of the Scopes Calibration Index. Independent benchmarks, verification and reference data for prediction markets, administered under published rules.

Research & information only, not investment advice.

Collecting since July 2026 · 227,524 snapshots and counting.

451 divergences resolved. Who was right?

When market price and Scopes Fair Value disagreed:

Market right · 273Fair value right · 178
Statistically significant at the 95% level (n = 451) What does this split mean? →Explore the record →

Most resolved flags so far are energy. Other desks are live and accumulating. Full scorecard by desk →

conflicts · versioned · append-only · archived · verifiable · last ingest 4h ago

The Scopes Calibration Index

How well each prediction market's prices have actually been calibrated against outcomes, by venue, by desk and by horizon, measured under rules published in advance from a timestamped record. Published monthly by Scopes Reference Data, LLC. The Scopes Fair Value is not scored by the index.

Methodologyv0.9 consultation draft, published October 1, 2026
First valuesNovember 5, 2026
SeriesBy venue, by desk, composite

Read the methodology · Licensing and use

The Read-off

Which frontier model reads the event best

For every open event there are two estimates of the same probability: the price a prediction market implies, and an independent fair value. We ask each frontier model the same question, before the event resolves: which one is better calibrated? When it settles, we record whose read matched the estimate that proved closer. The models grade the estimates; they do not forecast the outcome.

Tested on events that have not happened yet

Each read is recorded and locked before the event resolves, so a model cannot have trained on the answer. A live test, not a static benchmark.

Reproducible line by line

The exact prompt, the inputs each model saw, and its raw samples are published. Anyone can re-run a read and check the score by hand.

Scored mechanically

When the event settles, the estimate closer to the outcome is the better calibrated one. Append-only, no editorializing, and a model may abstain when the question is genuinely undecidable.

ModelMatched better est.DecidedBy choice
anthropic claude-opus-569.6% (16/23)23mkt 16/22 · fair 0/1
google gemini-3.8-flashearly14mkt 9/12 · fair 0/2
xai grok-4.6early1mkt 0/1
openai gpt-6-astraearly0—
deepseek deepseek-v4-pro-0813early0—

Below 20 decided reads a model shows "early", not a rating. Abstentions are recorded but never scored. Scoring begins as events resolve.

Point in time, with proof

Scopes Reference Reports

Name a contract and a moment. Get the market price, the Scopes Fair Value and the divergence as published, with the Merkle inclusion proof and its Bitcoin timestamp: a point-in-time reference mark anyone can verify, not take on faith.

Request a report → · How it works →

Read it · Rung 1 of the data ladder

Research access from $49 a month. Institutional licensing and benchmark administration for funds, administrators and data platforms.

Research access → · Institutions →

The Week in Review: a free weekly recap of the biggest market-vs-fair-value moves across every desk. Read the latest →