Scopes

The scorekeeperfor prediction markets

Everyone quotes the odds.
We keep score.

So you know how much to trust a number before you cite it or trade on it: the odds everyone quotes become something you can calibrate, not something you take on faith.

Today's widest live divergence: Utah St. v Boise St.. Kalshi says 7.5%, our independent benchmark says 10.9%. Yours to call:

Every figure is fingerprinted and timestamped to the Bitcoin blockchain. Verify the record →

For funds, administrators & researchers: Scopes Reference™: the point-in-time independent record →

Research & information only, not investment advice.

Collecting since July 2026 · 222,293 snapshots and counting.

451 divergences resolved. Who was right?

When market price and Scopes Fair Value disagreed:

Market right · 273Fair value right · 178
Statistically significant at the 95% level (n = 451) What does this split mean? →Explore the record →

Most resolved flags so far are energy. Other desks are live and accumulating. Full scorecard by desk →

conflicts · versioned · append-only · archived · verifiable · last ingest 4h ago

Where to start

One record. Three ways in. The public score stays free to cite.

ScorecardWeek in ReviewDocsBenchmarksInstitutions

How it works

  1. 1

    Record

    the market price

  2. 2

    Pair

    it with an independent Scopes Fair Value

  3. 3

    Resolve

    track the gap and publish who was right

Every reading is timestamped the moment we record it, and the record is append-only, so the score is one you can cite. How the method works →

RatesPollsSportsEnergyStocksCryptoWeatherCatastrophe

Filled = live today; outline = coming. One method across every desk.

The Read-off

Which frontier model reads the event best

For every open event there are two estimates of the same probability: the price a prediction market implies, and an independent fair value. We ask each frontier model the same question, before the event resolves: which one is better calibrated? When it settles, we record whose read matched the estimate that proved closer. The models grade the estimates; they do not forecast the outcome.

Tested on events that have not happened yet

Each read is recorded and locked before the event resolves, so a model cannot have trained on the answer. A live test, not a static benchmark.

Reproducible line by line

The exact prompt, the inputs each model saw, and its raw samples are published. Anyone can re-run a read and check the score by hand.

Scored mechanically

When the event settles, the estimate closer to the outcome is the better calibrated one. Append-only, no editorializing, and a model may abstain when the question is genuinely undecidable.

ModelMatched better est.DecidedBy choice
anthropic claude-opus-572.7% (16/22)22mkt 16/22
google gemini-3.8-flashearly13mkt 9/12 · fair 0/1
xai grok-4.6early1mkt 0/1
openai gpt-6-astraearly0—
deepseek deepseek-v4-pro-0813early0—

Below 20 decided reads a model shows "early", not a rating. Abstentions are recorded but never scored. Scoring begins as events resolve.

Used when you need to cite a measured answer, read the odds without trading them, or license the record. How institutions use Scopes → · Reference →

The Week in Review

A free weekly recap: the biggest gaps between the market and the Scopes Fair Value across every desk: what resolved, and what's ahead. No spam, unsubscribe anytime.

Read it · Rung 1 of the data ladder

Research access from $49 a month. Institutional licensing and benchmark administration for funds, administrators and data platforms.

Research access → · Institutions →