AGI odds vs evidence: the market prices announcements, we grade capability
Issue #1 — 2026-08-08
| Prediction market | Evidence layer (this site) | |
|---|---|---|
| Instrument | "OpenAI announces it has achieved AGI before 2027?" (Polymarket) | AGI-2027 verdict + Thesis Tracker |
| Question actually asked | Will OpenAI announce it has achieved AGI before 2027? | Do models do the work of an AI researcher by end-2027 (Aschenbrenner's own bar)? |
| Reading | 92% No / 8% Yes — machine-verified snapshot, 2026-08-24 04:12:59 UTC · $95,951 volume · market open | Verdict Open; thesis at 62.5/100 (not a probability — a mean of 8 graded verdicts) |
| Resolves on | An announcement (with or without the capability) | Pre-registered capability criteria, deadline 2028-01-01 (resolution page) |
Why the two numbers must not be conflated
A lab could announce "AGI" for a system that cannot do autonomous research — the market contract could pay out while the capability verdict stays Open or resolves Wrong. The reverse is also possible: genuine research autonomy demonstrated without anyone using the word AGI. Markets are excellent at pricing events and incentives; a graded evidence ledger is what tells you whether the substance happened. Traders on these markets are, in effect, this site's target reader: the resolution criteria they need are the ones we pre-registered.
Review log
| Date | Market side | Evidence side | Issue published? |
|---|---|---|---|
| 2026-08-24 | First machine-verified reading: 92% No / 8% Yes, $95,951 volume. Issue #1 recorded «≈11% Yes» from a hand-copied snapshot with no exact timestamp, so the ~2pt drift toward No is indicative, not a measured move — the baseline was not precise enough to claim one. | No verdict change; tracker held at 62.5/100 | No — a drift inside the noise of a vague prior baseline is not news. What changed is the method: odds are now fetched weekly on a runner and carry an ISO timestamp, so the next comparison will be measured rather than indicative. |
| 2026-08-10 | No further movement found beyond the issue #1 snapshot (~11% Yes) | No verdict change; tracker held at 62.5/100 | No — nothing moved, so there was nothing to say |
| 2026-08-08 | Issue #1 snapshot recorded | Verdict Open; tracker 62.5/100 | Yes — issue #1 |
This series updates when something moves, not on a calendar. A weekly slot that must be filled produces filler; an evidence layer that publishes noise to look busy is worth less than one that publishes nothing and says so. Reviews happen weekly and are logged above either way — including the weeks the answer is "neither side moved". When a verdict actually flips, subscribers hear the same day, before they would read it in the odds.
Odds snapshots are third-party market data quoted with their as-of date; this site takes no positions and this is not trading advice.
Frequently asked questions
No. 62.5/100 is the mean of 8 graded verdict weights — an auditable evidence composite, not a forecast. Prediction-market prices are crowd probabilities of specific contract wordings. The two answer different questions.
Because the gap is informative. An announcement-priced market and a capability-graded ledger diverging tells you the crowd expects labeling to run ahead of substance (or behind it). Traders need pre-registered resolution criteria; that is exactly what this site publishes.
Each issue quotes a dated snapshot with a link to the live market — never a 'current' price. Since 2026-08-24 the odds are fetched automatically once a week and carry an exact UTC timestamp and the market's traded volume, so a reading can no longer drift into being quoted as if it were live. The series is reviewed weekly but only publishes a new issue when one side actually moves; every review, including the quiet ones, is logged on the page.
The live scorecard updates as models ship and verdicts change.
View the live scorecard →