▸HOW V3 WOULD HAVE READ
V3 REPLAY · 1 Apr 2026 – 26 Sep 2026Every 3 hours from 1 Apr 2026: what the markets said about the next 72 hours and the next 7 days, read only from prices that existed at the time, against the 45 frontier launches that followed.
The chart as a table, by week
| Week of | Window | 72h mean | 72h range | 7d read mean | Launches (▲ priced) |
|---|---|---|---|---|---|
| 2026-04-01 | fit | 34% | 25%–49% | 24% | qwen3.6-plus; gemma-4-31b-it; gemma-4-26b-a4b-it |
| 2026-04-08 | fit | 53% | 29%–100% | 68% | none |
| 2026-04-15 | fit | 63% | 24%–98% | 86% | ▲ claude-opus-4.7 |
| 2026-04-22 | fit | 45% | 24%–99% | 30% | ▲ deepseek-v4-flash, deepseek-v4-pro; ▲ gpt-5.5, gpt-5.5-pro; qwen3.6-27b, qwen3.6-max-preview +3 |
| 2026-04-29 | fit | 26% | 24%–34% | 8% | grok-4.3; gpt-chat-latest |
| 2026-05-06 | fit | 35% | 26%–100% | 49% | ▲ gemini-3.1-flash-lite |
| 2026-05-13 | fit | 72% | 29%–100% | 94% | ▲ gemini-3.5-flash |
| 2026-05-20 | fit | 31% | 26%–48% | 13% | grok-build-0.1; qwen3.7-max |
| 2026-05-27 | fit | 40% | 26%–85% | 28% | claude-opus-4.8 |
| 2026-06-03 | fit | 43% | 32%–100% | 30% | qwen3.7-plus; ▲ claude-fable-5 |
| 2026-06-10 | fit | 34% | 31%–38% | 25% | none |
| 2026-06-17 | fit | 52% | 36%–66% | 60% | none |
| 2026-06-24 | fit | 56% | 32%–96% | 52% | ▲ claude-sonnet-5 |
| 2026-07-01 | fit | 50% | 25%–95% | 81% | none |
| 2026-07-08 | fit | 46% | 26%–100% | 35% | ▲ grok-4.5; ▲ gpt-5.6-sol, gpt-5.6-sol-pro +4 |
| 2026-07-15 | fit / held out | 59% | 38%–88% | 83% | muse-spark-1.1; gemini-3.5-flash-lite, gemini-3.6-flash |
| 2026-07-22 | held out | 53% | 32%–93% | 54% | ▲ claude-opus-5; qwen3.7-flash |
| 2026-07-29 | held out | 50% | 26%–89% | 79% | deepseek-v4-flash-0731 |
| 2026-08-05 | held out | 66% | 26%–91% | 81% | muse-spark-1.2; muse-glimmer-30b |
| 2026-08-12 | held out | 36% | 25%–100% | 19% | ▲ grok-4.6; deepseek-v4-pro-0813; qwen3.8-2.4t-a95b; ▲ gemini-3.7-flash; qwen3.8-27b |
| 2026-08-19 | held out | 53% | 39%–100% | 37% | deepseek-v4-flash-vision-exp; muse-spark-1.2-contributor |
| 2026-08-26 | held out | 74% | 56%–100% | 92% | qwen3.8-flash; ▲ claude-fable-5.1 |
| 2026-09-02 | held out | 58% | 31%–100% | 85% | ▲ gemini-3.8-flash; muse-spark-1.3, muse-spark-1.3-contributor; qwen3.8-max-0902; ▲ gpt-6-astra-pro, gpt-6-astra |
| 2026-09-09 | held out | 85% | 59%–100% | 99% | deepseek-v4.1-flash |
| 2026-09-16 | held out | 83% | 28%–100% | 91% | qwen3.8-omni-flash; ▲ grok-4.7; ▲ claude-opus-5.5; ▲ gpt-6-sol, gpt-6-sol-pro +2 |
| 2026-09-23 | held out | 28% | 23%–39% | 5% | qwen3.8-max-prime |
▸RESULTS: FORECAST VS BASE RATE
V3 REPLAY · HELD OUT 16 Jul 2026 – 26 Sep 2026One event throughout: some frontier lab (OpenAI, Anthropic, Google DeepMind, xAI, DeepSeek, Alibaba Qwen, Meta) lists a new text model on OpenRouter within the next 24 hours, 72 hours or 7 days. Twin listings of one model, and one lab's listings within 2 hours of each other, count once. Each formula was fitted on 1 Apr 2026 to 16 Jul 2026 and scored on the 1,638 held-out hours after it. The rule was fixed before the scores were seen.
| Formula | Next 72h · decidesrate 39% → 63% | Next 24hrate 17% → 28% | Next 7drate 73% → 94% |
|---|---|---|---|
| (a) constant base rate | 0 the reference | 0 the reference | 0 the reference |
| (b) max over labs | −0.309−0.83 to +0.06fails sanity | −0.027−0.24 to +0.15fails sanity | −1.481−3.36 to −0.28fails sanity |
| (c) max(market, base) | −0.074−0.61 to +0.26fails sanity | +0.077−0.14 to +0.24sane | +0.037−0.32 to +0.51fails sanity |
| (d) noisy-OR with unpriced rate u | −0.002−0.57 to +0.34fails sanity | +0.032−0.25 to +0.23fails sanity | +0.054−0.36 to +0.67fails sanity |
| (e) (d) + logistic recalibration | −0.035−0.33 to +0.15fails sanity | +0.073−0.04 to +0.19fails sanity | −0.015−0.24 to +0.33fails sanity |
Skill is 1 − Brier ÷ Brier of (a), the fit window's base rate, on the held-out hours; above zero beats it. The small line is the 95% interval from a block bootstrap (2,000 resamples of 168-hour blocks). "Sane" means the mean forecast is within 0.05 of what happened and every bin with at least 100 hours is within 0.2. Rates are the share of hours with a frontier launch inside the horizon, fit window → held out.
PRE-REGISTERED RULEShip a probability only if the best formula beats the base rate at 72h by more than +0.02 with sane reliability. The best was (d) noisy-OR with unpriced rate u: −0.002 (95% interval −0.566 to +0.344), and its worst bin was off by 0.47. Verdict: keep the hand-weighted lead score; the probability is context only.
Rolling origin at 72h: fit on everything before each origin, test to the next
| Test window | Launches | Rate | (d) | (b) | (c) | (e) |
|---|---|---|---|---|---|---|
| 06-11 → 07-07 | 3 | 20% | −0.075 | +0.507 | +0.048 | +0.073 |
| 07-07 → 08-03 | 8 | 65% | +0.187 | −0.120 | +0.045 | +0.060 |
| 08-03 → 08-30 | 11 | 63% | +0.122 | −0.306 | +0.141 | +0.074 |
| 08-30 → 09-26 | 11 | 57% | −0.426 | −0.645 | −0.527 | −0.311 |
(d) scored −0.075, +0.187, +0.122, −0.426 across the 4 folds: the sign flips with the window, which is what a forecast with no stable edge looks like.
Sensitivity at 72h (checked after the fact; none fed the verdict)
| Variant | (d) skill | 95% interval | Sanity |
|---|---|---|---|
| no-bucketsDay-bucket floors dropped: the low end of the bid proxy | +0.190 | −0.04 to +0.36 | fails |
| trustedOnly reads the live trusted-read rule keeps here (no bucket floors, no lower bounds) | +0.187 | +0.01 to +0.35 | fails |
| no-settleMarkets read until they close instead of stopping at the launch | +0.021 | −0.55 to +0.38 | fails |
| uncensoredTest ends 14 days before the pull, so no late rungs are missing | +0.197 | −0.02 to +0.38 | fails |
The best variant reaches +0.197 against the primary read's −0.002, and none passes the sanity check. Dropping day-bucket prices alone moves it by +0.19: they stand in for bids, and with no historical bid or ask the history can't say which read is right.
Reliability: when a forecast says X%, does X% happen?
Left, the v3 forecast on held-out hours. Across forecasts from 27% to 98%, a launch followed in 49% to 81% of hours: the rate hardly tracks the forecast, and the top bin came true 50% of the time. Right, Polymarket on its own questions, from wave 1 and still valid: each rung's price 72 hours before its deadline against how that rung resolved (Brier skill +0.512 over 325 rungs). The markets answer the question they ask, whether one named model ships by a date; that is not a forecast of the next frontier launch.
The plotted bins as a table
| Forecast | hours | Mean forecast | Came true |
|---|---|---|---|
| 20%–30% | 163 | 26.5% | 48.5% |
| 30%–40% | 154 | 35.3% | 59.7% |
| 40%–50% | 239 | 45.5% | 63.6% |
| 50%–60% | 248 | 55.1% | 75.0% |
| 60%–70% | 259 | 64.8% | 78.0% |
| 70%–80% | 197 | 74.4% | 51.8% |
| 80%–90% | 101 | 85.3% | 81.2% |
| 90%–100% | 277 | 97.5% | 50.2% |
| Forecast | rungs | Mean forecast | Came true |
|---|---|---|---|
| 0%–20% | 269 | 2.6% | 1.1% |
| 20%–40% | 31 | 29.7% | 19.4% |
| 40%–60% | 11 | 47.7% | 36.4% |
| 60%–80% | 8 | 70.7% | 87.5% |
| 80%–100% | 6 | 89.6% | 100.0% |
Wave 1 · Polymarket's own calibration
| Priced before deadline | Skill | Brier | Rungs | Resolved Yes |
|---|---|---|---|---|
| 24h | +0.447 | 0.0375 | 328 | 7% |
| 72h | +0.512 | 0.0360 | 325 | 8% |
Skill against always forecasting that horizon's Yes rate, over every frontier-lab rung in the wave-1 pull. On their own questions the markets beat their base rate. The trouble is the question: a rung asks whether one named model ships by a date, so it sits near zero while other frontier models list, and it moves after leaks rather than ahead of them.
By lab: where named markets exist
| Lab | Skill | Priced first | Rate, fit → held out | Hours with a read |
|---|---|---|---|---|
| Anthropic | +0.484 | 3 of 3 | 12% → 13% | 637 |
| OpenAI | +0.305 | 2 of 2 | 9% → 9% | 372 |
| Meta | +0.038 | 0 of 4 | 0% → 18% | 93 |
| Google DeepMind | +0.037 | 2 of 3 | 8% → 13% | 729 |
| DeepSeek | −0.056 | 0 of 4 | 3% → 18% | 0 |
| Alibaba Qwen | −0.109 | 0 of 7 | 10% → 28% | 342 |
| xAI | −0.708 | 2 of 2 | 9% → 9% | 585 |
Each lab's own 72-hour read against its own launches, scored on the held-out hours against that lab's fit-window rate. "Priced first" counts held-out launches whose lab's read reached 50% in the 72 hours before the listing (9 of 25 overall). "Hours with a read" counts held-out hours with a read above 5%. With 2 to 7 launches per lab, a single launch moves these numbers a lot, so read them as where the markets exist, not as a ranking.
Had the level been the 72-hour probability
| Level · band | Launch within 72h | Mean forecast | Held-out hours |
|---|---|---|---|
| 1 · ≥95% | 40% | 99% | 210 (13%) |
| 2 · 70–95% | 66% | 81% | 365 (22%) |
| 3 · 45–70% | 72% | 57% | 655 (40%) |
| 4 · 30–45% | 65% | 38% | 245 (15%) |
| 5 · <30% | 49% | 27% | 163 (10%) |
Bands sit at the fit window's 30%, 45%, 70%, 95% edges. On held-out hours the launch rate did not rise with the level: level 1 hours were followed by a launch 40% of the time, level 3 hours 72%. A level that doesn't rank outcomes can't be sold as a probability.
Limits of the replay
- The sample is launches, not hours. 1,638 held-out hours hold only 25 launches, and one launch marks up to 72 hours in a row as a hit. That is why every interval above is wide.
- The release rate moved. A launch followed within 72 hours in 39% of fit-window hours and 63% of held-out hours, so every constant fitted on the first window is off on the second.
- No historical bid or ask. Every past quote is treated as a tight book, so the live thin-book gate can't be replayed, and day-bucket prices stand in for bids.
- Survivorship. The outcome counts only models OpenRouter still lists. All 21 of 21 hand-curated frontier launches in the window still are, and the delistings OpenRouter had scheduled on 2026-09-26 would erase 0 of its launches. Every xAI listing before 2026-03-31 is gone, so the window starts 2026-04-01. Inside it, a release event disappears only when all of its listings are delisted.
- Censoring. Rungs due after the pull weren't pulled, so market reads in the last 14 days are censored low. Ending the test 14 days early gives (d) +0.197, still failing the sanity check; it did not feed the verdict.
- Launched markets are cut off. Each market is read only until its model launched. Live, a launched model's market kept trading near 1 for a median 2.6 hours (up to 63.3 hours, over 39 markets); the dashboard drops a family for 4 days after its model lists.
Every limitation as the replay records it
- Hourly rows are heavily autocorrelated: one release event sets y=1 for h consecutive hours, so the effective sample size is the number of distinct release events (test: tens), not the number of hours (thousands).
- The release rate is not stationary: the 72h rate was 0.39 in the train window and 0.631 in the test window, so every constant fitted on train is miscalibrated on test.
- No historical bid/ask: every quote is treated as a tight book, so the live thin-book gate (spread > 10¢) cannot be replayed; date-bucket floors use the price in place of the best bid (an upper bound on the live floor; no-buckets is the lower bound). PAV weights are uniform (no historical liquidity).
- The CLOB pull kept the last 15 days of each rung, so a rung is admitted only within 14 days of its deadline. The live trusted-read rule mainly guards reads extrapolated from rungs further out; those cannot be replayed without leaking outcomes, so the trusted variant here only drops lower-bound and bucket-floor reads.
- Rungs due after the pull are missing (unresolved rungs were not pulled), so market reads in the last 14 days are censored low; the uncensored variant stops the test there.
- OpenRouter survivorship: the outcome counts only models OpenRouter still lists (see survivorship).
- Markets are read only until their model launched (announcement, else first YES resolution). Live markets keep trading near 1 until they resolve (median 2.6h after the announcement, up to 63.3h, over 39 markets); the no-settle variant scores an unguarded reader.
▸LEAD TIMES
WAVE 1 · 23 LAUNCHES TIMED BY HANDThe announcement is the first Hacker News story linking the lab's own site or official account and naming the model; teasers ("launching soon", "this Thursday") don't count, and neither do the few earlier first-party links the table marks "not the launch", each with its reason. Availability is OpenRouter's created time. Market crossings are measured against the announcement.
| Launch | Announced (first first-party HN story) | OpenRouter | Before launch | Market read | 0.5 | 0.7 | 0.9 |
|---|---|---|---|---|---|---|---|
| Qwen3.5Alibaba Qwen | 2026-02-16 09:32ZQwen3.5: Towards Native Multimodal AgentsHN | −3.1h | no HN precursorNo pre-launch story on HN. The transformers architecture merged 6.9 days earlier (see below).transformers #43830 | no market | — | — | — |
| Gemma 4Google DeepMind | 2026-04-02 16:00ZGemma 4: Byte for byte, the most capable open modelsHN | +0.8h | no HN precursorNo pre-launch story on HN; the transformers merge landed 36 minutes before the blog.transformers #45192 | no market | — | — | — |
| Claude Opus 4.7Anthropic | 2026-04-16 14:23ZClaude Opus 4.7HN | +0.5h | no HN precursorNo pre-launch story on HN. | On or prior to April 16 | +19.4h | +18.4h | −0.6h |
| GPT-5.5OpenAI | 2026-04-23 18:01ZGPT-5.5HN | +23.5h | leakedA Codex build exposed it 38 hours early; the API (and OpenRouter) followed a day after the post.HN: "GPT 5.5 Released in Codex", 2026-04-22 04:12Z | on April 23 | +9.6d | +9.1d | +35.0h |
| DeepSeek V4DeepSeek | 2026-04-24 02:55ZDeepSeek-V4 Technical Report [pdf]HN | +0.4h | no HN precursorNo pre-launch story on HN; weights and API went live together. | by April 24 | +7.9h | −0.1h | −2.1h |
| Gemini 3.5 FlashGoogle DeepMind | 2026-05-19 17:25ZGemini 3.5 FlashHN | −4.9h | no HN precursorLaunched on stage at the Google I/O keynote, with no model-specific story on HN beforehand. It settled the "Gemini 3.2" and "Gemini 3.5" ladders. | on May 19 | +10.6d | +10.3d | +6.7d |
| Claude Opus 4.8Anthropic | 2026-05-28 16:49ZClaude Opus 4.8HN | −22.7h | leakedA same-morning rumour. OpenRouter lists it 22.7 hours before the post.HN: "Claude Opus 4.8 coming today?", 2026-05-28 09:52Z | only "May 31" (due +3.5d after) | — | — | — |
| Claude Fable 5 / Mythos 5Anthropic | 2026-06-09 16:58ZClaude Fable 5HN | −4.7h | leakedAn unsourced "releasing tomorrow" post the evening before.HN: "Claude Fable 5 by Anthropic, releasing tomorrow", 2026-06-08 19:36Z | by June 9 | +10.0h* | +2.0h | 0.0h |
| Claude Sonnet 5Anthropic | 2026-06-30 15:01ZClaude Sonnet 5HN | +3.2h | no HN precursorNo pre-launch story on HN. | by June 30 | +5.3d | +5.3d | −3.0h |
| Grok 4.5xAI | 2026-07-08 18:00ZGrok 4.5HN | −2.9h | no HN precursorSettled the "Grok 4.4 released by" ladder. | by July 8 | +5.0h | +5.0h | 0.0h |
| GPT-5.6 Sol/Terra/LunaOpenAI | 2026-07-09 17:03ZGPT 5.6 System CardHN | −7.2h | pre-announcedOpenAI named the day 37 hours ahead; a YouTube placeholder went up the evening before.HN: OpenAI "will launch publicly this Thursday", 2026-07-08 04:12ZHN: "GPT-5.6 Sol Ultra will be in Codex", 2026-07-06 01:04Z | on July 9 | +4.3d | +3.7d | +36.1h |
| Kimi K3Moonshot Kimi | 2026-07-16 14:46ZKimi K3: Open Frontier IntelligenceHN | +0.7h | no HN precursorWent live in the Kimi app about an hour before the blog. Not a frontier lab in lab.ts.HN: "Kimi K3 released on web and app", 2026-07-16 13:48Z | only "July 31" (due +15.5d after) | — | — | — |
| Claude Opus 5Anthropic | 2026-07-24 16:55ZClaude Opus 5HN | +0.1h | leakedAn Artificial Analysis model page for it was on HN 23 hours before launch.HN: artificialanalysis.ai/models/claude-opus-5, 2026-07-23 18:01Z | on July 24 | −0.1h | −0.1h | −0.1h |
| Grok 4.6xAI | 2026-08-12 15:32ZGrok 4.6HN | +0.1h | no HN precursorNo pre-launch story on HN. | on August 12 | −0.5h | −0.5h | −0.5h |
| Gemini 3.7 FlashGoogle DeepMind | 2026-08-13 17:02ZGemini 3.7 FlashHN | 0.0h | no HN precursorNo pre-launch story on HN. | by August 13 | +1.0h* | +1.0h* | −1.0h |
| Qwen3.8-FlashAlibaba Qwen | 2026-08-26 12:36ZQwen/Qwen3.8-Flash-NextHN | +7.0h | pre-announcedQwen’s ModelScope page said "releasing tomorrow" 25 hours ahead.HN: "Qwen 3.8-Flash-Next releasing tomorrow" (modelscope.cn), 2026-08-25 11:49Z | only "September 30" (due +35.6d after) | — | — | — |
| Claude Fable 5.1 / Mythos 5.1Anthropic | 2026-09-01 17:53ZPrompting Claude Fable 5.1 and Claude Mythos 5.1HN | +0.2h | no HN precursorNo pre-launch story on HN. | on September 1 | +29.9h | −0.1h | −0.1h |
| Gemini 3.8 FlashGoogle DeepMind | 2026-09-02 15:12ZGemini 3.8 Flash and 3.8 Flash CyberHN | 0.0h | leakedA WSJ scoop 16 hours before launch.HN: WSJ "New Google AI Model Said to Narrow Gap", 2026-09-01 23:23Z | on September 2 | +31.2h | +30.2h | +7.2h |
| GPT-6 AstraOpenAI | 2026-09-03 18:12ZPlayco cut manual fixes 50% prototyping games with GPT‑6 AstraHN | +26.0h | pre-announcedOpenAI published "Path to Astra" two days ahead and a wordless teaser video three hours before the launch. Announced Sep 3, generally available Sep 4; the markets resolved on Sep 4.OpenAI RSS: "Path to Astra", 2026-09-01 13:00ZHN: help.openai.com "OpenAI Astra Launching Soon", 2026-09-03 03:02ZNot the launch: "Path to Astra: critical capabilities and frontier safeguards", the pre-launch safety post, two days earlyNot the launch: "Open AI X post on Astra", an @OpenAI post with no words, only a 12-second video (tweet 2095527557924082061)Not the launch: "OpenAI Releases GPT Astra", the generic models page; commenters found no Astra on it yet | on September 4 | −24.8h | −24.8h | −25.8h |
| DeepSeek V4.1 FlashDeepSeek | 2026-09-10 05:57ZDeepSeek-v4.1-ExpHN | +0.4h | leakedBeta leaks two days out, then a Vercel AI Gateway beta listing 14.5 hours ahead.HN: "v4.1 Flash is now available for internal beta testing", 2026-09-08 08:04ZHN: "DeepSeek v4.1 Flash Beta on Vercel", 2026-09-09 15:28Z | only "September 30" (due +20.9d after) | — | — | — |
| Grok 4.7xAI | 2026-09-21 15:50ZGrok 4.7HN | +0.5h | leakedA 1-point "launching soon" tweet 3.4 days out. xAI has no leading coverage on this site.HN: "Grok 4.7 Launching Soon", 2026-09-18 06:42Z | on September 21 | +0.8h | −0.2h | −0.2h |
| Claude Opus 5.5Anthropic | 2026-09-22 16:27ZClaude Opus 5.5HN | +0.1h | leakedNothing on HN, but TestingCatalog reported testing 31.7 hours ahead; the day market moved 35 minutes after it.TestingCatalog: "Anthropic tests Fable 5.2 and Opus 5.5", 2026-09-21 08:45Z | on September 22 | +24.5h | +21.5h | +0.5h |
| GPT-6 Sol/LunaOpenAI | 2026-09-22 17:58ZGPT-6 SolHN | +0.2h | leakedA Reddit post saw the id on the API 11 days early; TestingCatalog called the day that morning.HN: "GPT-6-sol appeared on OpenAI API", 2026-09-11 20:43ZTestingCatalog: "prepares to launch GPT-6 Sol and Luna today", 2026-09-22 12:07Z | On or prior to September 22 | +7.0h | +6.0h | 0.0h |
OpenRouter is hours from the announcement to the model's created time (negative: listed first). Market leads are hours before the announcement at which the price reached the threshold and held it for two hours; negative means after. An asterisk marks a rung that was already above the threshold at its first usable price, so the lead is the rung's age, not a move.
▸PRE-ANNOUNCED, LEAKED OR NO HN PRECURSOR
WAVE 1By what was public beforehand
| Group | Launches with a dated market | Held 0.5 before | Held 0.9 before | Median 0.5 lead | Undated only |
|---|---|---|---|---|---|
| Pre-announced by the lab | 2 | 1 | 1 | +39.2h | 1 |
| Leaked or reported | 7 | 6 | 3 | +10.0h | 2 |
| No HN precursor | 8 | 7 | 1 | +13.7h | 1 |
By lab
| Group | Launches with a dated market | Held 0.5 before | Held 0.9 before | Median 0.5 lead | Undated only |
|---|---|---|---|---|---|
| Anthropic | 6 | 5 | 1 | +22.0h | 1 |
| OpenAI | 4 | 3 | 2 | +2.3d | 0 |
| DeepSeek | 1 | 1 | 0 | +7.9h | 1 |
| Google DeepMind | 3 | 3 | 2 | +31.2h | 0 |
| xAI | 3 | 2 | 0 | +0.8h | 0 |
| Moonshot Kimi | 0 | 0 | 0 | — | 1 |
| Alibaba Qwen | 0 | 0 | 0 | — | 1 |
"Undated only" launches were priced only on rungs due weeks after they shipped, which say the launch is coming but not when. Of the 5 launches whose market held 0.9 before the announcement, 4 followed a leak, a press report or the lab's own pre-announcement (the exception: Gemini 3.5 Flash, launched at a scheduled keynote or with no public trail). The markets aggregate that news rather than foresee it.
The "above 50% is usually real" claim
| Reading | With a day or more to spare | Inside the last day |
|---|---|---|
| Research note (script not kept) | 11 / 24 (46%) | 33 / 46 (72%) |
| First rung to fire per ladder | 12 / 23 (52%) | 15 / 18 (83%) |
| Every rung that fired | 51 / 68 (75%) | 16 / 19 (84%) |
A rung "fires" when a frontier lab's "released by" rung trades above 0.5 within 7 days of its deadline, before the model launched; it counts as real when the rung resolved Yes. Counting the first rung per ladder (55 ladders, 260 rungs) gives 12 of 23, close to the research note's 11 of 24. Counting every rung inflates the rate, because neighbouring rungs of one ladder move together. The note's 33 of 46 near launch does not reproduce under either count, and its script was not kept. On the per-ladder count, a reading above 50% with a day to spare was right 52% of the time.
▸FALSE ALARMS
WAVE 1| Market type | Threshold | No rungs that crossed | Per market-day | Yes rungs that crossed before launch | Crossings that resolved Yes |
|---|---|---|---|---|---|
| "Released by" rungs | 0.5 | 37 / 126 | 0.0265 | 129 / 134 | 78% |
| "Released by" rungs | 0.7 | 16 / 126 | 0.0114 | 127 / 134 | 89% |
| "Released by" rungs | 0.9 | 4 / 126 | 0.0029 | 105 / 134 | 96% |
| "Released on" day buckets | 0.5 | 17 / 426 | 0.0052 | 8 / 10 | 32% |
| "Released on" day buckets | 0.7 | 6 / 426 | 0.0018 | 5 / 10 | 46% |
| "Released on" day buckets | 0.9 | 1 / 426 | 0.0003 | 5 / 10 | 83% |
Market-days count the traded life of every No rung (1399 days for "by" rungs). A Yes rung counts only prices from before its model was announced, so a post-launch settlement at 0.99 is not a correct call. Every rung's price counts from the end of its first hour, and only once it has moved off the quote the book opened at (often exactly 0.5 before anyone trades).
Loudest false alarms (No rungs that peaked at 0.5 or more)
| Market | Rung | Peak | Peak at | Launch came |
|---|---|---|---|---|
| OpenAI's Astra released on...? | on September 3 | 0.966 | 2026-09-03 16:00Z | 1d later† |
| OpenAI’s Astra released by…? | by September 3 | 0.964 | 2026-09-03 16:00Z | 1d later† |
| Next Claude Opus released by...? | by July 23 | 0.920 | 2026-07-23 13:00Z | 1d later |
| Next Google Gemini Pro Model released by...? | by August 14 | 0.915 | 2026-08-03 18:00Z | not yet |
| GPT-5.6 released by...? | by June 30 | 0.900 | 2026-06-19 19:00Z | 9d later |
| New Gemini reasoning flagship released by...? | by May 22 | 0.895 | 2026-05-11 19:00Z | not yet |
| GPT-5.6 released by...? | by July 8 | 0.890 | 2026-07-03 05:00Z | 1d later |
| Next Claude Opus released on...? | on July 23 | 0.890 | 2026-07-23 13:00Z | 1d later |
| Next Google Gemini Pro Model released by...? | by August 7 | 0.885 | 2026-07-24 15:00Z | not yet |
| Next Grok Model (4.7+) released by...? | by September 18 | 0.885 | 2026-09-07 23:00Z | not yet |
| New Gemini reasoning flagship released by...? | by June 30 | 0.879 | 2026-06-16 17:00Z | not yet |
| Next Google Gemini Pro Model released by...? | by June 30 | 0.850 | 2026-06-16 12:00Z | not yet |
54 No rungs peaked at 0.5 or more; the 12 highest are listed. "Launch came" is the gap to the ladder's first Yes deadline. † The launch was announced inside this rung, which still resolved No under the market's own release rule (OpenAI's Astra released on...?, OpenAI’s Astra released by…?). Those are rule misses more than timing misses.
▸STEALTH REVEALS
EARLY WARNING · NEVER SCORED| Slot | Days to official | Revealed as | Stealth listing | Named by |
|---|---|---|---|---|
| Union Alpha | 1.3 | Pareto · UnbiasedA priced third-party model, not OpenRouter's openrouter/pareto-code router. | 2026-09-16 | OpenRouter page |
| Ox Alpha | 5.7 | GLM-5.3-Flash · Z.ai | 2026-08-20 | OpenRouter page |
| Owl Alpha | 82.8 | LongCat-2.0 · MeituanNo notice on the OpenRouter page. longcatai.org says the LongCat team confirmed it after about two months; the official OpenRouter listing came 83 days in. | 2026-04-28 | third partyunverified, left out of the medians |
| Hunter/Healer Alpha | 7.0 | MiMo-V2-Pro and MiMo-V2-Omni · Xiaomi | 2026-03-11 | OpenRouter page |
| Pony Alpha | 5.0 | GLM-5 · Z.ai | 2026-02-06 | OpenRouter page |
| Bert-Nebulon Alpha | 7.2 | Mistral Large 3 · Mistral | 2025-11-24 | slot description |
| Sherlock Think/Dash Alpha | 4.1 | Grok 4.1 Fast · xAIfrontier lab | 2025-11-15 | slot description |
| Polaris Alpha | 7.0 | GPT-5.1 · OpenAIfrontier lab | 2025-11-06 | slot description |
| Andromeda Alpha | 6.9 | Nemotron Nano 2 VL · NVIDIA | 2025-10-21 | slot description |
| Sonoma Sky/Dusk Alpha | 13.3 | Grok 4 Fast · xAIfrontier lab | 2025-09-05 | OpenRouter page |
| Horizon Alpha/Beta | 7.8 | GPT-5 · OpenAIfrontier lab | 2025-07-30 | OpenRouter blog |
| Quasar/Optimus Alpha | 11.9 | GPT-4.1 · OpenAIfrontier lab | 2025-04-02 | OpenRouter blog |
OpenRouter sometimes lists a lab's next model under a codename before launch. Days run from the slot's first OpenRouter created time to the official listing's, both re-read from the API; created can move after the fact, so treat single rows as approximate. Only verified rows feed the medians: Owl Alpha is shown but left out. A slot that never revealed itself is not in the table, so the medians describe slots that were revealed, not every slot. The dashboard lists live slots as an early warning with this track record; they never move the level.
▸SCHEDULED BROADCASTS
OpenAI sometimes publishes an upcoming-livestream page hours before a launch. 4 of these 5 have an archived copy proving it was up before the stream; GPT-5.6 is timed from YouTube's own publishedAt, with no archive. 4 other OpenAI launches had no scheduled stream at all.
| Launch | Stream | Placeholder published | Starts | Warning | Evidence |
|---|---|---|---|---|---|
| GPT-4.5 | Introduction to GPT-4.5Placeholder named the model. | 2025-02-27 17:00Z | 2025-02-27 20:00Z | +3.0h | Wayback, upcoming |
| GPT-4.1 | New models in the APIPlaceholder title did not name the model. | 2025-04-14 14:49Z | 2025-04-14 17:00Z | +2.2h | Wayback, upcoming |
| o3 / o4-mini | Introduction to new o-series modelsThe longest lead: scheduled two days out. | 2025-04-14 20:48Z | 2025-04-16 17:00Z | +44.2h | Wayback, upcoming |
| GPT-5 | Introducing GPT-5Placeholder named the model. | 2025-08-07 12:07Z | 2025-08-07 17:00Z | +4.9h | Wayback, upcoming |
| GPT-5.6 | Introducing the next chapter for ChatGPTStart is the actual air time; OpenRouter listed the models 13.5 hours after the placeholder. | 2026-07-08 20:25Z | 2026-07-09 16:53Z | +20.5h | no capture |
Launches a broadcast watch would have missed
- GPT-5.4 No scheduled stream.
- GPT-5.5 No scheduled stream.
- GPT-6 Astra No scheduled stream.
- GPT-6 Sol/Luna No scheduled stream.
- Every Anthropic launch Uploads coincide with or trail the post.
- Google DeepMind Uploads trail the launch.
- xAI The @xai handle resolves to an unrelated channel.
▸ARCHITECTURE MERGES
Open-weight labs often add their model code to Hugging Face transformers before release. The merge time comes from the pull request; the announcement is the first first-party HN story.
The merge led the launch for 7 of 16 shipped families, coincided within 6h for 5, and trailed it for 4. 1 module is merged and still unshipped.
| Family | Module | Merged | Announced | Merge lead | Verdict |
|---|---|---|---|---|---|
| Qwen3Qwen | qwen3 #36878 | 2025-03-31 07:50Z | 2025-04-28 20:44Z | +28.5d | leads |
| Qwen3-VLQwen | qwen3_vl #40795 | 2025-09-15 10:46Z | 2025-09-23 20:59Z | +8.4d | leads |
| GLM-4.5Z.ai | glm4_moe #39393 | 2025-07-21 11:24Z | 2025-07-28 14:15Z | +7.1d | leads |
| Qwen3.5Qwen | qwen3_5 #43830 | 2026-02-09 11:21Z | 2026-02-16 09:32Z | +6.9d | leads |
| GLM-5Z.ai | glm_moe_dsa #43858 | 2026-02-09 12:06Z | 2026-02-11 13:42Z | +2.1d | leads |
| Qwen3-NextQwen | qwen3_next #40771 | 2025-09-09 21:46Z | 2025-09-11 17:38Z | +43.9h | leads |
| Qwen3-OmniQwen | qwen3_omni_moe #41025 | 2025-09-21 21:46Z | 2025-09-22 17:50Z | +20.1h | leads |
| Mistral Small 4Mistral | mistral4 #44760 | 2026-03-16 19:39Z | 2026-03-16 20:40Z | +1.0h | coincident |
| gpt-ossOpenAI | gpt_oss #39923 | 2025-08-05 16:02Z | 2025-08-05 17:00Z | +1.0h | coincident |
| Gemma 4Google | gemma4 #45192 | 2026-04-02 15:24Z | 2026-04-02 16:00Z | +0.6h | coincident |
| GLM-5.3-FlashZ.ai | glm5_next #48342 | 2026-08-26 14:26Z | 2026-08-26 14:08Z | −0.3h | coincident |
| Llama 4Meta | llama4 #37307 | 2025-04-05 20:02Z | 2025-04-05 18:33Z | −1.5h | coincident |
| DeepSeek V4DeepSeek | deepseek_v4 #45643 | 2026-05-02 11:41Z | 2026-04-24 02:55Z | −8.4d | lags |
| MiniMax-M2MiniMax | minimax_m2 #42028 | 2026-01-09 15:25Z | 2025-10-27 05:12Z | −74.4d | lags |
| MiniMax-Text-01MiniMax | minimax #35831 | 2025-06-04 07:38Z | 2025-01-14 19:32Z | −140.5d | lags |
| Kimi LinearMoonshot | kimi_linear #48250 | 2026-09-05 17:04Z | 2025-10-30 15:44Z | −310.1d | lags |
| Qwen4-ExpQwen | qwen4_exp #48337 | 2026-08-26 12:03Z | not yet | 31d pending | pending |
- Qwen3.8 Reused the qwen3_5 classes; no qwen3_8 module exists (registry read 2026-09-26).
- OpenAI, Anthropic, xAI frontier models Closed weights never touch transformers.
▸TESTED AND DIDN'T LEAD
| Signal and verdict | Finding |
|---|---|
| Release cadence ("overdue" labs)no better than chance | Walk-forward: labs at z ≥ 0.5 past their mean gap shipped within 7 days 11.3% of the time, against a 20% base rate. Releases are bursty (CV 0.9–1.6). OpenRouter models API |
| Cross-lab clusteringno better than chance | Another frontier lab shipped within 3d/7d 37.3%/66.7% of the time, against 37.3%/66.1% for a weekday-preserving shuffle (p = 0.51/0.44). OpenRouter models API |
| SDK and OpenAPI commit feedslags | Coincident at best; 5 of 7 cases lag. The Opus 5.5 SDK release went out 2 minutes before the post. anthropic-sdk-python v1.8.0 |
| sglang commitslags | Led 1 of 7 launches, and only echoed a transformers merge that came first. sglang commits |
| litellm and vercel/ai model listslags | Trail OpenRouter by 20 minutes to 1h51m; litellm mirrors OpenRouter through a bot. litellm commits |
| Lab docs catalog pagescoincident | Across 5 launches, no Wayback snapshot shows a docs page carrying a model before its announcement. GPT-6 Astra was still absent 6.4h after launch. Wayback Machine |
| AWS What’s Newfakes a lead | pubDate is a backdated editorial time; read at face value it fakes a lead on 13 of 15 Claude launches. AWS What’s New feed |
| Status pages (Claude, OpenAI)lags | A model is first mentioned 6 hours to about 20 days after its listing. xAI’s returns 403. status.claude.com |
| Wikipedia infoboxeslags | Updated after launch in 6 of 6 cases. MediaWiki API |
| Hugging Face collectionslags | Lagged or coincided in 9 of 9 releases. Hugging Face API |
| Vertex AI release noteslags | The legacy feed has been stale since March; its successor has day-level dates and lags. Vertex AI release notes |
| Design Arena codenameslags | The public registry purged its anonymous codenames in June 2026, and every entry since appeared after launch. One anecdote survives: "Radon" was Muse Spark, 6d14h early. Design Arena |
| Manifoldnot independent | Pegged to Polymarket, 30 minutes to 3 hours behind it, on thin books. Manifold API |
| Kalshinot independent | GPT-6 moved 0.19 → 0.83 in the same hour as Polymarket (Sep 2, 15–17Z). No markets existed for Opus 5.5 or Sol/Luna. Kalshi API |
Findings from the 2026-09-26 research pass. pnpm backtest carries them with their sources; it does not recompute them.
▸TIMESTAMP TRAPS
Each of these timestamps looks like an early warning and isn't. Including one this site got wrong.
| Trap | Example | Fake lead | Fix |
|---|---|---|---|
| Git commit dates | The anthropic-sdk-python commit adding claude-opus-5-5 is dated 2026-09-20 22:54:59Z; it shipped in v1.8.0 at 2026-09-22 16:25:23Z, two minutes before the launch post. commit b5cc700 | +41.5h | Use the release or push time, never the author or committer date. |
| AWS What’s New pubDate | The GovCloud Opus 5.5 item was created 2h30m after its own pubDate (research pass; not re-probed here). AWS What’s New feed | +2.5h | Record when the item was first seen. |
| Hugging Face createdAt | createdAt is when the repo was made, while still private: google/gemma-4-E4B-it says 2026-03-02 19:57Z, 30.8 days before the launch post. google/gemma-4-E4B-it | +30.8d | Only a first-seen time from polling counts. |
| Date-only fields pinned to midnight | xAI’s release notes date Grok 4.7 "September 21, 2026" with no time. Read as 00:00Z, that is 15h50m before its HN story. xAI release notes | +15.8h | Mark day-precision items and keep them out of lead analysis. |
| Rounded RSS pubDates | openai.com’s RSS dates "GPT-6 Astra: A new generation of intelligence" to 2026-09-03 11:00:00Z; its HN story is 18:41Z. openai.com RSS | +7.7h | Treat whole-hour and midnight pubDates as day precision. |
| OpenRouter created before the announcement | Listed before the announcement: Claude Opus 4.8 -22.7h, GPT-5.6 Sol/Terra/Luna -7.2h, Gemini 3.5 Flash -4.9h, Claude Fable 5 / Mythos 5 -4.7h, Qwen3.5 -3.1h, Grok 4.5 -2.9h. OpenRouter also rewrites it: Space Bunny Alpha moved from 2026-09-22 10:58Z to 2026-09-23 14:48Z. OpenRouter models API | +22.7h | Use it as availability, measured against the announcement, never as a lead. |
| This site’s own "her" story | The History timeline said Altman posted "her" the night before GPT-4o. The tweet id decodes to 2024-05-13T17:45:09Z; the GPT-4o announcement hit HN at 2024-05-13T17:28:00Z, so the tweet came 17 minutes after it. sama/status/1790075827666796666 | — | Decode the tweet id rather than trusting a remembered order. |
▸REPRODUCE IT
Every computed table is rebuilt from raw pulls committed to the repo. The curated tables (stealth reveals, broadcasts, architecture merges, what didn't lead, timestamp traps) are carried as sourced data, not recomputed. The rebuild is deterministic: run it twice and git diff stays empty.
pnpm install
pnpm backtest # rebuild data/backtest/*.json from data/backtest/raw
pnpm backtest:replay # rerun the v3 replay into data/backtest/v3-replay.json
pnpm backtest --refresh # re-pull Gamma, CLOB, HN and OpenRouter firstRaw pulls last refreshed 2026-09-26 00:46Z. Curated inputs (release ids, sources, the tables that are not recomputed) live in scripts/backtest/curated.ts; the replay, its pre-registered rule and every constant it fits are in scripts/backtest/replay.ts.
Or check any single number by hand:
# Closed AI events, 50 per page; pass next_cursor back as after_cursor
curl 'https://gamma-api.polymarket.com/events/keyset?tag_slug=ai&closed=true&limit=50'
# Claude Opus 5.5: "September 22" price history (windows of 15 days or less)
curl 'https://clob.polymarket.com/prices-history?market=109495115084615347788619140162263908414615168053426928181800069578275691598743&startTs=1789597495&endTs=1790102053&fidelity=60'
# The announcement: first first-party story in the window (Algolia wants the filter URL-encoded)
curl 'https://hn.algolia.com/api/v1/search_by_date?query=Opus%205.5&tags=story&hitsPerPage=1000&numericFilters=created_at_i%3E%3D1789776000%2Ccreated_at_i%3C%3D1790208000'
# Availability
curl -s 'https://openrouter.ai/api/v1/models' | jq '.data[] | select(.id == "anthropic/claude-opus-5.5") | .created'
# The pending architecture
curl -s 'https://cdn.jsdelivr.net/gh/huggingface/transformers@main/src/transformers/models/__init__.py' | grep qwen4_exp