◀ WHENMODEL DASHBOARD

BACKTEST

Can public signals see a frontier model launch coming?

METHODPolymarket's release odds, replayed hour by hour from 1 Apr 2026 to 26 Sep 2026 through the dashboard's own code: 5 forecasts fitted on the first 20 frontier launches and scored on the next 25. Plus 23 launches timed by hand against Hacker News, lab feeds, TestingCatalog and the other sources in the tables below.

Best 72h skill
−0.002
needed +0.02 · 95% −0.57 to +0.34
Priced before listing
9 / 25
held-out launches, own lab ≥50%
Held-out launches
25
the real sample size · fitted on 20
Level's input, 7d
−1.48
95% −3.36 to −0.28

VERDICTNo formula beat the base rate at 72 hours, so DROPCON stays a hand-weighted lead score, not a probability. Anthropic (3 launches) and OpenAI (2) scored highest on their own launches, +0.48 and +0.30, and the other 5 labs −0.71 to +0.04; with 2 to 7 launches per lab these are anecdotes, not a ranking. Most held-out launches, 16 of 25, had no market at 50% or more beforehand.

▸HOW V3 WOULD HAVE READ

V3 REPLAY · 1 Apr 2026 – 26 Sep 2026

Every 3 hours from 1 Apr 2026: what the markets said about the next 72 hours and the next 7 days, read only from prices that existed at the time, against the 45 frontier launches that followed.

◀ FIT · 20 LAUNCHESHELD OUT · 25 LAUNCHES ▶NEXT 72H · FORECAST (d)0%50%100%fit rate 39%held out 63%NEXT 7D · BEST LAB READ (P7)0%50%100%fit rate 73%held out 94%≥90% for 4.6 days, nothing listedAprMayJunJulAugSep2026-04-02 · Alibaba Qwen · qwen3.6-plus. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-04-02 · Google DeepMind · gemma-4-31b-it. Its own lab's 72-hour read peaked at 1% in the 72 hours before it listed (not priced ≥50%).2026-04-03 · Google DeepMind · gemma-4-26b-a4b-it. Its own lab's 72-hour read peaked at 1% in the 72 hours before it listed (not priced ≥50%).2026-04-16 · Anthropic · claude-opus-4.7. Its own lab's 72-hour read peaked at 100% in the 72 hours before it listed (priced ≥50%).2026-04-24 · DeepSeek · deepseek-v4-flash, deepseek-v4-pro. Its own lab's 72-hour read peaked at 99% in the 72 hours before it listed (priced ≥50%).2026-04-24 · OpenAI · gpt-5.5, gpt-5.5-pro. Its own lab's 72-hour read peaked at 98% in the 72 hours before it listed (priced ≥50%).2026-04-27 · Alibaba Qwen · qwen3.6-27b, qwen3.6-max-preview +3. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-04-30 · xAI · grok-4.3. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-05-05 · OpenAI · gpt-chat-latest. Its own lab's 72-hour read peaked at 2% in the 72 hours before it listed (not priced ≥50%).2026-05-07 · Google DeepMind · gemini-3.1-flash-lite. Its own lab's 72-hour read peaked at 100% in the 72 hours before it listed (priced ≥50%).2026-05-19 · Google DeepMind · gemini-3.5-flash. Its own lab's 72-hour read peaked at 100% in the 72 hours before it listed (priced ≥50%).2026-05-20 · xAI · grok-build-0.1. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-05-21 · Alibaba Qwen · qwen3.7-max. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-05-27 · Anthropic · claude-opus-4.8. Its own lab's 72-hour read peaked at 46% in the 72 hours before it listed (not priced ≥50%).2026-06-03 · Alibaba Qwen · qwen3.7-plus. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-06-09 · Anthropic · claude-fable-5. Its own lab's 72-hour read peaked at 100% in the 72 hours before it listed (priced ≥50%).2026-06-30 · Anthropic · claude-sonnet-5. Its own lab's 72-hour read peaked at 94% in the 72 hours before it listed (priced ≥50%).2026-07-08 · xAI · grok-4.5. Its own lab's 72-hour read peaked at 97% in the 72 hours before it listed (priced ≥50%).2026-07-09 · OpenAI · gpt-5.6-sol, gpt-5.6-sol-pro +4. Its own lab's 72-hour read peaked at 100% in the 72 hours before it listed (priced ≥50%).2026-07-16 · Meta · muse-spark-1.1. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-07-21 · Google DeepMind · gemini-3.5-flash-lite, gemini-3.6-flash. Its own lab's 72-hour read peaked at 16% in the 72 hours before it listed (not priced ≥50%).2026-07-24 · Anthropic · claude-opus-5. Its own lab's 72-hour read peaked at 93% in the 72 hours before it listed (priced ≥50%).2026-07-27 · Alibaba Qwen · qwen3.7-flash. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-07-31 · DeepSeek · deepseek-v4-flash-0731. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-08-05 · Meta · muse-spark-1.2. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-08-09 · Meta · muse-glimmer-30b. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-08-12 · xAI · grok-4.6. Its own lab's 72-hour read peaked at 100% in the 72 hours before it listed (priced ≥50%).2026-08-12 · DeepSeek · deepseek-v4-pro-0813. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-08-12 · Alibaba Qwen · qwen3.8-2.4t-a95b. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-08-13 · Google DeepMind · gemini-3.7-flash. Its own lab's 72-hour read peaked at 91% in the 72 hours before it listed (priced ≥50%).2026-08-14 · Alibaba Qwen · qwen3.8-27b. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-08-21 · DeepSeek · deepseek-v4-flash-vision-exp. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-08-21 · Meta · muse-spark-1.2-contributor. Its own lab's 72-hour read peaked at 13% in the 72 hours before it listed (not priced ≥50%).2026-08-26 · Alibaba Qwen · qwen3.8-flash. Its own lab's 72-hour read peaked at 31% in the 72 hours before it listed (not priced ≥50%).2026-09-01 · Anthropic · claude-fable-5.1. Its own lab's 72-hour read peaked at 92% in the 72 hours before it listed (priced ≥50%).2026-09-02 · Google DeepMind · gemini-3.8-flash. Its own lab's 72-hour read peaked at 100% in the 72 hours before it listed (priced ≥50%).2026-09-02 · Meta · muse-spark-1.3, muse-spark-1.3-contributor. Its own lab's 72-hour read peaked at 16% in the 72 hours before it listed (not priced ≥50%).2026-09-03 · Alibaba Qwen · qwen3.8-max-0902. Its own lab's 72-hour read peaked at 9% in the 72 hours before it listed (not priced ≥50%).2026-09-04 · OpenAI · gpt-6-astra-pro, gpt-6-astra. Its own lab's 72-hour read peaked at 100% in the 72 hours before it listed (priced ≥50%).2026-09-10 · DeepSeek · deepseek-v4.1-flash. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-09-21 · Alibaba Qwen · qwen3.8-omni-flash. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-09-21 · xAI · grok-4.7. Its own lab's 72-hour read peaked at 77% in the 72 hours before it listed (priced ≥50%).2026-09-22 · Anthropic · claude-opus-5.5. Its own lab's 72-hour read peaked at 97% in the 72 hours before it listed (priced ≥50%).2026-09-22 · OpenAI · gpt-6-sol, gpt-6-sol-pro +2. Its own lab's 72-hour read peaked at 100% in the 72 hours before it listed (priced ≥50%).2026-09-23 · Alibaba Qwen · qwen3.8-max-prime. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).◀ FITHELD OUT ▶NEXT 72H · FORECAST (d)0%50%100%fit 39%held out 63%NEXT 7D · BEST LAB READ (P7)0%50%100%fit 73%held out 94%AprMayJunJulAugSep2026-04-02 · Alibaba Qwen · qwen3.6-plus. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-04-02 · Google DeepMind · gemma-4-31b-it. Its own lab's 72-hour read peaked at 1% in the 72 hours before it listed (not priced ≥50%).2026-04-03 · Google DeepMind · gemma-4-26b-a4b-it. Its own lab's 72-hour read peaked at 1% in the 72 hours before it listed (not priced ≥50%).2026-04-16 · Anthropic · claude-opus-4.7. Its own lab's 72-hour read peaked at 100% in the 72 hours before it listed (priced ≥50%).2026-04-24 · DeepSeek · deepseek-v4-flash, deepseek-v4-pro. Its own lab's 72-hour read peaked at 99% in the 72 hours before it listed (priced ≥50%).2026-04-24 · OpenAI · gpt-5.5, gpt-5.5-pro. Its own lab's 72-hour read peaked at 98% in the 72 hours before it listed (priced ≥50%).2026-04-27 · Alibaba Qwen · qwen3.6-27b, qwen3.6-max-preview +3. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-04-30 · xAI · grok-4.3. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-05-05 · OpenAI · gpt-chat-latest. Its own lab's 72-hour read peaked at 2% in the 72 hours before it listed (not priced ≥50%).2026-05-07 · Google DeepMind · gemini-3.1-flash-lite. Its own lab's 72-hour read peaked at 100% in the 72 hours before it listed (priced ≥50%).2026-05-19 · Google DeepMind · gemini-3.5-flash. Its own lab's 72-hour read peaked at 100% in the 72 hours before it listed (priced ≥50%).2026-05-20 · xAI · grok-build-0.1. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-05-21 · Alibaba Qwen · qwen3.7-max. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-05-27 · Anthropic · claude-opus-4.8. Its own lab's 72-hour read peaked at 46% in the 72 hours before it listed (not priced ≥50%).2026-06-03 · Alibaba Qwen · qwen3.7-plus. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-06-09 · Anthropic · claude-fable-5. Its own lab's 72-hour read peaked at 100% in the 72 hours before it listed (priced ≥50%).2026-06-30 · Anthropic · claude-sonnet-5. Its own lab's 72-hour read peaked at 94% in the 72 hours before it listed (priced ≥50%).2026-07-08 · xAI · grok-4.5. Its own lab's 72-hour read peaked at 97% in the 72 hours before it listed (priced ≥50%).2026-07-09 · OpenAI · gpt-5.6-sol, gpt-5.6-sol-pro +4. Its own lab's 72-hour read peaked at 100% in the 72 hours before it listed (priced ≥50%).2026-07-16 · Meta · muse-spark-1.1. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-07-21 · Google DeepMind · gemini-3.5-flash-lite, gemini-3.6-flash. Its own lab's 72-hour read peaked at 16% in the 72 hours before it listed (not priced ≥50%).2026-07-24 · Anthropic · claude-opus-5. Its own lab's 72-hour read peaked at 93% in the 72 hours before it listed (priced ≥50%).2026-07-27 · Alibaba Qwen · qwen3.7-flash. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-07-31 · DeepSeek · deepseek-v4-flash-0731. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-08-05 · Meta · muse-spark-1.2. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-08-09 · Meta · muse-glimmer-30b. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-08-12 · xAI · grok-4.6. Its own lab's 72-hour read peaked at 100% in the 72 hours before it listed (priced ≥50%).2026-08-12 · DeepSeek · deepseek-v4-pro-0813. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-08-12 · Alibaba Qwen · qwen3.8-2.4t-a95b. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-08-13 · Google DeepMind · gemini-3.7-flash. Its own lab's 72-hour read peaked at 91% in the 72 hours before it listed (priced ≥50%).2026-08-14 · Alibaba Qwen · qwen3.8-27b. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-08-21 · DeepSeek · deepseek-v4-flash-vision-exp. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-08-21 · Meta · muse-spark-1.2-contributor. Its own lab's 72-hour read peaked at 13% in the 72 hours before it listed (not priced ≥50%).2026-08-26 · Alibaba Qwen · qwen3.8-flash. Its own lab's 72-hour read peaked at 31% in the 72 hours before it listed (not priced ≥50%).2026-09-01 · Anthropic · claude-fable-5.1. Its own lab's 72-hour read peaked at 92% in the 72 hours before it listed (priced ≥50%).2026-09-02 · Google DeepMind · gemini-3.8-flash. Its own lab's 72-hour read peaked at 100% in the 72 hours before it listed (priced ≥50%).2026-09-02 · Meta · muse-spark-1.3, muse-spark-1.3-contributor. Its own lab's 72-hour read peaked at 16% in the 72 hours before it listed (not priced ≥50%).2026-09-03 · Alibaba Qwen · qwen3.8-max-0902. Its own lab's 72-hour read peaked at 9% in the 72 hours before it listed (not priced ≥50%).2026-09-04 · OpenAI · gpt-6-astra-pro, gpt-6-astra. Its own lab's 72-hour read peaked at 100% in the 72 hours before it listed (priced ≥50%).2026-09-10 · DeepSeek · deepseek-v4.1-flash. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-09-21 · Alibaba Qwen · qwen3.8-omni-flash. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).2026-09-21 · xAI · grok-4.7. Its own lab's 72-hour read peaked at 77% in the 72 hours before it listed (priced ≥50%).2026-09-22 · Anthropic · claude-opus-5.5. Its own lab's 72-hour read peaked at 97% in the 72 hours before it listed (priced ≥50%).2026-09-22 · OpenAI · gpt-6-sol, gpt-6-sol-pro +2. Its own lab's 72-hour read peaked at 100% in the 72 hours before it listed (priced ≥50%).2026-09-23 · Alibaba Qwen · qwen3.8-max-prime. Its own lab's 72-hour read peaked at 0% in the 72 hours before it listed (not priced ≥50%).
Both lines are market reads alone, replayed as a reader could have seen them at the time through the dashboard's own curve code. The top line is the 72-hour forecast the calibration study scored: the markets combined with a 23% chance of a launch no market prices. The bottom line is the best 7-day read across labs, the shape of the input behind 80 of DROPCON's 100 points. The level itself can't be replayed: its 30-day term reads rungs up to a month out, and the price history holds each rung only from 14 days before its deadline. The rate rose from 39% to 63% of 72-hour windows, and 16 of the 25 held-out launches had no market at 50% or more beforehand. The longest false alarm: the forecast held 90% or more from 12 Sep to 16 Sep and nothing frontier listed in the 72 hours after.
The chart as a table, by week
Weekly summary of the replay: the 72-hour forecast, the best 7-day read and the frontier launches listed that week.
Week ofWindow72h mean72h range7d read meanLaunches (▲ priced)
2026-04-01fit34%25%–49%24%qwen3.6-plus; gemma-4-31b-it; gemma-4-26b-a4b-it
2026-04-08fit53%29%–100%68%none
2026-04-15fit63%24%–98%86%▲ claude-opus-4.7
2026-04-22fit45%24%–99%30%▲ deepseek-v4-flash, deepseek-v4-pro; ▲ gpt-5.5, gpt-5.5-pro; qwen3.6-27b, qwen3.6-max-preview +3
2026-04-29fit26%24%–34%8%grok-4.3; gpt-chat-latest
2026-05-06fit35%26%–100%49%▲ gemini-3.1-flash-lite
2026-05-13fit72%29%–100%94%▲ gemini-3.5-flash
2026-05-20fit31%26%–48%13%grok-build-0.1; qwen3.7-max
2026-05-27fit40%26%–85%28%claude-opus-4.8
2026-06-03fit43%32%–100%30%qwen3.7-plus; ▲ claude-fable-5
2026-06-10fit34%31%–38%25%none
2026-06-17fit52%36%–66%60%none
2026-06-24fit56%32%–96%52%▲ claude-sonnet-5
2026-07-01fit50%25%–95%81%none
2026-07-08fit46%26%–100%35%▲ grok-4.5; ▲ gpt-5.6-sol, gpt-5.6-sol-pro +4
2026-07-15fit / held out59%38%–88%83%muse-spark-1.1; gemini-3.5-flash-lite, gemini-3.6-flash
2026-07-22held out53%32%–93%54%▲ claude-opus-5; qwen3.7-flash
2026-07-29held out50%26%–89%79%deepseek-v4-flash-0731
2026-08-05held out66%26%–91%81%muse-spark-1.2; muse-glimmer-30b
2026-08-12held out36%25%–100%19%▲ grok-4.6; deepseek-v4-pro-0813; qwen3.8-2.4t-a95b; ▲ gemini-3.7-flash; qwen3.8-27b
2026-08-19held out53%39%–100%37%deepseek-v4-flash-vision-exp; muse-spark-1.2-contributor
2026-08-26held out74%56%–100%92%qwen3.8-flash; ▲ claude-fable-5.1
2026-09-02held out58%31%–100%85%▲ gemini-3.8-flash; muse-spark-1.3, muse-spark-1.3-contributor; qwen3.8-max-0902; ▲ gpt-6-astra-pro, gpt-6-astra
2026-09-09held out85%59%–100%99%deepseek-v4.1-flash
2026-09-16held out83%28%–100%91%qwen3.8-omni-flash; ▲ grok-4.7; ▲ claude-opus-5.5; ▲ gpt-6-sol, gpt-6-sol-pro +2
2026-09-23held out28%23%–39%5%qwen3.8-max-prime

▸RESULTS: FORECAST VS BASE RATE

V3 REPLAY · HELD OUT 16 Jul 2026 – 26 Sep 2026

One event throughout: some frontier lab (OpenAI, Anthropic, Google DeepMind, xAI, DeepSeek, Alibaba Qwen, Meta) lists a new text model on OpenRouter within the next 24 hours, 72 hours or 7 days. Twin listings of one model, and one lab's listings within 2 hours of each other, count once. Each formula was fitted on 1 Apr 2026 to 16 Jul 2026 and scored on the 1,638 held-out hours after it. The rule was fixed before the scores were seen.

Out-of-sample Brier skill of each candidate formula against the fit-window base rate, with its 95% block-bootstrap interval and whether its reliability passed the sanity check, at each horizon.
FormulaNext 72h · decidesrate 39% → 63%Next 24hrate 17% → 28%Next 7drate 73% → 94%
(a) constant base rate0 the reference0 the reference0 the reference
(b) max over labs−0.309−0.83 to +0.06fails sanity−0.027−0.24 to +0.15fails sanity−1.481−3.36 to −0.28fails sanity
(c) max(market, base)−0.074−0.61 to +0.26fails sanity+0.077−0.14 to +0.24sane+0.037−0.32 to +0.51fails sanity
(d) noisy-OR with unpriced rate u−0.002−0.57 to +0.34fails sanity+0.032−0.25 to +0.23fails sanity+0.054−0.36 to +0.67fails sanity
(e) (d) + logistic recalibration−0.035−0.33 to +0.15fails sanity+0.073−0.04 to +0.19fails sanity−0.015−0.24 to +0.33fails sanity

Skill is 1 − Brier ÷ Brier of (a), the fit window's base rate, on the held-out hours; above zero beats it. The small line is the 95% interval from a block bootstrap (2,000 resamples of 168-hour blocks). "Sane" means the mean forecast is within 0.05 of what happened and every bin with at least 100 hours is within 0.2. Rates are the share of hours with a frontier launch inside the horizon, fit window → held out.

PRE-REGISTERED RULEShip a probability only if the best formula beats the base rate at 72h by more than +0.02 with sane reliability. The best was (d) noisy-OR with unpriced rate u: −0.002 (95% interval −0.566 to +0.344), and its worst bin was off by 0.47. Verdict: keep the hand-weighted lead score; the probability is context only.

Rolling origin at 72h: fit on everything before each origin, test to the next

Test windowLaunchesRate(d)(b)(c)(e)
06-11 → 07-07320%−0.075+0.507+0.048+0.073
07-07 → 08-03865%+0.187−0.120+0.045+0.060
08-03 → 08-301163%+0.122−0.306+0.141+0.074
08-30 → 09-261157%−0.426−0.645−0.527−0.311

(d) scored −0.075, +0.187, +0.122, −0.426 across the 4 folds: the sign flips with the window, which is what a forecast with no stable edge looks like.

Sensitivity at 72h (checked after the fact; none fed the verdict)

Variant(d) skill95% intervalSanity
no-bucketsDay-bucket floors dropped: the low end of the bid proxy+0.190−0.04 to +0.36fails
trustedOnly reads the live trusted-read rule keeps here (no bucket floors, no lower bounds)+0.187+0.01 to +0.35fails
no-settleMarkets read until they close instead of stopping at the launch+0.021−0.55 to +0.38fails
uncensoredTest ends 14 days before the pull, so no late rungs are missing+0.197−0.02 to +0.38fails

The best variant reaches +0.197 against the primary read's −0.002, and none passes the sanity check. Dropping day-bucket prices alone moves it by +0.19: they stand in for bids, and with no historical bid or ask the history can't say which read is right.

Reliability: when a forecast says X%, does X% happen?

Left, the v3 forecast on held-out hours. Across forecasts from 27% to 98%, a launch followed in 49% to 81% of hours: the rate hardly tracks the forecast, and the top bin came true 50% of the time. Right, Polymarket on its own questions, from wave 1 and still valid: each rung's price 72 hours before its deadline against how that rung resolved (Brier skill +0.512 over 325 rungs). The markets answer the question they ask, whether one named model ships by a date; that is not a forecast of the next frontier launch.

v3 · (d) next 72h · held-out hours
00252550507575100100perfectheld-out rate 63%Forecast 20%–30%: 163 hours, mean forecast 27%, came true 49% of the timeForecast 30%–40%: 154 hours, mean forecast 35%, came true 60% of the timeForecast 40%–50%: 239 hours, mean forecast 46%, came true 64% of the timeForecast 50%–60%: 248 hours, mean forecast 55%, came true 75% of the timeForecast 60%–70%: 259 hours, mean forecast 65%, came true 78% of the timeForecast 70%–80%: 197 hours, mean forecast 74%, came true 52% of the timeForecast 80%–90%: 101 hours, mean forecast 85%, came true 81% of the timeForecast 90%–100%: 277 hours, mean forecast 98%, came true 50% of the timeForecast 20%–30%: 163 hours, mean forecast 27%, came true 49% of the time163Forecast 30%–40%: 154 hours, mean forecast 35%, came true 60% of the time154Forecast 40%–50%: 239 hours, mean forecast 46%, came true 64% of the time239Forecast 50%–60%: 248 hours, mean forecast 55%, came true 75% of the time248Forecast 60%–70%: 259 hours, mean forecast 65%, came true 78% of the time259Forecast 70%–80%: 197 hours, mean forecast 74%, came true 52% of the time197Forecast 80%–90%: 101 hours, mean forecast 85%, came true 81% of the time101Forecast 90%–100%: 277 hours, mean forecast 98%, came true 50% of the time277forecast, %came true, %hours
Wave 1 · Polymarket rungs 72h before deadline
00252550507575100100perfectresolved Yes 8%Forecast 0%–20%: 269 rungs, mean forecast 3%, came true 1% of the timeForecast 20%–40%: 31 rungs, mean forecast 30%, came true 19% of the timeForecast 40%–60%: 11 rungs, mean forecast 48%, came true 36% of the timeForecast 60%–80%: 8 rungs, mean forecast 71%, came true 88% of the timeForecast 80%–100%: 6 rungs, mean forecast 90%, came true 100% of the timeForecast 0%–20%: 269 rungs, mean forecast 3%, came true 1% of the time269Forecast 20%–40%: 31 rungs, mean forecast 30%, came true 19% of the time31Forecast 40%–60%: 11 rungs, mean forecast 48%, came true 36% of the time11Forecast 60%–80%: 8 rungs, mean forecast 71%, came true 88% of the time8Forecast 80%–100%: 6 rungs, mean forecast 90%, came true 100% of the time6forecast, %came true, %rungs
The plotted bins as a table
v3 · (d) next 72h · held-out hours · 1638 hours
ForecasthoursMean forecastCame true
20%–30%16326.5%48.5%
30%–40%15435.3%59.7%
40%–50%23945.5%63.6%
50%–60%24855.1%75.0%
60%–70%25964.8%78.0%
70%–80%19774.4%51.8%
80%–90%10185.3%81.2%
90%–100%27797.5%50.2%
Wave 1 · Polymarket rungs 72h before deadline · 325 rungs
ForecastrungsMean forecastCame true
0%–20%2692.6%1.1%
20%–40%3129.7%19.4%
40%–60%1147.7%36.4%
60%–80%870.7%87.5%
80%–100%689.6%100.0%

Wave 1 · Polymarket's own calibration

Polymarket's own calibration: each rung's price a fixed time before its deadline against how the rung resolved.
Priced before deadlineSkillBrierRungsResolved Yes
24h+0.4470.03753287%
72h+0.5120.03603258%

Skill against always forecasting that horizon's Yes rate, over every frontier-lab rung in the wave-1 pull. On their own questions the markets beat their base rate. The trouble is the question: a rung asks whether one named model ships by a date, so it sits near zero while other frontier models list, and it moves after leaks rather than ahead of them.

By lab: where named markets exist

Each frontier lab's own 72-hour market read scored against that lab's own launches on the held-out window.
LabSkillPriced firstRate, fit → held outHours with a read
Anthropic+0.4843 of 312% → 13%637
OpenAI+0.3052 of 29% → 9%372
Meta+0.0380 of 40% → 18%93
Google DeepMind+0.0372 of 38% → 13%729
DeepSeek−0.0560 of 43% → 18%0
Alibaba Qwen−0.1090 of 710% → 28%342
xAI−0.7082 of 29% → 9%585

Each lab's own 72-hour read against its own launches, scored on the held-out hours against that lab's fit-window rate. "Priced first" counts held-out launches whose lab's read reached 50% in the 72 hours before the listing (9 of 25 overall). "Hours with a read" counts held-out hours with a read above 5%. With 2 to 7 launches per lab, a single launch moves these numbers a lot, so read them as where the markets exist, not as a ranking.

Had the level been the 72-hour probability

Held-out hours per probability band, their mean forecast, and how often a frontier launch followed within 72 hours.
Level · bandLaunch within 72hMean forecastHeld-out hours
1 · ≥95%40%99%210 (13%)
2 · 70–95%66%81%365 (22%)
3 · 45–70%72%57%655 (40%)
4 · 30–45%65%38%245 (15%)
5 · <30%49%27%163 (10%)

Bands sit at the fit window's 30%, 45%, 70%, 95% edges. On held-out hours the launch rate did not rise with the level: level 1 hours were followed by a launch 40% of the time, level 3 hours 72%. A level that doesn't rank outcomes can't be sold as a probability.

Limits of the replay

  • The sample is launches, not hours. 1,638 held-out hours hold only 25 launches, and one launch marks up to 72 hours in a row as a hit. That is why every interval above is wide.
  • The release rate moved. A launch followed within 72 hours in 39% of fit-window hours and 63% of held-out hours, so every constant fitted on the first window is off on the second.
  • No historical bid or ask. Every past quote is treated as a tight book, so the live thin-book gate can't be replayed, and day-bucket prices stand in for bids.
  • Survivorship. The outcome counts only models OpenRouter still lists. All 21 of 21 hand-curated frontier launches in the window still are, and the delistings OpenRouter had scheduled on 2026-09-26 would erase 0 of its launches. Every xAI listing before 2026-03-31 is gone, so the window starts 2026-04-01. Inside it, a release event disappears only when all of its listings are delisted.
  • Censoring. Rungs due after the pull weren't pulled, so market reads in the last 14 days are censored low. Ending the test 14 days early gives (d) +0.197, still failing the sanity check; it did not feed the verdict.
  • Launched markets are cut off. Each market is read only until its model launched. Live, a launched model's market kept trading near 1 for a median 2.6 hours (up to 63.3 hours, over 39 markets); the dashboard drops a family for 4 days after its model lists.
Every limitation as the replay records it
  • Hourly rows are heavily autocorrelated: one release event sets y=1 for h consecutive hours, so the effective sample size is the number of distinct release events (test: tens), not the number of hours (thousands).
  • The release rate is not stationary: the 72h rate was 0.39 in the train window and 0.631 in the test window, so every constant fitted on train is miscalibrated on test.
  • No historical bid/ask: every quote is treated as a tight book, so the live thin-book gate (spread > 10¢) cannot be replayed; date-bucket floors use the price in place of the best bid (an upper bound on the live floor; no-buckets is the lower bound). PAV weights are uniform (no historical liquidity).
  • The CLOB pull kept the last 15 days of each rung, so a rung is admitted only within 14 days of its deadline. The live trusted-read rule mainly guards reads extrapolated from rungs further out; those cannot be replayed without leaking outcomes, so the trusted variant here only drops lower-bound and bucket-floor reads.
  • Rungs due after the pull are missing (unresolved rungs were not pulled), so market reads in the last 14 days are censored low; the uncensored variant stops the test there.
  • OpenRouter survivorship: the outcome counts only models OpenRouter still lists (see survivorship).
  • Markets are read only until their model launched (announcement, else first YES resolution). Live markets keep trading near 1 until they resolve (median 2.6h after the announcement, up to 63.3h, over 39 markets); the no-settle variant scores an unguarded reader.

▸LEAD TIMES

WAVE 1 · 23 LAUNCHES TIMED BY HAND

The announcement is the first Hacker News story linking the lab's own site or official account and naming the model; teasers ("launching soon", "this Thursday") don't count, and neither do the few earlier first-party links the table marks "not the launch", each with its reason. Availability is OpenRouter's created time. Market crossings are measured against the announcement.

−10d−3d−1d−6h0+6h+1d+2dANNOUNCEDMARKET GOT THERE FIRSTAFTERQwen3.5no marketQwen3.5: listed on OpenRouter 2026-02-16 06:23Z (qwen/qwen3.5-397b-a17b), −3.1h after the announcementGemma 4no marketGemma 4: listed on OpenRouter 2026-04-02 16:48Z (google/gemma-4-31b-it), +0.8h after the announcementClaude Opus 4.7Claude Opus 4.7: cumulative rung "On or prior to April 16" held 0.5 from 2026-04-15 19:00Z, +19.4h relative to the announcement (+ = before)Claude Opus 4.7: cumulative rung "On or prior to April 16" held 0.7 from 2026-04-15 20:00Z, +18.4h relative to the announcement (+ = before)Claude Opus 4.7: cumulative rung "On or prior to April 16" held 0.9 from 2026-04-16 15:00Z, −0.6h relative to the announcement (+ = before)Claude Opus 4.7: listed on OpenRouter 2026-04-16 14:51Z (anthropic/claude-opus-4.7), +0.5h after the announcementGPT-5.5GPT-5.5: day market "April 23" held 0.5 from 2026-04-14 04:00Z, +9.6d relative to the announcement (+ = before)GPT-5.5: day market "April 23" held 0.7 from 2026-04-14 16:00Z, +9.1d relative to the announcement (+ = before)GPT-5.5: day market "April 23" held 0.9 from 2026-04-22 07:00Z, +35.0h relative to the announcement (+ = before)GPT-5.5: listed on OpenRouter 2026-04-24 17:31Z (openai/gpt-5.5), +23.5h after the announcementDeepSeek V4DeepSeek V4: cumulative rung "April 24" held 0.5 from 2026-04-23 19:00Z, +7.9h relative to the announcement (+ = before)DeepSeek V4: cumulative rung "April 24" held 0.7 from 2026-04-24 03:00Z, −0.1h relative to the announcement (+ = before)DeepSeek V4: cumulative rung "April 24" held 0.9 from 2026-04-24 05:00Z, −2.1h relative to the announcement (+ = before)DeepSeek V4: listed on OpenRouter 2026-04-24 03:17Z (deepseek/deepseek-v4-flash), +0.4h after the announcementGemini 3.5 FlashGemini 3.5 Flash: day market "May 19" held 0.5 from 2026-05-09 04:00Z, +10.6d relative to the announcement (+ = before)Gemini 3.5 Flash: day market "May 19" held 0.7 from 2026-05-09 11:00Z, +10.3d relative to the announcement (+ = before)Gemini 3.5 Flash: day market "May 19" held 0.9 from 2026-05-13 00:00Z, +6.7d relative to the announcement (+ = before)Gemini 3.5 Flash: listed on OpenRouter 2026-05-19 12:30Z (google/gemini-3.5-flash), −4.9h after the announcementClaude Opus 4.8no dated marketClaude Opus 4.8: listed on OpenRouter 2026-05-27 18:04Z (anthropic/claude-opus-4.8), −22.7h after the announcementClaude Fable 5 / Mythos 5Claude Fable 5 / Mythos 5: cumulative rung "June 9" held 0.5 from 2026-06-09 07:00Z, +10.0h relative to the announcement (+ = before); already above at its first usable priceClaude Fable 5 / Mythos 5: cumulative rung "June 9" held 0.7 from 2026-06-09 15:00Z, +2.0h relative to the announcement (+ = before)Claude Fable 5 / Mythos 5: cumulative rung "June 9" held 0.9 from 2026-06-09 17:00Z, 0.0h relative to the announcement (+ = before)Claude Fable 5 / Mythos 5: listed on OpenRouter 2026-06-09 12:18Z (anthropic/claude-fable-5), −4.7h after the announcementClaude Sonnet 5Claude Sonnet 5: cumulative rung "June 30" held 0.5 from 2026-06-25 08:00Z, +5.3d relative to the announcement (+ = before)Claude Sonnet 5: cumulative rung "June 30" held 0.7 from 2026-06-25 08:00Z, +5.3d relative to the announcement (+ = before)Claude Sonnet 5: cumulative rung "June 30" held 0.9 from 2026-06-30 18:00Z, −3.0h relative to the announcement (+ = before)Claude Sonnet 5: listed on OpenRouter 2026-06-30 18:11Z (anthropic/claude-sonnet-5), +3.2h after the announcementGrok 4.5Grok 4.5: cumulative rung "July 8" held 0.5 from 2026-07-08 13:00Z, +5.0h relative to the announcement (+ = before)Grok 4.5: cumulative rung "July 8" held 0.7 from 2026-07-08 13:00Z, +5.0h relative to the announcement (+ = before)Grok 4.5: cumulative rung "July 8" held 0.9 from 2026-07-08 18:00Z, 0.0h relative to the announcement (+ = before)Grok 4.5: listed on OpenRouter 2026-07-08 15:05Z (x-ai/grok-4.5), −2.9h after the announcementGPT-5.6 Sol/Terra/LunaGPT-5.6 Sol/Terra/Luna: day market "July 9" held 0.5 from 2026-07-05 10:00Z, +4.3d relative to the announcement (+ = before)GPT-5.6 Sol/Terra/Luna: day market "July 9" held 0.7 from 2026-07-06 01:00Z, +3.7d relative to the announcement (+ = before)GPT-5.6 Sol/Terra/Luna: day market "July 9" held 0.9 from 2026-07-08 05:00Z, +36.1h relative to the announcement (+ = before)GPT-5.6 Sol/Terra/Luna: listed on OpenRouter 2026-07-09 09:54Z (openai/gpt-5.6-sol), −7.2h after the announcementKimi K3no dated marketKimi K3: listed on OpenRouter 2026-07-16 15:30Z (moonshotai/kimi-k3), +0.7h after the announcementClaude Opus 5Claude Opus 5: day market "July 24" held 0.5 from 2026-07-24 17:00Z, −0.1h relative to the announcement (+ = before)Claude Opus 5: day market "July 24" held 0.7 from 2026-07-24 17:00Z, −0.1h relative to the announcement (+ = before)Claude Opus 5: day market "July 24" held 0.9 from 2026-07-24 17:00Z, −0.1h relative to the announcement (+ = before)Claude Opus 5: listed on OpenRouter 2026-07-24 17:02Z (anthropic/claude-opus-5), +0.1h after the announcementGrok 4.6Grok 4.6: day market "August 12" held 0.5 from 2026-08-12 16:00Z, −0.5h relative to the announcement (+ = before)Grok 4.6: day market "August 12" held 0.7 from 2026-08-12 16:00Z, −0.5h relative to the announcement (+ = before)Grok 4.6: day market "August 12" held 0.9 from 2026-08-12 16:00Z, −0.5h relative to the announcement (+ = before)Grok 4.6: listed on OpenRouter 2026-08-12 15:35Z (x-ai/grok-4.6), +0.1h after the announcementGemini 3.7 FlashGemini 3.7 Flash: cumulative rung "August 13" held 0.5 from 2026-08-13 16:00Z, +1.0h relative to the announcement (+ = before); already above at its first usable priceGemini 3.7 Flash: cumulative rung "August 13" held 0.7 from 2026-08-13 16:00Z, +1.0h relative to the announcement (+ = before); already above at its first usable priceGemini 3.7 Flash: cumulative rung "August 13" held 0.9 from 2026-08-13 18:00Z, −1.0h relative to the announcement (+ = before)Gemini 3.7 Flash: listed on OpenRouter 2026-08-13 17:03Z (google/gemini-3.7-flash), 0.0h after the announcementQwen3.8-Flashno dated marketQwen3.8-Flash: listed on OpenRouter 2026-08-26 19:37Z (qwen/qwen3.8-flash), +7.0h after the announcementClaude Fable 5.1 / Mythos 5.1Claude Fable 5.1 / Mythos 5.1: day market "September 1" held 0.5 from 2026-08-31 12:00Z, +29.9h relative to the announcement (+ = before)Claude Fable 5.1 / Mythos 5.1: day market "September 1" held 0.7 from 2026-09-01 18:00Z, −0.1h relative to the announcement (+ = before)Claude Fable 5.1 / Mythos 5.1: day market "September 1" held 0.9 from 2026-09-01 18:00Z, −0.1h relative to the announcement (+ = before)Claude Fable 5.1 / Mythos 5.1: listed on OpenRouter 2026-09-01 18:03Z (anthropic/claude-fable-5.1), +0.2h after the announcementGemini 3.8 FlashGemini 3.8 Flash: day market "September 2" held 0.5 from 2026-09-01 08:00Z, +31.2h relative to the announcement (+ = before)Gemini 3.8 Flash: day market "September 2" held 0.7 from 2026-09-01 09:00Z, +30.2h relative to the announcement (+ = before)Gemini 3.8 Flash: day market "September 2" held 0.9 from 2026-09-02 08:00Z, +7.2h relative to the announcement (+ = before)Gemini 3.8 Flash: listed on OpenRouter 2026-09-02 15:14Z (google/gemini-3.8-flash), 0.0h after the announcementGPT-6 AstraGPT-6 Astra: day market "September 4" held 0.5 from 2026-09-04 19:00Z, −24.8h relative to the announcement (+ = before)GPT-6 Astra: day market "September 4" held 0.7 from 2026-09-04 19:00Z, −24.8h relative to the announcement (+ = before)GPT-6 Astra: day market "September 4" held 0.9 from 2026-09-04 20:00Z, −25.8h relative to the announcement (+ = before)GPT-6 Astra: listed on OpenRouter 2026-09-04 20:13Z (openai/gpt-6-astra-pro), +26.0h after the announcementDeepSeek V4.1 Flashno dated marketDeepSeek V4.1 Flash: listed on OpenRouter 2026-09-10 06:21Z (deepseek/deepseek-v4.1-flash), +0.4h after the announcementGrok 4.7Grok 4.7: day market "September 21" held 0.5 from 2026-09-21 15:00Z, +0.8h relative to the announcement (+ = before)Grok 4.7: day market "September 21" held 0.7 from 2026-09-21 16:00Z, −0.2h relative to the announcement (+ = before)Grok 4.7: day market "September 21" held 0.9 from 2026-09-21 16:00Z, −0.2h relative to the announcement (+ = before)Grok 4.7: listed on OpenRouter 2026-09-21 16:19Z (x-ai/grok-4.7), +0.5h after the announcementClaude Opus 5.5Claude Opus 5.5: day market "September 22" held 0.5 from 2026-09-21 16:00Z, +24.5h relative to the announcement (+ = before)Claude Opus 5.5: day market "September 22" held 0.7 from 2026-09-21 19:00Z, +21.5h relative to the announcement (+ = before)Claude Opus 5.5: day market "September 22" held 0.9 from 2026-09-22 16:00Z, +0.5h relative to the announcement (+ = before)Claude Opus 5.5: listed on OpenRouter 2026-09-22 16:32Z (anthropic/claude-opus-5.5), +0.1h after the announcementGPT-6 Sol/LunaGPT-6 Sol/Luna: cumulative rung "On or prior to September 22" held 0.5 from 2026-09-22 11:00Z, +7.0h relative to the announcement (+ = before)GPT-6 Sol/Luna: cumulative rung "On or prior to September 22" held 0.7 from 2026-09-22 12:00Z, +6.0h relative to the announcement (+ = before)GPT-6 Sol/Luna: cumulative rung "On or prior to September 22" held 0.9 from 2026-09-22 18:00Z, 0.0h relative to the announcement (+ = before)GPT-6 Sol/Luna: listed on OpenRouter 2026-09-22 18:12Z (openai/gpt-6-sol), +0.2h after the announcement−10d−1d0+1dANNOUNCEDEARLIERAFTERQwen3.5no marketQwen3.5: listed on OpenRouter 2026-02-16 06:23Z (qwen/qwen3.5-397b-a17b), −3.1h after the announcementGemma 4no marketGemma 4: listed on OpenRouter 2026-04-02 16:48Z (google/gemma-4-31b-it), +0.8h after the announcementClaude Opus 4.7Claude Opus 4.7: cumulative rung "On or prior to April 16" held 0.5 from 2026-04-15 19:00Z, +19.4h relative to the announcement (+ = before)Claude Opus 4.7: cumulative rung "On or prior to April 16" held 0.7 from 2026-04-15 20:00Z, +18.4h relative to the announcement (+ = before)Claude Opus 4.7: cumulative rung "On or prior to April 16" held 0.9 from 2026-04-16 15:00Z, −0.6h relative to the announcement (+ = before)Claude Opus 4.7: listed on OpenRouter 2026-04-16 14:51Z (anthropic/claude-opus-4.7), +0.5h after the announcementGPT-5.5GPT-5.5: day market "April 23" held 0.5 from 2026-04-14 04:00Z, +9.6d relative to the announcement (+ = before)GPT-5.5: day market "April 23" held 0.7 from 2026-04-14 16:00Z, +9.1d relative to the announcement (+ = before)GPT-5.5: day market "April 23" held 0.9 from 2026-04-22 07:00Z, +35.0h relative to the announcement (+ = before)GPT-5.5: listed on OpenRouter 2026-04-24 17:31Z (openai/gpt-5.5), +23.5h after the announcementDeepSeek V4DeepSeek V4: cumulative rung "April 24" held 0.5 from 2026-04-23 19:00Z, +7.9h relative to the announcement (+ = before)DeepSeek V4: cumulative rung "April 24" held 0.7 from 2026-04-24 03:00Z, −0.1h relative to the announcement (+ = before)DeepSeek V4: cumulative rung "April 24" held 0.9 from 2026-04-24 05:00Z, −2.1h relative to the announcement (+ = before)DeepSeek V4: listed on OpenRouter 2026-04-24 03:17Z (deepseek/deepseek-v4-flash), +0.4h after the announcementGemini 3.5 FlashGemini 3.5 Flash: day market "May 19" held 0.5 from 2026-05-09 04:00Z, +10.6d relative to the announcement (+ = before)Gemini 3.5 Flash: day market "May 19" held 0.7 from 2026-05-09 11:00Z, +10.3d relative to the announcement (+ = before)Gemini 3.5 Flash: day market "May 19" held 0.9 from 2026-05-13 00:00Z, +6.7d relative to the announcement (+ = before)Gemini 3.5 Flash: listed on OpenRouter 2026-05-19 12:30Z (google/gemini-3.5-flash), −4.9h after the announcementClaude Opus 4.8no dated marketClaude Opus 4.8: listed on OpenRouter 2026-05-27 18:04Z (anthropic/claude-opus-4.8), −22.7h after the announcementClaude Fable 5 / Mythos 5Claude Fable 5 / Mythos 5: cumulative rung "June 9" held 0.5 from 2026-06-09 07:00Z, +10.0h relative to the announcement (+ = before); already above at its first usable priceClaude Fable 5 / Mythos 5: cumulative rung "June 9" held 0.7 from 2026-06-09 15:00Z, +2.0h relative to the announcement (+ = before)Claude Fable 5 / Mythos 5: cumulative rung "June 9" held 0.9 from 2026-06-09 17:00Z, 0.0h relative to the announcement (+ = before)Claude Fable 5 / Mythos 5: listed on OpenRouter 2026-06-09 12:18Z (anthropic/claude-fable-5), −4.7h after the announcementClaude Sonnet 5Claude Sonnet 5: cumulative rung "June 30" held 0.5 from 2026-06-25 08:00Z, +5.3d relative to the announcement (+ = before)Claude Sonnet 5: cumulative rung "June 30" held 0.7 from 2026-06-25 08:00Z, +5.3d relative to the announcement (+ = before)Claude Sonnet 5: cumulative rung "June 30" held 0.9 from 2026-06-30 18:00Z, −3.0h relative to the announcement (+ = before)Claude Sonnet 5: listed on OpenRouter 2026-06-30 18:11Z (anthropic/claude-sonnet-5), +3.2h after the announcementGrok 4.5Grok 4.5: cumulative rung "July 8" held 0.5 from 2026-07-08 13:00Z, +5.0h relative to the announcement (+ = before)Grok 4.5: cumulative rung "July 8" held 0.7 from 2026-07-08 13:00Z, +5.0h relative to the announcement (+ = before)Grok 4.5: cumulative rung "July 8" held 0.9 from 2026-07-08 18:00Z, 0.0h relative to the announcement (+ = before)Grok 4.5: listed on OpenRouter 2026-07-08 15:05Z (x-ai/grok-4.5), −2.9h after the announcementGPT-5.6 Sol/Terra/LunaGPT-5.6 Sol/Terra/Luna: day market "July 9" held 0.5 from 2026-07-05 10:00Z, +4.3d relative to the announcement (+ = before)GPT-5.6 Sol/Terra/Luna: day market "July 9" held 0.7 from 2026-07-06 01:00Z, +3.7d relative to the announcement (+ = before)GPT-5.6 Sol/Terra/Luna: day market "July 9" held 0.9 from 2026-07-08 05:00Z, +36.1h relative to the announcement (+ = before)GPT-5.6 Sol/Terra/Luna: listed on OpenRouter 2026-07-09 09:54Z (openai/gpt-5.6-sol), −7.2h after the announcementKimi K3no dated marketKimi K3: listed on OpenRouter 2026-07-16 15:30Z (moonshotai/kimi-k3), +0.7h after the announcementClaude Opus 5Claude Opus 5: day market "July 24" held 0.5 from 2026-07-24 17:00Z, −0.1h relative to the announcement (+ = before)Claude Opus 5: day market "July 24" held 0.7 from 2026-07-24 17:00Z, −0.1h relative to the announcement (+ = before)Claude Opus 5: day market "July 24" held 0.9 from 2026-07-24 17:00Z, −0.1h relative to the announcement (+ = before)Claude Opus 5: listed on OpenRouter 2026-07-24 17:02Z (anthropic/claude-opus-5), +0.1h after the announcementGrok 4.6Grok 4.6: day market "August 12" held 0.5 from 2026-08-12 16:00Z, −0.5h relative to the announcement (+ = before)Grok 4.6: day market "August 12" held 0.7 from 2026-08-12 16:00Z, −0.5h relative to the announcement (+ = before)Grok 4.6: day market "August 12" held 0.9 from 2026-08-12 16:00Z, −0.5h relative to the announcement (+ = before)Grok 4.6: listed on OpenRouter 2026-08-12 15:35Z (x-ai/grok-4.6), +0.1h after the announcementGemini 3.7 FlashGemini 3.7 Flash: cumulative rung "August 13" held 0.5 from 2026-08-13 16:00Z, +1.0h relative to the announcement (+ = before); already above at its first usable priceGemini 3.7 Flash: cumulative rung "August 13" held 0.7 from 2026-08-13 16:00Z, +1.0h relative to the announcement (+ = before); already above at its first usable priceGemini 3.7 Flash: cumulative rung "August 13" held 0.9 from 2026-08-13 18:00Z, −1.0h relative to the announcement (+ = before)Gemini 3.7 Flash: listed on OpenRouter 2026-08-13 17:03Z (google/gemini-3.7-flash), 0.0h after the announcementQwen3.8-Flashno dated marketQwen3.8-Flash: listed on OpenRouter 2026-08-26 19:37Z (qwen/qwen3.8-flash), +7.0h after the announcementClaude Fable 5.1 / Mythos 5.1Claude Fable 5.1 / Mythos 5.1: day market "September 1" held 0.5 from 2026-08-31 12:00Z, +29.9h relative to the announcement (+ = before)Claude Fable 5.1 / Mythos 5.1: day market "September 1" held 0.7 from 2026-09-01 18:00Z, −0.1h relative to the announcement (+ = before)Claude Fable 5.1 / Mythos 5.1: day market "September 1" held 0.9 from 2026-09-01 18:00Z, −0.1h relative to the announcement (+ = before)Claude Fable 5.1 / Mythos 5.1: listed on OpenRouter 2026-09-01 18:03Z (anthropic/claude-fable-5.1), +0.2h after the announcementGemini 3.8 FlashGemini 3.8 Flash: day market "September 2" held 0.5 from 2026-09-01 08:00Z, +31.2h relative to the announcement (+ = before)Gemini 3.8 Flash: day market "September 2" held 0.7 from 2026-09-01 09:00Z, +30.2h relative to the announcement (+ = before)Gemini 3.8 Flash: day market "September 2" held 0.9 from 2026-09-02 08:00Z, +7.2h relative to the announcement (+ = before)Gemini 3.8 Flash: listed on OpenRouter 2026-09-02 15:14Z (google/gemini-3.8-flash), 0.0h after the announcementGPT-6 AstraGPT-6 Astra: day market "September 4" held 0.5 from 2026-09-04 19:00Z, −24.8h relative to the announcement (+ = before)GPT-6 Astra: day market "September 4" held 0.7 from 2026-09-04 19:00Z, −24.8h relative to the announcement (+ = before)GPT-6 Astra: day market "September 4" held 0.9 from 2026-09-04 20:00Z, −25.8h relative to the announcement (+ = before)GPT-6 Astra: listed on OpenRouter 2026-09-04 20:13Z (openai/gpt-6-astra-pro), +26.0h after the announcementDeepSeek V4.1 Flashno dated marketDeepSeek V4.1 Flash: listed on OpenRouter 2026-09-10 06:21Z (deepseek/deepseek-v4.1-flash), +0.4h after the announcementGrok 4.7Grok 4.7: day market "September 21" held 0.5 from 2026-09-21 15:00Z, +0.8h relative to the announcement (+ = before)Grok 4.7: day market "September 21" held 0.7 from 2026-09-21 16:00Z, −0.2h relative to the announcement (+ = before)Grok 4.7: day market "September 21" held 0.9 from 2026-09-21 16:00Z, −0.2h relative to the announcement (+ = before)Grok 4.7: listed on OpenRouter 2026-09-21 16:19Z (x-ai/grok-4.7), +0.5h after the announcementClaude Opus 5.5Claude Opus 5.5: day market "September 22" held 0.5 from 2026-09-21 16:00Z, +24.5h relative to the announcement (+ = before)Claude Opus 5.5: day market "September 22" held 0.7 from 2026-09-21 19:00Z, +21.5h relative to the announcement (+ = before)Claude Opus 5.5: day market "September 22" held 0.9 from 2026-09-22 16:00Z, +0.5h relative to the announcement (+ = before)Claude Opus 5.5: listed on OpenRouter 2026-09-22 16:32Z (anthropic/claude-opus-5.5), +0.1h after the announcementGPT-6 Sol/LunaGPT-6 Sol/Luna: cumulative rung "On or prior to September 22" held 0.5 from 2026-09-22 11:00Z, +7.0h relative to the announcement (+ = before)GPT-6 Sol/Luna: cumulative rung "On or prior to September 22" held 0.7 from 2026-09-22 12:00Z, +6.0h relative to the announcement (+ = before)GPT-6 Sol/Luna: cumulative rung "On or prior to September 22" held 0.9 from 2026-09-22 18:00Z, 0.0h relative to the announcement (+ = before)GPT-6 Sol/Luna: listed on OpenRouter 2026-09-22 18:12Z (openai/gpt-6-sol), +0.2h after the announcement
Each row is one launch, in date order. Circles mark when the market price first reached 0.5, 0.7 and 0.9 and stayed there for two hours. They are read from the day market when there is one, otherwise from the tightest "released by" rung that resolved Yes and was due within two days of the launch. The time axis is compressed, linear within a few hours of the announcement and logarithmic beyond. OpenRouter availability is plotted separately because a listing is not a forecast.
Release events with announcement, availability and market crossing times. Leads are hours before the announcement; negative means after.
LaunchAnnounced (first first-party HN story)OpenRouterBefore launchMarket read0.50.70.9
Qwen3.5Alibaba Qwen2026-02-16 09:32ZQwen3.5: Towards Native Multimodal AgentsHN−3.1hno HN precursorNo pre-launch story on HN. The transformers architecture merged 6.9 days earlier (see below).transformers #43830no market———
Gemma 4Google DeepMind2026-04-02 16:00ZGemma 4: Byte for byte, the most capable open modelsHN+0.8hno HN precursorNo pre-launch story on HN; the transformers merge landed 36 minutes before the blog.transformers #45192no market———
Claude Opus 4.7Anthropic2026-04-16 14:23ZClaude Opus 4.7HN+0.5hno HN precursorNo pre-launch story on HN.On or prior to April 16+19.4h+18.4h−0.6h
GPT-5.5OpenAI2026-04-23 18:01ZGPT-5.5HN+23.5hleakedA Codex build exposed it 38 hours early; the API (and OpenRouter) followed a day after the post.HN: "GPT 5.5 Released in Codex", 2026-04-22 04:12Zon April 23+9.6d+9.1d+35.0h
DeepSeek V4DeepSeek2026-04-24 02:55ZDeepSeek-V4 Technical Report [pdf]HN+0.4hno HN precursorNo pre-launch story on HN; weights and API went live together.by April 24+7.9h−0.1h−2.1h
Gemini 3.5 FlashGoogle DeepMind2026-05-19 17:25ZGemini 3.5 FlashHN−4.9hno HN precursorLaunched on stage at the Google I/O keynote, with no model-specific story on HN beforehand. It settled the "Gemini 3.2" and "Gemini 3.5" ladders.on May 19+10.6d+10.3d+6.7d
Claude Opus 4.8Anthropic2026-05-28 16:49ZClaude Opus 4.8HN−22.7hleakedA same-morning rumour. OpenRouter lists it 22.7 hours before the post.HN: "Claude Opus 4.8 coming today?", 2026-05-28 09:52Zonly "May 31" (due +3.5d after)———
Claude Fable 5 / Mythos 5Anthropic2026-06-09 16:58ZClaude Fable 5HN−4.7hleakedAn unsourced "releasing tomorrow" post the evening before.HN: "Claude Fable 5 by Anthropic, releasing tomorrow", 2026-06-08 19:36Zby June 9+10.0h*+2.0h0.0h
Claude Sonnet 5Anthropic2026-06-30 15:01ZClaude Sonnet 5HN+3.2hno HN precursorNo pre-launch story on HN.by June 30+5.3d+5.3d−3.0h
Grok 4.5xAI2026-07-08 18:00ZGrok 4.5HN−2.9hno HN precursorSettled the "Grok 4.4 released by" ladder.by July 8+5.0h+5.0h0.0h
GPT-5.6 Sol/Terra/LunaOpenAI2026-07-09 17:03ZGPT 5.6 System CardHN−7.2hpre-announcedOpenAI named the day 37 hours ahead; a YouTube placeholder went up the evening before.HN: OpenAI "will launch publicly this Thursday", 2026-07-08 04:12ZHN: "GPT-5.6 Sol Ultra will be in Codex", 2026-07-06 01:04Zon July 9+4.3d+3.7d+36.1h
Kimi K3Moonshot Kimi2026-07-16 14:46ZKimi K3: Open Frontier IntelligenceHN+0.7hno HN precursorWent live in the Kimi app about an hour before the blog. Not a frontier lab in lab.ts.HN: "Kimi K3 released on web and app", 2026-07-16 13:48Zonly "July 31" (due +15.5d after)———
Claude Opus 5Anthropic2026-07-24 16:55ZClaude Opus 5HN+0.1hleakedAn Artificial Analysis model page for it was on HN 23 hours before launch.HN: artificialanalysis.ai/models/claude-opus-5, 2026-07-23 18:01Zon July 24−0.1h−0.1h−0.1h
Grok 4.6xAI2026-08-12 15:32ZGrok 4.6HN+0.1hno HN precursorNo pre-launch story on HN.on August 12−0.5h−0.5h−0.5h
Gemini 3.7 FlashGoogle DeepMind2026-08-13 17:02ZGemini 3.7 FlashHN0.0hno HN precursorNo pre-launch story on HN.by August 13+1.0h*+1.0h*−1.0h
Qwen3.8-FlashAlibaba Qwen2026-08-26 12:36ZQwen/Qwen3.8-Flash-NextHN+7.0hpre-announcedQwen’s ModelScope page said "releasing tomorrow" 25 hours ahead.HN: "Qwen 3.8-Flash-Next releasing tomorrow" (modelscope.cn), 2026-08-25 11:49Zonly "September 30" (due +35.6d after)———
Claude Fable 5.1 / Mythos 5.1Anthropic2026-09-01 17:53ZPrompting Claude Fable 5.1 and Claude Mythos 5.1HN+0.2hno HN precursorNo pre-launch story on HN.on September 1+29.9h−0.1h−0.1h
Gemini 3.8 FlashGoogle DeepMind2026-09-02 15:12ZGemini 3.8 Flash and 3.8 Flash CyberHN0.0hleakedA WSJ scoop 16 hours before launch.HN: WSJ "New Google AI Model Said to Narrow Gap", 2026-09-01 23:23Zon September 2+31.2h+30.2h+7.2h
GPT-6 AstraOpenAI2026-09-03 18:12ZPlayco cut manual fixes 50% prototyping games with GPT‑6 AstraHN+26.0hpre-announcedOpenAI published "Path to Astra" two days ahead and a wordless teaser video three hours before the launch. Announced Sep 3, generally available Sep 4; the markets resolved on Sep 4.OpenAI RSS: "Path to Astra", 2026-09-01 13:00ZHN: help.openai.com "OpenAI Astra Launching Soon", 2026-09-03 03:02ZNot the launch: "Path to Astra: critical capabilities and frontier safeguards", the pre-launch safety post, two days earlyNot the launch: "Open AI X post on Astra", an @OpenAI post with no words, only a 12-second video (tweet 2095527557924082061)Not the launch: "OpenAI Releases GPT Astra", the generic models page; commenters found no Astra on it yeton September 4−24.8h−24.8h−25.8h
DeepSeek V4.1 FlashDeepSeek2026-09-10 05:57ZDeepSeek-v4.1-ExpHN+0.4hleakedBeta leaks two days out, then a Vercel AI Gateway beta listing 14.5 hours ahead.HN: "v4.1 Flash is now available for internal beta testing", 2026-09-08 08:04ZHN: "DeepSeek v4.1 Flash Beta on Vercel", 2026-09-09 15:28Zonly "September 30" (due +20.9d after)———
Grok 4.7xAI2026-09-21 15:50ZGrok 4.7HN+0.5hleakedA 1-point "launching soon" tweet 3.4 days out. xAI has no leading coverage on this site.HN: "Grok 4.7 Launching Soon", 2026-09-18 06:42Zon September 21+0.8h−0.2h−0.2h
Claude Opus 5.5Anthropic2026-09-22 16:27ZClaude Opus 5.5HN+0.1hleakedNothing on HN, but TestingCatalog reported testing 31.7 hours ahead; the day market moved 35 minutes after it.TestingCatalog: "Anthropic tests Fable 5.2 and Opus 5.5", 2026-09-21 08:45Zon September 22+24.5h+21.5h+0.5h
GPT-6 Sol/LunaOpenAI2026-09-22 17:58ZGPT-6 SolHN+0.2hleakedA Reddit post saw the id on the API 11 days early; TestingCatalog called the day that morning.HN: "GPT-6-sol appeared on OpenAI API", 2026-09-11 20:43ZTestingCatalog: "prepares to launch GPT-6 Sol and Luna today", 2026-09-22 12:07ZOn or prior to September 22+7.0h+6.0h0.0h

OpenRouter is hours from the announcement to the model's created time (negative: listed first). Market leads are hours before the announcement at which the price reached the threshold and held it for two hours; negative means after. An asterisk marks a rung that was already above the threshold at its first usable price, so the lead is the rung's age, not a move.

▸PRE-ANNOUNCED, LEAKED OR NO HN PRECURSOR

WAVE 1

By what was public beforehand

GroupLaunches with a dated marketHeld 0.5 beforeHeld 0.9 beforeMedian 0.5 leadUndated only
Pre-announced by the lab211+39.2h1
Leaked or reported763+10.0h2
No HN precursor871+13.7h1

By lab

GroupLaunches with a dated marketHeld 0.5 beforeHeld 0.9 beforeMedian 0.5 leadUndated only
Anthropic651+22.0h1
OpenAI432+2.3d0
DeepSeek110+7.9h1
Google DeepMind332+31.2h0
xAI320+0.8h0
Moonshot Kimi000—1
Alibaba Qwen000—1

"Undated only" launches were priced only on rungs due weeks after they shipped, which say the launch is coming but not when. Of the 5 launches whose market held 0.9 before the announcement, 4 followed a leak, a press report or the lab's own pre-announcement (the exception: Gemini 3.5 Flash, launched at a scheduled keynote or with no public trail). The markets aggregate that news rather than foresee it.

The "above 50% is usually real" claim

ReadingWith a day or more to spareInside the last day
Research note (script not kept)11 / 24 (46%)33 / 46 (72%)
First rung to fire per ladder12 / 23 (52%)15 / 18 (83%)
Every rung that fired51 / 68 (75%)16 / 19 (84%)

A rung "fires" when a frontier lab's "released by" rung trades above 0.5 within 7 days of its deadline, before the model launched; it counts as real when the rung resolved Yes. Counting the first rung per ladder (55 ladders, 260 rungs) gives 12 of 23, close to the research note's 11 of 24. Counting every rung inflates the rate, because neighbouring rungs of one ladder move together. The note's 33 of 46 near launch does not reproduce under either count, and its script was not kept. On the per-ladder count, a reading above 50% with a day to spare was right 52% of the time.

▸FALSE ALARMS

WAVE 1
False alarms: frontier-lab rungs that resolved No but traded at or above each threshold, and the share of crossings that resolved Yes.
Market typeThresholdNo rungs that crossedPer market-dayYes rungs that crossed before launchCrossings that resolved Yes
"Released by" rungs0.537 / 1260.0265129 / 13478%
"Released by" rungs0.716 / 1260.0114127 / 13489%
"Released by" rungs0.94 / 1260.0029105 / 13496%
"Released on" day buckets0.517 / 4260.00528 / 1032%
"Released on" day buckets0.76 / 4260.00185 / 1046%
"Released on" day buckets0.91 / 4260.00035 / 1083%

Market-days count the traded life of every No rung (1399 days for "by" rungs). A Yes rung counts only prices from before its model was announced, so a post-launch settlement at 0.99 is not a correct call. Every rung's price counts from the end of its first hour, and only once it has moved off the quote the book opened at (often exactly 0.5 before anyone trades).

Loudest false alarms (No rungs that peaked at 0.5 or more)

MarketRungPeakPeak atLaunch came
OpenAI's Astra released on...?on September 30.9662026-09-03 16:00Z1d later†
OpenAI’s Astra released by…?by September 30.9642026-09-03 16:00Z1d later†
Next Claude Opus released by...?by July 230.9202026-07-23 13:00Z1d later
Next Google Gemini Pro Model released by...?by August 140.9152026-08-03 18:00Znot yet
GPT-5.6 released by...?by June 300.9002026-06-19 19:00Z9d later
New Gemini reasoning flagship released by...?by May 220.8952026-05-11 19:00Znot yet
GPT-5.6 released by...?by July 80.8902026-07-03 05:00Z1d later
Next Claude Opus released on...?on July 230.8902026-07-23 13:00Z1d later
Next Google Gemini Pro Model released by...?by August 70.8852026-07-24 15:00Znot yet
Next Grok Model (4.7+) released by...?by September 180.8852026-09-07 23:00Znot yet
New Gemini reasoning flagship released by...?by June 300.8792026-06-16 17:00Znot yet
Next Google Gemini Pro Model released by...?by June 300.8502026-06-16 12:00Znot yet

54 No rungs peaked at 0.5 or more; the 12 highest are listed. "Launch came" is the gap to the ladder's first Yes deadline. † The launch was announced inside this rung, which still resolved No under the market's own release rule (OpenAI's Astra released on...?, OpenAI’s Astra released by…?). Those are rule misses more than timing misses.

▸STEALTH REVEALS

EARLY WARNING · NEVER SCORED
Revealed slots
11
verified, both timestamps checked
Median days in stealth
7.0
stealth listing to official listing
Frontier-lab slots
5
median 7.8 days
Revealed OpenRouter stealth slots: when each appeared, what it turned out to be, and how many days passed before the official listing.
SlotDays to officialRevealed asStealth listingNamed by
Union Alpha1.3Pareto · UnbiasedA priced third-party model, not OpenRouter's openrouter/pareto-code router.2026-09-16OpenRouter page
Ox Alpha5.7GLM-5.3-Flash · Z.ai2026-08-20OpenRouter page
Owl Alpha82.8LongCat-2.0 · MeituanNo notice on the OpenRouter page. longcatai.org says the LongCat team confirmed it after about two months; the official OpenRouter listing came 83 days in.2026-04-28third partyunverified, left out of the medians
Hunter/Healer Alpha7.0MiMo-V2-Pro and MiMo-V2-Omni · Xiaomi2026-03-11OpenRouter page
Pony Alpha5.0GLM-5 · Z.ai2026-02-06OpenRouter page
Bert-Nebulon Alpha7.2Mistral Large 3 · Mistral2025-11-24slot description
Sherlock Think/Dash Alpha4.1Grok 4.1 Fast · xAIfrontier lab2025-11-15slot description
Polaris Alpha7.0GPT-5.1 · OpenAIfrontier lab2025-11-06slot description
Andromeda Alpha6.9Nemotron Nano 2 VL · NVIDIA2025-10-21slot description
Sonoma Sky/Dusk Alpha13.3Grok 4 Fast · xAIfrontier lab2025-09-05OpenRouter page
Horizon Alpha/Beta7.8GPT-5 · OpenAIfrontier lab2025-07-30OpenRouter blog
Quasar/Optimus Alpha11.9GPT-4.1 · OpenAIfrontier lab2025-04-02OpenRouter blog

OpenRouter sometimes lists a lab's next model under a codename before launch. Days run from the slot's first OpenRouter created time to the official listing's, both re-read from the API; created can move after the fact, so treat single rows as approximate. Only verified rows feed the medians: Owl Alpha is shown but left out. A slot that never revealed itself is not in the table, so the medians describe slots that were revealed, not every slot. The dashboard lists live slots as an early warning with this track record; they never move the level.

▸SCHEDULED BROADCASTS

OpenAI sometimes publishes an upcoming-livestream page hours before a launch. 4 of these 5 have an archived copy proving it was up before the stream; GPT-5.6 is timed from YouTube's own publishedAt, with no archive. 4 other OpenAI launches had no scheduled stream at all.

OpenAI scheduled livestreams: when the upcoming-stream placeholder was published and when it started.
LaunchStreamPlaceholder publishedStartsWarningEvidence
GPT-4.5Introduction to GPT-4.5Placeholder named the model.2025-02-27 17:00Z2025-02-27 20:00Z+3.0hWayback, upcoming
GPT-4.1New models in the APIPlaceholder title did not name the model.2025-04-14 14:49Z2025-04-14 17:00Z+2.2hWayback, upcoming
o3 / o4-miniIntroduction to new o-series modelsThe longest lead: scheduled two days out.2025-04-14 20:48Z2025-04-16 17:00Z+44.2hWayback, upcoming
GPT-5Introducing GPT-5Placeholder named the model.2025-08-07 12:07Z2025-08-07 17:00Z+4.9hWayback, upcoming
GPT-5.6Introducing the next chapter for ChatGPTStart is the actual air time; OpenRouter listed the models 13.5 hours after the placeholder.2026-07-08 20:25Z2026-07-09 16:53Z+20.5hno capture

Launches a broadcast watch would have missed

  • GPT-5.4 No scheduled stream.
  • GPT-5.5 No scheduled stream.
  • GPT-6 Astra No scheduled stream.
  • GPT-6 Sol/Luna No scheduled stream.
  • Every Anthropic launch Uploads coincide with or trail the post.
  • Google DeepMind Uploads trail the launch.
  • xAI The @xai handle resolves to an unrelated channel.

▸ARCHITECTURE MERGES

Open-weight labs often add their model code to Hugging Face transformers before release. The merge time comes from the pull request; the announcement is the first first-party HN story.

The merge led the launch for 7 of 16 shipped families, coincided within 6h for 5, and trailed it for 4. 1 module is merged and still unshipped.

Hugging Face transformers architecture merges against the first first-party Hacker News story for each model family.
FamilyModuleMergedAnnouncedMerge leadVerdict
Qwen3Qwenqwen3 #368782025-03-31 07:50Z2025-04-28 20:44Z+28.5dleads
Qwen3-VLQwenqwen3_vl #407952025-09-15 10:46Z2025-09-23 20:59Z+8.4dleads
GLM-4.5Z.aiglm4_moe #393932025-07-21 11:24Z2025-07-28 14:15Z+7.1dleads
Qwen3.5Qwenqwen3_5 #438302026-02-09 11:21Z2026-02-16 09:32Z+6.9dleads
GLM-5Z.aiglm_moe_dsa #438582026-02-09 12:06Z2026-02-11 13:42Z+2.1dleads
Qwen3-NextQwenqwen3_next #407712025-09-09 21:46Z2025-09-11 17:38Z+43.9hleads
Qwen3-OmniQwenqwen3_omni_moe #410252025-09-21 21:46Z2025-09-22 17:50Z+20.1hleads
Mistral Small 4Mistralmistral4 #447602026-03-16 19:39Z2026-03-16 20:40Z+1.0hcoincident
gpt-ossOpenAIgpt_oss #399232025-08-05 16:02Z2025-08-05 17:00Z+1.0hcoincident
Gemma 4Googlegemma4 #451922026-04-02 15:24Z2026-04-02 16:00Z+0.6hcoincident
GLM-5.3-FlashZ.aiglm5_next #483422026-08-26 14:26Z2026-08-26 14:08Z−0.3hcoincident
Llama 4Metallama4 #373072025-04-05 20:02Z2025-04-05 18:33Z−1.5hcoincident
DeepSeek V4DeepSeekdeepseek_v4 #456432026-05-02 11:41Z2026-04-24 02:55Z−8.4dlags
MiniMax-M2MiniMaxminimax_m2 #420282026-01-09 15:25Z2025-10-27 05:12Z−74.4dlags
MiniMax-Text-01MiniMaxminimax #358312025-06-04 07:38Z2025-01-14 19:32Z−140.5dlags
Kimi LinearMoonshotkimi_linear #482502026-09-05 17:04Z2025-10-30 15:44Z−310.1dlags
Qwen4-ExpQwenqwen4_exp #483372026-08-26 12:03Znot yet31d pendingpending
  • Qwen3.8 Reused the qwen3_5 classes; no qwen3_8 module exists (registry read 2026-09-26).
  • OpenAI, Anthropic, xAI frontier models Closed weights never touch transformers.

▸TESTED AND DIDN'T LEAD

Signals that were tested and did not lead a launch.
Signal and verdictFinding
Release cadence ("overdue" labs)no better than chanceWalk-forward: labs at z ≥ 0.5 past their mean gap shipped within 7 days 11.3% of the time, against a 20% base rate. Releases are bursty (CV 0.9–1.6). OpenRouter models API
Cross-lab clusteringno better than chanceAnother frontier lab shipped within 3d/7d 37.3%/66.7% of the time, against 37.3%/66.1% for a weekday-preserving shuffle (p = 0.51/0.44). OpenRouter models API
SDK and OpenAPI commit feedslagsCoincident at best; 5 of 7 cases lag. The Opus 5.5 SDK release went out 2 minutes before the post. anthropic-sdk-python v1.8.0
sglang commitslagsLed 1 of 7 launches, and only echoed a transformers merge that came first. sglang commits
litellm and vercel/ai model listslagsTrail OpenRouter by 20 minutes to 1h51m; litellm mirrors OpenRouter through a bot. litellm commits
Lab docs catalog pagescoincidentAcross 5 launches, no Wayback snapshot shows a docs page carrying a model before its announcement. GPT-6 Astra was still absent 6.4h after launch. Wayback Machine
AWS What’s Newfakes a leadpubDate is a backdated editorial time; read at face value it fakes a lead on 13 of 15 Claude launches. AWS What’s New feed
Status pages (Claude, OpenAI)lagsA model is first mentioned 6 hours to about 20 days after its listing. xAI’s returns 403. status.claude.com
Wikipedia infoboxeslagsUpdated after launch in 6 of 6 cases. MediaWiki API
Hugging Face collectionslagsLagged or coincided in 9 of 9 releases. Hugging Face API
Vertex AI release noteslagsThe legacy feed has been stale since March; its successor has day-level dates and lags. Vertex AI release notes
Design Arena codenameslagsThe public registry purged its anonymous codenames in June 2026, and every entry since appeared after launch. One anecdote survives: "Radon" was Muse Spark, 6d14h early. Design Arena
Manifoldnot independentPegged to Polymarket, 30 minutes to 3 hours behind it, on thin books. Manifold API
Kalshinot independentGPT-6 moved 0.19 → 0.83 in the same hour as Polymarket (Sep 2, 15–17Z). No markets existed for Opus 5.5 or Sol/Luna. Kalshi API

Findings from the 2026-09-26 research pass. pnpm backtest carries them with their sources; it does not recompute them.

▸TIMESTAMP TRAPS

Each of these timestamps looks like an early warning and isn't. Including one this site got wrong.

Timestamps that fake a lead, with an example of each and the fix.
TrapExampleFake leadFix
Git commit datesThe anthropic-sdk-python commit adding claude-opus-5-5 is dated 2026-09-20 22:54:59Z; it shipped in v1.8.0 at 2026-09-22 16:25:23Z, two minutes before the launch post. commit b5cc700+41.5hUse the release or push time, never the author or committer date.
AWS What’s New pubDateThe GovCloud Opus 5.5 item was created 2h30m after its own pubDate (research pass; not re-probed here). AWS What’s New feed+2.5hRecord when the item was first seen.
Hugging Face createdAtcreatedAt is when the repo was made, while still private: google/gemma-4-E4B-it says 2026-03-02 19:57Z, 30.8 days before the launch post. google/gemma-4-E4B-it+30.8dOnly a first-seen time from polling counts.
Date-only fields pinned to midnightxAI’s release notes date Grok 4.7 "September 21, 2026" with no time. Read as 00:00Z, that is 15h50m before its HN story. xAI release notes+15.8hMark day-precision items and keep them out of lead analysis.
Rounded RSS pubDatesopenai.com’s RSS dates "GPT-6 Astra: A new generation of intelligence" to 2026-09-03 11:00:00Z; its HN story is 18:41Z. openai.com RSS+7.7hTreat whole-hour and midnight pubDates as day precision.
OpenRouter created before the announcementListed before the announcement: Claude Opus 4.8 -22.7h, GPT-5.6 Sol/Terra/Luna -7.2h, Gemini 3.5 Flash -4.9h, Claude Fable 5 / Mythos 5 -4.7h, Qwen3.5 -3.1h, Grok 4.5 -2.9h. OpenRouter also rewrites it: Space Bunny Alpha moved from 2026-09-22 10:58Z to 2026-09-23 14:48Z. OpenRouter models API+22.7hUse it as availability, measured against the announcement, never as a lead.
This site’s own "her" storyThe History timeline said Altman posted "her" the night before GPT-4o. The tweet id decodes to 2024-05-13T17:45:09Z; the GPT-4o announcement hit HN at 2024-05-13T17:28:00Z, so the tweet came 17 minutes after it. sama/status/1790075827666796666—Decode the tweet id rather than trusting a remembered order.

▸REPRODUCE IT

Every computed table is rebuilt from raw pulls committed to the repo. The curated tables (stealth reveals, broadcasts, architecture merges, what didn't lead, timestamp traps) are carried as sourced data, not recomputed. The rebuild is deterministic: run it twice and git diff stays empty.

pnpm install
pnpm backtest            # rebuild data/backtest/*.json from data/backtest/raw
pnpm backtest:replay     # rerun the v3 replay into data/backtest/v3-replay.json
pnpm backtest --refresh  # re-pull Gamma, CLOB, HN and OpenRouter first

Raw pulls last refreshed 2026-09-26 00:46Z. Curated inputs (release ids, sources, the tables that are not recomputed) live in scripts/backtest/curated.ts; the replay, its pre-registered rule and every constant it fits are in scripts/backtest/replay.ts.

Or check any single number by hand:

# Closed AI events, 50 per page; pass next_cursor back as after_cursor
curl 'https://gamma-api.polymarket.com/events/keyset?tag_slug=ai&closed=true&limit=50'
# Claude Opus 5.5: "September 22" price history (windows of 15 days or less)
curl 'https://clob.polymarket.com/prices-history?market=109495115084615347788619140162263908414615168053426928181800069578275691598743&startTs=1789597495&endTs=1790102053&fidelity=60'
# The announcement: first first-party story in the window (Algolia wants the filter URL-encoded)
curl 'https://hn.algolia.com/api/v1/search_by_date?query=Opus%205.5&tags=story&hitsPerPage=1000&numericFilters=created_at_i%3E%3D1789776000%2Ccreated_at_i%3C%3D1790208000'
# Availability
curl -s 'https://openrouter.ai/api/v1/models' | jq '.data[] | select(.id == "anthropic/claude-opus-5.5") | .created'
# The pending architecture
curl -s 'https://cdn.jsdelivr.net/gh/huggingface/transformers@main/src/transformers/models/__init__.py' | grep qwen4_exp