October 5, 2026
| Claude model | Training cutoff | Obs. | Blind | Blind, with outcome |
|---|---|---|---|---|
| Haiku 4.5 | 31 July 2025 | 60 | 60 | 18 |
| Sonnet 5 | 31 January 2026 | 60 | 60 | 18 |
| Fable 5.0 | 31 January 2026 | 60 | 60 | 18 |
| Opus 5 | 31 May 2026 | 60 | 20 | 2 |
| Sonnet 5.5 | 30 June 2026 | 60 | 20 | 2 |
| Opus 5.5 | 30 June 2026 | 60 | 20 | 2 |
Notes: Each model forecasts 60 observations over the 3 rounds: four variables at five horizons in each round. A forecast is blind when the round's information date follows the model's training cutoff, so that nothing the model was trained on postdates the data it was given. Opus 5, Sonnet 5.5 and Opus 5.5 have the fewest because their cutoffs are the latest: the 2026Q1 and 2026Q2 rounds predate them.
Live track: forecasts against the published outcomes
| Variable | Target | Obs. | Outcome | Claude error | SPF error | Winner |
|---|---|---|---|---|---|---|
| CPI inflation | 2026Q1 | 3 | 3.19 | 0.369 | 0.541 | Claude |
| CPI inflation | 2026Q2 | 6 | 6.07 | 3.283 | 1.735 | SPF |
| Real GDP growth | 2026Q1 | 3 | 1.99 | 0.085 | 0.660 | Claude |
| Real GDP growth | 2026Q2 | 6 | 1.50 | 0.519 | 0.601 | Claude |
| Unemployment rate | 2026Q1 | 3 | 4.33 | 0.006 | 0.067 | Claude |
| Unemployment rate | 2026Q2 | 6 | 4.27 | 0.106 | 0.153 | Claude |
| Unemployment rate | 2026Q3 | 12 | 4.13 | 0.165 | 0.206 | Claude |
| 3-month Treasury bill | 2026Q1 | 3 | 3.59 | 0.008 | 0.057 | Claude |
| 3-month Treasury bill | 2026Q2 | 6 | 3.62 | 0.109 | 0.045 | SPF |
| 3-month Treasury bill | 2026Q3 | 12 | 3.80 | 0.254 | 0.184 | SPF |
| All pooled | 60 | 0.509 | 0.398 | SPF | ||
| By model, the same observations | ||||||
| Haiku 4.5 | 18 | 0.595 | 0.435 | SPF | ||
| Sonnet 5 | 18 | 0.552 | 0.435 | SPF | ||
| Fable 5.0 | 18 | 0.534 | 0.435 | SPF | ||
| Opus 5 | 2 | 0.038 | 0.067 | Claude | ||
| Sonnet 5.5 | 2 | 0.050 | 0.067 | Claude | ||
| Opus 5.5 | 2 | 0.042 | 0.067 | Claude | ||
Notes: Forecasts with a published outcome, from the 2026Q1, 2026Q2 and 2026Q3 rounds. Each observation is one model's forecast of the target from one round, the three draws averaged. A forecast counts only when the round's information date follows the model's training cutoff. Errors are mean absolute errors against the first published estimate. Claude is closer than the SPF median on 7 of the 10 targets, and pooled over the 60 observations its error is 0.509 against the SPF's 0.398. Outcome convention: 2025Q4 averaged over its 2 published months, as FRED aggregates it; October 2025 was never collected.