Forecast check · 2026 season
How good are the forecasts? AI vs physics
We check every storm forecast on Stormlines against where the storm actually went. Here is how the new compare with the forecasters have relied on for decades.
1 day ahead, AIFS ENS missed the storm’s centre by 45 mi on average and ECMWF ENS by 43 mi, on the same 34 forecasts of 20 storms; AIFS ENS was closer in 54% of them.
Updated Mon 28 Sep, 23:44 UTC · checked through the Mon 28 Sep 18Z runs · up to 148 forecasts from 41 storms per score so far
Every storm we checked. Storms with an agency best track are checked against it; the rest against the models’ own analysis (see How we check).
AI vs physics, head to head
Each centre runs an AI model next to its physics model. Here they are scored only on the same storms, from the same start time and at the same lead time, so neither gets the easier cases. The dots are the average distance between the forecast centre and where the storm really was; closer to the left is better.
European models: AIFS ENS (AI) vs ECMWF ENS (physics)
1 day ahead, AIFS ENS missed the storm’s centre by 45 mi on average and ECMWF ENS by 43 mi, on the same 34 forecasts of 20 storms; AIFS ENS was closer in 54% of them.
- ECMWF ENS physics
- AIFS ENS AI
- 48 h, 72 h, 96 h, 120 h collecting: no forecasts checked yet (needs 30 from 5 storms)
American models: AIGEFS (AI) vs GEFS (physics)
At the start of each forecast, AIGEFS placed the storm’s centre 19 mi from the verifying position on average and GEFS 22 mi, on the same 90 forecasts of 33 storms; AIGEFS was closer in 53% of them.
At the start, the model-analysis storms are checked against the models’ own middle position, so this mostly measures how far each model sits from the others.
- GEFS physics
- AIGEFS AI
- 24 h collecting: 15 cases from 15 storms so far (needs 30 from 5)
- 48 h, 72 h, 96 h, 120 h collecting: no forecasts checked yet (needs 30 from 5 storms)
How far off, by lead time
Lead time is how far ahead a forecast was made: a 48-hour forecast was made two days before the moment it describes. Errors grow with lead time, so compare models at the same lead. Each model is scored here on every forecast it made, so the samples differ a little; the head-to-head above is the fair comparison.
- GEFS physics
- AIGEFS AI
- ECMWF ENS physics
- AIFS ENS AI
- GEPS physics
- All models together
- 48 h collecting: 14 cases from 14 storms so far (needs 30 from 5)
- 72 h, 96 h, 120 h none checked yet: a forecast is checked once the moment it describes has passed, and only if the storm is still being followed then.
All five, on the same forecasts
Only the cases where all five ensembles forecast the same storm from the same start time. The fairest comparison of all five, and the slowest to fill.
- GEFS physics
- AIGEFS AI
- ECMWF ENS physics
- AIFS ENS AI
- GEPS physics
- 24 h collecting: no forecasts checked yet (needs 30 from 5 storms)
- 48 h collecting: no forecasts checked yet (needs 30 from 5 storms)
- 72 h collecting: no forecasts checked yet (needs 30 from 5 storms)
- 96 h collecting: no forecasts checked yet (needs 30 from 5 storms)
- 120 h collecting: no forecasts checked yet (needs 30 from 5 storms)
Does the storm stay inside the cone?
Each model’s is drawn to hold the storm’s centre in about 2 of every 3 forecasts (67%). If the storm lands inside more often than that, the cone was wider than it needed to be; less often, and it was too sure of itself. The “all models together” line is checked against the cones on our storm maps.
- GEFS physics
- AIGEFS AI
- ECMWF ENS physics
- AIFS ENS AI
- GEPS physics
- All models together
- 48 h collecting: 14 cases from 14 storms so far (needs 30 from 5)
- 72 h, 96 h, 120 h none checked yet: a forecast is checked once the moment it describes has passed, and only if the storm is still being followed then.
The numbers
Every score in a table, with the counts behind it
Distances in mi. Scores need 30 forecasts from 5 storms; cells with fewer show the counts so far. Along: + means the forecasts ran ahead of the storm, − behind. Cross: + means the forecast lay to the right of the real position, looking along the forecast’s heading. Cone: share inside the 67% cone. Envelope: share inside the middle 8 in 10 of the members, along and across the track.
| Model | Lead | Cases | Storms | Mean | Median | SD | Along | Cross | Cone | Env. |
|---|---|---|---|---|---|---|---|---|---|---|
| GEFS | Start | 132 | 36 | 19 | 12 | 31 | −1 | 0 | 93% | 81% |
| 24 h | 30 | 18 | 49 | 39 | 42 | −13 | −3 | 87% | 70% | |
| 48 h | 0 | 0 | collecting | |||||||
| 72 h | 0 | 0 | collecting | |||||||
| 96 h | 0 | 0 | collecting | |||||||
| 120 h | 0 | 0 | collecting | |||||||
| AIGEFS | Start | 101 | 37 | 25 | 14 | 38 | −4 | 0 | 92% | 74% |
| 24 h | 40 | 27 | 43 | 36 | 34 | −14 | +2 | 95% | 80% | |
| 48 h | 14 | 14 | collecting | |||||||
| 72 h | 0 | 0 | collecting | |||||||
| 96 h | 0 | 0 | collecting | |||||||
| 120 h | 0 | 0 | collecting | |||||||
| ECMWF ENS | Start | 122 | 38 | 14 | 9 | 31 | −4 | −3 | 100% | 92% |
| 24 h | 34 | 20 | 43 | 25 | 59 | −11 | −6 | 88% | 74% | |
| 48 h | 0 | 0 | collecting | |||||||
| 72 h | 0 | 0 | collecting | |||||||
| 96 h | 0 | 0 | collecting | |||||||
| 120 h | 0 | 0 | collecting | |||||||
| AIFS ENS | Start | 125 | 39 | 15 | 10 | 35 | −6 | −3 | 99% | 94% |
| 24 h | 34 | 20 | 45 | 25 | 65 | −15 | −2 | 94% | 85% | |
| 48 h | 0 | 0 | collecting | |||||||
| 72 h | 0 | 0 | collecting | |||||||
| 96 h | 0 | 0 | collecting | |||||||
| 120 h | 0 | 0 | collecting | |||||||
| GEPS | Start | 69 | 35 | 25 | 17 | 25 | −5 | −2 | 88% | 46% |
| 24 h | 16 | 16 | collecting | |||||||
| 48 h | 0 | 0 | collecting | |||||||
| 72 h | 0 | 0 | collecting | |||||||
| 96 h | 0 | 0 | collecting | |||||||
| 120 h | 0 | 0 | collecting | |||||||
| All models together | Start | 148 | 41 | 9 | 9 | 25 | −5 | +1 | 99% | 98% |
| 24 h | 33 | 20 | 40 | 24 | 59 | −16 | −3 | 94% | 91% | |
| 48 h | 0 | 0 | collecting | |||||||
| 72 h | 0 | 0 | collecting | |||||||
| 96 h | 0 | 0 | collecting | |||||||
| 120 h | 0 | 0 | collecting | |||||||
AI vs physics, same cases
| Pair | Lead | Cases | Storms | AI mean | Physics mean | AI closer |
|---|---|---|---|---|---|---|
| AIFS ENS vs ECMWF ENS | Start | 121 | 37 | 14 | 14 | 56% |
| 24 h | 34 | 20 | 45 | 43 | 54% | |
| 48 h | 0 | 0 | collecting | |||
| 72 h | 0 | 0 | collecting | |||
| 96 h | 0 | 0 | collecting | |||
| 120 h | 0 | 0 | collecting | |||
| AIGEFS vs GEFS | Start | 90 | 33 | 19 | 22 | 53% |
| 24 h | 15 | 15 | collecting | |||
| 48 h | 0 | 0 | collecting | |||
| 72 h | 0 | 0 | collecting | |||
| 96 h | 0 | 0 | collecting | |||
| 120 h | 0 | 0 | collecting | |||
All five on the same cases (mean error)
| Lead | Cases | Storms | GEFS | AIGEFS | ECMWF ENS | AIFS ENS | GEPS |
|---|---|---|---|---|---|---|---|
| Start | 40 | 27 | 19 | 19 | 13 | 15 | 23 |
| 24 h | 0 | 0 | collecting | ||||
| 48 h | 0 | 0 | collecting | ||||
| 72 h | 0 | 0 | collecting | ||||
| 96 h | 0 | 0 | collecting | ||||
| 120 h | 0 | 0 | collecting | ||||
The data: verify/summary.json (this season), and one file per storm under /data/verify/storms/.
Named storms checked so far
Each storm page shows what its forecasts said and what happened, once enough of them can be checked.
- Hurricane Polo41 forecasts checked
- Former Tropical Storm Odalys41 forecasts checked
- Tropical Storm Rachel41 forecasts checked
- Typhoon Surigae39 forecasts checked
- Hurricane Nolo40 forecasts checked
- September Nor'easter22 forecasts checked
How we check
- What is scored. For each model run, the forecast position of the storm’s centre is the middle (median) of that model’s ensemble members. A run counts only when at least 4 in 10 of its members have the storm. “All models together” pools every model’s forecasts and takes the middle of them all; its cone check uses the cones our storm maps draw.
- The truth. Where a storm is tracked by a hurricane or typhoon agency (NOAA’s National Hurricane Center or Central Pacific Hurricane Center, or the Joint Typhoon Warning Center), the truth is the agency’s best track: its record of where the centre was, every 6 hours. These are preliminary until the agency’s post-season review. Every other storm (nor’easters, European windstorms, most lows) has no official track, so the truth is where the models agreed it was: the middle of their starting positions at each time. That is not an observation, and it favours models that agree with the others; the “Checked against” switch keeps the two apart.
- The error. The distance along the Earth’s surface between the forecast centre and the truth at the same time, split into along-track (too fast or too slow) and cross-track (off to one side).
- Lead time. How long before the moment it describes the forecast was made. “Start” (0 h) is the forecast’s first position. A forecast is checked once its moment has passed, so scores at 5 days fill in 5 days after the storms.
- Enough to say anything. A score appears only when it rests on at least 30 forecasts from at least 5 storms. Until then we show the counts so far. Forecasts made 6 hours apart for the same storm are not independent, which is why storms count as well as forecasts. Storm pages use a lower bar for their own storm (4 checked runs of a model) and say so.
- Fair comparisons. “Head to head” and “All five” use only cases where every model in the comparison forecast the same storm from the same start at the same lead. The by-lead chart uses every case each model has, so its samples differ.
- What it is not. Track only: not strength, rain or wind. One season of the storms Stormlines follows is a small sample next to the official multi-year verifications by NHC and ECMWF. Not an official forecast or verification.
Model names: the glossary. Reading the storm maps: how to read a spaghetti map.