Skip to content
Climate & the Price of Bread

Step 3 · Backtests

Did the climate input beat the baseline on data it never saw?

Every month from 1990 onwards, a SARIMA model forecasts next month’s return. Its parameters are re-estimated every 12 months on the previous 120 months, and between re-estimations it is updated with each new month of data. Its climate twin is identical except for one extra regressor: a climate index read k months before the forecast month. Both see exactly the same history.

The evaluation period is split in half. The first half picks the lag for each index (and the best index); the second half scores that choice. Grid cells in between are shown for transparency, but with 72 index-lag combinations per commodity some will look good by luck, which is why the test half and the Diebold-Mariano p-value carry the weight.

Commodity
Prices
Scored on
Baseline model
SARIMA(0,0,1)
chosen by BIC on 359 months before the evaluation
Evaluation
Jan 1990 – Jun 2026
438 one-step-ahead forecasts
Lag chosen on
Jan 1990 – Mar 2008
219 months
Judged on (test)
Apr 2008 – Jun 2026
219 months, never used for choices

Wheat HRW: forecast skill of every climate index and lag

Test half. Skill = 1 − RMSE(climate) / RMSE(baseline); blue means the climate input reduced the error. Outlined: the lag each index would pick on the validation half.
Worse than baselineBetter(−4.0% … +4.0%)
Diebold-Mariano p < 0.05
MEI.v2ONIAOPNASAMDMI123456789101112Lag (months between the index reading and the forecast month)
Show the numbers
Skill by index and lag
LagMEI.v2ONIAOPNASAMDMI
1−0.38%+0.33%−0.24%−0.44%−1.34%−0.40%
2−0.54%+0.12%−0.10%−0.17%−1.86%−0.85%
3−0.68%−0.28%−0.43%−0.63%−0.84%−0.03%
4−0.90%−1.01%−0.46%−1.28%−0.42%−0.59%
5−0.74%−2.05%−1.27%−0.29%−0.03%−0.29%
6−1.06%−3.25%−1.19%−0.79%−0.16%−0.31%
7−2.47%−3.99%−1.20%−0.49%−0.65%−0.17%
8−2.97%−3.98%−1.15%−0.65%−1.14%−0.47%
9−3.06%−3.23%−0.46%−0.88%−0.37%−0.80%
10−3.46%−2.67%−0.26%−0.22%−0.76%−1.38%
11−2.63%−2.11%−0.18%−0.36%−0.68%−1.02%
12−1.75%−1.42%−0.39%−0.41%−0.77%−0.58%
Baseline on this split: RMSE 6.58 percentage points over 219 months; its 80% and 95% intervals covered 84% and 94% of outcomes.

Lag chosen on the validation half, scored on the test half

One row per index. The highlighted row is the index whose chosen lag did best on validation: the single model this procedure would have shipped.
Validation-selected lag and test performance per climate index
IndexLagSkillReading
3−0.68%No clear difference
2+0.12%No clear difference
8−1.15%No clear difference
7−0.49%No clear difference
2−1.86%Climate hurts
11−1.02%Climate hurts

RMSEs, p-values and 95% coverage are shown on wider screens and in the CSV downloads on the method page.

Decade by decade: MEI.v2 at lag 3

RMSE of one-step-ahead forecasts (percentage points of monthly log-return). Closer dots, smaller difference.
  • Baseline SARIMA
  • With MEI.v2
5.405.605.806.006.206.40RMSE (percentage points)1990s2000s2010s2020s
Decades mix the validation and test halves, so read them as description rather than as a test. A ring marks a decade where the two RMSEs almost coincide.

Across commodities

The model each commodity would ship (best validation index and lag), on its test half.
  • MEI.v2, lag 3−0.7%No clear difference
  • PNA, lag 7−1.3%Climate hurts
  • PNA, lag 12−0.9%No clear difference
  • DMI, lag 1−1.4%Climate hurts
  • PNA, lag 6−0.8%No clear difference
  • DMI, lag 7−8.2%Climate hurts
  • ONI, lag 5+1.0%No clear difference
  • PNA, lag 10+0.1%No clear difference

See how these models forecast month by month on the forecasts page.