MAST90106 / MAST90107 · The University of Melbourne · 2023 capstone, rebuilt
When the Pacific warms, does bread get dearer?
El Niño, La Niña and their atmospheric cousins reshuffle rain and heat across the world’s breadbaskets. This site asks a sharper question: do those climate signals help forecast next month’s price change for wheat, maize, rice, soybeans, palm oil and sugar, once you compare honestly against a model that ignores the climate?
46 years of ENSO, and the price of bread wheat
Monthly, January 1980 to Sep 2026. Shading marks El Niño and La Niña episodes (NOAA’s five-season rule).
- El Niño episode
- La Niña episode
Oceanic Niño Index, °C anomaly
Wheat (US HRW), US$/mt
Out-of-sample forecasts scored
460,192
one month ahead, from 1,168 rolling SARIMAX set-ups and 38,982 estimations
Commodities where climate won on held-out data
2 of 8
best index and lag chosen on the first half, judged on the second
…and significantly so
0 of 8
Diebold-Mariano test, 5% level
All index pairings that helped significantly
0 of 48
6 were significantly worse than the baseline
The short answer
Mostly, climate adds little month-ahead skill.
For each commodity a SARIMA baseline, estimated on the previous ten years and updated as each month arrives, forecasts the next month’s log-return. The climate version gets one extra input: a climate index from k months earlier. The index and lag are picked on the first half of the evaluation period and then scored, untouched, on the second half.
Skill is the share of forecast error removed: +2% means the climate model’s root-mean-square error is 2% smaller than the baseline’s. Monthly commodity returns are noisy, so even real climate effects show up as small numbers here.
Why can a climate input make things worse? An extra regressor with little real signal still has to be estimated, on ten noisy years at a time, and that estimation noise leaks into every forecast. In 3 of 8 commodities the model picked on the first half was significantly less accurate than the plain baseline on the second.
| Commodity | Chosen input | Test skill | p | Reading |
|---|---|---|---|---|
| Wheat HRWMEI.v2 · 3 mo · p 0.543 | −0.7% | No clear difference | ||
| Wheat SRWPNA · 7 mo · p 0.036 | −1.3% | Climate hurts | ||
| MaizePNA · 12 mo · p 0.185 | −0.9% | No clear difference | ||
| SoybeansDMI · 1 mo · p 0.029 | −1.4% | Climate hurts | ||
| Rice ThaiPNA · 6 mo · p 0.169 | −0.8% | No clear difference | ||
| Rice Viet NamDMI · 7 mo · p 0.009 | −8.2% | Climate hurts | ||
| Palm oilONI · 5 mo · p 0.413 | +1.0% | No clear difference | ||
| SugarPNA · 10 mo · p 0.890 | +0.1% | No clear difference |
Test skill = 1 − RMSE(climate) / RMSE(baseline) on one-step-ahead forecasts from Apr 2008 to Jun 2026 (Viet Nam rice: a shorter test half). p: two-sided Diebold-Mariano test on squared errors.
Right now
El Niño conditions are building fast.
The latest Oceanic Niño Index reading (season ending September 2026) is 2.16 °C, 4 overlapping seasons past the ±0.5 °C threshold (NOAA's historical rule counts an episode from five), and the multivariate ENSO index for September 2026 is 2.37. The forecasts page shows what the models make of that for the next twelve months, with and without the climate input, and how wide the honest uncertainty is.
How the rebuild works
Four steps, all reproducible from public data
- 01
Fetch
World Bank Pink Sheet prices (to Sep 2026), six NOAA climate indices and US CPI, re-downloaded by one Python script.
- 02
Look for lags
Correlate each month's price change with each index 0 to 24 months earlier, over the whole record and decade by decade.
Open - 03
Backtest honestly
Rolling 120-month SARIMAX fits, one-step-ahead, for every index and lag 1-12, against the same model without climate.
Open - 04
Forecast
Simulate the next twelve months from the latest data, so the intervals reflect the model's own uncertainty.
Open
Evaluation runs from Jan 1990 to Jun 2026 for the long series; the test half starts in Apr 2008.
About this project
A capstone, rebuilt in the open
- Subject
- MAST90106 / MAST90107 Data Science Project (Parts 1 and 2)
- Where
- The University of Melbourne, Master of Data Science
- When
- Semesters 1 and 2, 2023
- Setting
- Group 36, an industry capstone hosted by CSIRO
- Team
- Sunchuangyu (Rin) Huang, Ritwik Giri, Jiaqi Hu, Xiangyi (Emma) He, Jihang (Jonathan) Yu
In 2023 our team was asked whether large-scale climate patterns carry useful information about future movements in staple food prices, a question that matters for food security and, through it, for social stability. We explored the price history, built rolling SARIMAX forecasting experiments with lagged climate indices as extra inputs, searched over lags, and tested whether the climate-aware models beat ones without them, decade by decade.
This site is a clean-room rebuild. The original project used data and materials provided by the industry host, and the code was written jointly by the team. None of that is used here: every number on this site is re-derived from openly published World Bank, NOAA and FRED data by code written from scratch for this rebuild. The research question and the general approach are the same; the data, implementation and therefore the results are new. The original submission is not reproduced, and this rebuild is not affiliated with or endorsed by CSIRO.
Original stack (2023)
- Python and Jupyter
- pandas
- statsmodels SARIMAX
- pmdarima
- SQLite
- LaTeX
Revived stack
- Python (uv, statsmodels) for the offline pipeline
- SQLite analytics database
- Next.js 16 and React 19
- TypeScript ports of the statistics
- Tailwind CSS v4 and hand-drawn SVG charts
- A Web Worker for the in-browser scan