Method
How every number on this site is made
One Python script (scripts/build_analytics.py) downloads the public data, runs every model and writes a small read-only SQLite database. The website only reads that database, except in the lab, where TypeScript ports of the statistics run in your browser.
Data
Prices. World Bank Commodity Price Data (the “Pink Sheet”), monthly averages in nominal US dollars, updated 2 October 2026. Each price is turned into a monthly log-return in percent, 100 × (ln Pt − ln Pt−1). For the “CPI-deflated” view prices are first divided by US CPI (all urban consumers, seasonally adjusted) and expressed in August 2026 dollars.
Climate indices. Six standard teleconnection indices from NOAA. Each value is labelled by the last month it covers (the three-month ONI season ending in March is the March value), so a forecast for month t made with lag k ≥ 1 only uses index values labelled t − k or earlier. These are the latest revised values, and some indices are published well after the month they describe (the HadISST-based DMI runs months behind), so the backtests are pseudo-real-time rather than a record of what was known at the time.
| Series | Unit | From | To |
|---|---|---|---|
| Wheat (US hard red winter) | $/mt | Jan 1960 | Sep 2026 |
| Wheat (US soft red winter) | $/mt | Jan 1979 | Sep 2026 |
| Maize (US No. 2 yellow) | $/mt | Jan 1960 | Sep 2026 |
| Soybeans (US) | $/mt | Jan 1960 | Sep 2026 |
| Rice (Thai 5% broken) | $/mt | Jan 1960 | Sep 2026 |
| Rice (Viet Nam 5% broken) | $/mt | Apr 2009 | Sep 2026 |
| Palm oil (Malaysian) | $/mt | Jan 1960 | Sep 2026 |
| Sugar (world, ISA) | $/kg | Jan 1960 | Sep 2026 |
Viet Nam rice is used only from April 2009, where its record becomes unbroken. The Pink Sheet monthly values are averages of daily or weekly quotes, which smooths some within-month moves.
Forecasting design
- Lag screen. Pearson correlation of each month’s return with each index 0 to 24 months earlier, over the whole record and decade by decade. The grey band in the charts is ±1.96/√n, the rough 95% range under no association.
- Baseline. For each commodity a SARIMA(p,0,q)(P,0,Q)12 model with a constant is chosen by BIC over p, q ≤ 2 and one seasonal term, using only the returns observed before the evaluation starts.
- Climate model. The same SARIMA, plus one regressor: an index read k months before the forecast month (regression with SARIMA errors, i.e. SARIMAX). Six indices × lags 1 to 12 = 72 climate models per commodity.
- Rolling origin. Every month from January 1990 (or once 120 months of history exist), each model forecasts the next month. Parameters are re-estimated every 12 months on the trailing 120 months; between re-estimations the Kalman filter updates the model with each new month, so every forecast uses all data up to the month before it, and nothing after.
- Selection, then test. The evaluation months are split in half. For each index the lag with the lowest RMSE on the first half is kept, and the index whose kept lag did best is the commodity’s “shipped” model. The second half scores those choices without any further tuning.
- Significance. Diebold-Mariano test on squared errors with the Harvey-Leybourne-Newbold small-sample correction; two-sided p-values from Student’s t. “Climate helps” needs lower RMSE and p < 0.05.
- Forecasts. Each model is refitted on the latest 120 months and simulated 4,000 times for 12 months. Price paths compound the simulated returns from the last observed price. When a forecast needs an index value that is not published yet, it comes from an AR(1) fitted to that index’s recent history; the uncertainty of that projection is not added to the bands.
| Commodity | Model | Pre-sample | Evaluation | Test half from |
|---|---|---|---|---|
| Wheat HRW | SARIMA(0,0,1) | 359 months | Jan 1990 – Jun 2026 | Apr 2008 |
| Wheat SRW | SARIMA(0,0,0)(0,0,1)12 | 131 months | Jan 1990 – Jun 2026 | Apr 2008 |
| Maize | SARIMA(1,0,0) | 359 months | Jan 1990 – Jun 2026 | Apr 2008 |
| Soybeans | SARIMA(1,0,0) | 359 months | Jan 1990 – Jun 2026 | Apr 2008 |
| Rice Thai | SARIMA(1,0,0) | 359 months | Jan 1990 – Jun 2026 | Apr 2008 |
| Rice Viet Nam | SARIMA(2,0,1) | 120 months | May 2019 – Jun 2026 | Dec 2022 |
| Palm oil | SARIMA(0,0,1) | 359 months | Jan 1990 – Jun 2026 | Apr 2008 |
| Sugar | SARIMA(0,0,1) | 359 months | Jan 1990 – Jun 2026 | Apr 2008 |
In total 38,982 SARIMAX estimations produced 460,192 one-step-ahead forecasts (statsmodels 0.15.0, Python 3.13.9).
Python and TypeScript agree
The browser code is a port, not a reimplementation from memory. The build script writes a fixture of inputs and reference outputs computed with NumPy, SciPy and statsmodels, and the TypeScript unit tests check the ports against it on every CI run:
- Lagged correlations for lags 0 to 24 (
lib/stats.ts) matchlagged_corrto 1e-10. - The rolling ARX backtest (
lib/arx.ts, Householder least squares inlib/linalg.ts) reproduces NumPy’s predictions, residual standard deviations and RMSE for several orders, lags and windows. - The Diebold-Mariano statistic and p-value, and Student’s t CDF, match SciPy; RMSE and interval coverage recomputed from the stored SARIMAX predictions match the values the pipeline stored.
Caveats
- Many comparisons. 6 indices × 25 lags in the lag screen, 72 climate models per commodity in the backtests. Some will look good by chance; that is what the held-out half is for.
- Monthly returns are mostly noise. Even a genuine climate effect on harvests is a small share of month-to-month price variance, which is driven by stocks, energy, currencies, policy and speculation. Small skill numbers are expected.
- Final, not real-time, index values. Every backtest uses today’s revised index files. A forecaster in, say, 2005 would have seen earlier versions, and for short lags some values (notably the DMI) would not yet have been published, so the climate models are given a slight head start.
- Regimes change. The 2007–08 food crisis, the 2010–11 spike and 2022 dominate the test half; intervals that assume calm months under-cover in those periods.
- Not a trading tool. This is a research demonstration of an evaluation design, not investment advice.
Downloads
Derived tables from the analytics database, as CSV:
- Monthly prices and log-returns
- Climate index values (as aligned for forecasting)
- Lagged cross-correlations
- SARIMAX backtest metrics for every index and lag
- Validation-selected lag per commodity and index, scored on the test half
- One-step-ahead out-of-sample predictions
- Twelve-month forecasts (simulated quantiles)
Prices derive from World Bank data (CC BY 4.0, attribution: World Bank Commodity Price Data). Climate indices are NOAA products in the US public domain.
Reproduce it
The repository is private until the project is published; these are the steps once it is open.
git clone https://github.com/rNLKJA/UoM-MDS-Capstone-ENSO-Commodity-Forecasting cd UoM-MDS-Capstone-ENSO-Commodity-Forecasting uv run scripts/build_analytics.py # fetch, model, write web/data/analytics.db cd web && pnpm install && pnpm test && pnpm dev
Clean-room note
The 2023 capstone used data and documents supplied by its industry host and code written jointly by the team. This rebuild uses none of them. It was written from scratch from the research question alone, on openly published data, and its results are its own. It is not affiliated with or endorsed by CSIRO or the University of Melbourne.
Build Oct 2026 · window 120 months · refit every 12 months · seed 20231022 · 1,168 backtest configurations