Demand BenchForecast model benchmarking Start forecasting

Self-service demand forecasting

Compare forecasts on your own demand history

Start forecasting Compare the models first or

Free public preview · No account or payment required. At least 10 consecutive months; 24+ recommended.

How the model is picked
 
01Paste or uploadOne SKU, or a sheet of hundreds
02Up to 11 methodsEligible models compared for each SKU
03One-month-ahead rankingBased on backtest error and bias
04Excel and PDF reportsDownload the forecast, ranges and model comparison
05Nothing is uploadedEvery calculation runs in your browser

Croston · SBA · TSB · ADIDA · Holt-Winters · Holt · SES · ARIMA · naive · moving average · seasonal naive

Also here Models compared ADI and CV2 calculator Croston Excel template Intermittent demand guide

Forecast workspace

Monthly demand, oldest month first. Models are ranked on one-month-ahead backtests; longer-horizon forecasts are projections, not separately ranked winners.

ActualForecast P10–P90Before scenario

No benchmark yetPick a sample or paste your own monthly demand, then run the benchmark.

Not sure which model fits your demand?

Eleven models in four families. See how each one behaves on the same demand, where it is strong and where it lets you down. No data needed.

3Baselines · 3 modelsNaive, seasonal naive and moving average. The yardstick every other model has to beat.
3Exponential smoothing · 3 modelsSimple smoothing, Holt and Holt-Winters. Level, then trend, then season.
1Autoregressive · 1 modelARIMA. Forecasts the month-to-month change from the previous change.
4Sparse demand · 4 modelsCroston, SBA, TSB and ADIDA. Built for slow movers with many zero months.

Compare the models

How it picks

Four rules, all of them checkable. Open the full working below when you want it — every model, every formula, the coverage study.

01Classified firstADI and CV² decide which models may compete for a SKU. Intermittent demand never gets a model built for smooth demand.
02Judged on months it never sawRolling-origin backtesting. A model that fits history beautifully and forecasts badly loses here.
03Compared like for likeEvery eligible model starts from the same month, so the accuracy figures belong in the same column. The fold count is shown next to each.
04Understand the uncertaintyP10–P90 is a nominal 80% range. Synthetic testing shows that coverage varies by demand pattern and forecast horizon.
Read the full methodHide the method

How the benchmark works

Enough to judge whether the number is worth using. The rules, not the recipe.

The comparison

Scope of the test. Demand classification and automatic season detection use the supplied history before the backtest. Each model is fitted only on its training block, but the scores are conditional on that earlier candidate selection. They are not an independent evaluation of the complete selection process or a guarantee of future accuracy.

Every model starts from the same month of history, so the accuracy figures in the table were measured on the same data.

At each forecast origin a model sees only the history available at that moment and predicts forward. Move the origin along and you get a set of honest out-of-sample errors.

What usually goes wrong elsewhere is that different models need different amounts of history to start, so each gets scored from its own earliest origin — and accuracy measured on different data ends up in the same column. Here every eligible model starts at the same origin. If including seasonal models would leave too few evaluation folds, they are excluded outright and the page says so. The fold count sits next to every model so you can check.

Classification

ADI and CV² decide which models are allowed to compete for a SKU. They never decide the winner.

Average demand interval and the variability of the non-zero order sizes place each SKU in one of four quadrants — smooth, erratic, intermittent, lumpy — on the Syntetos–Boylan–Croston (2005) scheme. Check a single row with the ADI and CV² calculator. Smooth and erratic series go to the continuous-demand family; intermittent and lumpy go to the intermittent estimators, and a continuous-demand model cannot win there however good its error looks.

The eleven models

Seven for regular demand, four for intermittent. All published, all named.

Regular demand: Naive · Moving average (3) · SES / ETS(A,N,N) · Holt linear trend · Holt-Winters additive · ARIMA(1,1,0) with drift · seasonal naive with drift.

Intermittent demand: Croston · SBA · TSB · ADIDA.

Smoothing constants are fitted to each SKU rather than pinned at textbook defaults, and the fitting only ever sees the training block of the fold it is in. Ranking uses one-month-ahead rolling-origin error with a penalty for persistent bias; the score is in the table beside the error. The 3-, 6- and 12-month forecasts are not separate horizon-specific model selections.

The range

A P10–P90 that has been tested for coverage against held-out data, not asserted.

The spread of the rolling error is measured at each horizon from the folds that genuinely reach it, and grown beyond the last horizon the data supports — the chart marks that switch with hollow bars. Intermittent and lumpy series get a bootstrap instead of a normal curve, because their demand has a large lump at zero. The page says which method produced the numbers on screen.

Calibration. In the supplied reproducible study, 120 synthetic histories in each of six demand families — flat, trending, seasonal, random-walk, lumpy and intermittent — use automatic cycle detection, are fitted on 36 months and scored against the following 12. The latest run contained about 79–88% of monthly outcomes and 71–93% of cumulative 3- and 12-month outcomes, depending on the demand family. P10–P90 is a nominal 80% interval; these synthetic results do not guarantee coverage on an individual real-world series.

What the sheet should look like

SKU in column A, one column per month after it, oldest first. Zeros are data; keep them.

One row per SKU. Column A is the SKU code. Every other column represents one month, oldest first. Date headers (2025-01, Jan 2025, or Excel dates) must be consecutive, with no duplicates or missing months. Undated headers must be M1, M2, M3… (or 1, 2, 3…), and require a last actual month in the workspace. Headerless sheets also require this date. No season column is needed. Older templates with a Season column still import; that column is ignored.

Sheet layout
SKUOct 2023Nov 2023Dec 2023…
CHEM-1181,6231,8402,003…
CHEM-350307443251…
CHEM-41105140…

Zeros are data. A zero means no demand that month and the classification depends on it, so leave them in rather than deleting the cell — deleting it shifts every later month back by one.

Blanks. Missing demand cells anywhere in a row are reported, including the first and last periods. They are never removed or converted to zero. Resolve missing observations in the source data. For a newly launched SKU, analyse its actual history separately with the correct last actual month; do not invent pre-launch zero demand.

Negatives stop that SKU from being analysed. Resolve returns and credits in the source data before importing; the tool never silently replaces them with zero.

How much history. Ten months minimum, twenty-four before the backtest is worth much. A seasonal model needs two full cycles plus four evaluation folds — about 28 months for an annual cycle — or it is excluded and the page says so.

Automatic detection. The tool removes a linear trend, then checks for a repeating 3-, 4-, 6- or 12-month cycle. A cycle needs two full repetitions plus four evaluation periods and must clear the correlation checks. Short, noisy or changing histories can have no reliable cycle detected. The tool then compares non-seasonal methods; sparse-demand methods are used for intermittent and lumpy histories.

Known limitations

Where the numbers get thin, stated up front.

The seasonal entry is a baseline, not an order-searched SARIMA. The calibration study uses synthetic histories with clean structure; a real SKU with a regime change part-way through will be harder than anything in it. Intervals inherit the length of your history — with only a handful of folds the range is indicative, and the fold count is shown so you can judge it. Promotions, price changes, cannibalisation and new-product effects are not modelled; the scenario tool is where you put what you know and the history does not.

Sources

The cut-offs, the estimators, the error measures and the backtest design all come from published work, not house opinion.

  1. Croston, J. D. (1972). Forecasting and stock control for intermittent demands. Operational Research Quarterly, 23(3), 289–303. The original size-and-interval method. Journal record
  2. Syntetos, A. A. & Boylan, J. E. (2001). On the bias of intermittent demand estimates. International Journal of Production Economics, 71(1–3), 457–466. Croston's bias and the SBA correction. Journal record
  3. Syntetos, A. A., Boylan, J. E. & Croston, J. D. (2005). On the categorization of demand patterns. Journal of the Operational Research Society, 56(5), 495–503. The ADI / CV² quadrants and the 1.32 and 0.49 cut-offs. Journal record
  4. Kostenko, A. V. & Hyndman, R. J. (2006). A note on the categorization of demand patterns. Journal of the Operational Research Society, 57(10), 1256–1257. A refined boundary between Croston and SBA. Paper (PDF)
  5. Hyndman, R. J. & Koehler, A. B. (2006). Another look at measures of forecast accuracy. International Journal of Forecasting, 22(4), 679–688. MASE, the scaled error in the model table. Paper (PDF)
  6. Teunter, R. H., Syntetos, A. A. & Babai, M. Z. (2011). Intermittent demand: linking forecasting to inventory obsolescence. European Journal of Operational Research, 214(3), 606–615. TSB, which keeps updating through zero-demand runs. Journal record
  7. Hyndman, R. J. & Athanasopoulos, G. (2021). Forecasting: Principles and Practice, 3rd ed., section 5.10. OTexts. Rolling-origin (time-series) cross-validation. Online text