Inventory

Forecast Accuracy Formula: Check Whether Your SKU Forecasts Are Good Enough to Reorder

Which forecast accuracy formula should you use before a reorder?

Use mean absolute percentage error (MAPE) as the headline forecast accuracy formula, then check forecast bias before you place the next purchase order. MAPE tells you how far the forecast was from actual sales. Bias tells you whether the historical miss was mostly high or mostly low.

The two numbers a Shopify merchant needs are a forecast accuracy convention and a bias measure. Forecast accuracy (%) = 100 − MAPE is a derived convention, where MAPE is the average of |Forecast − Actual| / Actual × 100. It is not a universally standardized accuracy measure, and it is not bounded at zero: when MAPE exceeds 100%, the result is negative. Forecast bias (%) = average of (Forecast − Actual) / Actual × 100. A positive bias means historical forecasts were high on average. A negative bias means they were low on average. The sign describes that past series. It does not tell you which way the next forecast will miss.

MAPE is the metric most inventory teams can explain in a buying meeting. Weighted MAPE (WMAPE) is the better store-level number when a few bestsellers dominate revenue, because it weights error by actual units instead of treating every SKU equally. Mean absolute error (MAE) is useful when a product had zero sales and a percentage would break.

Do not trust a single-period accuracy score. A SKU can look 95% accurate in a quiet week and still stock you out in the week you reorder.

Step 1: Compare forecasts and actual sales over the same period

Pull the forecast you would have used to reorder, and the units actually sold, for the same SKU and the same date range. If the forecast was weekly, sum actual sales by week. Mixing a monthly forecast with daily sales makes the error look worse than the buying decision really was.

Score the forecast against net units sold, not gross orders. Refunds and cancelled orders inflate demand if you leave them in. Exclude periods where the SKU was out of stock for most of the window, or the actuals are censored: you sold what you had, not what customers wanted.

Shopify's inventory reports are a useful place to pull ending inventory and a quantity-sold figure, but do not treat that quantity as net units. Shopify's own description of the inventory report says Quantity sold does not reflect returns. Before you score a forecast, subtract returned units for the same SKU and date range, or reconcile sold units against a sales or returns report that includes returned quantities. Cancelled orders that never shipped should also come out of the actuals. If you skip that adjustment, the forecast will look worse than the demand you can actually replenish.

  • One row per SKU per period: forecast units, actual units sold, days in stock.
  • Same calendar window for both columns. Do not compare a 30-day forecast to 28 days of sales.
  • Adjust Quantity sold from the inventory report for returns before you call the column net units.
  • Drop or flag periods with stockouts longer than a day or two. Those actuals understate demand.
  • Keep at least 8 to 12 periods if you can. One month is a story, not a track record.

If you do not store old forecasts, start now. Write down the number you believed when you last reordered, even in a spreadsheet column next to the PO date. Accuracy cannot be reconstructed from today's sales chart alone.

Step 2: Calculate percentage error with a worked SKU example

Take one SKU across four weeks. Forecasts were 40, 50, 45, and 60 units. Actual sales were 32, 55, 30, and 70.

WeekForecastActualAbsolute errorAPE
14032825.0%
2505559.1%
345301550.0%
460701014.3%

Absolute percentage error (APE) for week 1 is |40 − 32| / 32 = 25%. Repeat that for each week, then average the four APEs: (25.0 + 9.1 + 50.0 + 14.3) / 4 = 24.6% MAPE. Using the 100 − MAPE convention, that prints 75.4%. The same convention is not a 0 to 100 score: a MAPE above 100% would produce a negative result.

WMAPE on the same SKU is total absolute error divided by total actuals: (8 + 5 + 15 + 10) / (32 + 55 + 30 + 70) = 38 / 187 = 20.3%. Applying the same 100 − error convention to WMAPE gives about 79.7%, with the same limitation if error exceeds 100%. The gap between MAPE and WMAPE is the point. Week 3 was a large percentage miss on a smaller week, so unweighted MAPE punishes it harder.

There is no universal pass mark. If you need a starting sketch on your own sheet, treat roughly 20 to 30% MAPE as a review flag for bestsellers at the horizon you buy against, and roughly 40% MAPE over the last 8 to 12 periods as a point where you stop treating that forecast as a purchase quantity. Those cutoffs are illustrative operating heuristics, not published benchmarks. Long-tail SKUs will usually look worse. Write your own thresholds down and revisit them.

Step 3: Handle zero-sales products and misleading averages

MAPE divides by actual sales. A week with zero units sold makes the formula undefined. Do not replace zero with 1 to force a percentage. That invents a 100% or larger error that has nothing to do with how you buy.

For intermittent SKUs, report MAE in units instead: average of |Forecast − Actual|. A MAE of 2 units on a product that sells 0 to 3 units a week is a different decision than a MAE of 2 on a product that sells 80. Say both numbers when you review the catalog: percentage error for steady sellers, unit error for sparse sellers.

  • If actuals are zero and the forecast was also zero, count the period as a hit, not as undefined.
  • If actuals are zero and the forecast was not, record the absolute unit error and exclude the period from MAPE.
  • Do not average MAPE across the whole catalog and call it store accuracy. A 90% score on 200 dead SKUs hides a 40% miss on the 15 products that fund the business.
  • Prefer WMAPE, or MAPE on your top revenue SKUs only, when you report one number to the owner.

A store-level average also hides mix. Ten products at 10% error and two at 80% can still print a flattering mean. Sort by revenue, then by error. The reorder review should start with high-revenue, high-error SKUs, not with the catalog average.

Step 4: Check whether forecasts consistently over- or under-predict

Accuracy without direction will stock you out or bury you in inventory. On the same four weeks, signed percentage error is (Forecast − Actual) / Actual.

Week 1: (40 − 32) / 32 = +25%. Week 2: (50 − 55) / 55 = −9.1%. Week 3: (45 − 30) / 30 = +50%. Week 4: (60 − 70) / 70 = −14.3%. Average bias is (+25 − 9.1 + 50 − 14.3) / 4 = +12.9%. That +12.9% is the mean of the four weekly percentage biases, not the unit overforecast. Totals are 195 forecast units and 187 actual units, about a 4% aggregate overforecast. Two weeks ran high and two ran low. A positive average is a signal to investigate the series. It does not mean the next forecast is more likely to land high.

Tracking signal is a simple way to see if that bias is sticking. Add the signed errors in units (forecast minus actual): +8, −5, +15, −10. The running sum is +8. Mean absolute deviation is 9.5 units. Tracking signal = running sum of errors / MAD = 8 / 9.5 = 0.84. If you want a starting sketch, flag the series for investigation when the absolute tracking signal is around 4. Treat that cutoff as an illustrative operating heuristic, not a standard control limit. On four weeks you are nowhere near that sketch. Keep the series. A run of positive errors is a reason to inspect the forecast and any manual edits. It is not evidence that the next forecast will be high.

Under-forecast bias is the more expensive miss for a bestseller with a long supplier lead time, because the lost sale happens before the next container lands. Over-forecast bias is the more expensive miss for a seasonal or perishable SKU, because the extra units sit after demand drops. Read the sign against the product, not as a universal good or bad.

How to use the result before placing the next reorder

The forecast is an input to the buy quantity, not the buy quantity. Before you raise a PO, run this check on the SKU you are about to replenish.

  1. Confirm the accuracy window matches the lead time. If the supplier takes 6 weeks, score the forecast at a 6-week horizon, not next week's number.
  2. If MAPE on recent comparable periods is under your own heuristic bar and bias is near zero, use the forecast as the demand input in your reorder point.
  3. If MAPE is acceptable but bias is steadily negative, investigate the shortfall before you raise the demand input or safety stock. Repeated under-forecasting is a signal to check, not proof the next period will also run short.
  4. If MAPE is acceptable but bias is steadily positive, investigate before you cut the demand input. A tidy accuracy score can still sit next to over-ordering, but past positive bias does not prove the next forecast is high.
  5. If MAPE is poor, do not average your way out of it. Buy to a shorter cover, split the PO, or use recent actuals plus a manual adjustment until the error comes down.
  6. Recalculate after the receipt lands. The next accuracy row is how you learn whether the override helped.

A common reorder point is (average daily demand × lead time in days) + safety stock. Forecast accuracy belongs in both terms. Noisy error means you need more safety stock to hold the same service level. Systematic bias means the average daily demand you plugged in is already wrong, and extra safety stock will not fix a forecast that always runs 15% low.

On the worked SKU, do not turn +12.9% into a buy-quantity discount. Dividing a 200-unit lead-time forecast by 1.129, about 177 units, treats the mean of four weekly percentage biases as a correction factor. The unit totals do not support that factor, and a positive historical average does not establish that the next forecast will be high. Investigate what drove the misses, then set the purchase quantity from lead time, recent actuals, and safety stock. If you still apply a manual adjustment, label it as a judgment and write down why, so the next accuracy review can test that judgment rather than a formula.

Pair this check with a reorder formula you already trust, and with a demand forecast built from recent sales rather than last year's hope. If you are deciding which SKUs deserve this review first, start with products that already behave like a Shopify bestseller. Those are the forecasts that move cash.

What good enough looks like on a small catalog

You do not need a statistical forecast engine to run this. A sheet with forecast, actual, absolute error, APE, and signed error is enough for the SKUs you reorder every month. Review it on the same day you cut POs.

Three rules keep the review honest. Score the horizon you buy against. Separate percentage error from bias. Ignore catalog-wide averages when a handful of SKUs produce most of the revenue. If those three are in place, the forecast accuracy formula stops being a textbook exercise and becomes a reason to change the next order quantity.

Skymetrics stock intelligence is built for the reorder side of this problem: watching sell-through and flagging stockout risk before the bestseller is gone. The accuracy sheet is still your job if you want to know whether the number you are buying against has earned that trust.