Forecasting · 6 min read

Why forecast overrides quietly destroy accuracy — and how to measure it

Most planning teams let people adjust the statistical forecast every cycle, and never check whether those adjustments actually help. Often, they make the forecast worse. Here's the one number that tells you the truth.

Every demand planning process has the same quiet ritual. A statistical engine produces a baseline forecast. Then people — planners, sales, marketing, the regional manager who "knows the market" — reach in and change the numbers. A bump here for a promotion. A cut there because "that customer always over-orders." By the time the forecast reaches the S&OP table, it has been touched by a dozen hands.

Here's the uncomfortable question almost no one asks: did any of that touching make the forecast more accurate, or less?

The instinct is to assume human judgment adds value. Sometimes it does. But study after study — and most Health Checks we run — find the same thing: a meaningful share of manual overrides degrade accuracy versus simply leaving the statistical baseline alone. Teams are spending hours every cycle making their forecast worse, and they can't see it, because they never measure it.

The concept: Forecast Value-Add

Forecast Value-Add (FVA) is a deceptively simple idea borrowed from lean thinking: for every step in your forecasting process, ask whether it improved the forecast compared to the step before it. If a step doesn't add accuracy, it's waste — and waste in a forecast isn't just wasted time, it's the wrong stock in the wrong place.

You compare accuracy across a chain of forecasts:

Then you measure accuracy for each against actuals, and look at the deltas:

If your consensus forecast isn't more accurate than your statistical forecast, your overrides are destroying value. If your statistical forecast isn't more accurate than the naïve one, your engine (or its setup) is the problem.

How to measure it, concretely

You don't need a fancy system. You need three things stored side by side for each SKU-channel-period: the statistical forecast, the final forecast, and the actual. Most teams already have all three — they're just never lined up.

For accuracy, use whatever error metric you already trust — MAPE, WMAPE, or absolute error weighted by volume. Then compute two comparisons across your portfolio:

  1. Statistical vs naïve: is the engine earning its keep?
  2. Consensus vs statistical: are the overrides earning theirs?

The result that stops the room is usually the second one, sliced by who and where. When you show that overrides on one channel add three points of accuracy while overrides on another destroy five, you've turned a vague debate about "trusting the forecast" into a specific, fixable operational finding.

What to do once you can see it

The point of FVA isn't to ban human judgment — it's to aim it. Once you know where overrides help and where they hurt:

Teams that run FVA for a few cycles routinely find they can cut forecasting effort and improve accuracy at the same time — because they stop doing the work that was quietly hurting them.

The forecast you commit to sets your production, your purchasing, and your inventory. If a chunk of it is being made worse by hands that mean well, that error doesn't stay on a spreadsheet — it shows up as write-offs on the slow movers and lost sales on the ones you under-called. FVA is how you find it. It's usually the first thing we quantify in a Health Check, and it's frequently the finding that pays for the whole engagement.

Want to know whether your overrides help or hurt? That's exactly what the Health Check measures.

Book a call