The same situation
The procedure sees exactly the state the planner saw — not the cleaned-up hindsight of tomorrow.
A procedure must not decide because it sounds plausible, but because it has been measured against reality. Three building blocks belong to it: run alongside before intervening · learn from human objection · and show uncertainty instead of rounding it away.
A new procedure gets the same data as dispatch, makes the same call — and writes it down without carrying it out. Nobody notices it in day-to-day business, and yet every order produces a line of evidence.
The procedure sees exactly the state the planner saw — not the cleaned-up hindsight of tomorrow.
It does not award, it does not reserve, it sends nothing. A shadow run that changes something is not a shadow run.
What is stored is not only the proposal but the quantities behind it: distance, waiting time, standing days, cost. Otherwise it cannot be clarified later why it was wrong.
A model that looks good on held-out data has only proved that it knows the past. The question is a different one: would the decision have been better than the one actually taken?
In how many cases does the procedure propose anything at all? A procedure that only answers the easy cases is no relief.
How often does it make the same choice as the human? High agreement does not mean good — at first it only means inconspicuous.
Where it would have decided differently: how many empty kilometres, standing days or euros lie between the two paths? That is the only figure that counts.
If the procedure sees data while computing that was not yet available at the moment of decision, it always wins — and loses immediately in live operation.
A shadow run over the well-documented lanes says nothing about the difficult ones. Measurement covers the whole portfolio, or the figure does not hold.
The comparison assumes that everything else stays the same. Had the shipment been awarded differently, the next trailer would have stood elsewhere too — that limits every statement and belongs said out loud.
A procedure does not leave the shadow because it has run long enough, but because it has shown a measurable gap to previous practice across the whole portfolio — and because it can be explained where the gap comes from. If the gap fails to appear, the shadow run was not a failure but a wrong decision saved.
The shadow remains afterwards as well: every later change to the procedure runs alongside again before it takes effect. Whoever abolishes that notices a deterioration only in the monthly evaluation.
A suggestion without its context is worthless: you see that someone objected, but not to what. That is why every decision is filed as a complete case.
Which candidates were up for choice, with which quantities — distance, waiting time, standing days, price? Exactly as the procedure saw them.
Whom did the rule put in first place, and because of which quantity? Not only the result, but the reason.
Adopted, changed or rejected — and if changed: in favour of what. Without the alternative the override is only a no.
How did it turn out: actual empty kilometres, deadline kept, margin? Entered afterwards, once the trip has been driven.
It gets interesting when objections pile up — always the same region, the same lane, the same time of day. Then it is not an operating error but a rule that does not fit there.
“In the southern region, the radius is overridden in eight out of ten cases; the follow-on loads actually chosen were between 90 and 110 kilometres.” That is a finding.
Every change proposal names the cases it rests on. Whoever rejects the proposal can look up why it came about.
Proposing happens automatically, changing by hand — with date, before and after in the log. A rule that adjusts itself can no longer be explained.
Some overrides are habit, some an arrangement the system knows nothing about, some simply a mistake. That is why the outcome is measured too: whoever overrides regularly and regularly comes off worse changes no rule — they get to see the figures.
And the other way round: if the machine is consistently right, that is not a victory over the planner but the hint that their attention is needed elsewhere. The routine belongs to the machine, the exception to the human — that division is the goal, not being right.
The point value forces a decision the data cannot support: it states a time although reality is a window. Take it seriously and you plan too tightly; do not take it seriously and you ignore the forecast altogether.
A range that reaches beyond the customer's time window is a reason to act. A point value inside the window conceals exactly that.
On a well-run lane the range is narrow, on a new one wide. Exactly that difference is the information — it disappears as soon as you average.
If 80 per cent are promised, then across many cases roughly 80 per cent must also occur. A range can be held to account; a single point in time remains a matter of taste.
Not “when does it arrive”, but “how likely is it to keep the window”. That is the question the loading point asks.
How much capacity will be needed next week is never one number. Whoever plans only the mean stands without a trailer in half the weeks.
A target price with a bandwidth says where there is room to negotiate. A model price to the exact euro pretends there is none.
Measured beats calculated: a position confirmed by GPS and one derived from the plan never sit in the same field, and whatever a model contributed is marked as a model value. A range is the continuation of the same thought — it says not only where the number comes from, but also how firmly it stands.
And the touchstone remains the same as with every prediction: did it come true, and was it better than the previous planning? A forecast without follow-up measurement is an assertion — with or without a bandwidth.
We show how a shadow run is set up, what belongs in the record along the way, and which figures carry a decision afterwards.