Metrics and AI

Every figure comes with what it means

A report shows 18 per cent empty running — what follows from it is written nowhere. Here it stands next to it: the metric on the left, the reading on the right, the conclusion below, and the slider moves the cut-off date.

Try it yourself

The same metric, read twelve times

The empty-run rate of one lane over a year. The slider moves the cut-off date; drawing changes, and the reading with it. It is not written but calculated from the figures — level, direction, breaking point, cause, money.

Empty running 24 % +4.3 points against the previous quarter
Break detected Jul there it still counted as unremarkable
Above the half-year +10 pts 14 % was the flat level until June
Cost a year €11,500 on this one lane
The reading

Dec stands at 24 %. That is critical, and the series is rising.

Against the previous quarter that is +4.3 points — measured over three months, not against last month.

It began in Jul. Back then the rate stood at 17 % and counted as unremarkable — the direction was already visible.

The growth sits in one cause: the missing return load, +5 points in three months. The other two are standing still.

Above the first half-year level that is roughly €11,500 a year on this one lane.

The conclusion

Recalculate the price of the lane.

At this level the agreed price no longer carries the empty leg. Either a return load comes with it or the lane gets renegotiated — both are decisions, not reports.

Who acts: Sales and purchasing

What July showed

17 per cent. Below every threshold, no colour, no alarm. A report would have had nothing to say that day.

What July already held

Three months against the three before: +1.7 points, and the movement sits entirely in one of the three causes. Four months before the traffic light changes.

A report answers “how much”. In operations the question is “and now?” — in between lies the work a machine can take over.

Explainer · about 35 secondsFrom the number to the reading to the conclusion

    The report shows 18 per cent. It says no more, and it can say no more.

    Starts on its own · pause any time
    The conclusion

    One number, four recipients

    24 per cent in December. It means something different to everyone in the building — and to one of them it means nothing. That belongs in the answer, or everybody reads everything.

    Planning

    Not the rate, the lane. Which trips get planned differently over the next two weeks — and whether there is a return load nobody is asking for today.

    Purchasing

    Buy a return load on this lane. The shortfall is larger than the effort of finding it; that is the whole comparison.

    Sales

    The agreed price no longer carries the empty leg. Either a return load comes with it or the lane gets renegotiated.

    Management

    One lane, some eleven thousand euro a year. It becomes interesting when ten lanes show the same pattern — and that is exactly what the next reading asks about.

    A finding without a recipient reaches everyone and therefore nobody. That is why dashboards go unopened after three months.

    How the machine becomes capable of it

    Shadow mode

    A model that starts planning on day one has no yardstick. Whether it was any good shows up in the invoice — three months later.

    In shadow mode it runs alongside and decides nothing: it files a proposal for every case, the human decides as always, and the two are laid side by side. That costs no sign-off and no risk — it costs compute.

    Agreement is the boring half. What matters are the differences, both ways: if the model is wrong, it is missing something nobody ever wrote down. If the human was wrong, a mistake is lying in the open that nobody had seen — nobody expects that one, and it is the more valuable.

    Where the model stands today — and where it does notProposal against decision, per case class, over three months. “Standard lane”: known customer, known route.
    release from 90% Standard lane Standard lane: 96 % 96 % Return load from stock Return load from stock: 93 % 93 % New delivery point New delivery point: 88 % 88 % Dangerous goods Dangerous goods: 74 % 74 % Special transport Special transport: 61 % 61 %
    released — runs without a querystays with a humanrelease from 90%

    Example data. Two classes are ready, three are not. The gap between the third and the fourth is why “everything from now on” would be wrong — release happens per class.

    StageWhat the model may doWhat is measuredWhen the next stage comes
    1 · Running alongNothing. It proposes, visible only for the evaluation.Agreement per case class, every difference with its reason.When one class agrees for three months.
    2 · ProposingThe field is pre-filled, the human confirms or changes it.How often it is changed — and at which point.When the change rate falls and stays down.
    3 · Acting in one classIt decides the simple class alone.Errors that get through. Not the hit rate — the ones that slip past.When none slips through over a period.
    4 · Acting with a sampleIt decides. A share is still put in front of a human.The same sample, indefinitely.Never. This stage does not end — the reason is below.

    The trap at the end: once everything is released, learning stops. The model only sees its own decisions, confirms itself and slowly gets worse without anyone noticing. That is why a sample stays with a human: it is not the control, it is the lesson.

    Shadow mode is not the cautious variant, it is the only one that produces a yardstick. Without it, “better than the human” is a claim.

    Explainer · about 35 secondsRun along, compare, release

      The human decides as always. Nothing about the process changes.

      Starts on its own · pause any time
      What runs back

      The loop

      The metric is not the end of the chain but its entrance: reading, proposal, decision, changed process — and that changes the metric it all starts with again.

      Difference = lesson Metric Reading Proposal Decision Process Effect
      The dashed way back is the effect: what was decided shows up in next month’s figure. The branch downwards is the lesson — every difference goes back into the model.

      And it closes measurably: the reading in July instead of November is four months on one lane — some €4,600 here. Not through a better model, but because somebody was asked earlier.

      A process does not get more efficient because the machine calculates faster, but because the decision falls earlier.

      Explainer · about 35 secondsWhat runs back from the decision

        The metric is not the end of the chain but its entrance.

        Starts on its own · pause any time
        The limit

        What the rule can do, where the model begins

        What you saw above is calculated by no language model but by a handful of rules: two thresholds, a three-month mean, a comparison of causes. That is the honest and the useful answer — most readings need no model, just somebody to write the thresholds down once.

        What the rule does

        Level, direction, breaking point, cause, money. Everything the series yields by itself. Checkable, repeatable, with no connection to the outside.

        Where the model begins

        At the cause nobody anticipated. At the sentence in the mail traffic explaining why the return load fell away. At the question nobody wrote a threshold for.

        Where the human stays

        Where a decision costs money and somebody answers for it. The machine prepares, it does not sign.

        Best practice · bench learning

        Where the best solution sits today

        What is named is the pattern, not the vendor — a product name ages, an idea does not.

        Shadow mode before release

        The pattern behind every roll-out that held: run alongside, propose, release — per case class and never in one go.

        The reading with the number

        Not in a footnote and not in the next meeting. A metric that does not bring its reading along gets read and forgotten.

        The sample stays

        Even after release. It costs little and is the only thing that stops a model from confirming itself.