Nothing went wrong. That’s the problem.
Every alarm you own is waiting for something to break. The expensive problems never break.
Ask an operator what their systems tell them and you will hear a list of exceptions. A run outside tolerance. An invoice past sixty days. A tank below the trigger. Every one of those is a threshold somebody set in advance, and every one of them fires the moment the threshold is crossed.
It is a reasonable way to build software and a poor way to protect a margin. Because the problems that actually cost real money almost never cross a line. They drift.
A scrap rate that moves from 2.1 per cent to 3.4 per cent over four months is inside tolerance on every single run. A packaging line that takes 34 days to turn over instead of 19 has not failed to deliver anything. A farm drawing feed eleven per cent faster than usual still has feed. Nobody gets an alert, because on any given day nothing is wrong.
What makes drift expensive is not the rate itself. It is that by the time it is visible in a monthly review, the cause is four changes back. Was it the resin supplier we swapped in June, or the tool that came back from refurbishment in May? Both are plausible. Neither is provable from memory, and the argument about which one it was costs more than the fix.
The alternative is not a better threshold. It is a system that knows what normal looks like for this operation and can say that something has been moving in one direction for long enough to matter — while it is still cheap.
That requires two things most tools do not have. The first is history: not industry benchmarks, but the last three winters on this farm, the last four hundred runs on this line, the last two years of turn rates on this product. The second is the ability to read several records against each other, because the cause is almost never in the same system as the symptom. Scrap sits in QA. The tool refurbishment sits in maintenance. The supplier change sits in procurement.
When those come together, the useful output is not an alert. It is a sentence with the records attached — the climb starts in the first week of May, before the supplier changed — and an operator who says May. Tool 4 came back from refurb then.
That exchange is the whole point. The system holds the numbers; the operator holds the context. Neither gets there alone, and the version where software tries to do both is the version nobody trusts.
The measure of a system like this is not how many notifications it sends. It is how few — and whether the ones it does send arrive early enough that the decision can be made calmly.
More notes.
By sector.
The rest of it is easier to show than describe.
A conversation