An andon that fires too often gets ignored. An andon that fires too late is just an expensive record of a line that already stopped. The engineering problem in andon escalation logistics integration is finding the narrow band between those two failures — and that band is defined almost entirely by the quality of your logistics signals. A shortage andon raised at the moment a line-side rack empties is a stoppage. The same andon raised when predicted time-to-stockout drops below replenishment lead time is a recoverable event that nobody outside the aisle ever notices. This guide covers the trigger logic, the shortage prediction maths, tiered escalation design, the state machine that makes response measurable, and the suppression techniques that stop false andons from destroying trust in the system. Talk to a solutions engineer about your trigger design, or start a free trial to see equipment-driven events become tracked actions.
ENTERPRISE GUIDE · ANDON & ESCALATION DESIGN
Andon Escalation Driven by Logistics Signals
From shortage prediction to trigger, through tiered escalation, to closing the loop — with the suppression logic that keeps false stoppages out of the system.
THE SHORT VERSION
Predict, don't detect. Trigger on time-to-stockout versus replenishment lead time, not on an empty rack.
Tier by response, not severity. Every tier owns a clock and an owner; escalation is automatic, never discretionary.
Suppress before you escalate. Confirmation windows, hysteresis and changeover masking prevent most false andons.
Measure acknowledgement. Time-to-acknowledge tells you whether the system is trusted; time-to-resolve tells you whether it works.
01 — TRIGGER TAXONOMY
Not Every Problem Deserves the Same Andon
Andon fatigue comes from treating all conditions as equal. Classify every candidate trigger on two axes before it goes anywhere near an escalation rule: how confident the signal is, and how fast the consequence arrives. Only the top-right quadrant should ever be allowed to stop a line automatically.
CONSEQUENCE SPEED
FAST · LOW CONFIDENCE
Alert a human, do not act
Predicted shortage from a noisy consumption estimate. Raise to a team leader for a look — never auto-stop on this.
FAST · HIGH CONFIDENCE
Auto-andon, immediate tier
Confirmed empty sequenced rack, scanner-verified wrong part, conveyor jam with photo-eye confirmation. Act now.
SLOW · LOW CONFIDENCE
Log and trend only
Drifting cycle times, marginal consumption variance. These belong on a dashboard, not on an andon board.
SLOW · HIGH CONFIDENCE
Scheduled action
Confirmed slow depletion, approaching PM threshold, known supplier delay. Creates a task with a due time, not a call.
SIGNAL CONFIDENCE →
02 — SHORTAGE PREDICTION
The Only Maths You Need to Get Right
Shortage prediction is not machine learning in most plants. It is a subtraction problem with honest inputs. The mistake is comparing stock to a fixed minimum quantity instead of comparing time to time.
03 — ESCALATION DESIGN
Tiers That Escalate Themselves
The purpose of a tier is to answer one question: if nobody responds, what happens next and when? Escalation that depends on somebody deciding to escalate is not escalation. Every tier gets an owner, a clock and an automatic promotion rule.
T00–2 min
Operator
Self-resolve window. Signal visible at the station only. Most triggers die here, and that is the goal.
T12–8 min
Team Leader
Andon board lights the zone. Leader acknowledges, which starts the resolution clock and stops the promotion timer.
T28–20 min
Logistics Supervisor
Material issue confirmed. Expedite triggered, alternate rack or emergency milk run dispatched, ERP notified.
T320–40 min
Area / Shift Manager
Stoppage now probable. Decisions widen: resequence, build-out, or planned stop with recovery plan attached.
T440+ min
Plant & Supply Chain
Supplier and network escalation. Every T4 gets a post-event review — no exceptions, regardless of outcome.
Acknowledgement pauses promotion; it does not close the event.
Every tier transition writes a timestamped record with identity.
Skipping tiers is allowed upward, never downward.
04 — EVENT LIFECYCLE
The Andon State Machine
An andon without a defined lifecycle produces a count of alarms and nothing else. Define the states, who can move between them, and which clock is running in each — that is what turns andon data into a measurable response process.
StateEntered byExits whenClock running
RaisedRule engine or operatorSomeone acknowledgesTime to acknowledge
AcknowledgedNamed responderWork begins on the causeTime to acknowledge stops
In ProgressResponderCondition physically clearedTime to resolve
ResolvedResponderSignal confirms clear stateVerification window
VerifiedSystem, from live signalRecord closedNone
SuppressedRule — changeover or maintenanceMask window expiresAudit clock only
False PositiveResponder, with reason codeFeeds threshold tuningNone
Resolution must be confirmed by the signal, not by the responder clicking done. Self-reported closure is how a plant discovers that a rack has been empty for forty minutes while the board showed green.
05 — FALSE STOPPAGES
Filtering Before Escalating
Raw plant signals are far noisier than their tag names suggest. The gap between a signal changing and an andon deserving to exist is filled by suppression logic. Alarm rationalisation practice — the thinking behind ISA-18.2 and EEMUA 191 — applies directly here: every alarm must be actionable by someone, right now.
Raw signal changeseverything the floor emits
After debounce & hysteresischatter and boundary flapping removed
After context maskingchangeover, break, planned maintenance excluded
After confirmation windowcondition persisted long enough to be real
Andon raisedactionable, owned, on a clock
Confirmation window
Require the condition to hold for N seconds before raising. Catches transient sensor states and momentary blockages that clear themselves.
Hysteresis on thresholds
Raise at one threshold, clear at a different one. Prevents an andon flapping on and off as a value sits on the boundary.
Context masking
Suppress shortage logic during planned changeover, breaks and maintenance windows. Masked events are logged, never silently discarded.
Deduplication
One root condition produces one andon. Five sensors reporting the same jam is one event with five contributing signals.
Rate limiting per zone
Cap concurrent andons per area. Breaching the cap is itself an escalation — it means something systemic, not five separate problems.
False-positive feedback loop
Every event closed as false positive carries a reason code that feeds weekly threshold tuning. Without this, thresholds never improve.
06 — ARCHITECTURE
Inputs, Engine, Outputs
Keep the rule engine separate from both the signal sources and the notification targets. Coupling them is what makes a threshold change require a controls download — and a threshold you cannot change is a threshold nobody tunes.
INPUTS
Line-side presence and weight sensors
Scan and consumption events
PLC and conveyor state via PLC and SCADA signals
WMS stock and open replenishment calls
ERP order, sequence and supplier data
Takt, model mix and schedule
▶
RULE ENGINE
Prediction calculation per part and zone
Confidence scoring and classification
Suppression, masking and dedupe
Tier assignment and promotion timers
State machine and audit trail
Threshold tuning from outcomes
▶
OUTPUTS
Andon boards and station lights
Mobile and radio notification by tier
Expedite or emergency milk-run call
ERP notification and maintenance order
Live response dashboard
Post-event records and analytics
07 — STRATEGY COMPARISON
Four Ways to Trigger, Honestly Compared
Most plants end up with a hybrid, but should choose deliberately rather than accumulate one by accident.
Tune triggers against your own history, not a vendor default
A solutions engineering session reviews your signal sources, shortage prediction inputs, tier timings and suppression rules — and identifies which andons should be creating maintenance work rather than repeat callouts.
Trigger classificationThreshold tuningTier & SLA designFalse-positive review
08 — MEASUREMENT
Six Numbers That Tell You If It Works
Andon counts alone measure nothing. These six, tracked per zone and per shift, tell you whether the system is trusted, accurate and actually preventing stoppages.
Mean Time to Acknowledge
Raised → first acknowledgement
Rising MTTA means responders have stopped believing the signal.
Mean Time to Resolve
Acknowledged → signal-verified clear
Measures the response process, not the alarm quality.
False Positive Ratio
False closures ÷ total andons
The single strongest predictor of whether the system survives its first year.
Escalation Leakage
Events promoted past target tier
High leakage means tier timings or staffing are wrong, not that people are slow.
Prevented Stoppages
Resolved before predicted stockout
The actual return on the whole programme — track it from day one.
Andon Rate per Shift
Events ÷ shift, by zone
A ceiling everyone agrees on beforehand keeps alarm load humane.
09 — ROLLOUT
Deployment in Four Phases
Weeks 1–3
Observe Only
Run the rule engine silently against live signals. No boards, no notifications. Compare what it would have raised against what actually happened.
Weeks 4–6
Shadow Tiers
Notify one pilot zone's team leader only. Tune thresholds and confirmation windows against real false-positive reason codes.
Weeks 7–10
Live Pilot
Full tier escalation in one zone with boards and SLAs active. Weekly tuning review. Manual path stays available throughout.
Weeks 11+
Scale & Govern
Zone-by-zone rollout with per-zone thresholds. Monthly rationalisation review to retire triggers nobody acts on.
10 — QUESTIONS
Frequently Asked Questions
Should an andon ever stop the line automatically?
Only for high-confidence, fast-consequence conditions where continuing causes quality or safety damage — confirmed wrong part in a sequenced build, for example. Predicted shortages should never auto-stop, because the whole point of prediction is to create time for a human to prevent the stop.
How do we stop andon fatigue?
Set a per-zone rate ceiling before go-live and treat breaching it as a defect in the rules, not a busy day. Require every trigger to name the person who will act on it and what they will do. Triggers that fail that test get logged to a dashboard instead of raised.
What accuracy is realistic for shortage prediction?
It depends far more on consumption-rate quality than on algorithm sophistication. Stable takt with scan-confirmed consumption predicts well; high model-mix variability with inferred consumption does not. Measure your own prediction error against actuals during the observe-only phase before committing to thresholds.
Where does this sit relative to SAP and MES?
The rule engine typically consumes stock and order context from ERP and real-time state from the plant floor, then pushes notifications and documents back. Keeping it as a separate layer means threshold changes do not require an ERP transport or a PLC download — which is what makes ongoing tuning realistic.
Do we still need the physical andon cord?
Yes. Operators see conditions no sensor covers, and removing the manual path signals that the automated system is considered infallible. Keep the cord, and treat every manual pull that the rule engine missed as a gap worth investigating.
How do andons connect to maintenance?
Close the Loop Between Andon and Action
FleetRabbit turns recurring equipment events into tracked work orders, usage-based PM schedules and complete asset history — so the andon that fires three times a week becomes a repair with an owner and a due date.
September 12, 2026
By Alex Rowan
All Articles