Andon Escalation Driven by Logistics Signals Best Practices

andon-escalation-logistics-integration-patterns-guide

An andon that fires too often gets ignored. An andon that fires too late is just an expensive record of a line that already stopped. The engineering problem in andon escalation logistics integration is finding the narrow band between those two failures — and that band is defined almost entirely by the quality of your logistics signals. A shortage andon raised at the moment a line-side rack empties is a stoppage. The same andon raised when predicted time-to-stockout drops below replenishment lead time is a recoverable event that nobody outside the aisle ever notices. This guide covers the trigger logic, the shortage prediction maths, tiered escalation design, the state machine that makes response measurable, and the suppression techniques that stop false andons from destroying trust in the system. Talk to a solutions engineer about your trigger design, or start a free trial to see equipment-driven events become tracked actions.


ENTERPRISE GUIDE · ANDON & ESCALATION DESIGN
Andon Escalation Driven by Logistics Signals
From shortage prediction to trigger, through tiered escalation, to closing the loop — with the suppression logic that keeps false stoppages out of the system.
THE SHORT VERSION
Predict, don't detect. Trigger on time-to-stockout versus replenishment lead time, not on an empty rack.
Tier by response, not severity. Every tier owns a clock and an owner; escalation is automatic, never discretionary.
Suppress before you escalate. Confirmation windows, hysteresis and changeover masking prevent most false andons.
Measure acknowledgement. Time-to-acknowledge tells you whether the system is trusted; time-to-resolve tells you whether it works.
01 — TRIGGER TAXONOMY

Not Every Problem Deserves the Same Andon

Andon fatigue comes from treating all conditions as equal. Classify every candidate trigger on two axes before it goes anywhere near an escalation rule: how confident the signal is, and how fast the consequence arrives. Only the top-right quadrant should ever be allowed to stop a line automatically.

CONSEQUENCE SPEED
FAST · LOW CONFIDENCE
Alert a human, do not act
Predicted shortage from a noisy consumption estimate. Raise to a team leader for a look — never auto-stop on this.
FAST · HIGH CONFIDENCE
Auto-andon, immediate tier
Confirmed empty sequenced rack, scanner-verified wrong part, conveyor jam with photo-eye confirmation. Act now.
SLOW · LOW CONFIDENCE
Log and trend only
Drifting cycle times, marginal consumption variance. These belong on a dashboard, not on an andon board.
SLOW · HIGH CONFIDENCE
Scheduled action
Confirmed slow depletion, approaching PM threshold, known supplier delay. Creates a task with a due time, not a call.
SIGNAL CONFIDENCE →
02 — SHORTAGE PREDICTION

The Only Maths You Need to Get Right

Shortage prediction is not machine learning in most plants. It is a subtraction problem with honest inputs. The mistake is comparing stock to a fixed minimum quantity instead of comparing time to time.

CORE CALCULATION
TimeToStockout = QtyAtLine ÷ ConsumptionRate
TriggerWhen: TimeToStockout ≤ ReplenLeadTime + SafetyMargin
Consumption rate comes from actual takt and scan consumption, not from the standard BOM rate. Replenishment lead time is measured from your own milk-run history, not from the route plan.
WORKED EXAMPLE
Qty at line180 pcs
Consumption rate3.0 pcs/min
Time to stockout60 min
Replenishment lead time42 min
Safety margin15 min
Trigger threshold57 min
Andon raises in 3 minutes — with 57 minutes of runway, not zero.
WHERE EACH INPUT COMES FROM
QtyAtLineScan events, rack weight or presence sensors via OPC UA, reconciled against WMS
ConsumptionRateRolling average of actual line takt and part usage, recalculated per shift and per model mix
ReplenLeadTimeMeasured p85 of historical call-to-delivery times, segmented by zone and shift
SafetyMarginTuned per part family — high for sequenced and single-source parts, low for bulk hardware
03 — ESCALATION DESIGN

Tiers That Escalate Themselves

The purpose of a tier is to answer one question: if nobody responds, what happens next and when? Escalation that depends on somebody deciding to escalate is not escalation. Every tier gets an owner, a clock and an automatic promotion rule.

T00–2 min
Operator
Self-resolve window. Signal visible at the station only. Most triggers die here, and that is the goal.
T12–8 min
Team Leader
Andon board lights the zone. Leader acknowledges, which starts the resolution clock and stops the promotion timer.
T28–20 min
Logistics Supervisor
Material issue confirmed. Expedite triggered, alternate rack or emergency milk run dispatched, ERP notified.
T320–40 min
Area / Shift Manager
Stoppage now probable. Decisions widen: resequence, build-out, or planned stop with recovery plan attached.
T440+ min
Plant & Supply Chain
Supplier and network escalation. Every T4 gets a post-event review — no exceptions, regardless of outcome.
Acknowledgement pauses promotion; it does not close the event. Every tier transition writes a timestamped record with identity. Skipping tiers is allowed upward, never downward.
Escalation only works if the resulting action is tracked to completion — the same discipline as work order management. Talk to a Solutions Engineer
04 — EVENT LIFECYCLE

The Andon State Machine

An andon without a defined lifecycle produces a count of alarms and nothing else. Define the states, who can move between them, and which clock is running in each — that is what turns andon data into a measurable response process.

StateEntered byExits whenClock running
RaisedRule engine or operatorSomeone acknowledgesTime to acknowledge
AcknowledgedNamed responderWork begins on the causeTime to acknowledge stops
In ProgressResponderCondition physically clearedTime to resolve
ResolvedResponderSignal confirms clear stateVerification window
VerifiedSystem, from live signalRecord closedNone
SuppressedRule — changeover or maintenanceMask window expiresAudit clock only
False PositiveResponder, with reason codeFeeds threshold tuningNone
Resolution must be confirmed by the signal, not by the responder clicking done. Self-reported closure is how a plant discovers that a rack has been empty for forty minutes while the board showed green.
05 — FALSE STOPPAGES

Filtering Before Escalating

Raw plant signals are far noisier than their tag names suggest. The gap between a signal changing and an andon deserving to exist is filled by suppression logic. Alarm rationalisation practice — the thinking behind ISA-18.2 and EEMUA 191 — applies directly here: every alarm must be actionable by someone, right now.

Raw signal changeseverything the floor emits
After debounce & hysteresischatter and boundary flapping removed
After context maskingchangeover, break, planned maintenance excluded
After confirmation windowcondition persisted long enough to be real
Andon raisedactionable, owned, on a clock
Confirmation window
Require the condition to hold for N seconds before raising. Catches transient sensor states and momentary blockages that clear themselves.
Hysteresis on thresholds
Raise at one threshold, clear at a different one. Prevents an andon flapping on and off as a value sits on the boundary.
Context masking
Suppress shortage logic during planned changeover, breaks and maintenance windows. Masked events are logged, never silently discarded.
Deduplication
One root condition produces one andon. Five sensors reporting the same jam is one event with five contributing signals.
Rate limiting per zone
Cap concurrent andons per area. Breaching the cap is itself an escalation — it means something systemic, not five separate problems.
False-positive feedback loop
Every event closed as false positive carries a reason code that feeds weekly threshold tuning. Without this, thresholds never improve.
06 — ARCHITECTURE

Inputs, Engine, Outputs

Keep the rule engine separate from both the signal sources and the notification targets. Coupling them is what makes a threshold change require a controls download — and a threshold you cannot change is a threshold nobody tunes.

INPUTS
Line-side presence and weight sensors Scan and consumption events PLC and conveyor state via PLC and SCADA signals WMS stock and open replenishment calls ERP order, sequence and supplier data Takt, model mix and schedule
▶
RULE ENGINE
Prediction calculation per part and zone Confidence scoring and classification Suppression, masking and dedupe Tier assignment and promotion timers State machine and audit trail Threshold tuning from outcomes
▶
OUTPUTS
Andon boards and station lights Mobile and radio notification by tier Expedite or emergency milk-run call ERP notification and maintenance order Live response dashboard Post-event records and analytics
07 — STRATEGY COMPARISON

Four Ways to Trigger, Honestly Compared

Most plants end up with a hybrid, but should choose deliberately rather than accumulate one by accident.

StrategyWarning TimeFalse PositivesData NeededTuning EffortBest Fit
Manual pull / cordNoneVery lowNoneNoneAlways keep as a parallel path
Fixed minimum quantityShortLowStock count onlyLowStable demand, bulk parts
Predictive time-to-stockoutLongMediumConsumption + lead timeMediumSequenced and JIT parts
Hybrid predictive + confirmedLongLowAll of the aboveHighAutomotive OEM lines
Tune triggers against your own history, not a vendor default
A solutions engineering session reviews your signal sources, shortage prediction inputs, tier timings and suppression rules — and identifies which andons should be creating maintenance work rather than repeat callouts.
Trigger classificationThreshold tuningTier & SLA designFalse-positive review
08 — MEASUREMENT

Six Numbers That Tell You If It Works

Andon counts alone measure nothing. These six, tracked per zone and per shift, tell you whether the system is trusted, accurate and actually preventing stoppages.

Mean Time to Acknowledge
Raised → first acknowledgement
Rising MTTA means responders have stopped believing the signal.
Mean Time to Resolve
Acknowledged → signal-verified clear
Measures the response process, not the alarm quality.
False Positive Ratio
False closures ÷ total andons
The single strongest predictor of whether the system survives its first year.
Escalation Leakage
Events promoted past target tier
High leakage means tier timings or staffing are wrong, not that people are slow.
Prevented Stoppages
Resolved before predicted stockout
The actual return on the whole programme — track it from day one.
Andon Rate per Shift
Events ÷ shift, by zone
A ceiling everyone agrees on beforehand keeps alarm load humane.
09 — ROLLOUT

Deployment in Four Phases

Weeks 1–3
Observe Only
Run the rule engine silently against live signals. No boards, no notifications. Compare what it would have raised against what actually happened.
Weeks 4–6
Shadow Tiers
Notify one pilot zone's team leader only. Tune thresholds and confirmation windows against real false-positive reason codes.
Weeks 7–10
Live Pilot
Full tier escalation in one zone with boards and SLAs active. Weekly tuning review. Manual path stays available throughout.
Weeks 11+
Scale & Govern
Zone-by-zone rollout with per-zone thresholds. Monthly rationalisation review to retire triggers nobody acts on.
10 — QUESTIONS

Frequently Asked Questions

Should an andon ever stop the line automatically?
Only for high-confidence, fast-consequence conditions where continuing causes quality or safety damage — confirmed wrong part in a sequenced build, for example. Predicted shortages should never auto-stop, because the whole point of prediction is to create time for a human to prevent the stop.
How do we stop andon fatigue?
Set a per-zone rate ceiling before go-live and treat breaching it as a defect in the rules, not a busy day. Require every trigger to name the person who will act on it and what they will do. Triggers that fail that test get logged to a dashboard instead of raised.
What accuracy is realistic for shortage prediction?
It depends far more on consumption-rate quality than on algorithm sophistication. Stable takt with scan-confirmed consumption predicts well; high model-mix variability with inferred consumption does not. Measure your own prediction error against actuals during the observe-only phase before committing to thresholds.
Where does this sit relative to SAP and MES?
The rule engine typically consumes stock and order context from ERP and real-time state from the plant floor, then pushes notifications and documents back. Keeping it as a separate layer means threshold changes do not require an ERP transport or a PLC download — which is what makes ongoing tuning realistic.
Do we still need the physical andon cord?
Yes. Operators see conditions no sensor covers, and removing the manual path signals that the automated system is considered infallible. Keep the cord, and treat every manual pull that the rule engine missed as a gap worth investigating.
How do andons connect to maintenance?
Recurring equipment-caused andons are maintenance demand wearing a different hat. Routing them into work orders with the asset attached turns repeat callouts into a fix — the approach in our preventive maintenance guide and predictive maintenance guide.
Close the Loop Between Andon and Action
FleetRabbit turns recurring equipment events into tracked work orders, usage-based PM schedules and complete asset history — so the andon that fires three times a week becomes a repair with an owner and a due date.
September 12, 2026 By Alex Rowan
All Articles

Share This Story, Choose Your Platform!

Latest Articles

Scroll