Micro-stops: the losses your downtime log does not contain
Short stops are not small stops. They are the ones nobody writes down, which is why they are usually the largest single loss on a line and the hardest to argue about.
What counts as a micro-stop
A micro-stop is a stop short enough that nobody records it. That is the working definition, and it is more useful than a number, because the threshold is set by your logging system, not by physics. A stop the operator clears in forty seconds and a stop that needs a maintenance technician are different events, and only the second one reaches the downtime log.
The numeric conventions in common use come from three places, and none of them is a standard.
- The loss taxonomy. Classical TPM loss accounting separates breakdowns from idling and minor stoppages. The split is behavioural: a breakdown needs an intervention and a repair, a minor stoppage needs a hand. That is why the boundary sits in the region of a few minutes rather than at a round number.
- The logging system's own resolution. If the shop-floor terminal asks for a reason code only when a stop exceeds five minutes, the plant's micro-stop threshold is five minutes whether anyone decided that or not.
- The machine's restart behaviour. On a line where a jam clears itself and the machine resumes without an operator, the stop may never appear in any counter. On a line that needs a reset and a reference run, the same fault produces a longer, visible event.
Thresholds you will meet in practice are two, five and ten minutes. Automated monitoring systems publish their own detection floors — TeepTrak's product page, for example, states detection down to two-minute micro-stops. The number matters less than agreeing one and keeping it fixed, because every figure downstream is a function of it. Write the threshold into the measurement definition before anyone collects data, and never change it in the middle of a comparison. Disclosure: these sites are published by TEEPTRAK SAS, which sells production-monitoring and OEE software. Other vendors publish other floors, and a PLC counter or a spreadsheet can do the first pass for nothing.
Why manual logging misses them systematically
Not occasionally. Systematically, and in the same direction every time. Four mechanisms do the work.
| Mechanism | What happens | Effect on the number |
|---|---|---|
| Shift-report granularity | The report is a sheet with a line per event and a total that must add up to the shift. Its smallest usable unit is a few minutes. | Anything shorter than the smallest unit cannot be written down. It does not become a small number; it becomes zero. |
| The operator's decision rule | The operator is at the machine, clears the jam, and the machine is running again before a terminal could be reached. Logging the stop would cost more time than the stop did. | The operator is behaving correctly. The loss disappears anyway. Any method that depends on the person who fixed the fault also recording it will under-count. |
| Rounding to 5 or 15 minutes | Durations are entered in whole units. A 90-second stop is rounded to zero, or, if it is written down at all, inflated to the smallest unit. | Two opposite errors that do not cancel. Short events vanish; the few that survive are overstated, which makes the log look plausible while being wrong. |
| The balancing rule | Logged time must reconcile to the shift length. Whatever is unexplained is pushed into the largest named bucket or into a residual. | The log becomes internally consistent and externally false. Micro-stop time reappears as changeover, as maintenance, or as nothing at all. |
Observed behaviour of paper and terminal-based downtime logging; no survey figure is claimed here.
There is a fifth, quieter mechanism. Reason-code lists are written from maintenance history, so they contain codes for failures. They rarely contain a code for "cleared it and carried on". A loss with no code is a loss with no owner.
How micro-stops show up in the numbers
Because micro-stops are not logged as downtime, they do not reduce availability. They reduce the rate at which parts come off the machine while it is recorded as running. That is why the symptom is almost always the same complaint: the availability figure looks fine and the output does not.
- Performance loss, not availability loss. The lost time is inside recorded running time, so it lands in the performance term. A plant with clean availability and a stubborn performance gap is describing micro-stops, whatever it calls them.
- Cycle-time spread. Take the time between consecutive parts. A machine free of micro-stops gives a tight distribution around the standard cycle. A machine full of them gives a right-skewed distribution: the same median, a long tail, and a mean well above the median. The gap between mean and median is the cleanest single indicator you can compute from data you already have.
- Theoretical rate versus actual rate. The difference between parts you should have made in the recorded running minutes and parts you actually made, once scrap and logged stops are removed, is the residual. Micro-stops and slow running are the only two things left in it.
- Standards drifting upward. When a standard time is recalculated from recent actuals, the micro-stops inside those actuals are baked into the new standard. The loss then becomes invisible for a second reason: the plan already assumes it.
- Shift-to-shift variation without a cause. Two shifts, the same machine, the same product, five per cent apart in output and identical downtime logs. The difference is in the stops nobody wrote down.
Size the loss before you buy anything
You do not need a monitoring system to find out whether you have a problem worth solving. You need one shift of arithmetic and, at most, two hours of somebody's attention. Do this first; it also gives you the baseline any funding application will ask for.
- Fix the reference cycleEvery number below depends on it. Use the engineering standard if you trust it. If you do not, take the best sustained rate the machine has actually achieved over a clean hour in the last quarter, and use that until the standard is confirmed. State which one you used. A wrong reference cycle produces a confident wrong answer.
- Compute the residualFor one shift: planned run time, minus logged downtime, equals recorded running time. Multiply recorded running time by the reference rate to get expected parts. Subtract the parts actually produced, scrap included, because scrap consumed cycle time too. Convert the shortfall back into minutes at the reference cycle. That is the residual: time the machine was recorded as running and produced nothing.
- Split the residualThe residual is micro-stops plus slow running, and the two have different fixes. Separate them with a short observation study: one person, one machine, two hours, a phone timer, one line per stop with start time, duration and what was done. Count and mean duration give you the micro-stop share. Whatever is left is slow running.
- Cross-check with a counterIf the machine has a part counter, read it at a fixed interval — fifteen minutes is enough — for two shifts. Intervals that fall below the expected count locate the losses in time, and often in product or in operator. This costs nothing and survives scrutiny better than recollection.
- Decide with a rule, not a feelingSet the threshold before you see the answer. A sensible one: if the residual is under about 3% of recorded running time, the machine is not your problem — go and measure a different one. Between roughly 3% and 10%, fix the top two causes with standard work and a sensor check before spending money. Above that, and on a machine that constrains the plant, automatic capture pays for itself quickly, because the loss is distributed across hundreds of events a shift and cannot be reconstructed by hand.
A worked example
Illustration with assumed numbers. One assembly cell, one shift.
- Shift 480 min; planned breaks and planned maintenance 40 min; planned run time 440 min.
- Reference cycle 12 s per part, i.e. 5 parts per minute.
- Logged downtime: one 20-minute changeover and one 15-minute fault = 35 min. Recorded running time 405 min.
- Output: 1,730 good parts and 40 scrap = 1,770 parts produced.
- Expected at the reference rate: 405 × 5 = 2,025 parts.
- Shortfall 2,025 − 1,770 = 255 parts → 255 × 12 s = 3,060 s = 51 minutes.
- 51 / 405 = 12.6% of recorded running time, inside a shift whose downtime log balanced perfectly.
Now the two-hour observation study on the same cell: 14 stops, mean duration 55 seconds. That is 7 stops per hour. Over 405 minutes (6.75 hours) it projects to 47 stops × 55 s = 2,585 s = 43 minutes of micro-stops. The remaining 51 − 43 = 8 minutes is slow running — the machine cycling above 12 seconds, which is a different problem with a different owner.
In output terms: 43 minutes × 5 parts per minute = 215 parts per shift. On two shifts and 230 operating days, that is 460 shifts, 98,900 parts and about 330 machine-hours a year — roughly 41 eight-hour shifts of capacity. Whether that is worth money depends on whether this cell is the constraint and whether the parts can be sold; that arithmetic belongs on the cost-of-downtime pages, not here. What the example shows is the shape of the thing: 47 events, none of them individually worth writing down, together larger than both logged stops combined.
The five causes that dominate in discrete manufacturing
In the order they usually rank by total lost time — not by count, which is a different and less useful ranking.
| Cause | What it looks like | What to do | How you know it worked |
|---|---|---|---|
| Material presentation and feed | Misfeeds, double feeds, jams in bowl feeders, magazines, label and blister webs. Usually the largest single group. Correlates with lot, supplier and humidity rather than with the machine. | Treat it as an incoming-quality problem, not a machine problem. Check orientation and burr condition at the feeder, derate the feed rate until the stop rate falls, add a physical guide or an orientation poka-yoke on the infeed. | Stop rate by lot number. If the improvement is real, the between-lot spread narrows as well as the mean falling. |
| Sensor and detection faults | Photo-eyes fouled by dust, oil or coolant; drifting thresholds; false part-present; light curtains tripped by a chip or an operator's sleeve. Rises through the shift and resets after cleaning. | Put sensor cleaning into the shift start-up standard, check switching margin rather than just function, rigidify mounting so vibration cannot shift alignment, and swap optical for inductive or capacitive where the environment is dirty. | The within-shift trend. A cleaning-and-margin fix flattens the rising curve; if the curve still rises, the cause is elsewhere. |
| Blocking and starving | The machine stops because the next machine is full or the previous one is empty. It shows on the machine you are watching, but it belongs to a neighbour or to the buffer between them. | Record the stop state separately — blocked, starved, faulted — before doing anything else. Then size the buffer or fix the neighbour. Optimising a starved machine is wasted work. | The share of stops coded blocked or starved. It should fall to near zero on the constraint and can legitimately stay high elsewhere. |
| Product and tool variation | Stop rate climbs with tool wear, then falls sharply after a change. Also adhesive viscosity, cap torque, film thickness, weld tip condition. | Plot stop rate against tool life rather than against the calendar, and move the change point to where the curve turns up. Tie the same plot to lot and to set-point so the variable is identified, not guessed. | The curve flattens between changes, and total stop time falls even though you change tools more often. |
| Operator interventions that are not changeovers | Reloads, clearing, minor adjustment, in-process checks that require the machine to stop. Individually legitimate, collectively large, and invisible because nobody calls them downtime. | Separate the check from the stop where the process allows it, fix reach and access so a reload takes seconds rather than a minute, and standardise the intervention so it is the same length every time. | Mean duration of this class of stop, and its spread. Standard work shows up as a narrower distribution before it shows up as a lower mean. |
Ordered by typical share of lost time in discrete assembly and machining. Your own Pareto may differ — build it before acting on this one.
Two rules cut across all five. Rank by total lost time, not by event count: sixty one-minute stops beat four five-minute stops, and a count-based Pareto will send you to the wrong machine. And fix the measurement before the cause: if stops are not separated into blocked, starved, faulted and intervention, every Pareto you build is a mixture of four different problems.
If you decide to automate the capture
Three honest options, in increasing cost. A PLC or counter tap into a spreadsheet, which is free and adequate for one machine and a few weeks. An existing MES or SCADA historian, if you already have one and it timestamps parts rather than shifts. A dedicated monitoring system, which is the only option that scales to a plant without adding a job. Whichever you choose, agree the micro-stop threshold, the state definitions and the retention period in advance — and if the project is going to carry a funding application, write the baseline down before the sensors go on, because a baseline reconstructed afterwards from a shift report will not survive the verification visit.
One sheet per country with the rate, the caps and the status, an eligibility checklist, and a budget model that applies the cap for you. Excel, no macros.
- Funding map: Western Europe and Türkiye, status 22 September 2026
- Funding stack calculator: net cost after the instrument
- Western Europe Funding & Tax-Credit Pack (xlsx)
Questions
- What length of stop is a micro-stop?
- There is no standard that fixes a number. In practice plants use two, five or ten minutes, and the threshold is usually set by whatever the logging system can record rather than by a decision. The useful definition is behavioural: a stop the operator clears without calling anyone, which therefore never reaches the downtime log. Pick a threshold, write it into the measurement definition and keep it fixed, because every figure downstream depends on it.
- Why do micro-stops reduce performance rather than availability?
- Because they happen inside time that the system has already recorded as running. Availability is computed from logged stops; if a stop is not logged, availability does not move. The lost time shows up instead as a gap between the parts you should have produced in the recorded running minutes and the parts you actually produced. That is why a plant can report healthy availability and still miss its schedule every week.
- Can I measure micro-stops without buying anything?
- Yes, for a first pass. Compute the residual for one shift — recorded running time × reference rate, minus parts produced — then run a two-hour observation study with a timer to split that residual into micro-stops and slow running. Read the machine's part counter every fifteen minutes for two shifts to locate the losses in time. That is enough to decide whether the problem is worth spending money on, and it produces the baseline a funding application will ask for.
- Our downtime log balances to the shift. Does that mean it is right?
- No, and a log that balances perfectly is mildly suspicious. Logged time is usually reconciled to the shift length, so any unexplained minutes are absorbed into the largest named category or into a residual. The log becomes internally consistent while the short stops it cannot represent are silently redistributed. Compare the log against the parts counter rather than against itself.
- Should I chase every micro-stop?
- No. Rank causes by total lost time, not by how often they occur, and start with the top two on a machine that actually constrains the plant. A machine that is blocked or starved should not be optimised at all until the neighbour or the buffer is fixed — the stops belong to something else. If the residual is under about 3% of recorded running time, go and measure a different machine.
Sources
- TeepTrak PerfTrak — stated detection down to 2-minute micro-stops
- OEE definitions, benchmarks and cost-of-downtime arithmetic
Published by TEEPTRAK SAS, which makes production-monitoring and OEE software. Every figure is sourced on the page. Funding rules, standards and reporting duties change: check the official documents before you budget or commit.