10 Questions Before Blaming Bad Luck on Failures
A conveyor motor fails for the third time this year. Someone shrugs, calls it bad luck, orders the replacement, and moves on. Bad luck and “it’s just old” are cousins. Both let a team stop looking a beat too soon. A landmark reliability study found that most equipment failures aren’t actually driven by simple age or wear at all.[1] They have real, findable causes, a wrong interval, a skipped diagnostic step, an operating condition nobody wrote down, that “bad luck” quietly skips past. A few questions before writing off a repeat failure usually reveal which one it actually is.
![]()
10 Questions to Ask Before Blaming “Bad Luck” for Repeat Equipment Failures
1. Has This Exact Failure Happened on This Asset Before?
Why ask this: A single failure can genuinely be random. A second failure of the same type on the same asset is a pattern, not chance.
If the answer is yes: Two occurrences of the same failure mode point to a specific cause, wear, a design limitation, an operating condition, not general unpredictability.
Next step: Pull the asset’s full failure history inside your CMMS software before ordering the replacement part, and check whether this exact failure mode already has a name in the record.
2. Do Other Identical Assets Show the Same Failure?
Why ask this: A failure that looks random on one machine can be highly consistent across ten identical machines.
If the answer is yes: The problem isn’t this specific unit. It’s something the whole fleet shares, a spec, an installation detail, or an environment.
Next step: Compare failure history across every identical asset inside asset management records, not just the one that just failed.
3. Did the Repair Actually Fix the Cause, or Just the Symptom?
Why ask this: Swapping the failed part restarts the equipment. It doesn’t necessarily explain why the part failed in the first place.
If the answer is “just the symptom”: The underlying cause is still there, waiting to produce the same failure again under the same conditions.
Next step: Require a documented root cause, not just a documented fix, on any failure that’s happened more than once.
4. Has Anything Changed in How This Equipment Actually Runs?
Why ask this: Equipment fails differently under a different load, a different raw material batch, or a different ambient condition than it did when it was installed.
If the answer is yes: The “random” failure might track perfectly against whatever changed, once someone actually checks.
Next step: Log operating conditions at the time of failure, temperature, load, shift, recent changes, so a real pattern has something to show up against.
5. Is the PM Interval Actually Matched to This Equipment’s Real Wear?
Why ask this: Many PM intervals still come from a generic OEM manual, not from how this specific asset wears under these specific conditions.
If the answer is no: The equipment could be failing right in the gap between scheduled checks, on a wear pattern the manual’s interval was never built for.
Next step: Compare failure dates against the preventive maintenance interval, and adjust the schedule if failures keep clustering right before the next PM.
6. Was the Replacement Part Actually Equivalent to the Original?
Why ask this: A generic or off-spec replacement part can look identical and still fail faster than the original did.
If the answer is uncertain: The “bad luck” might just be a part substitution nobody flagged as a variable worth tracking.
Next step: Log the specific part and supplier used in each repair, so a pattern tied to a particular batch or substitute becomes visible over time.
7. Did the Technician Have Everything Needed to Diagnose This Properly?
Why ask this: A rushed diagnosis under time pressure sometimes treats the obvious symptom and moves on, without confirming the actual cause.
If the answer is no: The “fix” may have addressed a guess, not a confirmed diagnosis, which means the real cause is still sitting there.
Next step: Check whether the closed work order documents a confirmed failure mode, or just a repair action.
8. Is the Failure Clustering Around a Specific Condition?
Why ask this: A failure that seems to strike at random can actually track tightly against a specific shift, process, or production run, once someone lines the dates up.
If the answer is yes: That’s not bad luck. That’s a variable nobody’s been tracking as a suspect.
Next step: Plot failure dates against shift schedules, production changes, and seasonal patterns before concluding there’s no pattern to find.
9. Would This Survive an Actual Root Cause Review?
Why ask this: “Bad luck” rarely comes out of a structured investigation. It’s usually the answer that gets accepted when nobody investigates at all.
If the answer is uncertain: That uncertainty is itself the signal that a real review hasn’t happened yet.
Next step: Set a threshold, based on the asset’s criticality, for when a repeat failure automatically triggers a root cause review instead of just another repair.
10. Has Anyone Looked at the Full History, or Just This One Incident?
Why ask this: Three failures spread across two years, three technicians, and three different notes can look unrelated, right up until someone puts them side by side.
If the answer is “just this one”: The pattern might already exist in the records. Nobody’s connected it yet.
Next step: Review the complete asset history before writing off any repeat failure, since the pattern sometimes already sits in the data, unread.
Where a CMMS Fits Into Answering These Questions
Every question above depends on the same thing: seeing the full history, across time, across identical assets, across technicians, instead of relying on whoever remembers the last three calls. Platforms like eWorkOrders keep that history in one searchable place, so answering these questions takes minutes instead of a week of asking around.
- Complete failure history by asset, searchable by failure mode, not just by date
- Fleet-wide comparison across identical assets, so a pattern invisible on one unit shows up across ten
- Work orders that capture confirmed failure mode and root cause, not just the repair action
- Preventive maintenance intervals that can be checked against actual failure dates, not just followed blindly
- Reporting and KPIs that support downtime reduction and equipment reliability reviews across the full asset history
Frequently Asked Questions
How many repeat failures justify a root cause investigation?
There’s no universal number. A safety-critical or high-cost asset may warrant a review after the first failure, while a low-criticality asset might reasonably wait for a second or third. The threshold should match the asset’s criticality, not a fixed rule applied everywhere.
Is “bad luck” ever actually the right explanation?
Occasionally, yes. Some failures are genuinely random and don’t track against any variable a team can identify. The difference is that this conclusion should come after checking, not instead of checking.
What’s the fastest way to check for a pattern across repeat failures?
Pull the full failure history for the asset and for identical assets, then look for anything that repeats: a failure mode, a part, an operating condition, or a time period. A pattern usually shows up within the first few data points once they’re actually lined up side by side.