Skip to content

Implementation Guide

5 Whys.

5 whys is the entry-level root cause method: starting from one verified problem, ask why it happened, verify the answer at the process, and repeat until you reach a cause the management system can fix. Five is a heuristic, not a rule. If your chain ends at operator error or a broken part, you stopped too early: the real root cause is usually a missing standard, an unowned check, or a skipped training.

What 5 whys is

5 whys is the simplest structured root cause method in the Lean toolbox. Starting from a precisely described problem, you ask why it happened, verify the answer at the process, then ask why again about that answer, and repeat until the chain reaches a cause the management system can actually fix. It takes a small team under an hour, needs no software and no statistics, and it is where nearly every plant's problem-solving capability either starts or quietly fails to.

Five is a heuristic, not a rule. Toyota practitioners noticed that chains typically need about five iterations to travel from a symptom to a weakness in the management system, but some need three and some need seven. What matters is the stopping rule: a finished chain ends at something like a missing standard, an unowned check or a skipped training, because that is the level where a countermeasure prevents recurrence. A chain that stops at operator error or a broken part has described the event, not explained it.

Scope note: 5 whys is the entry level of the problem-solving ladder, sized for deviations a team can analyze the same day. Chronic problems worth weeks of study belong in an A3, problems that must formally prevent recurrence in an RCCA cycle, and customer-facing escapes in an 8D report. This page covers the everyday tool; those guides cover the heavier machinery, and all of them use 5 whys chains inside their analysis steps.

Where 5 whys sits in the transformation roadmap

On the TeamGuru deployment roadmap, 5 whys is the entry door of the structured problem solving practice in the Run stage. It is fed by daily management: the tier meeting cascade surfaces deviations every morning and assigns the analysis, and the countermeasures the chains produce land in standard work, where leader routines keep them checked. Without that daily stream of well-described deviations, 5 whys has nothing real to work on; without 5 whys, the daily system identifies the same problems forever.

When 5 whys is not enough

5 whys earns its popularity by being fast, and it gets misused for the same reason. It is built for problems with one dominant causal path that a small team can verify the same day. Move up the ladder when:

  • Several causes interact. When machine condition, material variation and method could all plausibly contribute, a single chain will braid them into a story. Map the candidates with a fishbone diagram first, then run 5 whys on the strongest one or two.
  • The problem is chronic and has survived previous fixes. If the same deviation has been countermeasured twice and returned, the quick tool has had its chance. Move to the structured rigor of an A3 or a full RCCA with real data collection.
  • A customer received the defect. Escapes need two chains, why it occurred and why it was not detected, plus containment and formal closure. That is an 8D report, not a huddle at the machine.
  • Someone was hurt, or nearly was. Serious safety events need a formal investigation to its own standard. A 5 whys can feed that investigation, but it must never be the whole analysis.

How to run one properly

A good 5 whys takes 30 to 45 minutes, three or four people, and happens where the problem happened. Nothing in the sequence is difficult; everything in it is easy to skip, which is why most bad chains are bad in the same few ways.

Start within 24 hours, at the process

Evidence evaporates fast: parts get scrapped, logs roll over, memories smooth themselves out. The usual trigger is the morning tier meeting, which assigns the 5 whys and expects the result the next day. The meeting never runs the analysis itself; the analysis happens at the machine.

Bring the people who were there

The operator who saw it, the team leader, and the maintenance tech if equipment is involved. Three or four people for 30 to 45 minutes. Anyone who was not there and does not own a likely countermeasure is an audience, and audiences turn analysis into performance.

Write a problem statement with numbers

What, where, when, how much. Machine M-14 stopped at 10:20, 34 minutes lost, 17 units behind takt beats machine trouble on line 2. A vague problem statement guarantees a vague chain, because nobody can verify a why against a symptom that was never pinned down.

Walk the chain, one verified why at a time

Ask why, then go look before you write the answer down: open the machine, pull the log, watch the step. The five rules below are the discipline. Expect to leave the huddle at least once to check something; a 5 whys that never moves from where it started is probably guessing.

Countermeasure the system cause

The countermeasure targets the last verified why, has an owner and a date, and changes something checkable: a standard, a route, a device, a check. If the proposed countermeasure is retraining and nothing else, treat that as a flag that the chain stopped at a person instead of the system around them.

Schedule the effectiveness check

Two to six weeks later, someone named looks at data, not at opinions: did the deviation recur, did the measurement move. A closed 5 whys without a dated check is a hope with paperwork. If the check fails, the chain reopens without shame; a wrong first chain is normal.

The five rules of the method

These rules are the whole difference between a root cause analysis and a conversation that ends in a plausible sentence. Print them next to the form.

  1. Verify each answer before asking the next why

    Evidence means something you can hold, read or observe at the process: the blown fuse, the drive log, the empty reservoir. An answer accepted by nodding is a guess, and every why built on it inherits the guess.

  2. Follow one causal path at a time

    If a why has two verified answers, branch the chain and run each path separately. Braiding parallel causes into one line produces a story that reads well and fixes nothing.

  3. Ask why the process allowed it, not who did it

    Names end chains. Process questions keep them moving: not why did he skip the check, but what made skipping the check possible and easy.

  4. Stop when you reach a system cause you can fix

    A missing standard, an unowned check, a skipped training, an impossible layout. That is the level where a countermeasure prevents recurrence. Stopping earlier fixes a symptom; digging further usually produces philosophy.

  5. Check the chain backward with therefore

    Read from root cause up to the problem, inserting therefore between every step. Each sentence must hold on its own. Any therefore you would not defend out loud marks the weak link.

Notice what the sequence does not contain: a projector, a template debate, or a vote. A root cause is not the answer most people in the room prefer. It is the answer the evidence survives.

Worked example: the same stop, two chains

The event is illustrative but realistic, set in the same 450-person components manufacturer used across the transformation roadmap: 456 units per day, two shifts, a takt of 118 seconds. At 10:20, machine M-14 stops mid-shift; by the time it runs again the line has lost 34 minutes, roughly 17 units against takt. Two teams analyze the same stop. One writes a capital request. The other fixes the management system.

Step The bad chain The good chain
Why 1 The fuse blew. The fuse blew on overload. Verified: the blown fuse is in the electrician's hand, and the drive log shows current climbing for 40 minutes before the stop.
Why 2 The circuit was overloaded. The spindle bearing had seized, overloading the motor. Verified: the bearing does not turn freely by hand and shows heat discoloration.
Why 3 The machine is old. The bearing was not getting enough lubrication. Verified: the reservoir was nearly empty; the last entry on the lubrication sheet is three weeks old.
Why 4 No fourth why was asked. Old is an attribute, not a cause, and the chain died there. A lubrication schedule exists, but it is not part of anyone's standard work. Verified: the schedule hangs in the maintenance office. No route, no named owner, no check that it happens.
Countermeasure Request a new machine. The capital request waits months, the line keeps stopping, and every sister machine keeps running on the same neglected schedule. Lubrication added to the operator's standard work with a visual level check; the weekly leader standard work audit verifies it. Effectiveness check in four weeks: bearing temperatures and drive current on M-14 and its two sister machines.

Both chains start from the same verified problem statement: M-14 stopped at 10:20, 34 minutes lost, 17 units behind takt. The good chain reached a fixable system cause in four whys; five is a heuristic, not a quota.

Where the bad chain went wrong

  • Why 2 was never verified. Overloaded was inferred from the fuse rating; nobody opened the machine, so the seized bearing stayed invisible and the chain drifted toward the machine's age.
  • Why 3 answers a different question. Old describes the machine; it does not explain why this bearing seized this week. When an answer is an attribute instead of an event, the chain has left the causal path.
  • The countermeasure buys a new machine and keeps the system that starved this one. The replacement inherits the same unowned lubrication schedule and fails the same way, later and more expensively.

Now run the backward test on the good chain: the schedule sits outside standard work, therefore lubrication was missed, therefore the bearing ran dry and seized, therefore the motor overloaded and blew the fuse, therefore the machine stopped. Every therefore holds. Try the same on the bad chain and it fails immediately: the machine is old, therefore the circuit overloaded is not a sentence anyone would defend out loud.

Notice where the good countermeasure lands: in two documents that already have owners. The operator's standardized work gains a lubrication step with a visual level check, and the weekly leader standard work audit verifies it keeps happening. Nothing new was invented. The system absorbed the lesson, which is exactly what the fifth why is for.

The human-error trap

Somewhere around the third why, most chains meet a person: the operator skipped the check, maintenance missed the route, the planner keyed the wrong number. This is the most important fork in the method. Blaming the person is always available and never sufficient. Available, because a human touches every process in the plant, so every chain can be ended with a name if you want it to end. Insufficient, because the same weakness will recruit the next human: a check that exists only in someone's memory will be skipped again, at the worst possible moment, by the most conscientious person on the crew.

The question that keeps the chain alive is: what made the error possible, and what made it easy? A lubrication route with no owner and no check does not depend on anyone being careless to fail; it only depends on time. Asking why the process allowed the error is not politeness toward the operator. It is precision about where recurrence actually lives.

There is also a system consequence. A plant where 5 whys chains end in names becomes a plant where very few problems get reported, and the daily management system quietly loses its raw material. Punishing the third why buys silence at the first one.

Common failure modes

5 whys fails quietly, and almost always in one of four ways. Each has a countermeasure that works better as a standing rule than as a one-time correction.

Meeting-room archaeology

The analysis happens three days later, in a conference room, from memory. The parts are gone, the logs have rolled, and the chain is built from what people can defend rather than what happened. The countermeasure is a standing rule, not a reminder: within 24 hours, at the process, or not at all.

The predecided answer

Someone wants a new machine, a bigger buffer or a pet gadget, and the chain is built backward to arrive there. The tell is a chain where no step carries evidence, because evidence would resist. Cure: evidence per why, and someone from outside the area reads the chain backward before the countermeasure is approved.

Five parallel whys in one chain

Each why gets answered with two or three causes, and the chain becomes a tree drawn as a line. The analysis feels thorough and proves nothing. Branch explicitly when a why has two verified answers, or map the candidates with a fishbone first and drill the strongest path.

Retraining by default

Every chain ends with retrain the operator, because it is cheap to write and nobody has to change a process. This is the human-error trap wearing a countermeasure. If training keeps being the answer, the causes were never verified: ask what the training would compete against and fix that instead.

What happens next

One verified chain fixes one problem. The system effect comes from what happens to the chains afterward. Countermeasures change standards or they were theater. Problems that come back after a countermeasure graduate to a full RCCA with containment, verified corrective action and a formal effectiveness check. And when the fourth or fifth why keeps finding the same missing routine across different chains, that is not a coincidence. That is the next improvement theme, announced by the data.

The bookkeeping is where good intentions leak: chains on whiteboards get erased, photos get buried in chat threads, countermeasures lose their owners, and the four-week check never happens because nothing was scheduled. This is the point where TeamGuru fits the practice: the root cause analysis use case keeps the problem statement, the chain, its evidence and the countermeasures with their owners in one record, with the effectiveness check scheduled instead of remembered.

On the roadmap, structured problem solving runs alongside leader standard work and feeds kaizen: the same discipline of verify the cause, change the standard, check that it held, pointed at opportunities instead of deviations. A team that can run an honest 5 whys has already learned the hard half of both.

5 Whys implementation diagram (TeamGuru guide)
Take this with you: free to reuse in internal training and workshops. Download PNG

Frequently asked questions

Why five whys and not three or seven?
Five is a heuristic from Toyota practice: most chains need roughly that many iterations to travel from a symptom to a cause in the management system. Some chains reach a fixable system cause in three whys, some need seven. The stopping rule matters more than the count: stop when you reach a cause you can fix with a standard, an owned check or a process change, not when you hit five.
Who invented the 5 whys method?
It comes from Toyota. The idea is attributed to Sakichi Toyoda and was developed into an everyday problem-solving habit inside the Toyota Production System, where Taiichi Ohno taught it with the seized-machine example most trainings still quote. It spread worldwide with Lean manufacturing from the 1980s onward.
Can a problem have more than one root cause?
Yes, and pretending otherwise is how chains get braided into mush. When a why has two verified answers, branch the chain and follow each path separately to its own system cause and countermeasure. Customer escapes usually need two chains by design: why the defect occurred, and why it was not detected before it left.
What is the difference between 5 whys and a fishbone diagram?
A fishbone maps candidate causes broadly across categories like machine, method, material and people; 5 whys drills one verified path deep. They work best in sequence: when many factors could plausibly contribute, collect and rank candidates on the fishbone, then run 5 whys on the strongest one or two. A fishbone without the drill-down produces suspects, not causes.
When in the day should a 5 whys happen?
Within 24 hours of the event, at the process, while the parts, the logs and the memories still exist. The daily tier meeting is the usual trigger: it assigns the analysis and hears the result the next morning. The analysis itself happens at the machine with the people involved, never inside the meeting.
How should a 5 whys be documented?
One record with four things: the problem statement with numbers, the chain with its evidence per step, the countermeasure with an owner and a date, and the effectiveness check date and result. A photo of a whiteboard satisfies none of this for longer than a week. Keep the format simple enough that a team fills it in ten minutes; the discipline lives in the evidence, not the template.

Make the root cause stick

See how TeamGuru keeps every chain, its evidence and its countermeasures in one record with owners, dates and a scheduled effectiveness check.