Skip to content

Implementation Guide

TPM.

Total productive maintenance (TPM) is a system for eliminating equipment losses by putting daily machine care in the hands of the people who run the machines, backed by planned maintenance and engineering. When operators clean, inspect and lubricate their own equipment, deterioration gets caught early and breakdowns stop being weather and start being decisions. Start with autonomous maintenance on the constraint machine, not with eight pillars on a poster.

What TPM is, and what it is for

Total productive maintenance is a system for eliminating equipment losses: breakdowns, minor stops, speed losses, and defects caused by equipment condition. It works through three engines. Autonomous maintenance puts daily care, cleaning, inspection, lubrication and tightening, in the hands of the operators who run the machine. Planned maintenance gives the professional maintenance team a schedule driven by condition and criticality instead of by whoever shouts loudest. Focused improvement attacks the biggest recorded losses one at a time, cross-functionally.

The classic model wraps these in eight pillars, adding quality maintenance, early equipment management, education and training, safety and environment, and administrative TPM. Mention the pillars once, then put the poster away. Launching eight parallel programs with subcommittees is how TPM becomes a bureaucracy that never touches a machine. Autonomous maintenance plus planned maintenance plus focused improvement carries most of the value, 90 percent is a fair working number, and the remaining pillars only make sense once those three visibly run.

The reason operator ownership is the pillar that matters is mechanical, not cultural. Most breakdowns do not arrive; they develop. A bearing runs dry, a bolt loosens, swarf packs into a slide, and for weeks the machine announces it through noise, heat, vibration and leaks. The only people positioned to hear the announcement every day are the operators. When they clean the machine, they touch it; when they touch it, they find what is loosening and leaking; and when finding it triggers a fix, breakdowns stop being weather that happens to the plant and start being decisions the plant made or failed to make.

Where TPM sits in the transformation roadmap

TPM belongs to the equipment methods practice in the Improve stage of the deployment roadmap, alongside SMED, FMEA and poka-yoke. The placement is deliberate: TPM pays off where daily management already holds gains and produces honest numbers. The daily boards record downtime by reason every shift, and the availability losses inside the OEE capture name the machines and failure modes worth the effort. Start TPM before that data exists and machine selection is politics; start after, and it is arithmetic.

When TPM stalls before it starts

Most failed TPM launches were lost before the first deep clean, and the causes are management's to fix, not the operators'. Check these three before announcing anything:

  • A maintenance backlog so deep that operators find only broken promises. If the machine has ten known defects that maintenance has not fixed in a year, a tag system just documents the eleventh. Burn down the critical backlog on the pilot machine first, so the deep clean starts on a machine whose known faults are fixed.
  • No time budgeted for the routines. If the shift plan assumes every minute is production, the daily clean-inspect-lubricate routine is theft, and it will be skipped on the first busy day and every day after. Plan the minutes into the schedule and defend them.
  • No loss data. Without downtime recorded by reason there is no way to pick the right machine, no baseline, and no proof six months later. A few weeks of honest capture at the daily board is enough to start.

How to implement autonomous maintenance

Do not start with a program. Start with one machine, and make it the constraint. The constraint is where availability losses are throughput losses: an hour of breakdown there is an hour of plant output, which makes the payback visible and the case for the second machine easy. At the illustrative 450-person components plant used across this site (456 units per day, two shifts), that machine is the machining line, which loses 33 minutes a day to breakdowns and another 16 to minor stops out of 840 available minutes. Those two rows of the loss ledger are the pilot's target.

The method is the five steps of autonomous maintenance, run in order, with maintenance and the supervisor working alongside the operators. The durations are practitioner guidance for one machine, not a corporate schedule.

Step What happens Output Typical duration
1. Initial deep clean as inspection Operators, maintenance and the supervisor restore the machine to base condition together. Every abnormality found while cleaning, leaks, loose fasteners, worn guards, missing covers, gets a defect tag on the spot. Machine at base condition, plus a tag list of everything wrong with it 1 to 2 days
2. Eliminate contamination sources Fix what makes the machine dirty and hard to inspect: seal leaks at the source, add guards and covers, relocate lubrication points and gauges to positions a person can actually reach. Cleaning and inspection time cut sharply, often by half or more 4 to 8 weeks, alongside production
3. Provisional CIL standards The team writes cleaning, inspection and lubrication standards: what, how, how often, how long, with photos of the correct state. The routine must fit a stated time budget per shift. A provisional daily routine the crew helped write and can run in the time given 1 to 2 weeks
4. Train inspection skills Operators learn to inspect subsystem by subsystem: fasteners, lubrication, pneumatics, hydraulics, drives. What a failing bearing sounds like, what a frayed belt or an abnormal gauge reading looks like. Operators qualified to catch deterioration early; standards upgraded from provisional to final 2 to 3 months, one subsystem at a time
5. Operator-run routine with visual controls The daily routine runs on marked gauges, labeled sight glasses, match-marked fasteners and point-of-use photos. Abnormalities flow into defect tags, and the routine is audited weekly. A self-sustaining daily routine with completion and findings visible Ongoing; stable within 4 to 6 weeks

If the area already runs 5S, step 1 will feel familiar: it is the shine step done with inspection meaning, and the two practices reinforce each other. The CIL standard in step 3 is standard work for machine care, and it lives by the same rules: written with the people who execute it, fitted to a time budget, revised when reality disagrees with it.

Visual controls that make inspection fast

Inspection that requires judgment gets skipped on a bad day. Inspection that requires a glance survives. Mark the normal operating range on every gauge so abnormal is visible in a second, label minimum and maximum on sight glasses, match-mark critical fasteners so a loosened bolt shows itself, and put a photo of the correct state at the point of use, not in a binder. The test of a good visual standard: a new operator can complete the routine correctly, and notice what is wrong, without asking anyone.

Defect tags: the trust meter

The deep clean in step 1 surfaces everything the machine has been quietly accumulating; on a neglected machine, expect dozens of tags, commonly 50 to 150. What happens to those tags decides the fate of the whole effort. Run a two-color system: one color for defects operators can fix themselves, another for maintenance. Every tag records what was found, where, by whom and when. Triage weekly with maintenance, close the quick ones within days, and put the backlog trend and its median age on the area board.

The tag backlog is the trust meter, and it deserves that name. Operators are being asked to notice deterioration on management's promise that noticing leads to fixing. If tags sit open for months, the promise is broken in public, operators learn that noticing is decorative, and autonomous maintenance dies politely while the checklists stay green. Closing tags fast is what buys engagement for the second machine. The weekly audit of the routine and the tag review belong in the plant's layered audit program, so the check on the system survives busy weeks.

Measuring whether it works

TPM produces its own proof, provided the measures were running before the pilot started. Take the baseline from the capture the daily system already runs, then trend each machine against itself:

  • Availability losses from the OEE capture: breakdown minutes and minor stops per day, by reason. On the pilot machining line the illustrative baseline is 33 and 16 minutes; the breakdown row is the first target.
  • MTBF, the mean time between failures, trended per machine. The absolute number matters less than the direction.
  • Tags raised and tags closed per week, plus the median age of open tags. Tags raised falling toward zero is not success; it usually means people stopped looking.
  • Routine completion from the audited checklist, reported honestly. A missed routine is data about the time budget, not a disciplinary matter.

Review downtime daily at the board, Pareto the losses monthly, and let the Pareto pick the next focused improvement and the next machine. How the losses roll up into OEE is covered in its own guide, and the baseline discipline is the same as for every other number on the plant's KPI baseline: defined once, captured at the process, trended against itself.

Who owns what on the machine

Autonomous maintenance triggers two fears, usually unspoken. Maintenance hears that operators will take over their work; operators hear that maintenance work is being pushed onto them unpaid. Both are answered the same way: a written split of responsibilities per machine, agreed by all three parties before step 1 starts. The split below is a workable default for one machine. Adjust the lines to your equipment, but keep the principle that every task has exactly one owner.

Role Owns day to day Escalates when
Operator Daily clean, inspect and lubricate to the CIL standard; tighten what is match-marked; first response on minor stops; raise a defect tag for anything beyond the standard. Anything not fixable within minutes goes on a tag the same shift. A safety-relevant finding stops the machine and calls maintenance immediately.
Maintenance Planned overhauls and time-based replacements; predictive checks where they earn their cost; repairs beyond first response; defect tag triage and closure. A failure mode that returns after a correct repair goes to engineering with the failure history attached.
Engineering Eliminating recurring failure modes by design: component upgrades, contamination source removal, specification changes, maintainability improvements. Fixes that need capital or a production window beyond the area's authority go to the monthly review for a decision.

The escalation path is part of the split, not an afterthought. A tag raised the same shift, a stopped machine for safety findings, and a documented handover to engineering for repeat failures: when those three routes work, the roles stay clean. When they do not, everything quietly becomes maintenance's problem again within a quarter.

Common mistakes

What bad looks like

  • Autonomous maintenance introduced as unpaid extra work on top of unchanged production targets
  • Cleaning without inspection meaning: the machine shines and breaks down exactly as often
  • Planned maintenance windows that production refuses week after week, until the backlog forces a breakdown at the worst time
  • Eight pillars, a steering committee and a master plan before a single machine reaches base condition
  • Defect tags from the first deep clean still open three months later

What good looks like

  • Minutes for the daily routine planned into every shift, visibly
  • Every cleaning step doubles as an inspection with a known failure mode behind it
  • Maintenance windows protected like customer orders, because they are
  • One constraint machine through all five steps before any talk of rollout
  • Tag backlog trending down, reviewed weekly, with median open age on the board

What stable equipment enables next

The payoff of TPM is not shinier machines; it is a stream you can design with. Equipment that runs when the schedule needs it is what allows batches to shrink and pull systems to hold, because every supermarket size and every changeover plan silently assumes the machine will be available. That is why TPM and changeover reduction usually run as siblings on the same constraint: one attacks the 33 breakdown minutes, the other the 47 changeover minutes, and together they buy the availability that flow is built on.

The routines themselves need a home that survives shift changes and busy quarters. In TeamGuru, the daily CIL routines run as scheduled checklists with photo findings that become owned actions, and downtime, MTBF and the tag backlog live on the same daily boards as the rest of the plant's numbers, so the machine's health is reviewed in the same five minutes as safety and delivery.

On the roadmap, equipment methods run in parallel with the kaizen pipeline and feed it: every monthly loss Pareto hands the next focused improvement to a team, and when the constraint machine holds its gains, the same five steps move to the next machine the data points at.

TPM implementation diagram (TeamGuru guide)
Take this with you: free to reuse in internal training and workshops. Download PNG

Frequently asked questions

What does TPM stand for?
TPM stands for total productive maintenance. The approach was developed in Japan in the 1970s, building on American preventive maintenance practice, and is closely associated with Seiichi Nakajima and the Japan Institute of Plant Maintenance. The word total refers to participation: everyone from operator to plant manager has a defined role in equipment care.
What is the difference between TPM and preventive maintenance?
Preventive maintenance is scheduled service performed by the maintenance department. TPM is a wider system that adds daily operator care of the equipment, focused improvement of the biggest recorded losses, and engineering work to remove recurring failure modes. Preventive maintenance is one component inside TPM, not a synonym for it.
Where should a plant start with TPM?
Start with autonomous maintenance on one constraint machine, because availability losses there are throughput losses and the payback is visible. Clear the machine's critical maintenance backlog first, then run the five autonomous maintenance steps until the daily routine is stable and audited. Roll out to the next machine only after the first one works.
How much operator time does autonomous maintenance take per shift?
After contamination sources are removed, a typical daily clean-inspect-lubricate routine takes 10 to 15 minutes per shift. The early steps cost more: the initial deep clean takes a day or two, and a new routine runs slower for its first weeks. The time must be planned into the shift, because a routine that competes with the production target loses every time.
Do we need CMMS software before starting TPM?
No. Autonomous maintenance runs on visual standards, defect tags and a daily routine, none of which need software. A CMMS earns its place later, for planned maintenance scheduling, spare parts and failure history. Buying software first is a common way to delay the part that actually changes anything.
How long does it take for breakdowns to fall?
On a single pilot machine, restoring basic condition and running the daily routine usually shows a visible drop in breakdowns and minor stops within three to six months. Plant-wide, TPM is a multi-year effort. Track the trend of each machine against its own baseline rather than waiting for a dramatic step change.

Make machine care part of the daily system

See how TeamGuru runs autonomous maintenance routines, defect tags and downtime KPIs inside the same boards the plant already reviews.