Problem Solving & Quality · Contain
Evidence Preservation
Keep the failed part, the sample, the logs and the settings exactly as found — tagged, photographed and stored before anyone cleans up.
- Time30 min
- FormatSolo
- StageContain
Evidence Preservation: what it is and why it works
Evidence preservation secures the physical and digital traces of a failure before cleanup, repair or restart destroys them. It covers failed parts and fragments, fluids and deposits, photographs of the scene, control-system trends, alarm lists, settings and maintenance records. Each item is tagged, located, photographed and stored in one place with a custody log, and any laboratory work is planned before destructive testing starts, so that one test does not destroy what another needed.
Many failure investigations are decided by evidence that existed for only a few hours: a fracture face before it corrodes, a deposit before it is washed off, a control-system buffer before it is overwritten. Without it, root-cause analysis falls back on opinion and the loudest theory wins. Preservation is quick and cheap compared with waiting for the failure to happen again. It is part of the first-hour protocol and provides the raw material for the event timeline, fault tree analysis and metallurgical or chemical lab work. It must always respect site safety rules: securing evidence never justifies entering an unsafe area, working on equipment that is not isolated, or bypassing a permit.
What you need
- Authority to delay cleanup and restart
- A camera or phone with timestamps, and a scale reference
- Tags, clean bags and containers, markers, gloves
- Access to the control-system historian, alarm logs and the CMMS
- A secure storage location and a custody log template
What you get
- Tagged and bagged parts, fragments and samples with notes on where each was found
- A photo set from wide views to close-ups
- Exported trends, alarm lists and records covering the hours before the event
- A custody log listing every item and who handled it
- A lab plan agreed before any destructive test
When to use it
When the broken bearing is already in the scrap bin and the PLC log has been overwritten.
How to do it, step by step
- Stop anyone from cleaning, repairing or restarting until the evidence is secured.
- Photograph the scene from wide to close-up, with a scale and a timestamp.
- Tag and bag failed parts, fragments, fluids and samples; note where each was found.
- Export control-system trends, alarm lists and maintenance records covering the hours before the event.
- Store everything in one labeled place with a custody log, and plan any lab analysis before destructive tests.
Worked example: Seized gearbox on a kiln drive
Illustrative scenario — figures are realistic but not from a real company.
A cement plant's rotary kiln trips when the main drive gearbox seizes. Every hour of downtime costs roughly $25,000 in lost clinker, and the maintenance crew's first instinct is to pull the gearbox and fit the spare as fast as possible.
- Once the drive was locked out and verified at zero energy, the shift manager held the teardown for 30 minutes so the reliability engineer could secure evidence.
- The engineer photographed the gearbox, coupling and lube system from wide to close-up with a ruler in frame, and drew oil samples from the sump and the filter housing into clean bottles.
- During removal, the crew bagged the filter element and the debris from the sump magnet, noting positions; the failed bearing was wrapped to protect its raceways, not washed.
- The control engineer exported 72 hours of vibration, bearing-temperature and lube-pressure trends, plus the lube-pump alarm history, before the short-term buffer rolled over.
- Everything went into a locked cage with a custody log. The lab was asked for oil analysis first, then sectioning of the bearing.
Result. The trends showed lube-oil pressure falling slowly over two days, and oil analysis found high water content. The bearing surfaces pointed to lubrication failure rather than a material defect. The fixes, repairing a leaking oil cooler and adding a low-pressure trip, would likely have been missed had the parts been washed and scrapped. The 30-minute hold added little to an 18-hour repair.
Common pitfalls and how to avoid them
- Washing or reassembling failed parts.Keep parts as found, protect fracture faces from handling and moisture, and let the specialist decide on any cleaning.
- Waiting to export data.Export historian trends and alarm lists immediately, and know in advance how long each system keeps high-resolution data.
- No custody log.Record each item, where it was found, who took it and where it is stored, so the evidence stays credible in a dispute or claim.
- Running destructive tests in the wrong order.Agree a lab plan before cutting or dissolving anything; non-destructive and surface examinations usually come first.
Frequently asked questions
What evidence should be kept after an equipment failure?
Keep the failed component and any fragments, lubricants, fluids and deposits, filters, photographs of the scene, control-system trends and alarm lists covering the hours or days before the event, operating settings, and maintenance and inspection records. When in doubt, keep it: storage costs little compared with repeating the failure to learn what the evidence would have shown.
Should you clean a failed part before sending it to the lab?
No. Cleaning can remove deposits, corrosion products and fracture features that show how the part failed. Protect it from further damage and moisture, record its condition and let the failure analyst decide what cleaning is appropriate and when. If contamination is a safety or environmental concern, follow site rules and tell the lab exactly what was done.
How long do control systems keep data?
It varies widely with the system and its configuration. Some controller buffers and event logs keep only hours or days of high-resolution data before overwriting it, while plant historians may keep compressed data for years. Check the retention settings of each system before an incident happens, and include the export steps in your incident protocol.
Origin
Failure analysis and incident investigation practice; no single author.
Used in these playbooks
Quality alert: the first 24 hours 1 day
One day to take control of a fresh incident: make safe, protect the customer, block every suspect lot, keep the evidence intact and pin down where the problem is — and is not.
Related methods
- Event Timeline ReconstructionRebuild minute by minute what happened before the failure from logs, historian data, shift notes and…
- First-Hour Incident ProtocolA fixed sequence for the first hour: make safe, stop the spread, preserve evidence, inform, appoint a leader.…
- Fault Tree AnalysisStart from the undesired top event and break it down with AND / OR gates into the combinations of failures…
More in “Contain”
Protect people, the environment and the customer while the real fix is found.