The core sequence
The framework follows twelve connected steps. The order may be adapted to the system and risk, but the investigation should retain clear boundaries, evidence, ownership and restoration control.
- Define the reported symptom.
- Verify the symptom independently.
- Establish system boundaries.
- Identify what has changed.
- Separate cause from consequence.
- Review energy, signal, control, process and mechanical paths.
- Form and rank credible hypotheses.
- Test the simplest credible causes first.
- Record evidence and rejected hypotheses.
- Control temporary interventions.
- Confirm restoration under realistic conditions.
- Identify root cause and prevent recurrence.
1. Define the reported symptom
Begin with the condition that was reported, using the language of the person who observed it. Do not immediately translate the report into a technical diagnosis.
“Pump failed,” “breaker tripped” or “control system fault” may describe an operational consequence rather than the initiating failure. Record what was seen, heard, measured or lost before deciding what it means.
Questions to establish:
- What function was expected?
- What actually happened?
- When did it happen?
- Was the condition continuous, intermittent or repeatable?
- What alarms, trips, indications or physical effects were observed?
- Who witnessed the event?
- What operating condition existed at the time?
A clearly defined symptom gives the investigation a stable starting point. A presumed cause does not.
2. Verify the symptom independently
Where it is safe and authorised, confirm the reported behaviour using independent evidence. This may include control-system history, alarm logs, trends, instrument readings, physical inspection, test records or a controlled attempt to reproduce the condition.
Verification should answer three questions:
- Does the reported condition exist?
- Under what conditions does it occur?
- Which parts of the original report are confirmed, uncertain or incorrect?
Do not dismiss witness information because it is incomplete. Witness accounts often identify timing, sequence or unusual conditions that are absent from automated records. Treat them as evidence to be checked rather than conclusions to be accepted or rejected without review.
3. Establish system boundaries and safe conditions
Define what is inside the investigation boundary and what is outside it. A system fault may involve power, controls, mechanical equipment, process conditions, utilities, communications or interfaces with another system.
The investigation boundary should identify:
- the affected function;
- supplying and receiving systems;
- energy sources;
- control and protection interfaces;
- isolations and permits;
- operating authority;
- equipment that must remain available;
- conditions that require the work to stop.
Before testing begins, confirm how the system will be left safe if the investigation is interrupted or produces an unexpected result.
4. Build an evidence timeline
Create a factual timeline before developing a detailed cause theory. Combine control-system records, alarms, maintenance history, photographs, test results, operating logs and witness accounts in chronological order.
For each item, distinguish between:
- observed fact;
- recorded value;
- witness recollection;
- technical interpretation;
- unverified assumption.
Preserve original timestamps and source information. Do not rewrite the sequence to make it fit a preferred explanation.
Evidence record fields:
- Time
- System condition
- Observation or event
- Evidence source
- Confidence level
- Follow-up required
A good timeline often reveals that the apparent failure began earlier than the final alarm or loss of function.
5. Identify what changed
Review recent changes to the equipment, system, environment and operating method. Many faults appear after maintenance, construction, software work, setting changes, isolation, testing or a change in operating mode.
Check for changes involving:
- maintenance or component replacement;
- wiring, tubing or mechanical reconnection;
- software, logic or configuration;
- alarm, trip or protection settings;
- valve, damper or selector position;
- temporary equipment or bypasses;
- construction completion or system energisation;
- utilities or upstream services;
- environmental conditions;
- operating sequence or loading;
- personnel, procedure or shift handover.
The most recent change is not automatically the root cause, but it is a disciplined place to begin checking.
6. Separate cause from consequence
Alarms, trips, low readings and loss of function may be protective responses or downstream effects. Treating the final consequence as the initiating cause can send the investigation in the wrong direction.
Ask:
- What happened first?
- Which indications appeared only after the system had already degraded?
- Which protection operated as designed?
- Which condition could explain the full sequence?
- Which observations cannot be explained by the current theory?
A credible cause should explain the timeline and the important evidence, not merely one visible alarm.
7. Trace the required paths
Review every path required for the intended function. Different systems use different combinations, but most investigations should consider the following.
Energy path
Confirm that the required electrical, hydraulic, pneumatic, mechanical, thermal or stored energy is available at the correct condition and reaches the intended equipment.
Command and signal path
Trace the command from its origin through selectors, communications, input and output devices, relays, controllers and field equipment.
Control and protection path
Check permissives, interlocks, trips, inhibits, modes, setpoints and protection logic without assuming that an active trip is the original fault.
Mechanical and process path
Confirm alignment, freedom of movement, lubrication, flow paths, pressure conditions, valve positions, loading and the physical condition required for operation.
The investigation should follow the function from request to final response rather than examining individual components in isolation.
8. Form and rank credible hypotheses
Develop a short list of causes that are technically possible and consistent with the evidence. Avoid creating a long unranked list that gives every possibility equal weight.
A useful hypothesis should:
- explain the verified symptom;
- fit the event timeline;
- be consistent with system design and operating conditions;
- identify evidence that would support or reject it;
- be testable without uncontrolled risk.
Rank hypotheses using:
- likelihood;
- consequence;
- ease and safety of testing;
- strength of existing evidence;
- ability of the proposed test to distinguish one cause from another.
The purpose of a hypothesis is to guide evidence collection. It is not a conclusion to defend.
9. Test the simplest credible causes first
Start with checks that are safe, reversible and capable of clearly confirming or rejecting a credible cause. Simple does not mean superficial; it means avoiding unnecessary intervention before basic conditions have been established.
Before each test, define:
- the hypothesis being tested;
- the expected result if it is correct;
- the expected result if it is incorrect;
- the system state and isolation requirements;
- the measuring method;
- the acceptance criteria;
- the authorised person;
- the restoration method.
Do not change several settings, components or conditions simultaneously. Multiple uncontrolled changes may restore the system while destroying the evidence needed to understand the failure.
10. Record evidence and rejected hypotheses
Record what was tested, under which conditions, by whom and with what result. Include evidence that rejects a hypothesis as well as evidence that supports one.
A rejected cause is useful progress when the test was valid and the result is recorded. Without that record, teams may repeat the same checks or revive an explanation that has already been disproved.
The investigation record should distinguish between:
- confirmed facts;
- open questions;
- rejected hypotheses;
- temporary assumptions;
- actions completed;
- actions still required.
11. Control temporary interventions
Temporary jumpers, bypasses, overrides, substituted signals, altered settings and provisional repairs may sometimes support fault isolation or controlled recovery. They must not become informal or invisible changes to the system.
Every authorised temporary intervention should define:
- purpose;
- owner;
- approval;
- risk controls;
- identification;
- start time;
- operating limitation;
- monitoring requirement;
- expiry or removal condition;
- final restoration evidence.
Temporary recovery must not conceal an unresolved hazard or allow a system to be accepted as complete when it remains dependent on an uncontrolled modification.
12. Confirm restoration under realistic conditions
Clearing an alarm or achieving one successful start does not necessarily prove that the system is restored.
Verification should reproduce the operating conditions that matter, which may include:
- repeated start and stop cycles;
- normal operating load;
- expected process demand;
- automatic and manual modes;
- control and protection functions;
- upstream and downstream interfaces;
- standby changeover;
- communication with connected systems;
- observation over an appropriate period.
Record the final configuration, readings, settings, test conditions, witnesses and acceptance result.
Restoration is confirmed when the intended function is stable, repeatable and supported by evidence.
Root cause and prevention
The failed component or incorrect setting may be the direct cause, but a complete investigation should also consider why the condition was introduced, why it was not detected and why existing controls did not prevent the event.
Direct cause
The immediate technical condition that produced the failure.
Contributing conditions
Factors that increased the likelihood or consequence of the failure.
Latent weakness
A weakness in design, procedure, maintenance, supervision, documentation, competence or change control that allowed the condition to exist.
Prevention actions may include:
- permanent repair or replacement;
- design modification;
- setting or logic control;
- inspection or maintenance changes;
- procedure revision;
- commissioning or reinstatement checklist changes;
- training or competence action;
- spare-parts improvement;
- alarm or monitoring improvement;
- documentation correction;
- wider fleet or asset review.
A root-cause statement should lead to specific prevention action. If it only restates the failure, the investigation is incomplete.
Worked example: equipment fails after maintenance
The following simplified example illustrates the method without replacing equipment-specific procedures or competent technical investigation.
Reported symptom
A motor-driven auxiliary pump does not start after planned maintenance.
Verified evidence
The start request is present, but the equipment remains unavailable. No mechanical damage is visible and the supplying system is healthy.
Boundary and change review
The pump, starter controls, local selector, protection status and maintenance reinstatement are included in the investigation. The equipment had been isolated and locally operated during maintenance.
Initial hypotheses
- power supply unavailable;
- protection not reset;
- local selector not returned to the required operating mode;
- control connection not reinstated;
- mechanical seizure.
Controlled checks
The supplying system and protection state are confirmed first. The selector is then found in the maintenance position, which prevents the remote start command from progressing.
Restoration
Under the authorised reinstatement process, the selector is returned to the required operating mode. The pump is tested through repeated controlled starts, normal indication and connected-system response.
Root cause
The immediate cause is incomplete reinstatement of the operating selector. The contributing weakness is that the maintenance completion check did not explicitly require confirmation of the final selector position.
Prevention
The reinstatement checklist is revised to include operating-mode confirmation, independent verification and recording of the final equipment state.
The example demonstrates why restoring the equipment is not the end of the investigation. The control weakness that allowed the condition to pass through maintenance close-out must also be corrected.
Field troubleshooting checklist
- Define the reported symptom without assuming the cause.
- Verify the condition using independent evidence.
- Establish the investigation and isolation boundaries.
- Confirm authority, safe state and restoration requirements.
- Build a factual event timeline.
- Review maintenance, construction, software and operational changes.
- Separate initiating causes from alarms, trips and consequences.
- Trace energy, signal, control, process and mechanical paths.
- Form a short, ranked list of credible hypotheses.
- Define the expected result before each test.
- Make one controlled change at a time.
- Record supporting and rejecting evidence.
- Control every temporary intervention.
- Test restoration under realistic conditions.
- Record final configuration and acceptance evidence.
- Identify direct cause, contributing conditions and latent weaknesses.
- Assign prevention actions with owners and completion criteria.
- Communicate relevant learning to other systems, assets or projects.
Minimum investigation close-out
A completed troubleshooting record should provide enough information for another competent person to understand what happened, what was tested, what changed and why the system was accepted.
Minimum close-out evidence:
- verified symptom;
- system and investigation boundaries;
- event timeline;
- evidence reviewed;
- hypotheses considered;
- tests performed;
- rejected causes;
- temporary interventions;
- repair or correction;
- restoration test;
- final settings and configuration;
- root-cause statement;
- prevention actions;
- outstanding limitations;
- responsible review and acceptance.
This framework supports structured technical investigation. It does not replace equipment-manufacturer instructions, approved procedures, safe systems of work, competent technical judgement, site rules or applicable legal and regulatory requirements. Testing, isolation, bypassing and restoration must be authorised and controlled for the specific equipment and operating environment.