Why investigations stall
Commissioning faults often arrive as conclusions: “the controller is faulty,” “the motor is underpowered,” or “the software has changed.” These statements may be plausible, but they are not yet evidence. Time pressure encourages parts replacement and repeated tests before the operating state, sequence and recent changes are understood.
A structured method protects the equipment and the investigation. It defines the symptom precisely, preserves the initial condition and changes one thing at a time. It also makes the reasoning visible so that vendors, engineering, construction and operations can contribute without creating parallel uncontrolled experiments.
Protect the quality of the evidence
Evidence can be degraded quickly during a fault response. Resetting alarms, cycling power, changing settings or disconnecting a suspect component may remove the condition that would have identified the cause. Before intervention, capture the operating mode, loads, indications, alarm order, environmental conditions and relevant configuration. Where it is safe, preserve logs and photographs of indications.
The investigation team should distinguish correlation from cause. A warning that appears immediately before a trip may be a consequence of the same initiating fault rather than the cause of it. A replacement component that temporarily restores operation may point to a connection disturbed during replacement rather than prove the original component failed.
Useful review questions include: What is directly observed? What is inferred? When was the system last known to work? What changed? Which measurement could disprove the leading hypothesis? Can the test be performed without bypassing protection or risking further damage? This discipline reduces confirmation bias and makes escalation to vendors or engineering more productive.
A practical method
- Make the condition safe. Stop if continued operation could damage equipment or defeat a protection function.
- Describe the symptom. Record what happened, when, under what load or sequence, and what was expected instead.
- Verify the evidence. Check instruments, timestamps, alarms, trends, protection targets and witness accounts.
- Review changes. Identify modifications, settings, wiring changes, maintenance and environmental differences since the last known good state.
- Divide the system. Use functional boundaries and simple tests to narrow the fault without masking it.
- Test a hypothesis. State the expected observation before changing a setting, connection or component.
- Prove recovery. Repeat the relevant operating sequence, verify protections and document the final configuration.
Using the investigation record
Keep a simple chronology of observations, hypotheses, actions and results. Record settings before and after any authorised change and identify the person directing the test. When a hypothesis is rejected, retain that result; it prevents repeated work and helps the next specialist understand why the investigation moved in a different direction.
Short investigation checklist
- Safe state confirmed
- Symptom written factually
- Expected behaviour defined
- Test instruments verified
- Alarm and event timing captured
- Recent changes reviewed
- Drawings and settings current
- Hypothesis recorded
- One controlled change at a time
- Original state recoverable
- Recovery test agreed
- Final evidence retained
Common failure patterns
Common traps are replacing the suspected component before checking its inputs, adjusting several settings together, ignoring measurement uncertainty, and declaring success after one unloaded run. Another is allowing the investigation log to record actions but not the reason, expected result or actual observation.
Boundary: Troubleshooting must remain within authorised competence, isolation rules and approved test procedures. Protection bypasses, energised work and changes to safety functions require project-specific control.
See the fuller troubleshooting framework, the electrical reference, and safe handover and start-up preparation.