The switch passed every test. The integrator signed off. Six months later, the network failed during a restart. The test never simulated a restart.


Commissioning engineer with tablet showing FAT PASS and red fault LED in background

The Problem Acceptance Tests Miss

Factory acceptance tests pass. Site acceptance tests pass. The integrator signs off. Everyone agrees the system is ready.

But FATs test ideal conditions. They don't test restarts. They don't simulate power loss. They don't push traffic to the breaking point. The switch that recovers in 50ms during a test may take 500ms when every device on the network reconnects simultaneously.

A 450ms delay may seem insignificant. To a PLC scan cycle, protection relay, or motion control system, it can be the difference between continuity and fault. The gap between test and reality is where failures live.

Testing functionality is not the same as testing behaviour.

What the Test Lab Cannot Replicate

A commissioning environment contains known devices, controlled traffic and stable power conditions. Operations introduces variable loads, unexpected events, maintenance activities and human intervention. The behaviour of a network under those conditions is rarely measured during acceptance testing.

The test that passed in commissioning does not reflect what happens during an actual event. The test lab is controlled. The field is not. The field introduces restart storms, traffic peaks, environmental stress, and legacy device interactions that no single test case can simulate.

Diagram showing checklist of tests passed and field failure

The test passed. The field failed.

The Tests That Are Missing

A FAT verifies that equipment functions as specified. It does not verify behaviour under stress. It does not measure what happens during a restart. It does not test failover under load. It does not simulate a corrupted frame.

Recovery behaviour is rarely tested. Failover timing is rarely measured. Interaction with legacy devices is rarely validated. Environmental stress is rarely applied. Cumulative load is rarely simulated. Each of these factors can cause a system to fail in ways that no single test will reveal.

The tests that are missing are the tests that matter most. The system passes acceptance. The system fails in the field. The gap is not in the equipment. It is in the test procedure.

Where Failures Begin

Failures begin with the assumption that the test is sufficient. The test passes. The assumption is that the system is ready. The assumption is false.

The failure begins during a restart. A switch reconverges. A drive faults. A SCADA master times out. The test never simulated a restart. The test never simulated a failover under load. The test never introduced an edge case.

The failure is not the restart. The failure is the assumption that the test was sufficient. The engineer who designed the test did not anticipate the field conditions. The test was incomplete.

How Leading Organisations Approach Testing

Successful organisations test behaviour, not just functionality. They write test procedures that include restarts, failover, and edge cases. They run the tests themselves. They document the behaviour they observe.

They test recovery behaviour. They measure network latency during a scheduled restart. They compare it to steady-state latency. The difference is what their control system experiences but their monitoring never shows. They test failover under load. They pull the primary link while the network is busy. They see if the secondary takes over seamlessly.

They also test environmental stress. They simulate temperature cycling. They introduce power sags. They test the system under conditions that are likely to occur in the field.

Testing behaviour is the difference between a system that works and a system that survives.

Technology In Practice

Many industrial recovery tests are performed on lightly loaded networks where recovery appears almost instantaneous. In operation, every controller, HMI, RTU and historian may reconnect simultaneously. Recovery behaviour under load is often very different to recovery behaviour in a test environment.

Westermo designs switches specifically to maintain predictable recovery characteristics during these conditions, allowing engineers to test the behaviour they are likely to experience in the field rather than the behaviour they hope to see in a laboratory.

The technology exists. The question is whether the test procedure includes it.

What the Test Should Include

A complete test procedure includes three elements. The restart test simulates a complete power loss and measures recovery time. The failover test pulls the primary link under load and measures the secondary takeover time. The edge case test introduces corrupted frames, simulates a noisy link, and observes how the system responds.

Each test measures behaviour, not just functionality. Each test reveals conditions that the FAT would miss. Each test reduces the gap between test and reality.

The Test That Matters

The test that matters is not the FAT. It is the field test. The FAT verifies functionality. The field test verifies behaviour. The FAT is necessary. The field test is essential.

Next time, test for behaviour, not just compliance. Test restarts. Test failover under load. Test edge cases. Document the behaviour you observe.

The engineer who tests for behaviour finds the root cause. The engineer who tests for functionality will chase the same fault again.


Continue Exploring


TEST FOR BEHAVIOUR, NOT JUST COMPLIANCE

Throughput advises on testing procedures that catch what FATs miss—restarts, failover, edge cases, and environmental stress.

What test did your last upgrade pass – but should have failed?

Fill out the online form.

You May Also Be Interested In ...