Redundancy Is Not The Same As Resilience
Redundancy adds paths. It does not guarantee predictable recovery. Here is why redundancy is not resilience.
The network was commissioned in 2019. Every test passed. The integrator signed off. Three months later, production stopped. Nothing was installed incorrectly. The operation simply behaved differently.
Most industrial projects measure success by whether systems eventually start. Few ask whether the network was genuinely ready before the first energisation sequence began.
Commissioning is the final act of a project. It is the moment when equipment is energised, tests are run, and signatures are collected. It is the moment when the installation is declared complete.
But commissioning is not operational readiness. Commissioning proves the installation works. Operational readiness proves the operation will remain predictable after handover. Those are completely different objectives.
Most industrial projects treat commissioning as the moment operational readiness begins. The truth is exactly the opposite. Commissioning simply exposes whether the design, installation, and configuration were already ready.
Factory acceptance tests pass. Site acceptance tests pass. The integrator signs off. Everyone agrees the system is ready. Three months later, production stops. Nothing was installed incorrectly. Nothing was configured incorrectly. The operation simply behaved differently.
The difference is not in the equipment. It is in the definition of success. FAT and SAT are necessary. They are not sufficient. They verify that equipment functions as specified. They do not verify behaviour under the conditions that occur after handover. The gap between commissioning and operations is where failures hide.
Industrial projects routinely omit the tests that matter most. Restart testing. Failover under load. Recovery sequencing. Environmental stress. Each omission is a future failure mode.
What if the test procedure included restart testing? A complete power loss. Every device reconnects simultaneously. The network reconverges. The control system restores. What if the test procedure measured recovery time under load, not idle? What if it simulated a communication loss between two critical devices?
These tests rarely happen. The FAT schedule is tight. The budget is fixed. The assumption is that if everything works individually, it will work together. The assumption is often wrong.
"The restart that was never simulated becomes the failure that no one can explain."
Engineers often discover that the network behaviour during a restart is fundamentally different from behaviour during a single-link failure. The network that recovered in 50ms during commissioning takes 500ms when the entire system powers back on. The difference is not a fault. It is a design assumption that was never tested.
Commissioning documented the start. Operations documented the failures.
The operation that exists after handover is not the operation that was commissioned. Maintenance begins. Software updates are applied. Operators change. The network grows. Assets age.
Modern industrial networks rarely fail because of a single switch. They fail because the environment changed. Summer temperatures exceed design limits. Power quality degrades. Dust accumulates. Connectors corrode. The equipment that worked perfectly during commissioning begins to behave differently.
The specification that was sufficient during testing becomes insufficient in operations. The system that passed acceptance is not the system that operates three years later. Organisations that treat commissioning as the end of readiness discover this pattern repeatedly.
Airports don't open because construction finished. They open because the operation has been rehearsed. Rail should think the same. Utilities should think the same. Mining should think the same.
Operational readiness is not a checklist. It is an engineering discipline. It requires understanding how the system will behave under stress. It requires testing recovery procedures. It requires documenting failure modes. It requires validating that the operation remains stable after handover.
The technology exists. The architecture supports readiness. It does not create it.
Industrial networking platforms such as Westermo are designed to support deterministic behaviour under stress. Forensic visibility platforms such as Micromedia ALERT provide the correlation layer that reveals hidden failure modes. Secure remote access solutions such as Secomea enable controlled maintenance without introducing exposure.
Partners provide tools. Engineers provide discipline. Operational readiness is designed, tested, and validated. It cannot be purchased.
Commissioning proves today's installation. Operational readiness proves tomorrow's operation.
The integrator hands over the project. The acceptance certificate is signed. Three months later, production stops. Nothing was installed incorrectly. Nothing was configured incorrectly. The operation simply behaved differently.
The pattern is consistent. The equipment changes. The cause does not. Engineers who recognise the pattern stop asking "What failed?" They start asking "What was never tested?"
Throughput advises on commissioning, operational readiness, and lifecycle planning – aligning testing with real-world conditions and proven industrial technologies.
Does your commissioning plan test what happens after handover?
Redundancy adds paths. It does not guarantee predictable recovery. Here is why redundancy is not resilience.
Your FAT passed. Your site failed. The test never simulated a restart. Here is what commissioning misses.
Industrial failures rarely have one source. The evidence exists. It is simply scattered across different systems.