Two paths to the same device. Both reconnected at the same moment. Neither worked. Redundancy is not resilience.


Two fibre optic cables entering the same cabinet through the same opening – primary and secondary paths with a common vulnerability

The Decision Most Organisations Face

A substation loses communication with a protection relay. The network has dual paths. Both paths are live. Both paths fail during a power event. The redundancy was designed. It did not work.

Engineers assume that if you add a second path, the network will survive a failure. The assumption is often wrong. The second path is not independent. It shares common infrastructure. A common power supply. A common cable tray. A common switch. The failure that takes out the primary path often takes out the secondary path.

The decision is whether to add redundancy or build resilience. Most organisations choose redundancy. They add paths, switches, and power supplies. They assume the components will work when needed. The assumption is often false. Redundancy is necessary. It is not sufficient.

Redundancy adds paths. It does not guarantee predictable recovery. Independence matters more than path count. A second path that shares the same vulnerability as the first path is not redundancy. It is duplication.

What The Industry Usually Recommends

The standard industry advice is "add redundancy." Add a second path. Add a second switch. Add a second power supply. The advice is sound. It is also incomplete. Redundancy is necessary. It is not sufficient.

Redundancy assumes that the second path is independent. The assumption is often false. The second path shares the same cable tray. The second switch shares the same power supply. The second power supply shares the same panel. The redundancy is duplicated. It is not independent.

The assumption is built into the architecture. Designers specify dual power supplies, dual switches, and dual paths. The system passes acceptance. The operator trusts the architecture. The trust is misplaced.

Common-mode failure is the reason redundancy fails. The two paths share a common vulnerability. The vulnerability is exposed. The redundancy disappears. The common-mode failure can be a power supply, a cable tray, a cooling fan, or a software bug. It is almost always something that was overlooked during design.

Diagram showing two paths sharing a common power supply – a single vulnerability

Two paths. One vulnerability.

Where That Advice Falls Short

The advice to add redundancy falls short because it does not address independence. A voltage sag causes both power supplies to droop. Both switches reboot. Both paths fail simultaneously. The protection relay loses communication. The breaker does not trip.

The redundancy was designed. It did not work. The two paths shared a common power supply. The common-mode failure was not considered. The investigation traces the cause to the common-mode failure. The engineer had assumed the two paths were independent. The assumption was wrong.

Redundancy is about adding components. Resilience is about behaviour under stress. Redundancy asks: "What happens if this component fails?" Resilience asks: "What happens when the system is under stress?" The questions are different. The answers are different.

The Factors That Actually Matter

The factors that actually determine whether your network survives a failure are not the number of paths. They are the independence of the paths and the behaviour during recovery.

Independence – Are the paths truly independent? Do they share common power? Common infrastructure? Common vulnerability? The failure that takes out one path will take out the other. Design accordingly.

Recovery behaviour – What happens when a failure occurs? Does the network recover predictably? Or does it recover sometimes and fail other times? The behaviour is the difference between redundancy and resilience.

A redundant network has spare components. A resilient network recovers predictably. The first is about presence. The second is about performance. The first is about cost. The second is about behaviour. The difference is the difference between surviving a failure and failing gracefully.

What Changes The Outcome

The outcome changes when organisations move from adding components to designing behaviour. Resilience is not about avoiding failure. It is about recovering predictably.

The network that recovers in 50ms every time is resilient. The network that recovers in 50ms sometimes and 500ms other times is not. The resilience is not in the component. It is in the behaviour.

Resilient architecture is the difference between a network that survives and a network that fails gracefully. The first is luck. The second is design. Resilience is the ability to fail without failing. It is the ability to recover within defined time windows. It is the ability to maintain predictable behaviour under stress.

Resilience is not about eliminating failure. It is about controlling the consequences of failure. The network that recovers in 50ms is resilient. The network that recovers in 50ms sometimes and 500ms other times is not. The network that loses visibility during recovery is not resilient. The network that preserves forensic evidence during recovery is resilient.

Technology In Practice

Resilience depends on how a network behaves during failure and recovery, not simply the presence of redundant components. Westermo's WeOS architecture and FRNT recovery technologies are designed to minimise recovery times and maintain predictable communications in demanding operational environments.

Westermo's FRNT is tested under full load, not idle conditions. The recovery time specified is what you get when the network is busy – not when it is empty. That difference matters when your network is actually running.

Decision Checklist

Validate your redundancy. Are the paths truly independent? Do they share common power? Common infrastructure? Common vulnerability? The failure that takes out one path will take out the other.

Test recovery behaviour before the next restart tests it for you. Simulate simultaneous power loss across multiple cabinets, not isolated single-device failures. Real events do not occur one device at a time.

Plan for common-mode failures. The failure that takes out one path will take out the other. Design accordingly. Independence is the foundation of resilience.


Related Solutions


DESIGN FOR RESILIENCE, NOT JUST REDUNDANCY

Throughput advises on network architectures that recover predictably, preserve forensic evidence, and reveal root cause.

Where is your redundancy hiding a common-mode failure?

Fill out the online form.

You May Also Be Interested In ...