Validate the whole failover path
Review HA state together with routing, upstream, management and application dependencies.
Damocles Network Resilience & High Availability Review assesses HA peer state, routing and failover paths, upstream dependencies, management access, maintenance design and recovery controls for critical network-security infrastructure.
Resilience failures often sit outside the HA pair itself. Upstream routes, asymmetric return paths, state synchronisation, management access, DNS, authentication, virtualisation, power or carrier dependencies can leave a technically healthy standby device unable to carry production traffic.
Damocles reviews the end-to-end failover model and the operational steps required during maintenance or failure so the customer can identify single points of failure and validation gaps before an outage exposes them.
Review HA state together with routing, upstream, management and application dependencies.
Identify links, routes, services and management paths that undermine the intended resilient design.
Define failover checks, decision points and recovery procedures before planned upgrades or changes.
The review can focus on a firewall pair, a data-centre edge, remote-access service or a broader resilient network design.
Peer health, state synchronisation, control links, election/preference and failover behaviour.
Static and dynamic routes, path preference, convergence and return-path dependencies.
Carrier, switch, virtualisation, cloud, DNS or other external dependencies needed for failover.
Administrative access to both active and standby components during degraded operation.
Upgrade, reboot, change and failover sequence used during planned work.
Backups, configuration recovery, spares, documentation and operator decision paths.
Review topology, dependencies and high-availability controls, then validate agreed failure scenarios to expose single points of failure and unsafe recovery assumptions.
Redundancy, quorum and dependency assumptions.
Representative device, link and service-loss scenarios.
Traffic behaviour as paths fail and recover.
Impact on connections and security enforcement.
Power, DNS, identity, time and management dependencies.
Useful status and notifications during degradation.
Runbooks, authority, rollback and restoration evidence.
These are governed procedures from the service assurance profile—not generic marketing categories. The complete matrix below records applicability, access requirements and limitations for every coverage area.
Damocles maps peers, links, failure domains, traffic ownership and single points of failure from diagrams and configuration.
Damocles compares peer configuration, session and state synchronisation status and records material divergence.
Damocles reviews or observes approved interface and link failure scenarios, traffic movement, alerts and recovery.
Damocles reviews or observes route withdrawal and convergence against the selected scenario and success criteria.
Damocles traces upstream, downstream, identity, DNS, monitoring and provider dependencies for each resilient service path.
Damocles assesses expected stateful-session handling and records continuity only when an authorised scenario is observed.
Showing 6 representative areas. The full governed matrix contains 12 coverage areas.
References are shown according to the role they play in the engagement. Methodologies guide testing, taxonomies classify findings, severity methods support consistent scoring, and control mappings connect scoped evidence to broader assurance work.
Testing is structured by the governed service profile and the procedures expressly included in the authorised scope.
Findings are reported against the governed service coverage and customer impact. Separate classification references are shown only where they are part of the approved profile.
The engagement produces point-in-time technical evidence about the configured and observed security behaviour within the authorised Network resilience and high availability scope.
A resilience report with topology findings, scenario results, observed impact and prioritised recovery improvements.
What this means: technical evidence may support risk, compliance and audit activity where the mapping is applicable. It does not by itself certify the organisation, establish complete compliance or assess controls outside the authorised scope.
Damocles reviews topology, state synchronisation and dependencies and observes only customer-approved failover scenarios during a controlled window. Design evidence cannot prove operational recovery without an exercise, and testing stops at customer rollback criteria.
Directly assessed means the engagement tests the relevant behaviour. Supporting evidence means the result can contribute to a broader control assessment. Contextual references explain relevance without claiming that the control was tested.
Records state, traffic, alerts and recovery during an explicitly approved resilience scenario.
Reviews topology, synchronisation, convergence, session and VPN behaviour and controls the exercise.
Confirms service dependencies, acceptable interruption and observed business impact.
| Coverage area | What Damocles tests | Perspective and access | Evidence produced | References and limits |
|---|---|---|---|---|
| High-availability topology | Damocles maps peers, links, failure domains, traffic ownership and single points of failure from diagrams and configuration. | Network or security engineer The current HA topology or architecture diagram, peer configuration, interface and link relationships, traffic ownership or active/standby model, routing and upstream or downstream dependencies, expected failure domains, and relevant service ownership. | A topology or dependency map, peer and interface configuration or status, traffic-ownership state, identified shared dependency or single point of failure, and the configuration or status evidence supporting each conclusion. | Security and Privacy Controls for Information Systems and Organizations Release 5.2.0 Applicability: Applies whenever HA topology is included in design or configuration review, whether or not a live failover scenario is authorised. Limit: Topology and configuration review cannot by itself prove live failover behaviour or recovery time; a live result is claimed only where a scenario is separately authorised and observed, and excluded peers, links and dependent services remain outside the conclusion. |
| Peer state and synchronisation | Damocles compares peer configuration, session and state synchronisation status and records material divergence. | Failover observer Assessment requires topology, peer status, configurations, dependencies, scenario plan, monitoring and rollback authority, selected specifically for peer state and synchronisation. | Evidence records the configuration, status output, alert, traffic or session observation, timeline and recovery result relevant to peer state and synchronisation. | Security and Privacy Controls for Information Systems and Organizations Release 5.2.0 Applicability: Applies to Peer state and synchronisation when the design review or explicitly approved scenario includes this failure condition. Limit: The conclusion is limited to the sampled peer state and synchronisation; no production failover is implied without observation; exercises stop at agreed safety and rollback criteria. |
| Interface and link failure | Damocles reviews or observes approved interface and link failure scenarios, traffic movement, alerts and recovery. | Failover observer Architecture, dependency list, monitoring view, authorised failover window and rollback owner; customer inputs must identify the approved interface and link failure targets and expected behaviour. | A timestamped failover observation showing state, convergence, service checks, monitoring and recovery or rollback; the record names the tested interface and link failure object, path or control and its observed result. | Security and Privacy Controls for Information Systems and Organizations Release 5.2.0 Applicability: Performed only when a customer-approved recovery or failover scenario, observer and rollback window are available. Limit: The result covers the exercised failure mode; activity stops at the rollback threshold and does not predict every compound failure. |
| Routing convergence | Damocles reviews or observes route withdrawal and convergence against the selected scenario and success criteria. | Dependent service owner Named source and destination test points, expected flow matrix and a safe test window. | Source-to-destination results linked to rule, route or boundary evidence. | Security and Privacy Controls for Information Systems and Organizations Release 5.2.0 Applicability: Performed when both ends of the network path and the enforcing device are owned or expressly authorised for testing. Limit: Results cover the tested source, destination, protocol and route; testing stops on instability and does not authorise third-party or denial-of-service activity. |
| Upstream and downstream dependencies | Damocles traces upstream, downstream, identity, DNS, monitoring and provider dependencies for each resilient service path. | Dependent service owner Assessment requires topology, peer status, configurations, dependencies, scenario plan, monitoring and rollback authority, selected specifically for upstream and downstream dependencies. | Evidence records the configuration, status output, alert, traffic or session observation, timeline and recovery result relevant to upstream and downstream dependencies. | Security and Privacy Controls for Information Systems and Organizations Release 5.2.0 Applicability: Applies to Upstream and downstream dependencies when the design review or explicitly approved scenario includes this failure condition. Limit: The conclusion is limited to the sampled upstream and downstream dependencies; no production failover is implied without observation; exercises stop at agreed safety and rollback criteria. |
| Stateful session continuity | Damocles assesses expected stateful-session handling and records continuity only when an authorised scenario is observed. | Network or security engineer Assessment requires topology, peer status, configurations, dependencies, scenario plan, monitoring and rollback authority, selected specifically for stateful session continuity. | Evidence records the configuration, status output, alert, traffic or session observation, timeline and recovery result relevant to stateful session continuity. | Security and Privacy Controls for Information Systems and Organizations Release 5.2.0 Applicability: Applies to Stateful session continuity when the design review or explicitly approved scenario includes this failure condition. Limit: The conclusion is limited to the sampled stateful session continuity; no production failover is implied without observation; exercises stop at agreed safety and rollback criteria. |
| VPN failover | Damocles reviews VPN peer, tunnel, routing and session failover and observes it only in an approved test window. | Failover observer Architecture, dependency list, monitoring view, authorised failover window and rollback owner. | A timestamped failover observation showing state, convergence, service checks, monitoring and recovery or rollback; the record names the tested vpn failover object, path or control and its observed result. | Security and Privacy Controls for Information Systems and Organizations Release 5.2.0 Applicability: Performed when both ends of the network path and the enforcing device are owned or expressly authorised for testing. Limit: Results cover the tested source, destination, protocol and route; testing stops on instability and does not authorise third-party or denial-of-service activity. |
| Power and location dependencies | Damocles reviews evidenced power, rack, site and carrier diversity and records shared dependencies or unverified assumptions. | Dependent service owner Assessment requires topology, peer status, configurations, dependencies, scenario plan, monitoring and rollback authority, selected specifically for power and location dependencies. | Evidence records the configuration, status output, alert, traffic or session observation, timeline and recovery result relevant to power and location dependencies. | Security and Privacy Controls for Information Systems and Organizations Release 5.2.0 Applicability: Applies to Power and location dependencies when the design review or explicitly approved scenario includes this failure condition. Limit: The conclusion is limited to the sampled power and location dependencies; no production failover is implied without observation; exercises stop at agreed safety and rollback criteria. |
| Monitoring and alerting | Damocles confirms scenario alerts, status signals, escalation paths and recovery notification using supplied monitoring evidence. | Network or security engineer, Failover observer The HA monitoring or status view, expected peer, interface, routing and service alerts, notification destinations, a representative existing event or approved scenario where available, and the responsible network or security operations contact. | A peer, link or state alert; status-change timestamp; monitoring state before, during and after an approved scenario where performed; notification or recovery alert; and evidence showing whether the expected signal was generated. | Security and Privacy Controls for Information Systems and Organizations Release 5.2.0 Applicability: Configuration review applies wherever HA monitoring is in scope; live alert validation applies only where a representative event or approved scenario is available. Limit: Sampled monitoring evidence does not prove every future failure will generate an alert; no production failure is induced solely to test monitoring unless expressly authorised, and this row does not assess SOC investigation, threat detection or attack-detection coverage. |
| Capacity and failure domains | Damocles compares capacity assumptions and failure-domain load with current evidence without generating production stress. | Failover observer Architecture, dependency list, monitoring view, authorised failover window and rollback owner. | A timestamped failover observation showing state, convergence, service checks, monitoring and recovery or rollback; the record names the tested capacity and failure domains object, path or control and its observed result. | Security and Privacy Controls for Information Systems and Organizations Release 5.2.0 Applicability: Performed only when a customer-approved recovery or failover scenario, observer and rollback window are available. Limit: The result covers the exercised failure mode; activity stops at the rollback threshold and does not predict every compound failure. |
| Planned validation scenarios | Damocles defines representative failure scenarios, prerequisites, success measures, observers and stopping conditions. | Failover observer Assessment requires topology, peer status, configurations, dependencies, scenario plan, monitoring and rollback authority, selected specifically for planned validation scenarios. | Evidence records the configuration, status output, alert, traffic or session observation, timeline and recovery result relevant to planned validation scenarios. | Security and Privacy Controls for Information Systems and Organizations Release 5.2.0 Applicability: Applies to Planned validation scenarios when the design review or explicitly approved scenario includes this failure condition. Limit: The conclusion is limited to the sampled planned validation scenarios; no production failover is implied without observation; exercises stop at agreed safety and rollback criteria. |
| Production safety and rollback | Damocles confirms rollback authority, triggers, restoration steps and post-test health checks before any live scenario. | Failover observer Assessment requires topology, peer status, configurations, dependencies, scenario plan, monitoring and rollback authority, selected specifically for production safety and rollback. | Evidence records the configuration, status output, alert, traffic or session observation, timeline and recovery result relevant to production safety and rollback. | Security and Privacy Controls for Information Systems and Organizations Release 5.2.0 Applicability: Applies to Production safety and rollback when the design review or explicitly approved scenario includes this failure condition. Limit: The conclusion is limited to the sampled production safety and rollback; no production failover is implied without observation; exercises stop at agreed safety and rollback criteria. |
| Framework and controls | Mapping type | What Damocles assesses | Evidence produced | Applicability and limits |
|---|---|---|---|---|
| Security and Privacy Controls for Information Systems and Organizations Release 5.2.0 Framework-level context | Contextual | The engagement produces point-in-time technical evidence about the configured and observed security behaviour within the authorised Network resilience and high availability scope. | Coverage status, procedure results and findings can inform the customer’s control assessment and risk treatment records. | Applies: Framework-level context is provided when the customer uses NIST SP 800-53 to organise its security control program. Limit: No individual NIST control is published as verified; organisational implementation, continuous operation, governance and complete catalogue coverage remain outside the engagement. |
Assurance boundary: Damocles maps assessed coverage and observations to agreed objectives as traceable technical evidence. The review is not a certification, does not establish complete compliance, and does not confirm controls outside the authorised scope.
Records authorised scope and rules of engagement produced from the authorised Network resilience and high availability work.
Records coverage matrix produced from the authorised Network resilience and high availability work.
Records retest and residual-risk record produced from the authorised Network resilience and high availability work.
Records dependency and failure-domain map produced from the authorised Network resilience and high availability work.
Records failover observation timeline produced from the authorised Network resilience and high availability work.
Records recovery and rollback record produced from the authorised Network resilience and high availability work.
A resilience report with topology findings, scenario results, observed impact and prioritised recovery improvements.
The review separates device HA from end-to-end service resilience.
Both HA peers depending on the same upstream, link, virtual host, power or management service.
Routing changes that create state, NAT or inspection problems after the active path moves.
Session or configuration state not synchronised as the service design assumes.
Operators unable to reach or diagnose the standby component during a degraded event.
Failover and maintenance documented but not validated against the current environment.
Configuration, credential, backup or spare-hardware assumptions that delay restoration.
The scope identifies devices, sites, routing protocols, upstreams, management and maintenance procedures.
Review a firewall pair and the routing, upstream and management dependencies around it.
Review redundant edge paths, firewalls, routing and external connectivity.
Review gateways, portals, identity dependencies, routing and failover for remote-access services.
Review the maintenance plan and validate failover assumptions before a major upgrade or change.
The output supports architecture improvement and planned maintenance.
HA peers, links, routes, upstreams, management and service dependencies in the reviewed path.
Dependencies that can defeat or constrain the intended failover design.
Routing, state, NAT or management conditions affecting degraded operation.
Risk in upgrade, failover, reboot and restoration procedures.
Architecture and operational changes required to improve real resilience.
Controlled tests for confirming peer, routing, access and application behaviour during failover.
Live failover testing is planned around business impact, rollback, application validation and vendor/provider dependencies before production activity begins.
Architecture changes such as routing redesign, link changes, HA reconfiguration or management-path improvements are scoped with their own change and validation plan.
Application clustering, server/storage resilience, cloud-service availability and carrier contractual resilience are separate unless included because they are part of the tested network path.
A design review does not imply disruptive live failover testing unless expressly authorised.
We will define the devices, routing, upstream dependencies, maintenance procedures, live-test boundaries and deliverables required.