Infrastructure resilience guide

Design infrastructure around the service that must continue—not the hardware you happen to own.

Resilience comes from knowing which services matter, what they depend on, how long each dependency can be unavailable, what capacity remains during failure, and who has authority to recover or degrade service. This guide turns those questions into an operating model.

Substantively reviewed . The current source baseline uses NIST cyber-resiliency and contingency-planning guidance plus CISA resilient-power, infrastructure-dependency, and emergency-communications resources.

Executive summary

Infrastructure resilience is the ability to anticipate disruption, continue the most important service functions, recover within an acceptable window, and adapt after failure. NIST SP 800-160 Volume 2 Rev. 1 frames cyber resiliency around the ability to anticipate, withstand, recover from, and adapt to adverse conditions.NIST SP 800-160 Vol. 2 Rev. 1 NIST contingency-planning guidance similarly begins with business-impact analysis, recovery requirements, plan development, testing, and maintenance.NIST SP 800-34 Rev. 1

The practical implication is simple: resilience should not be organized around a list of servers, generators, carriers, or vendors. It should be organized around services and their dependency chains. CISA's Infrastructure Dependency Primer emphasizes that energy, communications, IT, transportation, water, and other systems are interdependent; failure in one can disrupt several others.CISA Infrastructure Dependency Primer

This guide uses seven operating layers: service criticality, dependency mapping, resilient power and communications, recovery architecture, maintenance and spares, supplier/concentration risk, and exercises/evidence.

1. Define the minimum service that must survive

Start with the outcome users or operators need, then identify what can degrade. A service that normally uses ten application nodes, two network paths, and a full analytics stack may only need a smaller transaction path to continue safely during a disruption.

Record four decisions for each important service

  • Minimum viable service: what must still work during a serious disruption?
  • Maximum tolerable outage: how long can that minimum service be unavailable before consequences become unacceptable?
  • Data and state requirement: how much data loss, transaction replay, or synchronization delay can be tolerated?
  • Recovery authority: who can declare degraded mode, fail over, suspend nonessential work, or accept temporary risk?

Do not automatically copy contractual SLA numbers into engineering recovery targets. A contractual availability percentage, a recovery-time objective, a safety requirement, and a user-experience objective answer different questions.

2. Build a dependency map that includes facilities and external services

For each important service, map the dependencies that must be available for the minimum viable service to work:

Technology dependencies

  • identity and privileged access
  • DNS, certificates, secrets, time, and configuration services
  • network paths, carriers, VPNs, and cloud control/data planes
  • storage, databases, queues, backups, and replication
  • monitoring, alerting, ticketing, and communications tools

Physical and organizational dependencies

  • utility power, backup power, fuel, cooling, and site access
  • replacement hardware, field technicians, and logistics
  • critical suppliers and managed-service providers
  • decision makers, communications staff, and external contacts
  • facilities, public safety, or other operational partners

Mark single points of failure and hidden common-mode dependencies. Two carriers that share the same local conduit, two cloud services that depend on the same identity tenant, or two backup systems that require the same administrator account are not truly independent.

3. Engineer power and communications to the required resilience level

CISA's Resilient Power Best Practices recommends defining resilience requirements, performing a gap analysis, planning operations and maintenance, testing readiness, and considering telecommunications together with power continuity.CISA Resilient Power Best Practices

Power decision record

  • Critical and noncritical loads, including startup/inrush behavior.
  • Required autonomy for UPS, batteries, generators, or other backup sources.
  • Fuel or energy replenishment assumptions during a regional disruption.
  • Transfer-system behavior and manual recovery procedure.
  • Maintenance/load-test cadence and known deferred maintenance.
  • Remote monitoring and local manual control if networks are unavailable.

Communications decision record

  • Primary and alternate network paths and their physical/common-provider dependencies.
  • Emergency communications if normal collaboration, email, or VoIP services fail.
  • Carrier escalation contacts and restoration-priority arrangements where applicable.
  • Out-of-band access for critical administration.

CISA's National Emergency Communications Plan emphasizes operable, interoperable, resilient, and secure communications during response and recovery.CISA National Emergency Communications Plan Even organizations outside public safety can use the underlying principle: do not make the same communications system you are recovering the only way responders can coordinate.

4. Make recovery architecture explicit

For each service, document which recovery pattern is actually implemented rather than assuming the word "redundant" is sufficient.

  • Local redundancy: component failure can be absorbed within the same site or failure domain.
  • Alternate site/zone: the service can move when a facility, availability zone, or local network is unavailable.
  • Regional recovery: state, data, identity, networking, and dependencies can be restored in another region.
  • Manual degraded mode: staff can continue the most important work when automation is unavailable.
  • Rebuild: infrastructure can be reconstructed from controlled configuration, backups, and documented dependencies.

Recovery plans should state what is intentionally not recovered first. Capacity, people, network paths, and supplier availability may not support restoring every service simultaneously.

5. Treat maintenance, spares, and configuration as resilience controls

Many outages are made worse by maintenance debt, undocumented changes, failed batteries, unavailable spares, or recovery equipment that was never tested under load. Maintain evidence for:

  • preventive maintenance and load testing;
  • battery/fuel/UPS/generator condition where applicable;
  • critical spare quantities, location, shelf life, and replacement lead time;
  • firmware and configuration baselines for infrastructure devices;
  • protected backups of network, hypervisor, storage, and facility-control configurations;
  • vendor support status and end-of-support dates.

A recovery design that depends on a part with a six-month lead time is not a recovery design unless that lead time has been accepted explicitly.

6. Manage supplier and concentration risk as architecture

Supplier resilience is not solved by collecting a questionnaire once. Maintain an inventory of dependencies whose failure can materially affect the service and record:

  • which service and data each supplier supports;
  • whether the dependency is substitutable and how long substitution would take;
  • shared concentration across cloud, carrier, identity, hardware, support, or logistics providers;
  • support and incident-notification paths;
  • exit data, configuration, and credential requirements;
  • reassessment triggers such as acquisition, service redesign, geography change, repeated incidents, or financial distress.

For detailed supplier governance, use Third-Party Technology Governance and Public-Sector Third-Party Cyber Risk.

7. Exercise the dependency chain, not just the backup file

NIST contingency-planning guidance includes testing, training, exercises, and plan maintenance as part of the lifecycle.NIST SP 800-34 Rev. 1 Build exercises around realistic loss of dependency rather than isolated component tests.

Useful scenarios

  • loss of the primary site plus degraded communications;
  • identity provider unavailable while administrators need recovery access;
  • regional cloud outage with replication lag;
  • generator starts but fuel replenishment is delayed;
  • critical carrier failure during a cyber incident;
  • major supplier unavailable during restoration;
  • backup restores technically but the application cannot reconcile data.

Capture actual recovery time, manual workarounds, missing access, missing contacts, data reconciliation issues, and decisions that required executive approval. Turn those findings into owned remediation rather than a presentation-only after-action report.

Decision-ready resilience metrics

Readiness

  • critical services with current dependency maps
  • services with tested recovery procedures
  • critical infrastructure with current maintenance evidence
  • known single points of failure without approved treatment

Performance

  • actual versus target recovery time in exercises/incidents
  • recovery failures caused by access, configuration, or dependency gaps
  • critical supplier/concentration findings past due
  • repeat incidents where the same resilience weakness reappeared

30-day infrastructure resilience reset

  1. Week 1: select the most consequential services, define minimum viable service, and map dependencies.
  2. Week 2: validate power, communications, identity, backup, site, and supplier assumptions against actual architecture.
  3. Week 3: run one cross-dependency recovery exercise and capture actual recovery evidence.
  4. Week 4: assign remediation owners, document accepted exceptions, and establish a quarterly resilience review.

For remote sites and distributed estates, continue with Edge Infrastructure Resilience. For telemetry and service-health design, use Cloud Observability.

Continue learning

Related guides after Infrastructure Resilience

Follow the next implementation topic without returning to search.

Put this guide to work

Turn Infrastructure Resilience Operating Guide into a decision-ready next step.

Use the source-backed research to pressure-test assumptions, then build a reusable evaluation brief before you compare products, scope implementation, or request a fit review.