Data Center Resilience: A Buyer Guide
Data center resilience is the ability to keep important services within an acceptable operating state when equipment, power, networks, suppliers, people, or facilities fail. Buyers should compare dependencies and recovery evidence, not only facility labels.
What does data center resilience include?
Resilience includes facility design, power, cooling, network connectivity, physical security, operations, supplier dependencies, software, people, and recovery. A resilient facility can still support an unreliable service if the application, identity system, or carrier path has a single point of failure.
Start with the service the data center must support. Define the business or clinical consequence of interruption, the data that must be recovered, the acceptable delay, and the people who must make decisions during an event.
The NIST Cybersecurity Framework is useful for the information and security side of resilience. It should sit beside facility, application, continuity, and supplier assessments rather than replacing them.
Resilience test: draw the service path from user to data and mark every shared dependency.
Which dependencies should a buyer map?
Dependency mapping reveals the difference between a resilient component and a resilient service. Buyers should map the facility, power feeds, cooling, network carriers, DNS, identity, storage, backup, monitoring, support, and the people who can change or restore them.
Ask providers to show the failure boundary. If two sites use the same carrier, control plane, region, staff, or supplier, they may not be independent for the event the buyer is trying to survive.
The map should be readable by operations and leadership. A technical architecture that cannot explain the customer impact of a failure will not guide a good recovery decision.
- Facility: location, hazards, access, maintenance, and shared building systems.
- Power and cooling: feeds, backup duration, testing, fuel, capacity, and failure transfer.
- Connectivity: carriers, routes, edge points, DNS, and external control planes.
- Technology: compute, storage, backup, identity, monitoring, and management systems.
- People and suppliers: skills, response coverage, contracts, spares, and escalation.
How should recovery evidence be evaluated?
A recovery point or recovery time statement is only useful when the service, data, dependencies, and test conditions are defined. Ask what was tested, when, with which version, and what users experienced. A tabletop discussion and a technical failover answer different questions.
Request the test summary and the unresolved actions. Mature operations do not claim that every exercise is perfect. They show what failed, what was accepted, who owns the fix, and when it will be retested.
Recovery also includes data integrity and access. A system that starts quickly but restores incomplete data, loses audit history, or cannot authenticate the right users is not recovered for the business.
| Evidence | What it answers | Question to ask |
|---|---|---|
| Design documentation | What should happen? | Which dependency and failure boundary does the design cover? |
| Tabletop exercise | Who decides and communicates? | Are roles, contacts, and escalation clear? |
| Technical test | Can the system fail over or restore? | What version, data, and user path were tested? |
| Live exercise | Does the service work under stress? | What did users see and what remained unresolved? |
What should capacity and maintenance cover?
Resilience depends on capacity headroom and the ability to maintain systems without creating avoidable risk. Buyers should understand current utilisation, growth assumptions, replacement cycles, maintenance windows, and the effect of a component failure on the remaining capacity.
Ask how the provider handles a failure during planned maintenance, a heat event, a carrier outage, or a supplier delay. The answer should include decision thresholds and communication, not just a statement that redundancy exists.
Maintenance evidence matters because backup equipment that is not tested, fuel that is not managed, or spares that are not available can create a false sense of protection.
- Headroom: how much capacity remains after a component or site is unavailable?
- Maintenance: can planned work occur without sharing the same failure risk?
- Replacement: how are ageing components identified and funded?
- Supplier: what happens when a spare, contractor, or carrier is delayed?
- Communication: who receives the decision and service-impact update?
Data center resilience models compared
The right model depends on service criticality, geography, regulatory requirements, budget, and the organisation’s recovery capability. More sites do not automatically mean more resilience if they share dependencies.
Compare the operating burden as well as the architecture. A multi-site design can improve continuity but creates more data, identity, network, support, and testing work.
| Model | Best fit | Strength | Trade-off |
|---|---|---|---|
| Single resilient site | Moderate criticality with strong local design | Simpler operations and support | Large site or regional failure remains material |
| Active and standby sites | Services with planned recovery | Balances cost and recovery needs | Standby readiness needs regular tests |
| Active across sites | Low tolerance for interruption | Can continue through some failures | More complex data and traffic management |
| Cloud and facility mix | Variable workloads and risk | Flexible placement and capacity | Shared dependencies need careful mapping |
How should a buyer evaluate a provider?
Use a structured evidence request, then verify the parts that matter to the service. Ask for architecture boundaries, test summaries, maintenance practice, incident communication, supplier controls, security evidence, and recovery responsibilities. Separate provider commitments from buyer responsibilities.
Run a scenario workshop with operations, security, application owners, finance, and the service owner. Use realistic events: a carrier failure, a power event, a software mistake, a cyber incident, a building restriction, or a supplier outage. The workshop should end with named decisions and open gaps.
Compare the result with the questions in the digital infrastructure market insight. Resilience should be evaluated as part of the investment case, not as a facilities-only checkbox.
- Define service impact and recovery objectives.
- Map people, technology, facility, network, and supplier dependencies.
- Request evidence under normal, maintenance, and failure conditions.
- Test one representative service end to end.
- Record residual risk, owner, review date, and decision.
What does not matter as much as buyers think?
A tier label or a long list of redundant components does not describe the resilience of the service the buyer cares about. Uptime history alone also misses the question of how the provider will respond to a new failure mode.
The useful signal is evidence linked to an operating path. Buyers should know what fails, what continues, who acts, how users are informed, and how the service returns with trustworthy data.
One-page buyer worksheet
Use this worksheet for one important service. Draw the user path and mark facility, power, cooling, network, identity, storage, backup, supplier, people, and communication dependencies. For each, record the failure state and recovery decision.
Keep the worksheet with the research brief, procurement record, or operating review. It turns a broad market question into a set of checks that can be answered, assigned, and revisited when new evidence arrives.
- Service impact and acceptable operating state
- Shared dependency and independence test
- Normal, maintenance, and failure evidence
- Recovery data integrity and access check
- Owner, residual risk, and next exercise date
FAQ
What is the difference between redundancy and resilience?
Redundancy adds alternate capacity or components. Resilience also requires dependency mapping, operating decisions, testing, recovery, and learning.
How should a buyer verify recovery claims?
Ask what service, data, version, dependencies, and users were included in the test, then review unresolved actions and retest evidence.
Does a second data center remove all risk?
No. Sites may share carriers, identity, control planes, people, suppliers, or regional hazards. Independence must be tested against the failure the buyer wants to survive.
Who owns resilience?
The service owner is accountable for the outcome, with facility, infrastructure, application, security, supplier, and continuity teams owning defined parts.
What is the first resilience exercise to run?
Map one important service end to end and run a scenario workshop that includes technical failure, communication, user impact, and recovery decisions.
Bottom line
For more context on digital infrastructure investment, read the infrastructure market insight or request a buyer brief.