operational resilience

Operational Resilience: Definition and Use in Compliance

Published: Last updated:

Operational resilience is a risk management discipline that lets a financial institution keep delivering its important business services through disruptions like cyberattacks, outages, or supplier failures, and recover within tolerable limits when those services break.

What is Operational Resilience?

Operational resilience is a firm's ability to keep delivering its important business services through disruption, and to recover within limits it has agreed in advance. The shift from older thinking is subtle but real: instead of asking whether a system might fail, you assume it will, and you design around the consequences.

Regulators anchor the whole discipline on important business services. These are the services that, if interrupted, would cause intolerable harm to customers or threaten market integrity. For a retail bank that means access to deposits, card payments, and online banking. For a custodian it means settlement. The test is harm to the outside world, not inconvenience to the firm. Supervisors expect firms to map each service, identify every resource it depends on, then test what breaks when any of those dependencies fail. That extends past core banking platforms to questions like whether the MLRO can still file suspicious activity reports during a major incident.

Each important service gets an impact tolerance. This is a hard number: the maximum disruption the firm can absorb before the harm becomes unacceptable. "Payments restored within four hours" is an impact tolerance. "We aim to fix things quickly" is not. The board owns these numbers and has to defend them to supervisors.

Take a mid-sized bank whose card authorization platform goes down on a Friday evening. Without resilience planning, the team scrambles. With it, they already know the tolerance is two hours, they know which vendor and which data feed sit underneath the service, and they have a tested failover. The difference shows up in whether customers can buy groceries that night. Operational resilience connects directly to a firm's risk appetite and its broader control environment, since the controls protecting a service determine how much disruption it can survive.

The concept sits at the intersection of IT risk management, business continuity, and prudential supervision. Unlike a standard disaster recovery drill, it's designed to expose real failure modes: cloud provider outages, payment system breakdowns, data corruption events, third-party vendor collapses. The goal isn't to avoid disruption entirely. It's to confirm that when disruption happens, the institution stays inside its own pre-defined tolerances for downtime, data loss, and customer impact.

The testing side of the discipline appears under several labels: scenario testing, impact tolerance testing, business continuity testing, and operational risk stress testing. All are components of the same overarching requirement.


Operational Resilience in regulatory context

The regulatory picture solidified fast between 2021 and 2025. The UK led with binding rules, the EU followed with DORA, and the Basel Committee set global principles. Firms operating across borders now juggle several overlapping regimes, and what was once voluntary best practice now carries enforcement consequences.

In the UK, the Bank of England, Prudential Regulation Authority, and Financial Conduct Authority published their final policy in PS21/3 and SS1/21 in March 2021. Firms had to identify important business services and set impact tolerances by March 2022, then prove they could remain within tolerance by March 2025. Testing is mandatory: firms must test under a range of severe but plausible disruption scenarios and document the results.

The EU's Digital Operational Resilience Act applies from 17 January 2025. DORA is more prescriptive on technology, covering ICT risk management, incident reporting, resilience testing, and oversight of critical ICT third parties. It mandates threat-led penetration testing for the largest institutions and recurring resilience tests for all in-scope entities, which include banks, payment institutions, crypto asset service providers, and their critical ICT third-party providers. It even brings major cloud providers under direct EU supervisory scrutiny, a first for the sector.

In the US, the OCC's Bulletin 2020-10 and the Federal Reserve's SR 21-5 set out supervisory expectations for operational resilience across the banking sector. The FFIEC Business Continuity Management booklet requires scenario-based testing with documented results reviewed by boards.

Globally, the Basel Committee on Banking Supervision issued its Principles for Operational Resilience in March 2021, aligning the concept with operational risk management and defining it as one of seven core requirements. These principles cover governance, mapping, third-party dependency, and incident management.

The financial crime dimension is easy to miss. Compliance with FATF Recommendation 1 requires that financial crime controls are tested as part of a firm's risk-based approach, so sanctions screening platforms and transaction detection controls are explicitly in scope for resilience validation. When these controls degrade or go offline during an operational incident, firms face compounded exposure across two separate enforcement regimes, prudential and financial crime, asking the same uncomfortable question.

There's overlap with adjacent regimes too. A firm's third-party risk management program and its business continuity plan both feed the resilience picture, and supervisors increasingly expect to see them joined up rather than run as separate silos.


How is Operational Resilience used in practice?

In practice, resilience runs as a yearly cycle with continuous monitoring underneath it. Teams identify services, set tolerances, map dependencies, test against scenarios, fix gaps, and report to the board. Then they do it again as the business changes.

Mapping is where most of the hard work lives. For each important business service, the team documents every person, process, system, piece of data, and third party that the service depends on. A single payment service might touch a core banking system, a fraud engine, a sanctions screening tool, two cloud regions, and four vendors. Miss one dependency and your resilience picture is fiction.

Scenario testing is the proof. Firms run severe but plausible events: a prolonged cloud outage, a ransomware attack, the sudden failure of a critical supplier. The point is to find the dependency that breaks tolerance before a real incident does. A bank might discover its sanctions screening provider has no viable backup, which means screening stops the moment that vendor goes dark.

For compliance teams specifically, the work overlaps with transaction monitoring resilience. If monitoring fails, alerts queue, investigators fall behind, and regulatory filing deadlines come under pressure. One useful exercise: simulate 48 hours of monitoring downtime and count how many alerts and reports back up, then decide whether that volume is survivable. The honest answer often drives investment in redundancy.


What do regulators expect to see?

On exam day, examiners want evidence that resilience testing is real, documented, and acted upon. Vague policy frameworks and self-assessments don't pass.

The PRA's SS1/21 is specific: firms must maintain a mapping from important business services to the people, processes, technology, facilities, and third parties that deliver them. Examiners will ask to see that mapping, and they'll trace specific test results back to it.

Board-approved governance documentation. A resilience framework with clear executive ownership (typically CRO or COO), defined impact tolerances for each important business service, and evidence of at least annual board review.

Scenario test plans and results. Not theoretical scenarios. Documented tests of severe but plausible events: cyber attacks, third-party outages, mass staff unavailability, data loss events. Results must show whether the firm stayed inside its tolerances or breached them, and what remediation followed.

Third-party dependency testing. Firms can't disclaim responsibility because a vendor failed. Examiners expect evidence that critical third-party relationships are included in resilience testing, with contracts that give the firm the right to audit.

AML and financial crime continuity plans. MLROs face direct scrutiny on whether suspicious activity report filing can continue during a major incident. If transaction monitoring goes down and the firm continues processing transactions without detection capability, that's a material gap requiring immediate escalation, not a footnote in an incident report.

Board MI. Meeting minutes showing test results were presented to the board, with clear escalation paths when tolerances were breached.

Remediation tracking. Findings from tests must feed into a tracked register with closure dates. Examiners look for closure rates and item aging, not just a list of open issues.


What does good Operational Resilience Testing look like?

Best practice goes well beyond annual disaster recovery drills. The Basel Committee's 2021 Principles for Operational Resilience and the FCA's final rules both describe a continuous, iterative process built around real service outcomes, not just system availability metrics.

A mature resilience testing programme works through these steps:

  1. Map important business services to resource dependencies. Document every person, process, technology, data set, and third party that an IBS depends on. Update the map whenever the service changes.

  2. Set quantified impact tolerances. For each IBS, define a maximum tolerable duration of disruption (for example: "payment processing offline for no more than 4 hours"), a data loss threshold, and a customer impact ceiling.

  3. Test against tolerances, not just recovery time objectives. A recovery time objective (RTO) measures when a system comes back online. An impact tolerance measures whether the business service stayed within acceptable bounds during the outage. These are different questions, and regulators want evidence of the latter.

  4. Test the worst case. The FCA expects "severe but plausible" scenarios: simultaneous outage of primary and backup data centres, compromise of the firm's cloud provider, prolonged loss of key personnel under a pandemic-type scenario.

  5. Include financial crime controls in scope. Customer due diligence processes and transaction detection systems must be resilience-tested. A firm that can process payments but can't screen them isn't operationally resilient, regardless of what its IT recovery metrics show.

  6. Test third-party dependencies directly. Review SLAs, but also run tabletop exercises or simulation tests with critical vendors. Evidence of actual vendor testing is increasingly a standard examiner request.

  7. Report findings to the board with quantified gap assessments. Each test should produce a pass/fail verdict per IBS, a root cause analysis for any breach, and a tracked remediation plan with named owners and deadlines.


Common challenges and how to address them

The hardest part of operational resilience is honesty about dependencies. Most firms underestimate how many systems and vendors sit behind a single service, and the gaps surface only during a real incident or a rigorous test.

Concentration risk is the sharpest example. When dozens of banks run on the same handful of cloud providers, an outage at one provider can take down multiple firms at once. The fix isn't simple, since rebuilding on a second cloud is expensive and slow. Practical responses include negotiating stronger exit and failover terms, holding offline backups of critical data, and stress-testing the assumption that a major provider can vanish for 24 hours. This connects directly to concentration risk and fourth-party risk, where your vendor's vendor becomes your problem.

Impact tolerances are the second trap. Teams set them too soft to avoid hard conversations, then discover during testing that they can't meet even the generous number. The answer is to set tolerances based on customer harm, validate them with real recovery data, and let the board feel the discomfort early rather than during a crisis.

A third challenge is treating resilience as a documentation exercise. A binder full of dependency maps means nothing if no one has tested it. Run tabletop exercises with the people who'd actually respond, inject surprises, and capture what broke. One bank found its incident bridge call had no compliance representative, so sanctions and reporting decisions stalled for hours during a simulated outage. They fixed the runbook before it cost them in reality.


Common audit findings and exam citations

Regulators have cited institutions for operational resilience failures across four recurring themes.

Untested controls. The most common finding: controls existed on paper but hadn't been tested under realistic stress conditions. In the Danske Bank Estonia case, examiners found that AML controls were technically present but functionally inoperable under the actual transaction volumes the branch was processing. No resilience testing caught the gap before it became a €200 billion scandal. A rule that hasn't been stress-tested against real volumes isn't a working control.

Impact tolerances set to pass, not to challenge. Firms that set tolerances so wide that they're trivially met attract supervisory scrutiny. The FCA has been explicit in multiple Dear CEO letters: tolerances must be set to protect customers, not to avoid exam failures.

Poor governance of test results. Multiple citations describe boards that approved resilience frameworks but never received test results. Without board-level oversight, remediation doesn't happen at pace, and the same findings recur year over year.

Third-party blind spots. The Deutsche Bank mirror trades case demonstrated what happens when oversight of complex counterparty relationships is weak. Regulators now expect firms to test whether they can operate through a critical vendor failure, not just whether the vendor has an acceptable recovery plan on file.

AML control outages treated as technical incidents, not compliance failures. Multiple FCA reviews since 2020 have flagged firms that continued processing transactions while PEP screening or detection systems were offline. Processing without screening is a regulatory violation regardless of cause.

The FCA's 2021 multi-firm review of operational resilience found that fewer than 30% of surveyed firms could demonstrate scenario testing that genuinely challenged their impact tolerances. That's where examiner attention concentrates in every subsequent review.


Metrics and KPIs

Measuring control health for operational resilience testing requires going beyond whether tests were completed on schedule. The goal is to know whether the firm can actually stay inside its tolerances when it matters.

IBS coverage rate. What percentage of important business services have been tested in the last 12 months? A mature programme targets 100%. Below 80% is a gap regulators will flag.

Scenario severity distribution. What percentage of tests were "severe but plausible" versus low-risk or theoretical? Track the ratio. Programmes weighted toward easy scenarios produce reassurance, not assurance.

Impact tolerance breach rate per test cycle. In each scenario test, how many IBS breached their defined tolerance? Track this over time. A declining breach rate signals genuine improvement. A flat one signals untested controls or tolerances set too wide.

Remediation closure rate and item age. Open findings from resilience tests should be tracked to closure. Average age above 90 days is a governance warning sign. Regulators treat aged, unresolved findings as evidence of weak board oversight.

Third-party testing coverage. Of the firm's critical third-party providers, what percentage have been subject to resilience testing or exercise-based simulation in the past 12 months?

AML control downtime incidents. Track every instance where transaction monitoring or screening systems experienced unplanned downtime, with duration, impact assessment, and any regulatory notification made. This metric sits at the intersection of resilience and AML compliance, and it's one regulators ask about directly.

Time to restore versus tolerance. Track actual time to restore (TTR) against the stated impact tolerance for each IBS. Consistent overruns mean the tolerance is wrong or the recovery capability is insufficient.

Board reporting should cover all IBS with traffic-light status, scenario test outcomes, and a tracked remediation register presented at least quarterly.


Related terms and concepts

Operational resilience sits inside a web of related risk and continuity concepts, and understanding the neighbors sharpens the term itself.

The closest relatives are the building blocks regulators name explicitly. A critical business service (often called an important business service) is the unit resilience protects, and impact tolerance is the limit you commit to staying within. These two terms carry the regulatory weight; everything else supports them.

On the continuity side, disaster recovery and business continuity planning predate operational resilience and feed into it. Disaster recovery focuses on restoring technology after an event. Resilience is broader: it covers people, processes, and third parties, and it starts from the assumption that disruption will happen rather than treating it as an exception.

Incident handling links in through incident management, which governs how a firm detects, escalates, and resolves disruptions in real time. Strong incident management is how you actually stay inside an impact tolerance when something breaks.

The most direct financial crime connection is with transaction detection. Any resilience program that doesn't treat monitoring system availability as a critical service metric is leaving a material exposure unaddressed, because regulators treat detection control outages as AML failures rather than IT incidents. Sanctions exposure follows the same pattern: if screening systems go offline during an incident and transactions are processed without checks, the firm has breached its obligations regardless of intent.

The typologies that exploit degraded controls are worth naming. Layering and smurfing and structuring activity tends to accelerate during periods when detection systems are impaired. Sophisticated criminal networks time high-volume, rapid-fire transactions to coincide with known maintenance windows or incident response periods when monitoring is reduced.

FATF Recommendation 11 record-keeping requirements apply during resilience events too. Firms must maintain records even when primary systems are down. Testing whether backup record-keeping is adequate should be a standard scenario component in every annual test cycle.

Resilience also overlaps with the wider risk vocabulary. It connects to residual risk (what's left after controls), the three lines of defense governance model, and standards like ISO 31000 for risk management. For compliance leaders, the most useful link is to financial crime compliance, because a sanctions or monitoring system that goes dark is both a compliance failure and a resilience failure at the same time.

A firm that can demonstrate its financial crime controls remain operational under severe disruption is one regulators trust. One that can't explain what happens to its AML controls during a cloud outage is going to struggle in any supervisory review that asks the question, and they're all asking it now.


How FluxForce supports Operational Resilience Testing

FluxForce agents continuously monitor financial crime control availability, alert throughput, and detection rates in real time. When a control degrades or goes offline, the platform flags the gap immediately and creates an auditable incident record with timestamps, affected services, and downstream exposure.

Automated evidence capture means every test run, every alert processed, and every detection decision is logged with a full audit trail available for examiner review. The reporting layer surfaces IBS coverage metrics, remediation status, and board-level dashboards without manual extraction.

For compliance teams running scenario tests, FluxForce tracks whether controls stayed inside their impact tolerances throughout each test cycle.

Book a demo to see how FluxForce supports resilience testing.

Where does the term come from?

The phrase gained regulatory weight through the UK. The Bank of England, PRA, and FCA published a joint discussion paper in 2018 and final policy in March 2021, requiring firms to identify important business services and set impact tolerances by March 2022, with full compliance by March 2025.

The concept borrows from earlier work on operational risk under Basel II and from business continuity practice, but it shifted the frame from preventing failure to surviving it. The Basel Committee issued its own Principles for Operational Resilience in 2021. The EU then codified the idea for digital risk in the Digital Operational Resilience Act (DORA), which applies from January 2025. The term keeps expanding as cloud concentration and supply chain attacks reshape what "disruption" means.

How FluxForce handles operational resilience

FluxForce AI agents monitor operational resilience-related patterns in real time, flag anomalies for analyst review, and generate evidence-backed decisions with full audit trails.

← Back to Glossary