Imagine an AI-assisted claims management process.
The system receives documentation, retrieves customer information, classifies the case, assigns a priority and routes it to the appropriate team.
One morning, the model continues to operate normally. It responds within the expected timeframes and maintains its performance metrics.
However, a change in permissions prevents it from accessing part of the documentary context.
The AI is still working.
The operation can no longer function in the same way.
At that point, the issue is no longer the model’s performance. It becomes a question of the operational capability the organisation needs to preserve.
The team can still process claims. But it now needs to know which claims can proceed, in what order, with what information and for how long it can continue without compromising the service.
The situation reveals something increasingly relevant as AI becomes embedded in critical processes: operational resilience in AI depends on more than the reliability of the model an organisation uses.
When a tool becomes a dependency
The Financial Stability Board (FSB) documents the use of AI in credit, fraud detection, AML/KYC, customer service, risk management, back-office operations, document review and compliance.
Some of these applications have already reached considerable scale. Among the cases documented by the FSB is a system that analyses more than 80 million fraud-related signals every day, alongside an insurer that eliminated approximately 400,000 manual interventions during 2025.
The significance of these figures lies in the level of integration they reveal.
When AI becomes a stable part of a process, a chain of dependencies develops around the model: data, infrastructure, APIs, identity management, permissions, connectors, rules, workflows, legacy systems and the people responsible for specific decisions.
Each layer affects the operational capability available.
A data source may become inaccessible. A permission may change. An integration may be interrupted. An external update may alter expected behaviour.
The Bank for International Settlements (BIS) warns of growing dependence on specialised hardware, cloud services, third-party data providers and pretrained models, with significant concentration among certain providers. It also notes that these interdependencies can increase the intensity, speed and complexity with which disruptions spread.
This requires a broader perspective.
Evaluating the model remains necessary. But when that model supports part of the operation, we also need to understand which business capabilities depend on it and which other components must remain operational to sustain them.
Operating when conditions change
Let us return to the claims process.
The loss of documentary context forces the organisation to decide how the process should continue.
Some cases can proceed.
Others require human review.
Certain decisions can wait.
And some operations must stop until the necessary conditions have been restored.
The organisation enters a degraded mode of operation: it maintains part of its capability under conditions that differ from normal operations.
This raises a decision that should be made before it becomes necessary:
What level of operation do we want to preserve when normal functioning can no longer be maintained?
At Cognodata, we see that when a process begins to depend on several of these layers, one question significantly changes the conversation:
If we removed one of these dependencies tomorrow, would the team know exactly what it could continue doing?
Answering this requires understanding more than the technology architecture.
Which services must be maintained.
What can wait.
Which decisions require information that is no longer available.
Who can change priorities.
Who has the authority to stop part of the workflow.
And under what conditions normal operations can be restored.
Business continuity and operational resilience have long addressed these questions. The introduction of AI creates new combinations of dependencies and extends this discussion to processes with increasing levels of automation.
The FSB recommends incorporating AI-related scenarios into technology risk testing and paying particular attention to third parties, concentration, supply chains and continuity when these systems support critical or material functions.
In Europe, DORA has strengthened the management and oversight of technology and third-party dependencies since 2025, alongside operational resilience testing.
These developments point to an important shift: the ability to continue operating must be defined and tested before a disruption occurs.
How much operational capability do we need to preserve?
The answer depends on the process.
An internal assistant used to prepare meetings may tolerate several hours of unavailability.
A system involved in prioritising fraud cases, sensitive claims or critical incidents requires different conditions.
That is why the starting point should be the business capability the organisation needs to protect and its materiality.
Five questions help make this explicit:
- Which business capability are we making dependent on this system?
- What minimum level of operation must we preserve if one of its dependencies becomes unavailable?
- What can continue, what must be reduced and what must stop?
- Who has the authority to change the operating mode?
- What conditions must be met before normal operations can resume?
The answers help distinguish use cases that can wait from those that need to operate in a controlled degraded mode.
They also help identify which dependencies require redundancy, which alternatives must remain available and where the authority to act should sit.
Maturity therefore begins to incorporate another dimension: understanding the minimum operational capability the organisation needs to preserve before deciding how much it wants to automate.
Dependency is also a decision
There is another consequence.
When we introduce AI into a process, we decide which capabilities we want to gain: greater speed, increased scale, new analytical possibilities or automation.
At the same time, we are creating a new architecture of dependencies.
Some of the capability that previously resided in particular people, systems and procedures becomes supported by a different combination of models, data, providers, permissions, integrations and internal knowledge.
Every decision to introduce AI does more than add a new capability. It also redefines the dependencies on which that capability will rely.
This means that scaling AI also involves deciding which dependencies we are prepared to accept in order to sustain a particular capability.
And that decision can be made explicit.
We can identify what must remain available.
Define how far a service can be degraded.
Assign authority to stop it.
Establish the conditions for restoring it.
And test whether those decisions work before they are needed.
Maturity then becomes visible in something very concrete: the organisation knows which capabilities it needs to preserve, what they depend on and what it is prepared to lose temporarily in order to maintain control over what really matters.
Frequently Asked Questions About Operational Resilience in AI
What is operational resilience in artificial intelligence?
Operational resilience in AI is an organisation’s ability to maintain essential business functions when an AI system or one of its dependencies can no longer operate under normal conditions. It involves understanding which operations can continue, which must be restricted and who has the authority to intervene.
Can an AI model work correctly while still affecting business continuity?
Yes. A model may maintain its performance metrics while losing access to the data, permissions or integrations required to execute the complete process. This is why evaluating model performance is not the same as assessing operational continuity.
What does operating in degraded mode mean when using AI?
It means maintaining part of a service under conditions that differ from normal operations. This may involve reducing operational volumes, introducing human reviews, changing priorities or suspending certain decisions until the necessary conditions have been restored.
How can an organisation prepare for failures in AI dependencies?
By identifying critical business capabilities, their dependencies, the minimum level of operation that must be preserved and the people authorised to modify or stop the process. The organisation should also establish and test the conditions required to restore normal operations.