Intuit Uses AI to Coordinate Disaster Recovery: What Can Businesses Learn from Its AWS Assistant?
Intuit's AWS AI assistant shows how businesses can coordinate disaster recovery faster. Learn key lessons for automation, resilience, response and planning now.

Intuit's AWS-based disaster-recovery assistant shows that AI can coordinate recovery without being given unrestricted control of production systems. For Sydney and NSW organisations, the practical lesson is to keep failover execution deterministic, encode recovery knowledge, verify readiness, preserve approval gates and use people for consequential decisions. AI disaster recovery works best as supervised operational coordination, not as autonomous emergency authority.
Disaster recovery is often described as a technology problem: replicate the data, maintain backup infrastructure, define a recovery environment and restore the service. At enterprise scale, however, the difficult part can occur between those components.
Somebody still has to determine which recovery process applies, establish whether the affected service is ready to move, understand dependencies, navigate change controls, record the decision, trigger the correct workflow, follow its progress and decide what to do when the normal procedure encounters an exception.
Intuit's recent work with Amazon Web Services provides an unusually useful example of how artificial intelligence can enter that coordination layer without replacing the deterministic systems that actually execute a production failover.
According to AWS, Intuit built an AI-powered disaster-recovery assistant known as EWOK Agent on top of its existing Ecosystem Wide Orchestrator Kit, or EWOK. The underlying platform already standardised recovery execution across infrastructure including compute, databases, networking, caches and asynchronous workloads. AWS says the established orchestration system reduced recovery times from several hours to about 20 minutes for supported workloads.
The remaining problem was operational knowledge. Experienced engineers still needed to know which recovery workflow to use, whether an asset was ready and what to do when policy or operating conditions interrupted the normal path.
That is where the AI assistant becomes strategically interesting.
The Important Innovation Is Not Autonomous Failover
It would be easy to describe the Intuit system as an AI agent capable of failing over production infrastructure. That description misses the architecture that makes the example useful.
The large language model does not appear to have been made the final execution authority. Instead, Intuit separates interpretation and coordination from the tested machinery that changes production systems.
In practical terms, an engineer can make a plain-language recovery request. The assistant can determine the relevant capability, resolve the asset, identify recovery workflows, check readiness and policy conditions, and coordinate subsequent steps. State-changing operations are then executed through conventional APIs and the existing recovery platform.
This creates a useful distinction for any organisation considering AI disaster recovery:
- Interpret an operator's request
- Suitable role for AI: Identify intent and relevant recovery capability.
- Suitable role for deterministic systems or people: Validate the resulting action against known assets and permissions.
- Select between known procedures
- Suitable role for AI: Reason over approved recovery options.
- Suitable role for deterministic systems or people: Execute only registered and tested procedures.
- Check operating context
- Suitable role for AI: Assemble readiness, policy and incident information.
- Suitable role for deterministic systems or people: Apply authoritative policy gates.
- Handle routine coordination
- Suitable role for AI: Move through approved steps and report status.
- Suitable role for deterministic systems or people: Stop on defined errors or exceptions.
- Resolve consequential exceptions
- Suitable role for AI: Present context and available paths.
- Suitable role for deterministic systems or people: Authorised human decides where judgement is required.
- Change production infrastructure
- Suitable role for AI: Request an approved operation.
- Suitable role for deterministic systems or people: Tested code executes it under controlled credentials.
That boundary matters because disaster recovery is one of the worst environments in which to allow a probabilistic system to improvise.
The organisation may already be operating with degraded visibility, elevated time pressure, unusual infrastructure conditions and incomplete information. Automation should remove coordination burden without creating another uncontrolled source of variation.
Intuit Has Turned Recovery Knowledge Into An Operational Asset
One of the more significant ideas in the AWS case study is Intuit's treatment of recovery procedures as structured, versioned skills rather than relying solely on conventional runbooks.
Traditional disaster-recovery documents can be comprehensive and still perform badly during an incident. The problem is not necessarily that the instructions are wrong. It is that the organisation must locate the right document, identify the applicable section, understand the current environment, translate prose into actions and remember what happens when the procedure reaches an exception.
At scale, this creates dependency on what businesses often call tribal knowledge. The technically documented recovery process is supplemented by information held in the heads of the employees who have previously performed it.
Intuit's approach is important because it attempts to encode more of that operating knowledge into machine-consumable capabilities with defined inputs, outputs, rules and stopping conditions.
This is related to, but more operationally demanding than, ordinary workflow documentation. A recovery skill must answer questions such as:
- What asset is being acted upon?
- Which environments are valid?
- Which recovery workflows are available?
- What must be true before execution begins?
- Which policy restrictions can block execution?
- Which error should stop the process immediately?
- What information must be returned after each stage?
- When must an authorised person intervene?
That shift can make disaster recovery less dependent on who happens to be on call. It also creates a stronger foundation for training, testing and post-incident review.
A Recovery Assistant Should Know When It Is Not Allowed To Continue
Intuit's change-freeze example is particularly valuable for business operators.
A technically possible failover may still be operationally prohibited. During a restricted change period, the system can encounter a policy gate requiring an incident reference or authorised emergency justification before proceeding.
The distinction is subtle but important.
The restriction is not treated simply as a technical failure for an AI model to work around. It becomes a defined branch in the operating process.
For Sydney businesses, the same principle can apply far beyond cloud infrastructure. A business-continuity workflow may encounter a valid restriction because:
- A customer communication requires management approval.
- A payment cannot be redirected without finance authority.
- A replacement supplier has not completed procurement checks.
- A building access change requires strata or facilities approval.
- A regulated record must be preserved before a system is restored.
- A recovery environment has not passed its readiness test.
- Another critical dependency has not yet recovered.
- An emergency action would exceed the operator's delegated authority.
Strong automation does not make these boundaries disappear. It makes them harder to overlook during an incident.
This is distinct from the security issue examined in Elyment's analysis of AI-agent tool-call controls and inspection.
Inspection asks whether an action should be allowed through a technical boundary. Disaster-recovery coordination asks whether the entire recovery sequence remains authorised, ready and operationally correct.
The Engineer Moves From Manual Orchestrator To Supervisor
AWS describes a change in the engineer's role that may ultimately be more important than the model itself.
Previously, the engineer coordinated a sequence of runbook lookups, consoles and API calls. With the assistant, routine coordination is handled by the system while the engineer remains involved in judgement calls and approvals.
That is a more credible model for high-consequence enterprise automation than the common idea of removing people from the recovery process.
Human expertise is expensive during an incident precisely because it is scarce. Senior staff should not spend their most valuable minutes copying identifiers between systems, checking repetitive statuses or manually navigating a procedure the organisation has already approved.
Their attention should concentrate on questions automation cannot safely settle:
- Is the current incident materially different from the tested scenario?
- Should an operating restriction be overridden?
- What customer or business impact is acceptable during recovery?
- Is a dependency sufficiently healthy to proceed?
- Should the organisation continue failover, stop or adopt an alternative arrangement?
- Who must be informed before the next consequential action?
Australia's current Guidance for AI Adoption follows a similar risk-based principle: meaningful human oversight should increase with the autonomy and consequence of the AI system, while organisations should retain intervention and alternative operating pathways for critical functions.
What Sydney Businesses Can Apply Without Building Intuit's Infrastructure
Most NSW organisations do not operate thousands of microservices across multiple cloud regions. They should not copy Intuit's architecture component by component.
They can copy the operating principles.
- Identify the business service before the technology. Define what must continue: taking payments, receiving customer enquiries, accessing project information, processing claims, dispatching field staff, managing bookings or communicating with customers.
- Map the dependencies required to deliver that service. Include identity systems, cloud applications, databases, telecommunications, payment services, external providers, staff access and physical facilities.
- Define a known recovery path. AI should coordinate a procedure the organisation understands rather than inventing one after the outage begins.
- Turn important checks into explicit gates. Recovery environment readiness, approvals, incident status, data integrity and dependency health should be observable conditions.
- Separate interpretation from execution. Let AI assemble context and select among approved capabilities. Keep consequential changes behind constrained APIs, rules and permissions.
- Define exceptions before production. Decide which conditions stop the process and which person or team receives the case.
- Create an execution record. Preserve what was requested, what was approved, what actions occurred, when they occurred and whether each stage succeeded.
- Exercise the workflow. A recovery procedure that has never been rehearsed is still largely theoretical.
Businesses beginning this work may need process mapping before automation. Elyment's workflow automation approach for Sydney operations teams focuses on system handovers, approvals, exceptions and audit trails rather than treating the AI model as the complete workflow.
For Property And Project Operators, Recovery Is Also A Delivery Problem
The relevance extends beyond software companies.
Consider a Sydney property-services operator managing multiple active renovation projects. Its delivery environment may depend on a CRM, cloud document storage, accounting software, scheduling, email, messaging, supplier systems and mobile access for project staff.
If one core system becomes unavailable on the morning of a major strip-out or flooring programme, the immediate problem is not simply restoring a server.
The business may need to establish:
- Which crews are already authorised to attend.
- Which site addresses and access instructions remain trustworthy.
- Which strata restrictions or working hours apply.
- Whether the latest scope variation was accepted.
- Which purchase orders have already been issued.
- Whether waste, parking or lift bookings remain confirmed.
- Which customer communications must be sent.
- How new project records created during the outage will later be reconciled.
A backup may restore the database. It does not automatically coordinate those operational dependencies.
This is why AI disaster recovery should be considered alongside business process continuity rather than being left solely inside infrastructure engineering.
Australian Guidance Is Moving Towards The Same Operating Discipline
The Intuit example is an enterprise engineering case study, not an Australian compliance standard. Its architecture nevertheless aligns with several resilience themes already visible in Australia.
The Australian Signals Directorate recommends that organisations maintain incident response arrangements aligned with business continuity and disaster recovery, establish recovery procedures and regularly test them. Its Essential Eight maturity model also treats restoration testing as part of effective backup practice.
For APRA-regulated entities, the regulatory expectations are stronger. The current CPS 230 operational-risk framework requires critical operations, disruption tolerance, business-continuity capabilities, disaster-recovery planning for critical information assets, dependency assessment and systematic testing.
NSW Government organisations operate under additional state cyber-security and AI assurance frameworks. Those requirements do not automatically apply to ordinary private businesses, but they reinforce a wider direction of travel: operational resilience depends on named accountability, tested processes, controlled AI use and demonstrable assurance.
Organisations considering AI in sensitive workflows can also review AI readiness for Sydney operations before connecting reasoning systems to business-critical applications.
The Critical Metric Is Not Whether The Agent Completed Its Task
A disaster-recovery assistant should not be judged by the number of prompts it answers or how convincingly it explains an outage.
The relevant measures are operational.
- Time to establish the correct recovery path
- What it reveals: Whether coordination has actually become faster.
- Time spent by senior engineers on routine orchestration
- What it reveals: Whether specialist capacity is being preserved for judgement.
- Policy-gate exception rate
- What it reveals: How often the normal recovery path encounters governance constraints.
- Human intervention rate
- What it reveals: Whether authority boundaries are correctly placed.
- Failed or aborted recovery stages
- What it reveals: Where underlying procedures remain fragile.
- Recovery within business tolerance
- What it reveals: Whether the critical service was restored quickly enough.
- Audit completeness
- What it reveals: Whether management can reconstruct the incident afterwards.
These measurements can also reveal a difficult truth: AI cannot repair an undefined operating model.
If teams disagree about which system is authoritative, nobody owns an exception, failover dependencies have never been tested or a recovery procedure exists only in one employee's memory, adding an agent will expose those weaknesses rather than resolve them.
Disaster-Recovery Automation Should Be Tested As A Complete Sequence
The most important practical lesson may be to stop testing individual recovery components in isolation.
A database can restore successfully while the business service still fails because authentication is unavailable. A secondary application can start while its upstream integration remains unreachable. A cloud environment can be technically healthy while staff do not have the credentials or communications channels required to operate it.
ASD's incident-response guidance encourages organisations to regularly review and test response plans, including the people, processes, technology and supporting procedures required for recovery.
An AI-assisted recovery workflow therefore needs exercises that include imperfect conditions:
- One recovery dependency is unavailable.
- The normal approver cannot be reached.
- The procedure encounters a policy restriction.
- A third-party provider responds slowly.
- One data source disagrees with another.
- The AI assistant itself becomes unavailable.
- A failover is started but must be safely stopped.
- Operations must continue temporarily using a manual alternative.
The final scenario is especially important. AI should improve resilience, not become another dependency without which the organisation can no longer recover.
That principle complements Elyment's earlier analysis of shutdown and recovery controls for production AI agents.
A business should be able to recover its core operation even when the automation intended to assist that recovery is unavailable.
Map The Recovery Workflow Before AI Coordinates It
Review critical processes, system dependencies, recovery sequencing, approval boundaries, exception paths, audit requirements and human responsibilities before AI is connected to a business-continuity workflow.
Request An Operational Workflow Review
What Intuit's Example Really Says About AI Disaster Recovery
Intuit's AWS assistant is significant because it avoids one of the most tempting ideas in enterprise AI: that giving a capable model more authority automatically creates better automation.
The opposite pattern is visible here.
The underlying recovery actions remain structured. Assets are resolved. Readiness is checked. Policy gates exist. Production changes are recorded. Credentials sit outside the model. Significant exceptions can require human intervention. Execution status is returned through a traceable workflow.
AI is valuable because it connects those elements and reduces the manual coordination required from an experienced operator.
That is the lesson Sydney and NSW businesses can apply regardless of their scale.
Do not begin by asking how much disaster-recovery authority an AI agent can be given. Begin by asking whether the business already has a recovery process precise enough to be coordinated safely.
Once the service, dependencies, procedures, gates, exceptions, permissions and human responsibilities are explicit, AI can become a useful supervisory layer.
Until then, automation risks making an undocumented recovery process move faster without making it more reliable.
Map The Recovery Workflow Before AI Coordinates It
Review critical processes, system dependencies, recovery sequencing, approval boundaries, exception paths, audit requirements and human responsibilities before AI is connected to a business-continuity workflow.
Request a Workflow Review