A few checkpoint tiles forming a clean route around a chrome clock, one pause tile and one forward tile.

Long-Running Foundry Agents: Recovery and Approval Planning

How should you plan recovery for a long-running Foundry agent?

Define the business task independently of the connection that starts it. Record meaningful progress, decide which actions can safely repeat and provide a clear status for the user. Recovery should preserve the intended outcome without silently duplicating an external action or losing an approval decision.

Microsoft's September 24, 2026 announcement describes long-running resilience in Foundry Agent Service as public preview. Assess it as a specific capability with its own constraints. The general availability of another Foundry component does not establish the status of this feature.

Why is background execution different from recovery?

A task can continue after a client disconnects yet still lose useful progress if the process running it stops. Those are different failure conditions. The architecture needs an answer for each, along with an explanation of what the user sees while the result is uncertain.

Microsoft's resilience documentation separates platform execution metadata from application-owned progress. Recovery can re-enter a handler; it does not automatically restore every intermediate business state. Use that distinction when deciding what your team must persist and verify.

Start with a timeline of the task. Mark reads, calculations, approvals and external writes. Ask what would happen if execution stopped immediately before or after each step. This exposes assumptions that a successful end-to-end demonstration can leave untested.

What progress should the application preserve?

Preserve enough information to determine what has completed and what remains. The useful checkpoint is a business boundary, not simply a record that the agent was active. For example, distinguish “analysis prepared,” “review requested” and “approved result submitted” when those states have different consequences.

  • Task identity: a stable identifier for the business operation.
  • Progress: the last completed phase and the evidence for it.
  • Approval state: what was approved, by whom and for which version.
  • External effect: whether a requested write was accepted.
  • Recovery decision: resume, safely repeat, request review or stop.

Keep the source of truth clear. If several components maintain conflicting status fields, an operator may not know which one to trust after an interruption. Define where each state belongs and how the application reconciles uncertain results.

How do you prevent duplicate actions after a crash?

Identify actions that would cause a problem if repeated. Creating a customer record, sending a message or submitting an instruction may need a different recovery rule from reading a document. Classify those actions before selecting a retry strategy.

Use the capabilities of the downstream system to establish whether an operation already succeeded. Where supported, a stable operation identifier can help recognize a repeated request. Where the outcome remains uncertain, route it for reconciliation instead of assuming that another attempt is harmless.

The current Foundry recovery guide provides technical implementation details. Keep implementation aligned with the current preview APIs. This article's business-state checklist is a design aid, not a substitute for reviewing the runtime contract and testing the actual integration.

How should human approval survive an interruption?

Record the exact proposed action that a person reviewed. An approval of one set of inputs should not silently authorize a different action after the task resumes. If relevant information changes, define whether the application must request a new review.

Make approval status visible to both the reviewer and the workflow owner. A task waiting for a person is different from a failed task or one that has completed. Clear states reduce repeated submissions and support requests caused by an ambiguous “still working” message.

Set a rule for expired or abandoned approvals. Identify who can cancel the task and what happens to its data afterward. A long-running process needs an end condition even when the original requester never returns.

Which failure tests should a pilot include?

Test interruption before an external write, after a write is accepted and while a human decision is pending. Include a client disconnect, an unavailable dependency and a cancellation request. Record the expected business state after recovery for each case.

Use a controlled environment and data appropriate for testing. Do not discover duplicate-action behavior by sending repeated instructions into a live business process. The purpose is to create evidence that the recovery design respects the intended boundaries.

Review both system records and user-visible behavior. An internal log may show that the task recovered while the user still sees a failed status. Conversely, a reassuring progress message may conceal a stalled operation. Both views must describe the same business situation.

What should operations receive before the workflow expands?

Provide a short runbook with common failure states, their owners and permitted recovery actions. Include how to find the task, inspect its status and determine whether external work already occurred. Operators should not have to infer a business outcome from a long sequence of raw tool messages.

Define monitoring around the process: tasks that exceed an expected duration, repeated recoveries, unresolved approvals and uncertain external effects. Set thresholds from observed behavior and business needs. Avoid presenting an arbitrary technical timeout as an accepted service commitment.

Use the existing Foundry production-readiness guide and multi-agent workflow patterns for the broader architecture. Recovery should fit the workflow's ownership and evaluation model rather than becoming a separate mechanism nobody owns.

Frequently asked questions

Does long-running execution guarantee that a task finishes?

No. Dependencies, approvals and application errors can still prevent completion. Define acceptable terminal states, including safe cancellation and a request for human intervention.

Does stream replay restore the entire workflow?

Do not treat replayed output as proof that application state and external effects are restored. Review the documented behavior of the selected runtime and preserve the business progress your workflow requires.

Who owns checkpoints and approval records?

The architecture should assign those responsibilities explicitly. Platform execution support does not remove the application's responsibility for its business state and the meaning of an approval.

Is this preview appropriate for a critical production dependency?

Review the preview terms, limitations and alternatives before making that decision. Use a controlled pilot to learn, and avoid treating a successful test as a production service guarantee.

Make long-running work recoverable

Explore Agentic Platforms and AI Native Product Engineering. If an interrupted agent workflow loses progress or repeats actions, Say Hello. Share the task and the point where recovery becomes uncertain.

Manish Surapaneni

A visionary leader passionately committed to AI innovation and driving business transformation.

Share:

Ready to apply this to your business?

Tell us the workflow you want to improve. Our consulting team will review your enquiry and discuss a practical next step. We aim to respond within one business day.

Discuss This Challenge
Discuss This Challenge

Insights & resources

Frequently Asked Questions
No items found.
Technology