An AI workflow is not finished when it produces a convincing answer. It is finished when the intended work has happened, the result can be checked, and the people responsible know what remains unresolved. Reliability comes from designing the transitions around the model as carefully as the model’s output.
Define completion in the destination system
Consider a workflow that turns a meeting handover into a project task and notifies its owner. The model may extract a clear action, the project system may accept a write, and the notification service may fail. Describing the whole run as either “success” or “failure” hides information the operator needs. Record the task and notification outcomes separately.
Write down the business completion condition: the correct project contains one assigned task with an agreed due date, and the intended owner has been notified or an explicit notification failure has been handed over. Decide whether notification failure blocks completion or creates follow-up work. That is a product decision, not a detail for the exception handler to invent.
Use durable workflow records that survive a process restart. Store an operation identifier, intended action, relevant record versions, approval state, attempt history, and the destination’s result identifier. Keep unnecessary private content out of operational logs. An operator should be able to establish the state without reconstructing it from a model’s prose.
Make approval specific and time bounded
Before an external change, present the exact proposal to an authorized reviewer: the project, assignee, due date, task text, and supporting handover. Show missing information instead of silently filling it with a plausible value. Give the reviewer a useful choice to approve, correct, or decline the proposal.
Approval should refer to a stored version of that proposal. If the assignee or destination changes after review, return it for the appropriate approval. Set an expiry where stale authorization could matter, and recheck the reviewer’s current permission before execution. A long-running workflow should not treat an old browser session as permanent authority.
OWASP’s excessive-agency guidance recommends limiting tools and permissions and involving people in consequential actions. Translate that into controls the workflow actually enforces: the execution service checks the approved operation and refuses mismatched changes. A prompt asking the model to “always get approval” is not sufficient evidence that every execution path respects the boundary.
Make an uncertain outcome visible.
Validate the proposed change and required evidence.
Bind permission to the specific approved action.
Record attempts using a stable operation identity.
Check the destination, then complete or route the unresolved work.
Make the callWhen an external action has an unknown outcome, reconcile it using a stable operation identity before deciding whether a retry is safe.
Give every consequential operation a stable identity
The project API can create the task and then lose the response before your workflow receives it. Retrying with a new request may create a second task. First establish what the provider guarantees: whether it accepts an idempotency key, what scope that key has, how long it is retained, and how a repeated request is reported.
AWS describes caller-provided request identifiers as a way to distinguish a retry of the same intent from a new operation. Where the destination supports that contract, preserve the same identifier across retries of the approved action. A genuinely new action needs a new identity. Reusing a key with changed parameters is not a safe way to revise an approved task.
If the destination has no suitable mechanism, keep the outcome unresolved until you can reconcile it through a reliable destination identifier or an operator. A local “already sent” flag cannot alone prevent duplicates across a crash between the remote change and the local update. Nor does a search followed by a write automatically protect against concurrent attempts. Document the remaining limitation before enabling automatic replay.
Retry a classified failure, not an entire story
An invalid assignee, an expired credential, a temporary service outage, and an unknown write outcome require different responses. Define a policy for each operation using the provider’s actual error contract. Google Workflows, for instance, documents different default retry policies for idempotent and non-idempotent steps. A generic “retry on error” setting does not establish safe behavior.
Put a limit on attempts and total elapsed time. Use the supported backoff behavior and respect service retry instructions. AWS explains that retries can worsen overload and that adding jitter spreads attempts over time. Check whether the SDK already retries before adding another loop around it; layered retries can multiply the requests during an outage.
The workflow also needs a useful exhausted state. Preserve the last meaningful error, affected operation, attempts, and next owner. Avoid turning exhaustion into an empty successful output. For the project task, a failed notification can wait in a notification queue; it should not cause the successfully created task to be submitted again.
A timeout tells you that you stopped waiting. It does not tell you whether the work happened.
FROM THE AYAN STUDIO
Keep recovery boundaries aligned with side effects
Split the workflow at meaningful changes: prepare the proposal, record approval, create the task, confirm its identity, send the notification, and record the notification result. Each boundary should make a partial outcome understandable. Do not put several independent external writes inside one opaque step if replaying that step can repeat earlier successful writes.
Cloudflare’s workflow guidance describes persisted step results and advises designing retried operations to be idempotent. That is a useful engineering model, but a durable workflow engine does not automatically give every connected service the same guarantee. Review the contract at the external boundary, including how the workflow resumes after an interrupted response.
Keep the task-creation result available to the later notification step. If a restart occurs, the workflow should continue from its recorded state rather than ask the model to reconstruct what it probably did. Where the engine’s recovery semantics differ from your business requirement, add an explicit reconciliation state and test it. The state map should explain recovery as clearly as the ordinary path.
Assign recovery work before an incident
Not every action can be undone. A notification may already have been read. A task may have acquired comments or downstream dependencies. Decide which corrections are allowed and which require the process owner. Automatically deleting a task to compensate for a failed notification may remove useful work that another person has already started.
Microsoft’s compensating-transaction guidance explains that recovery follows business rules, may not restore the original state exactly, and can itself fail. Treat compensation as another operation with its own evidence and failure handling. Keep sufficient state to resume it, and identify when manual intervention is the appropriate outcome.
Give the operator a compact recovery view: intended change, confirmed effects, unknown effects, supporting identifiers, and permitted next actions. Separate “retry notification” from “create task again.” Make the owner and escalation path explicit, including coverage when the usual reviewer is unavailable. A recovery queue without anyone responsible is only a different place to hide unfinished work.
Release only after exercising interrupted work
Test the handover workflow with a lost response after task creation, a duplicated trigger, an expired approval, a changed assignee, an unavailable notification service, and a restart between steps. For each run, inspect the actual project record and notification state. Confirm that authorized work can still finish and that unresolved work remains visible.
Then ask an operator who did not build the workflow to recover one interrupted run using the provided information. Record where they need undocumented context. That exercise checks the operating design as well as the code. Use the result to improve the recovery view, ownership, and instructions before increasing the workflow’s reach.
- Verify the destination state, not only the completion message.
- Confirm repeated delivery does not create unintended duplicate work.
- Ensure a changed proposal cannot inherit an unrelated approval.
- Keep unknown outcomes separate from confirmed rejection or success.
- Record which person owns every state requiring intervention.