Resolve incidents

Understand why a flow instance stopped, choose the right resolution for each kind of incident, and retry or resolve incidents in bulk.

An incident is a token that cannot go on without someone's help. The instance stays Running with an incident badge until the incident is resolved or the instance is cancelled. Flow Monitor offers only the resolutions that are safe for the kind of incident, and only those you are allowed to run: Plain resolutions such as Retry need pulse.edit. Override resolutions (marked "override" below) need pulse.override, ask for a Reason of at least 10 characters, and are audited with the state before and after. People without pulse.override do not see them. Find open incidents# The Needs attention line on every Flow Monitor tab counts the open incidents; select the figure to list them. The Incidents tab groups incidents with the same element, kind and message, with their Count, Versions, First seen and Last seen. Select a group to list its instances. Resolve an incident on an instance# Open the instance. Each open incident is shown above the diagram: what happened and at which element, the error, why, and what to do. Select one of the buttons on the incident. For an override, enter a Reason and confirm. Incident kinds and their resolutions# Incident Why Resolutions The job did not start (admission refused) The inputs are invalid, or no robot can run the job. Fix the inputs or add a robot, then Retry. The job did not start (subscription inactive) The tenant's subscription lapsed. Nothing: it retries by itself once the subscription is restored. The job did not start (start permission missing) The person who started the instance no longer has jobs.create in the workspace. Grant it, then Retry; or Re-home to me (override), which continues the instance as you. The job did not start (input schema mismatch) A robot task set to Latest at dispatch now has a different input schema. Publish a fixed version, then Migrate (override). The job failed The job ended with an error. Retry, only when no code ran or the task has a retry policy; Retry with variable edits (override); Take the error path (override). The robot stopped mid-run The robot stopped after the job started; its side effects are unknown. No plain retry. Mark succeeded with the outputs the step would have produced (override); Mark failed (override); Retry anyway — I verified there was no side effect (override). A condition could not be evaluated A value has the wrong type at a path in a condition. Set the variable (override), then Retry. Too many steps without a wait The model has a loop without a wait (over 1,000 steps). Fix the model; Cancel instance. The instance stopped (pinned version changed) The process now runs another package version than the flow version pinned. Publish a version against the current package version and select Migrate at the top of the instance page (override); or point the process back to the pinned version, then Retry. The instance stopped (error not caught) An error end event or a failed called flow has no error boundary event or event sub-process. Add a catch to the model, publish, and select Migrate at the top of the instance page (override); or Cancel instance. What each override does: Take the error path and Mark failed count the step as failed. An error boundary event on the step, or an error event sub-process around it, takes over; when nothing catches the error, the instance ends as Failed. Mark succeeded asks for the step's Outputs as a JSON object. For a robot task they are checked against the output schema of the package version the job ran. Retry with variable edits opens the variables, saves your changes and runs the step again. Retry anyway runs the step again from the start, whatever ran before. Use it only after you checked that the robot's side effects did not happen. Retry only what is safe to repeatA job that failed after the automation started may already have changed something, for example posted half an invoice. That is why Retry is offered only when no code ran or when the robot task declares a retry policy (and the author confirmed The automation is safe to run again). Give a robot task a retry policy in Flow Studio only when running it twice is harmless. Run a bulk operation# You can retry, cancel or migrate many instances in one operation. Every bulk operation runs in two steps: a dry run shows how many instances are affected, skipped and refused, then you confirm. Select the instances: On the Incidents tab, tick one or more incident groups and select Retry incidents, or On the Instances tab, tick running instances and select Retry incidents, Migrate or Cancel instances. Read the preview: affected, skipped (for example an incident that is not safe to retry) and refused, with the reasons. Confirm. Bulk migration also needs pulse.override, a Target version and a reason. Each instance is changed in its own transaction, at most 1,000 per operation. Follow the progress under Batch operations (the button at the top of Flow Monitor): Next steps# Migrate and change running instances Jobs

Find open incidents

Resolve an incident on an instance

Incident kinds and their resolutions

Run a bulk operation

Next steps