Incidents, retries and message waits
Diagnose a stalled process and recover without repeating completed work.
Identify the instance first
Keep the process instance key, job key and external business identifier in your logs. Use the instance key to inspect engine state; use the business identifier to reconcile external effects. Avoid logging secrets or full personal payloads.
Distinguish failure from uncertainty
| Situation | Next action |
|---|---|
| Invalid input | Correct the input or policy; blind retries will repeat the error. |
| Temporary external failure | Fail the job with retries set to job.retries - 1. It is offered again at once, with no delay, so wait in the worker if the service needs time. |
| External action succeeded, completion unknown | Reconcile first; avoid repeating the action. A 404 on complete means the job is gone. |
| Retries exhausted | The instance becomes INCIDENT. It cannot be resumed over REST today: fix the cause, then start a new instance. PATCH /api/v1/incidents/{id} only closes the incident record. |
Find stuck instances with GET /api/v1/process-instances?filter.state=INCIDENT.
Message waits are jobs
Priostack has no hosted route that publishes or correlates messages. A receive task waits as a job of type message:<messageName>, and a message catch event as event:message:<elementId>; a worker activates and completes that job to move the instance on. There is no correlation-key matching, so the worker must check that the instance it completes is the one the message is for.
To find an instance by your own identifier, store it as a variable and filter with GET /api/v1/process-instances?filter.variable=correlationId:<value>.
Jobs have no lock period
An activated job stays with its worker until it is completed or failed; nothing takes it back at its deadline. A business deadline belongs in the model, as a timer boundary event. The two solve different problems and need separate tests.
Job policies reference Incidents reference