Skip to content

Architecture Article

Durable Workflow Engines: A Queue Is Not the Source of Truth

Follow one five-day automation to see what queues, idempotency, and durable workflow engines each solve when processes crash.

Published 5 Mar 2026Updated 1 Oct 202611 min read
Distributed Systems · Workflow Automation · Durable Execution · Temporal · Postgres
One five-day automation run crossing durable waits, an AI activity, a human approval, external effects, and recovery after failure.
On this page (11)

A workflow engine exists for work that outlives the request that started it. This first part follows one automation from Friday to Thursday. At each failure, it asks a simple question: which fact did the system need to remember, and where was that fact stored?

That question is the key to the whole subject. A queue can tell a worker that there is work to try. It cannot, by itself, prove that a run started only once, remember a timer after the queue loses it, join an approval to the right run, or know whether an email provider acted before a timeout. Those facts need a durable home.

The example uses a common automation shape: long waits, outside providers, model calls, and people who may answer hours later. The design lesson also applies to onboarding, fulfilment, billing, data syncs, and agents that pause for human approval.

How to read the diagrams

  • The engine and its records: runs, steps, and state changes
  • Time: timers, waits, and wake-ups
  • The outside world: model calls, email, and CRM writes
  • People and signals: approvals and replies
  • Failure: crashes, duplicates, and lost wake-ups

Read this page in four passes:

  1. Follow Maya's run without thinking about implementation.
  2. Replay the same run on workers and a queue, and name the missing fact at each failure.
  3. Compare event-history replay with a persisted workflow graph.
  4. Use the engine map and decision table to decide whether the problem needs durable execution.

The running example

A marketing team builds this automation, "Demo follow-up", in a visual editor and publishes it. Maya, who lives in Berlin, fills in the demo form on a Friday afternoon, and her run starts.

Demo follow-up · one run for one contact

  1. Fri 14:02

    Trigger: demo form submitted

    The form provider sends a webhook with Maya's details.

    trigger event evt_7Hq2 · contact maya · zone Europe/Berlin

  2. Fri 14:02

    Wait until 09:00 on the next weekday, in her time zone

    Friday afternoon becomes Monday morning. For almost three days the run does nothing at all.

    due 2026-10-05 07:00 UTC (09:00 in Berlin, summer time)

  3. Mon 09:00

    Draft a follow-up email with a model

    One model call, about twenty seconds, billed per token.

    lane SLOW · effect key run_81f3/draft_email

  4. Mon 09:00

    Ask Sam, the account owner, to approve the draft

    The run parks. Sam approves in Slack at 11:40.

    approval expires Tue 09:00

  5. Mon 11:40

    Send the email

    lane IO · effect key run_81f3/send_email

  6. Mon 11:40

    Create a follow-up task in the CRM

    lane IO · effect key run_81f3/create_crm_task

  7. Mon 11:40

    Wait up to three days for a reply

    Two ways forward: a reply arrives, or the deadline passes.

    signal email.replied · deadline Thu 11:40

  8. Thu 11:40

    No reply: remind Sam

    lane IO · effect key run_81f3/notify_owner

The rest of the chapter uses a small vocabulary. An automation is the published graph, here "Demo follow-up". Each publish freezes a new version, and a run stays on the version it started with, which matters when the team edits an automation while runs are in progress. A run is one execution of it for one contact; Maya's run has the id run_81f3. A run is made of steps. When a step reaches outside the engine, to call a model, send an email, or write to a CRM, that call is an effect. A run waiting for time to pass is waiting on a timer, and a run waiting for something to happen, such as an approval or a reply, is waiting on a signal.

Nothing about this automation is unusual, and it is still hard to run correctly. It lasts five days, touches four outside systems, and waits on a clock and on two people. During those five days the code that runs it will be deployed many times, the machines under it will be replaced, and at least one message will arrive twice. A workflow engine is the part of the system whose job is to make all of that invisible to the person who built the automation.

A queue can deliver work; it cannot be the workflow record

A queue is a good way to distribute work, retry a failed attempt, and apply backpressure. It can absolutely be part of a correct workflow system. The weak design is narrower: each step is a job, the job enqueues its successor, delayed jobs are the only timers, and a status column is the only other record of progress. In that design, the queue has quietly become the workflow record even though it cannot commit atomically with the database or an outside provider.

The failures below are therefore not an argument against queues. They show the extra guarantees a queue-backed workflow engine must add.

The same run when jobs are the only coordinator

  1. Fri 14:02

    The webhook arrives twice

    The form provider's first delivery timed out, so it retried.

    Without a unique trigger identity, two deliveries create two runs. Idempotency at run creation fixes this.

  2. Fri → Mon

    The due time exists only as a delayed job

    If that job is lost during a restore, flush, or damaged entry, no independent record says the run became due. Idempotency cannot recreate missing work.

  3. Mon 09:00

    The draft activity is delivered again

    The first worker misses its renewal but may still be running. A heartbeat reduces this risk and fencing rejects a stale result, but avoiding a second billed call requires provider idempotency. Without it, duplicate execution can be reduced, not eliminated.

  4. Mon 11:40

    The approval is saved before the send job is published

    A restart between the database write and queue publish leaves a valid approval with no wake-up. Idempotency prevents duplicates; it does not publish a command that never existed.

  5. Mon 11:40

    The email provider accepts the request, but its reply is lost

    The retry cannot tell whether the email exists. A provider idempotency key solves this when supported; otherwise the worker must reconcile or stop for review.

  6. Mon 11:41

    The email result is stored before the CRM job is published

    The email is safely recorded, but a crash can still strand the next step. The database must remember that the CRM step remains runnable, and an outbox or recovery scan must publish it.

Read each row as cause → missing fact → visible consequence:

  1. Duplicate webhook. The provider did not receive an acknowledgement in time, so it sent the same event again. A unique constraint on a stable trigger identity makes run creation idempotent. This one is solved by idempotency, as long as both deliveries carry or derive the same identity.
  2. Lost Monday wake-up. Redis contains the delayed job, but Postgres only says the run is waiting. If the entry disappears, no durable record says “this run became due at Monday 09:00.” Nothing wakes it again. Redis can be configured for persistence and noeviction, and a serious queue deployment should do that. The design problem remains: if the queue is the only copy of the due time, a queue restore, flush, or damaged entry has no independent backstop. Idempotency only recognises repeated work; it cannot detect work that disappeared. A timer row plus a due sweep supplies that missing detection.
  3. Two workers receive one activity. Redelivery after a lost lease is expected queue behaviour. A fencing token prevents the older worker from committing stale state, but it cannot undo an outside call already made. Heartbeats make unnecessary redelivery less likely. If the model provider accepts an idempotency key, both attempts can share one effect identity. If it does not, the engine can commit only one result but cannot guarantee that the provider bills only one call.
  4. Approval saved, send never published. Updating the database and publishing a queue job are two separate writes. A restart between them leaves a true approval with no command to continue the run. An idempotency key on the send would make a repeated send safer, but nothing is being repeated here. A transactional outbox closes the database-to-queue gap.
  5. Email outcome unknown. The provider accepts the email and the response is lost before success is recorded. A stable idempotency key fixes this when the provider honours it. If the provider offers no such contract, the worker must reconcile by a client reference or park the run for a person; no queue setting can manufacture knowledge held only by the provider.
  6. Next step stranded. Even after the email result is safely stored, a separate enqueue CRM write can be lost. The run store must say the CRM step is runnable, and an outbox relay or recovery scan must keep trying to publish its wake-up.

There is also an operational gap. When Sam asks on Wednesday why Maya never got a reminder, the answer may be spread across a queue, a status column, and the logs of whichever workers handled each attempt. A durable run history is not required to enqueue work, but it is required to explain the run reliably.

which guarantee repairs each boundary
failure window
Webhook delivered again
does idempotency solve it?
Yes, at run creation
full repair
unique trigger identity + database constraint
failure window
Delayed job disappears
does idempotency solve it?
No; there is no retry to deduplicate
full repair
durable timer row + due sweep
failure window
Lease expires while work continues
does idempotency solve it?
Only if the external service honours the same key
full repair
heartbeat + fence; accept possible duplicate cost otherwise
failure window
Approval saved before queue publish
does idempotency solve it?
No; the command is missing
full repair
transactional outbox
failure window
Provider response is lost
does idempotency solve it?
Yes when the provider supports it; otherwise reconcile
full repair
effect ledger + stable key + reconcile-or-park policy
failure window
Result saved before next-step publish
does idempotency solve it?
No; the wake-up is missing
full repair
runnable database state + outbox or recovery scan
failure window
Someone asks what happened
does idempotency solve it?
No
full repair
one durable run history

These failures cross three boundaries: trigger to database, database to queue, and worker to outside provider. A queue guarantees what its contract says about job delivery; it does not make those separate systems one transaction. Idempotency handles repeated intent. It does not recover missing intent, retain a due time, or reveal an ambiguous provider outcome when the provider offers no idempotency or lookup contract.

That is the recommendation in precise terms. Workers and a queue are often the right execution layer. A long-running workflow also needs durable state around them: constraints, timers, an outbox, leases and fencing, effect identities, reconciliation, and a run history. You can build those guarantees on Postgres and a queue. The next chapter does exactly that. You can also run a durable workflow system such as Temporal.

What Temporal does for this run

Temporal separates coordination from outside work. A workflow coordinates waits, branches, and signals. An activity performs fallible work such as a model call or email send. The Temporal service stores event history, timers, retry state, and signals.

The Monday draft on Temporal

  1. Fri 14:02Temporal

    Record start and Monday timer

  2. Mon 09:00Temporal → Workflow worker

    Timer fired; create workflow task

  3. Workflow worker → Temporal

    Schedule draft activity

  4. Temporal → Activity worker

    Deliver activity task

  5. Activity worker → Model

    Generate draft

  6. Model → Activity worker

    Return draft d_42

  7. Activity worker → Temporal

    Record activity result

  8. Temporal → Workflow worker

    Resume workflow with d_42

Workers may stop; the service keeps the history and creates another task from recorded state.

If a workflow worker stops, another reconstructs coordination state from the history. If an activity worker stops, Temporal retries according to the activity policy. The email activity still needs an idempotency key because a retry cannot know whether an earlier request reached the provider.

Two ways to remember where a run is

After a crash, every durable engine has to answer the same question: where was this run, and what has it already done? There are two established answers, and the choice between them shapes everything else.

Record and replay

In the first, the workflow is ordinary code:

Sketch: the example as workflow code
export async function demoFollowUp(contact: Contact) {
  await sleepUntil(nextWeekdayAt(9, contact.timeZone));
  const draft = await draftFollowUp(contact);    // model call
  await requestApproval(contact.owner, draft);    // waits on Sam
  const message = await sendEmail(contact, draft);
  await createCrmTask(contact, message);
  const replied = await waitForReply(message, { days: 3 });
  if (!replied) await notifyOwner(contact);
}

The runtime records the result of every call that leaves the function (the timer, the model call, the approval) in an append-only history. If the worker running Maya's workflow dies on Monday at 11:40, another worker loads her history and runs the function again from the top. This time sleepUntil returns at once because the history says the timer already fired, and draftFollowUp returns the recorded draft without calling the model. Execution reaches sendEmail with exactly the state it had before the crash. This is how Temporal works.

The price is determinism. Replay only rebuilds the same state if the code makes the same decisions from the same history. In Temporal's TypeScript sandbox, Date.now() and Math.random() are made deterministic. In a Go workflow, however, branching on time.Now().Hour() reads the process clock: the first execution may take the morning path and a later replay the afternoon path. Anything that can differ between executions has to come from the SDK's deterministic APIs or through a recorded activity.

Walking a published graph

In the second, the workflow is data. The visual editor saves a graph of steps and edges, and publishing freezes it as a version. A run is a row that points at that version and records its position. Advancing a run is a function that takes the graph, the run's current state, and one event, and returns the new state plus a list of commands for the rest of the engine to carry out:

Sketch: advancing a run is a pure function
type RunEvent =
  | { type: "timer.fired"; stepId: string; at: Date }
  | { type: "step.completed"; stepId: string; output: unknown; at: Date }
  | { type: "signal.received"; name: string; payload: unknown; at: Date };
 
type Command =
  | { type: "schedule"; stepId: string; lane: Lane }
  | { type: "wake"; stepId: string; at: Date }
  | { type: "park"; stepId: string; until: "approval" | "signal" }
  | { type: "finish" };
 
// No clock, no I/O, no randomness: time arrives inside the event,
// so the same inputs always give the same state and commands.
declare function advance(
  graph: PublishedGraph,
  run: RunState,
  event: RunEvent,
): { run: RunState; commands: Command[] };

Nothing is replayed, so the people writing steps don't have to follow determinism rules. In exchange, the engine has to keep everything else itself: which worker holds which step, when each timer is due, which signals have arrived. Most of this chapter is about that. Here is Maya's run as the engine stores it, at three moments:

runsone row per run; highlighted cells changed at that moment
idautomationversionpositionstatusupdated_at (UTC)
Fri 14:02:11 · run started, first step is a wait
run_81f3demo-follow-up3wait_until_morningwaiting_timer2026-10-02 12:02:11
Mon 09:00:00 · timer fired, draft step claimed
run_81f3demo-follow-up3draft_emailrunning2026-10-05 07:00:00
Mon 09:00:19 · draft done, approval requested
run_81f3demo-follow-up3approve_draftawaiting_approval2026-10-05 07:00:19

Visual automation builders fit this second model naturally, because the user has already drawn the graph. An interpreter can keep its graph walk pure, which makes every branch of a published graph testable without a database or queue.

Record and replay

State lives in
An append-only history of every recorded result
After a crash
Run the code again; recorded calls return their stored results
Rules for workflow code
Deterministic: no clock, randomness, or I/O outside recorded calls
Changing a live workflow
New code must still agree with existing histories
Examples
Temporal, Cadence, Azure Durable Functions

Walk a published graph

State lives in
Rows: the run, its steps, timers, signals, and effects
After a crash
Load the rows; the next worker continues from the recorded position
Rules for workflow code
The walk must be pure; steps reach the world only through effects
Changing a live workflow
Runs stay on the version they started with
Examples
AWS Step Functions, most visual automation builders

A map of the engine

Before going into each mechanism, here is the whole engine on one page. The numbers trace Maya's run through it.

Triggers

Webhooks and forms1

form.submitted · email.replied

API and schedules

manual starts · recurring automations

one transaction per state change

Run store · Postgres

runs26

where each run is

id · version · position · status

steps · effects45

who holds what, what happened

attempt · lease_until · effect key · result

timers3

what is due when

due_at · rule · zone

signals · approvals7

what has arrived

name · correlation key · consumed_at

outbox2

changes to announce

topic · payload · sent_at

after commit

Wake-up paths

Outbox relay2

publishes committed changes to the queue

Delayed jobs3

on-time wake-ups for timers

Due sweep3

backstop that finds anything the queue lost

FOR UPDATE SKIP LOCKED · per tenant

Lanes · one queue each

INLINE

branches · transforms

IO5

email · CRM · Slack

SLOW4

model calls · metered

HUMAN7

approvals · parked

Workers

Claim → advance → perform456

lease and fencing token · pure walk · effects under keys

Outside world

Model providers4

Email provider · CRM5

Slack7

The database is the source of truth. The queue only wakes workers sooner than the sweep would.
  1. The form's webhook arrives. A unique trigger identity makes both deliveries converge on one logical run.
  2. In the same transaction as the run, it writes an outbox row, which a relay publishes to the queue after the commit.
  3. The first step is a wait. The engine writes a timer row and schedules a delayed job for Monday at 07:00 UTC. If the job is lost, the due sweep finds the timer anyway.
  4. On Monday a worker claims the draft step from the SLOW lane under a lease and calls the model.
  5. After Sam approves, workers on the IO lane claim the email and CRM steps and perform each effect under its own key.
  6. Every result is written back to the run in one transaction, and the pure walk decides what comes next.
  7. Approvals and replies arrive as signals. They are stored first and matched to the run when it reaches the step that waits for them.

A useful test of this design is to disable delayed-job scheduling in a staging environment while leaving ordinary queue delivery and the due sweep running. Timers should wake no later than the sweep interval. If a run remains stranded, some due-state or recovery fact still lives only in the delayed-job system.

Continue from the map to the implementation

The map above is the conceptual model. The next two parts answer different questions:

Part 2

Build the guarantees

Trace the inbox, unique start key, transactional outbox, leases, fencing, effect ledger, timers, signals, and run log through concrete rows and failure windows.

Part 3

Use Temporal

See how event history and replay replace most engine machinery, which idempotency work remains yours, and how sagas and workflow versioning fit.

Continue with Part 2: Building a Durable Workflow Engine, or go directly to Part 3: Durable Workflows with Temporal.

Which to choose

For this example the recommendation is clear. A queue without a durable workflow record is not suitable, because queue delivery alone cannot keep a run's state, timers, and side effects consistent across crashes. A Postgres-and-queue design can provide those guarantees, but by then the database schema, relay, leases, fences, timer sweep, signal inbox, and reconciliation logic form a small workflow engine. Between building that engine and running one, Temporal is the better default for workflows like this one: runs that last days, wait on people, call several unreliable providers, and change while they are in flight. Building your own is reasonable in narrower cases: the workflows are short and few, the team can't operate new infrastructure, or you already have an engine that interprets published graphs and need only some of these guarantees.

Build on Postgres and a queueRun Temporal
You writeClaims, leases, timers, the sweep, signal delivery, recovery, a run logWorkflows and activities
You operateThe Postgres and Redis you already runA Temporal cluster, or Temporal Cloud
Workflow codeAnything, as long as the walk stays pureDeterministic, versioned, replay-tested
LanguagesThe engine's ownAny SDK language, split across task queues
Worth it whenWorkflows are simple, the team wants no new infrastructure, and it can afford to own the correctness workRuns are long, step types are many, several languages are involved, and the team would rather not own that work

Whichever side you choose, define the runtime boundary first: wake a run, give an external effect a stable identity, deliver a signal, and cancel. That keeps the workflow's meaning separate from the execution technology and makes a future runtime change less invasive.

A checklist before real traffic

Starting runs

  • Every run has a start key from its trigger, enforced by a unique constraint
  • The run and its first wake-up are written in one transaction (outbox)
  • Consumers deduplicate each stable command id or commit its state transition idempotently

Executing steps

  • Claims take a lease; long steps heartbeat
  • Every write back is fenced by the attempt number
  • All time comparisons use the database clock
  • Shutdown releases leases instead of letting them lapse

Side effects

  • Every effect is claimed under a stable key before it runs
  • Claimed-but-unsettled effects are treated as possibly done
  • Each provider has a plan: idempotency key, reconcile, or park
  • Consent and suppression are checked at send time

Time

  • Timers store the rule and zone, not only the instant
  • Both daylight-saving edge cases are decided and tested
  • A delayed job is the fast path; the sweep is the backstop
  • The sweep is fair between tenants and runs under a lease
  • Schedules built by people are spread out

People and signals

  • Every signal has a correlation key
  • Signals are stored first and consumed once
  • Approvals check permission at click time, expire, and escalate

Change and failure

  • Published versions are immutable; runs keep theirs
  • Step ids are stable across edits
  • Authorisation errors park the run instead of retrying
  • Runs with no way forward are queried for and alerted on
  • Redis never evicts; wake-up job ids include the attempt

When you don't need any of this

A background job with one step, no waits, no approvals, and modest retry needs is well served by a queue and an idempotency key. Durable execution starts earning its cost when a run has to outlive the process that started it, as Maya's run did from Friday afternoon to Thursday.