A workflow engine exists for work that outlives the request that started it. This first part follows one automation from Friday to Thursday. At each failure, it asks a simple question: which fact did the system need to remember, and where was that fact stored?
That question is the key to the whole subject. A queue can tell a worker that there is work to try. It cannot, by itself, prove that a run started only once, remember a timer after the queue loses it, join an approval to the right run, or know whether an email provider acted before a timeout. Those facts need a durable home.
The example uses a common automation shape: long waits, outside providers, model calls, and people who may answer hours later. The design lesson also applies to onboarding, fulfilment, billing, data syncs, and agents that pause for human approval.
How to read the diagrams
- The engine and its records: runs, steps, and state changes
- Time: timers, waits, and wake-ups
- The outside world: model calls, email, and CRM writes
- People and signals: approvals and replies
- Failure: crashes, duplicates, and lost wake-ups
Read this page in four passes:
- Follow Maya's run without thinking about implementation.
- Replay the same run on workers and a queue, and name the missing fact at each failure.
- Compare event-history replay with a persisted workflow graph.
- Use the engine map and decision table to decide whether the problem needs durable execution.
The running example
A marketing team builds this automation, "Demo follow-up", in a visual editor and publishes it. Maya, who lives in Berlin, fills in the demo form on a Friday afternoon, and her run starts.
Demo follow-up · one run for one contact
- Fri 14:02
Trigger: demo form submitted
The form provider sends a webhook with Maya's details.
trigger event evt_7Hq2 · contact maya · zone Europe/Berlin
- Fri 14:02
Wait until 09:00 on the next weekday, in her time zone
Friday afternoon becomes Monday morning. For almost three days the run does nothing at all.
due 2026-10-05 07:00 UTC (09:00 in Berlin, summer time)
- Mon 09:00
Draft a follow-up email with a model
One model call, about twenty seconds, billed per token.
lane SLOW · effect key run_81f3/draft_email
- Mon 09:00
Ask Sam, the account owner, to approve the draft
The run parks. Sam approves in Slack at 11:40.
approval expires Tue 09:00
- Mon 11:40
Send the email
lane IO · effect key run_81f3/send_email
- Mon 11:40
Create a follow-up task in the CRM
lane IO · effect key run_81f3/create_crm_task
- Mon 11:40
Wait up to three days for a reply
Two ways forward: a reply arrives, or the deadline passes.
signal email.replied · deadline Thu 11:40
- Thu 11:40
No reply: remind Sam
lane IO · effect key run_81f3/notify_owner
The rest of the chapter uses a small vocabulary. An automation is the published graph, here "Demo follow-up". Each publish freezes a new version, and a run stays on the version it started with, which matters when the team edits an automation while runs are in progress. A run is one execution of it for one contact; Maya's run has the id run_81f3. A run is made of steps. When a step reaches outside the engine, to call a model, send an email, or write to a CRM, that call is an effect. A run waiting for time to pass is waiting on a timer, and a run waiting for something to happen, such as an approval or a reply, is waiting on a signal.
Nothing about this automation is unusual, and it is still hard to run correctly. It lasts five days, touches four outside systems, and waits on a clock and on two people. During those five days the code that runs it will be deployed many times, the machines under it will be replaced, and at least one message will arrive twice. A workflow engine is the part of the system whose job is to make all of that invisible to the person who built the automation.
A queue can deliver work; it cannot be the workflow record
A queue is a good way to distribute work, retry a failed attempt, and apply backpressure. It can absolutely be part of a correct workflow system. The weak design is narrower: each step is a job, the job enqueues its successor, delayed jobs are the only timers, and a status column is the only other record of progress. In that design, the queue has quietly become the workflow record even though it cannot commit atomically with the database or an outside provider.
The failures below are therefore not an argument against queues. They show the extra guarantees a queue-backed workflow engine must add.
The same run when jobs are the only coordinator
- Fri 14:02
The webhook arrives twice
The form provider's first delivery timed out, so it retried.
Without a unique trigger identity, two deliveries create two runs. Idempotency at run creation fixes this.
- Fri → Mon
The due time exists only as a delayed job
If that job is lost during a restore, flush, or damaged entry, no independent record says the run became due. Idempotency cannot recreate missing work.
- Mon 09:00
The draft activity is delivered again
The first worker misses its renewal but may still be running. A heartbeat reduces this risk and fencing rejects a stale result, but avoiding a second billed call requires provider idempotency. Without it, duplicate execution can be reduced, not eliminated.
- Mon 11:40
The approval is saved before the send job is published
A restart between the database write and queue publish leaves a valid approval with no wake-up. Idempotency prevents duplicates; it does not publish a command that never existed.
- Mon 11:40
The email provider accepts the request, but its reply is lost
The retry cannot tell whether the email exists. A provider idempotency key solves this when supported; otherwise the worker must reconcile or stop for review.
- Mon 11:41
The email result is stored before the CRM job is published
The email is safely recorded, but a crash can still strand the next step. The database must remember that the CRM step remains runnable, and an outbox or recovery scan must publish it.
Read each row as cause → missing fact → visible consequence:
- Duplicate webhook. The provider did not receive an acknowledgement in time, so it sent the same event again. A unique constraint on a stable trigger identity makes run creation idempotent. This one is solved by idempotency, as long as both deliveries carry or derive the same identity.
- Lost Monday wake-up. Redis contains the delayed job, but Postgres only
says the run is waiting. If the entry disappears, no durable record says
“this run became due at Monday 09:00.” Nothing wakes it again.
Redis can be configured for persistence and
noeviction, and a serious queue deployment should do that. The design problem remains: if the queue is the only copy of the due time, a queue restore, flush, or damaged entry has no independent backstop. Idempotency only recognises repeated work; it cannot detect work that disappeared. A timer row plus a due sweep supplies that missing detection. - Two workers receive one activity. Redelivery after a lost lease is expected queue behaviour. A fencing token prevents the older worker from committing stale state, but it cannot undo an outside call already made. Heartbeats make unnecessary redelivery less likely. If the model provider accepts an idempotency key, both attempts can share one effect identity. If it does not, the engine can commit only one result but cannot guarantee that the provider bills only one call.
- Approval saved, send never published. Updating the database and publishing a queue job are two separate writes. A restart between them leaves a true approval with no command to continue the run. An idempotency key on the send would make a repeated send safer, but nothing is being repeated here. A transactional outbox closes the database-to-queue gap.
- Email outcome unknown. The provider accepts the email and the response is lost before success is recorded. A stable idempotency key fixes this when the provider honours it. If the provider offers no such contract, the worker must reconcile by a client reference or park the run for a person; no queue setting can manufacture knowledge held only by the provider.
- Next step stranded. Even after the email result is safely stored, a
separate
enqueue CRMwrite can be lost. The run store must say the CRM step is runnable, and an outbox relay or recovery scan must keep trying to publish its wake-up.
There is also an operational gap. When Sam asks on Wednesday why Maya never got a reminder, the answer may be spread across a queue, a status column, and the logs of whichever workers handled each attempt. A durable run history is not required to enqueue work, but it is required to explain the run reliably.
- failure window
- Webhook delivered again
- does idempotency solve it?
- Yes, at run creation
- full repair
- unique trigger identity + database constraint
- failure window
- Delayed job disappears
- does idempotency solve it?
- No; there is no retry to deduplicate
- full repair
- durable timer row + due sweep
- failure window
- Lease expires while work continues
- does idempotency solve it?
- Only if the external service honours the same key
- full repair
- heartbeat + fence; accept possible duplicate cost otherwise
- failure window
- Approval saved before queue publish
- does idempotency solve it?
- No; the command is missing
- full repair
- transactional outbox
- failure window
- Provider response is lost
- does idempotency solve it?
- Yes when the provider supports it; otherwise reconcile
- full repair
- effect ledger + stable key + reconcile-or-park policy
- failure window
- Result saved before next-step publish
- does idempotency solve it?
- No; the wake-up is missing
- full repair
- runnable database state + outbox or recovery scan
- failure window
- Someone asks what happened
- does idempotency solve it?
- No
- full repair
- one durable run history
These failures cross three boundaries: trigger to database, database to queue, and worker to outside provider. A queue guarantees what its contract says about job delivery; it does not make those separate systems one transaction. Idempotency handles repeated intent. It does not recover missing intent, retain a due time, or reveal an ambiguous provider outcome when the provider offers no idempotency or lookup contract.
That is the recommendation in precise terms. Workers and a queue are often the right execution layer. A long-running workflow also needs durable state around them: constraints, timers, an outbox, leases and fencing, effect identities, reconciliation, and a run history. You can build those guarantees on Postgres and a queue. The next chapter does exactly that. You can also run a durable workflow system such as Temporal.
What Temporal does for this run
Temporal separates coordination from outside work. A workflow coordinates waits, branches, and signals. An activity performs fallible work such as a model call or email send. The Temporal service stores event history, timers, retry state, and signals.
The Monday draft on Temporal
- Fri 14:02Temporal
Record start and Monday timer
- Mon 09:00Temporal → Workflow worker
Timer fired; create workflow task
- Workflow worker → Temporal
Schedule draft activity
- Temporal → Activity worker
Deliver activity task
- Activity worker → Model
Generate draft
- Model → Activity worker
Return draft d_42
- Activity worker → Temporal
Record activity result
- Temporal → Workflow worker
Resume workflow with d_42
If a workflow worker stops, another reconstructs coordination state from the history. If an activity worker stops, Temporal retries according to the activity policy. The email activity still needs an idempotency key because a retry cannot know whether an earlier request reached the provider.
Two ways to remember where a run is
After a crash, every durable engine has to answer the same question: where was this run, and what has it already done? There are two established answers, and the choice between them shapes everything else.
Record and replay
In the first, the workflow is ordinary code:
export async function demoFollowUp(contact: Contact) {
await sleepUntil(nextWeekdayAt(9, contact.timeZone));
const draft = await draftFollowUp(contact); // model call
await requestApproval(contact.owner, draft); // waits on Sam
const message = await sendEmail(contact, draft);
await createCrmTask(contact, message);
const replied = await waitForReply(message, { days: 3 });
if (!replied) await notifyOwner(contact);
}The runtime records the result of every call that leaves the function (the timer, the model call, the approval) in an append-only history. If the worker running Maya's workflow dies on Monday at 11:40, another worker loads her history and runs the function again from the top. This time sleepUntil returns at once because the history says the timer already fired, and draftFollowUp returns the recorded draft without calling the model. Execution reaches sendEmail with exactly the state it had before the crash. This is how Temporal works.
The price is determinism. Replay only rebuilds the same state if the code makes the same decisions from the same history. In Temporal's TypeScript sandbox, Date.now() and Math.random() are made deterministic. In a Go workflow, however, branching on time.Now().Hour() reads the process clock: the first execution may take the morning path and a later replay the afternoon path. Anything that can differ between executions has to come from the SDK's deterministic APIs or through a recorded activity.
Walking a published graph
In the second, the workflow is data. The visual editor saves a graph of steps and edges, and publishing freezes it as a version. A run is a row that points at that version and records its position. Advancing a run is a function that takes the graph, the run's current state, and one event, and returns the new state plus a list of commands for the rest of the engine to carry out:
type RunEvent =
| { type: "timer.fired"; stepId: string; at: Date }
| { type: "step.completed"; stepId: string; output: unknown; at: Date }
| { type: "signal.received"; name: string; payload: unknown; at: Date };
type Command =
| { type: "schedule"; stepId: string; lane: Lane }
| { type: "wake"; stepId: string; at: Date }
| { type: "park"; stepId: string; until: "approval" | "signal" }
| { type: "finish" };
// No clock, no I/O, no randomness: time arrives inside the event,
// so the same inputs always give the same state and commands.
declare function advance(
graph: PublishedGraph,
run: RunState,
event: RunEvent,
): { run: RunState; commands: Command[] };Nothing is replayed, so the people writing steps don't have to follow determinism rules. In exchange, the engine has to keep everything else itself: which worker holds which step, when each timer is due, which signals have arrived. Most of this chapter is about that. Here is Maya's run as the engine stores it, at three moments:
| id | automation | version | position | status | updated_at (UTC) |
|---|---|---|---|---|---|
| Fri 14:02:11 · run started, first step is a wait | |||||
| run_81f3 | demo-follow-up | 3 | wait_until_morning | waiting_timer | 2026-10-02 12:02:11 |
| Mon 09:00:00 · timer fired, draft step claimed | |||||
| run_81f3 | demo-follow-up | 3 | draft_email | running | 2026-10-05 07:00:00 |
| Mon 09:00:19 · draft done, approval requested | |||||
| run_81f3 | demo-follow-up | 3 | approve_draft | awaiting_approval | 2026-10-05 07:00:19 |
Visual automation builders fit this second model naturally, because the user has already drawn the graph. An interpreter can keep its graph walk pure, which makes every branch of a published graph testable without a database or queue.
Record and replay
- State lives in
- An append-only history of every recorded result
- After a crash
- Run the code again; recorded calls return their stored results
- Rules for workflow code
- Deterministic: no clock, randomness, or I/O outside recorded calls
- Changing a live workflow
- New code must still agree with existing histories
- Examples
- Temporal, Cadence, Azure Durable Functions
Walk a published graph
- State lives in
- Rows: the run, its steps, timers, signals, and effects
- After a crash
- Load the rows; the next worker continues from the recorded position
- Rules for workflow code
- The walk must be pure; steps reach the world only through effects
- Changing a live workflow
- Runs stay on the version they started with
- Examples
- AWS Step Functions, most visual automation builders
A map of the engine
Before going into each mechanism, here is the whole engine on one page. The numbers trace Maya's run through it.
Triggers
Webhooks and forms1
form.submitted · email.replied
API and schedules
manual starts · recurring automations
Run store · Postgres
runs26
where each run is
id · version · position · status
steps · effects45
who holds what, what happened
attempt · lease_until · effect key · result
timers3
what is due when
due_at · rule · zone
signals · approvals7
what has arrived
name · correlation key · consumed_at
outbox2
changes to announce
topic · payload · sent_at
Wake-up paths
Outbox relay2
publishes committed changes to the queue
Delayed jobs3
on-time wake-ups for timers
Due sweep3
backstop that finds anything the queue lost
FOR UPDATE SKIP LOCKED · per tenant
Lanes · one queue each
INLINE
branches · transforms
IO5
email · CRM · Slack
SLOW4
model calls · metered
HUMAN7
approvals · parked
Workers
Claim → advance → perform456
lease and fencing token · pure walk · effects under keys
Outside world
Model providers4
Email provider · CRM5
Slack7
- The form's webhook arrives. A unique trigger identity makes both deliveries converge on one logical run.
- In the same transaction as the run, it writes an outbox row, which a relay publishes to the queue after the commit.
- The first step is a wait. The engine writes a timer row and schedules a delayed job for Monday at 07:00 UTC. If the job is lost, the due sweep finds the timer anyway.
- On Monday a worker claims the draft step from the SLOW lane under a lease and calls the model.
- After Sam approves, workers on the IO lane claim the email and CRM steps and perform each effect under its own key.
- Every result is written back to the run in one transaction, and the pure walk decides what comes next.
- Approvals and replies arrive as signals. They are stored first and matched to the run when it reaches the step that waits for them.
A useful test of this design is to disable delayed-job scheduling in a staging environment while leaving ordinary queue delivery and the due sweep running. Timers should wake no later than the sweep interval. If a run remains stranded, some due-state or recovery fact still lives only in the delayed-job system.
Continue from the map to the implementation
The map above is the conceptual model. The next two parts answer different questions:
Part 2
Build the guarantees
Trace the inbox, unique start key, transactional outbox, leases, fencing, effect ledger, timers, signals, and run log through concrete rows and failure windows.
Part 3
Use Temporal
See how event history and replay replace most engine machinery, which idempotency work remains yours, and how sagas and workflow versioning fit.
Continue with Part 2: Building a Durable Workflow Engine, or go directly to Part 3: Durable Workflows with Temporal.
Which to choose
For this example the recommendation is clear. A queue without a durable workflow record is not suitable, because queue delivery alone cannot keep a run's state, timers, and side effects consistent across crashes. A Postgres-and-queue design can provide those guarantees, but by then the database schema, relay, leases, fences, timer sweep, signal inbox, and reconciliation logic form a small workflow engine. Between building that engine and running one, Temporal is the better default for workflows like this one: runs that last days, wait on people, call several unreliable providers, and change while they are in flight. Building your own is reasonable in narrower cases: the workflows are short and few, the team can't operate new infrastructure, or you already have an engine that interprets published graphs and need only some of these guarantees.
| Build on Postgres and a queue | Run Temporal | |
|---|---|---|
| You write | Claims, leases, timers, the sweep, signal delivery, recovery, a run log | Workflows and activities |
| You operate | The Postgres and Redis you already run | A Temporal cluster, or Temporal Cloud |
| Workflow code | Anything, as long as the walk stays pure | Deterministic, versioned, replay-tested |
| Languages | The engine's own | Any SDK language, split across task queues |
| Worth it when | Workflows are simple, the team wants no new infrastructure, and it can afford to own the correctness work | Runs are long, step types are many, several languages are involved, and the team would rather not own that work |
Whichever side you choose, define the runtime boundary first: wake a run, give an external effect a stable identity, deliver a signal, and cancel. That keeps the workflow's meaning separate from the execution technology and makes a future runtime change less invasive.
A checklist before real traffic
Starting runs
- Every run has a start key from its trigger, enforced by a unique constraint
- The run and its first wake-up are written in one transaction (outbox)
- Consumers deduplicate each stable command id or commit its state transition idempotently
Executing steps
- Claims take a lease; long steps heartbeat
- Every write back is fenced by the attempt number
- All time comparisons use the database clock
- Shutdown releases leases instead of letting them lapse
Side effects
- Every effect is claimed under a stable key before it runs
- Claimed-but-unsettled effects are treated as possibly done
- Each provider has a plan: idempotency key, reconcile, or park
- Consent and suppression are checked at send time
Time
- Timers store the rule and zone, not only the instant
- Both daylight-saving edge cases are decided and tested
- A delayed job is the fast path; the sweep is the backstop
- The sweep is fair between tenants and runs under a lease
- Schedules built by people are spread out
People and signals
- Every signal has a correlation key
- Signals are stored first and consumed once
- Approvals check permission at click time, expire, and escalate
Change and failure
- Published versions are immutable; runs keep theirs
- Step ids are stable across edits
- Authorisation errors park the run instead of retrying
- Runs with no way forward are queried for and alerted on
- Redis never evicts; wake-up job ids include the attempt
When you don't need any of this
A background job with one step, no waits, no approvals, and modest retry needs is well served by a queue and an idempotency key. Durable execution starts earning its cost when a run has to outlive the process that started it, as Maya's run did from Friday afternoon to Thursday.