Skip to content

Architecture Article

Durable Workflows with Temporal

Follow the same multi-day automation on Temporal to learn replay, activities, sagas, versioning, and the guarantees applications still own.

Published 12 May 2026Updated 1 Oct 20269 min read
Temporal · Durable Execution · Sagas · Idempotency · Workflow Automation
A Temporal event history replaying workflow state before scheduling an idempotent activity in the outside world.
On this page (6)

Temporal is a runtime for code that may pause for minutes or months and still continue after workers restart. It records a workflow's history, then replays that history to rebuild the program's local state. Work that reaches a database, model, payment service, or email provider runs separately as an activity.

The first part of this series explains why a queue alone does not retain every fact a multi-day run needs. The implementation chapter builds the missing guarantees with database rows, leases, timers, and recovery loops. This page runs the same example on Temporal. Reading those earlier parts helps, but the terms and example needed here are introduced again.

Temporal removes a large amount of coordination code, but distributed-systems constraints remain. Workflow code must replay deterministically. Activities may run more than once. An email or payment still needs a stable effect identity. Permissions still need checking when an action happens. This page separates what the runtime owns from what the application must continue to own.

How to read the diagrams

  • Workflow code and event history
  • Durable timers, replay, and runtime-owned recovery
  • Activities and outside effects
  • Signals and human decisions
  • Replay, versioning, and operational failure boundaries

Where Temporal comes from

Everything in the mechanics chapter can be built on Postgres and a queue, and some teams do. For workloads like this one, the better default is to run an engine that already solves it, and the best-known one is Temporal. It helps to know where it came from, because its design is the record-and-replay model, refined over more than a decade by the same two people.

From sagas to Temporal

  1. 1987

    "Sagas" (Garcia-Molina and Salem)

    A paper on long-lived transactions: split them into steps that each commit on their own, and undo completed steps with compensating actions when a later one fails.

  2. 2012

    Amazon Simple Workflow Service

    Maxim Fateev and Samar Abbas, who met at Amazon, work on SWF, a hosted service for coordinating long-running work. Its programming model lets developers write the coordination as code.

  3. Mid-2010s

    Durable Task Framework at Microsoft

    Abbas, now at Microsoft, builds an open-source orchestration framework that later becomes the basis of Azure Durable Functions.

  4. 2015

    Cadence at Uber

    The two reunite at Uber and build Cadence, an open-source workflow engine drawing on SWF and the Durable Task Framework.

  5. 2019

    Temporal

    They leave Uber, found Temporal Technologies in October 2019, and fork Cadence to start Temporal. Version 1.0 ships in 2020.

The Temporal server is written in Go. There are official SDKs for Go, Java, TypeScript, Python, .NET, PHP, and Ruby, and several of them (TypeScript, Python, .NET, and Ruby) are thin layers over a shared core written in Rust, so the difficult parts of replay are implemented once rather than once per language. That is where Rust enters the picture; the pattern Temporal implements is older than any of these languages.

The idea underneath it is event sourcing applied to a running program. Event sourcing stores a system's state as the sequence of events that produced it and rebuilds the state by replaying them. Temporal stores a workflow's history of events (timer started, activity completed, signal received) and rebuilds the workflow's local variables by replaying its code against that history. The determinism rules exist because the code is the function that turns the history back into state.

Sagas, and how Temporal relates to them

Temporal is often described as a saga engine, and the relationship is close enough to be worth getting exactly right.

The 1987 paper started from a practical problem. A transaction that lasts hours or days, such as booking a trip or fulfilling an order, cannot hold database locks for its whole duration without blocking everyone else. The authors proposed splitting such a transaction into a sequence of smaller transactions, each of which commits on its own, and pairing each one with a compensating transaction that semantically undoes it. If a later step fails, the saga runs the compensations for the completed steps in reverse order. The paper called this backward recovery, and also described forward recovery: retrying from a saved point instead of undoing.

A saga that fails at its third step

  1. T1 · Reserve stock

    Commits. Compensation registered: release the stock.

  2. T2 · Charge the card

    Commits. Compensation registered: refund the charge.

  3. T3 · Create the shipment

    Fails permanently: the address is undeliverable.

    Backward recovery starts. Compensations run newest first.

  4. C2 · Refund the charge

    A new transaction that reverses the effect of T2. It does not erase T2; the customer's statement shows both.

  5. C1 · Release the stock

    The saga ends in a consistent state: nothing reserved, nothing charged.

Sagas come in two styles. In a choreographed saga there is no coordinator: each service reacts to the previous service's events and emits its own. It is loosely coupled and hard to see, because no single place knows how far a given saga has got or what it would take to undo it. In an orchestrated saga, one coordinator tells each participant what to do and keeps track of the saga's state.

Here is the relationship. An orchestrated saga needs a coordinator that cannot lose its place. If the coordinator crashes after T2 and before starting C2, someone has to know that the card was charged and a refund is owed. That coordinator is a durable workflow engine: it needs the durable state, timers, retries, and leases from the mechanics chapter to make the saga's promise hold. Temporal is not a saga library. It is a durable execution runtime, and on top of it a saga becomes ordinary code:

Sketch: an orchestrated saga as a Temporal workflow (TypeScript SDK)
import { CancellationScope, proxyActivities } from "@temporalio/workflow";
import type * as activities from "./activities";
 
const {
  reserveStock, releaseStockIfReserved,
  chargeCard, refundIfCharged, createShipment,
} =
  proxyActivities<typeof activities>({
    startToCloseTimeout: "1 minute",
    retry: { maximumAttempts: 5 },   // the default is to retry forever
  });
 
export async function placeOrder(order: Order): Promise<void> {
  const compensations: Array<() => Promise<void>> = [];
  try {
    // Register a conditional, idempotent compensation before each attempt.
    // It first checks whether the forward effect exists, then converges on undone.
    compensations.unshift(() => releaseStockIfReserved(order.id));
    await reserveStock(order);
 
    compensations.unshift(() => refundIfCharged(order.paymentId));
    await chargeCard(order);
 
    await createShipment(order);
  } catch (err) {
    await CancellationScope.nonCancellable(async () => {
      for (const compensate of compensations) await compensate();
    });
    throw err;
  }
}

Because the workflow is durable, the list of compensations survives crashes like any other local variable. Registering each compensation before its forward attempt covers the case where the remote step succeeds but its response is lost. That compensation must be conditional and idempotent: “release if this order reserved stock” or “refund if this payment was charged.” Temporal's default activity retry policy retries until a non-retryable error or an attempt limit is reached. Compensations are activities too, and the non-cancellable scope prevents cancellation from interrupting the refund sequence.

Sagas also give up something that transactions have: isolation. Between T2 and C2, the rest of the system can see a charge that is about to be refunded. The usual countermeasures are to mark in-progress records (an order in a "pending" state that other processes treat carefully) and to design steps whose effects commute. And compensation is semantic, not a rollback. You can refund a charge, but you cannot unsend an email; the best you can do is send a correction. That is why Maya's run has no compensations at all. Almost every step in a marketing automation is impossible to undo, so it relies on forward recovery (retry until the step succeeds, or park it for a person) rather than backward recovery.

The same automation on Temporal

Here is Maya's automation written as a Temporal workflow. Every hard part in the mechanics chapter is still present, but most of it has become a single line:

Sketch: the demo follow-up as a Temporal workflow (TypeScript SDK)
import {
  condition, defineSignal, proxyActivities, setHandler, sleep,
} from "@temporalio/workflow";
import type * as activities from "./activities";
 
// Each lane becomes activity options (and, if needed, its own task queue).
const { nextSendTime, requestApproval, notifyOwner } = proxyActivities<typeof activities>({
  startToCloseTimeout: "30 seconds",
});
const { draftFollowUp } = proxyActivities<typeof activities>({
  startToCloseTimeout: "2 minutes",
  heartbeatTimeout: "20 seconds",
  retry: { maximumAttempts: 2 },                 // every attempt is billed
});
const { sendEmail, createCrmTask } = proxyActivities<typeof activities>({
  startToCloseTimeout: "30 seconds",
  retry: { maximumInterval: "5 minutes" },       // 429s back off, then retry
});
 
export const approved = defineSignal<[{ by: string; digest: string }]>("approved");
export const replied = defineSignal("replied");
 
export async function demoFollowUp(contact: Contact): Promise<void> {
  let approval: { by: string; digest: string } | undefined;
  let gotReply = false;
  setHandler(approved, (decision) => { approval = decision; });
  setHandler(replied, () => { gotReply = true; });     // works even if early
 
  const wakeAt = await nextSendTime(contact);    // zone rules, recorded
  await sleep(Math.max(0, wakeAt - Date.now())); // durable timer, no worker held
 
  const draft = await draftFollowUp(contact);
  const { digest } = await requestApproval(contact.owner, draft);
  if (!(await condition(() => approval?.digest === digest, "24 hours"))) {
    return;                                      // nobody approved in time
  }
 
  const message = await sendEmail(contact, draft);  // effect key inside
  await createCrmTask(contact, message);
 
  if (!(await condition(() => gotReply, "72 hours"))) {
    await notifyOwner(contact);
  }
}
Sketch: starting it once per form submission
await client.workflow.start(demoFollowUp, {
  taskQueue: "automations",
  workflowId: `demo-follow-up:${event.id}`,    // same id while this run is open
  args: [contact],
});

Read it against the mechanics chapter. The weekend wait is sleep, and no worker is held while it runs. Activity timeouts and heartbeats replace lease-renewal code, but they do not fence an outside provider call: a timed-out attempt can still finish remotely. The approval and reply are signals, and condition with a timeout is the race between “a person acted” and “the deadline passed.” The early reply is not a special case, because its handler sets gotReply whenever it arrives.

The workflow id collapses duplicate starts while that workflow is open. A duplicate that arrives after it closes depends on the configured workflow-id reuse policy, so a long duplicate window may still need a durable trigger receipt. Temporal removes the workflow's run store, timer sweep, and signal table from application code. An outbox may remain at the application-database to workflow-start boundary when those two writes must recover together.

What is still there is what the section on side effects said would be. sendEmail and createCrmTask are activities. An activity attempt may finish externally and then time out before Temporal records completion; with retries enabled, the activity can therefore run again. Each one claims its effect key, passes it to the provider where it can, and reconciles before retrying. The consent check belongs inside sendEmail, and the check for a revoked mailbox inside every activity that uses a connection.

At platform scale the natural shape is one workflow per contact per automation. A campaign that enrols 50,000 contacts starts 50,000 workflows; start them at a controlled rate, for example from a parent workflow that starts children in batches and continues-as-new between batches to keep its own history small. Recurring automations ("every Monday, email everyone whose trial ends this week") map onto Temporal schedules.

What Temporal takes off your hands

Mapping the custom engine onto Temporal shows which problems disappear and which stay yours:

custom engine → Temporal
in a custom engine
Start key and unique constraint
on Temporal
Workflow id: a second start with the id of a running workflow is rejected
who owns it
Temporal
in a custom engine
Run row and pure walk
on Temporal
Event history, replayed through deterministic workflow code
who owns it
Temporal, with rules
in a custom engine
Lease and heartbeat bookkeeping
on Temporal
Activity timeouts and heartbeats
who owns it
Temporal, configured by you
in a custom engine
Fencing an outside write
on Temporal
A timed-out attempt can still reach the provider
who owns it
You and the provider
in a custom engine
Delayed jobs and the due sweep
on Temporal
Durable timers
who owns it
Temporal
in a custom engine
Signals table
on Temporal
Signals, recorded in history even before the workflow waits for them
who owns it
Temporal
in a custom engine
Lanes and concurrency
on Temporal
Task queues and worker concurrency limits
who owns it
Temporal, configured by you
in a custom engine
Run log
on Temporal
Event history and the web UI
who owns it
Temporal
in a custom engine
Effect keys and reconciliation
on Temporal
Still needed: an activity can be attempted again after an ambiguous timeout
who owns it
You
in a custom engine
Fairness and the Monday 09:00 spread
on Temporal
Still needed
who owns it
You
in a custom engine
Consent and credential checks
on Temporal
Still needed, inside activities
who owns it
You
in a custom engine
Versions of an automation
on Temporal
Patching or worker versioning, plus replay tests
who owns it
Temporal, with rules

The rules that come with the replay model are specific to each SDK:

  • In Go, workflow code uses workflow.Now for the time, workflow.Sleep or workflow.NewTimer to wait, and workflow.SideEffect for random values. Native goroutines, channels, and select are replaced by their workflow equivalents, and ranging over a map is not deterministic, so sort the keys first.
  • In TypeScript, workflow code runs in a sandbox where Date.now() and Math.random() are made deterministic and the SDK's sleep creates a durable timer. Anything that does I/O belongs in an activity.
  • In Python, workflow code uses workflow.now(), workflow.random(), and workflow.uuid4(), and asyncio.sleep becomes a durable timer. Imports with side effects are restricted by the sandbox, and I/O belongs in activities.

Changing a workflow while executions are running is the Temporal version of the version-3-to-version-4 problem, and it bites in a quieter way. A harmless-looking change passes every test, ships, and then breaks executions that started on the old code, because replaying their histories now makes different decisions. The defences are to guard changes with the SDK's versioning calls (workflow.GetVersion in Go, patched in TypeScript and Python) or to pin running workflows to the worker build they started on, and to replay a sample of recent production histories against the new code in CI before every deploy. The Go, TypeScript, and Python SDK testing guides each document history replay for this check.

Activity timeouts are the next thing to get right. There are four: schedule-to-start (how long a task may wait in the queue), start-to-close (one attempt), schedule-to-close (all attempts together), and heartbeat. Temporal requires either start-to-close or schedule-to-close. The one people leave out is the heartbeat timeout on long activities: without it, a worker that dies ten seconds into a thirty-minute activity is only noticed when the thirty minutes are up.

Finally, histories have limits. Temporal's current programming-model limit is 51,200 events or 50 MB for one workflow execution, with warnings much earlier; Temporal Cloud also caps one request payload at 2 MB. Large results should be stored elsewhere and passed by reference, and long-lived workflows should periodically continue-as-new, which starts a fresh history carrying only the state that matters. These are service limits to verify again when designing a real deployment, not sizing targets. A visual automation builder on Temporal usually becomes one interpreter workflow that walks the published graph deterministically, so the published-graph model doesn't go away either.

What large teams learned running this

Durable workflows are no longer unusual infrastructure. Several companies have published what running them at scale taught them, and their lessons line up with the same boundaries: fairness, idempotency, versioning, and history size.

Uber · Cadence

Noisy neighbours

Scale
Over 12 billion executions and 270 billion actions a month, across 1,000+ services (2023)
Lesson
Month-end batch jobs used to consume whole clusters and delay interactive workflows, until task processing was made fair per tenant

Netflix · Temporal

Idempotent by force

Result
Deployments failing from transient cloud errors fell from 4% to 0.0001%
Lesson
Moving each step into an activity forced the team to make it idempotent; changing a function signature can break long-running workflows

Airbnb · Skipper

A library, not a cluster

Scale
Peaks of 10,000 workflows per second on DynamoDB, 15+ use cases
Lesson
Chose an embedded engine so a cluster outage couldn't stop every service; still calls workflow evolution 'the biggest friction point'

Uber runs Cadence, Temporal's predecessor, at a scale few systems reach, and its hardest problem was fairness rather than durability: "bursty traffic or workflow tasks from one customer usually consumes all the resources in a cluster … especially common and undesirable when batch processing jobs were triggered during month-ends" (Uber, 2021). That is the same problem as the per-tenant due sweep, one level down.

Netflix moved the steps of its deployment system into Temporal activities and reported that the share of deployments failing because of transient cloud errors "dropped from 4% to 0.0001%." The team also wrote that moving operations into activities forced the logic to become idempotent and warned that workflow changes can break long-running executions. A separate Netflix team building the Maestro workflow engine described the stale-worker race after garbage-collection pauses and network interruptions (Netflix, 2025).

Airbnb made the opposite platform choice from most: Skipper runs as a library inside each service rather than as a cluster, because "an orchestration cluster outage would mean every dependent service would lose the ability to start or advance workflows." It accepts the same trade-off as the side-effects section, "actions may execute more than once in edge cases (crash after execution but before checkpoint). Actions should be idempotent", and names the same pain: "workflow evolution remains the biggest friction point."

Versioning appears in almost every account. Uber notes that workflow changes are usually "only backward-compatible, but not forward-compatible", which makes rollbacks unsafe (Cadence, 2025). Microsoft warns that breaking changes to Durable Functions orchestrations can leave them "stuck indefinitely in a Running status" (Microsoft).

The platforms also have hard limits to understand before a design depends on them:

documented limits
platform
Temporal
limit
Event history of 51,200 events or 50 MB per workflow, with a warning at 10,240 events or 10 MB
what happens at the limit
the workflow is terminated; continue-as-new before it gets close
platform
Temporal Cloud
limit
2 MB per request payload; a 500-actions-per-second namespace floor under current on-demand capacity
what happens at the limit
large payloads are rejected; sustained throughput above the namespace limit is throttled
platform
Self-hosted Temporal
limit
The number of history shards is fixed when the cluster is created
what happens at the limit
can't be changed later: size it for growth
platform
AWS Step Functions (Standard)
limit
25,000 events per execution history; 256 KiB input or output; up to one year per execution
what happens at the limit
the execution fails
A dated planning snapshot from each platform's documentation, rechecked 1 October 2026. Account and deployment settings can differ; verify current limits before sizing.

The first two rows come from Temporal's current Cloud and programming-model limits; the self-hosted row comes from the numHistoryShards configuration reference. The table is here to expose design constraints, not to replace the provider's live limits page.

Datadog's engineers, who have run many self-hosted Temporal clusters, suggest designing workflows to stay well under the history limit (under 10,000 events), to make no more than about one state change per second per workflow, and to finish or continue-as-new within a day (Datadog, Replay 2026). A per-contact automation like Maya's fits those comfortably. A workflow that loops over every contact in a campaign does not, which is why fan-out belongs in child workflows started in batches.

The practical decision

Temporal is the stronger default when runs last for days, wait on people, cross several unreliable providers, or must survive deployments while they are in flight. A smaller Postgres-and-queue engine can still be sensible when the workflow set is narrow and the team is prepared to own every invariant in the mechanics chapter.

The dividing line is ownership. Temporal owns durable coordination; it does not own the correctness of the business effect. Activities still need stable identities, reconciliation, current permission checks, and careful payload boundaries.