Temporal is a runtime for code that may pause for minutes or months and still continue after workers restart. It records a workflow's history, then replays that history to rebuild the program's local state. Work that reaches a database, model, payment service, or email provider runs separately as an activity.
The first part of this series explains why a queue alone does not retain every fact a multi-day run needs. The implementation chapter builds the missing guarantees with database rows, leases, timers, and recovery loops. This page runs the same example on Temporal. Reading those earlier parts helps, but the terms and example needed here are introduced again.
Temporal removes a large amount of coordination code, but distributed-systems constraints remain. Workflow code must replay deterministically. Activities may run more than once. An email or payment still needs a stable effect identity. Permissions still need checking when an action happens. This page separates what the runtime owns from what the application must continue to own.
How to read the diagrams
- Workflow code and event history
- Durable timers, replay, and runtime-owned recovery
- Activities and outside effects
- Signals and human decisions
- Replay, versioning, and operational failure boundaries
Where Temporal comes from
Everything in the mechanics chapter can be built on Postgres and a queue, and some teams do. For workloads like this one, the better default is to run an engine that already solves it, and the best-known one is Temporal. It helps to know where it came from, because its design is the record-and-replay model, refined over more than a decade by the same two people.
From sagas to Temporal
- 1987
"Sagas" (Garcia-Molina and Salem)
A paper on long-lived transactions: split them into steps that each commit on their own, and undo completed steps with compensating actions when a later one fails.
- 2012
Amazon Simple Workflow Service
Maxim Fateev and Samar Abbas, who met at Amazon, work on SWF, a hosted service for coordinating long-running work. Its programming model lets developers write the coordination as code.
- Mid-2010s
Durable Task Framework at Microsoft
Abbas, now at Microsoft, builds an open-source orchestration framework that later becomes the basis of Azure Durable Functions.
- 2015
Cadence at Uber
The two reunite at Uber and build Cadence, an open-source workflow engine drawing on SWF and the Durable Task Framework.
- 2019
Temporal
They leave Uber, found Temporal Technologies in October 2019, and fork Cadence to start Temporal. Version 1.0 ships in 2020.
The Temporal server is written in Go. There are official SDKs for Go, Java, TypeScript, Python, .NET, PHP, and Ruby, and several of them (TypeScript, Python, .NET, and Ruby) are thin layers over a shared core written in Rust, so the difficult parts of replay are implemented once rather than once per language. That is where Rust enters the picture; the pattern Temporal implements is older than any of these languages.
The idea underneath it is event sourcing applied to a running program. Event sourcing stores a system's state as the sequence of events that produced it and rebuilds the state by replaying them. Temporal stores a workflow's history of events (timer started, activity completed, signal received) and rebuilds the workflow's local variables by replaying its code against that history. The determinism rules exist because the code is the function that turns the history back into state.
Sagas, and how Temporal relates to them
Temporal is often described as a saga engine, and the relationship is close enough to be worth getting exactly right.
The 1987 paper started from a practical problem. A transaction that lasts hours or days, such as booking a trip or fulfilling an order, cannot hold database locks for its whole duration without blocking everyone else. The authors proposed splitting such a transaction into a sequence of smaller transactions, each of which commits on its own, and pairing each one with a compensating transaction that semantically undoes it. If a later step fails, the saga runs the compensations for the completed steps in reverse order. The paper called this backward recovery, and also described forward recovery: retrying from a saved point instead of undoing.
A saga that fails at its third step
T1 · Reserve stock
Commits. Compensation registered: release the stock.
T2 · Charge the card
Commits. Compensation registered: refund the charge.
T3 · Create the shipment
Fails permanently: the address is undeliverable.
Backward recovery starts. Compensations run newest first.
C2 · Refund the charge
A new transaction that reverses the effect of T2. It does not erase T2; the customer's statement shows both.
C1 · Release the stock
The saga ends in a consistent state: nothing reserved, nothing charged.
Sagas come in two styles. In a choreographed saga there is no coordinator: each service reacts to the previous service's events and emits its own. It is loosely coupled and hard to see, because no single place knows how far a given saga has got or what it would take to undo it. In an orchestrated saga, one coordinator tells each participant what to do and keeps track of the saga's state.
Here is the relationship. An orchestrated saga needs a coordinator that cannot lose its place. If the coordinator crashes after T2 and before starting C2, someone has to know that the card was charged and a refund is owed. That coordinator is a durable workflow engine: it needs the durable state, timers, retries, and leases from the mechanics chapter to make the saga's promise hold. Temporal is not a saga library. It is a durable execution runtime, and on top of it a saga becomes ordinary code:
import { CancellationScope, proxyActivities } from "@temporalio/workflow";
import type * as activities from "./activities";
const {
reserveStock, releaseStockIfReserved,
chargeCard, refundIfCharged, createShipment,
} =
proxyActivities<typeof activities>({
startToCloseTimeout: "1 minute",
retry: { maximumAttempts: 5 }, // the default is to retry forever
});
export async function placeOrder(order: Order): Promise<void> {
const compensations: Array<() => Promise<void>> = [];
try {
// Register a conditional, idempotent compensation before each attempt.
// It first checks whether the forward effect exists, then converges on undone.
compensations.unshift(() => releaseStockIfReserved(order.id));
await reserveStock(order);
compensations.unshift(() => refundIfCharged(order.paymentId));
await chargeCard(order);
await createShipment(order);
} catch (err) {
await CancellationScope.nonCancellable(async () => {
for (const compensate of compensations) await compensate();
});
throw err;
}
}Because the workflow is durable, the list of compensations survives crashes like any other local variable. Registering each compensation before its forward attempt covers the case where the remote step succeeds but its response is lost. That compensation must be conditional and idempotent: “release if this order reserved stock” or “refund if this payment was charged.” Temporal's default activity retry policy retries until a non-retryable error or an attempt limit is reached. Compensations are activities too, and the non-cancellable scope prevents cancellation from interrupting the refund sequence.
Sagas also give up something that transactions have: isolation. Between T2 and C2, the rest of the system can see a charge that is about to be refunded. The usual countermeasures are to mark in-progress records (an order in a "pending" state that other processes treat carefully) and to design steps whose effects commute. And compensation is semantic, not a rollback. You can refund a charge, but you cannot unsend an email; the best you can do is send a correction. That is why Maya's run has no compensations at all. Almost every step in a marketing automation is impossible to undo, so it relies on forward recovery (retry until the step succeeds, or park it for a person) rather than backward recovery.
The same automation on Temporal
Here is Maya's automation written as a Temporal workflow. Every hard part in the mechanics chapter is still present, but most of it has become a single line:
import {
condition, defineSignal, proxyActivities, setHandler, sleep,
} from "@temporalio/workflow";
import type * as activities from "./activities";
// Each lane becomes activity options (and, if needed, its own task queue).
const { nextSendTime, requestApproval, notifyOwner } = proxyActivities<typeof activities>({
startToCloseTimeout: "30 seconds",
});
const { draftFollowUp } = proxyActivities<typeof activities>({
startToCloseTimeout: "2 minutes",
heartbeatTimeout: "20 seconds",
retry: { maximumAttempts: 2 }, // every attempt is billed
});
const { sendEmail, createCrmTask } = proxyActivities<typeof activities>({
startToCloseTimeout: "30 seconds",
retry: { maximumInterval: "5 minutes" }, // 429s back off, then retry
});
export const approved = defineSignal<[{ by: string; digest: string }]>("approved");
export const replied = defineSignal("replied");
export async function demoFollowUp(contact: Contact): Promise<void> {
let approval: { by: string; digest: string } | undefined;
let gotReply = false;
setHandler(approved, (decision) => { approval = decision; });
setHandler(replied, () => { gotReply = true; }); // works even if early
const wakeAt = await nextSendTime(contact); // zone rules, recorded
await sleep(Math.max(0, wakeAt - Date.now())); // durable timer, no worker held
const draft = await draftFollowUp(contact);
const { digest } = await requestApproval(contact.owner, draft);
if (!(await condition(() => approval?.digest === digest, "24 hours"))) {
return; // nobody approved in time
}
const message = await sendEmail(contact, draft); // effect key inside
await createCrmTask(contact, message);
if (!(await condition(() => gotReply, "72 hours"))) {
await notifyOwner(contact);
}
}await client.workflow.start(demoFollowUp, {
taskQueue: "automations",
workflowId: `demo-follow-up:${event.id}`, // same id while this run is open
args: [contact],
});Read it against the mechanics chapter.
The weekend wait is sleep, and no worker is held while it runs. Activity
timeouts and heartbeats replace lease-renewal code, but they do not fence an
outside provider call: a timed-out attempt can still finish remotely. The
approval and reply are signals, and condition with a timeout is the race
between “a person acted” and “the deadline passed.” The early reply is not a
special case, because its handler sets gotReply whenever it arrives.
The workflow id collapses duplicate starts while that workflow is open. A duplicate that arrives after it closes depends on the configured workflow-id reuse policy, so a long duplicate window may still need a durable trigger receipt. Temporal removes the workflow's run store, timer sweep, and signal table from application code. An outbox may remain at the application-database to workflow-start boundary when those two writes must recover together.
What is still there is what the section on side effects said would be. sendEmail and createCrmTask are activities. An activity attempt may finish externally and then time out before Temporal records completion; with retries enabled, the activity can therefore run again. Each one claims its effect key, passes it to the provider where it can, and reconciles before retrying. The consent check belongs inside sendEmail, and the check for a revoked mailbox inside every activity that uses a connection.
At platform scale the natural shape is one workflow per contact per automation. A campaign that enrols 50,000 contacts starts 50,000 workflows; start them at a controlled rate, for example from a parent workflow that starts children in batches and continues-as-new between batches to keep its own history small. Recurring automations ("every Monday, email everyone whose trial ends this week") map onto Temporal schedules.
What Temporal takes off your hands
Mapping the custom engine onto Temporal shows which problems disappear and which stay yours:
- in a custom engine
- Start key and unique constraint
- on Temporal
- Workflow id: a second start with the id of a running workflow is rejected
- who owns it
- Temporal
- in a custom engine
- Run row and pure walk
- on Temporal
- Event history, replayed through deterministic workflow code
- who owns it
- Temporal, with rules
- in a custom engine
- Lease and heartbeat bookkeeping
- on Temporal
- Activity timeouts and heartbeats
- who owns it
- Temporal, configured by you
- in a custom engine
- Fencing an outside write
- on Temporal
- A timed-out attempt can still reach the provider
- who owns it
- You and the provider
- in a custom engine
- Delayed jobs and the due sweep
- on Temporal
- Durable timers
- who owns it
- Temporal
- in a custom engine
- Signals table
- on Temporal
- Signals, recorded in history even before the workflow waits for them
- who owns it
- Temporal
- in a custom engine
- Lanes and concurrency
- on Temporal
- Task queues and worker concurrency limits
- who owns it
- Temporal, configured by you
- in a custom engine
- Run log
- on Temporal
- Event history and the web UI
- who owns it
- Temporal
- in a custom engine
- Effect keys and reconciliation
- on Temporal
- Still needed: an activity can be attempted again after an ambiguous timeout
- who owns it
- You
- in a custom engine
- Fairness and the Monday 09:00 spread
- on Temporal
- Still needed
- who owns it
- You
- in a custom engine
- Consent and credential checks
- on Temporal
- Still needed, inside activities
- who owns it
- You
- in a custom engine
- Versions of an automation
- on Temporal
- Patching or worker versioning, plus replay tests
- who owns it
- Temporal, with rules
The rules that come with the replay model are specific to each SDK:
- In Go, workflow code uses
workflow.Nowfor the time,workflow.Sleeporworkflow.NewTimerto wait, andworkflow.SideEffectfor random values. Native goroutines, channels, andselectare replaced by theirworkflowequivalents, and ranging over a map is not deterministic, so sort the keys first. - In TypeScript, workflow code runs in a sandbox where
Date.now()andMath.random()are made deterministic and the SDK'ssleepcreates a durable timer. Anything that does I/O belongs in an activity. - In Python, workflow code uses
workflow.now(),workflow.random(), andworkflow.uuid4(), andasyncio.sleepbecomes a durable timer. Imports with side effects are restricted by the sandbox, and I/O belongs in activities.
Changing a workflow while executions are running is the Temporal version of the version-3-to-version-4 problem, and it bites in a quieter way. A harmless-looking change passes every test, ships, and then breaks executions that started on the old code, because replaying their histories now makes different decisions. The defences are to guard changes with the SDK's versioning calls (workflow.GetVersion in Go, patched in TypeScript and Python) or to pin running workflows to the worker build they started on, and to replay a sample of recent production histories against the new code in CI before every deploy. The Go, TypeScript, and Python SDK testing guides each document history replay for this check.
Activity timeouts are the next thing to get right. There are four: schedule-to-start (how long a task may wait in the queue), start-to-close (one attempt), schedule-to-close (all attempts together), and heartbeat. Temporal requires either start-to-close or schedule-to-close. The one people leave out is the heartbeat timeout on long activities: without it, a worker that dies ten seconds into a thirty-minute activity is only noticed when the thirty minutes are up.
Finally, histories have limits. Temporal's current programming-model limit is 51,200 events or 50 MB for one workflow execution, with warnings much earlier; Temporal Cloud also caps one request payload at 2 MB. Large results should be stored elsewhere and passed by reference, and long-lived workflows should periodically continue-as-new, which starts a fresh history carrying only the state that matters. These are service limits to verify again when designing a real deployment, not sizing targets. A visual automation builder on Temporal usually becomes one interpreter workflow that walks the published graph deterministically, so the published-graph model doesn't go away either.
What large teams learned running this
Durable workflows are no longer unusual infrastructure. Several companies have published what running them at scale taught them, and their lessons line up with the same boundaries: fairness, idempotency, versioning, and history size.
Uber · Cadence
Noisy neighbours
- Scale
- Over 12 billion executions and 270 billion actions a month, across 1,000+ services (2023)
- Lesson
- Month-end batch jobs used to consume whole clusters and delay interactive workflows, until task processing was made fair per tenant
Netflix · Temporal
Idempotent by force
- Result
- Deployments failing from transient cloud errors fell from 4% to 0.0001%
- Lesson
- Moving each step into an activity forced the team to make it idempotent; changing a function signature can break long-running workflows
Airbnb · Skipper
A library, not a cluster
- Scale
- Peaks of 10,000 workflows per second on DynamoDB, 15+ use cases
- Lesson
- Chose an embedded engine so a cluster outage couldn't stop every service; still calls workflow evolution 'the biggest friction point'
Uber runs Cadence, Temporal's predecessor, at a scale few systems reach, and its hardest problem was fairness rather than durability: "bursty traffic or workflow tasks from one customer usually consumes all the resources in a cluster … especially common and undesirable when batch processing jobs were triggered during month-ends" (Uber, 2021). That is the same problem as the per-tenant due sweep, one level down.
Netflix moved the steps of its deployment system into Temporal activities and reported that the share of deployments failing because of transient cloud errors "dropped from 4% to 0.0001%." The team also wrote that moving operations into activities forced the logic to become idempotent and warned that workflow changes can break long-running executions. A separate Netflix team building the Maestro workflow engine described the stale-worker race after garbage-collection pauses and network interruptions (Netflix, 2025).
Airbnb made the opposite platform choice from most: Skipper runs as a library inside each service rather than as a cluster, because "an orchestration cluster outage would mean every dependent service would lose the ability to start or advance workflows." It accepts the same trade-off as the side-effects section, "actions may execute more than once in edge cases (crash after execution but before checkpoint). Actions should be idempotent", and names the same pain: "workflow evolution remains the biggest friction point."
Versioning appears in almost every account. Uber notes that workflow changes are usually "only backward-compatible, but not forward-compatible", which makes rollbacks unsafe (Cadence, 2025). Microsoft warns that breaking changes to Durable Functions orchestrations can leave them "stuck indefinitely in a Running status" (Microsoft).
The platforms also have hard limits to understand before a design depends on them:
- platform
- Temporal
- limit
- Event history of 51,200 events or 50 MB per workflow, with a warning at 10,240 events or 10 MB
- what happens at the limit
- the workflow is terminated; continue-as-new before it gets close
- platform
- Temporal Cloud
- limit
- 2 MB per request payload; a 500-actions-per-second namespace floor under current on-demand capacity
- what happens at the limit
- large payloads are rejected; sustained throughput above the namespace limit is throttled
- platform
- Self-hosted Temporal
- limit
- The number of history shards is fixed when the cluster is created
- what happens at the limit
- can't be changed later: size it for growth
- platform
- AWS Step Functions (Standard)
- limit
- 25,000 events per execution history; 256 KiB input or output; up to one year per execution
- what happens at the limit
- the execution fails
The first two rows come from Temporal's current Cloud and programming-model limits; the self-hosted row comes from the numHistoryShards configuration reference. The table is here to expose design constraints, not to replace the provider's live limits page.
Datadog's engineers, who have run many self-hosted Temporal clusters, suggest designing workflows to stay well under the history limit (under 10,000 events), to make no more than about one state change per second per workflow, and to finish or continue-as-new within a day (Datadog, Replay 2026). A per-contact automation like Maya's fits those comfortably. A workflow that loops over every contact in a campaign does not, which is why fan-out belongs in child workflows started in batches.
The practical decision
Temporal is the stronger default when runs last for days, wait on people, cross several unreliable providers, or must survive deployments while they are in flight. A smaller Postgres-and-queue engine can still be sensible when the workflow set is narrow and the team is prepared to own every invariant in the mechanics chapter.
The dividing line is ownership. Temporal owns durable coordination; it does not own the correctness of the business effect. Activities still need stable identities, reconciliation, current permission checks, and careful payload boundaries.