A chat reply usually ends after one model response. An agent keeps going: the model chooses a tool, reads the result, and chooses what to do next. If the agent drafts twenty-three emails and its process crashes, a useful question is not “can we restart it?” It is “can we continue at draft twenty-four without paying for the first twenty-three again or sending anything twice?”
This is the companion to the durable workflow chapter. The runtime keeps run state outside the worker process, records every completed result before advancing, and gives each outside action a stable identity. A crash can then resume from recorded facts, while a timeout at a side effect can be reconciled without blindly repeating it.
The values, people, IDs and times in the running example below are illustrative. They make each failure visible; they are not production measurements from Robin's work.
How to read the diagrams
- The runtime and its records: checkpoints, history, the loop
- Model calls and tool calls: the outside world, where money and side effects happen
- People: the approvals an agent waits for
- Budgets and time: step limits, spend limits, deadlines
- Failure: crashes, duplicated actions, runaway loops
When a workflow engine is unnecessary
Durable execution has a cost: another runtime to operate, histories or checkpoints to retain, and stricter rules around state and code changes. A short agent that finishes inside one request, performs only read-only calls, has no human wait, and is cheap to restart can stay in the application process. A queue is also enough for a bounded background task when one durable job record represents the whole attempt and repeating the task is safe.
The threshold changes when the run spans deploys, waits for a person, fans out, spends enough that restarting hurts, or calls tools that change the outside world. At that point the application needs durable state, stable identities for effects, deadlines, recovery, and an audit trail. A workflow engine provides those mechanics as a product; a database, queue, outbox, timers, leases, and reconciliation can provide them too, but together they are the application's own workflow runtime. The decision is about the guarantees the run needs, not whether its next step was chosen by a model.
Before the example: what an agent loop actually does
An agent loop alternates between a model and ordinary program code. The model does not send email or query a customer system itself. It emits a structured tool call: a tool name plus arguments. The application validates that call, runs the corresponding code, returns the result to the model, and asks the model for the next step.
Follow one contact through the loop. The model first asks for data, then uses the returned row to draft an email, then waits for a person before the send tool can run.
The basic agent loop · one illustrative contact
- 1Agent runtime → Model
What should happen next?
request + state so far
- 2Model → Agent runtime
search_contacts
{ eventId: 'evt_demo', noEmailSince: '2026-09-22' }
- 3Agent runtime → Tools
Run the search
- 4Tools → Agent runtime
Result
[{ contactId: 'ct_017', name: 'Maya', question: 'Can I export?' }]
- 5Agent runtime → Model
What next?
the tool result is now part of the state
- 6Model → Agent runtime
Proposed send_email
{ to, subject, body }
- 7Agent runtime → Reviewer
Approve these exact arguments?
A checkpoint is a saved copy of the run's state after a step. Idempotency means that repeating the same requested action has the same effect as doing it once; a retry of one email returns the original message ID instead of sending another email. A token budget is a hard limit on the model input and output a run may consume. These three controls solve different problems: checkpoints preserve progress, idempotency prevents duplicate effects, and budgets stop unbounded work.
The example
A user of a marketing product asks its assistant: “Find everyone who attended last Tuesday's webinar and hasn't heard from us since, draft a personal follow-up for each, and check with me before anything goes out.” The request becomes the following run. A repeated step such as draft_follow_up × 38 means the runtime makes the same shaped call once per contact; the row shows one representative call so the arguments and result stay concrete.
One agent run · all values are illustrative
- 10:00:02
Plan
The model proposes the search, draft, approval, send and CRM steps.
model result → { next: 'search_contacts', reason: 'find eligible attendees' }
- 10:00:05
Find eligible contacts
A read-only tool returns 38 illustrative rows.
search_contacts({ eventId: 'evt_demo', noEmailSince: '2026-09-22' }) → [{ contactId: 'ct_017', name: 'Maya', question: 'Can I export?' }, …]
- 10:00:09
Draft each follow-up
The model drafts from one contact row at a time.
draft_follow_up({ contactId: 'ct_017', question: 'Can I export?' }) → { draftId: 'draft_017', subject: 'Your export question', body: '…' } · repeated for 38 contacts
- 10:06:51
Pause for review
The runtime stores the exact proposed sends and releases the worker.
request_approval({ action: 'send_email_batch', argumentsHash: '…', draftIds: ['draft_001', …] }) → { approvalId: 'apr_91', status: 'pending' }
- 14:20:13
Record the decision
The reviewer approves 36 drafts and edits 2 in this illustration.
approval_result({ approvalId: 'apr_91' }) → { status: 'approved_with_edits', approvedArgumentsHash: '…' }
- 14:20:14
Send each approved email
Each call changes the outside world and carries a stable idempotency key.
38 logical sends use run_4b9/send/0 … run_4b9/send/37 · contact 17 uses run_4b9/send/16
- 14:21:02
Create follow-up tasks
The next side effect also has its own stable key.
create_crm_task({ contactId: 'ct_017', dueOn: '2026-09-30', idempotencyKey: 'run_4b9/crm/16' }) → { taskId: 'task_2e1' }
- 14:21:05
Finish
The model turns the recorded results into a short report.
model result → { status: 'complete', sent: 38, tasksCreated: 38 }
The values are invented, but the shape matters. Waiting for a person can take far longer than the computation. Drafting consumes model tokens. Sending email and creating tasks can create harm if a retry repeats them.
What goes wrong without durability
First place the same crash at one precise point: the worker dies after saving draft_023, before it starts draft_024. The difference between the two designs is which calls the replacement worker makes.
- design
- state only in memory
- calls the replacement worker makes
- plan → search_contacts → draft_follow_up for contacts 1–23 → draft_follow_up for contact 24
- result
- the search and first 23 drafts repeat; their cost is paid twice, and the new drafts may differ
- design
- checkpoint after every result
- calls the replacement worker makes
- load checkpoint after draft_023 → draft_follow_up for contact 24
- result
- recorded search results and drafts 1–23 are reused; work continues from the first unfinished call
Why can the runtime reuse a model result? It stored the result before advancing the run. On recovery, “step 23 completed” is a fact in durable storage. Asking the model again would create a new answer, so replay returns the recorded answer instead.
There is still one unavoidable window: the model may finish and bill the call, then the worker may lose the response before it is checkpointed. If the model API offers no idempotency or result lookup, recovery has to call it again. A durable runtime prevents replay of results it recorded; it cannot prove the outcome of a remote call whose response never arrived. For model calls, that usually means accepting a small risk of duplicate cost and keeping retry counts low.
A second failure happens at a more dangerous boundary. The email provider accepts send_email for Maya and returns msg_91c, but the network drops that response. The worker sees a timeout. A checkpoint alone cannot prove whether the provider acted, because the success result never reached the checkpoint. If the provider supports idempotency, retry with the same run_4b9/send/16 key so it can return the first result without sending again. If it supports lookup by a client reference, reconcile before retrying. If it supports neither, the honest state is outcome_unknown: automatic retry can duplicate the email, so the run needs a product-specific policy or human review.
- failure
- worker dies after draft 23
- what repeats or disappears
- completed model calls repeat
- protection
- checkpoint each result
- failure
- approval waits through a deploy
- what repeats or disappears
- the in-memory promise disappears
- protection
- durable interrupt or workflow signal
- failure
- send succeeds but its reply is lost
- what repeats or disappears
- the email may be sent again
- protection
- stable idempotency key
- failure
- model keeps choosing the same tool
- what repeats or disappears
- tokens and calls grow without a bound
- protection
- step, spend and repetition budgets
Anthropic's multi-agent engineering report describes agents that "can run for long periods of time, maintaining state across many tool calls." Its authors explain why recovery matters: "we can't just restart from the beginning: restarts are expensive and frustrating for users." In their measurements, agents used about four times the tokens of chat interactions, and multi-agent systems about fifteen times. Those figures are specific to Anthropic's research system, but the restart cost applies to any loop that repeats completed model calls.
An agent is a workflow whose next step is chosen at run time
Anthropic's “Building effective agents” calls workflows systems where models and tools follow "predefined code paths", while agents let models "dynamically direct their own processes and tool usage." Both still execute sequences of steps, and the runtime records each completed result before the next step starts.
- in the agent
- a model call
- as a durable step
- a recorded step whose result is stored
- why
- on resume, the stored answer is reused instead of paying for a new one that might differ
- in the agent
- a read-only tool call
- as a durable step
- a recorded step, retried under a bounded policy
- why
- a second read does not repeat a mutation, but it can cost money, hit a rate limit, or observe newer data; replay reuses the recorded result
- in the agent
- a tool call that changes something
- as a durable step
- a recorded step with an idempotency key
- why
- a retry after a crash must not send, charge, or delete twice
- in the agent
- waiting for a person
- as a durable step
- a durable wait with a deadline
- why
- no process, connection, or memory is held while waiting
- in the agent
- one turn of the loop
- as a durable step
- a checkpoint
- why
- a crash resumes from the last completed turn
- in the agent
- a sub-agent
- as a durable step
- a child run with its own history
- why
- its failure is reported to the parent instead of killing it
- in the agent
- the conversation so far
- as a durable step
- state stored outside the process
- why
- large transcripts and tool results are kept by reference, not copied into every step
Two ways to make an agent durable
Checkpoint the agent's state
- How it works
- After each step, the framework saves the whole agent state under a thread id. Resuming loads the latest checkpoint and continues
- Example
- LangGraph with a checkpointer and interrupt() for human steps
- You still own
- Retries and timeouts per call, idempotency of tools, what happens to a step that crashed halfway, timers for deadlines
- Good fit
- An agent inside an existing service; moderate run lengths; the team already uses the framework
Run the loop in a durable engine
- How it works
- The loop is workflow code. Every model call and tool call is an activity with a retry policy, recorded in the history; approvals are signals
- Example
- Temporal with the OpenAI Agents SDK or Vercel AI SDK integrations
- You still own
- Idempotency of tools, keeping history small, versioning workflow code
- Good fit
- Long runs, waits of hours or days, many tools, several languages, fleets of agents
Checkpointing with LangGraph
LangGraph's interrupt documentation, checked in September 2026, gives a pause four pieces. A persistent checkpointer stores graph state. A thread_id points to that saved state. interrupt(payload) pauses and exposes its payload in result.__interrupt__. Calling the graph again with the same thread ID and new Command({ resume: value }) supplies the value that interrupt() returns inside the node.
This sketch shows the round trip without describing any product implementation.
durableCheckpointer stands for a configured database-backed checkpointer; the
in-memory saver used in tutorials does not survive a process restart.
import { Command, interrupt } from "@langchain/langgraph";
function awaitApproval(state: RunState) {
const decision = interrupt({
approvalId: state.approvalId,
argumentsHash: state.argumentsHash,
});
return { decision };
}
const graph = builder.compile({ checkpointer: durableCheckpointer });
const config = { configurable: { thread_id: "run_4b9" } };
const paused = await graph.invoke(initialState, config);
console.log(paused.__interrupt__); // payload for the approval UI
const resumed = await graph.invoke(
new Command({ resume: { status: "approved" } }),
config,
);One detail causes duplicate side effects if it is missed: LangGraph starts an interrupted node again from the beginning when it resumes. Code before interrupt() runs again. If a node posts a notification and then pauses, approval makes that notification run a second time.
// Wrong: resuming after approval re-runs this whole node,
// so the Slack notification goes out a second time.
async function askApproval(state: RunState) {
await notifySlack(state.ownerId, "38 drafts are ready for review");
const decision = interrupt({ drafts: state.drafts });
return { decision };
}
// Right: the side effect has its own node before the pause,
// and the node that pauses does nothing else.
async function notifyReviewer(state: RunState) {
await notifySlack(state.ownerId, "38 drafts are ready for review", {
idempotencyKey: `${state.runId}/notify-review`,
});
return {};
}
async function awaitApproval(state: RunState) {
return { decision: interrupt({ drafts: state.drafts }) };
}Keep the pause in a node that has no earlier side effect, or make the earlier effect idempotent. The same page warns that multiple interrupts in one node are matched by position, so do not reorder or conditionally skip them between the initial run and the resume.
Running the loop in Temporal
The second approach makes the loop itself a workflow. In Temporal, workflow code holds the durable control flow: what step comes next, whether approval has arrived, and which budget remains. An activity performs uncertain outside work such as calling a model, sending an email, or reading a customer system. Temporal records each activity's outcome in workflow history, so replay rebuilds control state without repeating completed outside work.
Temporal's Durable AI documentation, checked in September 2026, lists integrations for LangGraph, the OpenAI Agents SDK, and Vercel AI SDK. The OpenAI Agents SDK integration runs model invocations as activities and can expose activities as tools. The Vercel AI SDK integration says its provider "automatically wraps every LLM call in a Temporal Activity"; tools that call external APIs still need their own activity boundaries.
OpenAI has described Temporal as "a critical part of the infrastructure powering Codex, responsible for executing our core control flows" (Temporal and OpenAI, 2025).
The trade is described in the Temporal chapter. Temporal owns retries, timers, waits, and history. The workflow code must remain deterministic and versioned. Model and tool calls belong in activities, where their answers can be recorded and reused during replay.
Tool calls are the dangerous part
Model calls cost money. Tool calls can change things. A duplicated model call wastes money; a duplicated send_email, issue_refund, or delete_contact can harm a user.
Sending email 17: the provider accepted it, the worker didn't hear back
send_email(contact 17)
key run_4b9/send/16
POST /messages with the key
timeout: the provider may or may not have sent it
retry send_email(contact 17)
same key run_4b9/send/16
look up the key before sending
already sent: msg_91c
return the original result
The runtime needs a policy for each tool before a model can call it. “Retry” is safe advice for a contact search. It is incomplete advice for an email send because the timeout may have happened after the provider accepted the message.
- tool kind
- read-only
- example
- search_contacts
- retry rule
- retry after transient failure; reuse a recorded success during replay
- evidence to store
- arguments hash, returned artifact reference, time
- tool kind
- state-changing with provider key
- example
- send_email
- retry rule
- retry with the same key derived from run and call position
- evidence to store
- key, provider message ID, final status
- tool kind
- state-changing with local ledger
- example
- create_crm_task
- retry rule
- claim the key locally, call once, then save the result; reconcile an uncertain attempt
- evidence to store
- claim state, remote task ID, attempt history
- tool kind
- irreversible without deduplication
- example
- an API that has no lookup or key
- retry rule
- stop after an uncertain result and ask a person
- evidence to store
- request bytes, timeout, operator decision
The key comes from the logical call, such as run_4b9/send/16, and stays the same across attempts. Including an attempt number would create a new identity on every retry and defeat deduplication.
A model turn that has already produced tool calls also needs care. Re-running that turn may produce a different set of calls. Record the proposed calls before execution, then execute those recorded calls through the tool layer. A provider failure before any proposal exists can retry the model call; a failure after the proposal exists resumes from the recorded proposal.
Route every tool through one application boundary that validates arguments against the tool schema, checks the current user's permission, assigns the idempotency key, and writes an audit record. Treat instructions found inside fetched web pages, email bodies, and documents as untrusted data. The authorisation decision belongs to application code, never to text the model saw.
Approvals must bind to the exact action
"Check with me before anything goes out" sounds simple. The details decide whether the approval means anything.
Hash a canonical representation: the same fields in the same order, encoded the same way every time. For one illustrative email, the bytes below produce the shown SHA-256 digest.
{"body":"Thanks for joining, Maya.","subject":"Your export question","to":"maya@example.test"}7292b1cf5c92933d867792003ffe175badcd8a036258d24e1e2a5c7f864be97eWatch the stored row as the reviewer approves, then the agent tries to change the body.
- action
- send_email to maya@example.test
- arguments hash
- 7292b1cf…64be97e
- decision
- pending
- execution
- blocked
- action
- send_email to maya@example.test
- arguments hash
- 7292b1cf…64be97e
- decision
- approved
- execution
- allowed only for this hash
- action
- send_email to maya@example.test
- arguments hash
- d110fc10…d8da963
- decision
- does not cover these bytes
- execution
- rejected; request approval again
The run asks for one batch decision over an ordered list of 38 immutable sends. Each item has its own argument digest; the batch digest covers that ordered list. The table isolates Maya's item so the byte change is visible. Editing two drafts produces a new batch digest, which is the value the reviewer approves. Each send then executes under its own send/0 through send/37 key. A later regeneration changes both the item and batch digests, so the old decision cannot authorise it.
The wait itself must be durable: a checkpoint or a workflow signal, not a process sitting on an open promise. It needs a deadline and a plan for when the deadline passes (remind, escalate, or cancel the run and say so). The decision must be checked against the person's permissions at the moment they decide, and a second decision on the same request must be refused. These are the same rules as approvals in any workflow, and the workflow-engine mechanics chapter covers them in more detail.
Budgets and stopping
An agent decides how many steps it takes, so something else has to decide how many it may take. Anthropic's agent-building guidance recommends explicit stopping conditions, including a maximum iteration count. In practice there are four limits worth enforcing, and all of them belong in the runtime rather than in the prompt:
Steps
A maximum number of model turns and tool calls per run. Hitting it ends the run with a clear message, not a silent stop.
Spend
Credits reserved before each turn and settled after, against a per-run ceiling as well as the organisation's balance.
Wall clock
A deadline for the whole run, separate from the approval deadline, so a run can't live forever.
Repetition
The same tool with the same arguments several times in a row is a loop, not progress. Stop it and report.
One more rule: do not put a semantic response cache in front of agent turns. Consecutive turns share most of their text, so their embeddings are nearly identical, and a cache can return the previous turn's answer. LiteLLM's documentation describes the visible symptom: "an agent repeating the same tool call over and over." A provider's exact-prefix prompt cache is different. It reuses computation for the stable prefix while the model still makes a new decision. The semantic caching chapter explains the failure in detail.
State grows
Every turn can add model messages, tool arguments, tool results and generated artifacts. Copying all of that into every checkpoint makes writes slower and retains customer data in more places. Putting it directly into a workflow history also spends a finite operational budget. Temporal's current limits documentation, rechecked 1 October 2026, caps one workflow history at 51,200 events or 50 MB and warns at 10,240 events or 10 MB. Treat those as hard ceilings to stay well below and verify again when sizing a deployment.
The saved control state should be enough to decide what happens next. Large evidence belongs in an artifact store, addressed by an immutable reference and checksum.
- kind
- control
- save directly
- nextContactIndex: 23 · status: drafting
- save by reference
- kept in checkpoint
- why
- the runtime needs these small values to resume
- kind
- budgets
- save directly
- stepsUsed: 25 · tokensUsed: 41,820
- save by reference
- kept in checkpoint
- why
- limits must survive a crash
- kind
- messages
- save directly
- recent turns needed for the next decision
- save by reference
- older transcript → artifact transcript_a8
- why
- old context can be loaded for audit without copying it into every checkpoint
- kind
- tool results
- save directly
- contact IDs and checksums
- save by reference
- 38 contact rows → artifact contacts_52
- why
- the checkpoint stays small and the exact source data remains recoverable
- kind
- drafts
- save directly
- draft IDs and approval status
- save by reference
- email bodies → artifact drafts_b1
- why
- approval and send steps refer to immutable bytes
Compaction means replacing older conversational detail with a smaller summary after the exact transcript has been stored elsewhere. The summary helps the model continue; it is not the audit record and should not erase the original evidence. Test compaction with questions that depend on old constraints, because an over-aggressive summary can change the agent's decision.
For a very long Temporal run, continue-as-new closes the current history and starts another with a small carry-forward state. A checkpointed graph needs an equivalent rollover or pruning policy. Both designs also need retention and deletion rules because checkpoints, transcripts and artifacts can contain customer data.
Many agents
Some tasks have independent branches: one agent can research pricing while another checks reliability and a third compares APIs. In durable terms, the lead agent is a parent run and each sub-agent is a child run with its own history, budget and result.
Bounded fan-out · three child runs
- 1Parent → Child A
research pricing
budget: 8 calls
- 1Parent → Child B
research reliability
budget: 8 calls
- 1Parent → Child C
research API
budget: 8 calls
- 2Child A → Parent
complete
artifact price_notes
- 3Child B → Parent
failed
source unavailable after retry limit
- 4Child C → Parent
complete
artifact api_notes
- 5Parent
decide
continue with a stated gap, replace Child B, or stop
The parent should specify an objective, output shape, allowed tools and budget for each child. Without distinct boundaries, children can repeat the same search and spend more without adding coverage. Bound the number running at once because concurrency multiplies provider rate pressure as well as token spend.
- child outcome
- transient tool failure
- parent choice
- child resumes from its own checkpoint
- use when
- the work is still valuable and its retry budget remains
- child outcome
- permanent failure
- parent choice
- record the gap and continue
- use when
- other branches can still answer the user's question honestly
- child outcome
- required branch failed
- parent choice
- start a replacement or stop the parent
- use when
- the final answer would be misleading without this evidence
- child outcome
- parent is cancelled
- parent choice
- request cancellation of every unfinished child
- use when
- no orphan should keep spending after its result is unwanted
Anthropic's research-system report found a 90.2% improvement over its single-agent baseline on one internal research evaluation. It also reports that its multi-agent runs used about fifteen times the tokens of chat interactions. Those results support parallel agents for broad, high-value research; they do not support adding agents to a sequential task where every branch depends on the previous one.
Deploying while agents are running
Prompts, tools, and models change often, and agent runs are long. A deploy that changes the tool list or the state shape can break every run that is paused on the old version.
Anthropic describes a rainbow deployment: new runs move to the new release while existing runs keep using the release they started on. The durable-engine equivalent is worker versioning, which pins each run to a compatible worker build. For checkpointed agents, store a version with each checkpoint and keep the code that can resume old versions until none are left, or write an explicit migration for the state shape. In every case, a run that started on version 3 should finish on version 3 unless someone has decided otherwise.
Seeing what an agent did
When a user asks why the assistant emailed someone twice, or never finished, the answer should be one screen:
Webinar follow-ups · illustrative run
requested by the user · 38 contacts · budget 150 steps · 500 credits
- 10:00:02turn.completedturn 1 · planned: search contacts, draft, ask for approval
- 10:00:05tool.search_contacts38 results · read-only
- 10:03:58turn.retriedturn 19 · provider 429 · resumed from checkpoint after 20 s
- 10:06:51approval.requested38 drafts · exact batch arguments hashed
- 14:20:13approval.decidedapproved with 2 edits · edited batch hashed again
- 14:20:43tool.send_emailcontact 17 · retry recognised by key, not re-sent
- 14:21:02tool.create_crm_task38 tasks
- 14:21:05run.completed40 turns · 64,112 tokens · 212 credits
Record every turn and every tool call with its arguments' hash, result size, token use and latency. Anthropic reports that production traces helped diagnose agent failures while examining decision patterns without reading individual conversation contents. That is a sound default: put structure and measurements in traces, and include content only where the customer has agreed to it.
A checklist
Durability
- Every model call and tool call is a recorded step
- A crash resumes from the last completed step
- Nothing waits in memory: approvals are durable waits
- A node that resumes from an interrupt cannot repeat an earlier unprotected side effect
Tools
- Read-only and state-changing tools are separated
- State-changing tools run under keys from run and call position
- Tool-calling turns are not replayed solely to retry the model
- One gate validates, authorises (fail closed), and audits every call
Approvals
- Approval binds to a hash of the exact arguments
- Any change after approval voids it
- Deadlines, escalation, and permission checks at decision time
Limits
- Step, spend, and wall-clock budgets per run
- Repeated identical tool calls stop the run
- No semantic response cache on agent turns
State
- Large tool results stored by reference
- Older turns compacted; long runs start fresh histories
- Checkpoints versioned and given a retention policy
Operations
- Sub-agents are child runs with bounded fan-out
- Deploys pin running agents to their version
- Traces record structure and numbers, not content
Sources
- Anthropic, How we built our multi-agent research system (June 2025) and Building effective agents (December 2024)
- LangGraph, Interrupts
- Temporal, Durable AI integrations, Production-ready agents with the OpenAI Agents SDK + Temporal, Building durable agents with Temporal and AI SDK by Vercel (January 2026), and a post on Codex and Temporal's Java SDK (May 2025)
- Temporal, Workflow execution limits
- LiteLLM, Semantic caching