Skip to content

Architecture Article

Receiving and Sending Webhooks Reliably

Follow webhooks in both directions, from signature verification and deduplication to durable sending, retries, endpoint isolation, and delivery logs.

Published 18 Jun 2026Updated 1 Oct 202610 min read
Webhooks · Integrations · Security · Reliability · Event-Driven
A signed webhook crosses a network boundary while one delivery route enters retry.
On this page (17)

A webhook is an HTTP request sent when something happens. Without it, the receiving system often has to poll: ask the sender's API repeatedly whether anything changed. Webhooks reduce that repeated work, but the sender and receiver can fail independently and rarely agree about what "delivered" means.

I have worked on both ends. At Vimeo, meeting-recording integrations received Microsoft Graph change notifications. The sending side in this chapter is an illustrative platform built from public protocols rather than a description of another product's deployment.

On receipt, verify the exact request, store the event before acknowledging it, deduplicate it, and reconcile later. On send, create the event durably, sign it with a timestamp, isolate each endpoint's retries, and give customers a delivery log.

How to read the diagrams

  • The other system: the provider sending to you, or the customer endpoint you send to
  • Your platform and its records
  • Time: deadlines, retry schedules, subscription lifetimes
  • Customers and the people debugging a delivery
  • Failure: forged requests, lost events, disabled subscriptions

The map below marks where the failures happen. The numbered badges correspond to the order in which a receiver or sender handles an event.

Sender

Provider or your outbox1

creates one event

event id · type · raw body

Signer2

authenticates bytes

timestamp + HMAC

timeouts, retries, duplicates, replay and SSRF live here

Network boundary

HTTPS request3

headers + body bytes

may arrive zero, one, or many times

authenticate, resolve tenant, persist, acknowledge

Receiver

Verifier4

raw bytes and freshness

signature · timestamp · delivery id

Durable inbox5

deduplicates

UNIQUE (source, event_id)

Background work6

can retry independently

queue · reconciliation · dead letters

Receiving moves top to bottom. When your platform sends, the same map applies with your outbox at the top and a customer endpoint at the bottom.

Three terms explain most of the design. At-least-once delivery means the sender may send the same event more than once, so the receiver needs an idempotency key, a stable identifier that makes a duplicate produce no second effect. A replay attack is different: someone captures a valid request and deliberately sends it again. Signatures prove who created bytes; timestamps and remembered delivery ids limit how long those bytes remain acceptable.

What the big senders promise

The best way to understand what a receiver must handle is to read what senders actually promise. They promise less than most receivers assume:

documented delivery behaviour
sender
Stripe
answer within
quickly; return 2xx before any slow work
retries
up to 3 days, exponential backoff (live mode)
order
not guaranteed
after repeated failure
email to the account; some billing waits up to 72 hours on invoice.created
sender
GitHub
answer within
10 seconds
retries
none automatic; manual redelivery for 3 days
order
not stated
after repeated failure
delivery recorded as failed
sender
Shopify
answer within
1 s to connect, 5 s in total
retries
8 times over 4 hours
order
not guaranteed
after repeated failure
subscription deleted (if created through the Admin API)
sender
Slack Events API
answer within
3 seconds
retries
3: immediately, after 1 minute, after 5 minutes
order
not stated
after repeated failure
subscription disabled if over 95% fail within 60 minutes
sender
Twitch EventSub
answer within
quickly
retries
at least once
order
not stated
after repeated failure
subscription revoked
sender
Svix (a webhook service)
answer within
15 seconds
retries
8 attempts over about 27 hours
order
not stated
after repeated failure
endpoint disabled after 5 days of failures
From each sender's documentation, checked September 2026.

The contracts require three design choices. Order cannot be assumed. Duplicates are normal under at-least-once delivery. A receiver that remains slow or unavailable can lose its subscription, after which events stop rather than accumulating forever.

Receiving

The running example is an illustrative Instagram comment delivered to a social inbox. Here is what happens between the request arriving and the 200 going back:

One inbound webhook, from bytes to acknowledgement

  1. 0 ms

    Capture the raw body before any parsing

    The signature covers the exact bytes Meta sent. Parsing and re-serialising JSON changes them.

    rawBody: 1,284 bytes

  2. 1 ms

    Verify the signature in constant time

    HMAC-SHA256 of the raw body with the app secret, compared byte by byte after checking the lengths match.

    X-Hub-Signature-256: sha256=9f2c… · valid

  3. 1 ms

    Accept authentic retries idempotently

    Meta's signature carries no timestamp. An identical retry is indistinguishable from a captured replay, so the receiver verifies it and relies on the event uniqueness key to make it a no-op.

    same external message id → 200 OK, no duplicate work

  4. 2 ms

    Find the tenant from your own records

    The payload names the Instagram account; your connection table says which organisation owns it. Nothing in the payload chooses the tenant.

    ig account 1784… → conn_91 → org_42

  5. 4 ms

    Store it with a uniqueness key

    The event, thread update, and an outbox command for slow work are written in one transaction, with a unique key on the provider's own id.

    UNIQUE (connection_id, external_message_id) · inserted

  6. 5 ms

    Acknowledge

    200 OK

  7. later

    Do the slow work in the background

    Classify the comment, update the conversation, and trigger any configured automation.

    queued: background handling

Verify the bytes you received

Every sender signs differently, so each needs its own verifier:

signature schemes
sender
Meta (Instagram, Facebook, WhatsApp)
header
X-Hub-Signature-256
signed content
HMAC-SHA256 of the raw body
replay window
none: accept authentic retries and deduplicate effects
sender
GitHub
header
X-Hub-Signature-256
signed content
HMAC-SHA256 of the raw body
replay window
none; X-GitHub-Delivery is a unique id
sender
Stripe
header
Stripe-Signature: t=…,v1=…
signed content
HMAC-SHA256 of timestamp.body
replay window
5 minutes by default
sender
Slack
header
X-Slack-Signature, X-Slack-Request-Timestamp
signed content
HMAC-SHA256 of v0:timestamp:body
replay window
5 minutes
sender
Twitch EventSub
header
Twitch-Eventsub-Message-Signature
signed content
HMAC-SHA256 of id + timestamp + body
replay window
10 minutes
sender
Discord interactions
header
X-Signature-Ed25519
signed content
Ed25519 over timestamp + body
replay window
timestamp header
sender
Twilio
header
X-Twilio-Signature
signed content
HMAC-SHA1 of the URL and parameters
replay window
none
sender
Standard Webhooks
header
webhook-id, webhook-timestamp, webhook-signature
signed content
HMAC-SHA256 (or Ed25519) of id.timestamp.body
replay window
timestamp header

An HMAC is a keyed digest: anyone with the shared secret can turn the same bytes into the same fixed-length value. A receiver calculates that value and compares it with the header. A constant-time comparison avoids returning faster for an early mismatch, which could leak information about a valid signature over many measurements.

This illustrative example uses Stripe's documented shape. The secret is deliberately fake, the Unix timestamp is 1790500447, and the raw UTF-8 body is 39 bytes. Stripe's real header is named Stripe-Signature and carries comma-separated t= and v1= fields.

Sketch: illustrative signature input, shown byte for byte
secret:      whsec_demo_42
timestamp:   1790500447
raw body:    {"id":"evt_demo","type":"lead.created"}
signed text: 1790500447.{"id":"evt_demo","type":"lead.created"}
HMAC-SHA256: bd11f7da27f13e04f9af4aa4625cc5271420199d22ab0567dc5e785572282202
 
Stripe-Signature: t=1790500447,v1=bd11f7da27f13e04f9af4aa4625cc5271420199d22ab0567dc5e785572282202

If a JSON parser reverses the fields and serialises {"type":"lead.created","id":"evt_demo"}, the data still means the same thing to an application. The signed bytes are different, so the digest becomes 537a78…c2433 and verification correctly fails. This is why the request body and headers must be captured before JSON parsing.

Two implementation details cause most verification bugs. First, verify before parsing: capture the raw body in the HTTP framework's body-parser hook, because a JSON parser followed by JSON.stringify produces different bytes and the signature will never match (or, worse, a hand-rolled check will pass on the parsed form). Second, compare in constant time and check lengths first. Node's crypto.timingSafeEqual throws if the two buffers have different lengths, so a header with multibyte characters or the wrong length turns a forged request into an unhandled exception and a 500, which some senders treat as "retry" and some attackers treat as a signal. Compare lengths, then compare bytes.

Where the scheme has a timestamp, reject requests outside the window. Where it does not, signature verification proves who could have created the bytes but cannot prove when they were sent. GitHub supplies a delivery id that can be deduplicated. For Meta, accept an authentic identical delivery, insert with a provider event or message uniqueness key, and answer 200 when the insert conflicts. This preserves Meta's own retry path while ensuring the repeated request causes no repeated effect.

The tenant comes from you, not the payload

A webhook endpoint is public, so everything in the body is attacker-controlled until proven otherwise, including any field that looks like your own organisation id. Resolve the tenant from something you created: the connection that owns the provider account, a per-subscription channel id you registered, or an unguessable secret bound to the subscription. A channel-to-connection mapping prevents an event from choosing another tenant, but a public channel id alone does not authenticate the sender. If a provider offers no signature, use every secret or token its protocol supplies, keep the callback identifier unguessable where the protocol permits it, restrict the accepted request shape, and reconcile against the provider API. Check consistency too: an event that arrives on the Instagram path but names a connection belonging to another provider should be refused and counted.

Acknowledge only what is safe

The sender's timeout forces you to answer quickly, and it is tempting to answer first and work later. Acknowledge only after the event is durable, meaning written to a database or a queue that survives a crash. An endpoint that verifies an event, hands it to an in-memory worker, answers 201, and then crashes has made a promise it cannot keep; the sender will not retry a successful response.

The opposite mistake also causes failures. Meta disables subscriptions that keep failing, so a valid event your platform doesn't understand (a new field, a new event type) should still be stored for inspection, marked unsupported, and acknowledged successfully instead of retried forever.

Handshakes prove that you control the URL

Before sending events, providers use several distinct challenge protocols. Microsoft Graph puts an opaque validationToken in the query string; the endpoint URL-decodes it and returns it as text/plain within ten seconds. Meta sends hub.mode, hub.verify_token, and hub.challenge; after checking the configured verify token, the callback returns the challenge. X sends a CRC token and expects a JSON response_token containing an HMAC-SHA256 result. LinkedIn's validation handshake likewise asks the endpoint to return its challenge plus an HMAC made with the client secret. Each handshake proves control of the callback URL. It does not authenticate later notifications, so the event path still applies that provider's signature or subscription-secret check.

Some protocols ask the callback to sign a challenge. Keep that operation narrow: accept the documented challenge format and length, use the challenge-specific response construction, and reject everything else. An unrestricted endpoint that returns an HMAC for arbitrary input becomes a signing oracle that an attacker can use to obtain valid signatures over chosen bytes.

Subscriptions expire

Some webhook subscriptions are permanent. Others have finite lifetimes: Microsoft Graph subscriptions have short maximum lifetimes that depend on the resource, YouTube's WebSub subscriptions are leases, and CRM "watch" channels expire. A renewal job that fails quietly produces the worst kind of outage, because nothing errors: events stop, and someone notices stale data days later.

Renewal needs its own schedule, well ahead of expiry, and its own alert. The most useful alert is on absence: a source that normally sends events every hour and has sent none for six hours is broken, whatever the renewal job's logs say.

Microsoft Graph can include encrypted resource data when a subscription sets includeResourceData. The subscription registers a public certificate and an identifier. A notification returns that identifier so the receiver can choose the matching private key, unwrap the per-message key, verify the data signature, and decrypt the content. This public Graph contract adds certificate rotation to the renewal design: keep an old private key available while notifications can still reference it.

Reconcile, because some events never come

Even with everything above, some events will be lost: the sender gives up after its retry window, a subscription lapses, an outage on your side lasts longer than their patience. Shopify says plainly that "webhook delivery isn't always guaranteed" and recommends reconciliation jobs. A periodic job that lists recent changes through the sender's API and compares them with what arrived is the only way to close that gap. Stripe's events API returns the last 30 days, which sets how long you can afford to wait before reconciling.

Here is the whole inbound path, including the failures it is built for:

Inbound: a duplicate, a replay, and a normal event

  1. 12:00:00.000Meta → Webhook endpoint

    POST comment event

  2. Webhook endpoint

    raw body · signature valid · not seen before

  3. 12:00:00.004Webhook endpoint → Postgres

    COMMIT event + thread update + outbox

    unique (connection, external id) · 1 row

  4. Webhook endpoint → Meta

    200 OK

  5. laterPostgres → Outbox relay + queue

    relay intent analysis

    retries until the queue accepts it

  6. 12:00:09.000Meta → Webhook endpoint

    same event again (sender retried)

  7. Webhook endpoint → Postgres

    INSERT … ON CONFLICT DO NOTHING

    0 rows: already stored

  8. Webhook endpoint → Meta

    200 OK: nothing new to do

  9. 12:03:00.000Meta → Webhook endpoint

    captured request replayed by someone else

  10. Webhook endpoint → Postgres

    same signed bytes and event id: conflict, no new work

  11. Webhook endpoint → Meta

    200 OK: an authentic repeat is safe to acknowledge

Sending

When your platform sends webhooks, you are the one whose promises customers read. The design problem changes from "accept safely" to "deliver fairly to endpoints you don't control, some of which are down".

Create the event once, durably

An event should be written to a transactional outbox in the same database transaction as the change that caused it, so a crash can't produce a change with no event or an event with no change (the workflow-engine mechanics chapter covers this dual-write problem in detail). The envelope is created once, with a stable id, and every delivery attempt to every endpoint sends that same id, so receivers can deduplicate:

Sketch: illustrative event envelope
{
  "id": "evt_4c1f9a2be07d53a1c6f0",
  "type": "lead.created",
  "createdAt": "2026-09-27T09:14:07.412Z",
  "data": { "leadId": "ld_8812", "formId": "fm_19", "email": "…" }
}

Sign it so receivers can check freshness

Use a timestamped signature so the receiver can reject old requests without keeping every request hash. The following illustrative product header uses the same t= and v1= field shape that Stripe documents:

Sketch: signature headers on one delivery attempt
X-Example-Event: lead.created
X-Example-Signature: t=1790500447,v1=5e1b…c09a
                     ^ unix time    ^ hex HMAC-SHA256 of "1790500447." + raw body

Re-sign every attempt with a fresh timestamp, as Stripe does, so a retry three hours later isn't rejected by the receiver's five-minute window. Support secret rotation by signing with both the old and the new secret for a period (Stripe allows up to 24 hours of overlap), sending both signatures so the receiver can accept either. The Standard Webhooks specification, started by Svix with contributors from Zapier, Twilio, Kong, and others, formalises the same ideas: webhook-id, webhook-timestamp, and a space-separated list of signatures over id.timestamp.body, so rotation is built into the format.

Customer URLs are an SSRF risk

A customer-supplied URL is a request your servers will make. Server-side request forgery (SSRF) happens when an attacker chooses that destination and reaches something your public users cannot reach. An attacker might register http://169.254.169.254/latest/meta-data/, the IPv4 address used by Amazon EC2's instance metadata service, its IPv6 counterpart fd00:ec2::254, or https://admin.internal.example/. A naive sender would attach the webhook body and follow the attacker's chosen network path.

endpoint registration and delivery
step
1 · parse
candidate
http://169.254.169.254/latest/meta-data/
result
reject: HTTPS required
step
2 · resolve every address
candidate
attacker.example → 169.254.169.254 or fd00:ec2::254
result
reject: metadata destination
step
3 · connect
candidate
https://hooks.customer.example → public IPv4 and IPv6
result
pin an approved address for this attempt
step
4 · redirect
candidate
302 Location: https://admin.internal.example/
result
refuse, or repeat every validation before following
step
5 · network policy
candidate
any missed private or metadata route
result
egress firewall denies it
Validate at registration and again for every attempt. DNS can change after the customer saves the endpoint.

Use a standard URL parser, allow public HTTPS destinations, resolve every IPv4 and IPv6 answer, and block loopback, private, link-local, and metadata ranges. Connect to an address that passed validation while preserving the intended hostname for TLS. Refuse redirects unless the new destination passes the full sequence. A network egress policy is the final containment layer if application validation misses an encoding or address form.

Retry for hours, with jitter, per endpoint

Senders disagree about how long to keep trying, and the range is wide:

How long each sender keeps retrying a failed delivery

Retry window

y: hours

Checked September 2026. Slack's default retries end after 5 minutes; Microsoft Graph and Shopify retry for up to 4 hours; Svix's published schedule totals about 27.6 hours; Stripe retries for up to 3 days in live mode.

A window of minutes turns a short customer deploy into lost events. A window of days means your queue has to hold days of backlog for a dead endpoint. Exponential backoff increases the delay after each failure. Jitter adds randomness so thousands of delayed deliveries do not wake at the same instant. Choose the window from the product's recovery promise and storage budget, then publish it.

After the final automatic attempt, move the delivery to a dead-letter queue or equivalent failed state: durable storage for work that needs inspection or manual replay. A dead letter is evidence and a recovery point; it should not become a second invisible retry loop.

The deeper lesson comes from Segment, which built a delivery system called Centrifuge to send about 400,000 outbound requests per second. Their write-up explains why a single shared queue fails: if each of 200+ endpoints has an hour of downtime a year, outages occur somewhere almost every day, and one queue lets each outage slow everyone. They also measured how much retries matter: "about 1.5% of all global data succeeds on a retry", and "about 50% of retries succeed only on the third through the tenth attempts." Retrying is a meaningful share of delivered data, and it has to happen without one broken endpoint delaying the rest.

So retry state belongs to the endpoint, not only to the delivery:

Isolate

Each endpoint gets its own concurrency limit and its own position in the retry schedule. A slow endpoint fills its own lane, not a shared one.

Be fair

The retry sweep takes due deliveries round-robin across endpoints and tenants, so one customer's backlog doesn't hold up everyone's first attempts.

Check state at send time

A delivery due for retry checks that its endpoint is still active. Pausing an endpoint should stop its retries as well as new events.

Disable deliberately, and tell someone

An endpoint that has failed every delivery for days consumes queue capacity and leaves the customer's data behind. Large senders publish different rules: Slack can temporarily disable event subscriptions when more than 95% of deliveries fail within 60 minutes, Shopify can delete an Admin API subscription after eight failed attempts, Svix disables an endpoint after its documented multi-day condition, and Standard Webhooks treats 410 Gone as a disable signal. Pick a rule, show it before it triggers, notify the customer, and preserve a clear replay path.

Follow one illustrative endpoint through a six-hour customer outage:

hooks.example.com is down for six hours

  1. 09:14

    First attempt fails

    Store the 503 response and schedule from this endpoint's retry state.

    evt_4c1f9a2be07d53a1c6f0 · attempt 1 · next attempt includes jitter

  2. 09:15–12:40

    Backoff grows

    Fresh deliveries for healthy endpoints continue. This endpoint's failures do not occupy their concurrency.

    attempts 2–6 · 503 or timeout

  3. 15:14

    Endpoint recovers

    The next due attempt returns 200. The event keeps its original id and receives a fresh signature timestamp.

    attempt 7 · 200 OK

  4. after recovery

    Customer reviews and replays

    The delivery log shows the outage window. Any delivery that exhausted the policy can be replayed from its dead-letter state.

    replay preserves event id for receiver deduplication

If the outage had crossed the platform's disable threshold, the endpoint would move to disabled, the customer would be notified, and no new attempts would run until re-enabled. Re-enabling does not silently discard the backlog; the customer chooses the documented replay window.

Give customers a delivery log

Most webhook support tickets are "did you send it?". A delivery log that answers that without a ticket needs, per attempt: the event id and type, the time, the response status, the response time, and a snippet of the response body, plus a replay button for one delivery and for everything since a point in time. Keep it long enough to cover the retry window and a weekend.

lead.created → https://hooks.example.com/events

endpoint ep_31 · org_42 · signing secret v2

delivered
  1. 09:14:07.9attempt 1503 Service Unavailable · 2,114 ms
  2. 09:15:21.3attempt 2timeout after 15 s
  3. 09:17:58.6attempt 3connection refused
  4. 09:26:40.2attempt 4503 Service Unavailable
  5. 09:58:11.0attempt 5timeout after 15 s
  6. 11:42:03.5attempt 6503 Service Unavailable
  7. 15:14:09.4attempt 7200 OK · 184 ms · endpoint recovered after six hours

Stripe's documentation mentions a load pattern worth planning for: deliveries spike "during the beginning of the month when all subscriptions renew." Any platform with scheduled business events has its own version, and the delivery system needs headroom for it.

A checklist

Receiving: verification

  • Raw body captured before parsing
  • One verifier per scheme, constant-time, lengths checked first
  • Timestamp window where available; stable delivery or event ids deduplicate effects
  • Challenge endpoints accept only the sender's exact format

Receiving: handling

  • Tenant resolved from your records, never the payload
  • Acknowledge only after the event is durable
  • Unique key on the provider's event id
  • Unknown but valid events acknowledged and logged

Receiving: over time

  • No reliance on order: versions or re-fetch
  • Subscription renewal scheduled well before expiry
  • Alerts on silence from sources that normally send
  • Reconciliation against the sender's API

Sending: events

  • Outbox in the same transaction as the change
  • One envelope with a stable id for every attempt
  • Timestamped signature, re-signed per attempt
  • Two signatures during secret rotation

Sending: delivery

  • Public HTTPS only, checked at save and at send, no redirects
  • Per-endpoint concurrency and retry state
  • Backoff with jitter over hours, fair across tenants
  • Endpoint state checked before each retry

Sending: customers

  • A visible disable rule and an email when it triggers
  • Delivery log with status, latency, and response snippet
  • Replay one delivery or everything since a time
  • Documented retry schedule and verification samples

Sources