A webhook is an HTTP request sent when something happens. Without it, the receiving system often has to poll: ask the sender's API repeatedly whether anything changed. Webhooks reduce that repeated work, but the sender and receiver can fail independently and rarely agree about what "delivered" means.
I have worked on both ends. At Vimeo, meeting-recording integrations received Microsoft Graph change notifications. The sending side in this chapter is an illustrative platform built from public protocols rather than a description of another product's deployment.
On receipt, verify the exact request, store the event before acknowledging it, deduplicate it, and reconcile later. On send, create the event durably, sign it with a timestamp, isolate each endpoint's retries, and give customers a delivery log.
How to read the diagrams
- The other system: the provider sending to you, or the customer endpoint you send to
- Your platform and its records
- Time: deadlines, retry schedules, subscription lifetimes
- Customers and the people debugging a delivery
- Failure: forged requests, lost events, disabled subscriptions
The map below marks where the failures happen. The numbered badges correspond to the order in which a receiver or sender handles an event.
Sender
Provider or your outbox1
creates one event
event id · type · raw body
Signer2
authenticates bytes
timestamp + HMAC
Network boundary
HTTPS request3
headers + body bytes
may arrive zero, one, or many times
Receiver
Verifier4
raw bytes and freshness
signature · timestamp · delivery id
Durable inbox5
deduplicates
UNIQUE (source, event_id)
Background work6
can retry independently
queue · reconciliation · dead letters
Three terms explain most of the design. At-least-once delivery means the sender may send the same event more than once, so the receiver needs an idempotency key, a stable identifier that makes a duplicate produce no second effect. A replay attack is different: someone captures a valid request and deliberately sends it again. Signatures prove who created bytes; timestamps and remembered delivery ids limit how long those bytes remain acceptable.
What the big senders promise
The best way to understand what a receiver must handle is to read what senders actually promise. They promise less than most receivers assume:
- sender
- Stripe
- answer within
- quickly; return 2xx before any slow work
- retries
- up to 3 days, exponential backoff (live mode)
- order
- not guaranteed
- after repeated failure
- email to the account; some billing waits up to 72 hours on invoice.created
- sender
- GitHub
- answer within
- 10 seconds
- retries
- none automatic; manual redelivery for 3 days
- order
- not stated
- after repeated failure
- delivery recorded as failed
- sender
- Shopify
- answer within
- 1 s to connect, 5 s in total
- retries
- 8 times over 4 hours
- order
- not guaranteed
- after repeated failure
- subscription deleted (if created through the Admin API)
- sender
- Slack Events API
- answer within
- 3 seconds
- retries
- 3: immediately, after 1 minute, after 5 minutes
- order
- not stated
- after repeated failure
- subscription disabled if over 95% fail within 60 minutes
- sender
- Twitch EventSub
- answer within
- quickly
- retries
- at least once
- order
- not stated
- after repeated failure
- subscription revoked
- sender
- Svix (a webhook service)
- answer within
- 15 seconds
- retries
- 8 attempts over about 27 hours
- order
- not stated
- after repeated failure
- endpoint disabled after 5 days of failures
The contracts require three design choices. Order cannot be assumed. Duplicates are normal under at-least-once delivery. A receiver that remains slow or unavailable can lose its subscription, after which events stop rather than accumulating forever.
Receiving
The running example is an illustrative Instagram comment delivered to a social inbox. Here is what happens between the request arriving and the 200 going back:
One inbound webhook, from bytes to acknowledgement
- 0 ms
Capture the raw body before any parsing
The signature covers the exact bytes Meta sent. Parsing and re-serialising JSON changes them.
rawBody: 1,284 bytes
- 1 ms
Verify the signature in constant time
HMAC-SHA256 of the raw body with the app secret, compared byte by byte after checking the lengths match.
X-Hub-Signature-256: sha256=9f2c… · valid
- 1 ms
Accept authentic retries idempotently
Meta's signature carries no timestamp. An identical retry is indistinguishable from a captured replay, so the receiver verifies it and relies on the event uniqueness key to make it a no-op.
same external message id → 200 OK, no duplicate work
- 2 ms
Find the tenant from your own records
The payload names the Instagram account; your connection table says which organisation owns it. Nothing in the payload chooses the tenant.
ig account 1784… → conn_91 → org_42
- 4 ms
Store it with a uniqueness key
The event, thread update, and an outbox command for slow work are written in one transaction, with a unique key on the provider's own id.
UNIQUE (connection_id, external_message_id) · inserted
- 5 ms
Acknowledge
200 OK
- later
Do the slow work in the background
Classify the comment, update the conversation, and trigger any configured automation.
queued: background handling
Verify the bytes you received
Every sender signs differently, so each needs its own verifier:
- sender
- Meta (Instagram, Facebook, WhatsApp)
- header
- X-Hub-Signature-256
- signed content
- HMAC-SHA256 of the raw body
- replay window
- none: accept authentic retries and deduplicate effects
- sender
- GitHub
- header
- X-Hub-Signature-256
- signed content
- HMAC-SHA256 of the raw body
- replay window
- none; X-GitHub-Delivery is a unique id
- sender
- Stripe
- header
- Stripe-Signature: t=…,v1=…
- signed content
- HMAC-SHA256 of timestamp.body
- replay window
- 5 minutes by default
- sender
- Slack
- header
- X-Slack-Signature, X-Slack-Request-Timestamp
- signed content
- HMAC-SHA256 of v0:timestamp:body
- replay window
- 5 minutes
- sender
- Twitch EventSub
- header
- Twitch-Eventsub-Message-Signature
- signed content
- HMAC-SHA256 of id + timestamp + body
- replay window
- 10 minutes
- sender
- Discord interactions
- header
- X-Signature-Ed25519
- signed content
- Ed25519 over timestamp + body
- replay window
- timestamp header
- sender
- Twilio
- header
- X-Twilio-Signature
- signed content
- HMAC-SHA1 of the URL and parameters
- replay window
- none
- sender
- Standard Webhooks
- header
- webhook-id, webhook-timestamp, webhook-signature
- signed content
- HMAC-SHA256 (or Ed25519) of id.timestamp.body
- replay window
- timestamp header
An HMAC is a keyed digest: anyone with the shared secret can turn the same bytes into the same fixed-length value. A receiver calculates that value and compares it with the header. A constant-time comparison avoids returning faster for an early mismatch, which could leak information about a valid signature over many measurements.
This illustrative example uses Stripe's documented shape. The secret is deliberately fake, the Unix timestamp is 1790500447, and the raw UTF-8 body is 39 bytes. Stripe's real header is named Stripe-Signature and carries comma-separated t= and v1= fields.
secret: whsec_demo_42
timestamp: 1790500447
raw body: {"id":"evt_demo","type":"lead.created"}
signed text: 1790500447.{"id":"evt_demo","type":"lead.created"}
HMAC-SHA256: bd11f7da27f13e04f9af4aa4625cc5271420199d22ab0567dc5e785572282202
Stripe-Signature: t=1790500447,v1=bd11f7da27f13e04f9af4aa4625cc5271420199d22ab0567dc5e785572282202If a JSON parser reverses the fields and serialises {"type":"lead.created","id":"evt_demo"}, the data still means the same thing to an application. The signed bytes are different, so the digest becomes 537a78…c2433 and verification correctly fails. This is why the request body and headers must be captured before JSON parsing.
Two implementation details cause most verification bugs. First, verify before parsing: capture the raw body in the HTTP framework's body-parser hook, because a JSON parser followed by JSON.stringify produces different bytes and the signature will never match (or, worse, a hand-rolled check will pass on the parsed form). Second, compare in constant time and check lengths first. Node's crypto.timingSafeEqual throws if the two buffers have different lengths, so a header with multibyte characters or the wrong length turns a forged request into an unhandled exception and a 500, which some senders treat as "retry" and some attackers treat as a signal. Compare lengths, then compare bytes.
Where the scheme has a timestamp, reject requests outside the window. Where it does not, signature verification proves who could have created the bytes but cannot prove when they were sent. GitHub supplies a delivery id that can be deduplicated. For Meta, accept an authentic identical delivery, insert with a provider event or message uniqueness key, and answer 200 when the insert conflicts. This preserves Meta's own retry path while ensuring the repeated request causes no repeated effect.
The tenant comes from you, not the payload
A webhook endpoint is public, so everything in the body is attacker-controlled until proven otherwise, including any field that looks like your own organisation id. Resolve the tenant from something you created: the connection that owns the provider account, a per-subscription channel id you registered, or an unguessable secret bound to the subscription. A channel-to-connection mapping prevents an event from choosing another tenant, but a public channel id alone does not authenticate the sender. If a provider offers no signature, use every secret or token its protocol supplies, keep the callback identifier unguessable where the protocol permits it, restrict the accepted request shape, and reconcile against the provider API. Check consistency too: an event that arrives on the Instagram path but names a connection belonging to another provider should be refused and counted.
Acknowledge only what is safe
The sender's timeout forces you to answer quickly, and it is tempting to answer first and work later. Acknowledge only after the event is durable, meaning written to a database or a queue that survives a crash. An endpoint that verifies an event, hands it to an in-memory worker, answers 201, and then crashes has made a promise it cannot keep; the sender will not retry a successful response.
The opposite mistake also causes failures. Meta disables subscriptions that keep failing, so a valid event your platform doesn't understand (a new field, a new event type) should still be stored for inspection, marked unsupported, and acknowledged successfully instead of retried forever.
Handshakes prove that you control the URL
Before sending events, providers use several distinct challenge protocols. Microsoft Graph puts an opaque validationToken in the query string; the endpoint URL-decodes it and returns it as text/plain within ten seconds. Meta sends hub.mode, hub.verify_token, and hub.challenge; after checking the configured verify token, the callback returns the challenge. X sends a CRC token and expects a JSON response_token containing an HMAC-SHA256 result. LinkedIn's validation handshake likewise asks the endpoint to return its challenge plus an HMAC made with the client secret. Each handshake proves control of the callback URL. It does not authenticate later notifications, so the event path still applies that provider's signature or subscription-secret check.
Some protocols ask the callback to sign a challenge. Keep that operation narrow: accept the documented challenge format and length, use the challenge-specific response construction, and reject everything else. An unrestricted endpoint that returns an HMAC for arbitrary input becomes a signing oracle that an attacker can use to obtain valid signatures over chosen bytes.
Subscriptions expire
Some webhook subscriptions are permanent. Others have finite lifetimes: Microsoft Graph subscriptions have short maximum lifetimes that depend on the resource, YouTube's WebSub subscriptions are leases, and CRM "watch" channels expire. A renewal job that fails quietly produces the worst kind of outage, because nothing errors: events stop, and someone notices stale data days later.
Renewal needs its own schedule, well ahead of expiry, and its own alert. The most useful alert is on absence: a source that normally sends events every hour and has sent none for six hours is broken, whatever the renewal job's logs say.
Microsoft Graph can include encrypted resource data when a subscription sets includeResourceData. The subscription registers a public certificate and an identifier. A notification returns that identifier so the receiver can choose the matching private key, unwrap the per-message key, verify the data signature, and decrypt the content. This public Graph contract adds certificate rotation to the renewal design: keep an old private key available while notifications can still reference it.
Reconcile, because some events never come
Even with everything above, some events will be lost: the sender gives up after its retry window, a subscription lapses, an outage on your side lasts longer than their patience. Shopify says plainly that "webhook delivery isn't always guaranteed" and recommends reconciliation jobs. A periodic job that lists recent changes through the sender's API and compares them with what arrived is the only way to close that gap. Stripe's events API returns the last 30 days, which sets how long you can afford to wait before reconciling.
Here is the whole inbound path, including the failures it is built for:
Inbound: a duplicate, a replay, and a normal event
- 12:00:00.000Meta → Webhook endpoint
POST comment event
- Webhook endpoint
raw body · signature valid · not seen before
- 12:00:00.004Webhook endpoint → Postgres
COMMIT event + thread update + outbox
unique (connection, external id) · 1 row
- Webhook endpoint → Meta
200 OK
- laterPostgres → Outbox relay + queue
relay intent analysis
retries until the queue accepts it
- 12:00:09.000Meta → Webhook endpoint
same event again (sender retried)
- Webhook endpoint → Postgres
INSERT … ON CONFLICT DO NOTHING
0 rows: already stored
- Webhook endpoint → Meta
200 OK: nothing new to do
- 12:03:00.000Meta → Webhook endpoint
captured request replayed by someone else
- Webhook endpoint → Postgres
same signed bytes and event id: conflict, no new work
- Webhook endpoint → Meta
200 OK: an authentic repeat is safe to acknowledge
Sending
When your platform sends webhooks, you are the one whose promises customers read. The design problem changes from "accept safely" to "deliver fairly to endpoints you don't control, some of which are down".
Create the event once, durably
An event should be written to a transactional outbox in the same database transaction as the change that caused it, so a crash can't produce a change with no event or an event with no change (the workflow-engine mechanics chapter covers this dual-write problem in detail). The envelope is created once, with a stable id, and every delivery attempt to every endpoint sends that same id, so receivers can deduplicate:
{
"id": "evt_4c1f9a2be07d53a1c6f0",
"type": "lead.created",
"createdAt": "2026-09-27T09:14:07.412Z",
"data": { "leadId": "ld_8812", "formId": "fm_19", "email": "…" }
}Sign it so receivers can check freshness
Use a timestamped signature so the receiver can reject old requests without keeping every request hash. The following illustrative product header uses the same t= and v1= field shape that Stripe documents:
X-Example-Event: lead.created
X-Example-Signature: t=1790500447,v1=5e1b…c09a
^ unix time ^ hex HMAC-SHA256 of "1790500447." + raw bodyRe-sign every attempt with a fresh timestamp, as Stripe does, so a retry three hours later isn't rejected by the receiver's five-minute window. Support secret rotation by signing with both the old and the new secret for a period (Stripe allows up to 24 hours of overlap), sending both signatures so the receiver can accept either. The Standard Webhooks specification, started by Svix with contributors from Zapier, Twilio, Kong, and others, formalises the same ideas: webhook-id, webhook-timestamp, and a space-separated list of signatures over id.timestamp.body, so rotation is built into the format.
Customer URLs are an SSRF risk
A customer-supplied URL is a request your servers will make. Server-side request forgery (SSRF) happens when an attacker chooses that destination and reaches something your public users cannot reach. An attacker might register http://169.254.169.254/latest/meta-data/, the IPv4 address used by Amazon EC2's instance metadata service, its IPv6 counterpart fd00:ec2::254, or https://admin.internal.example/. A naive sender would attach the webhook body and follow the attacker's chosen network path.
- step
- 1 · parse
- candidate
- http://169.254.169.254/latest/meta-data/
- result
- reject: HTTPS required
- step
- 2 · resolve every address
- candidate
- attacker.example → 169.254.169.254 or fd00:ec2::254
- result
- reject: metadata destination
- step
- 3 · connect
- candidate
- https://hooks.customer.example → public IPv4 and IPv6
- result
- pin an approved address for this attempt
- step
- 4 · redirect
- candidate
- 302 Location: https://admin.internal.example/
- result
- refuse, or repeat every validation before following
- step
- 5 · network policy
- candidate
- any missed private or metadata route
- result
- egress firewall denies it
Use a standard URL parser, allow public HTTPS destinations, resolve every IPv4 and IPv6 answer, and block loopback, private, link-local, and metadata ranges. Connect to an address that passed validation while preserving the intended hostname for TLS. Refuse redirects unless the new destination passes the full sequence. A network egress policy is the final containment layer if application validation misses an encoding or address form.
Retry for hours, with jitter, per endpoint
Senders disagree about how long to keep trying, and the range is wide:
How long each sender keeps retrying a failed delivery
Retry window
y: hours
A window of minutes turns a short customer deploy into lost events. A window of days means your queue has to hold days of backlog for a dead endpoint. Exponential backoff increases the delay after each failure. Jitter adds randomness so thousands of delayed deliveries do not wake at the same instant. Choose the window from the product's recovery promise and storage budget, then publish it.
After the final automatic attempt, move the delivery to a dead-letter queue or equivalent failed state: durable storage for work that needs inspection or manual replay. A dead letter is evidence and a recovery point; it should not become a second invisible retry loop.
The deeper lesson comes from Segment, which built a delivery system called Centrifuge to send about 400,000 outbound requests per second. Their write-up explains why a single shared queue fails: if each of 200+ endpoints has an hour of downtime a year, outages occur somewhere almost every day, and one queue lets each outage slow everyone. They also measured how much retries matter: "about 1.5% of all global data succeeds on a retry", and "about 50% of retries succeed only on the third through the tenth attempts." Retrying is a meaningful share of delivered data, and it has to happen without one broken endpoint delaying the rest.
So retry state belongs to the endpoint, not only to the delivery:
Isolate
Each endpoint gets its own concurrency limit and its own position in the retry schedule. A slow endpoint fills its own lane, not a shared one.
Be fair
The retry sweep takes due deliveries round-robin across endpoints and tenants, so one customer's backlog doesn't hold up everyone's first attempts.
Check state at send time
A delivery due for retry checks that its endpoint is still active. Pausing an endpoint should stop its retries as well as new events.
Disable deliberately, and tell someone
An endpoint that has failed every delivery for days consumes queue capacity and leaves the customer's data behind. Large senders publish different rules: Slack can temporarily disable event subscriptions when more than 95% of deliveries fail within 60 minutes, Shopify can delete an Admin API subscription after eight failed attempts, Svix disables an endpoint after its documented multi-day condition, and Standard Webhooks treats 410 Gone as a disable signal. Pick a rule, show it before it triggers, notify the customer, and preserve a clear replay path.
Follow one illustrative endpoint through a six-hour customer outage:
hooks.example.com is down for six hours
- 09:14
First attempt fails
Store the 503 response and schedule from this endpoint's retry state.
evt_4c1f9a2be07d53a1c6f0 · attempt 1 · next attempt includes jitter
- 09:15–12:40
Backoff grows
Fresh deliveries for healthy endpoints continue. This endpoint's failures do not occupy their concurrency.
attempts 2–6 · 503 or timeout
- 15:14
Endpoint recovers
The next due attempt returns 200. The event keeps its original id and receives a fresh signature timestamp.
attempt 7 · 200 OK
- after recovery
Customer reviews and replays
The delivery log shows the outage window. Any delivery that exhausted the policy can be replayed from its dead-letter state.
replay preserves event id for receiver deduplication
If the outage had crossed the platform's disable threshold, the endpoint would move to disabled, the customer would be notified, and no new attempts would run until re-enabled. Re-enabling does not silently discard the backlog; the customer chooses the documented replay window.
Give customers a delivery log
Most webhook support tickets are "did you send it?". A delivery log that answers that without a ticket needs, per attempt: the event id and type, the time, the response status, the response time, and a snippet of the response body, plus a replay button for one delivery and for everything since a point in time. Keep it long enough to cover the retry window and a weekend.
lead.created → https://hooks.example.com/events
endpoint ep_31 · org_42 · signing secret v2
- 09:14:07.9attempt 1503 Service Unavailable · 2,114 ms
- 09:15:21.3attempt 2timeout after 15 s
- 09:17:58.6attempt 3connection refused
- 09:26:40.2attempt 4503 Service Unavailable
- 09:58:11.0attempt 5timeout after 15 s
- 11:42:03.5attempt 6503 Service Unavailable
- 15:14:09.4attempt 7200 OK · 184 ms · endpoint recovered after six hours
Stripe's documentation mentions a load pattern worth planning for: deliveries spike "during the beginning of the month when all subscriptions renew." Any platform with scheduled business events has its own version, and the delivery system needs headroom for it.
A checklist
Receiving: verification
- Raw body captured before parsing
- One verifier per scheme, constant-time, lengths checked first
- Timestamp window where available; stable delivery or event ids deduplicate effects
- Challenge endpoints accept only the sender's exact format
Receiving: handling
- Tenant resolved from your records, never the payload
- Acknowledge only after the event is durable
- Unique key on the provider's event id
- Unknown but valid events acknowledged and logged
Receiving: over time
- No reliance on order: versions or re-fetch
- Subscription renewal scheduled well before expiry
- Alerts on silence from sources that normally send
- Reconciliation against the sender's API
Sending: events
- Outbox in the same transaction as the change
- One envelope with a stable id for every attempt
- Timestamped signature, re-signed per attempt
- Two signatures during secret rotation
Sending: delivery
- Public HTTPS only, checked at save and at send, no redirects
- Per-endpoint concurrency and retry state
- Backoff with jitter over hours, fair across tenants
- Endpoint state checked before each retry
Sending: customers
- A visible disable rule and an email when it triggers
- Delivery log with status, latency, and response snippet
- Replay one delivery or everything since a time
- Documented retry schedule and verification samples
Sources
- Stripe, Webhooks, Subscription webhooks, and Process undelivered events; Brandur Leach, idempotency design
- GitHub, Handling failed webhook deliveries and Redelivering webhooks
- Shopify, HTTPS webhooks and Webhook best practices
- Slack, Events API and Verifying requests
- Microsoft Graph, webhook delivery and validation and change notifications with resource data
- Twitch, Handling webhook events; Twilio, Webhooks security
- Standard Webhooks specification; Svix, Retries
- Amazon EC2, Instance Metadata Service
- Segment, Introducing Centrifuge (2018)