Skip to content

Webhooks done right: HMAC-signed, retried, and debuggable — Open Pulse

A webhook is a promise you make to somebody else's infrastructure: show up signed, keep showing up until somebody acknowledges you, and leave a trail when something breaks. The four parts of the sender's contract, what the Svix guide and the Standard Webhooks specification actually recommend, and exactly where Open Pulse's own delivery meets that bar and where it deliberately stops short.

What this page answers

What does this post cover?

A webhook is a promise you make to somebody else's infrastructure: show up signed, keep showing up until somebody acknowledges you, and leave a trail when something breaks. The four parts of the sender's contract, what the Svix guide and the Standard Webhooks specification actually recommend, and exactly where Open Pulse's own delivery meets that bar and where it deliberately stops short.

Who is it for?

Engineers building a webhook sender, and anyone evaluating one before they wire an endpoint to it

Why sign the timestamp as well as the body?

Signing only the body makes every delivery replayable forever. An attacker who captures one valid request can resend it a thousand times and each copy verifies, because nothing in the signed material says when it was made. Binding the timestamp into the signed string is what lets a receiver reject anything older than its tolerance window, which is conventionally five minutes.

Do I have to handle duplicate webhook deliveries?

Yes. At-least-once delivery is the whole design: a sender that retries until it sees a 2xx will send the same event twice whenever the first 2xx is lost in transit. Store the delivery ID on first sight and return 2xx for repeats without re-applying them. Any sender that does not give you a stable ID across retries has handed you a problem you cannot solve.

Should a webhook handler do the work before returning 200?

No. Return 2xx as soon as the event is durably accepted, then do the real work asynchronously. A handler that runs for thirty seconds before acknowledging will trip the sender's per-attempt timeout, which makes it look like a dead endpoint and triggers retries you did not need.

At a glance

Published
2026-09-26
Written for
Engineers building a webhook sender, and anyone evaluating one before they wire an endpoint to it
Category
informative
Reading time
10 minutes
Keywords
webhook best practices, hmac webhook signature, webhook retries, webhook idempotency, webhook delivery log

Overview

A webhook is a promise you make to somebody else's infrastructure. You are saying: when something happens over here, I will show up at your URL, at any hour, in a shape you can trust, and I will keep showing up until you confirm it.

Most of the webhooks that break in production break on one of those clauses. They arrive unverifiable, so the receiver cannot trust them. They arrive once and never again, so a single 500 loses the event forever. Or they fail silently, so nobody learns the endpoint died until a customer asks where their data went.

The sender's half of the contract has four parts: sign every delivery, retry with backoff and know when to stop, give every event an identity so retries are harmless, and log every attempt so failures are debuggable. None of this is novel. It is what the industry learned the hard way, written down.

Sign every delivery, and timestamp the signature

A webhook endpoint is a public URL that accepts POSTs and does something consequential. Anyone on the internet can find it and anyone can call it. The question the receiver has to answer for every delivery is twofold: did this come from the sender it claims, and is it fresh?

The standard answer to the first question is an HMAC signature, and the good implementations bind it to a timestamp to answer the second. Stripe's scheme is the clearest example and worth learning once, because a dozen providers copied its shape. Every delivery carries a Stripe-Signature header with two comma-separated parts, t=<timestamp> and v1=<signature>. The receiver recomputes the signature from the raw body and rejects anything that does not match. The timestamp is what turns a signature into replay protection: the receiver also rejects anything older than a tolerance window, conventionally 300 seconds. An attacker who captures a valid delivery cannot forge one, and cannot replay one indefinitely.

Three details separate careful implementations from careless ones. Sign the raw body, not a parsed and re-serialised one, because re-encoding JSON changes whitespace and key order and silently breaks verification. Compare in constant time, not with string equality. And during a secret rotation, accept both the old and the new signature if the sender transmits them together, or events get dropped in the rotation window.

Retry like the receiver is asleep, not like it is gone

Webhooks are at-least-once delivery by nature. The sender retries until the receiver confirms with a 2xx, which means the receiver must expect duplicates, and the sender must decide how aggressively to retry, for how long, and where to give up.

Svix, which sells webhook infrastructure, publishes the clearest checklist in the category. Retry any non-2xx response and any timeout with exponential backoff. Stop after a fixed number of attempts, so a permanently dead endpoint does not accumulate work forever. Move exhausted deliveries to a dead letter queue instead of dropping them. Three more rules cover the sender's own resources: send asynchronously rather than letting a slow receiver hold a worker, give each attempt a few seconds at most, and record every attempt with its response code and latency so you can tell a broken receiver from a broken sender.

Two design questions matter more than the exact backoff numbers. The first is where you give up, because a retry schedule without a terminal state means failed deliveries accumulate forever. The second is when you stop trying an endpoint entirely rather than a single event. Without some version of that circuit breaker, one abandoned customer integration generates traffic and log volume indefinitely.

For contrast, not every sender retries at all. GitHub's documentation states plainly that it does not automatically redeliver failed deliveries: a 5xx from your endpoint means the event sits there until somebody redelivers it by hand or a scheduled script sweeps the REST API for deliveries whose status is not OK. The point is not that one policy is right. It is that the sender's policy has to be knowable, and the receiver's design has to match it.

Give every event an identity

Retries make duplicates inevitable, so every delivery needs an identity the receiver can store and check. The Standard Webhooks specification, which Svix implements and a growing number of providers follow, ships three headers with each request: webhook-id, webhook-timestamp and webhook-signature. Svix's own deliveries send the same three under a svix- prefix.

Including the ID and the timestamp in the signed content is what makes the signature resistant to replay. An attacker cannot change the timestamp without invalidating the signature, so a receiver that rejects old timestamps has a bounded replay window.

The receiver's half of the contract is to dedupe on that identity and keep its downstream writes idempotent: store the event ID on first sight, and return 2xx for repeats without re-applying them. Our own integration guide recommends deduping on the delivery ID before doing anything else, because the cheapest duplicate is the one you never process.

Log every attempt, both sides of it

The failure mode of webhooks is rarely a dramatic outage. It is a slow drift: deliveries start failing at 3 a.m., nobody notices, and by Monday the question is whether the sender or the receiver is broken. The only way to answer that question is a log of attempts showing both sides of each delivery.

What a usable attempt log needs is small: the timestamp of each attempt, the response code and latency, the request body and the response body. That last pair is the part most senders skip, and it is the part that resolves arguments. "Your payload was malformed" is a claim. The payload, shown next to the 400 the receiver returned, is a fact.

A receiver can hold up its half cheaply: return a 2xx fast once the event is durably accepted, and do the real work asynchronously. A handler that processes for thirty seconds before acknowledging forces the sender's retry machinery to make decisions about an endpoint that might be fine.

How Open Pulse keeps its half of the contract

These four principles are the checklist we built our own delivery against. Here is what it actually does, with the numbers as they are in the code rather than as a marketing page would round them.

Signing. Every delivery carries four headers: x-openpulse-event, x-openpulse-delivery, x-openpulse-timestamp and x-openpulse-signature. The signature is sha256= followed by an HMAC-SHA256 hex digest over <unix seconds>.<raw body>, so the timestamp is inside the signed material rather than beside it. The scheme name is in the value rather than assumed, so the day it moves to another algorithm old subscribers fail loudly instead of verifying against the wrong thing. Secrets are prefixed whsec_ so they are recognisable in a log, and the tolerance window a subscriber should enforce is 300 seconds. We ship the constant-time comparison we use ourselves, because it is the half of the contract subscribers have to implement and shipping it means the documented example is code that is actually tested.

Retries. Three attempts, with 500 ms and then 2 seconds of backoff between them, and an 8 second timeout on each. A 4xx is not retried, because a 4xx means the subscriber understood us and said no: a wrong URL, a revoked route, a rejected signature. Repeating that produces three identical failures and delays the real answer. A 429, a 5xx and any transport failure are retried, because the request was fine and the far side was not. The retry loop lives inside the job rather than in the queue, so a retried delivery keeps its envelope and therefore its idempotency key: the subscriber sees two attempts at one event, never two events.

Identity. Every envelope carries a stable id, sent as x-openpulse-delivery and unchanged across retries. That is the value to dedupe on.

The log. Every attempt is stored with its response code, its latency, the request body and the response body, each truncated at 4,000 bytes. The response body is read even on success, because a 200 with an error body in it is a thing subscribers do and the log is where somebody works that out. The last twenty attempts are returned by default and up to a hundred on request. You can fire a test event while wiring up an endpoint, so setup does not require waiting for a real event to happen. A subscriber's endpoint being down never fails the run that produced the signal.

There are six events, and each maps to something you would act on:

EventFires
signal.createdOnce per kept signal, the first time it is seen.
signal.urgentWhen a signal clears your urgency bar. A strict subset of signal.created, so urgent signals fire both.
run.completedAfter a run finishes, with what it searched and kept.
run.failedWhen a run could not finish, with the error.
pipeline.stage_changedWhen somebody moves a deal, with where it moved from and to.
content.piece.readyWhen a generated draft finishes. One delivery per piece.

The plan difference is the number of endpoints, one on Go and unlimited on Pro, not a different delivery engine. The signature, the retries and the attempt log are the same for everyone, because a webhook that is unverifiable or silent is not a cheaper tier, it is a broken integration.

Where our delivery stops short of the checklist

It would be easy to end there. But the checklist above has six items and we meet four of them, so here is the honest accounting.

Our retry window is seconds, not hours. Svix's guide recommends spreading attempts over hours so a receiver that is briefly down still gets its event. Ours spends about two and a half seconds across three attempts. That is enough to survive a restart or a transient network fault and not enough to survive a deploy that takes a minute. If your endpoint has maintenance windows, treat our deliveries as best-effort during them and reconcile afterwards with the API, which is exactly what the history endpoint is for.

There is no dead letter queue and no replay. An exhausted delivery is recorded as failed with its request and response, which makes it debuggable, but there is no button that sends it again. Reconciliation is a pull, not a redelivery.

There is no auto-disable circuit breaker. An endpoint that has been dead for a month still gets three attempts per event. It costs us rather than you, and it is on the list, but it is not built.

Naming these is the point of the post. If you are evaluating any webhook sender, ours included, the questions are: are deliveries signed with a timestamped signature, what is the retry schedule and where does it give up, does every event carry a stable identity, and can you see every attempt with its request and its response. The answers are what you are buying, whether or not the marketing page mentions them.

References

  • Svix, "How to Build a Webhook Sender" (vendor-authored, and it recommends Svix at the end; the checklist is sound regardless).
  • Standard Webhooks, the open specification behind the webhook-id, webhook-timestamp and webhook-signature headers.
  • Ayi NEDJIMI, "Implementing Webhook Signature Verification (GitHub, Stripe, Slack)", dev.to, July 2026. Source for the Stripe t=/v1= format and the 300 second tolerance.
  • GitHub Docs, "Handling failed webhook deliveries". Source for GitHub not automatically redelivering.
  • Open Pulse, "Wire buying signals into your own stack", the working integration: verify a signed webhook, dedupe on the delivery id, route by event.

Frequently asked questions

Why sign the timestamp as well as the body?

Signing only the body makes every delivery replayable forever. An attacker who captures one valid request can resend it a thousand times and each copy verifies, because nothing in the signed material says when it was made. Binding the timestamp into the signed string is what lets a receiver reject anything older than its tolerance window, which is conventionally five minutes.

Do I have to handle duplicate webhook deliveries?

Yes. At-least-once delivery is the whole design: a sender that retries until it sees a 2xx will send the same event twice whenever the first 2xx is lost in transit. Store the delivery ID on first sight and return 2xx for repeats without re-applying them. Any sender that does not give you a stable ID across retries has handed you a problem you cannot solve.

Should a webhook handler do the work before returning 200?

No. Return 2xx as soon as the event is durably accepted, then do the real work asynchronously. A handler that runs for thirty seconds before acknowledging will trip the sender's per-attempt timeout, which makes it look like a dead endpoint and triggers retries you did not need.

What should I do when a webhook signature does not verify?

Reject with a 4xx and log the raw body you verified against. The overwhelmingly common cause is that the body was parsed and re-serialised by a framework before verification, which changes whitespace and key ordering. The second most common is a secret that was rotated on one side only. Neither is worth retrying, which is why a good sender treats your 4xx as final.

How long should a webhook sender keep retrying?

Long enough to cover a receiver restart, short enough that a dead endpoint does not accumulate work forever, and the schedule should be published either way. What matters more than the numbers is that the sender has a terminal state, that exhausted deliveries land somewhere you can inspect, and that you can tell from the log which of the two of you failed.

Next

  • Start the free trial — Creates a workspace. A card is taken up front and nothing is charged during the trial.
  • Pricing — What it costs after the trial.
  • All posts — The comparisons and the guides, grouped.

See also