On this page
Key takeaways
- Buyers experience webhooks as product, not plumbing: the contract decides whether an integration is trusted.
- A good event is signed, versioned, named stably, and replayable without a second API.
- Design for at-least-once delivery: duplicates and reordering are normal, so handlers must be idempotent.
- If you cannot explain what happens when a delivery fails twice, you do not have an integration.
Teams talk about webhooks as plumbing. Buyers feel them as product. A signed, versioned, replayable event is the difference between a gateway you can trust and a black box you poll.
This note lays out how we think about the event contract in StablePay: the lifecycle, what a delivery must guarantee, what the receiver must do, and the failure modes worth designing for on purpose.
Why the contract is the product
The host application owns the customer, the order, and the business logic. The gateway owns payment state. The webhook is the only place those two meet in real time. If it is vague, every integration invents its own interpretation, and every interpretation fails in a different way.
- Stable names. An event called
payment.confirmednever quietly changes meaning. - Explicit lifecycle. Every payment moves through defined states in a defined order.
- Verifiable origin. The receiver can prove an event came from the gateway without calling it.
- Recoverable. A missed delivery can be replayed without forging anything.
The lifecycle
StablePay treats a payment as five stages: create, track, listen, reconcile, settle. Webhooks cover the middle of that arc — the moments where the application needs to change something.
| Stage | What the gateway does | What the application does |
|---|---|---|
| Create | Records an intent, link, or invoice | Stores its own reference |
| Track | Watches the network for the payment | Nothing — waits |
| Listen | Emits signed lifecycle events | Verifies, updates the order, returns 2xx |
| Reconcile | Exposes state in dashboard and API | Compares its record to the gateway's |
| Settle | Marks settlement complete | Releases goods or credit |
Delivery semantics: assume at-least-once
Any honest webhook system is at-least-once. Networks fail, endpoints deploy, and timeouts happen after work was done. Exactly-once is a property you build in the receiver, not one you can ask the sender for.
What that means for the sender
- Retry with backoff when the receiver does not return 2xx.
- Give every event a stable identifier so duplicates can be recognised.
- Classify retries, so operators can tell a first attempt from a redelivery.
What that means for the receiver
- Verify the signature before doing anything else.
- Store the event identifier and ignore repeats.
- Acknowledge quickly, then process — do not do slow work inside the request.
- Never assume events arrive in order; compare state, not sequence.
Verifying a signature
The receiver should be able to verify without a network call. The general pattern is an HMAC over the raw request body with a shared secret, compared in constant time. The snippet below shows the pattern only — header and field names are illustrative, so check the product documentation for the real ones.
import { createHmac, timingSafeEqual } from "node:crypto";
export function verifySignature(rawBody: string, signature: string, secret: string) {
const expected = createHmac("sha256", secret).update(rawBody).digest("hex");
const a = Buffer.from(expected);
const b = Buffer.from(signature);
return a.length === b.length && timingSafeEqual(a, b);
}
export async function handle(req: Request) {
const raw = await req.text(); // verify against the RAW body, not re-serialised JSON
const sig = req.headers.get("x-signature") ?? "";
if (!verifySignature(raw, sig, process.env.WEBHOOK_SECRET!)) {
return new Response("invalid signature", { status: 401 });
}
const event = JSON.parse(raw);
// 1. dedupe on event id 2. enqueue for processing 3. return 2xx fast
return new Response(null, { status: 204 });
}The classic mistake
Verifying against a parsed-then-re-serialised body. Whitespace and key order change, the HMAC no longer matches, and the bug appears only in production.
Fragile handler
- Parses JSON, then verifies
- Does slow work before responding
- Assumes events arrive in order
- Treats duplicates as errors
Robust handler
- Verifies the raw body first
- Acknowledges, then processes
- Compares state, not sequence
- Returns 2xx for duplicates
Failure modes worth designing for
| Failure | What happens without a plan | What the contract should provide |
|---|---|---|
| Endpoint down for an hour | Events lost or duplicated unpredictably | Retries with backoff and a replay path |
| Handler crashes mid-way | Order half-updated | Idempotent handlers and stable event ids |
| Duplicate delivery | Customer charged or fulfilled twice | Dedupe on event id |
| Out-of-order events | State goes backwards | Compare state, not arrival order |
| Forged request | Attacker marks an order paid | Signature verification on the raw body |
| Silent gap | Nobody notices a missed event | Reconciliation against the gateway's state |
Versioning the contract
An event contract is a promise that lasts longer than the code that emits it. Three rules keep it honest:
- Add, do not repurpose. New fields are safe. A field that changes meaning is a new field.
- Tolerate the unknown. Receivers should ignore fields they do not recognise, so additive changes never break them.
- Announce removals early. Anything deprecated is written in the changelog with a date, before it disappears.
| Change | Safe for receivers? | How we ship it |
|---|---|---|
| New optional field | Yes | Any release, noted in the changelog |
| New event type | Yes, if unknown events are ignored | Documented before it is emitted |
| Field meaning changes | No | Do not — add a new field instead |
| Event removed | No | Deprecation notice, then removal |
Replay on purpose
Release 0.9.4 added something we had wanted for a long time: deliveries carry a retry class and a replay token that operators can use without forging events. The principle is that recovery should use the same contract as normal operation. There should be no second, privileged API that quietly bypasses verification.
The dashboard also groups silent confirmations — on-chain yes, application no — so the disagreement between the two records is visible rather than discovered.
A receiver checklist
- Verify the signature against the raw body.
- Reject anything older than your tolerance window.
- Deduplicate on the event identifier.
- Persist the event, then acknowledge with 2xx.
- Process asynchronously and idempotently.
- Reconcile periodically against the gateway state.
If you cannot explain what happens when a delivery fails twice, you do not have an integration. You have a hope with a retry counter.
Designing the retry schedule
Retries are a compromise between two failures: giving up too early and losing an event, or retrying so aggressively that you turn a receiver's outage into a self-inflicted denial of service. The standard answer is exponential backoff with jitter [3], and the reasoning matters more than the constants.
Exponential backoff spaces attempts further apart each time, giving a struggling receiver room to recover. Jitter randomises each delay so that thousands of failed deliveries do not all retry at the same instant. Without jitter, a brief outage produces a synchronised thundering herd on recovery.
function nextDelayMs(attempt: number, baseMs = 5_000, capMs = 6 * 60 * 60 * 1000) {
// exponential ceiling: 5s, 10s, 20s, 40s ... capped at 6 hours
const ceiling = Math.min(capMs, baseMs * 2 ** attempt);
// "full jitter": pick uniformly between 0 and the ceiling
return Math.floor(Math.random() * ceiling);
}
// Example schedule ceilings for attempts 0..8:
// 5s, 10s, 20s, 40s, 80s, 160s, 320s, 640s, 1280s| Decision | Common choice | Trade-off |
|---|---|---|
| Max attempts | 8–15 over 24–72 hours | Longer windows recover from long outages but delay dead-lettering |
| Timeout per attempt | 5–10 seconds | Short timeouts fail fast; too short and slow-but-healthy receivers look broken |
| Which statuses retry | Network errors, timeouts, 5xx, 429 | 4xx other than 429 usually will not succeed on retry |
| Honour Retry-After | Yes for 429 and 503 | Respecting the receiver's hint prevents making an outage worse |
| Dead-letter after exhaustion | Keep the event, expose it in the dashboard | Losing an exhausted event silently is the worst outcome |
Ordering, and why you should not rely on it
Events are emitted in order but delivered over an unreliable network with retries, so receivers can and will see them out of order: a payment.confirmed arriving before the payment.pending that logically preceded it, or a retried old event arriving after a newer one. Designing for this is cheaper than trying to guarantee ordering.
- Make handlers state-based. Compare the state carried in the event with the state you hold and apply only forward transitions.
- Carry a monotonic version or timestamp. A sequence or updated-at field lets the receiver discard stale events.
- Keep the payment state machine explicit. A payment cannot go from
confirmedback topending; a handler that enforces this is immune to reordering.
const order = ["created", "seen", "confirming", "confirmed", "settled"] as const;
type State = (typeof order)[number];
function applyEvent(current: State, incoming: State): State {
// ignore stale or duplicate events; never move backwards
return order.indexOf(incoming) > order.indexOf(current) ? incoming : current;
}Replay protection and timestamp tolerance
A valid signature proves an event came from the gateway. It does not prove the event is fresh. An attacker who captures a signed request could replay it later. The standard defence is to include a timestamp in the signed content and reject deliveries outside a tolerance window, commonly a few minutes [1][5].
import { createHmac, timingSafeEqual } from "node:crypto";
const TOLERANCE_SECONDS = 300;
export function verify(rawBody: string, timestamp: string, signature: string, secret: string) {
const age = Math.abs(Date.now() / 1000 - Number(timestamp));
if (!Number.isFinite(age) || age > TOLERANCE_SECONDS) return false; // stale or malformed
const expected = createHmac("sha256", secret)
.update(`${timestamp}.${rawBody}`) // sign the timestamp too, or it can be swapped
.digest("hex");
const a = Buffer.from(expected);
const b = Buffer.from(signature);
return a.length === b.length && timingSafeEqual(a, b); // constant-time compare [4]
}Do not confuse the two protections
Signature verification stops forgery. Timestamp tolerance stops replay of a real message. Event-id deduplication stops legitimate redelivery from being processed twice. You need all three, and each answers a different question.
Rotating secrets without downtime
Secrets leak and people leave, so rotation has to be a routine operation, not an incident. The difficulty is the overlap: at the moment of rotation, in-flight deliveries were signed with the old secret. The solution is to accept two secrets during a transition window.
- 01
Generate a new secret
Create it in the gateway alongside the current one. Both are now valid for verification.
- 02
Deploy the receiver accepting both
The handler tries the new secret, then the old one, and accepts either.
- 03
Switch signing to the new secret
New deliveries are signed only with the new secret.
- 04
Wait out the retry window
Any retried delivery signed with the old secret still verifies.
- 05
Remove the old secret
Only after the longest retry window has passed.
Thin events or fat events?
Should an event carry the whole payment object, or just an identifier and a type? The choice shapes security, freshness, and coupling.
| Thin events (id and type) | Fat events (full snapshot) | |
|---|---|---|
| Payload size | Small | Larger |
| Freshness | Receiver fetches current state, so always up to date | May be stale by the time it is processed |
| Coupling | Receiver depends on an extra API call | Receiver depends on the payload schema |
| Failure mode | API outage blocks processing | Schema change breaks parsing |
| Sensitive data | Less exposed in transit and logs | More exposed |
A pragmatic middle path is a snapshot with a version field, plus a fetch endpoint for the authoritative state. The event tells you something changed and roughly what; the API tells you the truth. Standard Webhooks [1] takes the same view of a consistent envelope with signing headers, and following an existing convention lowers the integration cost for every receiver.
Observability for deliveries
Every delivery should be answerable with one query: what was sent, when, to whom, with what result. Without that, debugging an integration is guesswork.
| Field | Why it is stored |
|---|---|
| event id and type | Deduplication and search |
| attempt number and scheduled time | Understanding retry behaviour |
| request headers (signature redacted) | Reproducing a failed delivery |
| response status and truncated body | Seeing what the receiver said |
| duration | Spotting slow receivers before they time out |
| retry class | Distinguishing first attempts from redeliveries and replays |
Dashboard views worth building
- Deliveries failing right now, grouped by endpoint.
- Events that exhausted retries and need attention.
- Silent confirmations: on-chain yes, application no.
- Median and p95 latency per endpoint.
- A one-click replay with an audit record of who replayed what.
Testing webhooks locally and in CI
- Record real payloads. Capture examples from staging and commit them as fixtures, with secrets removed. Your tests then exercise the true shape, not an imagined one.
- Test the signature path with bad input. A body altered by one byte, a stale timestamp, a missing header — each should be rejected with a clear status.
- Simulate duplicates and reordering. Feed the same event twice and events out of order, and assert the final state.
- Use a tunnel for manual testing. A local tunnel to your development machine lets the staging gateway deliver to your laptop; treat that endpoint as public and short-lived.
Closing
If you cannot explain what happens when a delivery fails twice, you do not have an integration. You have a hope with a retry counter. Treat the event contract as public surface, document it like one, and the night jobs disappear.
References & further reading
- 1Standard Webhooks — Standard Webhooks specificationA vendor-neutral convention for webhook signing and delivery.
- 2RFC 2104: HMAC: Keyed-Hashing for Message Authentication — H. Krawczyk, M. Bellare, R. Canetti, IETF, 1997
- 3Exponential Backoff And Jitter — Marc Brooker, AWS Architecture Blog
- 4crypto.timingSafeEqual(a, b) — Node.js documentation
- 5Validating webhook deliveries — GitHub DocsA widely used example of HMAC verification of a raw body.
- 6Receive Stripe events in your webhook endpoint — Stripe DocumentationCovers signature verification, tolerance, and handling duplicates.
The product behind this post
StablePay
Self-hosted stablecoin payment infrastructure.
From $4,800 one-time license
About the author
Product behaviour described here reflects what is implemented and tested; anything else is marked as planned. Code samples are illustrative.
All writing