Engineering · 9 Sep 2025

Webhooks are a product surface.

If the event contract is vague, the integration will be a night job forever.

Published
9 Sep 2025
Reading time
9 min
Readers
—
Topics
StablePayWebhooksAPI designReliability
On this page

Key takeaways

  • Buyers experience webhooks as product, not plumbing: the contract decides whether an integration is trusted.
  • A good event is signed, versioned, named stably, and replayable without a second API.
  • Design for at-least-once delivery: duplicates and reordering are normal, so handlers must be idempotent.
  • If you cannot explain what happens when a delivery fails twice, you do not have an integration.

Teams talk about webhooks as plumbing. Buyers feel them as product. A signed, versioned, replayable event is the difference between a gateway you can trust and a black box you poll.

This note lays out how we think about the event contract in StablePay: the lifecycle, what a delivery must guarantee, what the receiver must do, and the failure modes worth designing for on purpose.

Why the contract is the product

The host application owns the customer, the order, and the business logic. The gateway owns payment state. The webhook is the only place those two meet in real time. If it is vague, every integration invents its own interpretation, and every interpretation fails in a different way.

  • Stable names. An event called payment.confirmed never quietly changes meaning.
  • Explicit lifecycle. Every payment moves through defined states in a defined order.
  • Verifiable origin. The receiver can prove an event came from the gateway without calling it.
  • Recoverable. A missed delivery can be replayed without forging anything.

The lifecycle

StablePay treats a payment as five stages: create, track, listen, reconcile, settle. Webhooks cover the middle of that arc — the moments where the application needs to change something.

StageWhat the gateway doesWhat the application does
CreateRecords an intent, link, or invoiceStores its own reference
TrackWatches the network for the paymentNothing — waits
ListenEmits signed lifecycle eventsVerifies, updates the order, returns 2xx
ReconcileExposes state in dashboard and APICompares its record to the gateway's
SettleMarks settlement completeReleases goods or credit
CreateIntent, link, or invoice
TrackNetwork watched, state recorded
ListenSigned event fired
ReconcileCompare records
SettleFunds where treasury expects
Where the webhook sits in the payment lifecycle

Delivery semantics: assume at-least-once

Any honest webhook system is at-least-once. Networks fail, endpoints deploy, and timeouts happen after work was done. Exactly-once is a property you build in the receiver, not one you can ask the sender for.

What that means for the sender

  • Retry with backoff when the receiver does not return 2xx.
  • Give every event a stable identifier so duplicates can be recognised.
  • Classify retries, so operators can tell a first attempt from a redelivery.

What that means for the receiver

  • Verify the signature before doing anything else.
  • Store the event identifier and ignore repeats.
  • Acknowledge quickly, then process — do not do slow work inside the request.
  • Never assume events arrive in order; compare state, not sequence.
Signed event arrivesRaw body + signature header
VerifyHMAC over the raw body, constant-time compare
DedupeHave we seen this event id?
Persist, then 2xxAcknowledge fast
Process asyncIdempotent handler
One delivery, handled correctly

Verifying a signature

The receiver should be able to verify without a network call. The general pattern is an HMAC over the raw request body with a shared secret, compared in constant time. The snippet below shows the pattern only — header and field names are illustrative, so check the product documentation for the real ones.

Illustrative — the pattern, not the exact contractTypeScript
import { createHmac, timingSafeEqual } from "node:crypto";

export function verifySignature(rawBody: string, signature: string, secret: string) {
  const expected = createHmac("sha256", secret).update(rawBody).digest("hex");
  const a = Buffer.from(expected);
  const b = Buffer.from(signature);
  return a.length === b.length && timingSafeEqual(a, b);
}

export async function handle(req: Request) {
  const raw = await req.text(); // verify against the RAW body, not re-serialised JSON
  const sig = req.headers.get("x-signature") ?? "";
  if (!verifySignature(raw, sig, process.env.WEBHOOK_SECRET!)) {
    return new Response("invalid signature", { status: 401 });
  }
  const event = JSON.parse(raw);
  // 1. dedupe on event id  2. enqueue for processing  3. return 2xx fast
  return new Response(null, { status: 204 });
}

The classic mistake

Verifying against a parsed-then-re-serialised body. Whitespace and key order change, the HMAC no longer matches, and the bug appears only in production.

Fragile handler

  • Parses JSON, then verifies
  • Does slow work before responding
  • Assumes events arrive in order
  • Treats duplicates as errors

Robust handler

  • Verifies the raw body first
  • Acknowledges, then processes
  • Compares state, not sequence
  • Returns 2xx for duplicates

Failure modes worth designing for

FailureWhat happens without a planWhat the contract should provide
Endpoint down for an hourEvents lost or duplicated unpredictablyRetries with backoff and a replay path
Handler crashes mid-wayOrder half-updatedIdempotent handlers and stable event ids
Duplicate deliveryCustomer charged or fulfilled twiceDedupe on event id
Out-of-order eventsState goes backwardsCompare state, not arrival order
Forged requestAttacker marks an order paidSignature verification on the raw body
Silent gapNobody notices a missed eventReconciliation against the gateway's state

Versioning the contract

An event contract is a promise that lasts longer than the code that emits it. Three rules keep it honest:

  • Add, do not repurpose. New fields are safe. A field that changes meaning is a new field.
  • Tolerate the unknown. Receivers should ignore fields they do not recognise, so additive changes never break them.
  • Announce removals early. Anything deprecated is written in the changelog with a date, before it disappears.
ChangeSafe for receivers?How we ship it
New optional fieldYesAny release, noted in the changelog
New event typeYes, if unknown events are ignoredDocumented before it is emitted
Field meaning changesNoDo not — add a new field instead
Event removedNoDeprecation notice, then removal

Replay on purpose

Release 0.9.4 added something we had wanted for a long time: deliveries carry a retry class and a replay token that operators can use without forging events. The principle is that recovery should use the same contract as normal operation. There should be no second, privileged API that quietly bypasses verification.

The dashboard also groups silent confirmations — on-chain yes, application no — so the disagreement between the two records is visible rather than discovered.

A receiver checklist

  1. Verify the signature against the raw body.
  2. Reject anything older than your tolerance window.
  3. Deduplicate on the event identifier.
  4. Persist the event, then acknowledge with 2xx.
  5. Process asynchronously and idempotently.
  6. Reconcile periodically against the gateway state.
If you cannot explain what happens when a delivery fails twice, you do not have an integration. You have a hope with a retry counter.

Designing the retry schedule

Retries are a compromise between two failures: giving up too early and losing an event, or retrying so aggressively that you turn a receiver's outage into a self-inflicted denial of service. The standard answer is exponential backoff with jitter [3], and the reasoning matters more than the constants.

Exponential backoff spaces attempts further apart each time, giving a struggling receiver room to recover. Jitter randomises each delay so that thousands of failed deliveries do not all retry at the same instant. Without jitter, a brief outage produces a synchronised thundering herd on recovery.

Illustrative — full-jitter backoff with a capTypeScript
function nextDelayMs(attempt: number, baseMs = 5_000, capMs = 6 * 60 * 60 * 1000) {
  // exponential ceiling: 5s, 10s, 20s, 40s ... capped at 6 hours
  const ceiling = Math.min(capMs, baseMs * 2 ** attempt);
  // "full jitter": pick uniformly between 0 and the ceiling
  return Math.floor(Math.random() * ceiling);
}

// Example schedule ceilings for attempts 0..8:
// 5s, 10s, 20s, 40s, 80s, 160s, 320s, 640s, 1280s
DecisionCommon choiceTrade-off
Max attempts8–15 over 24–72 hoursLonger windows recover from long outages but delay dead-lettering
Timeout per attempt5–10 secondsShort timeouts fail fast; too short and slow-but-healthy receivers look broken
Which statuses retryNetwork errors, timeouts, 5xx, 4294xx other than 429 usually will not succeed on retry
Honour Retry-AfterYes for 429 and 503Respecting the receiver's hint prevents making an outage worse
Dead-letter after exhaustionKeep the event, expose it in the dashboardLosing an exhausted event silently is the worst outcome

Ordering, and why you should not rely on it

Events are emitted in order but delivered over an unreliable network with retries, so receivers can and will see them out of order: a payment.confirmed arriving before the payment.pending that logically preceded it, or a retried old event arriving after a newer one. Designing for this is cheaper than trying to guarantee ordering.

  • Make handlers state-based. Compare the state carried in the event with the state you hold and apply only forward transitions.
  • Carry a monotonic version or timestamp. A sequence or updated-at field lets the receiver discard stale events.
  • Keep the payment state machine explicit. A payment cannot go from confirmed back to pending; a handler that enforces this is immune to reordering.
Illustrative — apply an event only if it moves the payment forwardTypeScript
const order = ["created", "seen", "confirming", "confirmed", "settled"] as const;
type State = (typeof order)[number];

function applyEvent(current: State, incoming: State): State {
  // ignore stale or duplicate events; never move backwards
  return order.indexOf(incoming) > order.indexOf(current) ? incoming : current;
}

Replay protection and timestamp tolerance

A valid signature proves an event came from the gateway. It does not prove the event is fresh. An attacker who captures a signed request could replay it later. The standard defence is to include a timestamp in the signed content and reject deliveries outside a tolerance window, commonly a few minutes [1][5].

Illustrative — verify signature and freshness togetherTypeScript
import { createHmac, timingSafeEqual } from "node:crypto";

const TOLERANCE_SECONDS = 300;

export function verify(rawBody: string, timestamp: string, signature: string, secret: string) {
  const age = Math.abs(Date.now() / 1000 - Number(timestamp));
  if (!Number.isFinite(age) || age > TOLERANCE_SECONDS) return false; // stale or malformed

  const expected = createHmac("sha256", secret)
    .update(`${timestamp}.${rawBody}`) // sign the timestamp too, or it can be swapped
    .digest("hex");
  const a = Buffer.from(expected);
  const b = Buffer.from(signature);
  return a.length === b.length && timingSafeEqual(a, b); // constant-time compare [4]
}

Do not confuse the two protections

Signature verification stops forgery. Timestamp tolerance stops replay of a real message. Event-id deduplication stops legitimate redelivery from being processed twice. You need all three, and each answers a different question.

Rotating secrets without downtime

Secrets leak and people leave, so rotation has to be a routine operation, not an incident. The difficulty is the overlap: at the moment of rotation, in-flight deliveries were signed with the old secret. The solution is to accept two secrets during a transition window.

  1. 01

    Generate a new secret

    Create it in the gateway alongside the current one. Both are now valid for verification.

  2. 02

    Deploy the receiver accepting both

    The handler tries the new secret, then the old one, and accepts either.

  3. 03

    Switch signing to the new secret

    New deliveries are signed only with the new secret.

  4. 04

    Wait out the retry window

    Any retried delivery signed with the old secret still verifies.

  5. 05

    Remove the old secret

    Only after the longest retry window has passed.

Thin events or fat events?

Should an event carry the whole payment object, or just an identifier and a type? The choice shapes security, freshness, and coupling.

Thin events (id and type)Fat events (full snapshot)
Payload sizeSmallLarger
FreshnessReceiver fetches current state, so always up to dateMay be stale by the time it is processed
CouplingReceiver depends on an extra API callReceiver depends on the payload schema
Failure modeAPI outage blocks processingSchema change breaks parsing
Sensitive dataLess exposed in transit and logsMore exposed

A pragmatic middle path is a snapshot with a version field, plus a fetch endpoint for the authoritative state. The event tells you something changed and roughly what; the API tells you the truth. Standard Webhooks [1] takes the same view of a consistent envelope with signing headers, and following an existing convention lowers the integration cost for every receiver.

Observability for deliveries

Every delivery should be answerable with one query: what was sent, when, to whom, with what result. Without that, debugging an integration is guesswork.

FieldWhy it is stored
event id and typeDeduplication and search
attempt number and scheduled timeUnderstanding retry behaviour
request headers (signature redacted)Reproducing a failed delivery
response status and truncated bodySeeing what the receiver said
durationSpotting slow receivers before they time out
retry classDistinguishing first attempts from redeliveries and replays

Dashboard views worth building

  • Deliveries failing right now, grouped by endpoint.
  • Events that exhausted retries and need attention.
  • Silent confirmations: on-chain yes, application no.
  • Median and p95 latency per endpoint.
  • A one-click replay with an audit record of who replayed what.

Testing webhooks locally and in CI

  • Record real payloads. Capture examples from staging and commit them as fixtures, with secrets removed. Your tests then exercise the true shape, not an imagined one.
  • Test the signature path with bad input. A body altered by one byte, a stale timestamp, a missing header — each should be rejected with a clear status.
  • Simulate duplicates and reordering. Feed the same event twice and events out of order, and assert the final state.
  • Use a tunnel for manual testing. A local tunnel to your development machine lets the staging gateway deliver to your laptop; treat that endpoint as public and short-lived.

Closing

If you cannot explain what happens when a delivery fails twice, you do not have an integration. You have a hope with a retry counter. Treat the event contract as public surface, document it like one, and the night jobs disappear.

References & further reading

  1. 1
    Standard Webhooks — Standard Webhooks specificationA vendor-neutral convention for webhook signing and delivery.
  2. 2
    RFC 2104: HMAC: Keyed-Hashing for Message Authentication — H. Krawczyk, M. Bellare, R. Canetti, IETF, 1997
  3. 3
    Exponential Backoff And Jitter — Marc Brooker, AWS Architecture Blog
  4. 4
    crypto.timingSafeEqual(a, b) — Node.js documentation
  5. 5
    Validating webhook deliveries — GitHub DocsA widely used example of HMAC verification of a raw body.
  6. 6
    Receive Stripe events in your webhook endpoint — Stripe DocumentationCovers signature verification, tolerance, and handling duplicates.
Found this useful? Share it

Get the next essay in your inbox

Practical writing on payments infrastructure, operations software, and shipping real systems. No spam, no sales sequence.

We only use your email to send the studio's writing. See the privacy policy.

About the author

Product behaviour described here reflects what is implemented and tested; anything else is marked as planned. Code samples are illustrative.

All writing

Have a system like this to run?

Explore the catalog, or write down the problem and the constraints. We respond when the fit is real.

Free apps from the studio. Enter your email, get a private download link. Free for personal use.

Get them free

Have a product to sell? We review, list, and sell it for you — you keep 90% of every sale.

Apply to sell with us