Engineering · 2 Apr 2026

Building a webhook delivery system: the sender's side of the contract.

Outbox, queue, signing, retries, dead letters, and observability — how to send events reliably to receivers you do not control.

Published
2 Apr 2026
Reading time
7 min
Readers
—
Topics
StablePayWebhooksArchitectureReliabilityPostgreSQL
On this page

Key takeaways

  • Reliable delivery starts in the sender's database: write the event and the intent to send it in one transaction.
  • A delivery is a state machine with attempts, not a fire-and-forget HTTP call.
  • Sign what you send, bound what you retry, and never lose an exhausted event silently.
  • Give operators — and receivers — a delivery log and a safe replay path.

Most writing about webhooks is aimed at the receiver: how to verify, deduplicate, and respond quickly. This post is about the other side — the system that sends. If you run a gateway, an operations platform, or any product that tells other systems when something happened, you are a webhook sender, and the reliability of your customers' integrations depends on decisions made in your delivery pipeline.

We will design one from the database outward: how an event is recorded, how a delivery is scheduled and attempted, how retries and failures are handled, and how operators and receivers see what happened. The same design underpins the signed events StablePay emits, and much of it applies to any sender.

What a sender promises

PromiseWhat it means in practice
At-least-once deliveryEvery event is delivered at least once, possibly more; receivers deduplicate
DurabilityAn event that was committed is never lost, even if the sender crashes
AuthenticityReceivers can verify the event came from you and is fresh
Bounded impactA slow or broken receiver cannot harm other receivers or your own system
VisibilityAnyone can see what was sent, when, and what came back
RecoverabilityMissed events can be replayed without forging anything

No exactly-once

Do not promise exactly-once. Networks fail in ways that make it impossible to guarantee. Promise at-least-once, give every event a stable id, and document that receivers must be idempotent.

Step one: write the event and the intent together

The most common cause of lost events is the dual write: a service updates its database and then sends a message, and crashes in between. The transactional outbox [2] removes the gap by writing the event to an outbox table in the same transaction as the state change. A separate process then turns outbox rows into deliveries.

Business changePayment confirmed
Same transactionState row plus event row
DispatcherReads new events
Fan-outOne delivery per subscribed endpoint
AttemptHTTP POST to the receiver
The outbox pattern
Illustrative — events and deliveriesSQL
CREATE TABLE events (
  id          uuid PRIMARY KEY DEFAULT gen_random_uuid(),
  type        text NOT NULL,                 -- e.g. payment.confirmed
  payload     jsonb NOT NULL,
  created_at  timestamptz NOT NULL DEFAULT now()
);

CREATE TABLE endpoints (
  id       uuid PRIMARY KEY,
  url      text NOT NULL,
  secrets  text[] NOT NULL,                  -- current and previous, for rotation
  enabled  boolean NOT NULL DEFAULT true,
  types    text[] NOT NULL DEFAULT '{}'      -- subscribed event types (empty = all)
);

CREATE TYPE delivery_status AS ENUM ('pending','succeeded','failed','exhausted');

CREATE TABLE deliveries (
  id            uuid PRIMARY KEY DEFAULT gen_random_uuid(),
  event_id      uuid NOT NULL REFERENCES events(id),
  endpoint_id   uuid NOT NULL REFERENCES endpoints(id),
  status        delivery_status NOT NULL DEFAULT 'pending',
  attempts      int NOT NULL DEFAULT 0,
  next_attempt  timestamptz NOT NULL DEFAULT now(),
  last_status   int,
  last_error    text,
  UNIQUE (event_id, endpoint_id)             -- one delivery per event per endpoint
);
CREATE INDEX deliveries_due ON deliveries (next_attempt) WHERE status IN ('pending','failed');

Step two: a queue built on the database

You do not need a separate message broker to start. A table with a due-time and SELECT ... FOR UPDATE SKIP LOCKED [4] gives you a safe work queue: several workers can pull due deliveries concurrently without two ever taking the same row.

Illustrative — a worker claims and processes due deliveriesTypeScript
async function workOnce() {
  const batch = await db.transaction(async (tx) => {
    const { rows } = await tx.query(
      `SELECT d.id, d.event_id, d.attempts, e.url, e.secrets, ev.type, ev.payload, ev.created_at
         FROM deliveries d
         JOIN endpoints e ON e.id = d.endpoint_id AND e.enabled
         JOIN events ev   ON ev.id = d.event_id
        WHERE d.status IN ('pending','failed') AND d.next_attempt <= now()
        ORDER BY d.next_attempt
        LIMIT 20
        FOR UPDATE OF d SKIP LOCKED`,
    );
    // push next_attempt forward so a crashed worker's rows become due again
    await tx.query(`UPDATE deliveries SET next_attempt = now() + interval '2 minutes' WHERE id = ANY($1)`, [rows.map((r) => r.id)]);
    return rows;
  });
  await Promise.allSettled(batch.map(attempt));
}

The two-minute lease is a deliberate trick. If the worker dies mid-attempt, the row becomes due again by itself; no separate recovery job is needed. Choose a lease longer than your request timeout.

When you outgrow the table queueSymptomMove to
Very high event volumePolling load or lock contention on the deliveries tableA dedicated queue or stream, keeping the outbox as the source of truth
Many independent consumersFan-out logic getting complicatedA message broker with topics
Strict per-endpoint orderingHead-of-line concernsPartitioned queues keyed by endpoint

Step three: sign what you send

Receivers should be able to verify authenticity and freshness without calling you. The conventional approach is an HMAC [5] over a timestamp and the raw body, sent in a header, following an existing convention such as Standard Webhooks [1] where you can so that receivers can reuse familiar libraries.

Illustrative — build a signed requestTypeScript
import { createHmac } from "node:crypto";

function buildRequest(evt: { id: string; type: string; payload: unknown; createdAt: Date }, secret: string) {
  const body = JSON.stringify({ id: evt.id, type: evt.type, created: evt.createdAt.toISOString(), data: evt.payload });
  const ts = Math.floor(Date.now() / 1000).toString();
  const signature = createHmac("sha256", secret).update(`${ts}.${body}`).digest("hex");
  return {
    body,
    headers: {
      "content-type": "application/json",
      "x-event-id": evt.id,           // stable id for receiver deduplication
      "x-timestamp": ts,
      "x-signature": signature,        // illustrative header names
    },
  };
}
  • Send the exact bytes you signed. Serialise once and reuse the string; re-serialising can change key order or whitespace and break verification.
  • Support two secrets. During rotation, sign with the new secret and, if you choose, include a signature for the old one so receivers can migrate at their own pace.
  • Never include secrets in the payload or URL. Anything in a URL ends up in logs.

Step four: attempt, classify, schedule

An attempt sends the request with a timeout and interprets the result. Not every failure deserves a retry, and treating them all the same either wastes effort or drops recoverable events.

OutcomeClassificationAction
2xxSuccessMark succeeded
Network error or timeoutTransientRetry with backoff
5xxTransientRetry with backoff
429 or 503 with Retry-AfterThrottledRetry after the indicated delay
410 GoneEndpoint permanently removedDisable the endpoint; alert the owner
Other 4xxProbably permanentRetry a few times, then mark failed and surface it
RedirectDo not follow blindlyTreat as a misconfiguration; surface it
Illustrative — classify and schedule the next attemptTypeScript
async function attempt(d: DueDelivery) {
  const req = buildRequest(d, d.secrets[0]);
  let status = 0, error: string | null = null;
  try {
    const res = await fetch(d.url, { method: "POST", body: req.body, headers: req.headers, redirect: "manual", signal: AbortSignal.timeout(8_000) });
    status = res.status;
  } catch (e) { error = String(e); }

  if (status >= 200 && status < 300) return markSucceeded(d.id, status);
  if (status === 410) return disableEndpoint(d.endpointId, d.id);

  const attempts = d.attempts + 1;
  if (attempts >= MAX_ATTEMPTS) return markExhausted(d.id, status, error);
  const delayMs = fullJitter(attempts);            // exponential with jitter [3][6]
  return scheduleRetry(d.id, attempts, delayMs, status, error);
}

Do not follow redirects blindly

Following redirects to attacker-influenced locations is a classic server-side request forgery vector. Treat a redirect as a failure to surface, not a hop to follow, and consider restricting which hosts and address ranges you will deliver to.

Protecting yourself from your receivers

A sender that calls arbitrary URLs is a network client with a lot of trust. Two risks deserve attention.

RiskWhat could happenMitigation
Server-side request forgeryA customer registers an internal address as their endpointBlock private and link-local ranges; resolve and validate the address at send time
Slow receiver ties up workersOne endpoint that hangs stalls the pipelineShort timeouts, per-endpoint concurrency limits, a separate pool for slow endpoints
Noisy endpointRepeated failures create loadCircuit-break an endpoint after sustained failure and notify its owner
Large payloadsMemory pressure and slow deliveryCap payload size; use thin events with a fetch API

Step five: exhausted deliveries and replay

After the last attempt, a delivery becomes exhausted. The critical rule is that nothing is silently discarded. The event stays, the delivery is visible, and an authorised operator can replay it once the receiver is fixed.

  1. 01

    Surface it

    Exhausted deliveries appear in the dashboard and trigger an alert to the endpoint's owner.

  2. 02

    Show the evidence

    The last status, the response body excerpt, and the timeline of attempts.

  3. 03

    Let a person replay

    A one-click replay creates a new attempt with the same event id and a fresh timestamp and signature.

  4. 04

    Record who replayed

    The audit log stores the actor and the reason.

  5. 05

    Support bulk replay

    After an outage, replay every exhausted delivery for an endpoint within a time range.

Replay uses the normal contract

A replay should go through the same signing and delivery path as a first attempt, with the same event id. That way the receiver's deduplication and verification work unchanged, and there is no privileged back door to secure.

Observability: a delivery log receivers can read

FieldSender's useReceiver's use
event id and typeSearch and supportDeduplicate; correlate with logs
attempt number and timestampsUnderstand retry behaviourSee what was tried and when
request headers with signature redactedReproduce the requestCheck what was sent
response status and body excerptDiagnose failuresProve what they returned
durationSpot slow receiversTune handler performance
endpoint health scorePrioritise attentionSee their own reliability
  • Give receivers self-service. A page where they can see recent deliveries and trigger a replay reduces support load dramatically.
  • Alert receivers proactively. When an endpoint has been failing for a while, tell them before they lose data.
  • Publish your retry schedule. Receivers design better when they know how long you will keep trying.

Testing the sender

Tests worth having

  • A crash between writing the outbox row and dispatch loses nothing.
  • Two workers never attempt the same delivery concurrently.
  • A worker crash mid-attempt results in a retry, not a lost delivery.
  • Each failure class produces the documented behaviour.
  • The signature verifies against an independent reference implementation.
  • An endpoint returning redirects, huge bodies, or hanging is handled safely.
  • Replay produces a delivery that a real receiver treats as a duplicate, not a new event.

A webhook sender is a small distributed system with an unfriendly network on one side. Write events transactionally, deliver them from a durable queue, sign and bound every request, and never lose a failure silently. Do that, and the receivers — the people who trust your events enough to move money on them — get an integration they can forget about.

References & further reading

  1. 1
    Standard Webhooks — Standard Webhooks specification
  2. 2
    Pattern: Transactional outbox — Chris Richardson, microservices.io
  3. 3
    Exponential Backoff And Jitter — Marc Brooker, AWS Architecture Blog
  4. 4
    PostgreSQL Documentation: SELECT (The Locking Clause) — PostgreSQL Global Development GroupCovers FOR UPDATE and SKIP LOCKED, used to build safe work queues.
  5. 5
    RFC 2104: HMAC: Keyed-Hashing for Message Authentication — H. Krawczyk, M. Bellare, R. Canetti, IETF, 1997
  6. 6
    Timeouts, retries and backoff with jitter — Marc Brooker, Amazon Builders' Library
  7. 7
Found this useful? Share it

Get the next essay in your inbox

Practical writing on payments infrastructure, operations software, and shipping real systems. No spam, no sales sequence.

We only use your email to send the studio's writing. See the privacy policy.

About the author

Product behaviour described here reflects what is implemented and tested; anything else is marked as planned. Code samples are illustrative.

All writing

Have a system like this to run?

Explore the catalog, or write down the problem and the constraints. We respond when the fit is real.

Free apps from the studio. Enter your email, get a private download link. Free for personal use.

Get them free

Have a product to sell? We review, list, and sell it for you — you keep 90% of every sale.

Apply to sell with us