On this page
Key takeaways
- Verify first, acknowledge fast, process later.
- Assume duplicates, reordering, and delays; write handlers that shrug at all three.
- Log the event id everywhere, and make failures visible before customers notice.
- Test with bad input on purpose: altered bodies, stale timestamps, and repeated deliveries.
A webhook endpoint looks trivial: receive a POST, update a row, return 200. It is also one of the most common sources of production incidents in integrations, because it sits on an unreliable boundary and fails in ways you cannot reproduce on your laptop. These tips are the small habits that make the difference. They apply to any provider [2][3], and we mention StablePay where the specifics help.
Verify before anything else
Authenticity and freshness are the price of admission. Do nothing with an event until it passes both.
1. Verify the signature on the raw body
Parse JSON after verification. A body that is parsed and re-serialised can differ in whitespace and key order, and the HMAC [5] will no longer match. Use the exact bytes you received.
export async function POST(req: Request) {
const raw = await req.text(); // raw bytes as text, before any JSON.parse
const ok = verify(raw, req.headers.get("x-timestamp")!, req.headers.get("x-signature")!, SECRET);
if (!ok) return new Response("invalid", { status: 401 });
const event = JSON.parse(raw);
// ...
}2. Compare signatures in constant time
Ordinary string comparison can leak information through timing. Use a constant-time comparison such as Node's timingSafeEqual [4], and check lengths first, since it throws on unequal lengths.
3. Enforce a timestamp tolerance
Sign the timestamp with the body, and reject deliveries outside a window of a few minutes. A valid signature alone does not stop replay of a captured request [1].
4. Support two secrets during rotation
Try the new secret, then the old one. Rotation should never require a coordinated, zero-tolerance cutover.
5. Reject early and quietly
Return 401 or 400 for bad requests without echoing details. Do not tell an attacker which part failed.
Respond fast, process safely
Providers time out quickly and retry, so your response speed and processing design decide whether retries pile up.
6. Acknowledge first, work later
Verify, persist the raw event, return 2xx, and process asynchronously. A handler that does slow work inline invites timeouts and duplicate retries.
7. Persist the event before doing anything else
Store the event id and payload in your own table first. If processing crashes, you can replay from your record rather than depending on the sender.
8. Deduplicate on the event id
Delivery is at-least-once. Insert the event id with a unique constraint and treat a conflict as "already handled".
CREATE TABLE webhook_events (
event_id text PRIMARY KEY,
received_at timestamptz NOT NULL DEFAULT now(),
payload jsonb NOT NULL,
processed_at timestamptz
);
-- INSERT ... ON CONFLICT (event_id) DO NOTHING; if nothing inserted, return 2xx and stop.9. Return 2xx for duplicates
A duplicate is not an error. Returning an error makes the sender retry a delivery you have already handled.
10. Write the state change and the event record in one transaction
Recording that you handled the event, and the effect of handling it, should commit together or not at all [6].
11. Make handlers idempotent by design
Setting status = 'paid' is idempotent. Incrementing a counter is not. Prefer operations you can safely run twice.
Expect disorder
Networks reorder and delay. Design as if every event might arrive late, early, or twice.
12. Never assume ordering
A confirmed event can arrive before the pending that preceded it. Handle by comparing states, not by trusting arrival order.
13. Only move state forward
Encode allowed transitions and ignore events that would move a record backwards. See the state-machine post for a full design.
const rank = { created: 0, seen: 1, confirming: 2, confirmed: 3, settled: 4 } as const;
function advance(current: keyof typeof rank, incoming: keyof typeof rank) {
return rank[incoming] > rank[current] ? incoming : current; // stale events are no-ops
}14. Fetch the truth when in doubt
If an event seems inconsistent, call the provider's API for the current state instead of trusting the payload alone.
15. Tolerate unknown fields and event types
Ignore what you do not recognise. Providers add fields; strict parsers break on additions.
16. Handle very late deliveries
After an outage, retries may arrive hours later. Your handler should still be correct then.
Test like something will go wrong
The interesting bugs live in the cases you did not think to try.
17. Keep recorded payload fixtures
Capture real events from staging, redact secrets, and commit them. Tests against real shapes catch schema surprises.
18. Write tests for rejection
Alter one byte of the body, use a stale timestamp, omit the signature: each must be refused.
19. Test duplicates and disorder
Deliver the same event twice and events out of order, then assert the final state.
20. Test the slow path
Make processing slow and check you still acknowledge within the sender's timeout.
Operate it
A handler is only reliable if you know when it is not.
21. Log the event id on every line
One id lets you follow a delivery from receipt to effect, across services and retries.
22. Alert on failure rate, not single failures
A steady stream of non-2xx responses means retries are stacking up. One failure is noise; a trend is a page.
23. Reconcile against the provider
Periodically compare your records with the sender's. It is the safety net for events that never arrived.
24. Use replay, not hand edits
When you miss an event, replay it through the normal path. Editing records by hand hides the cause and skips your own safeguards.
25. Write the runbook entry now
Include the symptom, how to check deliveries, how to replay, and when to escalate. You will need it at the worst time.
A one-page summary
Before you ship a webhook endpoint
- Signature verified on the raw body, in constant time.
- Timestamp tolerance enforced; two secrets accepted.
- Event stored and deduplicated by id before processing.
- Acknowledged quickly; processed asynchronously and idempotently.
- Forward-only state transitions; unknown fields ignored.
- Tests for rejection, duplicates, and disorder.
- Failure-rate alerts, an event id in every log line, and a runbook entry.
Nothing here is clever. Reliable webhook handling is the accumulation of dull habits, applied every time. Put them in a template, and every new endpoint starts from a safe baseline.
References & further reading
- 1Standard Webhooks — Standard Webhooks specification
- 2Validating webhook deliveries — GitHub Docs
- 3Receive Stripe events in your webhook endpoint — Stripe Documentation
- 4crypto.timingSafeEqual(a, b) — Node.js documentation
- 5RFC 2104: HMAC: Keyed-Hashing for Message Authentication — H. Krawczyk, M. Bellare, R. Canetti, IETF, 1997
- 6Pattern: Transactional outbox — Chris Richardson, microservices.io
The product behind this post
StablePay
Self-hosted stablecoin payment infrastructure.
From $4,800 one-time license
About the author
Product behaviour described here reflects what is implemented and tested; anything else is marked as planned. Code samples are illustrative.
All writing