On this page
Key takeaways
- Runbooks exist for the case that appears when the author is not there.
- Write symptoms first, causes second, actions third — the reader starts from what they see.
- Every runbook entry needs a stop condition: when to escalate, and to whom.
- A handover is finished when someone on the client side has used the runbook without us.
The install guide describes the day everything works. The runbook describes the night it does not. Ours are written for one specific reader: someone competent, alone, and looking at a screen that does not match what they expected.
The case that started it
During Vellum Markets' handover we wrote the runbook for the case that would eventually happen: a payment that looks confirmed on-chain and is silent in the application. It is the classic disagreement between two records, and the reason people distrust systems.
It has a name now in our docs: the silent confirmation. Naming it matters. A named problem can be searched for at 2am.
Structure of a good entry
- Symptom. What the operator sees, in their words. Not the internal error name.
- What it usually is. The two or three likeliest causes, ordered.
- How to tell which. The check that distinguishes them, with where to look.
- What to do. Steps that are safe to perform alone.
- When to stop. The exact condition for escalating, and who to call.
SYMPTOM
Payment shows confirmed on the network but the order is still pending.
USUALLY
1. Webhook delivery failed and is queued for retry.
2. Your endpoint returned non-2xx after the event.
3. The application processed it but the update was rolled back.
CHECK
Dashboard > Exceptions > Silent confirmations
Look at delivery attempts and last response code.
DO
Replay the event using the replay token. Confirm the order updates.
STOP AND ESCALATE IF
Replay succeeds but the order does not change, or the amount
does not match the intent. Page the payments engineer.A runbook nobody reads
- Starts with architecture
- Explains causes at length
- Says "contact support" with no condition
- Written once, never drilled
A runbook that gets used
- Starts from what the operator sees
- One-line checks that tell causes apart
- A stop condition and a name to call
- Fire-drilled in staging
Rules for writing them
- Start from what the reader sees. Nobody at 2am searches for the internal cause.
- Prefer checks over explanations. A one-line check beats a paragraph of theory.
- Mark dangerous steps. Anything that moves value or deletes data gets a clear warning and a second person.
- Keep it short enough to read once. If it needs scrolling, the reader is already lost.
The stop condition
The most important line in any runbook is the one that says when to stop trying. Without it, a tired operator makes the situation worse.
What every handover includes
| Artifact | Purpose | Reader |
|---|---|---|
| Architecture note | How the parts fit and why | The next engineer |
| Runbook | What to do when something breaks | On-call operator |
| Upgrade notes | How to apply patch releases safely | Whoever owns the deploy |
| Decision log | Trade-offs written down | Future maintainers |
| Contact path | Who to ask, and what is out of scope | Everyone |
- 01
Pick an entry
Choose one runbook entry the client's team has never used.
- 02
Break something safely
In staging, reproduce the symptom: fail a webhook, expire a link, stall a provider.
- 03
Hand over the runbook, not the author
The client's operator follows it. We watch and stay quiet.
- 04
Fix what they stumbled on
Every hesitation is a sentence to rewrite or a check to add.
The test
A handover is finished when someone on the client side has used the runbook without us. We ask for a fire drill: pick an entry, break something safely in staging, and have their operator follow it. Every time, the runbook improves.
Handover is done when
- An operator on the client side has used the runbook without us.
- Every dangerous step has a warning and a second-person rule.
- Every entry has a stop condition and a contact.
- Upgrade notes explain how to apply a patch release safely.
- The decision log says why, not just what.
A starter library of runbook entries
The structure above is easier to see in concrete entries. These are illustrative starters for a self-hosted payment gateway. Adapt the names to your environment and delete anything that does not apply.
ENTRY Silent confirmation
SEVERITY Medium (customer waiting, funds safe)
SYMPTOM
Dashboard shows payment CONFIRMED. Application order is PENDING.
USUALLY
1. Webhook delivery failed (endpoint down or non-2xx).
2. Handler crashed after acknowledging.
3. Order update rolled back.
CHECK
Dashboard > Deliveries: last attempt status and response code.
Application logs for the event id.
DO
1. If deliveries failed: fix the endpoint, then replay with the replay token.
2. Confirm the order moved to PAID.
STOP AND ESCALATE IF
Replay succeeds but the order does not change, OR amount differs
from the intent. Page: payments engineer. Do not edit records by hand.ENTRY Wrong-network payment
SEVERITY High (customer funds not credited; do not improvise)
SYMPTOM
Customer says they paid. No matching intent activity.
CHECK
1. Ask for the transaction hash and network used.
2. Look up the hash in the relevant network explorer.
3. Compare network with the one shown at checkout.
DO
Record the case with hash and network. Do NOT resend funds.
Follow the recovery path in the network-support document.
STOP AND ESCALATE
Immediately. Involve the payments lead and finance. Legal review
if unsure. This entry is a triage, not a fix.ENTRY Indexer lag
SEVERITY High if lag exceeds the confirmation window
SYMPTOM
Payments appear late. Dashboard lag indicator is red.
CHECK
1. Indexer lag metric and last processed block.
2. RPC provider status page and response times.
3. Provider quota or rate-limit errors in logs.
DO
Switch to the secondary provider if configured. Restart the
indexer only after noting the last processed block.
STOP AND ESCALATE IF
Lag persists after switching providers, or blocks are skipped.Severity levels and who gets woken up
Not every problem justifies a phone call at 2am. Agreeing severity levels in advance prevents both over-escalation, which burns people out, and under-escalation, which burns customers. The framework below follows the shape used in mature incident response practice [3][4].
| Level | Definition | Response | Example |
|---|---|---|---|
| SEV-1 | Funds at risk or all payments failing | Page on-call immediately; incident lead named | Unexpected signing activity; gateway down |
| SEV-2 | Significant degradation, no funds at risk | Page on-call; respond within 30 minutes | Indexer lag beyond the confirmation window |
| SEV-3 | Localised problem, workaround exists | Handle in working hours unless it worsens | One silent confirmation |
| SEV-4 | Cosmetic or informational | Ticket only | A dashboard label is wrong |
- Escalate up, never down. If unsure, treat it as the higher level and downgrade after diagnosis.
- One incident lead. One person coordinates; others investigate. Shouting in three channels is how records get edited by hand.
- Write as you go. A running timeline in a shared note makes the review afterwards factual instead of remembered.
After the incident: blameless review
A runbook that never changes is a runbook nobody reads. The mechanism that keeps it alive is the review after each incident that mattered. The blameless format [1] treats an incident as a system failure rather than a personal one: the question is what allowed a reasonable person to make that decision, not who to blame.
- 01
Reconstruct the timeline
What happened, in order, from logs and the incident note. Facts first.
- 02
Find the contributing factors
Not a single root cause. Usually several small gaps lined up.
- 03
Ask what made the wrong step look right
If a decision looked reasonable at the time, the system, not the person, needs changing.
- 04
Write actions with owners and dates
Every action is a runbook edit, an alert, a test, or a design change.
- 05
Update the runbook first
The cheapest fix is often a sentence in the entry that would have saved twenty minutes.
Fire drills and game days
A drill turns a document into a habit. Schedule them; do not wait for a good time. The client's operator follows the runbook while someone who wrote it stays silent.
| Drill | How to run it safely | What you learn |
|---|---|---|
| Failed webhook | Point the staging endpoint at a failing handler; create a payment | Whether the operator finds the silent-confirmation entry unaided |
| Expired link then paid | Let a payment link expire in staging; send a test payment | Whether the closed-reason and refund path are clear |
| Provider outage | Block the primary RPC provider in staging | Whether failover is configured and observable |
| Restore from backup | Restore last night's backup into a scratch environment | How long recovery actually takes |
| Key rotation | Rotate a non-production signing key end to end | Whether the procedure survives contact with reality |
Runbook maintenance rhythm
- Review every entry at least twice a year.
- Update after every incident, in the same week.
- Fire-drill one new entry each quarter.
- Record the last-tested date on each entry.
- Retire entries for features you removed.
Closing
Software you operate means someone will operate it at the worst time. The runbook is the kindest thing we can leave them.
References & further reading
- 1Postmortem Culture: Learning from Failure — John Lunney, Sue Lueder, Google SRE Book
- 2Managing Incidents — Andrew Stribblehill, Google SRE Book
- 3PagerDuty Incident Response Documentation — PagerDutyAn open, practical guide to roles, severity, and communication during incidents.
- 4Incident Response — Google SRE Workbook
The product behind this post
StablePay
Self-hosted stablecoin payment infrastructure.
From $4,800 one-time license
About the authors
Daniel Whitfield
Staff Engineer, Platform & Reliability
Owns webhooks, PostgreSQL, deployment, and the operational habits that keep self-hosted software boring.
Emily Sinclair
Head of Product
Leads Northframe and Meridian, and the discovery work that turns shadow spreadsheets into products.
Product behaviour described here reflects what is implemented and tested; anything else is marked as planned. Code samples are illustrative.
All writing