Journal · 28 Jul 2026

Handover runbooks: writing for the 2am case.

A runbook is not documentation of the happy path. It is a letter to the person who is alone, tired, and looking at something odd.

Topics
HandoverRunbooksStablePayOperations
On this page

Key takeaways

  • Runbooks exist for the case that appears when the author is not there.
  • Write symptoms first, causes second, actions third — the reader starts from what they see.
  • Every runbook entry needs a stop condition: when to escalate, and to whom.
  • A handover is finished when someone on the client side has used the runbook without us.

The install guide describes the day everything works. The runbook describes the night it does not. Ours are written for one specific reader: someone competent, alone, and looking at a screen that does not match what they expected.

The case that started it

During Vellum Markets' handover we wrote the runbook for the case that would eventually happen: a payment that looks confirmed on-chain and is silent in the application. It is the classic disagreement between two records, and the reason people distrust systems.

It has a name now in our docs: the silent confirmation. Naming it matters. A named problem can be searched for at 2am.

NetworkPayment confirmed on-chain
WebhookDelivery fails or handler rolls back
ApplicationOrder still pending
OperatorFinds it under silent confirmations
ReplayReplay token, order updates
The silent confirmation — the case the first runbook was written for

Structure of a good entry

  1. Symptom. What the operator sees, in their words. Not the internal error name.
  2. What it usually is. The two or three likeliest causes, ordered.
  3. How to tell which. The check that distinguishes them, with where to look.
  4. What to do. Steps that are safe to perform alone.
  5. When to stop. The exact condition for escalating, and who to call.
Illustrative runbook entryPlain text
SYMPTOM
  Payment shows confirmed on the network but the order is still pending.

USUALLY
  1. Webhook delivery failed and is queued for retry.
  2. Your endpoint returned non-2xx after the event.
  3. The application processed it but the update was rolled back.

CHECK
  Dashboard > Exceptions > Silent confirmations
  Look at delivery attempts and last response code.

DO
  Replay the event using the replay token. Confirm the order updates.

STOP AND ESCALATE IF
  Replay succeeds but the order does not change, or the amount
  does not match the intent. Page the payments engineer.

A runbook nobody reads

  • Starts with architecture
  • Explains causes at length
  • Says "contact support" with no condition
  • Written once, never drilled

A runbook that gets used

  • Starts from what the operator sees
  • One-line checks that tell causes apart
  • A stop condition and a name to call
  • Fire-drilled in staging

Rules for writing them

  • Start from what the reader sees. Nobody at 2am searches for the internal cause.
  • Prefer checks over explanations. A one-line check beats a paragraph of theory.
  • Mark dangerous steps. Anything that moves value or deletes data gets a clear warning and a second person.
  • Keep it short enough to read once. If it needs scrolling, the reader is already lost.

The stop condition

The most important line in any runbook is the one that says when to stop trying. Without it, a tired operator makes the situation worse.

What every handover includes

ArtifactPurposeReader
Architecture noteHow the parts fit and whyThe next engineer
RunbookWhat to do when something breaksOn-call operator
Upgrade notesHow to apply patch releases safelyWhoever owns the deploy
Decision logTrade-offs written downFuture maintainers
Contact pathWho to ask, and what is out of scopeEveryone
  1. 01

    Pick an entry

    Choose one runbook entry the client's team has never used.

  2. 02

    Break something safely

    In staging, reproduce the symptom: fail a webhook, expire a link, stall a provider.

  3. 03

    Hand over the runbook, not the author

    The client's operator follows it. We watch and stay quiet.

  4. 04

    Fix what they stumbled on

    Every hesitation is a sentence to rewrite or a check to add.

The test

A handover is finished when someone on the client side has used the runbook without us. We ask for a fire drill: pick an entry, break something safely in staging, and have their operator follow it. Every time, the runbook improves.

Handover is done when

  • An operator on the client side has used the runbook without us.
  • Every dangerous step has a warning and a second-person rule.
  • Every entry has a stop condition and a contact.
  • Upgrade notes explain how to apply a patch release safely.
  • The decision log says why, not just what.

A starter library of runbook entries

The structure above is easier to see in concrete entries. These are illustrative starters for a self-hosted payment gateway. Adapt the names to your environment and delete anything that does not apply.

Illustrative entry — a payment is confirmed on the network but the order is still pendingPlain text
ENTRY  Silent confirmation
SEVERITY  Medium (customer waiting, funds safe)

SYMPTOM
  Dashboard shows payment CONFIRMED. Application order is PENDING.

USUALLY
  1. Webhook delivery failed (endpoint down or non-2xx).
  2. Handler crashed after acknowledging.
  3. Order update rolled back.

CHECK
  Dashboard > Deliveries: last attempt status and response code.
  Application logs for the event id.

DO
  1. If deliveries failed: fix the endpoint, then replay with the replay token.
  2. Confirm the order moved to PAID.

STOP AND ESCALATE IF
  Replay succeeds but the order does not change, OR amount differs
  from the intent. Page: payments engineer. Do not edit records by hand.
Illustrative entry — payment received on the wrong networkPlain text
ENTRY  Wrong-network payment
SEVERITY  High (customer funds not credited; do not improvise)

SYMPTOM
  Customer says they paid. No matching intent activity.

CHECK
  1. Ask for the transaction hash and network used.
  2. Look up the hash in the relevant network explorer.
  3. Compare network with the one shown at checkout.

DO
  Record the case with hash and network. Do NOT resend funds.
  Follow the recovery path in the network-support document.

STOP AND ESCALATE
  Immediately. Involve the payments lead and finance. Legal review
  if unsure. This entry is a triage, not a fix.
Illustrative entry — indexer is behind the networkPlain text
ENTRY  Indexer lag
SEVERITY  High if lag exceeds the confirmation window

SYMPTOM
  Payments appear late. Dashboard lag indicator is red.

CHECK
  1. Indexer lag metric and last processed block.
  2. RPC provider status page and response times.
  3. Provider quota or rate-limit errors in logs.

DO
  Switch to the secondary provider if configured. Restart the
  indexer only after noting the last processed block.

STOP AND ESCALATE IF
  Lag persists after switching providers, or blocks are skipped.

Severity levels and who gets woken up

Not every problem justifies a phone call at 2am. Agreeing severity levels in advance prevents both over-escalation, which burns people out, and under-escalation, which burns customers. The framework below follows the shape used in mature incident response practice [3][4].

LevelDefinitionResponseExample
SEV-1Funds at risk or all payments failingPage on-call immediately; incident lead namedUnexpected signing activity; gateway down
SEV-2Significant degradation, no funds at riskPage on-call; respond within 30 minutesIndexer lag beyond the confirmation window
SEV-3Localised problem, workaround existsHandle in working hours unless it worsensOne silent confirmation
SEV-4Cosmetic or informationalTicket onlyA dashboard label is wrong
  • Escalate up, never down. If unsure, treat it as the higher level and downgrade after diagnosis.
  • One incident lead. One person coordinates; others investigate. Shouting in three channels is how records get edited by hand.
  • Write as you go. A running timeline in a shared note makes the review afterwards factual instead of remembered.

After the incident: blameless review

A runbook that never changes is a runbook nobody reads. The mechanism that keeps it alive is the review after each incident that mattered. The blameless format [1] treats an incident as a system failure rather than a personal one: the question is what allowed a reasonable person to make that decision, not who to blame.

  1. 01

    Reconstruct the timeline

    What happened, in order, from logs and the incident note. Facts first.

  2. 02

    Find the contributing factors

    Not a single root cause. Usually several small gaps lined up.

  3. 03

    Ask what made the wrong step look right

    If a decision looked reasonable at the time, the system, not the person, needs changing.

  4. 04

    Write actions with owners and dates

    Every action is a runbook edit, an alert, a test, or a design change.

  5. 05

    Update the runbook first

    The cheapest fix is often a sentence in the entry that would have saved twenty minutes.

Fire drills and game days

A drill turns a document into a habit. Schedule them; do not wait for a good time. The client's operator follows the runbook while someone who wrote it stays silent.

DrillHow to run it safelyWhat you learn
Failed webhookPoint the staging endpoint at a failing handler; create a paymentWhether the operator finds the silent-confirmation entry unaided
Expired link then paidLet a payment link expire in staging; send a test paymentWhether the closed-reason and refund path are clear
Provider outageBlock the primary RPC provider in stagingWhether failover is configured and observable
Restore from backupRestore last night's backup into a scratch environmentHow long recovery actually takes
Key rotationRotate a non-production signing key end to endWhether the procedure survives contact with reality

Runbook maintenance rhythm

  • Review every entry at least twice a year.
  • Update after every incident, in the same week.
  • Fire-drill one new entry each quarter.
  • Record the last-tested date on each entry.
  • Retire entries for features you removed.

Closing

Software you operate means someone will operate it at the worst time. The runbook is the kindest thing we can leave them.

References & further reading

  1. 1
    Postmortem Culture: Learning from Failure — John Lunney, Sue Lueder, Google SRE Book
  2. 2
    Managing Incidents — Andrew Stribblehill, Google SRE Book
  3. 3
    PagerDuty Incident Response Documentation — PagerDutyAn open, practical guide to roles, severity, and communication during incidents.
  4. 4
    Incident Response — Google SRE Workbook
Found this useful? Share it

Get the next essay in your inbox

Practical writing on payments infrastructure, operations software, and shipping real systems. No spam, no sales sequence.

We only use your email to send the studio's writing. See the privacy policy.

About the authors

Product behaviour described here reflects what is implemented and tested; anything else is marked as planned. Code samples are illustrative.

All writing

Have a system like this to run?

Explore the catalog, or write down the problem and the constraints. We respond when the fit is real.

Free apps from the studio. Enter your email, get a private download link. Free for personal use.

Get them free

Have a product to sell? We review, list, and sell it for you — you keep 90% of every sale.

Apply to sell with us