Engineering · 19 Mar 2026

Indexing blockchain payments reliably: cursors, reorgs, and failover.

How to watch a network without missing a payment or counting one twice — polling versus subscriptions, cursor management, backfill, and surviving a bad RPC provider.

Topics
StablePayIndexingNetworksRPCReliability
On this page

Key takeaways

  • An indexer's job is to turn an unreliable stream of chain data into a durable, ordered, deduplicated record.
  • Persist a cursor and the block hash at that cursor, so a restart or a reorg is recoverable.
  • Prefer polling with a cursor as the source of truth; use subscriptions only as a low-latency hint.
  • Treat RPC providers as fallible: cross-check, fail over, and alert on lag.

A payment gateway that accepts tokens has to notice when value arrives. That sounds like a loop — ask the network for new transfers, match them to intents — and a first version really is that simple. The complexity arrives with the failure modes: the process restarts mid-block, the provider returns stale data, the chain reorganises, or a burst of activity outpaces your queries.

This post covers the design of the component that does the watching, usually called an indexer. The examples use an Ethereum-style account model and token Transfer events [2], but the principles — durable cursors, idempotent ingestion, and distrust of providers — apply to any network.

What an indexer promises

GuaranteeMeaningFailure if broken
No missed eventsEvery relevant transfer in a confirmed range is eventually recordedA customer paid and nothing happened
No double countingEach on-chain event is recorded exactly onceAn order fulfilled twice
Ordering by chain positionEvents are processed in block and log orderState machine sees effects before causes
RecoverableA crash or reorg leaves the index repairableManual database surgery
ObservableLag and errors are visibleSilent staleness

At-least-once plus idempotent

You cannot get exactly-once delivery from a network you do not control. You can get at-least-once ingestion combined with idempotent processing, which is equivalent in effect and much easier to reason about [6].

Polling or subscriptions?

Networks typically offer two ways to learn about new data: you can ask periodically, or you can open a subscription and be told. Subscriptions feel more elegant, but they are stateful connections that drop, and a dropped connection silently loses events.

Polling with a cursorSubscription (WebSocket)
Missed eventsImpossible if the cursor is durablePossible on any disconnect
LatencyBounded by the poll intervalLowest
LoadPredictableDepends on provider limits
RecoveryResume from the cursorMust detect the gap and backfill
ComplexityLowHigher: reconnects, heartbeats, gap detection

The robust pattern is to make polling the source of truth and use a subscription, if at all, only as a hint to poll sooner. If the hint is lost, nothing is lost; the next poll picks it up. Design so the system is correct without the subscription and merely faster with it.

The cursor: the most important row in the database

The cursor records how far you have processed. It must be durable and updated atomically with the events it covers. Store the block number and the block hash, because the number alone cannot tell you whether the chain beneath it has changed.

Illustrative — cursor and ingested eventsSQL
CREATE TABLE indexer_cursor (
  network      text PRIMARY KEY,
  block_number bigint NOT NULL,
  block_hash   text   NOT NULL,
  updated_at   timestamptz NOT NULL DEFAULT now()
);

CREATE TABLE chain_transfers (
  network      text   NOT NULL,
  tx_hash      text   NOT NULL,
  log_index    int    NOT NULL,
  block_number bigint NOT NULL,
  block_hash   text   NOT NULL,
  token        text   NOT NULL,
  to_address   text   NOT NULL,
  amount_units numeric(78,0) NOT NULL,
  PRIMARY KEY (network, tx_hash, log_index)   -- idempotent ingestion key
);

The primary key on (network, tx_hash, log_index) is what makes ingestion idempotent: re-reading a range after a crash inserts nothing new. The cursor and the inserted events should be written in the same transaction, so the cursor never claims progress the events do not reflect.

The polling loop

Illustrative — a cursor-driven loop with a safety windowTypeScript
async function pollOnce(net: Network, rpc: Rpc) {
  const cursor = await loadCursor(net);                       // { number, hash }
  const head = await rpc.latestBlockNumber();
  const target = head - BigInt(net.safetyDepth);              // do not read the very tip
  if (target <= cursor.number) return;

  // 1. reorg check: is the block at our cursor still the one we recorded?
  const onChain = await rpc.getBlockByNumber(cursor.number);
  if (onChain.hash !== cursor.hash) return handleReorg(net, cursor, rpc);

  // 2. read in bounded chunks so one poll cannot time out
  const from = cursor.number + 1n;
  const to = min(target, from + BigInt(net.maxRange) - 1n);
  const logs = await rpc.getLogs({ fromBlock: from, toBlock: to, address: net.tokens, topics: [TRANSFER_TOPIC] });

  // 3. persist events and advance the cursor in ONE transaction
  const last = await rpc.getBlockByNumber(to);
  await db.transaction(async (tx) => {
    for (const l of logs) await insertTransferIfNew(tx, net.name, l);
    await saveCursor(tx, net.name, { number: to, hash: last.hash });
  });
}
  • Read behind the tip. A small safety depth avoids reading blocks most likely to be replaced, at the cost of latency. Combine it with your confirmation policy rather than replacing it.
  • Bound the range. Providers limit how many blocks or logs one request may cover. Chunk the range and handle a "too many results" error by halving it.
  • Filter at the source. Ask only for the token contracts and event signatures you care about; scanning everything wastes quota and time.

Reorgs: rewinding the cursor

When the block at your cursor no longer matches the recorded hash, the chain has reorganised beneath you. The recovery is to walk back to the last block whose hash still matches, delete the events above it, move the cursor there, and resume.

DetectCursor hash differs from chain
Walk backFind the last matching block
RewindDelete events above it, in one transaction
NotifyTell the payment state machine
ResumeRe-ingest the new canonical blocks
Handling a reorganisation
Illustrative — find the fork point and rewindTypeScript
async function handleReorg(net: Network, cursor: Cursor, rpc: Rpc) {
  let n = cursor.number;
  while (n > 0n) {
    const recorded = await blockHashWeRecorded(net, n);       // from chain_transfers or a blocks table
    const onChain = (await rpc.getBlockByNumber(n)).hash;
    if (recorded === onChain) break;                          // fork point found
    n -= 1n;
  }
  await db.transaction(async (tx) => {
    const removed = await deleteTransfersAbove(tx, net.name, n);
    await saveCursor(tx, net.name, { number: n, hash: await hashAt(rpc, n) });
    await emitReorgEvents(tx, removed); // payments touching removed events step back
  });
}

The emitted reorg events feed the payment state machine, which demotes affected payments from confirming to seen or created. This is why the state machine treats those moves as reversible until the policy's finality point, and why on some networks you can key the policy on an explicit finalised marker rather than a block count [3].

Backfill: catching up after downtime

Any indexer will occasionally be off for hours. When it returns, it must catch up quickly without hammering the provider. The cursor makes this natural, because the loop simply keeps reading chunks until it reaches the head. A few refinements help.

  • Adaptive chunk size. Start large; halve on errors; grow again on success.
  • Rate limit yourself. Stay comfortably inside the provider's quota, with jittered retries [4].
  • Prioritise addresses you care about. During a large backfill, process ranges that contain pending payments first, so open orders resolve sooner.
  • Expose progress. A backfill that shows its percentage complete is a backfill nobody restarts in a panic.

Providers fail in creative ways

FailureSymptomDefence
Provider lags the networkIndexer looks healthy but data is oldCompare head to a second provider; alert on divergence
Provider returns stale or wrong dataA transaction seen and then vanishesVerify block hashes across two providers before acting
Rate limitingSporadic 429 errorsBackoff with jitter; lower the polling rate; upgrade the plan
Silent partial resultsgetLogs returns fewer logs than existReduce the range; cross-check counts on suspicious ranges
OutageTimeoutsFailover to a secondary provider with the same cursor
Illustrative — fail over between providers, verifying continuityTypeScript
async function withFailover<T>(providers: Rpc[], call: (r: Rpc) => Promise<T>): Promise<T> {
  let lastError: unknown;
  for (const rpc of providers) {
    try {
      return await withTimeout(call(rpc), 8_000);
    } catch (e) {
      lastError = e;
      metrics.increment("rpc.failover", { provider: rpc.name });
    }
  }
  throw lastError;
}

// Because the cursor stores a block HASH, switching providers is safe:
// the next provider must agree on the hash at the cursor or we treat it as a reorg.

Matching transfers to payments

Ingested transfers are raw facts. A second step matches them to payment intents by destination address and, where applicable, amount. Keeping ingestion and matching separate is valuable: you can re-run matching after fixing a bug without re-reading the chain.

StageInputOutputIdempotency key
IngestChain logschain_transfers rows(network, tx_hash, log_index)
MatchTransfers and open intentsObservation events for the state machine(transfer key, payment id)
ApplyObservation eventsPayment state changesEvent id per observation

Decimals and amounts

Store token amounts as integers in minor units with a wide numeric type, and read the token's decimals per network rather than assuming a value [2]. Formatting for humans happens at the edge.

Observability: what to measure

MetricWhy it mattersAlert when
Indexer lag (head − cursor)Directly delays confirmationsLag approaches the confirmation window
Time since last successful pollDetects a stuck loopBeyond a few intervals
Provider error and failover rateEarly sign of provider troubleRising trend
Reorg count and depthNetwork health and policy relevanceAny reorg deeper than your safety depth
Unmatched transfersPayments to addresses with no intentAny above a set age
Backfill progressRecovery statusStalls

Testing an indexer without a live network

  1. 01

    Build a fake chain

    A small in-memory chain that can append blocks, reorganise, and drop connections on command.

  2. 02

    Test crash recovery

    Kill the process between reading and committing; restart; assert no gaps and no duplicates.

  3. 03

    Test reorgs of varying depth

    Reorganise one block, then several, and assert the state machine steps back correctly.

  4. 04

    Test provider misbehaviour

    Return stale heads, partial logs, and timeouts; assert failover and alerts.

  5. 05

    Replay recorded history

    Feed a real historical range, including a known reorganisation, and compare the index to a trusted source.

Indexer readiness

  • The cursor stores number and hash and is updated atomically with events.
  • Ingestion is idempotent by a stable natural key.
  • Reorgs are detected and rewound, and the state machine is notified.
  • At least two providers are configured, with lag comparison.
  • Lag, failover rate, and reorg depth are monitored and alerted.
  • A backfill has been run in staging against real history.

An indexer is unglamorous infrastructure that quietly decides whether payments feel instant or unreliable. Get the cursor, the idempotency key, and the reorg path right, and everything downstream — the state machine, the webhooks, the reconciliation — has a solid foundation.

References & further reading

  1. 1
    JSON-RPC API — ethereum.org developer documentationMethods such as eth_getLogs, eth_getBlockByNumber, and the block tags.
  2. 2
    EIP-20: Token Standard — Fabian Vogelsteller, Vitalik Buterin, Ethereum Improvement ProposalsDefines the Transfer event that payment indexers watch.
  3. 3
    Proof-of-stake and finality (Gasper) — ethereum.org developer documentation
  4. 4
    Timeouts, retries and backoff with jitter — Marc Brooker, Amazon Builders' Library
  5. 5
    Pattern: Transactional outbox — Chris Richardson, microservices.io
  6. 6
    Designing Data-Intensive Applications — Martin Kleppmann, O'Reilly Media
Found this useful? Share it

Get the next essay in your inbox

Practical writing on payments infrastructure, operations software, and shipping real systems. No spam, no sales sequence.

We only use your email to send the studio's writing. See the privacy policy.

About the authors

Product behaviour described here reflects what is implemented and tested; anything else is marked as planned. Code samples are illustrative.

All writing

Have a system like this to run?

Explore the catalog, or write down the problem and the constraints. We respond when the fit is real.

Free apps from the studio. Enter your email, get a private download link. Free for personal use.

Get them free

Have a product to sell? We review, list, and sell it for you — you keep 90% of every sale.

Apply to sell with us