SolidRPC plans change on 28th October. Some prices are going up. Lock in today’s pricing and limits by activating before then. See what’s changing.
Skip to main content
Guides
Category
Operations
Reading time
13 minutes
Published
Updated

How to backfill EVM logs without gaps or duplicates

An EVM log backfill looks simple: call eth_getLogs from the contract's deployment block to the chain head, decode the events, and save them. The first production run usually exposes the missing details. Wide ranges time out, dense contracts overflow response limits, retries insert duplicates, and a reorganization can make a successfully stored block non-canonical.

A reliable backfill is therefore a small data-ingestion system, not one large RPC call. It needs bounded ranges, a durable cursor, idempotent writes, explicit chain boundaries, and a recovery path for both provider errors and reorganizations.

This guide takes that process from filter design to the handoff into live indexing. It does not prescribe one block-range size or confirmation count, because those depend on the chain, event density, endpoint, and correctness requirement. It shows how to discover and enforce the right values for your workload.

Section 01

What is the short answer?

Choose a fixed end block, split the inclusive interval into adjacent pages, and request each page with the narrowest correct address and topic filter. Commit the returned logs and the page cursor in one database transaction. If a page is too heavy, shrink it; after repeated fast successes, grow it cautiously. Never advance the cursor for a failed or partially processed page.

Store block identity with every log and make ingestion idempotent. For finalized-only data, backfill to a chain-appropriate finalized boundary. If you ingest unfinalized blocks, retain enough block hashes to detect a reorganization, delete or mark orphaned data, and replay from the common ancestor.

The official `eth_getLogs` specification defines the filter and returned log fields, including block hash, block number, transaction hash, log index, and the removed flag. Those fields are the raw material for a correct cursor and uniqueness model; a successful HTTP response alone is not a completeness guarantee.

Section 02

How should you define the backfill before sending requests?

Freeze the job's scope first. Record the chain ID, contract addresses, event signatures, deployment or starting block, and a fixed target block. Do not leave toBlock as latest throughout a long run: the finish line will move while the worker is trying to reach it, making progress and validation ambiguous.

Choose the target according to the data contract:

  • Use a finalized or otherwise chain-approved boundary when consumers only need stable history. Ethereum's JSON-RPC documentation lists safe and finalized among the supported block tags, but tag support and finality semantics must be tested on the actual EVM network and client.
  • Use an explicit recent block when low-latency consumers accept reorg repair. Save that block's hash and keep a configurable rollback window.
  • Separate historical backfill from the live tail. They have different range sizes, concurrency, latency targets, and failure risks.

Then make the filter as selective as the product permits. address may be one contract or a list. Topic positions are order-dependent: topics[0] normally identifies the event signature, a nested array means OR within one position, and null leaves a position unconstrained. An incorrect topic layout can return an empty array that looks like a successful empty page, so test the filter against several known transactions before starting the full range.

Section 03

How do you page through eth_getLogs safely?

Treat fromBlock and toBlock as an inclusive interval. If the current page is [start, end], the next page starts at end + 1; using end again creates an overlap, while end + 2 creates a gap. Clamp every end to the fixed job target.

A practical adaptive loop is:

  1. Start with a conservative page width that has already succeeded for a representative dense interval.
  2. Request [cursor, min(cursor + width - 1, target)].
  3. Validate the JSON-RPC envelope and every returned log before decoding application data.
  4. In one database transaction, upsert the page's logs, store the observed block identities, and advance the cursor to end + 1.
  5. After several fast, bounded responses, increase the width gradually up to a configured ceiling.
  6. On a timeout, response-size limit, or provider range error, reduce the width and retry the same cursor; do not skip forward.
  7. If even a single-block page fails repeatedly, stop and surface the exact block, filter, response, and endpoint. Range splitting cannot repair pruned history, a malformed filter, or an incapable upstream.

There is no universal safe width. The same number of blocks can return zero logs for one contract and an enormous response for another. Keep separate learned widths per chain and filter class, and consider a result-count or response-byte ceiling in addition to a latency threshold.

Section 04

Which errors should shrink the range, retry, or stop the job?

Classify the failure before acting. Blind retries can turn a bounded problem into a throttling incident.

  • Range, response-size, or deadline failure: reduce the page width and retry the same interval. Keep a minimum width and a total attempt budget.
  • HTTP 429: honor Retry-After when present, pause the affected worker, and resume with exponential backoff plus jitter. Lower concurrency before assuming the page itself is too wide.
  • Transient connection error or 5xx: retry within a bounded end-to-end deadline, preferably against a qualified fallback that retains the same log history.
  • Pruned history: move the job to an endpoint that retains the required receipts and logs. The current Execution API specifies a Pruned history unavailable error for eth_getLogs; making the range smaller does not restore deleted history.
  • Invalid parameters or unsupported method: fix the filter or route. Another retry of the same invalid request is waste.
  • Empty result: accept it as data. [] is the correct result when no matching logs exist; do not rotate providers merely because a page is empty.

Persist the normalized error class, original provider response, range, width, attempt count, and delay. That record explains whether progress stopped because of density, capacity, throttling, retention, or a client-specific response.

Section 05

How should logs and the cursor be stored?

Make every page safe to replay. Store at least chain ID, block number, block hash, transaction hash, transaction index when present, log index, contract address, topics, data, and decoded fields. Preserve the raw log so a decoder upgrade does not require another RPC backfill.

Use a uniqueness key that includes canonical identity rather than an application field such as token ID. One useful model is chain ID plus block hash plus log index; retaining the transaction hash as well makes diagnostics and joins clearer. A transaction can emit several logs, and a reorganization can place the same transaction on a different block, so transaction hash alone is insufficient.

Commit the logs, processed block identities, and next cursor atomically. If the process dies after inserting logs but before moving the cursor, an idempotent replay is harmless. If it moves the cursor before the inserts are durable, the page becomes a silent gap. Keep decoding failures in a dead-letter or error table and fail the page unless the product has explicitly approved skipping that event type.

The cursor should identify the job configuration as well as the next block. Include a filter or schema version so changing addresses, topics, or decoding rules cannot accidentally continue an incompatible run.

Section 06

How do you handle chain reorganizations?

Finalized-only ingestion reduces the repair surface, but any job that approaches the head needs an explicit fork policy. Store block number, hash, and parent hash for the unfinalized window. Before committing a new block or page, verify continuity with the last accepted block. On a mismatch, walk backward until both the stored chain and the endpoint agree, remove or mark data from orphaned blocks, reset the cursor to the next block after the common ancestor, and replay.

Do not rely only on the removed field from a range query. Geth documents that a live log subscription can resend old-chain logs with removed: true and may emit the same transaction's logs more than once during a reorganization. A polling worker or a disconnected subscriber can miss those removal notifications, so stored block hashes remain the durable source of reconciliation.

When checking or repairing one known block, prefer an eth_getLogs filter by blockHash where the endpoint supports it. EIP-234 added that option specifically to bind the returned logs to the intended block and avoid a race where a block number resolves to a different hash during a reorganization. blockHash is mutually exclusive with fromBlock and toBlock.

Section 07

How do you hand off from backfill to live indexing?

A safe handoff closes the gap between the fixed historical target and the live consumer. The exact sequence depends on the transport.

For polling, finish the historical job at its fixed target, resolve a new explicit upper block, and continue with smaller overlapping or adjacent ranges under the same cursor and reorg policy. Polling is often simpler operationally because recovery always uses the same deterministic range path.

For WebSocket subscriptions, establish the subscription and record the block observed at connection time, then backfill from the last committed cursor through that block before trusting the live stream. Deduplicate subscription messages against the database. A subscription begins with what the connection observes; it does not prove that no block was missed while the process was disconnected. Geth's real-time event documentation also warns that subscriptions are coupled to the connection and can deliver duplicate transaction logs around a reorganization.

Keep the range backfiller even after live indexing starts. It is the recovery path for deployments, disconnects, provider incidents, decoder bugs, and schema rebuilds.

Section 08

How can you verify that the backfill is complete?

Completeness is a property of the process, not the total row count. Record enough evidence to reproduce each run:

  • The first range begins at the configured start block, the final range ends at the fixed target, and every adjacent range satisfies next.fromBlock = previous.toBlock + 1.
  • Every committed page has its request filter, endpoint, attempt history, result count, response size, duration, and cursor transaction recorded.
  • Every stored log's block number falls inside the requested page and its address and topics satisfy the filter.
  • Known transactions near the beginning, dense periods, sparse periods, and the end of the range decode to expected events.
  • A sample of pages is repeated by blockHash; compare identities and raw fields, not only decoded business rows.
  • The unfinalized window's stored block hashes still match the canonical chain before the job is declared current.
  • Restart, timeout, duplicate-delivery, 429, provider failover, and reorg drills leave the same canonical dataset.

For high-value data, compare selected ranges through an independently qualified endpoint. Agreement is useful evidence, but it is not a substitute for deterministic boundaries and durable cursors: two endpoints can share the same client, pruning policy, or upstream source.

Section 09

How should you run a log backfill through SolidRPC?

Use an authenticated HTTPS endpoint in the form https://rpc.solidrpc.io/YOUR_API_KEY/evm/<chainId> for a sustained backfill. On activated networks, eligible paid accounts can use native WebSocket newHeads to trigger live reads. Fetch logs through ordinary RPC and reconcile missed blocks after reconnecting. Log subscriptions are not supported. HTTPS polling remains available. See the WebSocket guide for coverage, shared response-unit billing and the monetary SLA exclusion. Standard calls have a 30-second server-side budget; choose a client deadline that can receive the response, while still shrinking slow pages instead of treating the maximum timeout as a target.

The keyless public endpoints are for development and lightweight traffic, not production indexing. Their eth_getLogs policy bounds each inclusive range per chain — and grants a wider range to a query that carries an address or topics filter — up to a ceiling of 2,000 blocks, and applies shared capacity controls; the current per-chain limits and error shapes are documented in the public RPC policy. An authenticated route removes that public-policy range cap, but the practical safe width still depends on event density, response size, chain, and upstream capability.

SolidRPC routes each request to capacity qualified for that network and method, with external capacity covering failover and anything it does not serve directly. Check the current network catalog and test the oldest required range through the production route. Logs require retained receipt and log history, not automatically historical world state; the archive-node decision guide explains that distinction. A fallback is qualified for this job only if it returns the same required history and filter behavior.

Section 10

eth_getLogs backfill checklist

  • Record chain ID, addresses, topics, start block, fixed target block, and filter version.
  • Test the exact filter against known transactions before the full run.
  • Use inclusive, adjacent ranges with no overlaps or gaps.
  • Start conservatively and adapt width from observed latency, errors, result count, and bytes.
  • Limit concurrency separately from page width.
  • Treat empty arrays as valid results and classify errors before retrying.
  • Never advance the cursor for a failed or partially committed page.
  • Store raw logs, block identity, decoded data, and the cursor atomically.
  • Make inserts idempotent and safe to replay after a crash.
  • Stop on a repeatedly failing single block instead of silently skipping it.
  • Use a finalized boundary or implement block-hash continuity and rollback.
  • Use blockHash queries for single-block verification and reorg repair.
  • Close the historical-to-live gap explicitly; keep the backfiller as the recovery path.
  • Verify boundaries, representative known events, restart behavior, failover, and reorg recovery.
  • Monitor progress rate, head distance, width changes, retries, throttling, response bytes, and oldest unresolved error.

The durable design principle is simple: a page is either fully and reproducibly committed, or the cursor does not move. That rule turns eth_getLogs from a fragile loop into an auditable ingestion pipeline.

Workload review

Want us to benchmark your current RPC stack?

Send your providers, routing setup, method mix, or monthly bills. We will compare the complete stack with one SolidRPC integration across cost, useful completion rate, failover, and operational work.