SolidRPC plans change on 28th October. Some prices are going up. Lock in today’s pricing and limits by activating before then. See what’s changing.
Skip to main content
Guides
Category
Operations
Reading time
12 minutes
Published

Ethereum WebSocket vs HTTP polling for event indexing

WebSocket subscriptions look like the natural way to index EVM events in real time: open a connection, call eth_subscribe, and process each matching log as it arrives. HTTP polling can look wasteful by comparison because the indexer repeatedly asks for blocks or logs, including intervals with no matches.

That comparison misses the requirement that matters most. A live indexer must recover after a deployment, a broken connection, a slow consumer, an upstream incident, or a chain reorganization without silently losing or double-counting events. A subscription is a delivery channel, not a durable cursor. Polling can be low-latency enough for many products, and a WebSocket design still needs an HTTP range-reconciliation path.

This guide compares the two transports by latency, correctness, concurrency, recovery, cost, and operational ownership. It then gives you production designs for polling-only and subscription-assisted indexers, plus a test plan for choosing between managed RPC and self-hosted nodes.

Section 01

What is the short answer?

Use WebSocket subscriptions when the product benefits materially from push delivery between new blocks or needs an event family that is naturally connection-oriented. Use head-aware HTTP polling when block-level discovery latency is acceptable, operational simplicity matters, or the endpoint does not offer subscriptions.

For durable log indexing, neither transport changes the source of truth: an explicit canonical block range plus a persisted block-number-and-hash cursor. Geth's subscription documentation says notifications cover current rather than past events and disappear when the connection closes. The standardized `eth_getLogs` interface can query an explicit range or one block hash, which makes it suitable for catch-up and reconciliation.

The strongest general design is therefore not “WebSocket instead of polling.” It is “WebSocket as a low-latency signal, range reads as the recovery mechanism.” If the latency saved by push delivery does not change the user experience or business outcome, polling alone is usually the smaller system to operate.

Section 02

What does each transport actually guarantee?

A successful WebSocket connection gives the client a full-duplex channel. After eth_subscribe succeeds, the node can push matching notifications without a new HTTP request for each event. The subscription ID belongs to that connection; it is not a server-side checkpoint the application can resume on another socket.

An HTTP range request is discrete. The client supplies fromBlock and toBlock, or a blockHash, and receives the matching logs for that explicit scope. It does not push the next event automatically, but the request can be repeated, audited, split, and replayed from a durable application cursor.

Do not confuse eth_newFilter plus eth_getFilterChanges with stateless range polling. A filter ID is node-side state and can expire or disappear when the request reaches another node. For a recoverable indexer, keep the authoritative cursor in your own database and use explicit eth_getLogs ranges.

The practical guarantee is asymmetric: a subscription can tell you quickly that something happened, while a block-bound range query can prove what the application asked the node to return. Neither response proves finality, canonical continuity, successful decoding, or a durable database commit; the indexer must enforce those properties.

Section 03

Can an Ethereum WebSocket subscription miss or duplicate logs?

Yes. A disconnect creates an interval in which the client receives no subscription messages, and a new subscription starts with current events rather than replaying the missing interval. Geth also documents that subscriptions are removed when their connection closes.

Backpressure is another failure mode. Geth buffers notifications for a slow client and closes the connection when its current buffer limit is reached. That limit is implementation-specific, so do not treat one client's number as a portable capacity guarantee. Keep the socket reader fast, move decoding and database writes to a bounded queue, and reconnect before an overloaded consumer turns delay into an unobserved gap.

Reorganizations create valid duplicates and removals. Geth's logs subscription can resend old-chain logs with removed: true, then emit the new-chain logs; the same transaction can appear more than once. Its newHeads subscription can emit multiple headers at the same height, and when Geth imports several blocks together it may emit only the last header. A consumer that assumes one message per height or one delivery per transaction is incorrect even while the connection stays open.

Store chain ID, block number, block hash, parent hash, transaction hash, and log index. Deduplicate by canonical log identity, retain enough recent block history to find a common ancestor, and make rollback plus replay a normal operation rather than an incident-only script.

Section 04

How do you build a subscription-assisted indexer without gaps?

Use the subscription as an accelerator around a deterministic backfiller:

  1. Load the last committed block number and hash from durable storage.
  2. Open the WebSocket, create the narrowest correct logs or newHeads subscription, and start buffering notifications.
  3. Resolve an explicit catch-up target after the subscription is active. Save its number and hash.
  4. Fetch every uncommitted block from the cursor through that target over a qualified range-read path. Validate parent-hash continuity and upsert logs idempotently.
  5. Reconcile buffered notifications against the committed blocks, then process new heads in order. Fetch by block hash where supported instead of trusting notification arrival order.
  6. Commit application rows and the block cursor atomically. A crash must leave either the whole block committed or the cursor unchanged.
  7. On any close, stalled heartbeat, queue overflow, provider switch, or process restart, discard the ephemeral subscription ID, reconnect, resubscribe, and repeat the range reconciliation before declaring the stream current.

EIP-234 added the blockHash log filter because network failure during a reorganization can leave a subscription consumer unable to determine which block an empty result described. Use that hash-bound query for single-block verification and repair. It is mutually exclusive with fromBlock and toBlock.

Keep historical catch-up and live consumption on separate concurrency budgets. A large backlog should not starve the socket reader, and a flood of notifications should not prevent the range worker from closing a recovery gap.

Section 05

How do you make HTTP polling low-latency and correct?

Polling does not require a fixed sleep followed by one request per contract. Make it chain-tip-aware:

  1. Read an explicit target head or the finality tag your product accepts. Ethereum JSON-RPC defines latest, safe, and finalized, but other EVM chains and clients can implement different finality semantics, so qualify the actual route.
  2. If the target is ahead of the durable cursor, request adjacent eth_getLogs ranges with the narrowest correct address and topic filters. Use a conservative width and adapt it to density, latency, response size, and provider errors.
  3. Fetch or retain the block hashes that anchor the range, validate parent continuity, and commit logs plus the next cursor in one transaction.
  4. If no new block exists, wait according to the observed block cadence and latency objective rather than hammering the endpoint at a universal interval. Add jitter across workers so they do not synchronize their requests.
  5. Re-read and reconcile the unfinalized window until the chosen finality boundary passes it.

The detailed paging, retry, cursor, and completeness mechanics are covered in the gap-free `eth_getLogs` backfill guide. The live poller should use the same recovery code with smaller targets and a tighter schedule, not a separate ingestion path with different correctness rules.

Polling latency is bounded by the chain's block production, the polling schedule, request time, and commit time. Measure the application-visible delay from block timestamp or observation to durable event availability. Do not optimize request frequency before you know which part of that path dominates.

Section 06

How should reconnects, heartbeats, and failover work?

A WebSocket can fail without a clean close frame. Use the client library's ping/pong support or an application-level liveness check, enforce an idle deadline, and reconnect with exponential backoff plus jitter. RFC 6455 defines ping frames as a keepalive or responsiveness check, but a successful pong proves only that the peer answered—it does not prove that the execution client is synced or that new heads are advancing.

Track both transport liveness and chain freshness. Useful signals include connection age, reconnect count, close code, seconds since the last head, observed head number and hash, lag from a qualified reference, local queue depth, oldest uncommitted block, reconciliation duration, duplicate and removed-log counts, and rollback depth.

Do not expect a load balancer to make an existing subscription survive upstream replacement. Because the subscription is coupled to the connection, losing that connection means the application must create a new subscription and backfill the gap. Transparent failover is useful for new range requests, but it cannot turn an ephemeral subscription ID into a portable cursor.

Bound reconnect concurrency so a regional outage does not make every worker reconnect at once. Preserve capacity for range recovery, and stop admitting live work when the durable gap grows beyond the limit your system can safely buffer.

Section 07

Managed WebSocket RPC or a self-hosted node?

A managed WebSocket endpoint shifts client upgrades, node sync, public endpoint exposure, and much of the connection infrastructure to the provider. It does not shift the application's cursor, deduplication, reorg rollback, schema recovery, or missed-range reconciliation. Before buying, test the exact chain and subscription type, concurrent connection and subscription policy, idle timeout, message or bandwidth limits, slow-consumer behavior, historical range path, and capability after provider failover.

Self-hosting gives you control over the execution client, enabled namespaces, network boundary, buffering behavior, and capacity planning. It also makes you responsible for a synced execution and consensus stack where applicable, TLS and reverse proxies, authentication, origin policy, rate limiting, observability, upgrades, storage, rebuilds, and an independent recovery path. Geth's current RPC server documentation requires WebSocket to be enabled and its exposed APIs and allowed origins to be configured. Its security guidance warns that node APIs are not designed to face hostile public traffic directly and recommends protective proxies, filtering, rate limits, logging, TLS termination, and monitoring.

Neither model is inherently more reliable. A managed endpoint can still close connections or change limits; one self-hosted node can be a larger single point of failure than a managed fleet. Compare architectures that meet the same chain, method, history, freshness, concurrency, and recovery requirements—not one provider URL against one unprotected node process.

Section 08

Which option is cheaper at your workload shape?

There is no transport-independent cost winner. Polling cost grows with poll frequency, number of filters, empty intervals, range retries, and the provider's method-billing model. WebSocket cost can depend on connections, subscriptions, notification volume, egress, connection duration, or plan-specific units, and it still includes HTTP catch-up after every gap. Self-hosting replaces provider units with compute, storage, network, proxy capacity, observability, upgrades, recovery, and engineering time.

Model the same successful output in three scenarios:

  • Normal live indexing: connections, subscriptions, head checks, range requests, notifications, bytes, and durable events per day.
  • Recovery: a realistic disconnect or deployment gap, the backfill calls it creates, duplicate work, and time to become current.
  • Rebuild or incident: a schema replay, provider degradation, or node rebuild while live traffic continues.

Use provider documentation or a current quote for every billing unit; do not infer that one pushed event equals one free request. For self-hosting, include the redundant capacity and operator time required to achieve the same recovery target. Compare cost per canonical block or event committed within the service objective, not cost per socket message or advertised RPC call.

Section 09

When should you choose WebSocket, polling, or both?

Choose HTTP polling when block-level latency is acceptable, the endpoint is HTTPS-only, filters are selective, and a simple deterministic recovery path is more valuable than push delivery. It is especially natural for finalized accounting, analytics, portfolio data, and indexers whose downstream work is already measured in blocks.

Choose subscription-assisted indexing when faster discovery has a measured product value, the provider and chain expose the required subscription, and the team is prepared to operate a stateful connection plus a range backfiller. New-head notifications are often a cleaner wake-up signal than subscribing to a very broad log stream because the application can fetch and verify each intended block itself. Test that choice with the actual client: Geth documents client-specific behavior when several heads are imported together.

Choose a specialized stream or webhook product only when its transformed schema, delivery semantics, replay window, and vendor coupling fit the product better than raw JSON-RPC. Treat it as a data product, not as a transparent replacement for eth_subscribe.

Do not choose WebSocket only to avoid empty polls, and do not choose polling only because it is easier to prototype. Set a maximum event-discovery delay, a maximum missed-range recovery time, a permitted reorg exposure, and an operating-cost envelope. The smallest architecture that passes all four is the right one.

Section 10

How does live event indexing work with SolidRPC?

SolidRPC combines HTTPS JSON-RPC with native newHeads subscriptions and ordinary RPC over WSS on activated networks for eligible paid accounts. Check the WebSocket guide for live network availability. V1 does not offer log or pending-transaction subscriptions. Each notification written to the connection costs one existing response unit, shared with HTTP/WS calls and account rate limits. WebSocket traffic is excluded from monetary SLA credits. A live indexer should use heads as a wake-up signal and advance through bounded eth_getLogs ranges with a durable block-number-and-hash cursor. Reconnect, resubscribe and reconcile gaps through SolidRPC.

On supported networks, SolidRPC serves recent and historical workloads through one managed route. Qualified external capacity provides failover and covers networks or capabilities SolidRPC does not serve directly. That managed routing absorbs upstream selection and node incidents within the service boundary; the application still owns its filter, canonical cursor, atomic writes, reorg rollback, and completeness tests.

Use the live network catalog to verify the chain, node type, and method family before a backfill. Historical logs depend on retained receipt and log history, while old state and trace namespaces are separate requirements. If a product depends on subscriptions other than newHeads, confirm support before migrating. SolidRPC supports ordinary JSON-RPC calls and newHeads over WebSocket on selected networks. If block-aware polling meets the latency target, SolidRPC can provide one managed HTTPS integration without asking the application to operate its own provider fallback pool.

Section 11

EVM live-indexing decision checklist

  • Define the maximum acceptable delay from block observation to a durable application event.
  • Record required chains, addresses, topics, finality policy, and oldest recoverable block.
  • Keep the source-of-truth cursor in your database as block number plus hash.
  • Use explicit eth_getLogs ranges or blockHash; do not rely on a remote filter ID as the durable cursor.
  • If using WebSocket, subscribe before closing the catch-up gap and buffer live notifications during reconciliation.
  • Make block data, decoded events, and cursor advancement one atomic commit.
  • Deduplicate deliveries and implement rollback to a common ancestor.
  • Treat socket liveness and chain-head freshness as separate health checks.
  • Reconnect with backoff and jitter, resubscribe, and backfill every disconnected interval.
  • Bound socket-reader queues, HTTP concurrency, range width, retries, and recovery traffic independently.
  • Test slow consumers, clean and unclean closes, stale heads, provider failover, duplicate logs, removed logs, and reorgs.
  • Compare managed and self-hosted options against the same method, history, concurrency, security, and recovery contract.
  • Price normal, disconnect-recovery, and full-rebuild scenarios using current documented billing or infrastructure inputs.
  • Re-run the qualification after changing clients, providers, load balancers, pruning, finality policy, or chain support.

A reliable indexer does not trust the transport to remember history. It uses the transport to discover work and its own canonical cursor to prove that every block was processed exactly as intended.

Workload review

Want us to benchmark your current RPC stack?

Send your providers, routing setup, method mix, or monthly bills. We will compare the complete stack with one SolidRPC integration across cost, useful completion rate, failover, and operational work.