feat(beacon): time the state-writer hand-off as its own import phase - #651
Draft
MegaRedHand wants to merge 2 commits into
Draft
MegaRedHand wants to merge 2 commits into
MegaRedHand wants to merge 2 commits into
Conversation
Beacon post-states are persisted by one background thread behind a bounded queue, and Store::insert_state blocks on it while the queue is full. That call sits inside fork_choice::on_block, which the chain actor times as the stf phase, so when blocks import faster than the writer persists them the importer's wait was reported as state-transition time. StateWriterHandle::send now returns the instants around its blocking send, insert_state keeps the span (first start, last end) and Store::take_state_handoff hands it to the chain actor, which records it as writer_wait_start/end. The report gains a writer_wait row after stf, and the phase label is added to lean_block_import_phase_seconds. Semantics change: beacon stf now excludes the writer wait (saturating subtraction in the report), so the rows still sum to the end-to-end time. Lean never records a wait, so its report is unchanged.
The replay feeds blocks back to back, while a live node gets one per slot. When the build imports faster than the background state writer persists post-states, the writer's queue fills and the importer blocks in writer_wait, so the replay measures the writer instead of the import (one merged build waited 326 ms per block). --block-delay <MILLISECONDS> (default 0, today's behavior) sleeps after each block's import, warm-up included, once its phases and wall time are recorded, so the sleep sits in no per-block phase, no wall_seconds and no summary row. The delay is recorded as params.block_delay_ms in the JSON report and printed in the human header.
This was referenced Oct 2, 2026
MegaRedHand
added a commit
that referenced
this pull request
Oct 2, 2026
…3-636-638-gloas-live Conflicted files: - crates/storage/src/store.rs (the std imports: this branch's MutexGuard plus #651's Instant) No semantic adaptation was needed: the writer_wait phase, the state handoff timing and the replay --block-delay flag do not touch gloas-specific code, and the import timing rows are fork-agnostic. The branch sits on bci 125dce1, so this also brings main's #641 (one dedicated prover thread) with its crypto tests. Cargo.lock: the merge adds the zk_alloc dev-dependency to ethlambda-crypto, and the stale lockfile this branch inherits records zk_alloc under another crate. The lockfile is left exactly as on the first parent; builds here run without --locked.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
Beacon post-states are persisted by one background thread fed through a bounded queue.
StateWriterHandle::sendblocks while the queue is full, andStore::insert_statecalls it from insidefork_choice::on_block, which the chain actor times as thestfphase. If blocks import faster than the writer persists them, the importer's wait for the writer is reported as state-transition time. This PR reports the wait as its own phase, sostfmeasures state-transition work only.Measured on the import-replay benchmark (128 mainnet blocks, eth-5) with this change merged into each leg:
stfp50 (now exclusive)writer_waitp50With all five merged, blocks import faster than the writer persists their states, so the queue fills and the importer waits about 326 ms per block. Without this phase that wait was reported as
stf(about 530 ms instead of 202 ms), which hid most of the combined gain and pointed at the wrong place. The state writer, not the state transition, is now the replay's bottleneck.Changes
storage/state_writer.rsStateWriterHandle::sendreturns the(start, end)instants around its blockingsendstorage/store.rsinsert_staterecords the hand-off span (first start, last end if repeated) in a shared field; newStore::take_state_handoff()takes and clears itblockchain/import_timing.rswriter_wait_start/writer_wait_endonImportTimingsandStoreTimings, copied inabsorb_store; newwriter_waitrow afterstf; header documents the nestingblockchain/lib.rson_blockreturns (bothOkandErr); clears any stale span before the callblockchain/metrics.rswriter_waitadded toBLOCK_IMPORT_PHASES(thephaselabel list)docs/metrics.md,docs/benchmarking.mdstfnow excludes the waitThe replay benchmark builds its
phasesmap fromBLOCK_IMPORT_PHASESand the same histogram, so it picks upwriter_waitand the exclusivestfwith no code change.One beacon
on_blockmakes oneinsert_statecall. The store still merges repeated calls (first start, last end), the module's convention for repeated sections.Semantics change
Beacon
stf(the report row and thelean_block_import_phase_seconds{phase="stf"}series) now excludes the writer wait: it is theon_blockspan minuswriter_wait, saturating at zero. The rows still sum to the end-to-end time. Lean blocks never set the new instants, so theirstfand report are unchanged. A newphase="writer_wait"series appears on beacon only. Dashboards comparing beaconstfacross this change will see it drop by the wait.Testing
cargo fmt --allcargo clippy --workspace --all-targets --profile release-fast -- -D warningscargo test -p ethlambda-storage --profile release-fast --libcargo test -p ethlambda-blockchain --profile release-fast --libcargo check -p ethlambda --profile release-fastNew tests:
stfminuswriter_wait, wait absent, wait longer thanstfsaturating to zero, row order and presence (existing label-list test also covers the new label), and a storage test thatinsert_staterecords a start <= end pair and that taking it clears it. Spec suites and the replay benchmark were not run.Clippy ran without
--locked:Cargo.lockon the base (125dce1) is already stale (zk_allocis listed under the wrong crate), so--lockedfails there with or without this change. The lockfile is not part of this PR.