Skip to content

feat(beacon): time the state-writer hand-off as its own import phase - #651

Draft
MegaRedHand wants to merge 2 commits into
beacon-chain-integrationfrom
feat/beacon-import-writer-wait
Draft

MegaRedHand wants to merge 2 commits into
beacon-chain-integrationfrom
feat/beacon-import-writer-wait

Conversation

@MegaRedHand

@MegaRedHand MegaRedHand commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

Beacon post-states are persisted by one background thread fed through a bounded queue. StateWriterHandle::send blocks while the queue is full, and Store::insert_state calls it from inside fork_choice::on_block, which the chain actor times as the stf phase. If blocks import faster than the writer persists them, the importer's wait for the writer is reported as state-transition time. This PR reports the wait as its own phase, so stf measures state-transition work only.

Measured on the import-replay benchmark (128 mainnet blocks, eth-5) with this change merged into each leg:

leg stf p50 (now exclusive) writer_wait p50 block wall p50
base (beacon-chain-integration) 581 ms ~0 777 ms
each single performance PR (#647-#650, #633) 353-618 ms ~0 (max 4 us) 551-778 ms
all five performance PRs merged 202 ms 326 ms 559 ms

With all five merged, blocks import faster than the writer persists their states, so the queue fills and the importer waits about 326 ms per block. Without this phase that wait was reported as stf (about 530 ms instead of 202 ms), which hid most of the combined gain and pointed at the wrong place. The state writer, not the state transition, is now the replay's bottleneck.

Changes

Area Change
storage/state_writer.rs StateWriterHandle::send returns the (start, end) instants around its blocking send
storage/store.rs insert_state records the hand-off span (first start, last end if repeated) in a shared field; new Store::take_state_handoff() takes and clears it
blockchain/import_timing.rs writer_wait_start/writer_wait_end on ImportTimings and StoreTimings, copied in absorb_store; new writer_wait row after stf; header documents the nesting
blockchain/lib.rs Beacon arm takes the hand-off instants right after on_block returns (both Ok and Err); clears any stale span before the call
blockchain/metrics.rs writer_wait added to BLOCK_IMPORT_PHASES (the phase label list)
docs/metrics.md, docs/benchmarking.md Phase documented; beacon stf now excludes the wait

The replay benchmark builds its phases map from BLOCK_IMPORT_PHASES and the same histogram, so it picks up writer_wait and the exclusive stf with no code change.

One beacon on_block makes one insert_state call. The store still merges repeated calls (first start, last end), the module's convention for repeated sections.

Semantics change

Beacon stf (the report row and the lean_block_import_phase_seconds{phase="stf"} series) now excludes the writer wait: it is the on_block span minus writer_wait, saturating at zero. The rows still sum to the end-to-end time. Lean blocks never set the new instants, so their stf and report are unchanged. A new phase="writer_wait" series appears on beacon only. Dashboards comparing beacon stf across this change will see it drop by the wait.

Testing

Command Result
cargo fmt --all clean
cargo clippy --workspace --all-targets --profile release-fast -- -D warnings clean
cargo test -p ethlambda-storage --profile release-fast --lib 139 passed, 1 ignored
cargo test -p ethlambda-blockchain --profile release-fast --lib 210 passed, 0 ignored
cargo check -p ethlambda --profile release-fast clean

New tests: stf minus writer_wait, wait absent, wait longer than stf saturating to zero, row order and presence (existing label-list test also covers the new label), and a storage test that insert_state records a start <= end pair and that taking it clears it. Spec suites and the replay benchmark were not run.

Clippy ran without --locked: Cargo.lock on the base (125dce1) is already stale (zk_alloc is listed under the wrong crate), so --locked fails there with or without this change. The lockfile is not part of this PR.

Beacon post-states are persisted by one background thread behind a bounded
queue, and Store::insert_state blocks on it while the queue is full. That
call sits inside fork_choice::on_block, which the chain actor times as the
stf phase, so when blocks import faster than the writer persists them the
importer's wait was reported as state-transition time.

StateWriterHandle::send now returns the instants around its blocking send,
insert_state keeps the span (first start, last end) and
Store::take_state_handoff hands it to the chain actor, which records it as
writer_wait_start/end. The report gains a writer_wait row after stf, and the
phase label is added to lean_block_import_phase_seconds.

Semantics change: beacon stf now excludes the writer wait (saturating
subtraction in the report), so the rows still sum to the end-to-end time.
Lean never records a wait, so its report is unchanged.
@MegaRedHand MegaRedHand added the beacon Ethereum Beacon Chain client label Oct 2, 2026
The replay feeds blocks back to back, while a live node gets one per slot.
When the build imports faster than the background state writer persists
post-states, the writer's queue fills and the importer blocks in
writer_wait, so the replay measures the writer instead of the import (one
merged build waited 326 ms per block).

--block-delay <MILLISECONDS> (default 0, today's behavior) sleeps after each
block's import, warm-up included, once its phases and wall time are
recorded, so the sleep sits in no per-block phase, no wall_seconds and no
summary row. The delay is recorded as params.block_delay_ms in the JSON
report and printed in the human header.
MegaRedHand added a commit that referenced this pull request Oct 2, 2026
…3-636-638-gloas-live

Conflicted files:
- crates/storage/src/store.rs (the std imports: this branch's MutexGuard
  plus #651's Instant)

No semantic adaptation was needed: the writer_wait phase, the state handoff
timing and the replay --block-delay flag do not touch gloas-specific code, and
the import timing rows are fork-agnostic. The branch sits on bci 125dce1, so
this also brings main's #641 (one dedicated prover thread) with its crypto
tests.

Cargo.lock: the merge adds the zk_alloc dev-dependency to ethlambda-crypto,
and the stale lockfile this branch inherits records zk_alloc under another
crate. The lockfile is left exactly as on the first parent; builds here run
without --locked.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

beacon Ethereum Beacon Chain client

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant