Skip to content

perf(beacon): write epoch balances through an in-order copy-on-write cursor - #647

Draft
MegaRedHand wants to merge 2 commits into
beacon-chain-integrationfrom
perf/ssz-tree-write-cursor
Draft

MegaRedHand wants to merge 2 commits into
beacon-chain-integrationfrom
perf/ssz-tree-write-cursor

Conversation

@MegaRedHand

Copy link
Copy Markdown
Collaborator

Motivation

An epoch boundary writes every balance. get_mut on a tree list buffers a copy of each element in a dense map, sorts the map into a second Vec, then rebuilds every touched leaf and recomputes the root of every written element, changed or not. The altair rewards step calls it up to eight times per validator (four reward/penalty pairs, each as increase then decrease), and process_effective_balance_updates reads balances()[i] per validator, which was only cheap because the rewards pass had just filled the map.

This PR adds an in-order lazy copy-on-write write cursor to ethlambda-ssz-tree and moves the epoch balance writes onto it. Expected gain (estimate, not measured on a real chain): tens of ms off each epoch-start block at mainnet registry size. The import-replay benchmark was not run for this PR.

Changes

commit what
perf(ssz-tree): add an in-order copy-on-write write cursor List/Vector::iter_cow, IterCow::next_cow, ElemCow::{index, make_mut, set}, try_update_each; with_updated_leaves generalized over a leaf source; unit tests and model-based property tests
perf(beacon): apply epoch balance and effective-balance writes through the cursor BeaconState::registry_mut (disjoint borrows of validators and balances); apply_balance_deltas used by the altair+ and phase0 process_rewards_and_penalties; both process_effective_balance_updates; new tree_bench rows; docs

How the cursor works: it walks leaves in order and copies a leaf on the first write that reaches it. At the leaf's end it compares only the touched offsets with the original; if all are equal, the original leaf (its Arc and cached hash) is kept. Otherwise the new leaf carries over the element roots of unchanged elements and computes the changed ones. On drop, one descent swaps the changed leaves in and rebuilds only their paths. Pending buffered writes are applied before the pass starts.

Design decisions

  • Drop does the work, and the tree is untouched during the pass, so an early return or a panic leaves a consistent list holding the writes so far (tested with catch_unwind).
  • Names as in the design: iter_cow, next_cow, make_mut, set, try_update_each. No set_with.
  • Changed composite elements get their root computed eagerly, on the state-transition thread.
  • A kept Arc or carried root always follows an actual comparison. The cursor only creates fresh empty hash cells, or carried roots for a leaf it owns; it never fills a shared OnceLock.
  • Spec order for balances: per index, each pair is saturating_add(reward) then saturating_sub(penalty), pairs in order. Deltas are never netted. Each balance depends only on its own deltas, so index-by-index equals the spec's pair-by-pair loop.
  • apply_balance_deltas lives in helpers/mutators.rs and is shared by the phase0 and altair+ paths. It keeps the old error for a balances list shorter than the registry, and does not touch balances past the registry length.
  • Rewards and both effective-balance passes move together: once rewards bypass the map, the effective-balance reads would otherwise become tree descents.
  • Effective-balance updates only copy a leaf when the new effective balance differs from the old one (it can equal it when capped at the maximum).
  • No leaf-slice API, no parallel bulk update.

Not in this PR

Testing

command result
cargo test -p ethlambda-ssz-tree --profile release-fast 71 unit + 31 model tests passed (before the second commit)
cargo test -p ethlambda-types --profile release-fast --lib 197 passed
cargo test -p ethlambda-state-transition --profile release-fast --lib 324 passed, 16 ignored
beacon spec suite, mainnet preset, full 5734 passed, 4 failed
beacon spec suite, minimal preset, full 40038 passed, 4 failed
cargo clippy --locked --workspace --all-targets --profile release-fast -- -D warnings and the state-transition beacon-spec-tests variant clean
cargo check -p ethlambda --profile release-fast ok

The 4 spec failures in each preset are the gossip/* suites (gossip_beacon_block, gossip_data_column_sidecar, gossip_beacon_aggregate_and_proof, gossip_beacon_attestation): the local gossip fixture directory is intentionally empty, so they match no cases. Every other suite, including epoch_processing, rewards, sanity, finality, random and transition, passes.

tree_bench (2.4M validators, --ignored, single run)

Noisy: run while the machine was under a load average above 100 from other builds, so absolute numbers (and even the order of the old rows) are unreliable; read them as rough ratios only.

row ms
epoch_sweep_tree (old path) 333.8
epoch_sweep_tree_phases write 58.1, apply 70.0, hash 83.1
epoch_sweep_cursor_tree 26.8
epoch_sweep_noop_cursor_tree 90.4
epoch_sweep_rebuild_tree (rebuild baseline) 105.8
rewards_apply_4_pairs_get_mut 342.5
rewards_apply_4_pairs_cursor 158.9
eb_updates_indexed 253.8
eb_updates_cursor 115.4
lockstep_three_lists_cursor 136.2

The bench asserts the cursor rows produce the same roots as the old paths, and that a no-op sweep keeps the tree (ptr_eq).

Rewriting most of a list through get_mut buffers an entry per element,
sorts the buffer into a second Vec, and then rebuilds every touched leaf,
recomputing the root of every written element even when the value did not
change.

iter_cow / try_update_each walk the leaves in order instead. A leaf is
copied on the first write that reaches it, compared with the original at
the leaf's end (only the offsets make_mut handed out), and kept as it was
(Arc and cached hash) if nothing differs. Changed leaves carry over the
element roots of unchanged elements and compute the changed ones. The
changed leaves are swapped in on drop with one descent, so an early
return or a panic leaves a consistent list. Pending writes are applied
before the pass starts.

The with_updated_leaves descent is generalized over a leaf source so the
element-update path and the whole-leaf replacement share it. Covered by
unit tests (changed-path sharing, carried roots, early return,
catch_unwind, registry depth) and model-based property tests against
SszList.
…h the cursor

Both rewards-and-penalties passes wrote every balance through get_mut, up
to eight calls per validator on altair and later, each a map entry in the
dense pending-write buffer. They now apply every (rewards, penalties)
pair per index in one cursor pass, keeping the spec's per-pair
saturating_add then saturating_sub order (no netting).

Both process_effective_balance_updates variants read balances()[index]
per validator, which was a buffer hit only because the rewards pass had
just filled the buffer. With the rewards write bypassing the buffer those
reads would become tree descents, so they move in the same change: the
registry's cursor runs zipped with a balances iterator, through the new
BeaconState::registry_mut disjoint borrow, and copies a leaf only for a
validator whose effective balance actually changes.

Adds the tree_bench rows comparing the old and new write paths, and the
docs for the cursor.
@MegaRedHand

Copy link
Copy Markdown
Collaborator Author

Import-replay benchmark

Method. ethlambda benchmark import replay, release build, 128 consecutive mainnet blocks (slots 15279073..=15279200, 4 of them epoch-start blocks) on a 16-core host. This PR's leg and a base leg (beacon-chain-integration) were interleaved, 2 rounds each; the table shows the mean of the two. Both legs carry #651, so stf excludes the state-writer hand-off wait. Block roots were identical to the base in every leg. Base vs base noise floor: 1.004.

base this PR change
stf epoch-start block (mean of 4) 1505 ms 1436 ms −69 ms (−5%)
stf ordinary block p50 577 ms 575 ms noise
get_head p50 200 ms 200 ms
block wall p50 777 ms 771 ms −1%
user CPU (whole run) 214.7 s 214.1 s
paired stf, ratio of totals (round 1 / round 2) 0.998 / 0.988

Per epoch-start block (stf ms):

slot base this PR
15279104 1607 1547
15279136 1521 1446
15279168 1486 1411
15279200 1405 1341

The change is epoch-only, so ordinary blocks are unchanged, as expected. The epoch-start saving falls inside the estimated 50-150 ms. This corpus has only 4 epoch transitions in 128 blocks, so the run-level totals barely move.

All five performance PRs together (#647, #648, #649, #650, #633), measured with a 1 s delay between blocks so the state writer drains (--block-delay, #651):

  • ordinary stf 527 → 163 ms; epoch-start 1440 → 420 ms;
  • block wall p50 729 → 185 ms (3.9x); user CPU −30%.

In that build this cursor also carries #633's fused epoch pass: its inactivity-score, balance and validator writes go through iter_cow/try_update_each.

MegaRedHand added a commit that referenced this pull request Oct 2, 2026
…6-638-gloas-live

Conflicted files:
- CLAUDE.md
- crates/blockchain/state_transition/src/beacon/stf/epoch/altair.rs
- crates/blockchain/state_transition/src/beacon/stf/epoch/electra.rs
- crates/blockchain/state_transition/src/beacon/stf/epoch/mod.rs
- crates/blockchain/state_transition/src/beacon/stf/epoch/rewards.rs
- crates/common/ssz-tree/tests/model.rs

Gloas adaptations (gloas keeps validators and balances in
ProgressiveList, so the cursor and BeaconState::registry_mut as written
cannot serve it):
- ssz-tree: ProgressiveList gains iter_cow / try_update_each, backed by
  ProgressiveIterCow, one IterCow per subtree in turn. IterCow gets a base
  offset so ElemCow::index is the list-wide index, and its next_cow is
  split into advance/take so the chained pass holds no borrow across
  subtrees. A unit test crosses several subtrees against libssz's root.
- types: BeaconState::registry_mut (a struct of &mut Validators and
  &mut Balances) is replaced by two methods that dispatch over both list
  kinds: try_update_balances and try_update_validators_with_balances (the
  latter stops on a short balances list, zip semantics). The public
  RegistryMut struct is dropped; the private enum of the same name stays.

Other adaptations:
- mutators.rs apply_balance_deltas uses the element accessors
  (validator_count, iter_balances) and try_update_balances.
- rewards.rs: take #647's apply_balance_deltas call.
- stf/epoch/mod.rs and electra.rs process_effective_balance_updates:
  cursor pass with #633's checked leaves_hysteresis_band /
  updated_effective_balance; the early UnknownValidator check for a short
  balances list from this branch is kept before the pass.
- altair.rs apply_rewards_and_penalties: decided balance changes are
  applied through one try_update_balances pass instead of balance_mut per
  change; the delta arithmetic order is untouched.
- model.rs: keep both sides' tests (progressive list and write cursor).

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

beacon Ethereum Beacon Chain client

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant