perf(beacon): write epoch balances through an in-order copy-on-write cursor - #647
MegaRedHand wants to merge 2 commits into
Conversation
Rewriting most of a list through get_mut buffers an entry per element, sorts the buffer into a second Vec, and then rebuilds every touched leaf, recomputing the root of every written element even when the value did not change. iter_cow / try_update_each walk the leaves in order instead. A leaf is copied on the first write that reaches it, compared with the original at the leaf's end (only the offsets make_mut handed out), and kept as it was (Arc and cached hash) if nothing differs. Changed leaves carry over the element roots of unchanged elements and compute the changed ones. The changed leaves are swapped in on drop with one descent, so an early return or a panic leaves a consistent list. Pending writes are applied before the pass starts. The with_updated_leaves descent is generalized over a leaf source so the element-update path and the whole-leaf replacement share it. Covered by unit tests (changed-path sharing, carried roots, early return, catch_unwind, registry depth) and model-based property tests against SszList.
…h the cursor Both rewards-and-penalties passes wrote every balance through get_mut, up to eight calls per validator on altair and later, each a map entry in the dense pending-write buffer. They now apply every (rewards, penalties) pair per index in one cursor pass, keeping the spec's per-pair saturating_add then saturating_sub order (no netting). Both process_effective_balance_updates variants read balances()[index] per validator, which was a buffer hit only because the rewards pass had just filled the buffer. With the rewards write bypassing the buffer those reads would become tree descents, so they move in the same change: the registry's cursor runs zipped with a balances iterator, through the new BeaconState::registry_mut disjoint borrow, and copies a leaf only for a validator whose effective balance actually changes. Adds the tree_bench rows comparing the old and new write paths, and the docs for the cursor.
Import-replay benchmarkMethod.
Per epoch-start block (
The change is epoch-only, so ordinary blocks are unchanged, as expected. The epoch-start saving falls inside the estimated 50-150 ms. This corpus has only 4 epoch transitions in 128 blocks, so the run-level totals barely move. All five performance PRs together (#647, #648, #649, #650, #633), measured with a 1 s delay between blocks so the state writer drains (
In that build this cursor also carries #633's fused epoch pass: its inactivity-score, balance and validator writes go through |
…6-638-gloas-live Conflicted files: - CLAUDE.md - crates/blockchain/state_transition/src/beacon/stf/epoch/altair.rs - crates/blockchain/state_transition/src/beacon/stf/epoch/electra.rs - crates/blockchain/state_transition/src/beacon/stf/epoch/mod.rs - crates/blockchain/state_transition/src/beacon/stf/epoch/rewards.rs - crates/common/ssz-tree/tests/model.rs Gloas adaptations (gloas keeps validators and balances in ProgressiveList, so the cursor and BeaconState::registry_mut as written cannot serve it): - ssz-tree: ProgressiveList gains iter_cow / try_update_each, backed by ProgressiveIterCow, one IterCow per subtree in turn. IterCow gets a base offset so ElemCow::index is the list-wide index, and its next_cow is split into advance/take so the chained pass holds no borrow across subtrees. A unit test crosses several subtrees against libssz's root. - types: BeaconState::registry_mut (a struct of &mut Validators and &mut Balances) is replaced by two methods that dispatch over both list kinds: try_update_balances and try_update_validators_with_balances (the latter stops on a short balances list, zip semantics). The public RegistryMut struct is dropped; the private enum of the same name stays. Other adaptations: - mutators.rs apply_balance_deltas uses the element accessors (validator_count, iter_balances) and try_update_balances. - rewards.rs: take #647's apply_balance_deltas call. - stf/epoch/mod.rs and electra.rs process_effective_balance_updates: cursor pass with #633's checked leaves_hysteresis_band / updated_effective_balance; the early UnknownValidator check for a short balances list from this branch is kept before the pass. - altair.rs apply_rewards_and_penalties: decided balance changes are applied through one try_update_balances pass instead of balance_mut per change; the delta arithmetic order is untouched. - model.rs: keep both sides' tests (progressive list and write cursor).
Motivation
An epoch boundary writes every balance.
get_muton a tree list buffers a copy of each element in a dense map, sorts the map into a secondVec, then rebuilds every touched leaf and recomputes the root of every written element, changed or not. The altair rewards step calls it up to eight times per validator (four reward/penalty pairs, each as increase then decrease), andprocess_effective_balance_updatesreadsbalances()[i]per validator, which was only cheap because the rewards pass had just filled the map.This PR adds an in-order lazy copy-on-write write cursor to
ethlambda-ssz-treeand moves the epoch balance writes onto it. Expected gain (estimate, not measured on a real chain): tens of ms off each epoch-start block at mainnet registry size. The import-replay benchmark was not run for this PR.Changes
perf(ssz-tree): add an in-order copy-on-write write cursorList/Vector::iter_cow,IterCow::next_cow,ElemCow::{index, make_mut, set},try_update_each;with_updated_leavesgeneralized over a leaf source; unit tests and model-based property testsperf(beacon): apply epoch balance and effective-balance writes through the cursorBeaconState::registry_mut(disjoint borrows of validators and balances);apply_balance_deltasused by the altair+ and phase0process_rewards_and_penalties; bothprocess_effective_balance_updates; newtree_benchrows; docsHow the cursor works: it walks leaves in order and copies a leaf on the first write that reaches it. At the leaf's end it compares only the touched offsets with the original; if all are equal, the original leaf (its
Arcand cached hash) is kept. Otherwise the new leaf carries over the element roots of unchanged elements and computes the changed ones. On drop, one descent swaps the changed leaves in and rebuilds only their paths. Pending buffered writes are applied before the pass starts.Design decisions
Dropdoes the work, and the tree is untouched during the pass, so an early return or a panic leaves a consistent list holding the writes so far (tested withcatch_unwind).iter_cow,next_cow,make_mut,set,try_update_each. Noset_with.Arcor carried root always follows an actual comparison. The cursor only creates fresh empty hash cells, or carried roots for a leaf it owns; it never fills a sharedOnceLock.saturating_add(reward)thensaturating_sub(penalty), pairs in order. Deltas are never netted. Each balance depends only on its own deltas, so index-by-index equals the spec's pair-by-pair loop.apply_balance_deltaslives inhelpers/mutators.rsand is shared by the phase0 and altair+ paths. It keeps the old error for a balances list shorter than the registry, and does not touch balances past the registry length.Not in this PR
Balancesto a sparseBTreeMapupdate map: to be decided after measuring whether a dense writer remains.List::filled: its only consumer is the participation-list work, which is deferred.process_rewards_and_penalties.Testing
cargo test -p ethlambda-ssz-tree --profile release-fastcargo test -p ethlambda-types --profile release-fast --libcargo test -p ethlambda-state-transition --profile release-fast --libcargo clippy --locked --workspace --all-targets --profile release-fast -- -D warningsand the state-transitionbeacon-spec-testsvariantcargo check -p ethlambda --profile release-fastThe 4 spec failures in each preset are the
gossip/*suites (gossip_beacon_block,gossip_data_column_sidecar,gossip_beacon_aggregate_and_proof,gossip_beacon_attestation): the local gossip fixture directory is intentionally empty, so they match no cases. Every other suite, includingepoch_processing,rewards,sanity,finality,randomandtransition, passes.tree_bench (2.4M validators,
--ignored, single run)Noisy: run while the machine was under a load average above 100 from other builds, so absolute numbers (and even the order of the old rows) are unreliable; read them as rough ratios only.
epoch_sweep_tree(old path)epoch_sweep_tree_phasesepoch_sweep_cursor_treeepoch_sweep_noop_cursor_treeepoch_sweep_rebuild_tree(rebuild baseline)rewards_apply_4_pairs_get_mutrewards_apply_4_pairs_cursoreb_updates_indexedeb_updates_cursorlockstep_three_lists_cursorThe bench asserts the cursor rows produce the same roots as the old paths, and that a no-op sweep keeps the tree (
ptr_eq).