perf(beacon): fuse electra epoch steps 2-9 into one registry pass - #633
MegaRedHand wants to merge 3 commits into
Conversation
Electra and fulu walk the tree-backed registry once per epoch step: inactivity updates, rewards, registry updates, slashings, pending deposits and effective-balance updates, several of them with a validator(i) descent per active validator. At mainnet scale that is most of an epoch-start block. stf::epoch::single_pass visits each validator once, in index order, on a local copy of it and of its balance, and writes the changes back afterwards. What depends on more than one validator stays outside the loop: the exit-churn cursor is advanced locally in index order, the pending-deposit queue is planned before the loop (top-ups of existing validators inside it, deposits that create validators after it, one registry walk for all pubkey lookups), and pending consolidations run after it with the effective-balance update of every validator they name deferred until they are done. The total active balance is the summary's, which equals get_total_active_balance through step 8. The genesis epoch and states whose registry-sized lists differ in length fall back to the step-by-step path. The step functions stay public and callable in isolation for the fixtures; the per-step rules the loop shares with them (registry action, slashing penalty, effective-balance hysteresis, inactivity score, exit-cursor advance) were factored out so both paths run the same code. In debug builds a registry of up to 4096 validators is also run through the specification-shaped steps on a clone and the two full states must hash equal. A randomized test compares the fused and unfused paths on crafted electra and fulu states.
balance + DOWNWARD_THRESHOLD and effective_balance + UPWARD_THRESHOLD were unchecked in both effective-balance updates, so a balance near u64::MAX panicked in debug builds and, in release, wrapped into a small sum that moved the effective balance. The specification treats that overflow as an invalid transition, and Lighthouse's single pass rejects it with safe_add. Both updates now share one checked leaves_hysteresis_band, which keeps the specification's short-circuit: the upward sum is only taken when the downward comparison fails. The fused-pass test gave nearly every state a near-u64::MAX balance, which now fails the whole epoch in both paths, so only about one state in ten carries one. Its finalized epoch also stays behind the previous epoch, as in any reachable state, now that get_finality_delay fails past it.
Import-replay benchmark (re-measured on the current base)Method.
Per epoch-start block (
This reproduces the earlier measurement on the older base (epoch-start 1503 → 709 ms, ordinary 577 → 479 ms), including the small All five performance PRs together (#647, #648, #649, #650, this one), measured with a 1 s delay between blocks so the state writer drains (
In that build, the fused pass's inactivity-score, balance and validator writes go through #647's in-order write cursor, and the scores are #650's tree list. |
…6-638-gloas-live Conflicted files: - CLAUDE.md - crates/blockchain/state_transition/src/beacon/stf/epoch/altair.rs - crates/blockchain/state_transition/src/beacon/stf/epoch/electra.rs - crates/blockchain/state_transition/src/beacon/stf/epoch/mod.rs - crates/blockchain/state_transition/src/beacon/stf/epoch/rewards.rs - crates/common/ssz-tree/tests/model.rs Gloas adaptations (gloas keeps validators and balances in ProgressiveList, so the cursor and BeaconState::registry_mut as written cannot serve it): - ssz-tree: ProgressiveList gains iter_cow / try_update_each, backed by ProgressiveIterCow, one IterCow per subtree in turn. IterCow gets a base offset so ElemCow::index is the list-wide index, and its next_cow is split into advance/take so the chained pass holds no borrow across subtrees. A unit test crosses several subtrees against libssz's root. - types: BeaconState::registry_mut (a struct of &mut Validators and &mut Balances) is replaced by two methods that dispatch over both list kinds: try_update_balances and try_update_validators_with_balances (the latter stops on a short balances list, zip semantics). The public RegistryMut struct is dropped; the private enum of the same name stays. Other adaptations: - mutators.rs apply_balance_deltas uses the element accessors (validator_count, iter_balances) and try_update_balances. - rewards.rs: take #647's apply_balance_deltas call. - stf/epoch/mod.rs and electra.rs process_effective_balance_updates: cursor pass with #633's checked leaves_hysteresis_band / updated_effective_balance; the early UnknownValidator check for a short balances list from this branch is kept before the pass. - altair.rs apply_rewards_and_penalties: decided balance changes are applied through one try_update_balances pass instead of balance_mut per change; the delta arithmetic order is untouched. - model.rs: keep both sides' tests (progressive list and write cursor).
…3-64-633-636-638-gloas-live Conflicted files: - crates/blockchain/state_transition/src/beacon/stf/epoch/altair.rs - crates/blockchain/state_transition/src/beacon/stf/mod.rs - crates/common/ssz-tree/tests/model.rs - crates/common/types/src/beacon/containers/mod.rs - crates/common/types/src/beacon/containers/shared.rs Gloas adaptations: - tree_fields! gets a gloas line: validators and balances (progressive trees) plus block_roots, state_roots, historical_roots, eth1_data_votes, randao_mixes, slashings and historical_summaries, which gloas takes from the shared aliases. inactivity_scores is left out on purpose: gloas declares its own flat libssz ProgressiveList for it, which buffers nothing, so there is nothing to flush or rebase. - ssz-tree: ProgressiveList implements Buffered, so the field list can hold the progressive registry beside the bounded fields. - apply_pending_mutations / has_pending_mutations dispatch over all eight beacon forks through buffered()/buffered_mut(); the lean guard on has_pending_mutations is kept (the merge had dropped it with the old registry match). - rebase_on: the validators/balances list-kind match stays (so a gloas state still rebases onto a gloas base only); the other fields are shared types in every fork and rebase unconditionally; inactivity scores rebase only for a Tree/Tree pair. - Inactivity scores are a tree List before gloas and a flat slice-like list in gloas, so no &[u64] can serve both. altair_validator_lists now returns an InactivityScoresRef view (len, get, in-order iter, ptr_eq, has_pending_updates, PartialEq) as its third element, and inactivity_scores_mut (a &mut [u64]) becomes inactivity_score_mut(index), an element write that dispatches over both kinds. - stf/epoch/altair.rs update_inactivity_scores: decides over an in-order walk of the scores (in step with the summary) and writes only changed scores afterwards, which subsumes both #633's "write if different" and #650's no-op skip; the error order (first eligible validator without a score, then checked-add overflow) is unchanged. - stf/mod.rs tests: BEACON_FORKS now includes Gloas, the historical summaries write covers gloas, and the inactivity score case expects no pending write on gloas. process_slot's registry check uses the new BeaconState::registry_has_pending_updates, since the per-slot roots writes stay buffered by design. - storage store.rs test and the epoch-processing fixture runner comments follow the renamed accessors. Other: shared.rs imports both List/Vector and the progressive alias; model.rs keeps both sides' tests (the base side was empty).
One of three alternative implementations of the same follow-up to lambdaclass/ethlambda_private#37 (epoch participation without per-index registry reads): (a) #630, (b) #632, (c) this PR. Stacked on (b): this PR's diff is the fusion only, and it targets (b)'s branch.
Motivation
After (b), electra and fulu still walk the tree-backed registry once per remaining epoch step: registry updates, slashings, pending deposits and effective-balance updates. Several of those steps also recompute
get_total_active_balance, which descends once per active validator. This is option (c) of the plan, the Lighthouse-style single pass (process_epoch_single_pass).Change: fuse electra/fulu epoch steps 2-9 into one loop
stf::epoch::single_passvisits each validator once, in index order, working on a local copy of the validator and its balance. Per validator, in spec order:Changes are written back after the loop. What depends on more than one validator stays outside it, as in Lighthouse:
get_total_active_balancethrough step 8.The hysteresis test both effective-balance updates use is now one checked
leaves_hysteresis_band: a balance nearu64::MAXis anArithmeticOverflow, as in the spec and Lighthouse'ssafe_add, rather than a debug panic or a release wrap that moves the effective balance. It keeps the spec's short-circuitor.The genesis epoch, and states whose registry-sized lists differ in length, fall back to the step-by-step path. The step functions stay public and callable in isolation for the fixtures. The per-step rules both paths share were factored out so both run the same code: registry action, slashing penalty, effective-balance hysteresis, inactivity score and the exit-cursor advance.
Testing
hash_tree_root. A mismatch names the differing fields.u64::MAXbalance, since one now fails the whole epoch; those cases check that both paths fail together. As in (b), the finalized epoch stays at or behind the previous epoch.-D warningsand theethlambda-state-transitionlib tests pass, including the dev-profileTest minimal presetrun, where overflow checks are on.epoch_processing,rewards,sanity,finality,random,fork_choiceandtransitionpass on both presets. That run predates the rebase onto (b)'s finality-delay fix and the checked hysteresis; CI's beacon spec jobs cover both.Over (b), this saves about 200 ms per epoch-start block and nothing on other blocks. That is more than the plan estimated, probably because the fusion also removes the repeated total-active-balance calls and the per-deposit pubkey scans (not profiled). A per-epoch total cache (follow-up 01) and a pubkey-to-index map would recover part of that gain more simply.
Benchmark (all three options)
ethlambda benchmark import replayon ethlambda-5:cargo build --release, mainnet corpus of 128 blocks (slots 15279073..=15279200, 4 epoch-start blocks), legs interleavedbase a b cfor two rounds. Base against base: 0.996. Every leg imported all 128 blocks.stf, epoch-start block (mean of 4)stfp50, other blocksstfvs base (ratio of totals)get_headp50get_headrises asstffalls, although no option touches the head computation. This is probably contention with the background state writer, as seen in earlier import benchmarks; it is not verified. Block wall time still falls.Moved
beacon-chain-integrationlives on this repo.beacon-chain-integration@c79fabd5, is merged into the branch.