Skip to content

perf(beacon): fuse electra epoch steps 2-9 into one registry pass - #633

Draft
MegaRedHand wants to merge 3 commits into
perf/beacon-epoch-participation-bfrom
perf/beacon-epoch-participation-c
Draft

MegaRedHand wants to merge 3 commits into
perf/beacon-epoch-participation-bfrom
perf/beacon-epoch-participation-c

Conversation

@MegaRedHand

@MegaRedHand MegaRedHand commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

One of three alternative implementations of the same follow-up to lambdaclass/ethlambda_private#37 (epoch participation without per-index registry reads): (a) #630, (b) #632, (c) this PR. Stacked on (b): this PR's diff is the fusion only, and it targets (b)'s branch.

Motivation

After (b), electra and fulu still walk the tree-backed registry once per remaining epoch step: registry updates, slashings, pending deposits and effective-balance updates. Several of those steps also recompute get_total_active_balance, which descends once per active validator. This is option (c) of the plan, the Lighthouse-style single pass (process_epoch_single_pass).

Change: fuse electra/fulu epoch steps 2-9 into one loop

stf::epoch::single_pass visits each validator once, in index order, working on a local copy of the validator and its balance. Per validator, in spec order:

  1. inactivity score;
  2. rewards and penalties;
  3. registry action (queue, eject or activate);
  4. slashing penalty;
  5. planned deposit top-up;
  6. effective-balance update.

Changes are written back after the loop. What depends on more than one validator stays outside it, as in Lighthouse:

  • Exit churn: the cursor is advanced locally in index order and written back once.
  • Pending deposits: the queue is planned before the loop. A deposit counts as "exited" if the validator already has an exit or this epoch's registry update will eject it. Top-ups of existing validators apply inside the loop; deposits that create validators apply after it, and one registry walk serves all pubkey lookups.
  • Consolidations: pending consolidations run after the loop. The effective-balance update of every validator they name, and of every validator created by deposits, waits until then.
  • Total active balance: taken from the summary and reused for slashings and every churn limit. It equals get_total_active_balance through step 8.

The hysteresis test both effective-balance updates use is now one checked leaves_hysteresis_band: a balance near u64::MAX is an ArithmeticOverflow, as in the spec and Lighthouse's safe_add, rather than a debug panic or a release wrap that moves the effective balance. It keeps the spec's short-circuit or.

The genesis epoch, and states whose registry-sized lists differ in length, fall back to the step-by-step path. The step functions stay public and callable in isolation for the fixtures. The per-step rules both paths share were factored out so both run the same code: registry action, slashing penalty, effective-balance hysteresis, inactivity score and the exit-cursor advance.

Testing

  • In debug builds, a registry of up to 4096 validators also runs through the specification-shaped steps 1-9 on a clone, and the two full post-states must have equal hash_tree_root. A mismatch names the differing fields.
  • A randomized fused-vs-unfused test runs 800 electra and fulu cases. They cover ejections that move the churn cursor, activations and the activation queue, slashings, every pending-deposit branch (including new pubkeys with valid and invalid signatures), and processed, slashed and not-yet-withdrawable consolidations. Two deliberately injected bugs were caught within a few seeds. About one state in ten carries a near-u64::MAX balance, since one now fails the whole epoch; those cases check that both paths fail together. As in (b), the finalized epoch stays at or behind the previous epoch.
  • Clippy -D warnings and the ethlambda-state-transition lib tests pass, including the dev-profile Test minimal preset run, where overflow checks are on.
  • Beacon spec filters epoch_processing, rewards, sanity, finality, random, fork_choice and transition pass on both presets. That run predates the rebase onto (b)'s finality-delay fix and the checked hysteresis; CI's beacon spec jobs cover both.

Over (b), this saves about 200 ms per epoch-start block and nothing on other blocks. That is more than the plan estimated, probably because the fusion also removes the repeated total-active-balance calls and the per-deposit pubkey scans (not profiled). A per-epoch total cache (follow-up 01) and a pubkey-to-index map would recover part of that gain more simply.

Benchmark (all three options)

ethlambda benchmark import replay on ethlambda-5: cargo build --release, mainnet corpus of 128 blocks (slots 15279073..=15279200, 4 epoch-start blocks), legs interleaved base a b c for two rounds. Base against base: 0.996. Every leg imported all 128 blocks.

base (c913728) (a) iterator passes (b) summary pass (c) single pass
process wall 104.0 s 96.1 s 92.0 s 91.0 s
user CPU 215.1 s 208.4 s 205.1 s 203.8 s
stf, epoch-start block (mean of 4) 1503 ms 1126 ms 912 ms 709 ms
stf p50, other blocks 577 ms 517 ms 478 ms 479 ms
paired stf vs base (ratio of totals) 1 0.869 0.797 0.785
get_head p50 201 ms 214 ms 226 ms 225 ms
peak RSS 3.88 GiB 3.93 GiB 4.04 GiB 4.04 GiB
  • The per-block saving comes from the pulled-up tip, which reruns justification on every imported block.
  • (c) gains nothing over (b) on ordinary blocks, since their per-block path is the same.
  • get_head rises as stf falls, although no option touches the head computation. This is probably contention with the background state writer, as seen in earlier import benchmarks; it is not verified. Block wall time still falls.

Moved

Electra and fulu walk the tree-backed registry once per epoch step:
inactivity updates, rewards, registry updates, slashings, pending deposits
and effective-balance updates, several of them with a validator(i) descent per
active validator. At mainnet scale that is most of an epoch-start block.

stf::epoch::single_pass visits each validator once, in index order, on a
local copy of it and of its balance, and writes the changes back afterwards.
What depends on more than one validator stays outside the loop: the exit-churn
cursor is advanced locally in index order, the pending-deposit queue is planned
before the loop (top-ups of existing validators inside it, deposits that create
validators after it, one registry walk for all pubkey lookups), and pending
consolidations run after it with the effective-balance update of every
validator they name deferred until they are done. The total active balance is
the summary's, which equals get_total_active_balance through step 8.

The genesis epoch and states whose registry-sized lists differ in length fall
back to the step-by-step path. The step functions stay public and callable in
isolation for the fixtures; the per-step rules the loop shares with them
(registry action, slashing penalty, effective-balance hysteresis, inactivity
score, exit-cursor advance) were factored out so both paths run the same code.

In debug builds a registry of up to 4096 validators is also run through the
specification-shaped steps on a clone and the two full states must hash equal.
A randomized test compares the fused and unfused paths on crafted electra and
fulu states.
balance + DOWNWARD_THRESHOLD and effective_balance + UPWARD_THRESHOLD were
unchecked in both effective-balance updates, so a balance near u64::MAX
panicked in debug builds and, in release, wrapped into a small sum that
moved the effective balance. The specification treats that overflow as an
invalid transition, and Lighthouse's single pass rejects it with safe_add.
Both updates now share one checked leaves_hysteresis_band, which keeps the
specification's short-circuit: the upward sum is only taken when the
downward comparison fails.

The fused-pass test gave nearly every state a near-u64::MAX balance, which
now fails the whole epoch in both paths, so only about one state in ten
carries one. Its finalized epoch also stays behind the previous epoch, as
in any reachable state, now that get_finality_delay fails past it.
@MegaRedHand

Copy link
Copy Markdown
Collaborator Author

Import-replay benchmark (re-measured on the current base)

Method. ethlambda benchmark import replay, release build, 128 consecutive mainnet blocks (slots 15279073..=15279200, 4 of them epoch-start blocks) on a 16-core host. This PR's leg and a base leg (beacon-chain-integration) were interleaved, 2 rounds each; the table shows the mean of the two. Both legs carry #651, so stf excludes the state-writer hand-off wait. Block roots were identical to the base in every leg. Base vs base noise floor: 1.004.

base this PR (includes #632) change
stf epoch-start block (mean of 4) 1505 ms 713 ms −53%
stf ordinary block p50 577 ms 476 ms −18%
get_head p50 200 ms 223 ms +11%
block wall p50 777 ms 696 ms −10%
user CPU (whole run) 214.7 s 203.8 s −5%
peak RSS 3.90 GiB 4.33 GiB +0.43
paired stf, ratio of totals (round 1 / round 2) 0.789 / 0.776

Per epoch-start block (stf ms):

slot base this PR
15279104 1607 835
15279136 1521 727
15279168 1486 683
15279200 1405 609

This reproduces the earlier measurement on the older base (epoch-start 1503 → 709 ms, ordinary 577 → 479 ms), including the small get_head rise. That rise is probably contention with the background state writer's encode, which now overlaps get_head more; not verified.

All five performance PRs together (#647, #648, #649, #650, this one), measured with a 1 s delay between blocks so the state writer drains (--block-delay, #651):

  • ordinary stf 527 → 163 ms; epoch-start 1440 → 420 ms;
  • block wall p50 729 → 185 ms (3.9x); user CPU −30%.

In that build, the fused pass's inactivity-score, balance and validator writes go through #647's in-order write cursor, and the scores are #650's tree list.

MegaRedHand added a commit that referenced this pull request Oct 2, 2026
…6-638-gloas-live

Conflicted files:
- CLAUDE.md
- crates/blockchain/state_transition/src/beacon/stf/epoch/altair.rs
- crates/blockchain/state_transition/src/beacon/stf/epoch/electra.rs
- crates/blockchain/state_transition/src/beacon/stf/epoch/mod.rs
- crates/blockchain/state_transition/src/beacon/stf/epoch/rewards.rs
- crates/common/ssz-tree/tests/model.rs

Gloas adaptations (gloas keeps validators and balances in
ProgressiveList, so the cursor and BeaconState::registry_mut as written
cannot serve it):
- ssz-tree: ProgressiveList gains iter_cow / try_update_each, backed by
  ProgressiveIterCow, one IterCow per subtree in turn. IterCow gets a base
  offset so ElemCow::index is the list-wide index, and its next_cow is
  split into advance/take so the chained pass holds no borrow across
  subtrees. A unit test crosses several subtrees against libssz's root.
- types: BeaconState::registry_mut (a struct of &mut Validators and
  &mut Balances) is replaced by two methods that dispatch over both list
  kinds: try_update_balances and try_update_validators_with_balances (the
  latter stops on a short balances list, zip semantics). The public
  RegistryMut struct is dropped; the private enum of the same name stays.

Other adaptations:
- mutators.rs apply_balance_deltas uses the element accessors
  (validator_count, iter_balances) and try_update_balances.
- rewards.rs: take #647's apply_balance_deltas call.
- stf/epoch/mod.rs and electra.rs process_effective_balance_updates:
  cursor pass with #633's checked leaves_hysteresis_band /
  updated_effective_balance; the early UnknownValidator check for a short
  balances list from this branch is kept before the pass.
- altair.rs apply_rewards_and_penalties: decided balance changes are
  applied through one try_update_balances pass instead of balance_mut per
  change; the delta arithmetic order is untouched.
- model.rs: keep both sides' tests (progressive list and write cursor).
MegaRedHand added a commit that referenced this pull request Oct 2, 2026
…3-64-633-636-638-gloas-live

Conflicted files:
- crates/blockchain/state_transition/src/beacon/stf/epoch/altair.rs
- crates/blockchain/state_transition/src/beacon/stf/mod.rs
- crates/common/ssz-tree/tests/model.rs
- crates/common/types/src/beacon/containers/mod.rs
- crates/common/types/src/beacon/containers/shared.rs

Gloas adaptations:
- tree_fields! gets a gloas line: validators and balances (progressive
  trees) plus block_roots, state_roots, historical_roots, eth1_data_votes,
  randao_mixes, slashings and historical_summaries, which gloas takes from
  the shared aliases. inactivity_scores is left out on purpose: gloas
  declares its own flat libssz ProgressiveList for it, which buffers
  nothing, so there is nothing to flush or rebase.
- ssz-tree: ProgressiveList implements Buffered, so the field list can
  hold the progressive registry beside the bounded fields.
- apply_pending_mutations / has_pending_mutations dispatch over all eight
  beacon forks through buffered()/buffered_mut(); the lean guard on
  has_pending_mutations is kept (the merge had dropped it with the old
  registry match).
- rebase_on: the validators/balances list-kind match stays (so a gloas
  state still rebases onto a gloas base only); the other fields are shared
  types in every fork and rebase unconditionally; inactivity scores
  rebase only for a Tree/Tree pair.
- Inactivity scores are a tree List before gloas and a flat slice-like list
  in gloas, so no &[u64] can serve both. altair_validator_lists now returns
  an InactivityScoresRef view (len, get, in-order iter, ptr_eq,
  has_pending_updates, PartialEq) as its third element, and
  inactivity_scores_mut (a &mut [u64]) becomes inactivity_score_mut(index),
  an element write that dispatches over both kinds.
- stf/epoch/altair.rs update_inactivity_scores: decides over an in-order
  walk of the scores (in step with the summary) and writes only changed
  scores afterwards, which subsumes both #633's "write if different" and
  #650's no-op skip; the error order (first eligible validator without a
  score, then checked-add overflow) is unchanged.
- stf/mod.rs tests: BEACON_FORKS now includes Gloas, the historical
  summaries write covers gloas, and the inactivity score case expects no
  pending write on gloas. process_slot's registry check uses the new
  BeaconState::registry_has_pending_updates, since the per-slot roots
  writes stay buffered by design.
- storage store.rs test and the epoch-processing fixture runner comments
  follow the renamed accessors.

Other: shared.rs imports both List/Vector and the progressive alias; model.rs
keeps both sides' tests (the base side was empty).

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

beacon Ethereum Beacon Chain client

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant