Skip to content

perf(beacon): walk the registry once in epoch steps 1-3 instead of per-index reads - #630

Draft
MegaRedHand wants to merge 2 commits into
beacon-chain-integrationfrom
perf/beacon-epoch-participation-a
Draft

MegaRedHand wants to merge 2 commits into
beacon-chain-integrationfrom
perf/beacon-epoch-participation-a

Conversation

@MegaRedHand

@MegaRedHand MegaRedHand commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

One of three alternative implementations of the same follow-up to lambdaclass/ethlambda_private#37 (epoch participation without per-index registry reads): (a) this PR, (b) #632, (c) #633. They are meant to be compared, and at most one of (a) and (b) should merge; (c) is stacked on (b).

Motivation

Since lambdaclass/ethlambda_private#37 the registry is a persistent Merkle tree, so state.validator(i) is a tree descent rather than an array index, while validators().iter() walks leaf slices.

Epoch steps 1-3 (justification, inactivity updates, rewards) built index Vecs, then did one descent per participant and one binary search per eligible validator. The pulled-up-tip check in fork choice repeats justification on every imported block, so part of that cost is paid per block, not per epoch.

Change: option (a), iterator passes, spec-shaped

Every spec function keeps its shape and signature. Each now walks validators().iter().enumerate() and indexes the flat participation and inactivity-score lists by position:

  • get_unslashed_participating_indices is one pass, no descents.
  • New get_unslashed_participating_balance, plus crate-internal one-pass versions of the total active balance and the base reward per increment. The public get_total_active_balance and get_base_reward_per_increment are untouched.
  • Justification uses those three passes and builds no index Vec.
  • get_flag_index_deltas and get_inactivity_penalty_deltas decide eligibility and participation inline, with the leak flag computed once.
  • process_inactivity_updates makes one immutable pass, then one pass over the scores.
  • process_rewards_and_penalties is unchanged: it still builds eight delta Vecs and does 8N get_mut. That is the known remaining cost of this option.

Testing

  • The previous implementations are kept as a test-only reference.
  • Randomized tests (altair and electra, 100-300 validators) compare outputs, errors and post-step state against the reference. They cover short lists, off-increment and zero effective balances, scores near u64::MAX, epochs 0, 1 and 2, and leaking epochs.
  • Clippy -D warnings and the ethlambda-state-transition lib tests pass.
  • Beacon spec filters epoch_processing, rewards, sanity, finality, random, fork_choice and transition pass on both presets.

Benchmark (all three options)

ethlambda benchmark import replay on ethlambda-5: cargo build --release, mainnet corpus of 128 blocks (slots 15279073..=15279200, 4 epoch-start blocks), legs interleaved base a b c for two rounds. Base against base: 0.996. Every leg imported all 128 blocks.

base (c913728) (a) iterator passes (b) summary pass (c) single pass
process wall 104.0 s 96.1 s 92.0 s 91.0 s
user CPU 215.1 s 208.4 s 205.1 s 203.8 s
stf, epoch-start block (mean of 4) 1503 ms 1126 ms 912 ms 709 ms
stf p50, other blocks 577 ms 517 ms 478 ms 479 ms
paired stf vs base (ratio of totals) 1 0.869 0.797 0.785
get_head p50 201 ms 214 ms 226 ms 225 ms
peak RSS 3.88 GiB 3.93 GiB 4.04 GiB 4.04 GiB
  • The per-block saving comes from the pulled-up tip, which reruns justification on every imported block.
  • (c) gains nothing over (b) on ordinary blocks, since their per-block path is the same.
  • get_head rises as stf falls, although no option touches the head computation. This is probably contention with the background state writer, as seen in earlier import benchmarks; it is not verified. Block wall time still falls.

Moved

  • Moved from lambdaclass/ethlambda_private#60, now that beacon-chain-integration lives on this repo.
  • Based on beacon-chain-integration @ c79fabd5, merged into the branch.

…r-index reads

Justification, inactivity updates and rewards built index Vecs, then did a
tree descent (state.validator(i)) per participant and a binary search per
eligible validator, and the pulled-up-tip check repeats justification on
every imported block. Each of those reads is a tree walk; the registry's
iter() is a leaf-slice walk.

Every spec function keeps its shape and signature, but now iterates
validators().iter().enumerate() and indexes the flat participation and
score lists by position. New helpers compute the unslashed participating
balance and the total active balance (plus base reward per increment) in
one pass without an index Vec; the public get_total_active_balance and
get_base_reward_per_increment are untouched.

The previous implementations are kept as a test-only reference, and
randomized tests (altair and electra, short lists, off-grid balances,
scores near the top of the range, leaking epochs) compare outputs, errors
and post-step state against it.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

beacon Ethereum Beacon Chain client

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant