perf(beacon): cache the total active balance per epoch - #648
MegaRedHand wants to merge 3 commits into
Conversation
get_total_active_balance built an index Vec of the active validators and then read each one back with a validator(i) tree descent. Sum the effective balances in a single in-order walk over state.validators().iter() instead, with the same saturating add and one-increment floor as get_total_balance. Every caller (block, epoch, fork choice, churn) gets the speedup without a signature change.
An ordinary block asks for the total active balance about nine times (every attestation through get_base_reward_per_increment, twice in the sync aggregate), and each ask is a registry pass. Its value is fixed by the epoch and the block root at the last slot of the previous epoch: effective balances are written only by that epoch's effective-balance step, and every activation or exit lands at least MAX_SEED_LOOKAHEAD epochs ahead. Keep it in an ActiveBalanceCache held by the Store next to the committee cache, keyed by that (epoch, decision root), so sibling branches never share an entry they should not. The cache is external to BeaconState: no new state field, no SSZ change. The state-aware half (key derivation, one-pass computation, a debug cross-check on every hit) lives in helpers::accessors as ActiveBalanceCacheExt, mirroring CommitteeCacheExt. Block processing takes the cache next to the committee cache; on_block uses the store's. Spec runners, block production and unit tests pass a fresh default. Epoch processing, churn helpers and fork choice keep the uncached one-pass function. Adds lean_beacon_total_active_balance_lookups_total.
The three process_attestation variants (altair, deneb, electra) read the attester's effective balance with a validator(i) tree descent inside the per-flag loop, up to once per newly earned flag. Read it lazily once per attester instead: the first newly earned flag fetches it, the others reuse it. Staying lazy keeps the error behavior identical, since a validator that earns no flag is never read. The effective balance cannot change inside the read-only phase, so the arithmetic is unchanged.
Import-replay benchmarkMethod.
The ordinary-block saving (224 ms) sits just under the estimated 300-414 ms. Peak RSS rises although the cache holds a few All five performance PRs together (#647, #648, #649, #650, #633), measured with a 1 s delay between blocks so the state writer drains (
Without the delay, the combined build outpaces the state writer and waits ~326 ms per block on it. |
…63-64-633-636-638-gloas-live Conflicted files: - crates/blockchain/state_transition/src/beacon/stf/block.rs - crates/blockchain/state_transition/src/beacon/stf/electra.rs - crates/blockchain/state_transition/src/beacon/stf/fulu.rs - crates/blockchain/state_transition/src/beacon/stf/operations.rs - crates/blockchain/state_transition/tests/beacon_spec/operations.rs - crates/storage/src/store.rs - docs/beacon_stf.md Gloas adaptations: - stf/gloas.rs process_block, process_operations and process_attestation take the ActiveBalanceCache; the gloas block's sync aggregate goes through the shared altair process_sync_aggregate with it. The base reward per increment comes from the cache, and the attester's effective balance is read once per attester (it also feeds the pending builder payment weight), as in the other forks. - Soundness check for gloas: effective_balance has one writer outside the fork upgrades (queue_entire_balance_and_reset_validator, called only by the electra upgrade) and genesis, besides the epoch-end process_effective_balance_updates. Gloas block processing (parent payload application, withdrawals, bids, builder payments and deposits, execution requests, slashings, exits) writes balances, exit epochs, slashed flags and queues, never an effective balance, and exits and activations land at least MAX_SEED_LOOKAHEAD epochs ahead, so the cache is used for gloas too. A full or empty parent changes none of those inputs. The debug cross-check on every hit runs in the gloas sanity and transition fixtures. - The gloas arm of block::process_block keeps its engine-free call and gains the cache argument. - fork_choice.rs and the fixture runners (operations gloas arm) pass the cache; get_total_active_balance walks iter_validators(). Other: store.rs keeps both the liveness tracker and the cache field and accessor; fulu.rs keeps its own process_operations and pending-queue import with the cache argument; docs follow the element accessor names and no longer mention registry_mut.
…636-638-gloas-live
Conflicted files:
- crates/blockchain/src/lib.rs
- crates/blockchain/state_transition/src/beacon/fork_choice.rs
- crates/blockchain/state_transition/src/metrics.rs
- docs/metrics.md
Gloas: the precompute stays enabled, because the base state cannot differ.
In this branch a gloas block's post-state never holds its own payload: the
child applies it (process_parent_execution_payload) inside process_block,
after process_slots, and on_execution_payload_envelope only records and
checks, it stores no second state. The store keeps one state per root and
both on_block (get_state(parent_root)) and checkpoint_state read it, and the
worker clones get_state(head_root) and caches the result under
CheckpointState { E+1, head_root }, so the entry is process_slots of the very
state an import of a child of that root clones. The fulu-to-gloas boundary
needs no case either: the upgrade runs inside process_slots on both paths.
A new unit test (a_gloas_precompute_starts_from_the_state_an_import_would)
checks, for a gloas parent, that the worker's result equals
checkpoint_state's for the key, that an import finds it, and that advancing
the rest of a skipped first slot from it equals plain process_slots.
Other adaptations:
- transition_block and stf::apply_block take the ActiveBalanceCache that
#648 added to state_transition; on_block passes the store's own cache.
- on_block's gloas-specific code between the parent read and the
transition does not use the cloned state, so removing that early clone
(as #652 does) needed nothing else.
- lib.rs: the NewBlock handler keeps the envelope settling that follows the
import and triggers the precompute after it, so a head moved by a settled
envelope's block is seen.
- metrics.rs and docs/metrics.md: both sides' metrics kept.
- Test fixtures of #652 follow this branch's BlockInputs (operations) and
the extra argument.
- docs/beacon_stf.md notes why gloas shares the key.
Motivation
get_total_active_balancerescans the registry about ten times per ordinary block: once per attestation (throughget_base_reward_per_increment), twice in the sync aggregate, plus churn and fork-choice callers. Each call built an indexVecof the active validators and then read every index back with avalidator(i)tree descent. In the profile of lambdaclass/ethlambda_private#37 this was the largest cost left after hashing (the plan's estimate: about 414 ms per block at mainnet scale; an estimate, not measured here).Changes
perf(beacon): sum the total active balance in one registry passget_total_active_balanceis one in-ordervalidators().iter()pass with the samesaturating_addand one-increment floor asget_total_balance. No indexVec, no descents. Every caller gets it, cached or not. Randomized test (inline SplitMix64) againstget_total_balance(get_active_validator_indices(..)), plus an all-zero registry.perf(beacon): cache the total active balance per epoch in the storeActiveBalanceCache(ethlambda-storage), held by theStorebeside the committee cache (Store::active_balance_cache()). Key derivation and lookup inActiveBalanceCacheExt(helpers/accessors.rs). Attestation processing (altair, deneb, electra) andprocess_sync_aggregatetake it; plumbed up through block processing andstate_transition. New metriclean_beacon_total_active_balance_lookups_total{result}. Docs inbeacon_stf.mdandmetrics.md.perf(beacon): read each attester's validator once per attestationDesign decisions
CommitteeCacheover a non-SSZ cache field onBeaconState. That keeps the state type, its SSZ derive and its fork upgrades untouched: no libssz change, norebase_onadoption, no reset at upgrades.process_effective_balance_updates; the Electra upgrade zeroes only never-activated validators, which are not active at E. The active set at E cannot change inside E, because every activation and exit lands at leastMAX_SEED_LOOKAHEADepochs ahead; deposits append validators that are not yet active; slashing setsslashedbut leaves effective balances alone (the spec's total keeps slashed validators). So(E, block root at the last slot of E-1)identifies the history: two states agreeing on it agree on the total. Sibling branches that diverge during E-1 have different roots, so they get different entries.ActiveBalanceKey) keeps the two apart.unkeyable: computed, not cached.process_effective_balance_updatesand the slot increment a state is in epoch E-1 but already carries E's balances, so a key rooted at the end of E-2 would be wrong there. Block processing never runs in that window (process_slotshas already advanced the state), and epoch processing keeps the uncached one-pass function. Documented onactive_balance_key.debug_assert_eq!, so a missed writer fails the fixtures (release-fastkeeps debug assertions on).on_blockreads the store's cache itself (store.active_balance_cache()), so its signature and the chain actor are unchanged. Other callers (spec runners, block production, unit tests) passActiveBalanceCache::default().Not in this PR
get_proposer_scoreandcalculate_committee_fractionkeep calling the uncached function; another change replaces that total with a justified-balances snapshot.Testing
Measured on release-fast with debug assertions. The import-replay benchmark was not run for this PR, so there is no measured speedup to report.
cargo clippy --locked --workspace --all-targets --profile release-fast -- -D warningscargo clippy -p ethlambda-state-transition ... --features beacon-spec-tests -- -D warningscargo test -p ethlambda-storage --libcargo test -p ethlambda-state-transition --libcargo check -p ethlambda --profile release-fastThe 4 failures per preset are the gossip
matched_fixture_caseschecks: the gossip fixture tree is empty in this environment, so they match no cases. Every other suite passes.