Skip to content

perf(beacon): cache the total active balance per epoch - #648

Draft
MegaRedHand wants to merge 3 commits into
beacon-chain-integrationfrom
perf/beacon-total-active-balance-cache
Draft

MegaRedHand wants to merge 3 commits into
beacon-chain-integrationfrom
perf/beacon-total-active-balance-cache

Conversation

@MegaRedHand

Copy link
Copy Markdown
Collaborator

Motivation

get_total_active_balance rescans the registry about ten times per ordinary block: once per attestation (through get_base_reward_per_increment), twice in the sync aggregate, plus churn and fork-choice callers. Each call built an index Vec of the active validators and then read every index back with a validator(i) tree descent. In the profile of lambdaclass/ethlambda_private#37 this was the largest cost left after hashing (the plan's estimate: about 414 ms per block at mainnet scale; an estimate, not measured here).

Changes

Commit Change
perf(beacon): sum the total active balance in one registry pass get_total_active_balance is one in-order validators().iter() pass with the same saturating_add and one-increment floor as get_total_balance. No index Vec, no descents. Every caller gets it, cached or not. Randomized test (inline SplitMix64) against get_total_balance(get_active_validator_indices(..)), plus an all-zero registry.
perf(beacon): cache the total active balance per epoch in the store New ActiveBalanceCache (ethlambda-storage), held by the Store beside the committee cache (Store::active_balance_cache()). Key derivation and lookup in ActiveBalanceCacheExt (helpers/accessors.rs). Attestation processing (altair, deneb, electra) and process_sync_aggregate take it; plumbed up through block processing and state_transition. New metric lean_beacon_total_active_balance_lookups_total{result}. Docs in beacon_stf.md and metrics.md.
perf(beacon): read each attester's validator once per attestation The effective balance is read once per attester, lazily on the first newly earned flag, instead of once per flag.

Design decisions

  • External cache, not a state field. The design review chose a store-held table like CommitteeCache over a non-SSZ cache field on BeaconState. That keeps the state type, its SSZ derive and its fork upgrades untouched: no libssz change, no rebase_on adoption, no reset at upgrades.
  • Why the key is sound. The total for a state in epoch E is fixed by two inputs. Effective balances are written only by epoch E-1's process_effective_balance_updates; the Electra upgrade zeroes only never-activated validators, which are not active at E. The active set at E cannot change inside E, because every activation and exit lands at least MAX_SEED_LOOKAHEAD epochs ahead; deposits append validators that are not yet active; slashing sets slashed but leaves effective balances alone (the spec's total keeps slashed validators). So (E, block root at the last slot of E-1) identifies the history: two states agreeing on it agree on the total. Sibling branches that diverge during E-1 have different roots, so they get different entries.
  • Decision root differs from the shuffling key. The shuffling key uses the last slot of E-2, since the seed is fixed ahead of use; the total is not, so it needs the later root. A new key type (ActiveBalanceKey) keeps the two apart.
  • Epoch 0 keys on the genesis block root (slot 0), like the shuffling key does for its first epochs. A state that cannot name the root (the genesis state itself, or a root outside the history window) is unkeyable: computed, not cached.
  • Only block processing uses the cache. Between process_effective_balance_updates and the slot increment a state is in epoch E-1 but already carries E's balances, so a key rooted at the end of E-2 would be wrong there. Block processing never runs in that window (process_slots has already advanced the state), and epoch processing keeps the uncached one-pass function. Documented on active_balance_key.
  • Debug cross-check. Every hit is compared against a fresh computation under debug_assert_eq!, so a missed writer fails the fixtures (release-fast keeps debug assertions on).
  • No single-flight. A concurrent double compute on a miss is accepted: the value is one cheap registry pass and both callers compute the same number.
  • on_block reads the store's cache itself (store.active_balance_cache()), so its signature and the chain actor are unchanged. Other callers (spec runners, block production, unit tests) pass ActiveBalanceCache::default().
  • Capacity is a small constant with eviction of the lowest epoch; entries are tiny, so room for several forks is cheap.
  • Per-attester read stays lazy so that an attester earning no flag is still never read, which keeps error behavior identical.

Not in this PR

Testing

Measured on release-fast with debug assertions. The import-replay benchmark was not run for this PR, so there is no measured speedup to report.

Command Result
cargo clippy --locked --workspace --all-targets --profile release-fast -- -D warnings clean
cargo clippy -p ethlambda-state-transition ... --features beacon-spec-tests -- -D warnings clean
cargo test -p ethlambda-storage --lib 142 passed, 1 ignored (run before the store-sharing test was added; not re-run after)
cargo test -p ethlambda-state-transition --lib 330 passed, 16 ignored
beacon spec suite, mainnet (full) 5734 passed, 152 ignored, 4 failed
beacon spec suite, minimal (full) 40038 passed, 3692 ignored, 4 failed
cargo check -p ethlambda --profile release-fast clean

The 4 failures per preset are the gossip matched_fixture_cases checks: the gossip fixture tree is empty in this environment, so they match no cases. Every other suite passes.

get_total_active_balance built an index Vec of the active validators and then
read each one back with a validator(i) tree descent. Sum the effective
balances in a single in-order walk over state.validators().iter() instead,
with the same saturating add and one-increment floor as get_total_balance.
Every caller (block, epoch, fork choice, churn) gets the speedup without a
signature change.
An ordinary block asks for the total active balance about nine times (every
attestation through get_base_reward_per_increment, twice in the sync
aggregate), and each ask is a registry pass. Its value is fixed by the epoch
and the block root at the last slot of the previous epoch: effective balances
are written only by that epoch's effective-balance step, and every activation
or exit lands at least MAX_SEED_LOOKAHEAD epochs ahead.

Keep it in an ActiveBalanceCache held by the Store next to the committee
cache, keyed by that (epoch, decision root), so sibling branches never share
an entry they should not. The cache is external to BeaconState: no new state
field, no SSZ change. The state-aware half (key derivation, one-pass
computation, a debug cross-check on every hit) lives in helpers::accessors as
ActiveBalanceCacheExt, mirroring CommitteeCacheExt.

Block processing takes the cache next to the committee cache; on_block uses
the store's. Spec runners, block production and unit tests pass a fresh
default. Epoch processing, churn helpers and fork choice keep the uncached
one-pass function. Adds lean_beacon_total_active_balance_lookups_total.
The three process_attestation variants (altair, deneb, electra) read the
attester's effective balance with a validator(i) tree descent inside the
per-flag loop, up to once per newly earned flag. Read it lazily once per
attester instead: the first newly earned flag fetches it, the others reuse it.
Staying lazy keeps the error behavior identical, since a validator that earns
no flag is never read. The effective balance cannot change inside the
read-only phase, so the arithmetic is unchanged.
@MegaRedHand MegaRedHand added the beacon Ethereum Beacon Chain client label Oct 2, 2026
@MegaRedHand

Copy link
Copy Markdown
Collaborator Author

Import-replay benchmark

Method. ethlambda benchmark import replay, release build, 128 consecutive mainnet blocks (slots 15279073..=15279200, 4 of them epoch-start blocks) on a 16-core host. This PR's leg and a base leg (beacon-chain-integration) were interleaved, 2 rounds each; the table shows the mean of the two. Both legs carry #651, so stf excludes the state-writer hand-off wait. Block roots were identical to the base in every leg. Base vs base noise floor: 1.004.

base this PR change
stf ordinary block p50 577 ms 353 ms −39%
stf epoch-start block (mean of 4) 1505 ms 1250 ms −17%
get_head p50 200 ms 191 ms −5%
block wall p50 777 ms 554 ms −29%
user CPU (whole run) 214.7 s 190.9 s −11%
peak RSS 3.90 GiB 4.42 GiB +0.52
paired stf, ratio of totals (round 1 / round 2) 0.700 / 0.671

The ordinary-block saving (224 ms) sits just under the estimated 300-414 ms. Peak RSS rises although the cache holds a few u64s. The most likely reason is that faster imports leave more states queued for the background state writer at once; that is not verified.

All five performance PRs together (#647, #648, #649, #650, #633), measured with a 1 s delay between blocks so the state writer drains (--block-delay, #651):

  • ordinary stf 527 → 163 ms; epoch-start 1440 → 420 ms;
  • block wall p50 729 → 185 ms (3.9x); user CPU −30%.

Without the delay, the combined build outpaces the state writer and waits ~326 ms per block on it.

MegaRedHand added a commit that referenced this pull request Oct 2, 2026
…63-64-633-636-638-gloas-live

Conflicted files:
- crates/blockchain/state_transition/src/beacon/stf/block.rs
- crates/blockchain/state_transition/src/beacon/stf/electra.rs
- crates/blockchain/state_transition/src/beacon/stf/fulu.rs
- crates/blockchain/state_transition/src/beacon/stf/operations.rs
- crates/blockchain/state_transition/tests/beacon_spec/operations.rs
- crates/storage/src/store.rs
- docs/beacon_stf.md

Gloas adaptations:
- stf/gloas.rs process_block, process_operations and process_attestation
  take the ActiveBalanceCache; the gloas block's sync aggregate goes
  through the shared altair process_sync_aggregate with it. The base
  reward per increment comes from the cache, and the attester's
  effective balance is read once per attester (it also feeds the pending
  builder payment weight), as in the other forks.
- Soundness check for gloas: effective_balance has one writer outside the
  fork upgrades (queue_entire_balance_and_reset_validator, called only by
  the electra upgrade) and genesis, besides the epoch-end
  process_effective_balance_updates. Gloas block processing (parent
  payload application, withdrawals, bids, builder payments and deposits,
  execution requests, slashings, exits) writes balances, exit epochs,
  slashed flags and queues, never an effective balance, and exits and
  activations land at least MAX_SEED_LOOKAHEAD epochs ahead, so the cache
  is used for gloas too. A full or empty parent changes none of those
  inputs. The debug cross-check on every hit runs in the gloas sanity and
  transition fixtures.
- The gloas arm of block::process_block keeps its engine-free call and
  gains the cache argument.
- fork_choice.rs and the fixture runners (operations gloas arm) pass the
  cache; get_total_active_balance walks iter_validators().

Other: store.rs keeps both the liveness tracker and the cache field and
accessor; fulu.rs keeps its own process_operations and pending-queue
import with the cache argument; docs follow the element accessor names
and no longer mention registry_mut.
MegaRedHand added a commit that referenced this pull request Oct 2, 2026
…636-638-gloas-live

Conflicted files:
- crates/blockchain/src/lib.rs
- crates/blockchain/state_transition/src/beacon/fork_choice.rs
- crates/blockchain/state_transition/src/metrics.rs
- docs/metrics.md

Gloas: the precompute stays enabled, because the base state cannot differ.
In this branch a gloas block's post-state never holds its own payload: the
child applies it (process_parent_execution_payload) inside process_block,
after process_slots, and on_execution_payload_envelope only records and
checks, it stores no second state. The store keeps one state per root and
both on_block (get_state(parent_root)) and checkpoint_state read it, and the
worker clones get_state(head_root) and caches the result under
CheckpointState { E+1, head_root }, so the entry is process_slots of the very
state an import of a child of that root clones. The fulu-to-gloas boundary
needs no case either: the upgrade runs inside process_slots on both paths.
A new unit test (a_gloas_precompute_starts_from_the_state_an_import_would)
checks, for a gloas parent, that the worker's result equals
checkpoint_state's for the key, that an import finds it, and that advancing
the rest of a skipped first slot from it equals plain process_slots.

Other adaptations:
- transition_block and stf::apply_block take the ActiveBalanceCache that
  #648 added to state_transition; on_block passes the store's own cache.
- on_block's gloas-specific code between the parent read and the
  transition does not use the cloned state, so removing that early clone
  (as #652 does) needed nothing else.
- lib.rs: the NewBlock handler keeps the envelope settling that follows the
  import and triggers the precompute after it, so a head moved by a settled
  envelope's block is seen.
- metrics.rs and docs/metrics.md: both sides' metrics kept.
- Test fixtures of #652 follow this branch's BlockInputs (operations) and
  the extra argument.
- docs/beacon_stf.md notes why gloas shares the key.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

beacon Ethereum Beacon Chain client

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant