perf(beacon): weigh fork-choice votes from a justified-balances snapshot - #649
MegaRedHand wants to merge 2 commits into
Conversation
Since the validator registry moved onto a persistent tree, every `state.validator(i)` in the fork-choice vote loop is a tree descent, and `get_proposer_score` rescans the registry through `get_total_active_balance` (an active-index Vec, then one descent per active index). Both run on every `get_head`, so the head recompute went from flat-array reads to millions of descents per call at mainnet scale. Flatten the justified checkpoint state once into `JustifiedBalances`: a per-validator weight (zero when inactive or slashed) and the spec-exact total active balance (slashed-but-active validators count, floored at one increment), computed in the same in-order pass. It is keyed by the justified checkpoint itself, so every writer of that checkpoint is covered without a hook, and its only source is `checkpoint_state(justified)`, the state `get_weight` reads. `get_weight` stays spec-literal and is the oracle: the fork-choice fixture runner now compares `compute_weights` with it for every block of the filtered tree at each `checks` step. `calculate_committee_fraction` is split so `is_head_weak`, `is_parent_strong` and the boost share the snapshot's total. Adds lean_beacon_justified_balances_lookups_total and lean_beacon_justified_balances_build_seconds.
The vote loop walked a HashMap<u64, LatestMessage> in hash order and probed the equivocator set once per vote, then read the snapshot balance at a random offset. Store the votes as a Vec<Option<LatestMessage>> indexed by validator instead, so votes and snapshot balances are read in index order and the equivocator probe is skipped while the set is empty. Saturating sums of non-negative balances give the same weights in any order, so the result is unchanged. The table grows to the highest voting index, so set_latest_message asserts (debug builds) that the index is plausible; indices only come from validated attestations, and a wild one would otherwise allocate a table to match.
Import-replay benchmarkMethod.
This PR does not touch block processing, yet All five performance PRs together (#647, #648, #649, #650, #633), measured with a 1 s delay between blocks so the state writer drains (
|
…3-636-638-gloas-live Conflicted files: - crates/blockchain/state_transition/src/beacon/fork_choice.rs - crates/blockchain/state_transition/src/metrics.rs (rerere reused an earlier resolution: both metric sets kept) - crates/common/types/src/beacon/fork_choice.rs - crates/storage/src/store.rs - docs/metrics.md (same rerere resolution) Gloas adaptations: - compute_node_weights, the payload-aware vote loop every head computation uses (this branch's gloas fork choice has no compute_weights path in production), reads each voter's weight from the justified-balances snapshot instead of state.validator(i) per vote. A zero in the snapshot (past the registry, inactive, slashed) is a skip, as the old missing/slashed/inactive checks were; a zero effective balance adds nothing either way. The payload status, previous-slot rule, bound and the gated proposer boost are untouched: should_apply_proposer_boost_with still gets the justified state itself, and get_proposer_score uses the snapshot total, which is the number the justified state gave before. - The spec-literal gloas references (gloas_get_weight, gloas_get_attestation_score, gloas_get_head) are not changed. - is_head_weak and is_parent_strong keep this branch's versions (they score against the justified state, including the equivocator top-up); only the pre-gloas shapes #649 rewrote were dropped, so their thresholds still come from calculate_committee_fraction over that state. - Fixture runner: the compute_weights vs get_weight oracle #649 added is kept and runs on every case; a gloas counterpart compares compute_node_weights with gloas_get_weight for every node under the justified root (not only the leaves viable_for_head_roots_and_weights covers) at each checks step of a case whose current slot is gloas. - The dense latest_messages tests build the gloas LatestMessage (slot, payload_present); the snapshot builder walks iter_validators(). Other: types fork_choice.rs and the storage import list keep both sides' imports.
Motivation
Since the registry moved onto a persistent tree (lambdaclass/ethlambda_private#37),
state.validator(i)is a tree descent instead of an array read. Two places inget_headpay one per validator:compute_weightsdoes a descent per latest vote, then a hash-order probe into the equivocator set and a SipHash map.get_proposer_scorecallsget_total_active_balance: an active-indexVec, then one descent per active index.At mainnet scale (estimate: about 1M voters and 1M active validators in a 2.4M registry) that is of the order of two million descents per
get_head. Theget_headregression after the tree PR was measured at 126 ms to 198 ms p50 in the import-replay benchmark; how much of it is these two loops was inferred from the code, not profiled. The import-replay benchmark was not run as part of this work, so this PR carries no timing claims.Changes
Two commits, so the PR can be split along the commit boundary if preferred.
perf(beacon): weigh fork-choice votes from a justified-balances snapshotJustifiedBalances(types crate): per-validator vote weight (0 if inactive or slashed) plus the spec-exact total active balance, built in one in-order pass overcheckpoint_state(justified)and cached inBeaconScratch, keyed by its own checkpoint.compute_weights,get_proposer_score,is_head_weak,is_parent_strongread it.calculate_committee_fractionis split into a purecommittee_fraction(total, percent). Metricslean_beacon_justified_balances_lookups_total{result}andlean_beacon_justified_balances_build_seconds; docs indocs/metrics.mdanddocs/beacon_stf.md.perf(beacon): keep latest messages dense, indexed by validatorBeaconScratch.latest_messagesbecomesVec<Option<LatestMessage>>, so the vote loop walks votes and snapshot balances in index order and skips the equivocator probe while the set is empty. No API change.Tests:
get_weightstays spec-literal as the oracle. The weights-vs-spec unit test now includes slashed, not-yet-active, exited and past-the-registry voters while a boost is set.build_justified_balancesedge cases (activation and exit at the epoch, slashed-but-active, all-inactive floor, index past the end); a new justified checkpoint rebuilds the snapshot;get_proposer_scorecounts a slashed-but-active validator; dense storage table (index order, gaps, overwrite, equivocators).fork_choicefixture runner now asserts, at everychecksstep, thatcompute_weightsequalsget_weightfor every root of the filtered block tree. No fixture checks weights directly, so this puts the snapshot against the spec path on every case. It skips a root only whereget_weightitself errors because invalidation removed the voted block (thesync/*/optimistic/from_syncing_to_invalidcases), sincecompute_weightsdrops those votes by design. It is the only edit to the runner.Design decisions
saturating_add). Not the unslashed sum some other clients use.Box<[Gwei]>, notu16increments (avoids an invariant to assert).BeaconScratchfield, not a new module.checkpoint_state(justified). A later post-state of the same epoch is not exact (slashedcan change mid-epoch). Equivocations stay out of the snapshot and are filtered per vote.is_head_weak/is_parent_strongcheckpoint_state(justified), so it is the same state.debug_assert!that an index is below a plausible bound before the table grows; the table has one slot per index up to the highest seen, so only validated attestation indices may reach it. Release builds do not check.BeaconScratchlock on this path: nocrates/netcode calls the snapshot or vote accessors.Deviation from the original sketch: the builder does not
debug_assertthat the state's epoch equals the checkpoint's epoch. Two existing unit-test stores pair an epoch-0 checkpoint with a state one epoch in, and activity is evaluated at the state's own epoch exactly asget_weightdoes, so correctness does not depend on it.Not in this PR
get_headmax at epoch-start blocks after this PR.update_latest_messagesand a root-prefix hasher for the weights map.Open question: the replay harness never ticks, so the proposer boost never resets there and every replayed
get_headpays the boost path. Should it tick? Out of scope here.Testing
Run with
--profile release-fast; spec fixtures from the existing local copy.cargo clippy --locked --workspace --all-targets -- -D warningscargo clippy -p ethlambda-state-transition --all-targets --features beacon-spec-tests -- -D warningscargo test -p ethlambda-types --libcargo test -p ethlambda-storage --libcargo test -p ethlambda-state-transition --libgossip/*/matched_fixture_casesfail only because the gossip fixture tree is empty in this environmentcargo check -p ethlambda