Skip to content

perf(beacon): weigh fork-choice votes from a justified-balances snapshot - #649

Draft
MegaRedHand wants to merge 2 commits into
beacon-chain-integrationfrom
perf/beacon-justified-balances
Draft

MegaRedHand wants to merge 2 commits into
beacon-chain-integrationfrom
perf/beacon-justified-balances

Conversation

@MegaRedHand

Copy link
Copy Markdown
Collaborator

Motivation

Since the registry moved onto a persistent tree (lambdaclass/ethlambda_private#37), state.validator(i) is a tree descent instead of an array read. Two places in get_head pay one per validator:

  • compute_weights does a descent per latest vote, then a hash-order probe into the equivocator set and a SipHash map.
  • With a proposer boost set, get_proposer_score calls get_total_active_balance: an active-index Vec, then one descent per active index.

At mainnet scale (estimate: about 1M voters and 1M active validators in a 2.4M registry) that is of the order of two million descents per get_head. The get_head regression after the tree PR was measured at 126 ms to 198 ms p50 in the import-replay benchmark; how much of it is these two loops was inferred from the code, not profiled. The import-replay benchmark was not run as part of this work, so this PR carries no timing claims.

Changes

Two commits, so the PR can be split along the commit boundary if preferred.

Commit What
perf(beacon): weigh fork-choice votes from a justified-balances snapshot JustifiedBalances (types crate): per-validator vote weight (0 if inactive or slashed) plus the spec-exact total active balance, built in one in-order pass over checkpoint_state(justified) and cached in BeaconScratch, keyed by its own checkpoint. compute_weights, get_proposer_score, is_head_weak, is_parent_strong read it. calculate_committee_fraction is split into a pure committee_fraction(total, percent). Metrics lean_beacon_justified_balances_lookups_total{result} and lean_beacon_justified_balances_build_seconds; docs in docs/metrics.md and docs/beacon_stf.md.
perf(beacon): keep latest messages dense, indexed by validator BeaconScratch.latest_messages becomes Vec<Option<LatestMessage>>, so the vote loop walks votes and snapshot balances in index order and skips the equivocator probe while the set is empty. No API change.

Tests:

  • get_weight stays spec-literal as the oracle. The weights-vs-spec unit test now includes slashed, not-yet-active, exited and past-the-registry voters while a boost is set.
  • New unit tests: build_justified_balances edge cases (activation and exit at the epoch, slashed-but-active, all-inactive floor, index past the end); a new justified checkpoint rebuilds the snapshot; get_proposer_score counts a slashed-but-active validator; dense storage table (index order, gaps, overwrite, equivocators).
  • The beacon fork_choice fixture runner now asserts, at every checks step, that compute_weights equals get_weight for every root of the filtered block tree. No fixture checks weights directly, so this puts the snapshot against the spec path on every case. It skips a root only where get_weight itself errors because invalidation removed the voted block (the sync/*/optimistic/from_syncing_to_invalid cases), since compute_weights drops those votes by design. It is the only edit to the runner.

Design decisions

Question Decision
Boost total Spec-exact: slashed-but-active validators count, floored at the increment, computed in the same pass as the snapshot (saturating_add). Not the unslashed sum some other clients use.
Storage Box<[Gwei]>, not u16 increments (avoids an invariant to assert).
Where it lives A BeaconScratch field, not a new module.
Snapshot source Only checkpoint_state(justified). A later post-state of the same epoch is not exact (slashed can change mid-epoch). Equivocations stay out of the snapshot and are filtered per vote.
Cache key The checkpoint itself, checked on every read, so no hook on the code that moves the justified checkpoint.
is_head_weak / is_parent_strong Switched to the snapshot's total: both already read checkpoint_state(justified), so it is the same state.
Dense vote bound debug_assert! that an index is below a plausible bound before the table grows; the table has one slot per index up to the highest seen, so only validated attestation indices may reach it. Release builds do not check.
Lock Only the chain actor reaches the BeaconScratch lock on this path: no crates/net code calls the snapshot or vote accessors.

Deviation from the original sketch: the builder does not debug_assert that the state's epoch equals the checkpoint's epoch. Two existing unit-test stores pair an epoch-0 checkpoint with a state one epoch in, and activity is evaluated at the state's own epoch exactly as get_weight does, so correctness does not depend on it.

Not in this PR

  • Prefetching the snapshot for the unrealized justified checkpoint (so the epoch-start pull-up finds it ready): wait for get_head max at epoch-start blocks after this PR.
  • Per-root running vote totals (delta weights): only worth it if a profile after this PR still shows the weight loop.
  • A proto-array style tree: a separate fork-choice project.
  • Optional single lock per attestation in update_latest_messages and a root-prefix hasher for the weights map.

Open question: the replay harness never ticks, so the proposer boost never resets there and every replayed get_head pays the boost path. Should it tick? Out of scope here.

Testing

Run with --profile release-fast; spec fixtures from the existing local copy.

Command Result
cargo clippy --locked --workspace --all-targets -- -D warnings clean
cargo clippy -p ethlambda-state-transition --all-targets --features beacon-spec-tests -- -D warnings clean
cargo test -p ethlambda-types --lib 198 passed
cargo test -p ethlambda-storage --lib 139 passed, 1 ignored
cargo test -p ethlambda-state-transition --lib 327 passed, 16 ignored
beacon spec suite, mainnet preset, full 5734 passed, 152 ignored; 4 gossip/*/matched_fixture_cases fail only because the gossip fixture tree is empty in this environment
beacon spec suite, minimal preset, full 40038 passed, 3692 ignored; same 4 gossip trials
cargo check -p ethlambda ok

Since the validator registry moved onto a persistent tree, every
`state.validator(i)` in the fork-choice vote loop is a tree descent, and
`get_proposer_score` rescans the registry through `get_total_active_balance`
(an active-index Vec, then one descent per active index). Both run on every
`get_head`, so the head recompute went from flat-array reads to millions of
descents per call at mainnet scale.

Flatten the justified checkpoint state once into `JustifiedBalances`: a
per-validator weight (zero when inactive or slashed) and the spec-exact total
active balance (slashed-but-active validators count, floored at one
increment), computed in the same in-order pass. It is keyed by the justified
checkpoint itself, so every writer of that checkpoint is covered without a
hook, and its only source is `checkpoint_state(justified)`, the state
`get_weight` reads. `get_weight` stays spec-literal and is the oracle: the
fork-choice fixture runner now compares `compute_weights` with it for every
block of the filtered tree at each `checks` step.

`calculate_committee_fraction` is split so `is_head_weak`,
`is_parent_strong` and the boost share the snapshot's total.

Adds lean_beacon_justified_balances_lookups_total and
lean_beacon_justified_balances_build_seconds.
The vote loop walked a HashMap<u64, LatestMessage> in hash order and probed
the equivocator set once per vote, then read the snapshot balance at a
random offset. Store the votes as a Vec<Option<LatestMessage>> indexed by
validator instead, so votes and snapshot balances are read in index order
and the equivocator probe is skipped while the set is empty. Saturating sums
of non-negative balances give the same weights in any order, so the result is
unchanged.

The table grows to the highest voting index, so set_latest_message asserts
(debug builds) that the index is plausible; indices only come from validated
attestations, and a wild one would otherwise allocate a table to match.
@MegaRedHand

Copy link
Copy Markdown
Collaborator Author

Import-replay benchmark

Method. ethlambda benchmark import replay, release build, 128 consecutive mainnet blocks (slots 15279073..=15279200, 4 of them epoch-start blocks) on a 16-core host. This PR's leg and a base leg (beacon-chain-integration) were interleaved, 2 rounds each; the table shows the mean of the two. Both legs carry #651, so stf excludes the state-writer hand-off wait. Block roots were identical to the base in every leg. Base vs base noise floor: 1.004.

base this PR change
get_head p50 200 ms 15 ms −92%
block_atts p50 5.4 ms 2.2 ms −59%
stf ordinary block p50 577 ms 618 ms +7%
stf epoch-start block (mean of 4) 1505 ms 1583 ms +5%
block wall p50 777 ms 654 ms −16%
user CPU (whole run) 214.7 s 197.9 s −8%
peak RSS 3.90 GiB 4.39 GiB +0.49
paired block wall, ratio of totals (round 1 / round 2) 0.829 / 0.829

get_head lands well under the estimated ≤ 126 ms.

This PR does not touch block processing, yet stf rises 7%. The writer-wait phase shows the state-writer queue never fills here, so the cause is not that wait. A likely cause is contention with the writer thread's state encode: with get_head this much shorter, that encode overlaps stf instead of get_head. Supporting this, a 1 s delay between blocks lowers even the base's stf by ~50 ms. Not verified on this PR's leg.

All five performance PRs together (#647, #648, #649, #650, #633), measured with a 1 s delay between blocks so the state writer drains (--block-delay, #651):

  • ordinary stf 527 → 163 ms; epoch-start 1440 → 420 ms;
  • get_head 200 → 16 ms;
  • block wall p50 729 → 185 ms (3.9x); user CPU −30%.

MegaRedHand added a commit that referenced this pull request Oct 2, 2026
…3-636-638-gloas-live

Conflicted files:
- crates/blockchain/state_transition/src/beacon/fork_choice.rs
- crates/blockchain/state_transition/src/metrics.rs (rerere reused an earlier
  resolution: both metric sets kept)
- crates/common/types/src/beacon/fork_choice.rs
- crates/storage/src/store.rs
- docs/metrics.md (same rerere resolution)

Gloas adaptations:
- compute_node_weights, the payload-aware vote loop every head computation
  uses (this branch's gloas fork choice has no compute_weights path in
  production), reads each voter's weight from the justified-balances
  snapshot instead of state.validator(i) per vote. A zero in the snapshot
  (past the registry, inactive, slashed) is a skip, as the old
  missing/slashed/inactive checks were; a zero effective balance adds
  nothing either way. The payload status, previous-slot rule, bound and
  the gated proposer boost are untouched: should_apply_proposer_boost_with
  still gets the justified state itself, and get_proposer_score uses the
  snapshot total, which is the number the justified state gave before.
- The spec-literal gloas references (gloas_get_weight,
  gloas_get_attestation_score, gloas_get_head) are not changed.
- is_head_weak and is_parent_strong keep this branch's versions (they score
  against the justified state, including the equivocator top-up); only the
  pre-gloas shapes #649 rewrote were dropped, so their thresholds still
  come from calculate_committee_fraction over that state.
- Fixture runner: the compute_weights vs get_weight oracle #649 added is kept
  and runs on every case; a gloas counterpart compares compute_node_weights
  with gloas_get_weight for every node under the justified root (not only
  the leaves viable_for_head_roots_and_weights covers) at each checks step
  of a case whose current slot is gloas.
- The dense latest_messages tests build the gloas LatestMessage (slot,
  payload_present); the snapshot builder walks iter_validators().

Other: types fork_choice.rs and the storage import list keep both sides'
imports.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

beacon Ethereum Beacon Chain client

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant