Skip to content

perf(beacon): keep inactivity scores and the ring buffers on the tree - #650

Draft
MegaRedHand wants to merge 3 commits into
beacon-chain-integrationfrom
perf/beacon-more-state-fields-on-tree
Draft

MegaRedHand wants to merge 3 commits into
beacon-chain-integrationfrom
perf/beacon-more-state-fields-on-tree

Conversation

@MegaRedHand

Copy link
Copy Markdown
Collaborator

Motivation

After lambdaclass/ethlambda_private#37 (persistent trees for validators and balances) every cached state still holds a private deep copy of the other large fields and rehashes them in full on each state root.

Field(s) Mainnet size per state Flat cost today
inactivity_scores ~18 MiB copied per state, rehashed per root, rewritten for every eligible index each epoch
randao_mixes ~2 MiB same
block_roots, state_roots, slashings ~0.6 MiB same
eth1_data_votes, historical_roots, historical_summaries ~0.25 MiB same

Estimates, not measurements: ~26 MiB of flat fields per state and ~0.85M Merkle nodes rehashed per root, of which this PR moves the large majority of the nodes (participation lists and queues stay flat). The import-replay benchmark was not run for this PR, and no memory or node-count number was measured; the estimates come from the field sizes above.

Changes

Commit What
perf(ssz-tree): flush tree fields through one Buffered list per fork Object-safe Buffered trait (impls for List and Vector); tree_fields! lists each fork's tree fields once and generates both apply_pending_mutations and has_pending_mutations. H256 writes HashTreeRoot by hand and forwards is_basic_type (as H160 does). New ssz-tree properties: multi-leaf u8 lists, u64 and 32-byte-newtype vectors, decode parity, rebase of mostly-zero lists and vectors.
perf(beacon): keep inactivity scores on the tree and skip no-op writes InactivityScores is a tree list (BTreeMap update map). process_inactivity_updates computes the same value with the same checked arithmetic and writes only if new != score. Joins the flush list and BeaconState::rebase_on.
perf(beacon): keep the roots and randao ring buffers on the tree BlockRoots, StateRoots, RandaoMixes, Slashings become tree vectors; Eth1DataVotes, HistoricalRoots, HistoricalSummaries become tree lists (alias swap). All join the flush list and rebase_on. Block production now flushes before it hashes the post-state. Test-only sites that sliced or iterated in place use indexing.

SSZ bytes, roots and the storage delta format are unchanged.

Design decisions

  • H256 change done locally, not in libssz's transparent derive (no dependency bump). The lean chain shares H256, so ethlambda-types has unit tests that an SszList/SszVector of H256 hashes like one of [u8; 32], and that a container holding H256 collections keeps its root. The lean state's pinned genesis root test passes unchanged.
  • BTreeMap update map for inactivity scores. Outside a leak almost no score changes, so writes are sparse. A leak writes nearly every score, which costs about one B-tree insert per validator per epoch (~1M at the benchmark's registry size) until a bulk write path exists. A dense map (VecMap) was the alternative; it allocates 16 B per validator on the first push or write.
  • Skip loop written by hand in process_inactivity_updates: read the score, compute the new one exactly as before (overflow still raises ArithmeticOverflow before the recovery step), write only if different. A write rebuilds its leaf even for an equal value, which would unshare the list from the parent state's.
  • rebase_on pairs scores through the altair accessors, so it works across forks and skips phase0 (no scores).
  • process_slot leaves its two roots writes buffered until the next flush (the next slot's, or process_slots' last step). process_historical_summaries_update hashes the roots vectors with that write pending: correct, and the vectors are small. Tests pin this.
  • Block production flush: produce_block hashed the post-state without apply_pending_mutations, so the computed hashes were discarded; one added call.
  • Buffered and tree_fields! are additive in ssz-tree and containers/mod.rs, to keep a merge with other work in the same files easy.

Not in this PR

Item Why
Participation lists on the tree (previous_/current_epoch_participation) Blocked on a zip pass for get_unslashed_participating_indices (#632): its per-index reads would cost ~2M tree descents per block in the pulled-up tip. Also needs List::filled for an O(1) rotation.
pending_deposits on the tree Depends on the mainnet queue length, which has not been measured.
Sync committees (Arc + cached-root newtype) Small win (~1.5 MiB per state), separate design.
Rebase across an epoch boundary (decoded previous equals the base's current) Positional rebase misses it; one if in rebase_on, left for the participation PR.

Testing

Command Result
cargo test -p ethlambda-ssz-tree --profile release-fast 35 passed (properties), 0 failed
cargo test -p ethlambda-types --profile release-fast --lib 201 passed
cargo test -p ethlambda-storage --profile release-fast --lib 139 passed, 1 ignored
cargo test -p ethlambda-state-transition --profile release-fast --lib 329 passed, 16 ignored
cargo test -p ethlambda-rpc --profile release-fast --lib 110 passed
Beacon spec suite, mainnet preset (full) 5734 passed, 152 ignored; 4 failed
Beacon spec suite, minimal preset (full) 40038 passed, 3692 ignored; 4 failed
cargo clippy (workspace, and state-transition with beacon-spec-tests) with -D warnings, cargo fmt --check, cargo check -p ethlambda clean

The 4 failures in each spec run are the gossip/* trials, which report "matched no fixture cases": the gossip fixture directory is intentionally empty in the local setup. Every ssz_static, sanity, finality, random, epoch_processing and transition trial passes. The lean ssz_spectests and the import-replay benchmark were not run.

Moving more BeaconState fields onto the tree would otherwise mean editing
apply_pending_mutations and has_pending_mutations by hand for each one, and
the two could drift apart (a field flushed but not counted as pending, or the
reverse). An object-safe `Buffered` trait over `List` and `Vector`, and a
`tree_fields!` list per fork that generates both views, keep them in step.

H256 now forwards `is_basic_type`, as H160 already does: the transparent
derive drops it, so a tree-backed collection of roots was treated as
composite and kept a second cached copy of every root. The root is unchanged
(libssz packs 32-byte basic elements one per chunk); tests compare
`SszList`/`SszVector` of H256 against `[u8; 32]` and pin a lean container.

New ssz-tree properties cover multi-leaf u8 lists, u64 and 32-byte-newtype
vectors, their decode parity, and rebase of mostly-zero lists and vectors.
Inactivity scores are the largest flat field of a state (8 bytes per
validator) and every cached state held a private deep copy, rehashed in full
on each state root. On the tree, a derived state shares the list with its
parent and rehashes only the leaves it touched.

That only pays if epoch processing stops rewriting every eligible score: a
write rebuilds its leaf even for an equal value. Outside a leak a missed
epoch adds the bias and recovery takes it back, so nearly every score ends
where it started. process_inactivity_updates now computes the new score with
the same arithmetic as before (same checked add, same errors) and writes it
only if it differs.

The field joins the per-fork flush list, and BeaconState::rebase_on pairs the
two states' scores whenever both forks carry them.

Writes are buffered in a BTreeMap: sparse outside a leak. A leak writes about
every score of the registry, which costs one B-tree insert each until a bulk
write path exists.
block_roots, state_roots, randao_mixes and slashings are fixed-size ring
buffers that change by a handful of entries per block or epoch, yet every
cached state held a private deep copy of each and rehashed it in full on
every state root. As tree vectors a derived state shares the unchanged leaves
with its parent and rehashes only the touched paths. The small lists
(eth1_data_votes, historical_roots, historical_summaries) move by the same
alias swap, so no flat field is left that a derived state copies in bulk
apart from the participation lists and the queues.

All join the per-fork flush list and BeaconState::rebase_on, so a decoded
state shares them with its resident parent.

process_slot writes its two roots after its flush, so they stay buffered
until the next flush (the next slot's, or process_slots' last step); the
historical summaries update hashes them with that write pending, which is
correct and cheap.

Block production hashed the post-state without flushing it first, so every
node it computed was thrown away; it flushes now, as the state transition
does.

Test-only sites that sliced or iterated these fields in place (the RPC test
states, the storage delta fixtures) write through indexing instead.
@MegaRedHand

Copy link
Copy Markdown
Collaborator Author

Import-replay benchmark

Method. ethlambda benchmark import replay, release build, 128 consecutive mainnet blocks (slots 15279073..=15279200, 4 of them epoch-start blocks) on a 16-core host. This PR's leg and a base leg (beacon-chain-integration) were interleaved, 2 rounds each; the table shows the mean of the two. Both legs carry #651, so stf excludes the state-writer hand-off wait. Block roots were identical to the base in every leg. Base vs base noise floor: 1.004.

base this PR change
peak RSS 3.90 GiB 3.31 GiB −0.59 GiB
stf ordinary block p50 577 ms 530 ms −8%
stf epoch-start block (mean of 4) 1505 ms 1436 ms −5%
get_head p50 200 ms 197 ms
block wall p50 777 ms 728 ms −6%
user CPU (whole run) 214.7 s 212.2 s −1%
paired stf, ratio of totals (round 1 / round 2) 0.922 / 0.915

This matches the estimates: peak RSS was expected at 4.0 → 3.3 GiB and the ordinary-block saving at ~30 ms (47 ms measured). It is the only one of these PRs that lowers memory. RSS at the end of the import was not recorded, only the peak. The peak was consistent across both rounds (3.31 / 3.31 GiB).

All five performance PRs together (#647, #648, #649, #650, #633), measured with a 1 s delay between blocks so the state writer drains (--block-delay, #651):

  • ordinary stf 527 → 163 ms; epoch-start 1440 → 420 ms;
  • block wall p50 729 → 185 ms (3.9x); user CPU −30%.

In that build #633's fused epoch pass reads and writes the tree-backed inactivity scores in order (iter, and iter_cow from #647).

MegaRedHand added a commit that referenced this pull request Oct 2, 2026
…3-64-633-636-638-gloas-live

Conflicted files:
- crates/blockchain/state_transition/src/beacon/stf/epoch/altair.rs
- crates/blockchain/state_transition/src/beacon/stf/mod.rs
- crates/common/ssz-tree/tests/model.rs
- crates/common/types/src/beacon/containers/mod.rs
- crates/common/types/src/beacon/containers/shared.rs

Gloas adaptations:
- tree_fields! gets a gloas line: validators and balances (progressive
  trees) plus block_roots, state_roots, historical_roots, eth1_data_votes,
  randao_mixes, slashings and historical_summaries, which gloas takes from
  the shared aliases. inactivity_scores is left out on purpose: gloas
  declares its own flat libssz ProgressiveList for it, which buffers
  nothing, so there is nothing to flush or rebase.
- ssz-tree: ProgressiveList implements Buffered, so the field list can
  hold the progressive registry beside the bounded fields.
- apply_pending_mutations / has_pending_mutations dispatch over all eight
  beacon forks through buffered()/buffered_mut(); the lean guard on
  has_pending_mutations is kept (the merge had dropped it with the old
  registry match).
- rebase_on: the validators/balances list-kind match stays (so a gloas
  state still rebases onto a gloas base only); the other fields are shared
  types in every fork and rebase unconditionally; inactivity scores
  rebase only for a Tree/Tree pair.
- Inactivity scores are a tree List before gloas and a flat slice-like list
  in gloas, so no &[u64] can serve both. altair_validator_lists now returns
  an InactivityScoresRef view (len, get, in-order iter, ptr_eq,
  has_pending_updates, PartialEq) as its third element, and
  inactivity_scores_mut (a &mut [u64]) becomes inactivity_score_mut(index),
  an element write that dispatches over both kinds.
- stf/epoch/altair.rs update_inactivity_scores: decides over an in-order
  walk of the scores (in step with the summary) and writes only changed
  scores afterwards, which subsumes both #633's "write if different" and
  #650's no-op skip; the error order (first eligible validator without a
  score, then checked-add overflow) is unchanged.
- stf/mod.rs tests: BEACON_FORKS now includes Gloas, the historical
  summaries write covers gloas, and the inactivity score case expects no
  pending write on gloas. process_slot's registry check uses the new
  BeaconState::registry_has_pending_updates, since the per-slot roots
  writes stay buffered by design.
- storage store.rs test and the epoch-processing fixture runner comments
  follow the renamed accessors.

Other: shared.rs imports both List/Vector and the progressive alias; model.rs
keeps both sides' tests (the base side was empty).

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

beacon Ethereum Beacon Chain client

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant