Skip to content

fix: Use double precision and integer weight sums in rollout bucketing - #231

Open
tanderson-ld wants to merge 1 commit into
mainfrom
tanderson/sdk-1537-bucketing-precision
Open

tanderson-ld wants to merge 1 commit into
mainfrom
tanderson/sdk-1537-bucketing-precision

Conversation

@tanderson-ld

@tanderson-ld tanderson-ld commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Requirements

  • I have added test coverage for new or changed functionality
  • I have followed the repository's pull request submission guidelines
  • I have validated my changes against all supported platform versions

Related issues

Fixes #94 (internal tracking: SDK-1537)

Describe the solution you've provided

The evaluation spec (FLGEA §II.F steps 6d and 7a) requires double-precision floating-point arithmetic for bucketing, and requires rollout bucket boundaries to be computed by summing the integer weights and then dividing by 100000. This SDK was doing both in single precision, and was accumulating a running float sum of per-weight quotients. With long weight lists (e.g. mutually exclusive experiments in a shared layer), the accumulated rounding error shifts bucket boundaries differently per flag, so a context near a boundary can be placed in overlapping experiments — the behavior reported in #94.

Changes:

  • EvaluatorBucketing.computeBucketValue now returns double, and the hash division by 2^60-1 is performed in double precision.
  • Evaluator.getValueForVariationOrRollout sums the weights as a long and divides by 100000.0 at each comparison instead of accumulating a float sum.
  • The segment-rule weight path already divided in double precision; it now also receives a double-precision bucket value.

Verification:

  • New EvaluatorBucketingPrecisionTest reproduces the exact scenario from Non-exclusive experiment assignments on layer due to bucketization rounding error #94 (real 551-variation experiment data, seed, and context key). Both tests fail against the previous implementation (0.98308944702... float bucket value and wrong bucket assignment) and pass with this change.
  • RolloutRandomizationConsistencyTest's hard-coded cross-SDK values pass unchanged within their existing 1e-7 tolerances (float→double moves results by at most ~3e-8).
  • Full java-server-sdk test suite passes.

Describe alternatives you've considered

Making the bucketing algorithm configurable (a "BucketerV1"/"BucketerV2" selection, so customers could adopt the fix on their own schedule) was discussed on #94. That adds permanent complexity for a correction whose blast radius is small, and the SDK would remain out of spec by default; the direct fix follows the precedent of previous cross-SDK bucketing-consistency fixes.

Additional context

This is a behavioral change for contexts whose bucket value falls between the old (single-precision) and new (double-precision) bucket boundaries:

  • Typical rollouts (2–5 variations): reassignment probability ~1e-8 per context per flag; boundaries at exact binary fractions (50/50, 25×4) have zero boundary drift.
  • Many-variation experiment layers (the configurations affected by the bug): up to ~2.5e-4 (1 in ~4,000 contexts) for the real flag data in Non-exclusive experiment assignments on layer due to bucketization rounding error #94. These reassignments move contexts to the bucket that the spec — and the other server SDKs, which already use double precision — assign them to.

Contexts not near a bucket boundary (the overwhelming majority) are unaffected, and assignments remain deterministic after the one-time step at upgrade.

Note for release planning: merging with fix: will cut a patch release via release-please. The release-notes wording for the assignment discontinuity (and whether any additional comms are needed for experiment-heavy customers) is still pending the experimentation-team review, so hold the merge if that should land first.


Note

Overview
Aligns Java server SDK rollout bucketing with the evaluation spec (FLGEA) and other SDKs by fixing two precision bugs that could mis-assign contexts near bucket boundaries—especially in large experiment layers.

computeBucketValue now returns double and normalizes the hash with double arithmetic instead of single-precision float.

Rollout variation selection no longer accumulates a running float sum of weight/100000; it sums weights as a long and compares bucket < weightSum / 100000.0 per bucket so rounding cannot drift boundaries differently across flags with long weight lists.

Adds EvaluatorBucketingPrecisionTest reproducing #94 (551-variation experiment); existing bucketing tests are updated for double. Contexts not near a boundary stay stable; a small fraction near old float boundaries may get a one-time reassignment at upgrade.

Reviewed by Cursor Bugbot for commit 140211b. Bugbot is set up for automated code reviews on this repo. Configure here.

The bucket value hash division is now performed in double precision, and
rollout bucket boundaries are computed by summing the integer weights and
dividing at each comparison, per the evaluation spec. Accumulating a
single-precision sum allowed rounding error to shift bucket boundaries,
which could break mutual exclusivity of experiments sharing a layer.

Fixes #94 (SDK-1537)
@tanderson-ld
tanderson-ld requested a review from a team as a code owner October 1, 2026 17:35

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Non-exclusive experiment assignments on layer due to bucketization rounding error

1 participant