Skip to content

fix(io): write NaN values into identity-partitioned float columns - #4030

Merged
Fokko merged 1 commit into
apache:mainfrom
breken-ai:fix/nan-identity-partition-write
Oct 1, 2026
Merged

Fokko merged 1 commit into
apache:mainfrom
breken-ai:fix/nan-identity-partition-write

Conversation

@breken-ai

Copy link
Copy Markdown
Contributor

Rationale for this change

Appending a NaN to a table with an identity partition on a double or float column fails:

schema = Schema(NestedField(1, "id", IntegerType()), NestedField(2, "value", DoubleType()))
spec = PartitionSpec(PartitionField(source_id=2, field_id=1000, transform=IdentityTransform(), name="value"))
tbl = catalog.create_table("default.t", schema=schema, partition_spec=spec)
tbl.append(pa.table({"id": [1, 2, 3], "value": [1.0, float("nan"), None]}))
# ZeroDivisionError: division by zero  (bin_pack_arrow_table: tbl.nbytes / tbl.num_rows)

_determine_partitions groups the rows by partition value, which yields one NaN group, and then selects each group's rows with pc.field(name) == value. NaN never compares equal to itself, so the NaN partition selects zero rows and the empty table reaches bin_pack_arrow_table. Nulls are already handled with is_null(). This PR selects the NaN partition with pc.is_nan() in the same way.

After the fix the NaN rows are written to their own partition, and value is nan returns them.

Are these changes tested?

Yes.

  • tests/io/test_pyarrow.py: test_determine_partitions_identity_nan checks that the rows for 1.0, NaN and None land in their partitions. On main it returns {'nan': []}.
  • tests/catalog/test_catalog_behaviors.py: test_append_nan_to_identity_partitioned_table appends to a real table (memory and SQL catalogs), then checks that a full scan returns all 4 rows and value is nan returns the two NaN rows. On main the append raises ZeroDivisionError.

All 4 new tests fail on main and pass with the fix. tests/io/test_pyarrow.py and tests/catalog/test_catalog_behaviors.py pass. prek run --files (ruff, ruff-format, mypy, pydocstyle, codespell) passes. I also checked float columns by hand.

Are there any user-facing changes?

Yes. append()/overwrite() now write NaN values into identity-partitioned floating-point columns instead of failing.

AI disclosure: this bug was found, fixed and tested by an AI coding agent (Claude) running under the breken-ai account; the red/green runs above are its local results.

_determine_partitions selected the rows of each partition with
field == value. NaN never compares equal to itself, so the NaN partition
matched zero rows and append() failed with ZeroDivisionError in
bin_pack_arrow_table. Select the NaN partition with is_nan, as null is
already selected with is_null.

Generated-by: Claude Opus 5.5

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@Fokko Fokko left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is indeed a bug, thanks @breken-ai for fixing this 👍

@Fokko
Fokko added this pull request to the merge queue Oct 1, 2026
Merged via the queue into apache:main with commit 687d479 Oct 1, 2026
21 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants