Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,11 +12,11 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
- `vortex inspect` (text and `--html`) shows min/max for Rust-written files: bounds are now read from the zone-map table, not only array-level stats ([#416](https://github.com/dfa1/vortex-java/issues/416)).
- `vortex inspect --html` per-chunk min/max for dictionary columns showed the dictionary-code range instead of the value range ([#416](https://github.com/dfa1/vortex-java/issues/416)).
- vortex-jni no longer aborts the JVM on a filtered or row-range read of a vortex-java file: every array buffer declared 64-byte alignment, and Rust refuses to slice a buffer off a multiple of its declared alignment ([#418](https://github.com/dfa1/vortex-java/issues/418)).

- Filtered vortex-jni reads of zone-mapped vortex-java files returned too few rows, often none: the zone map declared `WriteOptions#chunkSize()` as its zone length whatever the real chunk sizes, and Rust maps rows to zones by that stride. It now declares the actual batch length, or none when batches differ ([#418](https://github.com/dfa1/vortex-java/issues/418)).

### Changed

- **Breaking:** removed `WriteOptions#chunkSize`. `writeChunk` never split batches by it; its only effect was the wrong zone-map stride fixed above. Each `writeChunk` call is one chunk ([#418](https://github.com/dfa1/vortex-java/issues/418)).
- **Breaking (custom encoders):** `EncodeResult` buffers are now `EncodedBuffer`s carrying their element alignment; `EncodeResult.simple` takes an `EncodedBuffer` (`EncodedBuffer.of(seg, ptype)` / `EncodedBuffer.bytes(seg)`) ([#418](https://github.com/dfa1/vortex-java/issues/418)).

### Added
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -108,7 +108,7 @@ void minMaxOnVarcharColumnRewritesToValuesFromStats(@TempDir Path localTmp) thro
// at all, so it would abandon regardless of this fix.
Path stringsFile = localTmp.resolve("strings.vortex");
DType.Struct stringsSchema = DType.structBuilder().field("symbol", DType.UTF8).build();
WriteOptions stringsOpts = new WriteOptions(4, true, 0.90, 0, false, false, MemorySize.ofMiB(256), Map.of());
WriteOptions stringsOpts = new WriteOptions(true, 0.90, 0, false, false, MemorySize.ofMiB(256), Map.of());
try (var ch = FileChannel.open(stringsFile, StandardOpenOption.CREATE, StandardOpenOption.WRITE);
var writer = VortexWriter.create(ch, stringsSchema, stringsOpts)) {
writer.writeChunk(Map.of(ColumnName.of("symbol"), new String[]{"AAPL", "MSFT", "NVDA", "TSLA"}));
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -50,7 +50,7 @@ static void writeFile() throws Exception {
// fixture's globalDict "symbol" column covers the other shape in the same test (issue #409).
Path stringsFile = tmp.resolve("strings.vortex");
DType.Struct stringsSchema = DType.structBuilder().field("symbol", DType.UTF8).build();
WriteOptions stringsOpts = new WriteOptions(4, true, 0.90, 0, false, false, MemorySize.ofMiB(256), Map.of());
WriteOptions stringsOpts = new WriteOptions(true, 0.90, 0, false, false, MemorySize.ofMiB(256), Map.of());
try (var ch = FileChannel.open(stringsFile, StandardOpenOption.CREATE, StandardOpenOption.WRITE);
var writer = VortexWriter.create(ch, stringsSchema, stringsOpts)) {
writer.writeChunk(Map.of(ColumnName.of("symbol"), new String[]{"AAPL", "MSFT", "NVDA", "TSLA"}));
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -50,7 +50,7 @@ private SchemaPlus tableOf(long[] values, boolean[] valid) throws IOException {
List.of(ColumnName.of("v")), List.of(new DType.Primitive(PType.I64, true)), false);
Path file = tmp.resolve("sum-nulls.vortex");
// Large chunk so the whole column is one chunk; zone maps on so the SUM stat is emitted.
WriteOptions opts = new WriteOptions(1024, true, 0.90, 0, false, false, MemorySize.ofMiB(256), Map.of());
WriteOptions opts = new WriteOptions(true, 0.90, 0, false, false, MemorySize.ofMiB(256), Map.of());
try (var ch = FileChannel.open(file, StandardOpenOption.CREATE, StandardOpenOption.WRITE);
var writer = VortexWriter.create(ch, schema, opts)) {
writer.writeChunk(Map.of(ColumnName.of("v"), new NullableData(values, valid)));
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -58,7 +58,7 @@ static void write() throws Exception {
.build();
// enableZoneMaps=true emits the per-chunk min/max/sum/null-count the tier-1 fold reads and the
// classify() step uses to find the boundary zones.
WriteOptions opts = new WriteOptions(CHUNK, true, 0.90, 0, true, false, MemorySize.ofMiB(256), Map.of());
WriteOptions opts = new WriteOptions(true, 0.90, 0, true, false, MemorySize.ofMiB(256), Map.of());
try (var ch = FileChannel.open(file, StandardOpenOption.CREATE, StandardOpenOption.WRITE);
VortexWriter writer = VortexWriter.create(ch, schema, opts)) {
for (int c = 0; c < CHUNKS; c++) {
Expand Down Expand Up @@ -459,7 +459,7 @@ private static Ground reduce(java.util.function.LongPredicate predicate) {
private static void writeChunks(Path file, DType.Struct schema, Map<ColumnName, Object> chunk0,
Map<ColumnName, Object> chunk1) throws Exception {
// chunkSize large so each writeChunk is exactly one chunk (one zone).
WriteOptions opts = new WriteOptions(1024, true, 0.90, 0, false, false, MemorySize.ofMiB(256), Map.of());
WriteOptions opts = new WriteOptions(true, 0.90, 0, false, false, MemorySize.ofMiB(256), Map.of());
try (var ch = FileChannel.open(file, StandardOpenOption.CREATE, StandardOpenOption.WRITE);
VortexWriter writer = VortexWriter.create(ch, schema, opts)) {
writer.writeChunk(chunk0);
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -57,7 +57,7 @@ static void write() throws Exception {
.field("val", DType.I64)
.build();
// enableZoneMaps=true emits the per-chunk min/max/sum/null-count the fold reads.
WriteOptions opts = new WriteOptions(CHUNK, true, 0.90, 0, true, false, MemorySize.ofMiB(256), Map.of());
WriteOptions opts = new WriteOptions(true, 0.90, 0, true, false, MemorySize.ofMiB(256), Map.of());
try (var ch = FileChannel.open(file, StandardOpenOption.CREATE, StandardOpenOption.WRITE);
VortexWriter writer = VortexWriter.create(ch, schema, opts)) {
for (int c = 0; c < CHUNKS; c++) {
Expand Down Expand Up @@ -319,7 +319,7 @@ void nanRowNeverMatchesAFloatingWhereFilterEvenAfterAbandoningToScan() throws Ex
List.of(new DType.Primitive(PType.F64, false), new DType.Primitive(PType.I64, false)),
false);
Path f = tmp.resolve("floating-nan.vortex");
WriteOptions opts = new WriteOptions(1024, true, 0.90, 0, false, false, MemorySize.ofMiB(256), Map.of());
WriteOptions opts = new WriteOptions(true, 0.90, 0, false, false, MemorySize.ofMiB(256), Map.of());
try (var ch = FileChannel.open(f, StandardOpenOption.CREATE, StandardOpenOption.WRITE);
VortexWriter writer = VortexWriter.create(ch, schema, opts)) {
writer.writeChunk(Map.of(
Expand Down Expand Up @@ -527,7 +527,7 @@ private static Path nullPartitionedFile(String name) throws Exception {
private static void writeChunks(Path file, DType.Struct schema, Map<ColumnName, Object> chunk0,
Map<ColumnName, Object> chunk1) throws Exception {
// chunkSize large so each writeChunk is exactly one chunk (one zone).
WriteOptions opts = new WriteOptions(1024, true, 0.90, 0, false, false, MemorySize.ofMiB(256), Map.of());
WriteOptions opts = new WriteOptions(true, 0.90, 0, false, false, MemorySize.ofMiB(256), Map.of());
try (var ch = FileChannel.open(file, StandardOpenOption.CREATE, StandardOpenOption.WRITE);
VortexWriter writer = VortexWriter.create(ch, schema, opts)) {
writer.writeChunk(chunk0);
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -21,7 +21,7 @@ private OhlcGenerator() {
}

static void write(Path file, int totalRows, int chunkSize) throws IOException {
WriteOptions opts = new WriteOptions(chunkSize, true, 0.90, 0, true, false, MemorySize.ofMiB(256), Map.of());
WriteOptions opts = new WriteOptions(true, 0.90, 0, true, false, MemorySize.ofMiB(256), Map.of());
try (var ch = FileChannel.open(file, StandardOpenOption.CREATE, StandardOpenOption.WRITE);
var writer = VortexWriter.create(ch, OhlcData.SCHEMA, opts)) {
for (OhlcData.Batch batch : OhlcData.generate(totalRows, chunkSize)) {
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -111,7 +111,7 @@ private Path write(String name, WriteOptions opts, int u32, long u64) throws Exc
private static WriteOptions noZoneMaps() {
// Same shape as the adapter coverage test's zone-maps-off options: the second flag disables
// zone maps so no per-zone SUM exists and VortexAggregates falls back to scanSum.
return new WriteOptions(65_536, false, 0.90, 0, true, false, MemorySize.ofMiB(256), Map.of());
return new WriteOptions(false, 0.90, 0, true, false, MemorySize.ofMiB(256), Map.of());
}

private static ReadRegistry registry() {
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -279,7 +279,7 @@ void nonNumericColumn_throws() throws Exception {
void noZoneMap_sumFallsBackToFullScan(@TempDir Path noStats) throws Exception {
// Given — a file written with zone maps off, so no per-zone SUM exists to fold
Path bare = noStats.resolve("nostats.vortex");
WriteOptions noZoneMaps = new WriteOptions(65_536, false, 0.90, 0, true, false, MemorySize.ofMiB(256), Map.of());
WriteOptions noZoneMaps = new WriteOptions(false, 0.90, 0, true, false, MemorySize.ofMiB(256), Map.of());
try (var ch = FileChannel.open(bare, StandardOpenOption.CREATE, StandardOpenOption.WRITE);
var w = VortexWriter.create(ch, SCHEMA, noZoneMaps)) {
w.writeChunk(Map.ofEntries(
Expand Down
4 changes: 2 additions & 2 deletions docs/reference.md
Original file line number Diff line number Diff line change
Expand Up @@ -168,11 +168,11 @@ Accepted array types per column `DType`:

### `WriteOptions` (`io.github.dfa1.vortex.writer.WriteOptions`)

Record: `(int chunkSize, boolean enableZoneMaps, double compressionRatioThreshold, int allowedCascading, boolean globalDict, boolean enableZstd, MemorySize globalDictMaxRetainedBytes, Map<EditionFamily, Edition> editions)`.
Record: `(boolean enableZoneMaps, double compressionRatioThreshold, int allowedCascading, boolean globalDict, boolean enableZstd, MemorySize globalDictMaxRetainedBytes, Map<EditionFamily, Edition> editions)`.

| Factory | Defaults |
|---------------------------------|---------------------------------------------------------------------------------------------------|
| `WriteOptions.defaults()` | `chunkSize=65_536`, `enableZoneMaps=true`, `compressionRatioThreshold=0.90`, `allowedCascading=0`, `globalDict=true`, `enableZstd=false`, `globalDictMaxRetainedBytes=MemorySize.ofGiB(2)`, `editions={CORE: Editions.CORE_2026_08_0}` |
| `WriteOptions.defaults()` | `enableZoneMaps=true`, `compressionRatioThreshold=0.90`, `allowedCascading=0`, `globalDict=true`, `enableZstd=false`, `globalDictMaxRetainedBytes=MemorySize.ofGiB(2)`, `editions={CORE: Editions.CORE_2026_08_0}` |
| `WriteOptions.cascading(depth)` | Same defaults, `allowedCascading=depth` |

| Method | Notes |
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -615,7 +615,7 @@ void javaWriter_jniReader_zoneMapped_multipleZones(@TempDir Path tmp) throws IOE
// zone-map with one zone per chunk. The Rust reader must parse that layout and still
// return every value (zones are a transparent pruning aux).
Path file = tmp.resolve("java_zoned.vtx");
WriteOptions zoneMapped = new WriteOptions(4, true, 0.90, 0, true, false, MemorySize.ofMiB(256), Map.of());
WriteOptions zoneMapped = new WriteOptions(true, 0.90, 0, true, false, MemorySize.ofMiB(256), Map.of());
long[] ids = new long[20];
double[] vals = new double[20];
for (int i = 0; i < 20; i++) {
Expand Down Expand Up @@ -649,7 +649,7 @@ void javaWriter_jniReader_zoneMapped_filteredScan_doesNotAbortAndKeepsEveryMatch
// of it (an i64 zone field sliced at row 2 is byte 16), aborting the whole JVM. Chunks are
// full-size here so zones line up with the stride Rust assumes and only alignment is tested.
Path file = tmp.resolve("java_zoned_filtered.vtx");
WriteOptions zoneMapped = new WriteOptions(4, true, 0.90, 0, true, false, MemorySize.ofMiB(256), Map.of());
WriteOptions zoneMapped = new WriteOptions(true, 0.90, 0, true, false, MemorySize.ofMiB(256), Map.of());
try (var ch = FileChannel.open(file, StandardOpenOption.CREATE, StandardOpenOption.WRITE);
var sut = VortexWriter.create(ch, SCHEMA, zoneMapped)) {
long next = 0;
Expand Down Expand Up @@ -724,7 +724,7 @@ void javaWriter_jniReader_zoneMapped_allNullChunkStillRoundTrips(@TempDir Path t
Path file = tmp.resolve("java_zoned_null_chunk.vtx");
DType.Struct schema = new DType.Struct(
List.of(ColumnName.of("v")), List.of(new DType.Primitive(PType.I64, true)), false);
WriteOptions zoneMapped = new WriteOptions(4, true, 0.90, 0, false, false, MemorySize.ofMiB(256), Map.of());
WriteOptions zoneMapped = new WriteOptions(true, 0.90, 0, false, false, MemorySize.ofMiB(256), Map.of());
Long[] data = {
0L, 1L, 2L, 3L,
null, null, null, null,
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -195,7 +195,7 @@ private static void writeFixture(Path file) throws IOException {
.field("val", DType.I64)
.build();
// enableZoneMaps=true emits the per-chunk min/max/sum/null-count the interior-zone fold reads.
WriteOptions opts = new WriteOptions(CHUNK_SIZE, true, 0.90, 0, true, false, MemorySize.ofMiB(256), Map.of());
WriteOptions opts = new WriteOptions(true, 0.90, 0, true, false, MemorySize.ofMiB(256), Map.of());
java.util.Random rng = new java.util.Random(SEED);
try (FileChannel ch = FileChannel.open(file,
StandardOpenOption.CREATE, StandardOpenOption.WRITE, StandardOpenOption.TRUNCATE_EXISTING);
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,6 @@

/// Tuning knobs for the Vortex writer.
///
/// @param chunkSize target row count per chunk (default 65 536)
/// @param enableZoneMaps write per-chunk min/max statistics for zone-map pruning
/// @param compressionRatioThreshold minimum compression ratio for an encoding to be accepted (0–1)
/// @param allowedCascading maximum recursive cascade depth; 0 = no cascading
Expand Down Expand Up @@ -49,7 +48,6 @@
/// `unstable`-family encoding is reached by enabling its edition explicitly
/// via [#withEdition(Edition)], not by opting out of the guard altogether.
public record WriteOptions(
int chunkSize,
boolean enableZoneMaps,
double compressionRatioThreshold,
int allowedCascading,
Expand Down Expand Up @@ -94,7 +92,7 @@ public record WriteOptions(
///
/// @return default `WriteOptions`
public static WriteOptions defaults() {
return new WriteOptions(65_536, true, 0.90, 0, true, false, DEFAULT_GLOBAL_DICT_MAX_RETAINED_BYTES,
return new WriteOptions(true, 0.90, 0, true, false, DEFAULT_GLOBAL_DICT_MAX_RETAINED_BYTES,
DEFAULT_EDITIONS);
}

Expand All @@ -104,7 +102,7 @@ public static WriteOptions defaults() {
/// @param depth maximum cascade depth
/// @return `WriteOptions` with cascading enabled at the given depth
public static WriteOptions cascading(int depth) {
return new WriteOptions(65_536, true, 0.90, depth, true, false, DEFAULT_GLOBAL_DICT_MAX_RETAINED_BYTES,
return new WriteOptions(true, 0.90, depth, true, false, DEFAULT_GLOBAL_DICT_MAX_RETAINED_BYTES,
DEFAULT_EDITIONS);
}

Expand All @@ -113,7 +111,7 @@ public static WriteOptions cascading(int depth) {
/// @param enabled `true` to write per-chunk min/max/sum statistics for zone-map pruning
/// @return a new `WriteOptions` with the zone-map flag updated
public WriteOptions withZoneMaps(boolean enabled) {
return new WriteOptions(chunkSize, enabled, compressionRatioThreshold, allowedCascading, globalDict, enableZstd,
return new WriteOptions(enabled, compressionRatioThreshold, allowedCascading, globalDict, enableZstd,
globalDictMaxRetainedBytes, editions);
}

Expand All @@ -122,7 +120,7 @@ public WriteOptions withZoneMaps(boolean enabled) {
/// @param enabled `true` to enable global dictionary encoding across chunks
/// @return a new `WriteOptions` with the global dict flag updated
public WriteOptions withGlobalDict(boolean enabled) {
return new WriteOptions(chunkSize, enableZoneMaps, compressionRatioThreshold, allowedCascading, enabled, enableZstd,
return new WriteOptions(enableZoneMaps, compressionRatioThreshold, allowedCascading, enabled, enableZstd,
globalDictMaxRetainedBytes, editions);
}

Expand All @@ -143,7 +141,7 @@ public WriteOptions withGlobalDict(boolean enabled) {
/// @return a new `WriteOptions` with the Zstd flag updated
/// @throws IllegalArgumentException if `enabled` is `true` and `allowedCascading()` is `0`
public WriteOptions withZstd(boolean enabled) {
return new WriteOptions(chunkSize, enableZoneMaps, compressionRatioThreshold, allowedCascading, globalDict, enabled,
return new WriteOptions(enableZoneMaps, compressionRatioThreshold, allowedCascading, globalDict, enabled,
globalDictMaxRetainedBytes, editions);
}

Expand All @@ -158,7 +156,7 @@ public WriteOptions withZstd(boolean enabled) {
/// @param budget aggregate retention budget for buffered global-dict candidate columns
/// @return a new `WriteOptions` with the global-dict retention budget updated
public WriteOptions withGlobalDictMaxRetainedBytes(MemorySize budget) {
return new WriteOptions(chunkSize, enableZoneMaps, compressionRatioThreshold, allowedCascading, globalDict,
return new WriteOptions(enableZoneMaps, compressionRatioThreshold, allowedCascading, globalDict,
enableZstd, budget, editions);
}

Expand All @@ -177,7 +175,7 @@ public WriteOptions withGlobalDictMaxRetainedBytes(MemorySize budget) {
public WriteOptions withEdition(Edition edition) {
Map<EditionFamily, Edition> updated = new HashMap<>(editions);
updated.put(edition.id().family(), edition);
return new WriteOptions(chunkSize, enableZoneMaps, compressionRatioThreshold, allowedCascading, globalDict,
return new WriteOptions(enableZoneMaps, compressionRatioThreshold, allowedCascading, globalDict,
enableZstd, globalDictMaxRetainedBytes, updated);
}
}
Original file line number Diff line number Diff line change
Expand Up @@ -34,7 +34,7 @@ class NullCountPruningTest {
// sizes and null patterns: 3 rows / 0 nulls, 2 rows / 1 null, 4 rows / all null.
private Path write() throws IOException {
Path file = tmp.resolve("nulls.vtx");
WriteOptions opts = new WriteOptions(1024, true, 0.90, 0, false, false, MemorySize.ofMiB(256), Map.of());
WriteOptions opts = new WriteOptions(true, 0.90, 0, false, false, MemorySize.ofMiB(256), Map.of());
try (var ch = FileChannel.open(file, StandardOpenOption.CREATE, StandardOpenOption.WRITE);
var sut = VortexWriter.create(ch, SCHEMA, opts)) {
sut.writeChunk(Map.of(ColumnName.of("v"), new NullableData(
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -61,7 +61,7 @@ void writeSegments_are64ByteAligned(@TempDir Path tmp) throws IOException {
// VortexWriter pads before each segment so every buffer starts 64-aligned (Arrow-compatible);
// a broken pad — wrong modulus arithmetic or a skipped writePadding — leaves a segment offset
// off a 64-byte boundary.
WriteOptions opts = new WriteOptions(3, false, 0.90, 0, false, false, MemorySize.ofMiB(256), Map.of());
WriteOptions opts = new WriteOptions(false, 0.90, 0, false, false, MemorySize.ofMiB(256), Map.of());
Path file = tmp.resolve("aligned.vtx");
try (var ch = FileChannel.open(file, StandardOpenOption.CREATE, StandardOpenOption.WRITE);
var sut = VortexWriter.create(ch, SCHEMA, opts)) {
Expand Down
Loading
Loading