[rust] Cache decoded BLOB indexes across readers - #897
Conversation
| if let Some(index) = cache.get(file_path) { | ||
| return Ok(index.clone()); |
There was a problem hiding this comment.
The same path can refer to different files when two FileIO instances use different storage roots. I tested this with two BLOB files of the same size and row count: opening the second reused the first file’s index and returned NULL instead of its actual value. The cache key needs to distinguish the storage namespace too.
There was a problem hiding this comment.
The same path can refer to different files when two
FileIOinstances use different storage roots. I tested this with two BLOB files of the same size and row count: opening the second reused the first file’s index and returnedNULLinstead of its actual value. The cache key needs to distinguish the storage namespace too.
Thanks. I hesitated because Paimon files use UUID names, so two files within one storage namespace are extremely unlikely to have the same path, and PyPaimon also keys by path. I’ll fix the cache key and add a regression test.
Purpose
Opening the same immutable
.blobfile with a new reader currently reloads its footer and index, adding two object-store range reads each time.This caches the decoded index across readers, following the bounded process-local cache used by PyPaimon in apache/paimon#9547.
Changes
Arc; readers and BLOB payloads are not cached.FileIOstorage context.FileIOand isolation across independent storage contexts.Trade-offs
Tests
cargo test -p paimon --lib blob(140 passed)cargo test -p paimon --lib file_io(48 passed)cargo clippy -p paimon --lib --tests -- -D warningscargo fmt --all --checkgit diff --check