Conversation
Move the input infos out of AnalysisDataProcessorBuilder so that it can be reused outside of it.
Prepare options and services first, then collect all the slicing cache entries
Drop the duplicated binding check isMissing() on the binding key for all Preslice kinds rename missingOptionalPreslice to missingPreslice
Introduce SliceInfo, which defines the layout of the slice-info tables and builds them from the sliced table ArrowTableSlicingCache keeps views into them The unsorted slice info now uses a counting pass and a single list array instead of per-group vectors
Add inputForEntry and matcherForEntry, which makes the InputSpec of the slice-info table for a slicing cache entry
Add the ArrowTableSlicer algorithm plugin, which builds the requested slice-info tables from the slice sources The slice inputs are grouped based on the provider of the sources, so that multiple slicers can be injected right after the providers to avoid loops in topology
For grouping process functions add the slice-info input for each argument related to the grouping table by an index with the control option of the process function and follow the origin replacement of the sliced table.
Non-optional Preslice declarations no longer fail when the column is missing A declaration that cannot be used is skipped with a warning: - the table is not an input and does not have the column - the table is not an input (non-optional only) - the table does not have the column (non-optional only)
The slicing cache picks up the slice-info tables produced by the internal slicer devices
Collaborator
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This is a major development, aimed to reduce the overhead of recomputing the table slices in each analysis task. In a given realistic workflow, several tasks will require tracks sliced by collision id, for example, thus providing it from an extra device once, instead of calculating everywhere, will reduce CPU and memory consumption for the large workflows.
For this purpose, the slicing information is now created and stored as an Arrow table in shared memory, that is then sent to consumers. The original slicing cache service now collects and views these tables instead of computing the slices itself, otherwise the interface is unchanged and the rest of analysis framework still queries that cache as before.
Preslice/PresliceOptional had to be reworked, previously, an ineffective Preslice, requiring a table that is not present in the task's input, would be silently skipped. I do not break this behavior, but such declaration will emit a runtime warning. Using PresliceOptional will silence the warning.
To avoid topology loops, the slicer device multiplies itself when its source tables are provided by a device that itself is a slice consumer. The slice requests are grouped by their source provider and corresponding devices are then inserted in the pre-sorted workflow right after that provider.
Since it is a major development, it requires testing with realistic hyperloop workflows.
@ktf, @dsekihat