Skip to content

DPL Analysis: detach table slicing from the analysis tasks - #15867

Open
aalkin wants to merge 11 commits into
AliceO2Group:devfrom
aalkin:centralized-slicing-experiment
Open

aalkin wants to merge 11 commits into
AliceO2Group:devfrom
aalkin:centralized-slicing-experiment

Conversation

@aalkin

@aalkin aalkin commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

This is a major development, aimed to reduce the overhead of recomputing the table slices in each analysis task. In a given realistic workflow, several tasks will require tracks sliced by collision id, for example, thus providing it from an extra device once, instead of calculating everywhere, will reduce CPU and memory consumption for the large workflows.

For this purpose, the slicing information is now created and stored as an Arrow table in shared memory, that is then sent to consumers. The original slicing cache service now collects and views these tables instead of computing the slices itself, otherwise the interface is unchanged and the rest of analysis framework still queries that cache as before.

Preslice/PresliceOptional had to be reworked, previously, an ineffective Preslice, requiring a table that is not present in the task's input, would be silently skipped. I do not break this behavior, but such declaration will emit a runtime warning. Using PresliceOptional will silence the warning.

To avoid topology loops, the slicer device multiplies itself when its source tables are provided by a device that itself is a slice consumer. The slice requests are grouped by their source provider and corresponding devices are then inserted in the pre-sorted workflow right after that provider.

Since it is a major development, it requires testing with realistic hyperloop workflows.

@ktf, @dsekihat

Move the input infos out of AnalysisDataProcessorBuilder so that it can be reused outside of it.
Prepare options and services first, then collect all the slicing cache
entries
Drop the duplicated binding
check isMissing() on the binding key for all Preslice kinds
rename missingOptionalPreslice to missingPreslice
Introduce SliceInfo, which defines the layout of the slice-info tables and builds them from the sliced table
ArrowTableSlicingCache keeps views into them

The unsorted slice info now uses a counting pass and a single list
array instead of per-group vectors
Add inputForEntry and matcherForEntry, which makes the InputSpec of the slice-info table for a slicing cache entry
Add the ArrowTableSlicer algorithm plugin, which builds the requested slice-info tables from the slice sources

The slice inputs are grouped based on the provider of the sources, so
that multiple slicers can be injected right after the providers to avoid
loops in topology
For grouping process functions add the slice-info input for each argument related to the grouping table by an index with the control option of the process function and follow the origin replacement of the sliced table.
Non-optional Preslice declarations no longer fail when the column is missing
A declaration that cannot be used is skipped with a warning:
- the table is not an input and does not have the column
- the table is not an input (non-optional only)
- the table does not have the column (non-optional only)
The slicing cache picks up the slice-info tables produced by the internal slicer devices
@alibuild

Copy link
Copy Markdown
Collaborator

Error while checking build/O2/fullCI_slc9 for 918b5cc at 2026-09-29 19:38:

No log files found

Full log here.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

2 participants