You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Is your feature request related to a problem or challenge? Please describe what you are trying to do.
The datafusion package root re-exports about 30 classes plus the common functions. Early on the tendency was to expose everything there, but it no longer tells users where things live. It isn't predictable:
WindowFrame is at the root, but Window is not, so the windows guide imports both from datafusion.expr anyway.
string_literal and str_lit are defined at the root but left out of __all__.
A user can't tell whether to look in datafusion or in the submodule. Each new API also raises the question again, as a review comment on #1763 did.
Actual use is concentrated in a small set. Counting from datafusion import ... across docs/source, examples, and python/tests: SessionContext 237, col 115, lit 64, functions 56, udf 25, Expr 18, SessionConfig 15, then the other UDF decorators. Most other root names appear 0 to 4 times, and several (InsertOp, ParquetColumnOptions, ParquetWriterOptions, DataFrameWriteOptions, Metric, MetricsSet, PhysicalPartitioning, RecordBatchStream, DFSchema, TableProviderFactory, SQLOptions) are essentially never imported from the root.
Describe the solution you'd like
Adopt the rule 55.0.0 already applied to the extension protocols: the root holds what a typical program constructs or calls on the common path, and a type that is only named (an enum, a protocol, an options or result type) is imported from its submodule. The 55.0.0 upgrade guide states it for SessionExtensionComponents, which "stays at the root, because a bundle constructs one rather than merely naming it".
Proposed split:
Keep at the root:
SessionContext, SessionConfig, RuntimeEnvBuilder
DataFrame, Expr, col / column, lit / literal
functions (and the other public submodules)
udf, udaf, udwf, udtf, Accumulator
read_avro, read_csv, read_json, read_parquet
SessionExtensionComponents
Deprecate at the root (each stays importable from the module in parentheses):
string_literal / str_lit: add to __all__, or move them to datafusion.expr alongside Expr.string_literal.
Mechanics:
A module-level __getattr__ in datafusion/__init__.py resolves each deprecated name, emits a DeprecationWarning that names the new import path, and is removed after one release.
Drop the names from __all__ immediately, so from datafusion import * and the API reference stop advertising them.
Add an upgrade-guide section with a before/after table.
Update docs/source, examples, and skills/datafusion_python/SKILL.md to the new import paths in the same change.
Record the rule in the contributor docs, so new APIs follow it.
Describe alternatives you've considered
Keep exporting everything and add new types to the root as they appear (the review suggestion on feat: close upstream coverage gaps for DataFusion 55.1.0 #1763). This is consistent in the narrow sense, but the root keeps growing and stops being a useful entry point.
Remove the names outright, without a deprecation period. Less work, but it breaks existing imports with no warning.
Leave the root as it is and only document the rule for new APIs. This avoids churn, but keeps the inconsistencies listed above indefinitely.
Additional context
The trade-off is discoverability: from datafusion import X is easy to guess and to autocomplete. The deprecation shim keeps old imports working through the transition, and a smaller root is easier to autocomplete, not harder.
Is your feature request related to a problem or challenge? Please describe what you are trying to do.
The
datafusionpackage root re-exports about 30 classes plus the common functions. Early on the tendency was to expose everything there, but it no longer tells users where things live. It isn't predictable:ExplainFormatis at the root, but its siblingsExplainAnalyzeLevelandExplainMetricCategory(added in feat: close upstream coverage gaps for DataFusion 55.1.0 #1763) are not.WindowFrameis at the root, butWindowis not, so the windows guide imports both fromdatafusion.expranyway.string_literalandstr_litare defined at the root but left out of__all__.A user can't tell whether to look in
datafusionor in the submodule. Each new API also raises the question again, as a review comment on #1763 did.Actual use is concentrated in a small set. Counting
from datafusion import ...acrossdocs/source,examples, andpython/tests:SessionContext237,col115,lit64,functions56,udf25,Expr18,SessionConfig15, then the other UDF decorators. Most other root names appear 0 to 4 times, and several (InsertOp,ParquetColumnOptions,ParquetWriterOptions,DataFrameWriteOptions,Metric,MetricsSet,PhysicalPartitioning,RecordBatchStream,DFSchema,TableProviderFactory,SQLOptions) are essentially never imported from the root.Describe the solution you'd like
Adopt the rule 55.0.0 already applied to the extension protocols: the root holds what a typical program constructs or calls on the common path, and a type that is only named (an enum, a protocol, an options or result type) is imported from its submodule. The 55.0.0 upgrade guide states it for
SessionExtensionComponents, which "stays at the root, because a bundle constructs one rather than merely naming it".Proposed split:
Keep at the root:
SessionContext,SessionConfig,RuntimeEnvBuilderDataFrame,Expr,col/column,lit/literalfunctions(and the other public submodules)udf,udaf,udwf,udtf,Accumulatorread_avro,read_csv,read_json,read_parquetSessionExtensionComponentsDeprecate at the root (each stays importable from the module in parentheses):
ExplainFormat,InsertOp,DataFrameWriteOptions,ParquetColumnOptions,ParquetWriterOptions(datafusion.dataframe)ExecutionPlan,LogicalPlan,Metric,MetricsSet,PhysicalPartitioning(datafusion.plan)Catalog,TableProviderFactory,TableProviderFactoryExportable(datafusion.catalog)DFSchema(datafusion.common)WindowFrame(datafusion.expr)SQLOptions(datafusion.context)RecordBatchStream(datafusion.record_batch)TableFunction(datafusion.user_defined)To decide here:
ScalarUDF,AggregateUDF,WindowUDF: usually reached through the decorators, but also constructed directly from FFI capsules.Table(7 root imports),CsvReadOptions,RecordBatch,configure_formatter.string_literal/str_lit: add to__all__, or move them todatafusion.expralongsideExpr.string_literal.Mechanics:
__getattr__indatafusion/__init__.pyresolves each deprecated name, emits aDeprecationWarningthat names the new import path, and is removed after one release.__all__immediately, sofrom datafusion import *and the API reference stop advertising them.docs/source,examples, andskills/datafusion_python/SKILL.mdto the new import paths in the same change.Describe alternatives you've considered
Additional context
The trade-off is discoverability:
from datafusion import Xis easy to guess and to autocomplete. The deprecation shim keeps old imports working through the transition, and a smaller root is easier to autocomplete, not harder.