Skip to main content

Module schema_union

Module schema_union 

Source
Expand description

The schema of a dataset made of many files, which drift over time: every column any file has, from the footers the row count already reads, with one type per column:

  • equal types, or types that widen losslessly (Int32 into Int64, Float32 into Float64, ms into ns), become the wider one;
  • types that do not, become the one most rows have. The column is not read from the other files, which is why DatasetSchema::omitted names them per file.

Column order is the newest file’s columns in its own order, then columns only older files have, in the order they first appear.

Structs§

ColumnDrift
What the footers said about one column.
Counted
What a count found.
DatasetSchema
One dataset’s schema, and what deciding it revealed.
Disagreement
How a directory’s files differed, as read: a column some files lack (ordinary drift), and a column held in two types (forcing the wider). Either, both or neither.
DriftGroup
What a file is missing relative to the schema; files missing the same share a group, so a row carries only its group.
FileFooter
What one file’s footer said, short of the data: one type for every location and pass.
FooterCount
A dataset’s row count from its footers, each read once: starting from those the open read and reading the rest. Once all are in, the set is handed back once for the shape cache, so a reopen reads none.
FooterProgress
How far a dataset’s footer pass has got, for the loading screen: a climbing count on a directory of thousands of files. Atomic: written only by reading threads, read by the render.
Listing
A listing counting found objects against a FooterProgress, which stops reporting however it ends.
Pass
A footer pass in progress. Says it has finished when dropped, panic or no panic.
ReadAs
The reader settings that change what a sample of a file’s columns returns, taken from the open’s actual options (from_args_and_config always fills infer_schema_length and parse_strings, so guessing from “did the user set anything” never works). Default matches Polars’ readers, as home opens use.
RowEstimate
A row count estimated from a sample of footers (mean rows per file read times the file count), shown as ~4.12B rows (est.) until counted.
Sampled
What a spread of a directory’s files says about whether they are one table.
ScanDrift
How files differ, as the scan needs it: what each lacks and where its rows begin, keyed by the path or URL the scan uses.
SkippedFiles
Files a dataset’s listing walked past. A file beside the data (a .csv beside Parquet) may have been meant as data and is worth saying; one elsewhere (Delta’s _delta_log/, Hudi’s .hoodie/, Iceberg’s metadata/, folder markers) is infrastructure: a directory with no Parquet is nobody’s table. Counts only what the walk saw.

Enums§

ColumnRange
Where a column not in every file sits among the dataset’s partitions: in one partition, or from some point onward (a field added to a feed).
SchemaOrigin
Where a dataset’s schema came from, shown in Info’s Schema tab so a missing column traces to the files looked at.

Constants§

COUNT_AT_ONCE
Footers a count reads at once: each is a round trip on a store.
DRIFT_COLUMN
The column the scan writes each row’s dataset position into, tracing a cell to its file and whether that file had the column. Never shown, filtered, sorted or exported: the display projects only the column order.
ESTIMATE_SAMPLE
Footers sampled for the first estimate: a mean within a few percent for alike files, in one wave or a few.
FOOTERS_AT_ONCE
Footers read at once: one wave. Reads wait on round trips rather than compute, so this exceeds the core count. Larger datasets open from their ends and read the rest behind.
MAX_FOOTER_READS
Footers read before a dataset opens; past this a spread sample stands in. A fixed threshold documented in docs/user-guide/large-datasets.md.

Functions§

can_read_as_text
Whether a column of this type can be shown as text. Polars cannot cast durations or lists to strings, and binary fails on non-UTF-8, and a failed cast fails the whole scan, so this is decided by type before offering. types_the_cast_agrees_with_are_exactly_the_ones_offered checks it against Polars.
column_bytes_per_row
Uncompressed bytes per row of each column over the footers read; files without a column count its rows as zero (they read null).
column_schema_of
The column names one data file holds, read as cheaply as its format allows: the header or first object’s keys, via Polars’ inference as the open will. For formats without footers (is_nested’s evidence). None when the schema needs the whole file (JSON documents) or the file will not parse: no evidence, so the directory keeps its name-based kind.
ends_of
The first and last files, which a dataset opens from while the rest of its footers are read; ascending, one index for one file. Last by name, not date (unpadded partition values sort oddly): a heuristic, with the rest arriving later.
fits
Whether a file storing from reads into a to column losslessly, matching lenient_scan’s cast policy; otherwise the column is left unread in that file.
footers_from_cache
The cached footers as a fresh pass would read them, with file_bytes from the listing. None for an unreadable footer, as the pass reports it. Refused whole when inconsistent with itself or the listing.
footers_to_cache
Footers as the shape cache keeps them: schemas tabled and referenced by index, since ten thousand files usually share one schema.
footers_to_read
Which of a dataset’s files footers to read: all of them, or — past MAX_FOOTER_READS — a sample spread evenly across them, always including the first and the newest. Indices are ascending.
is_nested
Whether every file’s columns are contained in the widest file’s: the shape schema evolution produces, read cleanly as a union. No score or threshold: a file either brings a column no other has or not (overlap ratios cannot tell drift from unrelated tables sharing a key). Failing is not a refusal: the row goes inside, and (all files) still opens the union. Fewer than two files, or files without columns, are one table.
lenient_scan
A scan of paths into schema across files written at different times: missing columns and fields fill with nulls, extras are ignored, and integers, floats and datetime units widen (scan_parquet alone only fills).
parquet_column_bytes
Uncompressed bytes of each column in a footer, summed over row groups and nested leaves: the only way to know a binary or string column’s size short of reading it.
partition_columns_of_key
Partition column names from one file’s path, in path order, each key once.
partitions_of_listing
A listed dataset’s partition columns and typing values, from its first and newest files’ /-separated keys (the newest names the columns, as later keys appear there). Shared by local and cloud listings so a tree reads the same.
random_sample
n random indices below files from seed, ascending; all when files <= n. The same seed draws the same sample.
readable_paths
The paths whose footers were read. An unreadable footer is a file Polars cannot read, and in the scan it would fail the first page for the whole dataset; it is still counted for the note.
sample_files
Read a spread of files (the ends and the middle, since sorted names group each table’s files) and say what they are. Bounded reads whatever the size: this runs in listing passes and on the way into an open.
text_schema
schema with every as_text column as text, for callers describing what they hold (lenient_scan does not need it). Columns keep their places.
top_level_columns
The top-level column names of Parquet leaf paths, in order, deduplicated. Leaf wrappers (inputs.list.element.address) differ by writer, so comparisons use the columns a reader sees.
union_file_schemas
Fold every file’s footer into one schema. files is in scan order (the last readable is the newest); None is an unreadable footer.
union_sampled
The schema of a dataset of files files whose footers at indices read were fetched. The union is over the footers read; DatasetSchema::files stays that count (datui knows nothing of unopened files). Only omitted, file_group and unreadable, indexed by the scan, are spread to the full length.
widen
The narrowest type a and b both read into losslessly, if any.
with_partition_columns
Put the partition columns, typed from their directory names, ahead of the file’s own.