Expand description
The schema of a dataset made of many files, which drift over time: every column any file has, from the footers the row count already reads, with one type per column:
- equal types, or types that widen losslessly (
Int32intoInt64,Float32intoFloat64,msintons), become the wider one; - types that do not, become the one most rows have. The column is not read from
the other files, which is why
DatasetSchema::omittednames them per file.
Column order is the newest file’s columns in its own order, then columns only older files have, in the order they first appear.
Structs§
- Column
Drift - What the footers said about one column.
- Counted
- What a count found.
- Dataset
Schema - One dataset’s schema, and what deciding it revealed.
- Disagreement
- How a directory’s files differed, as read: a column some files lack (ordinary drift), and a column held in two types (forcing the wider). Either, both or neither.
- Drift
Group - What a file is missing relative to the schema; files missing the same share a group, so a row carries only its group.
- File
Footer - What one file’s footer said, short of the data: one type for every location and pass.
- Footer
Count - A dataset’s row count from its footers, each read once: starting from those the open read and reading the rest. Once all are in, the set is handed back once for the shape cache, so a reopen reads none.
- Footer
Progress - How far a dataset’s footer pass has got, for the loading screen: a climbing count on a directory of thousands of files. Atomic: written only by reading threads, read by the render.
- Listing
- A listing counting found objects against a
FooterProgress, which stops reporting however it ends. - Pass
- A footer pass in progress. Says it has finished when dropped, panic or no panic.
- ReadAs
- The reader settings that change what a sample of a file’s columns returns, taken
from the open’s actual options (
from_args_and_configalways fillsinfer_schema_lengthandparse_strings, so guessing from “did the user set anything” never works).Defaultmatches Polars’ readers, as home opens use. - RowEstimate
- A row count estimated from a sample of footers (mean rows per file read times the
file count), shown as
~4.12B rows (est.)until counted. - Sampled
- What a spread of a directory’s files says about whether they are one table.
- Scan
Drift - How files differ, as the scan needs it: what each lacks and where its rows begin, keyed by the path or URL the scan uses.
- Skipped
Files - Files a dataset’s listing walked past. A file beside the data (a
.csvbeside Parquet) may have been meant as data and is worth saying; one elsewhere (Delta’s_delta_log/, Hudi’s.hoodie/, Iceberg’smetadata/, folder markers) is infrastructure: a directory with no Parquet is nobody’s table. Counts only what the walk saw.
Enums§
- Column
Range - Where a column not in every file sits among the dataset’s partitions: in one partition, or from some point onward (a field added to a feed).
- Schema
Origin - Where a dataset’s schema came from, shown in Info’s Schema tab so a missing column traces to the files looked at.
Constants§
- COUNT_
AT_ ONCE - Footers a count reads at once: each is a round trip on a store.
- DRIFT_
COLUMN - The column the scan writes each row’s dataset position into, tracing a cell to its file and whether that file had the column. Never shown, filtered, sorted or exported: the display projects only the column order.
- ESTIMATE_
SAMPLE - Footers sampled for the first estimate: a mean within a few percent for alike files, in one wave or a few.
- FOOTERS_
AT_ ONCE - Footers read at once: one wave. Reads wait on round trips rather than compute, so this exceeds the core count. Larger datasets open from their ends and read the rest behind.
- MAX_
FOOTER_ READS - Footers read before a dataset opens; past this a spread sample stands in. A fixed
threshold documented in
docs/user-guide/large-datasets.md.
Functions§
- can_
read_ as_ text - Whether a column of this type can be shown as text. Polars cannot cast durations
or lists to strings, and binary fails on non-UTF-8, and a failed cast fails the
whole scan, so this is decided by type before offering.
types_the_cast_agrees_with_are_exactly_the_ones_offeredchecks it against Polars. - column_
bytes_ per_ row - Uncompressed bytes per row of each column over the footers read; files without a column count its rows as zero (they read null).
- column_
schema_ of - The column names one data file holds, read as cheaply as its format allows: the
header or first object’s keys, via Polars’ inference as the open will. For
formats without footers (
is_nested’s evidence).Nonewhen the schema needs the whole file (JSON documents) or the file will not parse: no evidence, so the directory keeps its name-based kind. - ends_of
- The first and last files, which a dataset opens from while the rest of its footers are read; ascending, one index for one file. Last by name, not date (unpadded partition values sort oddly): a heuristic, with the rest arriving later.
- fits
- Whether a file storing
fromreads into atocolumn losslessly, matchinglenient_scan’s cast policy; otherwise the column is left unread in that file. - footers_
from_ cache - The cached footers as a fresh pass would read them, with
file_bytesfrom the listing.Nonefor an unreadable footer, as the pass reports it. Refused whole when inconsistent with itself or the listing. - footers_
to_ cache - Footers as the shape cache keeps them: schemas tabled and referenced by index, since ten thousand files usually share one schema.
- footers_
to_ read - Which of a dataset’s
filesfooters to read: all of them, or — pastMAX_FOOTER_READS— a sample spread evenly across them, always including the first and the newest. Indices are ascending. - is_
nested - Whether every file’s columns are contained in the widest file’s: the shape schema
evolution produces, read cleanly as a union. No score or threshold: a file either
brings a column no other has or not (overlap ratios cannot tell drift from
unrelated tables sharing a key). Failing is not a refusal: the row goes inside,
and
(all files)still opens the union. Fewer than two files, or files without columns, are one table. - lenient_
scan - A scan of
pathsintoschemaacross files written at different times: missing columns and fields fill with nulls, extras are ignored, and integers, floats and datetime units widen (scan_parquetalone only fills). - parquet_
column_ bytes - Uncompressed bytes of each column in a footer, summed over row groups and nested leaves: the only way to know a binary or string column’s size short of reading it.
- partition_
columns_ of_ key - Partition column names from one file’s path, in path order, each key once.
- partitions_
of_ listing - A listed dataset’s partition columns and typing values, from its first and newest
files’
/-separated keys (the newest names the columns, as later keys appear there). Shared by local and cloud listings so a tree reads the same. - random_
sample nrandom indices belowfilesfromseed, ascending; all whenfiles <= n. The same seed draws the same sample.- readable_
paths - The paths whose footers were read. An unreadable footer is a file Polars cannot read, and in the scan it would fail the first page for the whole dataset; it is still counted for the note.
- sample_
files - Read a spread of
files(the ends and the middle, since sorted names group each table’s files) and say what they are. Bounded reads whatever the size: this runs in listing passes and on the way into an open. - text_
schema schemawith everyas_textcolumn as text, for callers describing what they hold (lenient_scandoes not need it). Columns keep their places.- top_
level_ columns - The top-level column names of Parquet leaf paths, in order, deduplicated. Leaf
wrappers (
inputs.list.element.address) differ by writer, so comparisons use the columns a reader sees. - union_
file_ schemas - Fold every file’s footer into one schema.
filesis in scan order (the last readable is the newest);Noneis an unreadable footer. - union_
sampled - The schema of a dataset of
filesfiles whose footers at indicesreadwere fetched. The union is over the footers read;DatasetSchema::filesstays that count (datui knows nothing of unopened files). Onlyomitted,file_groupandunreadable, indexed by the scan, are spread to the full length. - widen
- The narrowest type
aandbboth read into losslessly, if any. - with_
partition_ columns - Put the partition columns, typed from their directory names, ahead of the file’s own.