Expand description
The schema of a dataset made of many files.
A dataset written over years is rarely uniform: a vendor adds a column one day, drops another for a week, or writes a number as text for a month. Reading the schema from one file makes those columns appear or vanish depending on which file is picked. This module takes every column any file has, from the footers the row count already reads, and decides one type per column:
- equal types, or types that widen losslessly (
Int32intoInt64,Float32intoFloat64,msintons), become the wider one; - types that do not, become the one most rows have. The column is not read from
the other files, which is why
DatasetSchema::omittednames them per file.
Column order is the newest file’s columns in its own order, then columns only older files have, in the order they first appear.
Structs§
- Column
Drift - What the footers said about one column.
- Counted
- What a count found.
- Dataset
Schema - One dataset’s schema, and what deciding it revealed.
- Disagreement
- How a directory’s files differed, as the read found them.
- Drift
Group - What a file is missing relative to the dataset’s schema. Files that are missing the same things share a group, so a row need only carry its group to know how to draw.
- File
Footer - What one file’s footer said, short of the data: the one footer type, wherever the file is and whichever pass read it.
- Footer
Count - A dataset’s row count from its footers, each read once.
- Footer
Progress - How far a dataset’s footer pass has got, for the loading screen to read.
- Listing
- A listing counting the objects it finds against a
FooterProgress, which stops saying so however it ends. - Pass
- A footer pass in progress. Says it has finished when dropped, panic or no panic.
- ReadAs
- The few reader settings that change what a sample of a file’s columns comes back as.
- RowEstimate
- A dataset’s row count from a sample of its files’ footers: the mean of the files
read, times the files there are. Said as
~4.12B rows (est.)until it is counted. - Sampled
- What a spread of a directory’s files says about whether they are one table.
- Scan
Drift - How a dataset’s files differ, in the form the scan needs: what each file is missing, and where its rows begin in the dataset, by the path or URL the scan names it by.
- Skipped
Files - What a dataset’s listing walked past: files under the directory that are not read.
Enums§
- Column
Range - Where a column that is not in every file sits, in the dataset’s own partitions.
- Schema
Origin - Where a dataset’s schema came from. Shown in the Info panel’s Schema tab, so a column that is missing is traceable to the files that were looked at.
Constants§
- COUNT_
AT_ ONCE - Footers read at once by a count: the exact count of a dataset of many files reads every footer it does not have, and on a store each is a round trip of waiting.
- DRIFT_
COLUMN - The column the scan writes each row’s position in the dataset into, so a cell can be traced back to the file it came from and told whether that file had the column at all. Never shown, filtered, sorted or exported: it is in the buffer the display is sliced from, and the display only ever projects the column order.
- ESTIMATE_
SAMPLE - Footers sampled at random for a dataset’s first row estimate: enough for a mean within a few percent on any dataset whose files are alike, and one wave or a few.
- FOOTERS_
AT_ ONCE - Footers read at once: one wave. A footer read is waiting, not computing — on a network mount or a store it is a few round trips — so this is not the core count. A dataset of more files than this opens from its ends and reads the rest behind.
- MAX_
FOOTER_ READS - Footers read before a dataset opens. Past this many files the reads cost more than
the schema is worth, so a spread sample stands in for the rest. Documented in
docs/user-guide/large-datasets.md; a fixed threshold, not a setting.
Functions§
- can_
read_ as_ text - Whether a column of this type can be shown as text.
- column_
bytes_ per_ row - Uncompressed bytes per row of each column over the footers read. A file without a column counts its rows at nothing, since they read as null there.
- column_
schema_ of - The column names one data file holds, read as cheaply as its format allows.
- ends_of
- The first file and the last, which is what a dataset opens from while the rest of its footers are still being read. Ascending, and one index when there is one file.
- fits
- Whether a file storing
fromcan be read into a column oftowithout loss. Mirrors the cast policy [crate::cloud_hive::lenient_scan] gives Polars; a type that does not fit would fail the scan, so the column is left unread in that file instead. - footers_
from_ cache - The footers a cache kept, as a fresh pass would have read them.
file_bytesis each file’s size from the listing, which the cache does not hold. - footers_
to_ cache - Footers in the form the shape cache keeps them: the schemas gathered into a table and referred to by index, since a dataset of ten thousand files usually has one schema, and writing each file’s columns out in full would make the cache larger than the footers it saves reading.
- footers_
to_ read - Which of a dataset’s
filesfooters to read: all of them, or — pastMAX_FOOTER_READS— a sample spread evenly across them, always including the first and the newest. Indices are ascending. - is_
nested - Whether every file’s columns are contained in the widest file’s.
- lenient_
scan - A scan of
pathsintoschemathat reads files written at different times: columns and nested fields a file lacks are filled with nulls, ones it has beyond the schema are ignored, and integers, floats and datetime units widen. Polars’scan_parquetoffers only the first of those, and a Bitcoin transactions file from 2015 fails against the 2026 schema without the rest. - parquet_
column_ bytes - Uncompressed bytes of each column in a Parquet footer, summed over the row groups and a nested column’s leaves. What a binary or string column holds is known only from here short of reading it.
- partition_
columns_ of_ key - Partition column names from one file’s path, in path order: every
key=valuesegment, each key once. - partitions_
of_ listing - A listed dataset’s partition columns and the values that type them, from its first
and newest files’ keys,
/-separated. The newest names the columns, since a key added later is in it; both give values. Local directories and cloud prefixes derive them here alike, so a tree is the same table from either. - random_
sample nof the indices belowfiles, ascending, drawn at random fromseed: every index when there are no more thann. The same seed draws the same sample.- readable_
paths - The paths whose footers were read, out of the paths given.
- sample_
files - Read a spread of
filesand say what they are. - text_
schema schemawith every column inas_textspelled as text.- top_
level_ columns - The top-level column names in a list of Parquet leaf paths.
- union_
file_ schemas - Fold every file’s footer into one schema.
filesis in scan order, so the last readable entry is the newest file;Noneis a file whose footer could not be read. - union_
sampled - The schema of a dataset of
filesfiles whose footers at the indicesreadwere fetched, in that order. - widen
- The narrowest type both
aandbread into without loss, orNonewhen there is none. - with_
partition_ columns - Put the partition columns, typed from their directory names, ahead of the file’s own.