Skip to main content

Module schema_union

Module schema_union 

Source
Expand description

The schema of a dataset made of many files.

A dataset written over years is rarely uniform: a vendor adds a column one day, drops another for a week, or writes a number as text for a month. Reading the schema from one file makes those columns appear or vanish depending on which file is picked. This module takes every column any file has, from the footers the row count already reads, and decides one type per column:

  • equal types, or types that widen losslessly (Int32 into Int64, Float32 into Float64, ms into ns), become the wider one;
  • types that do not, become the one most rows have. The column is not read from the other files, which is why DatasetSchema::omitted names them per file.

Column order is the newest file’s columns in its own order, then columns only older files have, in the order they first appear.

Structs§

ColumnDrift
What the footers said about one column.
Counted
What a count found.
DatasetSchema
One dataset’s schema, and what deciding it revealed.
Disagreement
How a directory’s files differed, as the read found them.
DriftGroup
What a file is missing relative to the dataset’s schema. Files that are missing the same things share a group, so a row need only carry its group to know how to draw.
FileFooter
What one file’s footer said, short of the data: the one footer type, wherever the file is and whichever pass read it.
FooterCount
A dataset’s row count from its footers, each read once.
FooterProgress
How far a dataset’s footer pass has got, for the loading screen to read.
Listing
A listing counting the objects it finds against a FooterProgress, which stops saying so however it ends.
Pass
A footer pass in progress. Says it has finished when dropped, panic or no panic.
ReadAs
The few reader settings that change what a sample of a file’s columns comes back as.
RowEstimate
A dataset’s row count from a sample of its files’ footers: the mean of the files read, times the files there are. Said as ~4.12B rows (est.) until it is counted.
Sampled
What a spread of a directory’s files says about whether they are one table.
ScanDrift
How a dataset’s files differ, in the form the scan needs: what each file is missing, and where its rows begin in the dataset, by the path or URL the scan names it by.
SkippedFiles
What a dataset’s listing walked past: files under the directory that are not read.

Enums§

ColumnRange
Where a column that is not in every file sits, in the dataset’s own partitions.
SchemaOrigin
Where a dataset’s schema came from. Shown in the Info panel’s Schema tab, so a column that is missing is traceable to the files that were looked at.

Constants§

COUNT_AT_ONCE
Footers read at once by a count: the exact count of a dataset of many files reads every footer it does not have, and on a store each is a round trip of waiting.
DRIFT_COLUMN
The column the scan writes each row’s position in the dataset into, so a cell can be traced back to the file it came from and told whether that file had the column at all. Never shown, filtered, sorted or exported: it is in the buffer the display is sliced from, and the display only ever projects the column order.
ESTIMATE_SAMPLE
Footers sampled at random for a dataset’s first row estimate: enough for a mean within a few percent on any dataset whose files are alike, and one wave or a few.
FOOTERS_AT_ONCE
Footers read at once: one wave. A footer read is waiting, not computing — on a network mount or a store it is a few round trips — so this is not the core count. A dataset of more files than this opens from its ends and reads the rest behind.
MAX_FOOTER_READS
Footers read before a dataset opens. Past this many files the reads cost more than the schema is worth, so a spread sample stands in for the rest. Documented in docs/user-guide/large-datasets.md; a fixed threshold, not a setting.

Functions§

can_read_as_text
Whether a column of this type can be shown as text.
column_bytes_per_row
Uncompressed bytes per row of each column over the footers read. A file without a column counts its rows at nothing, since they read as null there.
column_schema_of
The column names one data file holds, read as cheaply as its format allows.
ends_of
The first file and the last, which is what a dataset opens from while the rest of its footers are still being read. Ascending, and one index when there is one file.
fits
Whether a file storing from can be read into a column of to without loss. Mirrors the cast policy [crate::cloud_hive::lenient_scan] gives Polars; a type that does not fit would fail the scan, so the column is left unread in that file instead.
footers_from_cache
The footers a cache kept, as a fresh pass would have read them. file_bytes is each file’s size from the listing, which the cache does not hold.
footers_to_cache
Footers in the form the shape cache keeps them: the schemas gathered into a table and referred to by index, since a dataset of ten thousand files usually has one schema, and writing each file’s columns out in full would make the cache larger than the footers it saves reading.
footers_to_read
Which of a dataset’s files footers to read: all of them, or — past MAX_FOOTER_READS — a sample spread evenly across them, always including the first and the newest. Indices are ascending.
is_nested
Whether every file’s columns are contained in the widest file’s.
lenient_scan
A scan of paths into schema that reads files written at different times: columns and nested fields a file lacks are filled with nulls, ones it has beyond the schema are ignored, and integers, floats and datetime units widen. Polars’ scan_parquet offers only the first of those, and a Bitcoin transactions file from 2015 fails against the 2026 schema without the rest.
parquet_column_bytes
Uncompressed bytes of each column in a Parquet footer, summed over the row groups and a nested column’s leaves. What a binary or string column holds is known only from here short of reading it.
partition_columns_of_key
Partition column names from one file’s path, in path order: every key=value segment, each key once.
partitions_of_listing
A listed dataset’s partition columns and the values that type them, from its first and newest files’ keys, /-separated. The newest names the columns, since a key added later is in it; both give values. Local directories and cloud prefixes derive them here alike, so a tree is the same table from either.
random_sample
n of the indices below files, ascending, drawn at random from seed: every index when there are no more than n. The same seed draws the same sample.
readable_paths
The paths whose footers were read, out of the paths given.
sample_files
Read a spread of files and say what they are.
text_schema
schema with every column in as_text spelled as text.
top_level_columns
The top-level column names in a list of Parquet leaf paths.
union_file_schemas
Fold every file’s footer into one schema. files is in scan order, so the last readable entry is the newest file; None is a file whose footer could not be read.
union_sampled
The schema of a dataset of files files whose footers at the indices read were fetched, in that order.
widen
The narrowest type both a and b read into without loss, or None when there is none.
with_partition_columns
Put the partition columns, typed from their directory names, ahead of the file’s own.