Skip to main content

lenient_scan

Function lenient_scan 

Source
pub fn lenient_scan(
    paths: &[String],
    schema: Arc<Schema>,
    cloud_options: Option<CloudOptions>,
    drift: Option<&ScanDrift>,
    as_text: &[PlSmallStr],
) -> PolarsResult<LazyFrame>
Expand description

A scan of paths into schema that reads files written at different times: columns and nested fields a file lacks are filled with nulls, ones it has beyond the schema are ignored, and integers, floats and datetime units widen. Polars’ scan_parquet offers only the first of those, and a Bitcoin transactions file from 2015 fails against the 2026 schema without the rest.

When drift is given, every row carries its position in the dataset in DRIFT_COLUMN, which is what lets a cell be traced to its file and a null told from a column that file never had.

The scan splits only where it has to: a column a file stores in another type must be left out of that file’s read, so consecutive files omitting the same columns are one scan and the scans are concatenated in file order. A file merely missing a column needs no split — MissingColumnsPolicy::Insert already reads it as null — so the common case stays a single scan however many files disagree.

as_text names columns to read as text from every file instead of leaving them out of the ones that disagree. Such a column is read at the type each file disagrees in and cast to text after, which is the only way to see the values a conflict hides: Polars’ scan can widen an integer and change a datetime’s unit, but it has no policy for reading a number as a string, and no way to hand back a column it was not told the type of. So the split is finer here — a run is a stretch of files that agree on the type of every as_text column as well as on what they are missing.

A file whose type merely widens into the column’s is read at the column’s type, not its own, because only conflicting types are recorded: an integer in a column read as a float reads as 7.0. Nothing is hidden by that — a widened value was always on screen — but it is not the file’s own spelling. The same goes for a file whose footer could not be read, and for one outside the sample on a dataset too large to read every footer: datui does not know what those hold, so it asks for the column’s type and they are no better off than before.

A column can_read_as_text refuses, or that the schema does not have, is dropped from as_text and read as it was.