pub fn lenient_scan(
paths: &[String],
schema: Arc<Schema>,
cloud_options: Option<CloudOptions>,
drift: Option<&ScanDrift>,
as_text: &[PlSmallStr],
) -> PolarsResult<LazyFrame>Expand description
A scan of paths into schema that reads files written at different times:
columns and nested fields a file lacks are filled with nulls, ones it has beyond the
schema are ignored, and integers, floats and datetime units widen. Polars’
scan_parquet offers only the first of those, and a Bitcoin transactions file from
2015 fails against the 2026 schema without the rest.
When drift is given, every row carries its position in the dataset in
DRIFT_COLUMN, which is what lets a cell be traced to its file and a null told
from a column that file never had.
The scan splits only where it has to: a column a file stores in another type must be
left out of that file’s read, so consecutive files omitting the same columns are
one scan and the scans are concatenated in file order. A file merely missing a
column needs no split — MissingColumnsPolicy::Insert already reads it as null —
so the common case stays a single scan however many files disagree.
as_text names columns to read as text from every file instead of leaving them out
of the ones that disagree. Such a column is read at the type each file disagrees
in and cast to text after, which is the only way to see the values a conflict hides:
Polars’ scan can widen an integer and change a datetime’s unit, but it has no policy
for reading a number as a string, and no way to hand back a column it was not told
the type of. So the split is finer here — a run is a stretch of files that agree on
the type of every as_text column as well as on what they are missing.
A file whose type merely widens into the column’s is read at the column’s type,
not its own, because only conflicting types are recorded: an integer in a column
read as a float reads as 7.0. Nothing is hidden by that — a widened value was
always on screen — but it is not the file’s own spelling. The same goes for a file
whose footer could not be read, and for one outside the sample on a dataset too
large to read every footer: datui does not know what those hold, so it asks for the
column’s type and they are no better off than before.
A column can_read_as_text refuses, or that the schema does not have, is dropped
from as_text and read as it was.