pub fn union_sampled(
files: usize,
read: &[usize],
footers: &[Option<FileFooter>],
) -> DatasetSchemaExpand description
The schema of a dataset of files files whose footers at the indices read were
fetched, in that order.
The union is over the footers read; the scan is over every file, so the per-file findings are spread back across the full list. A file not read omits nothing.
DatasetSchema::files stays the number of footers read, not files, and every
count taken against it — the notes’ denominator, how many files hold no rows — is
therefore over the sample. That is the honest population: datui knows nothing about
a file it did not open. Only omitted, file_group and unreadable, which the
scan indexes by file, are spread to the full length.