Skip to main content

union_sampled

Function union_sampled 

Source
pub fn union_sampled(
    files: usize,
    read: &[usize],
    footers: &[Option<FileFooter>],
) -> DatasetSchema
Expand description

The schema of a dataset of files files whose footers at the indices read were fetched, in that order.

The union is over the footers read; the scan is over every file, so the per-file findings are spread back across the full list. A file not read omits nothing.

DatasetSchema::files stays the number of footers read, not files, and every count taken against it — the notes’ denominator, how many files hold no rows — is therefore over the sample. That is the honest population: datui knows nothing about a file it did not open. Only omitted, file_group and unreadable, which the scan indexes by file, are spread to the full length.