pub fn scan_dir_bounded(dir: &Path) -> ScanExpand description
List one directory, doing a bounded amount of work regardless of what is in it.
The cost is one read_dir and a stat per entry, bounded by
MAX_ENTRIES_PER_DIR, and nothing per subdirectory at all.
Nothing here is classified. Telling a hive dataset from a plain directory means
reading the directory, which is a round trip apiece on a share — so no listing pays
for it, however small. Every subdirectory comes back EntryKind::Unknown, which
claims nothing, and is looked into later from the viewport, a batch at a time, by
whoever is actually reading the rows.
That is what makes a row’s label a fact about the row. Classifying the first sixty-four subdirectories and calling every identical one after them a plain directory made it a fact about position instead; classifying them only when a listing is small enough moved the arbitrariness rather than removing it, since two directories holding the same subdirectories would still disagree about what to call them.