pub const CLASSIFIER_VERSION: u32 = 7;Expand description
Bumped whenever a build starts classifying something differently.
A cached kind is the only thing a remote row has to go on — it was never stat’ed, and
classifying it means reading it — so it is restored rather than re-derived. That makes
it a way for an answer this build would not give to come back: a Delta root measured
before lake tables were recognized was recorded as multifile, and restoring that
opens it as one table, which is the whole of #237 read back off disk.
So the kind is restored only when the build that wrote it classified the way this one does. Everything else in the record — rows, columns, cost — is a measurement rather than a judgement, and survives.
7: a file with no extension is a SQLite database when its first bytes say so.
6: on a local disk, a file with no extension is data when its first bytes carry a
Parquet, Arrow, Avro or ORC signature, so a directory of Spark part files a 5 called
dir is one dataset. Unidentified ones are unnamed rather than not_read.
5: a directory of CSV or NDJSON is judged by the names at the front of its files, the
way a directory of Parquet is judged by its footers — so one a 4 called multi on its
filenames alone may be a place to look inside. A cached kind is restored without
looking again, so a record written by 4 would keep the answer this build exists to
correct (#275 follow-up).
4: a directory’s row carries what one listing of it found, beside its kind, and the two are restored together — a record written by 3 carries the kind and not the count, and a row given a kind from the cache is never looked into again (#275, phase 2).
3: one listing instead of a probe of the first eight entries, formats instead of
extension strings, and one bookkeeping predicate. A directory of .arrow beside
.ipc was dir and is now one dataset; a directory whose ninth entry decided it was
answered by whatever the filesystem returned first (#275, phase 1).