Pre-built index over a gene-name vocabulary for fast marker→row matching.
Resolves a query gene in tiers, returning the first matching row:
0. a locus query (chr:start-end) matches only a locus row, by its
locus key with the chromosome case kept, and stops here: loci never
go through the case-insensitive or fuzzy tiers below,
Comma-separated case-insensitive substring filter, parsed once and matched
many times. Used by --select-row-type / --remove-row-type /
--hto-row-type so callers can pass e.g. "gene,peak" to match either
“Gene Expression” or “Peaks”.
Import boundary for feature rows: peak rows get their id and name
rewritten in the colon locus form, whatever spelling the producer used
(chr1-100-200 and chr1_100_200 become chr1:100-200). With feature
types, a row is a peak when its type names peaks or ATAC; without them,
the file is a peak list only when every id reads as an interval. A gene
id that happens to end in two numbers is therefore left alone. Returns
how many rows were rewritten.
The composite name of one row: id{ROW_SEP}name, or the ID alone when
the name is empty or already equals it (e.g. 10x ATAC peaks, where both
features/id and features/name are chr1:1000-2000).
Inverse-document-frequency marker weight ln(C / df): a gene claimed by
all C types gets weight 0 (removed from scoring), a type-exclusive gene
the maximum ln(C).
Inverse of compose_id_name: split a composite id{ROW_SEP}name display
name back into (id, name) on the first ROW_SEP. When there is no
separator (a bare symbol, or an id-only composite where name was empty or
equalled id) both parts are the whole string, so a 10x features.tsv still
gets a non-empty gene name.