Expand description
Optional, versioned file-content analysis.
Metadata-only scans allocate none of these structures and open no file content.
An analysis pass reads each file in 64 KiB chunks, schedules its candidates in
bounded batches, classifies code lines in pieces over a bounded window
(CodeAccumulator), and never truncates or size-skips an eligible file. One thing
is retained whole per worker, because an exact answer needs it: the source of a
Markdown file under the words unit, since the CommonMark parser resolves list
tightness, headings, and references across the whole document. That is bounded at
64 MiB (MARKDOWN_EXACT_BYTES, stated with its reason in the analysis module): a
larger Markdown file is counted as plain text and its record says so
(CoverageReason::TextOnly). A file of another type retains at most its 16 KiB
classification prefix and one chunk once the prefix has settled its type.
Structs§
- Analysis
Report - Operational counters from one content-analysis pass.
- Analysis
Request - Settings for one analysis pass.
- Analysis
Set - The set of content analyzers a request enables.
- Analyzer
Coverage - Operational and semantic outcomes for one requested analyzer unit.
- Analyzer
Id - Stable analyzer identity.
- Analyzer
Outcome - One analyzer unit’s explicit coverage and optional successful value.
- Analyzer
Tally - Additive metrics and coverage for one analyzer unit.
- Analyzer
Version - Version of an analyzer’s counting semantics.
- Basic
Accumulator - Stateful fused counter that accepts arbitrary byte chunks.
- Basic
Metrics - Metrics owned by the always-present line analyzer unit.
- Code
Accumulator - Streaming
code-sloc-v1counter for a supported language (analyzer version 3). - Code
Metrics - Metrics owned by the code analyzer unit.
- Content
Cache Load - Result of conditionally restoring one content sidecar.
- Content
Detection - Content-derived classification evidence retained separately from name grouping.
- Content
Index - Optional derived-data tier owned by an index only after analysis is enabled.
- Content
Provenance - Analyzer/rule/options identity attached to cached and reported content.
- Content
Roll Up - Content totals for one directory subtree.
- File
Analysis - Sparse analysis record for one regular file.
- Logical
Word Stats - Additive sufficient statistics for FlexDoc-style logical word volume.
- Metric
Def - Definition of one measured value exposed by content reports.
- Metric
Tally - Additive tally across records of one content-tier identity.
- Metric
Values - Fixed additive metric slots shipped by the first content schema.
- Options
Fingerprint - Fingerprint of semantic analyzer options; operational worker count is excluded.
- Word
Metrics - Metrics owned by the word analyzer unit.
Enums§
- Coverage
Reason - Why a requested file did or did not produce metrics.
- Text
Admission - Outcome of the basic text-admission pass.
Constants§
- CODE_
SLOC - Common-language code/comment/blank analyzer.
- CONTENT_
BASIC - Fused physical-line and raw-word analyzer.
- MARKDOWN_
PROSE - Reader-visible Markdown prose analyzer.
- METRICS
- The single registry of content metric names, owners, and definitions.
- TEXT_
LOGICAL - Plain-text logical word and paragraph analyzer.
Functions§
- analyze_
index - Analyze all regular files selected by
requestwith a fixed-size worker pool. - content_
cache_ path - Derive the analysis sibling from a metadata snapshot path.
- load_
content_ cache - Restore the records of a sidecar whose identity equals
wanted, returning a miss for an absent, corrupt, or foreign sidecar and for one of any other identity. - save_
content_ cache - Persist the content tier’s sparse records as a separately invalidated sidecar, under the identity the tier holds.