Skip to main content

Module content

Module content 

Source
Expand description

Optional, versioned file-content analysis.

Metadata-only scans allocate none of these structures and open no file content.

An analysis pass reads each file in 64 KiB chunks, schedules its candidates in bounded batches, classifies code lines in pieces over a bounded window (CodeAccumulator), and never truncates or size-skips an eligible file. One thing is retained whole per worker, because an exact answer needs it: the source of a Markdown file under the words unit, since the CommonMark parser resolves list tightness, headings, and references across the whole document. That is bounded at 64 MiB (MARKDOWN_EXACT_BYTES, stated with its reason in the analysis module): a larger Markdown file is counted as plain text and its record says so (CoverageReason::TextOnly). A file of another type retains at most its 16 KiB classification prefix and one chunk once the prefix has settled its type.

Structs§

AnalysisReport
Operational counters from one content-analysis pass.
AnalysisRequest
Settings for one analysis pass.
AnalysisSet
The set of content analyzers a request enables.
AnalyzerCoverage
Operational and semantic outcomes for one requested analyzer unit.
AnalyzerId
Stable analyzer identity.
AnalyzerOutcome
One analyzer unit’s explicit coverage and optional successful value.
AnalyzerTally
Additive metrics and coverage for one analyzer unit.
AnalyzerVersion
Version of an analyzer’s counting semantics.
BasicAccumulator
Stateful fused counter that accepts arbitrary byte chunks.
BasicMetrics
Metrics owned by the always-present line analyzer unit.
CodeAccumulator
Streaming code-sloc-v1 counter for a supported language (analyzer version 3).
CodeMetrics
Metrics owned by the code analyzer unit.
ContentCacheLoad
Result of conditionally restoring one content sidecar.
ContentDetection
Content-derived classification evidence retained separately from name grouping.
ContentIndex
Optional derived-data tier owned by an index only after analysis is enabled.
ContentProvenance
Analyzer/rule/options identity attached to cached and reported content.
ContentRollUp
Content totals for one directory subtree.
FileAnalysis
Sparse analysis record for one regular file.
LogicalWordStats
Additive sufficient statistics for FlexDoc-style logical word volume.
MetricDef
Definition of one measured value exposed by content reports.
MetricTally
Additive tally across records of one content-tier identity.
MetricValues
Fixed additive metric slots shipped by the first content schema.
OptionsFingerprint
Fingerprint of semantic analyzer options; operational worker count is excluded.
WordMetrics
Metrics owned by the word analyzer unit.

Enums§

CoverageReason
Why a requested file did or did not produce metrics.
TextAdmission
Outcome of the basic text-admission pass.

Constants§

CODE_SLOC
Common-language code/comment/blank analyzer.
CONTENT_BASIC
Fused physical-line and raw-word analyzer.
MARKDOWN_PROSE
Reader-visible Markdown prose analyzer.
METRICS
The single registry of content metric names, owners, and definitions.
TEXT_LOGICAL
Plain-text logical word and paragraph analyzer.

Functions§

analyze_index
Analyze all regular files selected by request with a fixed-size worker pool.
content_cache_path
Derive the analysis sibling from a metadata snapshot path.
load_content_cache
Restore the records of a sidecar whose identity equals wanted, returning a miss for an absent, corrupt, or foreign sidecar and for one of any other identity.
save_content_cache
Persist the content tier’s sparse records as a separately invalidated sidecar, under the identity the tier holds.