# folio
Section-scoped semantic search over a markdown corpus. The index holds
references, never bodies.
A *folio* is a numbered leaf reference — a pointer to where text sits, not the
text. That is what this index stores: for every markdown heading section, a path,
a line range, the heading trail that names it, and the document's frontmatter.
Ask it a question and it ranks the sections you should read. Reading them is
your next step, and it reads the file, so an index that has fallen behind costs
you a wasted candidate rather than a wrong quotation — and where that would cost
more than a candidate, a query notices and catches the index up first.
## Why it exists
`grep` fails on the words you did not think of. Searching a corpus for
`independent` finds nothing when every document writes `independence`; searching
for `commit` returns six directories and buries the one that answers the
question among five plausible decoys. Picking wrong there is how an agent reads
the wrong document and misunderstands a project.
Semantic ranking fixes that much. What it does not fix — and makes worse — is
precedence. A superseded statement and the statement that replaced it are
semantically alike, so a vector index surfaces them side by side: in a two-line
fixture the retired definition of revenue ranks at 0.763 directly under the live
one at 0.842. Corpora that record their own lifecycle already carry the answer in
frontmatter, as a `status` or a `supersedes`. folio indexes those fields as
filters, so a query can say which of two similar passages still holds.
## Install
```sh
cargo install --path .
```
folio calls an OpenAI-compatible embeddings endpoint and contains no inference
code, so the model is yours to choose. Any server exposing `/v1/embeddings`
works. One that has been measured:
```sh
llama-server -hf keisuke-miyako/gte-modernbert-base-gguf \
--hf-file gte-modernbert-base-Q8_0.gguf \
--embeddings --pooling cls -c 8192 -b 8192 -ub 8192 --port 8080
```
Two flags there are not optional. `--pooling cls` is what this model wants, and
the default is wrong for it — the Qwen3-Embedding family wants `--pooling last`
instead. And `-b`/`-ub` must be at least your longest section: an encoder needs
its whole input in one physical batch, so at the default 512 a longer section
comes back as an HTTP 500 rather than a truncated vector.
## Use
```sh
export FOLIO_ENDPOINT=http://127.0.0.1:8080/v1/embeddings
folio index # every .md under the working directory
folio query "does unfinished work count as a failure"
folio status
```
Output names sections, with the heading trail beneath each:
```
#1 0.693 s21-does-an-incomplete-session-count-as-a-failure/README.md:6-6
Does an incomplete session count as a failure?
```
### Filtering on frontmatter
Frontmatter is flattened to dotted keys and stored as written — no field is
built in, so any producer's schema is queryable.
```sh
folio query "who owns a decision's status" --where status=live
folio query "the workspace layout" --where supersedes
folio query "current guidance" --where type!=deprecated
```
`key=value` matches, and reads as membership when the value is a list, so
`--where tags=alpha` works. `key` alone tests presence. `key!=value` also passes
when the key is absent, so a filter never silently drops the documents nobody has
annotated yet.
### Dropping what something else replaced
A `--where` predicate reads one record. Precedence does not live in one record:
the pointer sits on the successor, and the record you want gone is the one it
points at. That needs a join.
```sh
folio query "when is revenue recognised" --exclude-pointed-by supersedes
```
Any section whose `id` appears in any other section's `supersedes` stops being a
candidate, and the count of what went is printed so the drop is never silent.
Neither key is built in — `--exclude-pointed-by` names the pointer and
`--identity` names the key holding a record's own identity, which defaults to
`id` only because most schemas spell it that way.
The values are collected from the whole index rather than from what the other
filters leave, because a superseded record is superseded whether or not the
record that replaced it also answers this query.
Defaults belong to the query, not the index. The
[Open Knowledge Format](https://github.com/GoogleCloudPlatform/knowledge-catalog/blob/main/okf/SPEC.md)
reads an absent `status` as `stable`; folio stores what the file says and leaves
that reading to you.
### Keeping it current
`folio index` lists every file and re-embeds only the ones whose contents
changed. Listing is what makes it cheap — a file whose length and modification
time are what the index recorded is never opened — so on 14,616 files a re-index
with nothing changed is 0.27 s, and one changed file is 0.47 s.
There is no watcher and no daemon. A query checks itself instead. Before
returning a row it stats the file behind it, and if that file has moved since it
was indexed, folio brings the index up to date and answers again:
```
$ folio query "when is revenue recognised"
(1 of the files behind this result had changed; 1 file(s) re-embedded before answering)
#1 0.812 finance/revenue.md:42-57
Recognition > Timing
```
This is the one place a stale index does harm rather than waste. A line range is
not a candidate; it is an instruction to read lines 40 to 55, and once two lines
are inserted above that section, following the instruction reads the wrong
lines.
It costs one stat per returned row, which is nothing: over 119,359 sections a
query with nothing changed is 0.12 s either way. It checks the rows it returns
and not the corpus — a file that changed without surfacing still costs you a
candidate, which is the trade this index makes everywhere else too. Walking the
whole tree to close that would cost 0.27 s on every query, twice what the query
costs.
`--no-refresh` turns off the writing, not the checking. A row whose file has
moved is still marked, because the caller who asked to be answered from the
index as it stands is the one who most needs to know where it does not:
```
$ folio query "when is revenue recognised" --no-refresh
#1 0.812 finance/revenue.md:40-55 (stale)
Recognition > Timing
```
A refreshing query takes the same write lock `folio index` takes, and declines
rather than waits when another folio holds it, since that one is already
producing an index at least as fresh. It refreshes at most once: a row still
stale afterwards means the files are moving while folio reads them, and saying
so beats looping.
### From an agent
`SKILL.md` is the same surface written for an agent to act on: the commands, the
filters, and when to reach for `rg` instead. Point a skill-loading agent at it,
or copy it into wherever that agent keeps skills.
## What it does not do
- **Exact matching.** Use `rg`. It is exhaustive and folio is not, and a lexical
route mixed into the ranking measurably buried correct answers.
- **Code structure.** folio indexes prose sections, not symbols or call graphs.
- **Anything but markdown.** PDF, office documents and source files are skipped.
## Sizing
One `f32` matrix, mapped and scanned end to end. No approximate index and no
recall parameter: a query over 548 sections is 15 ms and one over 119,359 is
132 ms, of which about 102 ms is the scan itself. The rest is one embedding
round trip and reading the list of rows that are still live.
Three quarters of a query is therefore the exhaustive arithmetic — which is the
part an approximate index replaces, and it would replace 102 ms with a graph to
build, a recall parameter to defend, and more bytes to load beside the matrix.
Pointer precision is set by your headings, not by the model. A file whose long
sections carry `###` subheadings returns 12-line ranges; the same content under
one `##` returns a 133-line range. If a result feels too coarse, add a heading
before you change models.
## License
MIT. See [LICENSE](LICENSE).