# coremlit feature map
The mono-crate restructure collapsed five crates into one crate with
feature-gated modules. This file is the authoritative **rename table** (old
per-crate feature → new flat feature) and the **curated CI feature-combination
list**. It is pinned by the golden test `tests/feature_map.rs`, which parses
`Cargo.toml` and fails if the declared feature set drifts from this table, so a
rename or a dropped feature cannot land silently.
**Two packages.** `coremlit` is the publishable crate. The three third-party
parity oracles — `speaker-oracle` (dia/ort DER), `clap-oracle` (textclap),
`vad-bundled` (the `silero` crate's ONNX stack) — are **no longer coremlit
features**: `dia` and `textclap` are unpublished rev-pinned git sources that
`cargo publish` rejects even behind an optional feature, so those features, their
dependencies and their nine test binaries live in the never-published
`coremlit-parity` package. Their NAMES are unchanged; only the package
you pass to `-p` changed. `align-oracle` did NOT move — it only turns on a
feature of `asry`, a dependency `coremlit` keeps either way, so relocating it
would move code without removing a git dependency. The golden test pins BOTH
manifests and BOTH CI matrices.
## Rename table (old crate feature → new flat feature)
| whisperkit | (crate) | `whisper` | the former unconditional deps (libc, mach2, rand, serde_json, tokenizers, unicode_categories) now ride this feature |
| whisperkit | `nl-recognizer` | `nl-recognizer` | kept; now implies `whisper` |
| whisperkit | `vadkit` | `whisper` + `vad` | the cross-crate feature becomes a composition |
| whisperkit / alignkit / speakerkit | `serde` | `serde` | unified cross-cutting |
| whisperkit / alignkit | `tracing` | `tracing` | unified cross-cutting |
| alignkit | (crate) | `align` | asry's `emissions` seam rides this |
| alignkit | `parity-oracle` | `align-oracle` | asry ONNX aligner oracle (DEV/TEST) |
| speakerkit | (crate) | `speaker` | the CoreML segmentation + embedding backends (module `audio::speaker`) ride this |
| speakerkit | `dia` | `speaker` | diaric's backend-free runtime clustering core (formerly the `dia` offline bridge) |
| speakerkit | `dia-oracle` | `speaker-oracle` | dia's ort DER oracle (DEV/TEST) — now a **`coremlit-parity`** feature |
| vadkit | (crate) | `vad` | the `zuoer` detector core rides this |
| vadkit | dev-dep `silero/bundled` | `vad-bundled` | the `silero` crate's ONNX cross-backend oracle (DEV/TEST) — now a **`coremlit-parity`** feature |
| clapkit | (crate) | `clap` | CLAP-HTSAT dual-tower audio+text encoders (module `embeddings::clap`) ride this; Rust mel front-end + shared `tokenizers`, no ort; the long-audio window geometry + aggregation ride the crates.io `windit` dep |
| clapkit | `parity-oracle` | `clap-oracle` | textclap model-level parity oracle (DEV/TEST) — now a **`coremlit-parity`** feature |
| clapkit | `serde` | `serde` | unified cross-cutting |
## Flat feature set
`default = []` (the bare CoreML runtime core). Additive features:
`whisper`, `nl-recognizer`, `align`, `align-oracle`, `speaker`, `vad`, `clap`,
`granite`, `siglip`, `ced`, `lid`, `identity`, `face`, `serde`, `tracing`, and
the one licence-gated feature `commercial-face-arcface` (⇒ `face`).
`coremlit-parity`'s own `default = []` plus three additive oracle features:
`speaker-oracle` (⇒ `coremlit/speaker`), `clap-oracle` (⇒ `coremlit/clap`),
`vad-bundled` (⇒ `coremlit/vad`). Each rides its own feature so one oracle's
`ort` build is not forced on the other two.
`granite` is not a former per-crate kit but a NEW module (`embeddings::granite`,
the embedkit phase): general text sentence-embeddings on CoreML, first model
`granite-embedding-97m-multilingual-r2`. Its parity oracle is COMMITTED
transformers-fp32 goldens, not a live crate, so it has NO `granite-oracle`
sibling and pulls no `ort` — hence it appears in the rename table below only as a
new-module note, not an old-crate row. Its long-input `embed_long` path pulls the
crates.io `windit` dep (with `windit/text` for content-aware chunking); the
single-text `embed` path does not depend on it.
`siglip` is likewise a NEW module (`embeddings::siglip`): SigLIP 2
(`siglip2-base-patch16-naflex`) dual-tower image+text embeddings on CoreML (a
shared 768-dim joint space). Same posture as `granite` — COMMITTED
transformers-fp32 goldens, so NO `siglip-oracle` sibling and no `ort` — and it
composes with nothing (a single leaf feature). NaFlex resizes natively to a fixed
patch budget, so it is a `windit` non-consumer (no windowing).
`ced` is likewise a NEW module (`audio::ced`): CED (tiny/mini/small/base)
AudioSet sound-event tagging on CoreML — 16 kHz mono waveform in, ranked
predictions over the 527 rated AudioSet classes out (`soundevents-dataset`,
the ort-free data crate; the ort-based `soundevents` crate is never a
dependency). A Rust log-mel
front-end (`rustfft`) feeds one fp16 mel→logits graph; long clips ride the
crates.io `windit` engine (geometry only). Same posture as `granite` —
COMMITTED fp32 goldens, so NO `ced-oracle` sibling and no `ort` — and it
composes with nothing (a single leaf feature).
`lid` is likewise a NEW module (`audio::lid`): spoken-language identification
on CoreML — 16 kHz mono waveform in, ranked languages out over a 107-language
roster. A Rust log-mel front-end (`rustfft`) feeds one fp16
mel→log-probabilities graph, so the feature adds `dep:rustfft` plus the
crates.io `windit` engine (geometry only, no `windit/text`). It is a
**backend-neutral door**: no public name spells the model behind it, and
today's backend is `aufklarer/SpeechBrain-ECAPA-VoxLingua107-21M-CoreML`
(Apache-2.0), an export of `speechbrain/lang-id-voxlingua107-ecapa`. Unlike
`granite` and `siglip` the label roster is COMMITTED in-crate
(`include_bytes!`, 10.7 kB) rather than read from the artifact, so nothing
about the vocabulary is a model gate. ONE prediction is capped at 30.01 s by
the graph's own frame range; a longer clip goes through
`Identifier::identify_long`, whose window geometry rides `windit` and whose
log-probability pooling — a third domain again, neither `ced`'s independent
sigmoids nor windit's unit vectors — is this module's own. It composes with
nothing: a single leaf feature.
`identity` is a NEW module (`audio::identity`): speaker-identity embedding on
CoreML — one fixed 6 s window of 16 kHz mono audio in, one RAW 192-d vector out.
A Rust log-mel front-end feeds one fp16 mel→embedding graph, so the feature adds
`dep:rustfft` and NOTHING else: unlike `ced`/`lid`/`granite` it does not take
`windit`, because it embeds ONE window and any windowing or averaging across
several of them belongs to the caller. Another **backend-neutral door** — no
public name spells the model behind it — today backed by a conversion of
`IDRnD/redimnet`'s ReDimNet-B5 `-vox2-` checkpoint. It is NOT
`audio::speaker`'s embedder and does not replace it: that one is the
diarization lane's batch-3, mask-taking, 256-d WeSpeaker graph, pinned by a DER
gate that reds on any change to it, and this is additive.
`face` is likewise a NEW module (`embeddings::face`): the identity half of the
video face path, as `speaker` is for voices. It carries the ArcFace 5-point
similarity alignment as an explicit, golden-tested step OUTSIDE the embedder,
plus a CoreML embedder whose preprocessing — channel order, scale, bias,
layout — is a per-artifact manifest value rather than a constant at a call
site. Its **one dependency is `sha2`**, already optional in the workspace and
already pulled by `granite` and `siglip`, so it adds no crate to the graph:
`FaceEmbedder::load` hashes the artifact directory it loads, and the digest
rides on every embedding, so a cosine cannot be returned across two different
sets of weights that happen to share a schema (issue #138 §3). Everything else
is dependency-free — the alignment is `f64` arithmetic and the embedder builds
a `MultiArray` the always-compiled runtime core already owns. That `f64` is
also where the alignment departs from InsightFace, whose
`skimage` solve stays in `f32` — a divergence the `align` module doc measures
and shows has no single target to be closed against, because `skimage`'s `f32`
path delegates to a BLAS and two correct builds of it disagree by more than
this crate disagrees with either.
`face` STAYS PLAIN while its artifact does not, and that split is the worked
example this file's `commercial-` section describes. Issue #115's census found
no face-embedding model that clears both the licence bar and the off-angle
accuracy bar; the owner's answer was that CI may use research-only weights
behind a `commercial-` feature *provided coremlit never redistributes them*.
What that produced is `commercial-face-arcface`: a conversion of InsightFace's
`w600k_r50`, research-only at BOTH the weights and the WebFace600K corpus
layer, published to a PRIVATE Hugging Face repository CI fetches with a read
token, gating a manifest module (`embeddings::face::arcface`) of four
constants and nothing else.
Gating `face` itself would have been wrong, and the register would have said so
in the crate's own documentation: two thirds of this feature is encumbered by
nothing — the alignment contains no weights and no photograph, and the embedder
loads a **caller-supplied** path, including an artifact whose licence permits a
product. A feature whose first documented sentence must read "requires a
commercial licence" cannot honestly gate code that requires nothing. So the KIT
is the artifact and the LANE is `face`, which is also why `MODELS_LOCK`'s table
is `kit = "arcface"`: a row's kit must equal the module its licence row's
`loader` names, and the module carrying the `#[cfg]` is `arcface`.
What the artifact buys, all of it measured in `conversion/face/README.md`:
cross-implementation parity against a committed fp32 ONNX reference (≥ 0.99 on
every compute arm, worst observed `1 − cos` 2.2 × 10⁻⁴), the 18 same-person /
135 different-person pairs at InsightFace's own 0.28 / 0.20 operating point, a
four-arm placement table, and 287 faces/s on `CpuAndNeuralEngine`. Its four
gate suites live in `tests/face/` behind `commercial-face-arcface`; the
alignment golden and the load-contract algebra stay hermetic under plain
`face`.
Compositions (pinned by the golden test): `nl-recognizer` → `whisper`;
`align-oracle` → `align`. Across the package boundary, `coremlit-parity`'s
`speaker-oracle` → `coremlit/speaker`, `clap-oracle` → `coremlit/clap`,
`vad-bundled` → `coremlit/vad`. (`granite`, `siglip`, `ced`, `lid`, `identity`
and `face` each compose with nothing — a single leaf feature. `identity` and
`speaker` are a curated CI PAIRING rather than a composition: neither implies
the other, and `Scoring::IdentityCosine`'s `const` assert that its row length
IS `audio::identity::EMBEDDING_DIM` compiles only when both are on, so a
combo that runs them together is what keeps the two dimensions from drifting.)
## Model-licence gating (the `commercial-` prefix)
coremlit is MIT OR Apache-2.0, but products built on it need not be. A model
artifact whose **weights** or whose **training corpus** forbids commercial use
is therefore disqualifying for the shipping path — while still being perfectly
legal for CI to fetch and test against, because this repository redistributes
no weight bytes (see `NOTICE`, "CI DOWNLOADS; IT DOES NOT REDISTRIBUTE").
Any such artifact rides a feature named with the **`commercial-` prefix**, which
is never in `default`. The prefix reads backwards — `commercial-face` looks like
"cleared for commercial use" — so **the feature's comment in `Cargo.toml` must
BEGIN with the warning that a commercial licence is required**. Begins with, not
contains: "This feature no longer requires a commercial license" and "Cleared
for commercial use! This feature requires a commercial license" both contain the
phrase and both say the opposite of it.
One such feature exists: `commercial-face-arcface`, and it is the register's
first row that is research-only at BOTH layers — so the second and third
directions below stopped being tripwires with nothing to bind and now bind it.
**What enabling it does, exactly: it WIRES coremlit's own registered copy of
that artifact.** It compiles the `embeddings::face::arcface` manifest module
(four constants naming the bundle, the path `MODELS_LOCK` stages it to, its I/O
contract and its measured compute placement), it turns on the `arcface`
`MODELS_LOCK` table that stages those bytes for CI, and it builds the four
gated suites in `tests/face/`. Manifested, staged, tested — that wiring, and
only that wiring, is what the register governs.
**What it does not do: it does not stop a caller loading their own copy of
these — or of any other — weights through the plain `face` feature.**
`FaceEmbedder::load` takes a caller-supplied path and a caller-written
`FaceModel`, so a product holding a commercially licensed ArcFace-shaped model
of its own must be able to write that manifest and load it under `face`, and
nothing but a digest would tell those bytes from InsightFace's. The licence of
bytes a caller supplies is between them and whoever published them — for these
weights, InsightFace. That residual is issue #138 §8's and is stated in
`model_licences.rs`'s module doc; the duty the `commercial-` prefix carries is
to TELL you that coremlit's registered artifact needs a licence, and it is
discharged by the feature's name, by its first documented sentence, and by the
fact that this repository hands nobody the bytes.
The mechanism is `coremlit/tests/model_licences.rs`, which holds the licence
table — keyed by
**artifact file + SHA-256, never by repository**, because a repository's own
tag can contradict the terms of bytes it merely re-hosts — and refuses all
three of
- a file `MODELS_LOCK` stages with no licence row, and a row naming a file no
table stages. At FILE granularity: a table with an explicit `files` list is
an exact bijection against the rows, and a globbed table's rows must be
selected by the glob and must cover their whole `.mlmodelc` through a
per-file SHA-256 manifest;
- a research-only artifact WIRED by a feature closure that is not a commercial
opt-in: its loader's `#[cfg(feature = ...)]` sits in **any** non-commercial
feature's closure — not just `default`'s, because
`speaker = ["commercial-face"]` ships it just as surely — or is not
`commercial-` prefixed. *Wired*, not *reachable*: what the check reads is
what this crate manifests, stages and tests, never a path a caller passes to
a public door;
- a `commercial-` feature gating artifacts whose rows are all clear (a gate
left standing after the restriction it protected went away), or that **no
`#[cfg(feature = ...)]` in `src/` names at all** — a feature that gates no
code is a name, not a gate.
Every one of those reads a repository fact rather than the row: the gate comes
from the `#[cfg]` on the module that loads the artifact, the closures from this
manifest, and the selectors from `MODELS_LOCK`. The row's own `gate` field is a
claim, cross-checked against the tree by
`every_rows_gate_matches_the_cfg_that_guards_its_loader`.
Adding a `commercial-` feature therefore means four edits, not one: the
`Cargo.toml` feature (with its first-sentence warning), its row in
`expected_features()` and the table above, **a `#[cfg(feature = ...)]` in `src/`
that actually gates the loader**, and its artifact rows in the licence table.
## Curated CI feature-combination list
The former per-crate `cargo hack --each-feature` powerset is replaced by this
curated combo list — each kit feature alone, all-on, and none. It is pinned here
and driven by the `features` job of CI (`.github/workflows/ci.yml`), which runs
`cargo test -p coremlit --features <combo>`:
| (none, `default = []`) | the bare core builds/tests dependency-lean |
| `whisper` | the STT pipeline alone |
| `align` | forced alignment alone (asry emissions, no ort) |
| `speaker` | diarization backends + diaric clustering core (no ort) |
| `speaker,serde` | named smoke for `speaker`'s own serde code (`WindowOptions`' validated `Deserialize`) — see below |
| `vad` | Silero model layer alone (`zuoer` detector core, no ort) |
| `whisper,vad` | the `silero_vad` composition (former `vadkit` feature) |
| `align-oracle` | + asry ONNX aligner (ort + whisper.cpp) |
| `clap` | CLAP audio+text encoders alone (Rust mel + tokenizers, no ort) |
| `granite` | granite text embeddings alone (artifact-sidecar tokenizer + committed transformers-fp32 goldens, no ort; `embed_long` rides the crates.io `windit` engine + `windit/text`) |
| `siglip` | SigLIP 2 image+text embeddings alone (artifact-sidecar tokenizer + committed transformers-fp32 goldens, no ort) |
| `ced` | CED (tiny/mini/small/base) sound-event tagging alone (Rust mel + `soundevents-dataset` + `windit`, no ort) |
| `lid` | spoken-language identification alone (Rust mel + committed 107-label roster, no ort) |
| `identity` | speaker-identity embedding alone (Rust mel + committed mel goldens cut from the conversion recipe's own oracle, no ort) |
| `identity,speaker` | named pairing: `Scoring::IdentityCosine` sits behind `speaker` while the door producing its rows sits behind `identity`, and the `const` assert tying their dimensions together compiles only when both are on |
| `face` | the face alignment + embedder surface alone (its one dep is `sha2`; the caller supplies the artifact path, so this half is hermetic by construction) |
| `commercial-face-arcface` | the ONE licence-gated combo: the staged ArcFace artifact's manifest plus its four gate suites, whose hermetic halves (the committed silero refusal, the ONNX-reference loader, the alignment solve over the committed landmarks) run here while the model-gated halves stay `#[ignore]`d for the `arcface` shard |
| `whisper,align,speaker,vad,clap,granite,siglip,ced,lid,identity,face,serde,tracing,nl-recognizer` | all non-oracle features on |
| `whisper,align-oracle,speaker,vad,clap,granite,siglip,ced,lid,identity,face,serde,tracing,nl-recognizer` | all-on (every coremlit feature bar the `commercial-` gate, `align-oracle` included) |
`commercial-face-arcface` is deliberately NOT folded into either all-on row. A
licence opt-in that the crate's own "everything on" configuration turns on is
not an opt-in, and the row above is what keeps its gate suites building.
`serde` and `tracing` are cross-cutting and covered by the all-on runs, with
one named exception: `speaker,serde` (issue #129). `WindowOptions`' validated
`Deserialize` (`serde(try_from = WindowOptionsRepr)`, module doc "Validated
deserialization") is `speaker`'s only `#[cfg(feature = "serde")]` code, and it
broke — silently — behind the all-on rows' other ten features for long enough
that three `#[ignore]`d tests built on its PRE-fix behavior
(`extract_serde_bypassed_*`) went stale on `main` with nothing red. Both
all-on rows still exercise it too; this row exists so a `speaker`+`serde`
regression is legible by name instead of buried in a twelve-feature diff. It
is one curated row for the one kit that has needed it, not a new per-kit ×
`serde` dimension — the list stays explicit and reviewable, not an implicit
powerset.
"Artifact-sidecar tokenizer" (`granite`, `siglip`) means the crate embeds no
`tokenizer.json` for those two — each is a multi-megabyte file the published
model artifact ships beside the `.mlmodelc`, which `TextEmbedder::load` reads
and hash-checks against a pinned SHA-256. Only `clap` and `align` still
`include_bytes!` their tokenizers. Consequence for this table: the `features`
job is hermetic, so the `granite`/`siglip` rows build and run everything EXCEPT
the tokenizer gates — those are `#[ignore]`d on a staged artifact and belong to
the `model-tests` job, which stages both via `MODELS_LOCK`. SigLIP's entry is
the 34 MB `tokenizer.json` ALONE, not the ~784 MB bundle: its two tokenizer
gates call no `Model::load`, so the towers would buy them nothing. The
tower-dependent siglip gates (`model_io`, `text_model_io`, `parity_embed`,
`placement`, `e2e`) still run only locally.
`ced` has no tokenizer, but the same split applies to its model gates: the
`features` job runs the hermetic `ced` suite, and the model-gated
`ced_model_io`/`ced_parity_logits`/`ced_placement`/`ced_e2e` targets belong to
`model-tests`, which stages the artifact via `MODELS_LOCK`. The entry is
`ced-tiny` ALONE — 10.64 MB of the repo's 234 MB, since the four sizes are
I/O-identical and `ced-base` at 163.62 MB is past GitHub's 100 MB file limit
that let vadkit be committed instead. Each target declares its gates once per
size, so the CI step filters on `tiny::`; the `mini`/`small`/`base` gates stay
local/dev gates against an owner-staged `CED_TEST_MODELS` tree.
`lid` splits the same way. The `features` job runs the hermetic lid suite (the
mel front end, the roster/asset agreement checks, every typed-error path)
under the `lid` row and both all-on rows, and the `lid` `model-tests` shard
stages the 41 MB bundle and runs `lid_model_io` (5 gates), `lid_e2e` (4 gates)
and `lid_long_clip` (8 gates). There is no `@lib` half — every lid model gate
is a `tests/lid/*` target, so `--features lid --lib -- --list --ignored` lists
zero, and a group that selected it would trip the runner's anti-vacuum guard.
That shard was blocked for a release, and on the graph sweep rather than the
download. `tests/fp16_guards.rs` walks `Models/` WHOLE, and this artifact is a
coremltools 9 export whose scalar consts use MIL's terse
`fp16 v = const()[val = fp16(…)]` spelling instead of the `tensor<fp16, []>`
form the reader knew, so all 36 of its guard sites came back unresolved; and
once they were readable its final `softmax -> log` turned out to carry
`epsilon = 0x1p-149`, the same vanishing-guard defect `alignkit` and
`speakerkit/Segmentation` are pinned for. Both halves are settled — the reader
takes either spelling, and `KNOWN_DEFECTS` carries a
`lid/SpeechBrainECAPAVoxLingua107.mlmodelc` entry whose note records what the
inert guard costs today, the exact change that would arm it, and the two
repairs already measured NOT to work. The shard is also the second kit
to declare `checksum-dir: none`: this repo publishes its digests in an
`artifact_manifest.json` that `shasum -c` cannot read, so the exemption is
recorded with that reason in `CHECKSUMLESS_KITS`
(`tests/whisper/models_lock.rs`), which refuses `none` from any kit not listed
there and fails if a listed kit stops declaring it.
`speaker` splits the same way, with one wrinkle no other kit has: its artifact
set is staged from TWO repositories into one directory, and the second overlays
the first (MODELS_LOCK's last two tables, and the ORDER box in its header).
`model-tests` runs `speaker_model_io`, `speaker_parity_seg` and
`speaker_parity_embed` there. Four speaker targets deliberately stay out:
`speaker_parity_diarize_wiring`, whose fixtures live in the sibling
`diarization` repository (`DIA_PARITY_FIXTURES`) that no runner has, and the
three argmax targets (`speaker_argmax_model_io`,
`speaker_parity_argmax_accuracy`, `speaker_parity_argmax_swift`), because the
argmax artifact repo declares no license — so this repository does not fetch
those graphs in CI at all (NOTICE records the reasoning).
## Curated CI parity-oracle list
The three third-party oracles get their own CI job (`parity`), which runs
`cargo test -p coremlit-parity --features <combo>`. Pinned by the same golden
test, per job, so a dropped row cannot silently stop building an oracle:
| `speaker-oracle` | dia's ort DER reference oracle |
| `clap-oracle` | textclap model-level parity oracle (ort) |
| `vad-bundled` | the `silero` crate's ONNX cross-backend oracle |
| `speaker-oracle,clap-oracle,vad-bundled` | all-on |