
Agentic-first cryptography in pure Rust, with a machine-readable ontology.
Zero dependencies. No C. No build scripts. no_std from the ground up,
cross-compiled for ARM Cortex-M, RISC-V, and WebAssembly. Implemented primitives
are checked against published vectors or independent reconstructions; see
test provenance. Usable as a rustls
provider, so it can carry TLS.
Zero dependencies means every crate that implements an algorithm depends on nothing outside this workspace, checked per crate on each build. The rustls adapter is the one exception and necessarily so; see Supply chain.
$ ic recommend encrypt-message --fips
use: aes-256-gcm
AES-256-GCM is the approved authenticated cipher and retains 128-bit strength
against a quantum adversary.
call: ic_cipher::Aes256Gcm
must observe:
[critical] Never reuse a (key, nonce) pair.
Reuse leaks the authentication subkey, allowing forgery of arbitrary
messages, and XORs the two plaintexts together.
[serious] Let ic_cipher::Sealer choose the nonce: a sender tag and a strictly
increasing counter, never reused within it. Otherwise derive the nonce
from a counter yourself, or draw 96 random bits and bound the number of
messages per key.
Random 96-bit nonces collide with meaningful probability past 2^32 messages.
considered and rejected:
chacha20-poly1305: Not approved for the FIPS approved mode of operation.
aes-ctr: Unauthenticated; it is the confidentiality half of GCM.
Why another crypto library
Most cryptographic failures in production are not broken primitives. They are
correct primitives used wrongly: a reused GCM nonce, an unauthenticated CBC
mode, a password fed to a fast hash, a MAC compared with ==.
Humans learn those rules from documentation. An autonomous agent cannot read your prose and reliably act on it — so it guesses, and every guess is a chance to ship something broken.
IronCrypto puts the rules in the same place as the code, as data:
$ ic ontology show aes-256-gcm --json | jq '.constraints[0]'
{
"consequence": "Reuse leaks the authentication subkey, allowing forgery of arbitrary messages, and XORs the two plaintexts together.",
"id": "unique-nonce-per-key",
"requirement": "Never reuse a (key, nonce) pair.",
"severity": "critical"
}
An agent can refuse to emit code that violates a critical constraint. A
reviewer can diff the constraint set. A CI job can assert that nothing in the
codebase uses an algorithm the ontology marks disallowed.
The property that matters most
Never substitute something the caller did not ask for. That has two halves, and the first is refusing to trade a stated requirement for availability:
$ ic recommend sign-data --fips
use: ecdsa-p256-sha256
ECDSA P-256 is the approved signature scheme. This implementation derives its
nonce per RFC 6979, so the usual ECDSA nonce-reuse failure cannot occur.
call: ic_ec::p256::EcdsaP256Sha256
considered and rejected:
ed25519: Not approved for the FIPS approved mode of operation.
ml-dsa-65: Post-quantum and implemented, but larger and slower. Pass
--post-quantum to choose it, and prefer a hybrid with the classical scheme
over either alone.
rsa-pss-sha256: Also approved, but slower and far larger at the same strength.
Choose it only to meet an existing interface.
Ed25519 is implemented, is a signature scheme, and is faster. A library that optimizes for "always return something" could have returned it, silently missing the requirement the caller actually stated. Here it is passed over, and named, so the caller can see the choice rather than infer it.
The second half is refusing outright when nothing fits:
use ;
// No hash on offer is this quantum-resistant.
let outcome = recommend;
assert_eq!;
An honest "no" is worth more to an autonomous caller than a plausible "yes".
That second example used to be post-quantum key agreement, which was declined because ML-KEM-768 was implemented and unverified. It is checked against NIST's ACVP vectors now, so the request has a real answer and the example had to move to one that still does not.
Install
[]
= "0.2.14"
The current release is 0.2.14, published for all nineteen crates on
crates.io. It makes the library harder for an agent to misuse: Sealer
chooses AEAD nonces itself, the always-in-force rules ship as data and reach
MCP clients when they connect, ic lint checks code against them, and MCP
gains tools to open, verify and derive. See CHANGELOG.md.
ironcrypto is the facade and re-exports the rest; depend on the primitives
directly if you want a smaller graph:
[]
= "0.2.14" # AES, ChaCha20, the AEADs
= "0.2.14" # SHA-2, SHA-3, SHAKE, BLAKE2
= "0.2.14" # the NIST curves, X25519, Ed25519
= "0.2.14" # HPKE, RFC 9180
$ cargo install ic-cli --version 0.2.14 # the `ic` CLI and MCP server
The facade was renamed in 0.2.8. It was iron-crypto, imported as
iron_crypto::, and that package stays at 0.2.7 with no further releases.
Depend on ironcrypto and import ironcrypto::; nothing else changed.
The repository's release procedure requires recording a BIS/NSA notification before publication. docs/RELEASING.md describes the procedure, and docs/EXPORT.md its export considerations. The notification record records the user's reported submission and the 0.2.14 publication.
The source repository is public.
Use
use *;
This example uses the default std feature for OS seeding. A sealer's counter
is unique only within that sealer: a key that outlives it -- across a restart,
or shared with a second sealer -- needs its counter persisted and restored with
Sealer::resume, or AES-256-GCM-SIV. Never reuse a (key, nonce) pair.
Connect an agent
ic mcp is a Model Context Protocol server on stdio:
{
"mcpServers": {
"ironcrypto": { "command": "ic", "args": ["mcp"] }
}
}
| tool | what it answers |
|---|---|
crypto_recommend |
"What should I use for X under constraints Y?" |
ontology_list |
"What algorithms serve this purpose?" |
ontology_show |
"What are the parameter bounds and failure modes?" |
ontology_errors |
"What does this error mean and can I retry?" |
crypto_capabilities |
"What can this build actually do?" |
crypto_rules |
"What must I never do, whichever algorithm I use?" Also sent in the server's instructions at connect time |
crypto_selftest |
"Is the module healthy?" |
crypto_digest / crypto_hmac / crypto_seal / crypto_open / crypto_random |
primitive operations; crypto_seal draws the nonce when none is given |
crypto_verify |
"Is this signature valid?" ECDSA, Ed25519, ML-DSA and RSA; an invalid signature is an answer, not an error |
crypto_derive |
HKDF, to turn a shared secret into a key |
crypto_lint |
"Does this code misuse the library?" Pattern checks for each rule; a finding can be wrong, and none is not proof |
TLS
ic-rustls presents IronCrypto to rustls as a
CryptoProvider:
let roots = empty;
let config = builder_with_provider
.with_safe_default_protocol_versions?
.with_root_certificates
.with_no_client_auth;
# let _ = config;
# Ok::
Or ic_rustls::provider().install_default() once, for every rustls
configuration in the process.
| rustls slot | what IronCrypto supplies |
|---|---|
| AEAD | AES-128-GCM, AES-256-GCM, ChaCha20-Poly1305 — TLS 1.3 and TLS 1.2 |
| Hash | SHA-256, SHA-384 |
| MAC | HMAC-SHA256, HMAC-SHA384 |
| KDF | HKDF, as rustls's HkdfUsingHmac over the above |
| Signatures | ECDSA P-256/SHA-256 and P-384/SHA-384; Ed25519; RSA PKCS#1 v1.5 and PSS over SHA-256/384/512 — all verified and produced |
| Key exchange | X25519, ECDH P-256, ECDH P-384 |
| Randomness | SP 800-90A HMAC_DRBG, OS-seeded |
HKDF is rustls's own extract-and-expand over IronCrypto's HMAC rather than a second HKDF written for the occasion — HKDF is a construction, HMAC is the primitive — and the composition is checked against RFC 5869 appendix A.
It also signs, so it can run a server or present a client certificate and not
only authenticate one. load_private_key reads a PKCS#8 or a bare SEC1 EC key,
and a key on a curve the provider cannot verify is refused rather than loaded
— signing under a scheme this crate cannot check would advertise a capability
it cannot complete, and the handshake would then fail at the peer instead of
here, which is a much worse place to learn it. Signatures go out as the DER
Ecdsa-Sig-Value TLS carries, encoded by the same ic-pkix that decodes them
on the way in, so the two directions cannot drift apart. The round-trip test
signs with this crate and verifies with this crate's verifier — separate
modules, opposite ic-pkix calls — and both halves were mutation-checked:
handing back the fixed-width r || s instead of the DER, and choosing a scheme
that was never offered, each break it.
ChaCha20-Poly1305 is offered for both TLS versions, after the AES suites — most hardware here has AES instructions, and on hardware that doesn't, having it in the list at all is what keeps the connection alive. TLS 1.2 frames it differently from AES-GCM, which is the part worth stating: RFC 7905 gives it the TLS 1.3 nonce construction and sends no explicit nonce. Get that wrong and records round-trip against this crate perfectly and interoperate with nothing, so a test measures the record length against the plaintext rather than against the encrypter's own promise, and pins that the two TLS 1.2 framings differ by exactly those eight bytes.
RSA is verified and signed, which is what makes the provider usable against
the public web: most certificate chains out there are RSA, and TLS 1.3 needs
PSS for the handshake signature while the certificates themselves are
overwhelmingly PKCS#1 v1.5, so both paddings are required to authenticate one
connection. The public key algorithm is rsaEncryption throughout, including
for PSS — that is TLS's rsa_pss_rsae_*, a PSS signature made with an ordinary
RSA key.
Signing builds the key from its primes rather than from n, e and d, for
two reasons that happen to agree. ic-rsa then derives d and the CRT values
itself instead of reading the file's, so a file whose stored dP, dQ or
qInv contradict its primes cannot make the implementation compute a wrong
half and leak the factorization from one signature. It is also the only path
that can use the CRT at all — the n/e/d constructor yields a key carrying
no primes, which signs several times slower — so that one stays as a fallback
for keys that genuinely lack them. The modulus the file states is checked
against the one the primes produce; a key that disagrees is not the key the
certificate names.
One consequence worth stating plainly: a modulus below 2048 bits is refused,
so a chain with a 1024-bit key fails here and would pass against some other
providers. That isn't an oversight. ic ontology show rsa-pkcs1-sha256
marks rsa-modulus-at-least-2048-bits as critical, ic-rsa enforces it, and
a test pins the behaviour so it stays a decision on record rather than becoming
a mysterious handshake failure someone later "fixes".
QUIC packet and header protection are implemented for AES-GCM and ChaCha20-Poly1305. Header masks are checked against RFC 9001 vectors, with a separate test distinguishing long and short headers. Multipath QUIC is not implemented.
Absent: the mismatched ECDSA pairings — a P-256
key signed with SHA-384, or the reverse. ic-ec has no such combination, and
assembling one inside the adapter, out of sight of that crate's vectors, would
be worse than declining the chain.
The record layer is where this crate's own risk lives: the cipher is already checked against the GCM vectors, and none of that says whether a record is framed right. Wrong additional data, a misplaced tag, a nonce built wrong from the sequence number — each produces something that round-trips against itself and interoperates with nothing. So the tests check framing: every single-bit mutation across a whole record refused, replay at another sequence number refused, and for TLS 1.2 the content type and version bound into the header.
What's implemented
Most of what is below is checked against published test vectors — FIPS 180-4,
FIPS 197, FIPS 202, SP 800-38A/B/D, SP 800-90A, RFC
2104/4231/5869/7748/8032/8439/8452/9001/9106. Not all of it: cSHAKE, ECDSA
P-521 and CTR_DRBG have no published vector wired in, and PBKDF2 is
reconstructed from its own definition because RFC 6070 publishes HMAC-SHA1
only. Those rows are checked against independent reconstructions instead, and
docs/FIPS.md says which is which per algorithm rather than
leaving the stronger claim to stand for all of them.
| class | algorithms |
|---|---|
| Hashes | SHA-224/256/384/512, SHA-512/224, SHA-512/256, SHA3-224/256/384/512, BLAKE2b |
| SP 800-185 | cSHAKE128/256, KMAC128/256, TupleHash128/256, ParallelHash128/256 |
| XOFs | SHAKE128, SHAKE256 |
| MACs | HMAC (SHA-2 and SHA-3), CMAC-AES-128/192/256, KMAC128/256, Poly1305 |
| Block ciphers | AES-128/192/256 |
| Modes | CBC, CTR, PKCS#7, AES Key Wrap (KW and KWP) |
| AEADs | AES-128/192/256-GCM, ChaCha20-Poly1305, AES-128/256-GCM-SIV |
| Post-quantum | ML-KEM-512, -768 and -1024 (FIPS 203), ML-DSA-44, -65 and -87 (FIPS 204), each ACVP-checked. SLH-DSA (FIPS 205) is not implemented; ic ontology show slh-dsa says so and why |
| KDFs | HKDF, PBKDF2, SP 800-108 counter mode, Argon2id/i/d |
| DRBGs | HMAC_DRBG, CTR_DRBG, plus an OS-seeded auto-reseeding Rng |
| Curves | P-256, P-384, and P-521 (ECDSA with RFC 6979 nonces, ECDH), X25519, Ed25519 |
| Secret sharing | Shamir over GF(2^8), k of n, for backing up a key; constant time, in the AES field |
| Public-key encryption | HPKE (RFC 9180) base mode with DHKEM(X25519, HKDF-SHA256) or DHKEM(P-384, HKDF-SHA384) -- MLS suites 1 and 7 -- and AES-128-GCM, AES-256-GCM or ChaCha20-Poly1305, with the exporter |
| RSA | RSASSA-PSS and PKCS#1 v1.5 over SHA-256/384/512; 2048/3072/4096-bit key generation; CRT private operations |
| Backends | portable, written constant-time; on x86-64 AES-NI, PCLMULQDQ, SHA-NI and AVX2, each behind runtime detection. On 32-bit RISC-V the curve and RSA arithmetic runs on 32-bit words, because the 64-bit form compiled to branches on secrets there; SECURITY.md has what was checked on which target |
| Encodings | DER and PEM for SubjectPublicKeyInfo, PKCS#8, SEC1, and ECDSA signatures |
| TLS and QUIC | a rustls CryptoProvider: TLS 1.2, TLS 1.3 and QUIC; AES-GCM and ChaCha20-Poly1305; ECDSA, Ed25519 and RSA, verified and produced; ECDH and X25519; HKDF and the TLS 1.2 PRF |
What isn't — and why that's written down
The ontology registers algorithms this library does not provide, marked
planned or excluded, so that querying for them yields an honest answer:
- RSA encryption — not offered at all. RSAES-PKCS1-v1_5 is a Bleichenbacher oracle waiting to happen, and key transport is better served by ECDH. RSA signatures are implemented, because certificate chains are made of them.
- ARMv8 crypto extensions — implemented, behind the off-by-default
aarch64-cryptofeature. Local checks cross-compile it foraarch64-apple-darwinand its round structure is checked against a software model of the instructions — which is what caught the decryption key schedule being wrong. CI defines a native ARM64 job that runs the cipher and facade tests with this feature; a cross-build alone does not establish execution. - X.509 certificate parsing — out of scope. Names, validity, extensions, and
path validation are a far larger surface than key encoding, and a partial
implementation is worse than none. Keys and signatures do parse: hand the
SubjectPublicKeyInfofrom any X.509 parser toic_pkix::PublicKeyInfo. Issuing certificates is in scope, for one profile -- a CA and the leaves it signs, with basic constraints, key usage, extended key usage, alternative names and key identifiers -- inic_pkix::cert. Writing only what it chooses to is a small surface, and every certificate it writes for Ed25519 and ML-DSA is checked byte for byte against OpenSSL's.
Timing
A dudect-style leakage detector: two input classes interleaved at random, then Welch's t-test on the timings. It is a tool, not a gating test — timing needs a quiet machine, and a test that fails when a laptop indexes its disk teaches people to ignore failures.
It ships with a positive control, a deliberately early-exiting comparison that must show leakage. A detector that has never detected anything proves nothing; if the control is quiet, the report says every other result in the run is meaningless.
A null result means this run found no evidence on this machine. That is not a proof of constant time, and the output says so rather than printing a tick.
The recorded host diagnostics ran under heavy CPU load. Quiet-host repetition and Cortex-M/RISC-V hardware timing measurements remain outstanding.
Compiled-code checks
python scripts/check-ct.py checks nineteen fixed-size probes on x86-64 Linux,
Cortex-M0, Cortex-M4 and RISC-V: 76 probe/target checks. It rejects branches,
integer division and calls in the selected core, ML-DSA and ML-KEM operations.
The full local gate and CI run the checker. These checks cover the compiled
probes, not whole algorithms, memory-access timing or every caller context;
CONSTANT_TIME.md gives the rules and limits.
Supply chain
Deliberately deterministic: no timestamp, no random serial number, so two runs over identical source produce byte-identical documents and anyone can regenerate and diff it. A bill of materials you cannot reproduce is one you have to trust. The component list is checked against the workspace manifest by a test, so a crate added without being listed fails the build — an SBOM that quietly omits a component is worse than none, since completeness is its entire purpose.
It describes the source, not a particular binary. Establishing that a binary
came from this source needs a reproducible build pipeline, which T1195.001
records as an open gap rather than papering over.
Security frameworks
The same machinery, applied to the frameworks people are audited against:
Agents get the same through the crypto_controls MCP tool.
Read this before quoting any of it. CMMC SC.L2-3.13.11 requires
FIPS-validated cryptography. IronCrypto has no CMVP certificate, so that
practice is not satisfied, and no amount of correctness evidence changes
it — validation is a process with a laboratory and a certificate number, not a
property of source code. The summary line names unsatisfied controls
explicitly rather than leaving them to be filtered out, the MCP response carries
fips_validated: false as its own field, and tests assert both stay that way.
A compliance view that can be quoted without its gaps is worse than none.
On CVE: a library does not comply with CVE — CVEs are instances, and what an implementation can do is avoid the weakness classes they belong to, which is what the CWE entries cover. The other half is supply chain, and there every crate implementing an algorithm depends on nothing outside this workspace, so no advisory against another crate can apply to it. That claim is narrow on purpose and says nothing about defects in IronCrypto's own code.
One crate is outside it. ic-rustls implements rustls's traits and so depends
on rustls, which brings six crates with it — rustls-pki-types,
rustls-webpki, subtle, untrusted, once_cell and zeroize, seven in
all. Anything depending on ic-rustls inherits their advisories, and nothing
else here does. scripts/no-third-party.sh enforces both halves: it lists by
name what rustls may bring, so that set cannot grow unnoticed, and it checks
every other crate individually rather than looking at the workspace as a whole
— which is what lets the narrower claim still mean something.
Those seven are also held to a floor. scripts/advisories.sh pins each to the
version that fixed what is known against it — rustls to 0.23.45 for
RUSTSEC-2026-0285, rustls-webpki to 0.103.15, the version in use, above the
0.103.13 that fixed RUSTSEC-2026-0104 — and fails
the build below it, so a downgrade into a known vulnerability cannot pass
quietly. It runs on every commit and needs no network, because it compares
compiled versions against a table rather than fetching a database. The table
carries the date it was last reviewed, since a floor can only go stale in one
direction and no check learns that by itself. cargo audit runs beside it in
CI and is not allowed to fail the build: a gate that breaks when a remote
service is slow teaches people to bypass gates.
A scanner will count more than seven. cargo audit, Dependabot and most
SCA tools read Cargo.lock, which lists what the resolver considered rather
than what the compiler builds — eighteen packages here are never compiled,
ring among them, an optional dependency of rustls that no feature in this
workspace enables. cargo tree -i ring returns nothing. So a report of 44
dependencies is not wrong about the lock file and is not describing what ships;
SECURITY.md says how to tell the two apart before acting on a finding.
SECURITY.md has the disclosure process.
The standards knowledgebase
The registry says what algorithms exist. ic_ontology::standards says what
documents define them and what those documents require:
Agents get the same through the crypto_standard and crypto_requirements
MCP tools.
It is coupled to the code rather than merely filed beside it. A requirement marked met names a file and a symbol, and the tests check both exist — rename the function and the knowledgebase fails the build instead of going quietly out of date. Every document a registry entry cites must be described, and every document described must be cited or say why not, so neither list can drift from the other. An algorithm the library offers cannot rest solely on a withdrawn document.
A met requirement means the code does what the document asks, as far as the
tests can show. It does not mean a laboratory has agreed, and nothing here
claims otherwise — see has("fips-validated"), which returns false.
- SHA-1, MD5, Triple DES —
excluded, permanently. They are in the registry only so that a request for them resolves to a refusal with a reason.
$ ic capabilities
[x] no-std
[x] zero-dependencies
[x] constant-time-symmetric
[x] hardware-acceleration # on this machine; portable elsewhere
[ ] fips-validated
[x] post-quantum
[x] approved-asymmetric
[x] tls-provider
[x] key-encoding
Supplying test vectors
An algorithm is experimental when it is implemented and nobody has checked it
against values produced by something other than itself. That is a missing
file, not missing code: drop an ACVP or RFC vector file into testvectors/
and the matching test starts running. Without one it skips and says so.
It worked three times, and there is nothing left waiting. ML-KEM-768 and
ML-DSA-65 were experimental until NIST's published ACVP vectors were dropped
in — 105 cases across key generation, encapsulation and signing. AES-GCM-SIV
followed, on RFC 8452 appendix C's 50 cases. Every case in each set rather than
a selection, and no code changed for any of them.
No algorithm in the registry is experimental. Every one that is
implemented has been checked against values produced by something other than
itself.
testvectors/README.md has the format, the field names, and jq recipes for
converting ACVP output.
Honest limits
This is not a CMVP-validated module. FIPS.md describes what
is implemented (approved-mode policy, pre-operational self-tests, 73 algorithm
known-answer tests, a latching error state, service indicators) and what
validation would still require. ic capabilities reports
fips-validated: false and will keep reporting it until a certificate exists.
Throughput depends on the CPU, and the library tells you which case you are
in. On x86-64 with AES-NI and PCLMULQDQ the accelerated backend is selected
automatically. Everywhere else — ARM, RISC-V, WebAssembly, any bare-metal target
— AES falls back to the portable path, which computes its S-box algebraically
and multiplies GHASH bit by bit so that neither indexes memory with a secret.
That closes the cache-timing channel table-driven AES leaves open, and it is
slow:
| portable | accelerated | |
|---|---|---|
| AES-256, raw blocks | ~1.5 MiB/s | ~11 GiB/s (AES-NI) |
| AES-256-GCM | ~1 MiB/s | ~2.3 GiB/s (AES-NI + PCLMULQDQ) |
| ChaCha20-Poly1305 | ~0.4 GiB/s | ~1.5 GiB/s (AVX2) |
| SHA-256 | ~0.3 GiB/s | ~2.3 GiB/s (SHA-NI) |
Indicative figures from a Ryzen 9 9900X, moving by a factor of two between runs with clocks, load and build profile — orders of magnitude, not benchmarks.
How that compares
These are historical developer-machine comparisons from bench/ against
RustCrypto and dalek, using the same buffers and the best of nine runs. They
have not been rerun for 0.2.14 and are not performance guarantees. Shared-machine
interference caused substantial variation between runs.
| operation | historical comparison against RustCrypto/dalek |
|---|---|
| P-256 public key | ~3.2x faster |
| ECDSA P-256, sign | ~2.6x faster |
| AES-256-GCM | ~2.5x faster |
| ECDSA P-256, verify | ~1.5x faster |
| X25519 agreement | ~1.3x faster |
| AES-256 blocks, AES-NI | ~1.25x faster |
| Ed25519, sign | ~1.17x faster |
| SHA3-256 | ~1.05x faster |
| AES-256 blocks, portable | level |
| ECDH P-256 | level |
| ChaCha20-Poly1305 | level |
| SHA-256 | level |
| HMAC-SHA256 | level |
| HMAC-SHA256, 200 bytes | ~1.15x slower |
| SHA-512 | ~1.1x slower |
| Ed25519, verify | ~1.25x slower |
The portable AES row is against RustCrypto's software AES, which is fixsliced and is the like-for-like comparison; against AES-NI it is ~133x, which measures the instruction set rather than the implementation.
ECDH P-256 was 2.5x behind until 0.2.1, when arbitrary-point multiplication moved from a bit-at-a-time ladder to four-bit windows -- the method RustCrypto uses, which is why it now lands either side of parity rather than ahead. Public-key derivation had been skipping the generator table altogether, a multiplication of the generator by the ladder while signing used the table; nothing compared it against anything, which is why these two rows exist now.
Signing and public-key derivation moved from ~2.0x and ~2.2x after the generator table's accumulator changed, after 0.2.4, from Jacobian coordinates to homogeneous projective ones with Renes, Costello and Batina's complete addition. Jacobian addition is not complete, so it had been computing an addition and a doubling and selecting one, and the table does 65 additions to four doublings. Moving every point to the complete formulas was tried first and measured slower overall, since their doubling costs more and ECDH is four doublings per addition; so the table uses them and nothing else does.
AES-256-GCM was 1.4x ahead until GHASH, measured on its own, turned out to be three quarters of the time: it reduced every product and passed each one back through memory. Eight products summed and reduced once took it from 2.6 to 6.6 GiB/s, and the AEAD from 1.8 to 3.1. Counter mode then got a kernel of its own -- counters built in place, keystream XORed from registers, one call per message rather than one per eight blocks -- and the AEAD reached 4.1.
Short HMAC is the one row still behind that wiping explains. RustCrypto's
hmac does not zeroize by default; this library wipes both SHA-256 states and
the padded key on every tag. Over 200 bytes that is most of the remaining
difference, and it is a price this library pays on purpose. It was 3.5x behind
ring in an IronSocketLayer measurement until SHA-256 stopped sending buffered
and final blocks past SHA-NI.
Historical SHA-512 best-of-nine measurements put that row at 1.07 to 1.09x behind, while earlier medians straddled parity. Small differences require controlled, repeated comparisons that clear the observed noise.
The more recent focused ECDH diagnostics retain twenty alternating pairs per run. The ranges overlap under heavy load, so they support parity within the observed noise. IronCrypto also validates the SEC1 peer per call while RustCrypto receives a pre-parsed peer; these are API comparisons, not a before/after measurement of a source change.
Finding where the time went mattered more than optimising. Every figure that moved did so because a measurement contradicted the obvious explanation, and in several cases the obvious explanation was tested and disproved first.
SHA3-256 was assumed to need a faster permutation; timed alone the permutation
was already ahead, and the cost was in absorb, walking the input a byte at a
time. SHA-512 was assumed to need a vectorised schedule, and vectorising the
schedule made it slower -- AVX2 has no 64-bit rotate; what pays is
interleaving it into the rounds, which are a latency chain with idle issue
slots. The portable AES was assumed to need a better S-box, and it needed
bitslicing. Ed25519 was assumed to need faster field arithmetic, and our field
arithmetic was never the problem: point addition is 1.11x faster than dalek's
and decompression 1.37x faster. It was doing more work with it, in three places:
- Scalar reduction was long division.
reduce_widewalked 512 bits, shifting a nine-limb accumulator and conditionally subtractingLat each one -- some fifteen thousand operations, three of them per signature.Lis2^252 + cwithconly 125 bits, so a limb above2^252folds down multiplied byc's six digits instead. - Additions took their right-hand side in the wrong form. Every table here
holds points only in order to add them, so they now store
(Y+X, Y-X, Z, 2d·T): four multiplications instead of nine, three whenZis one. - Doublings computed a coordinate nothing read. A doubling reads
X,YandZand neverT, so on the way to another doubling the fourth multiplication producingTwas wasted.
Together those took signing from 1.6x behind dalek to ahead of it, and the double-scalar multiplication from 2715 field multiplications to about 2300.
Verification is the row still behind, and it is behind evenly. Timing a scalar with a single bit set against a dense one separates the doublings from the additions, and doing each scalar alone separates the two tables. Every piece lands in the same narrow band: the chain of doublings is 1.09x behind, the additions against the eight-entry table 1.16x, the additions against the basepoint table 1.11x.
There is no one slow step left to find. The operation counts match dalek's, the coordinate systems match, the window widths match, and a doubling issues the same 135 multiply instructions either way -- four squarings of fifteen products and three multiplications of twenty-five. What differs is how much else is in the instruction stream competing with the multiplier for issue slots, and six rounds of trying to reduce it have each either done nothing or made it worse: three arrangements of the carry (the one here is the best of them, confirmed by instruction count and by measurement), dalek's own term ordering, a narrower basepoint window, fusing the conversion into the addition, and inlining the point operations.
An AVX2 field backend was built far enough to measure -- a four-wide multiply
is real, 24.35ns against 43.83ns for four scalar ones -- and not kept, because
77% of a doubling is its multiplications and the other 23% becomes lane
shuffles rather than disappearing, and it would cost ic-ec its
forbid(unsafe_code). crates/ic-ec/src/field.rs records that and the four
other things tried against this gap that made it worse.
Ed25519VerifyKey holds a decompressed public key, the way Ed25519Key holds
a derived one for signing. Recovering the point is a field exponentiation, and
a caller verifying more than one signature against a key should not pay it
twice; Ed25519::verify still takes bytes and builds one per call.
Where the acceleration comes from
| primitive | how it is accelerated |
|---|---|
| AES | AES-NI, and ARMv8 crypto extensions behind a feature |
| GHASH | PCLMULQDQ, eight blocks per group, summed unreduced and reduced once |
| SHA-256 | SHA-NI |
| ChaCha20 | AVX2, eight blocks at a time |
| Poly1305 | four blocks per group, reducing once instead of four times |
Each is selected by runtime detection with the portable path as fallback, each is differentially tested against that portable path, and each reports which backend is live — a published vector passes whichever path runs, so it cannot tell you the backend executed at all.
bench/ holds the harness. It is excluded from the workspace, because
comparing against RustCrypto means depending on it and the library's own
dependency graph stays empty:
RUSTFLAGS="--cfg aes_force_soft"
This is why recommend asks the CPU rather than assuming: without AES
instructions it steers you to ChaCha20-Poly1305, and with them it picks
AES-256-GCM, which is then the faster of the two.
ic_ontology::runtime::backend() reports which backend is live.
The accelerated paths are not independently trusted — they are differentially tested against the portable ones, block for block, and the portable ones are validated against the FIPS 197, SP 800-38A, and GCM specification vectors.
This code has not been independently audited. It is correct against its test vectors; that is not the same as being reviewed by cryptographers. See SECURITY.md.
How this compares
Against aws-lc-rs, BoringSSL, and OpenSSL, IronCrypto leads on portability
(no dependencies in any algorithm crate, no C toolchain, genuine bare-metal
no_std) and on the
agent-facing ontology, which none of them has. With the NIST curves and RSA
signatures in place it covers the algorithms most deployments actually reach
for, post-quantum included. It still trails on RSA encryption, on X.509
certificate handling, and — decisively — on validation status. Pick
accordingly, and note that the ontology will tell you which case you are in
without your having to read this paragraph.
Docs
| document | what is in it |
|---|---|
| ONTOLOGY.md | the vocabulary, the query model, the export formats |
| FIPS.md | what is implemented, what validation would require |
| ARCHITECTURE.md | crate layout and design decisions |
| CHANGELOG.md | what changed between releases |
| AGENTS.md | instructions for agents working in this repo |
| SECURITY.md | threat model, side-channel posture, reporting |
| CONSTANT_TIME.md | compiled probes, target coverage and limits |
| bench/README.md | benchmark reproduction and controlled comparisons |
| RELEASING.md | why nothing goes public before the export notification |
| EXPORT.md | that notification, ready to send, and what to record |
Test
$ cargo test --workspace
$ cargo build -p ironcrypto --no-default-features --target thumbv7em-none-eabihf
$ cargo clippy --workspace --all-targets
$ bash scripts/check.sh
The full gate also checks formatting, dependencies, advisory floors and cross-target builds. Install its targets first:
$ rustup target add x86_64-unknown-linux-gnu thumbv6m-none-eabi thumbv7em-none-eabihf riscv32imac-unknown-none-elf wasm32-unknown-unknown aarch64-apple-darwin
The compiled-code checker requires Python 3. bash scripts/check.sh --quick
skips cross-target checks. Cross-compilation does not run code on an embedded
board.
License
IronCrypto is dual-licensed:
- AGPL-3.0-or-later for open-source use. Note that the network clause has real reach for a crypto library: linking this into a service that terminates TLS, signs tokens, or encrypts customer data makes that service a derivative work.
- Commercial for proprietary, embedded, or SaaS use without AGPL obligations. Contact licensing@nervosys.ai.
Contributions require agreement to the CLA; see CONTRIBUTING.md.
Neither license is a statement about cryptographic assurance: there is no CMVP certificate and no independent audit. See docs/FIPS.md and SECURITY.md.