Veredictum grades openEHR servers. Point it at a running clinical data repository (CDR) and it tells you, with a specification citation on every finding, which parts of the released openEHR specifications that server actually implements, what load it sustains, and how fast it answers.
It ships as two products over one engine:
- The instrument CLI — the
veredictumcommand, installed from crates.io or taken as a signed release binary. Every verdict this repository speaks is a run of it. - The web console — the container image at ghcr.io/rubentalstra/veredictum, a browser frontend that drives the same pinned CLI underneath: connect a CDR, paste the vendor's claim, watch the run live, read the results and the verdicts. The image is the console, never the CLI — a static binary needs no container.
A hosted instance of the console runs at console.veredictum.eu, no install needed: browse the catalogue, the party statements and the vendored specification text, or connect a publicly reachable CDR and drive a real run from the browser. Runs there write the instance's ephemeral filesystem, so a record you want to keep belongs on your own machine, one section down. The delivery pipeline behind the hosted instance is documented in deploy/vercel/README.md.
Both are pre-1.0. The 0.1.x line publishes working releases and makes no API-stability claim yet; every claim a release does make is checked, signed and reproducible, which is the stability that matters for a verdict.
What it does
The instrument is one binary plus a data tree. The data tree is a
machine-readable catalogue of 1146 test cases. Each case cites the
specification section it enforces, and the released specification text is
vendored in this repository, so every citation resolves against text you can
read. The case and binding counts on this page are the line
veredictum validate prints over artifacts/, and a CI guard fails the
build when a count here disagrees with the catalogue.
One command surface answers four different questions about a server:
| Question | Commands |
|---|---|
| Does it conform? | validate checks the catalogue itself before any server is involved; zero findings is the only passing result. run drives the applicable cases against the server over its own REST wire and records every request and response. verdicts computes the verdict from those recordings, as a pure function, and renders the report and certificate documents. |
| What class does it sustain? | perf seeds the population-scale corpus and holds a class's offered load for the sustained window, open-loop, merging the measured record into the results. |
| Where does it break? | stress steps load up to the maximum sustainable throughput; stress-compare overlays two committed stress reports; aql-probe explores AQL behaviour over the seeded corpus with per-statement database attribution. All three are exploration and never produce a conformance record. |
| How fast is it? | bench runs an embedded benchmark pack against any reachable CDR from a base URL and credentials; bench-compare aligns committed results into one table; bench-packs writes the byte-deterministic description of what every pack runs. Comparative speed only, never a conformance verdict. |
Three more subcommands serve the records themselves: verify-record checks
a sealed bundle, emit-schemas writes the published JSON Schemas, and
perf-assets / conformance-assets render the published charts from
committed artifacts. veredictum --help is the authoritative list.
run and verdicts take --sign-key, which seals the documents they emit
with a SHA-256 digest manifest and a detached OpenPGP signature over it.
verify-record recomputes every digest and checks that signature against a
public key you supply, so a published record is tamper-evident to anyone who
has the key. The bundle is ordinary files, so gpg --verify and sha256sum
answer the same questions without this tool.
Quick start
The fastest path installs nothing:
console.veredictum.eu is the hosted console
described above. For everything you want to keep, run it yourself — fastest
first: the console with docker compose up, the CLI from cargo, the signed
bare-metal binaries. Grading a server end to end needs the catalogue, which
lives in this repository — that path closes the section.
docker compose up — the console
Open http://127.0.0.1:3210. The image carries the console and the pinned engine; started beside an empty directory it comes up and says what it is missing, and started beside a checkout of this repository it reads the catalogue, the specification oracle and the party declarations from the mount. The console has no login, so the compose file binds it to loopback; exposing it further is the operator's decision, behind their own gate. From the next release onward the same file also sits in the release assets, pinned to that release's image. The console chapter shows what it does today.
cargo install
That puts the veredictum command on your PATH. The library target is
published with the binary, so an integrator can consume the typed artifact
model and the published JSON Schemas directly instead of reimplementing the
format.
Bare-metal binaries
Prebuilt binaries for x86_64 and aarch64 Linux are attached to each
release, each with a
sha256sum, a CycloneDX dependency SBOM and a Sigstore bundle you can check:
Benchmark a CDR in one command
The benchmark needs no clone and no declaration files: the packs are
embedded in the binary, pinned by digest, and described by bench-packs.
# The credential is read from the environment. It never rides argv.
community-vitals reproduces the openEHR community's vital-signs harness
and then measures the same population a second way, open-loop, so a stall
shows up in the percentiles instead of quietly reducing the request count.
--with-baselines composes the two pinned reference CDRs, EHRbase and
FerroEHR, from image digests on your machine and drives the same pack at the
same seed against each, so the record carries one relative index per
reference — the only kind of number that means anything across machines. A
declared posture profile is checked by canaries on both sides of the
measured window, and a run whose deployment disagrees with its declaration
is refused rather than recorded.
benchmarks/SUBMITTING.md
takes the record from there to the public board.
Run the full conformance catalogue
A conformance run reads the catalogue and the vendored specification oracle as paths. The published crate carries the code; those two trees are over 300 MB of data no registry accepts, and this repository is where they live — so grading a server starts from a clone:
# 1. Check the catalogue itself. Zero findings is the only passing result.
# 2. Declare your deployment: copy an example and edit the endpoints, the
# credential variable names and the postures your server actually serves.
# 3. Drive the catalogue against your running server.
# 4. Compute the verdicts and render the submission documents.
No installed binary? cargo run -- <subcommand> … from the clone does the
same; the toolchain pins itself from rust-toolchain.toml, and the only
extra tool is cargo-nextest, only if you intend to run the test suite.
Why an independent instrument
A vendor's own test suite cannot answer the question a hospital procurement is asking. The suite and the server come from the same people, built on the same reading of the specification, and when the two disagree it is usually the suite that gets adjusted.
Veredictum is built so that adjustment has nowhere to happen. The released specifications are the only authority it accepts: every expectation in the catalogue names the section it comes from, so it can be refuted by a better reading of that text and by nothing else. No server's behaviour, no vendor's documentation, and no stalled upstream test suite ever sets an expected value. Where the released text is genuinely silent or contradicts itself, the gap goes to the ambiguity register with a typed disposition and is reported back upstream. A private resolution never happens.
Every failure is attributed before anything is changed. A red row has exactly three possible causes, and the instrument itself is a suspect ahead of the server:
| Suspect | Fix path |
|---|---|
| The server under test violates the specification | a defect report to that CDR, carrying the reproduced exchange and the citation |
| The instrument misdrove the case or misjudged the response | fix the runner; those rows were inconclusive, never failures |
| The catalogue expectation is wrong against the specification | fix the artifact, with a new cited source for the corrected expectation |
The first live triage attributed 7 of 7 diagnosed defects to the runner and none to the server under test. An instrument that presumes itself correct is worth nothing to the people who rely on its verdicts.
The public results registry
Published results live in this repository as one append-only tree, conformance runs on one board and benchmark runs on another. A submission is a pull request that adds one entry, CI validates it before anybody reads the numbers, and the merge is the publication. Every entry records who submitted it, their relationship to the system, the deployment with its image digests, the instrument version, the machine, and the artifacts it stands on by digest.
Every entry carries one of two tiers, and the tier is a property of who
performed the run. A reproduced entry was produced by this repository's
own workflow: it composed the deployment from a recipe committed under
registry/topologies/, drove the catalogue against it, and attested the
bundle. A self-reported entry was run and signed by its submitter; the
signature proves who submitted the file and that the bytes have not moved,
and it never proves the run happened as described. The tier is the
discriminant of the entry's provenance block, so it cannot be claimed
without the evidence its variant requires.
A test report is not a certificate. An entry says what happened when a
named version of a named system was driven by a named version of this
instrument on a named machine. Certification is the openEHR Foundation's to
grant, and the registry is deliberately shaped to hand over: the rules
(registry/RULES.md, versioned, changed prospectively)
are public, the entries carry their own evidence, and no step of the
pipeline is proprietary.
What is in the box
| 1146 case cores | artifacts/schedule/ — one small isolated case per behaviour, so a red row names one defect. Grouped by chapter: EHR, composition, content, contribution, directory, query, definition, demographic, admin, messaging, security, SMART, simplified formats, system. schedule/performance/ holds the four measured-workload journey definitions, which are their own family and are not case cores |
| 249 operation bindings | artifacts/bindings/ — a case core says what an operation means, in the Service Model's own vocabulary; a binding says how it reaches the wire. A case core carries no status code, header or media type, so a new protocol adds binding files, never a new catalogue |
| The vocabularies | artifacts/vocab/ — the capability matrix behind the CORE, STANDARD and OPTIONS profiles, the wire surface the coverage gate enumerates, the outcome and selector grammars, and the journey catalogue the measured workload decomposes through |
| The corpora | artifacts/corpus/ — payload fixtures with their adjudicated verdicts, plus breadth packs vendored verbatim from upstream clinical-model libraries. Every invalid shape is kept as its own negative case, so a lenient server that accepts it fails |
| The ambiguity register | artifacts/registers/ambiguities.yaml — every place the specification is silent or contradicts itself, each with a typed disposition and, where we reported it, the upstream issue |
| The results registry | registry/ and benchmarks/ — the versioned submission rules, the deployment topologies the reproduction lane composes, and the committed entries the public boards render from |
| The published schemas | schemas/ — JSON Schema for every artifact family, emitted by the instrument and drift-tested, so an integrator can author against the format |
| The verification pack | verification-pack/ — a recorded transcript with adjudicated verdicts. A runner claiming to implement this catalogue replays it and must reproduce every verdict, so no harness, this one included, is trusted on its word |
| The oracle | specs/openehr/ — the released specification text, vendored verbatim, plus the released XSD, JSON Schema and OpenAPI bundles a citation resolves against |
How a verdict is computed
A verdict is a pure function of four inputs: the party's statement (the capabilities the server claims), the recorded results, the catalogue, and the capability matrix. Nothing else enters. Two independent runners given the same four inputs must compute identical verdicts, and the verification pack exists to check exactly that. A certificate row a human typed is a defect.
Two verdict machineries share that discipline
(ARCHITECTURE.md §8):
- Conformance by assertion: the statement selects the applicable cases,
typed assertions judge each recorded exchange, and case results roll up
through capabilities to a profile verdict against the CORE / STANDARD /
OPTIONS matrix. Version selection lives in the same two documents: each
case declares the spec-version ranges it applies to, the statement declares
the versions the product implements, and a case outside the declared
versions is out of scope — the instrument is version-aware per case, never
fixed to one release.
not_evidencedandnot_claimedare printed as first-class results, so a thin claim is visible instead of silently green. - Conformance by measurement: a performance class is earned when every threshold holds in one measured run. The class verdict is re-derived from the HDR histograms embedded in the record, so a stored summary is tamper-checked rather than trusted.
Load is offered open-loop: arrivals follow a seeded schedule of planned instants, and latency is measured from the planned instant rather than the actual send. A stalled server therefore accumulates the delay it caused, which is what stops coordinated omission from hiding a stall behind a slowed-down client.
The performance classes anchor to population served rather than to a
concurrent-user guess, with the full derivation from OECD, Eurostat and NHS
activity statistics in ARCHITECTURE.md §8.14:
| Class | Population served | Corpus | Sustained arrival floor | p99 budget | Error rate |
|---|---|---|---|---|---|
| POC | demonstration | 10k EHRs | 2/s | ≤ 1 s | 0 |
| S | 100 thousand | 100k EHRs | 15/s | ≤ 1 s | 0 |
| L | 1 million | 1M EHRs | 150/s | ≤ 1 s | 0 |
| R | 10 million | 10M EHRs | 1,500/s | ≤ 1 s | 0 |
Coverage is a mandate
A green run over a thin catalogue proves nothing, so coverage is
machine-checked rather than asserted. The surface-coverage gate enumerates
the wire surface from the released sources alone, the Service Model's
platform interfaces crossed with their ITS-REST branches, and fails on any
operation, status-code branch, header rule, negotiation variant or error
family that has neither a covering case nor a cited exception. A behaviour
the specification defines and the catalogue misses is a gap to close or an
honest boundary in the register.
Cases are added. They are never removed to make a run go green.
Lineage
None of the vocabulary here is invented. ISO/IEC 9646 standardized this
architecture in 1991: a supplier's conformance statement (ICS) selects the
applicable cases from an Abstract Test Suite, the supplier's IXIT provides
the instance parameters to run them, and verdicts land in a standardized
report. ETSI, the Bluetooth SIG and USB-IF still run on it. In those terms
the catalogue is the ATS, statement.json is the ICS, and ixit.json is
the IXIT.
openEHR's own conformance component defined the right concepts and then stalled: its last content amendment is from March 2022, its assessment layer was never written, and it carries zero AQL test cases. That component remains the structural guide for which behaviours need covering. It is never the correctness authority; the released specifications are.
Origin of the name
Veredictum is medieval Latin for "truly spoken", vere dictum, and it is the word that became the English verdict. That is what this instrument produces: it runs the catalogue against a running CDR and speaks a verdict about what it observed. The seal above is the mark of that verdict.
Design record
ARCHITECTURE.md carries the reasoning rather than a
summary of it: the testable surface and the case-core field definitions, the
per-operation wire bindings, the outcome taxonomy and the ambiguity
register, the assertion vocabulary, the verdict computation, and the
population-anchored performance-class model with its hospital-simulation
journey decomposition. It also records the evidence base for why the
instrument exists in this shape: the state of the official openEHR CNF
component, how other standards run conformance, and the ISO/IEC 9646 and
CASCO vocabulary the scheme is built in.
Contributing
CONTRIBUTING.md has the gates and the review bar.
CLAUDE.md is the working discipline the project holds itself
to, including the attribution law above. Security reports go through
SECURITY.md, and questions through
SUPPORT.md.
If you maintain a CDR and want it graded, open an issue. A defect this instrument finds in your server arrives with the reproduced exchange and the citation, and a defect you find in this instrument is a first-class bug here.
The public roadmap board shows what is planned, in progress, and shipped — a view over the issue tracker, where milestones are releases.
Credits
The conformance work here stands on work other people did first. Each entry below says what that work contributed to this catalogue.
- The openEHR SEC and the CNF authors: the Conformance component, whose Conformance Guide and Platform Conformance Test Schedule set the SUT model, the profile matrix and the certificate shape, and say which behaviours a platform product has to be tested for. The schedule's amendment record names T Beale, B Naess, I McNicoll, C Chevalley, H Frankel, S Iancu, B Lah and W Wagner across its revisions, beside P Pazos. 349 of the 1146 case cores here cite one of its Test Schedule chapters.
- Pablo Pazos (CaboLabs): the fleshed EHR, COMPOSITION, CONTRIBUTION and DIRECTORY chapters of that schedule, which are its usable core. The amendment record names him as the raiser of Test Schedule revisions 0.8.0 (23 Nov 2021) through 0.8.6 (24 Mar 2022), and as co-author of the Conformance Guide's initial writing with T Beale. He wrote the original 2019 EHRbase conformance tests at Hannover Medical School, and his openEHR conformance verification framework is the expanded continuation of that work, carrying a conformance testing specification of its own. He has argued the case for openEHR conformance testing on the community forums for years. 127 of the 1146 case cores cite the four chapters those revisions wrote.
- The EHRbase and vitasystems team: the executable battery. The 223 Robot files the CNF component vendored name Wladislaw Wagner (Vitasystems GmbH), Pablo Pazos and Jake Smolka (Hannover Medical School) in their copyright headers, and the team maintains that set as its integration tests. 156 of the 462 corpus provenance records here name that set as the source of the entry's bytes or of its template skeleton, each one re-adjudicated against the released specifications.
- The openEHR Foundation: the released specifications every expectation in the catalogue cites, and the machine-readable artifacts the bindings resolve against: the ITS-XML and ITS-JSON schema bundles and the ITS-REST OpenAPI documents.
The Test Schedule chapters are cited as the structural guide to which behaviours need covering. The correctness authority is always the released specification a case cites.
License
Apache-2.0. Attribution travels with every copy and derivative through the
license and the NOTICE file, as its section 4 requires. The vendored
specification text and clinical models keep their upstream terms, recorded
per tree in PROVENANCE.md and declared machine-readably in REUSE.toml.
openEHR
openEHR® is the registered trademark of the openEHR Foundation. Veredictum is an independent, community-driven conformance instrument: it names openEHR descriptively, to say what is being tested against, and it is not an official openEHR Foundation product, not the Foundation's CNF program, and not endorsed by or affiliated with the Foundation. The released openEHR specifications are this instrument's oracle by its own choice, and every expectation cites them — that fidelity is a design discipline here, never a claim of official status.