Veredictum grades openEHR servers. Point it at a running clinical data repository (CDR) and it tells you, with a specification citation on every finding, which parts of the released openEHR specifications that server actually implements, and what load it sustains while doing so.
What it does
The instrument is one binary plus a data tree. The data tree is a
machine-readable catalogue of 1103 test cases. Each case cites the
specification section it enforces, and the released specification text is
vendored in this repository, so every citation resolves against text you can
read. The case and binding counts on this page are the line
veredictum validate prints over artifacts/.
A grading run is three commands:
validatechecks the catalogue itself before any server is involved: id uniqueness, citation resolution against the vendored specs, binding completeness, and coverage of the enumerated wire surface. Zero findings is the only passing result.rundrives the applicable cases against your server over its own REST wire and records every request and response.verdictscomputes the verdict from those recordings and renders the report and certificate documents.
Three further subcommands share the same catalogue and recordings
discipline: perf measures a hospital-simulation workload against the
performance-class thresholds, stress finds the knee of the throughput
curve under stepped load, and aql-probe explores a server's AQL behaviour.
veredictum --help lists everything.
Why an independent instrument
A vendor's own test suite cannot answer the question a hospital procurement is asking. The suite and the server come from the same people, built on the same reading of the specification, and when the two disagree it is usually the suite that gets adjusted.
Veredictum is built so that adjustment has nowhere to happen. The released specifications are the only authority it accepts: every expectation in the catalogue names the section it comes from, so it can be refuted by a better reading of that text and by nothing else. No server's behaviour, no vendor's documentation, and no stalled upstream test suite ever sets an expected value. Where the released text is genuinely silent or contradicts itself, the gap goes to the ambiguity register with a typed disposition and is reported back upstream. A private resolution never happens.
Every failure is attributed before anything is changed. A red row has exactly three possible causes, and the instrument itself is a suspect ahead of the server:
| Suspect | Fix path |
|---|---|
| The server under test violates the specification | a defect report to that CDR, carrying the reproduced exchange and the citation |
| The instrument misdrove the case or misjudged the response | fix the runner; those rows were inconclusive, never failures |
| The catalogue expectation is wrong against the specification | fix the artifact, with a new cited source for the corrected expectation |
The first live triage attributed 7 of 7 diagnosed defects to the runner and none to the server under test. An instrument that presumes itself correct is worth nothing to the people who rely on its verdicts.
Quick start
Work from a clone. The published crate carries the code; the catalogue and
the vendored specification oracle are 347 MB of data no registry accepts, so
veredictum reads both as paths you pass it, and this repository is where
they live.
# 1. Check the catalogue itself. Zero findings is the only passing result.
# 2. Declare your deployment: copy an example and edit the endpoints, the
# credential variable names and the postures your server actually serves.
# 3. Drive the catalogue against your running server.
# 4. Compute the verdicts and render the submission documents.
The toolchain pins itself from rust-toolchain.toml. The only extra tool is
cargo-nextest, and only if you intend to run the test suite.
Without a Rust toolchain
Prebuilt binaries for x86_64 and aarch64 Linux are attached to each
release, each with a
sha256sum, a CycloneDX dependency SBOM and a Sigstore bundle:
The web console
The container image is the web console: a browser frontend over the same instrument, served by its own binary. Start it against a clone and it serves on port 3000:
The catalogue and the specification oracle are deliberately not baked into the image: the instrument reads every root as a path, and a party may legitimately point at their own. The console has no login, so the publish flag binds it to loopback; exposing it further is the operator's decision, behind their own gate. The console is under construction — image tags published before its first release still carry the CLI as the payload.
With cargo
Installing from crates.io puts the command on your PATH, which is the path
to take if you already have a catalogue checkout to point it at:
The library target is published with the binary, so an integrator can consume the typed artifact model and the published JSON Schemas directly instead of reimplementing the format.
What is in the box
| 1103 case cores | artifacts/schedule/ — one small isolated case per behaviour, so a red row names one defect. Grouped by chapter: EHR, composition, content, contribution, directory, query, definition, demographic, admin, messaging, security, SMART, simplified formats, system. schedule/performance/ holds the four measured-workload journey definitions, which are their own family and are not case cores |
| 247 operation bindings | artifacts/bindings/ — a case core says what an operation means, in the Service Model's own vocabulary; a binding says how it reaches the wire. A case core carries no status code, header or media type, so a new protocol adds binding files, never a new catalogue |
| The vocabularies | artifacts/vocab/ — the capability matrix behind the CORE, STANDARD and OPTIONS profiles, the wire surface the coverage gate enumerates, the outcome and selector grammars, and the journey catalogue the measured workload decomposes through |
| The corpora | artifacts/corpus/ — payload fixtures with their adjudicated verdicts, plus breadth packs vendored verbatim from upstream clinical-model libraries. Every invalid shape is kept as its own negative case, so a lenient server that accepts it fails |
| The ambiguity register | artifacts/registers/ambiguities.yaml — every place the specification is silent or contradicts itself, each with a typed disposition and, where we reported it, the upstream issue |
| The published schemas | schemas/ — JSON Schema for every artifact family, emitted by the instrument and drift-tested, so an integrator can author against the format |
| The verification pack | verification-pack/ — a recorded transcript with adjudicated verdicts. A runner claiming to implement this catalogue replays it and must reproduce every verdict, so no harness, this one included, is trusted on its word |
| The oracle | specs/openehr/ — the released specification text, vendored verbatim, plus the released XSD, JSON Schema and OpenAPI bundles a citation resolves against |
How a verdict is computed
A verdict is a pure function of four inputs: the party's statement (the capabilities the server claims), the recorded results, the catalogue, and the capability matrix. Nothing else enters. Two independent runners given the same four inputs must compute identical verdicts, and the verification pack exists to check exactly that. A certificate row a human typed is a defect.
Two verdict machineries share that discipline
(ARCHITECTURE.md §8):
- Conformance by assertion: the statement selects the applicable cases,
typed assertions judge each recorded exchange, and case results roll up
through capabilities to a profile verdict against the CORE / STANDARD /
OPTIONS matrix.
NotEvidencedandNoCasesare printed as first-class results, so a thin claim is visible instead of silently green. - Conformance by measurement: a performance class is earned when every threshold holds in one measured run. The class verdict is re-derived from the HDR histograms embedded in the record, so a stored summary is tamper-checked rather than trusted.
Load is offered open-loop: arrivals follow a seeded schedule of planned instants, and latency is measured from the planned instant rather than the actual send. A stalled server therefore accumulates the delay it caused, which is what stops coordinated omission from hiding a stall behind a slowed-down client.
The performance classes anchor to population served rather than to a
concurrent-user guess, with the full derivation from OECD, Eurostat and NHS
activity statistics in ARCHITECTURE.md §8.14:
| Class | Population served | Corpus | Sustained arrival floor | p99 budget | Error rate |
|---|---|---|---|---|---|
| POC | demonstration | 10k EHRs | 2/s | ≤ 1 s | 0 |
| S | 100 thousand | 100k EHRs | 15/s | ≤ 1 s | 0 |
| L | 1 million | 1M EHRs | 150/s | ≤ 1 s | 0 |
| R | 10 million | 10M EHRs | 1,500/s | ≤ 1 s | 0 |
Coverage is a mandate
A green run over a thin catalogue proves nothing, so coverage is
machine-checked rather than asserted. The surface-coverage gate enumerates
the wire surface from the released sources alone, the Service Model's
platform interfaces crossed with their ITS-REST branches, and fails on any
operation, status-code branch, header rule, negotiation variant or error
family that has neither a covering case nor a cited exception. A behaviour
the specification defines and the catalogue misses is a gap to close or an
honest boundary in the register.
Cases are added. They are never removed to make a run go green.
Lineage
None of the vocabulary here is invented. ISO/IEC 9646 standardized this
architecture in 1991: a supplier's conformance statement (ICS) selects the
applicable cases from an Abstract Test Suite, the supplier's IXIT provides
the instance parameters to run them, and verdicts land in a standardized
report. ETSI, the Bluetooth SIG and USB-IF still run on it. In those terms
the catalogue is the ATS, statement.json is the ICS, and ixit.json is
the IXIT.
openEHR's own conformance component defined the right concepts and then stalled: its last content amendment is from March 2022, its assessment layer was never written, and it carries zero AQL test cases. That component remains the structural guide for which behaviours need covering. It is never the correctness authority; the released specifications are.
Origin of the name
Veredictum is medieval Latin for "truly spoken", vere dictum, and it is the word that became the English verdict. That is what this instrument produces: it runs the catalogue against a running CDR and speaks a verdict about what it observed. The seal above is the mark of that verdict.
Design record
ARCHITECTURE.md carries the reasoning rather than a
summary of it: the testable surface and the case-core field definitions, the
per-operation wire bindings, the outcome taxonomy and the ambiguity
register, the assertion vocabulary, the verdict computation, and the
population-anchored performance-class model with its hospital-simulation
journey decomposition. It also records the evidence base for why the
instrument exists in this shape: the state of the official openEHR CNF
component, how other standards run conformance, and the ISO/IEC 9646 and
CASCO vocabulary the scheme is built in.
Contributing
CONTRIBUTING.md has the gates and the review bar.
CLAUDE.md is the working discipline the project holds itself
to, including the attribution law above. Security reports go through
SECURITY.md, and questions through
SUPPORT.md.
If you maintain a CDR and want it graded, open an issue. A defect this instrument finds in your server arrives with the reproduced exchange and the citation, and a defect you find in this instrument is a first-class bug here.
License
Apache-2.0. Attribution travels with every copy and derivative through the
license and the NOTICE file, as its section 4 requires. The vendored
specification text and clinical models keep their upstream terms, recorded
per tree in PROVENANCE.md and declared machine-readably in REUSE.toml.