gedcomkit 0.1.7

A byte-preserving GEDCOM document model: decoding, parsing, readings, version conversion, plausibility checks, and the GEDZIP container, for GEDCOM 5.5 through 7.x.
Documentation

gedcomkit

pipeline status crates.io docs.rs

An archival-grade GEDCOM library: nothing dropped, nothing rewritten, every anomaly reported, safe on hostile input. A byte-preserving document model for GEDCOM 5.5 through 7.x, with encoding detection, read-only view projections, plausibility checks, explicit version conversion, and the GEDZIP container. Rust core; Python and JavaScript/WASM bindings in the same repository.

Install

cargo add gedcomkit

Python:

python -m pip install gedcomkit

Node.js / WASM:

npm install gedcomkit

Release checklist

  1. Choose a version and edit Cargo.toml, bindings/gedcomkit-py/Cargo.toml, and bindings/gedcomkit-wasm/Cargo.toml. Update the matching package entries in all three Cargo.lock files, CHANGELOG.md, and any affected README or binding documentation.

  2. Run the gates and release dry run:

    cargo fmt --all --check
    cargo clippy --all-features --all-targets -- -D warnings
    cargo test --all-features
    cargo test --no-default-features
    cargo check --manifest-path bindings/gedcomkit-wasm/Cargo.toml --target wasm32-unknown-unknown
    cargo check --manifest-path bindings/gedcomkit-py/Cargo.toml
    cargo publish --locked --dry-run
    bash scripts/check-release-tag.sh vX.Y.Z
    
  3. Commit and push the preparation, then wait for the exact main pipeline to pass:

    git add Cargo.toml Cargo.lock bindings/gedcomkit-py/Cargo.toml bindings/gedcomkit-py/Cargo.lock bindings/gedcomkit-wasm/Cargo.toml bindings/gedcomkit-wasm/Cargo.lock CHANGELOG.md README.md
    git commit -m "chore: prepare X.Y.Z release"
    git push origin main
    
  4. Create and push the protected annotated tag. The tag version must match all three manifests exactly:

    git tag -a vX.Y.Z -m "Release X.Y.Z"
    git push origin vX.Y.Z
    
  5. Watch verification, build, and publish jobs in order. Cargo publishes first; PyPI and npm use trusted publishing afterward. Verify every package from clean registry-only directories, and never rerun a successful registry upload or move a published tag.

Why another GEDCOM library

Every existing library makes the same trade: parse into a typed model, write canonical GEDCOM back out. That is a fine trade for an application — and the wrong one for an archive. A family history file is often the only machine-readable record of decades of work, full of constructs its next program will not recognize, and a round trip that normalizes it quietly discards evidence. gedcomkit trades the other way:

  • A node no caller touched is written back byte for byte. Preservation is a property of the representation, not a feature per construct: every line keeps its source text whenever re-rendering the parsed pieces would differ from what arrived, and the only way to change a node is through setters that surrender exactly that node's kept line. Golden round-trip tests over a 112-file corpus keep this true.
  • Nothing is dropped for being unrecognized. Unknown tags, vendor extensions, and odd constructs are read, kept, rendered in the views, and written back.
  • Nothing is decided in silence. Encoding disagreements — a file that declares UTF-8 and contains Windows-1252, ANSEL with no declaration at all — are reported, not guessed away. Validation reports contradictions with the numbers they were drawn from and never refuses a file. Version conversion names every construct it touched and keeps what the target cannot hold.
  • Hostile input is bounded. Decoding, parsing, and the GEDZIP reader run under explicit Limits — input size, line length, nesting depth, record and structure counts — so a malicious file is rejected rather than exhausting memory. A 30-minute, 3.3-million-case mutation-fuzz campaign and libFuzzer targets stand behind the claim.

Quickstart

use gedcomkit::view::IndividualView;

let bytes = std::fs::read("family.ged")?;
let (mut document, report) = gedcomkit::Document::from_bytes(&bytes)?;
println!("{}", report.summary()); // "Read as ANSEL, declared ANSEL"

let graph = gedcomkit::family::FamilyGraph::new(&document); // CHIL∪FAMC merged
for record in document.records_of("INDI") {
    let person = IndividualView::from_node(record);
    println!("{} {}", person.display_name(), person.life_years().unwrap_or_default());
}

document
    .record_mut("@I1@")
    .and_then(|record| record.first_mut("NAME"))
    .expect("name line")
    .set_logical_value("Ada /Lovelace/");
std::fs::write("family.ged", document.to_bytes())?; // CHAR now tells the truth

From Python or JavaScript, the same library through the same corpus-proven core (see bindings/):

import gedcomkit
doc = gedcomkit.Gedcom.parse(open("family.ged", "rb").read())
print(doc.encoding, doc.version, len(doc))
print(doc.parents_of("@I3@"))   # both directions of the family link

The pieces

Module What it does
decode Bytes to text: BOMs, UTF-8/16, ANSEL, Windows-1252, and disagreements between the declared and the actual encoding, all reported
Document / Node The document model: parse, address, edit, render — with the byte-preservation invariant enforced by the API, not by convention
dates, names, tags, view Read-only projections for display, sorting, and search — never a rewrite of the payload. Dates accept varied real-world shapes (BCE included) and can be written via DatePoint::to_gedcom
family The family graph: parents, children, partners merged from both directions of the link, because real files are often one-sided
validate Plausibility findings: contradictions and missing citations, reported with the numbers they were drawn from, never enforced
convert Explicit 5.5.1 ↔ 7.0 conversion that reports every construct it touched and keeps what the target cannot hold
schema GEDCOM 7 HEAD.SCHMA extension declarations
gedzip The GEDZIP container, read and written under the same bounds (default-on feature; turn it off for a zero-dependency parser core)

How it compares

ged_io is a good, actively maintained Rust choice when you want a normalizing model — parse into typed structs, write canonical GEDCOM out — and don't need the original bytes to survive.

gedcomkit ged_io (Rust) read-gedcom (JS) gedcom4j (Java)
Byte-preserving round-trip guaranteed no (normalizes) read-only no (normalizes)
GEDCOM 7 + GEDZIP yes yes 5.5.1 only 5.5.1 only
Encoding-mismatch reporting yes silent choice silent choice silent choice
Hostile-input bounds size, depth, counts file size — —
Version conversion reported per construct — — —
Python + browser bindings yes — JS only —
License AGPL-3.0-or-later MIT MIT MIT

How it is tested

  • The gedcom7code/test-files corpus, vendored and run on every cargo test, byte-for-byte.
  • The FamilySearch GEDCOM specification's own machine-readable tables (ABNF and generated TSVs, vendored at a pinned tag), driving conformance tests so a spec patch release becomes a failing diff — the first run caught twelve vocabulary gaps.
  • Differential validation: every corpus file's verdict compared against GED-inline and GedValidate. Zero files this reader refuses that either oracle reads clean.
  • Archive interop against ZIPs written by an independent implementation, Zip64 included; deterministic property suites; corpus-seeded mutation fuzzing (tests/endurance.rs) and libFuzzer targets (fuzz/).
  • Fetch-gated suites over the gedcom.io samples, the 5.5.1 torture files, the Eichmann character-set challenges, and sixteen real producer exports from Family Tree Maker, Ancestry, RootsMagic, Legacy, PAF, and others.

Extension tags

The library refuses to hard-code anyone's extension vocabulary as meaning. tags::vendor_label offers display labels for the vendor tags real files carry (_APID, _FSFTID, _FREL, …), grounded in a fourteen-producer corpus — but they stay badged as extensions, are never rewritten, and never validated. Callers label their own tags on top of tags, and pass their (tag, documentation URI) pairs to convert::to_version_7_with so version 7 output declares them; foreign extensions are declared as undocumented rather than dropped.

Features

  • gedzip — the GEDZIP container (on by default). With default-features = false the parser core has no dependencies at all.
  • serde — Serialize on the views and reports, for callers that send them to a webview or over a wire. Off by default.
  • fixture — a deterministic, corpus-shaped synthetic tree generator of any size, for performance and stress tests. Off by default.

AI disclosure

This project was created with the help of AI — specifically Anthropic's Claude. That includes the library code, bindings, tests, CI configuration, and documentation, with design decisions and validation made by a human. See AGENTS.md for the conventions AI contributors are expected to follow.

License

AGPL-3.0-or-later (see LICENSE). In short: use, study, and change it freely; if you distribute it — or run a modified copy as a network service — pass the same freedoms on, source included. The copyright holder can grant separate terms; open an issue to ask.

The library's code is original to this project. What it learned from elsewhere, it learned as published fact — the GEDCOM specifications' tag and date grammars, ANSEL's Appendix-D mapping, Windows-1252 — not as copied expression.

Acknowledgements

  • gedcom7code/test-files (Unlicense) — the vendored parser corpus this crate is proven against.
  • The FamilySearch GEDCOM specification and its machine-readable extracted-files (Apache-2.0), vendored under fixtures/vendored/gedcom-spec/ with their license text; conformance tests are generated from them, not hand-restated.
  • Heiner Eichmann's GEDCOM 5.5 torture tests and character-set challenges, and the gedcom.io sample files — used at test time only, fetched by script, and never redistributed here (their terms differ from this project's license).
  • GED-inline (Nigel Munro Parker, MIT) and GedValidate (Armidale Software, MIT) — the outside validators in the differential test gate.

GEDCOM is a trademark of Intellectual Reserve, Inc. This project is not affiliated with or endorsed by FamilySearch.