gedcomkit 0.1.11

A byte-preserving GEDCOM document model: decoding, parsing, readings, version conversion, plausibility checks, and the GEDZIP container, for GEDCOM 5.5 through 7.x.
Documentation
# Changelog

All notable changes to this project will be documented in this file. The
format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and
the project uses semantic versioning within the usual pre-1.0 compatibility
rules.

## [0.1.11] - 2026-09-28

### Added

- A `ts` feature (off by default; implies `serde`) that derives
  `ts_rs::TS` on every type the `serde` feature serializes — the views, the
  date, encoding, conversion, and validation reports — so a caller that
  sends them to a TypeScript frontend can generate its types from these
  definitions rather than keep hand-written copies in step. The definitions
  follow the JSON exactly: camelCase field names, `null` for an absent
  value, and unit enums as their variant names. `StructureView::kind`, a
  `&'static str`, is declared as the five words it can hold
  (`"record" | "event" | "attribute" | "structure" | "unknown"`). Pulls in
  `ts-rs` 12 (MIT) and its proc-macro only when enabled.

  Two things a caller should know. ts-rs declares a 64-bit integer as
  `bigint` by default, while the JSON carries a number: export with
  `Config::with_large_int("number")` (as `tests/typescript.rs` does) for
  definitions that match what `serde_json` writes. And a
  `#[non_exhaustive]` enum becomes a closed TypeScript union, so a variant
  added in a later patch release — additive in Rust — can break an
  exhaustive `switch` in TypeScript until the definitions are regenerated.

## [0.1.10] - 2026-09-27

### Added

- `Node::fit_line_limit`, `Document::fit_line_limit` (the same, addressed by
  a record's cross-reference identifier), and the `MAX_LINE_LENGTH` constant
  for GEDCOM 5.5.1's 255-byte line, for a caller to fit an edited record to
  a pre-7.0 document's line limit just before serializing:
  `set_logical_value` only ever splits a payload across `CONT` at the
  newlines it is given, so a long edited line was otherwise written over
  length. Only nodes with no verbatim source line are touched, checked
  recursively, so a node no caller reached into keeps its exact bytes; a
  pointer payload is never split either. The line is measured with a
  2-byte terminator reserve — wider than `wrap_long_lines`' 1 byte — because
  this is meant for a caller whose own pipeline may keep a source file's
  `\r\n` terminator rather than always writing `\n`. Bound in both language
  bindings as `fit_line_limit` (Python) and `fitLineLimit` (WASM), mirroring
  `wrap_long_lines`/`wrapLongLines`.

### Fixed

- `Document::next_xref` matched its prefix case-sensitively, so in a file
  using lowercase identifiers (`@i1@`) `next_xref("I")` minted `@I1@` —
  which `Document::record`'s case-insensitive lookup resolves to the
  existing `@i1@`, silently colliding with it. The prefix match is now
  case-insensitive, and the returned identifier is checked against every
  existing xref case-insensitively and advanced past any collision. Its
  doc comment claimed "the lowest unused" identifier while returning one
  past the highest; the comment now says that.
- The same method minted `highest + 1` with unchecked arithmetic: a hostile
  or fuzzed file carrying `@I18446744073709551615@` (`u64::MAX`) made it
  panic in a debug build and wrap around to `@I0@` in release. Arithmetic
  is now checked throughout, and once the series above the high-water mark
  is exhausted, `next_xref` falls back to the lowest unused number instead
  of overflowing.

## [0.1.9] - 2026-09-27

### Fixed

- `validate::inspect` drew years from dates the reader could not read: a
  year merely mentioned in an unreadable payload (`WFT Est 1804-1911`) was
  taken as the event's year, contrary to the module's own rule, and made
  contradictions out of nothing. Such dates are now left out of every check.
- A date that bounds a year rather than naming one (`BEF 1811`, `AFT 1850`,
  `BET 1800 AND 1810`) was compared as if it named its one mentioned year.
  Checks now compare the years a date admits, so a finding is reported only
  when no year both dates allow could be true, and the message says the
  bounds it was drawn from (`born 1807 and died before 1805`; a parent age
  of `at most 10 years old`).
- A birth dated after a death was reported twice, as `died-before-born` and
  as `event-after-death`; it is now reported once.
- A Hebrew or French Republican year was compared as if it were Gregorian;
  such dates are now left out of the checks. A range written backwards
  (`BET 1910 AND 1900`) admits the same years as one written forwards.
- `child-born-after-mother-died` was skipped whenever the mother's own birth
  year was unknown, although it needs only her death and the child's birth;
  it now runs whenever those two are known, and alongside a
  `parent-too-young` finding rather than instead of one. A child born before
  its parent is worded as such (`5 years before the parent was born`), and a
  lifespan between two exact years is stated exactly.

## [0.1.8] - 2026-09-27

### Added

- Python wheels for Windows (x86_64) and macOS (Apple silicon and Intel)
  on PyPI, beside the Linux wheel, so installing on those platforms no
  longer compiles Rust from the source distribution. All three are
  cross-compiled on the Linux release runner — macOS with zig, Windows with
  cargo-xwin — and, being abi3, each serves Python 3.9 onward. (PyO3 links
  `python3.dll` by raw-dylib, so the Windows build needs no Windows Python.)

## [0.1.7] - 2026-09-27

### Added

- `dates::GedcomDate::to_gedcom_7`, the reading written in GEDCOM 7.0's own
  date grammar (`about 5 January 1882` → `ABT 5 JAN 1882`), or `None` when
  that grammar cannot hold it without loss — no `INT`, no dual year, and the
  calendar as a bare word (`JULIAN 10 MAR 1700`).
- Python and JavaScript/WASM: `parse_date` / `parseDate`'s reading gained a
  `gedcom7` key alongside `gedcom551`.
- `convert::to_version_7_with`'s owned-extension-tags parameter, now exposed
  by both bindings: Python's `Gedcom.to_version7(owned=[(tag, uri), ...])`
  and WASM's `toVersion7([[tag, uri], ...])`, either omittable.

### Changed

- `convert::to_version_7` now rewrites every `DATE` payload the lenient
  reader understands but that is not version 7's own grammar — `29 May
  1823`, `about 1575`, `13. september 1900`, `<1873>`, `ABT Oct 1803`,
  `1740/41`, `@#DJULIAN@ 10 MAR 1700`, `44 B.C.`, and the like — instead of
  passing most of them through unreported. Each becomes the version 7
  reading (or, when none exists — an unparsed expression, a dual year, a
  calendar version 7 cannot spell — an empty `DATE`), with the original kept
  verbatim in a `PHRASE` substructure and the change named in the report. A
  payload already in version 7's grammar is left untouched, byte for byte.
- `convert::to_version_7` now converts `OBJE.FILE.FORM`'s file extension
  (`jpg`, `png`, `pdf`, `html`, `txt`, …) to a version 7 media type
  (`image/jpeg`, `image/png`, `application/pdf`, `text/html`,
  `text/plain`, …) and `FORM.MEDI`'s source medium to version 7's
  enumeration, upper case (`PHOTO`, `NEWSPAPER`, …), folding anything else
  into `OTHER` with the original kept as a `PHRASE`. An extension with no
  known media type is kept as written and reported as kept rather than
  guessed. Both are named in the report under new `FILE.FORM` and
  `FORM.MEDI` notes, alongside the existing `OBJE.FORM` note for the move
  from object to file. This covers 5.5.1's own layout, where `FORM` already
  sits under `FILE` and the source medium is `TYPE` (renamed `MEDI`), as
  well as 5.5's `FORM` on the object. Known extensions now include `heic`,
  `webp`, `svg`, `json`, `xml`, `eml` and `mp4`.
- `convert::to_version_7` percent-encodes a `FILE` path that is not a URI
  reference, as version 7 requires — a space, a backslash, anything outside
  ASCII — and reports it under a `FILE` note. A path that is already a URI
  reference, escapes included, is untouched.

### Fixed

- A `DATE` rewritten by the version 7 conversion no longer keeps the
  `CONT`/`CONC` lines its old payload was split across, which a later read
  folded back into the new payload. A `DATE` that already carries a
  `PHRASE` is left as it is and reported as kept, rather than given a
  second `PHRASE`.
- `GedcomDate::to_gedcom_7` returns `None` for a Hebrew or French
  Republican date with a month, rather than spelling that month with a
  Gregorian abbreviation (`FRENCH_R 22 MAR 5`). A bare year is still
  written.

## [0.1.6] - 2026-09-25

### Added

- `build`, for a producer writing its own data as GEDCOM rather than
  reading a file. `build::NodeSpec` describes a structure (tag, optional
  identifier and payload, substructures) and `NodeSpec::into_node` checks
  it under `Limits`, refusing with an error what `Node::new` would panic on;
  with the `serde` feature it deserializes, refusing unknown keys.
  `Document::append_records` adds records before the trailer, all or none,
  refusing a repeated identifier or a second header or trailer, and
  `Document::append_structures` adds substructures to any record or the
  header (which 5.5.1 wants a `SUBM` pointer in, and `new_v551` cannot
  supply), leaving every existing line as it was.
- `Document::wrap_long_lines`, which fits a 5.5.1 document to a line limit
  by splitting payloads across `CONC` lines — at character boundaries, and
  not beside a space where the text allows — and names every structure it
  rewrote. Lines that fit are untouched; a version 7 document is refused.
- `dates::GedcomDate::to_gedcom_5_5_1`, the reading written in 5.5.1's own
  date grammar (`about 5 January 1882` → `ABT 5 JAN 1882`), or `None` when
  that grammar cannot hold it without loss. A dual year it can spell
  (`1740/41`) is kept; a slashed pair it cannot makes an exact date an `INT`
  date carrying the original as its phrase.
- Python and JavaScript/WASM bindings: `Gedcom.new_v551` / `newV551`,
  `new_v7` / `newV7`, `append_records` / `appendRecords` (records as plain
  dicts/objects), `append_structures` / `appendStructures`, `wrap_long_lines` / `wrapLongLines`, and
  `parse_date` / `parseDate`, a date reading with its `display` and
  `gedcom551` renderings.
- Python only: `ArchiveFileWriter`, a GEDZIP writer whose sink is a file,
  with `add_file` streaming an entry from disk without the bytes passing
  through Python — so a multi-gigabyte archive no longer has to fit in
  memory twice. And type stubs (`gedcomkit.pyi`), so the package is typed
  for mypy and pyright.

### Fixed

- The date reader accepts 5.5.1's `B.C.` epoch as well as 7.0's `BCE`.
- The date reader matched its keywords by ASCII case only, so `après 12
  juillet 1670` and `før 1679` went unread although both words were in its
  tables. Keywords now compare case in every script, and may run straight
  into the date (`abt1810`, `c1700`).
- `April 9, 1950` read its `9` as September and failed; a month written as a
  word now decides the order. `10 January 1566/67`, a dual year ending a full
  date, now reads as 1566 as the bare `1566/67` already did.
- More spellings: `vers`, `cir`, `c.`, `around`, `probably`, `omtrent`,
  `cirka` (about); `før` (before); `etter`, `efter` (after); `entre … et …`,
  `mellom … og …` (between); `du … au …`, `fra … til …` (from … to);
  and the months `déc`, `févr`, `janv`, `avr`, `juil`, `desember`.

## [0.1.5] - 2026-09-18

### Added

- `gedzip::ArchiveWriter`, a streaming GEDZIP writer: entries are added one
  at a time and written straight to a sink, so peak memory is a small
  multiple of the largest single entry plus the growing central directory,
  never the archive's total. Measured, one `add` runs from roughly 2x the
  entry (already-compressed media, stored) to a little over 3.5x (highly
  compressible text, which briefly holds the source, the deflate output and
  the encoder's state together) — never a function of how many entries came
  before. A hash set of every lowercased path added so far is resident for
  the writer's lifetime too, and at the default 100,000-entry limit the
  central directory is bounded but not small.
  It produces Zip64 structures once the entry count, an entry's size or the
  running offset needs them, so it has no archive-size ceiling; the local
  header's Zip64 extra carries both the original and compressed size
  together whenever either overflows, as APPNOTE 4.5.3 requires for readers
  that resolve local headers while streaming. Entry names carry
  general-purpose bit 11 (UTF-8). A write failure to the sink poisons the
  writer, so no entry is ever written next to bytes a failed write left
  truncated, and `finish` flushes the sink before returning it rather than
  leaving the last bytes to a `Drop` that cannot report failure.
  `ArchiveWriter` implements `Debug` — entry count, bytes written, the
  active limits and whether it is poisoned, never the sink or the directory
  buffer's contents.
- `gedzip::Archive::total_size`, the sum of an opened archive's entries'
  claimed decompressed sizes, read from the central directory alone — so a
  caller can ask what extracting will cost before decompressing anything.
- Python and JavaScript/WASM bindings for both: `ArchiveWriter` (`new`,
  `add`, `finish`) and `archive_total_size`. Both bindings write to an
  in-memory `Vec<u8>` sink and copy it again when `finish` hands the bytes
  to the host language, so — unlike the Rust API given a real sink such as
  a file — they do not get the streaming memory benefit yet; their
  docstrings say so.

### Fixed

Both entries below change `gedzip::write_archive`, which has shipped since
0.1.0. Everything under **Added** is new in this release and never shipped
broken; these two are what moves under an existing caller.

- **`write_archive` now sets general-purpose bit 11 (UTF-8 names)**, so a
  non-ASCII path round-trips as itself instead of being read back as IBM
  code page 437 mush (confirmed against Python 3.12 `zipfile` and Java 17
  `ZipFile`; .NET's `ZipArchive` happened to decode the unflagged form
  correctly, and 7-Zip guesses). Directory entries also now carry the MS-DOS
  directory external attribute (`0x10`) instead of `0`. **This changes the
  bytes `write_archive` emits**: on the same five-entry input, exactly
  eleven bytes differ from 0.1.4 — the flags field on every local and
  central record, and the external-attribute byte on each directory entry.
  Versions, timestamps, CRCs, sizes, offsets and names are unchanged. No
  fixture or test pinned the previous bytes.
- **`write_archive` now refuses one value earlier**, at the Zip64 sentinel
  rather than past it. It rejected a count or size only once it *exceeded*
  what a plain ZIP field holds, so it accepted a count of exactly 65,535
  entries, or a size, compressed size, directory offset or directory size of
  exactly 4,294,967,295 — but those values *are* the sentinels a reader
  takes as "see the Zip64 record", which `write_archive` never writes. It
  emitted the bare sentinel, and this crate's own `Archive::open` then
  refused the result ("the archive needs a Zip64 directory and has none")
  while Python, Java, 7-Zip and `bsdtar` accepted the broken archive.
  **An input that produced a malformed archive in 0.1.4 is now an error**,
  so a caller writing exactly 65,535 entries sees a refusal where it
  previously saw a file. The design is unchanged — `write_archive` still
  refuses rather than emit Zip64 — only the boundary moved from `>` the
  sentinel to `>=` it, for the entry count and every 32-bit size and offset
  field alike. Confirmed at 65,535 (refused), 65,534 and 4,294,967,294
  (accepted and read back) against this crate's reader, Python 3.12, Java 17
  `ZipFile` and `ZipInputStream`, 7-Zip 26.02, `bsdtar` and .NET's
  `ZipArchive`.

## [0.1.4] - 2026-09-16

### Changed

- Migrated Cargo release publishing to crates.io Trusted Publishing through
  GitLab OIDC, retaining the API token only as a transition fallback.

## [0.1.3] - 2026-09-16

### Changed

- Constrained PyPI Trusted Publishing to the protected `release` environment.
- Recorded the completed npm OIDC migration and the legacy token removal.

## [0.1.2] - 2026-09-16

### Changed

- Switched npm release publishing to GitLab OIDC Trusted Publishing.
- Documented the synchronized multi-registry release checklist.

## [0.1.1] - 2026-09-16

### Changed

- Synchronized the Rust core, Python binding, and JavaScript/WASM binding at
  version 0.1.1 for the next registry release.
- Added protected-tag publishing for Cargo, PyPI, and npm, with clean
  registry-install verification documented for all three packages.

## [0.1.0] - 2026-09-08

### Added

- Archival GEDCOM 5.5 through 7.x parsing, decoding, views, validation, and
  reported version conversion.
- Bounded GEDZIP reading and writing.
- Python and JavaScript/WASM bindings over the same Rust core.
- Byte-exact vendored corpus, conformance, property, interoperability, and
  adversarial test suites.

### Security

- Refuse traversal, platform-special, duplicate, and central/local-mismatched
  GEDZIP paths before returning entry bytes.
- Stage and checksum external fixture downloads before making them visible.
- Reject line injection through programmatically constructed GEDCOM fields.