# Changelog
All notable changes to this project will be documented in this file. The
format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and
the project uses semantic versioning within the usual pre-1.0 compatibility
rules.
## [0.1.11] - 2026-09-28
### Added
- A `ts` feature (off by default; implies `serde`) that derives
`ts_rs::TS` on every type the `serde` feature serializes — the views, the
date, encoding, conversion, and validation reports — so a caller that
sends them to a TypeScript frontend can generate its types from these
definitions rather than keep hand-written copies in step. The definitions
follow the JSON exactly: camelCase field names, `null` for an absent
value, and unit enums as their variant names. `StructureView::kind`, a
`&'static str`, is declared as the five words it can hold
(`"record" | "event" | "attribute" | "structure" | "unknown"`). Pulls in
`ts-rs` 12 (MIT) and its proc-macro only when enabled.
Two things a caller should know. ts-rs declares a 64-bit integer as
`bigint` by default, while the JSON carries a number: export with
`Config::with_large_int("number")` (as `tests/typescript.rs` does) for
definitions that match what `serde_json` writes. And a
`#[non_exhaustive]` enum becomes a closed TypeScript union, so a variant
added in a later patch release — additive in Rust — can break an
exhaustive `switch` in TypeScript until the definitions are regenerated.
## [0.1.10] - 2026-09-27
### Added
- `Node::fit_line_limit`, `Document::fit_line_limit` (the same, addressed by
a record's cross-reference identifier), and the `MAX_LINE_LENGTH` constant
for GEDCOM 5.5.1's 255-byte line, for a caller to fit an edited record to
a pre-7.0 document's line limit just before serializing:
`set_logical_value` only ever splits a payload across `CONT` at the
newlines it is given, so a long edited line was otherwise written over
length. Only nodes with no verbatim source line are touched, checked
recursively, so a node no caller reached into keeps its exact bytes; a
pointer payload is never split either. The line is measured with a
2-byte terminator reserve — wider than `wrap_long_lines`' 1 byte — because
this is meant for a caller whose own pipeline may keep a source file's
`\r\n` terminator rather than always writing `\n`. Bound in both language
bindings as `fit_line_limit` (Python) and `fitLineLimit` (WASM), mirroring
`wrap_long_lines`/`wrapLongLines`.
### Fixed
- `Document::next_xref` matched its prefix case-sensitively, so in a file
using lowercase identifiers (`@i1@`) `next_xref("I")` minted `@I1@` —
which `Document::record`'s case-insensitive lookup resolves to the
existing `@i1@`, silently colliding with it. The prefix match is now
case-insensitive, and the returned identifier is checked against every
existing xref case-insensitively and advanced past any collision. Its
doc comment claimed "the lowest unused" identifier while returning one
past the highest; the comment now says that.
- The same method minted `highest + 1` with unchecked arithmetic: a hostile
or fuzzed file carrying `@I18446744073709551615@` (`u64::MAX`) made it
panic in a debug build and wrap around to `@I0@` in release. Arithmetic
is now checked throughout, and once the series above the high-water mark
is exhausted, `next_xref` falls back to the lowest unused number instead
of overflowing.
## [0.1.9] - 2026-09-27
### Fixed
- `validate::inspect` drew years from dates the reader could not read: a
year merely mentioned in an unreadable payload (`WFT Est 1804-1911`) was
taken as the event's year, contrary to the module's own rule, and made
contradictions out of nothing. Such dates are now left out of every check.
- A date that bounds a year rather than naming one (`BEF 1811`, `AFT 1850`,
`BET 1800 AND 1810`) was compared as if it named its one mentioned year.
Checks now compare the years a date admits, so a finding is reported only
when no year both dates allow could be true, and the message says the
bounds it was drawn from (`born 1807 and died before 1805`; a parent age
of `at most 10 years old`).
- A birth dated after a death was reported twice, as `died-before-born` and
as `event-after-death`; it is now reported once.
- A Hebrew or French Republican year was compared as if it were Gregorian;
such dates are now left out of the checks. A range written backwards
(`BET 1910 AND 1900`) admits the same years as one written forwards.
- `child-born-after-mother-died` was skipped whenever the mother's own birth
year was unknown, although it needs only her death and the child's birth;
it now runs whenever those two are known, and alongside a
`parent-too-young` finding rather than instead of one. A child born before
its parent is worded as such (`5 years before the parent was born`), and a
lifespan between two exact years is stated exactly.
## [0.1.8] - 2026-09-27
### Added
- Python wheels for Windows (x86_64) and macOS (Apple silicon and Intel)
on PyPI, beside the Linux wheel, so installing on those platforms no
longer compiles Rust from the source distribution. All three are
cross-compiled on the Linux release runner — macOS with zig, Windows with
cargo-xwin — and, being abi3, each serves Python 3.9 onward. (PyO3 links
`python3.dll` by raw-dylib, so the Windows build needs no Windows Python.)
## [0.1.7] - 2026-09-27
### Added
- `dates::GedcomDate::to_gedcom_7`, the reading written in GEDCOM 7.0's own
date grammar (`about 5 January 1882` → `ABT 5 JAN 1882`), or `None` when
that grammar cannot hold it without loss — no `INT`, no dual year, and the
calendar as a bare word (`JULIAN 10 MAR 1700`).
- Python and JavaScript/WASM: `parse_date` / `parseDate`'s reading gained a
`gedcom7` key alongside `gedcom551`.
- `convert::to_version_7_with`'s owned-extension-tags parameter, now exposed
by both bindings: Python's `Gedcom.to_version7(owned=[(tag, uri), ...])`
and WASM's `toVersion7([[tag, uri], ...])`, either omittable.
### Changed
- `convert::to_version_7` now rewrites every `DATE` payload the lenient
reader understands but that is not version 7's own grammar — `29 May
1823`, `about 1575`, `13. september 1900`, `<1873>`, `ABT Oct 1803`,
`1740/41`, `@#DJULIAN@ 10 MAR 1700`, `44 B.C.`, and the like — instead of
passing most of them through unreported. Each becomes the version 7
reading (or, when none exists — an unparsed expression, a dual year, a
calendar version 7 cannot spell — an empty `DATE`), with the original kept
verbatim in a `PHRASE` substructure and the change named in the report. A
payload already in version 7's grammar is left untouched, byte for byte.
- `convert::to_version_7` now converts `OBJE.FILE.FORM`'s file extension
(`jpg`, `png`, `pdf`, `html`, `txt`, …) to a version 7 media type
(`image/jpeg`, `image/png`, `application/pdf`, `text/html`,
`text/plain`, …) and `FORM.MEDI`'s source medium to version 7's
enumeration, upper case (`PHOTO`, `NEWSPAPER`, …), folding anything else
into `OTHER` with the original kept as a `PHRASE`. An extension with no
known media type is kept as written and reported as kept rather than
guessed. Both are named in the report under new `FILE.FORM` and
`FORM.MEDI` notes, alongside the existing `OBJE.FORM` note for the move
from object to file. This covers 5.5.1's own layout, where `FORM` already
sits under `FILE` and the source medium is `TYPE` (renamed `MEDI`), as
well as 5.5's `FORM` on the object. Known extensions now include `heic`,
`webp`, `svg`, `json`, `xml`, `eml` and `mp4`.
- `convert::to_version_7` percent-encodes a `FILE` path that is not a URI
reference, as version 7 requires — a space, a backslash, anything outside
ASCII — and reports it under a `FILE` note. A path that is already a URI
reference, escapes included, is untouched.
### Fixed
- A `DATE` rewritten by the version 7 conversion no longer keeps the
`CONT`/`CONC` lines its old payload was split across, which a later read
folded back into the new payload. A `DATE` that already carries a
`PHRASE` is left as it is and reported as kept, rather than given a
second `PHRASE`.
- `GedcomDate::to_gedcom_7` returns `None` for a Hebrew or French
Republican date with a month, rather than spelling that month with a
Gregorian abbreviation (`FRENCH_R 22 MAR 5`). A bare year is still
written.
## [0.1.6] - 2026-09-25
### Added
- `build`, for a producer writing its own data as GEDCOM rather than
reading a file. `build::NodeSpec` describes a structure (tag, optional
identifier and payload, substructures) and `NodeSpec::into_node` checks
it under `Limits`, refusing with an error what `Node::new` would panic on;
with the `serde` feature it deserializes, refusing unknown keys.
`Document::append_records` adds records before the trailer, all or none,
refusing a repeated identifier or a second header or trailer, and
`Document::append_structures` adds substructures to any record or the
header (which 5.5.1 wants a `SUBM` pointer in, and `new_v551` cannot
supply), leaving every existing line as it was.
- `Document::wrap_long_lines`, which fits a 5.5.1 document to a line limit
by splitting payloads across `CONC` lines — at character boundaries, and
not beside a space where the text allows — and names every structure it
rewrote. Lines that fit are untouched; a version 7 document is refused.
- `dates::GedcomDate::to_gedcom_5_5_1`, the reading written in 5.5.1's own
date grammar (`about 5 January 1882` → `ABT 5 JAN 1882`), or `None` when
that grammar cannot hold it without loss. A dual year it can spell
(`1740/41`) is kept; a slashed pair it cannot makes an exact date an `INT`
date carrying the original as its phrase.
- Python and JavaScript/WASM bindings: `Gedcom.new_v551` / `newV551`,
`new_v7` / `newV7`, `append_records` / `appendRecords` (records as plain
dicts/objects), `append_structures` / `appendStructures`, `wrap_long_lines` / `wrapLongLines`, and
`parse_date` / `parseDate`, a date reading with its `display` and
`gedcom551` renderings.
- Python only: `ArchiveFileWriter`, a GEDZIP writer whose sink is a file,
with `add_file` streaming an entry from disk without the bytes passing
through Python — so a multi-gigabyte archive no longer has to fit in
memory twice. And type stubs (`gedcomkit.pyi`), so the package is typed
for mypy and pyright.
### Fixed
- The date reader accepts 5.5.1's `B.C.` epoch as well as 7.0's `BCE`.
- The date reader matched its keywords by ASCII case only, so `après 12
juillet 1670` and `før 1679` went unread although both words were in its
tables. Keywords now compare case in every script, and may run straight
into the date (`abt1810`, `c1700`).
- `April 9, 1950` read its `9` as September and failed; a month written as a
word now decides the order. `10 January 1566/67`, a dual year ending a full
date, now reads as 1566 as the bare `1566/67` already did.
- More spellings: `vers`, `cir`, `c.`, `around`, `probably`, `omtrent`,
`cirka` (about); `før` (before); `etter`, `efter` (after); `entre … et …`,
`mellom … og …` (between); `du … au …`, `fra … til …` (from … to);
and the months `déc`, `févr`, `janv`, `avr`, `juil`, `desember`.
## [0.1.5] - 2026-09-18
### Added
- `gedzip::ArchiveWriter`, a streaming GEDZIP writer: entries are added one
at a time and written straight to a sink, so peak memory is a small
multiple of the largest single entry plus the growing central directory,
never the archive's total. Measured, one `add` runs from roughly 2x the
entry (already-compressed media, stored) to a little over 3.5x (highly
compressible text, which briefly holds the source, the deflate output and
the encoder's state together) — never a function of how many entries came
before. A hash set of every lowercased path added so far is resident for
the writer's lifetime too, and at the default 100,000-entry limit the
central directory is bounded but not small.
It produces Zip64 structures once the entry count, an entry's size or the
running offset needs them, so it has no archive-size ceiling; the local
header's Zip64 extra carries both the original and compressed size
together whenever either overflows, as APPNOTE 4.5.3 requires for readers
that resolve local headers while streaming. Entry names carry
general-purpose bit 11 (UTF-8). A write failure to the sink poisons the
writer, so no entry is ever written next to bytes a failed write left
truncated, and `finish` flushes the sink before returning it rather than
leaving the last bytes to a `Drop` that cannot report failure.
`ArchiveWriter` implements `Debug` — entry count, bytes written, the
active limits and whether it is poisoned, never the sink or the directory
buffer's contents.
- `gedzip::Archive::total_size`, the sum of an opened archive's entries'
claimed decompressed sizes, read from the central directory alone — so a
caller can ask what extracting will cost before decompressing anything.
- Python and JavaScript/WASM bindings for both: `ArchiveWriter` (`new`,
`add`, `finish`) and `archive_total_size`. Both bindings write to an
in-memory `Vec<u8>` sink and copy it again when `finish` hands the bytes
to the host language, so — unlike the Rust API given a real sink such as
a file — they do not get the streaming memory benefit yet; their
docstrings say so.
### Fixed
Both entries below change `gedzip::write_archive`, which has shipped since
0.1.0. Everything under **Added** is new in this release and never shipped
broken; these two are what moves under an existing caller.
- **`write_archive` now sets general-purpose bit 11 (UTF-8 names)**, so a
non-ASCII path round-trips as itself instead of being read back as IBM
code page 437 mush (confirmed against Python 3.12 `zipfile` and Java 17
`ZipFile`; .NET's `ZipArchive` happened to decode the unflagged form
correctly, and 7-Zip guesses). Directory entries also now carry the MS-DOS
directory external attribute (`0x10`) instead of `0`. **This changes the
bytes `write_archive` emits**: on the same five-entry input, exactly
eleven bytes differ from 0.1.4 — the flags field on every local and
central record, and the external-attribute byte on each directory entry.
Versions, timestamps, CRCs, sizes, offsets and names are unchanged. No
fixture or test pinned the previous bytes.
- **`write_archive` now refuses one value earlier**, at the Zip64 sentinel
rather than past it. It rejected a count or size only once it *exceeded*
what a plain ZIP field holds, so it accepted a count of exactly 65,535
entries, or a size, compressed size, directory offset or directory size of
exactly 4,294,967,295 — but those values *are* the sentinels a reader
takes as "see the Zip64 record", which `write_archive` never writes. It
emitted the bare sentinel, and this crate's own `Archive::open` then
refused the result ("the archive needs a Zip64 directory and has none")
while Python, Java, 7-Zip and `bsdtar` accepted the broken archive.
**An input that produced a malformed archive in 0.1.4 is now an error**,
so a caller writing exactly 65,535 entries sees a refusal where it
previously saw a file. The design is unchanged — `write_archive` still
refuses rather than emit Zip64 — only the boundary moved from `>` the
sentinel to `>=` it, for the entry count and every 32-bit size and offset
field alike. Confirmed at 65,535 (refused), 65,534 and 4,294,967,294
(accepted and read back) against this crate's reader, Python 3.12, Java 17
`ZipFile` and `ZipInputStream`, 7-Zip 26.02, `bsdtar` and .NET's
`ZipArchive`.
## [0.1.4] - 2026-09-16
### Changed
- Migrated Cargo release publishing to crates.io Trusted Publishing through
GitLab OIDC, retaining the API token only as a transition fallback.
## [0.1.3] - 2026-09-16
### Changed
- Constrained PyPI Trusted Publishing to the protected `release` environment.
- Recorded the completed npm OIDC migration and the legacy token removal.
## [0.1.2] - 2026-09-16
### Changed
- Switched npm release publishing to GitLab OIDC Trusted Publishing.
- Documented the synchronized multi-registry release checklist.
## [0.1.1] - 2026-09-16
### Changed
- Synchronized the Rust core, Python binding, and JavaScript/WASM binding at
version 0.1.1 for the next registry release.
- Added protected-tag publishing for Cargo, PyPI, and npm, with clean
registry-install verification documented for all three packages.
## [0.1.0] - 2026-09-08
### Added
- Archival GEDCOM 5.5 through 7.x parsing, decoding, views, validation, and
reported version conversion.
- Bounded GEDZIP reading and writing.
- Python and JavaScript/WASM bindings over the same Rust core.
- Byte-exact vendored corpus, conformance, property, interoperability, and
adversarial test suites.
### Security
- Refuse traversal, platform-special, duplicate, and central/local-mismatched
GEDZIP paths before returning entry bytes.
- Stage and checksum external fixture downloads before making them visible.
- Reject line injection through programmatically constructed GEDCOM fields.