1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
//! Recency resolution over `balls/tasks` history (§2/§9, § id generation) — the
//! ONE walk every dead-ball lookup shares. A closed/dropped ball deletes its
//! `tasks/<id>.md` (§2, no archive dir); the deletion is not a tombstone but
//! older CONTENT, recoverable most-recent-down from `git log`.
//!
//! Both `bl show <id>` (a live miss) and `bl list --status closed/--all` reach the dead
//! set through here, so the discipline is factored once: for one id, find the
//! NEWEST commit that deleted it, reconstruct its frontmatter from that
//! deletion's PARENT (the last tree that still held the file), and derive the
//! retirement and its date from the deletion commit itself (§5).
//! Taking the newest deletion makes a reused id unambiguous — at most one
//! incarnation is ever live, so the most recent dead one is "the" dead ball.
//!
//! **The plural read is BATCHED, and that is the whole performance story
//! (bl-4c08).** Singular and plural share the discipline, not the plumbing: one
//! id costs one `git log`, but N ids must not cost N of them. A per-id
//! `git log -1 -- tasks/<id>.md` walks from HEAD until it reaches that ball's
//! deletion, so its cost grows with the ball's AGE, and history length grows
//! with ball count — N such walks is quadratic. [`dead_balls`] instead pays ONE
//! walk (which already names every deletion's sha) and ONE `cat-file --batch`
//! for every pre-deletion blob: O(history + dead), two subprocesses, whatever N
//! is. Measured on this repo's own store (395 dead over 1193 commits): 7.4s →
//! 0.087s. Nothing is stored, cached, or indexed to get that — the redundant
//! re-derivation was simply deleted (§0 derive-don't-store is untouched; there
//! is no second representation to drift).
//!
//! **A deletion is a deletion, never half a rename (bl-ae74).** Git detects
//! renames by default, so ONE commit that deletes `tasks/<a>.md` and adds a
//! similar-enough `tasks/<b>.md` reports the pair as a single `R` — and
//! `--diff-filter=D` never sees `<a>` die. Task files are near-identical by
//! construction (same frontmatter keys, kindred bodies), so the similarity
//! threshold is trivially crossed, and the cost of missing one is the worst
//! failure this module has: the id would then name NOTHING, live or dead, and
//! absence is exactly how balls spells "resolved" — a swallowed deletion is
//! indistinguishable from a ball that never existed. Both walks therefore pass
//! `--no-renames`, though only [`newest_deletions`] can actually be bitten:
//! [`resolve_dead`]'s pathspec is the single file, and git pairs renames only
//! among paths that survived pathspec filtering, so the sibling add is not
//! there to pair with. That immunity is incidental — it depends on git's
//! filter-then-pair order, not on anything this module states — so the flag is
//! pinned on both, and the two reads agree on the dead set by construction
//! rather than by coincidence.
use HashSet;
use Write;
use io;
use Path;
use Catalog;
use crategit;
use crateTask;
use crateinvalid;
/// A ball reconstructed from history: its id, the frontmatter+body as it stood
/// the instant before deletion, and its deletion-commit date — the one fact the
/// gone file cannot carry. The deleting op is *not* reconstructed: every
/// retirement — a `close`, or a legacy `drop` deletion from before the verb
/// was deleted — projects as `closed`; the op stays git bedrock alone (§5).
pub
/// The `\x1f` field separator the reconstruction `git log` format uses — a
/// control byte that cannot appear in a sha or unix timestamp, so the two
/// fields split unambiguously.
const SEP: char = '\u{1f}';
/// Reconstruct one dead ball by id, or `None` when `tasks/<id>.md` was never
/// deleted on this branch (so the id names nothing — live OR dead). The recency
/// walk's single id→content step: the caller checks the LIVE set first (§9).
pub
/// One enumerated deletion, before its content is read: the ball's id, the
/// `<sha>^` revision still holding its last live bytes, and the deletion date.
/// The enumeration walk already knows all three, so nothing here is re-derived
/// per id — the object name the batch reads is that revision plus the path.
/// Every currently-dead ball, newest-deletion first — the `list --status closed/--all`
/// set (§9). Two subprocesses whatever the store's size: [`newest_deletions`]
/// enumerates, then ONE `cat-file --batch` resolves every pre-deletion blob in
/// the order the enumeration fixed, so the reply stream and the deletion list
/// zip positionally.
///
/// This does NOT go through [`resolve_dead`] — that is the point (see the module
/// header). The two paths still share the one reconstruction DISCIPLINE (newest
/// deletion wins, content from the deletion's parent, date from the deletion
/// commit); what they no longer share is a per-id `git log`.
pub
/// The newest deletion of each id ever deleted under `tasks/`, newest first,
/// with ids that are live again dropped (a reused id resolves live, §9).
///
/// `--format` + `--name-only` interleaves one `<sha>\x1f<ct>` header per commit
/// with the paths it deleted; a path line can never contain [`SEP`], so the two
/// kinds of line tell themselves apart and the header in force is the one that
/// deleted the paths under it.
/// Split one `cat-file --batch` reply off the front of `stream`: a
/// `<sha> <type> <size>` header line, then exactly `size` bytes of content, then
/// a newline. Returns the content and the rest of the stream.
///
/// Framed on BYTES, not chars: the size is a byte count, so decoding before
/// splitting would let one invalid byte (rewritten as a 3-byte `U+FFFD`) desync
/// every reply after it. Every object name fed to the batch came from git's own
/// deletion log, so `missing`/`ambiguous` replies — the only ones without a size
/// — cannot occur, and the framing is total.