tga 10.2.0

Developer productivity analytics — git commit collection, classification, and reporting
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
# tga v1.3.0 Regression Snapshot — May 27, 2026

**Date**: 2026-05-27
**Version**: v1.3.0
**Binary**: `~/.cargo/bin/tga` (installed via `cargo install tga`)
**Corpus**: Duetto `~/Duetto/cto` — 72,608-commit production corpus (20 JIRA projects)
**First published benchmark** for trusty-git-analytics.

## Hardware & Environment

| Field | Value |
|-------|-------|
| Machine | Apple M4 Max |
| RAM | 128 GB |
| OS | macOS 26.5 (Build 25F71) |
| Logical CPUs | 16 |
| tga version | 1.3.0 |
| SQLite DB | `~/Duetto/cto/tga.db` |
| Rules file | `/tmp/bench-rules.yaml` (stripped copy — see Methodology) |
| JIRA tenant | duettoresearch.atlassian.net |
| JIRA projects | SRE, PM, GC, BB, ESS, ML, IC, PLA, BI, UIARCH, DP, ADV, IS, DSP, AE, DE, GEN, OS, PA, DRE |

## Methodology

### Corpus

The Duetto `~/Duetto/cto` repository is a private 72,608-commit production codebase
spanning multiple Duetto engineering projects. Commits range from the project's
inception to mid-2026 and reference JIRA tickets in 20 project keys. This corpus
represents a real-world, mixed-origin commit history — not a synthetic or cleaned
benchmark dataset.

### Rules file preparation

The source rules file (`~/Duetto/cto/configs/tga-classification-rules.yaml`, 344 lines,
37 categories) contains `method:` keys which are invalid under tga 1.2.2+ strict schema
validation (see tga strict schema changes in 1.3.0 release notes). A stripped copy was
prepared for all benchmark runs:

```bash
sed '/^    method:/d' ~/Duetto/cto/configs/tga-classification-rules.yaml > /tmp/bench-rules.yaml
```

The stripped file (307 lines) retains all `pattern:`, `category:`, `priority:`, and
`confidence:` fields. No rules were removed — only the unsupported `method:` annotation
keys were stripped.

### Timing method

All runs used `/usr/bin/time -l` (macOS `time` with `-l` flag) which reports:
- `real` — wall-clock time from process start to exit
- `user` + `sys` — CPU time (process + kernel)
- `maximum resident set size` — peak RSS in bytes
- `peak memory footprint` — macOS-specific working set peak

### Benchmark variants

| ID | Name | Flags | Purpose |
|----|------|-------|---------|
| B1 | Full classify (with JIRA) | `--force --rules /tmp/bench-rules.yaml` | Primary benchmark — all tiers including JIRA external source |
| B5 | No-external classify | `--force --rules /tmp/bench-rules.yaml --no-external` | Control — skips all JIRA HTTP; measures pure rule cascade throughput |

B5 ran first (synchronous, 0.64s). B1 ran after as the primary benchmark (22.8 minutes,
JIRA-bound). All runs used `--force` to re-classify every commit regardless of existing
classification state.

**Note on B1 DB state**: After B1 completed, the SQLite WAL state showed an unexpected
discrepancy (the classifications table reflected B5 data, not B1 data). This is a known
investigation item — see [Anomalies](#anomalies). The B1 numbers cited in this report
come from `tga classify`'s stdout classification summary (logged via `tee`), which is
authoritative for method/category counts. Post-run DB queries (B3, B6 category breakdown)
were taken from the DB in its B5 state and are labelled accordingly.

---

## Raw Numbers Summary

| Metric | B1 (with JIRA) | B5 (--no-external) | Delta |
|--------|---------------|-------------------|-------|
| Wall time | 1,367.15 s (22.8 min) | 0.64 s | +1,366.5 s |
| CPU user time | 19.28 s | 3.86 s | +15.4 s |
| CPU sys time | 1.42 s | 0.15 s | +1.3 s |
| Peak RSS | 235.0 MB | 232.2 MB | +2.8 MB |
| Classified commits | 49,143 / 72,608 | 46,654 / 72,608 | +2,489 |
| Coverage | 67.7% | 64.3% | +3.4 pp |
| Throughput (wall) | 53.1 commits/s | 113,450 commits/s | n/a (JIRA-bound) |
| CPU throughput | 3,508 commits/s | ~113,000 commits/s | — |

### Key headline numbers

- **CPU-local throughput**: ~113,000 commits/sec (no external sources) on Apple M4 Max
- **Overall coverage (with JIRA)**: 67.7% of 72,608 commits classified
- **JIRA contribution**: +3.4 pp coverage gain over rule-only baseline
- **Peak RSS**: 235 MB for full classify pass (72k-commit corpus)
- **JIRA wall time cost**: 22.8 minutes for 2,489 additional classified commits

---

## Method Distribution

### B1 (Full run — with JIRA external source)

Source: `tga classify` stdout summary from B1 run log.

| Method | Commits | % of Total | % of Classified |
|--------|---------|------------|-----------------|
| `regex_rule` | 33,746 | 46.5% | 68.7% |
| `fuzzy_match` | 23,465 | 32.3% | — (uncategorized) |
| `external_source` | 8,459 | 11.7% | 17.2% |
| `weighted_sum` | 6,938 | 9.6% | 14.1% |
| **Classified total** | **49,143** | **67.7%** | **100%** |
| **Uncategorized** | **23,465** | **32.3%** | — |
| **Grand total** | **72,608** | **100%** | — |

### B5 (No-external — rule cascade only)

Source: `tga classify` stdout summary from B5 run log; also confirmed by DB query.

| Method | Commits | % of Total | % of Classified |
|--------|---------|------------|-----------------|
| `regex_rule` | 38,673 | 53.3% | 82.9% |
| `weighted_sum` | 7,981 | 11.0% | 17.1% |
| `fuzzy_match` | 25,954 | 35.7% | — (uncategorized) |
| **Classified total** | **46,654** | **64.3%** | **100%** |
| **Uncategorized** | **25,954** | **35.7%** | — |

**Observation**: In B5, `regex_rule` accounts for 82.9% of classified commits vs. 68.7%
in B1. The B1 JIRA external source classified 8,459 commits that otherwise fell through
to `fuzzy_match` (uncategorized). This means JIRA classification is reclassifying commits
that regex rules cannot reach — it provides complementary coverage, not duplicate coverage.

**Observation**: `weighted_sum` (Tier 2.5, new in 1.3.0) contributes 7,981 commits (11.0%)
in B5 and 6,938 (9.6%) in B1. The lower B1 count suggests JIRA external source wins
some commits that would otherwise fall to weighted_sum (external_source outranks weighted_sum
in the cascade). The tier is active and meaningful — see [Weighted-sum Tier Analysis](#weighted-sum-tier-analysis-b6) below.

---

## Per-Category Breakdown (B1)

Source: `tga classify` stdout summary from B1 run log.

Top 15 categories (excluding `uncategorized`):

| Rank | Category | Commits | % of Total | % of Classified |
|------|----------|---------|------------|-----------------|
| 1 | `devops` | 10,973 | 15.1% | 22.3% |
| 2 | `bug_fix` | 8,121 | 11.2% | 16.5% |
| 3 | `new_feature` | 8,036 | 11.1% | 16.3% |
| 4 | `tech_debt_refactoring` | 8,022 | 11.0% | 16.3% |
| 5 | `merge` | 4,485 | 6.2% | 9.1% |
| 6 | `qa` | 3,761 | 5.2% | 7.7% |
| 7 | `integration` | 1,769 | 2.4% | 3.6% |
| 8 | `feature` | 909 | 1.3% | 1.8% |
| 9 | `platform_infrastructure` | 824 | 1.1% | 1.7% |
| 10 | `chore` | 768 | 1.1% | 1.6% |
| 11 | `security` | 681 | 0.9% | 1.4% |
| 12 | `refactor` | 531 | 0.7% | 1.1% |
| 13 | `bugfix` | 199 | 0.3% | 0.4% |
| 14 | `platform` | 39 | 0.1% | 0.1% |
| 15 | `contracted_work` | 23 | 0.0% | 0.0% |
| 16 | `docs` | 2 | 0.0% | 0.0% |

**Note**: `bug_fix` (8,121) and `bugfix` (199) are separate categories from distinct
rule sets. Both exist in the Duetto rules file. Same observation applies to `feature` vs
`new_feature`. This is an artifact of the Duetto-specific rules file, not a tga bug.

---

## External-Source Attribution (B3)

Source: B1 run log.

JIRA external source produced 8,459 classifications (11.7% of total corpus,
17.2% of classified). These are commits that referenced JIRA ticket keys in
their commit messages where the JIRA API returned a successful issue lookup.

**Per-project JIRA classification breakdown**: Not captured from this run
(B3 requires querying the DB after B1, which reflected inconsistent state — see
[Anomalies](#anomalies)). Recommend re-running and capturing with:
```sql
SELECT category, COUNT(*) FROM classifications WHERE method='external_source'
GROUP BY category;
```

**JIRA project keys configured**: SRE, PM, GC, BB, ESS, ML, IC, PLA, BI, UIARCH,
DP, ADV, IS, DSP, AE, DE, GEN, OS, PA, DRE (20 projects).

---

## JIRA API Efficiency (B4)

Source: B1 run log parsed at WARN level.

| Metric | Value |
|--------|-------|
| JIRA failures in log | 11 |
| — HTTP 404 Not Found | 10 |
| — Connection error | 1 (ESS-3357) |
| External-source classifications produced | 8,459 |
| JIRA wall time (approx) | ~1,366 s (22.8 min) |
| Cache hits logged | 0 (cache hits are not WARN-level; not visible in log) |

**Logged failure keys** (all unique): UIARCH-3457, UIARCH-32885, SRE-2060, ESS-3357,
IC-2123, IS-2039, IS-14455, IS-1395, PLA-2479, PLA-634, PLA-339.

**Interpretation**: Only failures appear at WARN level. The 8,459 external-source
classifications imply many successful JIRA lookups that did not appear in the log.
The 22.8-minute wall time is almost entirely network-bound JIRA API latency — CPU
time (19.28s user + 1.42s sys = 20.7s CPU) confirms this.

The 404 Not Found errors indicate deleted or archived JIRA issues still referenced
in commit messages. This is expected in a long-running corpus — tickets are
occasionally deleted after commits are written.

**Effective JIRA utilization rate**: Cannot compute cache hit ratio from WARN-only
logs. The tga JIRA source uses an in-process cache; running with `--log debug` would
expose `jira cache hit key=...` entries. For this benchmark pass, debug logging was
not enabled.

**JIRA latency per unique ticket (rough estimate)**: The log shows 11 failures spread
across 22.8 minutes = ~2 minutes between logged events on average. This suggests either:
1. Most of the time is spent on successful JIRA fetches (not logged at WARN), or
2. Large batches of commits are processed at rule-tier speed between JIRA lookups

Given that CPU time is only 20.7s total while wall time is 1,367s, the JIRA HTTP
round-trip latency dominates. Typical Atlassian Cloud API latency ranges from 100ms
to 2,000ms per request; the aggregate 1,366s across presumably hundreds of successful
lookups implies hundreds of unique JIRA tickets in the corpus.

---

## Weighted-Sum Tier Analysis (B6)

Source: DB query against B5 state; B1 log for B1 totals.

The weighted-sum tier (Tier 2.5, introduced in tga 1.3.0) classifies commits via
scoring across multiple weak signals when higher tiers (exact_rule, regex_rule,
external_source) yield no match.

**B5 weighted_sum count**: 7,981 commits (11.0% of total, 17.1% of classified)
**B1 weighted_sum count**: 6,938 commits (9.6% of total, 14.1% of classified)

Weighted-sum category breakdown (B5 DB state):

| Category | Commits via weighted_sum | % of w_sum total |
|----------|--------------------------|------------------|
| `merge` | 4,813 | 60.3% |
| `feature` | 1,369 | 17.2% |
| `chore` | 832 | 10.4% |
| `refactor` | 652 | 8.2% |
| `bugfix` | 240 | 3.0% |
| `platform` | 50 | 0.6% |
| `integration` | 16 | 0.2% |
| `docs` | 9 | 0.1% |
| **Total** | **7,981** | **100%** |

**Assessment**: The weighted-sum tier is highly active — it classifies 7,981–11% of the
corpus in the no-external run. The `merge` category dominates (60.3%) which makes sense:
merge commits often have generic messages ("Merge branch 'x' into 'y'") that don't match
specific regex patterns but have a consistent multi-signal profile.

The default weights are untuned for this corpus. `feature` (17.2%) and `chore` (10.4%)
via weighted-sum suggest borderline commits where no single regex fires but combined
signals are sufficient. Tuning the weights for the Duetto corpus could meaningfully
improve precision.

**No-tune baseline**: The current 7,981 weighted-sum classifications represent the
out-of-box tier behavior. This is a useful floor; future tuning should compare against
these numbers.

---

## Reachability Coverage (B8)

Source: DB query against `fact_commit_reachability`.

| Metric | Value | % of 72,608 |
|--------|-------|-------------|
| Total rows in reachability table | 72,608 | 100% |
| On any tag | 58,324 | 80.3% |
| On default branch | 0 | 0.0% |
| On release branch | 58,338 | 80.3% |

Reachability data is fully populated for all 72,608 commits. The `on_any_tag` and
`on_release_branch` figures are nearly identical (58,324 vs 58,338), suggesting
release branches and tags track closely in the Duetto corpus.

**`on_default_branch = 0`**: This is unexpected and may indicate that the default
branch reachability scan did not complete, or that the `main`/`master` branch is
configured differently in the Duetto repos collected. This is not a tga regression
but a corpus configuration observation.

---

## `--no-external` Baseline Comparison (B5 vs B1)

| Metric | B5 (no JIRA) | B1 (with JIRA) | Delta |
|--------|-------------|----------------|-------|
| Wall time | 0.64 s | 1,367.15 s | +2134× |
| CPU time | 4.01 s | 20.70 s | +5.2× |
| Peak RSS | 232.2 MB | 235.0 MB | +2.8 MB |
| Commits classified | 46,654 | 49,143 | +2,489 |
| Coverage | 64.3% | 67.7% | +3.4 pp |
| `regex_rule` count | 38,673 | 33,746 | −4,927 |
| `weighted_sum` count | 7,981 | 6,938 | −1,043 |
| `external_source` count | 0 | 8,459 | +8,459 |
| Uncategorized | 25,954 | 23,465 | −2,489 |

**JIRA trade-off summary**: JIRA adds 3.4 percentage points of coverage (+2,489 commits
newly classified) at the cost of 22.8 minutes of wall-clock time. The CPU cost is minimal
(+16.7s) — essentially all of the 1,366-second overhead is JIRA network latency.

**When to use `--no-external`**: For rapid iteration on rule changes, CI environments
without JIRA credentials, or when 64% coverage is sufficient. The 3.4 pp coverage gain
from JIRA is meaningful for coverage completeness but expensive in wall time. For
production scheduled runs, the full B1 profile is appropriate.

---

## Anomalies and Notes

### A1 — B1 classification results not persisted to DB (SQLite WAL discrepancy)

After B1 completed (wall time 1,367.15s), querying `~/Duetto/cto/tga.db` showed
B5 method/category distributions rather than B1's. The WAL file (`tga.db-wal`) was
last modified at 11:31 EDT (B1 completion time) and is 12 MB, but `PRAGMA
wal_checkpoint(FULL)` returned error 11 ("database disk image is malformed").

The integrity check (`PRAGMA integrity_check`) returned `ok`, and the DB is at the
consistent B5 state. The most likely cause: B1 wrote to the WAL successfully but the
final WAL-to-main-file checkpoint encountered a partial page corruption in the WAL.
The DB fell back to the last clean checkpoint (B5 state) rather than applying the
corrupted WAL.

**Impact on benchmark**: B1 log output (captured via `tee`) is authoritative for
method/category counts. Post-run DB queries for B3 (external_source category
breakdown) and B6 (weighted_sum by category with B1 data) were not available.
B6 weighted_sum breakdown uses B5 DB data; B3 was not captured.

**Recommendation**: File a bug against tga's SQLite WAL handling. The WAL should be
checkpointed atomically before classify exits (currently it may rely on SQLite's
automatic checkpoint-on-close, which can fail under certain conditions). Adding an
explicit `PRAGMA wal_checkpoint(TRUNCATE)` at the end of classify would harden this.

### A2 — `regex_rule` count differs between B1 and B5 (−4,927 commits)

B5 classified 38,673 commits via `regex_rule`; B1 classified 33,746 (−4,927).
This is expected: JIRA `external_source` wins before `regex_rule` in the cascade for
commits that reference JIRA ticket keys. With JIRA enabled, ~4,927 commits that
matched regex patterns were superseded by JIRA external lookups (which produce
higher-confidence verdicts). The difference is 4,927 ≠ 8,459 external_source count,
suggesting ~3,532 external_source commits were from `weighted_sum` fallback,
not `regex_rule`.

### A3 — Duplicate category names (`bug_fix` vs `bugfix`, `feature` vs `new_feature`)

The Duetto rules file defines both `bug_fix` (7,441–8,121 commits) and `bugfix` (199–240
commits) as separate categories, as well as `feature` (909–1,369) and `new_feature`
(7,134–8,036). These duplicates are in the user's rules file and do not indicate a tga
defect. However, downstream report consumers should merge these categories or the rules
file should be cleaned up to use a single canonical name per work type.

### A4 — `on_default_branch = 0` in reachability table

All 72,608 commits in `fact_commit_reachability` show `on_default_branch = 0`. This
is anomalous — at minimum the commits on the active production branch should have
`on_default_branch = 1`. Possible causes: the default branch column was not populated
during the last `tga collect` run, or the branch name does not match the configured
`dora.production_branch` value. This is not a throughput regression but a data quality
gap worth investigating.

### A5 — JIRA ticket 404s clustered around `UIARCH`, `IS`, `PLA` projects

Of the 11 logged JIRA API failures, 3 are from `PLA`, 3 from `IS`, 2 from `UIARCH`,
and 1 each from `SRE`, `ESS`, `IC`. 404 Not Found indicates deleted/archived issues.
The `UIARCH` project (UI Architecture) has one issue with key `32885` — an unusually
high ticket number. This may indicate a migrated project or a key numbering anomaly in
the Atlassian tenant.

---

## Reproducibility

### Commands (exact)

```bash
# Prepare stripped rules file
sed '/^    method:/d' ~/Duetto/cto/configs/tga-classification-rules.yaml > /tmp/bench-rules.yaml

# Load credentials
cd ~/Duetto/cto
set -a; source .env.local; set +a

# B5 — no-external baseline (control run)
/usr/bin/time -l tga classify --force --rules /tmp/bench-rules.yaml --no-external 2>&1 | tee /tmp/bench-noext.log

# B1 — full run with JIRA
/usr/bin/time -l tga classify --force --rules /tmp/bench-rules.yaml 2>&1 | tee /tmp/bench-classify-full.log

# Post-run DB queries (reflects whichever classify ran last)
sqlite3 tga.db "SELECT method, COUNT(*) FROM classifications GROUP BY method ORDER BY 2 DESC;"
sqlite3 tga.db "SELECT category, COUNT(*) FROM classifications GROUP BY category ORDER BY 2 DESC LIMIT 30;"
sqlite3 tga.db "SELECT COUNT(*), SUM(on_any_tag), SUM(on_default_branch), SUM(on_release_branch) FROM fact_commit_reachability;"
```

### Required environment variables

| Variable | Purpose |
|----------|---------|
| `JIRA_API_TOKEN` | Atlassian Cloud API token |
| `JIRA_EMAIL` | Atlassian account email |

Both are read from `~/Duetto/cto/.env.local` (not committed to version control).

### Corpus pre-requisite

`~/Duetto/cto/tga.db` must have been populated by a prior `tga collect` run covering
the full 72,608-commit history. No `tga collect` was run as part of this benchmark —
the existing DB was used as-is.

---

## Caveats

1. **Single-corpus, single-machine results.** All numbers are from one run on one machine
   (Apple M4 Max, 128 GB RAM). CPU throughput will vary on different hardware. JIRA wall
   time will vary significantly with Atlassian tenant load and network conditions.

2. **Duetto-specific corpus.** The 72,608-commit corpus is a private production codebase.
   Numbers (especially coverage percentages and category distributions) are not portable
   to other organizations with different commit message conventions.

3. **JIRA latency is not stable.** The 22.8-minute JIRA runtime reflects Atlassian Cloud
   API latency at the time of measurement. Results will vary by time of day, tenant load,
   and number of unique ticket keys in the corpus.

4. **No prior baseline to compare against.** This is the first published tga benchmark.
   There is no v1.2.x result to compare throughput or coverage against. Future snapshots
   will provide version-over-version delta analysis.

5. **Weighted-sum tier weights are untuned defaults.** The 1.3.0 default weights have
   not been calibrated for the Duetto corpus. Tuned weights could shift 7,981 borderline
   classifications toward more precise categories.

6. **Rules file was stripped for compatibility.** The benchmark used a `method:`-stripped
   copy of the Duetto rules file. A native 1.3.0 rules file (without `method:` keys) may
   produce slightly different results due to schema differences.

7. **B1 DB state not queryable post-run.** The SQLite WAL corruption (A1) prevented
   per-category breakdown queries against the B1 classification results. B3 external-source
   category breakdown was not captured in this snapshot.

8. **`fastembed` warmup not measured.** No LLM or embedding operations were invoked in
   these runs (`--use-llm` was not passed). The fastembed model warmup cost is not
   included in any of these timings.

---

## Operational Notes

- tga binary: `/Users/masa/.cargo/bin/tga` (1.3.0, installed via `cargo install`)
- Run date: 2026-05-27
- B5 start time: 15:07:58 UTC; B5 end time: 15:07:58 UTC (+0.64s)
- B1 start time: 15:09:59 UTC; B1 end time: 15:31:05 UTC (+1367.15s)
- Log files: `/tmp/bench-noext.log` (B5), `/tmp/bench-classify-full.log` (B1)