fp-conformance 0.4.0

Differential conformance harness for frankenpandas against a live pandas oracle — packet fixtures + fuzz seed tests + live-oracle parity gates.
# Known Conformance Divergences

> Every intentional divergence from pandas behavior is documented here.
> Format: DISC-NNN, status (ACCEPTED/INVESTIGATING/WILL-FIX/RESOLVED), affected tests.
>
> **DISC numbers must be UNIQUE — check before you claim one.** On 2026-08-16 three
> IDs were each defined twice (DISC-018, DISC-019, DISC-020), so a citation could
> resolve to the wrong entry depending on which heading a reader hit first, and one
> real citation did. The duplicates were renumbered to DISC-022/023/024, keeping the
> number on whichever entry the existing in-tree citations actually meant. The next
> free number is DISC-032. To see every ID in use:
> `grep -n '^### DISC-' crates/fp-conformance/DISCREPANCIES.md`

## Active Divergences

### DISC-001: Integer division by zero promotes to Float64 with NaN/inf
- **Reference:** pandas `int64 // int64` with zero divisor returns `float64` with `inf`
- **Our impl:** Same behavior - promotes to Float64, returns `inf` for floor division, `nan` for mod
- **Impact:** Dtype promotion matches, values match
- **Resolution:** ACCEPTED - exact pandas parity achieved
- **Tests affected:** `int64_mod_floordiv_with_zero_promotes_to_float`
- **Review date:** 2026-04-15

### DISC-002: Unicode width tables version
- **Reference:** pandas uses system's ICU or Python's unicodedata (varies by install)
- **Our impl:** Uses `unicode-width` crate (Unicode 15.1 tables)
- **Impact:** Some emoji/CJK width calculations may differ by 1 column
- **Resolution:** ACCEPTED - newer Unicode tables are more correct
- **Tests affected:** None currently - string display width not yet tested
- **Review date:** 2026-04-15

### DISC-003: Error message text differs
- **Reference:** pandas error messages vary by version and locale
- **Our impl:** Custom error messages with consistent format
- **Impact:** Error semantics match, exact text differs
- **Resolution:** ACCEPTED - tests check error category, not message text
- **Tests affected:** All error-expecting tests use `expected_error_contains`
- **Review date:** 2026-04-15

### DISC-004: CSV NA value handling default differs from pandas 1.x
- **Reference:** pandas 2.x treats "None" as NA by default; pandas 1.x did not
- **Our impl:** Follows pandas 2.x behavior with `keep_default_na=true`
- **Impact:** Users migrating from pandas 1.x may see different behavior
- **Resolution:** ACCEPTED - aligning with current pandas 2.x
- **Tests affected:** `csv_none_is_default_na`
- **Review date:** 2026-04-15

### DISC-006: Row MultiIndex is scaffolded, not full pandas parity
- **Reference:** pandas `MultiIndex` supports arbitrary-level hierarchical row labels with full slicing, `xs`, `droplevel`, `swaplevel`, `reindex`, `sort_index`, etc.
- **Our impl:** Row-MultiIndex first slice ships struct + constructor + level access. Full slicing / xs / droplevel / swaplevel coverage lands in subsequent slices (umbrella tracked by br-frankenpandas-1zzp).
- **Impact:** DataFrames built with a row MultiIndex may reject operations that pandas accepts, or return partial results. Error messages identify which operation is pending.
- **Resolution:** INVESTIGATING - slices land under br-1zzp child beads until coverage parity is reached.
- **Tests affected:** `live_oracle_dataframe_row_multiindex_*` suite (scoped to shipped operations).
- **Update 2026-09-06 (br-frankenpandas-wfkzm):** the last open symptom in this file is closed —
  multi-key `groupby(...).sum()` / `agg_multi` / `size` flat index labels now use the oracle's
  `tuple_label_to_flat_string` pipe-joined spelling (`"north|apple"`) instead of `", "`-joined,
  matching the `row_multiindex` levels the aggregation already attached. All 9
  `live_oracle_dataframe_row_multiindex` tests pass against live pandas 2.2.3. Broader
  slicing/xs/droplevel coverage remains under br-frankenpandas-1zzp.
- **Review date:** 2026-04-23

### DISC-007: SQL IO is SQLite-only; pandas supports multiple backends
- **Reference:** pandas `read_sql` / `to_sql` accept any SQLAlchemy-compatible backend (SQLite, PostgreSQL, MySQL, Oracle, MSSQL, etc.).
- **Our impl:** fp-io's `read_sql` / `write_sql` only accept a `rusqlite::Connection`. PostgreSQL / MySQL / Oracle connectors not shipped.
- **Impact:** Users whose pipelines depend on non-SQLite backends cannot drop-in replace pandas IO calls.
- **Resolution:** INVESTIGATING - tracked by br-frankenpandas-fd90 (SQL backend epic, 7 slices). SQLite remains the supported scope until slices 2+ land.
- **Tests affected:** `live_oracle_sql_*` suite (SQLite only).
- **Review date:** 2026-04-23

### DISC-008: No Python bindings shipped; pandas' Python-level drop-in positioning differs
- **Reference:** pandas IS a Python library. Users `import pandas`.
- **Our impl:** frankenpandas is a Rust library. Users `use frankenpandas::*` from Rust code. Python bindings (e.g. via PyO3) are not shipped.
- **Impact:** README's "drop-in pandas replacement" positioning applies at the API-shape level, not at the import-statement level. A Python pandas user cannot adopt frankenpandas without first porting to Rust.
- **Resolution:** ACCEPTED - README was updated to qualify the claim (br-frankenpandas-diic closed by wording rather than bindings). PyO3 bindings remain a future-epic candidate, not in scope for 0.1.0.
- **Tests affected:** N/A - positioning / documentation divergence, not behavioral.
- **Review date:** 2026-04-23

### DISC-009: Sparse dtype descriptor exists before compressed sparse storage
- **Reference:** pandas `SparseDtype` pairs an underlying value dtype with a fill value and stores only non-fill positions in `SparseArray`.
- **Our impl:** `fp-types::SparseDType` records the dtype/fill-value contract and `DType::Sparse` marks the logical dtype. fp-columnar still stores columns densely and IO falls back to textual sparse markers until a compressed sparse column representation lands.
- **Impact:** Code can now describe sparse dtype intent, but memory usage and `Series.sparse` accessor parity still differ from pandas.
- **Resolution:** WILL-FIX - remaining storage/accessor work tracked by br-frankenpandas-0xcm follow-up slices.
- **Tests affected:** Sparse storage/accessor conformance tests not yet enabled.
- **Review date:** 2026-04-24

### DISC-010: Rust GroupBy.apply uses explicit output-shape APIs
- **Reference:** pandas `DataFrameGroupBy.apply` dynamically dispatches scalar, Series, and DataFrame return values from one Python callable.
- **Our impl:** Rust's static return types expose the same shape families as explicit methods: `apply_scalar`, `apply_series`, `apply_series_stacked`, and DataFrame-returning `apply`. DataFrame-returning apply retains group-key row MultiIndex metadata; stacked Series output is represented as a one-column DataFrame until Series row MultiIndex metadata lands.
- **Impact:** Shape semantics are available, but Rust callers choose the expected output family at compile time instead of receiving a dynamic Python object.
- **Resolution:** ACCEPTED for Rust (static return types keep the shape-explicit methods); RESOLVED in Python. The binding's `DataFrameGroupBy.apply` / `SeriesGroupBy.apply` infer the shape per call as pandas' `_wrap_applied_output`: scalars become a Series over the group keys (NaN for None); Series sharing one index become a frame, a row per group; other Series and frames are concatenated under the group-key levels when `group_keys` (the default), else back in the original row order when every result kept its group's rows, else in group order; `include_groups` (default True, with pandas' DeprecationWarning for column keys) decides whether func sees the grouping columns (br-frankenpandas-rc0923-epic-rust-parity-bugs-4qg5w.10).
- **Tests affected:** `dataframe_groupby_apply`, `dataframe_groupby_apply_scalar_returns_series_indexed_by_keys`, `dataframe_groupby_apply_series_unions_sparse_result_columns`, `dataframe_groupby_apply_series_stacked_preserves_variable_labels`; Python: the `apply *` / `sgb apply *` cases of `crates/fp-python/tests/test_differential_packets.py::test_centered_windows_reindex_fill_dayfirst_and_groupby_apply_match_pandas`.
- **Review date:** 2026-09-25

### DISC-015: memory_usage exact bytes differ from pandas (structural divergence)
- **Reference:** pandas `DataFrame.memory_usage()` reports exact bytes consumed by numpy-backed columns. For the test frame in `FP-P2D-364`, pandas returns 234 bytes (index + column overhead + numpy array backing).
- **Our impl:** FrankenPandas uses `Vec<Scalar>` storage which has structurally different memory characteristics. The same frame reports 32 bytes — a 7x difference reflecting heap-allocated scalars vs numpy's contiguous buffer layout.
- **Impact:** Conformance packet `FP-P2D-364` (`dataframe_memory_usage_with_nulls_hardened`) fails with `actual=32, expected=234`. This is NOT a bug but a fundamental structural difference.
- **Resolution:** ACCEPTED — exact-byte parity is impossible without adopting numpy's physical layout. Documented in README Memory Model section. Relative/shape assertions remain valid (larger frames use more memory monotonically). Excluded from parity-score numerator via fixture waiver.
- **Tests affected:** `FP-P2D-364`, any exact memory_usage comparison tests.
- **Review date:** 2026-05-25
- **Waiver:** Signed by user request per br-frankenpandas-rg8ys.5.2.

### DISC-011: the Rust constructors infer ints with a missing value as an `Int64` holding it (pandas: float64 NaN, or the nullable `Int64`)
- **Reference:** Pandas (since v0.24) has a nullable `Int64` extension dtype (capital I) that preserves the integer encoding via a separate validity mask. When a non-nullable `int64` column receives a null (e.g. via index alignment introducing rows with no source data, or via `concat(axis=1)` aligning over a non-matching index), pandas can either preserve `Int64` (extension) or promote to `float64` depending on dtype. The conformance oracle uses extension `Int64` where the column was originally `int64`.
- **Our impl:** No nullable extension Int64 dtype yet. Int64 columns that gain null values are promoted to `Float64` with `NaN`. Downstream IO (JSON, CSV) then serializes the integer values with a trailing `.0` (`1.0` rather than `1`).
- **Impact:** Several conformance packets exhibit `actual=Float64(1.0), expected=Int64(1)` mismatches:
  - `FP-P2D-028` (dataframe_concat_axis1): 5 of 10 cases fail because alignment over a wider index introduces nulls into formerly-Int64 columns.
  - `FP-P2D-433` (dataframe_to_json_records): JSON output writes `"a":1.0` instead of `"a":1` for integer columns that were promoted via null introduction.
  - Plus other downstream packets where alignment + nulls hit Int64 columns.
- **⚠️ CORRECTION 2026-08-06 (br-frankenpandas-fixture-divergence-triage-9s0c4): the "Our impl" line above is NOT true of every path, and the difference decides whether ~97 fixtures get regenerated.** On the **merge and concat** paths FrankenPandas does NOT promote — it keeps `Int64` and carries the missing value in the validity mask, i.e. exactly the `int64 + null` the fixtures pin. Measured: `packet_filter_runs_dataframe_merge_sort_packet` (FP-P2D-037, a `how="right"` merge with an unmatched row) PASSES against a fixture pinning `left_v -> int64 [10, null, 30]`, and the whole `fp-conformance --lib` suite is green against fixtures of this shape. Live pandas 2.2.3 on the identical input returns `left_v -> float64 [10.0, NaN, 30.0]`; same story for `concat(axis=0, join="outer")`, where pandas gives `float64 [NaN, NaN, 300.0]` and the fixture pins `int64 [null, null, 300]`. So the divergence is real and is this DISC, but the mechanism is "FP represents Int64-with-nulls where pandas cannot" rather than "FP promotes like pandas does". Whoever implements the nullable-Int64 epic should re-derive which paths promote and which do not before trusting the original sentence.
- **Fixture-corpus impact, and why these must NOT be regenerated:** the live-pandas differ attributes **97 of its 181 divergent rows** to this one cause — 46 spelled `int64` vs pandas `float64`, 51 spelled `null` vs pandas `NaN`; they are the same phenomenon seen once in the dtype and once in the missing-value marker. Those fixtures pin the extension-`Int64` behavior this DISC is WILL-FIXing toward. Regenerating them to `float64 + NaN` would make the divergence vanish from the corpus and delete the pinned evidence of a tracked architectural gap — the golden-regeneration reflex at scale. They stay as they are, now with a named cause instead of an unexplained bucket.
- **⚠️ THE HARNESS SIDE, added 2026-08-15 (br-frankenpandas-vprpg) — this is what makes `FP-P2D-056` red, and two agents have now misdiagnosed it.** The gap is not only in FrankenPandas' storage; it is also in how the conformance harness BUILDS a fixture payload. `pandas_oracle.py::series_dtype_for_payload_values` maps an `{int64} + null` payload to the nullable `Int64`, while the Rust harness calls `Column::from_values`, whose `infer_dtype` skips nulls and yields plain `Int64`. So the two sides are handed **different dtypes for the same payload**. `fp_p2d_056_dataframe_merge_asof_backward_nan_left_key_strict` expects `incompatible merge keys` and FP merges instead — not because fp-join is wrong (it has a real `DType::Int64Nullable` and `br-frankenpandas-sopel` at `870a1003c` already implements and tests that exact rejection) but because the key lane never becomes nullable on FP's side. **Do not "fix" this in fp-join.** Mirroring the oracle's chooser in the harness was MEASURED and REJECTED on 2026-08-15: it left `FP-P2D-056` still failing (that fixture's data is in `frame`/`frame_right`, so `build_series` is never called for it) and regressed four other cases, because FP's kernels genuinely answer differently for `Int64Nullable` than for `Int64` — `series_clip_with_nulls_hardened` moved even though it had been re-banked on the `Int64` lane hours earlier under `br-frankenpandas-p5nku`. Full row in `docs/NEGATIVE_EVIDENCE.md`. `FP-P2D-056` therefore stays red as a named consequence of this DISC rather than an unexplained failure.
- **NARROWED 2026-09-27 (br-frankenpandas-rc0923-epic-rust-parity-bugs-4qg5w.11), measured against live pandas 2.2.3.** Every pandas-observable path now promotes as pandas does: a numpy int64 column that gains a missing value is float64 with NaN, and a nullable Int64 / Float64 / boolean keeps its dtype with an NA. Python binding, int64 and Int64 each: constructors (Series; DataFrame from a dict, tuples, a generator, records, list rows), reindex, alignment add, `align(join="outer")`, shift, concat (axis 0 and 1), merge outer / left, `read_csv`, `read_json` records, `where` / `mask`, `unstack`, `pivot`, setting with enlargement, `diff`, `pct_change`, `rolling`, groupby `shift` (the 2026-09-27 probe: 33 of 34 rows, the 34th `read_csv(skip_blank_lines=False)`, a refused parameter, not a promotion; the DataFrame constructor variants, `align`, concat axis 1, merge left and `pivot` from d55343759's probes, 26 of 27 and 21 of 21). The nullable dtypes keep their NA through concat, `where` / `mask` (Series and DataFrame), `diff` (boolean as xor, Int64 wrapping as numpy), `pct_change` (Float64; boolean raises pandas' `NotImplementedError`), `unstack`, `read_csv(dtype=...)` and setting with enlargement, and a write the dtype cannot hold raises pandas' `TypeError: Invalid value '2.5' for dtype Int64` (these came back float64, a plain int64 holding the NA, or upcast). Rust API: reindex, outer alignment, shift, concat and outer merge widen int64 to float64 NaN and keep `Int64Nullable`. **What remains is the Rust constructor**: `Column::from_values` / `Series::from_values` / `DataFrame::from_dict` infer ints with a missing value as `Int64` holding it, a column pandas never produces (`pd.Series([1, None])` is float64). It is the representation the harness builds fixture payloads with and the ~97 fixtures above pin, so changing it is one decision across the constructor, the harness and the corpus, not a per-path fix: br-frankenpandas-ih6ho. Residual edge: a masked Float64 keeps `0/0` as a NaN that is not NA (`pct_change` over two zeros); `Float64Nullable` cannot hold that, so FrankenPandas answers NA.
- **Resolution:** NARROWED (was WILL-FIX per br-frankenpandas-mywg / fd90.76): the promotion policy is pandas' on every observable path; the Rust constructors' int64 + missing inference remains, tracked by br-frankenpandas-ih6ho.
- **Tests affected:** `packet_filter_runs_dataframe_concat_axis1_packet`, `packet_filter_runs_dataframe_to_json_records_packet`, `fuzz_json_io_bytes_accepts_records_seed_fixture` (the records seed has `[{"temp":72},{"temp":null}]` — read promotes to Float64, write emits `72.0` instead of `72`, reparse + diff detects the drift), plus other downstream packets that hit the same root cause. The per-path parity is pinned by `null_introduction_per_path_matches_pandas_4qg5w_11` and `nullable_dtypes_keep_their_na_4qg5w_11` (fp-frame), `merge_outer_gap_widens_int64_and_keeps_int64_nullable_4qg5w_11` (fp-join), `csv_nullable_dtype_keeps_the_na_4qg5w_11` (fp-io) and pytest `test_nullable_dtypes_keep_their_na_like_pandas` (50 cases).
- **Review date:** 2026-09-27 (narrowed to the constructor; was 2026-08-15, harness-side cause and the rejected chooser added; 2026-08-06, itself a correction of 2026-04-26)

### DISC-018: Timedelta/Timestamp arithmetic overflow surfaces as NaT (pandas raises)
- **Reference:** pandas 2.2.3 is **not uniform** here, which is the load-bearing fact. Probed live against the installed oracle:
  - **RAISES** — `Series + Timedelta`, `Series - Series`, `Series.diff` → `OverflowError: Overflow in int64 addition`; `Series.sum()`, `Series.std()` → `ValueError: overflow in timedelta operation`; `Timedelta.max + Timedelta('1ns')` → `OverflowError`; `Timestamp.max + Timedelta('1ns')` → `OutOfBoundsDatetime`; `Series([tmin,tmax]).max() - .min()` (the ptp expression) → `OverflowError`.
  - **WRAPS SILENTLY** — `Series.cumsum()` and `Series * 2` wrap modulo 2^64 with no exception (raw numpy int64), e.g. `Timedelta.max * 2` → `-1 days +23:59:59.999999998`. `Period` arithmetic likewise wraps with no check at all.
  - **Returns NaT** — `Series([tmax, tmax]).mean()` and `.median()` and `.quantile(0.5)` (pandas' own internal float cast overflows); `Series * 2.0` with a **float** multiplier (contrast the int multiplier one line above, which wraps); `Series / 0`, `Series // 0`, `Series / 0.0`, and the overflowing `Series / 0.5` — the whole vectorized division family; `Series([0, tmax]).sum()`, where the exact answer *is* representable but pandas' f64 detour rounds `2^63 - 1` up to `2^63` and the cast back to int64 lands on the NaT sentinel.
  - **AGREES WITH US BY SENTINEL COLLISION** — `Timedelta.min - Timedelta('1ns')` → NaT, scalar *and* vectorized, and likewise `Timestamp.min - 1ns`. This is not a policy: `Timedelta.min.value` is `i64::MIN + 1`, so one step below it lands exactly on `i64::MIN`, which pandas also reserves as its NaT sentinel. **Two** steps below (`min - 2ns`) raises `OverflowError`. Any test that pins only the 1 ns case proves nothing about this divergence — it passes for a correct and an incorrect implementation alike.
  - **FAILS THE WHOLE OPERATION, NOT THE OFFENDING ELEMENT** (br-frankenpandas-lrgdu, measured 2026-08-16) — `pd.Series([0ns, Timedelta.max]) + Timedelta('1ns')` **raises**, even though row 0 is perfectly representable. pandas aborts the entire column operation and returns nothing; FrankenPandas returns `[1ns, NaT]`, recovering element-locally. So on the elementwise surface the divergence is not merely "NaT where pandas raises" — it is **element-local recovery vs whole-operation failure**, and the two produce different *shapes* of answer, not just different values in one slot. This is the row that constrains the STRICT arm below: a per-element raise inside the column map cannot reproduce pandas, because pandas never gets as far as producing a partial column. Applies to the `Series ± Timedelta` / `Series ± Series` / `Series.diff` group; it does not arise for the wrapping or NaT-returning groups, which have no failure to localise.
  - **NOT AN OVERFLOW AT ALL** — `-Timedelta.min`, `abs(Timedelta.min)`, and their vectorized forms return `Timedelta.max`. pandas' range is symmetric about zero precisely because `i64::MIN` is spent on NaT. FrankenPandas matches, and `neg`/`abs` carried a doc comment claiming the opposite until br-frankenpandas-fyr1z corrected it.
- **Our impl:** FrankenPandas returns **NaT** wherever the exact result is unrepresentable, uniformly, and keeps the helpers infallible (`#[must_use]`, no `Result`). This covers the scalar boundary (`Timedelta::{add,sub,mul_scalar}`, `Timestamp::{add_timedelta,sub_timedelta}`, `Timedelta::from_unit`) and the vectorized reduction surface (`nansum`, `nanptp`, `nancumsum`). The representable range is `[i64::MIN + 1, i64::MAX]`, since `i64::MIN` is the NaT sentinel. On the **elementwise** column paths the recovery is **per element**: `crates/fp-columnar/src/lib.rs:30810` and `:13952` call the infallible `fp_types::Timedelta::add`/`sub` (`try_add(..).unwrap_or(NAT)`), so a single overflowing row becomes NaT while its neighbours keep their exact values. The fallible carrier `Timedelta::try_add` and `OverflowPolicy` exist from `br-frankenpandas-fyr1z-policy-carrier-f29nz` but nothing in the elementwise paths consults them yet.
- **Impact:** A caller who overflows gets a missing value where pandas would raise, and NaT is indistinguishable from a genuinely missing input. This is a deliberate trade, not an oversight: the alternative previously in the tree was **saturation**, which fabricated a finite ~106751-day Timedelta and presented it as real data — fail-open, and strictly worse than either pandas behavior. Fail-closed-as-missing beats fail-open-as-plausible-data. Where pandas *wraps*, we deliberately do not chase the wrap (wrapping is not better behavior than NaT, and matching it would need its own justification).
- **Resolution:** ACCEPTED for the current API shape. The observability decision (`br-frankenpandas-fyr1z`) is now **DECIDED: a per-op strict/hardened mode split**, with one correction that the decision turns on — **a blanket "STRICT raises" is NOT pandas parity and would make things worse.** pandas refuses only on part of this surface; on the rest it returns NaT itself (division family, float multipliers, mean/median/quantile, the 1 ns sentinel step) or silently wraps (int multiplier, `cumsum`). Making STRICT raise uniformly would *introduce* divergence in every row of the "Returns NaT" and "agrees by sentinel collision" groups above, where FrankenPandas is already bit-for-bit correct. STRICT must therefore be specified **per operation** against the measured table above, not as a global policy; HARDENED keeps today's uniform NaT plus an audit-log entry. Two consequences worth stating plainly: STRICT parity for `Series * 2` and `cumsum` means **reproducing a silent two's-complement wrap**, which is fail-open behavior the repo otherwise forbids and which needs its own explicit sign-off; and the divergence surface that actually needs new code is much smaller than "every overflow". Escalating the fp-types helpers to `Result` (option (b)) is rejected as the primary mechanism — it would force ~12 elementwise call sites to re-litigate raise-vs-propagate while fixing none of the vectorized rows, which is where users meet this. Precedent: `Timedelta::div_scalar` returns NaT on divide-by-zero, which the scalar pandas API raises on (`ZeroDivisionError`) but the vectorized API agrees with.
- **Tests affected:** `nansum_timedelta_overflow_is_nat_not_fabricated_max_opz27`, `nansum_timedelta_negative_overflow_is_nat_opz27`, `nanptp_timedelta_overflow_is_nat_opz27`, `nancumsum_timedelta_overflow_recovers_exact_value_opz27`, `nancumsum_timedelta_exact_path_unchanged_opz27`, `arithmetic_overflow_never_fabricates_a_finite_value_lgyy8` (the divergence half), `overflow_divergence_surface_is_only_where_pandas_refuses_fyr1z` (the agreement half plus the 1-step/2-step sentinel pair), `neg_and_abs_at_the_representable_floor_are_not_overflow_fyr1z` (all in `fp-types`). Beads: `br-frankenpandas-lgyy8`, `br-frankenpandas-8v92m`, `br-frankenpandas-opz27`; decision recorded by `br-frankenpandas-fyr1z`, implementation decomposed into its children.
- **Review date:** 2026-08-06

### DISC-019: Datetime64 mean/median/quantile are exact in FrankenPandas; pandas snaps to its f64 grid
- **Reference:** pandas 2.2.3 does not compute datetime64 `mean`/`median`/`quantile` on the int64 nanosecond backing — it detours through float64. At realistic timestamps this loses precision outright, because the f64 ULP at 2020-01-01 (1.578e18 ns) is **256 ns**. Probed live:
  - `Series([B+1, B+1]).mean()` → `B+0`. A **constant** series, where no averaging is needed and the answer is unambiguous, still comes back wrong by 1 ns.
  - `Series([B, B+129, B+258]).median()` → `B+256`. An odd-count median is a pure **selection** of an existing element, yet the returned Timestamp **is not one of the inputs**.
  - `quantile(0.5)` on the same input → `B+256`.
  - Rounding on that 256 ns grid is ties-to-even (`B+128` → `B+0`, `B+129` → `B+256`).
  - Root cause reproducible directly: `s.values.view('i8').mean()` returns `1.5778368e+18`, dropping the low bits, while `i8.sum() // 2` is exact. By contrast `min`/`max` are exact, since they are integer operations.
- **Our impl:** FrankenPandas computes `nanmean`/`nanmedian` for Datetime64 in **i128 nanoseconds, exactly**, and narrows with Rust's toward-zero integer division. Where pandas is exact (small magnitudes), FP matches it bit-for-bit **including the rounding rule**, which was probed and is truncation toward zero — `[0,1]→0`, `[0,3]→1`, `[-1,0]→0`, `[-3,0]→-1`, `[1,2]→1`. Both `floor` (would give `-1` for `[-1,0]`) and ties-to-even (would give `2` for `[1,2]`) are refuted. Where pandas snaps to its f64 grid, FP stays exact. `nanquantile` deliberately keeps the f64 hop its Timedelta64 sibling already used, because quantile interpolates by a fractional weight; its Datetime64 arm inherits the same accuracy contract rather than inventing a stricter one for a single dtype.
- **Impact:** For datetime columns at modern timestamps, FP's mean/median can differ from pandas by up to ~128 ns. FP is the exact one in every such case. The divergence is invisible at second/day granularity, which is where the overwhelming majority of real datetime data sits.
- **Resolution:** ACCEPTED — deliberate, and consistent with precedent already in the tree. `nanptp`'s Timedelta64 arm carries the comment "Converting to f64 before subtracting loses ranges below one f64 ULP beyond 104 days", i.e. FrankenPandas had already rejected the f64 detour for temporal data; this extends the same choice to Datetime64. Reproducing pandas here would mean shipping a median that returns a timestamp the user never supplied. Same governing principle as DISC-018: where pandas is lossy, do the correct thing and document it rather than chase the defect.
- **Related, and NOT implemented:** pandas **raises** `TypeError: 'DatetimeArray' with dtype datetime64[ns] does not support reduction '<name>'` for datetime64 `sum`, `prod`, `var`, `sem`, and `skew`. FrankenPandas returns a missing value for these rather than inventing one; it deliberately does not add support pandas itself refuses.
- **`std` follows pandas exactly and is NOT part of this divergence** (br-frankenpandas-40ujm): pandas *does* support datetime64 `std`, returning a **Timedelta** — the dispersion of a set of instants is a duration. FP matches it bit-for-bit on the raw ns, including `ddof` handling (`n - ddof <= 0` → NaT) and pandas' toward-zero truncation of the sqrt. Verified: pandas returns the identical value for a Datetime64 series and a Timedelta64 series built from the same nanoseconds, which is why both families route through one computation. No exactness question arises here because a standard deviation involves a square root, so unlike mean/median there is no exact integer answer to prefer — FP and pandas are both computing in f64 and agree.
- **Tests affected:** `datetime64_averaging_reductions_match_pandas_adv58`, `datetime64_mean_median_round_toward_zero_adv58`, `datetime64_averaging_skips_nat_adv58`, `datetime64_averaging_is_exact_where_pandas_snaps_adv58` (all in `fp-types`). Beads: `br-frankenpandas-axhhk` (ordering reductions), `br-frankenpandas-adv58` (this).
- **Review date:** 2026-08-06

### DISC-020: STRICT mode deliberately reproduces pandas' silent timedelta int64 wrap
- **Reference:** pandas 2.2.3 wraps silently on the integer-multiplier and cumulative-sum timedelta paths — raw numpy int64 wraparound, no exception and no NaT. Probed: `Series([Timedelta.max]) * 2` → `-2` ns (displayed `-1 days +23:59:59.999999998`); `Series([max, max]).cumsum()` → `[9223372036854775807, -2]`; `Series([max, max, max]).cumsum()` → `[max, -2, max - 2]`, which shows the accumulator keeps wrapping rather than sticking. A **float** multiplier behaves differently: `Series * 2.0` returns NaT, so the wrap is integer-multiplier only.
- **Our impl:** under `OverflowPolicy::Strict` FrankenPandas **reproduces this exactly**, via `Timedelta::mul_scalar_with_policy` and `nancumsum_with_policy`. Rust's `wrapping_mul`/`wrapping_add` match numpy's int64 wraparound bit-for-bit, so no emulation is required. Under **every other policy — including the default `SurfaceNat`** — the result stays NaT, per DISC-018.
- **Impact:** this is the one place where FrankenPandas deliberately reproduces a behaviour it would otherwise classify as fail-open: multiplying a positive duration by 2 yields a *negative* duration presented as real data. That is the exact shape `br-frankenpandas-lgyy8` removed from this codebase. It is acceptable **only** because it is unreachable by default and requires a caller to explicitly ask for STRICT incumbent parity.
- **Resolution:** ACCEPTED — maintainer decision on `br-frankenpandas-fyr1z-wrap-signoff-s4lkx`, recorded verbatim: *"STRICT means bit-for-bit observable parity with the incumbent including its quirks, so reproduce the wrap and add a test that names it as a deliberately reproduced pandas behavior rather than a bug of ours."* This is a **port**, and STRICT's job is to be indistinguishable from the incumbent, not to be better than it. Note this deliberately differs in direction from DISC-019, where FrankenPandas chose exactness over pandas' f64 grid — the distinction is that DISC-019 concerns the *default* path, whereas this quirk is opt-in.
- **Tests affected:** `strict_reproduces_pandas_silent_timedelta_wrap_s4lkx` (pins the oracle values and states in its doc comment that a failure means pandas changed or a non-STRICT caller was misrouted — **not** that the arithmetic should be corrected), `non_strict_policies_still_refuse_to_wrap_s4lkx` (asserts the default and HARDENED still refuse to wrap, and that HARDENED still records the recovery).
- **Review date:** 2026-08-06

### DISC-027: Series.argsort uses stable sort order for ties; pandas delegates to numpy quicksort (unstable)
- **Reference:** pandas `Series.argsort()` delegates to numpy `kind='quicksort'`, which is unstable and leaves tie order dependent on numpy's internal quicksort partition implementation. On duplicate elements (e.g. `[2.5, 1.0, 2.5, 3.0, 1.0]`), pandas emits `[1, 4, 2, 0, 3]`, placing position 2 before position 0 for the 2.5 tie.
- **Our impl:** FrankenPandas `Series::argsort` uses stable pair-sorting (`Vec::sort_by`), preserving encounter order for tied elements (`[1, 4, 0, 2, 3]`), matching numpy `kind='stable'`.
- **Impact:** Duplicate values in argsort produce deterministic original index order rather than reproducing arbitrary numpy quicksort artifacts.
- **Resolution:** ACCEPTED (2026-09-08, `br-frankenpandas-dxkbb`). Option (a) (pinning numpy's undocumented quicksort artifact) is rejected as fragile tech debt. Conformance fixtures for argsort avoid ambiguous tie order across engines.
- **Tests affected:** `series_argsort` conformance fixtures avoid ties.
- **Review date:** 2026-09-08

### DISC-029: rank keeps a missing value missing where pandas 2.2.3 ranks what lies under it
- **Reference:** MEASURED, live pandas 2.2.3, two shapes. (1) GroupBy rank of a `datetime64` / `timedelta64` column ranks `NaT` as the SMALLEST value instead of as missing once some group KEY is missing: `DataFrame({"k": ["x", "x", "x", None], "v": to_datetime(["2020-01-03", None, "2020-01-02", "2020-01-01"])}).groupby("k")["v"].rank()` → `[3.0, 1.0, 2.0, NaN]`, and `na_option` has no effect on it; with every key present the same rank gives `[2.0, NaN, 1.0]`, as `Series.rank` does. (2) `Series.rank` of a nullable `Int64` column ranks each `<NA>` by the placeholder stored under its mask: `Series([3, None, 2, 2, 1, 0], dtype="Int64").rank()` → `[6.0, 2.5, 4.5, 4.5, 2.5, 1.0]` (the NA ties with the 1); the same column's groupby rank masks it (`<NA>`).
- **Our impl:** a missing value is missing in every rank: `NaT` and `<NA>` take `na_option` (NaN under `keep`), in Series, SeriesGroupBy and DataFrameGroupBy rank alike (`rank_sort_keys` / `rank_row_keys`, fp-frame).
- **Impact:** those two shapes only; every other rank of datetime, timedelta and nullable data agrees (probe `p12/probe_rank_dtypes.py` of `br-frankenpandas-rc0923-epic-python-honest-dropin-fvsao.71`: the remaining rows are these and the nullable groupby OUTPUT dtype, `Float64` in pandas and `float64` here, a separate masked-dtype gap). Reproducing (1) would rank a missing instant ahead of every real one, against pandas' own `Series.rank` and its NaT semantics; reproducing (2) needs the placeholder pandas happened to store, which depends on how the array was built.
- **Resolution:** ACCEPTED (2026-09-27, `br-frankenpandas-rc0923-epic-python-honest-dropin-fvsao.71`). Revisit if a pinned pandas release masks both.
- **Tests affected:** `test_rank_of_a_missing_value_pandas_does_not_mask` (its two strict xfail cases) and `test_rank_keeps_the_missing_values_pandas_ranks_missing` (fp-python pytest).
- **Review date:** 2026-09-27

### DISC-030: `to_string(max_cols=, formatters={...})` formats the column each key names
- **Reference:** MEASURED, live pandas 2.2.3: a formatters MAPPING is looked up by the column's position in the truncated frame, mapped through the FULL column list (`DataFrameFormatter._get_formatter`: `i = self.columns[i]` when the position is not itself a label). Past `max_cols` a formatter lands on the wrong column: `DataFrame({f"c{i}": [i] for i in range(6)}).to_string(max_cols=2, formatters={"c1": "<{}>".format})` prints `<5>` under `c5`, and `formatters={"c5": ...}` formats nothing.
- **Our impl:** fp applies each formatter to the column its key names, truncated or not (a formatters LIST is truncated with the columns, as pandas does). Without `max_cols` truncation the two agree.
- **Impact:** only a formatters mapping together with a `max_cols` that truncates: the named column prints formatted, its header without the numeric sign space (as pandas prints a formatted column).
- **Resolution:** ACCEPTED (2026-09-27, br-frankenpandas-xn05q). Reproducing it would format one column's values with another column's formatter.
- **Tests affected:** `test_to_string_formats_the_named_column_past_max_cols` pins fp's behavior (the differential table `test_to_string_keywords_like_pandas` covers a truncated formatters LIST instead).
- **Review date:** 2026-09-27

### DISC-031: `query` and `combine_first` answer over repeated column names, where pandas 2.2.3 fails
- **Reference:** MEASURED, live pandas 2.2.3, `DataFrame([[1.5, 9.0, 3.0, 1], ...], columns=['a', 'a', 'b', 'k'])`: `.query('k > 1')` raises `TypeError: dtype 'a float64 a float64 dtype: object' not understood` (its resolvers hand the repeated name a DataFrame where a Series is meant), and `.combine_first(d.fillna(0))` raises `AttributeError: 'DataFrame' object has no attribute 'dtype'` - internal failures, not a documented refusal.
- **Our impl:** fp keeps each repeated column apart (br-frankenpandas-i17d4) and answers both as pandas answers the same frame with its names made unique, the names given back.
- **Impact:** code that fails in pandas answers in fp; nothing that works in pandas answers differently.
- **Resolution:** ACCEPTED (2026-09-29, br-frankenpandas-i17d4). Reproducing an internal pandas failure would buy no compatibility.
- **Tests affected:** `test_repeated_names_where_pandas_fails_answer_as_unique_names` checks fp against pandas on the uniquely named frame and asserts pandas' own failure, so a pandas fix shows there.
- **Review date:** 2026-09-29

## Resolved Divergences

### DISC-025: `str.encode` returns byte LENGTHS, where pandas returns bytes objects
- **⚠️ RESOLVED 2026-09-27 (`br-frankenpandas-rc0923-epic-rust-parity-bugs-4qg5w.8`).** `str.encode` returns real bytes: fp-types has a bytes object cell, `ObjectValue::Bytes`, which the Python binding reads and writes as Python `bytes` (they were host objects). `StringAccessor::encode` gives each str as its bytes, keeps a missing value as it is and the name, NaN for any other value; `decode` reads bytes cells and gives NaN for anything else, a str included (measured, live pandas 2.2.3: `Series(['ab']).str.decode('utf-8')` is float64 `[nan]` - the AttributeError quoted below is not what 2.2.3 raises for an object Series). The core encodes UTF-8, ASCII and Latin-1 strictly with Python's error texts; the binding uses Python's own codecs (any encoding, `errors=`). The oracle had been answering FrankenPandas' question (`.str.encode('utf-8').str.len()`, decode as identity); it now asks pandas', and the four encode / decode fixtures were regenerated per fixture under `artifacts/audits/qg5w8_disc025_str_encode_decode_attribution_2026-09-27.json` (0 WIN / 4 LOSE against the pre-change code, 4 WIN after). The historical entry below is kept as written.
- **Reference:** MEASURED, live pandas 2.2.3 on `pd.Series(['foo','hello world',None], dtype=object)`: `s.str.encode('utf-8')` → `[b'foo', b'hello world', None]`, dtype `object`, element types `bytes`/`bytes`/`NoneType` — a bytes object per element, with a supplied `None` preserved. For contrast `s.str.len()` → `[3.0, 11.0, nan]`, dtype `float64`.
- **Our impl:** `StringAccessor::encode` (`crates/fp-frame/src/lib.rs`) is `apply_str_int(|s| s.len() as i64)` — it returns the UTF-8 byte LENGTH, `Int64` with no nulls and `Float64`/NaN when any null is present. Rust strings are always valid UTF-8, so the encoding argument is a genuine no-op; what is not a no-op is returning a number where pandas returns bytes. It is `str.len` under another name, and the two results agree on nothing except arity.
- **Impact:** any caller treating `encode` as pandas' `encode` gets a length, not the bytes. The null MARKER is correct given this implementation — a numeric result cannot hold a `None`, so NaN is right by the `br-frankenpandas-lufpu` result-dtype rule — which is why `fp_p2d_413_series_str_encode_with_nulls_hardened` legitimately pins float64 values. The fixture is therefore consistent, but it pins a divergence as expected behaviour, the same shape as DISC-009 and `br-frankenpandas-0g9m9`.
- **Resolution:** WILL-FIX / INVESTIGATING — decision tracked by `br-frankenpandas-rw01l`. Three options: return real bytes (needs somewhere to put them; FP has no bytes dtype, so this leans on the DISC-005 object bucket or a new one), refuse `encode` explicitly as unsupported (honest, and loses nothing since the current answer cannot be consumed as bytes), or keep byte lengths as a documented divergence. This entry is the third option's minimum and stands regardless of which is chosen.
- **Found by:** verifying `74910f3eb`'s re-bank under `br-frankenpandas-q0ktc`. That re-bank is correct about the null markers; the implementation it reasons from is what nobody had checked. The in-code comment asserted "matches pandas", which is what let it sit unnoticed — now corrected in place.
- **Tests affected:** `fp_p2d_413_series_str_encode_with_nulls_hardened` (consistent with the current implementation, not with pandas). `decode()` is the sibling op: in pandas 2.2.3, `s.str.decode('utf-8')` requires a Series of `bytes` objects (raising `AttributeError: Can only use .str.decode with 'bytes' dtype!` on string inputs) and returns `str`; FrankenPandas currently treats `decode()` on `Utf8` strings as an identity clone.
- **Review date:** 2026-08-16 (updated 2026-09-09)

### DISC-005: Mixed string/numeric constructors now preserve pandas object semantics
- **Reference:** `pd.Series(["x", 1])` and `pd.concat([pd.Series(["x", 1])], axis=1)` preserve heterogeneous values under pandas `object` dtype
- **Our impl:** Constructor inference now uses the existing `Utf8` storage bucket for pandas-style object columns while preserving heterogeneous `Scalar` payloads in order
- **Impact:** `Series::from_values` and `DataFrame::from_series` now match pandas for mixed string/numeric constructor inputs
- **Resolution:** ACCEPTED - parity achieved and covered by live-oracle plus fixture-backed tests
- **Tests affected:** `live_oracle_series_constructor_mixed_utf8_numeric_reports_object_values`, `live_oracle_dataframe_from_series_mixed_utf8_numeric_matches_object_values`, `series_constructor_utf8_numeric_object_strict`, `dataframe_from_series_utf8_numeric_object_strict`
- **Review date:** 2026-04-15

### DISC-013: Series + Series union alignment does not sort the result index
- **Reference:** Pandas `Series.add(other)` (and `series + other` operator) on differently-indexed Series performs an outer-join alignment that returns a sorted result index by default.
- **Our impl:** RESOLVED - unique-label Series arithmetic now uses a sorted outer union for `+` / `-` / `*` / `/` and fill-value arithmetic, while preserving the duplicate-aware cross-product path tracked separately by DISC-014.
- **Impact:** The `FP-P2C-001 series_add_alignment_union_strict` fallback fixture has been refreshed to pandas 2.2.3 output: result index `[1, 2, 3]`, values `[NaN, NaN, 34.0]`.
- **Resolution:** RESOLVED in br-frankenpandas-cod1d13 by routing unique-label Series arithmetic through sorted union alignment in fp-frame instead of changing the generic fp-index discovery-order helper. NB: the listed strict test still fails today, but for a different root cause (DISC-011 nullable-Int64 dtype promotion); the sort-order issue this entry tracked is no longer present.
- **Tests affected:** `series_add_aligns_on_union_index`, `series_add_fill_sorts_unique_outer_union_index`, `FP-P2C-001/series_add_alignment_union_strict`.
- **Review date:** 2026-04-28

### DISC-014: Series + Series duplicate-label arithmetic Int64 promotion (prior WILL-FIX premise was incorrect)
- **Reference:** Pandas `Series + Series` with duplicate labels performs cross-product alignment per label. The result stays `int64` when every label matches on both sides (no unmatched pairing, so no NaN is introduced); it promotes to `float64` only when an unmatched label actually injects a NaN.
- **Our impl:** Matches pandas exactly — the duplicate-aware cross-product keeps `Int64` when no NaN is generated and promotes to `Float64` when alignment leaves a position unmatched.
- **Impact:** None. The earlier entry claimed pandas *always* promotes duplicate-label results to `Float64` even with no NaN; that is false. Verified against the pandas 2.2.3 live oracle: `Series([1,2,3], index=['a','a','b']) + Series([3,4,5], index=['a','a','b'])` returns `int64 [4,6,8]`, while a partial match (`index=['a','a']` + `index=['a','a','b']`) returns `float64` with a trailing `NaN`. The fixture `fp_p2c_001_duplicate_hardened.json` already expects `int64 [4,5]` and the conformance test `conformance_series_add_duplicate_labels` passes. The previously-proposed "always promote to Float64 when alignment *can* introduce NaN" fix would have *broken* parity for the fully-matched case and must not be implemented.
- **Resolution:** RESOLVED - no code change required; FP already matches pandas. This entry corrects the prior incorrect WILL-FIX premise, which conflated this case with the genuinely-open DISC-011 (Int64 column that actually *receives* a null). Oracle-verified 2026-06-01.
- **Tests affected:** `conformance_series::conformance_series_add_duplicate_labels` (passing).
- **Review date:** 2026-06-01

### DISC-016: RangeIndex set-operation result ordering (all four ops RESOLVED)
- **Reference:** pandas 2.2.3 `RangeIndex` set operations use different default `sort` semantics:
  - `union` / `difference` / `symmetric_difference`: `sort=None` — result is ascending-sorted EXCEPT when an operand is empty or the two operands are value-equal, where the surviving operand passes through unchanged (order preserved).
  - `intersection`: `sort=False` — result is ascending EXCEPT when BOTH operands are descending (`step < 0`), where it is descending. (Verified vs pandas 2.2.3 over 150k random pairs.)
- **Our impl:** RESOLVED for all four. Previously fp returned every set op in self/discovery order (matching pandas only where the affine fast paths happened to already be in pandas' order), diverging for both-non-empty operands with descending or non-aligned (interleaving) lattices — union/difference/symmetric_difference for ~any descending/interleaved input, and intersection for self-descending ∩ other-ascending (≈23.5% of multi-element intersections). fp now normalizes operands to their ascending equivalent so the fast paths and fallbacks produce pandas' order, sorts the interleaving union/symmetric_difference fallbacks, and routes both-descending intersection through the operands as-is (self-order yields descending). Empty-operand and value-equal passthrough preserved; lazy affine / two-affine-run backing retained for the aligned common case (only reordered cases materialize).
- **Impact:** all four ops now bit-match pandas across a 263-case randomized differential (descending, non-aligned, empty, value-equal, disjoint, subset), plus 150k-pair Python sweeps confirming the ordering rules.
- **Resolution:** RESOLVED (DustySummit, br-frankenpandas). union/difference/symmetric_difference in commit 88e2a8487; intersection ordering in the follow-up commit.
- **Tests affected:** `range_index_set_ops_match_pandas_2_2_3_differential_dustysummit` (data-driven from `testdata_rangeset_pandas_cases.rs`, all four ops); `range_index_set_ops_use_direct_values_b7nxg`, `range_index_set_ops_closed_form_membership_preserves_order_iatnc`, `range_index_set_ops_return_affine_spans_iatnc` (refreshed to pandas-correct ordering).
- **Review date:** 2026-07-23

### DISC-017: RangeIndex.slice_locs / searchsorted reject monotonic-decreasing (step < 0) ranges
- **Reference:** pandas 2.2.3 permits `RangeIndex(10,0,-1).slice_locs(8,3)` → `(2, 8)` and `RangeIndex(10,0,-1).searchsorted(5)` → `0` on descending ranges.
- **Our impl:** ACCEPTED divergence — fp returns an explicit `InvalidArgument` error for `slice_locs` and `searchsorted` when `step < 0` ("requires a monotonic[ally-]increasing RangeIndex"), rather than replicating pandas' behavior. get_loc / get_indexer / get_indexer_non_unique / reindex (direction-agnostic value→position lookups) DO work correctly on descending ranges; only the two ordering-boundary ops refuse.
- **Impact:** Rationale for refusing rather than matching: pandas' `searchsorted` on a descending index is a numpy artifact — numpy `searchsorted` assumes an ascending array, so `desc.searchsorted(5) = 0` is a meaningless insertion point, not a correct one. pandas' descending `slice_locs` inherits the same ascending-assuming `get_slice_bound`, producing quirky edge results (e.g. `[7].slice_locs(8,-9) = (1,0)` — an EMPTY slice for a value that is in `[-9, 8]`). A randomized 40k-pair differential found no simple closed form matching pandas' descending slice_locs (≈7.6% of a naive `value<=start` / `value<end` closed form diverged, all in boundary/single-element cases). fp's explicit error avoids silently reproducing these numpy quirks; callers needing a descending slice can reverse the index first.
- **Resolution:** ACCEPTED (DustySummit, 2026-07-23). Revisit only if a use case needs pandas-bug-compatible descending `slice_locs`; it would require replicating pandas' `get_slice_bound` direction handling, not a closed form.
- **Tests affected:** none (fp's error path is covered by existing RangeIndex tests; no descending slice_locs/searchsorted parity test is asserted).
- **Review date:** 2026-07-23

### DISC-022: `oracle_attestation` — what a provenance stamp on an EXPECTED-ERROR fixture claims
> **Renumbered 2026-08-16 (CalmMink).** This entry was a SECOND `DISC-018`, colliding
> with "Timedelta/Timestamp arithmetic overflow surfaces as NaT" above. All 13
> in-tree citations of `DISC-018` mean the overflow entry, so that one kept the
> number and this one moved. No citation anywhere referred to this entry.
- **Reference:** `fixture_provenance.oracle_script_sha256` normally asserts *"running this oracle script on this input reproduced these expected VALUES"*. An expected-error fixture pins no values, so the same stamp cannot mean the same thing there.
- **Our impl:** error fixtures whose refusal came from pandas now carry an explicit `fixture_provenance.oracle_attestation: "error_agreement"`, meaning only *"running this oracle script on this input made PANDAS raise too"*. **Absence of the key keeps the strong value-reproduction reading**, so nothing about the other 977 stamps changed. The message text is deliberately NOT part of the claim: `expected_error_contains` pins FrankenPandas's wording (checked by the Rust harness) while the oracle surfaces pandas' own English, and the two were never meant to match.
- **Impact:** 49 of the 85 expected-error fixtures qualify. The other **36 are deliberately left unstamped and still count as red**: 31 were refused by the oracle *adapter's own argument validation* (e.g. `dataframe_concat concat_axis must be 0 or 1, got 2`) and 5 escaped as `unexpected`, so **pandas was never invoked** and there is no agreement to attest — the claim would be true and vacuous. The split is structural, from `pandas_oracle.oracle_error_origin` (an `OracleError` raised `from exc` wrapped an engine call; a bare one is adapter validation), never a substring match on the message. The 31 adapter refusals are better read as oracle coverage gaps.
- **Resolution:** ACCEPTED (BlueRobin, 2026-08-08, br-frankenpandas-fixture-corpus-stale-vs-oracle-p6srr). The alternative — restamping all 85 under the bare key — was rejected as silently widening what `oracle_script_sha256` means corpus-wide to make a number go down. `scripts/check_fixture_freshness.sh` is untouched and reads only the three oracle keys via `.get()`, so a provenance **superset is already gate-legal**; that is also why the generator fixture `fp_generated_tn6qb2_...` now restamps in place with its `generation_command` / `input_matrix` / `intentional_divergence_notes` preserved instead of being refused.
- **Tests affected:** `oracle/tests/test_error_origin.py` (7 cases: origin classification, fail-closed default, provenance on the error path); `oracle/tests/test_regenerate_fixtures.py` (superset faithful / dropped-extras refused / changed-extra refused / stale-sha refused / undeclared-key refused / attestation required / text insertion).
- **Review date:** 2026-08-08

### DISC-023: the oracle builds NULLABLE dtypes for int+null / bool+null payloads, unlike pandas' own constructor
> **Renumbered 2026-08-16 (CalmMink).** This entry was a SECOND `DISC-019`, colliding
> with "Datetime64 mean/median/quantile are exact" above. Three of the four in-tree
> citations of `DISC-019` mean the Datetime64 entry, so that one kept the number.
>
> *(The former citation in `crates/fp-frame/src/lib.rs` has since been updated to `DISC-023` at line 57361).*
- **Reference:** pandas 2.2.3 infers `pd.Series([1, None, 3])` as **float64** (values `[1.0, nan, 3.0]`) and `pd.Series([True, None])` as **object**. Nullable `Int64` / `boolean` are reached only by asking for them explicitly.
- **Our impl:** `pandas_oracle.series_dtype_for_payload_values` returns `"Int64"` for an all-int payload containing a null, and `"boolean"` for an all-bool payload containing a null — so the oracle constructs a column pandas' own constructor would never build from the same data.
- **Impact:** ACCEPTED, and it is **load-bearing rather than a defect**, which is the opposite of how it reads. The fixture format tags **every value** with its own `kind`; a float64 column would rewrite each `{"kind":"int64"}` into `{"kind":"float64"}` on the way out, so the nullable dtype is what preserves the payload's kinds across the round trip. Measured 2026-08-08 over the whole corpus, switching both arms to pandas' inference: `agree` 977 → 947 (−30), `moved, unattributed` 151 → 181 (+30), and the `KIND int64->float64` move class 57 → 86 (+29). The change makes the corpus strictly worse and **grows the very class it was expected to shrink**.
- **Consequences worth knowing:** (a) an int+null column is nullable `Int64`, so `fillna` with an incompatible scalar RAISES rather than promoting to object — that is why `fp_p2d_050_dataframe_fillna_cast_error_strict` records an error pandas-native would not produce, and why the fillna half of `br-frankenpandas-fp-stricter-than-pandas-rejections-gtkz1` cannot be settled by relaxing FrankenPandas alone; (b) `kinds <= {bool,int64,float64}` returns `float64`, while pandas infers **object** for a genuine `bool`+`int` mix (`pd.Series([True, 2])`), a third arm with the same shape and no fixture exercising it.
- **Resolution:** ACCEPTED (BlueRobin, 2026-08-08, br-frankenpandas-9ooer). The real question is not "which dtype should the oracle pick" but **whether the fixture format should carry a column dtype instead of per-value kinds** — the current format cannot express pandas' constructor promotion at all, and the dtype forcing is the compensation. Re-opening this should start from that, not from the dtype table.
- **Tests affected:** none changed. The negative result is recorded in `series_dtype_for_payload_values`'s own docstring so the next reader does not repeat the experiment.
- **Review date:** 2026-08-08

### DISC-024: `str.index` / `str.rindex` report a missing position where pandas raises — and therefore do NOT take the float64 promotion
> **Renumbered 2026-08-16 (CalmMink).** This entry was a SECOND `DISC-020`, colliding
> with "STRICT mode deliberately reproduces pandas' silent timedelta int64 wrap"
> above. Both in-tree citations of `DISC-020` name `mul_scalar_with_policy` and
> `nancumsum_with_policy`, which is the timedelta entry, so that one kept the
> number. No citation referred to this entry.
- **Reference:** every OTHER int-returning `.str` accessor follows a three-way result-dtype rule, measured on live pandas 2.2.3 over `pd.Series([...], dtype=object)`: no gaps → `int64`; a `None`/`nan`/non-string element → **`float64`** with the gap as `nan`; a `NaT` → **object**, ints unpromoted and the `NaT` preserved, because float64 cannot hold a NaT. `pd.Series.str.index(sub)` does not participate: it **raises `ValueError: substring not found`** rather than returning anything for an absent needle.
- **Our impl:** `StringAccessor::index_of` / `rindex_of` return a missing value at an absent needle instead of raising — an FP-defined nullable-int contract for an operation pandas has no total equivalent of. Because they are the only int-returning accessors that can invent a gap from *present* input, applying the promotion rule to them would make a found-position `float64` purely because some other row's needle was absent. They therefore keep `Int64` at every found position with only the gaps `NaN`. The oracle already models exactly this (`pandas_oracle._index_of_result` builds an OBJECT series of ints-and-NaN and says so in its docstring); this entry records the same contract on the FrankenPandas side so the two cannot drift apart silently.
- **Impact:** two accessors, `str.index` / `str.rindex`. Their sibling `str.find` / `str.rfind` are unaffected — pandas defines those totally (`-1` when absent), they never invent a gap, and they do take the promotion rule. The other ten int-returning accessors (`len`, `find`, `rfind`, `find_with_bounds`, `rfind_with_bounds`, `encode`, `count`, `count_literal`, `count_matches`, `split_count`) are on the pandas rule via `StringAccessor::apply_str_int_scalar`.
- **Resolution:** ACCEPTED (WindyHare, 2026-08-08, `br-frankenpandas-lwvet`). The alternative considered and rejected was routing these two through the promoting helper for uniformity: it would have redefined an already-documented FP contract as a side effect of a pandas-parity change, and moved `fp_p2d_299` / `fp_p2d_300` further from the oracle rather than closer.
- **Tests affected:** `str_int_returning_ops_promote_to_float64_unless_the_null_is_nat` (fp-frame, final block pins the exception); fixtures `fp_p2d_299_series_str_index_of_null_hardened`, `fp_p2d_300_series_str_rindex_of_null_hardened`.
- **Review date:** 2026-08-08

### DISC-021: `series[int_series]` against a NON-integer index — FrankenPandas takes by label where pandas 2.2.3 still falls back to POSITION
- **Reference:** `Series.__getitem__` only filters for a boolean key. Given anything else it stops filtering and looks the key's VALUES up as labels — measured, pandas 2.2.3: `Series([11,22,33], index=[0,1,2])[Series([1,0,1])]` → values `[22,11,22]`, index `[1,0,1]`, duplicates preserved in selector order, and a missing label raises `KeyError: '[5] not in index'`. FrankenPandas matches that exactly as of `b01a16a15`. The divergence is confined to ONE branch: when the selector is INTEGER and the index is NOT, pandas 2.2.3 indexes by POSITION — `Series([11,22,33], index=['x','y','z'])[Series([1,0,1])]` → `[22,11,22]` at index `['y','x','y']` — while emitting `FutureWarning: Series.__getitem__ treating keys as positions is deprecated. In a future version, integer keys will always be treated as labels (consistent with DataFrame behavior).`
- **Our impl:** `Series::take_by_label_selector` converts the selector's values to `IndexLabel`s and delegates to `Series::loc`, so an integer selector against a Utf8 index fails closed as a missing label. That is what pandas itself will do once the deprecation completes; we do not implement the positional fallback.
- **Impact:** ACCEPTED. Reproducing a branch upstream has already warned it is removing would be building tech debt to spec, and it is the one shape where the two answers differ — every non-deprecated selector/index combination agrees. The cost is that a caller relying on today's positional behaviour gets a `loc label not found` error instead of a row; the benefit is that our answer does not change under a pandas bump.
- **Consequence for the corpus:** `fp_p2c_010_series_filter_non_boolean_mask_strict` originally carried a Utf8 index with an integer mask, i.e. exactly this branch, and pinned FrankenPandas' older blanket refusal (`boolean mask required for filter`). Its payload was re-pointed to an Int64 index so the case exercises the surviving label take rather than the deprecated fallback; the oracle supplies its expected values and live pandas agrees (`[22,11,22]` at `[1,0,1]`). The deprecated branch is therefore recorded here rather than asserted by a fixture — deliberately, because a fixture pinning it would have to be rewritten the moment pandas 3.x lands.
- **Resolution:** ACCEPTED (OliveLark, 2026-08-16, `br-frankenpandas-75i7h`). Two alternatives were considered and rejected: (a) making the ORACLE refuse a non-boolean mask, which cannot work — a bare adapter `OracleError` has no `__cause__`, so `oracle_error_origin` classes it NOT-ATTESTABLE and the row merely changes bucket without the gate improving, and no error-agreement is possible anyway because pandas does not raise here; (b) implementing the positional fallback, i.e. the tech debt above. Revisit when the project pins a pandas release in which integer keys are always labels — at that point this entry closes with no code change.
- **Tests affected:** `series_filter_with_a_non_boolean_mask_takes_by_label_75i7h` and `series_filter_rejects_non_boolean_mask` (fp-frame); fixture `fp_p2c_010_series_filter_non_boolean_mask_strict`.
- **Review date:** 2026-08-16

### DISC-026: `Sparse` dtype NAME omits its subtype and fill value
- **Reference:** MEASURED, live pandas 2.2.3 — `pd.Series(pd.arrays.SparseArray([1,0,0,2], fill_value=0)).dtype` renders as **`Sparse[int64, 0]`**: the name carries the subtype and the fill value, because both are part of the dtype's identity (`Sparse[int64, 0]` and `Sparse[float64, nan]` are different dtypes).
- **Our impl:** `fp_types::DType::name()` returns the bare string `"Sparse"` (`crates/fp-types/src/lib.rs`). This is a NAME-rendering gap, not a modelling gap: `fp_types::SparseDType` already records the subtype/fill contract (see DISC-009), so the information exists and simply is not rendered. `DType::name()` is documented as matching numpy's `dtype.name` property, so this is a divergence from its own stated contract rather than an undecided design question.
- **⚠️ RESOLVED 2026-08-16 (br-frankenpandas-3gxc6, dup br-frankenpandas-8pg11).** FrankenPandas now renders the parameterized name. `SparseDType::pandas_name()` builds `Sparse[<subtype>, <fill>]` from the descriptor that already held both, `Series::dtype_name()` returns it for a sparse Series, and the conformance `series_dtype_check` handler routes onto that at BOTH the strict and hardened sites. The fill is rendered by the existing `scalar_to_string_for_astype`, which already spells every measured case the way pandas does (`True`/`False` capitalised, bare integers, lowercase `nan`, trailing `.0` preserved) — no second formatter was written. `fp_p2d_017_series_dtype_sparse_strict` is repinned to `Sparse[int64, 0]` IN THE SAME COMMIT, as the Resolution line below required. Gate: fp-types 344/0, fp-conformance packet suite green. **The historical note below is kept because the reasoning is still correct for the tree it described.**
- **Impact (historical):** `fp_p2d_017_series_dtype_sparse_strict` pinned `"Sparse"` and the oracle answered `"Sparse[int64, 0]"`, which is one of the three MOVED-unattributed rows in the freshness residue (`br-frankenpandas-nvnvr`). It became visible only once `br-frankenpandas-62d1s` taught the oracle to answer `series_dtype_check` at all — before that nothing asked pandas the question. **Do not re-bank the fixture to `Sparse[int64, 0]`**: it correctly pins what FrankenPandas currently answers, so re-banking would turn a documented divergence into a red fixture without fixing anything.
- **Distinct from DISC-009,** which covers sparse STORAGE and `Series.sparse` accessor parity. A fix here does not need compressed storage — only `name()` and whatever renders it.
- **Resolution:** ~~WILL-FIX~~ **DONE 2026-08-16.** Was tracked by `br-frankenpandas-8pg11`, implemented under the duplicate `br-frankenpandas-3gxc6`. Note `DType::name()` feeds the conformance dtype checks (`br-frankenpandas-62d1s` made those compare `name()` rather than Rust's `Debug`), so changing it moves any fixture asserting a sparse dtype and the two must land together.
- **Review date:** 2026-08-16

### DISC-012: Mixed naive / tz-aware CSV parse_dates normalizes per value
- **Reference:** Pandas handles a CSV column with mixed naive + tz-aware datetime strings by parsing each row independently (the naive rows produce `Timestamp` without tz; the aware rows produce `Timestamp` with tz). When converted to strings, both forms are reformatted into pandas' canonical `YYYY-MM-DD HH:MM:SS[±HH:MM]` shape.
- **Our impl:** fp-io now parses `read_csv(parse_dates=[...])` mixed naive + tz-aware columns per value by calling `to_datetime_values_with_options` with `infer_mixed_timezone=false` and `mixed_tz_as_object=true`. The column remains object-like (`Utf8`) because pandas cannot unify mixed tz-naive/tz-aware values into one `datetime64[ns]` dtype, but each value is normalized to the pandas object-string form.
- **Impact:** Conformance packet `FP-P2D-429` (`csv_read_frame_parse_dates_mixed_timezone_strict`) now matches the fixture: the aware row is normalized from `2024-01-15T10:30:00Z` to `2024-01-15 10:30:00+00:00`.
- **Resolution:** RESOLVED — covered by fp-io test `csv_parse_dates_mixed_naive_and_aware_strings_normalizes_per_value`; the stale accepted-divergence note was superseded by the per-value parse path used by `parse_csv_datetime_values`.
- **Tests affected:** none expected; historical coverage remains `packet_filter_runs_csv_read_frame_parse_dates_mixed_timezone_packet`.
- **Review date:** 2026-06-17

### DISC-028: Concat axis=1 duplicate column label support (prior rejection resolved)
- **Reference:** In pandas 2.2.3, `pd.concat([left, right], axis=1)` succeeds when input frames share column names, preserving all columns in order (e.g. `['dup', 'dup']`). It only rejects when `verify_integrity=True` is explicitly passed.
- **Our impl:** `concat_dataframes_axis1` now constructs output frames via `ColumnStore::from_pairs`, preserving duplicate column names in order when `allows_duplicate_labels` is true (the pandas default). When `allows_duplicate_labels=false`, it rejects with `FrameError::CompatibilityRejected`.
- **Impact:** Previous rejection contract was a legacy artifact where the store used `BTreeMap<String, Column>`. Four packet fixtures (`fp_p2d_028` strict/hardened and `fp_p2d_029` strict/hardened) asserted an artificial error contract (`duplicate column`) that neither pandas nor current FrankenPandas raises. They are marked `retired` with documented provenance per `br-frankenpandas-8b4d4`.
- **Resolution:** RESOLVED (br-frankenpandas-ih4t0 / br-frankenpandas-8b4d4).
- **Tests affected:** `fp_p2d_028` (strict, hardened), `fp_p2d_029` (strict, hardened), `concat_dataframes_axis1_duplicate_columns_succeeds`.
- **Review date:** 2026-09-12

## Rules

1. Every divergence gets a sequential ID (DISC-NNN)
2. Must state whether ACCEPTED, INVESTIGATING, or WILL-FIX
3. Must list affected test cases
4. Must include review date
5. Tests for ACCEPTED divergences use XFAIL markers where applicable