Expand description
The diff engine: Myers’ O(ND) algorithm in linear space, hunk grouping,
token-level diffs for intra-line emphasis and the diff -u text format.
Elements are interned to integers first, a common prefix and suffix are
trimmed, and regions with no element in common are resolved without a
search, so large inputs with few changes stay fast. Past an edit cost of
COST_LIMIT the search splits at its furthest-reaching point instead of
the exact middle snake (git’s heuristic): the script stays valid, it just
may not be minimal.
Structs§
- Hunk
- A group of changes with surrounding context, as
diff -uprints them. - Text
Diff - A line-level diff of two texts.
Enums§
- Op
- One run of an edit script. Ranges index the compared sequences: lines for
diff_lines, bytes of the input strings fordiff_wordsanddiff_chars. The side a run does not touch has an empty range at the position where the run applies.
Constants§
- COST_
LIMIT - The edit cost past which the middle-snake search takes a heuristic split.
Functions§
- diff_
chars - Diff two strings character by character; ranges are byte offsets.
- diff_
lines - Diff two sequences of lines. Include line terminators in the elements when
a missing final newline should count as a change, as
diffdoes. - diff_
slices - Diff two sequences of any hashable elements.
- diff_
words - Diff two strings by
tokenized words; ranges are byte offsets. - group_
hunks - Group an edit script into hunks with
contextunchanged elements around every change; changes closer than twice the context share a hunk. - hunk_
header @@ -a,b +c,d @@for 0-based line ranges, withdiff -u’s conventions: a single line omits,1, and an empty range names the line before it.- tokenize
- Split
textinto word, whitespace and punctuation tokens, as byte ranges. A word is a run of alphanumerics and_; whitespace runs are one token; every other character is its own token.