1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
// Copyright (c) 2022-2025 R3BL LLC. Licensed under Apache License, Version 2.0.
//! Rust uses `UTF-8` to represent text in [String]. `UTF-8` is a variable width encoding,
//! so each character can take up a different number of bytes, between 1 and 4, and 1 byte
//! is 8 bits; this is why we use [Vec] of [u8] to represent a [String].
//!
//! For example, the character `H` takes up 1 byte. `UTF-8` is also backward compatible
//! with `ASCII`, meaning that the first 128 characters (the ASCII characters) are
//! represented using the same single byte as in ASCII. So the character `H` is
//! represented by the same byte value in `UTF-8` as it is in `ASCII`. This is why `UTF-8`
//! is so popular, as it allows for the representation of all the characters in the
//! Unicode standard, while still being able to represent `ASCII` characters in the same
//! way.
//!
//! A grapheme cluster is a user-perceived character. Grapheme clusters can take up many
//! more bytes, eg 4 bytes or 2 or 3, etc. Here are some examples:
//! - `😃` takes up 4 bytes.
//! - `📦` also takes up 4 bytes.
//! - `🙏🏽` takes up 4 bytes, but it is a compound grapheme cluster.
//! - `H` takes up only 1 byte.
//!
//! Videos:
//!
//! - [Live coding video on Rust String](https://youtu.be/7I11degAElQ?si=xPDIhITDro7Pa_gq)
//! - [UTF-8 encoding video](https://youtu.be/wIVmDPc16wA?si=D9sTt_G7_mBJFLmc)
//!
//! Docs:
//!
//! - [Grapheme clusters](https://medium.com/flutter-community/working-with-unicode-and-grapheme-clusters-in-dart-b054faab5705)
//! - [UTF-8 String](https://doc.rust-lang.org/book/ch08-02-strings.html)
//!
//! There is a discrepancy between how a [String] that contains grapheme clusters is
//! represented in memory and how it is rendered in a terminal. When writing an TUI editor
//! it is necessary to have a caret (cursor position) that the user can move by pressing
//! up, down, left, right, etc. For left, this is assumed to move the caret or cursor one
//! position to the left, regardless of how wide that character may be. Let's unpack that.
//!
//! - If we use byte boundaries in the [String] we can move the cursor one byte to the
//! left.
//! - This falls apart when we have a grapheme cluster.
//! - A grapheme cluster can take up more than one byte, and they don't fall cleanly into
//! byte boundaries.
//!
//! To complicate things further, the size that a grapheme cluster takes up is not the
//! same as its byte size in memory. Let's unpack that.
//!
//! | Character | Byte size | Grapheme cluster size | Compound |
//! |:----------|:----------|:----------------------|:---------|
//! | `H` | 1 | 1 | No |
//! | `😃` | 4 | 2 | No |
//! | `📦` | 4 | 2 | No |
//! | `🙏🏽` | 4 | 2 | Yes |
//!
//! > **Note**: For input parsing of UTF-8 byte sequences from terminal input, see
//! > [`mod@crate::core::ansi::vt_100_terminal_input_parser::utf8`]. That module
//! > handles byte-level decoding (converting raw bytes to characters), while this
//! > module handles display width calculation (determining how many terminal columns
//! > a character occupies for rendering).
//!
//! Here are examples of compound grapheme clusters.
//!
//! ```text
//! 🏽 + 🙏 = 🙏🏽
//! 🏾 + 👨 + 🤝 + 👨 + 🏿 = 👨🏾🤝👨🏿
//! ```
//!
//! Let's say you're browsing this source file in `VSCode`. The `UTF-8` string this Rust
//! source file is rendered by `VSCode` correctly. But this is not how it looks in a
//! terminal. And the size of the string in memory isn't clear either from looking at the
//! string in `VSCode`. It isn't apparent that you can't just index into the string at
//! byte boundaries.
//!
//! To further complicate things, the output looks different on different terminals &
//! OSes. The function `test_crossterm_grapheme_cluster_width_calc()` (shown below) uses
//! crossterm commands to try and figure out what the width of a grapheme cluster is. When
//! you run this in an SSH session to a macOS machine from Linux, it will work the same
//! way it would locally on Linux. However, if you run the same program in locally via
//! Terminal.app on macOS it works differently! So there are some serious issues.
//!
//! ```no_run
//! use crossterm::{self, *, terminal::*, style::*, cursor::*, event::*};
//! use std::io::*;
//! use std::collections::*;
//!
//! pub fn test_crossterm_grapheme_cluster_width_calc() -> Result<()> {
//! // Enter raw mode, clear screen.
//! enable_raw_mode()?;
//! execute!(stdout(), EnterAlternateScreen)?;
//! execute!(stdout(), Clear(ClearType::All))?;
//! execute!(stdout(), MoveTo(0, 0))?;
//!
//! // Perform test of grapheme cluster width.
//! #[derive(Default, Debug, Clone, Copy)]
//! struct Positions {
//! orig_pos: (u16, u16),
//! new_pos: (u16, u16),
//! col_width: u16,
//! }
//!
//! let mut map = HashMap::<&str, Positions>::new();
//! map.insert("Hi", Positions::default());
//! map.insert(" ", Positions::default());
//! map.insert("😃", Positions::default());
//! map.insert("📦", Positions::default());
//! map.insert("🙏🏽", Positions::default());
//! map.insert("👨🏾🤝👨🏿", Positions::default());
//! map.insert(".", Positions::default());
//!
//! fn process_map(map: &mut HashMap<&str, Positions>) -> Result<()> {
//! for (index, (key, value)) in map.iter_mut().enumerate() {
//! let orig_pos: (u16, u16) = (/* col: */ 0, /* row: */ index as u16);
//! execute!(stdout(), MoveTo(orig_pos.0, orig_pos.1))?;
//! execute!(stdout(), Print(key))?;
//! let new_pos = cursor::position()?;
//! value.new_pos = new_pos;
//! value.orig_pos = orig_pos;
//! value.col_width = new_pos.0 - orig_pos.0;
//! }
//! Ok(())
//! }
//!
//! process_map(&mut map)?;
//!
//! // Just blocking on user input.
//! {
//! execute!(stdout(), Print("... Press any key to continue ..."))?;
//! if let Event::Key(_) = read()? {
//! execute!(stdout(), terminal::Clear(ClearType::All))?;
//! execute!(stdout(), cursor::MoveTo(0, 0))?;
//! }
//! }
//!
//! // Exit raw mode, clear screen.
//! execute!(stdout(), terminal::Clear(ClearType::All))?;
//! execute!(stdout(), cursor::MoveTo(0, 0))?;
//! execute!(stdout(), LeaveAlternateScreen)?;
//! disable_raw_mode().expect("Unable to disable raw mode");
//! println!("map:{:#?}", map);
//!
//! Ok(())
//! }
//! ```
//!
//! The basic problem arises from the fact that it isn't possible to treat the "logical"
//! index into the string (which isn't byte boundary based) as a "display" (or "physical")
//! index into the rendered output of the string in a terminal.
//!
//! - Some parsing is necessary to get "logical" index into the string that is grapheme
//! cluster based (not byte boundary based).
//! - This is where [`unicode-segmentation`](https://crates.io/crates/unicode-segmentation)
//! crate comes in and allows us to split our string into a vector of grapheme
//! clusters.
//! - Some translation is necessary to get from the "logical" index to the physical index
//! and back again. This is where we can apply one of the following approaches:
//! - We can use the [`unicode-width`](https://crates.io/crates/unicode-width) crate to
//! calculate the width of the grapheme cluster. This works on Linux, but doesn't work
//! very well on macOS & I haven't tested it on Windows. This crate will (on Linux)
//! reliably tell us what the displayed width of a grapheme cluster is.
//! - We can take the approach from the [`reedline`](https://crates.io/crates/reedline) crate's
//! [`repaint_buffer()`](https://github.com/nazmulidris/reedline/blob/79e7d8da92cd5ae4f8e459f901189d7419c3adfd/src/painting/painter.rs#L129)
//! where we split the string based on the "logical" index into the vector of grapheme
//! clusters. And then we print the 1st part of the string, then call `SavePosition`
//! to save the cursor at this point, then print the 2nd part of the string, then call
//! `RestorePosition` to restore the cursor to where it "should" be.
//!
//! Please take a look at [`crate::graphemes::GCStringOwned`] for the following
//! items:
//! - Methods in [`mod@crate::graphemes::gc_string`] for more details on how the
//! conversion between "display" (or `display_col_index`), ie, [`crate::ColIndex`] and
//! "logical" or "segment", ie, [`SegIndex`] is done.
//! - The choices that were made in the design of the [`GCStringOwned`] struct for
//! performance to minimize memory latency (for access and allocation). The results
//! might surprise you, as intuition around performance is often not reliable.
//!
//! # The Three Types of Indices
//!
//! When working with Unicode text in a terminal-based editor, we need three distinct
//! types of indices to handle text correctly. This is because there's a fundamental
//! mismatch between how text is stored in memory, how it's logically organized, and
//! how it's displayed on screen.
//!
//! ## 1. `ByteIndex` - Memory Position
//!
//! [`ByteIndex`](crate::ByteIndex) represents the raw byte offset in a UTF-8 encoded
//! string. This is crucial for:
//! - String slicing operations (Rust strings must be sliced at valid UTF-8 boundaries)
//! - Memory access and manipulation
//! - Efficient storage and retrieval
//!
//! Example: In the string "H😀!", 'H' starts at byte 0, '😀' starts at byte 1,
//! and '!' starts at byte 5 (since '😀' takes 4 bytes).
//!
//! ## 2. `SegIndex` - Logical Position (Grapheme Clusters)
//!
//! [`SegIndex`] represents the index of a grapheme cluster (user-perceived character).
//! This is essential for:
//! - Cursor movement (users expect to move by visible characters)
//! - Text editing operations (insert/delete should work on whole characters)
//! - Logical text manipulation
//!
//! Example: In "H😀!", there are 3 segments: seg\[0\]='H', seg\[1\]='😀', seg\[2\]='!'
//!
//! ## 3. `ColIndex` - Display Position
//!
//! [`ColIndex`](crate::ColIndex) represents the column position on the terminal screen.
//! This is necessary because:
//! - Some characters are wider than others (emojis typically take 2 columns)
//! - Terminal rendering requires knowing exact column positions
//! - Cursor positioning and selection highlighting need visual coordinates
//!
//! Example: In "H😀!", 'H' is at col 0, '😀' spans cols 1-2, '!' is at col 3
//!
//! ## Visual Example
//!
//! ```text
//! String: "H😀!"
//!
//! ByteIndex: 0 1 2 3 4 5
//! Content: [H][😀----][!]
//!
//! SegIndex: 0 1 2
//! Segments: [H] [😀] [!]
//!
//! ColIndex: 0 1 2 3
//! Display: [H][😀--] [!]
//! ```
//!
//! ## Conversion Between Index Types
//!
//! The [`GCStringOwned`] struct provides conversion operators to translate between these
//! index types:
//!
//! - `&GCStringOwned + ByteIndex → Option<SegIndex>`: Find which segment contains a byte
//! - `&GCStringOwned + ColIndex → Option<SegIndex>`: Find which segment is at a display
//! column
//! - `&GCStringOwned + SegIndex → Option<ColIndex>`: Find the display column of a segment
//!
//! These conversions can return `None` when indices are out of bounds or fall between
//! characters. For example, a `ByteIndex` in the middle of a multi-byte character would
//! return `None`.
// Attach sources.
// Re-export.
pub use *;
pub use *;
pub use *;
pub use *;