pub fn utf8_character_position(
path: &Path,
line: u64,
byte: u64,
) -> Result<Option<u64>, Error>Expand description
Maps a 1-based byte position to a 1-based Unicode code-point position.
Every byte of a multibyte character maps to the same character. Immediately after an existing line, returns the next character position. CR and LF each count as a code point; these positions are not visual editor columns.
Validates the entire selected line, returning None if any part is invalid
UTF-8, even after the requested byte. An absent line permits byte position
1 and returns None. Other lines’ encodings do not affect the result.
Reopens the path and requires file stability.
§Errors
Returns Error::InvalidLineNumber for line zero and
Error::InvalidBytePosition for byte zero or a position more than one
past the line. Coordinate validation also applies to invalid UTF-8 lines.
Returns Error::Io if metadata lookup, opening, or reading fails,
Error::NotRegularFile for a non-regular input, or Error::FileTooLarge
if the next character position cannot fit in u64.
§Examples
Both bytes of é map to character 4. The terminating LF counts as a
separate code point, and the position immediately after it is also valid.
use paircomp_core::utf8_character_position;
use std::fs;
let directory = std::env::temp_dir()
.join(format!("paircomp-utf8-example-{}", std::process::id()));
fs::create_dir(&directory)?;
let path = directory.join("sample.txt");
fs::write(&path, b"caf\xc3\xa9\nvalid prefix\xff\n")?;
assert_eq!(utf8_character_position(&path, 1, 4)?, Some(4));
assert_eq!(utf8_character_position(&path, 1, 5)?, Some(4));
assert_eq!(utf8_character_position(&path, 1, 6)?, Some(5)); // LF
assert_eq!(utf8_character_position(&path, 1, 7)?, Some(6)); // After LF
// Invalid UTF-8 later in line 2 suppresses even its first position.
assert_eq!(utf8_character_position(&path, 2, 1)?, None);
// A trailing LF does not create a third line.
assert_eq!(utf8_character_position(&path, 3, 1)?, None);
fs::remove_dir_all(&directory)?;