1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
//! The platform's own text: how the bytes a program writes are read back.
//!
//! On unix this module is the lossy UTF-8 reading the shell has always had, and
//! nothing about it changes. On Windows the machine's own programs — the
//! interpreter's built-ins, the console utilities, a localised system message —
//! write in the code page the machine uses for console output rather than in
//! UTF-8, so reading their bytes as UTF-8 alone turns that text into replacement
//! characters: a localised error message is the visible case. [`decode`] reads
//! such output with the code page the machine itself reports (its OEM code page,
//! `GetOEMCP`) through the platform's own converter (`MultiByteToWideChar`), so
//! every code page a machine can be set to is covered rather than a fixed table
//! of the common ones.
//!
//! The reading is per line, and the two guarantees it is built for are: output
//! that is valid UTF-8 is untouched — a native toolchain that writes UTF-8 is
//! read exactly as it was before — and one line that is not UTF-8 does not make
//! the rest of the same output unreadable, because only that line is handed to
//! the code page. Splitting at a line break is safe under either reading: a
//! `\n` is never a byte inside a multi-byte character of UTF-8, and none of the
//! code pages the platform converts has one inside a character either.
//!
//! Three residuals are stated rather than repaired. A line whose bytes happen to
//! be valid UTF-8 *and* valid text in the machine's code page reads as UTF-8. A
//! line that is mostly UTF-8 but carries one byte that is not — the capture cap
//! cutting a character in half is the case that reaches this in practice — is
//! handed over whole, so its readable part is read through the machine's page too
//! instead of as the text it was. And a program that writes some encoding of its
//! own still reads back wrong: a text file's contents, a dump in a legacy code
//! page the machine does not use for its own output, or a program that formats its
//! message in the ANSI code page, which is the visible case on a machine whose
//! regional code page differs from its console's. Only output written the way the
//! machine writes it is read here; nothing guesses a per-program encoding.
//!
//! Reading an arbitrary file an agent or the owner points at is a different
//! question, and stays the read tool's own reading; so is a datum the product reads
//! back for a purpose of its own — git's plumbing, a managed runtime's version
//! probe — which keeps the lossy reading it has always had.
/// Decode one program's output bytes into the text the agent reads.
///
/// The lines that are valid UTF-8 are kept byte for byte; each line that is not
/// is handed to this platform's fallback whole (`decode_fallback` on Windows, a
/// lossy UTF-8 read everywhere else — where the result is exactly the one a
/// whole-buffer lossy read gives, since a replacement never reaches across a line
/// break).
pub
/// The per-line reading, with the fallback as an argument so both halves of it —
/// which lines are kept and which are handed over, and that the kept ones really
/// are untouched — are driven from any host's test lane.
/// The reading of a line that is not UTF-8, for this platform.
/// The reading of a line that is not UTF-8, for this platform: the machine's own
/// console code page, as the machine reports it.
/// The reading of a line no converter of this platform's describes: the whole
/// reading where there is no converter (unix, where the result is exactly the
/// whole-buffer lossy read), and on Windows the backstop for a line the converter
/// was not asked about — a length past its own `int` (never the shell's: its
/// capture cap is far below `i32::MAX`) or a call that reports nothing. This is
/// the display of output a program has already written, so nothing here may drop
/// the line.
/// Convert one line with `code_page` through the platform's own converter
/// (`MultiByteToWideChar`) — the reason no table of code pages lives here: the
/// machine converts with the page it actually has, installed pages included, and
/// a byte the page does not define comes back as that converter's own substitute
/// character rather than as a refusal.