Skip to main content

rucc_lex/
number.rs

1//! Numeric constants: the value, and the type the standard's table walk gives it.
2//!
3//! Design: `spec/06-lexer-and-parser.md` section 6.1.
4//!
5//! This is the second piece of phase 7. A preprocessing number is a loose thing, deliberately
6//! looser than a constant, so `1.2.3` and `0x1p+3` are both one pp-token and only here does
7//! anyone ask what they mean. What comes back is a value and a type, and both of them are
8//! places a compiler quietly goes wrong.
9//!
10//! There are two entry points, [`integer`] and [`floating`], and either of them hands the
11//! spelling to the other rather than reporting an error when it turns out to belong there. The
12//! split is not where a reader expects it: `1e` is a floating constant with no exponent digits
13//! and `1f` is an integer constant with a suffix that does not exist, and both compilers agree
14//! on that, because the exponent marker is part of a preprocessing number and the suffix letter
15//! is not part of anything.
16//!
17//! The value is accumulated in a `u128` with every step checked, so a constant too large to
18//! represent is a diagnostic rather than a number the program did not write. gcc 13.3 does not
19//! do that: its accumulator is sixty four bits, and `18446744073709551616` compiles to zero of
20//! type `int` after a warning nobody reads. That is not a behaviour worth reproducing, so ours
21//! is the only measured difference here that is deliberate: past a hundred and twenty eight
22//! bits the constant is refused. clang refuses it too, one bit earlier.
23//!
24//! The type is the standard's table walk, 6.4.4.1p5: a candidate list chosen by the base and
25//! the suffix, walked in order, and the first type that holds the value wins. The list is not
26//! the same in every dialect. C89 puts `unsigned long` in the list for a decimal constant with
27//! no suffix, which is what makes `18446744073709551615` an `unsigned long` under `-std=c89`
28//! and something wider under `-std=c99`, and gcc says so in as many words: "this decimal
29//! constant is unsigned only in ISO C90". Both compilers keep `long long` out of the C89 lists
30//! and accept it when the suffix asks for it.
31//!
32//! `__int128` is on the end of every list, which is what gcc does and clang does not.
33//! `9223372036854775808` is an `__int128` in gcc 13.3 and an `unsigned long long` in clang,
34//! and the difference is visible to a program: negate it and gcc gives a negative number.
35//! We follow gcc, because the alternative silently turns a signed constant unsigned.
36//!
37//! The rest was measured the same way, by writing the constant and asking `_Generic` what it
38//! is, on gcc 13.3 on x86-64 Linux and on clang:
39//!
40//! The suffix letters may be in either case but not both, so `1ll` and `1LL` are constants and
41//! `1lL` is not, and the same rule holds for `wb`. The unsigned suffix may come before or after
42//! the length suffix. `wb` does not combine with `l` at all.
43//!
44//! Binary constants are accepted in every dialect by both compilers, as an extension before
45//! C23. Digit separators are C23 only in both. `_BitInt` constants are C23 in the standard and
46//! clang accepts them in every dialect.
47//!
48//! A `wb` constant has the narrowest type that holds it, which for a signed one includes the
49//! sign bit and is never less than two: `1wb` is `_BitInt(2)`, `42wb` is `_BitInt(7)`, `255uwb`
50//! is `unsigned _BitInt(8)` and `0uwb` is `unsigned _BitInt(1)`. Measured against gcc 16 and
51//! clang, which agree on every one of them.
52//!
53//! # Floating constants
54//!
55//! A floating constant has none of the table walk about it: the suffix names the type outright,
56//! and with no suffix it is a `double`. What it has instead is a list of suffixes that is much
57//! longer than the standard's three, and a conversion that has to be exactly right.
58//!
59//! The conversion is [`rucc_base::float`], correctly rounded and done in software, so that the
60//! bits do not depend on the machine the compiler runs on. The type decides the format and the
61//! target decides what some of the types are: `long double` is the x87 eighty bit format on
62//! x86-64 Linux and true quad precision on AArch64 Linux, and both of them say 128 bits wide,
63//! which is why [`TargetInfo::long_double_format`] exists. The target also decides which of the
64//! types exist at all: `f16` is not a suffix on a machine with no `_Float16`, and gcc turns it
65//! away there as a suffix it does not support rather than as one it has never heard of.
66//!
67//! The suffixes were measured on gcc 13.3 on x86-64 Linux, with `_Generic` for the type and by
68//! printing the bytes for the format. `f` is `float` and `l` is `long double`, and then the
69//! extensions: `q` is `__float128`, `w` is the x87 `__float80`, `d` is a `double` written the
70//! long way, `f16` `f32` `f64` `f128` are the `_FloatN` types and `f32x` `f64x` the `_FloatNx`
71//! ones, and `i` or `j` in either position makes the constant imaginary. `_Float32x` turns out
72//! to be plain `double` and `_Float64x` the x87 format, which is not what the names suggest:
73//! `0.1f32x` is `0x3fb999999999999a` and `0.1f64x` is `0x3ffbcccccccccccccccd`.
74//!
75//! The case rules are their own small grammar. The `f` of a `_FloatN` suffix may be either case
76//! and the trailing `x` may not, so `F64x` is a constant and `f64X` is not. A decimal float
77//! suffix is two letters that have to agree, so `df` and `DF` are constants and `Df` is not.
78//! Every one of these is accepted in every dialect, C89 included, and every one of them is
79//! worth a remark when the dialect did not ask for it.
80//!
81//! Decimal floating constants are refused rather than converted, because there is no decimal
82//! float value anywhere in this compiler to put one in. That is a gap and the error says so.
83
84use rucc_base::float::{Float, Format, ParseError, Status};
85use rucc_session::Std;
86use rucc_target::TargetInfo;
87use rucc_tuple::Arch;
88use rucc_types::{IntKind, int_width};
89
90use crate::remarks::Remarks;
91
92/// A converted integer constant.
93#[derive(Debug, Clone, Copy, PartialEq, Eq)]
94pub struct IntConstant {
95    /// The value, which is never negative: a minus sign is an operator and not part of the
96    /// constant, which is why `-2147483648` is a `long` on a 32-bit `int` and the reason
97    /// `INT_MIN` is spelled the way it is in `limits.h`.
98    pub value: u128,
99    /// The type the table walk arrived at, which is the type of a half where the constant is
100    /// imaginary and the type of the whole thing where it is not.
101    pub ty: IntConstantType,
102    /// Whether an `i` or a `j` made this the imaginary part of a complex constant. The value is
103    /// the whole number that was written either way, so the caller builds the complex one.
104    pub imaginary: bool,
105    /// What is worth saying about the constant, for the caller that holds the span.
106    pub remarks: Remarks,
107}
108
109/// The type of an integer constant.
110#[derive(Debug, Clone, Copy, PartialEq, Eq)]
111pub enum IntConstantType {
112    /// One of the integer kinds, chosen by the table walk.
113    Standard(IntKind),
114    /// A `_BitInt` of exactly the width it takes to hold the value.
115    BitInt {
116        /// Whether the type is signed, which it is when the `u` suffix was not there.
117        signed: bool,
118        /// The width in bits, including the sign bit when there is one.
119        width: u32,
120    },
121}
122
123impl IntConstantType {
124    /// A suffix that names this type, for a printer putting the constant back.
125    ///
126    /// The value decides the rest. Nothing spells `__int128`, because no suffix does and none is
127    /// needed: a constant only gets that type by being too large for every other one, so writing
128    /// the value back gets the same answer out of the table walk.
129    #[must_use]
130    pub const fn suffix(self) -> &'static str {
131        match self {
132            IntConstantType::Standard(kind) => match kind {
133                IntKind::UInt | IntKind::UInt128 => "u",
134                IntKind::Long => "l",
135                IntKind::ULong => "ul",
136                IntKind::LongLong => "ll",
137                IntKind::ULongLong => "ull",
138                _ => "",
139            },
140            IntConstantType::BitInt { signed: true, .. } => "wb",
141            IntConstantType::BitInt { signed: false, .. } => "uwb",
142        }
143    }
144}
145
146/// Why a preprocessing number is not an integer constant.
147///
148/// [`IntError::Floating`] is not a diagnostic. It means the spelling belongs to the floating
149/// path, and it is an error here so that the caller cannot forget to ask.
150#[derive(Debug, Clone, Copy, PartialEq, Eq)]
151pub enum IntError {
152    /// This is a floating constant. Nothing is wrong with it.
153    Floating,
154    /// The characters after the digits are not a suffix.
155    InvalidSuffix,
156    /// An `8` or a `9` in a constant that started with `0`.
157    InvalidOctalDigit,
158    /// `0x` or `0b` with no digits after it.
159    NoDigits,
160    /// Larger than any integer type, or than the hundred and twenty eight bits the value is
161    /// accumulated in.
162    TooLarge,
163}
164
165impl IntError {
166    /// What to print, in GCC's words where GCC has any.
167    ///
168    /// The offending character is not in the message, because the caller has the spelling and
169    /// the span and can say `invalid suffix "ux" on integer constant` the way GCC does.
170    #[must_use]
171    pub const fn message(self) -> &'static str {
172        match self {
173            IntError::Floating => "not an integer constant",
174            IntError::InvalidSuffix => "invalid suffix on integer constant",
175            IntError::InvalidOctalDigit => "invalid digit in octal constant",
176            IntError::NoDigits => "no digits in integer constant",
177            IntError::TooLarge => "integer constant is too large to be represented in any type",
178        }
179    }
180}
181
182/// A converted floating constant.
183#[derive(Debug, Clone, Copy, PartialEq, Eq)]
184pub struct FloatConstant {
185    /// The value, correctly rounded into the format of its type.
186    pub value: Float,
187    /// The type the suffix named, which is `double` when there was no suffix.
188    pub ty: FloatConstantType,
189    /// Whether an `i` or a `j` made this the imaginary part of a complex constant. The value is
190    /// the real number that was written either way, so the caller builds the complex one.
191    pub imaginary: bool,
192    /// What is worth saying about the constant, for the caller that holds the span.
193    pub remarks: Remarks,
194}
195
196/// The type of a floating constant.
197///
198/// This is not [`rucc_types::FloatKind`], which has the three types C has. The suffixes reach
199/// further than that, and a constant knows exactly which type it was written as long before
200/// anything has to decide what that type is on this target.
201#[derive(Debug, Clone, Copy, PartialEq, Eq)]
202pub enum FloatConstantType {
203    /// `f` or `F`.
204    Float,
205    /// No suffix, or the `d` and `D` that GCC also accepts for it.
206    Double,
207    /// `l` or `L`.
208    LongDouble,
209    /// `f16` or `F16`, which is `_Float16`.
210    Float16,
211    /// `f32` or `F32`, which is `_Float32` and is the same format as `float`.
212    Float32,
213    /// `f64` or `F64`, which is `_Float64` and is the same format as `double`.
214    Float64,
215    /// `f128` or `F128`, and `q` or `Q`, which is `_Float128` and `__float128`. GCC keeps the
216    /// two spellings as distinct types and they are the same format, which is all this says.
217    Float128,
218    /// `f32x` or `F32x`, which is `_Float32x`. The name suggests something wider than
219    /// `_Float32` and on every target here it is exactly `double`.
220    Float32x,
221    /// `f64x` or `F64x`, which is `_Float64x`: the widest format the target has beyond
222    /// `_Float64`, so the x87 one on x86 and quad precision elsewhere.
223    Float64x,
224    /// `w` or `W`, which is GCC's `__float80`. It is the x87 format whatever `long double` is,
225    /// which is the reason it is not the same thing as [`FloatConstantType::LongDouble`], and
226    /// GCC has it on x86 only.
227    Float80,
228}
229
230impl FloatConstantType {
231    /// The format a constant of this type is converted in.
232    ///
233    /// The two that depend on the target are the two that have to: `long double` is the x87
234    /// format on x86-64 Linux and quad precision on AArch64 Linux, and `_Float64x` is whatever
235    /// the target has above `_Float64`, which is the same split.
236    #[must_use]
237    pub fn format(self, target: &TargetInfo) -> Format {
238        match self {
239            FloatConstantType::Float | FloatConstantType::Float32 => Format::Single,
240            FloatConstantType::Double
241            | FloatConstantType::Float64
242            | FloatConstantType::Float32x => Format::Double,
243            FloatConstantType::LongDouble => target.long_double_format,
244            FloatConstantType::Float16 => Format::Half,
245            FloatConstantType::Float128 => Format::Quad,
246            FloatConstantType::Float64x if target.tuple.arch() == Arch::X86_64 => {
247                Format::X87Extended
248            }
249            FloatConstantType::Float64x => Format::Quad,
250            FloatConstantType::Float80 => Format::X87Extended,
251        }
252    }
253
254    /// The C spelling of the type, which a diagnostic naming it has to print.
255    ///
256    /// GCC says "floating constant exceeds range of 'double'" and puts the type in the message,
257    /// so the type has to be able to say what it is called.
258    #[must_use]
259    pub const fn name(self) -> &'static str {
260        match self {
261            FloatConstantType::Float => "float",
262            FloatConstantType::Double => "double",
263            FloatConstantType::LongDouble => "long double",
264            FloatConstantType::Float16 => "_Float16",
265            FloatConstantType::Float32 => "_Float32",
266            FloatConstantType::Float64 => "_Float64",
267            FloatConstantType::Float128 => "_Float128",
268            FloatConstantType::Float32x => "_Float32x",
269            FloatConstantType::Float64x => "_Float64x",
270            FloatConstantType::Float80 => "__float80",
271        }
272    }
273
274    /// The suffix that names this type, for a printer putting the constant back.
275    ///
276    /// One spelling per type rather than the one that was written, so `0.1q` and `0.1f128` both
277    /// come back as `f128`. They are the same type, and the printer's job is to mean the same
278    /// thing rather than to look the same.
279    #[must_use]
280    pub const fn suffix(self) -> &'static str {
281        match self {
282            FloatConstantType::Float => "f",
283            FloatConstantType::Double => "",
284            FloatConstantType::LongDouble => "l",
285            FloatConstantType::Float16 => "f16",
286            FloatConstantType::Float32 => "f32",
287            FloatConstantType::Float64 => "f64",
288            FloatConstantType::Float128 => "f128",
289            FloatConstantType::Float32x => "f32x",
290            FloatConstantType::Float64x => "f64x",
291            FloatConstantType::Float80 => "w",
292        }
293    }
294}
295
296/// Why a preprocessing number is not a floating constant.
297///
298/// [`FloatError::Integer`] is not a diagnostic, in the same way [`IntError::Floating`] is not:
299/// it means the spelling belongs to the other path.
300#[derive(Debug, Clone, Copy, PartialEq, Eq)]
301pub enum FloatError {
302    /// This is an integer constant. Nothing is wrong with it.
303    Integer,
304    /// The characters after the number are not a suffix.
305    InvalidSuffix,
306    /// A hexadecimal floating constant with no `p` exponent. The exponent is required there and
307    /// not optional as it is in a decimal one, because `f` is a hexadecimal digit and there
308    /// would be no way to tell a suffix from the number.
309    MissingExponent,
310    /// An `e` or a `p` with no digits after it.
311    NoExponentDigits,
312    /// A constant with a point and no digits at all.
313    NoDigits,
314    /// More than one point, which is a preprocessing number and not a constant.
315    TooManyPoints,
316    /// A `df`, `dd` or `dl` suffix. The constant is well formed and this compiler has nowhere
317    /// to put a decimal floating value yet.
318    DecimalFloat,
319    /// A suffix naming a type this target does not have, which is `w` anywhere but x86, `f128x`
320    /// everywhere, and `f16`, `f128`, `q` and `f64x` on the machines that have no such type.
321    UnsupportedType,
322}
323
324impl FloatError {
325    /// What to print, in GCC's words where GCC has any.
326    #[must_use]
327    pub const fn message(self) -> &'static str {
328        match self {
329            FloatError::Integer => "not a floating constant",
330            FloatError::InvalidSuffix => "invalid suffix on floating constant",
331            FloatError::MissingExponent => "hexadecimal floating constants require an exponent",
332            FloatError::NoExponentDigits => "exponent has no digits",
333            FloatError::NoDigits => "no digits in floating constant",
334            FloatError::TooManyPoints => "too many decimal points in number",
335            FloatError::DecimalFloat => "decimal floating constants are not supported yet",
336            // gcc's words, and it says them about the suffix rather than about the type. It is
337            // the same sentence for `1.0w` on AArch64 and `1.0f128` on armv7 and `1.0f128x`
338            // anywhere, which was measured rather than guessed.
339            FloatError::UnsupportedType => "unsupported non-standard suffix on floating constant",
340        }
341    }
342}
343
344/// Converts the spelling of a preprocessing number into an integer constant.
345///
346/// # Errors
347///
348/// [`IntError`], one case of which is that the spelling is a floating constant rather than a
349/// malformed integer one.
350pub fn integer(text: &str, std: Std, target: &TargetInfo) -> Result<IntConstant, IntError> {
351    let bytes = text.as_bytes();
352    let (base, start) = base_of(bytes);
353    if is_floating(bytes, base) {
354        return Err(IntError::Floating);
355    }
356    let mut remarks = Remarks::NONE;
357    if base == 2 && std < Std::C23 {
358        remarks = remarks.with(Remarks::BINARY);
359    }
360
361    let mut value: u128 = 0;
362    let mut digits = 0;
363    let mut index = start;
364    while index < bytes.len() {
365        let byte = bytes[index];
366        if byte == b'\'' {
367            // A separator is only a separator between two digits. The scanner keeps one in the
368            // number only when an identifier character follows, so a trailing one arrives here
369            // as a suffix instead and is refused as one.
370            if digits == 0 || index + 1 >= bytes.len() || digit(bytes[index + 1], base).is_none() {
371                return Err(IntError::InvalidSuffix);
372            }
373            if std < Std::C23 {
374                remarks = remarks.with(Remarks::SEPARATORS);
375            }
376            index += 1;
377            continue;
378        }
379        let Some(digit) = digit(byte, base) else {
380            break;
381        };
382        value = value
383            .checked_mul(u128::from(base))
384            .and_then(|shifted| shifted.checked_add(u128::from(digit)))
385            .ok_or(IntError::TooLarge)?;
386        digits += 1;
387        index += 1;
388    }
389    if digits == 0 {
390        // `0x` with nothing after it, which GCC reports as an invalid suffix because it read
391        // the `0` as the constant. The distinction is not worth a worse message than this.
392        return Err(IntError::NoDigits);
393    }
394    if base == 8 && bytes[start..index].iter().any(|&byte| byte == b'8' || byte == b'9') {
395        return Err(IntError::InvalidOctalDigit);
396    }
397
398    let suffix = suffix_of(&bytes[index..])?;
399    if suffix.length == Some(Length::LongLong) && std == Std::C89 {
400        remarks = remarks.with(Remarks::LONG_LONG);
401    }
402    if suffix.imaginary {
403        remarks = remarks.with(Remarks::IMAGINARY);
404    }
405    if suffix.length == Some(Length::BitInt) {
406        if std < Std::C23 {
407            remarks = remarks.with(Remarks::BIT_INT);
408        }
409        let ty = bit_int(value, suffix.unsigned);
410        return Ok(IntConstant { value, ty, imaginary: false, remarks });
411    }
412
413    let candidates = candidates(base, suffix, std);
414    let kind = candidates
415        .iter()
416        .copied()
417        .find(|&kind| fits(value, kind, target))
418        .ok_or(IntError::TooLarge)?;
419    if base == 10 && !suffix.unsigned && !signed_standard(kind) {
420        remarks = remarks.with(Remarks::UNSIGNED);
421    }
422    Ok(IntConstant {
423        value,
424        ty: IntConstantType::Standard(kind),
425        imaginary: suffix.imaginary,
426        remarks,
427    })
428}
429
430/// The base a spelling is written in, and where its digits start.
431///
432/// A leading `0` means octal only when a digit follows, so `0u` is a decimal zero with a
433/// suffix and `08` is an octal constant with a digit that does not exist. That is the split
434/// GCC makes, and it is what turns `08` into a message about octal rather than about a suffix.
435fn base_of(bytes: &[u8]) -> (u32, usize) {
436    match bytes {
437        [b'0', b'x' | b'X', ..] => (16, 2),
438        [b'0', b'b' | b'B', ..] => (2, 2),
439        [b'0', next, ..] if next.is_ascii_digit() => (8, 1),
440        _ => (10, 0),
441    }
442}
443
444/// Whether the spelling is a floating constant rather than an integer one.
445///
446/// A point anywhere, an `e` exponent in a decimal constant, or a `p` exponent in a hexadecimal
447/// one. `1e` and `1e+` are floating constants with no exponent digits, which is a diagnostic
448/// the floating path gives, and `1f` is an integer constant with a suffix that does not exist,
449/// which is one this path gives. Both compilers split them exactly there.
450///
451/// A leading zero does not survive an exponent: `08e5` is the floating constant eight hundred
452/// thousand and not an octal constant with a digit that does not exist.
453fn is_floating(bytes: &[u8], base: u32) -> bool {
454    let exponent = if base == 16 { *b"pP" } else { *b"eE" };
455    bytes.iter().any(|&byte| byte == b'.' || exponent.contains(&byte))
456}
457
458/// The value of a digit in the given base, and [`None`] when the byte is not one.
459///
460/// An octal constant reads `8` and `9` as digits, so that a constant holding one ends at the
461/// suffix and the error can name the digit rather than complain about the suffix.
462fn digit(byte: u8, base: u32) -> Option<u32> {
463    char::from(byte).to_digit(if base == 8 { 10 } else { base })
464}
465
466/// The length part of a suffix.
467#[derive(Debug, Clone, Copy, PartialEq, Eq)]
468enum Length {
469    /// `l` or `L`.
470    Long,
471    /// `ll` or `LL`.
472    LongLong,
473    /// `wb` or `WB`.
474    BitInt,
475}
476
477/// A parsed suffix.
478#[derive(Debug, Clone, Copy, PartialEq, Eq)]
479struct Suffix {
480    /// Whether `u` or `U` was there.
481    unsigned: bool,
482    /// The length part, when there was one.
483    length: Option<Length>,
484    /// Whether `i`, `j`, `I` or `J` was there.
485    imaginary: bool,
486}
487
488/// Reads the suffix, which may hold each part once and in any order.
489fn suffix_of(mut rest: &[u8]) -> Result<Suffix, IntError> {
490    let mut suffix = Suffix { unsigned: false, length: None, imaginary: false };
491    while let Some(&byte) = rest.first() {
492        let taken = match byte {
493            b'u' | b'U' if !suffix.unsigned => {
494                suffix.unsigned = true;
495                1
496            }
497            // The imaginary suffix, which sits on either side of the rest of the suffix the way
498            // it does on a floating constant: `1li` and `1il` are both `_Complex long`.
499            b'i' | b'j' | b'I' | b'J' if !suffix.imaginary => {
500                suffix.imaginary = true;
501                1
502            }
503            // The two letters have to agree about case, so `1ll` and `1LL` are constants and
504            // `1lL` is not. Both compilers refuse the mixed spelling in every dialect.
505            b'l' | b'L' if suffix.length.is_none() => {
506                if rest.get(1) == Some(&byte) {
507                    suffix.length = Some(Length::LongLong);
508                    2
509                } else {
510                    suffix.length = Some(Length::Long);
511                    1
512                }
513            }
514            b'w' | b'W' if suffix.length.is_none() => {
515                let second = if byte == b'w' { b'b' } else { b'B' };
516                if rest.get(1) != Some(&second) {
517                    return Err(IntError::InvalidSuffix);
518                }
519                suffix.length = Some(Length::BitInt);
520                2
521            }
522            _ => return Err(IntError::InvalidSuffix),
523        };
524        rest = &rest[taken..];
525    }
526    // `_BitInt` is the one length the imaginary suffix cannot join, because there is no complex
527    // `_BitInt` for the constant to have: gcc refuses `3wbi` as an invalid suffix rather than as
528    // a type it does not support, and refusing it here is what gives that message.
529    if suffix.imaginary && suffix.length == Some(Length::BitInt) {
530        return Err(IntError::InvalidSuffix);
531    }
532    Ok(suffix)
533}
534
535/// The type of a `wb` constant, which is the narrowest one that holds the value.
536///
537/// The sign bit counts, so a signed one is never narrower than two bits: `1wb` is
538/// `_BitInt(2)`. An unsigned zero is `unsigned _BitInt(1)`, because a width of zero is not a
539/// type. Measured against gcc 16 and clang.
540fn bit_int(value: u128, unsigned: bool) -> IntConstantType {
541    let used = 128 - value.leading_zeros();
542    let width = if unsigned { used.max(1) } else { used + 1 };
543    IntConstantType::BitInt { signed: !unsigned, width: width.max(if unsigned { 1 } else { 2 }) }
544}
545
546/// Whether `kind` is one of the standard signed types, which is what decides the remark about
547/// a decimal constant having gone unsigned.
548fn signed_standard(kind: IntKind) -> bool {
549    matches!(kind, IntKind::Int | IntKind::Long | IntKind::LongLong)
550}
551
552/// Whether the value fits in `kind` on this target.
553fn fits(value: u128, kind: IntKind, target: &TargetInfo) -> bool {
554    let width = int_width(kind, target);
555    // Signedness here never depends on what plain `char` is, because no candidate list holds a
556    // character type.
557    let bits = if kind.is_signed(false) { width - 1 } else { width };
558    // `unsigned __int128` holds every value the accumulator can, and shifting a `u128` by all
559    // of its bits is not a shift, so the widest type is answered without one.
560    bits >= 128 || value >> bits == 0
561}
562
563/// The candidate list for a base and a suffix, in the order the standard walks it.
564///
565/// `__int128` and `unsigned __int128` are on the end of every list, which is what gcc does:
566/// `9223372036854775808` is an `__int128` there and an `unsigned long long` in clang. Both
567/// compilers put `long long` out of reach in C89 unless the suffix asks for it, and C89 is
568/// also the dialect that offers `unsigned long` for a decimal constant with no suffix at all.
569fn candidates(base: u32, suffix: Suffix, std: Std) -> &'static [IntKind] {
570    use IntKind::{Int, Int128, Long, LongLong, UInt, UInt128, ULong, ULongLong};
571
572    let decimal = base == 10;
573    let c89 = std == Std::C89;
574    match (suffix.unsigned, suffix.length) {
575        (false, None) if decimal && c89 => &[Int, Long, ULong, Int128, UInt128],
576        (false, None) if decimal => &[Int, Long, LongLong, Int128],
577        (false, None) if c89 => &[Int, UInt, Long, ULong, Int128, UInt128],
578        (false, None) => &[Int, UInt, Long, ULong, LongLong, ULongLong, Int128, UInt128],
579
580        (true, None) if c89 => &[UInt, ULong, UInt128],
581        (true, None) => &[UInt, ULong, ULongLong, UInt128],
582
583        (false, Some(Length::Long)) if decimal && c89 => &[Long, ULong, Int128, UInt128],
584        (false, Some(Length::Long)) if decimal => &[Long, LongLong, Int128],
585        (false, Some(Length::Long)) if c89 => &[Long, ULong, Int128, UInt128],
586        (false, Some(Length::Long)) => &[Long, ULong, LongLong, ULongLong, Int128, UInt128],
587
588        (true, Some(Length::Long)) if c89 => &[ULong, UInt128],
589        (true, Some(Length::Long)) => &[ULong, ULongLong, UInt128],
590
591        (false, Some(Length::LongLong)) if decimal => &[LongLong, Int128],
592        (false, Some(Length::LongLong)) => &[LongLong, ULongLong, Int128, UInt128],
593        (true, Some(Length::LongLong)) => &[ULongLong, UInt128],
594
595        // A `wb` constant never reaches here: its type comes from the value alone.
596        (_, Some(Length::BitInt)) => &[],
597    }
598}
599
600/// Converts the spelling of a preprocessing number into a floating constant.
601///
602/// # Errors
603///
604/// [`FloatError`], one case of which is that the spelling is an integer constant rather than a
605/// malformed floating one.
606pub fn floating(text: &str, std: Std, target: &TargetInfo) -> Result<FloatConstant, FloatError> {
607    let bytes = text.as_bytes();
608    let (base, _) = base_of(bytes);
609    if !is_floating(bytes, base) {
610        return Err(FloatError::Integer);
611    }
612    // A leading zero means nothing to a floating constant, so there are two bases here and not
613    // four: `08e5` is eight hundred thousand rather than an octal constant with a bad digit.
614    let hex = base == 16;
615    let base = if hex { 16 } else { 10 };
616    let mut remarks = Remarks::NONE;
617    if hex && std < Std::C99 {
618        remarks = remarks.with(Remarks::HEX_FLOAT);
619    }
620
621    let mut index = if hex { 2 } else { 0 };
622    let mut digits = 0;
623    let mut point = false;
624    let mut separators = false;
625    while index < bytes.len() {
626        let byte = bytes[index];
627        if byte == b'\'' {
628            if digits == 0 || !next_is_digit(bytes, index, base) {
629                return Err(FloatError::InvalidSuffix);
630            }
631            separators = true;
632        } else if byte == b'.' {
633            if point {
634                return Err(FloatError::TooManyPoints);
635            }
636            point = true;
637        } else if digit(byte, base).is_some() {
638            digits += 1;
639        } else {
640            break;
641        }
642        index += 1;
643    }
644    if digits == 0 {
645        return Err(FloatError::NoDigits);
646    }
647
648    let marker = if hex { *b"pP" } else { *b"eE" };
649    if index < bytes.len() && marker.contains(&bytes[index]) {
650        index += 1;
651        if matches!(bytes.get(index), Some(b'+' | b'-')) {
652            index += 1;
653        }
654        let mut exponent_digits = 0;
655        while index < bytes.len() {
656            let byte = bytes[index];
657            if byte == b'\'' {
658                if exponent_digits == 0 || !next_is_digit(bytes, index, 10) {
659                    return Err(FloatError::InvalidSuffix);
660                }
661                separators = true;
662            } else if byte.is_ascii_digit() {
663                exponent_digits += 1;
664            } else {
665                break;
666            }
667            index += 1;
668        }
669        if exponent_digits == 0 {
670            return Err(FloatError::NoExponentDigits);
671        }
672    } else if hex {
673        // The exponent is not optional in a hexadecimal constant, because `f` is a digit there
674        // and `0x1.8f` would otherwise be a number and a suffix at the same time.
675        return Err(FloatError::MissingExponent);
676    }
677    if separators && std < Std::C23 {
678        remarks = remarks.with(Remarks::SEPARATORS);
679    }
680
681    let suffix = float_suffix(&bytes[index..], target)?;
682    remarks = remarks.with(suffix.remarks);
683    let (value, status) =
684        Float::parse(&text[..index], suffix.ty.format(target)).map_err(|error| match error {
685            // The scan above has already ruled all three of these out, and mapping them is
686            // still better than an unwrap that a later change could reach.
687            ParseError::NoDigits => FloatError::NoDigits,
688            ParseError::NoExponentDigits => FloatError::NoExponentDigits,
689            ParseError::Invalid => FloatError::InvalidSuffix,
690        })?;
691    if status.has(Status::OVERFLOW) {
692        remarks = remarks.with(Remarks::OUT_OF_RANGE);
693    }
694    // Underflow on its own is a subnormal, which is a number the program can use. Losing the
695    // value entirely is the part worth a word.
696    if status.has(Status::UNDERFLOW) && value.is_zero() {
697        remarks = remarks.with(Remarks::TRUNCATED);
698    }
699    Ok(FloatConstant { value, ty: suffix.ty, imaginary: suffix.imaginary, remarks })
700}
701
702/// Whether the byte after `index` is a digit in `base`, which is what makes a separator one.
703fn next_is_digit(bytes: &[u8], index: usize, base: u32) -> bool {
704    bytes.get(index + 1).is_some_and(|&next| digit(next, base).is_some())
705}
706
707/// A parsed floating suffix.
708struct FloatSuffix {
709    /// The type it named, which is `double` when it named none.
710    ty: FloatConstantType,
711    /// Whether it held an `i` or a `j`.
712    imaginary: bool,
713    /// What the suffix alone is worth saying about.
714    remarks: Remarks,
715}
716
717/// Reads the suffix, which may name a type once and mark the constant imaginary once, in either
718/// order.
719///
720/// Everything past `f` and `l` is an extension, and the extensions are where the case rules stop
721/// being uniform: the `f` of `_FloatN` may be either case and the `x` of `_FloatNx` may not, and
722/// the two letters of a decimal suffix have to agree. All of it measured on gcc 13.3.
723fn float_suffix(mut rest: &[u8], target: &TargetInfo) -> Result<FloatSuffix, FloatError> {
724    let mut ty = None;
725    let mut imaginary = false;
726    let mut remarks = Remarks::NONE;
727    while let Some(&byte) = rest.first() {
728        let taken = match byte {
729            b'i' | b'j' | b'I' | b'J' if !imaginary => {
730                imaginary = true;
731                remarks = remarks.with(Remarks::IMAGINARY);
732                1
733            }
734            // One type per constant, so `1.0fl` is not a constant and neither is `1.0ff`.
735            _ if ty.is_some() => return Err(FloatError::InvalidSuffix),
736            b'f' | b'F' => {
737                let (named, taken, extra) = float_n(rest, target)?;
738                ty = Some(named);
739                remarks = remarks.with(extra);
740                taken
741            }
742            b'l' | b'L' => {
743                ty = Some(FloatConstantType::LongDouble);
744                1
745            }
746            b'q' | b'Q' => {
747                // The same type the `f128` suffix names, so it goes where that type goes.
748                if !target.has_float128 {
749                    return Err(FloatError::UnsupportedType);
750                }
751                ty = Some(FloatConstantType::Float128);
752                remarks = remarks.with(Remarks::EXTENDED_SUFFIX);
753                1
754            }
755            b'w' | b'W' => {
756                // `__float80` is the x87 format, which only x86 has.
757                if target.tuple.arch() != Arch::X86_64 {
758                    return Err(FloatError::UnsupportedType);
759                }
760                ty = Some(FloatConstantType::Float80);
761                remarks = remarks.with(Remarks::EXTENDED_SUFFIX);
762                1
763            }
764            b'd' | b'D' => {
765                let second = rest.get(1).copied();
766                let decimal = if byte == b'd' {
767                    matches!(second, Some(b'f' | b'd' | b'l'))
768                } else {
769                    matches!(second, Some(b'F' | b'D' | b'L'))
770                };
771                if decimal {
772                    return Err(FloatError::DecimalFloat);
773                }
774                ty = Some(FloatConstantType::Double);
775                remarks = remarks.with(Remarks::DOUBLE_SUFFIX);
776                1
777            }
778            _ => return Err(FloatError::InvalidSuffix),
779        };
780        rest = &rest[taken..];
781    }
782    Ok(FloatSuffix { ty: ty.unwrap_or(FloatConstantType::Double), imaginary, remarks })
783}
784
785/// Reads a suffix that starts with `f`, which is `float` on its own and one of the `_FloatN` or
786/// `_FloatNx` types when digits follow.
787///
788/// Returns the type, how many bytes it took and what is worth saying about it.
789fn float_n(
790    rest: &[u8],
791    target: &TargetInfo,
792) -> Result<(FloatConstantType, usize, Remarks), FloatError> {
793    let mut end = 1;
794    while rest.get(end).is_some_and(u8::is_ascii_digit) {
795        end += 1;
796    }
797    if end == 1 {
798        return Ok((FloatConstantType::Float, 1, Remarks::NONE));
799    }
800    // The `x` is lower case in every spelling gcc accepts, so `F64x` is a constant and `f64X`
801    // is not, however odd that looks next to the `F` being free.
802    let extended = rest.get(end) == Some(&b'x');
803    let ty = match (&rest[1..end], extended) {
804        // A suffix naming a type the target does not have is a suffix the target does not have,
805        // which is one message in gcc and not two: the constant is turned away where it is
806        // written rather than given the type and refused later.
807        (b"16", false) if !target.has_float16 => return Err(FloatError::UnsupportedType),
808        (b"128", false) if !target.has_float128 => return Err(FloatError::UnsupportedType),
809        (b"64", true) if target.float64x_format.is_none() => {
810            return Err(FloatError::UnsupportedType);
811        }
812        (b"16", false) => FloatConstantType::Float16,
813        (b"32", false) => FloatConstantType::Float32,
814        (b"64", false) => FloatConstantType::Float64,
815        (b"128", false) => FloatConstantType::Float128,
816        (b"32", true) => FloatConstantType::Float32x,
817        (b"64", true) => FloatConstantType::Float64x,
818        // gcc knows the name `_Float128x` and has the type on no target here, and says so
819        // rather than calling the suffix invalid. `_Float16x` is not a type at all.
820        (b"128", true) => return Err(FloatError::UnsupportedType),
821        _ => return Err(FloatError::InvalidSuffix),
822    };
823    Ok((ty, end + usize::from(extended), Remarks::EXTENDED_SUFFIX))
824}
825
826#[cfg(test)]
827mod tests {
828    use rucc_target::Triple;
829
830    use super::*;
831
832    fn linux() -> TargetInfo {
833        TargetInfo::new("x86_64-unknown-linux-gnu".parse::<Triple>().expect("a known triple"))
834    }
835
836    fn aarch64() -> TargetInfo {
837        TargetInfo::new("aarch64-unknown-linux-gnu".parse::<Triple>().expect("a known triple"))
838    }
839
840    /// The value and the type of a constant in the default dialect.
841    fn c23(text: &str) -> Result<IntConstant, IntError> {
842        integer(text, Std::C23, &linux())
843    }
844
845    /// The type of a constant in the given dialect, on x86-64 Linux.
846    fn kind(text: &str, std: Std) -> IntKind {
847        match integer(text, std, &linux()).expect("a valid constant").ty {
848            IntConstantType::Standard(kind) => kind,
849            IntConstantType::BitInt { .. } => panic!("{text} is a _BitInt constant"),
850        }
851    }
852
853    #[test]
854    fn a_constant_in_each_base_has_the_value_it_says() {
855        assert_eq!(c23("0").expect("zero").value, 0);
856        assert_eq!(c23("42").expect("decimal").value, 42);
857        assert_eq!(c23("0777").expect("octal").value, 0o777);
858        assert_eq!(c23("0xdeadBEEF").expect("hex").value, 0xdead_beef);
859        assert_eq!(c23("0b1010").expect("binary").value, 0b1010);
860        assert_eq!(c23("0X10").expect("upper case prefix").value, 16);
861        // A leading zero with nothing after it is a decimal zero rather than an octal one with
862        // no digits, which is the split that lets `0u` through and stops `08`.
863        assert_eq!(c23("0u").expect("zero with a suffix").value, 0);
864    }
865
866    #[test]
867    fn digit_separators_are_stripped_and_reported_before_c23() {
868        let value = c23("1'000'000").expect("a C23 constant");
869        assert_eq!(value.value, 1_000_000);
870        assert!(value.remarks.is_none());
871        assert_eq!(c23("0x1'0").expect("hex with a separator").value, 16);
872
873        let older = integer("1'000", Std::C17, &linux()).expect("still converted");
874        assert!(older.remarks.has(Remarks::SEPARATORS));
875        assert_eq!(older.value, 1000);
876    }
877
878    #[test]
879    fn the_type_of_a_decimal_constant_walks_the_signed_types_only() {
880        // Measured with `_Generic` on gcc 13.3, x86-64 Linux.
881        assert_eq!(kind("2147483647", Std::C23), IntKind::Int);
882        assert_eq!(kind("2147483648", Std::C23), IntKind::Long);
883        assert_eq!(kind("4294967295", Std::C23), IntKind::Long);
884        assert_eq!(kind("9223372036854775807", Std::C23), IntKind::Long);
885        // Past `long long` gcc reaches for `__int128` rather than for an unsigned type, and
886        // says so: the constant is so large that it is unsigned.
887        assert_eq!(kind("9223372036854775808", Std::C23), IntKind::Int128);
888        assert_eq!(kind("18446744073709551615", Std::C23), IntKind::Int128);
889        let large = c23("18446744073709551615").expect("fits __int128");
890        assert!(large.remarks.has(Remarks::UNSIGNED));
891    }
892
893    #[test]
894    fn a_constant_in_another_base_may_be_unsigned_without_saying_so() {
895        // This is the split that surprises people: `4294967295` is a `long` and `0xffffffff`
896        // is an `unsigned int`, because only the decimal list is signed types alone.
897        assert_eq!(kind("0xffffffff", Std::C23), IntKind::UInt);
898        assert_eq!(kind("0x7fffffff", Std::C23), IntKind::Int);
899        assert_eq!(kind("0x80000000", Std::C23), IntKind::UInt);
900        assert_eq!(kind("0x100000000", Std::C23), IntKind::Long);
901        assert_eq!(kind("0xffffffffffffffff", Std::C23), IntKind::ULong);
902        assert_eq!(kind("0777", Std::C23), IntKind::Int);
903        assert_eq!(kind("0b1010", Std::C23), IntKind::Int);
904        // And no remark, because nothing about it is surprising enough to say.
905        assert!(c23("0xffffffff").expect("a constant").remarks.is_none());
906    }
907
908    #[test]
909    fn c89_has_unsigned_long_in_the_decimal_list_and_no_long_long_in_any() {
910        // gcc under `-std=c89 -pedantic`: "this decimal constant is unsigned only in ISO C90",
911        // and eight bytes rather than sixteen.
912        assert_eq!(kind("18446744073709551615", Std::C89), IntKind::ULong);
913        assert_eq!(kind("18446744073709551615", Std::C99), IntKind::Int128);
914        let old = integer("18446744073709551615", Std::C89, &linux()).expect("a C89 constant");
915        assert!(old.remarks.has(Remarks::UNSIGNED));
916        // The suffix still reaches `long long`, with the remark gcc prints for it.
917        let long_long = integer("1ll", Std::C89, &linux()).expect("an extension");
918        assert!(long_long.remarks.has(Remarks::LONG_LONG));
919        assert_eq!(kind("1ll", Std::C89), IntKind::LongLong);
920        assert!(integer("1ll", Std::C99, &linux()).expect("standard").remarks.is_none());
921    }
922
923    #[test]
924    fn a_suffix_narrows_the_list_it_does_not_pick_the_type() {
925        assert_eq!(kind("1u", Std::C23), IntKind::UInt);
926        assert_eq!(kind("1l", Std::C23), IntKind::Long);
927        assert_eq!(kind("1ul", Std::C23), IntKind::ULong);
928        assert_eq!(kind("1ll", Std::C23), IntKind::LongLong);
929        assert_eq!(kind("1llu", Std::C23), IntKind::ULongLong);
930        // The suffix is a floor rather than an answer: `4294967296u` is an `unsigned long`
931        // because `unsigned int` cannot hold it.
932        assert_eq!(kind("4294967296u", Std::C23), IntKind::ULong);
933        assert_eq!(kind("0xffffffffu", Std::C23), IntKind::UInt);
934    }
935
936    #[test]
937    fn the_letters_of_a_suffix_may_be_in_either_case_but_not_both() {
938        for text in ["1u", "1U", "1l", "1L", "1ll", "1LL", "1ul", "1lu", "1uL", "1LLU", "1llu"] {
939            assert!(c23(text).is_ok(), "{text} is a constant in both compilers");
940        }
941        for text in ["1lL", "1Ll", "1uu", "1lul", "1z", "1uz", "1f", "1x", "1_000"] {
942            assert_eq!(c23(text), Err(IntError::InvalidSuffix), "{text} is not");
943        }
944    }
945
946    #[test]
947    fn a_bit_int_constant_has_the_narrowest_type_that_holds_it() {
948        // Measured against clang, which is the only one of the two that has the type.
949        let cases = [
950            ("0wb", true, 2),
951            ("1wb", true, 2),
952            ("3wb", true, 3),
953            ("42wb", true, 7),
954            ("255wb", true, 9),
955            ("0uwb", false, 1),
956            ("1uwb", false, 1),
957            ("255uwb", false, 8),
958            ("256uwb", false, 9),
959            ("0xffffffffffffffffuwb", false, 64),
960        ];
961        for (text, signed, width) in cases {
962            let constant = c23(text).expect("a _BitInt constant");
963            assert_eq!(
964                constant.ty,
965                IntConstantType::BitInt { signed, width },
966                "{text} is the wrong width"
967            );
968        }
969        // Either order, either case, and never with a length suffix.
970        for text in ["1uwb", "1wbu", "1UWB", "1WBu", "1uWB"] {
971            assert!(c23(text).is_ok(), "{text} is a constant in clang");
972        }
973        for text in ["1wB", "1Wb", "1lwb", "1wbl", "1wbwb"] {
974            assert_eq!(c23(text), Err(IntError::InvalidSuffix), "{text} is not");
975        }
976        // Before C23 it is still converted, and still worth a word.
977        let older = integer("1wb", Std::C17, &linux()).expect("clang accepts it everywhere");
978        assert!(older.remarks.has(Remarks::BIT_INT));
979    }
980
981    #[test]
982    fn a_binary_constant_is_an_extension_before_c23() {
983        assert!(c23("0b1").expect("standard in C23").remarks.is_none());
984        let older = integer("0b1", Std::C17, &linux()).expect("both compilers accept it");
985        assert!(older.remarks.has(Remarks::BINARY));
986    }
987
988    #[test]
989    fn an_octal_constant_names_the_digit_that_is_not_one() {
990        assert_eq!(c23("08"), Err(IntError::InvalidOctalDigit));
991        assert_eq!(c23("0778"), Err(IntError::InvalidOctalDigit));
992        assert_eq!(c23("09"), Err(IntError::InvalidOctalDigit));
993        // A `9` elsewhere is fine, and the message is only for constants that began with `0`.
994        assert_eq!(c23("9").expect("decimal").value, 9);
995    }
996
997    #[test]
998    fn a_prefix_with_no_digits_after_it_is_not_a_constant() {
999        assert_eq!(c23("0x"), Err(IntError::NoDigits));
1000        assert_eq!(c23("0b"), Err(IntError::NoDigits));
1001    }
1002
1003    #[test]
1004    fn a_constant_larger_than_any_type_is_refused_rather_than_wrapped() {
1005        // gcc accumulates in sixty four bits and silently gives this the value zero and the
1006        // type `int` after a warning. That is the one measured behaviour here we refuse to
1007        // reproduce, and clang refuses it too.
1008        assert_eq!(c23("340282366920938463463374607431768211456"), Err(IntError::TooLarge));
1009        assert_eq!(c23("0x100000000000000000000000000000000"), Err(IntError::TooLarge));
1010        // 2^127 fits in the accumulator and in no signed type, and the decimal list has no
1011        // unsigned one to fall back to.
1012        assert_eq!(c23("170141183460469231731687303715884105728"), Err(IntError::TooLarge));
1013        // The same value written in hex reaches `unsigned __int128`, because that list has it.
1014        assert_eq!(kind("0x80000000000000000000000000000000", Std::C23), IntKind::UInt128);
1015        assert_eq!(kind("0xffffffffffffffffffffffffffffffff", Std::C23), IntKind::UInt128);
1016    }
1017
1018    #[test]
1019    fn a_floating_constant_is_handed_back_rather_than_refused() {
1020        for text in ["1.0", ".5", "1.", "1e5", "1E-5", "1e", "0x1p3", "0x1.8p+1", "1.5e3"] {
1021            assert_eq!(c23(text), Err(IntError::Floating), "{text} belongs to the other path");
1022        }
1023        // A leading zero does not make this an octal constant with a digit that does not
1024        // exist. gcc compiles it, as eight hundred thousand.
1025        assert_eq!(c23("08e5"), Err(IntError::Floating));
1026        // A hexadecimal `e` is a digit, not an exponent, and `1f` is an integer with a suffix
1027        // that does not exist rather than a float. Both compilers split them there.
1028        assert_eq!(c23("0xe5").expect("hex digits").value, 0xe5);
1029        assert_eq!(c23("1f"), Err(IntError::InvalidSuffix));
1030    }
1031
1032    #[test]
1033    fn the_type_comes_from_the_target_and_not_from_the_host() {
1034        // `4294967295` is a `long` where `long` is sixty four bits and a `long long` where it
1035        // is thirty two. A compiler that asked its own platform gets one of these wrong.
1036        let windows =
1037            TargetInfo::new("x86_64-pc-windows-msvc".parse::<Triple>().expect("a known triple"));
1038        let on_windows = integer("4294967295", Std::C23, &windows).expect("a constant");
1039        assert_eq!(on_windows.ty, IntConstantType::Standard(IntKind::LongLong));
1040        assert_eq!(kind("4294967295", Std::C23), IntKind::Long);
1041    }
1042
1043    /// A floating constant in the default dialect, on x86-64 Linux.
1044    fn float(text: &str) -> Result<FloatConstant, FloatError> {
1045        floating(text, Std::C23, &linux())
1046    }
1047
1048    /// The bits of a constant's value, which is the form every measured row here was taken in.
1049    fn bits(text: &str) -> u128 {
1050        float(text).expect("a valid constant").value.to_bits()
1051    }
1052
1053    #[test]
1054    fn a_constant_with_no_suffix_is_a_double() {
1055        let constant = float("1.5").expect("a constant");
1056        assert_eq!(constant.ty, FloatConstantType::Double);
1057        assert!(!constant.imaginary);
1058        assert!(constant.remarks.is_none());
1059        assert_eq!(constant.value.to_bits(), 0x3ff8_0000_0000_0000);
1060        assert_eq!(bits("0.1"), 0x3fb9_9999_9999_999a);
1061        assert_eq!(bits(".5"), 0x3fe0_0000_0000_0000);
1062        assert_eq!(bits("1."), 0x3ff0_0000_0000_0000);
1063        assert_eq!(bits("1e5"), 0x40f8_6a00_0000_0000);
1064        assert_eq!(bits("0x1p3"), 0x4020_0000_0000_0000);
1065        // A leading zero is not an octal prefix once there is an exponent, so this is eight
1066        // hundred thousand and gcc compiles it as one.
1067        assert_eq!(bits("08e5"), 0x4128_6a00_0000_0000);
1068    }
1069
1070    #[test]
1071    fn the_suffix_names_the_type_rather_than_narrowing_a_list() {
1072        // Measured with `_Generic` on gcc 13.3, x86-64 Linux.
1073        let cases = [
1074            ("1.0", FloatConstantType::Double),
1075            ("1.0f", FloatConstantType::Float),
1076            ("1.0F", FloatConstantType::Float),
1077            ("1.0l", FloatConstantType::LongDouble),
1078            ("1.0L", FloatConstantType::LongDouble),
1079            ("1.0d", FloatConstantType::Double),
1080            ("1.0q", FloatConstantType::Float128),
1081            ("1.0w", FloatConstantType::Float80),
1082            ("1.0f16", FloatConstantType::Float16),
1083            ("1.0F16", FloatConstantType::Float16),
1084            ("1.0f32", FloatConstantType::Float32),
1085            ("1.0f64", FloatConstantType::Float64),
1086            ("1.0f128", FloatConstantType::Float128),
1087            ("1.0f32x", FloatConstantType::Float32x),
1088            ("1.0F64x", FloatConstantType::Float64x),
1089        ];
1090        for (text, ty) in cases {
1091            assert_eq!(float(text).expect("a constant").ty, ty, "{text} has the wrong type");
1092        }
1093    }
1094
1095    #[test]
1096    fn each_type_is_converted_in_the_format_the_target_has_for_it() {
1097        // Every row measured by printing the bytes of the constant on gcc 13.3, x86-64 Linux.
1098        // The two that surprise are `_Float32x`, which is plain `double`, and `_Float64x`,
1099        // which is the x87 format and so the same bits as `long double` and `__float80`.
1100        assert_eq!(bits("0.1f"), 0x3dcc_cccd);
1101        assert_eq!(bits("0.1f16"), 0x2e66);
1102        assert_eq!(bits("0.1f32x"), 0x3fb9_9999_9999_999a);
1103        assert_eq!(bits("0.1f64x"), 0x3ffb_cccc_cccc_cccc_cccd);
1104        assert_eq!(bits("0.1w"), 0x3ffb_cccc_cccc_cccc_cccd);
1105        assert_eq!(bits("0.1l"), 0x3ffb_cccc_cccc_cccc_cccd);
1106        assert_eq!(bits("0.1q"), 0x3ffb_9999_9999_9999_9999_9999_9999_999a);
1107        assert_eq!(bits("0.1f128"), 0x3ffb_9999_9999_9999_9999_9999_9999_999a);
1108        assert_eq!(bits("1.0l"), 0x3fff_8000_0000_0000_0000);
1109    }
1110
1111    #[test]
1112    fn the_format_comes_from_the_target_and_not_from_the_host() {
1113        // `long double` is 128 bits wide on both of these and it is not the same type on both,
1114        // which is the whole reason the target carries a format and not only a width.
1115        let arm = floating("1.0l", Std::C23, &aarch64()).expect("a constant");
1116        assert_eq!(arm.value.to_bits(), 0x3fff_0000_0000_0000_0000_0000_0000_0000);
1117        assert_eq!(bits("1.0l"), 0x3fff_8000_0000_0000_0000);
1118        // And `_Float64x` follows it, being whatever the target has above `_Float64`.
1119        let arm_wide = floating("0.1f64x", Std::C23, &aarch64()).expect("a constant");
1120        assert_eq!(arm_wide.value.to_bits(), 0x3ffb_9999_9999_9999_9999_9999_9999_999a);
1121        // Where `long double` is `double` the constant is a `double` too.
1122        let windows =
1123            TargetInfo::new("x86_64-pc-windows-msvc".parse::<Triple>().expect("a known triple"));
1124        let on_windows = floating("1.0l", Std::C23, &windows).expect("a constant");
1125        assert_eq!(on_windows.value.to_bits(), 0x3ff0_0000_0000_0000);
1126    }
1127
1128    #[test]
1129    fn the_case_rules_of_a_floating_suffix_are_not_uniform() {
1130        // The `f` of a `_FloatN` suffix is free and the `x` of a `_FloatNx` one is not, and the
1131        // two letters of a decimal suffix have to agree. Measured on gcc 13.3, which accepts
1132        // every one of these in every dialect.
1133        for text in ["1.0f", "1.0F", "1.0L", "1.0Q", "1.0W", "1.0F32", "1.0f64x", "1.0F64x"] {
1134            assert!(float(text).is_ok(), "{text} is a constant in gcc");
1135        }
1136        for text in ["1.0F32X", "1.0f32X", "1.0f16x", "1.0ff", "1.0fl", "1.0lf", "1.0fF", "1.0LL"] {
1137            assert_eq!(float(text), Err(FloatError::InvalidSuffix), "{text} is not");
1138        }
1139    }
1140
1141    #[test]
1142    fn an_imaginary_suffix_may_sit_on_either_side_of_the_type() {
1143        for text in ["1.0i", "1.0j", "1.0I", "1.0J", "1.0if", "1.0fi", "1.0Li", "1.0iL", "1.0f16i"]
1144        {
1145            let constant = float(text).expect("a constant in gcc");
1146            assert!(constant.imaginary, "{text} is imaginary");
1147            assert!(constant.remarks.has(Remarks::IMAGINARY));
1148        }
1149        assert_eq!(float("1.0ii"), Err(FloatError::InvalidSuffix));
1150        assert_eq!(float("1.0ij"), Err(FloatError::InvalidSuffix));
1151        assert!(!float("1.0f").expect("a constant").imaginary);
1152    }
1153
1154    #[test]
1155    fn a_decimal_floating_constant_is_recognised_and_refused() {
1156        // The constant is well formed and there is nowhere in this compiler to put its value.
1157        for text in ["1.0df", "1.0dd", "1.0dl", "1.0DF", "1.0DD", "1.0DL"] {
1158            assert_eq!(float(text), Err(FloatError::DecimalFloat), "{text} is a decimal float");
1159        }
1160        // The letters have to agree about case, so these are not decimal floats and not
1161        // constants either.
1162        for text in ["1.0Df", "1.0dF", "1.0dD", "1.0Dl"] {
1163            assert_eq!(float(text), Err(FloatError::InvalidSuffix), "{text} is neither");
1164        }
1165        // A `d` on its own is a `double` written the long way, which gcc allows everywhere.
1166        let long_way = float("1.0d").expect("a GCC extension");
1167        assert_eq!(long_way.ty, FloatConstantType::Double);
1168        assert!(long_way.remarks.has(Remarks::DOUBLE_SUFFIX));
1169    }
1170
1171    #[test]
1172    fn a_type_the_target_does_not_have_is_refused_by_name() {
1173        // gcc calls this one unsupported rather than invalid, and it is unsupported on every
1174        // target: `_Float128x` is a type nothing here has.
1175        assert_eq!(float("1.0f128x"), Err(FloatError::UnsupportedType));
1176        // `__float80` is the x87 format, which only x86 has.
1177        assert_eq!(floating("1.0w", Std::C23, &aarch64()), Err(FloatError::UnsupportedType));
1178        assert!(float("1.0w").is_ok());
1179    }
1180
1181    #[test]
1182    fn a_suffix_goes_where_the_type_it_names_goes() {
1183        // The rows are gcc 13's, measured with the cross compilers. armv7 has no type wider than
1184        // a `double`, so all three of these are refused there, and i686 has the quad and not the
1185        // half, so one of them is.
1186        // The three field triple cannot spell either of these machines, so they come from the
1187        // tuple, which spells all forty two rows.
1188        let arm =
1189            TargetInfo::for_tuple("armv7-linux-gnueabihf".parse().expect("a row in the table"));
1190        let i686 = TargetInfo::for_tuple("i686-linux-gnu".parse().expect("a row in the table"));
1191        for text in ["1.0f16", "1.0f128", "1.0q", "1.0f64x"] {
1192            assert_eq!(
1193                floating(text, Std::C23, &arm),
1194                Err(FloatError::UnsupportedType),
1195                "{text} on armv7"
1196            );
1197            assert!(float(text).is_ok(), "{text} on x86-64");
1198        }
1199        assert_eq!(
1200            floating("1.0f16", Std::C23, &i686),
1201            Err(FloatError::UnsupportedType),
1202            "the half is the one i686 has no format for"
1203        );
1204        assert!(floating("1.0f128", Std::C23, &i686).is_ok());
1205        // The interchange types every machine has keep working on the machine that has least.
1206        for text in ["1.0f32", "1.0f64", "1.0f32x"] {
1207            assert!(floating(text, Std::C23, &arm).is_ok(), "{text} on armv7");
1208        }
1209    }
1210
1211    #[test]
1212    fn a_hexadecimal_constant_needs_an_exponent_and_a_decimal_one_does_not() {
1213        // `f` is a hexadecimal digit, so without the exponent there is no telling the number
1214        // from the suffix. Both compilers require it.
1215        assert_eq!(float("0x1.8"), Err(FloatError::MissingExponent));
1216        assert_eq!(bits("0x1.8p0"), 0x3ff8_0000_0000_0000);
1217        assert_eq!(bits("0x.8p1"), 0x3ff0_0000_0000_0000);
1218        assert_eq!(bits("1.5"), 0x3ff8_0000_0000_0000);
1219        for text in ["1.0e", "1e+", "1e-", "0x1p", "0x1p+"] {
1220            assert_eq!(float(text), Err(FloatError::NoExponentDigits), "{text} has no exponent");
1221        }
1222        assert_eq!(float("1.2.3"), Err(FloatError::TooManyPoints));
1223    }
1224
1225    #[test]
1226    fn an_integer_constant_is_handed_back_rather_than_refused() {
1227        for text in ["1", "0", "0x10", "1u", "0777", "1wb", "0b1", "0xe5", "1f"] {
1228            assert_eq!(float(text), Err(FloatError::Integer), "{text} belongs to the other path");
1229        }
1230    }
1231
1232    #[test]
1233    fn a_value_past_the_range_of_its_type_is_still_a_constant() {
1234        let large = float("1e400").expect("a constant gcc compiles");
1235        assert!(large.value.is_infinite());
1236        assert!(large.remarks.has(Remarks::OUT_OF_RANGE));
1237        let small = float("1e-400").expect("a constant gcc compiles");
1238        assert!(small.value.is_zero());
1239        assert!(small.remarks.has(Remarks::TRUNCATED));
1240        // The same two in the format the suffix asked for rather than in `double`.
1241        assert!(float("1e39f").expect("a constant").remarks.has(Remarks::OUT_OF_RANGE));
1242        assert!(float("1e-46f").expect("a constant").remarks.has(Remarks::TRUNCATED));
1243        assert!(float("1e-4951l").expect("a constant").remarks.has(Remarks::TRUNCATED));
1244        // A subnormal is a number the program can use, and gcc says nothing about it.
1245        let subnormal = float("1e-320").expect("a constant");
1246        assert!(!subnormal.value.is_zero());
1247        assert!(subnormal.remarks.is_none());
1248    }
1249
1250    #[test]
1251    fn the_dialect_decides_what_a_constant_is_worth_saying_about() {
1252        // "use of C99 hexadecimal floating constant", which gcc says under C89 and not after.
1253        let old = floating("0x1p3", Std::C89, &linux()).expect("gcc compiles it anyway");
1254        assert!(old.remarks.has(Remarks::HEX_FLOAT));
1255        assert!(floating("0x1p3", Std::C99, &linux()).expect("standard").remarks.is_none());
1256        // Separators are C23 in both compilers, in the number and in the exponent.
1257        assert_eq!(bits("1'0.5"), 0x4025_0000_0000_0000);
1258        assert_eq!(bits("1.0e1'0"), 0x4202_a05f_2000_0000);
1259        assert!(float("0x1'0p0").expect("a C23 constant").remarks.is_none());
1260        let older = floating("1'0.5", Std::C17, &linux()).expect("still converted");
1261        assert!(older.remarks.has(Remarks::SEPARATORS));
1262        // And every extension suffix is accepted in every dialect, with a word about it.
1263        for text in ["1.0q", "1.0w", "1.0f16", "1.0f32x"] {
1264            let constant = floating(text, Std::C89, &linux()).expect("gcc accepts it in C89");
1265            assert!(constant.remarks.has(Remarks::EXTENDED_SUFFIX), "{text} is not standard");
1266        }
1267        assert!(float("1.0f").expect("a constant").remarks.is_none());
1268    }
1269}