Skip to main content

rucc_lex/
number.rs

1//! Numeric constants: the value, and the type the standard's table walk gives it.
2//!
3//! Design: `spec/06-lexer-and-parser.md` section 6.1.
4//!
5//! This is the second piece of phase 7. A preprocessing number is a loose thing, deliberately
6//! looser than a constant, so `1.2.3` and `0x1p+3` are both one pp-token and only here does
7//! anyone ask what they mean. What comes back is a value and a type, and both of them are
8//! places a compiler quietly goes wrong.
9//!
10//! There are two entry points, [`integer`] and [`floating`], and either of them hands the
11//! spelling to the other rather than reporting an error when it turns out to belong there. The
12//! split is not where a reader expects it: `1e` is a floating constant with no exponent digits
13//! and `1f` is an integer constant with a suffix that does not exist, and both compilers agree
14//! on that, because the exponent marker is part of a preprocessing number and the suffix letter
15//! is not part of anything.
16//!
17//! The value is accumulated in a `u128` with every step checked, so a constant too large to
18//! represent is a diagnostic rather than a number the program did not write. gcc 13.3 does not
19//! do that: its accumulator is sixty four bits, and `18446744073709551616` compiles to zero of
20//! type `int` after a warning nobody reads. That is not a behaviour worth reproducing, so ours
21//! is the only measured difference here that is deliberate: past a hundred and twenty eight
22//! bits the constant is refused. clang refuses it too, one bit earlier.
23//!
24//! The type is the standard's table walk, 6.4.4.1p5: a candidate list chosen by the base and
25//! the suffix, walked in order, and the first type that holds the value wins. The list is not
26//! the same in every dialect. C89 puts `unsigned long` in the list for a decimal constant with
27//! no suffix, which is what makes `18446744073709551615` an `unsigned long` under `-std=c89`
28//! and something wider under `-std=c99`, and gcc says so in as many words: "this decimal
29//! constant is unsigned only in ISO C90". Both compilers keep `long long` out of the C89 lists
30//! and accept it when the suffix asks for it.
31//!
32//! `__int128` is on the end of every list, which is what gcc does and clang does not.
33//! `9223372036854775808` is an `__int128` in gcc 13.3 and an `unsigned long long` in clang,
34//! and the difference is visible to a program: negate it and gcc gives a negative number.
35//! We follow gcc, because the alternative silently turns a signed constant unsigned.
36//!
37//! The rest was measured the same way, by writing the constant and asking `_Generic` what it
38//! is, on gcc 13.3 on x86-64 Linux and on clang:
39//!
40//! The suffix letters may be in either case but not both, so `1ll` and `1LL` are constants and
41//! `1lL` is not, and the same rule holds for `wb`. The unsigned suffix may come before or after
42//! the length suffix. `wb` does not combine with `l` at all.
43//!
44//! Binary constants are accepted in every dialect by both compilers, as an extension before
45//! C23. Digit separators are C23 only in both. `_BitInt` constants are C23 in the standard,
46//! clang accepts them in every dialect, and gcc 13.3 has no `_BitInt` at all.
47//!
48//! A `wb` constant has the narrowest type that holds it, which for a signed one includes the
49//! sign bit and is never less than two: `1wb` is `_BitInt(2)`, `42wb` is `_BitInt(7)`, `255uwb`
50//! is `unsigned _BitInt(8)` and `0uwb` is `unsigned _BitInt(1)`. Measured against clang, since
51//! gcc 13.3 cannot say.
52//!
53//! # Floating constants
54//!
55//! A floating constant has none of the table walk about it: the suffix names the type outright,
56//! and with no suffix it is a `double`. What it has instead is a list of suffixes that is much
57//! longer than the standard's three, and a conversion that has to be exactly right.
58//!
59//! The conversion is [`rucc_base::float`], correctly rounded and done in software, so that the
60//! bits do not depend on the machine the compiler runs on. The type decides the format and the
61//! target decides what some of the types are: `long double` is the x87 eighty bit format on
62//! x86-64 Linux and true quad precision on AArch64 Linux, and both of them say 128 bits wide,
63//! which is why [`TargetInfo::long_double_format`] exists.
64//!
65//! The suffixes were measured on gcc 13.3 on x86-64 Linux, with `_Generic` for the type and by
66//! printing the bytes for the format. `f` is `float` and `l` is `long double`, and then the
67//! extensions: `q` is `__float128`, `w` is the x87 `__float80`, `d` is a `double` written the
68//! long way, `f16` `f32` `f64` `f128` are the `_FloatN` types and `f32x` `f64x` the `_FloatNx`
69//! ones, and `i` or `j` in either position makes the constant imaginary. `_Float32x` turns out
70//! to be plain `double` and `_Float64x` the x87 format, which is not what the names suggest:
71//! `0.1f32x` is `0x3fb999999999999a` and `0.1f64x` is `0x3ffbcccccccccccccccd`.
72//!
73//! The case rules are their own small grammar. The `f` of a `_FloatN` suffix may be either case
74//! and the trailing `x` may not, so `F64x` is a constant and `f64X` is not. A decimal float
75//! suffix is two letters that have to agree, so `df` and `DF` are constants and `Df` is not.
76//! Every one of these is accepted in every dialect, C89 included, and every one of them is
77//! worth a remark when the dialect did not ask for it.
78//!
79//! Decimal floating constants are refused rather than converted, because there is no decimal
80//! float value anywhere in this compiler to put one in. That is a gap and the error says so.
81
82use rucc_base::float::{Float, Format, ParseError, Status};
83use rucc_session::Std;
84use rucc_target::{Arch, TargetInfo};
85use rucc_types::{IntKind, int_width};
86
87use crate::remarks::Remarks;
88
89/// A converted integer constant.
90#[derive(Debug, Clone, Copy, PartialEq, Eq)]
91pub struct IntConstant {
92    /// The value, which is never negative: a minus sign is an operator and not part of the
93    /// constant, which is why `-2147483648` is a `long` on a 32-bit `int` and the reason
94    /// `INT_MIN` is spelled the way it is in `limits.h`.
95    pub value: u128,
96    /// The type the table walk arrived at.
97    pub ty: IntConstantType,
98    /// What is worth saying about the constant, for the caller that holds the span.
99    pub remarks: Remarks,
100}
101
102/// The type of an integer constant.
103#[derive(Debug, Clone, Copy, PartialEq, Eq)]
104pub enum IntConstantType {
105    /// One of the integer kinds, chosen by the table walk.
106    Standard(IntKind),
107    /// A `_BitInt` of exactly the width it takes to hold the value.
108    BitInt {
109        /// Whether the type is signed, which it is when the `u` suffix was not there.
110        signed: bool,
111        /// The width in bits, including the sign bit when there is one.
112        width: u32,
113    },
114}
115
116impl IntConstantType {
117    /// A suffix that names this type, for a printer putting the constant back.
118    ///
119    /// The value decides the rest. Nothing spells `__int128`, because no suffix does and none is
120    /// needed: a constant only gets that type by being too large for every other one, so writing
121    /// the value back gets the same answer out of the table walk.
122    #[must_use]
123    pub const fn suffix(self) -> &'static str {
124        match self {
125            IntConstantType::Standard(kind) => match kind {
126                IntKind::UInt | IntKind::UInt128 => "u",
127                IntKind::Long => "l",
128                IntKind::ULong => "ul",
129                IntKind::LongLong => "ll",
130                IntKind::ULongLong => "ull",
131                _ => "",
132            },
133            IntConstantType::BitInt { signed: true, .. } => "wb",
134            IntConstantType::BitInt { signed: false, .. } => "uwb",
135        }
136    }
137}
138
139/// Why a preprocessing number is not an integer constant.
140///
141/// [`IntError::Floating`] is not a diagnostic. It means the spelling belongs to the floating
142/// path, and it is an error here so that the caller cannot forget to ask.
143#[derive(Debug, Clone, Copy, PartialEq, Eq)]
144pub enum IntError {
145    /// This is a floating constant. Nothing is wrong with it.
146    Floating,
147    /// The characters after the digits are not a suffix.
148    InvalidSuffix,
149    /// An `8` or a `9` in a constant that started with `0`.
150    InvalidOctalDigit,
151    /// `0x` or `0b` with no digits after it.
152    NoDigits,
153    /// Larger than any integer type, or than the hundred and twenty eight bits the value is
154    /// accumulated in.
155    TooLarge,
156}
157
158impl IntError {
159    /// What to print, in GCC's words where GCC has any.
160    ///
161    /// The offending character is not in the message, because the caller has the spelling and
162    /// the span and can say `invalid suffix "ux" on integer constant` the way GCC does.
163    #[must_use]
164    pub const fn message(self) -> &'static str {
165        match self {
166            IntError::Floating => "not an integer constant",
167            IntError::InvalidSuffix => "invalid suffix on integer constant",
168            IntError::InvalidOctalDigit => "invalid digit in octal constant",
169            IntError::NoDigits => "no digits in integer constant",
170            IntError::TooLarge => "integer constant is too large to be represented in any type",
171        }
172    }
173}
174
175/// A converted floating constant.
176#[derive(Debug, Clone, Copy, PartialEq, Eq)]
177pub struct FloatConstant {
178    /// The value, correctly rounded into the format of its type.
179    pub value: Float,
180    /// The type the suffix named, which is `double` when there was no suffix.
181    pub ty: FloatConstantType,
182    /// Whether an `i` or a `j` made this the imaginary part of a complex constant. The value is
183    /// the real number that was written either way, so the caller builds the complex one.
184    pub imaginary: bool,
185    /// What is worth saying about the constant, for the caller that holds the span.
186    pub remarks: Remarks,
187}
188
189/// The type of a floating constant.
190///
191/// This is not [`rucc_types::FloatKind`], which has the three types C has. The suffixes reach
192/// further than that, and a constant knows exactly which type it was written as long before
193/// anything has to decide what that type is on this target.
194#[derive(Debug, Clone, Copy, PartialEq, Eq)]
195pub enum FloatConstantType {
196    /// `f` or `F`.
197    Float,
198    /// No suffix, or the `d` and `D` that GCC also accepts for it.
199    Double,
200    /// `l` or `L`.
201    LongDouble,
202    /// `f16` or `F16`, which is `_Float16`.
203    Float16,
204    /// `f32` or `F32`, which is `_Float32` and is the same format as `float`.
205    Float32,
206    /// `f64` or `F64`, which is `_Float64` and is the same format as `double`.
207    Float64,
208    /// `f128` or `F128`, and `q` or `Q`, which is `_Float128` and `__float128`. GCC keeps the
209    /// two spellings as distinct types and they are the same format, which is all this says.
210    Float128,
211    /// `f32x` or `F32x`, which is `_Float32x`. The name suggests something wider than
212    /// `_Float32` and on every target here it is exactly `double`.
213    Float32x,
214    /// `f64x` or `F64x`, which is `_Float64x`: the widest format the target has beyond
215    /// `_Float64`, so the x87 one on x86 and quad precision elsewhere.
216    Float64x,
217    /// `w` or `W`, which is GCC's `__float80`. It is the x87 format whatever `long double` is,
218    /// which is the reason it is not the same thing as [`FloatConstantType::LongDouble`], and
219    /// GCC has it on x86 only.
220    Float80,
221}
222
223impl FloatConstantType {
224    /// The format a constant of this type is converted in.
225    ///
226    /// The two that depend on the target are the two that have to: `long double` is the x87
227    /// format on x86-64 Linux and quad precision on AArch64 Linux, and `_Float64x` is whatever
228    /// the target has above `_Float64`, which is the same split.
229    #[must_use]
230    pub fn format(self, target: &TargetInfo) -> Format {
231        match self {
232            FloatConstantType::Float | FloatConstantType::Float32 => Format::Single,
233            FloatConstantType::Double
234            | FloatConstantType::Float64
235            | FloatConstantType::Float32x => Format::Double,
236            FloatConstantType::LongDouble => target.long_double_format,
237            FloatConstantType::Float16 => Format::Half,
238            FloatConstantType::Float128 => Format::Quad,
239            FloatConstantType::Float64x if target.triple.arch == Arch::X86_64 => {
240                Format::X87Extended
241            }
242            FloatConstantType::Float64x => Format::Quad,
243            FloatConstantType::Float80 => Format::X87Extended,
244        }
245    }
246
247    /// The C spelling of the type, which a diagnostic naming it has to print.
248    ///
249    /// GCC says "floating constant exceeds range of 'double'" and puts the type in the message,
250    /// so the type has to be able to say what it is called.
251    #[must_use]
252    pub const fn name(self) -> &'static str {
253        match self {
254            FloatConstantType::Float => "float",
255            FloatConstantType::Double => "double",
256            FloatConstantType::LongDouble => "long double",
257            FloatConstantType::Float16 => "_Float16",
258            FloatConstantType::Float32 => "_Float32",
259            FloatConstantType::Float64 => "_Float64",
260            FloatConstantType::Float128 => "_Float128",
261            FloatConstantType::Float32x => "_Float32x",
262            FloatConstantType::Float64x => "_Float64x",
263            FloatConstantType::Float80 => "__float80",
264        }
265    }
266
267    /// The suffix that names this type, for a printer putting the constant back.
268    ///
269    /// One spelling per type rather than the one that was written, so `0.1q` and `0.1f128` both
270    /// come back as `f128`. They are the same type, and the printer's job is to mean the same
271    /// thing rather than to look the same.
272    #[must_use]
273    pub const fn suffix(self) -> &'static str {
274        match self {
275            FloatConstantType::Float => "f",
276            FloatConstantType::Double => "",
277            FloatConstantType::LongDouble => "l",
278            FloatConstantType::Float16 => "f16",
279            FloatConstantType::Float32 => "f32",
280            FloatConstantType::Float64 => "f64",
281            FloatConstantType::Float128 => "f128",
282            FloatConstantType::Float32x => "f32x",
283            FloatConstantType::Float64x => "f64x",
284            FloatConstantType::Float80 => "w",
285        }
286    }
287}
288
289/// Why a preprocessing number is not a floating constant.
290///
291/// [`FloatError::Integer`] is not a diagnostic, in the same way [`IntError::Floating`] is not:
292/// it means the spelling belongs to the other path.
293#[derive(Debug, Clone, Copy, PartialEq, Eq)]
294pub enum FloatError {
295    /// This is an integer constant. Nothing is wrong with it.
296    Integer,
297    /// The characters after the number are not a suffix.
298    InvalidSuffix,
299    /// A hexadecimal floating constant with no `p` exponent. The exponent is required there and
300    /// not optional as it is in a decimal one, because `f` is a hexadecimal digit and there
301    /// would be no way to tell a suffix from the number.
302    MissingExponent,
303    /// An `e` or a `p` with no digits after it.
304    NoExponentDigits,
305    /// A constant with a point and no digits at all.
306    NoDigits,
307    /// More than one point, which is a preprocessing number and not a constant.
308    TooManyPoints,
309    /// A `df`, `dd` or `dl` suffix. The constant is well formed and this compiler has nowhere
310    /// to put a decimal floating value yet.
311    DecimalFloat,
312    /// A suffix naming a type this target does not have, which is `w` anywhere but x86 and
313    /// `f128x` everywhere.
314    UnsupportedType,
315}
316
317impl FloatError {
318    /// What to print, in GCC's words where GCC has any.
319    #[must_use]
320    pub const fn message(self) -> &'static str {
321        match self {
322            FloatError::Integer => "not a floating constant",
323            FloatError::InvalidSuffix => "invalid suffix on floating constant",
324            FloatError::MissingExponent => "hexadecimal floating constants require an exponent",
325            FloatError::NoExponentDigits => "exponent has no digits",
326            FloatError::NoDigits => "no digits in floating constant",
327            FloatError::TooManyPoints => "too many decimal points in number",
328            FloatError::DecimalFloat => "decimal floating constants are not supported yet",
329            FloatError::UnsupportedType => {
330                "the type of this floating constant is not supported on this target"
331            }
332        }
333    }
334}
335
336/// Converts the spelling of a preprocessing number into an integer constant.
337///
338/// # Errors
339///
340/// [`IntError`], one case of which is that the spelling is a floating constant rather than a
341/// malformed integer one.
342pub fn integer(text: &str, std: Std, target: &TargetInfo) -> Result<IntConstant, IntError> {
343    let bytes = text.as_bytes();
344    let (base, start) = base_of(bytes);
345    if is_floating(bytes, base) {
346        return Err(IntError::Floating);
347    }
348    let mut remarks = Remarks::NONE;
349    if base == 2 && std < Std::C23 {
350        remarks = remarks.with(Remarks::BINARY);
351    }
352
353    let mut value: u128 = 0;
354    let mut digits = 0;
355    let mut index = start;
356    while index < bytes.len() {
357        let byte = bytes[index];
358        if byte == b'\'' {
359            // A separator is only a separator between two digits. The scanner keeps one in the
360            // number only when an identifier character follows, so a trailing one arrives here
361            // as a suffix instead and is refused as one.
362            if digits == 0 || index + 1 >= bytes.len() || digit(bytes[index + 1], base).is_none() {
363                return Err(IntError::InvalidSuffix);
364            }
365            if std < Std::C23 {
366                remarks = remarks.with(Remarks::SEPARATORS);
367            }
368            index += 1;
369            continue;
370        }
371        let Some(digit) = digit(byte, base) else {
372            break;
373        };
374        value = value
375            .checked_mul(u128::from(base))
376            .and_then(|shifted| shifted.checked_add(u128::from(digit)))
377            .ok_or(IntError::TooLarge)?;
378        digits += 1;
379        index += 1;
380    }
381    if digits == 0 {
382        // `0x` with nothing after it, which GCC reports as an invalid suffix because it read
383        // the `0` as the constant. The distinction is not worth a worse message than this.
384        return Err(IntError::NoDigits);
385    }
386    if base == 8 && bytes[start..index].iter().any(|&byte| byte == b'8' || byte == b'9') {
387        return Err(IntError::InvalidOctalDigit);
388    }
389
390    let suffix = suffix_of(&bytes[index..])?;
391    if suffix.length == Some(Length::LongLong) && std == Std::C89 {
392        remarks = remarks.with(Remarks::LONG_LONG);
393    }
394    if suffix.length == Some(Length::BitInt) {
395        if std < Std::C23 {
396            remarks = remarks.with(Remarks::BIT_INT);
397        }
398        return Ok(IntConstant { value, ty: bit_int(value, suffix.unsigned), remarks });
399    }
400
401    let candidates = candidates(base, suffix, std);
402    let kind = candidates
403        .iter()
404        .copied()
405        .find(|&kind| fits(value, kind, target))
406        .ok_or(IntError::TooLarge)?;
407    if base == 10 && !suffix.unsigned && !signed_standard(kind) {
408        remarks = remarks.with(Remarks::UNSIGNED);
409    }
410    Ok(IntConstant { value, ty: IntConstantType::Standard(kind), remarks })
411}
412
413/// The base a spelling is written in, and where its digits start.
414///
415/// A leading `0` means octal only when a digit follows, so `0u` is a decimal zero with a
416/// suffix and `08` is an octal constant with a digit that does not exist. That is the split
417/// GCC makes, and it is what turns `08` into a message about octal rather than about a suffix.
418fn base_of(bytes: &[u8]) -> (u32, usize) {
419    match bytes {
420        [b'0', b'x' | b'X', ..] => (16, 2),
421        [b'0', b'b' | b'B', ..] => (2, 2),
422        [b'0', next, ..] if next.is_ascii_digit() => (8, 1),
423        _ => (10, 0),
424    }
425}
426
427/// Whether the spelling is a floating constant rather than an integer one.
428///
429/// A point anywhere, an `e` exponent in a decimal constant, or a `p` exponent in a hexadecimal
430/// one. `1e` and `1e+` are floating constants with no exponent digits, which is a diagnostic
431/// the floating path gives, and `1f` is an integer constant with a suffix that does not exist,
432/// which is one this path gives. Both compilers split them exactly there.
433///
434/// A leading zero does not survive an exponent: `08e5` is the floating constant eight hundred
435/// thousand and not an octal constant with a digit that does not exist.
436fn is_floating(bytes: &[u8], base: u32) -> bool {
437    let exponent = if base == 16 { *b"pP" } else { *b"eE" };
438    bytes.iter().any(|&byte| byte == b'.' || exponent.contains(&byte))
439}
440
441/// The value of a digit in the given base, and [`None`] when the byte is not one.
442///
443/// An octal constant reads `8` and `9` as digits, so that a constant holding one ends at the
444/// suffix and the error can name the digit rather than complain about the suffix.
445fn digit(byte: u8, base: u32) -> Option<u32> {
446    char::from(byte).to_digit(if base == 8 { 10 } else { base })
447}
448
449/// The length part of a suffix.
450#[derive(Debug, Clone, Copy, PartialEq, Eq)]
451enum Length {
452    /// `l` or `L`.
453    Long,
454    /// `ll` or `LL`.
455    LongLong,
456    /// `wb` or `WB`.
457    BitInt,
458}
459
460/// A parsed suffix.
461#[derive(Debug, Clone, Copy, PartialEq, Eq)]
462struct Suffix {
463    /// Whether `u` or `U` was there.
464    unsigned: bool,
465    /// The length part, when there was one.
466    length: Option<Length>,
467}
468
469/// Reads the suffix, which may hold each part once and in either order.
470fn suffix_of(mut rest: &[u8]) -> Result<Suffix, IntError> {
471    let mut suffix = Suffix { unsigned: false, length: None };
472    while let Some(&byte) = rest.first() {
473        let taken = match byte {
474            b'u' | b'U' if !suffix.unsigned => {
475                suffix.unsigned = true;
476                1
477            }
478            // The two letters have to agree about case, so `1ll` and `1LL` are constants and
479            // `1lL` is not. Both compilers refuse the mixed spelling in every dialect.
480            b'l' | b'L' if suffix.length.is_none() => {
481                if rest.get(1) == Some(&byte) {
482                    suffix.length = Some(Length::LongLong);
483                    2
484                } else {
485                    suffix.length = Some(Length::Long);
486                    1
487                }
488            }
489            b'w' | b'W' if suffix.length.is_none() => {
490                let second = if byte == b'w' { b'b' } else { b'B' };
491                if rest.get(1) != Some(&second) {
492                    return Err(IntError::InvalidSuffix);
493                }
494                suffix.length = Some(Length::BitInt);
495                2
496            }
497            _ => return Err(IntError::InvalidSuffix),
498        };
499        rest = &rest[taken..];
500    }
501    Ok(suffix)
502}
503
504/// The type of a `wb` constant, which is the narrowest one that holds the value.
505///
506/// The sign bit counts, so a signed one is never narrower than two bits: `1wb` is
507/// `_BitInt(2)`. An unsigned zero is `unsigned _BitInt(1)`, because a width of zero is not a
508/// type. Measured against clang.
509fn bit_int(value: u128, unsigned: bool) -> IntConstantType {
510    let used = 128 - value.leading_zeros();
511    let width = if unsigned { used.max(1) } else { used + 1 };
512    IntConstantType::BitInt { signed: !unsigned, width: width.max(if unsigned { 1 } else { 2 }) }
513}
514
515/// Whether `kind` is one of the standard signed types, which is what decides the remark about
516/// a decimal constant having gone unsigned.
517fn signed_standard(kind: IntKind) -> bool {
518    matches!(kind, IntKind::Int | IntKind::Long | IntKind::LongLong)
519}
520
521/// Whether the value fits in `kind` on this target.
522fn fits(value: u128, kind: IntKind, target: &TargetInfo) -> bool {
523    let width = int_width(kind, target);
524    // Signedness here never depends on what plain `char` is, because no candidate list holds a
525    // character type.
526    let bits = if kind.is_signed(false) { width - 1 } else { width };
527    // `unsigned __int128` holds every value the accumulator can, and shifting a `u128` by all
528    // of its bits is not a shift, so the widest type is answered without one.
529    bits >= 128 || value >> bits == 0
530}
531
532/// The candidate list for a base and a suffix, in the order the standard walks it.
533///
534/// `__int128` and `unsigned __int128` are on the end of every list, which is what gcc does:
535/// `9223372036854775808` is an `__int128` there and an `unsigned long long` in clang. Both
536/// compilers put `long long` out of reach in C89 unless the suffix asks for it, and C89 is
537/// also the dialect that offers `unsigned long` for a decimal constant with no suffix at all.
538fn candidates(base: u32, suffix: Suffix, std: Std) -> &'static [IntKind] {
539    use IntKind::{Int, Int128, Long, LongLong, UInt, UInt128, ULong, ULongLong};
540
541    let decimal = base == 10;
542    let c89 = std == Std::C89;
543    match (suffix.unsigned, suffix.length) {
544        (false, None) if decimal && c89 => &[Int, Long, ULong, Int128, UInt128],
545        (false, None) if decimal => &[Int, Long, LongLong, Int128],
546        (false, None) if c89 => &[Int, UInt, Long, ULong, Int128, UInt128],
547        (false, None) => &[Int, UInt, Long, ULong, LongLong, ULongLong, Int128, UInt128],
548
549        (true, None) if c89 => &[UInt, ULong, UInt128],
550        (true, None) => &[UInt, ULong, ULongLong, UInt128],
551
552        (false, Some(Length::Long)) if decimal && c89 => &[Long, ULong, Int128, UInt128],
553        (false, Some(Length::Long)) if decimal => &[Long, LongLong, Int128],
554        (false, Some(Length::Long)) if c89 => &[Long, ULong, Int128, UInt128],
555        (false, Some(Length::Long)) => &[Long, ULong, LongLong, ULongLong, Int128, UInt128],
556
557        (true, Some(Length::Long)) if c89 => &[ULong, UInt128],
558        (true, Some(Length::Long)) => &[ULong, ULongLong, UInt128],
559
560        (false, Some(Length::LongLong)) if decimal => &[LongLong, Int128],
561        (false, Some(Length::LongLong)) => &[LongLong, ULongLong, Int128, UInt128],
562        (true, Some(Length::LongLong)) => &[ULongLong, UInt128],
563
564        // A `wb` constant never reaches here: its type comes from the value alone.
565        (_, Some(Length::BitInt)) => &[],
566    }
567}
568
569/// Converts the spelling of a preprocessing number into a floating constant.
570///
571/// # Errors
572///
573/// [`FloatError`], one case of which is that the spelling is an integer constant rather than a
574/// malformed floating one.
575pub fn floating(text: &str, std: Std, target: &TargetInfo) -> Result<FloatConstant, FloatError> {
576    let bytes = text.as_bytes();
577    let (base, _) = base_of(bytes);
578    if !is_floating(bytes, base) {
579        return Err(FloatError::Integer);
580    }
581    // A leading zero means nothing to a floating constant, so there are two bases here and not
582    // four: `08e5` is eight hundred thousand rather than an octal constant with a bad digit.
583    let hex = base == 16;
584    let base = if hex { 16 } else { 10 };
585    let mut remarks = Remarks::NONE;
586    if hex && std < Std::C99 {
587        remarks = remarks.with(Remarks::HEX_FLOAT);
588    }
589
590    let mut index = if hex { 2 } else { 0 };
591    let mut digits = 0;
592    let mut point = false;
593    let mut separators = false;
594    while index < bytes.len() {
595        let byte = bytes[index];
596        if byte == b'\'' {
597            if digits == 0 || !next_is_digit(bytes, index, base) {
598                return Err(FloatError::InvalidSuffix);
599            }
600            separators = true;
601        } else if byte == b'.' {
602            if point {
603                return Err(FloatError::TooManyPoints);
604            }
605            point = true;
606        } else if digit(byte, base).is_some() {
607            digits += 1;
608        } else {
609            break;
610        }
611        index += 1;
612    }
613    if digits == 0 {
614        return Err(FloatError::NoDigits);
615    }
616
617    let marker = if hex { *b"pP" } else { *b"eE" };
618    if index < bytes.len() && marker.contains(&bytes[index]) {
619        index += 1;
620        if matches!(bytes.get(index), Some(b'+' | b'-')) {
621            index += 1;
622        }
623        let mut exponent_digits = 0;
624        while index < bytes.len() {
625            let byte = bytes[index];
626            if byte == b'\'' {
627                if exponent_digits == 0 || !next_is_digit(bytes, index, 10) {
628                    return Err(FloatError::InvalidSuffix);
629                }
630                separators = true;
631            } else if byte.is_ascii_digit() {
632                exponent_digits += 1;
633            } else {
634                break;
635            }
636            index += 1;
637        }
638        if exponent_digits == 0 {
639            return Err(FloatError::NoExponentDigits);
640        }
641    } else if hex {
642        // The exponent is not optional in a hexadecimal constant, because `f` is a digit there
643        // and `0x1.8f` would otherwise be a number and a suffix at the same time.
644        return Err(FloatError::MissingExponent);
645    }
646    if separators && std < Std::C23 {
647        remarks = remarks.with(Remarks::SEPARATORS);
648    }
649
650    let suffix = float_suffix(&bytes[index..], target)?;
651    remarks = remarks.with(suffix.remarks);
652    let (value, status) =
653        Float::parse(&text[..index], suffix.ty.format(target)).map_err(|error| match error {
654            // The scan above has already ruled all three of these out, and mapping them is
655            // still better than an unwrap that a later change could reach.
656            ParseError::NoDigits => FloatError::NoDigits,
657            ParseError::NoExponentDigits => FloatError::NoExponentDigits,
658            ParseError::Invalid => FloatError::InvalidSuffix,
659        })?;
660    if status.has(Status::OVERFLOW) {
661        remarks = remarks.with(Remarks::OUT_OF_RANGE);
662    }
663    // Underflow on its own is a subnormal, which is a number the program can use. Losing the
664    // value entirely is the part worth a word.
665    if status.has(Status::UNDERFLOW) && value.is_zero() {
666        remarks = remarks.with(Remarks::TRUNCATED);
667    }
668    Ok(FloatConstant { value, ty: suffix.ty, imaginary: suffix.imaginary, remarks })
669}
670
671/// Whether the byte after `index` is a digit in `base`, which is what makes a separator one.
672fn next_is_digit(bytes: &[u8], index: usize, base: u32) -> bool {
673    bytes.get(index + 1).is_some_and(|&next| digit(next, base).is_some())
674}
675
676/// A parsed floating suffix.
677struct FloatSuffix {
678    /// The type it named, which is `double` when it named none.
679    ty: FloatConstantType,
680    /// Whether it held an `i` or a `j`.
681    imaginary: bool,
682    /// What the suffix alone is worth saying about.
683    remarks: Remarks,
684}
685
686/// Reads the suffix, which may name a type once and mark the constant imaginary once, in either
687/// order.
688///
689/// Everything past `f` and `l` is an extension, and the extensions are where the case rules stop
690/// being uniform: the `f` of `_FloatN` may be either case and the `x` of `_FloatNx` may not, and
691/// the two letters of a decimal suffix have to agree. All of it measured on gcc 13.3.
692fn float_suffix(mut rest: &[u8], target: &TargetInfo) -> Result<FloatSuffix, FloatError> {
693    let mut ty = None;
694    let mut imaginary = false;
695    let mut remarks = Remarks::NONE;
696    while let Some(&byte) = rest.first() {
697        let taken = match byte {
698            b'i' | b'j' | b'I' | b'J' if !imaginary => {
699                imaginary = true;
700                remarks = remarks.with(Remarks::IMAGINARY);
701                1
702            }
703            // One type per constant, so `1.0fl` is not a constant and neither is `1.0ff`.
704            _ if ty.is_some() => return Err(FloatError::InvalidSuffix),
705            b'f' | b'F' => {
706                let (named, taken, extra) = float_n(rest)?;
707                ty = Some(named);
708                remarks = remarks.with(extra);
709                taken
710            }
711            b'l' | b'L' => {
712                ty = Some(FloatConstantType::LongDouble);
713                1
714            }
715            b'q' | b'Q' => {
716                ty = Some(FloatConstantType::Float128);
717                remarks = remarks.with(Remarks::EXTENDED_SUFFIX);
718                1
719            }
720            b'w' | b'W' => {
721                // `__float80` is the x87 format, which only x86 has.
722                if target.triple.arch != Arch::X86_64 {
723                    return Err(FloatError::UnsupportedType);
724                }
725                ty = Some(FloatConstantType::Float80);
726                remarks = remarks.with(Remarks::EXTENDED_SUFFIX);
727                1
728            }
729            b'd' | b'D' => {
730                let second = rest.get(1).copied();
731                let decimal = if byte == b'd' {
732                    matches!(second, Some(b'f' | b'd' | b'l'))
733                } else {
734                    matches!(second, Some(b'F' | b'D' | b'L'))
735                };
736                if decimal {
737                    return Err(FloatError::DecimalFloat);
738                }
739                ty = Some(FloatConstantType::Double);
740                remarks = remarks.with(Remarks::DOUBLE_SUFFIX);
741                1
742            }
743            _ => return Err(FloatError::InvalidSuffix),
744        };
745        rest = &rest[taken..];
746    }
747    Ok(FloatSuffix { ty: ty.unwrap_or(FloatConstantType::Double), imaginary, remarks })
748}
749
750/// Reads a suffix that starts with `f`, which is `float` on its own and one of the `_FloatN` or
751/// `_FloatNx` types when digits follow.
752///
753/// Returns the type, how many bytes it took and what is worth saying about it.
754fn float_n(rest: &[u8]) -> Result<(FloatConstantType, usize, Remarks), FloatError> {
755    let mut end = 1;
756    while rest.get(end).is_some_and(u8::is_ascii_digit) {
757        end += 1;
758    }
759    if end == 1 {
760        return Ok((FloatConstantType::Float, 1, Remarks::NONE));
761    }
762    // The `x` is lower case in every spelling gcc accepts, so `F64x` is a constant and `f64X`
763    // is not, however odd that looks next to the `F` being free.
764    let extended = rest.get(end) == Some(&b'x');
765    let ty = match (&rest[1..end], extended) {
766        (b"16", false) => FloatConstantType::Float16,
767        (b"32", false) => FloatConstantType::Float32,
768        (b"64", false) => FloatConstantType::Float64,
769        (b"128", false) => FloatConstantType::Float128,
770        (b"32", true) => FloatConstantType::Float32x,
771        (b"64", true) => FloatConstantType::Float64x,
772        // gcc knows the name `_Float128x` and has the type on no target here, and says so
773        // rather than calling the suffix invalid. `_Float16x` is not a type at all.
774        (b"128", true) => return Err(FloatError::UnsupportedType),
775        _ => return Err(FloatError::InvalidSuffix),
776    };
777    Ok((ty, end + usize::from(extended), Remarks::EXTENDED_SUFFIX))
778}
779
780#[cfg(test)]
781mod tests {
782    use rucc_target::Triple;
783
784    use super::*;
785
786    fn linux() -> TargetInfo {
787        TargetInfo::new("x86_64-unknown-linux-gnu".parse::<Triple>().expect("a known triple"))
788    }
789
790    fn aarch64() -> TargetInfo {
791        TargetInfo::new("aarch64-unknown-linux-gnu".parse::<Triple>().expect("a known triple"))
792    }
793
794    /// The value and the type of a constant in the default dialect.
795    fn c23(text: &str) -> Result<IntConstant, IntError> {
796        integer(text, Std::C23, &linux())
797    }
798
799    /// The type of a constant in the given dialect, on x86-64 Linux.
800    fn kind(text: &str, std: Std) -> IntKind {
801        match integer(text, std, &linux()).expect("a valid constant").ty {
802            IntConstantType::Standard(kind) => kind,
803            IntConstantType::BitInt { .. } => panic!("{text} is a _BitInt constant"),
804        }
805    }
806
807    #[test]
808    fn a_constant_in_each_base_has_the_value_it_says() {
809        assert_eq!(c23("0").expect("zero").value, 0);
810        assert_eq!(c23("42").expect("decimal").value, 42);
811        assert_eq!(c23("0777").expect("octal").value, 0o777);
812        assert_eq!(c23("0xdeadBEEF").expect("hex").value, 0xdead_beef);
813        assert_eq!(c23("0b1010").expect("binary").value, 0b1010);
814        assert_eq!(c23("0X10").expect("upper case prefix").value, 16);
815        // A leading zero with nothing after it is a decimal zero rather than an octal one with
816        // no digits, which is the split that lets `0u` through and stops `08`.
817        assert_eq!(c23("0u").expect("zero with a suffix").value, 0);
818    }
819
820    #[test]
821    fn digit_separators_are_stripped_and_reported_before_c23() {
822        let value = c23("1'000'000").expect("a C23 constant");
823        assert_eq!(value.value, 1_000_000);
824        assert!(value.remarks.is_none());
825        assert_eq!(c23("0x1'0").expect("hex with a separator").value, 16);
826
827        let older = integer("1'000", Std::C17, &linux()).expect("still converted");
828        assert!(older.remarks.has(Remarks::SEPARATORS));
829        assert_eq!(older.value, 1000);
830    }
831
832    #[test]
833    fn the_type_of_a_decimal_constant_walks_the_signed_types_only() {
834        // Measured with `_Generic` on gcc 13.3, x86-64 Linux.
835        assert_eq!(kind("2147483647", Std::C23), IntKind::Int);
836        assert_eq!(kind("2147483648", Std::C23), IntKind::Long);
837        assert_eq!(kind("4294967295", Std::C23), IntKind::Long);
838        assert_eq!(kind("9223372036854775807", Std::C23), IntKind::Long);
839        // Past `long long` gcc reaches for `__int128` rather than for an unsigned type, and
840        // says so: the constant is so large that it is unsigned.
841        assert_eq!(kind("9223372036854775808", Std::C23), IntKind::Int128);
842        assert_eq!(kind("18446744073709551615", Std::C23), IntKind::Int128);
843        let large = c23("18446744073709551615").expect("fits __int128");
844        assert!(large.remarks.has(Remarks::UNSIGNED));
845    }
846
847    #[test]
848    fn a_constant_in_another_base_may_be_unsigned_without_saying_so() {
849        // This is the split that surprises people: `4294967295` is a `long` and `0xffffffff`
850        // is an `unsigned int`, because only the decimal list is signed types alone.
851        assert_eq!(kind("0xffffffff", Std::C23), IntKind::UInt);
852        assert_eq!(kind("0x7fffffff", Std::C23), IntKind::Int);
853        assert_eq!(kind("0x80000000", Std::C23), IntKind::UInt);
854        assert_eq!(kind("0x100000000", Std::C23), IntKind::Long);
855        assert_eq!(kind("0xffffffffffffffff", Std::C23), IntKind::ULong);
856        assert_eq!(kind("0777", Std::C23), IntKind::Int);
857        assert_eq!(kind("0b1010", Std::C23), IntKind::Int);
858        // And no remark, because nothing about it is surprising enough to say.
859        assert!(c23("0xffffffff").expect("a constant").remarks.is_none());
860    }
861
862    #[test]
863    fn c89_has_unsigned_long_in_the_decimal_list_and_no_long_long_in_any() {
864        // gcc under `-std=c89 -pedantic`: "this decimal constant is unsigned only in ISO C90",
865        // and eight bytes rather than sixteen.
866        assert_eq!(kind("18446744073709551615", Std::C89), IntKind::ULong);
867        assert_eq!(kind("18446744073709551615", Std::C99), IntKind::Int128);
868        let old = integer("18446744073709551615", Std::C89, &linux()).expect("a C89 constant");
869        assert!(old.remarks.has(Remarks::UNSIGNED));
870        // The suffix still reaches `long long`, with the remark gcc prints for it.
871        let long_long = integer("1ll", Std::C89, &linux()).expect("an extension");
872        assert!(long_long.remarks.has(Remarks::LONG_LONG));
873        assert_eq!(kind("1ll", Std::C89), IntKind::LongLong);
874        assert!(integer("1ll", Std::C99, &linux()).expect("standard").remarks.is_none());
875    }
876
877    #[test]
878    fn a_suffix_narrows_the_list_it_does_not_pick_the_type() {
879        assert_eq!(kind("1u", Std::C23), IntKind::UInt);
880        assert_eq!(kind("1l", Std::C23), IntKind::Long);
881        assert_eq!(kind("1ul", Std::C23), IntKind::ULong);
882        assert_eq!(kind("1ll", Std::C23), IntKind::LongLong);
883        assert_eq!(kind("1llu", Std::C23), IntKind::ULongLong);
884        // The suffix is a floor rather than an answer: `4294967296u` is an `unsigned long`
885        // because `unsigned int` cannot hold it.
886        assert_eq!(kind("4294967296u", Std::C23), IntKind::ULong);
887        assert_eq!(kind("0xffffffffu", Std::C23), IntKind::UInt);
888    }
889
890    #[test]
891    fn the_letters_of_a_suffix_may_be_in_either_case_but_not_both() {
892        for text in ["1u", "1U", "1l", "1L", "1ll", "1LL", "1ul", "1lu", "1uL", "1LLU", "1llu"] {
893            assert!(c23(text).is_ok(), "{text} is a constant in both compilers");
894        }
895        for text in ["1lL", "1Ll", "1uu", "1lul", "1z", "1uz", "1f", "1x", "1_000"] {
896            assert_eq!(c23(text), Err(IntError::InvalidSuffix), "{text} is not");
897        }
898    }
899
900    #[test]
901    fn a_bit_int_constant_has_the_narrowest_type_that_holds_it() {
902        // Measured against clang, which is the only one of the two that has the type.
903        let cases = [
904            ("0wb", true, 2),
905            ("1wb", true, 2),
906            ("3wb", true, 3),
907            ("42wb", true, 7),
908            ("255wb", true, 9),
909            ("0uwb", false, 1),
910            ("1uwb", false, 1),
911            ("255uwb", false, 8),
912            ("256uwb", false, 9),
913            ("0xffffffffffffffffuwb", false, 64),
914        ];
915        for (text, signed, width) in cases {
916            let constant = c23(text).expect("a _BitInt constant");
917            assert_eq!(
918                constant.ty,
919                IntConstantType::BitInt { signed, width },
920                "{text} is the wrong width"
921            );
922        }
923        // Either order, either case, and never with a length suffix.
924        for text in ["1uwb", "1wbu", "1UWB", "1WBu", "1uWB"] {
925            assert!(c23(text).is_ok(), "{text} is a constant in clang");
926        }
927        for text in ["1wB", "1Wb", "1lwb", "1wbl", "1wbwb"] {
928            assert_eq!(c23(text), Err(IntError::InvalidSuffix), "{text} is not");
929        }
930        // Before C23 it is still converted, and still worth a word.
931        let older = integer("1wb", Std::C17, &linux()).expect("clang accepts it everywhere");
932        assert!(older.remarks.has(Remarks::BIT_INT));
933    }
934
935    #[test]
936    fn a_binary_constant_is_an_extension_before_c23() {
937        assert!(c23("0b1").expect("standard in C23").remarks.is_none());
938        let older = integer("0b1", Std::C17, &linux()).expect("both compilers accept it");
939        assert!(older.remarks.has(Remarks::BINARY));
940    }
941
942    #[test]
943    fn an_octal_constant_names_the_digit_that_is_not_one() {
944        assert_eq!(c23("08"), Err(IntError::InvalidOctalDigit));
945        assert_eq!(c23("0778"), Err(IntError::InvalidOctalDigit));
946        assert_eq!(c23("09"), Err(IntError::InvalidOctalDigit));
947        // A `9` elsewhere is fine, and the message is only for constants that began with `0`.
948        assert_eq!(c23("9").expect("decimal").value, 9);
949    }
950
951    #[test]
952    fn a_prefix_with_no_digits_after_it_is_not_a_constant() {
953        assert_eq!(c23("0x"), Err(IntError::NoDigits));
954        assert_eq!(c23("0b"), Err(IntError::NoDigits));
955    }
956
957    #[test]
958    fn a_constant_larger_than_any_type_is_refused_rather_than_wrapped() {
959        // gcc accumulates in sixty four bits and silently gives this the value zero and the
960        // type `int` after a warning. That is the one measured behaviour here we refuse to
961        // reproduce, and clang refuses it too.
962        assert_eq!(c23("340282366920938463463374607431768211456"), Err(IntError::TooLarge));
963        assert_eq!(c23("0x100000000000000000000000000000000"), Err(IntError::TooLarge));
964        // 2^127 fits in the accumulator and in no signed type, and the decimal list has no
965        // unsigned one to fall back to.
966        assert_eq!(c23("170141183460469231731687303715884105728"), Err(IntError::TooLarge));
967        // The same value written in hex reaches `unsigned __int128`, because that list has it.
968        assert_eq!(kind("0x80000000000000000000000000000000", Std::C23), IntKind::UInt128);
969        assert_eq!(kind("0xffffffffffffffffffffffffffffffff", Std::C23), IntKind::UInt128);
970    }
971
972    #[test]
973    fn a_floating_constant_is_handed_back_rather_than_refused() {
974        for text in ["1.0", ".5", "1.", "1e5", "1E-5", "1e", "0x1p3", "0x1.8p+1", "1.5e3"] {
975            assert_eq!(c23(text), Err(IntError::Floating), "{text} belongs to the other path");
976        }
977        // A leading zero does not make this an octal constant with a digit that does not
978        // exist. gcc compiles it, as eight hundred thousand.
979        assert_eq!(c23("08e5"), Err(IntError::Floating));
980        // A hexadecimal `e` is a digit, not an exponent, and `1f` is an integer with a suffix
981        // that does not exist rather than a float. Both compilers split them there.
982        assert_eq!(c23("0xe5").expect("hex digits").value, 0xe5);
983        assert_eq!(c23("1f"), Err(IntError::InvalidSuffix));
984    }
985
986    #[test]
987    fn the_type_comes_from_the_target_and_not_from_the_host() {
988        // `4294967295` is a `long` where `long` is sixty four bits and a `long long` where it
989        // is thirty two. A compiler that asked its own platform gets one of these wrong.
990        let windows =
991            TargetInfo::new("x86_64-pc-windows-msvc".parse::<Triple>().expect("a known triple"));
992        let on_windows = integer("4294967295", Std::C23, &windows).expect("a constant");
993        assert_eq!(on_windows.ty, IntConstantType::Standard(IntKind::LongLong));
994        assert_eq!(kind("4294967295", Std::C23), IntKind::Long);
995    }
996
997    /// A floating constant in the default dialect, on x86-64 Linux.
998    fn float(text: &str) -> Result<FloatConstant, FloatError> {
999        floating(text, Std::C23, &linux())
1000    }
1001
1002    /// The bits of a constant's value, which is the form every measured row here was taken in.
1003    fn bits(text: &str) -> u128 {
1004        float(text).expect("a valid constant").value.to_bits()
1005    }
1006
1007    #[test]
1008    fn a_constant_with_no_suffix_is_a_double() {
1009        let constant = float("1.5").expect("a constant");
1010        assert_eq!(constant.ty, FloatConstantType::Double);
1011        assert!(!constant.imaginary);
1012        assert!(constant.remarks.is_none());
1013        assert_eq!(constant.value.to_bits(), 0x3ff8_0000_0000_0000);
1014        assert_eq!(bits("0.1"), 0x3fb9_9999_9999_999a);
1015        assert_eq!(bits(".5"), 0x3fe0_0000_0000_0000);
1016        assert_eq!(bits("1."), 0x3ff0_0000_0000_0000);
1017        assert_eq!(bits("1e5"), 0x40f8_6a00_0000_0000);
1018        assert_eq!(bits("0x1p3"), 0x4020_0000_0000_0000);
1019        // A leading zero is not an octal prefix once there is an exponent, so this is eight
1020        // hundred thousand and gcc compiles it as one.
1021        assert_eq!(bits("08e5"), 0x4128_6a00_0000_0000);
1022    }
1023
1024    #[test]
1025    fn the_suffix_names_the_type_rather_than_narrowing_a_list() {
1026        // Measured with `_Generic` on gcc 13.3, x86-64 Linux.
1027        let cases = [
1028            ("1.0", FloatConstantType::Double),
1029            ("1.0f", FloatConstantType::Float),
1030            ("1.0F", FloatConstantType::Float),
1031            ("1.0l", FloatConstantType::LongDouble),
1032            ("1.0L", FloatConstantType::LongDouble),
1033            ("1.0d", FloatConstantType::Double),
1034            ("1.0q", FloatConstantType::Float128),
1035            ("1.0w", FloatConstantType::Float80),
1036            ("1.0f16", FloatConstantType::Float16),
1037            ("1.0F16", FloatConstantType::Float16),
1038            ("1.0f32", FloatConstantType::Float32),
1039            ("1.0f64", FloatConstantType::Float64),
1040            ("1.0f128", FloatConstantType::Float128),
1041            ("1.0f32x", FloatConstantType::Float32x),
1042            ("1.0F64x", FloatConstantType::Float64x),
1043        ];
1044        for (text, ty) in cases {
1045            assert_eq!(float(text).expect("a constant").ty, ty, "{text} has the wrong type");
1046        }
1047    }
1048
1049    #[test]
1050    fn each_type_is_converted_in_the_format_the_target_has_for_it() {
1051        // Every row measured by printing the bytes of the constant on gcc 13.3, x86-64 Linux.
1052        // The two that surprise are `_Float32x`, which is plain `double`, and `_Float64x`,
1053        // which is the x87 format and so the same bits as `long double` and `__float80`.
1054        assert_eq!(bits("0.1f"), 0x3dcc_cccd);
1055        assert_eq!(bits("0.1f16"), 0x2e66);
1056        assert_eq!(bits("0.1f32x"), 0x3fb9_9999_9999_999a);
1057        assert_eq!(bits("0.1f64x"), 0x3ffb_cccc_cccc_cccc_cccd);
1058        assert_eq!(bits("0.1w"), 0x3ffb_cccc_cccc_cccc_cccd);
1059        assert_eq!(bits("0.1l"), 0x3ffb_cccc_cccc_cccc_cccd);
1060        assert_eq!(bits("0.1q"), 0x3ffb_9999_9999_9999_9999_9999_9999_999a);
1061        assert_eq!(bits("0.1f128"), 0x3ffb_9999_9999_9999_9999_9999_9999_999a);
1062        assert_eq!(bits("1.0l"), 0x3fff_8000_0000_0000_0000);
1063    }
1064
1065    #[test]
1066    fn the_format_comes_from_the_target_and_not_from_the_host() {
1067        // `long double` is 128 bits wide on both of these and it is not the same type on both,
1068        // which is the whole reason the target carries a format and not only a width.
1069        let arm = floating("1.0l", Std::C23, &aarch64()).expect("a constant");
1070        assert_eq!(arm.value.to_bits(), 0x3fff_0000_0000_0000_0000_0000_0000_0000);
1071        assert_eq!(bits("1.0l"), 0x3fff_8000_0000_0000_0000);
1072        // And `_Float64x` follows it, being whatever the target has above `_Float64`.
1073        let arm_wide = floating("0.1f64x", Std::C23, &aarch64()).expect("a constant");
1074        assert_eq!(arm_wide.value.to_bits(), 0x3ffb_9999_9999_9999_9999_9999_9999_999a);
1075        // Where `long double` is `double` the constant is a `double` too.
1076        let windows =
1077            TargetInfo::new("x86_64-pc-windows-msvc".parse::<Triple>().expect("a known triple"));
1078        let on_windows = floating("1.0l", Std::C23, &windows).expect("a constant");
1079        assert_eq!(on_windows.value.to_bits(), 0x3ff0_0000_0000_0000);
1080    }
1081
1082    #[test]
1083    fn the_case_rules_of_a_floating_suffix_are_not_uniform() {
1084        // The `f` of a `_FloatN` suffix is free and the `x` of a `_FloatNx` one is not, and the
1085        // two letters of a decimal suffix have to agree. Measured on gcc 13.3, which accepts
1086        // every one of these in every dialect.
1087        for text in ["1.0f", "1.0F", "1.0L", "1.0Q", "1.0W", "1.0F32", "1.0f64x", "1.0F64x"] {
1088            assert!(float(text).is_ok(), "{text} is a constant in gcc");
1089        }
1090        for text in ["1.0F32X", "1.0f32X", "1.0f16x", "1.0ff", "1.0fl", "1.0lf", "1.0fF", "1.0LL"] {
1091            assert_eq!(float(text), Err(FloatError::InvalidSuffix), "{text} is not");
1092        }
1093    }
1094
1095    #[test]
1096    fn an_imaginary_suffix_may_sit_on_either_side_of_the_type() {
1097        for text in ["1.0i", "1.0j", "1.0I", "1.0J", "1.0if", "1.0fi", "1.0Li", "1.0iL", "1.0f16i"]
1098        {
1099            let constant = float(text).expect("a constant in gcc");
1100            assert!(constant.imaginary, "{text} is imaginary");
1101            assert!(constant.remarks.has(Remarks::IMAGINARY));
1102        }
1103        assert_eq!(float("1.0ii"), Err(FloatError::InvalidSuffix));
1104        assert_eq!(float("1.0ij"), Err(FloatError::InvalidSuffix));
1105        assert!(!float("1.0f").expect("a constant").imaginary);
1106    }
1107
1108    #[test]
1109    fn a_decimal_floating_constant_is_recognised_and_refused() {
1110        // The constant is well formed and there is nowhere in this compiler to put its value.
1111        for text in ["1.0df", "1.0dd", "1.0dl", "1.0DF", "1.0DD", "1.0DL"] {
1112            assert_eq!(float(text), Err(FloatError::DecimalFloat), "{text} is a decimal float");
1113        }
1114        // The letters have to agree about case, so these are not decimal floats and not
1115        // constants either.
1116        for text in ["1.0Df", "1.0dF", "1.0dD", "1.0Dl"] {
1117            assert_eq!(float(text), Err(FloatError::InvalidSuffix), "{text} is neither");
1118        }
1119        // A `d` on its own is a `double` written the long way, which gcc allows everywhere.
1120        let long_way = float("1.0d").expect("a GCC extension");
1121        assert_eq!(long_way.ty, FloatConstantType::Double);
1122        assert!(long_way.remarks.has(Remarks::DOUBLE_SUFFIX));
1123    }
1124
1125    #[test]
1126    fn a_type_the_target_does_not_have_is_refused_by_name() {
1127        // gcc says "'_Float128x' is not supported on this target" rather than calling the
1128        // suffix invalid, and it is supported on no target here.
1129        assert_eq!(float("1.0f128x"), Err(FloatError::UnsupportedType));
1130        // `__float80` is the x87 format, which only x86 has.
1131        assert_eq!(floating("1.0w", Std::C23, &aarch64()), Err(FloatError::UnsupportedType));
1132        assert!(float("1.0w").is_ok());
1133    }
1134
1135    #[test]
1136    fn a_hexadecimal_constant_needs_an_exponent_and_a_decimal_one_does_not() {
1137        // `f` is a hexadecimal digit, so without the exponent there is no telling the number
1138        // from the suffix. Both compilers require it.
1139        assert_eq!(float("0x1.8"), Err(FloatError::MissingExponent));
1140        assert_eq!(bits("0x1.8p0"), 0x3ff8_0000_0000_0000);
1141        assert_eq!(bits("0x.8p1"), 0x3ff0_0000_0000_0000);
1142        assert_eq!(bits("1.5"), 0x3ff8_0000_0000_0000);
1143        for text in ["1.0e", "1e+", "1e-", "0x1p", "0x1p+"] {
1144            assert_eq!(float(text), Err(FloatError::NoExponentDigits), "{text} has no exponent");
1145        }
1146        assert_eq!(float("1.2.3"), Err(FloatError::TooManyPoints));
1147    }
1148
1149    #[test]
1150    fn an_integer_constant_is_handed_back_rather_than_refused() {
1151        for text in ["1", "0", "0x10", "1u", "0777", "1wb", "0b1", "0xe5", "1f"] {
1152            assert_eq!(float(text), Err(FloatError::Integer), "{text} belongs to the other path");
1153        }
1154    }
1155
1156    #[test]
1157    fn a_value_past_the_range_of_its_type_is_still_a_constant() {
1158        let large = float("1e400").expect("a constant gcc compiles");
1159        assert!(large.value.is_infinite());
1160        assert!(large.remarks.has(Remarks::OUT_OF_RANGE));
1161        let small = float("1e-400").expect("a constant gcc compiles");
1162        assert!(small.value.is_zero());
1163        assert!(small.remarks.has(Remarks::TRUNCATED));
1164        // The same two in the format the suffix asked for rather than in `double`.
1165        assert!(float("1e39f").expect("a constant").remarks.has(Remarks::OUT_OF_RANGE));
1166        assert!(float("1e-46f").expect("a constant").remarks.has(Remarks::TRUNCATED));
1167        assert!(float("1e-4951l").expect("a constant").remarks.has(Remarks::TRUNCATED));
1168        // A subnormal is a number the program can use, and gcc says nothing about it.
1169        let subnormal = float("1e-320").expect("a constant");
1170        assert!(!subnormal.value.is_zero());
1171        assert!(subnormal.remarks.is_none());
1172    }
1173
1174    #[test]
1175    fn the_dialect_decides_what_a_constant_is_worth_saying_about() {
1176        // "use of C99 hexadecimal floating constant", which gcc says under C89 and not after.
1177        let old = floating("0x1p3", Std::C89, &linux()).expect("gcc compiles it anyway");
1178        assert!(old.remarks.has(Remarks::HEX_FLOAT));
1179        assert!(floating("0x1p3", Std::C99, &linux()).expect("standard").remarks.is_none());
1180        // Separators are C23 in both compilers, in the number and in the exponent.
1181        assert_eq!(bits("1'0.5"), 0x4025_0000_0000_0000);
1182        assert_eq!(bits("1.0e1'0"), 0x4202_a05f_2000_0000);
1183        assert!(float("0x1'0p0").expect("a C23 constant").remarks.is_none());
1184        let older = floating("1'0.5", Std::C17, &linux()).expect("still converted");
1185        assert!(older.remarks.has(Remarks::SEPARATORS));
1186        // And every extension suffix is accepted in every dialect, with a word about it.
1187        for text in ["1.0q", "1.0w", "1.0f16", "1.0f32x"] {
1188            let constant = floating(text, Std::C89, &linux()).expect("gcc accepts it in C89");
1189            assert!(constant.remarks.has(Remarks::EXTENDED_SUFFIX), "{text} is not standard");
1190        }
1191        assert!(float("1.0f").expect("a constant").remarks.is_none());
1192    }
1193}