Skip to main content

rucc_lex/
number.rs

1//! Numeric constants: the value, and the type the standard's table walk gives it.
2//!
3//! Design: `spec/06-lexer-and-parser.md` section 6.1.
4//!
5//! This is the second piece of phase 7. A preprocessing number is a loose thing, deliberately
6//! looser than a constant, so `1.2.3` and `0x1p+3` are both one pp-token and only here does
7//! anyone ask what they mean. What comes back is a value and a type, and both of them are
8//! places a compiler quietly goes wrong.
9//!
10//! There are two entry points, [`integer`] and [`floating`], and either of them hands the
11//! spelling to the other rather than reporting an error when it turns out to belong there. The
12//! split is not where a reader expects it: `1e` is a floating constant with no exponent digits
13//! and `1f` is an integer constant with a suffix that does not exist, and both compilers agree
14//! on that, because the exponent marker is part of a preprocessing number and the suffix letter
15//! is not part of anything.
16//!
17//! The value is accumulated in a `u128` with every step checked, so a constant too large to
18//! represent is a diagnostic rather than a number the program did not write. gcc 13.3 does not
19//! do that: its accumulator is sixty four bits, and `18446744073709551616` compiles to zero of
20//! type `int` after a warning nobody reads. That is not a behaviour worth reproducing, so ours
21//! is the only measured difference here that is deliberate: past a hundred and twenty eight
22//! bits the constant is refused. clang refuses it too, one bit earlier.
23//!
24//! The type is the standard's table walk, 6.4.4.1p5: a candidate list chosen by the base and
25//! the suffix, walked in order, and the first type that holds the value wins. The list is not
26//! the same in every dialect. C89 puts `unsigned long` in the list for a decimal constant with
27//! no suffix, which is what makes `18446744073709551615` an `unsigned long` under `-std=c89`
28//! and something wider under `-std=c99`, and gcc says so in as many words: "this decimal
29//! constant is unsigned only in ISO C90". Both compilers keep `long long` out of the C89 lists
30//! and accept it when the suffix asks for it.
31//!
32//! `__int128` is on the end of every list, which is what gcc does and clang does not.
33//! `9223372036854775808` is an `__int128` in gcc 13.3 and an `unsigned long long` in clang,
34//! and the difference is visible to a program: negate it and gcc gives a negative number.
35//! We follow gcc, because the alternative silently turns a signed constant unsigned.
36//!
37//! The rest was measured the same way, by writing the constant and asking `_Generic` what it
38//! is, on gcc 13.3 on x86-64 Linux and on clang:
39//!
40//! The suffix letters may be in either case but not both, so `1ll` and `1LL` are constants and
41//! `1lL` is not, and the same rule holds for `wb`. The unsigned suffix may come before or after
42//! the length suffix. `wb` does not combine with `l` at all.
43//!
44//! Binary constants are accepted in every dialect by both compilers, as an extension before
45//! C23. Digit separators are C23 only in both. `_BitInt` constants are C23 in the standard and
46//! clang accepts them in every dialect.
47//!
48//! A `wb` constant has the narrowest type that holds it, which for a signed one includes the
49//! sign bit and is never less than two: `1wb` is `_BitInt(2)`, `42wb` is `_BitInt(7)`, `255uwb`
50//! is `unsigned _BitInt(8)` and `0uwb` is `unsigned _BitInt(1)`. Measured against gcc 16 and
51//! clang, which agree on every one of them.
52//!
53//! # Floating constants
54//!
55//! A floating constant has none of the table walk about it: the suffix names the type outright,
56//! and with no suffix it is a `double`. What it has instead is a list of suffixes that is much
57//! longer than the standard's three, and a conversion that has to be exactly right.
58//!
59//! The conversion is [`rucc_base::float`], correctly rounded and done in software, so that the
60//! bits do not depend on the machine the compiler runs on. The type decides the format and the
61//! target decides what some of the types are: `long double` is the x87 eighty bit format on
62//! x86-64 Linux and true quad precision on AArch64 Linux, and both of them say 128 bits wide,
63//! which is why [`TargetInfo::long_double_format`] exists.
64//!
65//! The suffixes were measured on gcc 13.3 on x86-64 Linux, with `_Generic` for the type and by
66//! printing the bytes for the format. `f` is `float` and `l` is `long double`, and then the
67//! extensions: `q` is `__float128`, `w` is the x87 `__float80`, `d` is a `double` written the
68//! long way, `f16` `f32` `f64` `f128` are the `_FloatN` types and `f32x` `f64x` the `_FloatNx`
69//! ones, and `i` or `j` in either position makes the constant imaginary. `_Float32x` turns out
70//! to be plain `double` and `_Float64x` the x87 format, which is not what the names suggest:
71//! `0.1f32x` is `0x3fb999999999999a` and `0.1f64x` is `0x3ffbcccccccccccccccd`.
72//!
73//! The case rules are their own small grammar. The `f` of a `_FloatN` suffix may be either case
74//! and the trailing `x` may not, so `F64x` is a constant and `f64X` is not. A decimal float
75//! suffix is two letters that have to agree, so `df` and `DF` are constants and `Df` is not.
76//! Every one of these is accepted in every dialect, C89 included, and every one of them is
77//! worth a remark when the dialect did not ask for it.
78//!
79//! Decimal floating constants are refused rather than converted, because there is no decimal
80//! float value anywhere in this compiler to put one in. That is a gap and the error says so.
81
82use rucc_base::float::{Float, Format, ParseError, Status};
83use rucc_session::Std;
84use rucc_target::TargetInfo;
85use rucc_tuple::Arch;
86use rucc_types::{IntKind, int_width};
87
88use crate::remarks::Remarks;
89
90/// A converted integer constant.
91#[derive(Debug, Clone, Copy, PartialEq, Eq)]
92pub struct IntConstant {
93    /// The value, which is never negative: a minus sign is an operator and not part of the
94    /// constant, which is why `-2147483648` is a `long` on a 32-bit `int` and the reason
95    /// `INT_MIN` is spelled the way it is in `limits.h`.
96    pub value: u128,
97    /// The type the table walk arrived at.
98    pub ty: IntConstantType,
99    /// What is worth saying about the constant, for the caller that holds the span.
100    pub remarks: Remarks,
101}
102
103/// The type of an integer constant.
104#[derive(Debug, Clone, Copy, PartialEq, Eq)]
105pub enum IntConstantType {
106    /// One of the integer kinds, chosen by the table walk.
107    Standard(IntKind),
108    /// A `_BitInt` of exactly the width it takes to hold the value.
109    BitInt {
110        /// Whether the type is signed, which it is when the `u` suffix was not there.
111        signed: bool,
112        /// The width in bits, including the sign bit when there is one.
113        width: u32,
114    },
115}
116
117impl IntConstantType {
118    /// A suffix that names this type, for a printer putting the constant back.
119    ///
120    /// The value decides the rest. Nothing spells `__int128`, because no suffix does and none is
121    /// needed: a constant only gets that type by being too large for every other one, so writing
122    /// the value back gets the same answer out of the table walk.
123    #[must_use]
124    pub const fn suffix(self) -> &'static str {
125        match self {
126            IntConstantType::Standard(kind) => match kind {
127                IntKind::UInt | IntKind::UInt128 => "u",
128                IntKind::Long => "l",
129                IntKind::ULong => "ul",
130                IntKind::LongLong => "ll",
131                IntKind::ULongLong => "ull",
132                _ => "",
133            },
134            IntConstantType::BitInt { signed: true, .. } => "wb",
135            IntConstantType::BitInt { signed: false, .. } => "uwb",
136        }
137    }
138}
139
140/// Why a preprocessing number is not an integer constant.
141///
142/// [`IntError::Floating`] is not a diagnostic. It means the spelling belongs to the floating
143/// path, and it is an error here so that the caller cannot forget to ask.
144#[derive(Debug, Clone, Copy, PartialEq, Eq)]
145pub enum IntError {
146    /// This is a floating constant. Nothing is wrong with it.
147    Floating,
148    /// The characters after the digits are not a suffix.
149    InvalidSuffix,
150    /// An `8` or a `9` in a constant that started with `0`.
151    InvalidOctalDigit,
152    /// `0x` or `0b` with no digits after it.
153    NoDigits,
154    /// Larger than any integer type, or than the hundred and twenty eight bits the value is
155    /// accumulated in.
156    TooLarge,
157}
158
159impl IntError {
160    /// What to print, in GCC's words where GCC has any.
161    ///
162    /// The offending character is not in the message, because the caller has the spelling and
163    /// the span and can say `invalid suffix "ux" on integer constant` the way GCC does.
164    #[must_use]
165    pub const fn message(self) -> &'static str {
166        match self {
167            IntError::Floating => "not an integer constant",
168            IntError::InvalidSuffix => "invalid suffix on integer constant",
169            IntError::InvalidOctalDigit => "invalid digit in octal constant",
170            IntError::NoDigits => "no digits in integer constant",
171            IntError::TooLarge => "integer constant is too large to be represented in any type",
172        }
173    }
174}
175
176/// A converted floating constant.
177#[derive(Debug, Clone, Copy, PartialEq, Eq)]
178pub struct FloatConstant {
179    /// The value, correctly rounded into the format of its type.
180    pub value: Float,
181    /// The type the suffix named, which is `double` when there was no suffix.
182    pub ty: FloatConstantType,
183    /// Whether an `i` or a `j` made this the imaginary part of a complex constant. The value is
184    /// the real number that was written either way, so the caller builds the complex one.
185    pub imaginary: bool,
186    /// What is worth saying about the constant, for the caller that holds the span.
187    pub remarks: Remarks,
188}
189
190/// The type of a floating constant.
191///
192/// This is not [`rucc_types::FloatKind`], which has the three types C has. The suffixes reach
193/// further than that, and a constant knows exactly which type it was written as long before
194/// anything has to decide what that type is on this target.
195#[derive(Debug, Clone, Copy, PartialEq, Eq)]
196pub enum FloatConstantType {
197    /// `f` or `F`.
198    Float,
199    /// No suffix, or the `d` and `D` that GCC also accepts for it.
200    Double,
201    /// `l` or `L`.
202    LongDouble,
203    /// `f16` or `F16`, which is `_Float16`.
204    Float16,
205    /// `f32` or `F32`, which is `_Float32` and is the same format as `float`.
206    Float32,
207    /// `f64` or `F64`, which is `_Float64` and is the same format as `double`.
208    Float64,
209    /// `f128` or `F128`, and `q` or `Q`, which is `_Float128` and `__float128`. GCC keeps the
210    /// two spellings as distinct types and they are the same format, which is all this says.
211    Float128,
212    /// `f32x` or `F32x`, which is `_Float32x`. The name suggests something wider than
213    /// `_Float32` and on every target here it is exactly `double`.
214    Float32x,
215    /// `f64x` or `F64x`, which is `_Float64x`: the widest format the target has beyond
216    /// `_Float64`, so the x87 one on x86 and quad precision elsewhere.
217    Float64x,
218    /// `w` or `W`, which is GCC's `__float80`. It is the x87 format whatever `long double` is,
219    /// which is the reason it is not the same thing as [`FloatConstantType::LongDouble`], and
220    /// GCC has it on x86 only.
221    Float80,
222}
223
224impl FloatConstantType {
225    /// The format a constant of this type is converted in.
226    ///
227    /// The two that depend on the target are the two that have to: `long double` is the x87
228    /// format on x86-64 Linux and quad precision on AArch64 Linux, and `_Float64x` is whatever
229    /// the target has above `_Float64`, which is the same split.
230    #[must_use]
231    pub fn format(self, target: &TargetInfo) -> Format {
232        match self {
233            FloatConstantType::Float | FloatConstantType::Float32 => Format::Single,
234            FloatConstantType::Double
235            | FloatConstantType::Float64
236            | FloatConstantType::Float32x => Format::Double,
237            FloatConstantType::LongDouble => target.long_double_format,
238            FloatConstantType::Float16 => Format::Half,
239            FloatConstantType::Float128 => Format::Quad,
240            FloatConstantType::Float64x if target.tuple.arch() == Arch::X86_64 => {
241                Format::X87Extended
242            }
243            FloatConstantType::Float64x => Format::Quad,
244            FloatConstantType::Float80 => Format::X87Extended,
245        }
246    }
247
248    /// The C spelling of the type, which a diagnostic naming it has to print.
249    ///
250    /// GCC says "floating constant exceeds range of 'double'" and puts the type in the message,
251    /// so the type has to be able to say what it is called.
252    #[must_use]
253    pub const fn name(self) -> &'static str {
254        match self {
255            FloatConstantType::Float => "float",
256            FloatConstantType::Double => "double",
257            FloatConstantType::LongDouble => "long double",
258            FloatConstantType::Float16 => "_Float16",
259            FloatConstantType::Float32 => "_Float32",
260            FloatConstantType::Float64 => "_Float64",
261            FloatConstantType::Float128 => "_Float128",
262            FloatConstantType::Float32x => "_Float32x",
263            FloatConstantType::Float64x => "_Float64x",
264            FloatConstantType::Float80 => "__float80",
265        }
266    }
267
268    /// The suffix that names this type, for a printer putting the constant back.
269    ///
270    /// One spelling per type rather than the one that was written, so `0.1q` and `0.1f128` both
271    /// come back as `f128`. They are the same type, and the printer's job is to mean the same
272    /// thing rather than to look the same.
273    #[must_use]
274    pub const fn suffix(self) -> &'static str {
275        match self {
276            FloatConstantType::Float => "f",
277            FloatConstantType::Double => "",
278            FloatConstantType::LongDouble => "l",
279            FloatConstantType::Float16 => "f16",
280            FloatConstantType::Float32 => "f32",
281            FloatConstantType::Float64 => "f64",
282            FloatConstantType::Float128 => "f128",
283            FloatConstantType::Float32x => "f32x",
284            FloatConstantType::Float64x => "f64x",
285            FloatConstantType::Float80 => "w",
286        }
287    }
288}
289
290/// Why a preprocessing number is not a floating constant.
291///
292/// [`FloatError::Integer`] is not a diagnostic, in the same way [`IntError::Floating`] is not:
293/// it means the spelling belongs to the other path.
294#[derive(Debug, Clone, Copy, PartialEq, Eq)]
295pub enum FloatError {
296    /// This is an integer constant. Nothing is wrong with it.
297    Integer,
298    /// The characters after the number are not a suffix.
299    InvalidSuffix,
300    /// A hexadecimal floating constant with no `p` exponent. The exponent is required there and
301    /// not optional as it is in a decimal one, because `f` is a hexadecimal digit and there
302    /// would be no way to tell a suffix from the number.
303    MissingExponent,
304    /// An `e` or a `p` with no digits after it.
305    NoExponentDigits,
306    /// A constant with a point and no digits at all.
307    NoDigits,
308    /// More than one point, which is a preprocessing number and not a constant.
309    TooManyPoints,
310    /// A `df`, `dd` or `dl` suffix. The constant is well formed and this compiler has nowhere
311    /// to put a decimal floating value yet.
312    DecimalFloat,
313    /// A suffix naming a type this target does not have, which is `w` anywhere but x86 and
314    /// `f128x` everywhere.
315    UnsupportedType,
316}
317
318impl FloatError {
319    /// What to print, in GCC's words where GCC has any.
320    #[must_use]
321    pub const fn message(self) -> &'static str {
322        match self {
323            FloatError::Integer => "not a floating constant",
324            FloatError::InvalidSuffix => "invalid suffix on floating constant",
325            FloatError::MissingExponent => "hexadecimal floating constants require an exponent",
326            FloatError::NoExponentDigits => "exponent has no digits",
327            FloatError::NoDigits => "no digits in floating constant",
328            FloatError::TooManyPoints => "too many decimal points in number",
329            FloatError::DecimalFloat => "decimal floating constants are not supported yet",
330            FloatError::UnsupportedType => {
331                "the type of this floating constant is not supported on this target"
332            }
333        }
334    }
335}
336
337/// Converts the spelling of a preprocessing number into an integer constant.
338///
339/// # Errors
340///
341/// [`IntError`], one case of which is that the spelling is a floating constant rather than a
342/// malformed integer one.
343pub fn integer(text: &str, std: Std, target: &TargetInfo) -> Result<IntConstant, IntError> {
344    let bytes = text.as_bytes();
345    let (base, start) = base_of(bytes);
346    if is_floating(bytes, base) {
347        return Err(IntError::Floating);
348    }
349    let mut remarks = Remarks::NONE;
350    if base == 2 && std < Std::C23 {
351        remarks = remarks.with(Remarks::BINARY);
352    }
353
354    let mut value: u128 = 0;
355    let mut digits = 0;
356    let mut index = start;
357    while index < bytes.len() {
358        let byte = bytes[index];
359        if byte == b'\'' {
360            // A separator is only a separator between two digits. The scanner keeps one in the
361            // number only when an identifier character follows, so a trailing one arrives here
362            // as a suffix instead and is refused as one.
363            if digits == 0 || index + 1 >= bytes.len() || digit(bytes[index + 1], base).is_none() {
364                return Err(IntError::InvalidSuffix);
365            }
366            if std < Std::C23 {
367                remarks = remarks.with(Remarks::SEPARATORS);
368            }
369            index += 1;
370            continue;
371        }
372        let Some(digit) = digit(byte, base) else {
373            break;
374        };
375        value = value
376            .checked_mul(u128::from(base))
377            .and_then(|shifted| shifted.checked_add(u128::from(digit)))
378            .ok_or(IntError::TooLarge)?;
379        digits += 1;
380        index += 1;
381    }
382    if digits == 0 {
383        // `0x` with nothing after it, which GCC reports as an invalid suffix because it read
384        // the `0` as the constant. The distinction is not worth a worse message than this.
385        return Err(IntError::NoDigits);
386    }
387    if base == 8 && bytes[start..index].iter().any(|&byte| byte == b'8' || byte == b'9') {
388        return Err(IntError::InvalidOctalDigit);
389    }
390
391    let suffix = suffix_of(&bytes[index..])?;
392    if suffix.length == Some(Length::LongLong) && std == Std::C89 {
393        remarks = remarks.with(Remarks::LONG_LONG);
394    }
395    if suffix.length == Some(Length::BitInt) {
396        if std < Std::C23 {
397            remarks = remarks.with(Remarks::BIT_INT);
398        }
399        return Ok(IntConstant { value, ty: bit_int(value, suffix.unsigned), remarks });
400    }
401
402    let candidates = candidates(base, suffix, std);
403    let kind = candidates
404        .iter()
405        .copied()
406        .find(|&kind| fits(value, kind, target))
407        .ok_or(IntError::TooLarge)?;
408    if base == 10 && !suffix.unsigned && !signed_standard(kind) {
409        remarks = remarks.with(Remarks::UNSIGNED);
410    }
411    Ok(IntConstant { value, ty: IntConstantType::Standard(kind), remarks })
412}
413
414/// The base a spelling is written in, and where its digits start.
415///
416/// A leading `0` means octal only when a digit follows, so `0u` is a decimal zero with a
417/// suffix and `08` is an octal constant with a digit that does not exist. That is the split
418/// GCC makes, and it is what turns `08` into a message about octal rather than about a suffix.
419fn base_of(bytes: &[u8]) -> (u32, usize) {
420    match bytes {
421        [b'0', b'x' | b'X', ..] => (16, 2),
422        [b'0', b'b' | b'B', ..] => (2, 2),
423        [b'0', next, ..] if next.is_ascii_digit() => (8, 1),
424        _ => (10, 0),
425    }
426}
427
428/// Whether the spelling is a floating constant rather than an integer one.
429///
430/// A point anywhere, an `e` exponent in a decimal constant, or a `p` exponent in a hexadecimal
431/// one. `1e` and `1e+` are floating constants with no exponent digits, which is a diagnostic
432/// the floating path gives, and `1f` is an integer constant with a suffix that does not exist,
433/// which is one this path gives. Both compilers split them exactly there.
434///
435/// A leading zero does not survive an exponent: `08e5` is the floating constant eight hundred
436/// thousand and not an octal constant with a digit that does not exist.
437fn is_floating(bytes: &[u8], base: u32) -> bool {
438    let exponent = if base == 16 { *b"pP" } else { *b"eE" };
439    bytes.iter().any(|&byte| byte == b'.' || exponent.contains(&byte))
440}
441
442/// The value of a digit in the given base, and [`None`] when the byte is not one.
443///
444/// An octal constant reads `8` and `9` as digits, so that a constant holding one ends at the
445/// suffix and the error can name the digit rather than complain about the suffix.
446fn digit(byte: u8, base: u32) -> Option<u32> {
447    char::from(byte).to_digit(if base == 8 { 10 } else { base })
448}
449
450/// The length part of a suffix.
451#[derive(Debug, Clone, Copy, PartialEq, Eq)]
452enum Length {
453    /// `l` or `L`.
454    Long,
455    /// `ll` or `LL`.
456    LongLong,
457    /// `wb` or `WB`.
458    BitInt,
459}
460
461/// A parsed suffix.
462#[derive(Debug, Clone, Copy, PartialEq, Eq)]
463struct Suffix {
464    /// Whether `u` or `U` was there.
465    unsigned: bool,
466    /// The length part, when there was one.
467    length: Option<Length>,
468}
469
470/// Reads the suffix, which may hold each part once and in either order.
471fn suffix_of(mut rest: &[u8]) -> Result<Suffix, IntError> {
472    let mut suffix = Suffix { unsigned: false, length: None };
473    while let Some(&byte) = rest.first() {
474        let taken = match byte {
475            b'u' | b'U' if !suffix.unsigned => {
476                suffix.unsigned = true;
477                1
478            }
479            // The two letters have to agree about case, so `1ll` and `1LL` are constants and
480            // `1lL` is not. Both compilers refuse the mixed spelling in every dialect.
481            b'l' | b'L' if suffix.length.is_none() => {
482                if rest.get(1) == Some(&byte) {
483                    suffix.length = Some(Length::LongLong);
484                    2
485                } else {
486                    suffix.length = Some(Length::Long);
487                    1
488                }
489            }
490            b'w' | b'W' if suffix.length.is_none() => {
491                let second = if byte == b'w' { b'b' } else { b'B' };
492                if rest.get(1) != Some(&second) {
493                    return Err(IntError::InvalidSuffix);
494                }
495                suffix.length = Some(Length::BitInt);
496                2
497            }
498            _ => return Err(IntError::InvalidSuffix),
499        };
500        rest = &rest[taken..];
501    }
502    Ok(suffix)
503}
504
505/// The type of a `wb` constant, which is the narrowest one that holds the value.
506///
507/// The sign bit counts, so a signed one is never narrower than two bits: `1wb` is
508/// `_BitInt(2)`. An unsigned zero is `unsigned _BitInt(1)`, because a width of zero is not a
509/// type. Measured against gcc 16 and clang.
510fn bit_int(value: u128, unsigned: bool) -> IntConstantType {
511    let used = 128 - value.leading_zeros();
512    let width = if unsigned { used.max(1) } else { used + 1 };
513    IntConstantType::BitInt { signed: !unsigned, width: width.max(if unsigned { 1 } else { 2 }) }
514}
515
516/// Whether `kind` is one of the standard signed types, which is what decides the remark about
517/// a decimal constant having gone unsigned.
518fn signed_standard(kind: IntKind) -> bool {
519    matches!(kind, IntKind::Int | IntKind::Long | IntKind::LongLong)
520}
521
522/// Whether the value fits in `kind` on this target.
523fn fits(value: u128, kind: IntKind, target: &TargetInfo) -> bool {
524    let width = int_width(kind, target);
525    // Signedness here never depends on what plain `char` is, because no candidate list holds a
526    // character type.
527    let bits = if kind.is_signed(false) { width - 1 } else { width };
528    // `unsigned __int128` holds every value the accumulator can, and shifting a `u128` by all
529    // of its bits is not a shift, so the widest type is answered without one.
530    bits >= 128 || value >> bits == 0
531}
532
533/// The candidate list for a base and a suffix, in the order the standard walks it.
534///
535/// `__int128` and `unsigned __int128` are on the end of every list, which is what gcc does:
536/// `9223372036854775808` is an `__int128` there and an `unsigned long long` in clang. Both
537/// compilers put `long long` out of reach in C89 unless the suffix asks for it, and C89 is
538/// also the dialect that offers `unsigned long` for a decimal constant with no suffix at all.
539fn candidates(base: u32, suffix: Suffix, std: Std) -> &'static [IntKind] {
540    use IntKind::{Int, Int128, Long, LongLong, UInt, UInt128, ULong, ULongLong};
541
542    let decimal = base == 10;
543    let c89 = std == Std::C89;
544    match (suffix.unsigned, suffix.length) {
545        (false, None) if decimal && c89 => &[Int, Long, ULong, Int128, UInt128],
546        (false, None) if decimal => &[Int, Long, LongLong, Int128],
547        (false, None) if c89 => &[Int, UInt, Long, ULong, Int128, UInt128],
548        (false, None) => &[Int, UInt, Long, ULong, LongLong, ULongLong, Int128, UInt128],
549
550        (true, None) if c89 => &[UInt, ULong, UInt128],
551        (true, None) => &[UInt, ULong, ULongLong, UInt128],
552
553        (false, Some(Length::Long)) if decimal && c89 => &[Long, ULong, Int128, UInt128],
554        (false, Some(Length::Long)) if decimal => &[Long, LongLong, Int128],
555        (false, Some(Length::Long)) if c89 => &[Long, ULong, Int128, UInt128],
556        (false, Some(Length::Long)) => &[Long, ULong, LongLong, ULongLong, Int128, UInt128],
557
558        (true, Some(Length::Long)) if c89 => &[ULong, UInt128],
559        (true, Some(Length::Long)) => &[ULong, ULongLong, UInt128],
560
561        (false, Some(Length::LongLong)) if decimal => &[LongLong, Int128],
562        (false, Some(Length::LongLong)) => &[LongLong, ULongLong, Int128, UInt128],
563        (true, Some(Length::LongLong)) => &[ULongLong, UInt128],
564
565        // A `wb` constant never reaches here: its type comes from the value alone.
566        (_, Some(Length::BitInt)) => &[],
567    }
568}
569
570/// Converts the spelling of a preprocessing number into a floating constant.
571///
572/// # Errors
573///
574/// [`FloatError`], one case of which is that the spelling is an integer constant rather than a
575/// malformed floating one.
576pub fn floating(text: &str, std: Std, target: &TargetInfo) -> Result<FloatConstant, FloatError> {
577    let bytes = text.as_bytes();
578    let (base, _) = base_of(bytes);
579    if !is_floating(bytes, base) {
580        return Err(FloatError::Integer);
581    }
582    // A leading zero means nothing to a floating constant, so there are two bases here and not
583    // four: `08e5` is eight hundred thousand rather than an octal constant with a bad digit.
584    let hex = base == 16;
585    let base = if hex { 16 } else { 10 };
586    let mut remarks = Remarks::NONE;
587    if hex && std < Std::C99 {
588        remarks = remarks.with(Remarks::HEX_FLOAT);
589    }
590
591    let mut index = if hex { 2 } else { 0 };
592    let mut digits = 0;
593    let mut point = false;
594    let mut separators = false;
595    while index < bytes.len() {
596        let byte = bytes[index];
597        if byte == b'\'' {
598            if digits == 0 || !next_is_digit(bytes, index, base) {
599                return Err(FloatError::InvalidSuffix);
600            }
601            separators = true;
602        } else if byte == b'.' {
603            if point {
604                return Err(FloatError::TooManyPoints);
605            }
606            point = true;
607        } else if digit(byte, base).is_some() {
608            digits += 1;
609        } else {
610            break;
611        }
612        index += 1;
613    }
614    if digits == 0 {
615        return Err(FloatError::NoDigits);
616    }
617
618    let marker = if hex { *b"pP" } else { *b"eE" };
619    if index < bytes.len() && marker.contains(&bytes[index]) {
620        index += 1;
621        if matches!(bytes.get(index), Some(b'+' | b'-')) {
622            index += 1;
623        }
624        let mut exponent_digits = 0;
625        while index < bytes.len() {
626            let byte = bytes[index];
627            if byte == b'\'' {
628                if exponent_digits == 0 || !next_is_digit(bytes, index, 10) {
629                    return Err(FloatError::InvalidSuffix);
630                }
631                separators = true;
632            } else if byte.is_ascii_digit() {
633                exponent_digits += 1;
634            } else {
635                break;
636            }
637            index += 1;
638        }
639        if exponent_digits == 0 {
640            return Err(FloatError::NoExponentDigits);
641        }
642    } else if hex {
643        // The exponent is not optional in a hexadecimal constant, because `f` is a digit there
644        // and `0x1.8f` would otherwise be a number and a suffix at the same time.
645        return Err(FloatError::MissingExponent);
646    }
647    if separators && std < Std::C23 {
648        remarks = remarks.with(Remarks::SEPARATORS);
649    }
650
651    let suffix = float_suffix(&bytes[index..], target)?;
652    remarks = remarks.with(suffix.remarks);
653    let (value, status) =
654        Float::parse(&text[..index], suffix.ty.format(target)).map_err(|error| match error {
655            // The scan above has already ruled all three of these out, and mapping them is
656            // still better than an unwrap that a later change could reach.
657            ParseError::NoDigits => FloatError::NoDigits,
658            ParseError::NoExponentDigits => FloatError::NoExponentDigits,
659            ParseError::Invalid => FloatError::InvalidSuffix,
660        })?;
661    if status.has(Status::OVERFLOW) {
662        remarks = remarks.with(Remarks::OUT_OF_RANGE);
663    }
664    // Underflow on its own is a subnormal, which is a number the program can use. Losing the
665    // value entirely is the part worth a word.
666    if status.has(Status::UNDERFLOW) && value.is_zero() {
667        remarks = remarks.with(Remarks::TRUNCATED);
668    }
669    Ok(FloatConstant { value, ty: suffix.ty, imaginary: suffix.imaginary, remarks })
670}
671
672/// Whether the byte after `index` is a digit in `base`, which is what makes a separator one.
673fn next_is_digit(bytes: &[u8], index: usize, base: u32) -> bool {
674    bytes.get(index + 1).is_some_and(|&next| digit(next, base).is_some())
675}
676
677/// A parsed floating suffix.
678struct FloatSuffix {
679    /// The type it named, which is `double` when it named none.
680    ty: FloatConstantType,
681    /// Whether it held an `i` or a `j`.
682    imaginary: bool,
683    /// What the suffix alone is worth saying about.
684    remarks: Remarks,
685}
686
687/// Reads the suffix, which may name a type once and mark the constant imaginary once, in either
688/// order.
689///
690/// Everything past `f` and `l` is an extension, and the extensions are where the case rules stop
691/// being uniform: the `f` of `_FloatN` may be either case and the `x` of `_FloatNx` may not, and
692/// the two letters of a decimal suffix have to agree. All of it measured on gcc 13.3.
693fn float_suffix(mut rest: &[u8], target: &TargetInfo) -> Result<FloatSuffix, FloatError> {
694    let mut ty = None;
695    let mut imaginary = false;
696    let mut remarks = Remarks::NONE;
697    while let Some(&byte) = rest.first() {
698        let taken = match byte {
699            b'i' | b'j' | b'I' | b'J' if !imaginary => {
700                imaginary = true;
701                remarks = remarks.with(Remarks::IMAGINARY);
702                1
703            }
704            // One type per constant, so `1.0fl` is not a constant and neither is `1.0ff`.
705            _ if ty.is_some() => return Err(FloatError::InvalidSuffix),
706            b'f' | b'F' => {
707                let (named, taken, extra) = float_n(rest)?;
708                ty = Some(named);
709                remarks = remarks.with(extra);
710                taken
711            }
712            b'l' | b'L' => {
713                ty = Some(FloatConstantType::LongDouble);
714                1
715            }
716            b'q' | b'Q' => {
717                ty = Some(FloatConstantType::Float128);
718                remarks = remarks.with(Remarks::EXTENDED_SUFFIX);
719                1
720            }
721            b'w' | b'W' => {
722                // `__float80` is the x87 format, which only x86 has.
723                if target.tuple.arch() != Arch::X86_64 {
724                    return Err(FloatError::UnsupportedType);
725                }
726                ty = Some(FloatConstantType::Float80);
727                remarks = remarks.with(Remarks::EXTENDED_SUFFIX);
728                1
729            }
730            b'd' | b'D' => {
731                let second = rest.get(1).copied();
732                let decimal = if byte == b'd' {
733                    matches!(second, Some(b'f' | b'd' | b'l'))
734                } else {
735                    matches!(second, Some(b'F' | b'D' | b'L'))
736                };
737                if decimal {
738                    return Err(FloatError::DecimalFloat);
739                }
740                ty = Some(FloatConstantType::Double);
741                remarks = remarks.with(Remarks::DOUBLE_SUFFIX);
742                1
743            }
744            _ => return Err(FloatError::InvalidSuffix),
745        };
746        rest = &rest[taken..];
747    }
748    Ok(FloatSuffix { ty: ty.unwrap_or(FloatConstantType::Double), imaginary, remarks })
749}
750
751/// Reads a suffix that starts with `f`, which is `float` on its own and one of the `_FloatN` or
752/// `_FloatNx` types when digits follow.
753///
754/// Returns the type, how many bytes it took and what is worth saying about it.
755fn float_n(rest: &[u8]) -> Result<(FloatConstantType, usize, Remarks), FloatError> {
756    let mut end = 1;
757    while rest.get(end).is_some_and(u8::is_ascii_digit) {
758        end += 1;
759    }
760    if end == 1 {
761        return Ok((FloatConstantType::Float, 1, Remarks::NONE));
762    }
763    // The `x` is lower case in every spelling gcc accepts, so `F64x` is a constant and `f64X`
764    // is not, however odd that looks next to the `F` being free.
765    let extended = rest.get(end) == Some(&b'x');
766    let ty = match (&rest[1..end], extended) {
767        (b"16", false) => FloatConstantType::Float16,
768        (b"32", false) => FloatConstantType::Float32,
769        (b"64", false) => FloatConstantType::Float64,
770        (b"128", false) => FloatConstantType::Float128,
771        (b"32", true) => FloatConstantType::Float32x,
772        (b"64", true) => FloatConstantType::Float64x,
773        // gcc knows the name `_Float128x` and has the type on no target here, and says so
774        // rather than calling the suffix invalid. `_Float16x` is not a type at all.
775        (b"128", true) => return Err(FloatError::UnsupportedType),
776        _ => return Err(FloatError::InvalidSuffix),
777    };
778    Ok((ty, end + usize::from(extended), Remarks::EXTENDED_SUFFIX))
779}
780
781#[cfg(test)]
782mod tests {
783    use rucc_target::Triple;
784
785    use super::*;
786
787    fn linux() -> TargetInfo {
788        TargetInfo::new("x86_64-unknown-linux-gnu".parse::<Triple>().expect("a known triple"))
789    }
790
791    fn aarch64() -> TargetInfo {
792        TargetInfo::new("aarch64-unknown-linux-gnu".parse::<Triple>().expect("a known triple"))
793    }
794
795    /// The value and the type of a constant in the default dialect.
796    fn c23(text: &str) -> Result<IntConstant, IntError> {
797        integer(text, Std::C23, &linux())
798    }
799
800    /// The type of a constant in the given dialect, on x86-64 Linux.
801    fn kind(text: &str, std: Std) -> IntKind {
802        match integer(text, std, &linux()).expect("a valid constant").ty {
803            IntConstantType::Standard(kind) => kind,
804            IntConstantType::BitInt { .. } => panic!("{text} is a _BitInt constant"),
805        }
806    }
807
808    #[test]
809    fn a_constant_in_each_base_has_the_value_it_says() {
810        assert_eq!(c23("0").expect("zero").value, 0);
811        assert_eq!(c23("42").expect("decimal").value, 42);
812        assert_eq!(c23("0777").expect("octal").value, 0o777);
813        assert_eq!(c23("0xdeadBEEF").expect("hex").value, 0xdead_beef);
814        assert_eq!(c23("0b1010").expect("binary").value, 0b1010);
815        assert_eq!(c23("0X10").expect("upper case prefix").value, 16);
816        // A leading zero with nothing after it is a decimal zero rather than an octal one with
817        // no digits, which is the split that lets `0u` through and stops `08`.
818        assert_eq!(c23("0u").expect("zero with a suffix").value, 0);
819    }
820
821    #[test]
822    fn digit_separators_are_stripped_and_reported_before_c23() {
823        let value = c23("1'000'000").expect("a C23 constant");
824        assert_eq!(value.value, 1_000_000);
825        assert!(value.remarks.is_none());
826        assert_eq!(c23("0x1'0").expect("hex with a separator").value, 16);
827
828        let older = integer("1'000", Std::C17, &linux()).expect("still converted");
829        assert!(older.remarks.has(Remarks::SEPARATORS));
830        assert_eq!(older.value, 1000);
831    }
832
833    #[test]
834    fn the_type_of_a_decimal_constant_walks_the_signed_types_only() {
835        // Measured with `_Generic` on gcc 13.3, x86-64 Linux.
836        assert_eq!(kind("2147483647", Std::C23), IntKind::Int);
837        assert_eq!(kind("2147483648", Std::C23), IntKind::Long);
838        assert_eq!(kind("4294967295", Std::C23), IntKind::Long);
839        assert_eq!(kind("9223372036854775807", Std::C23), IntKind::Long);
840        // Past `long long` gcc reaches for `__int128` rather than for an unsigned type, and
841        // says so: the constant is so large that it is unsigned.
842        assert_eq!(kind("9223372036854775808", Std::C23), IntKind::Int128);
843        assert_eq!(kind("18446744073709551615", Std::C23), IntKind::Int128);
844        let large = c23("18446744073709551615").expect("fits __int128");
845        assert!(large.remarks.has(Remarks::UNSIGNED));
846    }
847
848    #[test]
849    fn a_constant_in_another_base_may_be_unsigned_without_saying_so() {
850        // This is the split that surprises people: `4294967295` is a `long` and `0xffffffff`
851        // is an `unsigned int`, because only the decimal list is signed types alone.
852        assert_eq!(kind("0xffffffff", Std::C23), IntKind::UInt);
853        assert_eq!(kind("0x7fffffff", Std::C23), IntKind::Int);
854        assert_eq!(kind("0x80000000", Std::C23), IntKind::UInt);
855        assert_eq!(kind("0x100000000", Std::C23), IntKind::Long);
856        assert_eq!(kind("0xffffffffffffffff", Std::C23), IntKind::ULong);
857        assert_eq!(kind("0777", Std::C23), IntKind::Int);
858        assert_eq!(kind("0b1010", Std::C23), IntKind::Int);
859        // And no remark, because nothing about it is surprising enough to say.
860        assert!(c23("0xffffffff").expect("a constant").remarks.is_none());
861    }
862
863    #[test]
864    fn c89_has_unsigned_long_in_the_decimal_list_and_no_long_long_in_any() {
865        // gcc under `-std=c89 -pedantic`: "this decimal constant is unsigned only in ISO C90",
866        // and eight bytes rather than sixteen.
867        assert_eq!(kind("18446744073709551615", Std::C89), IntKind::ULong);
868        assert_eq!(kind("18446744073709551615", Std::C99), IntKind::Int128);
869        let old = integer("18446744073709551615", Std::C89, &linux()).expect("a C89 constant");
870        assert!(old.remarks.has(Remarks::UNSIGNED));
871        // The suffix still reaches `long long`, with the remark gcc prints for it.
872        let long_long = integer("1ll", Std::C89, &linux()).expect("an extension");
873        assert!(long_long.remarks.has(Remarks::LONG_LONG));
874        assert_eq!(kind("1ll", Std::C89), IntKind::LongLong);
875        assert!(integer("1ll", Std::C99, &linux()).expect("standard").remarks.is_none());
876    }
877
878    #[test]
879    fn a_suffix_narrows_the_list_it_does_not_pick_the_type() {
880        assert_eq!(kind("1u", Std::C23), IntKind::UInt);
881        assert_eq!(kind("1l", Std::C23), IntKind::Long);
882        assert_eq!(kind("1ul", Std::C23), IntKind::ULong);
883        assert_eq!(kind("1ll", Std::C23), IntKind::LongLong);
884        assert_eq!(kind("1llu", Std::C23), IntKind::ULongLong);
885        // The suffix is a floor rather than an answer: `4294967296u` is an `unsigned long`
886        // because `unsigned int` cannot hold it.
887        assert_eq!(kind("4294967296u", Std::C23), IntKind::ULong);
888        assert_eq!(kind("0xffffffffu", Std::C23), IntKind::UInt);
889    }
890
891    #[test]
892    fn the_letters_of_a_suffix_may_be_in_either_case_but_not_both() {
893        for text in ["1u", "1U", "1l", "1L", "1ll", "1LL", "1ul", "1lu", "1uL", "1LLU", "1llu"] {
894            assert!(c23(text).is_ok(), "{text} is a constant in both compilers");
895        }
896        for text in ["1lL", "1Ll", "1uu", "1lul", "1z", "1uz", "1f", "1x", "1_000"] {
897            assert_eq!(c23(text), Err(IntError::InvalidSuffix), "{text} is not");
898        }
899    }
900
901    #[test]
902    fn a_bit_int_constant_has_the_narrowest_type_that_holds_it() {
903        // Measured against clang, which is the only one of the two that has the type.
904        let cases = [
905            ("0wb", true, 2),
906            ("1wb", true, 2),
907            ("3wb", true, 3),
908            ("42wb", true, 7),
909            ("255wb", true, 9),
910            ("0uwb", false, 1),
911            ("1uwb", false, 1),
912            ("255uwb", false, 8),
913            ("256uwb", false, 9),
914            ("0xffffffffffffffffuwb", false, 64),
915        ];
916        for (text, signed, width) in cases {
917            let constant = c23(text).expect("a _BitInt constant");
918            assert_eq!(
919                constant.ty,
920                IntConstantType::BitInt { signed, width },
921                "{text} is the wrong width"
922            );
923        }
924        // Either order, either case, and never with a length suffix.
925        for text in ["1uwb", "1wbu", "1UWB", "1WBu", "1uWB"] {
926            assert!(c23(text).is_ok(), "{text} is a constant in clang");
927        }
928        for text in ["1wB", "1Wb", "1lwb", "1wbl", "1wbwb"] {
929            assert_eq!(c23(text), Err(IntError::InvalidSuffix), "{text} is not");
930        }
931        // Before C23 it is still converted, and still worth a word.
932        let older = integer("1wb", Std::C17, &linux()).expect("clang accepts it everywhere");
933        assert!(older.remarks.has(Remarks::BIT_INT));
934    }
935
936    #[test]
937    fn a_binary_constant_is_an_extension_before_c23() {
938        assert!(c23("0b1").expect("standard in C23").remarks.is_none());
939        let older = integer("0b1", Std::C17, &linux()).expect("both compilers accept it");
940        assert!(older.remarks.has(Remarks::BINARY));
941    }
942
943    #[test]
944    fn an_octal_constant_names_the_digit_that_is_not_one() {
945        assert_eq!(c23("08"), Err(IntError::InvalidOctalDigit));
946        assert_eq!(c23("0778"), Err(IntError::InvalidOctalDigit));
947        assert_eq!(c23("09"), Err(IntError::InvalidOctalDigit));
948        // A `9` elsewhere is fine, and the message is only for constants that began with `0`.
949        assert_eq!(c23("9").expect("decimal").value, 9);
950    }
951
952    #[test]
953    fn a_prefix_with_no_digits_after_it_is_not_a_constant() {
954        assert_eq!(c23("0x"), Err(IntError::NoDigits));
955        assert_eq!(c23("0b"), Err(IntError::NoDigits));
956    }
957
958    #[test]
959    fn a_constant_larger_than_any_type_is_refused_rather_than_wrapped() {
960        // gcc accumulates in sixty four bits and silently gives this the value zero and the
961        // type `int` after a warning. That is the one measured behaviour here we refuse to
962        // reproduce, and clang refuses it too.
963        assert_eq!(c23("340282366920938463463374607431768211456"), Err(IntError::TooLarge));
964        assert_eq!(c23("0x100000000000000000000000000000000"), Err(IntError::TooLarge));
965        // 2^127 fits in the accumulator and in no signed type, and the decimal list has no
966        // unsigned one to fall back to.
967        assert_eq!(c23("170141183460469231731687303715884105728"), Err(IntError::TooLarge));
968        // The same value written in hex reaches `unsigned __int128`, because that list has it.
969        assert_eq!(kind("0x80000000000000000000000000000000", Std::C23), IntKind::UInt128);
970        assert_eq!(kind("0xffffffffffffffffffffffffffffffff", Std::C23), IntKind::UInt128);
971    }
972
973    #[test]
974    fn a_floating_constant_is_handed_back_rather_than_refused() {
975        for text in ["1.0", ".5", "1.", "1e5", "1E-5", "1e", "0x1p3", "0x1.8p+1", "1.5e3"] {
976            assert_eq!(c23(text), Err(IntError::Floating), "{text} belongs to the other path");
977        }
978        // A leading zero does not make this an octal constant with a digit that does not
979        // exist. gcc compiles it, as eight hundred thousand.
980        assert_eq!(c23("08e5"), Err(IntError::Floating));
981        // A hexadecimal `e` is a digit, not an exponent, and `1f` is an integer with a suffix
982        // that does not exist rather than a float. Both compilers split them there.
983        assert_eq!(c23("0xe5").expect("hex digits").value, 0xe5);
984        assert_eq!(c23("1f"), Err(IntError::InvalidSuffix));
985    }
986
987    #[test]
988    fn the_type_comes_from_the_target_and_not_from_the_host() {
989        // `4294967295` is a `long` where `long` is sixty four bits and a `long long` where it
990        // is thirty two. A compiler that asked its own platform gets one of these wrong.
991        let windows =
992            TargetInfo::new("x86_64-pc-windows-msvc".parse::<Triple>().expect("a known triple"));
993        let on_windows = integer("4294967295", Std::C23, &windows).expect("a constant");
994        assert_eq!(on_windows.ty, IntConstantType::Standard(IntKind::LongLong));
995        assert_eq!(kind("4294967295", Std::C23), IntKind::Long);
996    }
997
998    /// A floating constant in the default dialect, on x86-64 Linux.
999    fn float(text: &str) -> Result<FloatConstant, FloatError> {
1000        floating(text, Std::C23, &linux())
1001    }
1002
1003    /// The bits of a constant's value, which is the form every measured row here was taken in.
1004    fn bits(text: &str) -> u128 {
1005        float(text).expect("a valid constant").value.to_bits()
1006    }
1007
1008    #[test]
1009    fn a_constant_with_no_suffix_is_a_double() {
1010        let constant = float("1.5").expect("a constant");
1011        assert_eq!(constant.ty, FloatConstantType::Double);
1012        assert!(!constant.imaginary);
1013        assert!(constant.remarks.is_none());
1014        assert_eq!(constant.value.to_bits(), 0x3ff8_0000_0000_0000);
1015        assert_eq!(bits("0.1"), 0x3fb9_9999_9999_999a);
1016        assert_eq!(bits(".5"), 0x3fe0_0000_0000_0000);
1017        assert_eq!(bits("1."), 0x3ff0_0000_0000_0000);
1018        assert_eq!(bits("1e5"), 0x40f8_6a00_0000_0000);
1019        assert_eq!(bits("0x1p3"), 0x4020_0000_0000_0000);
1020        // A leading zero is not an octal prefix once there is an exponent, so this is eight
1021        // hundred thousand and gcc compiles it as one.
1022        assert_eq!(bits("08e5"), 0x4128_6a00_0000_0000);
1023    }
1024
1025    #[test]
1026    fn the_suffix_names_the_type_rather_than_narrowing_a_list() {
1027        // Measured with `_Generic` on gcc 13.3, x86-64 Linux.
1028        let cases = [
1029            ("1.0", FloatConstantType::Double),
1030            ("1.0f", FloatConstantType::Float),
1031            ("1.0F", FloatConstantType::Float),
1032            ("1.0l", FloatConstantType::LongDouble),
1033            ("1.0L", FloatConstantType::LongDouble),
1034            ("1.0d", FloatConstantType::Double),
1035            ("1.0q", FloatConstantType::Float128),
1036            ("1.0w", FloatConstantType::Float80),
1037            ("1.0f16", FloatConstantType::Float16),
1038            ("1.0F16", FloatConstantType::Float16),
1039            ("1.0f32", FloatConstantType::Float32),
1040            ("1.0f64", FloatConstantType::Float64),
1041            ("1.0f128", FloatConstantType::Float128),
1042            ("1.0f32x", FloatConstantType::Float32x),
1043            ("1.0F64x", FloatConstantType::Float64x),
1044        ];
1045        for (text, ty) in cases {
1046            assert_eq!(float(text).expect("a constant").ty, ty, "{text} has the wrong type");
1047        }
1048    }
1049
1050    #[test]
1051    fn each_type_is_converted_in_the_format_the_target_has_for_it() {
1052        // Every row measured by printing the bytes of the constant on gcc 13.3, x86-64 Linux.
1053        // The two that surprise are `_Float32x`, which is plain `double`, and `_Float64x`,
1054        // which is the x87 format and so the same bits as `long double` and `__float80`.
1055        assert_eq!(bits("0.1f"), 0x3dcc_cccd);
1056        assert_eq!(bits("0.1f16"), 0x2e66);
1057        assert_eq!(bits("0.1f32x"), 0x3fb9_9999_9999_999a);
1058        assert_eq!(bits("0.1f64x"), 0x3ffb_cccc_cccc_cccc_cccd);
1059        assert_eq!(bits("0.1w"), 0x3ffb_cccc_cccc_cccc_cccd);
1060        assert_eq!(bits("0.1l"), 0x3ffb_cccc_cccc_cccc_cccd);
1061        assert_eq!(bits("0.1q"), 0x3ffb_9999_9999_9999_9999_9999_9999_999a);
1062        assert_eq!(bits("0.1f128"), 0x3ffb_9999_9999_9999_9999_9999_9999_999a);
1063        assert_eq!(bits("1.0l"), 0x3fff_8000_0000_0000_0000);
1064    }
1065
1066    #[test]
1067    fn the_format_comes_from_the_target_and_not_from_the_host() {
1068        // `long double` is 128 bits wide on both of these and it is not the same type on both,
1069        // which is the whole reason the target carries a format and not only a width.
1070        let arm = floating("1.0l", Std::C23, &aarch64()).expect("a constant");
1071        assert_eq!(arm.value.to_bits(), 0x3fff_0000_0000_0000_0000_0000_0000_0000);
1072        assert_eq!(bits("1.0l"), 0x3fff_8000_0000_0000_0000);
1073        // And `_Float64x` follows it, being whatever the target has above `_Float64`.
1074        let arm_wide = floating("0.1f64x", Std::C23, &aarch64()).expect("a constant");
1075        assert_eq!(arm_wide.value.to_bits(), 0x3ffb_9999_9999_9999_9999_9999_9999_999a);
1076        // Where `long double` is `double` the constant is a `double` too.
1077        let windows =
1078            TargetInfo::new("x86_64-pc-windows-msvc".parse::<Triple>().expect("a known triple"));
1079        let on_windows = floating("1.0l", Std::C23, &windows).expect("a constant");
1080        assert_eq!(on_windows.value.to_bits(), 0x3ff0_0000_0000_0000);
1081    }
1082
1083    #[test]
1084    fn the_case_rules_of_a_floating_suffix_are_not_uniform() {
1085        // The `f` of a `_FloatN` suffix is free and the `x` of a `_FloatNx` one is not, and the
1086        // two letters of a decimal suffix have to agree. Measured on gcc 13.3, which accepts
1087        // every one of these in every dialect.
1088        for text in ["1.0f", "1.0F", "1.0L", "1.0Q", "1.0W", "1.0F32", "1.0f64x", "1.0F64x"] {
1089            assert!(float(text).is_ok(), "{text} is a constant in gcc");
1090        }
1091        for text in ["1.0F32X", "1.0f32X", "1.0f16x", "1.0ff", "1.0fl", "1.0lf", "1.0fF", "1.0LL"] {
1092            assert_eq!(float(text), Err(FloatError::InvalidSuffix), "{text} is not");
1093        }
1094    }
1095
1096    #[test]
1097    fn an_imaginary_suffix_may_sit_on_either_side_of_the_type() {
1098        for text in ["1.0i", "1.0j", "1.0I", "1.0J", "1.0if", "1.0fi", "1.0Li", "1.0iL", "1.0f16i"]
1099        {
1100            let constant = float(text).expect("a constant in gcc");
1101            assert!(constant.imaginary, "{text} is imaginary");
1102            assert!(constant.remarks.has(Remarks::IMAGINARY));
1103        }
1104        assert_eq!(float("1.0ii"), Err(FloatError::InvalidSuffix));
1105        assert_eq!(float("1.0ij"), Err(FloatError::InvalidSuffix));
1106        assert!(!float("1.0f").expect("a constant").imaginary);
1107    }
1108
1109    #[test]
1110    fn a_decimal_floating_constant_is_recognised_and_refused() {
1111        // The constant is well formed and there is nowhere in this compiler to put its value.
1112        for text in ["1.0df", "1.0dd", "1.0dl", "1.0DF", "1.0DD", "1.0DL"] {
1113            assert_eq!(float(text), Err(FloatError::DecimalFloat), "{text} is a decimal float");
1114        }
1115        // The letters have to agree about case, so these are not decimal floats and not
1116        // constants either.
1117        for text in ["1.0Df", "1.0dF", "1.0dD", "1.0Dl"] {
1118            assert_eq!(float(text), Err(FloatError::InvalidSuffix), "{text} is neither");
1119        }
1120        // A `d` on its own is a `double` written the long way, which gcc allows everywhere.
1121        let long_way = float("1.0d").expect("a GCC extension");
1122        assert_eq!(long_way.ty, FloatConstantType::Double);
1123        assert!(long_way.remarks.has(Remarks::DOUBLE_SUFFIX));
1124    }
1125
1126    #[test]
1127    fn a_type_the_target_does_not_have_is_refused_by_name() {
1128        // gcc says "'_Float128x' is not supported on this target" rather than calling the
1129        // suffix invalid, and it is supported on no target here.
1130        assert_eq!(float("1.0f128x"), Err(FloatError::UnsupportedType));
1131        // `__float80` is the x87 format, which only x86 has.
1132        assert_eq!(floating("1.0w", Std::C23, &aarch64()), Err(FloatError::UnsupportedType));
1133        assert!(float("1.0w").is_ok());
1134    }
1135
1136    #[test]
1137    fn a_hexadecimal_constant_needs_an_exponent_and_a_decimal_one_does_not() {
1138        // `f` is a hexadecimal digit, so without the exponent there is no telling the number
1139        // from the suffix. Both compilers require it.
1140        assert_eq!(float("0x1.8"), Err(FloatError::MissingExponent));
1141        assert_eq!(bits("0x1.8p0"), 0x3ff8_0000_0000_0000);
1142        assert_eq!(bits("0x.8p1"), 0x3ff0_0000_0000_0000);
1143        assert_eq!(bits("1.5"), 0x3ff8_0000_0000_0000);
1144        for text in ["1.0e", "1e+", "1e-", "0x1p", "0x1p+"] {
1145            assert_eq!(float(text), Err(FloatError::NoExponentDigits), "{text} has no exponent");
1146        }
1147        assert_eq!(float("1.2.3"), Err(FloatError::TooManyPoints));
1148    }
1149
1150    #[test]
1151    fn an_integer_constant_is_handed_back_rather_than_refused() {
1152        for text in ["1", "0", "0x10", "1u", "0777", "1wb", "0b1", "0xe5", "1f"] {
1153            assert_eq!(float(text), Err(FloatError::Integer), "{text} belongs to the other path");
1154        }
1155    }
1156
1157    #[test]
1158    fn a_value_past_the_range_of_its_type_is_still_a_constant() {
1159        let large = float("1e400").expect("a constant gcc compiles");
1160        assert!(large.value.is_infinite());
1161        assert!(large.remarks.has(Remarks::OUT_OF_RANGE));
1162        let small = float("1e-400").expect("a constant gcc compiles");
1163        assert!(small.value.is_zero());
1164        assert!(small.remarks.has(Remarks::TRUNCATED));
1165        // The same two in the format the suffix asked for rather than in `double`.
1166        assert!(float("1e39f").expect("a constant").remarks.has(Remarks::OUT_OF_RANGE));
1167        assert!(float("1e-46f").expect("a constant").remarks.has(Remarks::TRUNCATED));
1168        assert!(float("1e-4951l").expect("a constant").remarks.has(Remarks::TRUNCATED));
1169        // A subnormal is a number the program can use, and gcc says nothing about it.
1170        let subnormal = float("1e-320").expect("a constant");
1171        assert!(!subnormal.value.is_zero());
1172        assert!(subnormal.remarks.is_none());
1173    }
1174
1175    #[test]
1176    fn the_dialect_decides_what_a_constant_is_worth_saying_about() {
1177        // "use of C99 hexadecimal floating constant", which gcc says under C89 and not after.
1178        let old = floating("0x1p3", Std::C89, &linux()).expect("gcc compiles it anyway");
1179        assert!(old.remarks.has(Remarks::HEX_FLOAT));
1180        assert!(floating("0x1p3", Std::C99, &linux()).expect("standard").remarks.is_none());
1181        // Separators are C23 in both compilers, in the number and in the exponent.
1182        assert_eq!(bits("1'0.5"), 0x4025_0000_0000_0000);
1183        assert_eq!(bits("1.0e1'0"), 0x4202_a05f_2000_0000);
1184        assert!(float("0x1'0p0").expect("a C23 constant").remarks.is_none());
1185        let older = floating("1'0.5", Std::C17, &linux()).expect("still converted");
1186        assert!(older.remarks.has(Remarks::SEPARATORS));
1187        // And every extension suffix is accepted in every dialect, with a word about it.
1188        for text in ["1.0q", "1.0w", "1.0f16", "1.0f32x"] {
1189            let constant = floating(text, Std::C89, &linux()).expect("gcc accepts it in C89");
1190            assert!(constant.remarks.has(Remarks::EXTENDED_SUFFIX), "{text} is not standard");
1191        }
1192        assert!(float("1.0f").expect("a constant").remarks.is_none());
1193    }
1194}