Skip to main content

rucc_base/
float.rs

1//! Binary floating point, in software, for every format the compiler has to produce.
2//!
3//! A compiler cannot ask the machine it is running on what a floating constant means. The host
4//! may not have the format at all, `long double` is eighty bits on x86-64 and a hundred and
5//! twenty eight on AArch64 Linux and sixty four on Apple, and `strtod` is the host's libc
6//! rather than the target's semantics. Reproducible output means the same source gives the same
7//! bits whoever compiles it, so the conversion is done here, exactly, in integer arithmetic.
8//!
9//! [`Float`] is a sign, a category, an exponent and a significand of up to a hundred and
10//! thirteen bits, which is every IEEE encoding in [`Format`] including the x87 eighty bit one
11//! with its stored leading bit. The value of a finite number is `significand * 2^(exponent -
12//! precision + 1)`, so the significand is an integer rather than a fraction and the exponent is
13//! that of its leading bit.
14//!
15//! [`Format::DoubleDouble`] is the one format that shape does not fit, because a double-double is
16//! a pair of doubles rather than one number with one exponent, and the two halves can sit two
17//! thousand bits apart. [`Float`] refuses it, at [`Format::is_ieee`], in every constructor rather
18//! than at the point some later arithmetic gives a wrong answer. Representing one is what a
19//! PowerPC backend will need and there is no PowerPC backend, so the format is here to be
20//! described by `rucc-abi` and named in a data layout, which is what the fifteen psABIs of
21//! `spec/cross-compile/06-abis.md` section 6.1 want from it today.
22//!
23//! Conversion from text is correctly rounded, round to nearest with ties to even, which is the
24//! only rounding mode a translation-time constant uses. The decimal path scales the number by
25//! powers of two until it is in `[1, 2)` and then reads the significand off it, using the exact
26//! decimal in `decimal.rs` so that no step ever loses a bit. A naive `mantissa * 10^exponent`
27//! in `f64` is wrong in the last place for a noticeable fraction of literals, and the last
28//! place is exactly what a differential test against another compiler notices. Hexadecimal
29//! constants are exact by construction and only have to be rounded once.
30//!
31//! ```
32//! use rucc_base::float::{Float, Format};
33//!
34//! let (value, status) = Float::parse("0.1", Format::Double).expect("a number");
35//! assert_eq!(value.to_bits(), (0.1f64).to_bits() as u128);
36//! assert!(status.has(rucc_base::float::Status::INEXACT));
37//! ```
38//!
39//! The arithmetic is in `arith.rs`, on the same terms: every operation is correctly rounded, to
40//! nearest with ties to even, in integer operations that the host cannot get wrong.
41
42use crate::decimal::{Decimal, Fraction};
43
44mod arith;
45mod dec;
46
47pub use crate::float::arith::Integral;
48
49/// A floating point format.
50///
51/// Six of the ten are IEEE 754 binary encodings and the other four are not, which is why
52/// [`Format::is_ieee`] exists and why most of the questions below are answerable for six of them.
53/// The four are the double-double and the three decimal formats.
54/// The split is the same one the psABIs make, so this is the enum `rucc-abi` describes a target's
55/// scalar types with as well as the one [`Float`] carries.
56#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)]
57pub enum Format {
58    /// IEEE binary16, which C spells `_Float16`.
59    Half,
60    /// The brain float, an IEEE binary32 with the low sixteen bits of its significand cut off,
61    /// which C spells `__bf16`. It has the range of a `float` and less than half its precision.
62    BFloat16,
63    /// IEEE binary32, which C spells `float`.
64    Single,
65    /// IEEE binary64, which C spells `double`.
66    Double,
67    /// The x87 eighty bit format, which is `long double` on x86. It is the one format here that
68    /// stores the leading significand bit rather than leaving it implied.
69    X87Extended,
70    /// IEEE binary128, which C spells `_Float128`, and which is `long double` on AArch64 Linux,
71    /// on s390x and on RISC-V.
72    Quad,
73    /// IBM double-double, a pair of `double`s whose sum is the value, which is `long double` on
74    /// 64-bit PowerPC.
75    ///
76    /// Not an IEEE encoding and not a binary floating point format in IEEE's sense. It has no
77    /// exponent field of its own, no significand field of its own, and no single precision: the
78    /// gap between the two halves is whatever the value needs, so the number of significand bits
79    /// between the top of the first and the bottom of the second is a hundred and six for some
80    /// values and two thousand for others. `__LDBL_MANT_DIG__` says 106 because a macro has to
81    /// say something, and 106 is the figure everyone quotes, but it is the precision you get near
82    /// the top of the significand rather than a property of the format.
83    ///
84    /// [`Float`] does not represent one, per [`Format::is_ieee`].
85    DoubleDouble,
86    /// IEEE decimal32 in the binary integer decimal encoding, which C spells `_Decimal32`.
87    ///
88    /// The three decimal formats are IEEE 754 formats but not binary ones, and nothing about a
89    /// binary format's precision, exponent or significand fields means anything for them, so
90    /// [`Format::is_ieee`] answers false for them as it does for the double-double and [`Float`]
91    /// refuses them. What a decimal constant is made of is in [`crate::dfp`].
92    Decimal32,
93    /// IEEE decimal64 in the binary integer decimal encoding, which C spells `_Decimal64`.
94    Decimal64,
95    /// IEEE decimal128 in the binary integer decimal encoding, which C spells `_Decimal128`.
96    Decimal128,
97}
98
99/// What every IEEE binary question on a double-double or a decimal format fails with.
100///
101/// A function rather than a `panic!` in each arm, because the same sentence in five places drifts
102/// into five sentences, and because `panic!` in a `const fn` takes a literal and will not take a
103/// constant.
104const fn not_ieee() -> ! {
105    panic!(
106        "the double-double format is a pair of doubles and the decimal formats count in tens, so \
107         neither has a binary precision, exponent range or significand field to ask about"
108    )
109}
110
111impl Format {
112    /// The short name this format is written under, which is its width in bits for all of them
113    /// but the two whose width does not tell them apart from something else.
114    #[must_use]
115    pub const fn name(self) -> &'static str {
116        match self {
117            Format::Half => "f16",
118            Format::BFloat16 => "bf16",
119            Format::Single => "f32",
120            Format::Double => "f64",
121            Format::X87Extended => "f80",
122            Format::Quad => "f128",
123            Format::DoubleDouble => "ppc-f128",
124            Format::Decimal32 => "d32",
125            Format::Decimal64 => "d64",
126            Format::Decimal128 => "d128",
127        }
128    }
129
130    /// The format of that name, and [`None`] for a word that is not one.
131    #[must_use]
132    pub fn from_name(name: &str) -> Option<Self> {
133        Some(match name {
134            "f16" => Format::Half,
135            "bf16" => Format::BFloat16,
136            "f32" => Format::Single,
137            "f64" => Format::Double,
138            "f80" => Format::X87Extended,
139            "f128" => Format::Quad,
140            "ppc-f128" => Format::DoubleDouble,
141            "d32" => Format::Decimal32,
142            "d64" => Format::Decimal64,
143            "d128" => Format::Decimal128,
144            _ => return None,
145        })
146    }
147
148    /// Whether the format is an IEEE 754 binary encoding, which is every one of them but the
149    /// double-double and the three decimal ones.
150    ///
151    /// This is the guard on the rest of this type and on [`Float`]. A number in an IEEE encoding
152    /// is a sign, an exponent and one significand, which is what [`Float`] stores, so every
153    /// format that answers true here has a precision, an exponent range and a bit layout and can
154    /// be parsed, encoded and folded. The double-double is a pair, so it has none of those and
155    /// [`Float`] refuses it rather than answering with the nominal figures, which are close
156    /// enough to right to be believed and wrong often enough to matter.
157    #[must_use]
158    pub const fn is_ieee(self) -> bool {
159        !matches!(
160            self,
161            Format::DoubleDouble | Format::Decimal32 | Format::Decimal64 | Format::Decimal128
162        )
163    }
164
165    /// The decimal format this is, and [`None`] for every binary one.
166    ///
167    /// ```
168    /// use rucc_base::dfp::Width;
169    /// use rucc_base::float::Format;
170    ///
171    /// assert_eq!(Format::Decimal64.decimal(), Some(Width::D64));
172    /// assert_eq!(Format::Double.decimal(), None);
173    /// ```
174    #[must_use]
175    pub const fn decimal(self) -> Option<crate::dfp::Width> {
176        match self {
177            Format::Decimal32 => Some(crate::dfp::Width::D32),
178            Format::Decimal64 => Some(crate::dfp::Width::D64),
179            Format::Decimal128 => Some(crate::dfp::Width::D128),
180            _ => None,
181        }
182    }
183
184    /// The number of significand bits, counting the leading one whether it is stored or not.
185    ///
186    /// # Panics
187    ///
188    /// If the format is not an IEEE encoding, per [`Format::is_ieee`].
189    #[must_use]
190    pub const fn precision(self) -> u32 {
191        match self {
192            Format::Half => 11,
193            Format::BFloat16 => 8,
194            Format::Single => 24,
195            Format::Double => 53,
196            Format::X87Extended => 64,
197            Format::Quad => 113,
198            Format::DoubleDouble | Format::Decimal32 | Format::Decimal64 | Format::Decimal128 => {
199                not_ieee()
200            }
201        }
202    }
203
204    /// The exponent of the largest finite number, which is also the exponent bias.
205    ///
206    /// # Panics
207    ///
208    /// If the format is not an IEEE encoding, per [`Format::is_ieee`].
209    #[must_use]
210    pub const fn max_exponent(self) -> i32 {
211        match self {
212            Format::Half => 15,
213            Format::BFloat16 | Format::Single => 127,
214            Format::Double => 1023,
215            Format::X87Extended | Format::Quad => 16383,
216            Format::DoubleDouble | Format::Decimal32 | Format::Decimal64 | Format::Decimal128 => {
217                not_ieee()
218            }
219        }
220    }
221
222    /// The exponent of the smallest normal number.
223    ///
224    /// # Panics
225    ///
226    /// If the format is not an IEEE encoding, per [`Format::is_ieee`].
227    #[must_use]
228    pub const fn min_exponent(self) -> i32 {
229        1 - self.max_exponent()
230    }
231
232    /// The width of the encoding in bits, which for x87 is the eighty bits that matter and not
233    /// the ninety six or hundred and twenty eight an ABI pads them out to.
234    ///
235    /// Answered for every format, the double-double included, because a width is the one fact a
236    /// pair of doubles does have: it is the two of them and nothing else, so it is a hundred and
237    /// twenty eight bits the same way binary128 is.
238    #[must_use]
239    pub const fn width(self) -> u32 {
240        match self {
241            Format::Half | Format::BFloat16 => 16,
242            Format::Single | Format::Decimal32 => 32,
243            Format::Double | Format::Decimal64 => 64,
244            Format::X87Extended => 80,
245            Format::Quad | Format::DoubleDouble | Format::Decimal128 => 128,
246        }
247    }
248
249    /// Whether the leading significand bit is stored rather than implied.
250    ///
251    /// # Panics
252    ///
253    /// If the format is not an IEEE encoding, per [`Format::is_ieee`].
254    #[must_use]
255    pub const fn has_explicit_integer_bit(self) -> bool {
256        match self {
257            Format::X87Extended => true,
258            Format::Half | Format::BFloat16 | Format::Single | Format::Double | Format::Quad => {
259                false
260            }
261            Format::DoubleDouble | Format::Decimal32 | Format::Decimal64 | Format::Decimal128 => {
262                not_ieee()
263            }
264        }
265    }
266
267    /// The width of the exponent field.
268    const fn exponent_bits(self) -> u32 {
269        self.width() - self.significand_bits() - 1
270    }
271
272    /// The width of the stored significand field.
273    const fn significand_bits(self) -> u32 {
274        if self.has_explicit_integer_bit() { self.precision() } else { self.precision() - 1 }
275    }
276
277    /// A decimal exponent above which every number is too large for the format.
278    ///
279    /// The value is at least `10^(point - 1)`, so a point past this cannot be finite. It is
280    /// deliberately loose: it exists to stop the scaling loop from walking a million powers of
281    /// ten, not to decide anything.
282    const fn max_decimal_exponent(self) -> i32 {
283        (self.max_exponent() + 1) * 30103 / 100000 + 2
284    }
285
286    /// A decimal exponent below which every number rounds to zero.
287    const fn min_decimal_exponent(self) -> i32 {
288        (self.min_exponent() - self.precision() as i32) * 30103 / 100000 - 2
289    }
290}
291
292/// What a conversion had to do to the number to fit it in the format.
293///
294/// A bitmask, so that one conversion can report several. The names are IEEE 754's exceptions,
295/// which is what the diagnostics are ultimately about: GCC warns that a floating constant
296/// exceeds the range of its type, or that it was truncated to zero.
297#[derive(Debug, Clone, Copy, PartialEq, Eq, Default)]
298pub struct Status(u8);
299
300impl Status {
301    /// The value is exactly what was written.
302    pub const NONE: Status = Status(0);
303    /// The value had to be rounded, so it is not what was written.
304    pub const INEXACT: Status = Status(1);
305    /// The value is too large for the format and became an infinity.
306    pub const OVERFLOW: Status = Status(2);
307    /// The value is too small for the format and became a subnormal or a zero.
308    pub const UNDERFLOW: Status = Status(4);
309    /// The operation has no answer at all, such as an infinity minus an infinity.
310    pub const INVALID: Status = Status(8);
311    /// A number that is not zero was divided by one that is, so the answer is an infinity.
312    pub const DIVIDE_BY_ZERO: Status = Status(16);
313
314    /// Whether every flag in `other` is set here.
315    #[inline]
316    #[must_use]
317    pub const fn has(self, other: Status) -> bool {
318        self.0 & other.0 == other.0
319    }
320
321    /// This set with `other` added.
322    #[inline]
323    #[must_use]
324    pub const fn with(self, other: Status) -> Status {
325        Status(self.0 | other.0)
326    }
327
328    /// Whether nothing happened to the number.
329    #[inline]
330    #[must_use]
331    pub const fn is_none(self) -> bool {
332        self.0 == 0
333    }
334}
335
336/// Why a spelling is not a number.
337///
338/// The caller is expected to have checked the shape of the token already, so these are the
339/// cases a lexer cannot rule out rather than a full grammar.
340#[derive(Debug, Clone, Copy, PartialEq, Eq)]
341pub enum ParseError {
342    /// There is no digit anywhere in it.
343    NoDigits,
344    /// There is an exponent marker with no digits after it.
345    NoExponentDigits,
346    /// There is a character in it that a number does not have.
347    Invalid,
348}
349
350/// What kind of number this is.
351#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)]
352enum Category {
353    Zero,
354    Finite,
355    Infinite,
356    Nan,
357}
358
359/// A floating point number in a given format.
360///
361/// A finite value is `significand * 2^(exponent - precision + 1)`. A normal number has its
362/// leading significand bit set, a subnormal does not and has the format's minimum exponent.
363#[derive(Debug, Clone, Copy, PartialEq, Eq)]
364pub struct Float {
365    format: Format,
366    category: Category,
367    sign: bool,
368    exponent: i32,
369    significand: u128,
370}
371
372/// The format a [`Float`] is being built in, or a panic naming why it cannot be.
373///
374/// Every way of making a [`Float`] goes through here, so the one format this type does not
375/// represent is rejected where it is asked for rather than several steps later where the reason
376/// is no longer in view.
377const fn ieee(format: Format) -> Format {
378    if format.is_ieee() { format } else { not_ieee() }
379}
380
381/// The format a [`Float`] is being built in where a decimal one is as good as a binary one, which
382/// is the zeros, the infinities and the encoding. See `float/dec.rs` for what a decimal is here.
383const fn carried(format: Format) -> Format {
384    if format.is_ieee() || format.decimal().is_some() { format } else { not_ieee() }
385}
386
387impl Float {
388    /// A zero of the given sign.
389    ///
390    /// # Panics
391    ///
392    /// If the format is not an IEEE encoding, per [`Format::is_ieee`].
393    #[must_use]
394    pub const fn zero(format: Format, sign: bool) -> Float {
395        Float {
396            format: carried(format),
397            category: Category::Zero,
398            sign,
399            exponent: 0,
400            significand: 0,
401        }
402    }
403
404    /// An infinity of the given sign.
405    ///
406    /// # Panics
407    ///
408    /// If the format is not an IEEE encoding, per [`Format::is_ieee`].
409    #[must_use]
410    pub const fn infinity(format: Format, sign: bool) -> Float {
411        Float {
412            format: carried(format),
413            category: Category::Infinite,
414            sign,
415            exponent: 0,
416            significand: 0,
417        }
418    }
419
420    /// The smallest normal number of the given sign, which is the boundary `isnormal` asks
421    /// about.
422    ///
423    /// A normal number is one whose leading significand bit is set, so the smallest of them is
424    /// that bit alone at the format's lowest exponent. Every value below it is a subnormal or a
425    /// zero, which is why the question can be a comparison against this rather than a mask and a
426    /// shift over the exponent field.
427    ///
428    /// # Panics
429    ///
430    /// If the format is not an IEEE encoding, per [`Format::is_ieee`].
431    #[must_use]
432    pub const fn smallest_normal(format: Format, sign: bool) -> Float {
433        Float {
434            format: ieee(format),
435            category: Category::Finite,
436            sign,
437            exponent: format.min_exponent(),
438            significand: 1u128 << (format.precision() - 1),
439        }
440    }
441
442    /// A nan with a payload, which is the one thing `__builtin_nan` and its family can spell
443    /// that nothing else in C can.
444    ///
445    /// The payload is the low bits of the significand and is cut to the bits there are below the
446    /// quiet bit, which is what gcc does with one that does not fit. A quiet nan is the payload
447    /// with that bit set. A signalling one is the payload without it, and a signalling nan with
448    /// nothing in it is an infinity rather than a nan, so a payload of zero becomes the highest
449    /// bit that is left, which is the value gcc gives `__builtin_nans("")`.
450    ///
451    /// A decimal nan takes no payload here and is the one gcc writes for `__builtin_nand32("")`,
452    /// with the exponent saying whether it signals the way [`Float::to_bits`] reads it.
453    ///
454    /// # Panics
455    ///
456    /// If the format is neither an IEEE encoding, per [`Format::is_ieee`], nor a decimal one.
457    #[must_use]
458    pub const fn nan_with(format: Format, sign: bool, quiet: bool, payload: u128) -> Float {
459        if format.decimal().is_some() {
460            let exponent = if quiet { 0 } else { 1 };
461            return Float { format, category: Category::Nan, sign, exponent, significand: 0 };
462        }
463        let format = ieee(format);
464        let mut significand = payload & (Float::quiet_bit(format) - 1);
465        if quiet {
466            significand |= Float::quiet_bit(format);
467        } else if significand == 0 {
468            significand = Float::quiet_bit(format) >> 1;
469        }
470        Float {
471            format,
472            category: Category::Nan,
473            sign,
474            exponent: 0,
475            significand: significand | Float::leading_bit(format),
476        }
477    }
478
479    /// The bit that tells a quiet nan from a signalling one, which is the highest bit of the
480    /// stored fraction in every format IEEE 754 defines.
481    const fn quiet_bit(format: Format) -> u128 {
482        1u128 << (format.precision() - 2)
483    }
484
485    /// The leading significand bit, in the one format that stores it rather than implying it. It
486    /// is set in every value of that format that is not a zero, a nan and an infinity included.
487    const fn leading_bit(format: Format) -> u128 {
488        if format.has_explicit_integer_bit() { 1u128 << (format.precision() - 1) } else { 0 }
489    }
490
491    /// The format this number is in.
492    #[must_use]
493    pub const fn format(self) -> Format {
494        self.format
495    }
496
497    /// Whether the number is negative, which a zero can be.
498    #[must_use]
499    pub const fn is_negative(self) -> bool {
500        self.sign
501    }
502
503    /// Whether the number is a zero.
504    #[must_use]
505    pub const fn is_zero(self) -> bool {
506        matches!(self.category, Category::Zero)
507    }
508
509    /// Whether the number is an infinity.
510    #[must_use]
511    pub const fn is_infinite(self) -> bool {
512        matches!(self.category, Category::Infinite)
513    }
514
515    /// Whether the number is finite, which a zero is and a nan is not.
516    #[must_use]
517    pub const fn is_finite(self) -> bool {
518        matches!(self.category, Category::Zero | Category::Finite)
519    }
520
521    /// Whether the number is normal, which is finite with the leading significand bit set.
522    ///
523    /// A zero is not, a subnormal is not, and an infinity and a nan are not, which is the five
524    /// way split `fpclassify` asks about with the subnormal case being whatever is left.
525    #[must_use]
526    pub fn is_normal(self) -> bool {
527        if let Some(width) = self.format.decimal() {
528            return self.decimal_is_normal(width);
529        }
530        matches!(self.category, Category::Finite)
531            && self.significand >> (self.format.precision() - 1) != 0
532    }
533
534    /// Converts a decimal or hexadecimal spelling into the nearest number in `format`, rounding
535    /// to nearest with ties to even.
536    ///
537    /// The spelling is the number alone: no suffix, because the suffix is what chose the
538    /// format, and no infinity or nan, because C has no spelling for those. A sign is accepted
539    /// even though a C constant never has one, since the value the constant evaluator folds
540    /// does. C23 digit separators are stripped here.
541    ///
542    /// # Errors
543    ///
544    /// [`ParseError`], for a spelling that is not a number at all.
545    ///
546    /// # Panics
547    ///
548    /// If the format is not an IEEE encoding, per [`Format::is_ieee`]. A bad format is the
549    /// caller's bug and a bad spelling is the program's, which is why one is a panic and the
550    /// other is an error.
551    pub fn parse(text: &str, format: Format) -> Result<(Float, Status), ParseError> {
552        if let Some(width) = format.decimal() {
553            let (bits, status) = crate::dfp::parse(text, width)?;
554            return Ok((Float::decimal_from_bits(format, width, bits), status));
555        }
556        let format = ieee(format);
557        let bytes = text.as_bytes();
558        let (sign, rest) = match bytes.first() {
559            Some(b'-') => (true, &bytes[1..]),
560            Some(b'+') => (false, &bytes[1..]),
561            _ => (false, bytes),
562        };
563        if rest.len() > 1 && rest[0] == b'0' && rest[1] | 32 == b'x' {
564            hexadecimal(&rest[2..], sign, format)
565        } else {
566            decimal(rest, sign, format)
567        }
568    }
569
570    /// The bits of the encoding, in the low [`Format::width`] bits.
571    ///
572    /// The x87 format keeps its leading significand bit, so its eightieth bit is the sign and
573    /// its sixty fourth is the one every other format leaves implied.
574    #[must_use]
575    pub fn to_bits(self) -> u128 {
576        let format = self.format;
577        if let Some(width) = format.decimal() {
578            return self.decimal_bits(width);
579        }
580        let significand_mask = (1u128 << format.significand_bits()) - 1;
581        let (exponent_field, significand_field) = match self.category {
582            Category::Zero => (0, 0),
583            Category::Infinite => (
584                (1u128 << format.exponent_bits()) - 1,
585                if format.has_explicit_integer_bit() {
586                    1u128 << (format.precision() - 1)
587                } else {
588                    0
589                },
590            ),
591            // The significand of a nan is the whole of what it is, since the quiet bit and the
592            // payload are both in it and the exponent is the same for every nan there is.
593            Category::Nan => ((1u128 << format.exponent_bits()) - 1, self.significand),
594            Category::Finite => {
595                let subnormal = self.significand >> (format.precision() - 1) == 0;
596                let field =
597                    if subnormal { 0 } else { (self.exponent + format.max_exponent()) as u128 };
598                (field, self.significand & significand_mask)
599            }
600        };
601        let sign = u128::from(self.sign) << (format.width() - 1);
602        sign | (exponent_field << format.significand_bits()) | significand_field
603    }
604
605    /// Reads a number back out of its encoding, which is what makes [`Float::to_bits`] testable
606    /// and what a constant folded in the IR is stored as.
607    ///
608    /// A nan comes back with the quiet bit and the payload it went in with, so a value that came
609    /// from `__builtin_nan` survives being written down and read back, which is the round trip
610    /// every constant in the IR takes.
611    ///
612    /// # Panics
613    ///
614    /// If the format is not an IEEE encoding, per [`Format::is_ieee`].
615    #[must_use]
616    pub fn from_bits(format: Format, bits: u128) -> Float {
617        if let Some(width) = format.decimal() {
618            return Float::decimal_from_bits(format, width, bits);
619        }
620        let format = ieee(format);
621        let significand_bits = format.significand_bits();
622        let sign = (bits >> (format.width() - 1)) & 1 == 1;
623        let exponent_field =
624            ((bits >> significand_bits) & ((1u128 << format.exponent_bits()) - 1)) as i32;
625        let stored = bits & ((1u128 << significand_bits) - 1);
626        if exponent_field == (1 << format.exponent_bits()) - 1 {
627            // The fraction is what tells an infinity from a nan, and in the x87 format the bit
628            // above the fraction is stored rather than implied and is set in both.
629            let fraction = stored & ((1u128 << (format.precision() - 1)) - 1);
630            if fraction == 0 {
631                return Float::infinity(format, sign);
632            }
633            return Float {
634                format,
635                category: Category::Nan,
636                sign,
637                exponent: 0,
638                significand: stored,
639            };
640        }
641        let implicit = if format.has_explicit_integer_bit() || exponent_field == 0 {
642            0
643        } else {
644            1u128 << (format.precision() - 1)
645        };
646        let significand = stored | implicit;
647        if significand == 0 {
648            return Float::zero(format, sign);
649        }
650        let exponent = if exponent_field == 0 {
651            format.min_exponent()
652        } else {
653            exponent_field - format.max_exponent()
654        };
655        Float { format, category: Category::Finite, sign, exponent, significand }
656    }
657
658    /// A hexadecimal spelling that [`Float::parse`] turns back into exactly this number.
659    ///
660    /// Hexadecimal rather than decimal, because a hexadecimal constant is exact by construction
661    /// and a decimal one is not: printing a number in decimal so that it reads back unchanged
662    /// needs a shortest-round-trip algorithm, and printing it in decimal without one silently
663    /// changes the program. A printer that changes a constant is worse than a printer whose
664    /// output is unfamiliar, so this is `0x1p+0` where a reader would rather see `1.0`.
665    ///
666    /// The significand is written as an integer and the exponent scales it, so the spelling is
667    /// `significand * 2^exponent` with no leading digit to argue about. Trailing zero digits are
668    /// taken off, which is what makes a round number short.
669    ///
670    /// An infinity has no spelling in C at all. What comes back for one is an exponent past the
671    /// top of the format, which converts back to an infinity with the overflow that a constant
672    /// only ever became an infinity by. A nan is spelled `nan` and does not read back, since
673    /// there is no exponent that gives one and no constant that is one.
674    #[must_use]
675    pub fn to_hex(self) -> String {
676        if self.format.decimal().is_some() {
677            return self.decimal_spelling();
678        }
679        let sign = if self.sign { "-" } else { "" };
680        match self.category {
681            Category::Nan => format!("{sign}nan"),
682            Category::Infinite => format!("{sign}0x1p+{}", self.format.max_exponent() + 1),
683            Category::Zero => format!("{sign}0x0p+0"),
684            Category::Finite => {
685                let mut significand = self.significand;
686                let mut exponent = self.exponent - (self.format.precision() as i32 - 1);
687                while significand & 0xf == 0 {
688                    significand >>= 4;
689                    exponent += 4;
690                }
691                format!("{sign}0x{significand:x}p{exponent:+}")
692            }
693        }
694    }
695}
696
697/// Converts a decimal spelling.
698fn decimal(bytes: &[u8], sign: bool, format: Format) -> Result<(Float, Status), ParseError> {
699    let mut digits = Vec::new();
700    let mut integer_digits = 0i32;
701    let mut seen_point = false;
702    let mut seen_digit = false;
703    let mut index = 0;
704    while index < bytes.len() {
705        match bytes[index] {
706            byte @ b'0'..=b'9' => {
707                digits.push(byte - b'0');
708                if !seen_point {
709                    integer_digits += 1;
710                }
711                seen_digit = true;
712            }
713            b'\'' => {}
714            b'.' if !seen_point => seen_point = true,
715            b'e' | b'E' => break,
716            _ => return Err(ParseError::Invalid),
717        }
718        index += 1;
719    }
720    if !seen_digit {
721        return Err(ParseError::NoDigits);
722    }
723    let mut point = integer_digits;
724    if index < bytes.len() {
725        point = point.saturating_add(exponent_of(&bytes[index + 1..])?);
726    }
727    Ok(convert(Decimal::new(digits, point), sign, format))
728}
729
730/// Converts a hexadecimal spelling, which is exact until the one rounding at the end.
731fn hexadecimal(bytes: &[u8], sign: bool, format: Format) -> Result<(Float, Status), ParseError> {
732    let mut significand: u128 = 0;
733    let mut exponent = 0i32;
734    let mut sticky = false;
735    let mut seen_point = false;
736    let mut seen_digit = false;
737    let mut index = 0;
738    while index < bytes.len() {
739        let byte = bytes[index];
740        let digit = match byte {
741            b'0'..=b'9' => byte - b'0',
742            b'a'..=b'f' => byte - b'a' + 10,
743            b'A'..=b'F' => byte - b'A' + 10,
744            b'\'' => {
745                index += 1;
746                continue;
747            }
748            b'.' if !seen_point => {
749                seen_point = true;
750                index += 1;
751                continue;
752            }
753            b'p' | b'P' => break,
754            _ => return Err(ParseError::Invalid),
755        };
756        seen_digit = true;
757        if significand.leading_zeros() >= 4 {
758            significand = (significand << 4) | u128::from(digit);
759            if seen_point {
760                exponent -= 4;
761            }
762        } else {
763            // Past a hundred and twenty eight bits the digits cannot change the value, only
764            // whether it is exactly halfway, which is what the sticky bit is for.
765            sticky |= digit != 0;
766            if !seen_point {
767                exponent += 4;
768            }
769        }
770        index += 1;
771    }
772    if !seen_digit {
773        return Err(ParseError::NoDigits);
774    }
775    if index < bytes.len() {
776        exponent = exponent.saturating_add(exponent_of(&bytes[index + 1..])?);
777    }
778    Ok(round(significand, exponent, sticky, sign, format))
779}
780
781/// Reads the digits of an exponent, which may be signed.
782fn exponent_of(bytes: &[u8]) -> Result<i32, ParseError> {
783    let (negative, digits) = match bytes.first() {
784        Some(b'-') => (true, &bytes[1..]),
785        Some(b'+') => (false, &bytes[1..]),
786        _ => (false, bytes),
787    };
788    if digits.is_empty() {
789        return Err(ParseError::NoExponentDigits);
790    }
791    let mut value = 0i32;
792    for &byte in digits {
793        if byte == b'\'' {
794            continue;
795        }
796        if !byte.is_ascii_digit() {
797            return Err(ParseError::Invalid);
798        }
799        // An exponent far past the format's range is the same as one at the edge of it, so it
800        // saturates rather than overflowing.
801        value = value.saturating_mul(10).saturating_add(i32::from(byte - b'0'));
802    }
803    Ok(if negative { -value } else { value })
804}
805
806/// Scales an exact decimal down to the format's significand and rounds it.
807fn convert(mut value: Decimal, sign: bool, format: Format) -> (Float, Status) {
808    if value.is_zero() {
809        return (Float::zero(format, sign), Status::NONE);
810    }
811    if value.point() > format.max_decimal_exponent() {
812        return (Float::infinity(format, sign), Status::OVERFLOW.with(Status::INEXACT));
813    }
814    if value.point() < format.min_decimal_exponent() {
815        return (Float::zero(format, sign), Status::UNDERFLOW.with(Status::INEXACT));
816    }
817
818    // Scale until the value is in `[1, 2)`, counting the powers of two taken out of it. Each
819    // step is an underestimate of the distance left, so no step overshoots and the loop always
820    // moves, which is what stops it oscillating.
821    let mut exponent = 0i32;
822    loop {
823        let point = value.point();
824        if point > 1 || (point == 1 && value.first_digit() >= 2) {
825            let step = binary_digits(point - 1).clamp(1, 60);
826            value.shift(-step);
827            exponent += step;
828        } else if point < 1 {
829            let step = (1 + binary_digits(-point)).clamp(1, 60);
830            value.shift(step);
831            exponent -= step;
832        } else {
833            break;
834        }
835    }
836
837    // The significand is the value scaled by this many powers of two, clamped so that a number
838    // below the smallest normal loses precision instead of exponent.
839    let precision = format.precision() as i32;
840    let scale = (exponent - precision + 1).max(format.min_exponent() - precision + 1);
841    value.shift(exponent - scale);
842    let (integer, fraction) = value.round_to_u128();
843    let rounded = match fraction {
844        Fraction::Zero | Fraction::BelowHalf => integer,
845        Fraction::Half => integer + (integer & 1),
846        Fraction::AboveHalf => integer + 1,
847    };
848    finish(rounded, scale, fraction != Fraction::Zero, sign, format)
849}
850
851/// Roughly how many binary digits a decimal one of this many digits has, never overestimating.
852const fn binary_digits(decimal: i32) -> i32 {
853    decimal * 33219 / 10000
854}
855
856/// Rounds `significand * 2^exponent` into the format, with `sticky` saying that something
857/// nonzero was already dropped below it.
858fn round(
859    significand: u128,
860    exponent: i32,
861    sticky: bool,
862    sign: bool,
863    format: Format,
864) -> (Float, Status) {
865    if significand == 0 {
866        return (Float::zero(format, sign), Status::NONE);
867    }
868    let precision = format.precision() as i32;
869    let leading = (128 - significand.leading_zeros()) as i32;
870    let scale = (exponent + leading - precision).max(format.min_exponent() - precision + 1);
871    let mut sticky = sticky;
872    let (integer, half) = if scale <= exponent {
873        (significand << (exponent - scale), false)
874    } else {
875        let drop = (scale - exponent) as u32;
876        if drop >= 128 {
877            sticky = true;
878            (0, false)
879        } else {
880            let half = (significand >> (drop - 1)) & 1 == 1;
881            sticky |= drop > 1 && significand & ((1u128 << (drop - 1)) - 1) != 0;
882            (significand >> drop, half)
883        }
884    };
885    let rounded = if half && (sticky || integer & 1 == 1) { integer + 1 } else { integer };
886    finish(rounded, scale, half || sticky, sign, format)
887}
888
889/// Turns a rounded significand and the power of two it is scaled by into a number, handling the
890/// carry out of the significand and the two ends of the format's range.
891fn finish(
892    significand: u128,
893    scale: i32,
894    inexact: bool,
895    sign: bool,
896    format: Format,
897) -> (Float, Status) {
898    let precision = format.precision();
899    let mut significand = significand;
900    let mut scale = scale;
901    if significand >> precision != 0 {
902        // Rounding up carried out of the top bit, which only ever gives a power of two.
903        significand >>= 1;
904        scale += 1;
905    }
906    let mut status = if inexact { Status::INEXACT } else { Status::NONE };
907    if significand == 0 {
908        return (Float::zero(format, sign), status.with(Status::UNDERFLOW));
909    }
910    let exponent = scale + precision as i32 - 1;
911    if exponent > format.max_exponent() {
912        return (
913            Float::infinity(format, sign),
914            status.with(Status::OVERFLOW).with(Status::INEXACT),
915        );
916    }
917    let normal = significand >> (precision - 1) != 0;
918    if !normal && inexact {
919        status = status.with(Status::UNDERFLOW);
920    }
921    let exponent = if normal { exponent } else { format.min_exponent() };
922    (Float { format, category: Category::Finite, sign, exponent, significand }, status)
923}
924
925#[cfg(test)]
926mod tests {
927    use super::*;
928
929    /// The bits a `double` conversion gives, next to what Rust's own parser gives.
930    fn double(text: &str) -> u128 {
931        Float::parse(text, Format::Double).expect("a number").0.to_bits()
932    }
933
934    /// The bits a `float` conversion gives.
935    fn single(text: &str) -> u128 {
936        Float::parse(text, Format::Single).expect("a number").0.to_bits()
937    }
938
939    #[test]
940    fn the_ordinary_numbers_land_where_the_host_would_put_them() {
941        for text in ["0", "1", "2", "0.5", "1.5", "3.14159", "2.718281828459045", "100", "1e10"] {
942            let host = text.parse::<f64>().expect("a number Rust reads too");
943            assert_eq!(double(text), u128::from(host.to_bits()), "{text}");
944        }
945    }
946
947    #[test]
948    fn a_number_that_needs_the_last_bit_rounded_gets_it_right() {
949        // Every one of these is a literal a naive `mantissa * 10^exponent` gets wrong, and the
950        // last is the longest one a `double` conversion has to read to round correctly.
951        let hard = [
952            "0.1",
953            "0.3",
954            "2.2250738585072011e-308",
955            "2.2250738585072014e-308",
956            "1.7976931348623157e308",
957            "4.9406564584124654e-324",
958            "5e-324",
959            "8.98846567431158e307",
960            "9007199254740993",
961            "123456789012345678901234567890",
962            "1.000000000000000000000000000000000000000000000000000000000000000001",
963            "7.8459735791271921e65",
964            "3.518437208883201171875e13",
965            "0.500000000000000166533453693773481063544750213623046875",
966        ];
967        for text in hard {
968            let host = text.parse::<f64>().expect("a number Rust reads too");
969            assert_eq!(double(text), u128::from(host.to_bits()), "{text}");
970        }
971    }
972
973    #[test]
974    fn the_number_that_takes_seven_hundred_and_sixty_seven_digits() {
975        // The exact decimal of a `double` halfway case. A conversion that truncates its input
976        // rounds this one the wrong way, which is the bug this buffer size exists to avoid.
977        let text = concat!(
978            "2.47032822920623272088284396434110686182529901307162382",
979            "35378852574870103599108683372845652890455735483022221802",
980            "58573249056416711547735232764105795166208503595426876755",
981            "62317084535693494535245273750735013572761315046354601316",
982            "12127849863326369238975694273040488011871029093711789936",
983            "42245692702737764465109076580131048946378905599180391359",
984            "70011386455512221706120629864144453927884519445934871524",
985            "63344875888932891414823975864211858166195965106373837732",
986            "34435703331457550505022232309998195892058070506176382679",
987            "16323484472119097902806154870514036458498974142754747141",
988            "39683784321102080606305920253373777969877864922227306716",
989            "01324339457879181214233820577228206278891620001855078759",
990            "16278352090142077553206262229158550205643778244387017277",
991            "94459649305087139089301871550805125768938177360937844105",
992            "63661045147381814281647890691181239104545396303476425117",
993            "7562185422741845851144691421326303120484712594187004993e-324"
994        );
995        let host = text.parse::<f64>().expect("a number Rust reads too");
996        assert_eq!(double(text), u128::from(host.to_bits()));
997    }
998
999    #[test]
1000    fn a_sweep_of_random_numbers_agrees_with_rust_in_every_bit() {
1001        // A conversion that is wrong in the last place is wrong on a small fraction of inputs,
1002        // so this is a sweep rather than a handful. The generator is a fixed sequence, so a
1003        // failure names the same number on every machine.
1004        let mut state = 0x2545_f491_4f6c_dd1du64;
1005        for _ in 0..4000 {
1006            state = state.wrapping_mul(6364136223846793005).wrapping_add(1442695040888963407);
1007            let digits = state >> 11;
1008            let exponent = (state % 600) as i32 - 300;
1009            let text = format!("{digits}e{exponent}");
1010            let host = text.parse::<f64>().expect("a number Rust reads too");
1011            assert_eq!(double(&text), u128::from(host.to_bits()), "{text}");
1012            let host = text.parse::<f32>().expect("a number Rust reads too");
1013            assert_eq!(single(&text), u128::from(host.to_bits()), "{text} as a float");
1014        }
1015    }
1016
1017    #[test]
1018    fn the_ends_of_the_range_are_an_infinity_and_a_zero() {
1019        let (value, status) = Float::parse("1e400", Format::Double).expect("a number");
1020        assert!(value.is_infinite() && status.has(Status::OVERFLOW));
1021        let (value, status) = Float::parse("1e-400", Format::Double).expect("a number");
1022        assert!(value.is_zero() && status.has(Status::UNDERFLOW) && status.has(Status::INEXACT));
1023        // The largest `double` is finite and the next number up is not.
1024        let (value, status) = Float::parse("1.7976931348623157e308", Format::Double).expect("one");
1025        assert!(value.is_finite() && !status.has(Status::OVERFLOW));
1026        let (value, _) = Float::parse("1.8e308", Format::Double).expect("a number");
1027        assert!(value.is_infinite());
1028        // Half the smallest subnormal rounds to zero, and just over half rounds up to it.
1029        assert_eq!(double("2.4e-324"), u128::from((0f64).to_bits()));
1030        assert_eq!(double("2.5e-324"), 1);
1031    }
1032
1033    #[test]
1034    fn a_number_that_is_exactly_what_was_written_says_so() {
1035        assert!(Float::parse("1", Format::Double).expect("a number").1.is_none());
1036        assert!(Float::parse("0.5", Format::Double).expect("a number").1.is_none());
1037        assert!(Float::parse("0.1", Format::Double).expect("a number").1.has(Status::INEXACT));
1038        // A number small enough to lose bits is inexact and underflowed, both.
1039        let (_, status) = Float::parse("1e-320", Format::Double).expect("a number");
1040        assert!(status.has(Status::INEXACT) && status.has(Status::UNDERFLOW));
1041    }
1042
1043    #[test]
1044    fn a_hexadecimal_constant_is_exact_and_needs_no_scaling() {
1045        assert_eq!(double("0x1p0"), u128::from((1f64).to_bits()));
1046        assert_eq!(double("0x1.8p1"), u128::from((3f64).to_bits()));
1047        assert_eq!(double("0x1p-1074"), 1);
1048        assert_eq!(double("0xa.bp-4"), u128::from((0.66796875f64).to_bits()));
1049        assert_eq!(double("0X1.FFFFFFFFFFFFFP+1023"), u128::from(f64::MAX.to_bits()));
1050        assert!(Float::parse("0x1p0", Format::Double).expect("a number").1.is_none());
1051        // Seventeen hexadecimal digits is more than a `double` has, so this one rounds.
1052        let (_, status) = Float::parse("0x1.00000000000008p0", Format::Double).expect("a number");
1053        assert!(status.has(Status::INEXACT));
1054        assert_eq!(double("0x1.00000000000008p0"), u128::from((1f64).to_bits()));
1055        assert_eq!(double("0x1.00000000000018p0"), u128::from((1f64).to_bits() + 2));
1056    }
1057
1058    #[test]
1059    fn digit_separators_are_not_part_of_the_number() {
1060        assert_eq!(double("1'000.000'1"), double("1000.0001"));
1061        assert_eq!(double("0x1'0p0"), double("16.0"));
1062        assert_eq!(double("1e1'0"), double("1e10"));
1063    }
1064
1065    #[test]
1066    fn a_spelling_that_is_not_a_number_says_which_way_it_is_wrong() {
1067        assert_eq!(Float::parse("", Format::Double), Err(ParseError::NoDigits));
1068        assert_eq!(Float::parse(".", Format::Double), Err(ParseError::NoDigits));
1069        assert_eq!(Float::parse("1e", Format::Double), Err(ParseError::NoExponentDigits));
1070        assert_eq!(Float::parse("1e+", Format::Double), Err(ParseError::NoExponentDigits));
1071        assert_eq!(Float::parse("0x1p", Format::Double), Err(ParseError::NoExponentDigits));
1072        assert_eq!(Float::parse("0xp1", Format::Double), Err(ParseError::NoDigits));
1073        assert_eq!(Float::parse("1x0", Format::Double), Err(ParseError::Invalid));
1074    }
1075
1076    #[test]
1077    fn a_sign_is_accepted_although_a_c_constant_never_has_one() {
1078        let (value, _) = Float::parse("-1.5", Format::Double).expect("a number");
1079        assert!(value.is_negative());
1080        assert_eq!(value.to_bits(), u128::from((-1.5f64).to_bits()));
1081        let (value, _) = Float::parse("-0.0", Format::Double).expect("a number");
1082        assert!(value.is_zero() && value.is_negative());
1083        assert_eq!(value.to_bits(), u128::from((-0.0f64).to_bits()));
1084    }
1085
1086    #[test]
1087    fn every_format_says_how_wide_its_fields_are() {
1088        for format in [
1089            Format::Half,
1090            Format::BFloat16,
1091            Format::Single,
1092            Format::Double,
1093            Format::X87Extended,
1094            Format::Quad,
1095        ] {
1096            assert_eq!(
1097                format.exponent_bits() + format.significand_bits() + 1,
1098                format.width(),
1099                "{format:?}"
1100            );
1101            assert_eq!(format.min_exponent(), 1 - format.max_exponent());
1102        }
1103        assert_eq!(Format::Half.exponent_bits(), 5);
1104        assert_eq!(Format::BFloat16.exponent_bits(), 8);
1105        assert_eq!(Format::Single.exponent_bits(), 8);
1106        assert_eq!(Format::Double.exponent_bits(), 11);
1107        assert_eq!(Format::X87Extended.exponent_bits(), 15);
1108        assert_eq!(Format::Quad.exponent_bits(), 15);
1109    }
1110
1111    #[test]
1112    fn a_number_survives_a_trip_through_its_encoding() {
1113        for format in [
1114            Format::Half,
1115            Format::BFloat16,
1116            Format::Single,
1117            Format::Double,
1118            Format::X87Extended,
1119            Format::Quad,
1120        ] {
1121            for text in ["0", "-0", "1", "-1.5", "3.14159", "1e-5", "65504", "0x1p-20"] {
1122                let (value, _) = Float::parse(text, format).expect("a number");
1123                let bits = value.to_bits();
1124                assert_eq!(Float::from_bits(format, bits).to_bits(), bits, "{text} in {format:?}");
1125            }
1126            assert_eq!(
1127                Float::from_bits(format, Float::infinity(format, false).to_bits()).to_bits(),
1128                Float::infinity(format, false).to_bits()
1129            );
1130        }
1131    }
1132
1133    #[test]
1134    fn a_hexadecimal_spelling_reads_back_as_the_number_it_came_from() {
1135        for format in [
1136            Format::Half,
1137            Format::BFloat16,
1138            Format::Single,
1139            Format::Double,
1140            Format::X87Extended,
1141            Format::Quad,
1142        ] {
1143            for text in [
1144                "0", "-0", "1", "-1", "0.5", "-1.5", "3.14159", "1e-5", "0x1p-20", "0.1", "255",
1145                "1e30",
1146            ] {
1147                let (value, _) = Float::parse(text, format).expect("a number");
1148                let spelling = value.to_hex();
1149                let (again, status) = Float::parse(&spelling, format).expect("a number");
1150                assert_eq!(again.to_bits(), value.to_bits(), "{text} as {spelling} in {format:?}");
1151                // Exact, except where the number was already an infinity, which reading the
1152                // spelling back has to overflow into rather than land on.
1153                let rounded = status.has(Status::INEXACT) || status.has(Status::OVERFLOW);
1154                assert_eq!(rounded, !value.is_finite(), "{spelling} in {format:?}");
1155            }
1156            // A subnormal, which has leading zeros where a normal number has its implied one.
1157            let tiny = Float::from_bits(format, 1);
1158            let (again, _) = Float::parse(&tiny.to_hex(), format).expect("a number");
1159            assert_eq!(again.to_bits(), tiny.to_bits(), "the smallest subnormal in {format:?}");
1160            // An infinity, which C cannot spell and which comes back by overflowing again.
1161            let huge = Float::infinity(format, true);
1162            let (again, status) = Float::parse(&huge.to_hex(), format).expect("a number");
1163            assert!(again.is_infinite() && again.is_negative(), "{format:?}");
1164            assert!(status.has(Status::OVERFLOW));
1165        }
1166    }
1167
1168    #[test]
1169    fn a_round_number_gets_a_short_spelling() {
1170        let hex = |text: &str| Float::parse(text, Format::Double).expect("a number").0.to_hex();
1171        assert_eq!(hex("1"), "0x1p+0");
1172        assert_eq!(hex("-1"), "-0x1p+0");
1173        assert_eq!(hex("0"), "0x0p+0");
1174        assert_eq!(hex("-0"), "-0x0p+0");
1175        assert_eq!(hex("2"), "0x1p+1");
1176        assert_eq!(hex("0.5"), "0x1p-1");
1177        assert_eq!(hex("0.1"), "0x1999999999999ap-56");
1178    }
1179
1180    #[test]
1181    fn the_narrow_formats_round_where_they_are_supposed_to() {
1182        // `_Float16` has eleven bits, so its largest finite number is 65504 and the next power
1183        // of two is an infinity. `__bf16` has eight, so it loses a `float`'s low bits and keeps
1184        // its range, which is the whole point of the format.
1185        let (value, status) = Float::parse("65504", Format::Half).expect("a number");
1186        assert!(value.is_finite() && status.is_none());
1187        assert_eq!(value.to_bits(), 0x7bff);
1188        let (value, _) = Float::parse("65536", Format::Half).expect("a number");
1189        assert!(value.is_infinite());
1190        assert_eq!(Float::parse("1", Format::Half).expect("one").0.to_bits(), 0x3c00);
1191        assert_eq!(Float::parse("1", Format::BFloat16).expect("one").0.to_bits(), 0x3f80);
1192        assert_eq!(Float::parse("1e30", Format::BFloat16).expect("big").0.to_bits(), 0x714a);
1193        // The smallest `_Float16` subnormal, and half of it.
1194        assert_eq!(Float::parse("0x1p-24", Format::Half).expect("tiny").0.to_bits(), 1);
1195        assert!(Float::parse("0x1p-26", Format::Half).expect("tinier").0.is_zero());
1196    }
1197
1198    #[test]
1199    fn the_x87_format_stores_the_bit_the_others_leave_implied() {
1200        // 1.0 is 0x3fff8000000000000000: the exponent field, then a significand whose top bit
1201        // is stored rather than implied. Every other format here would have zeros there.
1202        let one = Float::parse("1", Format::X87Extended).expect("one").0;
1203        assert_eq!(one.to_bits(), 0x3fff_8000_0000_0000_0000);
1204        assert_eq!(
1205            Float::parse("2", Format::X87Extended).expect("two").0.to_bits(),
1206            0x4000_8000_0000_0000_0000
1207        );
1208        // Sixty four bits of precision, so this is exact where a `double` would round it.
1209        let (value, status) = Float::parse("9007199254740993", Format::X87Extended).expect("one");
1210        assert!(status.is_none());
1211        assert_eq!(value.to_bits(), 0x4034_8000_0000_0000_0400);
1212        // Measured, by compiling the constant with gcc 13.3 on x86-64 and reading the ten
1213        // bytes back out of the program rather than trusting a table.
1214        assert_eq!(
1215            Float::parse("0.1", Format::X87Extended).expect("a tenth").0.to_bits(),
1216            0x3ffb_cccc_cccc_cccc_cccd
1217        );
1218        // A subnormal four thousand powers of ten down, which is three of the smallest number
1219        // the format has. gcc puts the same three there.
1220        assert_eq!(Float::parse("1e-4950", Format::X87Extended).expect("tiny").0.to_bits(), 3);
1221    }
1222
1223    #[test]
1224    fn the_quad_format_has_a_hundred_and_thirteen_bits_of_it() {
1225        assert_eq!(
1226            Float::parse("1", Format::Quad).expect("one").0.to_bits(),
1227            0x3fff_0000_0000_0000_0000_0000_0000_0000
1228        );
1229        // 0.1 in binary128, which is the same digits a `double` gets and then sixty more bits.
1230        assert_eq!(
1231            Float::parse("0.1", Format::Quad).expect("a tenth").0.to_bits(),
1232            0x3ffb_9999_9999_9999_9999_9999_9999_999a
1233        );
1234        // Also measured against gcc, through `__float128`.
1235        assert_eq!(
1236            Float::parse("3.14159", Format::Quad).expect("pi, roughly").0.to_bits(),
1237            0x4000_921f_9f01_b866_e43a_a79b_badc_0981
1238        );
1239        let (value, status) = Float::parse("1e5000", Format::Quad).expect("a number");
1240        assert!(value.is_infinite() && status.has(Status::OVERFLOW));
1241        let (value, _) = Float::parse("1e-5000", Format::Quad).expect("a number");
1242        assert!(value.is_zero());
1243    }
1244
1245    /// Every number here is what gcc 16 puts in the object for the `__builtin_nan` that spells
1246    /// it, read back out of the object rather than reasoned about.
1247    #[test]
1248    fn a_nan_with_a_payload_has_the_bits_gcc_gives_it() {
1249        let double = |quiet, payload| Float::nan_with(Format::Double, false, quiet, payload);
1250        assert_eq!(double(true, 0).to_bits(), 0x7ff8_0000_0000_0000, "__builtin_nan(\"\")");
1251        assert_eq!(double(true, 1).to_bits(), 0x7ff8_0000_0000_0001, "__builtin_nan(\"0x1\")");
1252        assert_eq!(double(true, 8).to_bits(), 0x7ff8_0000_0000_0008, "__builtin_nan(\"010\")");
1253        // A signalling nan with nothing in it would be an infinity, so the highest bit below the
1254        // quiet one goes in instead.
1255        assert_eq!(double(false, 0).to_bits(), 0x7ff4_0000_0000_0000, "__builtin_nans(\"\")");
1256        assert_eq!(double(false, 1).to_bits(), 0x7ff0_0000_0000_0001, "__builtin_nans(\"0x1\")");
1257        // A payload that fills the fraction, and one bit more than fits, which is cut.
1258        assert_eq!(double(true, 0xf_ffff_ffff_ffff).to_bits(), 0x7fff_ffff_ffff_ffff);
1259        assert_eq!(double(true, 1 << 52).to_bits(), 0x7ff8_0000_0000_0000);
1260        assert_eq!(
1261            Float::nan_with(Format::Single, false, true, 1).to_bits(),
1262            0x7fc0_0001,
1263            "__builtin_nanf(\"0x1\")"
1264        );
1265        assert_eq!(
1266            Float::nan_with(Format::Single, false, false, 0).to_bits(),
1267            0x7fa0_0000,
1268            "__builtin_nansf(\"\")"
1269        );
1270        // The x87 format stores the leading significand bit, which is set in a nan as in
1271        // everything else that is not a zero.
1272        assert_eq!(
1273            Float::nan_with(Format::X87Extended, false, true, 1).to_bits(),
1274            0x7fff_c000_0000_0000_0001,
1275            "__builtin_nanl(\"0x1\") on x86"
1276        );
1277        assert_eq!(
1278            Float::nan_with(Format::X87Extended, false, false, 0).to_bits(),
1279            0x7fff_a000_0000_0000_0000,
1280            "__builtin_nansl(\"\") on x86"
1281        );
1282    }
1283
1284    /// A payload is part of the value, so it has to survive being written down and read back.
1285    #[test]
1286    fn a_payload_comes_back_out_of_the_encoding_it_went_into() {
1287        for format in [Format::Half, Format::Single, Format::Double, Format::X87Extended] {
1288            for (quiet, payload) in [(true, 0), (true, 1), (false, 3), (true, 5)] {
1289                let nan = Float::nan_with(format, false, quiet, payload);
1290                assert!(nan.is_nan(), "{format:?}");
1291                assert_eq!(Float::from_bits(format, nan.to_bits()), nan, "{format:?} {payload}");
1292            }
1293            // The sign of a nan is its own, and negating one leaves the payload alone.
1294            let nan = Float::nan_with(format, true, true, 7);
1295            assert!(nan.is_negative() && nan.negated().negated() == nan, "{format:?}");
1296        }
1297    }
1298
1299    /// The smallest normal is what `isnormal` compares against, so it has to be the exact value
1300    /// the host calls `MIN_POSITIVE` and the bit below it has to be a subnormal.
1301    #[test]
1302    fn the_smallest_normal_is_the_number_below_which_nothing_is_normal() {
1303        assert_eq!(
1304            Float::smallest_normal(Format::Single, false).to_bits(),
1305            u128::from(f32::MIN_POSITIVE.to_bits())
1306        );
1307        assert_eq!(
1308            Float::smallest_normal(Format::Double, false).to_bits(),
1309            u128::from(f64::MIN_POSITIVE.to_bits())
1310        );
1311        // The x87 format stores its leading bit, so the smallest normal has the lowest exponent
1312        // field that is not the subnormal one and that bit alone.
1313        assert_eq!(
1314            Float::smallest_normal(Format::X87Extended, false).to_bits(),
1315            (1u128 << 64) | (1u128 << 63)
1316        );
1317        for format in [Format::Half, Format::BFloat16, Format::Single, Format::Double] {
1318            let normal = Float::smallest_normal(format, false);
1319            assert!(normal.is_finite() && !normal.is_zero(), "{format:?}");
1320            assert_eq!(Float::from_bits(format, normal.to_bits()), normal, "{format:?}");
1321            // One less in the encoding is the largest subnormal, which is what the boundary
1322            // being in the right place means.
1323            let below = Float::from_bits(format, normal.to_bits() - 1);
1324            assert_eq!(below.compare(normal), Some(std::cmp::Ordering::Less), "{format:?}");
1325            // And the negative one is the same number with the sign bit set.
1326            let negative = Float::smallest_normal(format, true);
1327            assert!(negative.is_negative() && negative.negated() == normal, "{format:?}");
1328        }
1329    }
1330
1331    /// Every format there is, so that a new one has to be added here and answered for below.
1332    const EVERY_FORMAT: [Format; 10] = [
1333        Format::Half,
1334        Format::BFloat16,
1335        Format::Single,
1336        Format::Double,
1337        Format::X87Extended,
1338        Format::Quad,
1339        Format::DoubleDouble,
1340        Format::Decimal32,
1341        Format::Decimal64,
1342        Format::Decimal128,
1343    ];
1344
1345    #[test]
1346    fn the_double_double_and_the_decimals_are_the_formats_that_are_not_binary_ieee() {
1347        for format in EVERY_FORMAT {
1348            let binary = format != Format::DoubleDouble && format.decimal().is_none();
1349            assert_eq!(format.is_ieee(), binary, "{format:?}");
1350        }
1351    }
1352
1353    #[test]
1354    fn every_format_has_a_name_that_reads_back_as_itself() {
1355        // The names are what a data layout is written in and what a diagnostic says, so a format
1356        // whose name does not round trip is a format something else will read as another one.
1357        for format in EVERY_FORMAT {
1358            assert_eq!(Format::from_name(format.name()), Some(format), "{format:?}");
1359        }
1360        assert_eq!(Format::from_name("f128"), Some(Format::Quad));
1361        assert_eq!(Format::from_name("ppc-f128"), Some(Format::DoubleDouble));
1362        assert_eq!(Format::from_name("f256"), None);
1363    }
1364
1365    #[test]
1366    fn a_width_is_the_one_question_the_double_double_answers() {
1367        // It is a hundred and twenty eight bits the same way binary128 is, which is why the two
1368        // cannot be told apart by width and why `spec/cross-compile/06-abis.md` section 6.2 item
1369        // 1 says the format is carried beside it.
1370        assert_eq!(Format::DoubleDouble.width(), 128);
1371        assert_eq!(Format::Quad.width(), Format::DoubleDouble.width());
1372        assert_ne!(Format::Quad, Format::DoubleDouble);
1373    }
1374
1375    #[test]
1376    #[should_panic(expected = "pair of doubles")]
1377    fn asking_a_double_double_for_a_precision_says_why_there_is_not_one() {
1378        let _ = Format::DoubleDouble.precision();
1379    }
1380
1381    #[test]
1382    #[should_panic(expected = "pair of doubles")]
1383    fn a_double_double_cannot_be_parsed_into() {
1384        // The refusal is at the format rather than at the spelling, so a well formed number in a
1385        // format this type does not represent fails, and fails saying which of the two is wrong.
1386        let _ = Float::parse("1.0", Format::DoubleDouble);
1387    }
1388
1389    #[test]
1390    #[should_panic(expected = "pair of doubles")]
1391    fn a_double_double_cannot_be_read_out_of_its_bits_either() {
1392        let _ = Float::from_bits(Format::DoubleDouble, 0);
1393    }
1394
1395    #[test]
1396    #[should_panic(expected = "pair of doubles")]
1397    fn not_even_a_double_double_zero_can_be_made() {
1398        // A zero looks harmless and is the one that would get through, because it needs no
1399        // precision and no exponent to build. Letting it through is how a value in a format
1400        // nothing here can encode reaches `to_bits`, which is several steps from the mistake.
1401        let _ = Float::zero(Format::DoubleDouble, false);
1402    }
1403}