pub struct Float { /* private fields */ }Expand description
A floating point number in a given format.
A finite value is significand * 2^(exponent - precision + 1). A normal number has its
leading significand bit set, a subnormal does not and has the format’s minimum exponent.
Implementations§
Source§impl Float
impl Float
Sourcepub const fn nan(format: Format) -> Float
pub const fn nan(format: Format) -> Float
A quiet nan, which is what an operation with no answer gives.
Sourcepub const fn negated(self) -> Float
pub const fn negated(self) -> Float
The number with its sign flipped, which is exact and which a zero and a nan both have.
Sourcepub fn sum(self, other: Float) -> (Float, Status)
pub fn sum(self, other: Float) -> (Float, Status)
self + other, rounded to nearest with ties to even.
A nan operand gives a nan and nothing else. Two infinities of opposite sign give a nan and
Status::INVALID, because the answer depends on how they got there. Two zeros give a
negative zero only when both of them are negative, which is the round to nearest rule and
the reason x + 0.0 is not a way to drop a sign.
§Panics
If the two numbers are not in the same format. The usual arithmetic conversions have already made them so, and converting here would be a conversion nobody asked for.
Sourcepub fn difference(self, other: Float) -> (Float, Status)
pub fn difference(self, other: Float) -> (Float, Status)
self - other, rounded to nearest with ties to even.
This is the sum of self and the negation of other, which is exactly what it is in IEEE
754, so a subtraction that cancels completely gives a positive zero and an infinity minus
itself gives a nan.
§Panics
If the two numbers are not in the same format.
Sourcepub fn product(self, other: Float) -> (Float, Status)
pub fn product(self, other: Float) -> (Float, Status)
self * other, rounded to nearest with ties to even.
A zero times an infinity gives a nan and Status::INVALID. The sign is the two signs
multiplied, which a zero and a nan have as much as any other number does.
§Panics
If the two numbers are not in the same format.
Sourcepub fn quotient(self, other: Float) -> (Float, Status)
pub fn quotient(self, other: Float) -> (Float, Status)
self / other, rounded to nearest with ties to even.
A finite number divided by zero gives an infinity and Status::DIVIDE_BY_ZERO. Zero
divided by zero and an infinity divided by an infinity both give a nan and
Status::INVALID, which is the difference between a division that has no answer and one
whose answer is only too large to be a number.
§Panics
If the two numbers are not in the same format.
Sourcepub fn compare(self, other: Float) -> Option<Ordering>
pub fn compare(self, other: Float) -> Option<Ordering>
How the two compare, or None if either is a nan and they do not compare at all.
This is the comparison C’s relational operators do, so a positive zero and a negative zero
are equal and the unordered case is the one that makes x < y and !(x >= y) different
questions.
§Panics
If the two numbers are not in the same format.
Sourcepub fn to_format(self, format: Format) -> (Float, Status)
pub fn to_format(self, format: Format) -> (Float, Status)
The nearest number to this one in another format, rounded to nearest with ties to even.
Widening is exact for every pair of formats here except a __bf16 widened to a
_Float16, which has more precision and less range. Narrowing is what a cast does, and it
reports what it had to do to make the number fit.
Sourcepub fn from_signed(value: i128, format: Format) -> (Float, Status)
pub fn from_signed(value: i128, format: Format) -> (Float, Status)
The nearest number in format to a signed integer.
Sourcepub fn from_unsigned(value: u128, format: Format) -> (Float, Status)
pub fn from_unsigned(value: u128, format: Format) -> (Float, Status)
The nearest number in format to an unsigned integer.
Sourcepub fn to_integer(self, width: u32, signed: bool) -> (i128, Status)
pub fn to_integer(self, width: u32, signed: bool) -> (i128, Status)
The number truncated toward zero into an integer of width bits.
What comes back is what an integer constant is stored as, which is the value sign extended
out of the type it has, so an unsigned conversion of a hundred and twenty eight bits comes
back with its top bit in the sign of the i128.
Converting a number that does not fit is undefined behaviour in C rather than a value, so
what comes back is the nearest end of the range together with Status::INVALID, which
is what the caller warns about. A nan comes back as zero, for the same reason and with the
same flag. Dropping a fraction is Status::INEXACT and nothing worse, since that is the
conversion doing what it is for.
§Panics
If width is zero or wider than a hundred and twenty eight bits.
Source§impl Float
impl Float
Sourcepub const fn is_negative(self) -> bool
pub const fn is_negative(self) -> bool
Whether the number is negative, which a zero can be.
Sourcepub const fn is_infinite(self) -> bool
pub const fn is_infinite(self) -> bool
Whether the number is an infinity.
Sourcepub const fn is_finite(self) -> bool
pub const fn is_finite(self) -> bool
Whether the number is finite, which a zero is and a nan is not.
Sourcepub fn parse(text: &str, format: Format) -> Result<(Float, Status), ParseError>
pub fn parse(text: &str, format: Format) -> Result<(Float, Status), ParseError>
Converts a decimal or hexadecimal spelling into the nearest number in format, rounding
to nearest with ties to even.
The spelling is the number alone: no suffix, because the suffix is what chose the format, and no infinity or nan, because C has no spelling for those. A sign is accepted even though a C constant never has one, since the value the constant evaluator folds does. C23 digit separators are stripped here.
§Errors
ParseError, for a spelling that is not a number at all.
Sourcepub fn to_bits(self) -> u128
pub fn to_bits(self) -> u128
The bits of the encoding, in the low Format::width bits.
The x87 format keeps its leading significand bit, so its eightieth bit is the sign and its sixty fourth is the one every other format leaves implied.
Sourcepub fn from_bits(format: Format, bits: u128) -> Float
pub fn from_bits(format: Format, bits: u128) -> Float
Reads a number back out of its encoding, which is what makes Float::to_bits testable
and what a constant folded in the IR is stored as.
A signalling nan comes back as a quiet one and a payload comes back as nothing, because nothing here has anywhere to put either and no C program can see the difference in a constant.
Sourcepub fn to_hex(self) -> String
pub fn to_hex(self) -> String
A hexadecimal spelling that Float::parse turns back into exactly this number.
Hexadecimal rather than decimal, because a hexadecimal constant is exact by construction
and a decimal one is not: printing a number in decimal so that it reads back unchanged
needs a shortest-round-trip algorithm, and printing it in decimal without one silently
changes the program. A printer that changes a constant is worse than a printer whose
output is unfamiliar, so this is 0x1p+0 where a reader would rather see 1.0.
The significand is written as an integer and the exponent scales it, so the spelling is
significand * 2^exponent with no leading digit to argue about. Trailing zero digits are
taken off, which is what makes a round number short.
An infinity has no spelling in C at all. What comes back for one is an exponent past the
top of the format, which converts back to an infinity with the overflow that a constant
only ever became an infinity by. A nan is spelled nan and does not read back, since
there is no exponent that gives one and no constant that is one.