Skip to main content

e4m3

Struct e4m3 

Source
pub struct e4m3(/* private fields */);
Expand description

A 8-bit floating point type with 4 exponent bits and 3 mantissa bits.

Follows table 1 of FP8 Formats for Deep Learning. E4M3 departs from IEEE 754 to widen its range: it has no infinities, and reserves only S.1111.111 for NaN. Reclaiming the rest of the special-value patterns is what puts its maximum at 448 rather than the 240 an IEEE-style reservation would give.

MAX, MIN, MAX_EXP and MANTISSA_DIGITS are spelled out here rather than delegated to float8, each for the reason given on the constant itself. Every other constant delegates and agrees. No infinity constant is exposed, because the format has none.

See also the minifloat overview.

Implementations§

Source§

impl e4m3

Source

pub const MAX: e4m3

Maximum representable value, S.1111.110 = 1.75 * 2^8 = 448.

Not taken from F8E4M3::MAX, which is 416 and one representable step low. E4M3 has no infinities and reserves only S.1111.111 for NaN, so 0x7E is finite. See table 1 of “FP8 Formats for Deep Learning” (arXiv:2209.05433), which also matches what from_f32 saturates to and what QuantValue::E4M3 bounds against.

Source

pub const MIN: e4m3

Minimum representable value, the negation of MAX.

Source

pub const EPSILON: e4m3

the difference between 1.0 and the next largest representable number.

Source

pub const MIN_POSITIVE: e4m3

Minimum representable value

Source

pub const DIGITS: u32 = F8E4M3::DIGITS

Approximate number of significant digits in base 10

Source

pub const MANTISSA_DIGITS: u32 = 4

Number of mantissa digits, the 3 stored bits plus the implicit leading one.

F8E4M3::MANTISSA_DIGITS counts only the stored bits, where f32 and half::f16 both count the implicit one.

Source

pub const MAX_10_EXP: i32 = F8E4M3::MAX_10_EXP

Maximum possible normal power of 10 exponent

Source

pub const MAX_EXP: i32 = 9

One greater than the maximum possible normal power of 2 exponent, matching f32::MAX_EXP.

MAX is 1.75 * 2^8, so this is 9. F8E4M3::MAX_EXP is 7: one lower because it reserves an exponent for infinities this format does not have, and another because it does not add the one that f32 and half::f16 do.

Source

pub const MIN_10_EXP: i32 = F8E4M3::MIN_10_EXP

Minimum possible normal power of 10 exponent

Source

pub const MIN_EXP: i32 = F8E4M3::MIN_EXP

Minimum possible normal power of 2 exponent

Source

pub const RADIX: u32 = 2

The radix, or base, of the floating-point representation.

Source

pub const NAN: e4m3

nan

Source

pub const ZERO: e4m3

Zero

Source

pub const NEG_ZERO: e4m3

Negative Zero

Source

pub const ONE: e4m3

One

Source

pub const fn from_bits(bits: u8) -> e4m3

Constructs a e4m3 value from the raw bits.

Source

pub const fn from_f32(value: f32) -> e4m3

Constructs a e4m3 value from a 32-bit floating point value.

This operation is lossy. Values too large to fit, infinities included, saturate to ±MAX, since the format has no infinities of its own. NaN values are preserved. Subnormal values that are too tiny to be represented will result in ±0. All other values are truncated and rounded to the nearest representable value.

Source

pub const fn from_f64(value: f64) -> e4m3

Constructs a e4m3 value from a 64-bit floating point value.

This operation is lossy. Values too large to fit, infinities included, saturate to ±MAX, since the format has no infinities of its own. NaN values are preserved. 64-bit subnormal values are too tiny to be represented and result in ±0. Exponents that underflow the minimum exponent will result in subnormals or ±0. All other values are truncated and rounded to the nearest representable value.

Source

pub const fn to_bits(self) -> u8

Converts a e4m3 into the underlying bit representation.

Source

pub const fn to_f32(self) -> f32

Converts a e4m3 value into an f32 value.

This conversion is lossless as all values can be represented exactly in f32.

Source

pub fn is_nan(self) -> bool

check if an e4m3 value is Nan

Source

pub const fn to_f64(self) -> f64

Converts a e4m3 value into an f64 value.

This conversion is lossless as all values can be represented exactly in f64.

Source

pub fn total_cmp(self, other: e4m3) -> Ordering

Compares e4m3 values

Trait Implementations§

Source§

impl Add for e4m3

Source§

type Output = e4m3

The resulting type after applying the + operator.
Source§

fn add(self, rhs: e4m3) -> <e4m3 as Add>::Output

Performs the + operation. Read more
Source§

impl AddAssign for e4m3

Source§

fn add_assign(&mut self, rhs: e4m3)

Performs the += operation. Read more
Source§

impl Clone for e4m3

Source§

fn clone(&self) -> e4m3

Returns a duplicate of the value. Read more
1.0.0 (const: unstable) · Source§

fn clone_from(&mut self, source: &Self)

Performs copy-assignment from source. Read more
Source§

impl Copy for e4m3

Source§

impl Debug for e4m3

Source§

fn fmt(&self, f: &mut Formatter<'_>) -> Result<(), Error>

Formats the value using the given formatter. Read more
Source§

impl Default for e4m3

Source§

fn default() -> e4m3

Returns the “default value” for a type. Read more
Source§

impl<'de> Deserialize<'de> for e4m3

Source§

fn deserialize<__D>( __deserializer: __D, ) -> Result<e4m3, <__D as Deserializer<'de>>::Error>
where __D: Deserializer<'de>,

Deserialize this value from the given Serde deserializer. Read more
Source§

impl Display for e4m3

Source§

fn fmt(&self, f: &mut Formatter<'_>) -> Result<(), Error>

Formats the value using the given formatter. Read more
Source§

impl Div for e4m3

Source§

type Output = e4m3

The resulting type after applying the / operator.
Source§

fn div(self, rhs: e4m3) -> <e4m3 as Div>::Output

Performs the / operation. Read more
Source§

impl DivAssign for e4m3

Source§

fn div_assign(&mut self, rhs: e4m3)

Performs the /= operation. Read more
Source§

impl From<F8E4M3> for e4m3

Source§

fn from(value: F8E4M3) -> e4m3

Converts to this type from the input type.
Source§

impl Mul for e4m3

Source§

type Output = e4m3

The resulting type after applying the * operator.
Source§

fn mul(self, rhs: e4m3) -> <e4m3 as Mul>::Output

Performs the * operation. Read more
Source§

impl MulAssign for e4m3

Source§

fn mul_assign(&mut self, rhs: e4m3)

Performs the *= operation. Read more
Source§

impl Neg for e4m3

Source§

type Output = e4m3

The resulting type after applying the - operator.
Source§

fn neg(self) -> <e4m3 as Neg>::Output

Performs the unary - operation. Read more
Source§

impl Num for e4m3

Source§

type FromStrRadixErr = ParseFloatError

Source§

fn from_str_radix( src: &str, radix: u32, ) -> Result<e4m3, <e4m3 as Num>::FromStrRadixErr>

Convert from a string and radix (typically 2..=36). Read more
Source§

impl NumCast for e4m3

Source§

fn from<T>(n: T) -> Option<e4m3>
where T: ToPrimitive,

Creates a number from another value that can be converted into a primitive via the ToPrimitive trait. If the source value cannot be represented by the target type, then None is returned. Read more
Source§

impl One for e4m3

Source§

fn one() -> e4m3

Returns the multiplicative identity element of Self, 1. Read more
Source§

fn is_one(&self) -> bool

Returns true if self is equal to the multiplicative identity. Read more
Source§

fn set_one(&mut self)

Sets self to the multiplicative identity element of Self, 1.
Source§

impl PartialEq for e4m3

Source§

fn eq(&self, other: &e4m3) -> bool

Equality operator ==. Read more
1.0.0 (const: unstable) · Source§

fn ne(&self, other: &Rhs) -> bool

Inequality operator !=. Read more
Source§

impl PartialOrd for e4m3

Source§

fn partial_cmp(&self, other: &e4m3) -> Option<Ordering>

This method returns an ordering between self and other values if one exists. Read more
1.0.0 (const: unstable) · Source§

fn lt(&self, other: &Rhs) -> bool

Tests less than (for self and other) and is used by the < operator. Read more
1.0.0 (const: unstable) · Source§

fn le(&self, other: &Rhs) -> bool

Tests less than or equal to (for self and other) and is used by the <= operator. Read more
1.0.0 (const: unstable) · Source§

fn gt(&self, other: &Rhs) -> bool

Tests greater than (for self and other) and is used by the > operator. Read more
1.0.0 (const: unstable) · Source§

fn ge(&self, other: &Rhs) -> bool

Tests greater than or equal to (for self and other) and is used by the >= operator. Read more
Source§

impl Pod for e4m3

Source§

impl Rem for e4m3

Source§

type Output = e4m3

The resulting type after applying the % operator.
Source§

fn rem(self, rhs: e4m3) -> <e4m3 as Rem>::Output

Performs the % operation. Read more
Source§

impl RemAssign for e4m3

Source§

fn rem_assign(&mut self, rhs: e4m3)

Performs the %= operation. Read more
Source§

impl Serialize for e4m3

Source§

fn serialize<__S>( &self, __serializer: __S, ) -> Result<<__S as Serializer>::Ok, <__S as Serializer>::Error>
where __S: Serializer,

Serialize this value into the given Serde serializer. Read more
Source§

impl StructuralPartialEq for e4m3

Source§

impl Sub for e4m3

Source§

type Output = e4m3

The resulting type after applying the - operator.
Source§

fn sub(self, rhs: e4m3) -> <e4m3 as Sub>::Output

Performs the - operation. Read more
Source§

impl SubAssign for e4m3

Source§

fn sub_assign(&mut self, rhs: e4m3)

Performs the -= operation. Read more
Source§

impl ToPrimitive for e4m3

Source§

fn to_i64(&self) -> Option<i64>

Converts the value of self to an i64. If the value cannot be represented by an i64, then None is returned.
Source§

fn to_u64(&self) -> Option<u64>

Converts the value of self to a u64. If the value cannot be represented by a u64, then None is returned.
Source§

fn to_f32(&self) -> Option<f32>

Converts the value of self to an f32. Overflows may map to positive or negative inifinity, otherwise None is returned if the value cannot be represented by an f32.
Source§

fn to_f64(&self) -> Option<f64>

Converts the value of self to an f64. Overflows may map to positive or negative inifinity, otherwise None is returned if the value cannot be represented by an f64. Read more
Source§

fn to_isize(&self) -> Option<isize>

Converts the value of self to an isize. If the value cannot be represented by an isize, then None is returned.
Source§

fn to_i8(&self) -> Option<i8>

Converts the value of self to an i8. If the value cannot be represented by an i8, then None is returned.
Source§

fn to_i16(&self) -> Option<i16>

Converts the value of self to an i16. If the value cannot be represented by an i16, then None is returned.
Source§

fn to_i32(&self) -> Option<i32>

Converts the value of self to an i32. If the value cannot be represented by an i32, then None is returned.
Source§

fn to_i128(&self) -> Option<i128>

Converts the value of self to an i128. If the value cannot be represented by an i128 (i64 under the default implementation), then None is returned. Read more
Source§

fn to_usize(&self) -> Option<usize>

Converts the value of self to a usize. If the value cannot be represented by a usize, then None is returned.
Source§

fn to_u8(&self) -> Option<u8>

Converts the value of self to a u8. If the value cannot be represented by a u8, then None is returned.
Source§

fn to_u16(&self) -> Option<u16>

Converts the value of self to a u16. If the value cannot be represented by a u16, then None is returned.
Source§

fn to_u32(&self) -> Option<u32>

Converts the value of self to a u32. If the value cannot be represented by a u32, then None is returned.
Source§

fn to_u128(&self) -> Option<u128>

Converts the value of self to a u128. If the value cannot be represented by a u128 (u64 under the default implementation), then None is returned. Read more
Source§

impl Zero for e4m3

Source§

fn zero() -> e4m3

Returns the additive identity element of Self, 0. Read more
Source§

fn is_zero(&self) -> bool

Returns true if self is equal to the additive identity.
Source§

fn set_zero(&mut self)

Sets self to the additive identity element of Self, 0.
Source§

impl Zeroable for e4m3

Source§

fn zeroed() -> Self

Auto Trait Implementations§

§

impl Freeze for e4m3

§

impl RefUnwindSafe for e4m3

§

impl Send for e4m3

§

impl Sync for e4m3

§

impl Unpin for e4m3

§

impl UnsafeUnpin for e4m3

§

impl UnwindSafe for e4m3

Blanket Implementations§

Source§

impl<T> Any for T
where T: 'static + ?Sized,

Source§

fn type_id(&self) -> TypeId

Gets the TypeId of self. Read more
Source§

impl<T> AnyBitPattern for T
where T: Pod,

Source§

impl<T> Borrow<T> for T
where T: ?Sized,

Source§

fn borrow(&self) -> &T

Immutably borrows from an owned value. Read more
Source§

impl<T> BorrowMut<T> for T
where T: ?Sized,

Source§

fn borrow_mut(&mut self) -> &mut T

Mutably borrows from an owned value. Read more
Source§

impl<ST, DT> CastableFrom<ST, Initialized, Initialized> for DT
where ST: ?Sized, DT: ?Sized,

Source§

impl<ST, DT> CastableFrom<ST, Uninit, Uninit> for DT
where ST: ?Sized, DT: ?Sized,

Source§

impl<T> CheckedBitPattern for T
where T: AnyBitPattern,

Source§

type Bits = T

Self must have the same layout as the specified Bits except for the possible invalid bit patterns being checked during is_valid_bit_pattern.
Source§

fn is_valid_bit_pattern(_bits: &T) -> bool

If this function returns true, then it must be valid to reinterpret bits as &Self.
Source§

impl<T> CloneToUninit for T
where T: Clone,

Source§

unsafe fn clone_to_uninit(&self, dest: *mut u8)

🔬This is a nightly-only experimental API. (clone_to_uninit)
Performs copy-assignment from self to dest. Read more
Source§

impl<T> DeserializeOwned for T
where T: for<'de> Deserialize<'de>,

Source§

impl<T> From<T> for T

Source§

fn from(t: T) -> T

Returns the argument unchanged.

Source§

impl<T, U> Into<U> for T
where U: From<T>,

Source§

fn into(self) -> U

Calls U::from(self).

That is, this conversion is whatever the implementation of From<T> for U chooses to do.

Source§

impl<T> NoUninit for T
where T: Pod,

Source§

impl<T> NumAssign for T
where T: Num + NumAssignOps,

Source§

impl<T, Rhs> NumAssignOps<Rhs> for T
where T: AddAssign<Rhs> + SubAssign<Rhs> + MulAssign<Rhs> + DivAssign<Rhs> + RemAssign<Rhs>,

Source§

impl<T, Rhs, Output> NumOps<Rhs, Output> for T
where T: Sub<Rhs, Output = Output> + Mul<Rhs, Output = Output> + Div<Rhs, Output = Output> + Add<Rhs, Output = Output> + Rem<Rhs, Output = Output>,

Source§

impl<T> Read<Exclusive, BecauseExclusive> for T
where T: ?Sized,

Source§

impl<T> ToOwned for T
where T: Clone,

Source§

type Owned = T

The resulting type after obtaining ownership.
Source§

fn to_owned(&self) -> T

Creates owned data from borrowed data, usually by cloning. Read more
Source§

fn clone_into(&self, target: &mut T)

Uses borrowed data to replace owned data, usually by cloning. Read more
Source§

impl<T> ToString for T
where T: Display + ?Sized,

Source§

fn to_string(&self) -> String

Converts the given value to a String. Read more
Source§

impl<T, U> TryFrom<U> for T
where U: Into<T>,

Source§

type Error = Infallible

The type returned in the event of a conversion error.
Source§

fn try_from(value: U) -> Result<T, <T as TryFrom<U>>::Error>

Performs the conversion.
Source§

impl<T, U> TryInto<U> for T
where U: TryFrom<T>,

Source§

type Error = <U as TryFrom<T>>::Error

The type returned in the event of a conversion error.
Source§

fn try_into(self) -> Result<U, <U as TryFrom<T>>::Error>

Performs the conversion.