Skip to main content

QuantPreview

Struct QuantPreview 

Source
pub struct QuantPreview<'model> { /* private fields */ }
Expand description

Ask llama.cpp what it would do, without writing a file.

crate::model_quantize is all-or-nothing: it reads a model, quantizes every tensor, and writes the result. This exposes the decision layer underneath — which tensors get quantized at all, and to which ggml type — so a tool can show the plan, or estimate output size, before committing to a run that may take minutes and tens of gigabytes.

The k-quant mixes are why this is not simply LlamaFtype::default_ggml_type: they deliberately keep attention and output tensors at higher precision, so the per-tensor answer differs from the ftype’s nominal type.

Wraps llama_quant_init / llama_quant_free.

Implementations§

Source§

impl<'model> QuantPreview<'model>

Source

pub fn new( model: &'model LlamaModel, params: &QuantizeParams, ) -> Result<Self, QuantPreviewError>

Build a preview for model under params.

§Errors

Returns QuantPreviewError::Init if llama.cpp could not build the quantization state, which happens for a model it cannot quantize.

Source

pub fn allows_quantization(&self, tensor: &GgmlTensor) -> bool

Whether this tensor would be quantized at all.

llama.cpp skips 1-D tensors, tensors below a size threshold, and ones whose name marks them as needing full precision — so a false here is the usual reason a tensor keeps its original type.

Requires the ggml feature, which is what exposes crate::ggml::GgmlTensor.

Source

pub fn compute_types( &self, tensors: &[&GgmlTensor], ftype: LlamaFtype, ) -> Result<Vec<Option<GgmlType>>, QuantPreviewError>

Compute the storage type each tensor would be assigned under ftype.

Every tensor passed must already satisfy Self::allows_quantization — upstream states the caller filters first, and does not re-check.

An entry is None when llama.cpp picks a ggml type this crate’s GgmlType does not know.

§Errors

Returns QuantPreviewError::NotQuantizable naming the first tensor that fails the filter, rather than letting llama.cpp decide what to do with it.

Requires the ggml feature.

Trait Implementations§

Source§

impl<'model> Debug for QuantPreview<'model>

Source§

fn fmt(&self, f: &mut Formatter<'_>) -> Result

Formats the value using the given formatter. Read more
Source§

impl Drop for QuantPreview<'_>

Source§

fn drop(&mut self)

Executes the destructor for this type. Read more
Source§

fn pin_drop(self: Pin<&mut Self>)

🔬This is a nightly-only experimental API. (pin_ergonomics)
Execute the destructor for this type, but different to Drop::drop, it requires self to be pinned. Read more

Auto Trait Implementations§

§

impl<'model> !Send for QuantPreview<'model>

§

impl<'model> !Sync for QuantPreview<'model>

§

impl<'model> Freeze for QuantPreview<'model>

§

impl<'model> RefUnwindSafe for QuantPreview<'model>

§

impl<'model> Unpin for QuantPreview<'model>

§

impl<'model> UnsafeUnpin for QuantPreview<'model>

§

impl<'model> UnwindSafe for QuantPreview<'model>

Blanket Implementations§

Source§

impl<T> Any for T
where T: 'static + ?Sized,

Source§

fn type_id(&self) -> TypeId

Gets the TypeId of self. Read more
Source§

impl<T> Borrow<T> for T
where T: ?Sized,

Source§

fn borrow(&self) -> &T

Immutably borrows from an owned value. Read more
Source§

impl<T> BorrowMut<T> for T
where T: ?Sized,

Source§

fn borrow_mut(&mut self) -> &mut T

Mutably borrows from an owned value. Read more
Source§

impl<T> From<T> for T

Source§

fn from(t: T) -> T

Returns the argument unchanged.

Source§

impl<T> Instrument for T

Source§

fn instrument(self, span: Span) -> Instrumented<Self>

Instruments this type with the provided Span, returning an Instrumented wrapper. Read more
Source§

fn in_current_span(self) -> Instrumented<Self>

Instruments this type with the current Span, returning an Instrumented wrapper. Read more
Source§

impl<T, U> Into<U> for T
where U: From<T>,

Source§

fn into(self) -> U

Calls U::from(self).

That is, this conversion is whatever the implementation of From<T> for U chooses to do.

Source§

impl<T, U> TryFrom<U> for T
where U: Into<T>,

Source§

type Error = !

The type returned in the event of a conversion error.
Source§

fn try_from(value: U) -> Result<T, !>

Performs the conversion.
Source§

impl<T, U> TryInto<U> for T
where U: TryFrom<T>,

Source§

type Error = <U as TryFrom<T>>::Error

The type returned in the event of a conversion error.
Source§

fn try_into(self) -> Result<U, <U as TryFrom<T>>::Error>

Performs the conversion.
Source§

impl<T> WithSubscriber for T

Source§

fn with_subscriber<S>(self, subscriber: S) -> WithDispatch<Self>
where S: Into<Dispatch>,

Attaches the provided Subscriber to this type, returning a WithDispatch wrapper. Read more
Source§

fn with_current_subscriber(self) -> WithDispatch<Self>

Attaches the current default Subscriber to this type, returning a WithDispatch wrapper. Read more