ruvllm 2.3.0

LLM serving runtime with Ruvector integration - Paged attention, KV cache, and SONA learning
Documentation

ruvllm

There is currently very little information to present on this page because Docs.rs has only limited support for extracting structured feature metadata from Cargo crates. This issue is tracked in Rust RFC #3416. Check this library's main docs, readme, and Cargo.toml in case its authors have documentation for features available there instead.

This version has 31 feature flags, 11 of them enabled by default.

default

async-runtime (default)

candle (default)

hub-download (default)

quantize (default)

This feature flag does not enable additional features.

routing-metrics (default)

This feature flag does not enable additional features.

tokio (default)

tokio-stream (default)

candle-core (default)

candle-nn (default)

candle-transformers (default)

tokenizers (default)

accelerate

This feature flag does not enable additional features.

attention

coreml

cuda

fused-act

gguf-mmap

gnn

graph

hybrid-ane

inference-cuda

inference-metal

inference-metal-native

metal

metal-compute

minimal

mmap

parallel

ruvector-full

wasm

This feature flag does not enable additional features.

wasm-simd

This feature flag does not enable additional features.