pub fn release_cached_device_memory() -> Result<(), Error>Expand description
Release cached but currently-unused device memory back to the driver.
Candle hands every dropped Tensor back to CUDA’s async memory
pool (via cuMemFreeAsync). The pool keeps the bytes around for
fast reuse, which is normally what you want — but when the encoder
model finishes embedding, its ~2 GB of ModernBert per-batch
caches stay committed to the pool even after the callers drop the
model. That’s enough to block a subsequent 3.47 GB Tensor
allocation on a 12 GB card and surface as CUDA_ERROR_OUT_OF_MEMORY.
This function asks CUDA to trim the default mempool to zero
retained bytes, handing the freed blocks back to the driver so
the next large allocation can grow into them. Without the cuda
feature (or when running on CPU) it’s a no-op.
Safe wrapper around cuDeviceGetDefaultMemPool +
cuMemPoolTrimTo. Only trims device 0 — docbert currently only
ever uses the default CUDA device.