pub fn load_gguf_raw<P>(path: P) -> Result<GgufRawLoadResult, AprenderError>Expand description
Load GGUF with raw quantized tensors (preserves Q4K for GPU inference)
This is essential for APR format to achieve 2x Ollama performance. The Q4K bytes are stored directly in APR and used by GPU kernels.