hf-fetch-model 0.12.1

Download, inspect, and compare HuggingFace models from Rust. Multi-connection parallel downloads plus safetensors, NPZ, GGUF, and PyTorch .pth header inspection via HTTP Range. No weight data downloaded.
Documentation

hf-fetch-model

CI Crates.io docs.rs Rust License

A Rust library and CLI for downloading and inspecting HuggingFace models. Multi-connection parallel downloads, file filtering, checksum verification, retry — plus remote tensor-header inspection (.safetensors, NumPy .npz, .gguf) and structural comparison between models, all without downloading weight data.

Table of contents

New to hf-fm?

Common questions live in the FAQ; every flag is in the CLI Reference.

Install

cargo install hf-fetch-model --features cli

Verify the install:

hf-fm --version

Upgrading from a previous version

cargo install skips the build when any version of the binary is already on disk, even if crates.io has a newer release. Use --force to upgrade:

cargo install hf-fetch-model --features cli --force
hf-fm --version

Without --force, a stale local registry index can cause cargo install to exit 0 silently — the install command appears to succeed but the binary on PATH is unchanged. See FAQ → How do I upgrade hf-fm? for more.

Commands

Command Description
hf-fm <REPO_ID> (default) Download a model (multi-connection, auto-tuned)
hf-fm cache clean-partial Remove .chunked.part files from interrupted downloads
hf-fm cache delete <REPO_ID|N> Delete a cached model
hf-fm cache gc --older-than/--max-size Garbage-collect cached models by age and/or size budget
hf-fm cache path <REPO_ID|N> Print snapshot directory path (for scripting)
hf-fm cache verify <REPO_ID|N> Re-verify SHA256 digests of cached files against HF LFS metadata
hf-fm diff <REPO_A> <REPO_B> Compare tensor layouts between two models
hf-fm diff-config <REPO_A> <REPO_B> Compare config.json architecture fields between two models
hf-fm discover Find new model families on the Hub
hf-fm download-file <REPO_ID> <FILE> Download a single file (or glob pattern)
hf-fm du [REPO_ID|N] Show cache disk usage (by name or # index)
hf-fm inspect <REPO_ID> [FILE] Inspect tensor headers (names, shapes, dtypes) without downloading weights — safetensors/NPZ/GGUF/PTH remote or cached; add --check-gpu [--context N] for a GPU-fit verdict (with KV-cache budgeting), or --pick to choose the file interactively
hf-fm list-families List model families in local cache
hf-fm list-files <REPO_ID> List remote files (sizes, SHA256) without downloading
hf-fm peek <REPO_ID> <FILE> Print a small file's content (config, README, .gz sidecar) without downloading — --head/--tail bound the read, --gunzip decodes, --max caps the size (no tensor formats — use inspect for those)
hf-fm quants <REPO_ID> Aggregate a base model's quant sibling repos into one sorted table; add --fits <SIZE> [--reserve <SIZE>] for an offload-aware VRAM fit plan
hf-fm search <QUERY> Search the HuggingFace Hub for models
hf-fm status [REPO_ID] Per-repo: per-file download status (complete / partial / missing / excluded). With no REPO_ID: a table of all cached repos, each marked ok or PARTIAL.

See CLI Reference for all flags and output examples.

Try it

$ hf-fm search mistral,3B,instruct
Models matching "mistral,3B,instruct" (by downloads):

  hf-fm mistralai/Ministral-3-3B-Instruct-2512           (159,700 downloads)
  hf-fm mistralai/Ministral-3-3B-Instruct-2512-BF16      (62,600 downloads)
  hf-fm mistralai/Ministral-3-3B-Instruct-2512-GGUF      (32,700 downloads)
  ...

$ hf-fm search llama --tag gguf --limit 3
Models matching "llama" (by downloads):

  hf-fm bartowski/Llama-3.2-3B-Instruct-GGUF             (489,856 downloads)  [text-generation]
  hf-fm bartowski/Meta-Llama-3.1-8B-Instruct-GGUF        (237,791 downloads)  [text-generation]
  hf-fm MaziyarPanahi/Meta-Llama-3.1-8B-Instruct-GGUF    (184,847 downloads)  [text-generation]

$ hf-fm search fp4 --tag bitsandbytes --show tags,size --limit 3
Models matching "fp4" (by downloads):

  hf-fm HF-Quantization/Llama-3.2-1B-BNB-FP4-BF16     (5 downloads)  [transformers, text-generation]  1.50 GiB  tags: transformers, safetensors, llama, 4-bit, bitsandbytes
  hf-fm saxman/Qwen3-Coder-30B-A3B-Instruct-bnb-fp4   (4 downloads)  [transformers, text-generation]  18.20 GiB  tags: transformers, safetensors, qwen3, 4-bit, bitsandbytes
  hf-fm ema1234/qwen_mcqa_bnb_fp4                     (2 downloads)  [transformers, text-generation]  548.00 MiB  tags: transformers, safetensors, qwen3, 4-bit, bitsandbytes

$ hf-fm search mistralai/Ministral-3-3B-Instruct-2512 --exact
Exact match:

  hf-fm mistralai/Ministral-3-3B-Instruct-2512           (159,700 downloads)

  License:      apache-2.0
  Pipeline:     text-generation
  Library:      vllm
  Languages:    en, fr, es, de, it, pt, nl, zh, ja, ko, ar

$ hf-fm list-files mistralai/Ministral-3-3B-Instruct-2512 --preset safetensors
  File                                               Size      SHA256
  model-00001-of-00002.safetensors                 3.68 GiB    a1b2c3d4e5f6
  model-00002-of-00002.safetensors                 2.88 GiB    f6e5d4c3b2a1
  config.json                                        856 B     —
  ...
  7 files, 6.57 GiB total

$ hf-fm mistralai/Ministral-3-3B-Instruct-2512 --preset safetensors --dry-run
  Repo:     mistralai/Ministral-3-3B-Instruct-2512
  Revision: main

  File                                               Size      Status
  model-00001-of-00002.safetensors                 3.68 GiB    to download
  model-00002-of-00002.safetensors                 2.88 GiB    to download
  ...
  Total: 6.57 GiB (7 files, 0 cached, 7 to download)

  Recommended config:
    concurrency:        2
    connections/file:   8
    chunk threshold:  100 MiB

$ hf-fm mistralai/Ministral-3-3B-Instruct-2512 --preset safetensors
Downloaded to: ~/.cache/huggingface/hub/models--mistralai--Ministral-3-3B.../snapshots/...
  6.57 GiB in 18.2s (369.1 MiB/s)

# Download to flat layout (files directly in ./models/)
$ hf-fm mistralai/Ministral-3-3B-Instruct-2512 --preset safetensors --flat --output-dir ./models

# Download sharded PyTorch files by glob
$ hf-fm download-file org/model "pytorch_model-*.bin"

Inspect & compare

$ hf-fm inspect EleutherAI/pythia-1.4b model.safetensors --cached --filter "layers.0."
  Repo:     EleutherAI/pythia-1.4b
  File:     model.safetensors
  Source:   cached

  Tensor                                             Dtype    Shape                  Size     Params
  gpt_neox.layers.0.attention.dense.weight           F16      [2048, 2048]       8.00 MiB       4.2M
  gpt_neox.layers.0.mlp.dense_h_to_4h.weight         F16      [8192, 2048]      32.00 MiB      16.8M
  ...
  ────────────────────────────────────────────────────────────────────────────────────────────────
  Showing 15 of 364 tensors matching filter "layers.0.".
  Param counts: 54.6M matching filter, 1.52B total.

$ hf-fm inspect google/gemma-4-E2B-it model.safetensors --tree --filter "embed"
  Repo:     google/gemma-4-E2B-it
  File:     model.safetensors
  Source:   remote (4 range requests, 182.4 KiB fetched)

  └── model.
      ├── embed_audio.embedding_projection.weight   BF16  [1536, 1536]   4.50 MiB
      ├── embed_vision.embedding_projection.weight  BF16  [1536, 768]    2.25 MiB
      ├── language_model.
      │   ├── embed_tokens.weight            BF16  [262144, 1536]      768.00 MiB
      │   └── embed_tokens_per_layer.weight  BF16  [262144, 8960]        4.38 GiB
      └── vision_tower.patch_embedder.
          ├── input_proj.weight         BF16  [768, 768]        1.12 MiB
          └── position_embedding_table  BF16  [2, 10240, 768]  30.00 MiB
  Showing 6 of 2011 tensors matching filter "embed".
  Param counts: 2.77B matching filter, 5.12B total.

$ hf-fm inspect poolside/Laguna-XS-2.1-GGUF Q4_K_M.gguf --group-by 'blk.*.ffn_*_exps.weight'
  Group                    Tensors        Size  Percent
  MATCHED (blk.*.ffn_*_exps.weight)   117   17.68 GiB    93.6%
  OTHER                                561    1.20 GiB     6.4%
  ────────────────────────────────────────────────────────────
  TOTAL                                678   18.88 GiB   100.0%

  per-MoE-layer expert cost: 0.45 GiB (39 layers)

$ hf-fm diff RedHatAI/Llama-3.2-1B-Instruct-FP8 casperhansen/llama-3.2-1b-instruct-awq --cached --summary
  A: RedHatAI/Llama-3.2-1B-Instruct-FP8
  B: casperhansen/llama-3.2-1b-instruct-awq
  ──────────────────────────────────────────────────────────────────────────────────────────────
  A: 371 tensors | B: 370 tensors | only-A: 337 | only-B: 336 | differ: 34 | match: 0

$ hf-fm diff openai/gpt-oss-20b openai/gpt-oss-120b --dtypes
  A: openai/gpt-oss-20b
  B: openai/gpt-oss-120b

  Dtype  A Tensors     A Size  B Tensors      B Size      Δ Size
  U8           192  18.91 GiB        288  113.46 GiB  +94.55 GiB
  BF16         630   6.72 GiB        942    8.07 GiB   +1.35 GiB
  ──────────────────────────────────────────────────────────────
  A: 822 tensors, 25.63 GiB | B: 1230 tensors, 121.54 GiB | Δ: +408 tensors, +95.90 GiB

$ hf-fm diff openai/gpt-oss-20b openai/gpt-oss-120b --collapse --limit 3
  A: openai/gpt-oss-20b
  B: openai/gpt-oss-120b

  Only in B (408 tensors, 31 patterns):
  Pattern                                           Tensors       Bytes
  block.{N}.mlp.mlp{N}_weight.blocks                     24   17.80 GiB
  model.layers.{N}.mlp.experts.gate_up_proj_blocks       12   11.87 GiB
  model.layers.{N}.mlp.experts.down_proj_blocks          12    5.93 GiB
    … showing 3 of 31 (limit 3)

  Dtype/shape differences (384 tensors, 13 patterns):
  Pattern                                           Tensors     A Bytes     B Bytes      Δ Bytes
  block.{N}.mlp.mlp{N}_weight.blocks                     48    8.90 GiB   35.60 GiB   +26.70 GiB
  model.layers.{N}.mlp.experts.gate_up_proj_blocks       24    5.93 GiB   23.73 GiB   +17.80 GiB
  model.layers.{N}.mlp.experts.down_proj_blocks          24    2.97 GiB   11.87 GiB    +8.90 GiB
    … showing 3 of 13 (limit 3)

  ──────────────────────────────────────────────────────────────────────
  only-A: 0 tensors (0 patterns) | only-B: 408 tensors (31 patterns) | differ: 384 tensors (13 patterns)

$ hf-fm diff-config openai/gpt-oss-20b openai/gpt-oss-120b
  A: openai/gpt-oss-20b
  B: openai/gpt-oss-120b

  Field              A                                                              B
  num_hidden_layers  24                                                             36
  layer_types        sliding_attention, full_attention, sliding_attention, full_a…  sliding_attention, full_attention, sliding_attention, full_a…

  2 of 23 fields differ

$ hf-fm inspect meta-llama/Llama-3.2-3B --cached --check-gpu --context 32768
  ...existing tensor table...

  Model weights:  5.98 GiB  (BF16, 3.21B params)
  KV cache @ ctx=32768:  3.50 GiB  (BF16)
  Total:          9.48 GiB  (weights + KV)
  GPU 0:          NVIDIA GeForce RTX 5060 Ti — 15.93 GiB VRAM
                  free: 13.68 GiB, used: 2.25 GiB
  Fit:            ✓ 4.20 GiB headroom (weights + KV; runtime extra)
  Spilling:       not sampled (platform supports detection)

Inspect reads tensor metadata via HTTP Range requests — no weight data downloaded: 2 requests per .safetensors file, a handful (reported live on the Source: line, e.g. remote (6 range requests, 136.0 KiB fetched)) per NumPy .npz archive (remote NPZ since v0.11.0), a similar handful per .gguf file (remote GGUF since v0.11.2, e.g. remote (30 range requests, 1.75 MiB fetched) on an 84 MiB quantized model, --tree/--dtypes included), and likewise per .pth checkpoint (remote PTH since v0.11.4, e.g. remote (12 range requests, 113.7 KiB fetched) on a 364 MiB PyTorch state_dict — only the ZIP-archived data.pkl pickle stream is read, never the tensor-data files). When a repo has many tensor files, --list prints a numbered table (pass the # back as the FILE argument) and --pick chooses interactively, narrowing first by a case-insensitive substring — both cover every format inspect reads (.safetensors / .gguf / .npz / .pth). The --tree flag shows the hierarchical namespace with numeric sibling groups auto-collapsed to [0..N] for structural discovery. The --group-by <PATTERN> flag rolls tensors up into a MATCHED / OTHER byte split against a glob pattern instead — the tool for "what fraction of this file is the MoE expert weights" (blk.*.ffn_*_exps.weight on a GGUF checkpoint), with a per-MoE-layer expert cost line once the matched tensor names carry a single, unambiguous layer index. Add --cache-headers to persist a remote header to a local sidecar, keyed on (repo, revision, filename, etag), so a repeat inspect of the same file is free (Source: cached header (age: ...)) — off by default, so a plain remote inspect never touches local disk without it. The --check-gpu flag adds a one-line GPU-fit verdict using hypomnesis (NVML on Linux/Windows, DXGI on Windows); composes with --json. Add --context N to fold in the KV cache at a context length and get a real weights + KV verdict — the difference between "fits" and "OOM at token 8000" on a consumer card. The estimate is parameter-driven from the model's config.json (GQA, sliding-window, and hybrid Mamba/attention all handled; MLA is detected and skipped); see the FAQ entry on GPU fit for the formula and its limitations. Diff compares tensor names, dtypes, and shapes between any two models (remote or cached); --dtypes swaps the per-tensor body for a side-by-side per-dtype histogram with a signed Δ Size column — the high-leverage view for scaled-sibling pairs. --collapse groups only-A / only-B / dtype-shape-differences by numeric-segment pattern (model.layers.{N}.mlp.gate_proj.weight) into a Pattern / Tensors / Bytes table, the built-in counterpart to the jq recipe below. See the FAQ entry on comparing two models for that jq recipe, which uses the byte_count field in --json output for cases --collapse's digit-run heuristic doesn't cover. diff-config complements diff at the architecture level: a field-by-field comparison of config.json (layer counts, hidden size, GQA/sliding-window/hybrid-layout fields) instead of tensor headers — the tool to reach for when diff --dtypes shows a size jump and you want to know why (more layers? wider hidden dim? a different attention pattern?) without eyeballing two raw config.json files.

Quant fit planning

$ hf-fm quants google/gemma-2-2b-it
Searching for quant siblings of google/gemma-2-2b-it...
8 repos found, 0 verified via GGUF backlink

  ARTIFACT                                     SIZE  REPO                                      BITS
  gemma-2-2b-it.IQ1_S.gguf               793.61 MiB  MaziyarPanahi/gemma-2-2b-it-GGUF          ~1.6
  gemma-2-2b-it.Q2_K.gguf                  1.15 GiB  MaziyarPanahi/gemma-2-2b-it-GGUF          ~2.6
  gemma-2-2b-it-IQ3_M.gguf                 1.30 GiB  bartowski/gemma-2-2b-it-GGUF              ~3.7
  ...
  gemma-2-2b-it-f32.gguf                   9.74 GiB  bartowski/gemma-2-2b-it-GGUF              ~32.0

$ hf-fm quants google/gemma-2-2b-it --fits 9.5GiB
...
  gemma-2-2b-it.fp16.gguf                  4.88 GiB   4.88 GiB  full GPU
  gemma-2-2b-it-f32.gguf                   9.74 GiB          —  does not fit (no MoE experts to offload)

quants aggregates a base model's quant sibling repos — discovered by searching the Hub for the base model's short name and keeping results whose repo ID contains it, cross-checked against each .gguf candidate's own general.source.url / general.base_model.*.repo_url metadata backlink where present. No dedicated Hub endpoint exists for "sibling repos", so this is best-effort: a candidate with no backlink, or whose backlink check itself failed (network error, gated repo), is still listed rather than silently dropped — only an explicit backlink mismatch excludes a candidate. --fits <SIZE> adds a RESIDENT/PLAN column pair; candidates that already fit under budget are never inspected (full GPU, no header fetch), and only over-budget MoE GGUF files trigger a header fetch to compute a --n-cpu-moe N offload plan against the internal blk.*.*_exps.weight expert-tensor pattern — the plan a naive size-vs-budget comparison would miss, since a checkpoint 40% over the raw VRAM budget can still be viable once its expert tensors (touched only a few times per token) move to system RAM. --reserve <SIZE> carves out headroom (KV cache, runtime) from the budget before the comparison. See Pick a quant that fits before you download it for the walkthrough this feature was built from, and the FAQ for the underlying inspect --group-by rollup.

Disk usage

$ hf-fm du
   #        SIZE  REPO                                             FILES
   1    5.10 GiB  google/gemma-2-2b-it                                 8
   2    2.80 GiB  EleutherAI/pythia-1.4b                              12  ●
   3    1.20 GiB  google/gemma-scope-2b-pt-res                         3
  ─────────────────────────────────────────────────────────────────────────────
   9.10 GiB  total (3 repos, 23 files)
  ● = partial downloads

$ hf-fm du 2
  EleutherAI/pythia-1.4b:

   #        SIZE  FILE
   1    2.50 GiB  model-00001-of-00002.safetensors
   2    0.26 GiB  model-00002-of-00002.safetensors
   ...
  ──────────────────────────────────────────────────────────────────
   2.80 GiB  total (12 files)

  ● partial downloads — run `hf-fm status EleutherAI/pythia-1.4b` for details

$ hf-fm du --age
   #        SIZE  REPO                                             FILES  AGE
   1    5.10 GiB  google/gemma-2-2b-it                                 8  2 days ago
   2    2.80 GiB  EleutherAI/pythia-1.4b                              12  45 days ago     ●
   3    1.20 GiB  google/gemma-scope-2b-pt-res                         3  3 months ago
  ─────────────────────────────────────────────────────────────────────────────────────────
   9.10 GiB  total (3 repos, 23 files)
  ● = partial downloads

$ hf-fm du --tree
  ├── google/gemma-2-2b-it          5.10 GiB  (8 files)
  │   ├── model-00001-of-00002.safetensors  2.50 GiB
  │   ├── model-00002-of-00002.safetensors  2.60 GiB
  │   └── config.json                          856 B
  ├── EleutherAI/pythia-1.4b        2.80 GiB  (12 files)  ●
  │   └── ...
  └── google/gemma-scope-2b-pt-res  1.20 GiB  (3 files)
      └── ...
  ─────────────────────────────────────────────────────────
   9.10 GiB  total (3 repos, 23 files)
  ● = partial downloads

$ hf-fm cache path google/gemma-2-2b-it
/home/user/.cache/huggingface/hub/models--google--gemma-2-2b-it/snapshots/abc1234

Library quick start

let outcome = hf_fetch_model::download(
    "google/gemma-2-2b-it".to_owned(),
).await?;

println!("Model at: {}", outcome.inner().display());

Filter, progress, auth, and more via the builder — see Configuration.

Documentation

Topic
CLI Reference All subcommands, flags, and output examples
FAQ Common questions — installation, auth, cache location, discovery, errors
Inspect tutorial Walkthrough: read tensor metadata, size, and architecture without downloading weights
Cache tutorial Walkthrough: see what the cache holds, then reclaim disk space safely (dustatuscache gc)
Case studies Real investigations where inspect did the diagnostic work — per-layer shape variation, OOM forensics from a crash log
Search Comma filtering, --exact, model card metadata
Configuration Builder API, presets, progress callbacks
Architecture How hf-fetch-model relates to hf-hub and candle-mi
Diagnostics --verbose output, tracing setup for library users
Upstream differences Where hf-fetch-model diverges from Python huggingface_hub/hf_transfer
Candle example Inspect tensor layouts before downloading — for candle users
Changelog Release history and migration notes

Used by

  • candle-mi — Mechanistic interpretability toolkit for language models

License

Licensed under either of Apache License, Version 2.0 or MIT License at your option.

Development

  • Exclusively developed with Claude Code (dev)
  • Git workflow managed with Fork
  • All code follows CONVENTIONS.md, derived from Amphigraphic-Strict's Grit — a strict Rust subset designed to improve AI coding accuracy.
  • CI gates every push and PR: cargo fmt --check, clippy --all-targets --all-features -D warnings, and the test suite on Linux and Windows, plus a cargo audit security pass. Vulnerability disclosure: see SECURITY.md.