ratdmp 0.1.0

Fast streaming ASCII/UTF-16LE string extractor for raw memory dump (.dmp) files, with simple noise filtering for byte-fill/heap-fill patterns. Never loads the whole file into RAM.
Documentation

ratdmp

Fast, streaming ASCII / UTF-16LE string extractor for raw memory dump (.dmp) files — with simple, conservative noise filtering for byte-fill / heap-fill patterns. Written for malware analysis / DFIR / memory-forensics workflows.

Why

Running something like Unix strings on a multi-gigabyte memory dump usually means: (1) reading the whole file into RAM, and (2) getting flooded with junk like AAAAAAAA... or ABABABAB... from Windows heap padding / byte-fill patterns, alongside the IOCs you actually care about.

ratdmp fixes both:

  • Streaming. Reads the file in 32MB chunks (extract_strings_from_file / extract_strings_from_reader), never loading the entire dump into memory. A .dmp file can be many gigabytes — this crate's peak memory use does not scale with file size.
  • Single-pass, dual-encoding. ASCII and UTF-16LE runs are detected in the same scan over the buffer, not two separate passes.
  • Conservative noise filtering. Runs that are pure period-1 (aaaaaaaa...) or period-2 (ababab...) repeats, above a length threshold, are dropped as low-information byte-fill noise. Short repeats (like "0000") are deliberately kept, since they can be real data (PINs, years, etc.) — the filter only removes patterns long enough to be confidently noise, never sacrificing recall for short strings.

Usage

use ratdmp::extract_strings_from_file;

fn main() -> std::io::Result<()> {
    let strings = extract_strings_from_file("memory.dmp", /* min_len */ 4, /* max_strings */ 200_000)?;
    for s in &strings {
        println!("[{}] offset={} {:?}", s.encoding, s.offset, s.text);
    }
    Ok(())
}

Already have the bytes in memory instead of a file on disk? Use extract_strings(&data, min_len, max_strings). Reading from any std::io::Read (a socket, a pipe, a decompression stream, …)? Use extract_strings_from_reader(reader, min_len, max_strings, estimated_len).

Each result is an ExtractedString { offset: u64, encoding: &'static str, text: String } (encoding is "ascii" or "utf16le"), sorted by offset. The struct derives serde::Serialize if you want to hand it off as JSON.

What this crate does not do

  • No .dmp/MDMP structural parsing (no stream/module/thread table lookups) — it treats the file as a raw byte payload and scans the whole thing. This is intentional: it makes the extractor independent of the specific minidump format version or which tool produced it.
  • No IOC classification (URLs, IPs, wallet addresses, etc.) — this crate's only job is turning bytes into candidate strings. Feed its output into whatever classifier fits your pipeline.
  • No entropy-based or statistical filtering beyond the two explicit repeat patterns described above — the goal is to reject only what's unambiguously noise, not to be a general-purpose deduplication or scoring engine.

Minimum supported Rust version

1.70 (uses const fn with loops for the printable-byte lookup table).

License

Licensed under either of

at your option.

Contribution

Unless you explicitly state otherwise, any contribution intentionally submitted for inclusion in the work by you, as defined in the Apache-2.0 license, shall be dual licensed as above, without any additional terms or conditions.