Skip to main content

Crate cum_rs

Crate cum_rs 

Source
Expand description
cum-rs logo

§CUM

Crates.io Docs.rs npm PyPI CI Clippy MIT License MSRV

Claude Unmarking Machine: a multilanguage Rust crate that removes AI-provider watermarks from text, images, and documents. Works regardless of provider (Claude, OpenAI, Gemini, Grok, open-LLM). All processing is 100% local: no data leaves your machine.

crab dancing

The cum binary, cheerfully evicting zero-width gremlins from your prose.

§🤔 What is happening here, exactly?

So you copy-pasted some text from an AI. Totally normal. You are not doing anything illegal. Probably.

The bad news: every major LLM provider stuffs your output full of invisible Unicode graffiti so they can identify their own generation later. It is like spray-painting “CLAUDE WAS HERE” on every wall, except the paint is literally invisible and you cannot see it without specialized equipment.

The good news: we have the specialized equipment. And it is written in Rust, so it is blazingly fast.

LayerWhat lurks in the shadowsWhat we do about it
A: UnicodeZWSP, bidi controls, tag chars, variation selectors, private-use codepoints, dash homoglyphs (U+2011 non-breaking hyphen, en-dash, em-dash, etc.), punctuation homoglyphs (curly quotes, ellipsis U+2026, etc.), mathematical alphanumerics (𝑨→A), Braille blank U+2800Deterministic, lossless exorcism 🧹
File: MetadataC2PA manifests, EXIF, XMP, document properties, the digital equivalent of a tracking ankle braceletStripped from PNG, JPEG, WebP, SVG, PDF, DOCX, ODT, HTML, Markdown
B: StatisticalToken-sampling watermarks (SynthID-Text, KGW), watermarks baked into the actual word choicesBest-effort via stochastic synonym replacement (~30 000-entry Moby Thesaurus II + ES/FR/DE/AR multilingual support)
PixelSynthID-Image, StegaStamp, Tree-Ring, StableSignature: pixel-domain perturbations invisible to the eyeDecode→raw RGBA→lossless PNG re-encode via pixel-scrub feature

Fun fact: some of those invisible characters are technically in the Unicode “Tag” block, which was originally designed for plane tickets in 1997 and then deprecated. AI providers found a new use for them. The Unicode Consortium is presumably very proud.

§🖥️ CLI: for the terminal warriors

Install the cum binary. Yes, that is the name. Yes, the authors are aware. Yes, it compiles clean:

cargo install cum-rs --features rust-binary

Then run:

# Clean a Markdown file. The AI left crumbs everywhere.
cum clean report.md

# Clean a PNG: yes, even images can be watermarked now. We live in a society.
cum clean logo.png --output logo_clean.png

# Inspect your text for hidden nonsense, formatted as JSON for maximum nerd points
cum inspect --json article.txt

# Pipe from stdin like a true Unix philosopher
cat suspicious.txt | cum clean --stdin

# The aggressive mode: also replaces Cyrillic А with Latin A (sneaky!)
cum clean --aggressive sneaky_essay.txt

See CLI.md for the full command reference. It has tables and everything.

§🦀 Rust: the fast one

Available on crates.io. Because of course it is. Full API docs: RUST.md.

FeatureDescription
(default)Pure-Rust core: clean, inspect, all media formats. Zero drama.
cliClap CLI module required for the binary. Comes with an ASCII banner, because we have standards.
rust-binaryEnables the cum binary. Ship it.
pixel-scrubAdds pixel_scrub::scrub_pixels(): decode-then-re-encode any raster image to strip pixel-domain watermarks. Uses the image crate. Not available on wasm32.
pythonPython extension via PyO3. For the snake people. 🐍
nodeNode.js native add-on via napi-rs. For the node_modules enjoyers. 🟩
wasmWASM bindings. Run watermark removal in the browser. Why? Because we can.

§⚡ Quick Start

use cum_rs::cleaner::clean;
use cum_rs::types::MediaHint;

// "Hello​ world!": looks innocent, contains a ZWSP and a BOM. Rude.
let dirty = "Hello\u{200B} world\u{FEFF}!";
let output = clean(dirty.as_bytes(), Some(MediaHint::Text)).unwrap();

// Now it is just "Hello world!" like a normal person wrote it
assert_eq!(String::from_utf8(output.bytes).unwrap(), "Hello world!");
assert_eq!(output.stats.removed_count, 2); // Two gremlins evicted. You're welcome.

§🌐 WASM

Because if you are going to remove watermarks, you might as well do it at the speed of JavaScript. (Do not worry, the Rust core still does the actual work.) See WASM.md and the live demo: examples/unmark/.

§🐍 Python

import cum_rs

result = cum_rs.clean_text("Hello\u200b world\ufeff!")
print(result.cleaned)        # "Hello world!"
print(result.removed_count)  # 2
# The AI's fingerprints have been thoroughly wiped. You were never here.

See PYTHON.md for the full binding reference.

§🟩 Node.js

const { cleanText } = require("cum-rs");

const result = cleanText("Hello\u200b world\ufeff!");
console.log(result.cleaned); // "Hello world!"
console.log(result.removedCount); // 2
// node_modules is 9000 packages deep but THIS one actually does something useful

See NODE.md for the full binding reference.

§🎲 Stochastic Enhancer

CUM includes a best-effort countermeasure against Layer B statistical watermarks (SynthID, KGW) by stochastically replacing eligible words with semantically equivalent synonyms. This modifies the raw byte pairs chosen by the LLM’s token-sampler, disrupting the periodic watermark signal.

use cum_rs::stochastic::StochasticEnhancer;

let enhancer = StochasticEnhancer::new(0.5); // 50% substitution chance
let output = enhancer.enhance("The chaos governs the universe.");
println!("Substituted {} words", output.words_substituted);
println!("{}", output.text);

The English table is the full Moby Thesaurus II (~30 000 head words, ~2.5 M synonym tokens, public domain) embedded at compile time as a PHF map: O(1) lookups, zero runtime I/O. A two-tier fallback uses /usr/share/dict/ system wordlists for same-length substitution when no Moby entry exists.

§🌍 Multilingual Support

Pass a LanguageHint or let detect_language() auto-detect:

use cum_rs::stochastic::{StochasticEnhancer, LanguageHint, detect_language};

// Explicit language
let es = StochasticEnhancer::with_language(LanguageHint::Spanish);
let out = es.enhance("El texto contiene marcas invisibles");
println!("{}", out.text);
println!("Language: {}", out.language.as_bcp47()); // "es"

// Auto-detect
let lang = detect_language("يحتوي النص على علامات مائية");
assert_eq!(lang, LanguageHint::Arabic);
LanguageIdentifierEntries
EnglishLanguageHint::English~30 000 (Moby Thesaurus II, public domain)
SpanishLanguageHint::Spanish35
FrenchLanguageHint::French32
GermanLanguageHint::German32
ArabicLanguageHint::Arabic31

§🖼️ Pixel-Domain Scrubbing

Some AI image generators embed invisible watermarks by perturbing pixel values at a level imperceptible to humans but detectable by a matched neural decoder (SynthID-Image, StegaStamp, Tree-Ring, StableSignature).

Enable the pixel-scrub feature to strip them:

cum-rs = { version = "0.2.1", features = ["pixel-scrub"] }
#[cfg(feature = "pixel-scrub")]
{
    use cum_rs::pixel_scrub::scrub_pixels;

    let png_bytes = std::fs::read("watermarked.png").unwrap();
    let clean = scrub_pixels(&png_bytes).unwrap();
    std::fs::write("clean.png", &clean).unwrap();
    // Output is always lossless PNG regardless of input format
}

Supported input: PNG, JPEG, WebP. Output is always PNG (lossless). Not available on wasm32.

How it works: Decode any supported raster image to raw RGBA pixels, then re-encode from scratch as PNG. The watermark signal lives in the original compression stream’s state; a fresh encode from raw pixels cannot carry it.

§🚨 Disclaimer (the responsible adult part)

Layer A (Unicode scrubbing) and file metadata stripping are fully deterministic and lossless: every modification is logged in stats. You can see exactly what changed.

Layer B (statistical watermarks) lives inside the actual word choices. No tool can guarantee removal. The only real fix is to rewrite the content in your own words. Think of Layer B as the AI watermarking the vibes of the text, not just the characters.

This crate is for content you own: research, privacy hygiene, and understanding what AI providers are doing to your outputs. Read ETHICS.md before doing anything exciting.

§📄 License

MIT. Do what you want. Just do not be evil about it.

§Rust Usage Guide: cum-rs

§Installation

[dependencies]
cum-rs = "0.2.1"

§Text Cleaning (Layer A)

use cum_rs::unicode::{clean_text, CleanOpts};

let dirty = "Hello\u{200B} world\u{FEFF}!";
let opts = CleanOpts::safe();
let (clean, stats) = clean_text(dirty, &opts).unwrap();
assert_eq!(clean, "Hello world!");
println!("Removed: {}", stats.removed_count);

§Text Inspection

use cum_rs::unicode::{inspect_text, InspectOpts};

let report = inspect_text("Hello\u{200B}!", &InspectOpts::default()).unwrap();
for hit in &report.hits {
    println!("{}: {} occurrences ({})", hit.label, hit.count, hit.confidence.as_str());
}

§Stochastic Enhancement (Layer B)

use cum_rs::stochastic::StochasticEnhancer;

// Create an enhancer with 70% substitution probability
let enhancer = StochasticEnhancer::new(0.7);

let output = enhancer.enhance("The chaos governs the universe.");
println!("Enhanced text: {}", output.text);
println!("Words substituted: {}", output.words_substituted);

§Image Metadata Stripping

use cum_rs::image_meta::clean_image;
use std::fs;

let png = fs::read("input.png").unwrap();
let cleaned = clean_image(&png).unwrap();
fs::write("output.png", &cleaned).unwrap();

§Unified API (auto-detect format)

use cum_rs::cleaner::{clean, inspect};

let bytes = std::fs::read("draft.docx").unwrap();
let output = clean(&bytes, None).unwrap();
std::fs::write("draft.cleaned.docx", &output.bytes).unwrap();
println!("Chunks removed: {}", output.stats.metadata_chunks_removed);

§Feature Flags

FeatureEnables
pythonPyO3 extension module
nodenapi-rs Node.js add-on
wasmwasm-bindgen WASM bindings

§WASM Usage Guide: cum-rs

§Building

wasm-pack build --target web --features wasm
# or via Trunk (see examples/unmark/):
cd examples/unmark && trunk serve

§Browser Usage (vanilla JS)

<script type="module">
  import init, {
    clean_text_wasm,
    inspect_text_wasm,
    clean_bytes_wasm,
  } from "./pkg/cum_rs.js";

  await init();

  const result = JSON.parse(clean_text_wasm("Hello\u200b world\ufeff!"));
  console.log(result.cleaned); // "Hello world!"
  console.log(result.removed_count); // 2
</script>

§Yew / Rust front-end

See the live example in examples/unmark/: a split-panel Yew 0.22 CSR app that accepts text or file upload on the left and shows the cleaned output on the right with Copy / Download buttons.

§CORS

cum-rs operates entirely client-side (pure computation on input bytes). No network requests are made.

§cum-rs: Claude Unmarking Machine

A multilanguage watermark removal crate for AI-generated content. The core is pure Rust; Python, Node.js, and WASM bindings are built via Cargo feature flags.

Re-exports§

pub use container as container_meta;
pub use image::meta as image_meta;
pub use text::stochastic;
pub use text::unicode;
pub use image::pixel_scrub;pixel-scrub and non-WebAssembly

Modules§

cleaner
Unified Watermark Cleaner
clicli
CLI layer
container
Container Metadata Watermark Removal
error
Error Types
image
Image Watermark Removal
text
Text Watermark Removal
types
Public Types