spider_firewall 2.38.0

Firewall to use for Spider Web Crawler.
# spider_firewall

A Rust library to shield your system from malicious and unwanted websites by categorizing and blocking them.

## Installation

Add `spider_firewall` to your Cargo project with:

```sh
cargo add spider_firewall
```

## Size Tiers

The `small` tier is enabled by default. Enable `medium` or `large` for broader coverage — each tier includes all sources from the tier(s) below it.

| Tier | FST Size | Focus | Feature Flag |
|------|----------|-------|--------------|
| **small** (default) | ~13 MB | Ads, tracking, malware, phishing, scams, adult/porn | `small` |
| **medium** | ~26 MB | + ransomware, fraud, abuse, threat intel, extended phishing | `medium` |
| **large** | ~52 MB | + redirect/typosquatting, extended ads/tracking, full URLhaus | `large` |

```toml
# Default — small tier, all categories:
spider_firewall = "2.38"

# Medium tier:
spider_firewall = { version = "2.38", features = ["medium"] }

# Large tier:
spider_firewall = { version = "2.38", features = ["large"] }

# Small tier, only bad + ads (no tracking/gambling):
spider_firewall = { version = "2.38", default-features = false, features = ["default-tls", "bad", "ads", "small"] }
```

## Category Features

Categories can be toggled independently (all enabled by default):

| Feature | Description |
|---------|-------------|
| `bad` | Malware, phishing, scams, fraud, ransomware, abuse, plus the adult and category lists below |
| `ads` | Advertising domains |
| `tracking` | Tracking and analytics domains |
| `gambling` | Gambling domains |
| `ip` | Known-bad IPv4 network ranges (Spamhaus DROP) — opt-in, see [IP blocking](#ip-blocking) |

## Categories and feeds

Since 2.38 the `bad` feature fills three buckets instead of one. Only `CAT_BAD` is a threat verdict, and it is the only bit `is_bad_website_url` reads. `is_url_bad` still matches any bucket.

| Bucket | Read with | Feeds |
|--------|-----------|-------|
| `CAT_BAD` | `is_bad_website_url` | spider-rs/bad_websites; ShadowWhisperer Malware, Scam and Typo; Block List Project malware, phishing and scam; URLhaus filter; malware-filter phishing; CyberHost malware; romainmarcoux malicious-domains. Medium adds Block List Project ransomware, fraud and abuse; Phishing.Database; phishdestroy destroylist; durablenapkin scamblocklist; HaGeZi TIF mini; ThreatFox; CERT Polska; PhishIndex; malicious-domains tiers B and C; phishunt. Large adds HaGeZi TIF and the full URLhaus hostfile. |
| `CAT_ADULT` | `is_adult_website_url` | StevenBlack porn; ShadowWhisperer Adult |
| `CAT_LISTED` | `is_listed_website_url` | StevenBlack unified hosts, which mixes adware and malware; ShadowWhisperer AI, Apple, Chat, DNS, Dynamic, Junk, Remote, Risk, Shock, Top_Level, Tunnels, UrlShortener and the Wild_ lists other than ads and tracking. Medium adds maltrail suspicious and OISD small. Large adds the Block List Project redirect list. |

A feed entry that is an explicit ICANN public suffix (`com.cn`, `co.uk`) is dropped at build time, and lookups never test one, so a listing like that cannot block a whole zone.

## Usage

### Checking for Bad Websites

You can check if a website is part of the bad websites list using the `is_bad_website_url` function.

```rust
use spider_firewall::is_bad_website_url;

fn main() {
    let u = url::Url::parse("https://badwebsite.com").expect("parse");
    let blocked = is_bad_website_url(u.host_str().unwrap_or_default());
    println!("Is blocked: {}", blocked);
}
```

### Adding a Custom Firewall

You can add your own websites to the block list using the `define_firewall!` macro. This allows you to categorize new websites under a predefined or new category.

```rust
use spider_firewall::is_bad_website_url;

// Add "bad.com" to a custom category.
define_firewall!("unknown", "bad.com");

fn main() {
    let u = url::Url::parse("https://bad.com").expect("parse");
    let blocked = is_bad_website_url(u.host_str().unwrap_or_default());
    println!("Is blocked: {}", blocked);
}
```

### Example with Custom Ads List

You can specify websites to be blocked under specific categories such as "ads".

```rust
use spider_firewall::is_ad_website_url;

// Add "ads.com" to the ads category.
define_firewall!("ads", "ads.com");

fn main() {
    let u = url::Url::parse("https://ads.com").expect("parse");
    let blocked = is_ad_website_url(u.host_str().unwrap_or_default());
    println!("Is blocked: {}", blocked);
}
```

### IP blocking

Enable the opt-in `ip` feature to also block known-bad IPv4 network ranges. The ranges are
embedded at build time from the [Spamhaus DROP](https://www.spamhaus.org/drop/) list and matched
via longest-prefix (binary) search. IPv6 currently always returns `false`.

```toml
spider_firewall = { version = "2.38", features = ["ip"] }
```

```rust
use spider_firewall::{is_bad_ip, is_bad_ip_str};

fn main() {
    assert!(!is_bad_ip("8.8.8.8".parse().unwrap()));   // legitimate hosts are not blocked
    let blocked = is_bad_ip_str("1.2.3.4");            // convenience: parses the string
    println!("Is blocked: {}", blocked);
}
```

The feed is rate-limited (~1 download/day) and revocable, so it is fetched **non-fatally** at build
time — a failed or rate-limited fetch yields zero ranges rather than breaking the build, and emits a
`cargo:warning` reporting the embedded range count (or that IP blocking is inactive).

For production builds where IP blocking must not silently disable on a rate-limited fetch, set
`SPIDER_FIREWALL_IP_STRICT=1`: the build then **fails loudly** if the DROP fetch returns zero ranges
(instead of shipping with IP blocking inactive). Retry once the ~1/day limit resets.

**Attribution:** IP range data is provided by [The Spamhaus Project](https://www.spamhaus.org) under
the [Spamhaus DROP terms](https://www.spamhaus.org/drop/) (free for any use, attribution required).
© The Spamhaus Project.

## Blocklist Sources

### Small (default)

| Source | Categories | License |
|--------|-----------|---------|
| [ShadowWhisperer/BlockLists](https://github.com/ShadowWhisperer/BlockLists) | bad, ads, tracking, gambling | MIT |
| [badmojr/1Hosts Lite](https://github.com/badmojr/1Hosts) | ads, tracking | MPL-2.0 |
| [spider-rs/bad_websites](https://github.com/spider-rs/bad_websites) | bad | MIT |
| [Steven Black Unified Hosts](https://github.com/StevenBlack/hosts) | bad | MIT |
| [Block List Project — Malware](https://github.com/blocklistproject/Lists) | bad | MIT |
| [Block List Project — Phishing](https://github.com/blocklistproject/Lists) | bad | MIT |
| [Block List Project — Scam](https://github.com/blocklistproject/Lists) | bad | MIT |
| [URLhaus Filter (domains)](https://malware-filter.gitlab.io/malware-filter/urlhaus-filter-domains.txt) | bad | CC0/MIT |
| [Steven Black Hosts — Porn](https://github.com/StevenBlack/hosts) | bad (adult/porn) | MIT |
| [malware-filter — Phishing](https://gitlab.com/malware-filter/phishing-filter) | bad (phishing) | CC0/MIT |

### Medium (adds)

| Source | Categories | License |
|--------|-----------|---------|
| [Block List Project — Ransomware](https://github.com/blocklistproject/Lists) | bad | MIT |
| [Block List Project — Fraud](https://github.com/blocklistproject/Lists) | bad | MIT |
| [Block List Project — Abuse](https://github.com/blocklistproject/Lists) | bad | MIT |
| [Phishing.Database — Active Domains](https://github.com/mitchellkrogza/Phishing.Database) | bad | MIT |
| [Stamparm/maltrail — Suspicious](https://github.com/stamparm/maltrail) | bad | MIT |
| [phishdestroy/destroylist — Primary Active](https://github.com/phishdestroy/destroylist) | bad (phishing/scam) | MIT |

### Large (adds)

| Source | Categories | License |
|--------|-----------|---------|
| [Block List Project — Redirect](https://github.com/blocklistproject/Lists) | bad | MIT |
| [Block List Project — Tracking](https://github.com/blocklistproject/Lists) | tracking | MIT |
| [Block List Project — Ads](https://github.com/blocklistproject/Lists) | ads | MIT |
| [Stamparm/maltrail — Malware](https://github.com/stamparm/maltrail) | bad | MIT |
| [abuse.ch URLhaus Hostfile](https://urlhaus.abuse.ch/downloads/hostfile/) | bad | CC0 |

## Build Time

The initial build can take longer, approximately 5-10 minutes, as it may involve compiling dependencies and generating necessary data files.

## Contributing

Contributions and improvements are welcome. Feel free to open issues or submit pull requests on the GitHub repository.

## License

This project is licensed under the MIT License.