# urlsup ![CI][build_badge] [![Code Coverage][coverage_badge]][coverage_report]
`urlsup` (_urls up_) finds URLs in files and checks whether they are up by
making a `GET` request and checking the response status code. This tool is
useful for lists, repos or any type of project containing URLs that you want to
be up.
It's written in Rust (stable) and executes the requests async in multiple
threads, making it very fast. **Uses browser-like HTTP client behavior with
automatic protocol negotiation and reliable connection handling.** This in
combination with its ease of use makes it the perfect tool for your CI pipeline.
This project is a slim version of
[`awesome_bot`](https://github.com/dkhamsing/awesome_bot) but significantly faster.
<img src="banner.png" alt="Dotfiles Banner" width="100%" style="display: block; margin: 0 auto;">
## š What's New in v2.0
**urlsup v2.0** introduces a modern CLI design with breaking changes for better usability:
### š Renamed Flags (Breaking Changes)
- `--white-list` ā `--allowlist` (modern terminology)
- `--allow` ā `--allow-status` (clearer naming)
- `--threads` ā `--concurrency` (industry standard)
- `--file-types` ā `--include` (shorter, clearer)
### ⨠New Features
- **š Configuration Files**: TOML-based config with automatic discovery
- **š¤ Output Formats**: JSON support for automation (`--format json`)
- **š Progress Reporting**: Beautiful progress bars with real-time stats
- **š Advanced Filtering**: Regex-based URL exclusion patterns
- **š Retry Logic**: Configurable retry attempts with exponential backoff
- **ā±ļø Rate Limiting**: Built-in request throttling
- **š Quiet/Verbose Modes**: Better control over output verbosity
- **šØ Enhanced Error Handling**: Comprehensive error types with context
- **ā” Browser-like Behavior**: HTTP client that behaves like web browsers for maximum compatibility
- **š Performance Analysis**: Built-in memory monitoring and optimization suggestions
- **š HTML Dashboard**: Rich visual reporting with charts and performance metrics
## š Table of Contents
- [š Usage](#-usage)
- [š Examples](#-examples)
- [Basic File Checking](#basic-file-checking)
- [Directory Processing](#directory-processing)
- [File Type Filtering](#file-type-filtering)
- [Advanced Options](#advanced-options)
- [Git Integration](#git-integration)
- [š¦ Installation](#-installation)
- [š Shell Completions Installation](#-shell-completions-installation)
- [āļø Configuration File](#ļø-configuration-file)
- [Configuration Discovery](#configuration-discovery)
- [š§ Advanced Features](#-advanced-features)
- [šÆ Failure Threshold](#-failure-threshold)
- [Retry Logic & Rate Limiting](#retry-logic--rate-limiting)
- [URL Exclusion Patterns](#url-exclusion-patterns)
- [š Progress Reporting](#-progress-reporting)
- [š Custom User Agent & Proxy Support](#-custom-user-agent--proxy-support)
- [š¤ Output Formats](#-output-formats)
- [š Performance Analysis](#-performance-analysis)
- [š HTML Dashboard](#-html-dashboard)
- [Verbose Logging](#verbose-logging)
- [š Security Features](#-security-features)
- [š Performance Features](#-performance-features)
- [ā” Browser-like HTTP Client](#-browser-like-http-client)
- [šØ Error Handling](#-error-handling)
- [š GitHub Actions](#-github-actions)
- [š ļø Development](#ļø-development)
## š Usage
```bash
CLI to validate URLs in files
Usage: urlsup [OPTIONS] [FILES]... [COMMAND]
Commands:
completion-generate Generate shell completions
completion-install Install shell completions to standard location
help Print this message or the help of the given subcommand(s)
Arguments:
[FILES]... Files or directories to check
Options:
-h, --help Print help
-V, --version Print version
Core Options:
-r, --recursive Recursively process directories
-t, --timeout <SECONDS> Connection timeout in seconds (default: 30)
--concurrency <COUNT> Concurrent requests (default: CPU cores)
Filtering & Content:
--include <EXTENSIONS> File extensions to process (e.g., md,html,txt)
--allowlist <URLS> URLs to allow (comma-separated)
--allow-status <CODES> Status codes to allow (comma-separated)
--exclude-pattern <REGEX> URL patterns to exclude (regex)
Retry & Rate Limiting:
--retry <COUNT> Retry attempts for failed requests (default: 0)
--retry-delay <MS> Delay between retries in ms (default: 1000)
--rate-limit <MS> Delay between requests in ms (default: 0)
--allow-timeout Allow URLs that timeout
--failure-threshold <PERCENT> Failure threshold - fail only if more than X% of URLs are broken (0-100)
Output & Verbosity:
-q, --quiet Suppress progress output
-v, --verbose Enable verbose logging
--format <FORMAT> Output format [default: text] [possible values: text, json, minimal]
--no-progress Disable progress bars
Network & Security:
--user-agent <AGENT> Custom User-Agent header
--proxy <URL> HTTP/HTTPS proxy URL
--insecure Skip SSL certificate verification
Configuration:
--config <FILE> Use specific config file
--no-config Ignore config files
Performance Analysis:
--show-performance Show memory usage and optimization suggestions
--html-dashboard <PATH> Generate HTML dashboard report
```
## š Examples
### Basic File Checking
```bash
# Check a single file
$ urlsup README.md
# Check multiple files
$ urlsup README.md CHANGELOG.md
# Check files with wildcards
$ urlsup docs/*.md
```
### Directory Processing
**Important**: `urlsup` treats files and directories differently:
- **Files**: Directly processed (e.g., `urlsup README.md`)
- **Directories**: Must use `--recursive` flag (e.g., `urlsup --recursive docs/`)
```bash
# ā This will fail with an error
$ urlsup docs/
error: 'docs/' is a directory. Use --recursive to process directories.
# ā
Process all files in a directory recursively
$ urlsup --recursive docs/
# ā
Process only specific file types
$ urlsup --recursive --include md,txt docs/
# ā
Process current directory recursively
$ urlsup --recursive .
```
### File Type Filtering
```bash
# Only check markdown and text files
$ urlsup --recursive --include md,txt .
# Only check web files
$ urlsup --recursive --include html,css,js website/
# Multiple extensions
$ urlsup --recursive --include md,rst,txt docs/
```
### Advanced Options
```bash
# Allow specific status codes
$ urlsup README.md --allow-status 403,429
# Set timeout and allow timeouts
$ urlsup README.md --allow-timeout -t 5
# Allowlist URLs (partial matches)
$ urlsup README.md --allowlist rust,crates
# Failure threshold - only fail if more than X% of URLs are broken
$ urlsup --recursive docs/ --failure-threshold 10 # Allow up to 10% failures
$ urlsup README.md --failure-threshold 0 # Strict mode - fail on any broken URL
# Combine recursive with filtering and options
$ urlsup --recursive --include md --allow-status 403 --timeout 10 docs/
# Use quiet mode for scripts
$ urlsup --quiet --recursive docs/
# Enable verbose output for debugging
$ urlsup --verbose README.md
# Use JSON output format
$ urlsup --format json README.md
# Exclude URLs with patterns
$ urlsup --exclude-pattern ".*\.local$" --exclude-pattern "^http://localhost.*" docs/
# Performance analysis and reporting
$ urlsup --show-performance README.md
$ urlsup --html-dashboard report.html docs/
```
### Git Integration
When using `--recursive`, `urlsup` automatically respects your `.gitignore` files:
```bash
# This will skip files/directories listed in .gitignore
$ urlsup --recursive .
# Examples of automatically ignored paths:
# - node_modules/
# - target/
# - .git/
# - *.log files
# - Any patterns in your .gitignore
```
This means you don't need to manually exclude build artifacts, dependencies, or other generated files.
## š¦ Installation
Install with `cargo` to run `urlsup` on your local machine.
```bash
cargo install urlsup
```
## š Shell Completions Installation
`urlsup` supports shell completions for bash, zsh, and fish. You can generate completions manually or use the built-in installation command for automatic setup.
### Automatic Installation (Recommended)
The `completion-install` command automatically installs shell completions to standard directories and provides setup instructions:
```bash
# Install bash completions
$ urlsup completion-install bash
ā
Shell completions installed successfully!
Completion installed to: /Users/user/.local/share/bash-completion/completions/urlsup
To enable bash completions, add this to your ~/.bashrc or ~/.bash_profile:
if [[ -d ~/.local/share/bash-completion/completions ]]; then
for completion in ~/.local/share/bash-completion/completions/*; do
[[ -r "$completion" ]] && source "$completion"
done
fi
Then restart your shell or run: source ~/.bashrc
# Install zsh completions
$ urlsup completion-install zsh
ā
Shell completions installed successfully!
Completion installed to: /Users/user/.local/share/zsh/site-functions/_urlsup
To enable zsh completions, add this to your ~/.zshrc:
if [[ -d ~/.local/share/zsh/site-functions ]]; then
fpath=(~/.local/share/zsh/site-functions $fpath)
autoload -U compinit && compinit
fi
Then restart your shell or run: source ~/.zshrc
You may also need to clear the completion cache: rm -f ~/.zcompdump*
# Install fish completions
$ urlsup completion-install fish
ā
Shell completions installed successfully!
Completion installed to: /Users/user/.config/fish/completions/urlsup.fish
Fish completions are automatically loaded from ~/.config/fish/completions/
Restart your shell or run: fish -c 'complete --erase; source ~/.config/fish/config.fish'
```
### Manual Installation
For manual installation or unsupported shells, generate the completion script and add it yourself:
```bash
# Generate completions for your shell
$ urlsup completion-generate bash > urlsup_completion.bash
$ urlsup completion-generate zsh > _urlsup
$ urlsup completion-generate fish > urlsup.fish
$ urlsup completion-generate powershell > urlsup_completion.ps1
$ urlsup completion-generate elvish > urlsup_completion.elv
# Then add to your shell's configuration manually
```
### Supported Shells
| bash | ā
Yes | ā
Yes | `~/.local/share/bash-completion/completions/urlsup` |
| zsh | ā
Yes | ā
Yes | `~/.local/share/zsh/site-functions/_urlsup` |
| fish | ā
Yes | ā
Yes | `~/.config/fish/completions/urlsup.fish` |
| PowerShell | ā Manual only | ā
Yes | Add to `$PROFILE` manually |
| Elvish | ā Manual only | ā
Yes | Add to `~/.elvish/rc.elv` manually |
**Note**: The `completion-install` command creates directories as needed and handles path resolution automatically. For PowerShell and Elvish, use the manual `completion-generate` command and follow the provided instructions.
## āļø Configuration File
`urlsup` supports TOML configuration files for managing complex setups. Place a `.urlsup.toml` file in your project root:
```toml
# .urlsup.toml - Project configuration for urlsup
timeout = 30
threads = 8 # Number of concurrent threads (maps to --concurrency CLI option)
allow_timeout = false
file_types = ["md", "html", "txt"]
# URL patterns to exclude (regex)
exclude_patterns = [
"^https://example\\.com/private/.*",
".*\\.local$",
"^http://localhost.*"
]
# URLs to allowlist
allowlist = [
"https://api.github.com",
"https://docs.rs"
]
# HTTP status codes to allow
allowed_status_codes = [403, 429]
# Advanced network settings
user_agent = "MyBot/1.0"
retry_attempts = 3
retry_delay = 1000 # milliseconds
rate_limit_delay = 100 # milliseconds between requests
failure_threshold = 10.0 # Allow up to 10% of URLs to fail
# Performance settings
use_head_requests = false # Use HEAD instead of GET for faster validation
# Security settings
skip_ssl_verification = false
proxy = "http://proxy.company.com:8080"
# Output settings
output_format = "text" # or "json" or "minimal"
verbose = false
# Performance analysis
show_performance = false # Show memory usage and optimization suggestions
```
### Configuration Discovery
`urlsup` searches for configuration files in this order:
1. `.urlsup.toml` in current directory
2. `.urlsup.toml` in parent directories (up to 3 levels)
3. Default configuration if no file found
CLI arguments always override configuration file settings.
## š§ Advanced Features
### šÆ Failure Threshold
Control when `urlsup` should fail based on the percentage of broken URLs:
```bash
# Only fail if more than 20% of URLs are broken
$ urlsup --recursive docs/ --failure-threshold 20
# Strict mode - fail on any broken URL (default behavior)
$ urlsup docs/ --failure-threshold 0
# Lenient mode for large documentation sets
$ urlsup --recursive . --failure-threshold 5 # Allow up to 5% failures
```
**Configuration file:**
```toml
# In .urlsup.toml
failure_threshold = 10.0 # Allow up to 10% failures
```
**Use Cases:**
- **Large documentation**: Prevent CI failures for 1-2 stale external links out of hundreds
- **External API monitoring**: Allow some endpoints to be temporarily down
- **Migration periods**: Gradually improve link quality without breaking builds
- **Third-party content**: Handle external links that may be occasionally unreachable
**Example output:**
```bash
$ urlsup --recursive docs/ --failure-threshold 15
# When within threshold:
ā
Failure rate 12.5% is within threshold 15.0% (5/40 URLs failed)
# When exceeding threshold:
ā Failure rate 17.5% exceeds threshold 15.0% (7/40 URLs failed)
```
### Retry Logic & Rate Limiting
Handle flaky networks and respect server limits:
```bash
# Configure via CLI (basic)
$ urlsup --timeout 60 README.md
# Configure via .urlsup.toml (advanced)
retry_attempts = 5
retry_delay = 2000
rate_limit_delay = 500
```
### URL Exclusion Patterns
Exclude URLs matching regex patterns:
```toml
# In .urlsup.toml
exclude_patterns = [
"^https://internal\\.company\\.com/.*", # Skip internal URLs
".*\\.local$", # Skip .local domains
"^http://localhost.*", # Skip localhost
"https://example\\.com/api/.*" # Skip API endpoints
]
```
### š Progress Reporting
Beautiful progress bars for large operations:
```bash
# Progress bars are enabled automatically for TTY terminals
$ urlsup --recursive docs/
# Output includes:
# ā [00:01:23] [āāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāā] 150/150 files processed
# ā [00:00:45] [āāāāāāāāāāāāāāāāāāāāāāāāāā ] 245/320 URLs validated (76% successful)
```
### š Custom User Agent & Proxy Support
```toml
# In .urlsup.toml
user_agent = "MyCompany/URLChecker 2.0"
proxy = "http://proxy.company.com:8080"
skip_ssl_verification = false # Set to true for internal/dev environments
```
### š¤ Output Formats
```bash
# Text output (default) - clean, colorful, emoji-based with grouping
$ urlsup README.md
ā
No issues found!
# JSON output for scripts and automation
$ urlsup --format json README.md
{"status": "success", "issues": []}
# Minimal output for scripts (no colors, emojis, or config info)
$ urlsup --format minimal README.md
404 https://example.com/broken
500 https://api.broken.com
# Or configure in .urlsup.toml
output_format = "json"
```
#### š JSON Output Examples
JSON format is perfect for automation, CI/CD integration, and programmatic processing:
**Successful validation:**
```bash
$ urlsup --format json README.md
{"status": "success", "issues": []}
```
**Failed validation with issues:**
```bash
$ urlsup --format json docs/
{"status": "failure", "issues": [
{"url": "https://example.com/404", "file": "docs/api.md", "line": 23, "status_code": 404, "description": ""},
{"url": "https://broken.link", "file": "docs/guide.md", "line": 45, "status_code": null, "description": "connection timeout"}
]}
```
**Processing JSON with `jq`:**
```bash
# Extract all broken URLs
# Count issues by status code
# Find all timeout errors
# Get files with broken links
# Export issues to CSV for reporting
#### Output Format Comparison
| `text` | ā
Yes | ā
Yes | ā
Yes | ā
Yes | ā
Yes | Interactive use |
| `json` | ā No | ā No | ā No | ā No | ā No | Automation/scripts |
| `minimal` | ā No | ā No | ā No | ā No | ā No | Simple scripts/CI |
### Verbose Logging
```bash
# Enable verbose output via CLI
$ urlsup --verbose README.md
# Or configure in .urlsup.toml
verbose = true
# Quiet mode for scripts (minimal output)
$ urlsup --quiet README.md
```
Verbose mode provides detailed information about:
- Files being processed
- URLs found and filtered
- Request progress and timing
- Configuration settings used
### š Performance Analysis
Get detailed insights into memory usage and performance characteristics:
```bash
# Enable performance monitoring with optimization suggestions
$ urlsup --show-performance README.md
# Example output:
# ā” Performance Analysis
# Total execution time: 2.34s
# Peak memory usage: 45.2 MB
# Average CPU usage: 23.4%
#
# š Operation Breakdown:
# ⢠File processing: 0.12s (156 files)
# ⢠URL discovery: 0.89s (1,247 URLs found)
# ⢠URL validation: 1.33s (987 unique URLs validated)
#
# š” Optimization Suggestions:
# ⢠Consider using --concurrency 8 for better performance
# ⢠Enable HEAD requests for faster validation (use_head_requests = true)
# ⢠Add .gitignore patterns to reduce file processing overhead
```
**Configuration:**
```toml
# In .urlsup.toml
show_performance = true # Always show performance analysis
```
**Use Cases:**
- **Performance tuning**: Identify bottlenecks in large documentation sets
- **CI/CD optimization**: Monitor resource usage in automated pipelines
- **Capacity planning**: Understand resource requirements for scaling
- **Troubleshooting**: Debug slow validation issues
### š HTML Dashboard
Generate comprehensive visual reports with charts and detailed analysis:
```bash
# Generate HTML dashboard with performance metrics
$ urlsup --html-dashboard report.html --show-performance docs/
# Dashboard includes:
# ⢠Interactive charts showing validation results
# ⢠Performance metrics and timing breakdowns
# ⢠Detailed issue listings with file locations
# ⢠Configuration summary and recommendations
# ⢠Responsive design for desktop and mobile viewing
```
**Features:**
- **š Interactive Charts**: Doughnut charts showing success/failure rates by category
- **š Performance Metrics**: Memory usage, CPU utilization, and timing analysis
- **š Detailed Issue Tracking**: Line-by-line breakdown of broken URLs
- **š” Smart Recommendations**: Optimization suggestions based on actual usage patterns
- **š± Responsive Design**: Works perfectly on desktop, tablet, and mobile devices
- **šØ Modern UI**: Clean, professional styling with dark/light theme support
**Dashboard Sections:**
1. **Executive Summary**: Key metrics and success rates
2. **Validation Results**: Interactive visualization of URL status distribution
3. **Performance Analysis**: Detailed timing and resource usage breakdown
4. **Issue Details**: Comprehensive list of broken URLs with context
5. **Optimization Recommendations**: Actionable suggestions for improvement
**Example Usage in CI/CD:**
```yaml
# .github/workflows/urls.yml
- name: Validate URLs and generate report
run: |
urlsup --html-dashboard validation-report.html --show-performance docs/
- name: Upload report artifact
uses: actions/upload-artifact@v3
with:
name: url-validation-report
path: validation-report.html
```
**Sample Dashboard Output:**
The HTML dashboard provides a complete overview of your URL validation results with:
- Visual charts showing the health of your URLs
- Performance insights to optimize future runs
- Detailed breakdown of any issues found
- Professional presentation suitable for stakeholder reporting
## š Security Features
### SSL Certificate Verification
```toml
# Skip SSL verification for internal/development URLs
skip_ssl_verification = true
```
**ā ļø Warning**: Only disable SSL verification for trusted internal environments.
### Proxy Support
```toml
# HTTP/HTTPS proxy configuration
proxy = "http://username:password@proxy.company.com:8080"
```
Supports both HTTP and HTTPS proxies with optional authentication.
## š Performance Features
### HEAD Request Optimization
For even faster URL validation, enable HEAD requests instead of GET requests:
```toml
# In .urlsup.toml
use_head_requests = true # Use HEAD instead of GET for faster validation
```
**Benefits:**
- **Faster validation**: HEAD requests only fetch headers, not full content
- **Reduced bandwidth**: Minimal data transfer for each URL check
- **Better for CI/CD**: Faster pipeline execution for large documentation sets
**When to use:**
- ā
Internal documentation validation
- ā
Known-good server environments
- ā
CI/CD pipelines with trusted URL sets
- ā
Large-scale validation where speed is critical
**When NOT to use:**
- ā Public URL validation (some servers reject HEAD requests)
- ā Mixed server environments with unknown HEAD support
- ā First-time validation of unknown URLs
**Example usage:**
```bash
# Enable HEAD requests for faster CI validation
$ urlsup --config .urlsup-fast.toml --recursive docs/
# Where .urlsup-fast.toml contains:
use_head_requests = true
timeout = 15
threads = 16
```
## ā” Browser-like HTTP Client
`urlsup` uses a simplified HTTP client designed for maximum compatibility:
### š Browser-Compatible Behavior
- **Automatic Protocol Negotiation**: Lets the client and server automatically negotiate HTTP/1.1 or HTTP/2
- **Default Connection Management**: Uses reqwest's browser-like connection handling for reliability
- **Automatic Compression**: Leverages gzip, brotli, and deflate for reduced bandwidth (like browsers)
- **Reliable Error Handling**: Avoids complex optimizations that can cause connection issues
### šÆ Memory & Algorithm Improvements
- **Ultra-Fast Hashing**: Uses `FxHashSet` for 15-20% faster URL deduplication
- **Smart Pre-allocation**: File-type-aware capacity estimation (Markdown 2x, HTML 3x multipliers)
- **Optimized Deduplication**: O(n) hash-based instead of O(n²) sorting-based
- **Memory-Efficient Streaming**: Handles large URL sets without memory bloat
- **Adaptive Sizing**: Dynamic memory allocation based on file types and URL patterns
- **SIMD-Optimized String Processing**: Uses `memchr` for vectorized URL pattern detection
- **Vectorized Line Processing**: Chunked processing with cache-friendly memory access patterns
### š Concurrent Processing
- **Dynamic Batch Sizing**: Batch sizes adapt to URL count and system resources (2-100 range)
- **Connection Pooling**: Optimized HTTP connection reuse with configurable pool limits
- **Token Bucket Rate Limiting**: Smooth request distribution vs simple delays
- **Batched Progress Updates**: Reduced atomic operations for better concurrent performance
- **Static Resource Reuse**: Eliminates repeated allocations for parsing components
### š Performance Gains
- **Small workloads (10-100 URLs)**: 25-35% faster validation with optimized batch sizing
- **Large workloads (1000+ URLs)**: 45-65% faster with 60-80% less memory usage
- **Memory efficiency**: File-type-aware allocation reduces memory waste by 30-50%
- **Network optimization**: Connection pooling and token bucket rate limiting improve throughput
- **CI/CD pipelines**: Dramatically reduced execution time for documentation validation
## šØ Error Handling
Comprehensive error handling with specific error types:
- **Configuration errors**: Invalid TOML, missing files
- **Network errors**: Timeouts, connection failures, DNS resolution
- **Path errors**: Invalid file paths, permission issues
- **Validation errors**: Malformed URLs, regex compilation failures
All errors include helpful context and suggestions for resolution.
## š GitHub Actions
See [`urlsup-action`](https://github.com/simeg/urlsup-action).
## š ļø Development
This repo uses a Makefile as an interface for common operations.
1) Do code changes
2) Run `make build link` to build the project and create a symlink from the built binary to the root
of the project
3) Run `./urlsup` to execute the binary with your changes
4) Profit :star:
[build_badge]: https://github.com/simeg/urlsup/workflows/CI/badge.svg
[coverage_badge]: https://codecov.io/gh/simeg/urlsup/branch/master/graph/badge.svg?token=2bsQKkD1zg
[coverage_report]: https://codecov.io/gh/simeg/urlsup/branch/master