Expand description
The CodoSEO crawler. No database: the CLI, local MCP and the worker all use it directly.
Modules§
- crawl
- The crawl orchestrator: preflight, then a bounded, polite crawl one depth level at a
time, ending in a
CrawlOutput. - extract
- Pulls SEO fields out of an HTML page as it streams in, without building a DOM.
- fetch
- HTTP fetching with hand-followed redirects, body caps and timeouts.
- frontier
- The URL queue: depth levels, de-duplication and a hard cap on the total.
- guard
- Keeps the cloud crawler away from private and internal addresses.
- politeness
- Request pacing for one site: a minimum gap between requests, a cap on parallel connections, and a back-off when the site answers 429 or 503.
- preflight
- The first step of a crawl: settle the real start address, read robots.txt and the sitemaps, and decide the stop reasons that don’t need a crawl.
- robots
- robots.txt rules for CodoSEObot, following Google’s handling of groups, status codes and size.
- scope
- The internal/external rule: same host and port as the start address.
- sitemap
- Sitemap parsing (streamed, gzip-aware) and discovery through sitemap indexes.
Structs§
- Crawl
Output - Edge
- One internal link between two crawled pages.
- Link
Graph - Links between crawled pages, with anchor text interned so repeated navigation links don’t repeat their text.
- Progress
- A snapshot of a running crawl, for progress display.
Enums§
- Site
Fault - Whose fault a failed crawl was, read back from its stored
failure_reason. - Stop
Reason - Why a crawl ended.
Constants§
- BLOCKED_
REASON_ PREFIX - Start of a crawl’s
failure_reasonwhen the site refused our crawler (StopReason::Blocked). - UNREACHABLE_
REASON_ PREFIX - Start of a crawl’s
failure_reasonwhen the site never answered (StopReason::Unreachable).