Skip to main content

Crate codoseo_crawler

Crate codoseo_crawler 

Source
Expand description

The CodoSEO crawler. No database: the CLI, local MCP and the worker all use it directly.

Modules§

crawl
The crawl orchestrator: preflight, then a bounded, polite crawl one depth level at a time, ending in a CrawlOutput.
extract
Pulls SEO fields out of an HTML page as it streams in, without building a DOM.
fetch
HTTP fetching with hand-followed redirects, body caps and timeouts.
frontier
The URL queue: depth levels, de-duplication and a hard cap on the total.
guard
Keeps the cloud crawler away from private and internal addresses.
politeness
Request pacing for one site: a minimum gap between requests, a cap on parallel connections, and a back-off when the site answers 429 or 503.
preflight
The first step of a crawl: settle the real start address, read robots.txt and the sitemaps, and decide the stop reasons that don’t need a crawl.
robots
robots.txt rules for CodoSEObot, following Google’s handling of groups, status codes and size.
scope
The internal/external rule: same host and port as the start address.
sitemap
Sitemap parsing (streamed, gzip-aware) and discovery through sitemap indexes.

Structs§

CrawlOutput
Edge
One internal link between two crawled pages.
LinkGraph
Links between crawled pages, with anchor text interned so repeated navigation links don’t repeat their text.
Progress
A snapshot of a running crawl, for progress display.

Enums§

SiteFault
Whose fault a failed crawl was, read back from its stored failure_reason.
StopReason
Why a crawl ended.

Constants§

BLOCKED_REASON_PREFIX
Start of a crawl’s failure_reason when the site refused our crawler (StopReason::Blocked).
UNREACHABLE_REASON_PREFIX
Start of a crawl’s failure_reason when the site never answered (StopReason::Unreachable).