Skip to main content

Module crawl

Module crawl 

Source
Expand description

The crawl orchestrator: preflight, then a bounded, polite crawl one depth level at a time, ending in a CrawlOutput.

Every request goes through one Limiter. A level starts only when every fetch of the previous one has finished, so depth is exact even with parallel connections. Pages found only in sitemaps come after link exploration, with no depth.

Enums§

CrawlError
Why a crawl could not run at all. A crawl that ran and stopped early is a StopReason, not an error.

Functions§

crawl
Crawls cfg.start_url with its own in-flight cap of politeness.max_in_flight.
crawl_shared
Crawls cfg.start_url, taking a slot of global for every request, so several crawls can share one in-flight cap. on_progress is called after each recorded page.
inspect_page
Fetches cfg.start_url once, without reading robots.txt, and returns the record of the page it settles on, with every redirect hop in redirect_chain.