Skip to main content

Module preflight

Module preflight 

Source
Expand description

The first step of a crawl: settle the real start address, read robots.txt and the sitemaps, and decide the stop reasons that don’t need a crawl.

Order: robots.txt of the requested origin, then the start page (one retry after a 429 or 503), then robots.txt again when the start page moved to another origin. Every request goes through the limiter.

Constants§

BLOCKED_MSG
LOGIN_MSG