Skip to main content

Module extract

Module extract 

Source
Expand description

Pulls SEO fields out of an HTML page as it streams in, without building a DOM.

Encoding: a BOM wins, then the Content-Type charset, then <meta charset> (which lol_html switches to mid-stream), then UTF-8. Encodings that aren’t ASCII-compatible (UTF-16) are transcoded to UTF-8 before parsing.

Text comes from one document-wide handler and is routed by state that the element handlers keep (inside <title>, a heading, a link, a skipped element). Malformed input never panics: a parser error just ends extraction early.

Structs§

Extracted
Extractor
Link

Functions§

extract
Extracts everything from a complete body in one call.