Expand description
A high-performance HTML parser designed for document traversal.
Reader yields tokens in source order and selects HTML text states.
Tokenizer gives callers explicit control over those states. Document
retains lexical element scopes for repeated queries, with configurable Limits.
Input must already be decoded UTF-8. Unchanged strings borrow from that input. The parser does not implement browser tree construction.
Structs§
- Attribute
- An attribute. The first occurrence of each normalized name is retained.
- Doctype
- A document type declaration.
- Document
- An immutable source-order document with borrowed strings.
- Element
- A borrowed view of an element in a
Document. - Limits
- Resource limits for constructing a document.
- Reader
- A source-order HTML event reader that selects text modes from start tags.
- Tag
- A start or end tag.
- Tokenizer
- An HTML tokenizer over borrowed UTF-8 input.
Enums§
- Error
- A configured resource limit was exceeded.
- State
- A tokenizer state selected by the caller or a tree builder.
- Token
- An HTML token. Adjacent text tokens may be emitted separately.