Expand description
§Legible
Legible extracts the main article from an HTML document. It removes navigation, advertisements, sidebars, and other unrelated content. Legible is a Rust port of Mozilla’s Readability.js.
§Extract an article
Use parse for most applications:
use legible::parse;
let html = r#"
<html>
<head><title>My Article</title></head>
<body>
<nav>Navigation</nav>
<article>
<h1>Article Title</h1>
<p>This is the main content of the article.</p>
<p>This second paragraph contains more article text.</p>
</article>
<footer>Footer</footer>
</body>
</html>
"#;
match parse(html, Some("https://example.com/articles/1"), None) {
Ok(article) => {
println!("Title: {}", article.title);
println!("HTML: {}", article.content);
println!("Markdown: {}", article.markdown_content);
println!("Text: {}", article.text_content);
}
Err(error) => eprintln!("Error: {error}"),
}The optional URL must be absolute. Legible uses it as the base URL for relative
links and media URLs. Relative URLs stay relative if you pass None.
Article provides HTML, CommonMark, normalized plain text, and article metadata.
§Check a document before extraction
is_probably_readerable performs a quick content check. This check is a
heuristic. A true result does not guarantee successful extraction. A false
result does not prove that the document has no article.
Use Document if you want to run the check and then extract the article.
Document prevents a second HTML parse.
use legible::Document;
let text = "Article text. ".repeat(30);
let html = format!("<article><p>{text}</p></article>");
let document = Document::new(&html);
if document.is_probably_readerable(None) {
let result = document.parse(Some("https://example.com/articles/1"), None);
// Use the extraction result.
}The check borrows the document. Extraction consumes it because extraction changes the internal document tree.
§Configure extraction
Use Options to configure extraction. Use ReaderableOptions to configure the
quick content check.
use legible::{Options, parse};
let options = Options::new()
.char_threshold(250)
.keep_classes(true)
.disable_json_ld(true);
let result = parse(
"<html><body><article><p>Article text</p></article></body></html>",
Some("https://example.com/articles/1"),
Some(options),
);§Security
Do not render Article::content without sanitizing it.
Legible cleans article content, but it is not an HTML security sanitizer. The HTML can contain unsafe attributes, URLs, or other source markup. Apply a sanitizer that matches your security policy before you render the HTML.
Article::markdown_content does not contain raw HTML. It removes links and images
that have unsupported URI schemes. If you convert the Markdown to HTML, sanitize
that HTML according to your application’s security policy.
Structs§
- Article
- Extracted article content and metadata.
- Document
- A parsed HTML document.
- Options
- Options for
parse()andDocument::parse. - Readerable
Options - Options for
is_probably_readerableandDocument::is_probably_readerable.
Enums§
- Error
- Errors from article extraction.
Functions§
- is_
probably_ readerable - Checks if an HTML document probably contains readable article content.
- parse
- Extract article content and metadata from an HTML document.
Type Aliases§
- Result
- A result from article extraction.