unpdf
A high-performance Rust library for extracting content from PDF documents to structured Markdown, plain text, and JSON.
Features
- Comprehensive PDF support: PDF 1.0-2.0, including compressed object streams
- Encrypted PDF support: RC4 and AES-128 decryption (auto-tries empty password)
- Multiple output formats: Markdown, Plain Text, JSON (with full metadata)
- Structure preservation: Headings, paragraphs, lists, tables, inline formatting
- CJK text support: Smart spacing for Korean, Chinese, Japanese with Adobe CMap resources
- RTL text support: Arabic and Hebrew with Unicode BiDi reordering
- Form field extraction: AcroForm fields (text, checkbox, radio, dropdown) with values
- Multi-column layout: Recursive XY-Cut algorithm for N-column detection
- Asset extraction: Images, fonts, and embedded resources
- Extraction quality diagnostics: Automatic detection and reporting of extraction issues
- Text cleanup: Multiple presets for LLM training data preparation
- Self-update: Built-in update mechanism via GitHub releases
- WebAssembly / npm (0.7.0+):
@iyulab/unpdf— browser and Node.js via wasm-bindgen; live playground - C-ABI FFI: Native library for C#, Python, and other languages
- Parallel processing: Uses Rayon for multi-page documents
- Streaming pipeline (0.4.0+):
PdfParser::for_each_pageyields pages as they parse; peak memory bounded by window size regardless of document size - Deterministic page ordering: Parallel page parsing emits results in
page_numASC order via internal reorder buffer - Image deduplication (0.6.0+): Identical images across pages are written to disk only once; duplicate references are resolved to the canonical copy
Table of Contents
- Installation
- CLI Usage
- Rust Library Usage
- WebAssembly / JavaScript
- Python Integration
- C# / .NET Integration
- Output Formats
- Feature Flags
- License
Installation
Pre-built Binaries (Recommended)
Download the latest release from GitHub Releases.
Windows (x64)
# Download and extract (replace VERSION with the actual version, e.g. v0.7.0)
$VERSION = (Invoke-RestMethod "https://api.github.com/repos/iyulab/unpdf/releases/latest").tag_name
Invoke-WebRequest -Uri "https://github.com/iyulab/unpdf/releases/latest/download/unpdf-windows-x86_64-${VERSION}.zip" -OutFile "unpdf.zip"
Expand-Archive -Path "unpdf.zip" -DestinationPath "."
# Move to a directory in PATH (optional)
Move-Item -Path "unpdf.exe" -Destination "$env:LOCALAPPDATA\Microsoft\WindowsApps\"
# Verify installation
unpdf version
Linux (x64)
# Download and extract (replace VERSION with the actual version, e.g. v0.7.0)
VERSION=
# Install to /usr/local/bin (requires sudo)
# Or install to user directory
# Verify installation
macOS
# Intel Mac (replace VERSION with the actual version, e.g. v0.7.0)
VERSION=
# Apple Silicon (M1/M2/M3/M4)
# Install
# Verify
Available Binaries
| Platform | Architecture | File |
|---|---|---|
| Windows | x64 | unpdf-windows-x86_64-{version}.zip |
| Linux | x64 | unpdf-linux-x86_64-{version}.tar.gz |
| Linux | x64 (musl) | unpdf-linux-x86_64-musl-{version}.tar.gz |
| macOS | Intel | unpdf-macos-x86_64-{version}.tar.gz |
| macOS | Apple Silicon | unpdf-macos-aarch64-{version}.tar.gz |
Updating
unpdf includes a built-in self-update mechanism:
# Check for updates
# Update to latest version
# Force reinstall (even if on latest)
Install via Cargo
If you have Rust installed:
# Install CLI
# Add library to your project
CLI Usage
Quick Start
# Extract to Markdown + images (default)
# Specify output directory
# With text cleanup for LLM training
Output Structure
document_pdf_output/
├── extract.md # Markdown output with frontmatter
└── images/ # Extracted images (if any)
├── page1_img1.png
└── page2_img1.jpg
Use unpdf convert <file> --all to produce all three formats at once:
document_pdf_output/
├── extract.md # Markdown output with frontmatter
├── extract.txt # Plain text output
├── content.json # Full structured JSON
└── images/
Commands
Convert (multi-format streaming pipeline)
# Markdown only (default)
# All formats + images
# Specific formats
# Skip image extraction
# Custom image directory
# Drop images smaller than 128px (filter decorative icons)
# Tune streaming window size (pages in-flight, default: auto)
Convert Options
| Option | Description | Default |
|---|---|---|
-o, --output |
Output directory | <stem>_<ext>_output/ next to the input |
--formats |
Comma-separated formats: md,txt,json |
md |
--all |
Output all formats (MD + TXT + JSON) | false |
--no-images |
Skip image extraction | false |
--image-dir |
Custom image output directory | <out>/images |
--min-image-size |
Min pixel dimension; smaller images skipped | 64 |
--window |
Streaming window size (pages in-flight) | auto |
--keep-ocr-text |
Keep a scan's OCR text layer even when it recognised nothing readable | false |
--cleanup |
Text cleanup: minimal, standard, aggressive |
none |
--page-markers |
Insert <!-- page N --> markers |
false |
--ai-base-url, --ai-api-key, --ai-model, --ai-image-scope, --ai-refine |
See AI-assisted extraction. Configuring these switches convert to a buffered parse |
off |
-q, --quiet |
Suppress progress and warnings | false |
Convert to Markdown
# Basic conversion (output to stdout)
# Save to file
# With YAML frontmatter
# With text cleanup for LLM training
# Table rendering options
# Specify page range
# Insert page boundary markers for AI pipeline / RAG use
Markdown Options
| Option | Description | Default |
|---|---|---|
-o, --output |
Output file path | stdout |
-f, --frontmatter |
Include YAML frontmatter | false |
--table-mode |
Table rendering: markdown, html, ascii |
markdown |
--cleanup |
Text cleanup: minimal, standard, aggressive |
none |
--refine |
Apply the markdown shape-refinement pass | false |
--max-heading |
Maximum heading level (1-6) | 6 |
--pages |
Page range (e.g., 1-10, 1,3,5) |
all |
--page-markers |
Insert <!-- page N --> markers at page boundaries |
false |
--ai-base-url, --ai-api-key, --ai-model, --ai-image-scope, --ai-refine |
See AI-assisted extraction | off |
-q, --quiet |
Suppress quality warnings (root-level flag: unpdf --quiet markdown ...) |
false |
Convert to Plain Text
# Basic extraction
# With cleanup
# Specific pages
Convert to JSON
# Pretty-printed JSON
# Compact JSON
AI-assisted extraction
Optional, off by default, and entirely inert until an endpoint is configured — without these flags the output is byte-identical to a build that never had them.
Pointing unpdf at an OpenAI-compatible endpoint lets a vision model describe images
the text layer cannot represent: a full-page scan with no extractable text, a chart, a
photo. Tables recovered this way keep their merged-cell structure, which Markdown's own
table syntax cannot express.
# Describe every parsed image
# Only pages the low-confidence OCR gate flagged (cheaper, narrower)
# Additionally run the rendered markdown through an AI refine pass
| Option | Description | Default |
|---|---|---|
--ai-base-url |
Endpoint base URL, without a trailing /chat/completions |
— |
--ai-api-key |
Bearer token. Also read from UNPDF_AI_API_KEY |
— |
--ai-model |
Model name to request | — |
--ai-image-scope |
all, or low-confidence-only |
all |
--ai-refine |
Run an AI refine pass over the rendered markdown | false |
The first three go together: supplying only some of them is an error rather than a silent fallback, so a typo cannot look like a model that simply found nothing to say.
Availability. The passes run over a fully assembled document, so they are offered on
every command that produces one: convert, markdown, text and json. convert
normally streams pages to keep memory flat; configuring AI switches it to a buffered
parse, since a run making a vision-model call per page is not one whose bottleneck is
resident memory. Output is otherwise unchanged — same files, same bytes. --ai-refine
rewrites markdown, so it is offered only where markdown is rendered (convert and
markdown). When a call fails, extraction still succeeds with the original content and
the failure is counted in extraction_quality.ai_fallback_count.
Requires the ai cargo feature, which is on by default. Library consumers who want
neither the feature nor its HTTP dependency can opt out with default-features = false.
Show Document Information
Output:
Document Information
────────────────────────────────────────
File: document.pdf
Format: PDF 1.7
Pages: 42
Encrypted: No
Title: My Document
Author: John Doe
Creator: Microsoft Word
Producer: Adobe PDF Library
Created: 2025-01-15T10:30:00Z
Modified: 2025-01-20T14:45:00Z
Content Statistics
────────────────────────────────────────
Words: 12500
Characters: 75000
Images: 15
Extract Images
# Extract to current directory
# Extract to specific directory
# Extract specific pages
Self-Update
# Check for updates
# Update to latest version
# Force reinstall
Examples
# Convert PDF to Markdown with frontmatter
# Convert with aggressive cleanup for AI training
# Batch conversion (shell)
for; do ; done
# Batch conversion (PowerShell)
|
Rust Library Usage
Quick Start
use ;
Builder API
Unpdf provides a fluent builder for the common parse-then-render workflow:
use ;
let markdown = new
.lenient
.with_frontmatter
.with_cleanup
.with_table_fallback
.with_pages
.parse?
.to_markdown?;
Convenience Functions
// Parse from bytes or a reader
let doc = parse_bytes?;
let doc = parse_reader?;
// Password-protected PDF
let doc = parse_file_with_password?;
// One-shot conversions without building a Document first
let text = extract_text?;
let markdown = to_markdown?;
let json = to_json?;
Render Options
use ;
use PageMarkerStyle;
let options = new
.with_frontmatter
.with_table_fallback
.with_cleanup_preset
.with_max_heading
.with_page_range
.with_page_markers; // <!-- page N --> at each page boundary
let markdown = to_markdown?;
Working with Document Structure
use parse_file;
let doc = parse_file?;
// Access metadata
println!;
println!;
println!;
println!;
// Iterate pages
for in doc.pages.iter.enumerate
// Extract images
for in &doc.resources
Page Range Selection
use ;
// Parse only specific pages
let doc = parse_file_with_options?;
// Or parse all and render specific pages
let doc = parse_file?;
let options = new
.with_pages;
let markdown = to_markdown?;
Handling Encrypted PDFs
unpdf tries the empty user password first, which opens owner-password-only documents. When that
does not authenticate, the password you supplied is tried; if it does not work either, the error
is ErrorKind::InvalidPassword rather than ErrorKind::Encrypted, so "wrong password" and "no
password given" stay distinguishable.
use ;
// Auto-decrypts owner-password-only PDFs
let doc = parse_file?;
// Provide a password for user-password-protected PDFs
let options = new.with_password;
let doc = parse_file_with_options?;
// Check extraction quality
if let Some = doc.extraction_quality.warning_message
Working with Form Fields
use parse_file;
let doc = parse_file?;
for field in &doc.form_fields
Error Handling Defaults
Parsing is lenient by default. A PDF that is damaged in one place still yields the pages and text that could be read, rather than failing whole. That is the useful default for a format where partial damage is common and a caller usually wants whatever survived.
The cost of a lenient default is that a successful call can return less than the document
held, so every kind of loss is counted rather than swallowed — see
Detecting Incomplete Extraction below, and check
extraction_quality on any result you intend to index or archive.
Opt into failing instead:
use ;
let options = new.with_error_mode;
let doc = parse_file_with_options?;
Under Strict the same damage is an error: a content stream that cannot be decoded fails the
document instead of leaving a shorter page. A page that legitimately has no content at all is
not damage and stays a success in both modes.
Its sibling parsers answer this differently, because the formats do: unhwp defaults to
strict, and undoc has no error mode at all. Code that drives all three should not assume a
shared default.
Classifying Failures
Error::kind() returns a stable ErrorKind so you can branch on why a call failed
without matching on the message — useful when several failures reach the user as the
same "extraction failed":
use ;
match parse_file
There is one ErrorKind variant per Error variant. The discriminants are explicit
and part of the public contract — they cross the C ABI as unpdf_last_error_kind
return values, so existing values are never renumbered.
Rendering a Page
With the raster feature (included in the C ABI, .NET and Python packages), a page renders
to an image — painted from the same content interpretation text extraction reads, so it is
the same page, box and rotation:
use RasterOptions;
use PdfParser;
let parser = open?;
let page = parser.render_page?;
write?;
if !page.gaps.is_empty
Painted: text in embedded TrueType, OpenType and CFF fonts, paths, clipping, colors, and
images (Flate-family and JPEG, with soft and stencil masks). Not yet: text in fonts that are
not embedded or are Type 1/Type 3, JPEG 2000/JBIG2/CCITT images, inline images and shadings —
each is counted in gaps, and the rest of the page is still painted.
Detecting Incomplete Extraction
A damaged PDF does not always fail. When the cross-reference table survives but the
objects it points at do not, the parser recovers the pages it can read and returns
them — a success over an incomplete page set. extraction_quality reports that:
let doc = parse_file?;
let q = &doc.extraction_quality;
if q.pages_incomplete
| Field | Meaning |
|---|---|
pages_incomplete |
Pages are known to be missing. The one bit to branch on. |
declared_page_count |
Page count the document declares (/Count), or None if that was unreadable too. |
unresolved_page_nodes |
Unreadable page-tree nodes. Non-zero means incomplete — not a count of lost pages. |
skipped_object_count |
Objects that could not be loaded. Most cost no page (fonts, annotations), so this alone does not imply missing text. |
suppressed_text_runs |
Text runs the font decoder could not read and discarded. Non-zero means text is missing from an otherwise successful extraction. |
undecodable_content_streams |
Page content streams that could not be decoded. Lenient parsing leaves them out and keeps the rest of the page — an empty page when it was the page's only stream — so non-zero means content is missing. Strict parsing fails instead. |
unsupported_image_count |
Embedded images recognized as image XObjects but not extractable (unsupported color space/bit depth), so dropped. Distinguishes "no image" from "image present but couldn't be extracted". |
ai_fallback_count |
Times VLM image understanding fell back to the non-AI result (transport failure, non-success status, malformed response, or retries exhausted). Only ever non-zero when AI parse options are configured; extraction still succeeds, so treat it as a quality signal, not a failure. The render-side refine pass is not counted here. |
suppressed_text_runs covers a second way content goes missing without an error. When a
font's character codes cannot be resolved — a composite (Type0/CID) font with no usable
ToUnicode map, most often — the decoder discards the run rather than emit the raw bytes
as mojibake. That is the right call, but the discarded text is content the document had
and the output does not. The count is in runs (a Tj operand, or one element of a TJ
array), not characters: text that was never decoded has no knowable length. Treat any
non-zero value as "incomplete"; compare magnitudes between documents, not against a total.
Worth surfacing wherever extraction feeds an index or archive: a page that silently
never arrived is indistinguishable from a page that never existed, so the omission
shows up later as a search result that isn't there rather than as an error.
warning_message() already includes this case, ahead of the other warnings.
WebAssembly / JavaScript
Live playground → Drag and drop a PDF in the browser; no install required.
Browser / Bundler (webpack, vite)
import from '@iyulab/unpdf';
const response = await ;
const bytes = ;
const doc = ;
console.log;
console.log;
console.log;
Node.js
const = require;
const fs = require;
const bytes = ;
const doc = ;
console.log;
With Options
import from '@iyulab/unpdf';
const opts =
.
.
.;
const doc = ;
console.log;
API Reference
Functions
| Function | Signature | Description |
|---|---|---|
parse |
(data: Uint8Array) => PdfDocument |
Parse PDF bytes |
parseWithOptions |
(data: Uint8Array, opts: ParseOptions) => PdfDocument |
Parse with options |
PdfDocument
| Method | Returns | Description |
|---|---|---|
PdfDocument.fromBytes(data) |
PdfDocument |
Parse PDF bytes (static; same as parse) |
toMarkdown() |
string |
Convert to Markdown |
toText() |
string |
Convert to plain text |
toJson() |
string |
Convert to JSON |
pageCount() |
number |
Total page count |
metadata() |
string |
Metadata as JSON string |
extractionQuality() |
string |
Extraction diagnostics as JSON string — check pages_incomplete before indexing the result |
ParseOptions
| Method | Description |
|---|---|
new() |
Default options |
lenient() |
Ignore recoverable errors |
textOnly() |
Skip image extraction |
withPassword(pw: string) |
Set decryption password |
withPages(from: number, to: number) |
Page range (1-indexed) |
Python Integration
Install the Python package:
Basic Usage
# Convert PDF to Markdown
=
# Convert to plain text
=
# Convert to JSON
=
# Get document information
=
get_info returns only the keys it found: title and author are absent when the
document does not set them, and the page count is section_count.
Working with Paths and Bytes
Every function accepts a path (str or any os.PathLike, so pathlib.Path works) or
the PDF's own bytes. The two are told apart by type, so there is no ambiguity:
# str path
# pathlib.Path
# bytes — parsed in memory, no temp file
Bytes input goes through the native in-memory parser, so uploads and blobs need no detour through the filesystem.
Check PDF Validity
# Check if file is a valid PDF
=
Detecting Scanned (Image-only) PDFs
Empty extraction output can mean a scanned document (no text layer), a genuinely blank page, or a parse failure. The introspection surface tells them apart:
# Document level
=
# Page level (works for mixed documents too)
=
Note: a searchable scan (page image plus an invisible OCR text layer) reports
text_op_count > 0 — combine the check with ocr_text_suppressed, which flags
pages whose unreadable OCR layer was dropped.
The same call reports whether the document was damaged badly enough to lose pages — extraction can succeed over an incomplete page set:
=
unresolved_page_nodes counts unreadable page-tree nodes, not lost pages — one
unreadable node can cost a whole subtree, so treat any non-zero value as "incomplete"
and nothing more.
Handling Failures
The checks above need a parsed document. When parsing itself fails, UnpdfError
carries a kind so you can branch on the reason instead of matching on message text:
=
UnpdfError subclasses RuntimeError, so existing except RuntimeError handlers
keep working. ErrorKind values are part of the native ABI: new reasons take new
numbers and existing ones are never renumbered, so treat an unrecognised value as a
generic failure.
C# / .NET Integration
unpdf provides C-ABI compatible bindings for integration with C# and .NET applications.
Installation via NuGet
Or via Package Manager Console:
Install-Package Unpdf
Getting the Native Library (Manual)
Alternatively, download from GitHub Releases:
| Platform | Library File |
|---|---|
| Windows x64 | unpdf.dll |
| Linux x64 | libunpdf.so |
| macOS | libunpdf.dylib |
Or build from source:
C# Wrapper Usage
The API is handle-based: parse once into an UnpdfDocument, then read from it. The
handle owns native memory, so dispose it (using).
using Unpdf;
using var doc = UnpdfDocument.ParseFile("document.pdf");
string markdown = doc.ToMarkdown();
string text = doc.ToText();
string json = doc.ToJson(compact: false);
string plain = doc.PlainText();
// Document facts
Console.WriteLine($"Title: {doc.Title}, Pages: {doc.SectionCount}");
// Markdown options
string withFrontmatter = doc.ToMarkdown(new MarkdownOptions
{
IncludeFrontmatter = true,
EscapeSpecialChars = true,
PageMarkers = true,
});
// A single page (1-indexed)
string page1 = doc.PageToMarkdown(1);
// Embedded resources (images and the like), by id
foreach (var id in doc.GetResourceIds())
{
byte[]? bytes = doc.GetResourceData(id);
if (bytes is not null)
File.WriteAllBytes(Path.Combine("./images", id), bytes);
}
Parsing from memory works the same way: UnpdfDocument.ParseBytes(byte[]).
Note: SectionCount is the page count. ResourceCount counts embedded resources,
which is not the same as the number of images — filter with GetResourceInfo(id) if
you need images specifically. Its page field is the page a resource was collected
from (1-based, the numbering of the page markers).
Detecting Scanned (Image-only) PDFs
Empty extraction output can mean a scanned document (no text layer), a genuinely blank page, or a parse failure. The introspection surface tells them apart:
using Unpdf;
using var doc = UnpdfDocument.ParseFile("scan.pdf");
// Document level
var quality = doc.GetExtractionQuality();
if (quality.IsScanPdf)
Console.WriteLine("Scanned document - OCR required");
// Page level (works for mixed documents too)
var stats = doc.GetPageStats(1);
if (stats.TextOpCount == 0 && stats.ImageOpCount > 0)
Console.WriteLine("Page 1 is image-only (scanned)");
else if (stats.TextOpCount == 0)
Console.WriteLine("Page 1 is genuinely blank");
Note: a searchable scan (page image plus an invisible OCR text layer) reports
TextOpCount > 0 — combine the check with OcrTextSuppressed, which flags
pages whose unreadable OCR layer was dropped.
The same object reports whether the document was damaged badly enough to lose pages. Extraction can succeed over an incomplete page set, and a page that silently never arrived looks exactly like a page that never existed:
var quality = doc.GetExtractionQuality();
if (quality.PagesIncomplete)
Console.WriteLine(
$"incomplete: got {doc.SectionCount} page(s), " +
$"document declares {quality.DeclaredPageCount}");
UnresolvedPageNodes counts unreadable page-tree nodes, not lost pages — one
unreadable node can cost a whole subtree. Treat any non-zero value as "incomplete".
Handling Failures
The checks above need a parsed document. When parsing itself fails,
UnpdfException.Kind says why, so you can branch on the reason instead of matching
on Message:
try
{
using var doc = UnpdfDocument.ParseFile("document.pdf");
Console.WriteLine(doc.ToMarkdown());
}
catch (UnpdfException e)
{
switch (e.Kind)
{
case UnpdfErrorKind.Encrypted:
Console.WriteLine("Password required");
break;
case UnpdfErrorKind.Corrupted:
case UnpdfErrorKind.PdfParse:
Console.WriteLine("The file is damaged");
break;
case UnpdfErrorKind.UnknownFormat:
Console.WriteLine("Not a PDF");
break;
default:
Console.WriteLine($"Extraction failed ({e.Kind}): {e.Message}");
break;
}
}
UnpdfErrorKind values are part of the native ABI: new reasons take new numbers and
existing ones are never renumbered, so treat an unrecognised value as a generic
failure. A failure raised by the managed wrapper rather than the native library
reports UnpdfErrorKind.Other.
ASP.NET Core Example
[ApiController]
[Route("api/[controller]")]
public class PdfController : ControllerBase
{
[HttpPost("convert")]
public async Task<IActionResult> ConvertPdf(IFormFile file)
{
if (file == null) return BadRequest("No file");
// Parse straight from the uploaded bytes — no temp file needed.
using var buffer = new MemoryStream();
await file.CopyToAsync(buffer);
try
{
using var doc = UnpdfDocument.ParseBytes(buffer.ToArray());
var markdown = doc.ToMarkdown(new MarkdownOptions { IncludeFrontmatter = true });
// Success is not the same as complete: report a damaged page set instead of
// returning a short document as if it were whole.
var quality = doc.GetExtractionQuality();
return Ok(new
{
markdown,
pages = doc.SectionCount,
incomplete = quality.PagesIncomplete,
declaredPages = quality.DeclaredPageCount,
});
}
catch (UnpdfException ex)
{
return BadRequest(new { error = ex.Message, kind = ex.Kind.ToString() });
}
}
[HttpPost("extract-images")]
public async Task<IActionResult> ExtractImages(IFormFile file)
{
if (file == null) return BadRequest("No file");
using var buffer = new MemoryStream();
await file.CopyToAsync(buffer);
try
{
using var doc = UnpdfDocument.ParseBytes(buffer.ToArray());
var ids = doc.GetResourceIds();
return Ok(new { count = ids.Length, ids });
}
catch (UnpdfException ex)
{
return BadRequest(new { error = ex.Message, kind = ex.Kind.ToString() });
}
}
}
Output Formats
Markdown
Structured Markdown with preserved formatting:
- Headings: Document headings (detected from font size/style) ->
#,##,### - Paragraphs: Text blocks with proper spacing
- Lists: Detected ordered and unordered lists
- Tables: Markdown tables (with HTML/ASCII fallback for complex layouts)
- Inline styles: Bold (
**), italic (*) - Hyperlinks: Preserved as Markdown links
- Images: Reference-style image links
Plain Text
Pure text content without formatting markers.
JSON
Complete document structure with metadata:
Supported PDF Features
| Feature | Status |
|---|---|
| PDF 1.0 - 2.0 | Supported |
| Compressed object streams (ObjStm) | Supported |
| Cross-reference streams (XRef streams) | Supported |
| Linearized PDFs | Supported |
| Encrypted PDFs (RC4, AES-128 — revisions 2-4) | Supported |
| Encrypted PDFs (AES-256 — AESV3, revisions 5-6) | Not supported — refused as ErrorKind::UnsupportedVersion |
| Text extraction | Supported |
| CJK text (Korean, Chinese, Japanese) | Supported (Adobe CMap) |
| RTL text (Arabic, Hebrew) | Supported (BiDi) |
| CIDFont / ToUnicode CMap decoding | Supported |
| Embedded TrueType font decoding | Supported |
| Multi-column layout detection | Supported (XY-Cut) |
| Table detection | Supported |
| Form fields (AcroForms) | Supported |
| Image extraction (JPEG, JP2) | Supported |
| Bookmarks/Outlines | Supported |
| Extraction quality diagnostics | Supported |
| AES-256 encryption (R5-R6) | Not yet supported |
| Digital signatures | Metadata only |
| OCR (image-based PDFs) | Planned |
Feature Flags
| Feature | Description | Default |
|---|---|---|
fast-parse |
Enable optimised nom-based PDF tokeniser | Yes |
refine |
Markdown shape-refinement pass (RenderOptions::refine) |
Yes |
ai |
VLM image understanding and AI refine — see AI-assisted extraction. Inert without an endpoint; adds an HTTP dependency | Yes |
ffi |
C-ABI foreign function interface | No |
async |
Async I/O with Tokio | No |
Performance
- Custom zero-dependency PDF parser (no external C libraries)
- Parallel page processing with Rayon
- Memory-efficient handling of large documents
- Streaming support for very large files
License
MIT License - see LICENSE for details.
Contributing
Contributions are welcome! Please feel free to submit a Pull Request.