Skip to main content

Crate docling

Crate docling 

Source
Expand description

docling.rs: a Rust port of docling.

The public surface mirrors the Python SDK, kept deliberately small:

use docling::{DocumentConverter, SourceDocument};

let converter = DocumentConverter::new();
let result = converter
    .convert(SourceDocument::from_file("input.md").unwrap())
    .unwrap();
println!("{}", result.document.export_to_markdown());

For the PDF/image ML pipeline (pdfium + layout/TableFormer/OCR ONNX), reuse a Pipeline across documents to amortize model loading, instead of the per-call DocumentConverter. Deploying as a service: examples/Dockerfile is a 3-stage build that bakes the binary, native libs, and exported models (including the KV-cached TableFormer decoder) into a slim, Python-free runtime image — see the “Deploy in a container” section of the README.

See docs/MIGRATION.md for the architecture, format-by-format parity status, and how conformance against Python docling is measured.

Re-exports§

pub use archive::ArchiveLimits;
pub use archive::ArchiveOutcome;
pub use email_attachments::EmailAttachmentInfo;
pub use email_attachments::EmailAttachments;
pub use options::cli_flag;
pub use options::merge_options;
pub use options::ConvertOptions;
pub use options::OptionInfo;
pub use options::OptionsError;
pub use options::PipelineKind;
pub use options::OPTIONS;

Modules§

archive
ZIP archives as input (#557): every document inside converts on its own.
backend
Format backends.
base64
Minimal standard-alphabet Base64 codec (RFC 4648): encode for embedding image bytes as data: URIs, decode for reading them back out — avoids a dependency for the two things we need.
chunker
Document chunking for RAG pipelines — the Rust port of docling-core’s docling_core.transforms.chunker.
chunks
Chunk-record JSON export shared by the CLI (--to chunks) and the HTTP server (to=chunks): the hierarchical chunker’s records always, plus the hybrid chunker’s when a tokenizer is available — DOCLING_CHUNK_TOKENIZER, or .models/chunk/tokenizer.json as populated by scripts/install/download_dependencies.sh (requires the chunking build feature; DOCLING_CHUNK_MAX_TOKENS overrides the default budget of 256).
dclx
.dclx packaging: the DocLang OPC archive (doclang.pack counterpart).
email_attachments
Email attachments as input (#561): the payloads of an .eml / .msg.
options
The one set of conversion options every surface speaks (#577).
pandoc
Pandoc AST output (--to pandoc, #515): the document as the JSON serialization of Pandoc’s Pandoc type (pandoc -f json), which hands docling.rs’s parsing to every Pandoc writer — DOCX, ODT, EPUB, RST, Org, Typst, AsciiDoc, … — through docling-rs in.pdf --to pandoc | pandoc -f json -t docx -o out.docx.
video
Video frame sampling for InputFormat::Video (#138 Phase 2).
vlm
VLM pipeline (issue #77) — remote OpenAI-compatible vision endpoint.

Structs§

ConfidenceReport
The document-level report (docling’s ConfidenceReport): the four scores aggregated across pages, plus the per-page breakdown. Page keys are the real 1-based page numbers — the same numbering as the JSON export’s pages map (#171), --pages windows included. (docling keys by its 0-based internal page index; ours is the more useful spelling and the difference is documented in docs/MIGRATION.md.)
ContentLayers
A set of content layers, body included — docling-core’s set[ContentLayer] (HTMLParams.layers, export_to_html(included_content_layers=…)), where body is a member like any other. Default is docling’s DEFAULT_CONTENT_LAYERS: body only, so an export built with it is unchanged from before layers could be chosen (#499).
ConversionResult
The result of converting one crate::SourceDocument.
DoclingDocument
The unified, format-agnostic document produced by every backend.
DocumentConverter
Routes a SourceDocument to the backend for its format and returns a ConversionResult.
EnrichmentOptions
The opt-in enrichment passes, mirroring docling’s PdfPipelineOptions flags (do_picture_classification, do_code_enrichment, do_formula_enrichment). All off by default.
ErrorItem
One recorded problem of a conversion that still produced a document — docling’s ErrorItem (component_type, module_name, error_message). A result with any of these is a ConversionStatus::PartialSuccess.
HeadingHierarchyOptions
Options for the heading-hierarchy stage (docling’s HeadingHierarchyOptions, defaults included).
HtmlExportOptions
Options of the HTML export (DoclingDocument::export_to_html_with): docling-core’s HTMLParams subset the port honours. Default is upstream’s default export — placeholder images, artifacts as the referenced-image directory, the body layer only.
ImageOutput
Image outputs of the PDF/image pipeline (#519/#520): the scale picture crops are delivered at and whether each page’s render is kept as the document’s page image. See Pipeline::images_scale / Pipeline::generate_page_images.
MarkdownExportOptions
Options of the Markdown export (DoclingDocument::export_to_markdown_with_options, #599): the docling-core MarkdownParams the port honours beyond the image mode. Default is upstream’s default export — placeholder images, artifacts as the referenced-image directory, the body layer only, no picture traversal, HTML and underscore escaping on, <!-- image --> — so a document exported with it is byte-identical to DoclingDocument::export_to_markdown.
MarkdownStream
An iterator over a document’s Markdown, yielded in document order as conversion progresses. Each item is a chunk to write as-is; concatenating every Ok chunk reproduces the buffered Markdown byte-for-byte.
MarkdownStreamer
Incremental Markdown serializer: feed finalized, in-document-order batches of Nodes and receive Markdown chunks whose concatenation is byte-identical to [to_markdown_images] over the same nodes. This is the streaming counterpart of the buffered serializer — used to emit a document’s Markdown in chunks (e.g. page by page, as the parallel PDF pipeline finishes pages) instead of building the whole string up front.
ModelEntry
One resolved runtime asset — which file a stage would load right now, given the CWD, the env overrides and the int8/fp32 preference.
PictureImage
An extracted picture’s raw encoded bytes plus its mimetype and pixel size — the docling.rs analogue of docling-core’s ImageRef.
Pipeline
A reusable PDF pipeline. The primary worker runs its models on every core, so a single-page / small / image / METS input is converted at full intra-op speed with no pool to load. A document with enough pages instead fans out across a pool of narrower workers processed concurrently. Both load lazily and are cached for reuse, so a one-shot conversion only pays for what it uses.
RenderedPage
One rasterized page from render_pages (#243): the absolute 1-based page number in the source document, the pixel dimensions, and the PNG bytes.
SourceDocument
A loaded input document: its name, detected format, and raw bytes.
Table
A simple row-major table. By default rows[0] is the header row; a TableStructure overlay overrides that and adds column spans.

Enums§

ContentLayer
A DocLang content layer other than the default body (see Node::Furniture).
ConversionError
Anything that can go wrong while loading or converting a source document.
ConversionStatus
Outcome status of a conversion, mirroring docling.datamodel.base_models.ConversionStatus.
DocItemLabel
Semantic role of a document item, mirroring docling-core’s DocItemLabel.
ImageMode
How pictures are rendered (mirrors docling-core’s ImageRefMode).
InputFormat
A document format supported by docling.rs backends.
Node
A single piece of document content.
OcrEngine
Which OCR engine recognizes text (#460): the built-in PP-OCRv3 recognizer (+ the RapidOCR text detector) — the default and the engine every conformance baseline is pinned against — or the system tesseract binary (see [crate::tesseract]), docling’s TesseractCliOcrOptions counterpart. Both consume the same layout-region crops and produce the same cells; everything downstream is engine-agnostic.
OcrLang
OCR recognition language: which PP-OCRv3 model + dictionary pair runs when the PP-OCRv6 recognizer is not installed.
OcrMode
Which document regions feed the OCR — docling 2.116’s OcrMode (#254, upstream docling#3710). Upstream restructured its pipeline so OCR runs after layout, on layout regions filtered by the PDF text layer — the architecture this port has always had — and named the strategies:
QualityGrade
docling’s QualityGrade: a score bucketed for human consumption.

Constants§

DEFAULT_VIDEO_FRAMES
Default cap on sampled frames per video. Scene changes rarely exceed this in short clips, and uniform fallback at 8 keeps JSON/DCLX output (which embeds the PNGs) within sane bounds.
OUTPUT_FORMATS
The output formats a conversion can be written as — the --to values of the CLI and the to option of docling-serve, in the order the help text lists them (markdown is accepted as an alias of md on both). One list, so the validation, the error messages and --list-output-formats (#603) cannot drift apart.
PDF_ML_COMPILED
Which PDF conversion this build compiled in: the full ML pipeline (pdf feature), the pure-Rust text-layer path (pdf-text, the wasm32 build), or neither. Compile-time facts, exported so downstream crates (whose own features can’t see this crate’s) can branch — e.g. docling-wasm’s host tests, where workspace feature unification may pull pdf in.
PDF_TEXT_COMPILED
True when the pdf-text text-layer-only PDF path is compiled in.

Functions§

model_inventory
Resolve the whole runtime model set without loading anything — the exact selection each stage performs at load time (layout honors the int8/fp32 preference, TableFormer its decoder ranking, OCR the language pair), plus the pdfium library. docling-serve exposes this at /v1/config and logs it at startup, so “the server picked up different models” is one curl away instead of a mystery of dissolved tables. Resolution is CWD-relative with an exe-dir fallback, so the answer can legitimately differ between two working directories.
parse_page_range
Parse a user-facing page-range string (issue #80’s --pages): "A-B" for an inclusive 1-based window, or a single "N" for one page. Whitespace around the numbers is tolerated. Validation against the actual page count happens at convert time; this only checks the spelling (first >= 1, first <= last).
pdf_page_count
Number of pages in a PDF, without converting anything — what the CLI batch mode prints in its per-document start line.
pdf_text_layer_pages
convert_text_layer restricted to a 1-based inclusive page window (issue #80’s --pages); None converts everything. The window is validated the same way as Pipeline::pages: first <= last, 1-based, and it must select at least one existing page.
render_pdf_pages
Rasterize a PDF’s pages to PNG (#243) — the lean path behind serve’s to=images: pdfium render only, no text extraction, no models, and only one page bitmap resident at a time (each is PNG-encoded and dropped before the next renders). scale is pixels per PDF point — 2.0 matches the pipeline’s RENDER_SCALE (144 dpi). Unlike the pipeline’s render there is no 1.5× supersample + downsample pass: that dance exists only because TableFormer is pixel-pinned to docling’s bitmaps, and nothing downstream of this output is — a single render is nearly twice as fast.
tesseract_lang_arg
Tesseract’s -l argument for an ocr_lang value under this engine.