ferromark
Markdown to HTML with a secure default and every GFM extension included. The reproducible benchmark protocol and current CommonMark conformance result are documented below.
Quick start
let html = to_html;
One function call, no setup. Using Node.js instead? The same engine ships as a
native npm package: npm install ferromark — see the
Node.js package README.
When allocation pressure matters:
let mut buffer = Vecnew;
to_html_into;
// buffer survives across calls — zero repeated allocation
Benchmarks
Numbers, not adjectives. Apple Silicon (M-series), July 2026. All parsers run with GFM tables, strikethrough, and task lists enabled; ferromark's non-GFM extras (heading IDs, callouts) are disabled. Output buffers are reused where APIs allow and binaries are non-PGO. ferromark also keeps its secure default rendering in this published product lane, so it performs URL and raw-HTML safety work that pulldown-cmark does not.
CommonMark 5 KB (wiki-style, mixed content with tables)
| Parser | Throughput | vs ferromark |
|---|---|---|
| ferromark | 259.6 MiB/s | baseline |
| pulldown-cmark | 254.5 MiB/s | 0.98x |
| md4c (C) | 243.1 MiB/s | 0.94x |
| comrak | 67.9 MiB/s | 0.26x |
CommonMark 50 KB (same style, scaled)
| Parser | Throughput | vs ferromark |
|---|---|---|
| ferromark | 280.5 MiB/s | baseline |
| pulldown-cmark | 275.2 MiB/s | 0.98x |
| md4c (C) | 253.3 MiB/s | 0.90x |
| comrak | 71.8 MiB/s | 0.26x |
2% faster than pulldown-cmark. 11% faster than md4c. 4x faster than comrak. Competitor versions: pulldown-cmark 0.13.4, comrak 0.53, md4c @ 65c6c9d.
The fixtures are synthetic wiki-style documents with paragraphs, lists, code blocks, and tables. Nothing cherry-picked. The cross-parser harness is isolated from the library build and pins md4c at 65c6c9d for the published numbers:
MD4C_DIR=/path/to/md4c
MD4C_DIR is required deliberately; normal cargo build, cargo test, and package consumers never inspect or compile a sibling C checkout.
For strict, named feature intersections between the two closest Rust parsers, use the md4c-independent harness:
It provides CommonMark, GFM-overlap, and extended-overlap lanes with trusted raw-HTML semantics in both parsers. See the parity benchmark README for the exact feature matrix. Secure-default numbers remain separate because pulldown-cmark does not expose an equivalent trust boundary.
What you get
CommonMark conformance: With RenderPolicy::Trusted, Options::commonmark()
passes all 652 of 652 spec examples — enforced in CI by
commonmark_spec_trusted_full_conformance. The secure default
(RenderPolicy::Untrusted) intentionally escapes raw HTML and passes 577 of 652
(88.5%) — every failure is the safety boundary doing its job, not a parsing
gap. Run cargo test --test commonmark_spec -- --ignored --nocapture for the
per-section report of the secure default.
All five GFM extensions: Tables, strikethrough, task lists, autolink literals, disallowed raw HTML.
Beyond GFM: Reference and inline footnotes, definition lists, front matter extraction (---/+++), heading IDs (GitHub-compatible slugs), math spans ($/$$), highlight/mark syntax (==text==), superscript (^text^), subscript (~text~), and callouts (> [!NOTE], > [!WARNING], ...).
MDX support (opt-in via mdx feature): Segment and render .mdx files without a JavaScript toolchain. Covers 90%+ of real-world MDX patterns in Next.js, Docusaurus, and Astro.
Fine-grained options let you turn on exactly what you need:
allow_html · allow_link_refs · tables · merged_table_cells · table_column_widths · strikethrough · highlight · superscript · subscript · task_lists
autolink_literals · disallowed_raw_html · footnotes · inline_footnotes · front_matter
heading_ids · math · callouts · definition_lists · line_comments · indented_code_blocks · link_base_path
Syntax note: ferromark uses ~~text~~ for strikethrough, ~text~ for subscript, and ^text^ for superscript. Single-tilde strikethrough is intentionally not supported.
Merged table cells
merged_table_cells adds MultiMarkdown/iA-style horizontal spans to GFM pipe
tables. The number of directly adjacent pipes after a cell is its column span:
The last cell renders as <td colspan="2">0$</td>. ||| spans three
columns, and multiple cells in one row may be merged. Whitespace between pipes
preserves an explicit empty cell (| value | | next |). A merged cell uses
the alignment of its first covered column; body spans are clamped to the table
width and ragged rows are padded after the final span.
The flag requires tables and is disabled by default. With the flag off,
consecutive pipes retain standard GFM behavior and create empty cells.
Inline footnotes
inline_footnotes enables Pandoc-style ^[note text] independently of
reference footnotes:
The result needs context.^[This note can contain *inline Markdown*.]
The opening caret may be escaped as \^[literal]. Balanced brackets, links,
code spans, and soft line breaks are supported inside a note, but an inline
note is always one paragraph. The iA Presenter form [^Footnote text.] is not
accepted as an inline note because it is indistinguishable from ferromark's
existing [^label] reference syntax.
The HTML renderer numbers inline and reference notes together by first
appearance and emits their definitions in the document-end footnote section.
Presentation adapters should consume InlineEvent::InlineFootnote and flush
collected notes at their own slide boundary; the core HTML renderer does not
infer slides.
Definition lists
Enable definition_lists for PHP Markdown Extra-style terms and descriptions:
Term
: A definition with *inline Markdown*.
Markers may have up to three leading spaces and require whitespace after the colon. Continuation paragraphs and nested blocks must be indented to the description content; lazy continuation is supported only for paragraph text. The option is disabled by default and in every dialect constructor.
Line comments
Enable line_comments to omit source-only note lines from HTML:
Published text.
// Review this wording before publishing.
Only // at the physical line start (after at most three spaces) is a
comment. URLs, trailing //, code blocks, raw HTML blocks, and explicit
container-prefixed lines remain ordinary Markdown. Comment text remains in the
source and is not suitable for secrets.
Indented code blocks
Set indented_code_blocks: false for dialects that require fenced code blocks
and interpret four-space indentation as ordinary paragraph content. Fenced code
blocks remain available.
Markdown configuration
Start from the syntax contract you need, then enable individual extensions:
Options::minimal()keeps the smallest Markdown surface and disables raw HTML parsing, reference links, and every optional extension.Options::commonmark()enables CommonMark syntax, including reference links and raw HTML recognition.Options::gfm()adds the five GitHub Flavored Markdown extensions: tables, strikethrough, task lists, autolink literals, and disallowed raw HTML.
use Options;
let options = Options ;
let html = to_html_with_options;
table_column_widths is a separate, opt-in extension to GFM pipe tables. When
enabled, the relative number of dashes in each delimiter cell becomes a numeric
HTML column-width hint:
The example renders 25% and 75% <col> hints. Alignment colons are not counted.
No preset enables this interpretation because GFM otherwise treats delimiter
dash counts as formatting only. The extension accepts neither CSS nor arbitrary
HTML attributes. It composes with merged_table_cells: widths describe the
underlying table columns, while a merged cell spans those columns.
For documentation pipelines, parse() / parse_with_options() return the
rendered HTML together with the raw front matter block and a list of headings
(level, id, plain text) for table-of-contents rendering;
parse_with_renderer() adds the opt-in fenced-code renderer to the same pass.
link_base_path prefixes internal absolute link destinations (/…) for sites
deployed under a subpath; image sources and autolinks are not rewritten.
All constructors keep RenderPolicy::Untrusted. allow_html controls whether
raw HTML syntax is parsed; RenderPolicy independently controls whether parsed
HTML is preserved or escaped. Options::default() retains ferromark's
backward-compatible feature mix. Measure the configurations on your corpus with
cargo bench --bench options.
Trade-offs
ferromark is built for one job: turning Markdown into HTML as fast as possible. That focus means some things it deliberately skips:
- No AST access. You can't walk a syntax tree or write custom renderers against parsed nodes. If you need that, pulldown-cmark's iterator model or comrak's AST are better fits.
- No source maps. No byte-offset tracking for mapping HTML back to Markdown positions.
- HTML only. No XML, no CommonMark round-tripping, no alternative output formats.
These aren't planned. They'd compromise the streaming architecture that makes ferromark fast.
Rendering untrusted Markdown
The default RenderPolicy::Untrusted is the browser-facing safety boundary. It escapes all raw HTML and allows relative URLs plus a small set of non-script schemes (http, https, mailto, tel, and similar). URL schemes are checked after entity and control-character normalization, so spellings such as javascript: are blocked too.
let html = to_html;
Trusted documents and MDX can opt into passthrough explicitly:
use ;
let options = Options ;
let html = to_html_with_options;
disallowed_raw_html implements the narrower GFM tag filter in trusted mode. It is not a general-purpose HTML sanitizer and does not make arbitrary raw HTML safe by itself.
Upgrading from an older release? See the 0.2 migration guide for the rendering default and the fallible UTF-8 and MDX APIs, and the 0.3 migration guide for removed Cargo features and the integration APIs.
MDX support
MDX is the standard for component-driven docs in Next.js, Docusaurus, and Astro. Processing it usually requires a full JavaScript toolchain — Node.js, acorn, babel, the works.
ferromark takes a different approach: segment .mdx files into typed blocks and render them at native speed. No JS runtime. No AST.
Render — one call, full output
render() assembles the final output automatically: Markdown segments become HTML, JSX and expressions pass through unchanged, ESM and front matter are extracted separately.
use render;
let input = r#"import { Card } from './card'
---
title: Hello
---
# Hello World
<Card title="Example">
Markdown **inside** a component.
</Card>
{new Date().getFullYear()}
"#;
let output = render;
// output.body — HTML with JSX/expressions passed through
// output.esm — vec!["import { Card } from './card'\n"]
// output.front_matter — Some("title: Hello\n")
Use render_with_options() for custom Markdown settings (heading IDs, math, footnotes, etc.).
Component — ready-to-use JSX module
to_component() wraps the output as a complete JSX/TSX module with a named export. Works with React 19, Preact, Solid, and any JSX framework.
let output = render;
let tsx = output.to_component?;
import { Card } from './card'
export function HelloWorld() {
return (
<>
<h1 id="hello-world">Hello World</h1>
<Card title="Example">
<p>Markdown <strong>inside</strong> a component.</p>
</Card>
{new Date().getFullYear()}
</>
);
}
Segment — low-level control
When you need full control over each block, use segment() directly:
use ;
for seg in segment
The segmenter handles JSX attribute parsing (strings, expressions, spreads), brace-depth tracking (with string/comment/template-literal awareness), fragment syntax, member expressions (<Foo.Bar>), and multiline tags. Invalid constructs fall back to Markdown — no panics, always valid output.
For source locations, segment_spanned() returns the same zero-copy segments
with contiguous Range
values into the original UTF-8 input. A range covers the exact segment text,
including delimiters and a trailing newline when it belongs to that segment.
segment() remains deliberately permissive: malformed MDX falls back to a
Markdown segment. For content pipelines that must reject malformed structure,
use segment_strict(). It reports typed diagnostics with byte ranges; convert
an offset to a one-based line and Unicode column only when presenting it to a
user with source_location().
use ;
let input = "<Card bad=>\n";
let diagnostics = segment_strict.unwrap_err;
let location = source_location;
assert_eq!;
Strict mode checks MDX structure (flow-expression delimiters, JSX tag shape and nesting, and ESM placement). It intentionally does not parse or type-check JavaScript or TypeScript inside an otherwise well-delimited ESM block or expression.
Compiler and localization consumers can opt into a flat semantic event buffer
without going through HTML. parse_events() composes the existing MDX
segmenter, block parser, and MDX-aware inline parser; ranges in the returned
events point into the original input.
use InlineEvent;
use ;
let input = "# Hello {name}\n";
let stream = parse_events;
let prose = stream.events.iter.filter_map.;
assert_eq!;
The event path is fully opt-in and does not alter the normal Markdown or MDX
HTML renderer. parse_events_strict() applies the same structural diagnostics
as segment_strict() before producing events. Its strict validation covers
flow constructs; malformed inline MDX retains the documented text fallback.
Inside blockquotes and list items, a paragraph containing only one JSX tag or
expression is promoted to the corresponding flow event. Mixed prose remains
inline MDX, and the surrounding Markdown container events stay balanced.
Full example: cargo run --features mdx --example mdx_segment
The segmenter covers the block-level MDX patterns that make up 90%+ of real-world .mdx files: imports at the top, components wrapping content, expressions between paragraphs. This is what a typical Docusaurus, Next.js, or Astro page looks like — and it works out of the box.
What the segmenter deliberately skips — and why that's fine for most use cases:
| What | Our approach | When it matters |
|---|---|---|
Inline JSX (text <em>here</em>) |
Stays in segment() Markdown blocks; parse_events() and InlineParser::parse_mdx() expose typed MDX inline events |
Use the opt-in event APIs when a downstream consumer must distinguish prose and components |
| JS validation | Heuristic detection (keyword + brace counting) instead of acorn/swc | Only if you need to report syntax errors in user-authored MDX at parse time |
| Markdown grammar | Standard CommonMark/GFM rules | Official mdxjs disables indented code and HTML syntax — relevant if your content relies on <div> being JSX, not HTML |
| Container nesting | > <Component> stays Markdown to the renderer; parse_events() promotes tag-only or expression-only container paragraphs to semantic flow events |
Rendering-level container MDX, multiline constructs across prefixes, and container-local ESM remain out of scope |
| TypeScript generics | <Component<T>> not parsed |
Only relevant for TSX-heavy content pages — very rare in docs |
| Error reporting | Permissive fallback by default; opt-in structural diagnostics with segment_strict() |
Use strict mode when broken MDX must fail a content pipeline |
The full @mdx-js/mdx compiler exists to produce a React component tree from MDX. It needs a JavaScript parser because it compiles to JSX. ferromark's segmenter exists to answer a simpler question: where does the Markdown stop and the JSX start? That question doesn't need a JS runtime.
For the detailed technical spec, see src/mdx/mod.rs.
How it works
No AST. Block events stream from the scanner to the HTML writer with nothing in between.
Input bytes (&[u8])
│
▼
Block parser (line-oriented, memchr-driven)
│ emits BlockEvent stream
▼
Inline parser (mark collection → resolution → emit)
│ emits InlineEvent stream
▼
HTML writer (direct buffer writes)
│
▼
Output (Vec<u8>)
What makes this fast in practice:
- Block scanning runs on
memchrfor line boundaries. Container state is a compact stack, not a tree. - Inline parsing has three phases: collect delimiter marks, resolve precedence (code spans, math, links, emphasis, strikethrough, subscript, superscript, highlight), emit. No backtracking.
- Emphasis resolution uses the CommonMark modulo-3 rule with a delimiter stack instead of expensive rescans.
- SIMD scanning: NEON (AArch64) drives the inline specials scan; the HTML escaper uses SSE2 (x86-64) and NEON short scans; memchr covers the remaining byte searches on every architecture.
- Zero-copy references: events carry
Rangepointers into the input, not copied strings. - Compact events: 24 bytes each, cache-line friendly.
- Hot/cold annotation:
#[inline]on tight loops,#[cold]on error paths, table-driven byte classification.
Design principles
- Linear time. No regex, no backtracking, no quadratic blowup on adversarial input.
- Low allocation pressure. Compact events, range references, reusable output buffers.
- Operational safety. Enforced limits cap block nesting (32), inline marks (4,096), code-span backtick runs (32), link-destination parenthesis depth (32), ordered-list marker digits (9), and table columns (128). Footnote numbering has no arbitrary count cap; its definition-index lookup stays O(1) per reference.
- Small dependency surface. Minimal crates, straightforward integration.
How ferromark compares to the other three top-tier parsers across architecture, features, and output. Ratings use a 4-level heatmap focused on end-to-end Markdown-to-HTML throughput. Scoring is relative per row, so each row has at least one top mark.
Legend: 🟩 strongest 🟨 close behind 🟧 notable tradeoffs 🟥 weakest
Ferromark optimization backlog: docs/arch/ARCH-PLAN-001-performance-opportunities.md
Building
Project structure
src/
├── lib.rs # Public API (to_html, to_html_into, parse, Options)
├── main.rs # CLI binary
├── block/ # Block-level parser
│ ├── parser.rs # Line-oriented block parsing
│ └── event.rs # BlockEvent types
├── inline/ # Inline-level parser
│ ├── mod.rs # Three-phase inline parsing
│ ├── marks.rs # Mark collection + SIMD integration
│ ├── simd.rs # SIMD character scanning (NEON)
│ ├── event.rs # InlineEvent types
│ ├── code_span.rs
│ ├── emphasis.rs # Modulo-3 stack optimization
│ ├── strikethrough.rs # GFM strikethrough resolution
│ ├── subscript.rs # Subscript resolution (~text~)
│ ├── superscript.rs # Superscript resolution (^text^)
│ ├── math.rs # Math span resolution ($/$$ delimiters)
│ └── links.rs # Link/image/autolink parsing
├── mdx/ # MDX segmenter + renderer (feature = "mdx")
│ ├── mod.rs # Public API — Segment enum, segment(), render()
│ ├── render.rs # Assembly layer: segments → HTML body + ESM + front matter
│ ├── splitter.rs # Line-based state machine
│ ├── jsx_tag.rs # JSX tag boundary parser
│ └── expr.rs # Expression boundary parser (brace/string/comment tracking)
├── footnote.rs # Footnote store and rendering
├── link_ref.rs # Link reference definitions
├── cursor.rs # Pointer-based byte cursor
├── range.rs # Compact u32 range type
├── render.rs # HTML writer
├── escape.rs # HTML escaping (SSE2/NEON + memchr)
└── limits.rs # DoS prevention constants
License
MIT
The Ferramenta family
This project is part of Ferramenta — the family of Rust-native developer tools by Sebastian Software that keep the APIs the ecosystem already knows:
| Tool | Job |
|---|---|
| ferroni | Oniguruma-compatible regex engine |
| ferriki | Shiki-compatible syntax highlighting |
| ferromark | CommonMark/GFM Markdown to HTML |
| ferrovia | SVGO-compatible SVG optimizer |
| ferrocat | Translation catalog engine |
| ferrolex | Spell, dictionary, and brand validation |
| ferrugo | Rust-native PDF previews |