formualizer-parse 3.2.0

High-performance Excel/OpenFormula tokenizer + parser with a stable AST surface
Documentation

Formualizer banner

formualizer-parse

Arrow Powered

High-performance Excel and OpenFormula tokenizer, parser, and pretty-printer.

formualizer-parse turns raw formula strings into a structured AST that downstream crates use for evaluation, analysis, and transformation. It handles both Excel and OpenFormula dialects with source location tracking.

When to use this crate

Use formualizer-parse when you need formula analysis without evaluation:

  • Formula linting and validation
  • Static analysis of cell dependencies
  • AST transformation and rewriting
  • Pretty-printing formulas to canonical form
  • Building custom formula tooling

If you also need evaluation, use formualizer-workbook or formualizer-eval instead.

Quick start

use formualizer_parse::{FormulaDialect, Parser, canonical_formula, parse_with_dialect};

// One-shot parse
let ast = parse_with_dialect("=SUM(A1:B3)", FormulaDialect::Excel)?;

// Or use the stateful source-span parser directly
let mut parser = Parser::new("=SUM(A1:B3)")?;
let ast = parser.parse()?;

// Canonical form
assert_eq!(canonical_formula(&ast), "=SUM(A1:B3)");

Features

  • Tokenization — streaming tokenizer with dialect-aware classification, source location tracking, and operator metadata.
  • Pratt parser — precedence-climbing parser producing a stable AST with reference normalization.
  • Dialects — Excel (default) and OpenFormula syntax support through a single API.
  • Pretty-printing — canonicalize formulas or render diagnostic trees for debugging.
  • Source spans — every token and AST node carries byte positions for precise error reporting.
  • Fingerprinting — 64-bit structural hashes for formula identity comparison.

Resource limits

Default parsing admits at most 64 KiB of UTF-8 source bytes, 16,384 tokens, 8,192 AST nodes (including omitted arguments), 72 active Pratt frames and AST height 256. Height counts the root as one; parentheses do not add AST nodes. Limits apply independently, so a formula can hit one before another. Errors retain the existing parser/tokenizer error types.

Start from ParserLimits::default() and adjust budgets with the with_source_bytes, with_tokens, with_ast_nodes, with_pratt_frames and with_ast_height setters, then pass the value to Parser::builder().limits(limits) (or BatchParser::builder().limits(limits)). Values are used exactly as given; nothing is clamped. Raising pratt_frames or ast_height above the defaults requires a correspondingly larger stack on every thread that parses, clones, drops or evaluates the tree. These guarantees concern parser-produced trees, not manually constructed or externally deserialized ASTs.

use formualizer_parse::{Parser, ParserLimits};

let limits = ParserLimits::default().with_ast_nodes(1_024);
let ast = Parser::builder().limits(limits).parse("=SUM(A1:B3)").unwrap();

Stack cost

Parsing a left-associated chain (A1+A2+..., &, ^) is iterative. However, cloning, printing, hashing, dropping and evaluating the tree recurse once per level. Right-nested shapes (-(-(...)), 1+(1+(...)), nested IF) also recurse in the parser, which is why pratt_frames bounds them separately. Measured on x86-64 Linux release builds (Rust 1.93), per AST level:

Path Stack per level
Drop, hash/fingerprint, dependency collection 48–80 bytes
Clone, pretty_print, canonical_formula 0.3–0.7 KiB
Parse (right-nested shapes only) 2.7–4.2 KiB
formualizer-workbook set_formula (arena ingest and dependency analysis) 1.7 KiB (chains), up to 4.2 KiB (right-nested)
Workbook evaluation, Calamine XLSX load plus evaluation, XLSX cache recalculation 2.5 KiB (chains), up to 4.2 KiB (right-nested)

XLSX cache recalculation is the deepest path. A default-height (256) formula needs about 0.9 MiB of stack there, including roughly 190 KiB of fixed overhead. The defaults are therefore tested on a 1 MiB thread stack in release builds, which is the Windows main-thread and wasm default. Unoptimized (debug) builds use several times more stack per level: about 5 MiB natively for height 256, and the default 1 MiB wasm stack fits only about 100 levels. If you raise ast_height, budget about 2.5 KiB per extra level (4.2 KiB for right-nested shapes) on every thread that parses, clones, drops or evaluates the tree. Engine and workbook parse sites use the defaults.

parser::BatchParser retains lexical results in a FIFO cache bounded by 16,384 entries and 8 MiB of source/token payload (excluding container overhead). Hits do not reorder entries. cache_capacity(entries, bytes) configures retention; zero entries disables it. Entries exceeding the byte capacity still parse without retention. Eviction does not change formula semantics, but a working set larger than either bound may lose token-cache reuse. Tune the capacity for such workloads.

Externally supplied TokenStream.spans must be ordered, disjoint and valid UTF-8 byte ranges. Stream-to-parser and stream-to-owned-tokenizer admission validates these before copying. Infallible best-effort stream tokenization reports resource failures through its diagnostics; Tokenizer::new_best_effort and Tokenizer::from_token_stream expose them through admission_error(), with empty output on admission failure.

These are resource policies, not exact Excel compatibility limits: a valid Excel formula with more than 256 operands in a left-associated chain exceeds the default AST height even if its text fits Excel's character limit.

License

Dual-licensed under MIT or Apache-2.0, at your option.