Skip to main content

Crate fmt_lang

Crate fmt_lang 

Source
Expand description

§fmt_lang

A source formatter driven by declarative rules over a lossless syntax tree, rendering through pretty_lang.

A language gets a formatter from configuration, not code: describe how its nodes and tokens are laid out as Rules (data, keyed by kind names, the way a sketch names them), compile them against the language’s kinds once, and format() any syntax_lang tree of that language. Whatever the rules do not mention keeps its original whitespace.

§Quick start

use fmt_lang::{format, Indent, NodeRule, Rules, Space, TokenRule, Trailing};
use lang_forge::Language;

let json = Language::from_lsf(
    r#"
    [language]
    name = "json"

    [lexer]
    strings = ['"']
    line_comments = ["//"]

    [rules]
    document = "value"
    value    = "object | array | STRING | NUMBER | 'true' | 'false' | 'null'"
    object   = "'{' (member (',' member)* ','?)? '}'"
    member   = "STRING ':' value"
    array    = "'[' (value (',' value)* ','?)? ']'"
    "#,
)?;

let style = Rules::new()
    .indent(2)
    .verbatim("ERROR")
    .token(TokenRule::new(":").before(Space::None).after(Space::Single))
    .node(
        NodeRule::new("object")
            .group()
            .indent(Indent::Block)
            .delimiters("{", "}", Space::Line)
            .separator(",", Space::Line, Trailing::Never),
    )
    .node(
        NodeRule::new("array")
            .group()
            .indent(Indent::Block)
            .delimiters("[", "]", Space::SoftLine)
            .separator(",", Space::Line, Trailing::Never),
    )
    .compile(|name| json.kind(name))?;

let parse = json.parse("{\"a\":[1,2,],   \"b\" :{}}");
let wide = format(parse.tree(), parse.source(), &style, 80)?;
assert_eq!(wide, "{ \"a\": [1, 2], \"b\": {} }\n");

// Too wide for 16 columns: the object breaks, the array still fits.
let narrow = format(parse.tree(), parse.source(), &style, 16)?;
assert_eq!(narrow, "{\n  \"a\": [1, 2],\n  \"b\": {}\n}\n");

§The rule model

  • A NodeRule makes a node a group (flat if it fits, broken at its line opportunities otherwise), indents its body or its whole content, asks for spacing before and after the node, names the delimiters and separator of a list (with a Trailing separator policy), caps blank lines, and carries token rules for its direct child tokens.
  • A TokenRule asks for spacing before and after one token kind, at the top level (a default everywhere) or inside a node rule (that context only, and it wins).
  • Spacing is a Space; when several rules speak about one gap, the most generous answer wins. When no rule speaks, the gap keeps its original whitespace (with trailing spaces at line ends removed).
  • Rules::verbatim names node kinds written exactly as in the source, such as a parser’s ERROR nodes.

§Guarantees

Property-tested over random valid and invalid sources of forged languages and over arbitrary trees (see tests/):

  • Idempotent: formatting formatted output changes nothing.
  • Token-preserving: the significant token sequence is unchanged (the only exception is a trailing separator that a Trailing::Always or Trailing::Never policy adds or removes), and every comment’s text appears exactly once, in its original order. A line comment always ends its line. How comments are attached is documented on format().
  • Total: any tree whose tokens describe the source formats without panicking, error nodes included; a tree that does not match its source is a FormatError, never a panic.
  • Deterministic: the same input gives the same output.

For sources with lexical errors (an unterminated string or comment), the guarantees hold when the errors’ regions are passed to format_keeping: such a token’s extent depends on the whitespace after it, which the formatter cannot see from the tree.

The walk is iterative and linear in the size of the tree, so trees nested hundreds of thousands of levels deep format without exhausting the stack; indentation stops growing at Rules::max_indent, which bounds the output for hostile input.

§Limits

  • Whether two tokens may touch is decided by can_touch, a conservative heuristic, unless the style supplies an exact test (Style::with_touch).
  • A parser’s recovery that leaves no trace in the tree (an assumed missing token) is invisible here. Trailing::Always and Trailing::Never edit only lists with no error node and no unclosed delimited node inside, but a recovery these checks miss could still read an edited list differently; Trailing::Preserve never edits.
  • In languages whose line breaks are significant tokens, rules that break lines (Space::Line, Space::Hard) would add tokens; use them only where the language allows a line break.
  • Indentation is spaces; widths are counted in chars (pretty-lang’s measure), not display columns.

§Features

  • std (default): the standard library, forwarded to syntax-lang and pretty-lang. Without it the crate is no_std and needs only alloc.

Re-exports§

pub use pretty_lang;
pub use syntax_lang;

Structs§

NodeRule
Layout rules for one node kind.
Rules
A language’s formatting rules, as data.
Style
Formatting rules compiled against one language, ready for format().
TokenRule
Spacing rules for one token kind (or for every token), on either side.

Enums§

FormatError
Why format() could not format a tree.
Indent
How a node indents its contents.
RuleError
Why Rules::compile refused a set of rules.
Space
What goes between two neighbouring pieces of output.
Trailing
What to do with a separator after the last element of a delimited list.

Constants§

MAX_INDENT_STEP
The widest indentation step Rules::indent accepts.

Functions§

can_touch
The default test for whether two tokens may be written touching, with no whitespace between them: left is the earlier token’s text, right the later one’s.
format
Formats the source tree was parsed from, at width columns.
format_doc
Builds the Doc that format_keeping (or, with no keep regions, format()) renders, for callers that render it themselves: into an existing buffer or writer, or at several widths.
format_keeping
Formats like format(), but leaves the whitespace around and inside the keep regions exactly as written.