Expand description
§fmt_lang
A source formatter driven by declarative rules over a lossless syntax
tree, rendering through pretty_lang.
A language gets a formatter from configuration, not code: describe how
its nodes and tokens are laid out as Rules (data, keyed by kind names,
the way a sketch names them), compile them against the
language’s kinds once, and format() any syntax_lang tree of that
language. Whatever the rules do not mention keeps its original whitespace.
§Quick start
use fmt_lang::{format, Indent, NodeRule, Rules, Space, TokenRule, Trailing};
use lang_forge::Language;
let json = Language::from_lsf(
r#"
[language]
name = "json"
[lexer]
strings = ['"']
line_comments = ["//"]
[rules]
document = "value"
value = "object | array | STRING | NUMBER | 'true' | 'false' | 'null'"
object = "'{' (member (',' member)* ','?)? '}'"
member = "STRING ':' value"
array = "'[' (value (',' value)* ','?)? ']'"
"#,
)?;
let style = Rules::new()
.indent(2)
.verbatim("ERROR")
.token(TokenRule::new(":").before(Space::None).after(Space::Single))
.node(
NodeRule::new("object")
.group()
.indent(Indent::Block)
.delimiters("{", "}", Space::Line)
.separator(",", Space::Line, Trailing::Never),
)
.node(
NodeRule::new("array")
.group()
.indent(Indent::Block)
.delimiters("[", "]", Space::SoftLine)
.separator(",", Space::Line, Trailing::Never),
)
.compile(|name| json.kind(name))?;
let parse = json.parse("{\"a\":[1,2,], \"b\" :{}}");
let wide = format(parse.tree(), parse.source(), &style, 80)?;
assert_eq!(wide, "{ \"a\": [1, 2], \"b\": {} }\n");
// Too wide for 16 columns: the object breaks, the array still fits.
let narrow = format(parse.tree(), parse.source(), &style, 16)?;
assert_eq!(narrow, "{\n \"a\": [1, 2],\n \"b\": {}\n}\n");§The rule model
- A
NodeRulemakes a node a group (flat if it fits, broken at its line opportunities otherwise), indents its body or its whole content, asks for spacing before and after the node, names the delimiters and separator of a list (with aTrailingseparator policy), caps blank lines, and carries token rules for its direct child tokens. - A
TokenRuleasks for spacing before and after one token kind, at the top level (a default everywhere) or inside a node rule (that context only, and it wins). - Spacing is a
Space; when several rules speak about one gap, the most generous answer wins. When no rule speaks, the gap keeps its original whitespace (with trailing spaces at line ends removed). Rules::verbatimnames node kinds written exactly as in the source, such as a parser’sERRORnodes.
§Guarantees
Property-tested over random valid and invalid sources of forged languages
and over arbitrary trees (see tests/):
- Idempotent: formatting formatted output changes nothing.
- Token-preserving: the significant token sequence is unchanged (the
only exception is a trailing separator that a
Trailing::AlwaysorTrailing::Neverpolicy adds or removes), and every comment’s text appears exactly once, in its original order. A line comment always ends its line. How comments are attached is documented onformat(). - Total: any tree whose tokens describe the source formats without
panicking, error nodes included; a tree that does not match its source is
a
FormatError, never a panic. - Deterministic: the same input gives the same output.
For sources with lexical errors (an unterminated string or comment), the
guarantees hold when the errors’ regions are passed to format_keeping:
such a token’s extent depends on the whitespace after it, which the
formatter cannot see from the tree.
The walk is iterative and linear in the size of the tree, so trees nested
hundreds of thousands of levels deep format without exhausting the stack;
indentation stops growing at Rules::max_indent, which bounds the output
for hostile input.
§Limits
- Whether two tokens may touch is decided by
can_touch, a conservative heuristic, unless the style supplies an exact test (Style::with_touch). - A parser’s recovery that leaves no trace in the tree (an assumed missing
token) is invisible here.
Trailing::AlwaysandTrailing::Neveredit only lists with no error node and no unclosed delimited node inside, but a recovery these checks miss could still read an edited list differently;Trailing::Preservenever edits. - In languages whose line breaks are significant tokens, rules that break
lines (
Space::Line,Space::Hard) would add tokens; use them only where the language allows a line break. - Indentation is spaces; widths are counted in
chars (pretty-lang’s measure), not display columns.
§Features
std(default): the standard library, forwarded tosyntax-langandpretty-lang. Without it the crate isno_stdand needs onlyalloc.
Re-exports§
pub use pretty_lang;pub use syntax_lang;
Structs§
- Node
Rule - Layout rules for one node kind.
- Rules
- A language’s formatting rules, as data.
- Style
- Formatting rules compiled against one language, ready for
format(). - Token
Rule - Spacing rules for one token kind (or for every token), on either side.
Enums§
- Format
Error - Why
format()could not format a tree. - Indent
- How a node indents its contents.
- Rule
Error - Why
Rules::compilerefused a set of rules. - Space
- What goes between two neighbouring pieces of output.
- Trailing
- What to do with a separator after the last element of a delimited list.
Constants§
- MAX_
INDENT_ STEP - The widest indentation step
Rules::indentaccepts.
Functions§
- can_
touch - The default test for whether two tokens may be written touching, with no
whitespace between them:
leftis the earlier token’s text,rightthe later one’s. - format
- Formats the source
treewas parsed from, atwidthcolumns. - format_
doc - Builds the
Docthatformat_keeping(or, with nokeepregions,format()) renders, for callers that render it themselves: into an existing buffer or writer, or at several widths. - format_
keeping - Formats like
format(), but leaves the whitespace around and inside thekeepregions exactly as written.