Expand description
§lgwks_ast
A multi-language AST front end for tools that read source code.
A code tool needs the same three things: identify a source file’s language by
extension, by its #! line, or by explicitly trial-parsing a small candidate
set, select that language’s tree-sitter grammar, and walk the resulting syntax
tree safely. This crate provides those paths under explicit bounds. It does not
parse full Vue/Svelte containers; those extensions select the JavaScript
grammar as a filename heuristic.
It also provides the shared typed-diagnostic derive, so a parser and its
consumers report failures through one Display/source implementation rather
than each declaring thiserror.
§Usage
use lgwks_ast::{Language, inspect_ast, try_parse};
fn main() -> Result<(), lgwks_ast::ParseError> {
// Identify by extension — the cheap, always-on path.
let language = Language::of_path("src/lib.rs");
assert_eq!(language, Some(Language::Rust), "`.rs` resolves without parsing");
// Content sniffing is opt-in, because it costs one full parse per distinct
// candidate. Repeated candidates do not repeat parsing, and an unavailable or
// over-budget candidate returns an error instead of being treated as a failed
// syntax match. Its result distinguishes NoMatch, Unique(language), and
// Ambiguous:
// let detected = try_detect_content(source, &[Language::Rust, Language::Python])?;
// Checked parse: oversized bytes, recovery nodes, and over-budget trees are
// refused before any consumer sees the tree.
let parsed = try_parse("fn f() {}", Language::Rust)?;
let metrics = inspect_ast(&parsed.root(), None);
assert!(metrics.nodes > 1, "a `fn` item is more than one node");
Ok(())
}Every line of that block is compiled and run as a doctest, against this crate’s
own feature set. try_parse returns Result<Parsed, ParseError> — the ? is
required, not decorative; binding it directly and calling .root() does not
compile.
Consumers do not depend on ast-grep directly; the grammar types
(AstGrep, Node, StrDoc, SupportLang) are re-exported here as
Parsed, AstNode, and friends.
§Typed diagnostics
ParseError is built with this crate’s diagnostic derive, and an analyser
derives its own diagnostics from the same stack:
extern crate lgwks_ast as thiserror;
#[derive(thiserror::Error, Debug)]
enum Diagnostic {
#[error("parse refused: {0}")]
Refused(String),
#[error(transparent)]
Io(#[from] std::io::Error),
}#[derive(Error)] expands to absolute ::thiserror::__private<N>::… paths
resolved in the consuming crate, so the extern crate … as thiserror; line is
what makes the expansion resolve; the crate root re-exports the module it
needs.
§Grammar selection
One cargo feature per grammar forwards to ast-grep-language. The default
enables seven languages; every other grammar ast-grep-language ships is its own
opt-in feature, and full enables all 28:
[dependencies]
lgwks_ast = { version = "0.4.0", features = ["lang-c", "lang-cpp", "lang-scala", "lang-kotlin", "lang-tsx"] }
# or everything: features = ["full"]0.2.0 renames one variant. Language::C is now Language::CLang. The
workspace lint contract forbids single-character identifiers, and that lint is
forbid rather than deny, so it cannot be lowered from source by an
#[allow] on the variant. The variant itself had to change. Language::name()
still reports "c", which is the stable identity findings match on, and the
lang-c feature is unchanged, so only Rust code that names the variant is
affected.
| Feature | Language | Extensions | Default |
|---|---|---|---|
lang-rust | Rust | rs | yes |
lang-python | Python | py, pyi | yes |
lang-typescript | TypeScript | ts, mts, cts | yes |
lang-javascript | JavaScript | js, jsx, mjs, cjs, vue, svelte | yes |
lang-go | Go | go | yes |
lang-java | Java | java | yes |
lang-swift | Swift | swift | yes |
lang-tsx | TSX (parsed by the TypeScript grammar; shares it with lang-typescript) | tsx | no |
lang-c | C | c, h | no |
lang-cpp | C++ | cpp, hpp, cc, cxx, hh, hxx | no |
lang-csharp | C# | cs | no |
lang-scala | Scala | scala, sc | no |
lang-kotlin | Kotlin | kt, kts | no |
lang-bash | Bash / POSIX shell | sh, bash | no |
lang-css | CSS | css | no |
lang-dart | Dart | dart | no |
lang-elixir | Elixir | ex, exs | no |
lang-haskell | Haskell | hs | no |
lang-hcl | HCL / Terraform | hcl, tf | no |
lang-html | HTML | html, htm | no |
lang-json | JSON | json | no |
lang-lua | Lua | lua | no |
lang-md | Markdown | md, markdown | no |
lang-nix | Nix | nix | no |
lang-php | PHP | php | no |
lang-ruby | Ruby | rb | no |
lang-solidity | Solidity | sol | no |
lang-yaml | YAML | yaml, yml | no |
full | all of the above | — | no |
A language whose feature is off does not exist in Language::ALL and is never
returned by Language::of_path, so a consumer never compiles a grammar it
cannot select.
Extension lookup is a filename heuristic, not a syntax verdict. In particular,
.vue and .svelte select the JavaScript grammar; they do not enable dedicated
container grammars. An extensionless script is identified by its #! line
through Language::of_shebang, which detect does not consult.
try_detect_content returns ContentDetection::NoMatch, Unique(language), or
Ambiguous; non-syntax parser failures return Err because they leave a
required candidate uninspected.
§Custom languages
ast-grep ships a fixed built-in set; its documented extension point for
anything else is a caller-registered parser. Register one with CustomLang and
parse it through try_parse_with, under the same byte bound, node bound, and
recovery refusal as a built-in:
use lgwks_ast::{CustomLang, try_parse_with};
use some_sql_grammar::LANGUAGE; // the caller's own grammar crate
let sql = CustomLang::new("sql", LANGUAGE.into()).with_extensions(&["sql"]);
let parsed = try_parse_with("SELECT 1", &sql, sql.name())?;This block is ignored, and that is the honest encoding of what it shows:
LANGUAGE comes from a grammar crate the caller depends on, which this
crate deliberately does not, so no doctest here could compile it. What the
block does prove is compiled — try_parse_with accepts a CustomLang under
the same byte, node and recovery bounds as a built-in grammar, and the
surrounding #[test]s exercise it with an in-crate grammar.
TSLanguage is re-exported, so a consumer adds only the grammar crate, never
ast-grep-core directly. A grammar crate is a third-party edge like any
other: register it in contract/APPROVED.toml (owner, capability, reason)
before depending on it; lgwks-deps check refuses it otherwise.
A grammar that belongs upstream should be contributed to ast-grep-language
itself, following its
add-a-language guide:
it must be popular (TIOBE / GitHub Octoverse), use a maintained grammar
published on crates.io, and stay inside the budget that keeps ast-grep’s zipped
binary under 10 MB. Until it ships there, register it here with CustomLang;
when it does, delete the registration and select the built-in. This crate does
not fork upstream’s SupportLang tables, so every addition remains submittable
upstream as a pull request.
§Bounds
MAX_SOURCE_BYTES: 2 MiB per checked parse. This bounds the bytes handed to tree-sitter and keeps parse work linear in input.MAX_AST_NODES: 2,000,000 nodes; the boundary walk is capped and stops within one node of the limit, so refusal does not itself walk an unbounded tree. It is measured on the tree after tree-sitter builds it, so it bounds the validation walk, not the parser’s own allocation.MAX_DETECT_BYTES: 64 KiB per content-detection probe.try_detect_contenttries each distinct caller-named candidate in full, so the probe is bounded well belowMAX_SOURCE_BYTES;detectnever parses at all.MAX_SHEBANG_BYTES: 256 bytes — the#!line is the only part of a fileof_shebangreads, and a script whose first line is longer than that is treated as unidentified rather than scanned further.try_parserefuses anERROR/MISSINGrecovery node: a recoverable tree is not proof of valid syntax.parseis the unchecked escape hatch for diagnostics and tests that inspect malformed trees on purpose.- Checked syntax refusals carry at most
MAX_SYNTAX_DIAGNOSTICSrecovery-node diagnostics. Each identifiesERRORversusMISSINGand a half-open byte range into the original UTF-8 source; no source text is copied into the diagnostic. - A tree is not only accepted or refused.
diagnosticsandto_diagnosticturn a refused or suspect parse into aDiagnosticcarrying a byte offset, a 1-based line and column, and a caret span, so a caller reports where the source stopped making sense rather than only that it did. A refused parse points at its earliest recovery node; only whole-file refusals (size, parser, node budget) sit at the end of the file. inspect_astreports whether traversal completed, the applied node limit, and its stop reason. A partial walk’shas_syntax_issues == falseis not a clean-syntax result.
§The other crates
Four crates ship from this repository. They share a release process, not a
dependency graph: lgwks_bot and lgwks_deps depend on lgwks_std, and
lgwks_ast stands alone.
| Crate | What it gives you |
|---|---|
lgwks_std | Everyday primitives with no async runtime required: codecs, a blocking HTTP client, retry, default structured debugging, time, hashing, ids |
lgwks_bot | A runtime for bots that run for weeks: four verbs, capability-gated authority, change-triggered execution, supervised background work |
lgwks_deps | The audited storefront for third-party stacks, plus lgwks-deps check and lgwks-deps debug to prove dependency and debugger wiring |
The repository README indexes the design documents.
§License
Apache-2.0 — Copyright 2026 Logical Works Incorporated
lgwks_ast owns the single multi-language AST parser.
Code tools that consume this crate, a safety linter and a graph extractor among them, need the same three things: decide a source file’s language, select that language’s tree-sitter grammar, and walk the resulting syntax tree safely. Before this crate each rebuilt its own language enum, extension table, grammar mapping, and parse loop; that is one concept implemented twice, so it lives here instead.
Enforced invariant INV-AST-ONE-PARSER: a consumer identifies, selects,
and parses through this crate and does not depend on ast-grep directly.
§The README is compiled
This module documentation is generated from README.md, so the usage
example above is built and run as a doctest on every cargo test rather
than being prose that can quietly rot. cargo test -p lgwks_ast --doc
therefore fails if the example stops compiling or stops asserting what the
text above it claims.
§Code observability
Code observability here rests on two pieces: structural parsing, and the
shared typed-diagnostic derive in error. ParseError is built with
it, and downstream analysers derive Display/source/#[from] from the
same stack instead of each declaring thiserror. The crate root re-exports
the derive because #[derive(Error)] expands to absolute
::thiserror::__private<N>::… paths resolved in the consuming crate: a
consumer names this crate thiserror (extern crate lgwks_ast as thiserror;)
and derives normally, with no thiserror edge of its own.
§Grammar selection
One cargo feature per grammar forwards to ast-grep-language. The default
enables seven widely-used languages — Rust, Python, TypeScript, JavaScript,
Go, Java and Swift — and every remaining grammar ast-grep-language ships
(C#, CSS, Dart, Elixir, Haskell, HCL, HTML, JSON, Lua, Markdown, Nix, PHP,
Ruby, Solidity, YAML, and the rest) is its own opt-in lang-* feature, and
full enables all 28. A language whose feature is off is not in
Language::ALL and is never returned by Language::of_path, so no
consumer pays to compile a grammar it cannot select.
Two notes on reading that table. lang-tsx is a selection, not a separate
grammar: ast-grep parses .tsx with the TypeScript parser, so it forwards to
the same crate as lang-typescript and pulls nothing further. And a #!
line resolves a language where a filename cannot — see
Language::of_shebang, which is how an extensionless bin/deploy is
identified at all.
§Custom languages
ast-grep keeps its built-in set small on purpose; its documented extension
point for anything else is a caller-registered parser. CustomLang is
that registration: name the language, hand it a tree-sitter grammar, and
parse it through try_parse_with under the same bounds and recovery
refusal as a built-in. A grammar a consumer here needs should be contributed
to ast-grep-language upstream and the local registration deleted once it
ships. This crate never forks upstream’s language tables.
§Bounded parsing
A recoverable tree-sitter tree is not proof of valid syntax: recovery emits
ERROR and MISSING nodes. try_parse therefore refuses before any
detector sees the tree: oversized bytes, a parser that cannot produce a
tree, a tree past MAX_AST_NODES, or one carrying recovery nodes. The
unchecked parse exists for diagnostics and tests that inspect
malformed trees on purpose.
The two bounds are not interchangeable. MAX_SOURCE_BYTES bounds the
bytes handed to the parser and is what keeps parse work linear in input;
MAX_AST_NODES is measured on the tree after tree-sitter has built it,
so it bounds the validation walk and every downstream walk, not the
parser’s own allocation. Neither is a hard memory ceiling.
Content sniffing is opt-in for the same reason: try_detect_content
trial-parses each distinct candidate grammar in full, so the caller names
a small candidate set and the probe source is held to MAX_DETECT_BYTES. The
extension-only detect never parses.
§The README is compiled
README.md is attached to this crate root as its documentation, so the
usage example it opens with is built and run as a doctest on every
cargo test -p lgwks_ast --doc instead of being prose that can rot. When
the example above stops compiling — as it did while try_parse was
documented as if it were infallible — that doctest is what catches it.
Re-exports§
pub use diagnostic::Diagnostic;pub use diagnostic::Pos;pub use diagnostic::Severity;pub use diagnostic::Span;
Modules§
- diagnostic
- Spans, severities, and
diagnostics: what a tool reports on code, as opposed to what a parse refuses. Spans, severities, and the diagnostics a tool reports on code. - error
- Typed diagnostics:
ParseErrorand the shared error derive.errorowns this crate’s typed-diagnostic derive.
Structs§
- AstMetrics
- Node count, deepest depth, recovery state, and completeness from one traversal.
- Custom
Lang - A tree-sitter grammar registered from outside
ast-grep-language. - Node
- ’r represents root lifetime
- StrDoc
- Syntax
Diagnostic - A bounded recovery diagnostic with byte offsets into the original source.
- TSLanguage
- An opaque object that defines how to parse a particular language. The code
for each
Languageis generated by the Tree-sitter CLI.
Enums§
- Content
Detection - Complete content-detection result after every distinct candidate was checked.
- Inspection
Stop Reason - Why a bounded AST inspection stopped before visiting the full tree.
- Language
- Every language this build can select a grammar for.
- Parse
Error - A checked-parse refusal. None of these may be reported as clean.
- Support
Lang - Represents all built-in languages.
- Syntax
Issue Kind - Kind of tree-sitter recovery node reported by a checked parse.
Constants§
- MAX_
AST_ NODES - Largest concrete syntax tree admitted; bounds downstream walks, which a byte bound alone does not price. Measured on the tree after tree-sitter has built it, so it does not cap the parser’s own allocation.
- MAX_
DETECT_ BYTES - Largest probe source admitted to
try_detect_content(64 KiB). Content detection costs one full parse per distinct candidate grammar, so it is bounded far belowMAX_SOURCE_BYTES; a language probe only needs enough bytes to show one clean reading. - MAX_
SHEBANG_ BYTES - Longest
#!lineLanguage::of_shebangwill read (256 bytes). - MAX_
SOURCE_ BYTES - Largest source admitted to the checked parser (2 MiB). This is the input bound that keeps parse work linear in bytes.
- MAX_
SYNTAX_ DIAGNOSTICS - Maximum recovery diagnostics retained from one checked parse.
Functions§
- callee_
name - The callee a call node names, or
Nonewhennodeis not a call kind or names none ofname_kinds. Seedefinition_namefor why this trio is retained. - child_
text_ with_ kind - The owned text of the first direct child whose
kindequals one ofkinds, orNonewhen no direct child matches. Descendants are not searched, so a caller hunting a nested identifier must walk to that level first. The text is owned because the node’s source borrow is tied to the parsed tree. - definition_
name - The name a definition node declares, if a direct child kind in
name_kindsnames it. - detect
- Identify a language by file extension.
Nonemeans the name does not claim the file, so report unscanned, never clean. - has_
syntax_ issues - Whether the tree holds an
ERRORorMISSINGnode. Ask before reporting: on unreadable source, no finding means nothing parsed, not nothing wrong. - inspect_
ast - Node count, depth, recovery state, and traversal completeness in one walk.
- max_
depth - Deepest branch depth, root counting as 1.
- parse
- Unchecked parse for diagnostics and tests that intentionally inspect
malformed trees. Production call sites use
try_parse. - parse_
with - Unchecked parse of a caller-registered grammar. Production call sites use
try_parse_with. - tree_
diagnostics - Every recovery node in
tree, as a located diagnostic. - tree_
recovery_ count - How many recovery nodes
treecarries. - try_
detect_ content - Identify
sourceby trial-parsingcandidates, the opt-in content path. - try_
parse - Production boundary: refuse oversized bytes, then reject recovery nodes and over-budget trees in one traversal.
- try_
parse_ with - Checked parse of a caller-registered grammar, held to the same byte bound,
node bound, and recovery refusal as
try_parse.