README.md
README.md is a file in Code Parse. 97 lines of code and 0 definitions.
<!-- Auto-generated 2026-09-30T01:52Z v8 -->
# @govlab/code-parse
<!-- concern:overview -->
## Purpose
A dependency-light leaf that parses source code of any grammar-supported language into a concrete syntax tree using tree-sitter (`web-tree-sitter`, WASM). It ships pre-built grammar `.wasm` files, resolves the one for a file by extension or shebang, and exposes both an async (`parseCode`) and a preload-then-synchronous (`ensureLanguages` + `parseCodeSync`) parse path so callers bound to a synchronous contract can still parse. Generic AST helpers (`walk`, `commentNodes`) let consumers traverse the tree and locate node kinds deterministically. It is the shared parsing substrate for `@govlab/patterns`' code analysis and `@govlab/quality`'s comment cleaner.
<!-- /concern:overview -->
<!-- concern:use -->
## When to use
- Any tool that needs a deterministic syntax tree for source code without hand-rolling a per-language lexer, such as comment stripping, symbol extraction or structural analysis.
- Locating comment nodes (or any node type) with exact byte ranges for safe removal or rewriting, correct across string literals, heredocs, and regex literals by construction.
- Detecting a file's language from its extension or shebang, validated against the grammars actually built.
## When NOT to use
- A single-language, syntax-trivial scan where a small string check is enough and a WASM parser is overkill.
- Languages with no built tree-sitter grammar. `parseCode` returns null instead of guessing, so the grammar source goes into `configuration/configs/grammar.config.ts` first.
- Semantic analysis needing types or cross-file resolution. The package is syntax-only (a concrete syntax tree), not a type checker.
<!-- /concern:use -->
<!-- concern:charts -->
## Architecture charts
The structure, logical-flow and dependency diagrams derived from the source AST live in [_code.info.generated/mermaid-charts.generated.md](./_code.info.generated/mermaid-charts.generated.md).
<!-- /concern:charts -->
<!-- concern:install -->
## Install
Part of the Govlab monorepo, with no standalone install. From the repo root, `npm install` provisions `web-tree-sitter`. The grammar build writes the `.wasm` files under `core/generated/`, the `govlab.utils.codeParse.generated` path key. They are committed and shipped in `files`. Rebuild or extend the grammar set with `npm run build:grammars --workspace @govlab/code-parse`.
## Quick start
```js EXAMPLE: Detect language, parse, and collect comment nodes
import { detectLanguage, parseCode, commentNodes } from "@govlab/code-parse";
const lang = detectLanguage("main.go");
const root = lang ? await parseCode(source, lang) : null;
for (const node of root ? commentNodes(root) : []) {
console.log(node.type, node.startIndex, node.endIndex, node.text);
}
```
```js EXAMPLE: Preload grammars once, then parse synchronously inside a sync contract
import { ensureLanguages, parseCodeSync, commentNodes } from "@govlab/code-parse";
await ensureLanguages(["go", "typescript", "javascript"]);
function strip(source, lang) {
const root = parseCodeSync(source, lang);
return root ? commentNodes(root) : [];
}
```
<!-- /concern:install -->
<!-- concern:api -->
## API
- `function availableLanguages(): string[]` — the grammar names actually built under core/generated/, sorted.
- `interface CodeParseOptions`
- `function commentNodes(root: CstNode): CstNode[]` — walks a CstNode tree and returns every node whose type is a comment kind, with byte ranges for removal.
- `interface CstNode`
- `const DETECTABLE_LANGUAGES: ReadonlySet<string>`
- `function detectLanguage(filename: string, content?: string): string | null` — maps a filename extension (or shebang for extensionless scripts) to a tree-sitter grammar name, or null.
- `function ensureLanguages(languages: Iterable<string>): Promise<void>` — async. It initializes the runtime and loads the given languages' grammars once (memoized), so later parseCodeSync calls are synchronous.
- `function isCommentType(type: string): boolean`
- `function parseCode(source: string, language: string, options?: CodeParseOptions): Promise<CstNode | null>` — the async parse. It loads the grammar on demand and returns a CstNode root, or null when the language has no grammar or parsing fails.
- `function parseCodeSync(source: string, language: string, options?: CodeParseOptions): CstNode | null` — the synchronous parse, which requires the language preloaded through ensureLanguages and throws otherwise. It is the seam that keeps synchronous callers synchronous.
- `interface ParseLogger`
- `function resolveGrammarDir(): string`
- `function walk(root: CstNode, visit: (node: CstNode) => void): void` — depth-first CstNode traversal invoking a visitor per node.
<!-- /concern:api -->
<!-- concern:config -->
## Configuration
`parseCode` and `parseCodeSync` accept an options object with an optional `logger` whose `warn(message, detail?)` method receives parse diagnostics. Grammar filenames follow `<lang>.generated.wasm`, and the grammar folder resolves through the `govlab.utils.codeParse.generated` path key. `parseCodeSync` requires its language to have been preloaded through `ensureLanguages` and throws otherwise, while `parseCode` preloads on demand.
<!-- /concern:config -->
<!-- concern:deps -->
## Dependencies
- `@govlab/argv`
- `@govlab/canonical-write`
- `@ssot/paths`
<!-- /concern:deps -->
<!-- concern:ai-context -->
## AI context
- tree-sitter parses to a concrete syntax tree, so comment nodes are located precisely in any string, heredoc or regex context, with no guessing at string delimiters.
- The sync seam: `Parser.init` and `Language.load` are async and run once through `ensureLanguages`. The per-source `parseCodeSync` is then synchronous, so a synchronous caller, such as a native lint rule's `check` or `fix`, parses without going async.
- `CstNode` carries `startIndex` and `endIndex` (source offsets) for range-based edits, and `startPosition.row` for line reporting. It is an adapted, plain-object snapshot of the tree-sitter node, safe to hold after the tree is freed.
- Language coverage equals the grammars in `core/generated/` (`availableLanguages()`), and `detectLanguage` maps extensions and shebangs to those grammar names. Add a language by adding a source row to `configuration/configs/grammar.config.ts` and running `npm run build:grammars -w @govlab/code-parse`, which writes the extension map through `@govlab/canonical-write`.
- Its third-party runtime dependency is `web-tree-sitter`. It reads its locations through `@ssot/paths`, and the grammar build declares its command line through `@govlab/argv`.
<!-- /concern:ai-context -->
<!-- concern:domains -->
## Domains
This package serves these software domains, which `_manifest.json` declares in `domains` from the two-tier software-domain vocabulary (`meta → sub`):
- **developer-tooling** — linting-quality
- **platform** — utilities
<!-- /concern:domains -->
<!-- concern:quality-governance -->
## Quality governance
The canonical quality catalog resolves the quality concepts that govern this package. `_manifest.json` declares them in `governedBy`, and a lint package derives them from the concepts its own rules enforce. Each maps to the custom lint rules that enforce it:
- **separation-of-concerns** — _complexity_
- **type-safety** — _correctness_
<!-- /concern:quality-governance -->
<!-- concern:disposal -->
## Disposal
- Remove `govlab.root/govlab.utils/code-parse/`.
- Drop `"@govlab/code-parse": "*"` from `govlab.root/govlab.quality/package.json` and `govlab.root/govlab.patterns/package.json`.
- Remove every `import … from "@govlab/code-parse"` in consumers (`@govlab/patterns`' code ingestion, `@govlab/quality`'s comment cleaner) and restore their own parsing, or drop the feature.
<!-- /concern:disposal -->
<!-- concern:metrics -->
---
stable · 13 exports · 3 deps · 0 principles · 2 concepts
<!-- /concern:metrics -->