README.md

README.md is a file in Code Parse. 97 lines of code and 0 definitions.

<!-- Auto-generated 2026-09-30T01:52Z v8 -->

# @govlab/code-parse

<!-- concern:overview -->

## Purpose

A dependency-light leaf that parses source code of any grammar-supported language into a concrete syntax tree using tree-sitter (`web-tree-sitter`, WASM). It ships pre-built grammar `.wasm` files, resolves the one for a file by extension or shebang, and exposes both an async (`parseCode`) and a preload-then-synchronous (`ensureLanguages` + `parseCodeSync`) parse path so callers bound to a synchronous contract can still parse. Generic AST helpers (`walk`, `commentNodes`) let consumers traverse the tree and locate node kinds deterministically. It is the shared parsing substrate for `@govlab/patterns`' code analysis and `@govlab/quality`'s comment cleaner.
<!-- /concern:overview -->

<!-- concern:use -->

## When to use

- Any tool that needs a deterministic syntax tree for source code without hand-rolling a per-language lexer, such as comment stripping, symbol extraction or structural analysis.
- Locating comment nodes (or any node type) with exact byte ranges for safe removal or rewriting, correct across string literals, heredocs, and regex literals by construction.
- Detecting a file's language from its extension or shebang, validated against the grammars actually built.

## When NOT to use

- A single-language, syntax-trivial scan where a small string check is enough and a WASM parser is overkill.
- Languages with no built tree-sitter grammar. `parseCode` returns null instead of guessing, so the grammar source goes into `configuration/configs/grammar.config.ts` first.
- Semantic analysis needing types or cross-file resolution. The package is syntax-only (a concrete syntax tree), not a type checker.

<!-- /concern:use -->

<!-- concern:charts -->

## Architecture charts

The structure, logical-flow and dependency diagrams derived from the source AST live in [_code.info.generated/mermaid-charts.generated.md](./_code.info.generated/mermaid-charts.generated.md).
<!-- /concern:charts -->

<!-- concern:install -->

## Install

Part of the Govlab monorepo, with no standalone install. From the repo root, `npm install` provisions `web-tree-sitter`. The grammar build writes the `.wasm` files under `core/generated/`, the `govlab.utils.codeParse.generated` path key. They are committed and shipped in `files`. Rebuild or extend the grammar set with `npm run build:grammars --workspace @govlab/code-parse`.

## Quick start

```js EXAMPLE: Detect language, parse, and collect comment nodes
import { detectLanguage, parseCode, commentNodes } from "@govlab/code-parse";

const lang = detectLanguage("main.go");
const root = lang ? await parseCode(source, lang) : null;
for (const node of root ? commentNodes(root) : []) {
    console.log(node.type, node.startIndex, node.endIndex, node.text);
}
```

```js EXAMPLE: Preload grammars once, then parse synchronously inside a sync contract
import { ensureLanguages, parseCodeSync, commentNodes } from "@govlab/code-parse";

await ensureLanguages(["go", "typescript", "javascript"]);
function strip(source, lang) {
    const root = parseCodeSync(source, lang);
    return root ? commentNodes(root) : [];
}
```

<!-- /concern:install -->

<!-- concern:api -->

## API

- `function availableLanguages(): string[]` — the grammar names actually built under core/generated/, sorted.
- `interface CodeParseOptions`
- `function commentNodes(root: CstNode): CstNode[]` — walks a CstNode tree and returns every node whose type is a comment kind, with byte ranges for removal.
- `interface CstNode`
- `const DETECTABLE_LANGUAGES: ReadonlySet<string>`
- `function detectLanguage(filename: string, content?: string): string | null` — maps a filename extension (or shebang for extensionless scripts) to a tree-sitter grammar name, or null.
- `function ensureLanguages(languages: Iterable<string>): Promise<void>` — async. It initializes the runtime and loads the given languages' grammars once (memoized), so later parseCodeSync calls are synchronous.
- `function isCommentType(type: string): boolean`
- `function parseCode(source: string, language: string, options?: CodeParseOptions): Promise<CstNode | null>` — the async parse. It loads the grammar on demand and returns a CstNode root, or null when the language has no grammar or parsing fails.
- `function parseCodeSync(source: string, language: string, options?: CodeParseOptions): CstNode | null` — the synchronous parse, which requires the language preloaded through ensureLanguages and throws otherwise. It is the seam that keeps synchronous callers synchronous.
- `interface ParseLogger`
- `function resolveGrammarDir(): string`
- `function walk(root: CstNode, visit: (node: CstNode) => void): void` — depth-first CstNode traversal invoking a visitor per node.

<!-- /concern:api -->

<!-- concern:config -->

## Configuration

`parseCode` and `parseCodeSync` accept an options object with an optional `logger` whose `warn(message, detail?)` method receives parse diagnostics. Grammar filenames follow `<lang>.generated.wasm`, and the grammar folder resolves through the `govlab.utils.codeParse.generated` path key. `parseCodeSync` requires its language to have been preloaded through `ensureLanguages` and throws otherwise, while `parseCode` preloads on demand.
<!-- /concern:config -->

<!-- concern:deps -->

## Dependencies

- `@govlab/argv`
- `@govlab/canonical-write`
- `@ssot/paths`

<!-- /concern:deps -->

<!-- concern:ai-context -->

## AI context

- tree-sitter parses to a concrete syntax tree, so comment nodes are located precisely in any string, heredoc or regex context, with no guessing at string delimiters.
- The sync seam: `Parser.init` and `Language.load` are async and run once through `ensureLanguages`. The per-source `parseCodeSync` is then synchronous, so a synchronous caller, such as a native lint rule's `check` or `fix`, parses without going async.
- `CstNode` carries `startIndex` and `endIndex` (source offsets) for range-based edits, and `startPosition.row` for line reporting. It is an adapted, plain-object snapshot of the tree-sitter node, safe to hold after the tree is freed.
- Language coverage equals the grammars in `core/generated/` (`availableLanguages()`), and `detectLanguage` maps extensions and shebangs to those grammar names. Add a language by adding a source row to `configuration/configs/grammar.config.ts` and running `npm run build:grammars -w @govlab/code-parse`, which writes the extension map through `@govlab/canonical-write`.
- Its third-party runtime dependency is `web-tree-sitter`. It reads its locations through `@ssot/paths`, and the grammar build declares its command line through `@govlab/argv`.

<!-- /concern:ai-context -->

<!-- concern:domains -->

## Domains

This package serves these software domains, which `_manifest.json` declares in `domains` from the two-tier software-domain vocabulary (`meta → sub`):

- **developer-tooling** — linting-quality
- **platform** — utilities

<!-- /concern:domains -->

<!-- concern:quality-governance -->

## Quality governance

The canonical quality catalog resolves the quality concepts that govern this package. `_manifest.json` declares them in `governedBy`, and a lint package derives them from the concepts its own rules enforce. Each maps to the custom lint rules that enforce it:

- **separation-of-concerns** — _complexity_
- **type-safety** — _correctness_

<!-- /concern:quality-governance -->

<!-- concern:disposal -->

## Disposal

- Remove `govlab.root/govlab.utils/code-parse/`.
- Drop `"@govlab/code-parse": "*"` from `govlab.root/govlab.quality/package.json` and `govlab.root/govlab.patterns/package.json`.
- Remove every `import … from "@govlab/code-parse"` in consumers (`@govlab/patterns`' code ingestion, `@govlab/quality`'s comment cleaner) and restore their own parsing, or drop the feature.

<!-- /concern:disposal -->

<!-- concern:metrics -->

---

stable · 13 exports · 3 deps · 0 principles · 2 concepts
<!-- /concern:metrics -->