feat: gasm-devkit 0.1.0 — GAsm lexer, parser, linter, formatter, LSP and amd64 assembler
Assisted-by: Qwen 3.8 Max Preview
This commit is contained in:
@@ -0,0 +1,212 @@
|
||||
# Architecture
|
||||
|
||||
How gasm-devkit is put together and why.
|
||||
|
||||
## Design goals
|
||||
|
||||
1. **A real AST, not a grammar hack.** The linter, analyser, assembler and
|
||||
language server all need to *reason* about assembly — not just colour it.
|
||||
So the centre of the toolkit is a hand-written lexer and a parser that
|
||||
produce a typed AST with source positions on every node.
|
||||
2. **Architecture as data, not code.** Per-architecture differences (amd64,
|
||||
arm64, riscv64, loong64) live in register and instruction *tables* (`arch`),
|
||||
never in `if arch == …` branches scattered through the logic. The
|
||||
instruction tables are generated from the Go toolchain's own assembler
|
||||
source (`just gen`), so adding or refreshing an architecture is a data
|
||||
operation, not a coding one.
|
||||
3. **Open integration surface.** Everything the toolkit can do is reachable
|
||||
through two vendor-neutral interfaces: a CLI and an LSP server. No editor
|
||||
owns the toolkit; the toolkit is offered to editors on standard terms.
|
||||
|
||||
## Pipeline
|
||||
|
||||
```mermaid
|
||||
graph TD
|
||||
SRC["source .s"] --> LEX["lexer<br/>token stream"]
|
||||
LEX --> PAR["parser<br/>AST + diagnostics"]
|
||||
LEX --> FMT["format<br/>re-space tokens"]
|
||||
PAR --> LINT["lint<br/>static checks"]
|
||||
PAR --> LSP["lsp server"]
|
||||
LEX --> LSP
|
||||
ARCH["arch tables<br/>amd64 / arm64 / riscv64 / loong64"] --> LINT
|
||||
ARCH --> LSP
|
||||
LINT --> LSP
|
||||
FMT --> CLI["gasm CLI"]
|
||||
LINT --> CLI
|
||||
PAR --> CLI
|
||||
LEX --> CLI
|
||||
LSP --> EDITOR["any LSP editor"]
|
||||
```
|
||||
|
||||
The lexer is the shared foundation: the parser builds the AST from it, the
|
||||
formatter re-spaces its tokens directly, and the language server uses it for
|
||||
semantic highlighting.
|
||||
|
||||
## Components
|
||||
|
||||
### `token` and `lexer`
|
||||
|
||||
The scanner is hand-written and permissive: it never panics and maps anything
|
||||
it cannot classify to an `Illegal` token, so every downstream tool still works
|
||||
on malformed input. Newlines are significant tokens, because Plan 9 assembly
|
||||
is line-oriented and the parser relies on line structure.
|
||||
|
||||
The middle dot (`·`, U+00B7) is treated as an identifier character so that
|
||||
`·funcName(SB)` lexes as one symbol. Multi-character operators (`<<`, `>>`,
|
||||
`->`) are recognised so arm64 shift operands scan correctly. A backslash
|
||||
immediately before a newline is a C-preprocessor line continuation (used by
|
||||
`#define` macros in the runtime `.s` files); the lexer splices the lines
|
||||
together so a multi-line macro becomes one logical line the parser treats as an
|
||||
opaque preprocessor directive.
|
||||
|
||||
### `ast` and `parser`
|
||||
|
||||
The parser is **line-oriented**, matching how the Plan 9 assembler reads a
|
||||
file: it groups tokens into lines, classifies each line (directive, label,
|
||||
instruction, comment, preprocessor) and dispatches. A malformed line is
|
||||
reported and skipped; it never aborts the file.
|
||||
|
||||
Operands are parsed into a faithful, flat representation. The amd64
|
||||
addressing modes — `reg`, `$imm`, `(base)`, `off(base)`, `(base)(index*scale)`,
|
||||
`name+off(FP)`, `name<>(SB)` — are all captured structurally, and the original
|
||||
token text is retained for fidelity.
|
||||
|
||||
A deliberate boundary: the AST records **syntax only**. Whether a bare
|
||||
identifier is a register or a label is an *architecture* question, so it is
|
||||
left to `arch` and resolved in the lint/lsp layers. This keeps the parser
|
||||
arch-agnostic and its output deterministic.
|
||||
|
||||
### `arch`
|
||||
|
||||
Register files are generated programmatically (the regular `R8`–`R15`,
|
||||
`X0`–`X15`, `Y0`–`Y15`, `Z0`–`Z31`, `K0`–`K7` ranges) plus the irregularly
|
||||
named registers listed explicitly. Instruction names are **generated from the
|
||||
Go toolchain's own assembler source** (`cmd/internal/obj/<arch>/anames.go`,
|
||||
plus the common opcodes and the per-architecture front-end aliases such as the
|
||||
arm64 `B`/`BL` branches and the `.P`/`.W` load-store addressing suffixes) by
|
||||
`just gen`, so the tables always match what the real assembler accepts. Each
|
||||
mnemonic maps to a summary and an optional operand-count range; counts are
|
||||
recorded only where unambiguous (`-1` disables the operand-count lint for that
|
||||
instruction) so the linter stays silent rather than guess. For architectures
|
||||
with highly variable operand forms (arm64, riscv64, loong64) only a few
|
||||
fixed-arity instructions (`RET`, `NOP`, `JMP`, `CALL`) carry counts at all.
|
||||
|
||||
### `lint`
|
||||
|
||||
Rules are conservative by design — silence beats a false positive. The rules
|
||||
are `unknown-instruction`, `operand-count`, `undefined-label`,
|
||||
`duplicate-label`, `missing-ret`, `missing-textflag-include`, `abi-argsize` and
|
||||
`unreachable-code`. Every diagnostic carries a stable code so callers can
|
||||
disable rules individually, and arch-specific rules switch off entirely when
|
||||
the target architecture cannot be inferred from the file name.
|
||||
|
||||
Two things keep the rules honest on real-world code:
|
||||
|
||||
- **Pseudo-ops and macros are not instructions.** `unknown-instruction` knows
|
||||
the assembler pseudo-ops (`BYTE`, `WORD`, `FUNCDATA`, `PCDATA`, …) and
|
||||
recognises macro invocations — an in-file `#define` name, or any identifier
|
||||
containing an underscore (no Plan 9 mnemonic ever does).
|
||||
- **Macro-heavy files get the label/RET heuristics turned off.** Without a
|
||||
preprocessor, labels a macro defines are invisible, so `undefined-label` and
|
||||
`missing-ret` are suppressed for files that use macros (an in-file `#define`
|
||||
or a `#include` of anything other than `textflag.h`). `missing-ret` also
|
||||
treats a trailing unconditional jump and `UNDEF` as valid terminators.
|
||||
|
||||
The result is validated by `TestGoRuntimeCorpus`, which parses and lints every
|
||||
`src/runtime/*.s` file the toolchain ships for all four architectures and
|
||||
asserts zero parse errors and zero error-severity diagnostics.
|
||||
|
||||
Two deeper analyses sit on top of the AST:
|
||||
|
||||
- **`abi-argsize`.** Hand-written kernels document their signature in a
|
||||
`// func …` comment above the `TEXT`. The linter parses that signature with
|
||||
the standard library's Go parser, lays out the parameters and results under
|
||||
Go's ABI0 stack rules (results begin on a word boundary after the
|
||||
parameters), and checks the total against the argument size declared in the
|
||||
`TEXT` directive. It only runs for stack-argument functions (a non-zero
|
||||
declared arg area that is actually addressed through `FP`), and aborts
|
||||
silently on a type whose size it cannot determine — so it never guesses.
|
||||
- **`unreachable-code`.** Code after a `RET` and before the next label is
|
||||
dead. The check is suppressed for any function whose reachability cannot be
|
||||
decided statically: those using PC-relative jumps (`JMP 2(PC)`),
|
||||
register-indirect branches (`JALR`/`JR`/`JIRL`/`BR`/`BLR`), or living in a
|
||||
file with `#ifdef` conditionals. `UNDEF` is deliberately not a terminator —
|
||||
code after it is occasionally intentional metadata.
|
||||
- **`register-clobber` (register liveness).** The linter builds the function's
|
||||
control-flow graph (basic blocks split at labels and after branches, with
|
||||
fall-through and jump-target edges), computes a conservative per-instruction
|
||||
register def/use, and runs the standard backward liveness iteration to a fixed
|
||||
point. On top of that it flags a **callee-saved register that is written but
|
||||
never saved and restored** — the per-architecture callee-saved set is amd64
|
||||
`BX/BP/R12–R15`, arm64 `R19–R30`, riscv64 `X1/X8/X9/X18–X27`, loong64
|
||||
`R1/R22–R31`. This is an *audit*: the runtime's own assembly clobbers these
|
||||
registers freely (it controls both sides of the call), so the rule is
|
||||
advisory there, but in hand-written kernels called from ordinary Go code a
|
||||
clobber is a genuine ABI violation. It runs only on macro-free files, where
|
||||
no opaque macro can perform the save/restore.
|
||||
- **`funcdata-pcdata`.** `FUNCDATA $idx, sym(SB)` and `PCDATA $idx, $val` are
|
||||
checked for well-formed operands (arity, immediate index and value, symbol
|
||||
reference) and a literal index is range-checked; a named index constant such
|
||||
as `$PCDATA_StackMapIndex` is accepted without a range check.
|
||||
|
||||
### `format`
|
||||
|
||||
The formatter works on the **token stream, not the AST**, so it preserves
|
||||
every line — comments and blanks included. It only normalises indentation,
|
||||
operand spacing and per-function mnemonic alignment. It is idempotent and its
|
||||
output always round-trips through the parser.
|
||||
|
||||
### `lsp`
|
||||
|
||||
The server speaks JSON-RPC 2.0 with `Content-Length` framing over any
|
||||
`io.Reader`/`io.Writer` (normally stdin/stdout). It maintains an in-memory
|
||||
document store, republishes diagnostics on every change, and provides:
|
||||
|
||||
- **completion** — instructions, registers, pseudo-registers, textflag macros
|
||||
and local labels;
|
||||
- **hover** — instruction summaries and register descriptions from `arch`;
|
||||
- **document symbols** — `TEXT` functions with their labels, plus `GLOBL`/`DATA`;
|
||||
- **semantic tokens** — syntax highlighting delivered as LSP semantic tokens,
|
||||
classified with the lexer plus `arch` (instructions, registers by class,
|
||||
pseudo-registers, labels, immediates, comments, directives, textflag macros).
|
||||
|
||||
Semantic tokens are the key to editor-agnostic highlighting: the editor renders
|
||||
them from the standard LSP legend, so no editor-specific grammar is needed.
|
||||
|
||||
### `asm`
|
||||
|
||||
The standalone assembler (Phase 2). Its core is an amd64 instruction encoder:
|
||||
a REX/ModR-M/SIB/displacement/immediate engine plus the scalar instruction set,
|
||||
with the Plan 9 operand order (source first) mapped onto the x86 encoding.
|
||||
Every encoding is validated by decoding it again with `golang.org/x/arch` — the
|
||||
one module dependency, used in tests only and never linked into the binary.
|
||||
|
||||
On top of the encoder, `Assemble` walks a parsed `TEXT` body, converts each
|
||||
operand to an encoder operand, and lays the instructions out in two passes so
|
||||
local labels resolve to fixed rel32 jump offsets. The `FP`/`SP` pseudo-
|
||||
registers are translated onto the hardware stack pointer — `x+N(FP)` becomes
|
||||
`(N+8)(SP)` for a zero-frame function and `(N+frame+16)(SP)` once a frame
|
||||
pointer is set up, with the matching Go prologue/epilogue generated — so the
|
||||
output is byte-identical to the Go assembler for these cases. SIMD is handled
|
||||
SIMD is handled
|
||||
by a VEX (AVX/AVX2) encoder — the two- and three-byte VEX prefixes with XMM/YMM
|
||||
registers — across three operand forms (the three-operand NDS form, the
|
||||
two-operand reg/rm form, and the immediate-shift form), together covering the
|
||||
bulk of the integer SIMD set; each encoding is validated by round-trip
|
||||
decoding. This increment covers register / memory / immediate / FP-frame
|
||||
operands, local-label jumps and these VEX SIMD forms; the remaining SIMD forms
|
||||
(shuffles, extract/insert, permute, moves), EVEX / AVX-512, `SB` (global
|
||||
symbol) operands (relocations) and object-file emission are the rest of
|
||||
Phase 2.
|
||||
|
||||
## Extension points
|
||||
|
||||
- **New architecture:** add an entry to the generator in `_gen`, run
|
||||
`just gen`, and add a `buildXXX()` register file plus a case in `ForArch`.
|
||||
- **New lint rule:** add a function in `lint` and a rule-code constant.
|
||||
- **New LSP feature:** add a method case in `dispatch` and a handler.
|
||||
|
||||
The phases follow a dependency chain. Phase 1 (static analysis) builds only on
|
||||
the AST; Phase 2 (the standalone assembler) emits object code; Phases 3
|
||||
(dynamic analysis) and 4 (the debugger) both consume the execution substrate
|
||||
that the assembler provides.
|
||||
+79
@@ -0,0 +1,79 @@
|
||||
# Using gasm-devkit with Zed
|
||||
|
||||
This document is deliberately blunt, because the situation is a genuine
|
||||
conflict between two of the project's own commitments, and papering over it
|
||||
would be dishonest.
|
||||
|
||||
## The conflict
|
||||
|
||||
gasm-devkit is **pure Go, no C, no cgo, no JavaScript runtimes, no native
|
||||
binaries, no vendor lock-in, no platform-specific IDE internals.**
|
||||
|
||||
Zed's extension model, as verified against Zed's own documentation, is:
|
||||
|
||||
- Extensions are written in **Rust** and compiled to **WebAssembly**
|
||||
(`wasm32-wasip2`).
|
||||
- Syntax highlighting is provided by **Tree-sitter** grammars, which are
|
||||
**C** compiled to WebAssembly with the wasi-sdk, from a grammar written in a
|
||||
**JavaScript** DSL.
|
||||
- A *new* language cannot be registered through configuration alone. Defining
|
||||
a language requires an extension, and every language extension must name a
|
||||
Tree-sitter grammar. (Zed's `lsp` settings section configures
|
||||
already-registered servers; it does not register an arbitrary external binary
|
||||
for a brand-new language.)
|
||||
|
||||
There is therefore **no pure-Go path into Zed's extension host.** This is a
|
||||
property of Zed, not of gasm-devkit: no language tooling author can feed Zed a
|
||||
pure-Go highlighting grammar, because Zed's highlighting engine is Tree-sitter
|
||||
and its plugin runtime is Rust/WASM.
|
||||
|
||||
## What gasm-devkit gives Zed regardless
|
||||
|
||||
The toolkit's integration surface is the **Language Server Protocol**, an open
|
||||
standard. Through `gasm lsp` it provides, with zero editor-specific code:
|
||||
|
||||
- autocomplete (instructions, registers, pseudo-registers, labels),
|
||||
- hover documentation,
|
||||
- diagnostics (the linter, pushed as you type),
|
||||
- document outline (functions and labels),
|
||||
- **syntax highlighting, delivered as LSP semantic tokens.**
|
||||
|
||||
That last point matters: Zed can render highlighting entirely from LSP semantic
|
||||
tokens (`"semantic_tokens": "full"` replaces Tree-sitter highlighting for a
|
||||
language). So the highlighting *capability* exists in pure Go; what Zed needs
|
||||
is merely to be told that `.s` files are a language served by `gasm lsp`.
|
||||
|
||||
## The honest options
|
||||
|
||||
1. **Use an editor that registers an external LSP by configuration.**
|
||||
Neovim, Helix, VS Code and Sublime all let you associate `.s` with the
|
||||
`gasm lsp` binary and use its semantic tokens — no Rust, no C, no lock-in.
|
||||
This is the option that satisfies every stated constraint with no
|
||||
exception.
|
||||
|
||||
2. **Treat a Zed adapter as one quarantined exception.** A minimal Zed
|
||||
extension — a few lines of Rust that register the language and launch
|
||||
`gasm lsp` — plus either a Tree-sitter grammar or `"full"` semantic tokens
|
||||
for highlighting. Crucially, this adapter is the *editor's plugin format*;
|
||||
it is sandboxed inside Zed and never linked into, compiled into, or shipped
|
||||
with the Go toolkit. gasm-devkit itself stays pure Go. But producing it
|
||||
uses the Rust/wasi-sdk/Tree-sitter toolchain, which the project constraints
|
||||
forbid — so it must be a conscious, explicit decision, not a silent one.
|
||||
|
||||
The author's philosophy — digital sovereignty, no dependency on toolchains he
|
||||
does not control — is the tie-breaker, and it is a value judgement rather than
|
||||
a technical one. gasm-devkit is built so that **either** choice keeps the
|
||||
toolkit itself clean: the pure-Go core and the LSP are the product; a Zed
|
||||
adapter, if ever wanted, is a thin, separable leaf.
|
||||
|
||||
## Wiring the LSP (editor-agnostic)
|
||||
|
||||
Run the server and point an LSP client at it:
|
||||
|
||||
```sh
|
||||
go run ./cmd/gasm lsp # or: go install ./cmd/gasm && gasm lsp
|
||||
```
|
||||
|
||||
Associate the command with `*.s` (and `*_amd64.s` / `*_arm64.s`) in whichever
|
||||
editor you use. The server infers the target architecture from the file-name
|
||||
suffix and selects the amd64 or arm64 instruction tables accordingly.
|
||||
Reference in New Issue
Block a user