Files
gasm-sdk/docs/ARCHITECTURE.md
T

213 lines
11 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Architecture
How gasm-devkit is put together and why.
## Design goals
1. **A real AST, not a grammar hack.** The linter, analyser, assembler and
language server all need to *reason* about assembly — not just colour it.
So the centre of the toolkit is a hand-written lexer and a parser that
produce a typed AST with source positions on every node.
2. **Architecture as data, not code.** Per-architecture differences (amd64,
arm64, riscv64, loong64) live in register and instruction *tables* (`arch`),
never in `if arch == …` branches scattered through the logic. The
instruction tables are generated from the Go toolchain's own assembler
source (`just gen`), so adding or refreshing an architecture is a data
operation, not a coding one.
3. **Open integration surface.** Everything the toolkit can do is reachable
through two vendor-neutral interfaces: a CLI and an LSP server. No editor
owns the toolkit; the toolkit is offered to editors on standard terms.
## Pipeline
```mermaid
graph TD
SRC["source .s"] --> LEX["lexer<br/>token stream"]
LEX --> PAR["parser<br/>AST + diagnostics"]
LEX --> FMT["format<br/>re-space tokens"]
PAR --> LINT["lint<br/>static checks"]
PAR --> LSP["lsp server"]
LEX --> LSP
ARCH["arch tables<br/>amd64 / arm64 / riscv64 / loong64"] --> LINT
ARCH --> LSP
LINT --> LSP
FMT --> CLI["gasm CLI"]
LINT --> CLI
PAR --> CLI
LEX --> CLI
LSP --> EDITOR["any LSP editor"]
```
The lexer is the shared foundation: the parser builds the AST from it, the
formatter re-spaces its tokens directly, and the language server uses it for
semantic highlighting.
## Components
### `token` and `lexer`
The scanner is hand-written and permissive: it never panics and maps anything
it cannot classify to an `Illegal` token, so every downstream tool still works
on malformed input. Newlines are significant tokens, because Plan 9 assembly
is line-oriented and the parser relies on line structure.
The middle dot (`·`, U+00B7) is treated as an identifier character so that
`·funcName(SB)` lexes as one symbol. Multi-character operators (`<<`, `>>`,
`->`) are recognised so arm64 shift operands scan correctly. A backslash
immediately before a newline is a C-preprocessor line continuation (used by
`#define` macros in the runtime `.s` files); the lexer splices the lines
together so a multi-line macro becomes one logical line the parser treats as an
opaque preprocessor directive.
### `ast` and `parser`
The parser is **line-oriented**, matching how the Plan 9 assembler reads a
file: it groups tokens into lines, classifies each line (directive, label,
instruction, comment, preprocessor) and dispatches. A malformed line is
reported and skipped; it never aborts the file.
Operands are parsed into a faithful, flat representation. The amd64
addressing modes — `reg`, `$imm`, `(base)`, `off(base)`, `(base)(index*scale)`,
`name+off(FP)`, `name<>(SB)` — are all captured structurally, and the original
token text is retained for fidelity.
A deliberate boundary: the AST records **syntax only**. Whether a bare
identifier is a register or a label is an *architecture* question, so it is
left to `arch` and resolved in the lint/lsp layers. This keeps the parser
arch-agnostic and its output deterministic.
### `arch`
Register files are generated programmatically (the regular `R8`–`R15`,
`X0`–`X15`, `Y0`–`Y15`, `Z0`–`Z31`, `K0`–`K7` ranges) plus the irregularly
named registers listed explicitly. Instruction names are **generated from the
Go toolchain's own assembler source** (`cmd/internal/obj/<arch>/anames.go`,
plus the common opcodes and the per-architecture front-end aliases such as the
arm64 `B`/`BL` branches and the `.P`/`.W` load-store addressing suffixes) by
`just gen`, so the tables always match what the real assembler accepts. Each
mnemonic maps to a summary and an optional operand-count range; counts are
recorded only where unambiguous (`-1` disables the operand-count lint for that
instruction) so the linter stays silent rather than guess. For architectures
with highly variable operand forms (arm64, riscv64, loong64) only a few
fixed-arity instructions (`RET`, `NOP`, `JMP`, `CALL`) carry counts at all.
### `lint`
Rules are conservative by design — silence beats a false positive. The rules
are `unknown-instruction`, `operand-count`, `undefined-label`,
`duplicate-label`, `missing-ret`, `missing-textflag-include`, `abi-argsize` and
`unreachable-code`. Every diagnostic carries a stable code so callers can
disable rules individually, and arch-specific rules switch off entirely when
the target architecture cannot be inferred from the file name.
Two things keep the rules honest on real-world code:
- **Pseudo-ops and macros are not instructions.** `unknown-instruction` knows
the assembler pseudo-ops (`BYTE`, `WORD`, `FUNCDATA`, `PCDATA`, …) and
recognises macro invocations — an in-file `#define` name, or any identifier
containing an underscore (no Plan 9 mnemonic ever does).
- **Macro-heavy files get the label/RET heuristics turned off.** Without a
preprocessor, labels a macro defines are invisible, so `undefined-label` and
`missing-ret` are suppressed for files that use macros (an in-file `#define`
or a `#include` of anything other than `textflag.h`). `missing-ret` also
treats a trailing unconditional jump and `UNDEF` as valid terminators.
The result is validated by `TestGoRuntimeCorpus`, which parses and lints every
`src/runtime/*.s` file the toolchain ships for all four architectures and
asserts zero parse errors and zero error-severity diagnostics.
Two deeper analyses sit on top of the AST:
- **`abi-argsize`.** Hand-written kernels document their signature in a
`// func …` comment above the `TEXT`. The linter parses that signature with
the standard library's Go parser, lays out the parameters and results under
Go's ABI0 stack rules (results begin on a word boundary after the
parameters), and checks the total against the argument size declared in the
`TEXT` directive. It only runs for stack-argument functions (a non-zero
declared arg area that is actually addressed through `FP`), and aborts
silently on a type whose size it cannot determine — so it never guesses.
- **`unreachable-code`.** Code after a `RET` and before the next label is
dead. The check is suppressed for any function whose reachability cannot be
decided statically: those using PC-relative jumps (`JMP 2(PC)`),
register-indirect branches (`JALR`/`JR`/`JIRL`/`BR`/`BLR`), or living in a
file with `#ifdef` conditionals. `UNDEF` is deliberately not a terminator —
code after it is occasionally intentional metadata.
- **`register-clobber` (register liveness).** The linter builds the function's
control-flow graph (basic blocks split at labels and after branches, with
fall-through and jump-target edges), computes a conservative per-instruction
register def/use, and runs the standard backward liveness iteration to a fixed
point. On top of that it flags a **callee-saved register that is written but
never saved and restored** — the per-architecture callee-saved set is amd64
`BX/BP/R12–R15`, arm64 `R19–R30`, riscv64 `X1/X8/X9/X18–X27`, loong64
`R1/R22–R31`. This is an *audit*: the runtime's own assembly clobbers these
registers freely (it controls both sides of the call), so the rule is
advisory there, but in hand-written kernels called from ordinary Go code a
clobber is a genuine ABI violation. It runs only on macro-free files, where
no opaque macro can perform the save/restore.
- **`funcdata-pcdata`.** `FUNCDATA $idx, sym(SB)` and `PCDATA $idx, $val` are
checked for well-formed operands (arity, immediate index and value, symbol
reference) and a literal index is range-checked; a named index constant such
as `$PCDATA_StackMapIndex` is accepted without a range check.
### `format`
The formatter works on the **token stream, not the AST**, so it preserves
every line — comments and blanks included. It only normalises indentation,
operand spacing and per-function mnemonic alignment. It is idempotent and its
output always round-trips through the parser.
### `lsp`
The server speaks JSON-RPC 2.0 with `Content-Length` framing over any
`io.Reader`/`io.Writer` (normally stdin/stdout). It maintains an in-memory
document store, republishes diagnostics on every change, and provides:
- **completion** — instructions, registers, pseudo-registers, textflag macros
and local labels;
- **hover** — instruction summaries and register descriptions from `arch`;
- **document symbols** — `TEXT` functions with their labels, plus `GLOBL`/`DATA`;
- **semantic tokens** — syntax highlighting delivered as LSP semantic tokens,
classified with the lexer plus `arch` (instructions, registers by class,
pseudo-registers, labels, immediates, comments, directives, textflag macros).
Semantic tokens are the key to editor-agnostic highlighting: the editor renders
them from the standard LSP legend, so no editor-specific grammar is needed.
### `asm`
The standalone assembler (Phase 2). Its core is an amd64 instruction encoder:
a REX/ModR-M/SIB/displacement/immediate engine plus the scalar instruction set,
with the Plan 9 operand order (source first) mapped onto the x86 encoding.
Every encoding is validated by decoding it again with `golang.org/x/arch` — the
one module dependency, used in tests only and never linked into the binary.
On top of the encoder, `Assemble` walks a parsed `TEXT` body, converts each
operand to an encoder operand, and lays the instructions out in two passes so
local labels resolve to fixed rel32 jump offsets. The `FP`/`SP` pseudo-
registers are translated onto the hardware stack pointer — `x+N(FP)` becomes
`(N+8)(SP)` for a zero-frame function and `(N+frame+16)(SP)` once a frame
pointer is set up, with the matching Go prologue/epilogue generated — so the
output is byte-identical to the Go assembler for these cases. SIMD is handled
SIMD is handled
by a VEX (AVX/AVX2) encoder — the two- and three-byte VEX prefixes with XMM/YMM
registers — across three operand forms (the three-operand NDS form, the
two-operand reg/rm form, and the immediate-shift form), together covering the
bulk of the integer SIMD set; each encoding is validated by round-trip
decoding. This increment covers register / memory / immediate / FP-frame
operands, local-label jumps and these VEX SIMD forms; the remaining SIMD forms
(shuffles, extract/insert, permute, moves), EVEX / AVX-512, `SB` (global
symbol) operands (relocations) and object-file emission are the rest of
Phase 2.
## Extension points
- **New architecture:** add an entry to the generator in `_gen`, run
`just gen`, and add a `buildXXX()` register file plus a case in `ForArch`.
- **New lint rule:** add a function in `lint` and a rule-code constant.
- **New LSP feature:** add a method case in `dispatch` and a handler.
The phases follow a dependency chain. Phase 1 (static analysis) builds only on
the AST; Phase 2 (the standalone assembler) emits object code; Phases 3
(dynamic analysis) and 4 (the debugger) both consume the execution substrate
that the assembler provides.