# Architecture How gasm-devkit is put together and why. Repository: [sourcedock.dev/petrbalvin/gasm-devkit](https://sourcedock.dev/petrbalvin/gasm-devkit) ## Design goals 1. **A real AST, not a grammar hack.** The linter, analyser, assembler and language server all need to *reason* about assembly — not just colour it. So the centre of the toolkit is a hand-written lexer and a parser that produce a typed AST with source positions on every node. 2. **Architecture as data, not code.** Per-architecture differences (amd64, arm64, riscv64, loong64) live in register and instruction *tables* (`arch`), never in `if arch == …` branches scattered through the logic. The instruction tables are generated from the Go toolchain's own assembler source (`just gen`), so adding or refreshing an architecture is a data operation, not a coding one. 3. **Open integration surface.** Everything the toolkit can do is reachable through two vendor-neutral interfaces: a CLI and an LSP server. No editor owns the toolkit; the toolkit is offered to editors on standard terms. ## Pipeline ```mermaid graph TD SRC["source .s"] --> LEX["lexer
token stream"] LEX --> PAR["parser
AST + diagnostics"] LEX --> FMT["format
re-space tokens"] PAR --> LINT["lint
static checks"] PAR --> LSP["lsp server"] LEX --> LSP ARCH["arch tables
amd64 / arm64 / riscv64 / loong64"] --> LINT ARCH --> LSP LINT --> LSP FMT --> CLI["gasm CLI"] LINT --> CLI PAR --> CLI LEX --> CLI LSP --> EDITOR["any LSP editor"] ``` The lexer is the shared foundation: the parser builds the AST from it, the formatter re-spaces its tokens directly, and the language server uses it for semantic highlighting. ## Components ### `token` and `lexer` The scanner is hand-written and permissive: it never panics and maps anything it cannot classify to an `Illegal` token, so every downstream tool still works on malformed input. Newlines are significant tokens, because Plan 9 assembly is line-oriented and the parser relies on line structure. The middle dot (`·`, U+00B7) is treated as an identifier character so that `·funcName(SB)` lexes as one symbol. Multi-character operators (`<<`, `>>`, `->`) are recognised so arm64 shift operands scan correctly. A backslash immediately before a newline is a C-preprocessor line continuation (used by `#define` macros in the runtime `.s` files); the lexer splices the lines together so a multi-line macro becomes one logical line the parser treats as an opaque preprocessor directive. ### `ast` and `parser` The parser is **line-oriented**, matching how the Plan 9 assembler reads a file: it groups tokens into lines, classifies each line (directive, label, instruction, comment, preprocessor) and dispatches. A malformed line is reported and skipped; it never aborts the file. Operands are parsed into a faithful, flat representation. The amd64 addressing modes — `reg`, `$imm`, `(base)`, `off(base)`, `(base)(index*scale)`, `name+off(FP)`, `name<>(SB)` — are all captured structurally, and the original token text is retained for fidelity. A deliberate boundary: the AST records **syntax only**. Whether a bare identifier is a register or a label is an *architecture* question, so it is left to `arch` and resolved in the lint/lsp layers. This keeps the parser arch-agnostic and its output deterministic. ### `arch` Register files are generated programmatically (the regular `R8`–`R15`, `X0`–`X15`, `Y0`–`Y15`, `Z0`–`Z31`, `K0`–`K7` ranges) plus the irregularly named registers listed explicitly. Instruction names are **generated from the Go toolchain's own assembler source** (`cmd/internal/obj//anames.go`, plus the common opcodes and the per-architecture front-end aliases such as the arm64 `B`/`BL` branches and the `.P`/`.W` load-store addressing suffixes) by `just gen`, so the tables always match what the real assembler accepts. Each mnemonic maps to a summary and an optional operand-count range; counts are recorded only where unambiguous (`-1` disables the operand-count lint for that instruction) so the linter stays silent rather than guess. For architectures with highly variable operand forms (arm64, riscv64, loong64) only a few fixed-arity instructions (`RET`, `NOP`, `JMP`, `CALL`) carry counts at all. ### `lint` Rules are conservative by design — silence beats a false positive. The rules are `unknown-instruction`, `operand-count`, `undefined-label`, `duplicate-label`, `missing-ret`, `missing-textflag-include`, `abi-argsize` and `unreachable-code`. Every diagnostic carries a stable code so callers can disable rules individually, and arch-specific rules switch off entirely when the target architecture cannot be inferred from the file name. Two things keep the rules honest on real-world code: - **Pseudo-ops and macros are not instructions.** `unknown-instruction` knows the assembler pseudo-ops (`BYTE`, `WORD`, `FUNCDATA`, `PCDATA`, …) and recognises macro invocations — an in-file `#define` name, or any identifier containing an underscore (no Plan 9 mnemonic ever does). - **Macro-heavy files get the label/RET heuristics turned off.** Without a preprocessor, labels a macro defines are invisible, so `undefined-label` and `missing-ret` are suppressed for files that use macros (an in-file `#define` or a `#include` of anything other than `textflag.h`). `missing-ret` also treats a trailing unconditional jump and `UNDEF` as valid terminators. The result is validated by `TestGoRuntimeCorpus`, which parses and lints every `src/runtime/*.s` file the toolchain ships for all four architectures and asserts zero parse errors and zero error-severity diagnostics. Two deeper analyses sit on top of the AST: - **`abi-argsize`.** Hand-written kernels document their signature in a `// func …` comment above the `TEXT`. The linter parses that signature with the standard library's Go parser, lays out the parameters and results under Go's ABI0 stack rules (results begin on a word boundary after the parameters), and checks the total against the argument size declared in the `TEXT` directive. It only runs for stack-argument functions (a non-zero declared arg area that is actually addressed through `FP`), and aborts silently on a type whose size it cannot determine — so it never guesses. - **`unreachable-code`.** Code after a `RET` and before the next label is dead. The check is suppressed for any function whose reachability cannot be decided statically: those using PC-relative jumps (`JMP 2(PC)`), register-indirect branches (`JALR`/`JR`/`JIRL`/`BR`/`BLR`), or living in a file with `#ifdef` conditionals. `UNDEF` is deliberately not a terminator — code after it is occasionally intentional metadata. - **`register-clobber` (register liveness).** The linter builds the function's control-flow graph (basic blocks split at labels and after branches, with fall-through and jump-target edges), computes a conservative per-instruction register def/use, and runs the standard backward liveness iteration to a fixed point. On top of that it flags writes to the registers the **Go ABI** fixes across calls that are never saved and restored — calibrated from `cmd/compile/abi-internal.md`, *not* the platform ABI: Go's stack-based ABI0 has no System V style callee-saved registers (amd64 `BX`, `R12`–`R15` and the like are caller-saved or permanent scratch, and hand-written kernels may clobber them freely). The audited set is the frame pointer and the the frame pointer, the goroutine pointer per architecture (amd64 `BP`/`R14`, arm64 `R18`/`R28`/ `R29`, riscv64 `X27`, loong64 `R22`); the goroutine pointer is reported only when the function can reach the runtime — it is not `NOSPLIT` or makes a call — since the ABI0 transition machinery restores it on those paths, and NOSPLIT call-free leaves may use it (the runtime's own assembly does). It runs only on macro-free files, where no opaque macro can perform the save/restore. - **`funcdata-pcdata`.** `FUNCDATA $idx, sym(SB)` and `PCDATA $idx, $val` are checked for well-formed operands (arity, immediate index and value, symbol reference) and a literal index is range-checked; a named index constant such as `$PCDATA_StackMapIndex` is accepted without a range check. ### `format` The formatter works on the **token stream, not the AST**, so it preserves every line — comments and blanks included. It normalises indentation, operand spacing, per-function mnemonic alignment and blank-line layout: a new block (a label, `TEXT` or `GLOBL`) is preceded by exactly one blank line (comments leading a block stay with it), runs of blanks collapse to one, and a `RET` terminates the body so the next function's doc comment stays at column 0. It is idempotent and its output always round-trips through the parser. With a directory argument — or none — it reformats every `.s` file below it in place and lists the files changed, the way `go fmt` does (`.` and `_` directories are skipped). ### `lsp` The server speaks JSON-RPC 2.0 with `Content-Length` framing over any `io.Reader`/`io.Writer` (normally stdin/stdout). It maintains an in-memory document store, republishes diagnostics on every change, and provides: - **completion** — instructions, registers, pseudo-registers, textflag macros and local labels; - **hover** — instruction summaries and register descriptions from `arch`; - **document symbols** — `TEXT` functions with their labels, plus `GLOBL`/`DATA`; - **semantic tokens** — syntax highlighting delivered as LSP semantic tokens, classified with the lexer plus `arch` (instructions, registers by class, pseudo-registers, labels, immediates, comments, directives, textflag macros). Semantic tokens are the key to editor-agnostic highlighting: the editor renders them from the standard LSP legend, so no editor-specific grammar is needed. ### `asm` The standalone assembler (Phase 2). Its core is an amd64 instruction encoder: a REX/ModR-M/SIB/displacement/immediate engine plus the scalar instruction set, with the Plan 9 operand order (source first) mapped onto the x86 encoding. Every encoding is validated by decoding it again with `golang.org/x/arch` — the one module dependency, used in tests only and never linked into the binary. On top of the encoder, `Assemble` walks a parsed `TEXT` body, converts each operand to an encoder operand, and lays the instructions out so local labels resolve to relative jump offsets: jumps start in the short (rel8) form and expand to rel32 when the settled displacement does not fit, iterating to a fixed point, and jump-to-jump chains are folded (a conditional jump to a label whose only instruction is an unconditional jump is redirected to the ultimate target) exactly as the Go toolchain's linker does before it encodes branches. The `FP`/`SP` pseudo- registers are translated onto the hardware stack pointer — `x+N(FP)` becomes `(N+8)(SP)` for a zero-frame function and `(N+frame+16)(SP)` once a frame pointer is set up, with the matching Go prologue/epilogue generated — so the output is byte-identical to the Go assembler for these cases. SIMD is handled by a VEX (AVX/AVX2) encoder — the two- and three-byte VEX prefixes with XMM/YMM registers — across eight operand forms: the three-operand NDS form, the two-operand reg/rm form, the immediate-shift form (plus the variable-count shifts, which share the NDS shape with the count in an XMM register or memory), the immediate shuffle form (`VPSHUFD`, `VPERMQ`), the three-operand-plus-immediate form (`VSHUFPD`, `VPERM2I128`, `VINSERTI128`), the lane-extract form (`VEXTRACTI128`, `VEXTRACTF128`, where the YMM source occupies the reg field and the XMM or memory destination r/m), the direction-sensitive moves (`VMOVDQU`, `VMOVUPD`, `VMOVD`, `VMOVQ`, `VMOVSD`), the floating-point and FMA arithmetic — the packed double operations (`VADDPD`/`VSUBPD`/`VMULPD`/`VDIVPD`/`VMINPD`/ `VMAXPD`), the unpacks (`VUNPCKHPD`/`VUNPCKLPD`), the scalar SD and SS operations, `VMOVDDUP`, `VXORPD`, the width-changing conversions (`VCVTDQ2PS`, `VCVTPS2PD`, `VCVTDQ2PD`, and the `VCVTPD2DQX`/`Y` and `VCVTTPD2DQX`/`Y` spellings, whose length follows the wider source) and `VFMADD231PD` — and the no-operand `VZEROUPPER`, together with `VPERMD` and the scalar families (`CMOVcc`, `SETcc`, `LZCNT`/`TZCNT`, the extending moves, `CVTSx2SD`, `IMUL3`) and the EVEX (AVX-512) prefix — the four-byte prefix with 5-bit register fields (Z0–Z31, X/Y 16–31, with the mod=11 quirk that carries rm[4] in X̄), opmask registers (K0–K7 as operands, mask destinations and explicit merging/zeroing masks — written the way Go writes them, as a K operand among the operands plus a `.Z` mnemonic suffix), and the compressed disp8×N displacement, whose multiplier follows the memory operand's size — covering every instruction the go-flac and go-lz4 AVX2/AVX-512 kernels use, plus the common AVX-512 F/BW integer set, the floating-point and conversion set (the packed double and single arithmetic, the scalar SD/SS forms — whose EVEX encodings serve masked and zeroing use — `VMOVDDUP`, the replicating moves, and the width-changing conversions, including the `VCVTPD2DQ`/`VCVTTPD2DQ` family whose length follows the wider source operand), and the wider AVX-512 set: ternary logic, lane shuffles, inserts and extracts, compares with an opmask destination, the permutes, the expand/compress family, the broadcasts, the opmask-register instructions (KAND/KOR/KXNOR/KADD/KUNPCK/KNOT/KSHIFTL/KORTEST and KMOVQ), the aligned moves and the remaining extending/narrowing moves, the floating-point helper and conversion tail (VRCP14*, VRSQRT14*, VGETEXP*, VGETMANT*, VSCALEF*, VRNDSCALE*, VREDUCE*, VFIXUPIMM*, VRANGE*, VFPCLASS* with an opmask destination, and the VCVT* conversions — signed, unsigned and truncating, including the length-suffixed X/Y spellings and the mask/vector conversions VPMOVM2*/VPMOV*2M, and the scalar conversions between vector and general-purpose registers (VCVT{,T}S{D,S}2SI{,Q} and the unsigned forms, VCVTSI2*/VCVTUSI2*), and gather/scatter with VSIB addressing — both the VEX spelling with a vector mask register and the EVEX spelling with an explicit K mask, where the EVEX length follows the VSIB index register, not the data register. The EVEX mnemonic suffixes — rounding modes (.RN_SAE/.RD_SAE/.RU_SAE/.RZ_SAE), suppress-all-exceptions (.SAE) and memory broadcast (.BCST) — set the EVEX b bit and the L'L rounding-control field (broadcast keeps the vector length and scales disp8 by the element size), and combine with the .Z zeroing suffix. Every encoding is validated two ways: by round-trip decoding through `golang.org/x/arch`, and byte-for-byte against the machine code the real Go assembler emits — a comparison that holds for whole functions: all 27 functions of both kernels assemble to exactly the Go toolchain's bytes, the lone exception being the displacements of the static-constant loads, which the Go linker fills at link time. File-level assembly (`AssembleFile`) goes beyond single functions: it materialises the file's static symbols (`GLOBL`/`DATA`) in a data section behind the code and resolves references to them (`mask<>(SB)`) to RIP-relative loads whose displacements point inside the resulting image, so the bytes are self-consistent at any base address. References to symbols no `GLOBL` defines are kept as relocations on the function layout, and the object-file emitters turn the whole image into a linkable object: the ELF and Mach-O writers (`gasm asm --format elf|macho`) lay the code and data out as `.text`/`.data` (or `__text`/`__data`) sections, export a symbol per `TEXT` and `GLOBL` (the `<>` ones local, the rest global) and emit one PC-relative relocation per static-symbol reference — undefined external symbols included, so the output links with the system toolchain. The GOOBJ emitter (`gasm asm --format goobj`) writes the format the Go linker consumes directly: the functions as non-package symbols (the way `cmd/asm` records assembly symbols), the `GLOBL` data, one `FuncInfo` per function and the pc-value tables — `pcsp` built from the prologue and epilogue stack boundaries, plus flat `pcfile`, `pcline` and `pcinline` tables — so a gasm-assembled object drops into a `go build` in place of the toolchain's. The object preamble (the version-and-experiment header the linker compares verbatim) is captured from the installed `go tool asm`, so the output is always consistent with the toolchain that links it. External cross-package references and the implicit funcdata/DWARF symbols remain future work (the linker fills the latter's defaults); the rest of Phase 2 is those, the remaining EVEX forms and the other architectures. ### `verify` The dynamic-analysis substrate (Phase 3). It JIT-loads assembled images into executable memory and invokes them directly, enabling differential testing, runtime ABI checks and coverage profiling. The execution model is pure Go (stdlib only). `Map` copies machine code into an anonymous `syscall.Mmap` mapping and enforces W^X (write the bytes, then `mprotect` to read-execute). `Call` prepares a stack whose first word is the address of an assembly trampoline (`leaveJIT`), lays the ABI0 argument block after it, switches to that stack via `enterJIT` (which saves the Go stack pointer in a package global and jumps to the target), and recovers control when the function RETs into `leaveJIT` (which restores the Go stack and returns). A 64-byte pad below the return address accommodates the ABIInternal wrapper that the Go runtime interposes on assembly functions. `Load` / `LoadSource` / `LoadAST` parse, assemble and map a `.s` file in one step, returning a `Kernel` whose `CallFunc` method marshals the argument block by name. The image must be self-contained (no external relocations); the assembler’s `Image.Bytes()` provides the code-and-data concatenation. The `gasm verify` CLI subcommand exposes this: it loads a file, reports the available functions and (with `-smoke`) calls each NOSPLIT function with zeroed arguments to confirm the trampoline round-trips. ### `debug` The interactive debugger (Phase 4, linux/amd64). It launches the target function in a child process that maps the JIT code, calls `PTRACE_TRACEME`, and stops; the parent attaches via ptrace and controls execution. Breakpoints are patched as INT3 bytes through `/proc/pid/mem` (PTRACE_PEEKTEXT is unreliable with Go's multi-threaded runtime). The child pins its goroutine to the OS thread with `runtime.LockOSThread` so the traced thread is the one executing JIT code. The REPL provides single-step, register inspection, label resolution, and breakpoint management. ## Extension points - **New architecture:** add an entry to the generator in `_gen`, run `just gen`, and add a `buildXXX()` register file plus a case in `ForArch`. - **New lint rule:** add a function in `lint` and a rule-code constant. - **New LSP feature:** add a method case in `dispatch` and a handler. The phases follow a dependency chain. Phase 1 (static analysis) builds only on the AST; Phase 2 (the standalone assembler) emits object code; Phases 3 (dynamic analysis) and 4 (the debugger) both consume the execution substrate that the assembler provides.