docs: sync README, CHANGELOG and docs with the current state

This commit is contained in:
2026-08-30 10:44:18 +02:00
parent 6c1c8d9d96
commit 56f8babbce
7 changed files with 407 additions and 263 deletions
+89 -60
View File
@@ -7,7 +7,7 @@ Repository: [sourcedock.dev/petrbalvin/gasm-devkit](https://sourcedock.dev/petrb
## Design goals
1. **A real AST, not a grammar hack.** The linter, analyser, assembler and
language server all need to *reason* about assembly — not just colour it.
language server all need to *reason* about assembly, not just colour it.
So the centre of the toolkit is a hand-written lexer and a parser that
produce a typed AST with source positions on every node.
2. **Architecture as data, not code.** Per-architecture differences (amd64,
@@ -26,7 +26,7 @@ Repository: [sourcedock.dev/petrbalvin/gasm-devkit](https://sourcedock.dev/petrb
graph TD
SRC["source .s"] --> LEX["lexer<br/>token stream"]
LEX --> PAR["parser<br/>AST + diagnostics"]
LEX --> FMT["s<br/>re-space tokens"]
LEX --> FMT["format<br/>re-space tokens"]
PAR --> LINT["lint<br/>static checks"]
PAR --> LSP["lsp server"]
LEX --> LSP
@@ -69,8 +69,8 @@ instruction, comment, preprocessor) and dispatches. A malformed line is
reported and skipped; it never aborts the file.
Operands are parsed into a faithful, flat representation. The amd64
addressing modes — `reg`, `$imm`, `(base)`, `off(base)`, `(base)(index*scale)`,
`name+off(FP)`, `name<>(SB)` — are all captured structurally, and the original
addressing modes (`reg`, `$imm`, `(base)`, `off(base)`, `(base)(index*scale)`,
`name+off(FP)`, `name<>(SB)`) are all captured structurally, and the original
token text is retained for fidelity.
A deliberate boundary: the AST records **syntax only**. Whether a bare
@@ -80,8 +80,8 @@ arch-agnostic and its output deterministic.
### `arch`
Register files are generated programmatically (the regular `R8`–`R15`,
`X0`–`X15`, `Y0`–`Y15`, `Z0`–`Z31`, `K0`–`K7` ranges) plus the irregularly
Register files are generated programmatically (the regular `R8`-`R15`,
`X0`-`X15`, `Y0`-`Y15`, `Z0`-`Z31`, `K0`-`K7` ranges) plus the irregularly
named registers listed explicitly. Instruction names are **generated from the
Go toolchain's own assembler source** (`cmd/internal/obj/<arch>/anames.go`,
plus the common opcodes and the per-architecture front-end aliases such as the
@@ -95,11 +95,13 @@ fixed-arity instructions (`RET`, `NOP`, `JMP`, `CALL`) carry counts at all.
### `lint`
Rules are conservative by design — silence beats a false positive. The rules
Rules are conservative by design: silence beats a false positive. The rules
are `unknown-instruction`, `operand-count`, `undefined-label`,
`duplicate-label`, `missing-ret`, `missing-textflag-include`, `abi-argsize`,
`unreachable-code`, `register-clobber`, `funcdata-pcdata`, `unused-label`,
`invalid-textflag` and `stack-imbalance`. Every diagnostic carries a stable
`invalid-textflag`, `stack-imbalance`, `register-width-mismatch`,
`abi0-register-args`, `nonportable-register-name` and
`unencodable-instruction`. Every diagnostic carries a stable
code so callers can disable rules individually, and arch-specific rules switch
off entirely when the target architecture cannot be inferred from the file name.
@@ -107,7 +109,7 @@ Two things keep the rules honest on real-world code:
- **Pseudo-ops and macros are not instructions.** `unknown-instruction` knows
the assembler pseudo-ops (`BYTE`, `WORD`, `FUNCDATA`, `PCDATA`, …) and
recognises macro invocations — an in-file `#define` name, or any identifier
recognises macro invocations: an in-file `#define` name, or any identifier
containing an underscore (no Plan 9 mnemonic ever does).
- **Macro-heavy files get the label/RET heuristics turned off.** Without a
preprocessor, labels a macro defines are invisible, so `undefined-label` and
@@ -128,27 +130,27 @@ Two deeper analyses sit on top of the AST:
parameters), and checks the total against the argument size declared in the
`TEXT` directive. It only runs for stack-argument functions (a non-zero
declared arg area that is actually addressed through `FP`), and aborts
silently on a type whose size it cannot determine — so it never guesses.
silently on a type whose size it cannot determine, so it never guesses.
- **`unreachable-code`.** Code after a `RET` and before the next label is
dead. The check is suppressed for any function whose reachability cannot be
decided statically: those using PC-relative jumps (`JMP 2(PC)`),
register-indirect branches (`JALR`/`JR`/`JIRL`/`BR`/`BLR`), or living in a
file with `#ifdef` conditionals. `UNDEF` is deliberately not a terminator —
file with `#ifdef` conditionals. `UNDEF` is deliberately not a terminator:
code after it is occasionally intentional metadata.
- **`register-clobber` (register liveness).** The linter builds the function's
control-flow graph (basic blocks split at labels and after branches, with
fall-through and jump-target edges), computes a conservative per-instruction
register def/use, and runs the standard backward liveness iteration to a fixed
point. On top of that it flags writes to the registers the **Go ABI** fixes
across calls that are never saved and restored — calibrated from
across calls that are never saved and restored, calibrated from
`cmd/compile/abi-internal.md`, *not* the platform ABI: Go's stack-based ABI0
has no System V style callee-saved registers (amd64 `BX`, `R12`–`R15` and
has no System V style callee-saved registers (amd64 `BX`, `R12`-`R15` and
the like are caller-saved or permanent scratch, and hand-written kernels may
clobber them freely). The audited set is the frame pointer and the
the frame pointer, the goroutine pointer per architecture (amd64 `BP`/`R14`, arm64 `R18`/`R28`/
goroutine pointer per architecture (amd64 `BP`/`R14`, arm64 `R18`/`R28`/
`R29`, riscv64 `X27`, loong64 `R22`); the goroutine pointer is reported only
when the function can reach the runtime — it is not `NOSPLIT` or makes a
call — since the ABI0 transition machinery restores it on those paths, and
when the function can reach the runtime (it is not `NOSPLIT` or makes a
call), since the ABI0 transition machinery restores it on those paths, and
NOSPLIT call-free leaves may use it (the runtime's own assembly does). It
runs only on macro-free files, where
no opaque macro can perform the save/restore.
@@ -157,16 +159,16 @@ Two deeper analyses sit on top of the AST:
reference) and a literal index is range-checked; a named index constant such
as `$PCDATA_StackMapIndex` is accepted without a range check.
### `s`
### `format`
The formatter works on the **token stream, not the AST**, so it preserves
every line — comments and blanks included. It normalises indentation, operand
every line, comments and blanks included. It normalises indentation, operand
spacing, per-function mnemonic alignment and blank-line layout: a new block
(a label, `TEXT` or `GLOBL`) is preceded by exactly one blank line (comments
leading a block stay with it), runs of blanks collapse to one, and a `RET`
terminates the body so the next function's doc comment stays at column 0. It
is idempotent and its output always round-trips through the parser. With a
directory argument — or none — it reformats every `.s` file below it in
directory argument, or none, it reformats every `.s` file below it in
place and lists the files changed, the way `go fmt` does (`.` and `_`
directories are skipped).
@@ -176,13 +178,20 @@ The server speaks JSON-RPC 2.0 with `Content-Length` framing over any
`io.Reader`/`io.Writer` (normally stdin/stdout). It maintains an in-memory
document store, republishes diagnostics on every change, and provides:
- **completion** — instructions, registers, pseudo-registers, textflag macros
- **completion**: instructions, registers, pseudo-registers, textflag macros
and local labels;
- **hover** — instruction summaries and register descriptions from `arch`;
- **document symbols** — `TEXT` functions with their labels, plus `GLOBL`/`DATA`;
- **semantic tokens** — syntax highlighting delivered as LSP semantic tokens,
- **hover**: instruction summaries and register descriptions from `arch`;
- **document symbols**: `TEXT` functions with their labels, plus `GLOBL`/`DATA`;
- **semantic tokens**: syntax highlighting delivered as LSP semantic tokens,
classified with the lexer plus `arch` (instructions, registers by class,
pseudo-registers, labels, immediates, comments, directives, textflag macros).
pseudo-registers, labels, immediates, comments, directives, textflag macros);
- **navigation**: go-to-definition from a label reference to its definition,
find references, document highlights of every use of the symbol under the
cursor, rename, and workspace symbol search over the open documents;
- **assists**: document formatting through the `format` package, inlay hints
(the frame size after the TEXT argument area), signature help (the callee's
`// func` signature while the cursor is on a `CALL`), and code actions
offering quick fixes for the `missing-ret` and `unused-label` diagnostics.
Semantic tokens are the key to editor-agnostic highlighting: the editor renders
them from the standard LSP legend, so no editor-specific grammar is needed.
@@ -192,7 +201,7 @@ them from the standard LSP legend, so no editor-specific grammar is needed.
The standalone assembler (Phase 2). Its core is an amd64 instruction encoder:
a REX/ModR-M/SIB/displacement/immediate engine plus the scalar instruction set,
with the Plan 9 operand order (source first) mapped onto the x86 encoding.
Every encoding is validated by decoding it again with `golang.org/x/arch` — the
Every encoding is validated by decoding it again with `golang.org/x/arch`, the
one module dependency, used in tests only and never linked into the binary.
A **RISC-V encoder** (Phase 5, RV64IMAFDC + RVC compression) encodes the full
@@ -232,12 +241,12 @@ fixed point, and jump-to-jump chains are folded (a conditional jump to a label
whose only instruction is an unconditional jump is redirected to the ultimate
target) exactly as the Go toolchain's linker does before it encodes branches.
The `FP`/`SP` pseudo-
registers are translated onto the hardware stack pointer — `x+N(FP)` becomes
registers are translated onto the hardware stack pointer: `x+N(FP)` becomes
`(N+8)(SP)` for a zero-frame function and `(N+frame+16)(SP)` once a frame
pointer is set up, with the matching Go prologue/epilogue generated — so the
pointer is set up, with the matching Go prologue/epilogue generated, so the
output is byte-identical to the Go assembler for these cases. SIMD is handled
by a VEX (AVX/AVX2) encoder — the two- and three-byte VEX prefixes with XMM/YMM
registers — across eight operand forms: the three-operand NDS form, the
by a VEX (AVX/AVX2) encoder (the two- and three-byte VEX prefixes with XMM/YMM
registers) across eight operand forms: the three-operand NDS form, the
two-operand reg/rm form, the immediate-shift form (plus the variable-count
shifts, which share the NDS shape with the count in an XMM register or
memory), the immediate shuffle form (`VPSHUFD`, `VPERMQ`), the
@@ -245,24 +254,24 @@ three-operand-plus-immediate form (`VSHUFPD`,
`VPERM2I128`, `VINSERTI128`), the lane-extract form (`VEXTRACTI128`,
`VEXTRACTF128`, where the YMM source occupies the reg field and the XMM or
memory destination r/m), the direction-sensitive moves (`VMOVDQU`, `VMOVUPD`,
`VMOVD`, `VMOVQ`, `VMOVSD`), the floating-point and FMA arithmetic — the
`VMOVD`, `VMOVQ`, `VMOVSD`), the floating-point and FMA arithmetic: the
packed double operations (`VADDPD`/`VSUBPD`/`VMULPD`/`VDIVPD`/`VMINPD`/
`VMAXPD`), the unpacks (`VUNPCKHPD`/`VUNPCKLPD`), the scalar SD and SS
operations, `VMOVDDUP`, `VXORPD`, the width-changing conversions
(`VCVTDQ2PS`, `VCVTPS2PD`, `VCVTDQ2PD`, and the `VCVTPD2DQX`/`Y` and
`VCVTTPD2DQX`/`Y` spellings, whose length follows the wider source) and
`VFMADD231PD` — and the no-operand `VZEROUPPER`, together with `VPERMD` and
`VFMADD231PD`, and the no-operand `VZEROUPPER`, together with `VPERMD` and
the scalar families (`CMOVcc`, `SETcc`, `LZCNT`/`TZCNT`, the extending moves,
`CVTSx2SD`, `IMUL3`) and the EVEX (AVX-512) prefix — the four-byte prefix with
5-bit register fields (Z0–Z31, X/Y 16–31, with the mod=11 quirk that carries
rm[4] in X̄), opmask registers (K0–K7 as operands, mask destinations and
explicit merging/zeroing masks — written the way Go writes them, as a K
`CVTSx2SD`, `IMUL3`) and the EVEX (AVX-512) prefix, the four-byte prefix with
5-bit register fields (Z0-Z31, X/Y 16-31, with the mod=11 quirk that carries
rm[4] in X̄), opmask registers (K0-K7 as operands, mask destinations and
explicit merging/zeroing masks, written the way Go writes them, as a K
operand among the operands plus a `.Z` mnemonic suffix), and the compressed
disp8×N displacement, whose multiplier follows the memory operand's size —
disp8×N displacement, whose multiplier follows the memory operand's size,
covering every instruction the go-flac and go-lz4 AVX2/AVX-512 kernels use,
plus the common AVX-512 F/BW integer set, the floating-point and conversion
set (the packed double and single arithmetic, the scalar SD/SS forms —
whose EVEX encodings serve masked and zeroing use — `VMOVDDUP`, the
set (the packed double and single arithmetic, the scalar SD/SS forms
(whose EVEX encodings serve masked and zeroing use), `VMOVDDUP`, the
replicating moves, and the width-changing conversions, including the
`VCVTPD2DQ`/`VCVTTPD2DQ` family whose length follows the wider source
operand), and the wider AVX-512 set: ternary logic, lane shuffles, inserts
@@ -272,21 +281,21 @@ expand/compress family, the broadcasts, the opmask-register instructions
moves and the remaining extending/narrowing moves, the floating-point
helper and conversion tail (VRCP14*, VRSQRT14*, VGETEXP*, VGETMANT*,
VSCALEF*, VRNDSCALE*, VREDUCE*, VFIXUPIMM*, VRANGE*, VFPCLASS* with an
opmask destination, and the VCVT* conversions — signed, unsigned and
opmask destination, and the VCVT* conversions, signed, unsigned and
truncating, including the length-suffixed X/Y spellings and the
mask/vector conversions VPMOVM2*/VPMOV*2M, and the scalar conversions
between vector and general-purpose registers (VCVT{,T}S{D,S}2SI{,Q} and
the unsigned forms, VCVTSI2*/VCVTUSI2*), and gather/scatter with VSIB addressing — both the
the unsigned forms, VCVTSI2*/VCVTUSI2*), and gather/scatter with VSIB addressing, both the
VEX spelling with a vector mask register and the EVEX spelling with an
explicit K mask, where the EVEX length follows the VSIB index register,
not the data register. The EVEX mnemonic
suffixes — rounding modes (.RN_SAE/.RD_SAE/.RU_SAE/.RZ_SAE),
suppress-all-exceptions (.SAE) and memory broadcast (.BCST) — set the EVEX
suffixes (rounding modes (.RN_SAE/.RD_SAE/.RU_SAE/.RZ_SAE),
suppress-all-exceptions (.SAE) and memory broadcast (.BCST)) set the EVEX
b bit and the L'L rounding-control field (broadcast keeps the vector length
and scales disp8 by the element size), and combine with the .Z zeroing
suffix. Every encoding is validated two ways: by
round-trip decoding through `golang.org/x/arch`, and byte-for-byte against
the machine code the real Go assembler emits — a comparison that holds for
the machine code the real Go assembler emits, a comparison that holds for
whole functions: all 27 functions of both kernels assemble to exactly the Go
toolchain's bytes, the lone exception being the displacements of the
static-constant loads, which the Go linker fills at link time.
@@ -300,14 +309,17 @@ the bytes are self-consistent at any base address. References to symbols no
object-file emitters turn the whole image into a linkable object: the ELF
writer (`gasm asm --format elf`) lays the code and data out as `.text`/`.data`
sections, exports a symbol per
`TEXT` and `GLOBL` (the `<>` ones local, the rest global) and emit one
PC-relative relocation per static-symbol reference — undefined external
symbols included, so the output links with the system toolchain. The GOOBJ
`TEXT` and `GLOBL` (the `<>` ones local, the rest global), emits one
PC-relative relocation per static-symbol reference, undefined external
symbols included, and appends the DWARF5 debug sections
(`.debug_abbrev`, `.debug_info`, `.debug_line`, `.debug_line_str`, and a
`.debug_frame` CFI section on amd64) so `addr2line` and GDB/LLDB can debug
the output. The GOOBJ
emitter (`gasm asm --format goobj`) writes the format the Go linker consumes
directly: the functions as non-package symbols (the way `cmd/asm` records
assembly symbols), the `GLOBL` data, one `FuncInfo` per function and the
pc-value tables — `pcsp` built from the prologue and epilogue stack
boundaries, plus flat `pcfile`, `pcline` and `pcinline` tables — so a
pc-value tables (`pcsp` built from the prologue and epilogue stack
boundaries, plus flat `pcfile`, `pcline` and `pcinline` tables), so a
gasm-assembled object drops into a `go build` in place of the toolchain's.
The object preamble (the version-and-experiment header the linker compares
verbatim) is captured from the installed `go tool asm`, so the output is
@@ -315,18 +327,21 @@ always consistent with the toolchain that links it. RISC-V and LoongArch
GOOBJ emission share this emitter: the loong64 marker with
R_LOONG64_ADDR_HI/LO relocation types, and the riscv64 marker with a single
R_RISCV_PCREL_ITYPE/STYPE relocation per AUIPC pair (plus `R_RISCV_JAL` for
`CALL sym(SB)`) — the model `cmd/asm`
writes, not the ELF HI20/LO12 pair — and both link into a real `go build` for
`CALL sym(SB)`), the model `cmd/asm`
writes, not the ELF HI20/LO12 pair, and both link into a real `go build` for
their `GOARCH`. Per function, the emitter also writes the two DWARF
symbols the linker's DWARF pass reads verbatim — the subprogram DIE
symbols the linker's DWARF pass reads verbatim: the subprogram DIE
(`SDWARFFCN`) and the `.debug_line` state-machine program (`SDWARFLINES`),
both built the way `cmd/asm` builds them (the DIE carries the
R_DWTXTADDR_U4 address reference; the line program one row per source-line
change, in the same special-opcode encoding) — and the pc-value deltas are
change, in the same special-opcode encoding), and the pc-value deltas are
in the architecture's MinLC units, as the runtime's `pcvalue` expects.
External cross-package references remain future work (the amd64 and RISC-V
paths resolve them; LoongArch does not yet); the rest of Phase 2 is those
and the remaining EVEX forms.
Cross-package external references resolve on all four architectures: when
`img.Externals` is non-empty, the GOOBJ emitter locates the referenced
package's `.a` archive via `go list -json -export`, reads the GOOBJ symbol
definitions the linker reads, and wires the resolved package and symbol
indices into the emission, so a gasm-assembled object links against the
compiled packages it references.
### `verify`
@@ -343,6 +358,9 @@ stack pointer in a package global and jumps to the target), and recovers
control when the function RETs into `leaveJIT` (which restores the Go stack
and returns). A 64-byte pad below the return address accommodates the
ABIInternal wrapper that the Go runtime interposes on assembly functions.
Every supported architecture carries its own hand-written trampoline pair
(`trampoline_amd64.s`, `trampoline_arm64.s`, `trampoline_riscv64.s`,
`trampoline_loong64.s`), so `Call` works wherever the toolkit runs.
`Load` / `LoadSource` / `LoadAST` parse, assemble and map a `.s` file in one
step, returning a `Kernel` whose `CallFunc` method marshals the argument block
@@ -351,16 +369,20 @@ assembler’s `Image.Bytes()` provides the code-and-data concatenation.
The `gasm verify` CLI subcommand exposes this: it loads a file, reports the
available functions and (with `-smoke`) calls each NOSPLIT function with zeroed
arguments to confirm the trampoline round-trips. `gasm verify --fuzz` combines
arguments to confirm the trampoline round-trips. The `-smoke` and `-abi`
sweeps run in parallel and each inside a child process, so a function that
faults is reported without ending the sweep. `gasm verify --fuzz` combines
ABI checks (sentinel registers, canary, stack bounds) with differential fuzz
testing, comparing the JIT-assembled kernel against the portable Go reference
bit-for-bit while verifying the ABI contract on every iteration. When a fuzz
iteration crashes or mismatches, `FuzzResult.CrashInput` stores the exact input
for reproducibility. `gasm verify --call <func> --buf name:size:pattern`
invokes a single function with user-supplied buffers (patterns: zero, ones,
seq, or hex), printing the ABI0 argument block before and after the call —
seq, or hex), printing the ABI0 argument block before and after the call,
useful for partial functions (e.g. decoders) that crash on random input but
should succeed on valid data. `gasm verify --ground-truth` compares the
should succeed on valid data; `--args name=value,...` supplies scalar
arguments (decimal or `0x` hex) alongside the buffers. `gasm verify
--ground-truth` compares the
assembled machine code byte-for-byte against `go tool asm` (relocation sites
masked), reporting any encoding drift.
@@ -375,8 +397,15 @@ The child pins its goroutine to the OS thread with `runtime.LockOSThread`
so the traced thread is the one executing JIT code. The REPL provides
single-step, register inspection (GPR + YMM/XMM via `PTRACE_GETFPREGS`),
label resolution, named buffer allocation with pattern filling
(`--buf name:size:pattern` — zero, ones, seq, or hex), and breakpoint
management.
(`--buf name:size:pattern`: zero, ones, seq, or hex), and breakpoint
management. Breakpoints accept conditions
(`break <label> if <reg> <op> <val>`, including register-against-register
comparisons), and hardware watchpoints work on all four architectures.
For non-interactive use, `--script` runs REPL commands from a file (or
stdin) and exits, `--timeout` kills the debuggee when a run hangs (the
watchdog is armed before the ptrace attach, so a sandboxed debuggee cannot
block it), and `--cover` runs to completion with a breakpoint on every
label and reports which blocks executed.
## Extension points