docs: state the validation status and correct claims the material contradicts

Assisted-by: DeepSeek V4.1 Flash
This commit is contained in:
2026-09-20 01:40:51 +02:00
parent 2931bbd6b2
commit f0d5238c47
22 changed files with 356 additions and 191 deletions
+67 -39
View File
@@ -13,11 +13,13 @@ Three design goals shape everything below.
So the centre of the toolkit is a hand-written lexer and a parser that
produce a typed AST with source positions on every node.
2. **Architecture as data, not code.** Per-architecture differences (amd64,
arm64, riscv64, loong64) live in register and instruction *tables* (`arch`),
never in `if arch == …` branches scattered through the logic. The
instruction tables are generated from the Go toolchain's own assembler
source (`just gen`), so adding or refreshing an architecture is a data
operation, not a coding one.
arm64, riscv64, loong64) live in register and instruction *tables* (`arch`)
and per-architecture encoders, rather than in `if arch == …` branches
threaded through the analysis; the arch tests that remain are dispatch and
policy points, such as which encoder a file name selects and which
registers the liveness pass audits. The instruction tables are generated
from the Go toolchain's own assembler source (`just gen`), so refreshing an
architecture is a data operation, not a coding one.
3. **Open integration surface.** Everything the toolkit can do is reachable
through two vendor-neutral interfaces: a CLI and an LSP server. No editor
owns the toolkit; the toolkit is offered to editors on standard terms.
@@ -35,10 +37,19 @@ flowchart TD
ARCH["arch tables<br/>amd64 / arm64 / riscv64 / loong64"] --> LINT
ARCH --> LSP
LINT --> LSP
PAR --> ASM["asm<br/>encoders, image, object emitters"]
ASM --> VER["verify<br/>JIT mapping, ABI checks, fuzzing"]
ASM --> DBG["debug<br/>ptrace session"]
VER --> DBG
DIS["disasm<br/>golang.org/x/arch"] --> DBG
FMT --> CLI["gasm CLI"]
LINT --> CLI
PAR --> CLI
LEX --> CLI
ASM --> CLI
VER --> CLI
DBG --> CLI
DIS --> CLI
LSP --> EDITOR["any LSP editor"]
```
@@ -70,9 +81,11 @@ assembler provides.
The boundaries matter as much as the responsibilities: `ast` records syntax
only, and whether a name is a register or a label is left to `arch`, so the
parser stays architecture-agnostic. `asm` and `verify` are the only packages
that touch machine code and executable memory, and `cmd/gasm` owns no logic
beyond flags and output.
parser stays architecture-agnostic. `asm` produces the machine code, `verify`
and `debug` are the two packages that map it executable (read-execute in
`verify`, read-write-execute in the debuggee), and `cmd/gasm` is the CLI, with
the verify sweep orchestration and the audit, scaffold and unified-diff
helpers beside its flags and output.
### `token` and `lexer`
@@ -112,14 +125,17 @@ Register files are generated programmatically (the regular `R8`-`R15`,
`X0`-`X15`, `Y0`-`Y15`, `Z0`-`Z31`, `K0`-`K7` ranges) plus the irregularly
named registers listed explicitly. Instruction names are **generated from the
Go toolchain's own assembler source** (`cmd/internal/obj/<arch>/anames.go`,
plus the common opcodes and the per-architecture front-end aliases such as the
arm64 `B`/`BL` branches and the `.P`/`.W` load-store addressing suffixes) by
`just gen`, so the tables always match what the real assembler accepts. Each
mnemonic maps to a summary and an optional operand-count range; counts are
recorded only where unambiguous (`-1` disables the operand-count lint for that
instruction) so the linter stays silent rather than guess. For architectures
with highly variable operand forms (arm64, riscv64, loong64) only a few
fixed-arity instructions (`RET`, `NOP`, `JMP`, `CALL`) carry counts at all.
plus the common opcodes in `cmd/internal/obj/util.go`) by `just gen`, so the
tables always match what the real assembler accepts. The spellings the
toolchain's tables do not carry are hand-maintained instead: the front-end
alias lists in `arch/arm64.go`, `arch/amd64.go` and `arch/loong64.go` (the
arm64 `B`/`BL` branches among them), and the arm64 `.P`/`.W` load-store suffix
stripping in `arch/arch.go`. Each mnemonic maps to a summary and an optional
operand-count range; counts are recorded only where unambiguous (`-1`
disables the operand-count lint for that instruction) so the linter stays
silent rather than guess. For architectures with highly variable operand
forms (arm64, riscv64, loong64) `relaxCounts` clears those counts, leaving
`RET` and `NOP` with a range (`RET` alone on riscv64).
### `lint`
@@ -291,11 +307,11 @@ registers are translated onto the hardware stack pointer: `x+N(FP)` becomes
pointer is set up, with the matching Go prologue/epilogue generated, so the
output is byte-identical to the Go assembler for these cases. SIMD is handled
by a VEX (AVX/AVX2) encoder (the two- and three-byte VEX prefixes with XMM/YMM
registers) across eight operand forms: the three-operand NDS form, the
two-operand reg/rm form, the immediate-shift form (plus the variable-count
shifts, which share the NDS shape with the count in an XMM register or
memory), the immediate shuffle form (`VPSHUFD`, `VPERMQ`), the
three-operand-plus-immediate form (`VSHUFPD`,
registers) over nine operand forms plus a dedicated move encoder: the
three-operand NDS form, the two-operand reg/rm form, the immediate-shift form
(plus the variable-count shifts, which share the NDS shape with the count in
an XMM register or memory), the immediate shuffle form (`VPSHUFD`, `VPERMQ`),
the three-operand-plus-immediate form (`VSHUFPD`,
`VPERM2I128`, `VINSERTI128`), the lane-extract form (`VEXTRACTI128`,
`VEXTRACTF128`, where the YMM source occupies the reg field and the XMM or
memory destination r/m), the direction-sensitive moves (`VMOVDQU`, `VMOVUPD`,
@@ -340,10 +356,10 @@ b bit and the L'L rounding-control field (broadcast keeps the vector length
and scales disp8 by the element size), and combine with the .Z zeroing
suffix. Every encoding is validated two ways: by
round-trip decoding through `golang.org/x/arch`, and byte-for-byte against
the machine code the real Go assembler emits, a comparison that holds for
whole functions: all 27 functions of both kernels assemble to exactly the Go
toolchain's bytes, the lone exception being the displacements of the
static-constant loads, which the Go linker fills at link time.
the machine code the real Go assembler emits; the parity suites carry that
comparison over whole kernel files on all four architectures, with the
relocation fields masked because the Go linker fills those displacements at
link time.
File-level assembly (`AssembleFile`) goes beyond single functions: it
materialises the file's static symbols (`GLOBL`/`DATA`) in a data section
@@ -426,10 +442,13 @@ The `gasm verify` CLI subcommand exposes this: it loads a file, reports the
available functions and (with `-smoke`) calls each NOSPLIT function with zeroed
arguments to confirm the trampoline round-trips. The `-smoke` and `-abi`
sweeps run in parallel and each inside a child process, so a function that
faults is reported without ending the sweep. `gasm verify --fuzz` combines
ABI checks (sentinel registers, canary, stack bounds) with differential fuzz
testing, comparing the JIT-assembled kernel against the portable Go reference
bit-for-bit while verifying the ABI contract on every iteration. When a fuzz
faults is reported without ending the sweep; `-abi` is where the ABI check
lives, fuzzing each function with sentinel values in the registers the Go ABI
fixes across calls and a canary below `SP`, and reporting a violation on any
iteration. `gasm verify --fuzz` is the differential campaign instead: it
JIT-loads the kernel and the `go tool asm` build of the same kernel and
compares the output argument areas bit-for-bit, one child process per function
so a crash on a partial function is reported rather than fatal. When a fuzz
iteration crashes or mismatches, `FuzzResult.CrashInput` stores the exact input
for reproducibility. `gasm verify --call <func> --buf name:size:pattern`
invokes a single function with user-supplied buffers (patterns: zero, ones,
@@ -449,16 +468,25 @@ masked), reporting any encoding drift.
The interactive debugger (all four architectures). It launches the target
function in a child process that maps the JIT code, calls
`PTRACE_TRACEME`, and stops; the parent attaches via ptrace and controls
execution. Breakpoints are patched as INT3 bytes through `/proc/pid/mem`
(PTRACE_PEEKTEXT is unreliable with Go's multi-threaded runtime).
execution. Breakpoints are patched through `/proc/pid/mem`: the one-byte
`INT3` on amd64, the four-byte break instruction on the other three (arm64
`BRK #0`, riscv64 `ebreak`, loong64 `break 0`).
The child pins its goroutine to the OS thread with `runtime.LockOSThread`
so the traced thread is the one executing JIT code. The REPL provides
single-step, register inspection (GPR + YMM/XMM via `PTRACE_GETFPREGS`),
single-step, register inspection (the GPRs on every architecture; on amd64 the
XMM set through `PTRACE_GETFPREGS` and the YMM set through `PTRACE_GETREGSET`
on `NT_X86_XSTATE`; on the other three the FP/SIMD regset through
`PTRACE_GETREGSET` on `NT_PRFPREG`),
label resolution, named buffer allocation with pattern filling
(`--buf name:size:pattern`: zero, ones, seq, or hex), and breakpoint
management. Breakpoints accept conditions
(`break <label> if <reg> <op> <val>`, including register-against-register
comparisons), and hardware watchpoints work on all four architectures.
comparisons), and hardware watchpoints work on amd64 (the DR0-DR3 debug
registers), arm64 (`NT_ARM_HW_WATCH`) and loong64 (`NT_LOONGARCH_HW_WATCH`);
riscv64 reports that its kernel ptrace interface exposes no trigger regset.
The ptrace path is validated at run time on amd64, where the session tests are
built; arm64, riscv64 and loong64 compile and are covered by the
architecture-neutral units (label and line tables, the breakpoint manager).
For non-interactive use, `--script` runs REPL commands from a file (or
stdin) and exits, `--timeout` kills the debuggee when a run hangs (the
watchdog is armed before the ptrace attach, so a sandboxed debuggee cannot
@@ -499,9 +527,10 @@ sequenceDiagram
Errors are produced where the parse or the encoding fails and become values at
the CLI boundary: the parser returns a diagnostic list and never aborts a file,
`AssembleFile` returns an error, and `cmd/gasm` prints what it has to stderr
and returns a non-zero exit code. The formatter and the linter take the same
AST by a different route: `gasm fmt` re-spaces the token stream and `gasm lint`
walks the parsed file, so neither depends on an encoding.
and returns a non-zero exit code. The formatter and the linter take different
inputs from the assembler: `gasm fmt` re-spaces the token stream
(`format.Source` lexes the source text itself) and `gasm lint` walks the parsed
AST, so neither depends on an encoding.
## State and lifetime
@@ -522,9 +551,8 @@ walks the parsed file, so neither depends on an encoding.
## Dependencies
- **`golang.org/x/arch`** (v0.30.0) is the one module dependency: it is the
disassembler backend (`gasm dis` and the debugger's listings) and the source
of the register metadata the encoder consults (`asm/reg.go`, `asm/vex.go`).
The tests additionally decode through it to validate the encodings.
disassembler backend (`gasm dis` and the debugger's listings). The tests
additionally decode through it to validate the encodings.
- **The Go toolchain**, as an oracle and never as a library: `go tool asm`
supplies the object preamble and the ground truth for `gasm verify
--ground-truth`, `go list -json -export` locates the archives of the packages