docs: remove completed roadmap phases, fix licence description

This commit is contained in:
2026-08-20 15:01:57 +02:00
parent 7629963cab
commit 48334c4d5a
+8 -239
View File
@@ -2,8 +2,6 @@
Developer tooling for **GAsm** — Go's built-in Plan 9 assembler.
[sourcedock.dev/petrbalvin/gasm-devkit](https://sourcedock.dev/petrbalvin/gasm-devkit)
Go ships an assembler but no tooling for it. There is no syntax highlighting,
no autocomplete, no linter, no static analyser, no formatter, no standalone
assembler and no debugger for `.s` files. Developers write assembly blind,
@@ -18,27 +16,13 @@ gasm parse parse and report syntax errors
gasm fmt canonicalise formatting (gofmt for assembly)
gasm lint static checks
gasm lsp language server (completion, hover, symbols, diagnostics, highlighting)
gasm asm standalone assembler (Phase 2)
gasm verify dynamic analysis & verification (Phase 3)
gasm debug source-level debugger (Phase 4)
gasm asm standalone assembler
gasm verify dynamic analysis & verification
gasm debug source-level debugger
gasm diff compare machine code of two .s files
gasm profile show basic-block structure of functions
```
> **Status: Phase 5 — done.** Phase 1 (the language foundation, linter,
> formatter and language server) shipped in v0.1.0; Phase 2 (the standalone
> assembler — the full amd64 instruction set plus ELF and GOOBJ object
> emission) in v0.12.0; Phase 3 (dynamic analysis — JIT execution,
> differential testing, ABI checks and coverage profiling) in v0.25.0;
> Phase 4 (interactive debugger — ptrace-based, breakpoints, watchpoints,
> stepping, vector register display, named buffer allocation) in v0.27.0;
> RISC-V encoder (RV64IMAFDC + RVC, ELF emission, ground-truth, GOOBJ) in
> v0.28.0–v0.29.0; LoongArch encoder (the full instruction set with the MOV
> expansions, ELF and GOOBJ emission, and ground-truth verification) after
> v0.29.0; arm64 encoder (the full integer instruction set with the MOV
> expansions, bitmask immediates, ELF and GOOBJ emission, and ground-truth
> verification) completing Phase 5. See [Roadmap](#roadmap).
## Architecture support
gasm-devkit targets every architecture Go's assembler speaks. The instruction
@@ -68,222 +52,6 @@ cross-compiles the same four targets.
**FreeBSD support is planned for a future release.**
## Roadmap
The work is delivered in four phases. Each phase is completed and hardened
before the next begins. The ordering follows a dependency chain: understand
the code statically (Phase 1), make it runnable (Phase 2), then run it and
observe or control it (Phases 3–4).
### Phase 1 — language foundation, editor tooling and static analysis · *done*
Everything needed to read, understand, check, format and highlight GAsm —
without executing it.
| Capability | Status |
|------------|--------|
| Lexer — permissive, position-aware scanner for all four architectures | done |
| Parser — line-oriented, error-tolerant, full AST with source positions | done |
| Instruction + register tables for amd64, arm64, riscv64, loong64 (generated, complete) | done |
| Linter — `unknown-instruction`, `operand-count`, `undefined-label`, `duplicate-label`, `missing-ret`, `missing-textflag-include`, `abi-argsize`, `unreachable-code`, `register-clobber`, `funcdata-pcdata` | done |
| Formatter — idempotent, comment-preserving, per-function alignment; a `RET` terminates the body for indentation, so the next function's doc comment stays at column 0; exactly one blank line before every block (label, `TEXT`, `GLOBL`) and runs of blanks collapsed; directory / no-argument mode reformats every `.s` in place, `go fmt`-style | done |
| Language server — completion, hover, document symbols, diagnostics, semantic-token highlighting | done |
| CLI — `gasm tokens / parse / fmt / lint / lsp` | done |
| Real-world validation against production AVX2 / AVX-512 kernels | done |
| Lint hardening — zero false positives across the Go runtime corpus (90 files, all four architectures): macro-invocation handling, branch aliases (`B`/`BL`/`JAL`), addressing suffixes (`.P`/`.W`), terminal `UNDEF` | done |
| Static analysis — `abi-argsize` (argument/result area computed from the `// func` signature under Go's ABI0 layout and checked against the TEXT declaration) and `unreachable-code` (dead code after `RET`, suppressed where reachability is undecidable: PC-relative jumps, register-indirect branches, `#ifdef`) | done |
| Static analysis — register liveness (CFG construction + per-instruction def/use + iterative backward dataflow) driving `register-clobber`, calibrated to the **Go ABI** (not System V): flags writes to the registers Go fixes across calls — the frame pointer and the goroutine pointer (`R14` on amd64, `R28`/`R29` on arm64, `X27` on riscv64, `R22` on loong64, plus the OS-reserved `R18` on arm64) — that are never saved/restored; the goroutine pointer is reported only when the function can reach the runtime (not `NOSPLIT`, or makes calls), matching how the runtime's own assembly uses it. `funcdata-pcdata` structural validation of `FUNCDATA`/`PCDATA` operands and indices | done |
> **Limitation — macros.** gasm-devkit reads `.s` source as written; it does
> **not** run the C preprocessor, so `#define` macros are not expanded. Files
> that use macros (the runtime's `asm_*.s`, `race_*.s`, `sys_*.s`, …) parse
> cleanly, and macro *invocations* are recognised and never flagged, but the
> `undefined-label` and `missing-ret` heuristics are suppressed in macro-using
> files because labels a macro defines are invisible without expansion. Full
> macro expansion is future work (it pairs naturally with the Phase 2
> assembler). Hand-written, macro-free kernels — such as everything in
> `go-libraries` — are analysed in full.
### Phase 2 — standalone assembler · *done*
Assembly without the Go toolchain in the loop.
- **`gasm asm`:** a standalone assembler that turns a `.s` file into machine
code directly — pure Go, no `go build`, no external toolchain. Useful for
fast iteration, for environments without a full Go installation, and as the
execution substrate that Phases 3 and 4 build on.
Done so far:
- An amd64 (x86-64) **instruction encoder** — REX/ModR-M/SIB/displacement/
immediate machinery and the scalar instruction set (MOV, the ALU group, TEST,
LEA, INC/DEC/NEG/NOT, shifts, IMUL and IMUL3, PUSH/POP, JMP/CALL/Jcc,
CMOVcc, SETcc, LZCNT/TZCNT, the sign/zero-extending moves — MOVBLZX and
friends, MOVLQSX — and CVTSL2SD/CVTSQ2SD), validated by round-tripping
every encoding through `golang.org/x/arch`'s decoder and byte-for-byte
against the Go assembler.
- An **assembler** that drives the parser's AST into the encoder with local-
label resolution — jumps start in the short (rel8) form and expand to rel32
when the displacement does not fit, and jump-to-jump chains are folded the
way the Go toolchain folds them — so `gasm asm <file>` emits machine code
for each `TEXT` function.
- **File-level assembly with static data** — `GLOBL`/`DATA` symbols are laid
out in a data section behind the code and references to them (`mask<>(SB)`)
are encoded RIP-relative with the displacement resolved within the image,
so the output is self-consistent and position-independent. References to
symbols no `GLOBL` in the file defines are recorded as relocations and
carried into the object-file output.
- **GOOBJ emission** — `gasm asm --format goobj -p <pkgpath>` writes the Go
toolchain's own object format (the one `cmd/link` consumes directly), so
gasm-assembled kernels drop into a `go build` without the Go assembler:
the functions as non-package symbols, `GLOBL` data, one `FuncInfo` per
function and the pc-value tables (`pcsp` with the real prologue/epilogue
stack deltas, `pcfile`, `pcline`, `pcinline`). Verified end-to-end by
swapping a gasm-emitted object into a `go build` in place of the
toolchain's, linking and running — bit-identical behaviour.
- **Object-file emission** — `gasm asm --format elf` writes a relocatable
object (a `.text` and a `.data` section, a symbol table — file-local `<>`
symbols local, the rest global — and one `R_X86_64_PC32` relocation per
static-symbol reference) that links with the system toolchain: external
references resolve against undefined symbols, file-local ones against the
data section. Verified end-to-end by linking a gasm-emitted object with
a C driver and running it. RISC-V uses the equivalent `R_RISCV_PCREL_HI20`
/ `R_RISCV_PCREL_LO12_I` pair for AUIPC+JAL/JALR sequences.
- **`FP`/`SP` frame mapping** — the pseudo-registers are translated onto the
hardware stack pointer (`x+N(FP)` → `(N+8)(SP)` for a zero frame, `(N+frame+
16)(SP)` with a frame pointer; locals via `x-N(SP)`), and the Go-style
prologue/epilogue is generated for functions with a frame. The output is
**byte-identical to the Go assembler** for these cases (verified against
`go tool objdump`).
- **SIMD (VEX / AVX2)** — the VEX prefix machinery (2-byte C5 and 3-byte C4)
with XMM/YMM vector registers, validated by round-trip decoding **and**
byte-for-byte against the Go assembler's machine code, across eight operand
forms: the three-operand NDS form (VPADDD/Q, VPSUBD/Q, VPXOR, VPOR, VPAND/N,
VPCMPEQD, VPCMPGTQ, VPUNPCK*, VPMULLD, VPMULDQ, VPSHUFB, VPACKSSDW,
VPERMD), the two-operand reg/rm form (VPMOVSXWD/DQ, VPMOVZXDQ,
VPBROADCASTD/Q, VPMOVMSKB, VMOVMSKPS, VCVTDQ2PD), the immediate-shift and
variable-count shifts (VPSLLD/Q, VPSRAD, VPSRLD/Q with an immediate or an
XMM/memory count), the immediate shuffle (VPSHUFD, VPERMQ), the
three-operand-plus-immediate form (VSHUFPD, VPERM2I128, VINSERTI128), the
lane extract (VEXTRACTI128, VEXTRACTF128), the direction-sensitive moves
(VMOVDQU, VMOVUPD, VMOVD, VMOVQ, VMOVSD), the no-operand VZEROUPPER, and
the floating-point set: the packed double arithmetic
(VADDPD/VSUBPD/VMULPD/VDIVPD/VMINPD/VMAXPD), the unpacks
(VUNPCKHPD/VUNPCKLPD), the scalar SD and SS operations, VMOVDDUP, the
width-changing conversions (VCVTDQ2PS, VCVTPS2PD, VCVTDQ2PD and the
VCVTPD2DQX/Y / VCVTTPD2DQX/Y spellings, whose VEX.L follows the wider
source) and VFMADD231PD.
- **SIMD (EVEX / AVX-512)** — the four-byte EVEX prefix with the 5-bit
register fields (Z0–Z31, X/Y 16–31), opmask registers (K0–K7 as operands
and mask destinations, KMOVW, KTESTW) and the compressed disp8×N
displacement, covering every AVX-512 instruction the go-flac kernels use:
VPXORD/Q, VPADDD, VPSUBD/Q, VPUNPCK*DQ, VPMULLD/Q, VPERMD, VPSLLD/VPSRAD/
VPSRAQ, VALIGND, VPCMPEQD (with a K destination), VMOVDQU32, VMOVUPD,
VCVTQQ2PD, VPMOVSXDQ, the narrowing stores VPMOVDW/VPMOVQD, the lane
extracts VEXTRACTI64X4/VEXTRACTF64X4, VFMADD231PD, VADDPD, VMULPD,
VMOVDQU64 and the broadcasts VPBROADCASTD/Q from a GPR or memory, plus the
wider AVX-512 F/BW integer set (VPADDB/W, VPSUBB/W, VPANDD/Q/ND/NQ, VPMULLW,
VPMIN*/VPMAX* for B/W/D/Q elements, signed and unsigned, VPAVGB/W, the variable
shifts VPSLLV*/VPSRLV*/VPSRAV*, VMOVDQU8/16), the common floating-point
and conversion set (the packed double and single arithmetic
VADD/VSUB/VMUL/VDIV/VMIN/VMAX PD and PS, the scalar SD/SS operations —
whose EVEX forms exist for masked and zeroing use — the VUNPCK{L,H}PD
unpacks, VMOVDDUP, VMOVSLDUP/VMOVSHDUP and the VCVT* conversions), and
the wider AVX-512 set: ternary logic (VPTERNLOGD/Q), lane shuffles,
inserts and extracts (VSHUF{F,I}{32,64}X{2,4}, the VINSERT*/VEXTRACT*
{F,I}{32,64}X{2,4,8} family, VPALIGNR), compares with an opmask
destination (VCMPPD/PS/SD/SS), the permutes (VPERMB/W, VPERMI2/T2
D/Q/PD), the wider integer families (VPMADDWD/UBSW, VPMULHUW, VPACK*,
VPABS*, the VPROL*/VPROR* rotates and the word shifts), expand/compress
(VEXPAND*/VCOMPRESS*, VPEXPAND*/VPCOMPRESS*), the broadcasts
(VPBROADCASTB/W, VBROADCASTSS/SD), the opmask instructions (KAND/KOR/
KXNOR/KADD/KUNPCK/KNOT/KSHIFTL/KORTEST, KMOVQ), the aligned moves
(VMOVAPS/APD, VMOVDQA32/64, VMOVSS) and the remaining extending and
narrowing moves, the floating-point helper and conversion tail
(VRCP14*, VRSQRT14*, VGETEXP*, VGETMANT*, VSCALEF*, VRNDSCALE*,
VREDUCE*, VFIXUPIMM*, VRANGE*, VFPCLASS* with a K destination, and the
VCVT* conversions VCVTQQ2PS, VCVTPD2QQ/UQQ, VCVTPS2QQ, VCVTUDQ2PD/PS,
VCVTPH2PS, VCVTPS2PH), and gather/scatter with VSIB addressing
(VGATHER*/VPGATHER* in both the VEX mask-register spelling and the EVEX
K-mask spelling — where the L'L field follows the VSIB index — plus
VSCATTER*/VPSCATTER*). The EVEX mnemonic suffixes the Go assembler
accepts are honoured: rounding modes (.RN_SAE, .RD_SAE, .RU_SAE,
.RZ_SAE), suppress-all-exceptions (.SAE) and memory broadcast (.BCST,
with the element-sized disp8×N), each combinable with the .Z zeroing
suffix. Masking is supported the way
Go writes it — an explicit K1–K7 operand placed among the operands, and a
`.Z` mnemonic suffix for zeroing.
- **Legacy SSE moves** — `MOVOU`/`MOVO` (the Plan 9 names for MOVDQU/MOVDQA),
`MOVUPS`/`MOVAPS`/`MOVUPD`/`MOVAPD` and the scalar `MOVSD`/`MOVSS`.
- **Both go-flac kernels — all 17 AVX2 and all 10 AVX-512 functions —
assemble byte-identically to the Go toolchain's machine code**; the only
differing bytes are the displacements of the static-constant loads, which
the Go linker fills at link time and gasm resolves within its own image
(verified to reach the right constant bytes).
Remaining for Phase 2:
- External (cross-package) symbol references in the GOOBJ output —
**deferred** with a recorded decision and three options; see
[`docs/decisions.md`](docs/decisions.md). Single-package objects (no
cross-package references) work today, which covers the production
kernels. With that item deferred, the amd64 instruction set — scalar,
VEX/AVX2 and the full EVEX/AVX-512 set including GPR-interchanging
conversions — is complete, and RISC-V encoding (RV64IMAFDC + RVC)
including ELF and GOOBJ emission is complete.
### Phase 3 — dynamic analysis · *done*
Run the code and check what static analysis cannot. The oracle is the
portable Go implementation every kernel is derived from.
- **`gasm verify`:**
- **JIT execution substrate** — *done.* Assemble the kernel, map it into
executable memory (`syscall.Mmap`, W^X) and call it through an ABI0
trampoline; pure Go, no cgo, no external toolchain.
- **Differential testing** — *done.* The JIT-assembled kernel is fuzzed
against a portable Go reference, comparing the result bit-for-bit;
the automated form of the project's bit-identical contract.
- **Runtime ABI checks** — *done.* The ABI-checking trampoline sets
sentinels in BP and R14, verifies they survive the call, and fills a
128-byte red-zone canary below SP.
- **Coverage / basic-block profiling** — *done.* Static block enumeration
from the assembler's label map plus multi-input path-diversity
measurement: how many observationally distinct execution paths a
test corpus exercises.
### Phase 4 — debugger · *done*
- **`gasm debug`:** single-step a GAsm function, inspect registers (including
YMM vector registers), set breakpoints and watchpoints on addresses, write
memory, allocate and fill named buffers, disassemble at PC, and trace the
source-line mapping — the interactive counterpart to Phase 3's execution
substrate.
- ptrace-based debuggee subprocess (PTRACE_TRACEME + LockOSThread), entry
breakpoint (auto-run to function start), single-step, register inspection
(GPR + YMM/XMM via PTRACE_GETFPREGS), label resolution, breakpoint
management via `/proc/pid/mem`, named buffer allocation with pattern
filling (`--buf`), interactive REPL with conditional breakpoints, four
hardware watchpoints (DR0–DR3), step-over-CALL, run-to-return, backtrace,
memory read/write, disassembly at PC (x86asm), and source-line ↔ offset
mapping.
### Phase 5 — the other architectures · *done*
- **RISC-V encoding — done.** RV64IMAFDC instruction set, RVC compression,
MOV pseudo-instruction, SB/global symbols (AUIPC pairs), ELF64 and GOOBJ
emission, and ground-truth verification against `go tool asm`.
- **LoongArch encoding — done.** The LoongArch64 instruction set with the
MOV pseudo-instruction and its immediate-constant expansions, FP/SP frame
handling, SB/global symbol references (pcalau12i pairs), ELF64 and GOOBJ
emission, and ground-truth verification against `go tool asm` — the emitted
GOOBJ links into a real `go build` for `GOARCH=loong64`.
- **arm64 encoding — done.** The AArch64 integer instruction set with the
MOV pseudo-instruction and its immediate-constant expansions (MOVZ/MOVN/MOVK
and logical bitmask immediates), FP/SP frame handling, SB/global symbol
references (ADRP+ADD pairs), jump chain folding, ELF64 and GOOBJ emission,
and ground-truth verification against `go tool asm`.
## Principles
- **Pure Go and GAsm only.** No C, no cgo, no external toolchains, no native
@@ -316,8 +84,8 @@ portable Go implementation every kernel is derived from.
| `lint` | Conservative static checks. |
| `format` | A canonical formatter — `gofmt` for assembly. |
| `asm` | The standalone assembler: amd64, RISC-V and LoongArch encoders, linker, object-file emitters (ELF, GOOBJ). |
| `verify` | JIT execution substrate for dynamic analysis, combined ABI+fuzz differential testing (Phase 3). |
| `debug` | Interactive ptrace debugger with GPR/YMM register display and named buffer allocation (Phase 4). |
| `verify` | JIT execution substrate for dynamic analysis, combined ABI+fuzz differential testing. |
| `debug` | Interactive ptrace debugger with GPR/YMM register display and named buffer allocation. |
| `lsp` | Language Server Protocol server. |
| `cmd/gasm` | The `gasm` binary tying it all together. |
| `_gen` | The generator that rebuilds the instruction tables from the Go toolchain. |
@@ -370,6 +138,7 @@ binary and associate it with `.s` files. Syntax highlighting is delivered as
infers the target architecture from the file-name suffix
(`_amd64.s` / `_arm64.s` / `_riscv64.s` / `_loong64.s`).
## Licence
## License
BSD-3-Clause — the same licence as Go itself. See [`LICENSE`](LICENSE).
BSD-3-Clause — see [LICENSE](LICENSE).
Copyright © 2026 [Petr Balvín](https://petrbalvin.org)