Files
gasm-sdk/CHANGELOG.md
T
2026-09-21 19:20:01 +02:00

80 KiB
Raw Blame History

Changelog

All notable changes to gasm-devkit are documented here.

The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.

[development]

Added

  • GOOBJ format specification. docs/GOOBJ.md documents the Go object file format in full: both containers, the 96 byte header and all 19 blocks, every structure with its byte offsets, symbol kinds and flag bits, all 106 relocation types with the weak variants, aux symbols, the FuncInfo payload, the pc-value table encoding, the content hashes and the builtin table, all verified byte for byte against objects produced by Go 1.27.1's own tools.
  • Macro expansion and include splicing. gasm asm, gasm diff and gasm audit-instructions now preprocess assembly the way the toolchain does: object and parameterised #define macros expand at the point of use, #undef and the #ifdef/#ifndef/#else/ #endif family select branches, #include splices headers resolved through the source directory and the new repeatable -I flag, ; separates statements, and constant expressions left in operands ($(32-7), $~63, (index*4)(base)) fold at parse. Expansion happens only on the assembly path: gasm lint, gasm fmt and the language server keep reading the raw file.
  • Encoder coverage: the instruction families GOROOT's real code uses. The encoder now covers the instruction families GOROOT's real code uses that gasm lacked, byte-verified against go tool asm: on amd64 the carry ALU, the atomics (CMPXCHG, XADD, XCHG), AES-NI, SHA-1/256, PCLMULQDQ, CRC32, GFNI, ADX, BMI, the string primitives, the system set (CPUID, RDTSC, SYSCALL, fences, MXCSR) and the SSE/AVX/EVEX gaps; on arm64 the pair loads and stores (LDP/STP), acquire/release and LSE atomics, AES and SHA, the system operations, the bit ops and the NEON slice including structure loads and the literal-pool moves; on riscv64 the RV64A AMO family with aq/rl ordering, the Zbb pseudos with their RVC compressions, the FMA forms and the RVV slice with vsetvli/ vsetivli; on loong64 the AM atomics with acquire/release forms, the LSX/LASX slice, the VMOVQ/XVMOVQ transfer family and FSEL. Also fixed on the way: arm64 CASD/CASW lacked an opcode bit, and riscv64 VSETVLI with an immediate length now canonicalises to vsetivli as the toolchain does.
  • Encoder coverage: quad-register AVX-512 and floating-point immediates. The encoder gains the quad-register AVX-512 families (4FMAPS, 4FNMADD, 4VNNIW, VP4DPWSSD, VP4DPWSSDS) with the register list riding the inverted V'VVVV field, floating-point immediates on the SSE scalar moves and arithmetic (the constant lands in a synthesised read-only pool, a positive zero collapses to XORPS exactly as the toolchain does), accept-and-ignore FUNCDATA and PCDATA, three-operand double shifts, static-symbol operands for the legacy SSE moves, and the pooled 64-bit immediate materialisation on riscv64. The parser carries bracketed register ranges, index-only VSIB memory operands and bare trailing immediates; macro substitution reaches parameters used with element suffixes (A.S4), and ; separates statements in plain files.
  • gasm asm -GOOS. The go_asm.h generator type-checks per target GOOS, so the darwin-only and windows-only runtime files assemble with their own defines; the corpus audit derives the GOOS from the file name. DATA initialisers accept $symbol(SB) values (an absolute relocation at the data field, GOOBJ on all four architectures and ELF on amd64, arm64, riscv64 and loong64), and U+2215 is accepted inside symbol package paths.
  • The corpus audit measures honestly. Files named for Go ports gasm does not target (arm, 386, s390x, ...) are no longer attempted for the four supported architectures (no supported build compiles them), and the headline rate is reported over attemptable files: 136 of 433 on the full corpus (31.4 %), 135 of 383 on real code (35.2 %), from the 127 that the previous release measured. The probe battery that decides encodability gained the operand shapes the new families use.

[0.34.0] - 2026-09-20

Added

  • Indirect JMP and CALL on all four architectures. JMP AX, CALL AX, JMP (BX) and the memory forms encode at byte parity with the toolchain (FF /2 and FF /4 on amd64); arm64 lowers JMP (R0) to BR and accepts the raw BR/BLR spellings; riscv64 lowers JMP (X5) to JALR; loong64 accepts the raw JIRL rd, rj, off spelling the Go assembler cannot express. A frameless amd64 function containing a CALL now receives the toolchain's forced base-pointer frame. The riscv64 and loong64 verify trampolines join their ground-truth lists, and a lint check for control flow through registers and memory extends to the new forms.
  • gasm asm -GOARCH and gasm diff -GOARCH. The target architecture can be named explicitly instead of inferred from the file-name suffix, which is how the suffix-less majority of GOROOT's .s files (cpu_x86.s, stub.s, ...) become assemblable.
  • gasm audit-instructions --corpus [dir]. Assembles every .s file under a directory (default GOROOT/src) with the gasm encoder only: suffixed files for their architecture, suffix-less files for all four, as a GOARCH build would. Reports the headline number (127 of 627 GOROOT files, 20.3 %, assemble for every target architecture), the per-architecture pass rates and the most common failure reasons with a representative file each, which drive the encodability backlog by frequency.
  • The parser and the formatter are fuzzed. Two targets carry the guarantee: no input makes the parser panic, and every input yields a file the rest of the toolkit can work on; formatting twice equals formatting once, and clean input stays clean. They seed from the repository's own kernels, and just fuzz drives the mutation engine on demand.
  • Man pages. docs/man carries gasm(1) and a page for every command except version, which gasm(1) documents itself, written in roff: synopsis, description, every flag with its default, exit status, worked examples and cross-references. just install-man compresses them into ~/.local/share/man (MANDIR overrides) and just uninstall-man removes them. A test builds the binary and compares each page's flags and synopsis with its own -h output, so the pages cannot drift from the CLI.

Changed

  • Go 1.27.1 required. The module declares go 1.27.1, so building from source needs that patch release or newer.
  • Canonical just recipes. just gates is the definition of done (build, fmt-check, vet, test, race). install now builds and copies the binary into ~/.local/bin (BINDIR overrides) instead of downloading module dependencies, and install-bin is gone. The test gate sweeps the logic packages (arch through verify; the ptrace-bound debug and the thin cmd/gasm sit outside the coverage profile), so the coverage floor is computed over the product code and the number is identical locally and in CI; the two excluded packages' own tests run in the gate and in the pipelines, outside the floor. fuzz requires its target package.
  • The reported version comes from the build. gasm --version prints the version the toolchain recorded: the tag on a tagged checkout, a pseudo-version naming the commit below one, +dirty on a dirty tree and (devel) outside version control. Nothing is injected with -ldflags -X any more.
  • CI realigned with the gate set. The push pipeline runs the gates minus race in one job, in the gates order, with a cached Go setup and the module as the version source; a superseded run of the same branch is cancelled instead of queueing; every go test runs under a ten-minute bound that matches its job's; the race detector moved to a hand-dispatched workflow and runs in the local gate before a tag is cut, never on a push or a tag; the release builds without injection and its smoke test requires the recorded tag and rejects +dirty.
  • The documents follow the standard set. docs/ARCHITECTURE.md is organised as Overview, Packages, Data flow, State and lifetime and Dependencies, and carries a sequence diagram of the assembly path; docs/DEVELOPMENT.md lists every recipe in one table and documents the coverage floor, the CI and the release flow; docs/CLI.md gives the synopsis, the commands, every flag with its default, the exit codes and worked examples; CONTRIBUTING.md carries the Contributor terms and states the commit trailer form, the one-logical-change rule and the licence header rule; SECURITY.md states how a vulnerability is reported and what to expect. The repository's own assembly (the verify trampolines and the test kernels) is in gasm fmt canonical form.
  • The README states the project's purpose and status. It opens with a warning that the tool is an experiment under active development, version 0.x.x, free to change without warning, with 1.0.0 far off, and already in active use on real assembly work. It describes both goals (tooling for Plan 9 assembly, and Plan 9 assembly outside the Go toolchain), argues the case for the syntax in a new Why Plan 9 assembly section, and carries a Direction section: extended instruction support, full GOOBJ and ELF compilation, Linux and FreeBSD, and the four architectures. A Validation status section states what has been executed where: amd64 on real hardware, the other three architectures under qemu-user emulation, the encoding parity on the host for all four, and the debugger's ptrace path on amd64 only.

Fixed

  • The corpus audit attempts fewer files that no build would compile. Files named for Go ports gasm does not target (arm, 386, s390x, ...) are reported as other-port and never attempted, the headline rate is computed over attemptable files, and the audit searches the toolchain's shipped headers (funcdata.h and friends) automatically.
  • riscv64 JALR silently jumped to the wrong register. The trampoline form JALR X0, 0(X5) read the memory operand's base as the destination, encoding a jump to X0 with no diagnostic; the destination is the first operand. The leaf detection shared the confusion, so affected functions also grew a bogus prologue. JALR X0, 0(X1) as written in the verify trampoline was mis-encoded since its introduction.
  • DATA lines demanded their GLOBL first. collectData processed the declarations in file order, but the Plan 9 convention puts every DATA line before its symbol's GLOBL; correctly ordered files (most of GOROOT's) failed with "no matching GLOBL". Two passes: symbols are registered before initialisers are applied.
  • The formatter lost idempotency on degenerate lines. Illegal tokens survived into the output, a label sharing its line with a non-instruction split into a line the parser rejects, stray-operand lines entered the alignment width computation, and rendered / *, > > sequences re-lexed as comments and shifts. The label, width and spacing rules now agree between passes.
  • gasm fmt deleted the | separators from TEXT and GLOBL flag lists. The lexer had no token for |, the formatter dropped the resulting illegal token, and an in-place format silently rewrote NOSPLIT|DUPOK as NOSPLIT DUPOK, which the Go assembler rejects. The bars now round-trip byte-identically, and ·foo<ABIInternal>(SB) parses its ABI marker instead of swallowing ABIInternal and SB into the flags, which produced false lint warnings on the standard runtime spelling.
  • A malformed TEXT declaration crashed gasm lint and the language server. A TEXT line without a symbol left a nil name that lint and the LSP dereferenced; both now carry on with a diagnostic. A branch to a label at the end of a function body panicked the liveness analysis the same way. A real NUL byte truncated the token stream (everything after it was dropped); it is an illegal token now, invalid UTF-8 no longer inflates byte offsets, and CRLF files format to uniform LF.
  • The class-2 stack guard branched four bytes past its target. When the underflow branch relaxed to its 32-bit form, its displacement was still computed as if the branch were two bytes long, so it landed inside the morestack CALL instead of the compare that decides it. The long form is reachable once a large frame carries a body of roughly a hundred bytes.
  • Immediate operands wrapped silently on amd64. Shift counts, immediates beyond the operand's width and displacements beyond int32 truncated without a diagnostic (SHLQ $300 assembled as $44); they are range-checked now, matching go tool asm. EVEX scalar moves (VMOVSS Z1, Z2) accepted forms the toolchain rejects and emitted invalid encodings; PUSHW/POPW emit the 0x66-prefixed forms the toolchain emits; PUSHL is rejected as illegal in 64-bit mode; a bare zero-operand JE reports a diagnostic instead of panicking.
  • The arm64 shift and divide instructions encoded entirely different operations. LSL, LSR, ASR and ROR, immediate and register forms, all encoded as ORR; SDIV/UDIV sat in the wrong opcode space; MADD/MSUB never encoded the accumulate operand and silently read X0 for it. All now match the toolchain byte for byte (new differential kernels cover shifts, divides and multiplies), MADD takes its four operands in the toolchain's order, and shift amounts at or above the operand width are rejected.
  • arm64 multi-chunk immediates corrupted every branch that followed them. The size pass and the emitter disagreed on the expansion of constants with three or more non-zero 16-bit chunks and of MOVW $-1, so later label displacements were computed against the wrong offsets. The size now comes from the encoder itself. Large-frame stack guards (frames from roughly 64 KiB) branched to the wrong morestack entry, and the pcsp and DWARF CFA boundaries for materialised large frames are computed from the real prologue word counts.
  • arm64 immediates and addressing wrapped instead of erroring. Constants beyond the encodable range (ADD $0x100000000) wrapped to zero, memory offsets wrapped at 2^31, exclusive and atomic accesses silently ignored their offsets (LDXR 8(R1) read [R1]), and large register-based offsets were routed through SP instead of the operand's base. All four now either encode correctly or produce diagnostics.
  • riscv64 compressed stores with certain offsets wrote to the wrong address. The C.SD/C.SW/C.FSD immediate pattern dropped one bit, so any register-relative store with offset bit 4 or 5 set targeted a different address than the same-index load beside it. FENCE assembled as fence 0,0 instead of fence iorw, iorw. Branch and jump displacements beyond ±4 KiB / ±1 MiB wrapped silently; they are diagnostics now. The GOROOT width spellings (MOVW 4(SP), X9) compress to their C.LW/C.SW forms exactly as the toolchain lowers them, restoring byte parity for those shapes.
  • loong64 MOVW $c, Fd wrote a general register. The immediate was routed to the GPR of the F register's number (MOVW $2, F4 clobbered argument register R4), and the correct R30 + movgr2fr.w sequence was unreachable. Two-operand BLTU R4, label encoded as beqz (sometimes-taken where the toolchain's form is never-taken), and out-of-range FP constants and BSTRINS/BSTRPICK bit numbers wrapped silently; all are corrected or diagnosed.
  • ELF objects carried wrong relocation records. The amd64 stack-guard TLS load relocated as R_X86_64_PC32 against the null symbol (every non-NOSPLIT object mislinked); arm64 SB references applied HI21 twice instead of the HI21/LO12 pair; riscv64 PCREL_LO12 referenced the target instead of its AUIPC site, which the system linker rejects; riscv64 and loong64 e_flags declared the soft-float ABI, so standard linkers refused the merge.
  • ELF DWARF was unparseable. Eight abbrev-table constants were wrong, the version-5 line header carried DWARF2-shaped tables, the section count omitted .debug_frame (it sat past the section table, invisible to every tool), the CIE hardcoded one architecture's CFA and return-address registers for all four, and no DWARF address was ever relocated: the .rela.debug_info and .rela.debug_line records were computed and then discarded, so every address stayed zero after linking. The tables parse in readelf, the registers are per-architecture, .rela.debug_info, .rela.debug_line and .rela.debug_frame are emitted, and addresses resolve after the link; a data-only file emits a valid object instead of panicking, and the DWARF records the real source path.
  • GOOBJ cross-package references resolved against the wrong object. External package indices were zero-based against a table that reserves zero for the dummy invalid package, and symbol indices ignored the hashed definition blocks between the sections, so a reference into the first external package could bind to whatever object the loader saw first. An end-to-end cross-package link pins the chain. arm64 ADRP pairs now emit the toolchain's single 8-byte relocation (the previous twin 4-byte records were a hard link error), and symbols no longer claim the linkname flag the toolchain reserves for //go:linkname declarations.
  • The arm64 JIT trampolines saved a scratch register as the stack pointer. enterJIT and its checked twin stored R3, a plain caller-saved register on arm64, and restored RSP from it, so the first JIT call would have returned to a garbage stack. The loong64 trampoline hands its leave address through the raw-symbol pattern the arm64 one uses, avoiding the ABI wrapper's prologue.
  • The checked-ABI report flagged legal frames. The red-zone canary sat 64 bytes below the entry stack, so any kernel with a larger declared frame reported "stack below SP written"; the guard now sizes itself from the kernel's frame. The amd64 JIT tests are gated to amd64 hosts (the suite previously SIGILL-crashed on the other three architectures), fuzz signatures wider than the TEXT frame report instead of panicking, --buf specifications are validated strictly (a typo no longer verifies against a zeroed buffer), and ABI0 parameter sizes cover string and complex correctly.
  • amd64 hardware watchpoints never armed. The debug registers were poked at offsets inside user_regs_struct, corrupting five general registers while the REPL reported success; they now use the real u_debugreg window and stop on the watched address. loong64 watch goes through the kernel's HW_WATCH regset (riscv64 reports the kernel's interface as unsupported instead of failing obscurely).
  • gasm debug hung on the first faulting kernel. Genuine SIGSEGV/SIGBUS/SIGFPE/SIGILL stops were discarded as runtime noise and the faulting instruction restarted forever; faults now surface as reported stops. Conditional breakpoints with a false condition resumed mid-instruction, next and finish evaluated traps with stale registers and landed off instruction boundaries, and the breakpoint restore covered one byte of the four-byte traps (arm64 silently skipped the instruction under it); the trap PCs follow the kernel's reporting on every architecture. regs reports YMM from the xstate (a struct-size overrun crashed FP register reads before), V register halves print correctly on arm64, unwatch accepts the architecture's slot range, x <addr> -8 no longer crashes, break conditions accept memory operands, and session scratch directories are cleaned up.
  • The language server died on one malformed frame and corrupted sources on rename. A bad Content-Length or an unparsable JSON body terminated the process instead of answering -32700 and continuing; rename and references covered the stripped name instead of the full ·name token, so renaming produced ·helper minus its last letter; documentHighlight never matched middle-dot symbols. Positions are UTF-16 code units in both directions (astral characters no longer shift columns) and responses always carry an explicit result member.
  • The linter now recognises the g spelling of the goroutine register. MOVD R0, g clobbered R28 on arm64 (and the equivalents on the other architectures) unflagged, and the numeric spellings the linter did track are rejected by the toolchain there, so g was the one spelling that escaped the audit. FUNCDATA and PCDATA literal indices are validated against the ranges the runtime defines.
  • Usage errors exit 2 uniformly. audit-instructions and scaffold argument errors and an unknown asm --format exited 1 (or, for --format without -o, exited 0 silently); the documented exit-2 contract now holds, asm -o no longer prints the hex dump it claimed to replace, and verify --ground-truth works for amd64 kernels on non-amd64 hosts instead of refusing with JIT advice.
  • arm64 store-exclusive instructions read their operands in the toolchain's order. STXR treated the first register as the status register where go tool asm reads it as the data register, so the same source assembled to different code in the two assemblers; the pair forms (STXP, LDXP and their acquire/release variants) are accepted now, in the toolchain spelling.
  • Large arm64 frames matched the toolchain's sequences. A frame beyond the immediate range that is not a movcon constant (roughly 64 KiB and up) made gasm verify report a false mismatch: the toolchain splits the prologue subtraction into two 12-bit immediates and materialises the non-leaf epilogue addition through the temporary register; gasm emits the same sequences and the spadj boundaries follow the real word counts.
  • The width spellings GOROOT uses assemble. MOVLQZX (four uses in runtime/asm_amd64.s), MOVBQSX, MOVWQSX, MOVBLSX, MOVBWSX, MOVBWZX and PMOVMSKB (the bytealg kernels) encode byte-identically with go tool asm, and the linter reports them encodable; a MOVLQZX is the plain 32-bit move, exactly as the toolchain lowers it.
  • verify --ground-truth no longer reports a mismatch for functions whose size is not a multiple of 16. The toolchain pads text symbols to 16-byte boundaries; the comparison now checks the padding is zero instead of comparing it, the same rule the test suite applies.
  • riscv64 accepts the g spelling of the goroutine register, like the other architectures, and the abi kernels use it; every verify kernel is now ground-truth checkable (the numeric X27 spelling the kernels used is one go tool asm rejects).
  • Two more spellings GOROOT uses now assemble. riscv64 FCLASSD (classify a float64 into an integer mask) is encodable, and a displacement written as a product (0*8(X5), the toolchain's own spelling in several kernels) parses instead of being rejected.
  • loong64 JIT execution enabled. The loong64 trampoline is now validated end to end under qemu-user emulation (plain and ABI-checked calls, goroutine-clobber detection), so gasm verify runs the JIT checks on loong64 hosts instead of forcing every loong64 kernel down the ground-truth path. The arm64 and riscv64 trampolines carry the same validation; the arm64 ABI test now seeds its kernel arguments (a zeroed block made the passthrough check meaningless), and the loong64 basic kernel's branch maze terminates on every path so the smoke sweep cannot spin on leftover register values.

[0.33.0] - 2026-09-14

Added

  • Stack-split guards in gasm asm. Every framed function now gets the morestack prologue check and the trailing morestack block (CALL runtime.morestack_noctxt), byte-identical to the toolchain's stacksplit output on all four architectures: the small, medium and big frame classes, auto-NOSPLIT leaves, the materialised constants of large frames (arm64 R27, riscv64 X31, loong64 R30) and the arm64 extrasize rule. Assembled objects are therefore linkable for split functions, not only NOSPLIT leaves.
  • gasm dis. Standalone disassembly through golang.org/x/arch: a .s file is assembled and listed per TEXT function with local labels at their real offsets, or raw bytes from a file or stdin are disassembled linearly (-a selects the architecture). The debugger shares the same decoder instead of carrying its own.
  • gasm fmt -l and -d. Check mode lists files whose formatting differs; diff mode prints a unified diff from the project's own LCS-based differ, with GNU header semantics.
  • LSP cross-file navigation. Go-to-definition and find references fall back from local labels to the TEXT functions of every open document, and rename follows the same cross-file matching.
  • Large frame offsets on riscv64 and arm64. Frame-relative loads and stores beyond the signed 12-bit immediate range materialise the address through the toolchain temp register (riscv64 X31, arm64 R27) instead of silently truncating the offset (riscv64) or rejecting the instruction (arm64); arm64 frame sizes now add the toolchain's extrasize exactly (+8 when the frame leaves an alignment gap, +16 when it is already aligned).
  • Tail calls JMP sym(SB) on all four architectures (amd64 E9, arm64 B, riscv64 JAL X0, loong64 B) with the call relocation.
  • Live oracle-parity tests. Kernel files covering every guard class, large-offset pattern and tail call are assembled by gasm and by the installed go tool asm and compared byte-for-byte on all four architectures, alongside the existing pinned-byte tests.

Fixed

  • The v0.32.0 review findings: the parser rejects malformed TEXT frames and parses signed frame sizes; amd64 frame adjustments above 127 bytes encode with imm32; arm64 and loong64 relocation encodings match the toolchain; the linter guards unnamed TEXT directives and refreshes its textflag table; the LSP recovers from handler panics and decodes client URIs; watchpoint slot state moved into the debug session; dead verify code removed; em and en dashes replaced across sources.
  • gasm asm --format goobj: internal calls to TEXT symbols of the same file resolve on every architecture (the reference check accepted only the amd64 call kind).
  • gasm asm --format elf on loong64: branch relocations now map to R_LARCH_B26 instead of falling into R_LARCH_PCALA_HI20.
  • arm64 large-prologue ADD/SUB use the extended-register encoding the toolchain picks, and the morestack block saves the link register with the toolchain's OR form on loong64.

[0.32.0] - 2026-08-31

Added

  • Multi-architecture debugger. gasm debug carries per-architecture ptrace register access, disassemblers (golang.org/x/arch), register display, FP register views and stop-info handlers for arm64, riscv64 and loong64, and the REPL is arch-neutral. Sessions are runtime-validated on amd64; the other hosts execute through the now-working JIT trampolines, but their ptrace loops have not seen hardware yet.
  • Headless debugging. gasm debug --script runs REPL commands from a file (or stdin) and exits; --timeout kills the debuggee when a run hangs, with the watchdog armed before the ptrace attach.
  • Conditional breakpoints. break <label> if <reg> <op> <val> now also compares two registers (break loop if RAX > RBX), not only a register against an immediate.
  • Instruction-level coverage. gasm debug --cover now sets a breakpoint on every instruction (walked by disassembly length), counts the hits per instruction and reports the executed instructions with their hit counts, with the label coverage derived from the same run. Expect the run to slow to ptrace speed.
  • FP register display on riscv64 and loong64. The debugger regs command shows the 32 FP registers plus fcsr (and fcc on loong64) via PTRACE_GETREGSET.
  • JIT execution trampolines. verify.Call now works on all four architectures via hand-written assembly trampolines (trampoline_{arm64,riscv64,loong64}.s) that save the Go stack, switch to a prepared stack, and branch to the JIT function.
  • ABI checks on arm64, riscv64 and loong64. gasm verify -abi and the ABI half of -fuzz now cover the non-amd64 architectures via per-architecture checked trampolines: sentinels planted in the registers the Go ABI fixes across calls (arm64 R29/R28, riscv64 X27, loong64 R22; amd64 keeps BP/R14) are verified on return and the saved registers restored before Go code resumes, with the below-SP canary on every architecture. loong64 kernels are verified through the toolchain-comparison path only, pending hardware validation of their trampoline. verify.Load assembles each file with the encoder its name suffix calls for. The ABIReport fields are the architecture-neutral FPClobbered and GClobbered.
  • Hardware watchpoints on all architectures. arm64 uses DBGWVR/DBGWCR via PTRACE_SETREGSET with NT_ARM_HW_BREAK; riscv64 and loong64 use PTRACE_POKEUSER to access trigger/debug registers.
  • Cross-package GOOBJ resolution on arm64 and loong64. AssembleFileARM64 and AssembleFileLOONG64 mark external relocations and populate img.Externals, so GOOBJ output from those architectures resolves cross-package symbols like amd64 and riscv64 already do.
  • gasm verify --args. Scalar arguments (name=value, decimal or 0x hex) can now be supplied to a --call invocation alongside --buf buffers, closing the gap where only buffers could be supplied.
  • Fuzz corpus save and replay. gasm verify --fuzz --save-corpus dir records every input that crashes or mismatches as replayable JSON (buffer contents and scalars, not raw pointers), and gasm verify --replay dir re-runs the saved entries against the kernel in isolated child processes, reporting whether each one reproduces.
  • gasm audit-instructions. Black-box diff of a gasm encoder against the installed go tool asm, for amd64, arm64, riscv64 and loong64 (gasm audit-instructions <arch>): superset encodings (gasm-only, shippable via gasm asm --format goobj), known-but-unencodable names (the backlog) and go-only names (feature gaps).
  • gasm scaffold differential. Prints a differential test skeleton for every // func signature in a kernel file: random seed states, the kernel call and a portable reference (<name>Portable), compared byte-for-byte.
  • LSP: find references (textDocument/references).
  • LSP: rename symbol (textDocument/rename).
  • LSP: document formatting (textDocument/formatting) using the format package for canonical gofmt-style output.
  • LSP: inlay hints (textDocument/inlayHint): frame size hints after the TEXT directive's argument area.
  • LSP: workspace symbol search (workspace/symbol): substring search over the TEXT functions and GLOBL/DATA symbols of every open document.
  • LSP: code actions. Quick fixes for missing-ret (insert the RET) and unused-label (remove the label) diagnostics.
  • LSP: signature help (textDocument/signatureHelp): the callee's // func signature while the cursor is on a CALL.
  • LSP: document highlights: every reference to the function or label under the cursor is highlighted.
  • LSP: pull diagnostics (textDocument/diagnostic), #include document links (resolved against the document directory, then $GOROOT/pkg/include, so textflag.h opens) and folding ranges (one collapsible region per TEXT function body).
  • Lint: unused-label rule. Flags labels that are defined but never referenced by any jump (Hint severity).
  • Lint: invalid-textflag rule. Flags TEXT/GLOBL flags not in the known set from textflag.h (Warning severity).
  • Lint: stack-imbalance rule. Tracks SP changes and flags if the net delta at RET does not match the declared frame size.
  • Lint: register-width-mismatch rule. Flags amd64 operands whose register width does not match the width the mnemonic suffix prescribes (for example a 32-bit register in a MOVQ).
  • Lint: abi0-register-args rule. Flags kernels whose // func parameters are never read from the FP frame (for example arguments read from registers instead), which pass every test today and break on a toolchain upgrade. Shipped with the Go assembler's operand-width model.
  • Lint: nonportable-register-name rule. Flags the RAX/EAX-style register aliases gasm accepts but go tool asm rejects, so files using them only link through the gasm GOOBJ path.
  • Lint: unencodable-instruction rule. Flags mnemonics the architecture table knows but the encoder cannot yet emit, at edit time instead of at assembly time.
  • Lint: reserved-register-write rule. Flags writes to arm64 R18, the platform-reserved register the ABI checks cannot observe at runtime and the Go assembler cannot even spell. Reads and macro-using files are exempt.
  • DWARF5 debug sections in ELF output. All four ELF emitters now emit .debug_abbrev, .debug_info, .debug_line, and .debug_line_str sections, enabling addr2line and GDB/LLDB source-level debugging, and the amd64 emitter adds a .debug_frame CFI section for stack unwinding.
  • amd64: legacy SSE and conversion coverage. The encoder now handles the legacy (non-VEX) SSE packed binaries and immediate shuffles, the legacy SSE integer instructions and BSWAP, the scalar and packed double conversions (CVTSS2SD, CVTSD2SS, CVTPS2PD, CVTPD2PS) and the prefetch hints.
  • amd64: VPCMP and the full opmask set. The EVEX compare with an opmask destination and the remaining opmask-register instructions are encoded, byte for byte against the Go assembler.

Changed

  • Parallel verify sweeps. The -smoke and -abi per-function checks now run in parallel instead of sequentially.

Fixed

  • Reliable ptrace sessions. The debugger no longer mistakes runtime signal-delivery-stops (a Go tracee reports SIGURG preemption to the tracer) for its launch barrier, runs every ptrace request on the thread that forked the debuggee (requests from another thread fail with ESRCH), prefers the debuggee's reported code base over an RWX scan, and single-steps over a hit breakpoint so resuming cannot re-trap on the same instruction. A ptrace integration test (debug/ptrace_integration_test.go) drives a real session end to end.
  • arm64 BL relocation. BL sym(SB) now records a RelArm64Branch relocation instead of emitting a bare instruction with no relocation.
  • GOOBJ R_ADDRARM64 constant. Corrected from 9 (R_CALLARM64) to 3 (R_ADDRARM64).
  • Subprocess-isolated smoke and ABI sweeps. A function that faults during the -smoke or -abi sweep is reported without killing the parent; each sweep runs in a child process.
  • BSF, BSR and POPCNT encodings. Corrected to the bytes go tool asm emits.
  • MOVQ immediates. Compressed to the toolchain's imm32 forms.
  • Frame adjustments of 128 to 255 bytes. Now use the imm8 ADDQ stack adjustment.
  • amd64 encoding parity. A broad pass aligned the remaining encoder outputs and operand strictness with go tool asm.
  • asm help text. Updated to list arm64 as a supported architecture.

[0.31.1] - 2026-08-20

Fixed

  • Version stamp. The v0.31.0 release binary reported itself as 0.30.0 because the version variables in justfile and cmd/gasm/main.go were not bumped during the release commit.

[0.31.0] - 2026-08-20

The arm64 encoder ships with ELF64 and GOOBJ emission, verified byte-for-byte against GOARCH=arm64 go tool asm and linked into a real go build. The encoder covers the full integer instruction set, FP arithmetic, conditional select, CRC32, and the MOV pseudo-instruction with bitmask immediate encoding. The project now requires Go 1.27.

Added

  • arm64 encoder. gasm asm can now assemble _arm64.s files: the AArch64 integer instruction set with the MOV pseudo-instruction and its immediate-constant expansions (MOVZ/MOVN/MOVK for wide immediates, ORR with logical bitmask encoding for values like $1), data-processing (shifted register and immediate forms), load/store (scaled unsigned and unscaled9-bit immediate), conditional and unconditional branches, FP/SP frame mapping, SB/global symbol references (ADRP+ADD pairs with R_ADDRARM64 relocations), jump chain folding, and ELF64 emission (gasm asm --format elf). Ground-truth verification against GOARCH=arm64 go tool asm matches byte-for-byte. The encoder set for the remaining architectures (RISC-V, LoongArch, arm64) is complete.

Changed

  • Go 1.27 required. The project now requires Go 1.27 (toolchain go1.27.0). The R_DWTXTADDR_U4 relocation type is detected at runtime for backward compatibility.

[0.30.0] - 2026-08-13

The LoongArch encoder ships with ELF64 and GOOBJ emission, verified byte-for-byte against GOARCH=loong64 go tool asm and linked into a real go build; the shared GOOBJ emitter now writes the per-function DWARF symbols the linker's DWARF pass reads. The RISC-V encoder reaches byte-for-byte parity with go tool asm: the frame model, operand ordering, RVC compression, large-immediate and MOV $imm materialisation, branch/jump encodings, and CALL sym(SB) (now a JAL with an R_RISCV_JAL relocation). The debugger tracks four hardware watchpoint slots, and the toolkit is Linux-only.

Added

  • LoongArch encoder. gasm asm can now assemble _loong64.s files: the full LoongArch64 instruction set with the dual-form arithmetic mnemonics, the 16/21-bit branch families, the MOV pseudo-instruction and its immediate-constant expansions, FP/SP frame mapping, SB/global symbol references (pcalau12i pairs) and ELF64 emission (gasm asm --format elf). Ground-truth verification against GOARCH=loong64 go tool asm matches byte-for-byte; GOOBJ emission (gasm asm --format goobj) is proven end-to-end by linking the object into a cross-compiled go build.
  • GOOBJ DWARF symbols. The GOOBJ emitters now write the per-function DWARF symbols the linker requires (the subprogram DIE and the .debug_line program, byte-identical to cmd/asm's), and the pc-value table deltas are in the architecture's MinLC units as the runtime expects; the amd64 link test now genuinely substitutes the gasm object, and the amd64/loong64 end-to-end GOOBJ link tests pass.
  • RISC-V GOOBJ emission via the shared emitter. RISC-V GOOBJ output is now written by the same shared emitter as amd64 and LoongArch, modelling each AUIPC + second-instruction pair as a single R_RISCV_PCREL_ITYPE/STYPE relocation (the layout cmd/asm writes, not the ELF HI20/LO12 pair), so the object links into a cross-compiled go build for GOARCH=riscv64. An end-to-end link test substitutes the gasm object and reads the symbol back with go tool nm; the rewrite also corrects the relocation after field.

Fixed

  • RISC-V frame model and RVC encodings. The riscv64 frame layout now matches go tool asm: the prologue/epilogue save and restore the link register (LR) instead of S0, with the correct autosize (locals + 8) and the RVC-compressed prologue/epilogue instructions; RET emits the uncompressed JALR X0, 0(X1) the toolchain writes; the C.ADDI/C.LI/C.LUI/C.ADDIW opcode bit and the C.ADD CR-type encoding are fixed; and the LR/TMP register aliases now resolve to X1 and X31. The pcsp/pcfile/pcline tables are populated from the recorded stack-adjustment and source-line data, and a byte-exact ground-truth test compares framed and leaf functions against GOARCH=riscv64 go tool asm.
  • RISC-V operand ordering and RVC compression. R-type instructions now take rs2, rs1, rd and I-type arithmetic instructions take imm12, rs1, rd, matching the Go assembler's documented operand order (previously both were reversed, so non-commutative R-type instructions such as SUB encoded the wrong operation). The two-operand ternary forms (ADD rs2, rd, ADDI $imm, rd, SLLI $shamt, rd) are now accepted. RVC compression is completed for C.ADDI16SP, C.SLLI, C.SRLI, C.SRAI, C.ANDI, C.NOP, C.EBREAK, C.MV (from ADDI/ADD) and the commutative AND/OR/XOR forms; the byte-exact ground-truth test now covers these.
  • RISC-V compressed loads/stores and word arithmetic. RVC compression now also covers the register-relative C.LW/C.SW/C.LD/C.SD/ C.FLD/C.FSD forms (in addition to the stack-relative C.LWSP/C.SWSP/ C.LDSP/C.SDSP), plus C.ADDI4SPN, C.ADDW and C.SUBW. The byte-exact ground-truth test exercises these against GOARCH=riscv64 go tool asm.
  • RISC-V large-immediate materialisation. ADDI/ANDI/ORI/XORI with a 32-bit immediate that does not fit 12 bits now expand exactly as cmd/asm: two ADDIs for the small ADDI split range, and LUI+ADDIW+<op> otherwise, with the LUI and ADDIW compressed to C.LUI/C.ADDIW when their immediate fits six signed bits. The byte-exact ground-truth test covers positive, negative, and out-of-range immediates against GOARCH=riscv64 go tool asm.
  • RISC-V MOV $imm, rd materialisation. The immediate-loading pseudo-instruction now uses the toolchain's Split32BitImmediate split (previously it rounded the upper 20 bits, producing wrong results for negative and bit-11-set immediates) and compresses the emitted ADDI/LUI/ADDIW to C.LI/C.LUI/C.ADDIW when their immediate fits six signed bits. A byte-exact ground-truth test covers zero, small, negative, and 32-bit immediates against GOARCH=riscv64 go tool asm.
  • RISC-V branch/jump compression. JMP/JAL were being compressed to C.J and BEQ/BNE (with X0) to C.BEQZ/C.BNEZ, but go tool asm never emits these compressed forms. They now emit the 32-bit JAL and branch encodings the toolchain writes; the dead C.J/C.BEQZ/C.BNEZ encoders were removed, and the C.LUI direct-instruction compression now uses the correct six-bit signed range. A byte-exact ground-truth test covers the branch family and jumps against GOARCH=riscv64 go tool asm.
  • RISC-V CALL sym(SB). The call pseudo-instruction now emits the toolchain's JAL X1, sym(SB) with a single R_RISCV_JAL relocation (previously it emitted an AUIPC+JALR pair against a local branch label, a form go tool asm rejects). The GOOBJ and ELF emitters now map that relocation (Go objabi 59 / ELF R_RISCV_JAL 17, a 4-byte field), and relocation offsets are recorded relative to the function start (including the prologue). A byte-exact ground-truth test covers a call against GOARCH=riscv64 go tool asm.
  • Debugger watchpoint slots. gasm debug's watch command always used hardware watchpoint slot 0, so a second watch call silently overwrote the first. Watchpoint slots are now tracked in the Session (DR0-DR3); watch picks the first free slot and reports an error if all four are in use, and unwatch <slot> clears one (no argument clears all).

Changed

  • Linux only. The toolkit, its CI and the released binaries are now Linux-only; cross-compiled to linux/{amd64,arm64,riscv64,loong64}.
  • Debugger complete. README's "Remaining" list for the debugger is gone; disassembly at PC, memory-write, watchpoints, and source-line mapping are all shipped.

[0.29.0] - 2026-08-07

RISC-V GOOBJ emission, YMM vector register display, named buffer allocation in the debugger, two new CLI commands (diff, profile), go-to-definition in the LSP, combined ABI+fuzz verification, and did-you-mean label suggestions. A --map flag for diff and --call/--buf flags for verify extend the new CLI commands. A signature-parser fix corrects grouped Go parameters.

Added

  • RISC-V GOOBJ emission; gasm asm --format goobj for RISC-V produces linkable Go objects with funcdata, pc-value tables, and RISC-V relocation types (same format as amd64 GOOBJ, with the RISC-V architecture marker).
  • gasm diff; compare the machine code of two assembly files byte-for-byte; shows which functions differ and the first few differing bytes.
  • gasm profile; show the basic-block structure of each function: labels, offsets, frame size, and NOSPLIT flag.
  • LSP go-to-definition; textDocument/definition navigates from a label reference to its definition.
  • did-you-mean; when the RISC-V assembler encounters an undefined label, it suggests the closest existing label using Levenshtein distance.
  • YMM vector register display; regs in the debugger now shows YMM registers via PTRACE_GETFPREGS (falls back to XMM when XSAVE is unavailable).
  • Named buffer allocation; gasm debug --buf name:size:pattern allocates buffers in the debuggee filled with zero, ones, seq, or a hex pattern; buffer pointers are placed into the argument block at the matching positions.
  • Crash input storage; FuzzResult.CrashInput stores the input that caused a crash or mismatch for reproducibility.
  • ABI + fuzz combined; gasm verify --fuzz now runs ABI checks (sentinel registers, canary, stack bounds) alongside differential fuzz testing.
  • gasm diff --map; compare functions whose names differ between files (e.g. --map wideCopyAVX2=wideCopyAVX512 pairs two variants regardless of suffix). Unmapped functions fall back to the original name match.
  • gasm verify --call; invoke a single function with user-supplied buffers (--buf name:size:pattern) instead of the smoke/abi/fuzz sweeps. Patterns: zero, ones, seq, or a hex blob. Useful for partial functions (e.g. decoders) that crash on random input but should succeed on valid data. The arg block is printed before and after the call, showing return values.
  • gasm verify --ground-truth now documented in --help (was already a flag, just missing from the help text).

Fixed

  • Signature parser; grouped Go parameters like dst, src []byte are now parsed correctly (both get type []byte). Previously the first name was treated as its own type (dst with size 8), causing wrong ABI0 arg-block layout in both verify --call and the fuzzer.
  • Flaky JIT tests; runtime.KeepAlive guards and package-level buffers prevent GC from collecting heap objects whose addresses were passed to JIT code via unsafe.Pointer; all verify tests pass 100/100 under -race.

Changed

  • Removed external kernel test dependencies; the verify test suite no longer references production kernels from the separate go-libraries project. The remaining test suite uses only testdata/verify/*.s kernels, which are part of this repository. Coverage is identical locally and in CI (80.3 %).

[0.28.0] - 2026-08-03

RISC-V encoder: full RV64IMAFDC instruction set with RVC compression, MOV pseudo-instruction, SB/global symbol references, ELF64 object emission, and ground-truth verification against GOARCH=riscv64 go tool asm.

Added

  • RISC-V encoder; RV64I, RV64M, RV64A, RV64F/D, FMA, CSR, JALR.
  • MOV pseudo-instruction; load, store, reg-to-reg, immediate, frame mapping.
  • RVC compression; 22 compressed instruction types (C.LDSP, C.SDSP, C.FLDSP, C.FSDSP, C.ADDI, C.LI, C.LUI, C.ADDIW, C.MV, C.ADD, C.SUB, C.XOR, C.OR, C.AND, C.SLLI, C.SRLI, C.SRAI, C.ANDI, C.BEQZ, C.BNEZ, C.J, C.JR).
  • SB/global symbols; MOV $sym(SB), MOV sym(SB), MOV rd, sym(SB) encoded as AUIPC pairs with R_RISCV_PCREL_HI20/LO12 relocations.
  • GLOBL/DATA; data section layout in AssembleFileRISCV.
  • ELF64 emission; gasm asm --format elf produces EM_RISCV objects (.text, .data, .symtab, .rela.text).
  • gasm verify --ground-truth; byte-exact comparison against GOARCH=riscv64 go tool asm.
  • gasm verify --profile; function layout listing for RISC-V.
  • CALL; AUIPC + JALR pair encoding.

Fixed

  • Parser: bare-number offset before (SP) no longer misidentified as pseudo.
  • MOV: MOV $sym(FP/SP), rd now returns an explicit error instead of silent fallback.
  • RVC: C.LDSP/C.SDSP/FLDSP/FSDSP immediate encoding now matches Go toolchain (bit-interleaved format).

[0.27.0] - 2026-08-01

Subprocess isolation for --fuzz: each function is fuzzed in its own child process, so a partial function (decoder) that faults on random garbage is reported as "CRASH (partial function, use --ground-truth)" without killing the parent. CRASH is informational (exit 0); only MISMATCH is an error.

Fixed

  • gasm verify --fuzz no longer crashes the process on partial functions.

[0.26.0] - 2026-07-31

Universal differential fuzzing: gasm verify --fuzz needs no hand-written reference. It parses the // func signature from the assembly source, generates typed random inputs (slices with random content, ints, pointers to fixed arrays), JIT-executes BOTH the gasm-assembled and the go-tool-asm- assembled versions with independent buffer copies, and compares the result area bit-for-bit.

Added

  • verify: FuzzFunc / ExtractSignatures / parseFuncSig; universal differential fuzz driven by the conventional // func comment. Each version gets its own buffer set (deep copy) so functions that write to their arguments (histogram increments) don't corrupt the other's input.
  • gasm verify --fuzz [-n N]: runs the differential fuzz for every function with a parseable signature. Total functions (wideCopy, pack16, decorrelate, analyze, autocorr) pass; partial functions (decoders that fault on malformed input) should use --ground-truth instead.

Fixed

--fuzz crashes the process for partial functions (e.g. LZ4 decoders) whose over-copy paths read past the buffer on random garbage input. Subprocess isolation (fork per function) is planned. Use --ground-truth for decoders.

[0.25.0] - 2026-07-30

Universal ground-truth verification: gasm verify --ground-truth assembles any .s file with both gasm and go tool asm, then compares the machine code byte-for-byte per function (relocation sites masked). No hand-written reference needed; the Go toolchain IS the oracle.

Added

  • verify: GroundTruth; shells out to go tool asm, parses the GOOBJ output (minimal reader: block offsets, nonpkg symbol table, data index) and returns per-function code bytes.
  • gasm verify --ground-truth: compares gasm's output against the Go assembler's, reporting MATCH/MISMATCH per function with the first differing byte. Relocation disp32 fields (static-symbol references the linker fills) are masked before comparison.
  • Verified: go-lz4 AVX2 2/2, go-flac AVX2 17/17 functions byte-identical.

[0.24.0] - 2026-07-29

The full analyze family and stereo PCM decode are now differentially tested. 15 of 17 go-flac AVX2 kernels have bit-for-bit differential coverage; the two remaining (autocorrAVX2: FMA reassociation, lpcResidualAVX2: complex multi-arg) are deferred.

Added

  • verify: analyzeO3RangeAVX2 and analyzeO4RangeAVX2 differential tests (200 iterations each, same harness as O1/O2/Res).
  • verify: decodeStereo16AVX2 differential test (500 random interleaved stereo PCM buffers, both channels compared sample-by-sample).

[0.23.0] - 2026-07-28

The analyze family and 24-bit PCM decode join the differential suite.

Added

  • verify: analyzeO2RangeAVX2 and analyzeResRangeAVX2 differential tests (200 iterations each, shared harness with O1: zigzag fold, partial sum, overflow flag and Len32 histogram).
  • verify: decodeMono24AVX2 differential test (500 random 24-bit PCM buffers, sign-extension compared sample-by-sample).

[0.22.0] - 2026-07-27

The remaining go-flac encoder kernels join the differential suite.

Added

  • verify: analyzeO1RangeAVX2 differential test (300 random partitions: zigzag fold, partial sum, overflow flag and the 32-bin Len32 histogram compared element-by-element against the portable Go reference).
  • verify: fastStereoSumsAVX2 differential test (300 random stereo frames: the four zigzag-fold entropy sums compared against the scalar loop).

[0.21.0] - 2026-07-26

Differential testing extended to all four production kernels and the CLI exposes the full dynamic-analysis toolkit.

Added

  • verify: go-flac AVX2 differential tests; decodeMono16AVX2 (500 random PCM buffers), pack16AVX2 (500 random int32→int16 packings) and all four decorrelation kernels (200 iterations each: left-side, side-right, mid-side, interleave) compared bit-for-bit against the portable Go references.
  • verify: go-lz4 AVX-512 differential tests; decodeBlockAVX512 (3 000 fuzzed LZ4 blocks + known answers) and wideCopyAVX512 (0-1024 bytes) against the same portable oracle as the AVX2 suite.
  • gasm verify --abi: runs each NOSPLIT function with sentinel registers and a red-zone canary, reporting violations.
  • gasm verify --profile: lists the static basic-block count per function.

[0.20.0] - 2026-07-25

Coverage profiling. Static basic-block enumeration from the assembler's label map, combined with multi-input path diversity measurement; how many observationally distinct execution paths a test corpus exercises.

Added

  • verify: Kernel.Blocks / Kernel.BlockCount; enumerate basic blocks from the assembler's local-label map (every jump target is a block boundary; the function entry is always a block). decodeBlockAVX2 has 27 blocks.
  • verify: Kernel.ProfilePaths; run the function with a corpus of argument blocks and collect distinct output fingerprints (the result words); reports path diversity as a lower bound on code coverage.

Changed

INT3-based per-block hit counting was prototyped but deferred: Go's runtime signal management (sigaltstack, handler re-installation) makes raw rt_sigaction handlers fragile in a Go process. The static + path-diversity approach delivers the project's goal (proving the SIMD path and tail handling execute) without fighting the runtime.

[0.19.0] - 2026-07-24

Runtime ABI checks. The JIT trampoline now has an ABI-checking variant that sets sentinels in the callee-saved registers (BP, R14) before entering the assembled function and verifies they survive on return, plus a red-zone canary (128 bytes below SP filled with 0xA5) that detects any illegal write below the stack pointer.

Added

  • verify: CallChecked / Kernel.CallFuncChecked; ABI-checking JIT call with sentinel registers and red-zone canary; returns an ABIReport (BPClobbered, R14Clobbered, RedZoneHit).
  • verify: the raw leaveJITCheckedRaw trampoline; a TEXT symbol with no ABIInternal wrapper (address obtained via GLOBL/DATA), so the JIT function's RET lands directly in the check code and sees the registers exactly as the function left them.
  • Tests: deliberate BP/R14 clobberers detected; both go-lz4 kernels confirmed ABI-clean (BP preserved, R14 preserved, red zone intact).

[0.18.0] - 2026-07-23

Differential testing: the JIT-assembled go-lz4 decodeBlockAVX2 kernel is fuzzed against a portable Go reference; 5 000 valid LZ4 blocks compared bit-for-bit, plus 2 000 hostile (random garbage) inputs with matching error codes. This is the automated form of the project's bit-identical contract.

Added

  • verify: differential fuzz tests; a random LZ4 block generator produces valid blocks (literals, overlapping matches, extension bytes) and the JIT-assembled kernel's output is compared byte-for-byte against a portable Go decoder; a hostile-input suite confirms error-code agreement on random garbage (no crashes, same classification).

[0.17.0] - 2026-07-22

Dynamic analysis. A JIT execution substrate that assembles Plan 9 amd64 kernels into executable memory and calls them directly; pure Go (stdlib only, syscall.Mmap + an assembly trampoline), no cgo, no external toolchain.

Added

  • verify package: JIT infrastructure; Map copies machine code into a W^X memory mapping, Call invokes it through an ABI0 trampoline that switches to a prepared stack and back. Load/LoadSource/LoadAST parse, assemble and map a .s file in one step; Kernel.CallFunc marshals the argument block and returns results.
  • gasm verify subcommand: assembles a file, JIT-loads it and reports the available functions; with -smoke, calls each NOSPLIT function with zeroed arguments to confirm the trampoline works end-to-end.
  • Integration tests: the go-lz4 decodeBlockAVX2 and wideCopyAVX2 kernels (699 and 146 bytes) assemble, map and execute correctly; known-answer LZ4 blocks decode bit-for-bit, wide copies of 0-1024 bytes match, malformed input returns the correct error codes.

[0.16.0] - 2026-07-21

The scalar conversions between vector and general-purpose registers; the last of the amd64 EVEX instruction set.

Added

  • asm: the GPR-interchanging conversions, byte for byte against the Go assembler (28 ground-truth cases including memory sources and extended GPRs): vector to GPR; the signed and truncated VCVT{,T}S{D,S}2SI{,Q} in both VEX and EVEX, and the unsigned VCVT{,T}S{D,S}2USI{L,Q} (EVEX only); GPR to vector; VCVTSI2SD{L,Q}/VCVTSI2SS{L,Q} (VEX and EVEX) and VCVTUSI2SD{L,Q}/VCVTUSI2SS{L,Q} (EVEX only), whose preserved vector source sits in vvvv (three Plan 9 operands).

[0.15.0] - 2026-07-20

The last of the EVEX conversions and narrowing/extending moves; the EVEX instruction set is now complete save for the GPR-interchanging forms.

Added

  • asm: the unsigned and truncating conversions; VCVTPD2PS (and the X/Y spellings, whose length the spelling fixes), VCVTPD2UDQ (X/Y), VCVTTPD2UDQ (X/Y), VCVTTPD2UQQ, VCVTPS2UDQ, VCVTTPS2UDQ, VCVTPS2UQQ, VCVTTPS2UQQ, VCVTTPD2QQ, VCVTTPS2QQ, VCVTUQQ2PD, VCVTUQQ2PS (X/Y) and VCVTQQ2PS X/Y.
  • asm: the remaining sign/zero-extending moves (VPMOVSXBD/BQ/WQ and VPMOVZXBD/BQ/WD/WQ, VEX and EVEX) and the complete signed and unsigned narrowing stores (VPMOVS{DB,QB,DW,QW,QD,WB}, VPMOVUS{DB,QB,DW,QW,QD,WB}, VPMOVDB, VPMOVQW).
  • asm: the mask/vector conversions (VPMOVM2B/W/D/Q and VPMOVB2M/W2M/ D2M/Q2M), whose K register is a genuine operand rather than a mask and which therefore take no masking suffixes.

[0.14.0] - 2026-07-19

The floating-point helper and conversion tail of the AVX-512 set, plus gather and scatter with VSIB addressing; every encoding verified byte for byte against the Go assembler.

Added

  • asm: the floating-point helpers; reciprocals and reciprocal square roots (VRCP14/VRSQRT14 PD/PS/SD/SS), exponents and mantissas (VGETEXP*, VGETMANT*), scaling by powers of two (VSCALEF*), rounding (VRNDSCALE*), reduction (VREDUCE*), immediate fixup (VFIXUPIMM*) and range selection (VRANGE*), and floating-point class tests (VFPCLASSPD/PS X/Y/Z and VFPCLASSSD/SS; a new immediate form whose reg field carries the opmask destination).
  • asm: gather and scatter with VSIB addressing. The gathers take both Go spellings: the VEX form with a vector mask register (OP mask, vsib, dst) and the EVEX form with an explicit K mask (OP vsib, K, dst), where the EVEX L'L field follows the VSIB index register rather than the data register (a ZMM index with an YMM destination encodes L'L = 10, as the Go assembler emits). The scatters (VSCATTER*/VPSCATTER*) are EVEX only (OP src, K, vsib). All eight gather and eight scatter widths.
  • asm: the remaining conversions; VCVTQQ2PS (the 512-bit source sets the length), VCVTPD2QQ/UQQ, VCVTPS2QQ, VCVTUDQ2PD/PS, the half-precision VCVTPH2PS and VCVTPS2PH (the extract layout with an immediate).

[0.13.0] - 2026-07-18

The wider AVX-512 set: ternary logic, permutes, compares, expand/compress, the opmask instructions and the EVEX rounding/SAE/broadcast suffixes; every encoding verified byte for byte against the Go assembler.

Added

  • asm: the wider EVEX/AVX-512 set, across roughly sixty new ground-truth cases: ternary logic (VPTERNLOGD/Q), the lane shuffles/inserts/extracts (VSHUF{F,I}{32,64}X{2,4}, the VINSERT*/VEXTRACT* {F,I}{32,64}X{2,4,8} family, VPALIGNR), compares with an opmask destination (VCMPPD/PS/SD/SS; a new NDS-plus-immediate form with the K register in the reg field), the permutes (VPERMB/W, VPERMI2/T2 D/Q/PD), the wider integer families (VPMADDWD/UBSW, VPMULHUW, VPACKSSWB/USWB/SSDW/USDW, VPABS B/W/D/Q, the VPROL*/VPROR* rotates, the word shifts and the EVEX W1 qword shifts), expand/compress (VEXPANDPD/PS, VPEXPANDD/Q, VCOMPRESSPD/PS, VPCOMPRESSD/ Q), the broadcasts (VPBROADCASTB/W from a GPR or memory, VBROADCASTSS/ SD), the opmask-register instructions (KAND/KOR/KXNOR/KADD/KUNPCK/KNOT/ KSHIFTL/KORTEST B/W/D/Q and KMOVQ, whose width the L/W/pp bits select), the packed single arithmetic (VADD/VSUB/VMUL/VDIV/VMIN/VMAX PS), the aligned moves (VMOVAPS/APD, VMOVDQA32/64, VMOVSS), the replicating moves (VMOVSLDUP/VMOVSHDUP), the conversions (VCVTPS2DQ, VCVTTPS2DQ) and the remaining extending and narrowing moves (VPMOVSXBW, VPMOVZXBW, VPMOVWB, VPMOVQB).
  • asm: the EVEX mnemonic suffixes the Go assembler accepts; the rounding modes .RN_SAE, .RD_SAE, .RU_SAE, .RZ_SAE (the EVEX b bit with the rounding control in L'L), suppress-all-exceptions .SAE, and memory broadcast .BCST (the b bit, the vector length preserved, disp8×N scaled by the element size); each combinable with the .Z zeroing suffix, validated against the Go assembler's bytes, and rejected on instructions that do not support them.

[0.12.0] - 2026-07-17

GOOBJ emission: gasm-assembled functions drop into a go build without the Go assembler.

Added

  • asm: GOOBJ object output. gasm asm --format goobj -p <pkgpath> writes the Go toolchain's own object format; the one cmd/link consumes directly: the functions as non-package symbols qualified with the package path (exactly as cmd/asm records assembly symbols), the GLOBL data, one serialized FuncInfo per function (argument/frame sizes, the asm func flag, the start line, the file table) and the four pc-value tables (pcsp, pcfile, pcline, pcinline). The pcsp table carries the real stack deltas: the assembler now tracks every stack-adjustment boundary through the prologue (PUSHQ BP, SUBQ $frame, SP) and each RET's epilogue, so frame-pointer functions unwind correctly. The object preamble; the version-and-experiment header the linker compares verbatim; is captured from the installed go tool asm, so the output is always consistent with the toolchain that links it.
  • asm: relocations against file-local GLOBL symbols become R_PCREL entries in the GOOBJ output, with the instruction's displacement field left zero for the linker to fill (as cmd/asm leaves it).

Fixed

  • parser: 64-bit DATA literals above MaxInt64 (DATA mask<>+8(SB)/8, $0x800f…) parse as unsigned and keep their bit pattern, instead of being rejected as non-integer.

[0.11.0] - 2026-07-16

Linkable object output: external symbols and relocatable ELF / Mach-O objects.

Added

  • asm: object-file emission. gasm asm --format elf writes an ELF64 relocatable object and --format macho a Mach-O x86-64 MH_OBJECT: a code section (.text / __TEXT,__text) and a data section (.data / __DATA,__data), a symbol table with one symbol per TEXT and GLOBL (file-local <> symbols local, the rest global), and one PC-relative relocation per static-symbol reference (R_X86_64_PC32 / X86_64_RELOC_SIGNED, the −4 addend the form needs). The ELF output is verified end-to-end: a gasm-emitted object links with a C driver and runs, resolving both a file-local constant and an external symbol; the Mach-O output is verified structurally with debug/macho.
  • asm: external symbol references. A reference to a symbol no GLOBL in the file defines no longer aborts assembly; it is recorded as an external relocation (Image.Externals, FuncLayout.Relocs) and becomes an undefined global symbol in the object output. The raw image s (--format raw, the default) still reports them: only an object file can represent a reference the linker must resolve.

Changed

  • gasm asm takes a --format raw|elf|macho flag selecting what -o writes; without --format the behaviour is unchanged (the concatenated image).

[0.10.0] - 2026-07-15

The EVEX floating-point and conversion set: the packed-double arithmetic, the scalar SD/SS forms, VMOVDDUP and the width-changing conversions, each verified byte for byte against the Go assembler.

Added

  • asm: the rest of the common EVEX/VEX floating-point set; packed double arithmetic (VSUBPD, VDIVPD, VMINPD, VMAXPD, VUNPCKLPD and the EVEX form of VUNPCKHPD), the scalar double and single operations (VSUBSD, VDIVSD, VMINSD, VMAXSD and the full VADDSS/VSUBSS/VMULSS/VDIVSS/VMINSS/VMAXSS family in both VEX and EVEX; the EVEX scalar forms exist for masked and zeroing use), and VMOVDDUP (lane duplication, VEX and EVEX).
  • asm: the width-changing conversions; VCVTDQ2PS and VCVTPS2PD (VEX and EVEX; the destination sets the length for PS→PD), the EVEX form of VCVTDQ2PD, and the packed-double → dword family: VCVTPD2DQ/VCVTTPD2DQ (EVEX-512 only, a ZMM source and an XMM destination) and their X/Y spellings (VCVTPD2DQX/Y, VCVTTPD2DQX/Y), whose length follows the wider source; a new operand form, since the destination is always XMM while VEX.L / EVEX.L'L ride with the source (fixed by the spelling even for a memory source).
  • asm: masking and zeroing on every new form; the scalar SD/SS arithmetic, the unpacks, VMOVDDUP and the conversions all accept the explicit K1-K7 operand and the .Z suffix the way Go writes them.

Changed

  • VCVTPS2PD follows the Go assembler's encoding, which omits the F3 mandatory prefix (VEX.pp / EVEX.pp = 00) that Intel's maps prescribe; the Go toolchain's machine code is the project's byte-for-byte oracle, and gasm reproduces it exactly (and round-trips through the x86 decoder, which shares the convention).

[0.9.0] - 2026-07-14

AVX-512 masking and a wider EVEX integer set.

Added

  • asm: EVEX masking the way Go writes it; an explicit K1-K7 operand placed among the operands (merging mask), and a .Z mnemonic suffix for zeroing (VPADDD.Z Z1, Z2, K2, Z3). Supported across the NDS, reg/rm, immediate-shift, align, extract, convert and move forms, including masked comparisons with a K destination (VPCMPEQD Z0, Z3, K2, K1). K0 is rejected as an explicit mask, and .Z without a mask is an error, matching the Go assembler.
  • asm: the common AVX-512 F/BW integer set; VPADDB/W, VPSUBB/W, VPANDD/Q, VPANDND/Q, VPMULLW, VPAVGB/W, the signed/unsigned min/max family for B/W/D/Q elements, the variable shifts VPSLLVD/Q, VPSRLVD/Q, VPSRAVD/Q, the EVEX forms of VPSHUFD/VPSHUFB, and the VMOVDQU8/VMOVDQU16 move aliases. Register indices 16-31 encode correctly (the mod=11 quirk carries rm[4] in X̄). All verified byte for byte against the Go assembler.
  • lint: masked EVEX forms (.Z suffix, K operands) are recognised by unknown-instruction and exempted from operand-count.

Fixed

  • asm: EVEX register-register operands with indices 16-31 encoded rm[4] into B̄ instead of X̄ (the EVEX mod=11 extension quirk), producing wrong prefix bytes for X16+/Y16+ r/m operands.

[0.8.0] - 2026-07-13

Standard CLI ergonomics.

Added

  • gasm --help prints a proper top-level help (description, commands, flags, examples), and every subcommand now answers -h/--help with its own usage block (usage line, description, flag defaults), exiting 0. An unknown command points at gasm --help instead of dumping the whole usage.

Changed

  • The version is primarily available as the standard gasm --version / -V flag; the gasm version spelling remains as an alias.

[0.7.0] - 2026-07-12

The formatter behaves like go fmt and canonicalises block separation.

Added

  • gasm fmt now works like go fmt: with no arguments; or with a directory argument; it reformats every .s file below it in place and lists the changed files, skipping . and _ directories (.git, _refs, …). Explicit file arguments keep the -w / standard-output behaviour.

Changed

  • s: canonical blank-line layout; a new block (a label, TEXT or GLOBL) is preceded by exactly one blank line, neither more nor less. Comments leading a block stay with it (the blank line goes before them), stacked labels share their block, the function's first label keeps hugging its TEXT, and runs of blank lines collapse to one. The output remains idempotent and round-trips through the parser. All four go-flac/go-lz4 kernels were reformatted with this release and remain byte-identical when assembled.

[0.6.0] - 2026-07-11

Calibrated to the Go ABI: register-clobber stops reporting legal code, and the encoder learns the legacy SSE moves.

Changed

  • lint: register-clobber is now calibrated to the Go ABI (cmd/compile/abi-internal.md), not the platform ABI. Go's stack-based ABI0 has no System V style callee-saved registers; amd64 BX, R12-R15 and the arm64/riscv64/loong64 scratch sets are caller-saved or permanent scratch, and hand-written kernels may clobber them freely. The rule now audits only the registers Go fixes across calls: the frame pointer and the goroutine pointer (amd64 BP/R14, arm64 R18/R28/R29, riscv64 X27, loong64 R22), and the goroutine pointer is reported only when the function can reach the runtime (is not NOSPLIT or makes a call); the ABI0 transition restores it on those paths, and NOSPLIT call-free leaves may use it, exactly as the runtime's own assembly does. Both go-flac kernels now lint with zero diagnostics.

Fixed

  • lint: the liveness analysis took the destination operand to be the first operand on arm64, riscv64 and loong64; Plan 9 spelling puts it last on every architecture Go supports. The def/use and save/restore classification on those architectures was inverted.
  • s: a comment that follows a RET (typically the next function's doc comment) is no longer indented as if it were still inside the finished function body.

Added

  • asm: the legacy (non-VEX) SSE moves; MOVOU/MOVO (the Plan 9 names for MOVDQU/MOVDQA), MOVUPS/MOVAPS/MOVUPD/MOVAPD and the scalar MOVSD/MOVSS; and VMOVDQU64 in the EVEX set. All verified byte for byte against the Go assembler.

[0.5.0] - 2026-07-10

EVEX / AVX-512: the go-flac AVX-512 kernel now assembles, byte-identically to the Go toolchain, completing the production-kernel coverage.

Added

  • asm: EVEX (AVX-512) encoding; the four-byte EVEX prefix with the 5-bit register fields (Z0-Z31, X/Y 16-31, with the reg-r/m X̄ quirk and V'̄ shared between vvvv and the SIB index), opmask registers (K0-K7) as operands and as mask destinations, and the compressed disp8×N displacement (the multiplier follows the memory operand's size, as the Go assembler's opcode tables prescribe). Covers every AVX-512 instruction the go-flac kernels use: VPXORD/Q, VPADDD, VPSUBD/Q, VPUNPCK*DQ, VPMULLD/Q, VPERMD, VPSLLD/VPSRAD/VPSRAQ, VALIGND, VPCMPEQD (K destination), VMOVDQU32, VMOVUPD, VCVTQQ2PD, VPMOVSXDQ, the narrowing stores VPMOVDW/VPMOVQD, the extracts VEXTRACTI64X4/VEXTRACTF64X4, VFMADD231PD, VADDPD, VMULPD, the broadcasts VPBROADCASTD/Q (GPR and memory sources take different opcodes) and the mask moves KMOVW/KTESTW. Masking/zeroing suffixes are out of scope ; the kernels use neither.
  • asm: AssembleFile now accepts file-defined global (non-<>) symbols too; a reference is external only when no GLOBL in the file defines it.

Fixed

  • asm: registers X16-Y31 force the EVEX encoding of dual-form mnemonics; previously a VPBROADCASTD AX, Y30 fell into the VEX encoder, which cannot represent indices above 15 and silently truncated them.
  • asm: the VEX encoder now rejects vector register indices 16-31 instead of encoding a truncated (wrong) register.

[0.4.0] - 2026-07-09

The standalone assembler reaches the whole go-flac AVX2 kernel: static symbols assemble, and all 17 kernel functions now match the Go toolchain's machine code byte for byte.

Added

  • asm: file-level assembly; AssembleFile turns a parsed file into an Image: the function bodies in source order followed by a data section built from the file's GLOBL/DATA directives (each symbol 16-aligned).
  • asm: static-symbol (SB) operands; mask<>(SB) references encode as RIP-relative loads with a patched disp32, resolved against the image layout so the output is self-consistent and position-independent. External (non-file-local) symbols are rejected with a clear error: they need object-file emission.
  • gasm asm prints the data section and symbol map alongside the functions and writes the whole image (code + data) with -o.

[0.3.0] - 2026-07-08

The assembler reaches byte-identical parity with the Go toolchain on the production go-flac AVX2 kernels: every one of the 15 kernel functions that avoid global symbols now assembles to exactly the Go assembler's bytes (the two holdouts load a file-local constant through SB and wait on relocation support).

Added

  • asm: the scalar instruction families the kernels use; CMOVcc and SETcc (conditions spelled exactly like the jumps), LZCNT/TZCNT (legacy F3 0F BD/BC), the sign/zero-extending moves (MOVBLZX, MOVBQZX, MOVWLZX, MOVWQZX, MOVWLSX, MOVLQSX), CVTSL2SD/CVTSQ2SD (the legacy SSE encoding, as the Go assembler emits it), the traditional three-operand IMUL3{W,L,Q}, and the variable-count vector shifts (VPSRLQ X0, Y8, Y8; the count in an XMM register or memory takes the ordinary NDS form).
  • asm: jump relaxation; jumps start in the short (rel8) form and expand to rel32 when the settled displacement does not fit, iterating the layout to a fixed point (CALL is always rel32).
  • asm: jump-to-jump folding; a conditional jump to a label whose only instruction is an unconditional jump is redirected to the ultimate target, replicating the Go toolchain's linker, which chases such chains before it encodes branches.
  • parser: leading negative displacements with a base and index (LEAQ -4(DX)(R9*4), R9) parse into a fully populated address.

Fixed

  • asm: CMP with a register or memory operand computed second − first instead of first − second, silently inverting every condition that followed (CMPQ SI, R10; JGE tested R10 ≥ SI). The encoding now always records first − second; CMP r/m, r with the first operand in r/m, CMP r, r/m with the first operand in reg; and is byte-identical to the Go assembler.
  • asm: register-to-register MOV now uses the r/m ← r opcode (reg = source), the Go assembler's choice; the output is byte-identical.

[0.2.0] - 2026-07-07

The assembler grows the SIMD set: shuffles, extract/insert, permute and the moves, on top of the VEX forms of the first release.

Added

  • asm: four new VEX (AVX/AVX2) operand forms, each validated by round-trip decoding through golang.org/x/arch and byte-for-byte against the machine code the real Go assembler emits:
    • the immediate shuffle (VPSHUFD, VPERMQ),
    • the three-operand-plus-immediate form (VSHUFPD, VPERM2I128, VINSERTI128),
    • the lane extract (VEXTRACTI128, VEXTRACTF128; the YMM source occupies the ModRM.reg field, the XMM/memory destination the r/m field),
    • the direction-sensitive moves (VMOVDQU, VMOVUPD, VMOVD, VMOVQ, VMOVSD; each direction picks its own opcode and VEX.W; a vector→vector move uses the store-form layout, matching the Go assembler),
    • the no-operand VZEROUPPER, and VPERMD in the NDS form,
    • the floating-point and FMA set (VADDPD, VMULPD, VXORPD, VUNPCKHPD, the scalar VADDSD/VMULSD, VCVTDQ2PD, VFMADD231PD). With the scalar set and the earlier NDS / reg-rm / immediate-shift forms, the encoder now covers every integer, shuffle and FP instruction the go-flac AVX2 kernels use.
  • asm: CMP accepts the immediate in the second operand position (CMPL CX, $31), the spelling the Go assembler accepts, encoding it identically to the immediate-first form.

Fixed

  • asm: an unused VEX.vvvv field is now stored as 1111 (v̄vvv = 1111), as the hardware requires; the previous value (0000) made the two-operand reg/rm forms (VPMOVSXWD, VPBROADCASTD, VMOVMSKPS, …) raise #UD on real CPUs and differ from the Go assembler's bytes. The round-trip decoder ignores the field on these instructions, which is why the byte-for-byte Go comparison (added this release) is now part of the test suite.

[0.1.0] - 2026-07-06

Initial release: the foundation.

Added

  • token, lexer, ast, parser: a hand-written, error-tolerant front end for Plan 9 assembly. The lexer splices C-preprocessor line continuations (\ before a newline) so multi-line #define macros parse as one opaque directive. Validated against the production AVX2/AVX-512 kernels in go-libraries/go-flac and the Go runtime's src/runtime/*.s for all four architectures, with zero parse errors.
  • arch: register files and complete instruction tables for amd64, arm64, riscv64 and loong64, with the middle-dot symbol separator and static (<>) symbols. Instruction names are generated from the Go toolchain's own assembler source (just gen); the anames opcode lists plus the common opcodes and the per-architecture front-end aliases (arm64 B/BL, the .P/.W addressing suffixes, loong64 JAL, the x86 conditional-jump spellings); so every mnemonic the real assembler accepts is recognised.
  • lint: conservative rules; unknown-instruction, operand-count, undefined-label, duplicate-label, missing-ret, missing-textflag-include, abi-argsize and unreachable-code. Macro invocations are recognised (in-file #define names and underscore identifiers) and the label/RET heuristics are suppressed in macro-using files. abi-argsize parses the // func signature with the Go parser and checks the declared TEXT argument size against Go's ABI0 layout; unreachable-code flags dead code after RET, suppressed where reachability is undecidable (PC-relative jumps, register-indirect branches, #ifdef). Register liveness is computed by dataflow over the control-flow graph (basic blocks, def/use, iterative backward iteration) and drives register-clobber, an audit that flags a callee-saved register written but never saved/restored. funcdata-pcdata validates the structure of FUNCDATA/PCDATA directives. Zero error-severity diagnostics across the 90-file Go runtime corpus and the production go-flac kernels (the register-clobber audit additionally reports the go-flac kernels' unsaved callee-saved register use for review).
  • s: an idempotent canonical formatter (operand spacing and per-function mnemonic alignment) that preserves comments and round-trips through the parser.
  • lsp: a Language Server Protocol server over stdio providing completion, hover documentation, document symbols, publish-diagnostics and semantic-token highlighting.
  • asm: a standalone amd64 (x86-64) assembler; an instruction encoder (REX/ ModR-M/SIB/displacement/immediate plus the scalar instruction set, and VEX/ AVX2 SIMD across the three operand forms NDS, reg/rm and immediate-shift, covering the bulk of the integer SIMD set) validated by round-trip decoding against golang.org/x/arch, and an assembler that drives the parser's AST into the encoder with local-label resolution and FP/SP frame mapping (plus Go prologue/epilogue generation), producing output byte-identical to the Go assembler for the supported operand forms.
  • cmd/gasm: the gasm binary with tokens, parse, fmt, lint, asm and lsp subcommands.
  • _gen: the generator that rebuilds the architecture instruction tables from the Go toolchain source (just gen).