# Changelog All notable changes to gasm-devkit are documented here. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and this project adheres to [Conventional Commits](https://www.conventionalcommits.org/). ## [development] Unreleased changes on the `development` branch. ### Added - **arm64 encoder (Phase 5 — complete).** `gasm asm` can now assemble `_arm64.s` files: the AArch64 integer instruction set with the MOV pseudo-instruction and its immediate-constant expansions (MOVZ/MOVN/MOVK for wide immediates, ORR with logical bitmask encoding for values like `$1`), data-processing (shifted register and immediate forms), load/store (scaled unsigned and unscaled9-bit immediate), conditional and unconditional branches, FP/SP frame mapping, SB/global symbol references (ADRP+ADD pairs with `R_ADDRARM64` relocations), jump chain folding, and ELF64 emission (`gasm asm --format elf`). Ground-truth verification against `GOARCH=arm64 go tool asm` matches byte-for-byte. Phase 5 (the other architectures — RISC-V, LoongArch, arm64) is now complete. ## [0.30.0] — 2026-08-13 The LoongArch encoder (Phase 5) ships with ELF64 and GOOBJ emission, verified byte-for-byte against `GOARCH=loong64 go tool asm` and linked into a real `go build`; the shared GOOBJ emitter now writes the per-function DWARF symbols the linker's DWARF pass reads. The RISC-V encoder reaches byte-for-byte parity with `go tool asm`: the frame model, operand ordering, RVC compression, large-immediate and `MOV $imm` materialisation, branch/jump encodings, and `CALL sym(SB)` (now a `JAL` with an `R_RISCV_JAL` relocation). The debugger tracks four hardware watchpoint slots, and the toolkit is Linux-only. ### Added - **LoongArch encoder (Phase 5).** `gasm asm` can now assemble `_loong64.s` files: the full LoongArch64 instruction set with the dual-form arithmetic mnemonics, the 16/21-bit branch families, the MOV pseudo-instruction and its immediate-constant expansions, FP/SP frame mapping, SB/global symbol references (pcalau12i pairs) and ELF64 emission (`gasm asm --format elf`). Ground-truth verification against `GOARCH=loong64 go tool asm` matches byte-for-byte; GOOBJ emission (`gasm asm --format goobj`) is proven end-to-end by linking the object into a cross-compiled `go build`. - **GOOBJ DWARF symbols.** The GOOBJ emitters now write the per-function DWARF symbols the linker requires (the subprogram DIE and the `.debug_line` program, byte-identical to `cmd/asm`'s), and the pc-value table deltas are in the architecture's MinLC units as the runtime expects — the amd64 link test now genuinely substitutes the gasm object, and the amd64/loong64 end-to-end GOOBJ link tests pass. - **RISC-V GOOBJ emission via the shared emitter.** RISC-V GOOBJ output is now written by the same shared emitter as amd64 and LoongArch, modelling each AUIPC + second-instruction pair as a single R_RISCV_PCREL_ITYPE/STYPE relocation (the layout `cmd/asm` writes, not the ELF HI20/LO12 pair), so the object links into a cross-compiled `go build` for `GOARCH=riscv64`. An end-to-end link test substitutes the gasm object and reads the symbol back with `go tool nm`; the rewrite also corrects the relocation `after` field. ### Fixed - **RISC-V frame model and RVC encodings.** The riscv64 frame layout now matches `go tool asm`: the prologue/epilogue save and restore the link register (LR) instead of S0, with the correct autosize (locals + 8) and the RVC-compressed prologue/epilogue instructions; `RET` emits the uncompressed `JALR X0, 0(X1)` the toolchain writes; the `C.ADDI`/`C.LI`/`C.LUI`/`C.ADDIW` opcode bit and the `C.ADD` CR-type encoding are fixed; and the `LR`/`TMP` register aliases now resolve to X1 and X31. The pcsp/pcfile/pcline tables are populated from the recorded stack-adjustment and source-line data, and a byte-exact ground-truth test compares framed and leaf functions against `GOARCH=riscv64 go tool asm`. - **RISC-V operand ordering and RVC compression.** R-type instructions now take `rs2, rs1, rd` and I-type arithmetic instructions take `imm12, rs1, rd`, matching the Go assembler's documented operand order (previously both were reversed, so non-commutative R-type instructions such as `SUB` encoded the wrong operation). The two-operand ternary forms (`ADD rs2, rd`, `ADDI $imm, rd`, `SLLI $shamt, rd`) are now accepted. RVC compression is completed for `C.ADDI16SP`, `C.SLLI`, `C.SRLI`, `C.SRAI`, `C.ANDI`, `C.NOP`, `C.EBREAK`, `C.MV` (from `ADDI`/`ADD`) and the commutative `AND`/`OR`/`XOR` forms; the byte-exact ground-truth test now covers these. - **RISC-V compressed loads/stores and word arithmetic.** RVC compression now also covers the register-relative `C.LW`/`C.SW`/`C.LD`/`C.SD`/ `C.FLD`/`C.FSD` forms (in addition to the stack-relative `C.LWSP`/`C.SWSP`/ `C.LDSP`/`C.SDSP`), plus `C.ADDI4SPN`, `C.ADDW` and `C.SUBW`. The byte-exact ground-truth test exercises these against `GOARCH=riscv64 go tool asm`. - **RISC-V large-immediate materialisation.** `ADDI`/`ANDI`/`ORI`/`XORI` with a 32-bit immediate that does not fit 12 bits now expand exactly as `cmd/asm`: two `ADDI`s for the small `ADDI` split range, and `LUI`+`ADDIW`+`` otherwise, with the `LUI` and `ADDIW` compressed to `C.LUI`/`C.ADDIW` when their immediate fits six signed bits. The byte-exact ground-truth test covers positive, negative, and out-of-range immediates against `GOARCH=riscv64 go tool asm`. - **RISC-V `MOV $imm, rd` materialisation.** The immediate-loading pseudo-instruction now uses the toolchain's `Split32BitImmediate` split (previously it rounded the upper 20 bits, producing wrong results for negative and bit-11-set immediates) and compresses the emitted `ADDI`/`LUI`/`ADDIW` to `C.LI`/`C.LUI`/`C.ADDIW` when their immediate fits six signed bits. A byte-exact ground-truth test covers zero, small, negative, and 32-bit immediates against `GOARCH=riscv64 go tool asm`. - **RISC-V branch/jump compression.** `JMP`/`JAL` were being compressed to `C.J` and `BEQ`/`BNE` (with `X0`) to `C.BEQZ`/`C.BNEZ`, but `go tool asm` never emits these compressed forms. They now emit the 32-bit `JAL` and branch encodings the toolchain writes; the dead `C.J`/`C.BEQZ`/`C.BNEZ` encoders were removed, and the `C.LUI` direct-instruction compression now uses the correct six-bit signed range. A byte-exact ground-truth test covers the branch family and jumps against `GOARCH=riscv64 go tool asm`. - **RISC-V `CALL sym(SB)`.** The call pseudo-instruction now emits the toolchain's `JAL X1, sym(SB)` with a single `R_RISCV_JAL` relocation (previously it emitted an `AUIPC`+`JALR` pair against a local branch label, a form `go tool asm` rejects). The GOOBJ and ELF emitters now map that relocation (Go objabi 59 / ELF `R_RISCV_JAL` 17, a 4-byte field), and relocation offsets are recorded relative to the function start (including the prologue). A byte-exact ground-truth test covers a call against `GOARCH=riscv64 go tool asm`. - **Debugger watchpoint slots.** `gasm debug`'s `watch` command always used hardware watchpoint slot 0, so a second `watch` call silently overwrote the first. Watchpoint slots are now tracked in the `Session` (DR0–DR3); `watch` picks the first free slot and reports an error if all four are in use, and `unwatch ` clears one (no argument clears all). ### Changed - **Linux only.** The toolkit, its CI and the released binaries are now Linux-only; cross-compiled to linux/{amd64,arm64,riscv64,loong64}. - **Phase 4 closed.** README's "Remaining" list for the debugger is gone; disassembly at PC, memory-write, watchpoints, and source-line mapping are all shipped. ## [0.29.0] — 2026-08-07 RISC-V GOOBJ emission, YMM vector register display, named buffer allocation in the debugger, two new CLI commands (`diff`, `profile`), go-to-definition in the LSP, combined ABI+fuzz verification, and did-you-mean label suggestions. A `--map` flag for `diff` and `--call`/`--buf` flags for `verify` extend the new CLI commands. A signature-parser fix corrects grouped Go parameters. ### Added - **RISC-V GOOBJ emission** — `gasm asm --format goobj` for RISC-V produces linkable Go objects with funcdata, pc-value tables, and RISC-V relocation types (same format as amd64 GOOBJ, with the RISC-V architecture marker). - **`gasm diff`** — compare the machine code of two assembly files byte-for-byte; shows which functions differ and the first few differing bytes. - **`gasm profile`** — show the basic-block structure of each function: labels, offsets, frame size, and NOSPLIT flag. - **LSP go-to-definition** — `textDocument/definition` navigates from a label reference to its definition. - **did-you-mean** — when the RISC-V assembler encounters an undefined label, it suggests the closest existing label using Levenshtein distance. - **YMM vector register display** — `regs` in the debugger now shows YMM registers via `PTRACE_GETFPREGS` (falls back to XMM when XSAVE is unavailable). - **Named buffer allocation** — `gasm debug --buf name:size:pattern` allocates buffers in the debuggee filled with `zero`, `ones`, `seq`, or a hex pattern; buffer pointers are placed into the argument block at the matching positions. - **Crash input storage** — `FuzzResult.CrashInput` stores the input that caused a crash or mismatch for reproducibility. - **ABI + fuzz combined** — `gasm verify --fuzz` now runs ABI checks (sentinel registers, canary, stack bounds) alongside differential fuzz testing. - **`gasm diff --map`** — compare functions whose names differ between files (e.g. `--map wideCopyAVX2=wideCopyAVX512` pairs two variants regardless of suffix). Unmapped functions fall back to the original name match. - **`gasm verify --call`** — invoke a single function with user-supplied buffers (`--buf name:size:pattern`) instead of the smoke/abi/fuzz sweeps. Patterns: `zero`, `ones`, `seq`, or a hex blob. Useful for partial functions (e.g. decoders) that crash on random input but should succeed on valid data. The arg block is printed before and after the call, showing return values. - **`gasm verify --ground-truth`** now documented in `--help` (was already a flag, just missing from the help text). ### Fixed - **Signature parser** — grouped Go parameters like `dst, src []byte` are now parsed correctly (both get type `[]byte`). Previously the first name was treated as its own type (`dst` with size 8), causing wrong ABI0 arg-block layout in both `verify --call` and the fuzzer. - **Flaky JIT tests** — `runtime.KeepAlive` guards and package-level buffers prevent GC from collecting heap objects whose addresses were passed to JIT code via `unsafe.Pointer`; all verify tests pass 100/100 under `-race`. ### Changed - **Removed external kernel test dependencies** — the verify test suite no longer references production kernels from the separate go-libraries project. The remaining test suite uses only `testdata/verify/*.s` kernels, which are part of this repository. Coverage is identical locally and in CI (80.3 %). ## [0.28.0] — 2026-08-03 RISC-V encoder: full RV64IMAFDC instruction set with RVC compression, MOV pseudo-instruction, SB/global symbol references, ELF64 object emission, and ground-truth verification against `GOARCH=riscv64 go tool asm`. ### Added - **RISC-V encoder** — RV64I, RV64M, RV64A, RV64F/D, FMA, CSR, JALR. - **MOV pseudo-instruction** — load, store, reg-to-reg, immediate, frame mapping. - **RVC compression** — 22 compressed instruction types (C.LDSP, C.SDSP, C.FLDSP, C.FSDSP, C.ADDI, C.LI, C.LUI, C.ADDIW, C.MV, C.ADD, C.SUB, C.XOR, C.OR, C.AND, C.SLLI, C.SRLI, C.SRAI, C.ANDI, C.BEQZ, C.BNEZ, C.J, C.JR). - **SB/global symbols** — `MOV $sym(SB)`, `MOV sym(SB)`, `MOV rd, sym(SB)` encoded as AUIPC pairs with R_RISCV_PCREL_HI20/LO12 relocations. - **GLOBL/DATA** — data section layout in `AssembleFileRISCV`. - **ELF64 emission** — `gasm asm --format elf` produces EM_RISCV objects (.text, .data, .symtab, .rela.text). - **`gasm verify --ground-truth`** — byte-exact comparison against `GOARCH=riscv64 go tool asm`. - **`gasm verify --profile`** — function layout listing for RISC-V. - **CALL** — AUIPC + JALR pair encoding. ### Fixed - Parser: bare-number offset before `(SP)` no longer misidentified as pseudo. - MOV: `MOV $sym(FP/SP), rd` now returns an explicit error instead of silent fallback. - RVC: C.LDSP/C.SDSP/FLDSP/FSDSP immediate encoding now matches Go toolchain (bit-interleaved format). ## [0.27.0] — 2026-08-01 Subprocess isolation for `--fuzz`: each function is fuzzed in its own child process, so a partial function (decoder) that faults on random garbage is reported as "CRASH (partial function, use --ground-truth)" without killing the parent. CRASH is informational (exit 0); only MISMATCH is an error. ### Fixed - `gasm verify --fuzz` no longer crashes the process on partial functions. ## [0.26.0] — 2026-07-31 Universal differential fuzzing: `gasm verify --fuzz` needs no hand-written reference. It parses the `// func` signature from the assembly source, generates typed random inputs (slices with random content, ints, pointers to fixed arrays), JIT-executes BOTH the gasm-assembled and the go-tool-asm- assembled versions with independent buffer copies, and compares the result area bit-for-bit. ### Added - `verify`: `FuzzFunc` / `ExtractSignatures` / `parseFuncSig` — universal differential fuzz driven by the conventional `// func` comment. Each version gets its own buffer set (deep copy) so functions that write to their arguments (histogram increments) don't corrupt the other's input. - `gasm verify --fuzz [-n N]`: runs the differential fuzz for every function with a parseable signature. Total functions (wideCopy, pack16, decorrelate, analyze, autocorr) pass; partial functions (decoders that fault on malformed input) should use `--ground-truth` instead. ### Fixed `--fuzz` crashes the process for partial functions (e.g. LZ4 decoders) whose over-copy paths read past the buffer on random garbage input. Subprocess isolation (fork per function) is planned. Use `--ground-truth` for decoders. ## [0.25.0] — 2026-07-30 Universal ground-truth verification: `gasm verify --ground-truth` assembles any `.s` file with both gasm and `go tool asm`, then compares the machine code byte-for-byte per function (relocation sites masked). No hand-written reference needed — the Go toolchain IS the oracle. ### Added - `verify`: `GroundTruth` — shells out to `go tool asm`, parses the GOOBJ output (minimal reader: block offsets, nonpkg symbol table, data index) and returns per-function code bytes. - `gasm verify --ground-truth`: compares gasm's output against the Go assembler's, reporting MATCH/MISMATCH per function with the first differing byte. Relocation disp32 fields (static-symbol references the linker fills) are masked before comparison. - Verified: go-lz4 AVX2 2/2, go-flac AVX2 17/17 functions byte-identical. ## [0.24.0] — 2026-07-29 The full analyze family and stereo PCM decode are now differentially tested. 15 of 17 go-flac AVX2 kernels have bit-for-bit differential coverage; the two remaining (autocorrAVX2 — FMA reassociation, lpcResidualAVX2 — complex multi-arg) are deferred. ### Added - `verify`: `analyzeO3RangeAVX2` and `analyzeO4RangeAVX2` differential tests (200 iterations each, same harness as O1/O2/Res). - `verify`: `decodeStereo16AVX2` differential test (500 random interleaved stereo PCM buffers, both channels compared sample-by-sample). ## [0.23.0] — 2026-07-28 The analyze family and 24-bit PCM decode join the differential suite. ### Added - `verify`: `analyzeO2RangeAVX2` and `analyzeResRangeAVX2` differential tests (200 iterations each, shared harness with O1: zigzag fold, partial sum, overflow flag and Len32 histogram). - `verify`: `decodeMono24AVX2` differential test (500 random 24-bit PCM buffers, sign-extension compared sample-by-sample). ## [0.22.0] — 2026-07-27 The remaining go-flac encoder kernels join the differential suite. ### Added - `verify`: `analyzeO1RangeAVX2` differential test (300 random partitions: zigzag fold, partial sum, overflow flag and the 32-bin Len32 histogram compared element-by-element against the portable Go reference). - `verify`: `fastStereoSumsAVX2` differential test (300 random stereo frames: the four zigzag-fold entropy sums compared against the scalar loop). ## [0.21.0] — 2026-07-26 Differential testing extended to all four production kernels and the CLI exposes the full dynamic-analysis toolkit. ### Added - `verify`: go-flac AVX2 differential tests — `decodeMono16AVX2` (500 random PCM buffers), `pack16AVX2` (500 random int32→int16 packings) and all four decorrelation kernels (200 iterations each: left-side, side-right, mid-side, interleave) compared bit-for-bit against the portable Go references. - `verify`: go-lz4 AVX-512 differential tests — `decodeBlockAVX512` (3 000 fuzzed LZ4 blocks + known answers) and `wideCopyAVX512` (0–1024 bytes) against the same portable oracle as the AVX2 suite. - `gasm verify --abi`: runs each NOSPLIT function with sentinel registers and a red-zone canary, reporting violations. - `gasm verify --profile`: lists the static basic-block count per function. ## [0.20.0] — 2026-07-25 Coverage profiling: the third pillar of Phase 3. Static basic-block enumeration from the assembler's label map, combined with multi-input path diversity measurement — how many observationally distinct execution paths a test corpus exercises. ### Added - `verify`: `Kernel.Blocks` / `Kernel.BlockCount` — enumerate basic blocks from the assembler's local-label map (every jump target is a block boundary; the function entry is always a block). `decodeBlockAVX2` has 27 blocks. - `verify`: `Kernel.ProfilePaths` — run the function with a corpus of argument blocks and collect distinct output fingerprints (the result words); reports path diversity as a lower bound on code coverage. ### Changed INT3-based per-block hit counting was prototyped but deferred: Go's runtime signal management (sigaltstack, handler re-installation) makes raw rt_sigaction handlers fragile in a Go process. The static + path-diversity approach delivers the project's goal (proving the SIMD path and tail handling execute) without fighting the runtime. ## [0.19.0] — 2026-07-24 Runtime ABI checks: the second pillar of Phase 3. The JIT trampoline now has an ABI-checking variant that sets sentinels in the callee-saved registers (BP, R14) before entering the assembled function and verifies they survive on return, plus a red-zone canary (128 bytes below SP filled with 0xA5) that detects any illegal write below the stack pointer. ### Added - `verify`: `CallChecked` / `Kernel.CallFuncChecked` — ABI-checking JIT call with sentinel registers and red-zone canary; returns an `ABIReport` (BPClobbered, R14Clobbered, RedZoneHit). - `verify`: the raw `leaveJITCheckedRaw` trampoline — a TEXT symbol with no ABIInternal wrapper (address obtained via GLOBL/DATA), so the JIT function's RET lands directly in the check code and sees the registers exactly as the function left them. - Tests: deliberate BP/R14 clobberers detected; both go-lz4 kernels confirmed ABI-clean (BP preserved, R14 preserved, red zone intact). ## [0.18.0] — 2026-07-23 Differential testing: the JIT-assembled go-lz4 `decodeBlockAVX2` kernel is fuzzed against a portable Go reference — 5 000 valid LZ4 blocks compared bit-for-bit, plus 2 000 hostile (random garbage) inputs with matching error codes. This is the automated form of the project's bit-identical contract. ### Added - `verify`: differential fuzz tests — a random LZ4 block generator produces valid blocks (literals, overlapping matches, extension bytes) and the JIT-assembled kernel's output is compared byte-for-byte against a portable Go decoder; a hostile-input suite confirms error-code agreement on random garbage (no crashes, same classification). ## [0.17.0] — 2026-07-22 Phase 3 begins: dynamic analysis. A JIT execution substrate that assembles Plan 9 amd64 kernels into executable memory and calls them directly — pure Go (stdlib only, `syscall.Mmap` + an assembly trampoline), no cgo, no external toolchain. ### Added - `verify` package: JIT infrastructure — `Map` copies machine code into a W^X memory mapping, `Call` invokes it through an ABI0 trampoline that switches to a prepared stack and back. `Load`/`LoadSource`/`LoadAST` parse, assemble and map a `.s` file in one step; `Kernel.CallFunc` marshals the argument block and returns results. - `gasm verify` subcommand: assembles a file, JIT-loads it and reports the available functions; with `-smoke`, calls each NOSPLIT function with zeroed arguments to confirm the trampoline works end-to-end. - Integration tests: the go-lz4 `decodeBlockAVX2` and `wideCopyAVX2` kernels (699 and 146 bytes) assemble, map and execute correctly — known-answer LZ4 blocks decode bit-for-bit, wide copies of 0–1024 bytes match, malformed input returns the correct error codes. ## [0.16.0] — 2026-07-21 The scalar conversions between vector and general-purpose registers — the last of the amd64 EVEX instruction set. ### Added - `asm`: the GPR-interchanging conversions, byte for byte against the Go assembler (28 ground-truth cases including memory sources and extended GPRs): vector to GPR — the signed and truncated VCVT{,T}S{D,S}2SI{,Q} in both VEX and EVEX, and the unsigned VCVT{,T}S{D,S}2USI{L,Q} (EVEX only); GPR to vector — VCVTSI2SD{L,Q}/VCVTSI2SS{L,Q} (VEX and EVEX) and VCVTUSI2SD{L,Q}/VCVTUSI2SS{L,Q} (EVEX only), whose preserved vector source sits in vvvv (three Plan 9 operands). ## [0.15.0] — 2026-07-20 The last of the EVEX conversions and narrowing/extending moves — the EVEX instruction set is now complete save for the GPR-interchanging forms. ### Added - `asm`: the unsigned and truncating conversions — VCVTPD2PS (and the X/Y spellings, whose length the spelling fixes), VCVTPD2UDQ (X/Y), VCVTTPD2UDQ (X/Y), VCVTTPD2UQQ, VCVTPS2UDQ, VCVTTPS2UDQ, VCVTPS2UQQ, VCVTTPS2UQQ, VCVTTPD2QQ, VCVTTPS2QQ, VCVTUQQ2PD, VCVTUQQ2PS (X/Y) and VCVTQQ2PS X/Y. - `asm`: the remaining sign/zero-extending moves (VPMOVSXBD/BQ/WQ and VPMOVZXBD/BQ/WD/WQ, VEX and EVEX) and the complete signed and unsigned narrowing stores (VPMOVS{DB,QB,DW,QW,QD,WB}, VPMOVUS{DB,QB,DW,QW,QD,WB}, VPMOVDB, VPMOVQW). - `asm`: the mask/vector conversions (VPMOVM2B/W/D/Q and VPMOVB2M/W2M/ D2M/Q2M), whose K register is a genuine operand rather than a mask and which therefore take no masking suffixes. ## [0.14.0] — 2026-07-19 The floating-point helper and conversion tail of the AVX-512 set, plus gather and scatter with VSIB addressing — every encoding verified byte for byte against the Go assembler. ### Added - `asm`: the floating-point helpers — reciprocals and reciprocal square roots (VRCP14/VRSQRT14 PD/PS/SD/SS), exponents and mantissas (VGETEXP*, VGETMANT*), scaling by powers of two (VSCALEF*), rounding (VRNDSCALE*), reduction (VREDUCE*), immediate fixup (VFIXUPIMM*) and range selection (VRANGE*), and floating-point class tests (VFPCLASSPD/PS X/Y/Z and VFPCLASSSD/SS — a new immediate form whose reg field carries the opmask destination). - `asm`: **gather and scatter with VSIB addressing.** The gathers take both Go spellings: the VEX form with a vector mask register (OP mask, vsib, dst) and the EVEX form with an explicit K mask (OP vsib, K, dst), where the EVEX L'L field follows the VSIB index register rather than the data register (a ZMM index with an YMM destination encodes L'L = 10, as the Go assembler emits). The scatters (VSCATTER*/VPSCATTER*) are EVEX only (OP src, K, vsib). All eight gather and eight scatter widths. - `asm`: the remaining conversions — VCVTQQ2PS (the 512-bit source sets the length), VCVTPD2QQ/UQQ, VCVTPS2QQ, VCVTUDQ2PD/PS, the half-precision VCVTPH2PS and VCVTPS2PH (the extract layout with an immediate). ## [0.13.0] — 2026-07-18 The wider AVX-512 set: ternary logic, permutes, compares, expand/compress, the opmask instructions and the EVEX rounding/SAE/broadcast suffixes — every encoding verified byte for byte against the Go assembler. ### Added - `asm`: the wider EVEX/AVX-512 set, across roughly sixty new ground-truth cases: ternary logic (VPTERNLOGD/Q), the lane shuffles/inserts/extracts (VSHUF{F,I}{32,64}X{2,4}, the VINSERT*/VEXTRACT* {F,I}{32,64}X{2,4,8} family, VPALIGNR), compares with an opmask destination (VCMPPD/PS/SD/SS — a new NDS-plus-immediate form with the K register in the reg field), the permutes (VPERMB/W, VPERMI2/T2 D/Q/PD), the wider integer families (VPMADDWD/UBSW, VPMULHUW, VPACKSSWB/USWB/SSDW/USDW, VPABS B/W/D/Q, the VPROL*/VPROR* rotates, the word shifts and the EVEX W1 qword shifts), expand/compress (VEXPANDPD/PS, VPEXPANDD/Q, VCOMPRESSPD/PS, VPCOMPRESSD/ Q), the broadcasts (VPBROADCASTB/W from a GPR or memory, VBROADCASTSS/ SD), the opmask-register instructions (KAND/KOR/KXNOR/KADD/KUNPCK/KNOT/ KSHIFTL/KORTEST B/W/D/Q and KMOVQ, whose width the L/W/pp bits select), the packed single arithmetic (VADD/VSUB/VMUL/VDIV/VMIN/VMAX PS), the aligned moves (VMOVAPS/APD, VMOVDQA32/64, VMOVSS), the replicating moves (VMOVSLDUP/VMOVSHDUP), the conversions (VCVTPS2DQ, VCVTTPS2DQ) and the remaining extending and narrowing moves (VPMOVSXBW, VPMOVZXBW, VPMOVWB, VPMOVQB). - `asm`: the EVEX mnemonic suffixes the Go assembler accepts — the rounding modes `.RN_SAE`, `.RD_SAE`, `.RU_SAE`, `.RZ_SAE` (the EVEX b bit with the rounding control in L'L), suppress-all-exceptions `.SAE`, and memory broadcast `.BCST` (the b bit, the vector length preserved, disp8×N scaled by the element size) — each combinable with the `.Z` zeroing suffix, validated against the Go assembler's bytes, and rejected on instructions that do not support them. ## [0.12.0] — 2026-07-17 GOOBJ emission: gasm-assembled functions drop into a `go build` without the Go assembler. ### Added - `asm`: **GOOBJ object output.** `gasm asm --format goobj -p ` writes the Go toolchain's own object format — the one `cmd/link` consumes directly: the functions as non-package symbols qualified with the package path (exactly as `cmd/asm` records assembly symbols), the `GLOBL` data, one serialized `FuncInfo` per function (argument/frame sizes, the asm func flag, the start line, the file table) and the four pc-value tables (`pcsp`, `pcfile`, `pcline`, `pcinline`). The `pcsp` table carries the real stack deltas: the assembler now tracks every stack-adjustment boundary through the prologue (`PUSHQ BP`, `SUBQ $frame, SP`) and each `RET`'s epilogue, so frame-pointer functions unwind correctly. The object preamble — the version-and-experiment header the linker compares verbatim — is captured from the installed `go tool asm`, so the output is always consistent with the toolchain that links it. - `asm`: relocations against file-local `GLOBL` symbols become `R_PCREL` entries in the GOOBJ output, with the instruction's displacement field left zero for the linker to fill (as `cmd/asm` leaves it). ### Fixed - `parser`: 64-bit `DATA` literals above `MaxInt64` (`DATA mask<>+8(SB)/8, $0x800f…`) parse as unsigned and keep their bit pattern, instead of being rejected as non-integer. ## [0.11.0] — 2026-07-16 Linkable object output: external symbols and relocatable ELF / Mach-O objects. ### Added - `asm`: **object-file emission.** `gasm asm --format elf` writes an ELF64 relocatable object and `--format macho` a Mach-O x86-64 `MH_OBJECT`: a code section (`.text` / `__TEXT,__text`) and a data section (`.data` / `__DATA,__data`), a symbol table with one symbol per `TEXT` and `GLOBL` (file-local `<>` symbols local, the rest global), and one PC-relative relocation per static-symbol reference (`R_X86_64_PC32` / `X86_64_RELOC_SIGNED`, the −4 addend the form needs). The ELF output is verified end-to-end: a gasm-emitted object links with a C driver and runs, resolving both a file-local constant and an external symbol; the Mach-O output is verified structurally with `debug/macho`. - `asm`: **external symbol references.** A reference to a symbol no `GLOBL` in the file defines no longer aborts assembly — it is recorded as an external relocation (`Image.Externals`, `FuncLayout.Relocs`) and becomes an undefined global symbol in the object output. The raw image format (`--format raw`, the default) still reports them: only an object file can represent a reference the linker must resolve. ### Changed - `gasm asm` takes a `--format raw|elf|macho` flag selecting what `-o` writes; without `--format` the behaviour is unchanged (the concatenated image). ## [0.10.0] — 2026-07-15 The EVEX floating-point and conversion set: the packed-double arithmetic, the scalar SD/SS forms, VMOVDDUP and the width-changing conversions, each verified byte for byte against the Go assembler. ### Added - `asm`: the rest of the common EVEX/VEX floating-point set — packed double arithmetic (VSUBPD, VDIVPD, VMINPD, VMAXPD, VUNPCKLPD and the EVEX form of VUNPCKHPD), the scalar double and single operations (VSUBSD, VDIVSD, VMINSD, VMAXSD and the full VADDSS/VSUBSS/VMULSS/VDIVSS/VMINSS/VMAXSS family in both VEX and EVEX — the EVEX scalar forms exist for masked and zeroing use), and VMOVDDUP (lane duplication, VEX and EVEX). - `asm`: the width-changing conversions — VCVTDQ2PS and VCVTPS2PD (VEX and EVEX; the destination sets the length for PS→PD), the EVEX form of VCVTDQ2PD, and the packed-double → dword family: VCVTPD2DQ/VCVTTPD2DQ (EVEX-512 only, a ZMM source and an XMM destination) and their X/Y spellings (VCVTPD2DQX/Y, VCVTTPD2DQX/Y), whose length follows the wider source — a new operand form, since the destination is always XMM while VEX.L / EVEX.L'L ride with the source (fixed by the spelling even for a memory source). - `asm`: masking and zeroing on every new form — the scalar SD/SS arithmetic, the unpacks, VMOVDDUP and the conversions all accept the explicit K1–K7 operand and the `.Z` suffix the way Go writes them. ### Changed - VCVTPS2PD follows the Go assembler's encoding, which omits the F3 mandatory prefix (VEX.pp / EVEX.pp = 00) that Intel's maps prescribe; the Go toolchain's machine code is the project's byte-for-byte oracle, and gasm reproduces it exactly (and round-trips through the x86 decoder, which shares the convention). ## [0.9.0] — 2026-07-14 AVX-512 masking and a wider EVEX integer set. ### Added - `asm`: **EVEX masking** the way Go writes it — an explicit `K1`–`K7` operand placed among the operands (merging mask), and a `.Z` mnemonic suffix for zeroing (`VPADDD.Z Z1, Z2, K2, Z3`). Supported across the NDS, reg/rm, immediate-shift, align, extract, convert and move forms, including masked comparisons with a K destination (`VPCMPEQD Z0, Z3, K2, K1`). K0 is rejected as an explicit mask, and `.Z` without a mask is an error, matching the Go assembler. - `asm`: the common AVX-512 F/BW integer set — VPADDB/W, VPSUBB/W, VPANDD/Q, VPANDND/Q, VPMULLW, VPAVGB/W, the signed/unsigned min/max family for B/W/D/Q elements, the variable shifts VPSLLVD/Q, VPSRLVD/Q, VPSRAVD/Q, the EVEX forms of VPSHUFD/VPSHUFB, and the VMOVDQU8/VMOVDQU16 move aliases. Register indices 16–31 encode correctly (the mod=11 quirk carries rm[4] in X̄). All verified byte for byte against the Go assembler. - `lint`: masked EVEX forms (`.Z` suffix, K operands) are recognised by `unknown-instruction` and exempted from `operand-count`. ### Fixed - `asm`: EVEX register–register operands with indices 16–31 encoded rm[4] into B̄ instead of X̄ (the EVEX mod=11 extension quirk), producing wrong prefix bytes for X16+/Y16+ r/m operands. ## [0.8.0] — 2026-07-13 Standard CLI ergonomics. ### Added - `gasm --help` prints a proper top-level help (description, commands, flags, examples), and every subcommand now answers `-h`/`--help` with its own usage block (usage line, description, flag defaults), exiting 0. An unknown command points at `gasm --help` instead of dumping the whole usage. ### Changed - The version is primarily available as the standard `gasm --version` / `-V` flag; the `gasm version` spelling remains as an alias. ## [0.7.0] — 2026-07-12 The formatter behaves like `go fmt` and canonicalises block separation. ### Added - `gasm fmt` now works like `go fmt`: with no arguments — or with a directory argument — it reformats every `.s` file below it in place and lists the changed files, skipping `.` and `_` directories (`.git`, `_refs`, …). Explicit file arguments keep the `-w` / standard-output behaviour. ### Changed - `format`: canonical blank-line layout — a new block (a label, `TEXT` or `GLOBL`) is preceded by exactly one blank line, neither more nor less. Comments leading a block stay with it (the blank line goes before them), stacked labels share their block, the function's first label keeps hugging its `TEXT`, and runs of blank lines collapse to one. The output remains idempotent and round-trips through the parser. All four go-flac/go-lz4 kernels were reformatted with this release and remain byte-identical when assembled. ## [0.6.0] — 2026-07-11 Calibrated to the Go ABI: `register-clobber` stops reporting legal code, and the encoder learns the legacy SSE moves. ### Changed - `lint`: **`register-clobber` is now calibrated to the Go ABI** (`cmd/compile/abi-internal.md`), not the platform ABI. Go's stack-based ABI0 has no System V style callee-saved registers — amd64 `BX`, `R12`–`R15` and the arm64/riscv64/loong64 scratch sets are caller-saved or permanent scratch, and hand-written kernels may clobber them freely. The rule now audits only the registers Go fixes across calls: the frame pointer and the goroutine pointer (amd64 `BP`/`R14`, arm64 `R18`/`R28`/`R29`, riscv64 `X27`, loong64 `R22`), and the goroutine pointer is reported only when the function can reach the runtime (is not `NOSPLIT` or makes a call) — the ABI0 transition restores it on those paths, and NOSPLIT call-free leaves may use it, exactly as the runtime's own assembly does. Both go-flac kernels now lint with zero diagnostics. ### Fixed - `lint`: the liveness analysis took the destination operand to be the *first* operand on arm64, riscv64 and loong64; Plan 9 spelling puts it last on every architecture Go supports. The def/use and save/restore classification on those architectures was inverted. - `format`: a comment that follows a `RET` (typically the next function's doc comment) is no longer indented as if it were still inside the finished function body. ### Added - `asm`: the legacy (non-VEX) SSE moves — `MOVOU`/`MOVO` (the Plan 9 names for MOVDQU/MOVDQA), `MOVUPS`/`MOVAPS`/`MOVUPD`/`MOVAPD` and the scalar `MOVSD`/`MOVSS` — and `VMOVDQU64` in the EVEX set. All verified byte for byte against the Go assembler. ## [0.5.0] — 2026-07-10 EVEX / AVX-512: the go-flac AVX-512 kernel now assembles, byte-identically to the Go toolchain, completing the production-kernel coverage. ### Added - `asm`: **EVEX (AVX-512) encoding** — the four-byte EVEX prefix with the 5-bit register fields (Z0–Z31, X/Y 16–31, with the reg-r/m X̄ quirk and V'̄ shared between vvvv and the SIB index), opmask registers (K0–K7) as operands and as mask destinations, and the compressed disp8×N displacement (the multiplier follows the memory operand's size, as the Go assembler's opcode tables prescribe). Covers every AVX-512 instruction the go-flac kernels use: VPXORD/Q, VPADDD, VPSUBD/Q, VPUNPCK*DQ, VPMULLD/Q, VPERMD, VPSLLD/VPSRAD/VPSRAQ, VALIGND, VPCMPEQD (K destination), VMOVDQU32, VMOVUPD, VCVTQQ2PD, VPMOVSXDQ, the narrowing stores VPMOVDW/VPMOVQD, the extracts VEXTRACTI64X4/VEXTRACTF64X4, VFMADD231PD, VADDPD, VMULPD, the broadcasts VPBROADCASTD/Q (GPR and memory sources take different opcodes) and the mask moves KMOVW/KTESTW. Masking/zeroing suffixes are out of scope — the kernels use neither. - `asm`: `AssembleFile` now accepts file-defined global (`non-<>`) symbols too; a reference is external only when no `GLOBL` in the file defines it. ### Fixed - `asm`: registers X16–Y31 force the EVEX encoding of dual-form mnemonics; previously a `VPBROADCASTD AX, Y30` fell into the VEX encoder, which cannot represent indices above 15 and silently truncated them. - `asm`: the VEX encoder now rejects vector register indices 16–31 instead of encoding a truncated (wrong) register. ## [0.4.0] — 2026-07-09 The standalone assembler reaches the whole go-flac AVX2 kernel: static symbols assemble, and all 17 kernel functions now match the Go toolchain's machine code byte for byte. ### Added - `asm`: **file-level assembly** — `AssembleFile` turns a parsed file into an `Image`: the function bodies in source order followed by a data section built from the file's `GLOBL`/`DATA` directives (each symbol 16-aligned). - `asm`: **static-symbol (`SB`) operands** — `mask<>(SB)` references encode as RIP-relative loads with a patched disp32, resolved against the image layout so the output is self-consistent and position-independent. External (non-file-local) symbols are rejected with a clear error: they need object-file emission. - `gasm asm` prints the data section and symbol map alongside the functions and writes the whole image (code + data) with `-o`. ## [0.3.0] — 2026-07-08 The assembler reaches byte-identical parity with the Go toolchain on the production go-flac AVX2 kernels: every one of the 15 kernel functions that avoid global symbols now assembles to exactly the Go assembler's bytes (the two holdouts load a file-local constant through `SB` and wait on relocation support). ### Added - `asm`: the scalar instruction families the kernels use — `CMOVcc` and `SETcc` (conditions spelled exactly like the jumps), `LZCNT`/`TZCNT` (legacy `F3 0F BD/BC`), the sign/zero-extending moves (`MOVBLZX`, `MOVBQZX`, `MOVWLZX`, `MOVWQZX`, `MOVWLSX`, `MOVLQSX`), `CVTSL2SD`/`CVTSQ2SD` (the legacy SSE encoding, as the Go assembler emits it), the traditional three-operand `IMUL3{W,L,Q}`, and the variable-count vector shifts (`VPSRLQ X0, Y8, Y8` — the count in an XMM register or memory takes the ordinary NDS form). - `asm`: **jump relaxation** — jumps start in the short (rel8) form and expand to rel32 when the settled displacement does not fit, iterating the layout to a fixed point (CALL is always rel32). - `asm`: **jump-to-jump folding** — a conditional jump to a label whose only instruction is an unconditional jump is redirected to the ultimate target, replicating the Go toolchain's linker, which chases such chains before it encodes branches. - `parser`: leading negative displacements with a base and index (`LEAQ -4(DX)(R9*4), R9`) parse into a fully populated address. ### Fixed - `asm`: `CMP` with a register or memory operand computed **second − first** instead of first − second, silently inverting every condition that followed (`CMPQ SI, R10; JGE` tested R10 ≥ SI). The encoding now always records first − second — `CMP r/m, r` with the first operand in r/m, `CMP r, r/m` with the first operand in reg — and is byte-identical to the Go assembler. - `asm`: register-to-register `MOV` now uses the `r/m ← r` opcode (reg = source), the Go assembler's choice; the output is byte-identical. ## [0.2.0] — 2026-07-07 The Phase 2 assembler grows the SIMD set: shuffles, extract/insert, permute and the moves, on top of the Phase 1 VEX forms. ### Added - `asm`: four new VEX (AVX/AVX2) operand forms, each validated by round-trip decoding through `golang.org/x/arch` **and** byte-for-byte against the machine code the real Go assembler emits: - the immediate shuffle (`VPSHUFD`, `VPERMQ`), - the three-operand-plus-immediate form (`VSHUFPD`, `VPERM2I128`, `VINSERTI128`), - the lane extract (`VEXTRACTI128`, `VEXTRACTF128` — the YMM source occupies the ModRM.reg field, the XMM/memory destination the r/m field), - the direction-sensitive moves (`VMOVDQU`, `VMOVUPD`, `VMOVD`, `VMOVQ`, `VMOVSD` — each direction picks its own opcode and VEX.W; a vector→vector move uses the store-form layout, matching the Go assembler), - the no-operand `VZEROUPPER`, and `VPERMD` in the NDS form, - the floating-point and FMA set (`VADDPD`, `VMULPD`, `VXORPD`, `VUNPCKHPD`, the scalar `VADDSD`/`VMULSD`, `VCVTDQ2PD`, `VFMADD231PD`). With the scalar set and the earlier NDS / reg-rm / immediate-shift forms, the encoder now covers every integer, shuffle and FP instruction the go-flac AVX2 kernels use. - `asm`: `CMP` accepts the immediate in the second operand position (`CMPL CX, $31`) — the spelling the Go assembler accepts — encoding it identically to the immediate-first form. ### Fixed - `asm`: an unused VEX.vvvv field is now stored as `1111` (v̄vvv = 1111), as the hardware requires — the previous value (`0000`) made the two-operand reg/rm forms (VPMOVSXWD, VPBROADCASTD, VMOVMSKPS, …) raise #UD on real CPUs and differ from the Go assembler's bytes. The round-trip decoder ignores the field on these instructions, which is why the byte-for-byte Go comparison (added this release) is now part of the test suite. ## [0.1.0] — 2026-07-06 Initial release — the Phase 1 foundation. ### Added - `token`, `lexer`, `ast`, `parser`: a hand-written, error-tolerant front end for Plan 9 assembly. The lexer splices C-preprocessor line continuations (`\` before a newline) so multi-line `#define` macros parse as one opaque directive. Validated against the production AVX2/AVX-512 kernels in `go-libraries/go-flac` and the Go runtime's `src/runtime/*.s` for all four architectures, with zero parse errors. - `arch`: register files and **complete** instruction tables for amd64, arm64, riscv64 and loong64, with the middle-dot symbol separator and static (`<>`) symbols. Instruction names are generated from the Go toolchain's own assembler source (`just gen`) — the `anames` opcode lists plus the common opcodes and the per-architecture front-end aliases (arm64 `B`/`BL`, the `.P`/`.W` addressing suffixes, loong64 `JAL`, the x86 conditional-jump spellings) — so every mnemonic the real assembler accepts is recognised. - `lint`: conservative rules — `unknown-instruction`, `operand-count`, `undefined-label`, `duplicate-label`, `missing-ret`, `missing-textflag-include`, `abi-argsize` and `unreachable-code`. Macro invocations are recognised (in-file `#define` names and underscore identifiers) and the label/RET heuristics are suppressed in macro-using files. `abi-argsize` parses the `// func` signature with the Go parser and checks the declared TEXT argument size against Go's ABI0 layout; `unreachable-code` flags dead code after `RET`, suppressed where reachability is undecidable (PC-relative jumps, register-indirect branches, `#ifdef`). Register liveness is computed by dataflow over the control-flow graph (basic blocks, def/use, iterative backward iteration) and drives `register-clobber`, an audit that flags a callee-saved register written but never saved/restored. `funcdata-pcdata` validates the structure of `FUNCDATA`/`PCDATA` directives. Zero error-severity diagnostics across the 90-file Go runtime corpus and the production go-flac kernels (the `register-clobber` audit additionally reports the go-flac kernels' unsaved callee-saved register use for review). - `format`: an idempotent canonical formatter (operand spacing and per-function mnemonic alignment) that preserves comments and round-trips through the parser. - `lsp`: a Language Server Protocol server over stdio providing completion, hover documentation, document symbols, publish-diagnostics and semantic-token highlighting. - `asm`: a standalone amd64 (x86-64) assembler — an instruction encoder (REX/ ModR-M/SIB/displacement/immediate plus the scalar instruction set, and VEX/ AVX2 SIMD across three operand forms — NDS, reg/rm and immediate-shift — covering the bulk of the integer SIMD set) validated by round-trip decoding against `golang.org/x/arch`, and an assembler that drives the parser's AST into the encoder with local-label resolution and `FP`/`SP` frame mapping (plus Go prologue/epilogue generation), producing output byte-identical to the Go assembler for the supported operand forms. - `cmd/gasm`: the `gasm` binary with `tokens`, `parse`, `fmt`, `lint`, `asm` and `lsp` subcommands. - `_gen`: the generator that rebuilds the architecture instruction tables from the Go toolchain source (`just gen`).