# Changelog All notable changes to gasm-devkit are documented here. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and this project adheres to [Conventional Commits](https://www.conventionalcommits.org/). ## [development] Unreleased changes on the `development` branch. ## [0.29.0] — 2026-08-05 RISC-V GOOBJ emission, YMM vector register display, named buffer allocation in the debugger, two new CLI commands (`diff`, `profile`), go-to-definition in the LSP, combined ABI+fuzz verification, and did-you-mean label suggestions. ### Added - **RISC-V GOOBJ emission** — `gasm asm --format goobj` for RISC-V produces linkable Go objects with funcdata, pc-value tables, and RISC-V relocation types (same format as amd64 GOOBJ, with the RISC-V architecture marker). - **`gasm diff`** — compare the machine code of two assembly files byte-for-byte; shows which functions differ and the first few differing bytes. - **`gasm profile`** — show the basic-block structure of each function: labels, offsets, frame size, and NOSPLIT flag. - **LSP go-to-definition** — `textDocument/definition` navigates from a label reference to its definition. - **did-you-mean** — when the RISC-V assembler encounters an undefined label, it suggests the closest existing label using Levenshtein distance. - **YMM vector register display** — `regs` in the debugger now shows YMM registers via `PTRACE_GETFPREGS` (falls back to XMM when XSAVE is unavailable). - **Named buffer allocation** — `gasm debug --buf name:size:pattern` allocates buffers in the debuggee filled with `zero`, `ones`, `seq`, or a hex pattern; buffer pointers are placed into the argument block at the matching positions. - **Crash input storage** — `FuzzResult.CrashInput` stores the input that caused a crash or mismatch for reproducibility. - **ABI + fuzz combined** — `gasm verify --fuzz` now runs ABI checks (sentinel registers, canary, stack bounds) alongside differential fuzz testing. ### Verified - RISC-V GOOBJ output links correctly with the Go toolchain. - `gasm diff` detects byte-level differences in real assembly kernels. - LSP go-to-definition resolves labels across functions and files. - Debugger YMM display confirmed on AVX2-capable hardware. ## [0.28.0] — 2026-08-03 RISC-V encoder: full RV64IMAFDC instruction set with RVC compression, MOV pseudo-instruction, SB/global symbol references, ELF64 object emission, and ground-truth verification against `GOARCH=riscv64 go tool asm`. ### Added - **RISC-V encoder** — RV64I, RV64M, RV64A, RV64F/D, FMA, CSR, JALR. - **MOV pseudo-instruction** — load, store, reg-to-reg, immediate, frame mapping. - **RVC compression** — 22 compressed instruction types (C.LDSP, C.SDSP, C.FLDSP, C.FSDSP, C.ADDI, C.LI, C.LUI, C.ADDIW, C.MV, C.ADD, C.SUB, C.XOR, C.OR, C.AND, C.SLLI, C.SRLI, C.SRAI, C.ANDI, C.BEQZ, C.BNEZ, C.J, C.JR). - **SB/global symbols** — `MOV $sym(SB)`, `MOV sym(SB)`, `MOV rd, sym(SB)` encoded as AUIPC pairs with R_RISCV_PCREL_HI20/LO12 relocations. - **GLOBL/DATA** — data section layout in `AssembleFileRISCV`. - **ELF64 emission** — `gasm asm --format elf` produces EM_RISCV objects (.text, .data, .symtab, .rela.text). - **`gasm verify --ground-truth`** — byte-exact comparison against `GOARCH=riscv64 go tool asm`. - **`gasm verify --profile`** — function layout listing for RISC-V. - **CALL** — AUIPC + JALR pair encoding. ### Fixed - Parser: bare-number offset before `(SP)` no longer misidentified as pseudo. - MOV: `MOV $sym(FP/SP), rd` now returns an explicit error instead of silent fallback. - RVC: C.LDSP/C.SDSP/FLDSP/FSDSP immediate encoding now matches Go toolchain (bit-interleaved format). ### Verified - 118 RISC-V tests, asm coverage 83.3%. - Ground-truth: C.LDSP, C.SDSP, C.FLDSP, C.FSDSP byte-exact vs Go toolchain. ## [0.27.0] — 2026-08-01 Subprocess isolation for `--fuzz`: each function is fuzzed in its own child process, so a partial function (decoder) that faults on random garbage is reported as "CRASH (partial function, use --ground-truth)" without killing the parent. CRASH is informational (exit 0); only MISMATCH is an error. ### Fixed - `gasm verify --fuzz` no longer crashes the process on partial functions. ## [0.26.0] — 2026-07-31 Universal differential fuzzing: `gasm verify --fuzz` needs no hand-written reference. It parses the `// func` signature from the assembly source, generates typed random inputs (slices with random content, ints, pointers to fixed arrays), JIT-executes BOTH the gasm-assembled and the go-tool-asm- assembled versions with independent buffer copies, and compares the result area bit-for-bit. ### Added - `verify`: `FuzzFunc` / `ExtractSignatures` / `parseFuncSig` — universal differential fuzz driven by the conventional `// func` comment. Each version gets its own buffer set (deep copy) so functions that write to their arguments (histogram increments) don't corrupt the other's input. - `gasm verify --fuzz [-n N]`: runs the differential fuzz for every function with a parseable signature. Total functions (wideCopy, pack16, decorrelate, analyze, autocorr) pass; partial functions (decoders that fault on malformed input) should use `--ground-truth` instead. ### Known limitation `--fuzz` crashes the process for partial functions (e.g. LZ4 decoders) whose over-copy paths read past the buffer on random garbage input. Subprocess isolation (fork per function) is planned. Use `--ground-truth` for decoders. ## [0.25.0] — 2026-07-30 Universal ground-truth verification: `gasm verify --ground-truth` assembles any `.s` file with both gasm and `go tool asm`, then compares the machine code byte-for-byte per function (relocation sites masked). No hand-written reference needed — the Go toolchain IS the oracle. ### Added - `verify`: `GroundTruth` — shells out to `go tool asm`, parses the GOOBJ output (minimal reader: block offsets, nonpkg symbol table, data index) and returns per-function code bytes. - `gasm verify --ground-truth`: compares gasm's output against the Go assembler's, reporting MATCH/MISMATCH per function with the first differing byte. Relocation disp32 fields (static-symbol references the linker fills) are masked before comparison. - Verified: go-lz4 AVX2 2/2, go-flac AVX2 17/17 functions byte-identical. ## [0.24.0] — 2026-07-29 The full analyze family and stereo PCM decode are now differentially tested. 15 of 17 go-flac AVX2 kernels have bit-for-bit differential coverage; the two remaining (autocorrAVX2 — FMA reassociation, lpcResidualAVX2 — complex multi-arg) are deferred. ### Added - `verify`: `analyzeO3RangeAVX2` and `analyzeO4RangeAVX2` differential tests (200 iterations each, same harness as O1/O2/Res). - `verify`: `decodeStereo16AVX2` differential test (500 random interleaved stereo PCM buffers, both channels compared sample-by-sample). ## [0.23.0] — 2026-07-28 The analyze family and 24-bit PCM decode join the differential suite. ### Added - `verify`: `analyzeO2RangeAVX2` and `analyzeResRangeAVX2` differential tests (200 iterations each, shared harness with O1: zigzag fold, partial sum, overflow flag and Len32 histogram). - `verify`: `decodeMono24AVX2` differential test (500 random 24-bit PCM buffers, sign-extension compared sample-by-sample). ## [0.22.0] — 2026-07-27 The remaining go-flac encoder kernels join the differential suite. ### Added - `verify`: `analyzeO1RangeAVX2` differential test (300 random partitions: zigzag fold, partial sum, overflow flag and the 32-bin Len32 histogram compared element-by-element against the portable Go reference). - `verify`: `fastStereoSumsAVX2` differential test (300 random stereo frames: the four zigzag-fold entropy sums compared against the scalar loop). ### Verified - `gasm fmt` doc-comment indentation confirmed correct: comments before every TEXT are at column 0 (the RET-detection logic handles multi-exit functions). ## [0.21.0] — 2026-07-26 Differential testing extended to all four production kernels and the CLI exposes the full dynamic-analysis toolkit. ### Added - `verify`: go-flac AVX2 differential tests — `decodeMono16AVX2` (500 random PCM buffers), `pack16AVX2` (500 random int32→int16 packings) and all four decorrelation kernels (200 iterations each: left-side, side-right, mid-side, interleave) compared bit-for-bit against the portable Go references. - `verify`: go-lz4 AVX-512 differential tests — `decodeBlockAVX512` (3 000 fuzzed LZ4 blocks + known answers) and `wideCopyAVX512` (0–1024 bytes) against the same portable oracle as the AVX2 suite. - `gasm verify --abi`: runs each NOSPLIT function with sentinel registers and a red-zone canary, reporting violations. - `gasm verify --profile`: lists the static basic-block count per function. ## [0.20.0] — 2026-07-25 Coverage profiling: the third pillar of Phase 3. Static basic-block enumeration from the assembler's label map, combined with multi-input path diversity measurement — how many observationally distinct execution paths a test corpus exercises. ### Added - `verify`: `Kernel.Blocks` / `Kernel.BlockCount` — enumerate basic blocks from the assembler's local-label map (every jump target is a block boundary; the function entry is always a block). `decodeBlockAVX2` has 27 blocks. - `verify`: `Kernel.ProfilePaths` — run the function with a corpus of argument blocks and collect distinct output fingerprints (the result words); reports path diversity as a lower bound on code coverage. ### Note INT3-based per-block hit counting was prototyped but deferred: Go's runtime signal management (sigaltstack, handler re-installation) makes raw rt_sigaction handlers fragile in a Go process. The static + path-diversity approach delivers the project's goal (proving the SIMD path and tail handling execute) without fighting the runtime. ## [0.19.0] — 2026-07-24 Runtime ABI checks: the second pillar of Phase 3. The JIT trampoline now has an ABI-checking variant that sets sentinels in the callee-saved registers (BP, R14) before entering the assembled function and verifies they survive on return, plus a red-zone canary (128 bytes below SP filled with 0xA5) that detects any illegal write below the stack pointer. ### Added - `verify`: `CallChecked` / `Kernel.CallFuncChecked` — ABI-checking JIT call with sentinel registers and red-zone canary; returns an `ABIReport` (BPClobbered, R14Clobbered, RedZoneHit). - `verify`: the raw `leaveJITCheckedRaw` trampoline — a TEXT symbol with no ABIInternal wrapper (address obtained via GLOBL/DATA), so the JIT function's RET lands directly in the check code and sees the registers exactly as the function left them. - Tests: deliberate BP/R14 clobberers detected; both go-lz4 kernels confirmed ABI-clean (BP preserved, R14 preserved, red zone intact). ## [0.18.0] — 2026-07-23 Differential testing: the JIT-assembled go-lz4 `decodeBlockAVX2` kernel is fuzzed against a portable Go reference — 5 000 valid LZ4 blocks compared bit-for-bit, plus 2 000 hostile (random garbage) inputs with matching error codes. This is the automated form of the project's bit-identical contract. ### Added - `verify`: differential fuzz tests — a random LZ4 block generator produces valid blocks (literals, overlapping matches, extension bytes) and the JIT-assembled kernel's output is compared byte-for-byte against a portable Go decoder; a hostile-input suite confirms error-code agreement on random garbage (no crashes, same classification). ## [0.17.0] — 2026-07-22 Phase 3 begins: dynamic analysis. A JIT execution substrate that assembles Plan 9 amd64 kernels into executable memory and calls them directly — pure Go (stdlib only, `syscall.Mmap` + an assembly trampoline), no cgo, no external toolchain. ### Added - `verify` package: JIT infrastructure — `Map` copies machine code into a W^X memory mapping, `Call` invokes it through an ABI0 trampoline that switches to a prepared stack and back. `Load`/`LoadSource`/`LoadAST` parse, assemble and map a `.s` file in one step; `Kernel.CallFunc` marshals the argument block and returns results. - `gasm verify` subcommand: assembles a file, JIT-loads it and reports the available functions; with `-smoke`, calls each NOSPLIT function with zeroed arguments to confirm the trampoline works end-to-end. - Integration tests: the go-lz4 `decodeBlockAVX2` and `wideCopyAVX2` kernels (699 and 146 bytes) assemble, map and execute correctly — known-answer LZ4 blocks decode bit-for-bit, wide copies of 0–1024 bytes match, malformed input returns the correct error codes. ### Verified - `just test` (race, 84.6 % total coverage, verify 82.2 %). - `gasm verify` on both go-lz4 kernels: all functions JIT-load and smoke-test clean. ## [0.16.0] — 2026-07-21 The scalar conversions between vector and general-purpose registers — the last of the amd64 EVEX instruction set. ### Added - `asm`: the GPR-interchanging conversions, byte for byte against the Go assembler (28 ground-truth cases including memory sources and extended GPRs): vector to GPR — the signed and truncated VCVT{,T}S{D,S}2SI{,Q} in both VEX and EVEX, and the unsigned VCVT{,T}S{D,S}2USI{L,Q} (EVEX only); GPR to vector — VCVTSI2SD{L,Q}/VCVTSI2SS{L,Q} (VEX and EVEX) and VCVTUSI2SD{L,Q}/VCVTUSI2SS{L,Q} (EVEX only), whose preserved vector source sits in vvvv (three Plan 9 operands). ## [0.15.0] — 2026-07-20 The last of the EVEX conversions and narrowing/extending moves — the EVEX instruction set is now complete save for the GPR-interchanging forms. ### Added - `asm`: the unsigned and truncating conversions — VCVTPD2PS (and the X/Y spellings, whose length the spelling fixes), VCVTPD2UDQ (X/Y), VCVTTPD2UDQ (X/Y), VCVTTPD2UQQ, VCVTPS2UDQ, VCVTTPS2UDQ, VCVTPS2UQQ, VCVTTPS2UQQ, VCVTTPD2QQ, VCVTTPS2QQ, VCVTUQQ2PD, VCVTUQQ2PS (X/Y) and VCVTQQ2PS X/Y. - `asm`: the remaining sign/zero-extending moves (VPMOVSXBD/BQ/WQ and VPMOVZXBD/BQ/WD/WQ, VEX and EVEX) and the complete signed and unsigned narrowing stores (VPMOVS{DB,QB,DW,QW,QD,WB}, VPMOVUS{DB,QB,DW,QW,QD,WB}, VPMOVDB, VPMOVQW). - `asm`: the mask/vector conversions (VPMOVM2B/W/D/Q and VPMOVB2M/W2M/ D2M/Q2M), whose K register is a genuine operand rather than a mask and which therefore take no masking suffixes. ## [0.14.0] — 2026-07-19 The floating-point helper and conversion tail of the AVX-512 set, plus gather and scatter with VSIB addressing — every encoding verified byte for byte against the Go assembler. ### Added - `asm`: the floating-point helpers — reciprocals and reciprocal square roots (VRCP14/VRSQRT14 PD/PS/SD/SS), exponents and mantissas (VGETEXP*, VGETMANT*), scaling by powers of two (VSCALEF*), rounding (VRNDSCALE*), reduction (VREDUCE*), immediate fixup (VFIXUPIMM*) and range selection (VRANGE*), and floating-point class tests (VFPCLASSPD/PS X/Y/Z and VFPCLASSSD/SS — a new immediate form whose reg field carries the opmask destination). - `asm`: **gather and scatter with VSIB addressing.** The gathers take both Go spellings: the VEX form with a vector mask register (OP mask, vsib, dst) and the EVEX form with an explicit K mask (OP vsib, K, dst), where the EVEX L'L field follows the VSIB index register rather than the data register (a ZMM index with an YMM destination encodes L'L = 10, as the Go assembler emits). The scatters (VSCATTER*/VPSCATTER*) are EVEX only (OP src, K, vsib). All eight gather and eight scatter widths. - `asm`: the remaining conversions — VCVTQQ2PS (the 512-bit source sets the length), VCVTPD2QQ/UQQ, VCVTPS2QQ, VCVTUDQ2PD/PS, the half-precision VCVTPH2PS and VCVTPS2PH (the extract layout with an immediate). ## [0.13.0] — 2026-07-18 The wider AVX-512 set: ternary logic, permutes, compares, expand/compress, the opmask instructions and the EVEX rounding/SAE/broadcast suffixes — every encoding verified byte for byte against the Go assembler. ### Added - `asm`: the wider EVEX/AVX-512 set, across roughly sixty new ground-truth cases: ternary logic (VPTERNLOGD/Q), the lane shuffles/inserts/extracts (VSHUF{F,I}{32,64}X{2,4}, the VINSERT*/VEXTRACT* {F,I}{32,64}X{2,4,8} family, VPALIGNR), compares with an opmask destination (VCMPPD/PS/SD/SS — a new NDS-plus-immediate form with the K register in the reg field), the permutes (VPERMB/W, VPERMI2/T2 D/Q/PD), the wider integer families (VPMADDWD/UBSW, VPMULHUW, VPACKSSWB/USWB/SSDW/USDW, VPABS B/W/D/Q, the VPROL*/VPROR* rotates, the word shifts and the EVEX W1 qword shifts), expand/compress (VEXPANDPD/PS, VPEXPANDD/Q, VCOMPRESSPD/PS, VPCOMPRESSD/ Q), the broadcasts (VPBROADCASTB/W from a GPR or memory, VBROADCASTSS/ SD), the opmask-register instructions (KAND/KOR/KXNOR/KADD/KUNPCK/KNOT/ KSHIFTL/KORTEST B/W/D/Q and KMOVQ, whose width the L/W/pp bits select), the packed single arithmetic (VADD/VSUB/VMUL/VDIV/VMIN/VMAX PS), the aligned moves (VMOVAPS/APD, VMOVDQA32/64, VMOVSS), the replicating moves (VMOVSLDUP/VMOVSHDUP), the conversions (VCVTPS2DQ, VCVTTPS2DQ) and the remaining extending and narrowing moves (VPMOVSXBW, VPMOVZXBW, VPMOVWB, VPMOVQB). - `asm`: the EVEX mnemonic suffixes the Go assembler accepts — the rounding modes `.RN_SAE`, `.RD_SAE`, `.RU_SAE`, `.RZ_SAE` (the EVEX b bit with the rounding control in L'L), suppress-all-exceptions `.SAE`, and memory broadcast `.BCST` (the b bit, the vector length preserved, disp8×N scaled by the element size) — each combinable with the `.Z` zeroing suffix, validated against the Go assembler's bytes, and rejected on instructions that do not support them. ## [0.12.0] — 2026-07-17 GOOBJ emission: gasm-assembled functions drop into a `go build` without the Go assembler. ### Added - `asm`: **GOOBJ object output.** `gasm asm --format goobj -p ` writes the Go toolchain's own object format — the one `cmd/link` consumes directly: the functions as non-package symbols qualified with the package path (exactly as `cmd/asm` records assembly symbols), the `GLOBL` data, one serialized `FuncInfo` per function (argument/frame sizes, the asm func flag, the start line, the file table) and the four pc-value tables (`pcsp`, `pcfile`, `pcline`, `pcinline`). The `pcsp` table carries the real stack deltas: the assembler now tracks every stack-adjustment boundary through the prologue (`PUSHQ BP`, `SUBQ $frame, SP`) and each `RET`'s epilogue, so frame-pointer functions unwind correctly. The object preamble — the version-and-experiment header the linker compares verbatim — is captured from the installed `go tool asm`, so the output is always consistent with the toolchain that links it. - `asm`: relocations against file-local `GLOBL` symbols become `R_PCREL` entries in the GOOBJ output, with the instruction's displacement field left zero for the linker to fill (as `cmd/asm` leaves it). ### Fixed - `parser`: 64-bit `DATA` literals above `MaxInt64` (`DATA mask<>+8(SB)/8, $0x800f…`) parse as unsigned and keep their bit pattern, instead of being rejected as non-integer. ### Verified - End-to-end: a gasm-emitted GOOBJ swapped into a `go build` in place of the toolchain's assembly object links and runs with output identical to the baseline binary (stack-argument calls and a `GLOBL` relocation resolved by the Go linker). All 17 go-flac AVX2 kernel functions emit as a GOOBJ that `go tool nm` reads back with every symbol intact. ## [0.11.0] — 2026-07-16 Linkable object output: external symbols and relocatable ELF / Mach-O objects. ### Added - `asm`: **object-file emission.** `gasm asm --format elf` writes an ELF64 relocatable object and `--format macho` a Mach-O x86-64 `MH_OBJECT`: a code section (`.text` / `__TEXT,__text`) and a data section (`.data` / `__DATA,__data`), a symbol table with one symbol per `TEXT` and `GLOBL` (file-local `<>` symbols local, the rest global), and one PC-relative relocation per static-symbol reference (`R_X86_64_PC32` / `X86_64_RELOC_SIGNED`, the −4 addend the form needs). The ELF output is verified end-to-end: a gasm-emitted object links with a C driver and runs, resolving both a file-local constant and an external symbol; the Mach-O output is verified structurally with `debug/macho`. - `asm`: **external symbol references.** A reference to a symbol no `GLOBL` in the file defines no longer aborts assembly — it is recorded as an external relocation (`Image.Externals`, `FuncLayout.Relocs`) and becomes an undefined global symbol in the object output. The raw image format (`--format raw`, the default) still reports them: only an object file can represent a reference the linker must resolve. ### Changed - `gasm asm` takes a `--format raw|elf|macho` flag selecting what `-o` writes; without `--format` the behaviour is unchanged (the concatenated image). ## [0.10.0] — 2026-07-15 The EVEX floating-point and conversion set: the packed-double arithmetic, the scalar SD/SS forms, VMOVDDUP and the width-changing conversions, each verified byte for byte against the Go assembler. ### Added - `asm`: the rest of the common EVEX/VEX floating-point set — packed double arithmetic (VSUBPD, VDIVPD, VMINPD, VMAXPD, VUNPCKLPD and the EVEX form of VUNPCKHPD), the scalar double and single operations (VSUBSD, VDIVSD, VMINSD, VMAXSD and the full VADDSS/VSUBSS/VMULSS/VDIVSS/VMINSS/VMAXSS family in both VEX and EVEX — the EVEX scalar forms exist for masked and zeroing use), and VMOVDDUP (lane duplication, VEX and EVEX). - `asm`: the width-changing conversions — VCVTDQ2PS and VCVTPS2PD (VEX and EVEX; the destination sets the length for PS→PD), the EVEX form of VCVTDQ2PD, and the packed-double → dword family: VCVTPD2DQ/VCVTTPD2DQ (EVEX-512 only, a ZMM source and an XMM destination) and their X/Y spellings (VCVTPD2DQX/Y, VCVTTPD2DQX/Y), whose length follows the wider source — a new operand form, since the destination is always XMM while VEX.L / EVEX.L'L ride with the source (fixed by the spelling even for a memory source). - `asm`: masking and zeroing on every new form — the scalar SD/SS arithmetic, the unpacks, VMOVDDUP and the conversions all accept the explicit K1–K7 operand and the `.Z` suffix the way Go writes them. ### Documented - VCVTPS2PD follows the Go assembler's encoding, which omits the F3 mandatory prefix (VEX.pp / EVEX.pp = 00) that Intel's maps prescribe; the Go toolchain's machine code is the project's byte-for-byte oracle, and gasm reproduces it exactly (and round-trips through the x86 decoder, which shares the convention). ### Verified - 58 new ground-truth cases — every instruction extracted from the Go toolchain's own assembly (go build + an executable-segment dump), checked byte for byte and round-tripped through the decoder, covering disp8×N for the scalar (×8/×4), duplication (×8/×32/×64) and conversion (×8/×16/×32) memory operands, the 5-bit register fields and the masked/zeroing P2 byte. All four go-flac/go-lz4 kernels still assemble byte-identically and lint clean. ## [0.9.0] — 2026-07-14 AVX-512 masking and a wider EVEX integer set. ### Added - `asm`: **EVEX masking** the way Go writes it — an explicit `K1`–`K7` operand placed among the operands (merging mask), and a `.Z` mnemonic suffix for zeroing (`VPADDD.Z Z1, Z2, K2, Z3`). Supported across the NDS, reg/rm, immediate-shift, align, extract, convert and move forms, including masked comparisons with a K destination (`VPCMPEQD Z0, Z3, K2, K1`). K0 is rejected as an explicit mask, and `.Z` without a mask is an error, matching the Go assembler. - `asm`: the common AVX-512 F/BW integer set — VPADDB/W, VPSUBB/W, VPANDD/Q, VPANDND/Q, VPMULLW, VPAVGB/W, the signed/unsigned min/max family for B/W/D/Q elements, the variable shifts VPSLLVD/Q, VPSRLVD/Q, VPSRAVD/Q, the EVEX forms of VPSHUFD/VPSHUFB, and the VMOVDQU8/VMOVDQU16 move aliases. Register indices 16–31 encode correctly (the mod=11 quirk carries rm[4] in X̄). All verified byte for byte against the Go assembler. - `lint`: masked EVEX forms (`.Z` suffix, K operands) are recognised by `unknown-instruction` and exempted from `operand-count`. ### Fixed - `asm`: EVEX register–register operands with indices 16–31 encoded rm[4] into B̄ instead of X̄ (the EVEX mod=11 extension quirk), producing wrong prefix bytes for X16+/Y16+ r/m operands. ## [0.8.0] — 2026-07-13 Standard CLI ergonomics. ### Added - `gasm --help` prints a proper top-level help (description, commands, flags, examples), and every subcommand now answers `-h`/`--help` with its own usage block (usage line, description, flag defaults), exiting 0. An unknown command points at `gasm --help` instead of dumping the whole usage. ### Changed - The version is primarily available as the standard `gasm --version` / `-V` flag; the `gasm version` spelling remains as an alias. ## [0.7.0] — 2026-07-12 The formatter behaves like `go fmt` and canonicalises block separation. ### Added - `gasm fmt` now works like `go fmt`: with no arguments — or with a directory argument — it reformats every `.s` file below it in place and lists the changed files, skipping `.` and `_` directories (`.git`, `_refs`, …). Explicit file arguments keep the `-w` / standard-output behaviour. ### Changed - `format`: canonical blank-line layout — a new block (a label, `TEXT` or `GLOBL`) is preceded by exactly one blank line, neither more nor less. Comments leading a block stay with it (the blank line goes before them), stacked labels share their block, the function's first label keeps hugging its `TEXT`, and runs of blank lines collapse to one. The output remains idempotent and round-trips through the parser. All four go-flac/go-lz4 kernels were reformatted with this release and remain byte-identical when assembled. ## [0.6.0] — 2026-07-11 Calibrated to the Go ABI: `register-clobber` stops reporting legal code, and the encoder learns the legacy SSE moves. ### Changed - `lint`: **`register-clobber` is now calibrated to the Go ABI** (`cmd/compile/abi-internal.md`), not the platform ABI. Go's stack-based ABI0 has no System V style callee-saved registers — amd64 `BX`, `R12`–`R15` and the arm64/riscv64/loong64 scratch sets are caller-saved or permanent scratch, and hand-written kernels may clobber them freely. The rule now audits only the registers Go fixes across calls: the frame pointer and the goroutine pointer (amd64 `BP`/`R14`, arm64 `R18`/`R28`/`R29`, riscv64 `X27`, loong64 `R22`), and the goroutine pointer is reported only when the function can reach the runtime (is not `NOSPLIT` or makes a call) — the ABI0 transition restores it on those paths, and NOSPLIT call-free leaves may use it, exactly as the runtime's own assembly does. Both go-flac kernels now lint with zero diagnostics. ### Fixed - `lint`: the liveness analysis took the destination operand to be the *first* operand on arm64, riscv64 and loong64; Plan 9 spelling puts it last on every architecture Go supports. The def/use and save/restore classification on those architectures was inverted. - `format`: a comment that follows a `RET` (typically the next function's doc comment) is no longer indented as if it were still inside the finished function body. ### Added - `asm`: the legacy (non-VEX) SSE moves — `MOVOU`/`MOVO` (the Plan 9 names for MOVDQU/MOVDQA), `MOVUPS`/`MOVAPS`/`MOVUPD`/`MOVAPD` and the scalar `MOVSD`/`MOVSS` — and `VMOVDQU64` in the EVEX set. All verified byte for byte against the Go assembler. ## [0.5.0] — 2026-07-10 EVEX / AVX-512: the go-flac AVX-512 kernel now assembles, byte-identically to the Go toolchain, completing the production-kernel coverage. ### Added - `asm`: **EVEX (AVX-512) encoding** — the four-byte EVEX prefix with the 5-bit register fields (Z0–Z31, X/Y 16–31, with the reg-r/m X̄ quirk and V'̄ shared between vvvv and the SIB index), opmask registers (K0–K7) as operands and as mask destinations, and the compressed disp8×N displacement (the multiplier follows the memory operand's size, as the Go assembler's opcode tables prescribe). Covers every AVX-512 instruction the go-flac kernels use: VPXORD/Q, VPADDD, VPSUBD/Q, VPUNPCK*DQ, VPMULLD/Q, VPERMD, VPSLLD/VPSRAD/VPSRAQ, VALIGND, VPCMPEQD (K destination), VMOVDQU32, VMOVUPD, VCVTQQ2PD, VPMOVSXDQ, the narrowing stores VPMOVDW/VPMOVQD, the extracts VEXTRACTI64X4/VEXTRACTF64X4, VFMADD231PD, VADDPD, VMULPD, the broadcasts VPBROADCASTD/Q (GPR and memory sources take different opcodes) and the mask moves KMOVW/KTESTW. Masking/zeroing suffixes are out of scope — the kernels use neither. - `asm`: `AssembleFile` now accepts file-defined global (`non-<>`) symbols too; a reference is external only when no `GLOBL` in the file defines it. ### Fixed - `asm`: registers X16–Y31 force the EVEX encoding of dual-form mnemonics; previously a `VPBROADCASTD AX, Y30` fell into the VEX encoder, which cannot represent indices above 15 and silently truncated them. - `asm`: the VEX encoder now rejects vector register indices 16–31 instead of encoding a truncated (wrong) register. ### Verified - All 10 functions of the go-flac `avx512_amd64.s` kernel assemble byte-identically to the Go toolchain's machine code (the disp32 of the one `VMOVDQU32 idx16(SB), Z13` load is linker-filled in Go and resolved within gasm's own image — checked to reach the right constant bytes). The AVX2 kernel's 17 functions remain byte-identical. ## [0.4.0] — 2026-07-09 The standalone assembler reaches the whole go-flac AVX2 kernel: static symbols assemble, and all 17 kernel functions now match the Go toolchain's machine code byte for byte. ### Added - `asm`: **file-level assembly** — `AssembleFile` turns a parsed file into an `Image`: the function bodies in source order followed by a data section built from the file's `GLOBL`/`DATA` directives (each symbol 16-aligned). - `asm`: **static-symbol (`SB`) operands** — `mask<>(SB)` references encode as RIP-relative loads with a patched disp32, resolved against the image layout so the output is self-consistent and position-independent. External (non-file-local) symbols are rejected with a clear error: they need object-file emission. - `gasm asm` prints the data section and symbol map alongside the functions and writes the whole image (code + data) with `-o`. ### Verified - All 17 functions of the go-flac `avx2_amd64.s` kernel assemble byte-identically to the Go toolchain's machine code; the only differing bytes are the displacements of the two `VMOVDQU mask24<>(SB), X15` loads, which the Go linker fills at link time and gasm resolves within its own image (checked to reach the right constant bytes). ## [0.3.0] — 2026-07-08 The assembler reaches byte-identical parity with the Go toolchain on the production go-flac AVX2 kernels: every one of the 15 kernel functions that avoid global symbols now assembles to exactly the Go assembler's bytes (the two holdouts load a file-local constant through `SB` and wait on relocation support). ### Added - `asm`: the scalar instruction families the kernels use — `CMOVcc` and `SETcc` (conditions spelled exactly like the jumps), `LZCNT`/`TZCNT` (legacy `F3 0F BD/BC`), the sign/zero-extending moves (`MOVBLZX`, `MOVBQZX`, `MOVWLZX`, `MOVWQZX`, `MOVWLSX`, `MOVLQSX`), `CVTSL2SD`/`CVTSQ2SD` (the legacy SSE encoding, as the Go assembler emits it), the traditional three-operand `IMUL3{W,L,Q}`, and the variable-count vector shifts (`VPSRLQ X0, Y8, Y8` — the count in an XMM register or memory takes the ordinary NDS form). - `asm`: **jump relaxation** — jumps start in the short (rel8) form and expand to rel32 when the settled displacement does not fit, iterating the layout to a fixed point (CALL is always rel32). - `asm`: **jump-to-jump folding** — a conditional jump to a label whose only instruction is an unconditional jump is redirected to the ultimate target, replicating the Go toolchain's linker, which chases such chains before it encodes branches. - `parser`: leading negative displacements with a base and index (`LEAQ -4(DX)(R9*4), R9`) parse into a fully populated address. ### Fixed - `asm`: `CMP` with a register or memory operand computed **second − first** instead of first − second, silently inverting every condition that followed (`CMPQ SI, R10; JGE` tested R10 ≥ SI). The encoding now always records first − second — `CMP r/m, r` with the first operand in r/m, `CMP r, r/m` with the first operand in reg — and is byte-identical to the Go assembler. - `asm`: register-to-register `MOV` now uses the `r/m ← r` opcode (reg = source), the Go assembler's choice; the output is byte-identical. ## [0.2.0] — 2026-07-07 The Phase 2 assembler grows the SIMD set: shuffles, extract/insert, permute and the moves, on top of the Phase 1 VEX forms. ### Added - `asm`: four new VEX (AVX/AVX2) operand forms, each validated by round-trip decoding through `golang.org/x/arch` **and** byte-for-byte against the machine code the real Go assembler emits: - the immediate shuffle (`VPSHUFD`, `VPERMQ`), - the three-operand-plus-immediate form (`VSHUFPD`, `VPERM2I128`, `VINSERTI128`), - the lane extract (`VEXTRACTI128`, `VEXTRACTF128` — the YMM source occupies the ModRM.reg field, the XMM/memory destination the r/m field), - the direction-sensitive moves (`VMOVDQU`, `VMOVUPD`, `VMOVD`, `VMOVQ`, `VMOVSD` — each direction picks its own opcode and VEX.W; a vector→vector move uses the store-form layout, matching the Go assembler), - the no-operand `VZEROUPPER`, and `VPERMD` in the NDS form, - the floating-point and FMA set (`VADDPD`, `VMULPD`, `VXORPD`, `VUNPCKHPD`, the scalar `VADDSD`/`VMULSD`, `VCVTDQ2PD`, `VFMADD231PD`). With the scalar set and the earlier NDS / reg-rm / immediate-shift forms, the encoder now covers every integer, shuffle and FP instruction the go-flac AVX2 kernels use. - `asm`: `CMP` accepts the immediate in the second operand position (`CMPL CX, $31`) — the spelling the Go assembler accepts — encoding it identically to the immediate-first form. ### Fixed - `asm`: an unused VEX.vvvv field is now stored as `1111` (v̄vvv = 1111), as the hardware requires — the previous value (`0000`) made the two-operand reg/rm forms (VPMOVSXWD, VPBROADCASTD, VMOVMSKPS, …) raise #UD on real CPUs and differ from the Go assembler's bytes. The round-trip decoder ignores the field on these instructions, which is why the byte-for-byte Go comparison (added this release) is now part of the test suite. ## [0.1.0] — 2026-07-06 Initial release — the Phase 1 foundation. ### Added - `token`, `lexer`, `ast`, `parser`: a hand-written, error-tolerant front end for Plan 9 assembly. The lexer splices C-preprocessor line continuations (`\` before a newline) so multi-line `#define` macros parse as one opaque directive. Validated against the production AVX2/AVX-512 kernels in `go-libraries/go-flac` and the Go runtime's `src/runtime/*.s` for all four architectures, with zero parse errors. - `arch`: register files and **complete** instruction tables for amd64, arm64, riscv64 and loong64, with the middle-dot symbol separator and static (`<>`) symbols. Instruction names are generated from the Go toolchain's own assembler source (`just gen`) — the `anames` opcode lists plus the common opcodes and the per-architecture front-end aliases (arm64 `B`/`BL`, the `.P`/`.W` addressing suffixes, loong64 `JAL`, the x86 conditional-jump spellings) — so every mnemonic the real assembler accepts is recognised. - `lint`: conservative rules — `unknown-instruction`, `operand-count`, `undefined-label`, `duplicate-label`, `missing-ret`, `missing-textflag-include`, `abi-argsize` and `unreachable-code`. Macro invocations are recognised (in-file `#define` names and underscore identifiers) and the label/RET heuristics are suppressed in macro-using files. `abi-argsize` parses the `// func` signature with the Go parser and checks the declared TEXT argument size against Go's ABI0 layout; `unreachable-code` flags dead code after `RET`, suppressed where reachability is undecidable (PC-relative jumps, register-indirect branches, `#ifdef`). Register liveness is computed by dataflow over the control-flow graph (basic blocks, def/use, iterative backward iteration) and drives `register-clobber`, an audit that flags a callee-saved register written but never saved/restored. `funcdata-pcdata` validates the structure of `FUNCDATA`/`PCDATA` directives. Zero error-severity diagnostics across the 90-file Go runtime corpus and the production go-flac kernels (the `register-clobber` audit additionally reports the go-flac kernels' unsaved callee-saved register use for review). - `format`: an idempotent canonical formatter (operand spacing and per-function mnemonic alignment) that preserves comments and round-trips through the parser. - `lsp`: a Language Server Protocol server over stdio providing completion, hover documentation, document symbols, publish-diagnostics and semantic-token highlighting. - `asm`: a standalone amd64 (x86-64) assembler — an instruction encoder (REX/ ModR-M/SIB/displacement/immediate plus the scalar instruction set, and VEX/ AVX2 SIMD across three operand forms — NDS, reg/rm and immediate-shift — covering the bulk of the integer SIMD set) validated by round-trip decoding against `golang.org/x/arch`, and an assembler that drives the parser's AST into the encoder with local-label resolution and `FP`/`SP` frame mapping (plus Go prologue/epilogue generation), producing output byte-identical to the Go assembler for the supported operand forms. - `cmd/gasm`: the `gasm` binary with `tokens`, `parse`, `fmt`, `lint`, `asm` and `lsp` subcommands. - `_gen`: the generator that rebuilds the architecture instruction tables from the Go toolchain source (`just gen`).