44 KiB
Changelog
All notable changes to gasm-devkit are documented here.
The format is based on Keep a Changelog, and this project adheres to Conventional Commits.
[development]
Unreleased changes on the development branch.
Added
- LoongArch encoder (Phase 5).
gasm asmcan now assemble_loong64.sfiles: the full LoongArch64 instruction set with the dual-form arithmetic mnemonics, the 16/21-bit branch families, the MOV pseudo-instruction and its immediate-constant expansions, FP/SP frame mapping, SB/global symbol references (pcalau12i pairs) and ELF64 emission (gasm asm --format elf). Ground-truth verification againstGOARCH=loong64 go tool asmmatches byte-for-byte; GOOBJ emission (gasm asm --format goobj) is proven end-to-end by linking the object into a cross-compiledgo build. - GOOBJ DWARF symbols. The GOOBJ emitters now write the per-function
DWARF symbols the linker requires (the subprogram DIE and the
.debug_lineprogram, byte-identical tocmd/asm's), and the pc-value table deltas are in the architecture's MinLC units as the runtime expects — the amd64 link test now genuinely substitutes the gasm object, and the amd64/loong64 end-to-end GOOBJ link tests pass. - RISC-V GOOBJ emission via the shared emitter. RISC-V GOOBJ output is
now written by the same shared emitter as amd64 and LoongArch, modelling
each AUIPC + second-instruction pair as a single R_RISCV_PCREL_ITYPE/STYPE
relocation (the layout
cmd/asmwrites, not the ELF HI20/LO12 pair), so the object links into a cross-compiledgo buildforGOARCH=riscv64. An end-to-end link test substitutes the gasm object and reads the symbol back withgo tool nm; the rewrite also corrects the relocationafterfield.
Fixed
- RISC-V frame model and RVC encodings. The riscv64 frame layout now
matches
go tool asm: the prologue/epilogue save and restore the link register (LR) instead of S0, with the correct autosize (locals + 8) and the RVC-compressed prologue/epilogue instructions;RETemits the uncompressedJALR X0, 0(X1)the toolchain writes; theC.ADDI/C.LI/C.LUI/C.ADDIWopcode bit and theC.ADDCR-type encoding are fixed; and theLR/TMPregister aliases now resolve to X1 and X31. The pcsp/pcfile/pcline tables are populated from the recorded stack-adjustment and source-line data, and a byte-exact ground-truth test compares framed and leaf functions againstGOARCH=riscv64 go tool asm. - RISC-V operand ordering and RVC compression. R-type instructions now
take
rs2, rs1, rdand I-type arithmetic instructions takeimm12, rs1, rd, matching the Go assembler's documented operand order (previously both were reversed, so non-commutative R-type instructions such asSUBencoded the wrong operation). The two-operand ternary forms (ADD rs2, rd,ADDI $imm, rd,SLLI $shamt, rd) are now accepted. RVC compression is completed forC.ADDI16SP,C.SLLI,C.SRLI,C.SRAI,C.ANDI,C.NOP,C.EBREAK,C.MV(fromADDI/ADD) and the commutativeAND/OR/XORforms; the byte-exact ground-truth test now covers these. - RISC-V compressed loads/stores and word arithmetic. RVC compression
now also covers the register-relative
C.LW/C.SW/C.LD/C.SD/C.FLD/C.FSDforms (in addition to the stack-relativeC.LWSP/C.SWSP/C.LDSP/C.SDSP), plusC.ADDI4SPN,C.ADDWandC.SUBW. The byte-exact ground-truth test exercises these againstGOARCH=riscv64 go tool asm. - RISC-V large-immediate materialisation.
ADDI/ANDI/ORI/XORIwith a 32-bit immediate that does not fit 12 bits now expand exactly ascmd/asm: twoADDIs for the smallADDIsplit range, andLUI+ADDIW+<op>otherwise, with theLUIandADDIWcompressed toC.LUI/C.ADDIWwhen their immediate fits six signed bits. The byte-exact ground-truth test covers positive, negative, and out-of-range immediates againstGOARCH=riscv64 go tool asm. - Debugger watchpoint slots.
gasm debug'swatchcommand always used hardware watchpoint slot 0, so a secondwatchcall silently overwrote the first. Watchpoint slots are now tracked in theSession(DR0–DR3);watchpicks the first free slot and reports an error if all four are in use, andunwatch <slot>clears one (no argument clears all).
Changed
- Linux only. The toolkit, its CI and the released binaries are now Linux-only; cross-compiled to linux/{amd64,arm64,riscv64,loong64}.
- Phase 4 closed. README's "Remaining" list for the debugger is gone; disassembly at PC, memory-write, watchpoints, and source-line mapping are all shipped.
[0.29.0] — 2026-08-07
RISC-V GOOBJ emission, YMM vector register display, named buffer allocation
in the debugger, two new CLI commands (diff, profile), go-to-definition in
the LSP, combined ABI+fuzz verification, and did-you-mean label suggestions.
A --map flag for diff and --call/--buf flags for verify extend the
new CLI commands. A signature-parser fix corrects grouped Go parameters.
Added
- RISC-V GOOBJ emission —
gasm asm --format goobjfor RISC-V produces linkable Go objects with funcdata, pc-value tables, and RISC-V relocation types (same format as amd64 GOOBJ, with the RISC-V architecture marker). gasm diff— compare the machine code of two assembly files byte-for-byte; shows which functions differ and the first few differing bytes.gasm profile— show the basic-block structure of each function: labels, offsets, frame size, and NOSPLIT flag.- LSP go-to-definition —
textDocument/definitionnavigates from a label reference to its definition. - did-you-mean — when the RISC-V assembler encounters an undefined label, it suggests the closest existing label using Levenshtein distance.
- YMM vector register display —
regsin the debugger now shows YMM registers viaPTRACE_GETFPREGS(falls back to XMM when XSAVE is unavailable). - Named buffer allocation —
gasm debug --buf name:size:patternallocates buffers in the debuggee filled withzero,ones,seq, or a hex pattern; buffer pointers are placed into the argument block at the matching positions. - Crash input storage —
FuzzResult.CrashInputstores the input that caused a crash or mismatch for reproducibility. - ABI + fuzz combined —
gasm verify --fuzznow runs ABI checks (sentinel registers, canary, stack bounds) alongside differential fuzz testing. gasm diff --map— compare functions whose names differ between files (e.g.--map wideCopyAVX2=wideCopyAVX512pairs two variants regardless of suffix). Unmapped functions fall back to the original name match.gasm verify --call— invoke a single function with user-supplied buffers (--buf name:size:pattern) instead of the smoke/abi/fuzz sweeps. Patterns:zero,ones,seq, or a hex blob. Useful for partial functions (e.g. decoders) that crash on random input but should succeed on valid data. The arg block is printed before and after the call, showing return values.gasm verify --ground-truthnow documented in--help(was already a flag, just missing from the help text).
Fixed
- Signature parser — grouped Go parameters like
dst, src []byteare now parsed correctly (both get type[]byte). Previously the first name was treated as its own type (dstwith size 8), causing wrong ABI0 arg-block layout in bothverify --calland the fuzzer. - Flaky JIT tests —
runtime.KeepAliveguards and package-level buffers prevent GC from collecting heap objects whose addresses were passed to JIT code viaunsafe.Pointer; all verify tests pass 100/100 under-race.
Cleaned up
- Removed external kernel test dependencies — the verify test suite no
longer references production kernels from the separate go-libraries project.
The remaining test suite uses only
testdata/verify/*.skernels, which are part of this repository. Coverage is identical locally and in CI (80.3 %).
Verified
gasm diffdetects byte-level differences;--mappairs differently-named functions for comparison.gasm verify --callinvokes functions with user-supplied buffers; the arg block is printed before and after the call, showing return values.- LSP go-to-definition resolves labels across functions and files.
[0.28.0] — 2026-08-03
RISC-V encoder: full RV64IMAFDC instruction set with RVC compression, MOV
pseudo-instruction, SB/global symbol references, ELF64 object emission, and
ground-truth verification against GOARCH=riscv64 go tool asm.
Added
- RISC-V encoder — RV64I, RV64M, RV64A, RV64F/D, FMA, CSR, JALR.
- MOV pseudo-instruction — load, store, reg-to-reg, immediate, frame mapping.
- RVC compression — 22 compressed instruction types (C.LDSP, C.SDSP, C.FLDSP, C.FSDSP, C.ADDI, C.LI, C.LUI, C.ADDIW, C.MV, C.ADD, C.SUB, C.XOR, C.OR, C.AND, C.SLLI, C.SRLI, C.SRAI, C.ANDI, C.BEQZ, C.BNEZ, C.J, C.JR).
- SB/global symbols —
MOV $sym(SB),MOV sym(SB),MOV rd, sym(SB)encoded as AUIPC pairs with R_RISCV_PCREL_HI20/LO12 relocations. - GLOBL/DATA — data section layout in
AssembleFileRISCV. - ELF64 emission —
gasm asm --format elfproduces EM_RISCV objects (.text, .data, .symtab, .rela.text). gasm verify --ground-truth— byte-exact comparison againstGOARCH=riscv64 go tool asm.gasm verify --profile— function layout listing for RISC-V.- CALL — AUIPC + JALR pair encoding.
Fixed
- Parser: bare-number offset before
(SP)no longer misidentified as pseudo. - MOV:
MOV $sym(FP/SP), rdnow returns an explicit error instead of silent fallback. - RVC: C.LDSP/C.SDSP/FLDSP/FSDSP immediate encoding now matches Go toolchain (bit-interleaved format).
Verified
- 118 RISC-V tests, asm coverage 83.3%.
- Ground-truth: C.LDSP, C.SDSP, C.FLDSP, C.FSDSP byte-exact vs Go toolchain.
[0.27.0] — 2026-08-01
Subprocess isolation for --fuzz: each function is fuzzed in its own child
process, so a partial function (decoder) that faults on random garbage is
reported as "CRASH (partial function, use --ground-truth)" without killing
the parent. CRASH is informational (exit 0); only MISMATCH is an error.
Fixed
gasm verify --fuzzno longer crashes the process on partial functions.
[0.26.0] — 2026-07-31
Universal differential fuzzing: gasm verify --fuzz needs no hand-written
reference. It parses the // func signature from the assembly source,
generates typed random inputs (slices with random content, ints, pointers to
fixed arrays), JIT-executes BOTH the gasm-assembled and the go-tool-asm-
assembled versions with independent buffer copies, and compares the result
area bit-for-bit.
Added
verify:FuzzFunc/ExtractSignatures/parseFuncSig— universal differential fuzz driven by the conventional// funccomment. Each version gets its own buffer set (deep copy) so functions that write to their arguments (histogram increments) don't corrupt the other's input.gasm verify --fuzz [-n N]: runs the differential fuzz for every function with a parseable signature. Total functions (wideCopy, pack16, decorrelate, analyze, autocorr) pass; partial functions (decoders that fault on malformed input) should use--ground-truthinstead.
Known limitation
--fuzz crashes the process for partial functions (e.g. LZ4 decoders) whose
over-copy paths read past the buffer on random garbage input. Subprocess
isolation (fork per function) is planned. Use --ground-truth for decoders.
[0.25.0] — 2026-07-30
Universal ground-truth verification: gasm verify --ground-truth assembles
any .s file with both gasm and go tool asm, then compares the machine
code byte-for-byte per function (relocation sites masked). No hand-written
reference needed — the Go toolchain IS the oracle.
Added
verify:GroundTruth— shells out togo tool asm, parses the GOOBJ output (minimal reader: block offsets, nonpkg symbol table, data index) and returns per-function code bytes.gasm verify --ground-truth: compares gasm's output against the Go assembler's, reporting MATCH/MISMATCH per function with the first differing byte. Relocation disp32 fields (static-symbol references the linker fills) are masked before comparison.- Verified: go-lz4 AVX2 2/2, go-flac AVX2 17/17 functions byte-identical.
[0.24.0] — 2026-07-29
The full analyze family and stereo PCM decode are now differentially tested. 15 of 17 go-flac AVX2 kernels have bit-for-bit differential coverage; the two remaining (autocorrAVX2 — FMA reassociation, lpcResidualAVX2 — complex multi-arg) are deferred.
Added
verify:analyzeO3RangeAVX2andanalyzeO4RangeAVX2differential tests (200 iterations each, same harness as O1/O2/Res).verify:decodeStereo16AVX2differential test (500 random interleaved stereo PCM buffers, both channels compared sample-by-sample).
[0.23.0] — 2026-07-28
The analyze family and 24-bit PCM decode join the differential suite.
Added
verify:analyzeO2RangeAVX2andanalyzeResRangeAVX2differential tests (200 iterations each, shared harness with O1: zigzag fold, partial sum, overflow flag and Len32 histogram).verify:decodeMono24AVX2differential test (500 random 24-bit PCM buffers, sign-extension compared sample-by-sample).
[0.22.0] — 2026-07-27
The remaining go-flac encoder kernels join the differential suite.
Added
verify:analyzeO1RangeAVX2differential test (300 random partitions: zigzag fold, partial sum, overflow flag and the 32-bin Len32 histogram compared element-by-element against the portable Go reference).verify:fastStereoSumsAVX2differential test (300 random stereo frames: the four zigzag-fold entropy sums compared against the scalar loop).
Verified
gasm fmtdoc-comment indentation confirmed correct: comments before every TEXT are at column 0 (the RET-detection logic handles multi-exit functions).
[0.21.0] — 2026-07-26
Differential testing extended to all four production kernels and the CLI exposes the full dynamic-analysis toolkit.
Added
verify: go-flac AVX2 differential tests —decodeMono16AVX2(500 random PCM buffers),pack16AVX2(500 random int32→int16 packings) and all four decorrelation kernels (200 iterations each: left-side, side-right, mid-side, interleave) compared bit-for-bit against the portable Go references.verify: go-lz4 AVX-512 differential tests —decodeBlockAVX512(3 000 fuzzed LZ4 blocks + known answers) andwideCopyAVX512(0–1024 bytes) against the same portable oracle as the AVX2 suite.gasm verify --abi: runs each NOSPLIT function with sentinel registers and a red-zone canary, reporting violations.gasm verify --profile: lists the static basic-block count per function.
[0.20.0] — 2026-07-25
Coverage profiling: the third pillar of Phase 3. Static basic-block enumeration from the assembler's label map, combined with multi-input path diversity measurement — how many observationally distinct execution paths a test corpus exercises.
Added
verify:Kernel.Blocks/Kernel.BlockCount— enumerate basic blocks from the assembler's local-label map (every jump target is a block boundary; the function entry is always a block).decodeBlockAVX2has 27 blocks.verify:Kernel.ProfilePaths— run the function with a corpus of argument blocks and collect distinct output fingerprints (the result words); reports path diversity as a lower bound on code coverage.
Note
INT3-based per-block hit counting was prototyped but deferred: Go's runtime signal management (sigaltstack, handler re-installation) makes raw rt_sigaction handlers fragile in a Go process. The static + path-diversity approach delivers the project's goal (proving the SIMD path and tail handling execute) without fighting the runtime.
[0.19.0] — 2026-07-24
Runtime ABI checks: the second pillar of Phase 3. The JIT trampoline now has an ABI-checking variant that sets sentinels in the callee-saved registers (BP, R14) before entering the assembled function and verifies they survive on return, plus a red-zone canary (128 bytes below SP filled with 0xA5) that detects any illegal write below the stack pointer.
Added
verify:CallChecked/Kernel.CallFuncChecked— ABI-checking JIT call with sentinel registers and red-zone canary; returns anABIReport(BPClobbered, R14Clobbered, RedZoneHit).verify: the rawleaveJITCheckedRawtrampoline — a TEXT symbol with no ABIInternal wrapper (address obtained via GLOBL/DATA), so the JIT function's RET lands directly in the check code and sees the registers exactly as the function left them.- Tests: deliberate BP/R14 clobberers detected; both go-lz4 kernels confirmed ABI-clean (BP preserved, R14 preserved, red zone intact).
[0.18.0] — 2026-07-23
Differential testing: the JIT-assembled go-lz4 decodeBlockAVX2 kernel is
fuzzed against a portable Go reference — 5 000 valid LZ4 blocks compared
bit-for-bit, plus 2 000 hostile (random garbage) inputs with matching error
codes. This is the automated form of the project's bit-identical contract.
Added
verify: differential fuzz tests — a random LZ4 block generator produces valid blocks (literals, overlapping matches, extension bytes) and the JIT-assembled kernel's output is compared byte-for-byte against a portable Go decoder; a hostile-input suite confirms error-code agreement on random garbage (no crashes, same classification).
[0.17.0] — 2026-07-22
Phase 3 begins: dynamic analysis. A JIT execution substrate that assembles
Plan 9 amd64 kernels into executable memory and calls them directly — pure Go
(stdlib only, syscall.Mmap + an assembly trampoline), no cgo, no external
toolchain.
Added
verifypackage: JIT infrastructure —Mapcopies machine code into a W^X memory mapping,Callinvokes it through an ABI0 trampoline that switches to a prepared stack and back.Load/LoadSource/LoadASTparse, assemble and map a.sfile in one step;Kernel.CallFuncmarshals the argument block and returns results.gasm verifysubcommand: assembles a file, JIT-loads it and reports the available functions; with-smoke, calls each NOSPLIT function with zeroed arguments to confirm the trampoline works end-to-end.- Integration tests: the go-lz4
decodeBlockAVX2andwideCopyAVX2kernels (699 and 146 bytes) assemble, map and execute correctly — known-answer LZ4 blocks decode bit-for-bit, wide copies of 0–1024 bytes match, malformed input returns the correct error codes.
Verified
just test(race, 84.6 % total coverage, verify 82.2 %).gasm verifyon both go-lz4 kernels: all functions JIT-load and smoke-test clean.
[0.16.0] — 2026-07-21
The scalar conversions between vector and general-purpose registers — the last of the amd64 EVEX instruction set.
Added
asm: the GPR-interchanging conversions, byte for byte against the Go assembler (28 ground-truth cases including memory sources and extended GPRs): vector to GPR — the signed and truncated VCVT{,T}S{D,S}2SI{,Q} in both VEX and EVEX, and the unsigned VCVT{,T}S{D,S}2USI{L,Q} (EVEX only); GPR to vector — VCVTSI2SD{L,Q}/VCVTSI2SS{L,Q} (VEX and EVEX) and VCVTUSI2SD{L,Q}/VCVTUSI2SS{L,Q} (EVEX only), whose preserved vector source sits in vvvv (three Plan 9 operands).
[0.15.0] — 2026-07-20
The last of the EVEX conversions and narrowing/extending moves — the EVEX instruction set is now complete save for the GPR-interchanging forms.
Added
asm: the unsigned and truncating conversions — VCVTPD2PS (and the X/Y spellings, whose length the spelling fixes), VCVTPD2UDQ (X/Y), VCVTTPD2UDQ (X/Y), VCVTTPD2UQQ, VCVTPS2UDQ, VCVTTPS2UDQ, VCVTPS2UQQ, VCVTTPS2UQQ, VCVTTPD2QQ, VCVTTPS2QQ, VCVTUQQ2PD, VCVTUQQ2PS (X/Y) and VCVTQQ2PS X/Y.asm: the remaining sign/zero-extending moves (VPMOVSXBD/BQ/WQ and VPMOVZXBD/BQ/WD/WQ, VEX and EVEX) and the complete signed and unsigned narrowing stores (VPMOVS{DB,QB,DW,QW,QD,WB}, VPMOVUS{DB,QB,DW,QW,QD,WB}, VPMOVDB, VPMOVQW).asm: the mask/vector conversions (VPMOVM2B/W/D/Q and VPMOVB2M/W2M/ D2M/Q2M), whose K register is a genuine operand rather than a mask and which therefore take no masking suffixes.
[0.14.0] — 2026-07-19
The floating-point helper and conversion tail of the AVX-512 set, plus gather and scatter with VSIB addressing — every encoding verified byte for byte against the Go assembler.
Added
asm: the floating-point helpers — reciprocals and reciprocal square roots (VRCP14/VRSQRT14 PD/PS/SD/SS), exponents and mantissas (VGETEXP*, VGETMANT*), scaling by powers of two (VSCALEF*), rounding (VRNDSCALE*), reduction (VREDUCE*), immediate fixup (VFIXUPIMM*) and range selection (VRANGE*), and floating-point class tests (VFPCLASSPD/PS X/Y/Z and VFPCLASSSD/SS — a new immediate form whose reg field carries the opmask destination).asm: gather and scatter with VSIB addressing. The gathers take both Go spellings: the VEX form with a vector mask register (OP mask, vsib, dst) and the EVEX form with an explicit K mask (OP vsib, K, dst), where the EVEX L'L field follows the VSIB index register rather than the data register (a ZMM index with an YMM destination encodes L'L = 10, as the Go assembler emits). The scatters (VSCATTER*/VPSCATTER*) are EVEX only (OP src, K, vsib). All eight gather and eight scatter widths.asm: the remaining conversions — VCVTQQ2PS (the 512-bit source sets the length), VCVTPD2QQ/UQQ, VCVTPS2QQ, VCVTUDQ2PD/PS, the half-precision VCVTPH2PS and VCVTPS2PH (the extract layout with an immediate).
[0.13.0] — 2026-07-18
The wider AVX-512 set: ternary logic, permutes, compares, expand/compress, the opmask instructions and the EVEX rounding/SAE/broadcast suffixes — every encoding verified byte for byte against the Go assembler.
Added
asm: the wider EVEX/AVX-512 set, across roughly sixty new ground-truth cases: ternary logic (VPTERNLOGD/Q), the lane shuffles/inserts/extracts (VSHUF{F,I}{32,64}X{2,4}, the VINSERT*/VEXTRACT* {F,I}{32,64}X{2,4,8} family, VPALIGNR), compares with an opmask destination (VCMPPD/PS/SD/SS — a new NDS-plus-immediate form with the K register in the reg field), the permutes (VPERMB/W, VPERMI2/T2 D/Q/PD), the wider integer families (VPMADDWD/UBSW, VPMULHUW, VPACKSSWB/USWB/SSDW/USDW, VPABS B/W/D/Q, the VPROL*/VPROR* rotates, the word shifts and the EVEX W1 qword shifts), expand/compress (VEXPANDPD/PS, VPEXPANDD/Q, VCOMPRESSPD/PS, VPCOMPRESSD/ Q), the broadcasts (VPBROADCASTB/W from a GPR or memory, VBROADCASTSS/ SD), the opmask-register instructions (KAND/KOR/KXNOR/KADD/KUNPCK/KNOT/ KSHIFTL/KORTEST B/W/D/Q and KMOVQ, whose width the L/W/pp bits select), the packed single arithmetic (VADD/VSUB/VMUL/VDIV/VMIN/VMAX PS), the aligned moves (VMOVAPS/APD, VMOVDQA32/64, VMOVSS), the replicating moves (VMOVSLDUP/VMOVSHDUP), the conversions (VCVTPS2DQ, VCVTTPS2DQ) and the remaining extending and narrowing moves (VPMOVSXBW, VPMOVZXBW, VPMOVWB, VPMOVQB).asm: the EVEX mnemonic suffixes the Go assembler accepts — the rounding modes.RN_SAE,.RD_SAE,.RU_SAE,.RZ_SAE(the EVEX b bit with the rounding control in L'L), suppress-all-exceptions.SAE, and memory broadcast.BCST(the b bit, the vector length preserved, disp8×N scaled by the element size) — each combinable with the.Zzeroing suffix, validated against the Go assembler's bytes, and rejected on instructions that do not support them.
[0.12.0] — 2026-07-17
GOOBJ emission: gasm-assembled functions drop into a go build without the
Go assembler.
Added
asm: GOOBJ object output.gasm asm --format goobj -p <pkgpath>writes the Go toolchain's own object format — the onecmd/linkconsumes directly: the functions as non-package symbols qualified with the package path (exactly ascmd/asmrecords assembly symbols), theGLOBLdata, one serializedFuncInfoper function (argument/frame sizes, the asm func flag, the start line, the file table) and the four pc-value tables (pcsp,pcfile,pcline,pcinline). Thepcsptable carries the real stack deltas: the assembler now tracks every stack-adjustment boundary through the prologue (PUSHQ BP,SUBQ $frame, SP) and eachRET's epilogue, so frame-pointer functions unwind correctly. The object preamble — the version-and-experiment header the linker compares verbatim — is captured from the installedgo tool asm, so the output is always consistent with the toolchain that links it.asm: relocations against file-localGLOBLsymbols becomeR_PCRELentries in the GOOBJ output, with the instruction's displacement field left zero for the linker to fill (ascmd/asmleaves it).
Fixed
parser: 64-bitDATAliterals aboveMaxInt64(DATA mask<>+8(SB)/8, $0x800f…) parse as unsigned and keep their bit pattern, instead of being rejected as non-integer.
Verified
- End-to-end: a gasm-emitted GOOBJ swapped into a
go buildin place of the toolchain's assembly object links and runs with output identical to the baseline binary (stack-argument calls and aGLOBLrelocation resolved by the Go linker). All 17 go-flac AVX2 kernel functions emit as a GOOBJ thatgo tool nmreads back with every symbol intact.
[0.11.0] — 2026-07-16
Linkable object output: external symbols and relocatable ELF / Mach-O objects.
Added
asm: object-file emission.gasm asm --format elfwrites an ELF64 relocatable object and--format machoa Mach-O x86-64MH_OBJECT: a code section (.text/__TEXT,__text) and a data section (.data/__DATA,__data), a symbol table with one symbol perTEXTandGLOBL(file-local<>symbols local, the rest global), and one PC-relative relocation per static-symbol reference (R_X86_64_PC32/X86_64_RELOC_SIGNED, the −4 addend the form needs). The ELF output is verified end-to-end: a gasm-emitted object links with a C driver and runs, resolving both a file-local constant and an external symbol; the Mach-O output is verified structurally withdebug/macho.asm: external symbol references. A reference to a symbol noGLOBLin the file defines no longer aborts assembly — it is recorded as an external relocation (Image.Externals,FuncLayout.Relocs) and becomes an undefined global symbol in the object output. The raw image format (--format raw, the default) still reports them: only an object file can represent a reference the linker must resolve.
Changed
gasm asmtakes a--format raw|elf|machoflag selecting what-owrites; without--formatthe behaviour is unchanged (the concatenated image).
[0.10.0] — 2026-07-15
The EVEX floating-point and conversion set: the packed-double arithmetic, the scalar SD/SS forms, VMOVDDUP and the width-changing conversions, each verified byte for byte against the Go assembler.
Added
asm: the rest of the common EVEX/VEX floating-point set — packed double arithmetic (VSUBPD, VDIVPD, VMINPD, VMAXPD, VUNPCKLPD and the EVEX form of VUNPCKHPD), the scalar double and single operations (VSUBSD, VDIVSD, VMINSD, VMAXSD and the full VADDSS/VSUBSS/VMULSS/VDIVSS/VMINSS/VMAXSS family in both VEX and EVEX — the EVEX scalar forms exist for masked and zeroing use), and VMOVDDUP (lane duplication, VEX and EVEX).asm: the width-changing conversions — VCVTDQ2PS and VCVTPS2PD (VEX and EVEX; the destination sets the length for PS→PD), the EVEX form of VCVTDQ2PD, and the packed-double → dword family: VCVTPD2DQ/VCVTTPD2DQ (EVEX-512 only, a ZMM source and an XMM destination) and their X/Y spellings (VCVTPD2DQX/Y, VCVTTPD2DQX/Y), whose length follows the wider source — a new operand form, since the destination is always XMM while VEX.L / EVEX.L'L ride with the source (fixed by the spelling even for a memory source).asm: masking and zeroing on every new form — the scalar SD/SS arithmetic, the unpacks, VMOVDDUP and the conversions all accept the explicit K1–K7 operand and the.Zsuffix the way Go writes them.
Documented
- VCVTPS2PD follows the Go assembler's encoding, which omits the F3 mandatory prefix (VEX.pp / EVEX.pp = 00) that Intel's maps prescribe; the Go toolchain's machine code is the project's byte-for-byte oracle, and gasm reproduces it exactly (and round-trips through the x86 decoder, which shares the convention).
Verified
- 58 new ground-truth cases — every instruction extracted from the Go toolchain's own assembly (go build + an executable-segment dump), checked byte for byte and round-tripped through the decoder, covering disp8×N for the scalar (×8/×4), duplication (×8/×32/×64) and conversion (×8/×16/×32) memory operands, the 5-bit register fields and the masked/zeroing P2 byte. All four go-flac/go-lz4 kernels still assemble byte-identically and lint clean.
[0.9.0] — 2026-07-14
AVX-512 masking and a wider EVEX integer set.
Added
asm: EVEX masking the way Go writes it — an explicitK1–K7operand placed among the operands (merging mask), and a.Zmnemonic suffix for zeroing (VPADDD.Z Z1, Z2, K2, Z3). Supported across the NDS, reg/rm, immediate-shift, align, extract, convert and move forms, including masked comparisons with a K destination (VPCMPEQD Z0, Z3, K2, K1). K0 is rejected as an explicit mask, and.Zwithout a mask is an error, matching the Go assembler.asm: the common AVX-512 F/BW integer set — VPADDB/W, VPSUBB/W, VPANDD/Q, VPANDND/Q, VPMULLW, VPAVGB/W, the signed/unsigned min/max family for B/W/D/Q elements, the variable shifts VPSLLVD/Q, VPSRLVD/Q, VPSRAVD/Q, the EVEX forms of VPSHUFD/VPSHUFB, and the VMOVDQU8/VMOVDQU16 move aliases. Register indices 16–31 encode correctly (the mod=11 quirk carries rm[4] in X̄). All verified byte for byte against the Go assembler.lint: masked EVEX forms (.Zsuffix, K operands) are recognised byunknown-instructionand exempted fromoperand-count.
Fixed
asm: EVEX register–register operands with indices 16–31 encoded rm[4] into B̄ instead of X̄ (the EVEX mod=11 extension quirk), producing wrong prefix bytes for X16+/Y16+ r/m operands.
[0.8.0] — 2026-07-13
Standard CLI ergonomics.
Added
gasm --helpprints a proper top-level help (description, commands, flags, examples), and every subcommand now answers-h/--helpwith its own usage block (usage line, description, flag defaults), exiting 0. An unknown command points atgasm --helpinstead of dumping the whole usage.
Changed
- The version is primarily available as the standard
gasm --version/-Vflag; thegasm versionspelling remains as an alias.
[0.7.0] — 2026-07-12
The formatter behaves like go fmt and canonicalises block separation.
Added
gasm fmtnow works likego fmt: with no arguments — or with a directory argument — it reformats every.sfile below it in place and lists the changed files, skipping.and_directories (.git,_refs, …). Explicit file arguments keep the-w/ standard-output behaviour.
Changed
format: canonical blank-line layout — a new block (a label,TEXTorGLOBL) is preceded by exactly one blank line, neither more nor less. Comments leading a block stay with it (the blank line goes before them), stacked labels share their block, the function's first label keeps hugging itsTEXT, and runs of blank lines collapse to one. The output remains idempotent and round-trips through the parser. All four go-flac/go-lz4 kernels were reformatted with this release and remain byte-identical when assembled.
[0.6.0] — 2026-07-11
Calibrated to the Go ABI: register-clobber stops reporting legal code, and
the encoder learns the legacy SSE moves.
Changed
lint:register-clobberis now calibrated to the Go ABI (cmd/compile/abi-internal.md), not the platform ABI. Go's stack-based ABI0 has no System V style callee-saved registers — amd64BX,R12–R15and the arm64/riscv64/loong64 scratch sets are caller-saved or permanent scratch, and hand-written kernels may clobber them freely. The rule now audits only the registers Go fixes across calls: the frame pointer and the goroutine pointer (amd64BP/R14, arm64R18/R28/R29, riscv64X27, loong64R22), and the goroutine pointer is reported only when the function can reach the runtime (is notNOSPLITor makes a call) — the ABI0 transition restores it on those paths, and NOSPLIT call-free leaves may use it, exactly as the runtime's own assembly does. Both go-flac kernels now lint with zero diagnostics.
Fixed
lint: the liveness analysis took the destination operand to be the first operand on arm64, riscv64 and loong64; Plan 9 spelling puts it last on every architecture Go supports. The def/use and save/restore classification on those architectures was inverted.format: a comment that follows aRET(typically the next function's doc comment) is no longer indented as if it were still inside the finished function body.
Added
asm: the legacy (non-VEX) SSE moves —MOVOU/MOVO(the Plan 9 names for MOVDQU/MOVDQA),MOVUPS/MOVAPS/MOVUPD/MOVAPDand the scalarMOVSD/MOVSS— andVMOVDQU64in the EVEX set. All verified byte for byte against the Go assembler.
[0.5.0] — 2026-07-10
EVEX / AVX-512: the go-flac AVX-512 kernel now assembles, byte-identically to the Go toolchain, completing the production-kernel coverage.
Added
asm: EVEX (AVX-512) encoding — the four-byte EVEX prefix with the 5-bit register fields (Z0–Z31, X/Y 16–31, with the reg-r/m X̄ quirk and V'̄ shared between vvvv and the SIB index), opmask registers (K0–K7) as operands and as mask destinations, and the compressed disp8×N displacement (the multiplier follows the memory operand's size, as the Go assembler's opcode tables prescribe). Covers every AVX-512 instruction the go-flac kernels use: VPXORD/Q, VPADDD, VPSUBD/Q, VPUNPCK*DQ, VPMULLD/Q, VPERMD, VPSLLD/VPSRAD/VPSRAQ, VALIGND, VPCMPEQD (K destination), VMOVDQU32, VMOVUPD, VCVTQQ2PD, VPMOVSXDQ, the narrowing stores VPMOVDW/VPMOVQD, the extracts VEXTRACTI64X4/VEXTRACTF64X4, VFMADD231PD, VADDPD, VMULPD, the broadcasts VPBROADCASTD/Q (GPR and memory sources take different opcodes) and the mask moves KMOVW/KTESTW. Masking/zeroing suffixes are out of scope — the kernels use neither.asm:AssembleFilenow accepts file-defined global (non-<>) symbols too; a reference is external only when noGLOBLin the file defines it.
Fixed
asm: registers X16–Y31 force the EVEX encoding of dual-form mnemonics; previously aVPBROADCASTD AX, Y30fell into the VEX encoder, which cannot represent indices above 15 and silently truncated them.asm: the VEX encoder now rejects vector register indices 16–31 instead of encoding a truncated (wrong) register.
Verified
- All 10 functions of the go-flac
avx512_amd64.skernel assemble byte-identically to the Go toolchain's machine code (the disp32 of the oneVMOVDQU32 idx16(SB), Z13load is linker-filled in Go and resolved within gasm's own image — checked to reach the right constant bytes). The AVX2 kernel's 17 functions remain byte-identical.
[0.4.0] — 2026-07-09
The standalone assembler reaches the whole go-flac AVX2 kernel: static symbols assemble, and all 17 kernel functions now match the Go toolchain's machine code byte for byte.
Added
asm: file-level assembly —AssembleFileturns a parsed file into anImage: the function bodies in source order followed by a data section built from the file'sGLOBL/DATAdirectives (each symbol 16-aligned).asm: static-symbol (SB) operands —mask<>(SB)references encode as RIP-relative loads with a patched disp32, resolved against the image layout so the output is self-consistent and position-independent. External (non-file-local) symbols are rejected with a clear error: they need object-file emission.gasm asmprints the data section and symbol map alongside the functions and writes the whole image (code + data) with-o.
Verified
- All 17 functions of the go-flac
avx2_amd64.skernel assemble byte-identically to the Go toolchain's machine code; the only differing bytes are the displacements of the twoVMOVDQU mask24<>(SB), X15loads, which the Go linker fills at link time and gasm resolves within its own image (checked to reach the right constant bytes).
[0.3.0] — 2026-07-08
The assembler reaches byte-identical parity with the Go toolchain on the
production go-flac AVX2 kernels: every one of the 15 kernel functions that
avoid global symbols now assembles to exactly the Go assembler's bytes (the
two holdouts load a file-local constant through SB and wait on relocation
support).
Added
asm: the scalar instruction families the kernels use —CMOVccandSETcc(conditions spelled exactly like the jumps),LZCNT/TZCNT(legacyF3 0F BD/BC), the sign/zero-extending moves (MOVBLZX,MOVBQZX,MOVWLZX,MOVWQZX,MOVWLSX,MOVLQSX),CVTSL2SD/CVTSQ2SD(the legacy SSE encoding, as the Go assembler emits it), the traditional three-operandIMUL3{W,L,Q}, and the variable-count vector shifts (VPSRLQ X0, Y8, Y8— the count in an XMM register or memory takes the ordinary NDS form).asm: jump relaxation — jumps start in the short (rel8) form and expand to rel32 when the settled displacement does not fit, iterating the layout to a fixed point (CALL is always rel32).asm: jump-to-jump folding — a conditional jump to a label whose only instruction is an unconditional jump is redirected to the ultimate target, replicating the Go toolchain's linker, which chases such chains before it encodes branches.parser: leading negative displacements with a base and index (LEAQ -4(DX)(R9*4), R9) parse into a fully populated address.
Fixed
asm:CMPwith a register or memory operand computed second − first instead of first − second, silently inverting every condition that followed (CMPQ SI, R10; JGEtested R10 ≥ SI). The encoding now always records first − second —CMP r/m, rwith the first operand in r/m,CMP r, r/mwith the first operand in reg — and is byte-identical to the Go assembler.asm: register-to-registerMOVnow uses ther/m ← ropcode (reg = source), the Go assembler's choice; the output is byte-identical.
[0.2.0] — 2026-07-07
The Phase 2 assembler grows the SIMD set: shuffles, extract/insert, permute and the moves, on top of the Phase 1 VEX forms.
Added
asm: four new VEX (AVX/AVX2) operand forms, each validated by round-trip decoding throughgolang.org/x/archand byte-for-byte against the machine code the real Go assembler emits:- the immediate shuffle (
VPSHUFD,VPERMQ), - the three-operand-plus-immediate form (
VSHUFPD,VPERM2I128,VINSERTI128), - the lane extract (
VEXTRACTI128,VEXTRACTF128— the YMM source occupies the ModRM.reg field, the XMM/memory destination the r/m field), - the direction-sensitive moves (
VMOVDQU,VMOVUPD,VMOVD,VMOVQ,VMOVSD— each direction picks its own opcode and VEX.W; a vector→vector move uses the store-form layout, matching the Go assembler), - the no-operand
VZEROUPPER, andVPERMDin the NDS form, - the floating-point and FMA set (
VADDPD,VMULPD,VXORPD,VUNPCKHPD, the scalarVADDSD/VMULSD,VCVTDQ2PD,VFMADD231PD). With the scalar set and the earlier NDS / reg-rm / immediate-shift forms, the encoder now covers every integer, shuffle and FP instruction the go-flac AVX2 kernels use.
- the immediate shuffle (
asm:CMPaccepts the immediate in the second operand position (CMPL CX, $31) — the spelling the Go assembler accepts — encoding it identically to the immediate-first form.
Fixed
asm: an unused VEX.vvvv field is now stored as1111(v̄vvv = 1111), as the hardware requires — the previous value (0000) made the two-operand reg/rm forms (VPMOVSXWD, VPBROADCASTD, VMOVMSKPS, …) raise #UD on real CPUs and differ from the Go assembler's bytes. The round-trip decoder ignores the field on these instructions, which is why the byte-for-byte Go comparison (added this release) is now part of the test suite.
[0.1.0] — 2026-07-06
Initial release — the Phase 1 foundation.
Added
token,lexer,ast,parser: a hand-written, error-tolerant front end for Plan 9 assembly. The lexer splices C-preprocessor line continuations (\before a newline) so multi-line#definemacros parse as one opaque directive. Validated against the production AVX2/AVX-512 kernels ingo-libraries/go-flacand the Go runtime'ssrc/runtime/*.sfor all four architectures, with zero parse errors.arch: register files and complete instruction tables for amd64, arm64, riscv64 and loong64, with the middle-dot symbol separator and static (<>) symbols. Instruction names are generated from the Go toolchain's own assembler source (just gen) — theanamesopcode lists plus the common opcodes and the per-architecture front-end aliases (arm64B/BL, the.P/.Waddressing suffixes, loong64JAL, the x86 conditional-jump spellings) — so every mnemonic the real assembler accepts is recognised.lint: conservative rules —unknown-instruction,operand-count,undefined-label,duplicate-label,missing-ret,missing-textflag-include,abi-argsizeandunreachable-code. Macro invocations are recognised (in-file#definenames and underscore identifiers) and the label/RET heuristics are suppressed in macro-using files.abi-argsizeparses the// funcsignature with the Go parser and checks the declared TEXT argument size against Go's ABI0 layout;unreachable-codeflags dead code afterRET, suppressed where reachability is undecidable (PC-relative jumps, register-indirect branches,#ifdef). Register liveness is computed by dataflow over the control-flow graph (basic blocks, def/use, iterative backward iteration) and drivesregister-clobber, an audit that flags a callee-saved register written but never saved/restored.funcdata-pcdatavalidates the structure ofFUNCDATA/PCDATAdirectives. Zero error-severity diagnostics across the 90-file Go runtime corpus and the production go-flac kernels (theregister-clobberaudit additionally reports the go-flac kernels' unsaved callee-saved register use for review).format: an idempotent canonical formatter (operand spacing and per-function mnemonic alignment) that preserves comments and round-trips through the parser.lsp: a Language Server Protocol server over stdio providing completion, hover documentation, document symbols, publish-diagnostics and semantic-token highlighting.asm: a standalone amd64 (x86-64) assembler — an instruction encoder (REX/ ModR-M/SIB/displacement/immediate plus the scalar instruction set, and VEX/ AVX2 SIMD across three operand forms — NDS, reg/rm and immediate-shift — covering the bulk of the integer SIMD set) validated by round-trip decoding againstgolang.org/x/arch, and an assembler that drives the parser's AST into the encoder with local-label resolution andFP/SPframe mapping (plus Go prologue/epilogue generation), producing output byte-identical to the Go assembler for the supported operand forms.cmd/gasm: thegasmbinary withtokens,parse,fmt,lint,asmandlspsubcommands._gen: the generator that rebuilds the architecture instruction tables from the Go toolchain source (just gen).