Files
gasm-sdk/CHANGELOG.md
T

723 lines
34 KiB
Markdown
Raw Normal View History

# Changelog
All notable changes to gasm-devkit are documented here.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
and this project adheres to [Conventional Commits](https://www.conventionalcommits.org/).
## [development]
Unreleased changes on the `development` branch.
## [0.27.0] — 2026-08-01
Subprocess isolation for `--fuzz`: each function is fuzzed in its own child
process, so a partial function (decoder) that faults on random garbage is
reported as "CRASH (partial function, use --ground-truth)" without killing
the parent. CRASH is informational (exit 0); only MISMATCH is an error.
### Fixed
- `gasm verify --fuzz` no longer crashes the process on partial functions.
## [0.26.0] — 2026-07-31
Universal differential fuzzing: `gasm verify --fuzz` needs no hand-written
reference. It parses the `// func` signature from the assembly source,
generates typed random inputs (slices with random content, ints, pointers to
fixed arrays), JIT-executes BOTH the gasm-assembled and the go-tool-asm-
assembled versions with independent buffer copies, and compares the result
area bit-for-bit.
### Added
- `verify`: `FuzzFunc` / `ExtractSignatures` / `parseFuncSig` — universal
differential fuzz driven by the conventional `// func` comment. Each
version gets its own buffer set (deep copy) so functions that write to
their arguments (histogram increments) don't corrupt the other's input.
- `gasm verify --fuzz [-n N]`: runs the differential fuzz for every function
with a parseable signature. Total functions (wideCopy, pack16, decorrelate,
analyze, autocorr) pass; partial functions (decoders that fault on malformed
input) should use `--ground-truth` instead.
### Known limitation
`--fuzz` crashes the process for partial functions (e.g. LZ4 decoders) whose
over-copy paths read past the buffer on random garbage input. Subprocess
isolation (fork per function) is planned. Use `--ground-truth` for decoders.
## [0.25.0] — 2026-07-30
Universal ground-truth verification: `gasm verify --ground-truth` assembles
any `.s` file with both gasm and `go tool asm`, then compares the machine
code byte-for-byte per function (relocation sites masked). No hand-written
reference needed — the Go toolchain IS the oracle.
### Added
- `verify`: `GroundTruth` — shells out to `go tool asm`, parses the GOOBJ
output (minimal reader: block offsets, nonpkg symbol table, data index)
and returns per-function code bytes.
- `gasm verify --ground-truth`: compares gasm's output against the Go
assembler's, reporting MATCH/MISMATCH per function with the first
differing byte. Relocation disp32 fields (static-symbol references the
linker fills) are masked before comparison.
- Verified: go-lz4 AVX2 2/2, go-flac AVX2 17/17 functions byte-identical.
## [0.24.0] — 2026-07-29
The full analyze family and stereo PCM decode are now differentially tested.
15 of 17 go-flac AVX2 kernels have bit-for-bit differential coverage; the
two remaining (autocorrAVX2 — FMA reassociation, lpcResidualAVX2 — complex
multi-arg) are deferred.
### Added
- `verify`: `analyzeO3RangeAVX2` and `analyzeO4RangeAVX2` differential tests
(200 iterations each, same harness as O1/O2/Res).
- `verify`: `decodeStereo16AVX2` differential test (500 random interleaved
stereo PCM buffers, both channels compared sample-by-sample).
## [0.23.0] — 2026-07-28
The analyze family and 24-bit PCM decode join the differential suite.
### Added
- `verify`: `analyzeO2RangeAVX2` and `analyzeResRangeAVX2` differential
tests (200 iterations each, shared harness with O1: zigzag fold, partial
sum, overflow flag and Len32 histogram).
- `verify`: `decodeMono24AVX2` differential test (500 random 24-bit PCM
buffers, sign-extension compared sample-by-sample).
## [0.22.0] — 2026-07-27
The remaining go-flac encoder kernels join the differential suite.
### Added
- `verify`: `analyzeO1RangeAVX2` differential test (300 random partitions:
zigzag fold, partial sum, overflow flag and the 32-bin Len32 histogram
compared element-by-element against the portable Go reference).
- `verify`: `fastStereoSumsAVX2` differential test (300 random stereo
frames: the four zigzag-fold entropy sums compared against the scalar
loop).
### Verified
- `gasm fmt` doc-comment indentation confirmed correct: comments before
every TEXT are at column 0 (the RET-detection logic handles multi-exit
functions).
## [0.21.0] — 2026-07-26
Differential testing extended to all four production kernels and the CLI
exposes the full dynamic-analysis toolkit.
### Added
- `verify`: go-flac AVX2 differential tests — `decodeMono16AVX2` (500
random PCM buffers), `pack16AVX2` (500 random int32→int16 packings) and
all four decorrelation kernels (200 iterations each: left-side, side-right,
mid-side, interleave) compared bit-for-bit against the portable Go
references.
- `verify`: go-lz4 AVX-512 differential tests — `decodeBlockAVX512` (3 000
fuzzed LZ4 blocks + known answers) and `wideCopyAVX512` (0–1024 bytes)
against the same portable oracle as the AVX2 suite.
- `gasm verify --abi`: runs each NOSPLIT function with sentinel registers
and a red-zone canary, reporting violations.
- `gasm verify --profile`: lists the static basic-block count per function.
## [0.20.0] — 2026-07-25
Coverage profiling: the third pillar of Phase 3. Static basic-block
enumeration from the assembler's label map, combined with multi-input path
diversity measurement — how many observationally distinct execution paths a
test corpus exercises.
### Added
- `verify`: `Kernel.Blocks` / `Kernel.BlockCount` — enumerate basic blocks
from the assembler's local-label map (every jump target is a block
boundary; the function entry is always a block). `decodeBlockAVX2` has
27 blocks.
- `verify`: `Kernel.ProfilePaths` — run the function with a corpus of
argument blocks and collect distinct output fingerprints (the result
words); reports path diversity as a lower bound on code coverage.
### Note
INT3-based per-block hit counting was prototyped but deferred: Go's runtime
signal management (sigaltstack, handler re-installation) makes raw
rt_sigaction handlers fragile in a Go process. The static + path-diversity
approach delivers the project's goal (proving the SIMD path and tail handling
execute) without fighting the runtime.
## [0.19.0] — 2026-07-24
Runtime ABI checks: the second pillar of Phase 3. The JIT trampoline now
has an ABI-checking variant that sets sentinels in the callee-saved registers
(BP, R14) before entering the assembled function and verifies they survive on
return, plus a red-zone canary (128 bytes below SP filled with 0xA5) that
detects any illegal write below the stack pointer.
### Added
- `verify`: `CallChecked` / `Kernel.CallFuncChecked` — ABI-checking JIT call
with sentinel registers and red-zone canary; returns an `ABIReport`
(BPClobbered, R14Clobbered, RedZoneHit).
- `verify`: the raw `leaveJITCheckedRaw` trampoline — a TEXT symbol with no
ABIInternal wrapper (address obtained via GLOBL/DATA), so the JIT
function's RET lands directly in the check code and sees the registers
exactly as the function left them.
- Tests: deliberate BP/R14 clobberers detected; both go-lz4 kernels
confirmed ABI-clean (BP preserved, R14 preserved, red zone intact).
## [0.18.0] — 2026-07-23
Differential testing: the JIT-assembled go-lz4 `decodeBlockAVX2` kernel is
fuzzed against a portable Go reference — 5 000 valid LZ4 blocks compared
bit-for-bit, plus 2 000 hostile (random garbage) inputs with matching error
codes. This is the automated form of the project's bit-identical contract.
### Added
- `verify`: differential fuzz tests — a random LZ4 block generator produces
valid blocks (literals, overlapping matches, extension bytes) and the
JIT-assembled kernel's output is compared byte-for-byte against a portable
Go decoder; a hostile-input suite confirms error-code agreement on random
garbage (no crashes, same classification).
## [0.17.0] — 2026-07-22
Phase 3 begins: dynamic analysis. A JIT execution substrate that assembles
Plan 9 amd64 kernels into executable memory and calls them directly — pure Go
(stdlib only, `syscall.Mmap` + an assembly trampoline), no cgo, no external
toolchain.
### Added
- `verify` package: JIT infrastructure — `Map` copies machine code into a
W^X memory mapping, `Call` invokes it through an ABI0 trampoline that
switches to a prepared stack and back. `Load`/`LoadSource`/`LoadAST`
parse, assemble and map a `.s` file in one step; `Kernel.CallFunc`
marshals the argument block and returns results.
- `gasm verify` subcommand: assembles a file, JIT-loads it and reports the
available functions; with `-smoke`, calls each NOSPLIT function with
zeroed arguments to confirm the trampoline works end-to-end.
- Integration tests: the go-lz4 `decodeBlockAVX2` and `wideCopyAVX2`
kernels (699 and 146 bytes) assemble, map and execute correctly —
known-answer LZ4 blocks decode bit-for-bit, wide copies of 0–1024 bytes
match, malformed input returns the correct error codes.
### Verified
- `just test` (race, 84.6 % total coverage, verify 82.2 %).
- `gasm verify` on both go-lz4 kernels: all functions JIT-load and
smoke-test clean.
## [0.16.0] — 2026-07-21
The scalar conversions between vector and general-purpose registers — the
last of the amd64 EVEX instruction set.
### Added
- `asm`: the GPR-interchanging conversions, byte for byte against the Go
assembler (28 ground-truth cases including memory sources and extended
GPRs): vector to GPR — the signed and truncated VCVT{,T}S{D,S}2SI{,Q}
in both VEX and EVEX, and the unsigned VCVT{,T}S{D,S}2USI{L,Q}
(EVEX only); GPR to vector — VCVTSI2SD{L,Q}/VCVTSI2SS{L,Q} (VEX and
EVEX) and VCVTUSI2SD{L,Q}/VCVTUSI2SS{L,Q} (EVEX only), whose preserved
vector source sits in vvvv (three Plan 9 operands).
## [0.15.0] — 2026-07-20
The last of the EVEX conversions and narrowing/extending moves — the EVEX
instruction set is now complete save for the GPR-interchanging forms.
### Added
- `asm`: the unsigned and truncating conversions — VCVTPD2PS (and the X/Y
spellings, whose length the spelling fixes), VCVTPD2UDQ (X/Y),
VCVTTPD2UDQ (X/Y), VCVTTPD2UQQ, VCVTPS2UDQ, VCVTTPS2UDQ, VCVTPS2UQQ,
VCVTTPS2UQQ, VCVTTPD2QQ, VCVTTPS2QQ, VCVTUQQ2PD, VCVTUQQ2PS (X/Y) and
VCVTQQ2PS X/Y.
- `asm`: the remaining sign/zero-extending moves (VPMOVSXBD/BQ/WQ and
VPMOVZXBD/BQ/WD/WQ, VEX and EVEX) and the complete signed and unsigned
narrowing stores (VPMOVS{DB,QB,DW,QW,QD,WB}, VPMOVUS{DB,QB,DW,QW,QD,WB},
VPMOVDB, VPMOVQW).
- `asm`: the mask/vector conversions (VPMOVM2B/W/D/Q and VPMOVB2M/W2M/
D2M/Q2M), whose K register is a genuine operand rather than a mask and
which therefore take no masking suffixes.
## [0.14.0] — 2026-07-19
The floating-point helper and conversion tail of the AVX-512 set, plus
gather and scatter with VSIB addressing — every encoding verified byte for
byte against the Go assembler.
### Added
- `asm`: the floating-point helpers — reciprocals and reciprocal square
roots (VRCP14/VRSQRT14 PD/PS/SD/SS), exponents and mantissas (VGETEXP*,
VGETMANT*), scaling by powers of two (VSCALEF*), rounding (VRNDSCALE*),
reduction (VREDUCE*), immediate fixup (VFIXUPIMM*) and range selection
(VRANGE*), and floating-point class tests (VFPCLASSPD/PS X/Y/Z and
VFPCLASSSD/SS — a new immediate form whose reg field carries the opmask
destination).
- `asm`: **gather and scatter with VSIB addressing.** The gathers take
both Go spellings: the VEX form with a vector mask register (OP mask,
vsib, dst) and the EVEX form with an explicit K mask (OP vsib, K, dst),
where the EVEX L'L field follows the VSIB index register rather than the
data register (a ZMM index with an YMM destination encodes L'L = 10, as
the Go assembler emits). The scatters (VSCATTER*/VPSCATTER*) are EVEX
only (OP src, K, vsib). All eight gather and eight scatter widths.
- `asm`: the remaining conversions — VCVTQQ2PS (the 512-bit source sets
the length), VCVTPD2QQ/UQQ, VCVTPS2QQ, VCVTUDQ2PD/PS, the half-precision
VCVTPH2PS and VCVTPS2PH (the extract layout with an immediate).
## [0.13.0] — 2026-07-18
The wider AVX-512 set: ternary logic, permutes, compares, expand/compress,
the opmask instructions and the EVEX rounding/SAE/broadcast suffixes — every
encoding verified byte for byte against the Go assembler.
### Added
- `asm`: the wider EVEX/AVX-512 set, across roughly sixty new ground-truth
cases: ternary logic (VPTERNLOGD/Q), the lane shuffles/inserts/extracts
(VSHUF{F,I}{32,64}X{2,4}, the VINSERT*/VEXTRACT* {F,I}{32,64}X{2,4,8}
family, VPALIGNR), compares with an opmask destination (VCMPPD/PS/SD/SS —
a new NDS-plus-immediate form with the K register in the reg field), the
permutes (VPERMB/W, VPERMI2/T2 D/Q/PD), the wider integer families
(VPMADDWD/UBSW, VPMULHUW, VPACKSSWB/USWB/SSDW/USDW, VPABS B/W/D/Q, the
VPROL*/VPROR* rotates, the word shifts and the EVEX W1 qword shifts),
expand/compress (VEXPANDPD/PS, VPEXPANDD/Q, VCOMPRESSPD/PS, VPCOMPRESSD/
Q), the broadcasts (VPBROADCASTB/W from a GPR or memory, VBROADCASTSS/
SD), the opmask-register instructions (KAND/KOR/KXNOR/KADD/KUNPCK/KNOT/
KSHIFTL/KORTEST B/W/D/Q and KMOVQ, whose width the L/W/pp bits select),
the packed single arithmetic (VADD/VSUB/VMUL/VDIV/VMIN/VMAX PS), the
aligned moves (VMOVAPS/APD, VMOVDQA32/64, VMOVSS), the replicating moves
(VMOVSLDUP/VMOVSHDUP), the conversions (VCVTPS2DQ, VCVTTPS2DQ) and the
remaining extending and narrowing moves (VPMOVSXBW, VPMOVZXBW, VPMOVWB,
VPMOVQB).
- `asm`: the EVEX mnemonic suffixes the Go assembler accepts — the rounding
modes `.RN_SAE`, `.RD_SAE`, `.RU_SAE`, `.RZ_SAE` (the EVEX b bit with the
rounding control in L'L), suppress-all-exceptions `.SAE`, and memory
broadcast `.BCST` (the b bit, the vector length preserved, disp8×N scaled
by the element size) — each combinable with the `.Z` zeroing suffix,
validated against the Go assembler's bytes, and rejected on instructions
that do not support them.
## [0.12.0] — 2026-07-17
GOOBJ emission: gasm-assembled functions drop into a `go build` without the
Go assembler.
### Added
- `asm`: **GOOBJ object output.** `gasm asm --format goobj -p <pkgpath>`
writes the Go toolchain's own object format — the one `cmd/link` consumes
directly: the functions as non-package symbols qualified with the package
path (exactly as `cmd/asm` records assembly symbols), the `GLOBL` data,
one serialized `FuncInfo` per function (argument/frame sizes, the asm
func flag, the start line, the file table) and the four pc-value tables
(`pcsp`, `pcfile`, `pcline`, `pcinline`). The `pcsp` table carries the
real stack deltas: the assembler now tracks every stack-adjustment
boundary through the prologue (`PUSHQ BP`, `SUBQ $frame, SP`) and each
`RET`'s epilogue, so frame-pointer functions unwind correctly. The
object preamble — the version-and-experiment header the linker compares
verbatim — is captured from the installed `go tool asm`, so the output is
always consistent with the toolchain that links it.
- `asm`: relocations against file-local `GLOBL` symbols become `R_PCREL`
entries in the GOOBJ output, with the instruction's displacement field
left zero for the linker to fill (as `cmd/asm` leaves it).
### Fixed
- `parser`: 64-bit `DATA` literals above `MaxInt64`
(`DATA mask<>+8(SB)/8, $0x800f…`) parse as unsigned and keep their bit
pattern, instead of being rejected as non-integer.
### Verified
- End-to-end: a gasm-emitted GOOBJ swapped into a `go build` in place of
the toolchain's assembly object links and runs with output identical to
the baseline binary (stack-argument calls and a `GLOBL` relocation
resolved by the Go linker). All 17 go-flac AVX2 kernel functions emit as
a GOOBJ that `go tool nm` reads back with every symbol intact.
## [0.11.0] — 2026-07-16
Linkable object output: external symbols and relocatable ELF / Mach-O
objects.
### Added
- `asm`: **object-file emission.** `gasm asm --format elf` writes an
ELF64 relocatable object and `--format macho` a Mach-O x86-64
`MH_OBJECT`: a code section (`.text` / `__TEXT,__text`) and a data
section (`.data` / `__DATA,__data`), a symbol table with one symbol per
`TEXT` and `GLOBL` (file-local `<>` symbols local, the rest global), and
one PC-relative relocation per static-symbol reference
(`R_X86_64_PC32` / `X86_64_RELOC_SIGNED`, the −4 addend the form needs).
The ELF output is verified end-to-end: a gasm-emitted object links with
a C driver and runs, resolving both a file-local constant and an
external symbol; the Mach-O output is verified structurally with
`debug/macho`.
- `asm`: **external symbol references.** A reference to a symbol no
`GLOBL` in the file defines no longer aborts assembly — it is recorded
as an external relocation (`Image.Externals`, `FuncLayout.Relocs`) and
becomes an undefined global symbol in the object output. The raw image
format (`--format raw`, the default) still reports them: only an object
file can represent a reference the linker must resolve.
### Changed
- `gasm asm` takes a `--format raw|elf|macho` flag selecting what `-o`
writes; without `--format` the behaviour is unchanged (the concatenated
image).
## [0.10.0] — 2026-07-15
The EVEX floating-point and conversion set: the packed-double arithmetic,
the scalar SD/SS forms, VMOVDDUP and the width-changing conversions, each
verified byte for byte against the Go assembler.
### Added
- `asm`: the rest of the common EVEX/VEX floating-point set — packed double
arithmetic (VSUBPD, VDIVPD, VMINPD, VMAXPD, VUNPCKLPD and the EVEX form of
VUNPCKHPD), the scalar double and single operations (VSUBSD, VDIVSD,
VMINSD, VMAXSD and the full VADDSS/VSUBSS/VMULSS/VDIVSS/VMINSS/VMAXSS
family in both VEX and EVEX — the EVEX scalar forms exist for masked and
zeroing use), and VMOVDDUP (lane duplication, VEX and EVEX).
- `asm`: the width-changing conversions — VCVTDQ2PS and VCVTPS2PD (VEX and
EVEX; the destination sets the length for PS→PD), the EVEX form of
VCVTDQ2PD, and the packed-double → dword family: VCVTPD2DQ/VCVTTPD2DQ
(EVEX-512 only, a ZMM source and an XMM destination) and their X/Y
spellings (VCVTPD2DQX/Y, VCVTTPD2DQX/Y), whose length follows the wider
source — a new operand form, since the destination is always XMM while
VEX.L / EVEX.L'L ride with the source (fixed by the spelling even for a
memory source).
- `asm`: masking and zeroing on every new form — the scalar SD/SS
arithmetic, the unpacks, VMOVDDUP and the conversions all accept the
explicit K1–K7 operand and the `.Z` suffix the way Go writes them.
### Documented
- VCVTPS2PD follows the Go assembler's encoding, which omits the F3
mandatory prefix (VEX.pp / EVEX.pp = 00) that Intel's maps prescribe; the
Go toolchain's machine code is the project's byte-for-byte oracle, and
gasm reproduces it exactly (and round-trips through the x86 decoder, which
shares the convention).
### Verified
- 58 new ground-truth cases — every instruction extracted from the Go
toolchain's own assembly (go build + an executable-segment dump), checked
byte for byte and round-tripped through the decoder, covering disp8×N for
the scalar (×8/×4), duplication (×8/×32/×64) and conversion (×8/×16/×32)
memory operands, the 5-bit register fields and the masked/zeroing P2
byte. All four go-flac/go-lz4 kernels still assemble byte-identically
and lint clean.
## [0.9.0] — 2026-07-14
AVX-512 masking and a wider EVEX integer set.
### Added
- `asm`: **EVEX masking** the way Go writes it — an explicit `K1`–`K7`
operand placed among the operands (merging mask), and a `.Z` mnemonic
suffix for zeroing (`VPADDD.Z Z1, Z2, K2, Z3`). Supported across the NDS,
reg/rm, immediate-shift, align, extract, convert and move forms, including
masked comparisons with a K destination (`VPCMPEQD Z0, Z3, K2, K1`). K0 is
rejected as an explicit mask, and `.Z` without a mask is an error, matching
the Go assembler.
- `asm`: the common AVX-512 F/BW integer set — VPADDB/W, VPSUBB/W, VPANDD/Q,
VPANDND/Q, VPMULLW, VPAVGB/W, the signed/unsigned min/max family for
B/W/D/Q elements, the variable shifts VPSLLVD/Q, VPSRLVD/Q, VPSRAVD/Q, the
EVEX forms of VPSHUFD/VPSHUFB, and the VMOVDQU8/VMOVDQU16 move aliases.
Register indices 16–31 encode correctly (the mod=11 quirk carries rm[4]
in X̄). All verified byte for byte against the Go assembler.
- `lint`: masked EVEX forms (`.Z` suffix, K operands) are recognised by
`unknown-instruction` and exempted from `operand-count`.
### Fixed
- `asm`: EVEX register–register operands with indices 16–31 encoded rm[4]
into B̄ instead of X̄ (the EVEX mod=11 extension quirk), producing wrong
prefix bytes for X16+/Y16+ r/m operands.
## [0.8.0] — 2026-07-13
Standard CLI ergonomics.
### Added
- `gasm --help` prints a proper top-level help (description, commands,
flags, examples), and every subcommand now answers `-h`/`--help` with its
own usage block (usage line, description, flag defaults), exiting 0. An
unknown command points at `gasm --help` instead of dumping the whole usage.
### Changed
- The version is primarily available as the standard `gasm --version` / `-V`
flag; the `gasm version` spelling remains as an alias.
## [0.7.0] — 2026-07-12
The formatter behaves like `go fmt` and canonicalises block separation.
### Added
- `gasm fmt` now works like `go fmt`: with no arguments — or with a directory
argument — it reformats every `.s` file below it in place and lists the
changed files, skipping `.` and `_` directories (`.git`, `_refs`, …).
Explicit file arguments keep the `-w` / standard-output behaviour.
### Changed
- `format`: canonical blank-line layout — a new block (a label, `TEXT` or
`GLOBL`) is preceded by exactly one blank line, neither more nor less.
Comments leading a block stay with it (the blank line goes before them),
stacked labels share their block, the function's first label keeps hugging
its `TEXT`, and runs of blank lines collapse to one. The output remains
idempotent and round-trips through the parser. All four go-flac/go-lz4
kernels were reformatted with this release and remain byte-identical when
assembled.
## [0.6.0] — 2026-07-11
Calibrated to the Go ABI: `register-clobber` stops reporting legal code, and
the encoder learns the legacy SSE moves.
### Changed
- `lint`: **`register-clobber` is now calibrated to the Go ABI**
(`cmd/compile/abi-internal.md`), not the platform ABI. Go's stack-based
ABI0 has no System V style callee-saved registers — amd64 `BX`, `R12`–`R15`
and the arm64/riscv64/loong64 scratch sets are caller-saved or permanent
scratch, and hand-written kernels may clobber them freely. The rule now
audits only the registers Go fixes across calls: the frame pointer and the
goroutine pointer (amd64 `BP`/`R14`, arm64 `R18`/`R28`/`R29`, riscv64
`X27`, loong64 `R22`), and the goroutine pointer is reported only when the
function can reach the runtime (is not `NOSPLIT` or makes a call) — the
ABI0 transition restores it on those paths, and NOSPLIT call-free leaves
may use it, exactly as the runtime's own assembly does. Both go-flac
kernels now lint with zero diagnostics.
### Fixed
- `lint`: the liveness analysis took the destination operand to be the
*first* operand on arm64, riscv64 and loong64; Plan 9 spelling puts it last
on every architecture Go supports. The def/use and save/restore
classification on those architectures was inverted.
- `format`: a comment that follows a `RET` (typically the next function's doc
comment) is no longer indented as if it were still inside the finished
function body.
### Added
- `asm`: the legacy (non-VEX) SSE moves — `MOVOU`/`MOVO` (the Plan 9 names
for MOVDQU/MOVDQA), `MOVUPS`/`MOVAPS`/`MOVUPD`/`MOVAPD` and the scalar
`MOVSD`/`MOVSS` — and `VMOVDQU64` in the EVEX set. All verified byte for
byte against the Go assembler.
## [0.5.0] — 2026-07-10
EVEX / AVX-512: the go-flac AVX-512 kernel now assembles, byte-identically to
the Go toolchain, completing the production-kernel coverage.
### Added
- `asm`: **EVEX (AVX-512) encoding** — the four-byte EVEX prefix with the
5-bit register fields (Z0–Z31, X/Y 16–31, with the reg-r/m X̄ quirk and
V'̄ shared between vvvv and the SIB index), opmask registers (K0–K7) as
operands and as mask destinations, and the compressed disp8×N displacement
(the multiplier follows the memory operand's size, as the Go assembler's
opcode tables prescribe). Covers every AVX-512 instruction the go-flac
kernels use: VPXORD/Q, VPADDD, VPSUBD/Q, VPUNPCK*DQ, VPMULLD/Q, VPERMD,
VPSLLD/VPSRAD/VPSRAQ, VALIGND, VPCMPEQD (K destination), VMOVDQU32,
VMOVUPD, VCVTQQ2PD, VPMOVSXDQ, the narrowing stores VPMOVDW/VPMOVQD, the
extracts VEXTRACTI64X4/VEXTRACTF64X4, VFMADD231PD, VADDPD, VMULPD, the
broadcasts VPBROADCASTD/Q (GPR and memory sources take different opcodes)
and the mask moves KMOVW/KTESTW. Masking/zeroing suffixes are out of scope
— the kernels use neither.
- `asm`: `AssembleFile` now accepts file-defined global (`non-<>`) symbols
too; a reference is external only when no `GLOBL` in the file defines it.
### Fixed
- `asm`: registers X16–Y31 force the EVEX encoding of dual-form mnemonics;
previously a `VPBROADCASTD AX, Y30` fell into the VEX encoder, which cannot
represent indices above 15 and silently truncated them.
- `asm`: the VEX encoder now rejects vector register indices 16–31 instead of
encoding a truncated (wrong) register.
### Verified
- All 10 functions of the go-flac `avx512_amd64.s` kernel assemble
byte-identically to the Go toolchain's machine code (the disp32 of the one
`VMOVDQU32 idx16(SB), Z13` load is linker-filled in Go and resolved within
gasm's own image — checked to reach the right constant bytes). The AVX2
kernel's 17 functions remain byte-identical.
## [0.4.0] — 2026-07-09
The standalone assembler reaches the whole go-flac AVX2 kernel: static
symbols assemble, and all 17 kernel functions now match the Go toolchain's
machine code byte for byte.
### Added
- `asm`: **file-level assembly** — `AssembleFile` turns a parsed file into an
`Image`: the function bodies in source order followed by a data section
built from the file's `GLOBL`/`DATA` directives (each symbol 16-aligned).
- `asm`: **static-symbol (`SB`) operands** — `mask<>(SB)` references encode as
RIP-relative loads with a patched disp32, resolved against the image layout
so the output is self-consistent and position-independent. External
(non-file-local) symbols are rejected with a clear error: they need
object-file emission.
- `gasm asm` prints the data section and symbol map alongside the functions
and writes the whole image (code + data) with `-o`.
### Verified
- All 17 functions of the go-flac `avx2_amd64.s` kernel assemble
byte-identically to the Go toolchain's machine code; the only differing
bytes are the displacements of the two `VMOVDQU mask24<>(SB), X15` loads,
which the Go linker fills at link time and gasm resolves within its own
image (checked to reach the right constant bytes).
## [0.3.0] — 2026-07-08
The assembler reaches byte-identical parity with the Go toolchain on the
production go-flac AVX2 kernels: every one of the 15 kernel functions that
avoid global symbols now assembles to exactly the Go assembler's bytes (the
two holdouts load a file-local constant through `SB` and wait on relocation
support).
### Added
- `asm`: the scalar instruction families the kernels use — `CMOVcc` and
`SETcc` (conditions spelled exactly like the jumps), `LZCNT`/`TZCNT`
(legacy `F3 0F BD/BC`), the sign/zero-extending moves (`MOVBLZX`, `MOVBQZX`,
`MOVWLZX`, `MOVWQZX`, `MOVWLSX`, `MOVLQSX`), `CVTSL2SD`/`CVTSQ2SD` (the
legacy SSE encoding, as the Go assembler emits it), the traditional
three-operand `IMUL3{W,L,Q}`, and the variable-count vector shifts
(`VPSRLQ X0, Y8, Y8` — the count in an XMM register or memory takes the
ordinary NDS form).
- `asm`: **jump relaxation** — jumps start in the short (rel8) form and
expand to rel32 when the settled displacement does not fit, iterating the
layout to a fixed point (CALL is always rel32).
- `asm`: **jump-to-jump folding** — a conditional jump to a label whose only
instruction is an unconditional jump is redirected to the ultimate
target, replicating the Go toolchain's linker, which chases such chains
before it encodes branches.
- `parser`: leading negative displacements with a base and index
(`LEAQ -4(DX)(R9*4), R9`) parse into a fully populated address.
### Fixed
- `asm`: `CMP` with a register or memory operand computed **second − first**
instead of first − second, silently inverting every condition that followed
(`CMPQ SI, R10; JGE` tested R10 ≥ SI). The encoding now always records
first − second — `CMP r/m, r` with the first operand in r/m, `CMP r, r/m`
with the first operand in reg — and is byte-identical to the Go assembler.
- `asm`: register-to-register `MOV` now uses the `r/m ← r` opcode (reg =
source), the Go assembler's choice; the output is byte-identical.
## [0.2.0] — 2026-07-07
The Phase 2 assembler grows the SIMD set: shuffles, extract/insert, permute
and the moves, on top of the Phase 1 VEX forms.
### Added
- `asm`: four new VEX (AVX/AVX2) operand forms, each validated by round-trip
decoding through `golang.org/x/arch` **and** byte-for-byte against the
machine code the real Go assembler emits:
- the immediate shuffle (`VPSHUFD`, `VPERMQ`),
- the three-operand-plus-immediate form (`VSHUFPD`, `VPERM2I128`,
`VINSERTI128`),
- the lane extract (`VEXTRACTI128`, `VEXTRACTF128` — the YMM source occupies
the ModRM.reg field, the XMM/memory destination the r/m field),
- the direction-sensitive moves (`VMOVDQU`, `VMOVUPD`, `VMOVD`, `VMOVQ`,
`VMOVSD` — each direction picks its own opcode and VEX.W; a vector→vector
move uses the store-form layout, matching the Go assembler),
- the no-operand `VZEROUPPER`, and `VPERMD` in the NDS form,
- the floating-point and FMA set (`VADDPD`, `VMULPD`, `VXORPD`,
`VUNPCKHPD`, the scalar `VADDSD`/`VMULSD`, `VCVTDQ2PD`, `VFMADD231PD`).
With the scalar set and the earlier NDS / reg-rm / immediate-shift forms,
the encoder now covers every integer, shuffle and FP instruction the
go-flac AVX2 kernels use.
- `asm`: `CMP` accepts the immediate in the second operand position
(`CMPL CX, $31`) — the spelling the Go assembler accepts — encoding it
identically to the immediate-first form.
### Fixed
- `asm`: an unused VEX.vvvv field is now stored as `1111` (v̄vvv = 1111), as
the hardware requires — the previous value (`0000`) made the two-operand
reg/rm forms (VPMOVSXWD, VPBROADCASTD, VMOVMSKPS, …) raise #UD on real CPUs
and differ from the Go assembler's bytes. The round-trip decoder ignores
the field on these instructions, which is why the byte-for-byte Go
comparison (added this release) is now part of the test suite.
## [0.1.0] — 2026-07-06
Initial release — the Phase 1 foundation.
### Added
- `token`, `lexer`, `ast`, `parser`: a hand-written, error-tolerant front end
for Plan 9 assembly. The lexer splices C-preprocessor line continuations
(`\` before a newline) so multi-line `#define` macros parse as one opaque
directive. Validated against the production AVX2/AVX-512 kernels in
`go-libraries/go-flac` and the Go runtime's `src/runtime/*.s` for all four
architectures, with zero parse errors.
- `arch`: register files and **complete** instruction tables for amd64,
arm64, riscv64 and loong64, with the middle-dot symbol separator and static
(`<>`) symbols. Instruction names are generated from the Go toolchain's own
assembler source (`just gen`) — the `anames` opcode lists plus the common
opcodes and the per-architecture front-end aliases (arm64 `B`/`BL`, the
`.P`/`.W` addressing suffixes, loong64 `JAL`, the x86 conditional-jump
spellings) — so every mnemonic the real assembler accepts is recognised.
- `lint`: conservative rules — `unknown-instruction`, `operand-count`,
`undefined-label`, `duplicate-label`, `missing-ret`,
`missing-textflag-include`, `abi-argsize` and `unreachable-code`. Macro
invocations are recognised (in-file `#define` names and underscore
identifiers) and the label/RET heuristics are suppressed in macro-using
files. `abi-argsize` parses the `// func` signature with the Go parser and
checks the declared TEXT argument size against Go's ABI0 layout;
`unreachable-code` flags dead code after `RET`, suppressed where reachability
is undecidable (PC-relative jumps, register-indirect branches, `#ifdef`).
Register liveness is computed by dataflow over the control-flow graph (basic
blocks, def/use, iterative backward iteration) and drives `register-clobber`,
an audit that flags a callee-saved register written but never saved/restored.
`funcdata-pcdata` validates the structure of `FUNCDATA`/`PCDATA` directives.
Zero error-severity diagnostics across the 90-file Go runtime corpus and the
production go-flac kernels (the `register-clobber` audit additionally reports
the go-flac kernels' unsaved callee-saved register use for review).
- `format`: an idempotent canonical formatter (operand spacing and per-function
mnemonic alignment) that preserves comments and round-trips through the
parser.
- `lsp`: a Language Server Protocol server over stdio providing completion,
hover documentation, document symbols, publish-diagnostics and semantic-token
highlighting.
- `asm`: a standalone amd64 (x86-64) assembler — an instruction encoder (REX/
ModR-M/SIB/displacement/immediate plus the scalar instruction set, and VEX/
AVX2 SIMD across three operand forms — NDS, reg/rm and immediate-shift —
covering the bulk of the integer SIMD set) validated by round-trip decoding
against `golang.org/x/arch`, and an assembler that drives the parser's AST
into the encoder with local-label resolution and `FP`/`SP` frame mapping
(plus Go prologue/epilogue generation), producing output byte-identical to the
Go assembler for the supported operand forms.
- `cmd/gasm`: the `gasm` binary with `tokens`, `parse`, `fmt`, `lint`, `asm`
and `lsp` subcommands.
- `_gen`: the generator that rebuilds the architecture instruction tables from
the Go toolchain source (`just gen`).