816 lines
39 KiB
Markdown
816 lines
39 KiB
Markdown
# Changelog
|
||
|
||
All notable changes to gasm-devkit are documented here.
|
||
|
||
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
|
||
and this project adheres to [Conventional Commits](https://www.conventionalcommits.org/).
|
||
|
||
## [development]
|
||
|
||
Unreleased changes on the `development` branch.
|
||
|
||
## [0.29.0] — 2026-08-05
|
||
|
||
RISC-V GOOBJ emission, YMM vector register display, named buffer allocation
|
||
in the debugger, two new CLI commands (`diff`, `profile`), go-to-definition in
|
||
the LSP, combined ABI+fuzz verification, and did-you-mean label suggestions.
|
||
A `--map` flag for `diff` and `--call`/`--buf` flags for `verify` extend the
|
||
new CLI commands. A signature-parser fix corrects grouped Go parameters.
|
||
|
||
### Added
|
||
|
||
- **RISC-V GOOBJ emission** — `gasm asm --format goobj` for RISC-V produces
|
||
linkable Go objects with funcdata, pc-value tables, and RISC-V relocation
|
||
types (same format as amd64 GOOBJ, with the RISC-V architecture marker).
|
||
- **`gasm diff`** — compare the machine code of two assembly files byte-for-byte;
|
||
shows which functions differ and the first few differing bytes.
|
||
- **`gasm profile`** — show the basic-block structure of each function: labels,
|
||
offsets, frame size, and NOSPLIT flag.
|
||
- **LSP go-to-definition** — `textDocument/definition` navigates from a label
|
||
reference to its definition.
|
||
- **did-you-mean** — when the RISC-V assembler encounters an undefined label, it
|
||
suggests the closest existing label using Levenshtein distance.
|
||
- **YMM vector register display** — `regs` in the debugger now shows YMM
|
||
registers via `PTRACE_GETFPREGS` (falls back to XMM when XSAVE is unavailable).
|
||
- **Named buffer allocation** — `gasm debug --buf name:size:pattern` allocates
|
||
buffers in the debuggee filled with `zero`, `ones`, `seq`, or a hex pattern;
|
||
buffer pointers are placed into the argument block at the matching positions.
|
||
- **Crash input storage** — `FuzzResult.CrashInput` stores the input that caused
|
||
a crash or mismatch for reproducibility.
|
||
- **ABI + fuzz combined** — `gasm verify --fuzz` now runs ABI checks (sentinel
|
||
registers, canary, stack bounds) alongside differential fuzz testing.
|
||
- **`gasm diff --map`** — compare functions whose names differ between files
|
||
(e.g. `--map wideCopyAVX2=wideCopyAVX512` pairs AVX2 and AVX-512 variants
|
||
regardless of suffix). Unmapped functions fall back to the original name match.
|
||
- **`gasm verify --call`** — invoke a single function with user-supplied buffers
|
||
(`--buf name:size:pattern`) instead of the smoke/abi/fuzz sweeps. Patterns:
|
||
`zero`, `ones`, `seq`, or a hex blob. Useful for partial functions (e.g.
|
||
decoders) that crash on random input but should succeed on valid data.
|
||
The arg block is printed before and after the call, showing return values.
|
||
- **`gasm verify --ground-truth`** now documented in `--help` (was already a flag,
|
||
just missing from the help text).
|
||
|
||
### Fixed
|
||
|
||
- **Signature parser** — grouped Go parameters like `dst, src []byte` are now
|
||
parsed correctly (both get type `[]byte`). Previously the first name was
|
||
treated as its own type (`dst` with size 8), causing wrong ABI0 arg-block
|
||
layout in both `verify --call` and the fuzzer.
|
||
|
||
### Verified
|
||
|
||
- RISC-V GOOBJ output links correctly with the Go toolchain.
|
||
- `gasm diff` detects byte-level differences in real assembly kernels.
|
||
- `gasm diff --map` pairs AVX2 and AVX-512 kernels by mapped name.
|
||
- `gasm verify --call` invokes `wideCopyAVX2` and `decodeBlockAVX2` with
|
||
user-supplied buffers; the decoder returns its error code instead of crashing.
|
||
- LSP go-to-definition resolves labels across functions and files.
|
||
- Debugger YMM display confirmed on AVX2-capable hardware.
|
||
|
||
## [0.28.0] — 2026-08-03
|
||
|
||
RISC-V encoder: full RV64IMAFDC instruction set with RVC compression, MOV
|
||
pseudo-instruction, SB/global symbol references, ELF64 object emission, and
|
||
ground-truth verification against `GOARCH=riscv64 go tool asm`.
|
||
|
||
### Added
|
||
|
||
- **RISC-V encoder** — RV64I, RV64M, RV64A, RV64F/D, FMA, CSR, JALR.
|
||
- **MOV pseudo-instruction** — load, store, reg-to-reg, immediate, frame mapping.
|
||
- **RVC compression** — 22 compressed instruction types (C.LDSP, C.SDSP, C.FLDSP,
|
||
C.FSDSP, C.ADDI, C.LI, C.LUI, C.ADDIW, C.MV, C.ADD, C.SUB, C.XOR, C.OR, C.AND,
|
||
C.SLLI, C.SRLI, C.SRAI, C.ANDI, C.BEQZ, C.BNEZ, C.J, C.JR).
|
||
- **SB/global symbols** — `MOV $sym(SB)`, `MOV sym(SB)`, `MOV rd, sym(SB)`
|
||
encoded as AUIPC pairs with R_RISCV_PCREL_HI20/LO12 relocations.
|
||
- **GLOBL/DATA** — data section layout in `AssembleFileRISCV`.
|
||
- **ELF64 emission** — `gasm asm --format elf` produces EM_RISCV objects
|
||
(.text, .data, .symtab, .rela.text).
|
||
- **`gasm verify --ground-truth`** — byte-exact comparison against
|
||
`GOARCH=riscv64 go tool asm`.
|
||
- **`gasm verify --profile`** — function layout listing for RISC-V.
|
||
- **CALL** — AUIPC + JALR pair encoding.
|
||
|
||
### Fixed
|
||
|
||
- Parser: bare-number offset before `(SP)` no longer misidentified as pseudo.
|
||
- MOV: `MOV $sym(FP/SP), rd` now returns an explicit error instead of silent fallback.
|
||
- RVC: C.LDSP/C.SDSP/FLDSP/FSDSP immediate encoding now matches Go toolchain
|
||
(bit-interleaved format).
|
||
|
||
### Verified
|
||
|
||
- 118 RISC-V tests, asm coverage 83.3%.
|
||
- Ground-truth: C.LDSP, C.SDSP, C.FLDSP, C.FSDSP byte-exact vs Go toolchain.
|
||
|
||
## [0.27.0] — 2026-08-01
|
||
|
||
Subprocess isolation for `--fuzz`: each function is fuzzed in its own child
|
||
process, so a partial function (decoder) that faults on random garbage is
|
||
reported as "CRASH (partial function, use --ground-truth)" without killing
|
||
the parent. CRASH is informational (exit 0); only MISMATCH is an error.
|
||
|
||
### Fixed
|
||
|
||
- `gasm verify --fuzz` no longer crashes the process on partial functions.
|
||
|
||
## [0.26.0] — 2026-07-31
|
||
|
||
Universal differential fuzzing: `gasm verify --fuzz` needs no hand-written
|
||
reference. It parses the `// func` signature from the assembly source,
|
||
generates typed random inputs (slices with random content, ints, pointers to
|
||
fixed arrays), JIT-executes BOTH the gasm-assembled and the go-tool-asm-
|
||
assembled versions with independent buffer copies, and compares the result
|
||
area bit-for-bit.
|
||
|
||
### Added
|
||
|
||
- `verify`: `FuzzFunc` / `ExtractSignatures` / `parseFuncSig` — universal
|
||
differential fuzz driven by the conventional `// func` comment. Each
|
||
version gets its own buffer set (deep copy) so functions that write to
|
||
their arguments (histogram increments) don't corrupt the other's input.
|
||
- `gasm verify --fuzz [-n N]`: runs the differential fuzz for every function
|
||
with a parseable signature. Total functions (wideCopy, pack16, decorrelate,
|
||
analyze, autocorr) pass; partial functions (decoders that fault on malformed
|
||
input) should use `--ground-truth` instead.
|
||
|
||
### Known limitation
|
||
|
||
`--fuzz` crashes the process for partial functions (e.g. LZ4 decoders) whose
|
||
over-copy paths read past the buffer on random garbage input. Subprocess
|
||
isolation (fork per function) is planned. Use `--ground-truth` for decoders.
|
||
|
||
## [0.25.0] — 2026-07-30
|
||
|
||
Universal ground-truth verification: `gasm verify --ground-truth` assembles
|
||
any `.s` file with both gasm and `go tool asm`, then compares the machine
|
||
code byte-for-byte per function (relocation sites masked). No hand-written
|
||
reference needed — the Go toolchain IS the oracle.
|
||
|
||
### Added
|
||
|
||
- `verify`: `GroundTruth` — shells out to `go tool asm`, parses the GOOBJ
|
||
output (minimal reader: block offsets, nonpkg symbol table, data index)
|
||
and returns per-function code bytes.
|
||
- `gasm verify --ground-truth`: compares gasm's output against the Go
|
||
assembler's, reporting MATCH/MISMATCH per function with the first
|
||
differing byte. Relocation disp32 fields (static-symbol references the
|
||
linker fills) are masked before comparison.
|
||
- Verified: go-lz4 AVX2 2/2, go-flac AVX2 17/17 functions byte-identical.
|
||
|
||
## [0.24.0] — 2026-07-29
|
||
|
||
The full analyze family and stereo PCM decode are now differentially tested.
|
||
15 of 17 go-flac AVX2 kernels have bit-for-bit differential coverage; the
|
||
two remaining (autocorrAVX2 — FMA reassociation, lpcResidualAVX2 — complex
|
||
multi-arg) are deferred.
|
||
|
||
### Added
|
||
|
||
- `verify`: `analyzeO3RangeAVX2` and `analyzeO4RangeAVX2` differential tests
|
||
(200 iterations each, same harness as O1/O2/Res).
|
||
- `verify`: `decodeStereo16AVX2` differential test (500 random interleaved
|
||
stereo PCM buffers, both channels compared sample-by-sample).
|
||
|
||
## [0.23.0] — 2026-07-28
|
||
|
||
The analyze family and 24-bit PCM decode join the differential suite.
|
||
|
||
### Added
|
||
|
||
- `verify`: `analyzeO2RangeAVX2` and `analyzeResRangeAVX2` differential
|
||
tests (200 iterations each, shared harness with O1: zigzag fold, partial
|
||
sum, overflow flag and Len32 histogram).
|
||
- `verify`: `decodeMono24AVX2` differential test (500 random 24-bit PCM
|
||
buffers, sign-extension compared sample-by-sample).
|
||
|
||
## [0.22.0] — 2026-07-27
|
||
|
||
The remaining go-flac encoder kernels join the differential suite.
|
||
|
||
### Added
|
||
|
||
- `verify`: `analyzeO1RangeAVX2` differential test (300 random partitions:
|
||
zigzag fold, partial sum, overflow flag and the 32-bin Len32 histogram
|
||
compared element-by-element against the portable Go reference).
|
||
- `verify`: `fastStereoSumsAVX2` differential test (300 random stereo
|
||
frames: the four zigzag-fold entropy sums compared against the scalar
|
||
loop).
|
||
|
||
### Verified
|
||
|
||
- `gasm fmt` doc-comment indentation confirmed correct: comments before
|
||
every TEXT are at column 0 (the RET-detection logic handles multi-exit
|
||
functions).
|
||
|
||
## [0.21.0] — 2026-07-26
|
||
|
||
Differential testing extended to all four production kernels and the CLI
|
||
exposes the full dynamic-analysis toolkit.
|
||
|
||
### Added
|
||
|
||
- `verify`: go-flac AVX2 differential tests — `decodeMono16AVX2` (500
|
||
random PCM buffers), `pack16AVX2` (500 random int32→int16 packings) and
|
||
all four decorrelation kernels (200 iterations each: left-side, side-right,
|
||
mid-side, interleave) compared bit-for-bit against the portable Go
|
||
references.
|
||
- `verify`: go-lz4 AVX-512 differential tests — `decodeBlockAVX512` (3 000
|
||
fuzzed LZ4 blocks + known answers) and `wideCopyAVX512` (0–1024 bytes)
|
||
against the same portable oracle as the AVX2 suite.
|
||
- `gasm verify --abi`: runs each NOSPLIT function with sentinel registers
|
||
and a red-zone canary, reporting violations.
|
||
- `gasm verify --profile`: lists the static basic-block count per function.
|
||
|
||
## [0.20.0] — 2026-07-25
|
||
|
||
Coverage profiling: the third pillar of Phase 3. Static basic-block
|
||
enumeration from the assembler's label map, combined with multi-input path
|
||
diversity measurement — how many observationally distinct execution paths a
|
||
test corpus exercises.
|
||
|
||
### Added
|
||
|
||
- `verify`: `Kernel.Blocks` / `Kernel.BlockCount` — enumerate basic blocks
|
||
from the assembler's local-label map (every jump target is a block
|
||
boundary; the function entry is always a block). `decodeBlockAVX2` has
|
||
27 blocks.
|
||
- `verify`: `Kernel.ProfilePaths` — run the function with a corpus of
|
||
argument blocks and collect distinct output fingerprints (the result
|
||
words); reports path diversity as a lower bound on code coverage.
|
||
|
||
### Note
|
||
|
||
INT3-based per-block hit counting was prototyped but deferred: Go's runtime
|
||
signal management (sigaltstack, handler re-installation) makes raw
|
||
rt_sigaction handlers fragile in a Go process. The static + path-diversity
|
||
approach delivers the project's goal (proving the SIMD path and tail handling
|
||
execute) without fighting the runtime.
|
||
|
||
## [0.19.0] — 2026-07-24
|
||
|
||
Runtime ABI checks: the second pillar of Phase 3. The JIT trampoline now
|
||
has an ABI-checking variant that sets sentinels in the callee-saved registers
|
||
(BP, R14) before entering the assembled function and verifies they survive on
|
||
return, plus a red-zone canary (128 bytes below SP filled with 0xA5) that
|
||
detects any illegal write below the stack pointer.
|
||
|
||
### Added
|
||
|
||
- `verify`: `CallChecked` / `Kernel.CallFuncChecked` — ABI-checking JIT call
|
||
with sentinel registers and red-zone canary; returns an `ABIReport`
|
||
(BPClobbered, R14Clobbered, RedZoneHit).
|
||
- `verify`: the raw `leaveJITCheckedRaw` trampoline — a TEXT symbol with no
|
||
ABIInternal wrapper (address obtained via GLOBL/DATA), so the JIT
|
||
function's RET lands directly in the check code and sees the registers
|
||
exactly as the function left them.
|
||
- Tests: deliberate BP/R14 clobberers detected; both go-lz4 kernels
|
||
confirmed ABI-clean (BP preserved, R14 preserved, red zone intact).
|
||
|
||
## [0.18.0] — 2026-07-23
|
||
|
||
Differential testing: the JIT-assembled go-lz4 `decodeBlockAVX2` kernel is
|
||
fuzzed against a portable Go reference — 5 000 valid LZ4 blocks compared
|
||
bit-for-bit, plus 2 000 hostile (random garbage) inputs with matching error
|
||
codes. This is the automated form of the project's bit-identical contract.
|
||
|
||
### Added
|
||
|
||
- `verify`: differential fuzz tests — a random LZ4 block generator produces
|
||
valid blocks (literals, overlapping matches, extension bytes) and the
|
||
JIT-assembled kernel's output is compared byte-for-byte against a portable
|
||
Go decoder; a hostile-input suite confirms error-code agreement on random
|
||
garbage (no crashes, same classification).
|
||
|
||
## [0.17.0] — 2026-07-22
|
||
|
||
Phase 3 begins: dynamic analysis. A JIT execution substrate that assembles
|
||
Plan 9 amd64 kernels into executable memory and calls them directly — pure Go
|
||
(stdlib only, `syscall.Mmap` + an assembly trampoline), no cgo, no external
|
||
toolchain.
|
||
|
||
### Added
|
||
|
||
- `verify` package: JIT infrastructure — `Map` copies machine code into a
|
||
W^X memory mapping, `Call` invokes it through an ABI0 trampoline that
|
||
switches to a prepared stack and back. `Load`/`LoadSource`/`LoadAST`
|
||
parse, assemble and map a `.s` file in one step; `Kernel.CallFunc`
|
||
marshals the argument block and returns results.
|
||
- `gasm verify` subcommand: assembles a file, JIT-loads it and reports the
|
||
available functions; with `-smoke`, calls each NOSPLIT function with
|
||
zeroed arguments to confirm the trampoline works end-to-end.
|
||
- Integration tests: the go-lz4 `decodeBlockAVX2` and `wideCopyAVX2`
|
||
kernels (699 and 146 bytes) assemble, map and execute correctly —
|
||
known-answer LZ4 blocks decode bit-for-bit, wide copies of 0–1024 bytes
|
||
match, malformed input returns the correct error codes.
|
||
|
||
### Verified
|
||
|
||
- `just test` (race, 84.6 % total coverage, verify 82.2 %).
|
||
- `gasm verify` on both go-lz4 kernels: all functions JIT-load and
|
||
smoke-test clean.
|
||
|
||
## [0.16.0] — 2026-07-21
|
||
|
||
The scalar conversions between vector and general-purpose registers — the
|
||
last of the amd64 EVEX instruction set.
|
||
|
||
### Added
|
||
|
||
- `asm`: the GPR-interchanging conversions, byte for byte against the Go
|
||
assembler (28 ground-truth cases including memory sources and extended
|
||
GPRs): vector to GPR — the signed and truncated VCVT{,T}S{D,S}2SI{,Q}
|
||
in both VEX and EVEX, and the unsigned VCVT{,T}S{D,S}2USI{L,Q}
|
||
(EVEX only); GPR to vector — VCVTSI2SD{L,Q}/VCVTSI2SS{L,Q} (VEX and
|
||
EVEX) and VCVTUSI2SD{L,Q}/VCVTUSI2SS{L,Q} (EVEX only), whose preserved
|
||
vector source sits in vvvv (three Plan 9 operands).
|
||
|
||
## [0.15.0] — 2026-07-20
|
||
|
||
The last of the EVEX conversions and narrowing/extending moves — the EVEX
|
||
instruction set is now complete save for the GPR-interchanging forms.
|
||
|
||
### Added
|
||
|
||
- `asm`: the unsigned and truncating conversions — VCVTPD2PS (and the X/Y
|
||
spellings, whose length the spelling fixes), VCVTPD2UDQ (X/Y),
|
||
VCVTTPD2UDQ (X/Y), VCVTTPD2UQQ, VCVTPS2UDQ, VCVTTPS2UDQ, VCVTPS2UQQ,
|
||
VCVTTPS2UQQ, VCVTTPD2QQ, VCVTTPS2QQ, VCVTUQQ2PD, VCVTUQQ2PS (X/Y) and
|
||
VCVTQQ2PS X/Y.
|
||
- `asm`: the remaining sign/zero-extending moves (VPMOVSXBD/BQ/WQ and
|
||
VPMOVZXBD/BQ/WD/WQ, VEX and EVEX) and the complete signed and unsigned
|
||
narrowing stores (VPMOVS{DB,QB,DW,QW,QD,WB}, VPMOVUS{DB,QB,DW,QW,QD,WB},
|
||
VPMOVDB, VPMOVQW).
|
||
- `asm`: the mask/vector conversions (VPMOVM2B/W/D/Q and VPMOVB2M/W2M/
|
||
D2M/Q2M), whose K register is a genuine operand rather than a mask and
|
||
which therefore take no masking suffixes.
|
||
|
||
## [0.14.0] — 2026-07-19
|
||
|
||
The floating-point helper and conversion tail of the AVX-512 set, plus
|
||
gather and scatter with VSIB addressing — every encoding verified byte for
|
||
byte against the Go assembler.
|
||
|
||
### Added
|
||
|
||
- `asm`: the floating-point helpers — reciprocals and reciprocal square
|
||
roots (VRCP14/VRSQRT14 PD/PS/SD/SS), exponents and mantissas (VGETEXP*,
|
||
VGETMANT*), scaling by powers of two (VSCALEF*), rounding (VRNDSCALE*),
|
||
reduction (VREDUCE*), immediate fixup (VFIXUPIMM*) and range selection
|
||
(VRANGE*), and floating-point class tests (VFPCLASSPD/PS X/Y/Z and
|
||
VFPCLASSSD/SS — a new immediate form whose reg field carries the opmask
|
||
destination).
|
||
- `asm`: **gather and scatter with VSIB addressing.** The gathers take
|
||
both Go spellings: the VEX form with a vector mask register (OP mask,
|
||
vsib, dst) and the EVEX form with an explicit K mask (OP vsib, K, dst),
|
||
where the EVEX L'L field follows the VSIB index register rather than the
|
||
data register (a ZMM index with an YMM destination encodes L'L = 10, as
|
||
the Go assembler emits). The scatters (VSCATTER*/VPSCATTER*) are EVEX
|
||
only (OP src, K, vsib). All eight gather and eight scatter widths.
|
||
- `asm`: the remaining conversions — VCVTQQ2PS (the 512-bit source sets
|
||
the length), VCVTPD2QQ/UQQ, VCVTPS2QQ, VCVTUDQ2PD/PS, the half-precision
|
||
VCVTPH2PS and VCVTPS2PH (the extract layout with an immediate).
|
||
|
||
## [0.13.0] — 2026-07-18
|
||
|
||
The wider AVX-512 set: ternary logic, permutes, compares, expand/compress,
|
||
the opmask instructions and the EVEX rounding/SAE/broadcast suffixes — every
|
||
encoding verified byte for byte against the Go assembler.
|
||
|
||
### Added
|
||
|
||
- `asm`: the wider EVEX/AVX-512 set, across roughly sixty new ground-truth
|
||
cases: ternary logic (VPTERNLOGD/Q), the lane shuffles/inserts/extracts
|
||
(VSHUF{F,I}{32,64}X{2,4}, the VINSERT*/VEXTRACT* {F,I}{32,64}X{2,4,8}
|
||
family, VPALIGNR), compares with an opmask destination (VCMPPD/PS/SD/SS —
|
||
a new NDS-plus-immediate form with the K register in the reg field), the
|
||
permutes (VPERMB/W, VPERMI2/T2 D/Q/PD), the wider integer families
|
||
(VPMADDWD/UBSW, VPMULHUW, VPACKSSWB/USWB/SSDW/USDW, VPABS B/W/D/Q, the
|
||
VPROL*/VPROR* rotates, the word shifts and the EVEX W1 qword shifts),
|
||
expand/compress (VEXPANDPD/PS, VPEXPANDD/Q, VCOMPRESSPD/PS, VPCOMPRESSD/
|
||
Q), the broadcasts (VPBROADCASTB/W from a GPR or memory, VBROADCASTSS/
|
||
SD), the opmask-register instructions (KAND/KOR/KXNOR/KADD/KUNPCK/KNOT/
|
||
KSHIFTL/KORTEST B/W/D/Q and KMOVQ, whose width the L/W/pp bits select),
|
||
the packed single arithmetic (VADD/VSUB/VMUL/VDIV/VMIN/VMAX PS), the
|
||
aligned moves (VMOVAPS/APD, VMOVDQA32/64, VMOVSS), the replicating moves
|
||
(VMOVSLDUP/VMOVSHDUP), the conversions (VCVTPS2DQ, VCVTTPS2DQ) and the
|
||
remaining extending and narrowing moves (VPMOVSXBW, VPMOVZXBW, VPMOVWB,
|
||
VPMOVQB).
|
||
- `asm`: the EVEX mnemonic suffixes the Go assembler accepts — the rounding
|
||
modes `.RN_SAE`, `.RD_SAE`, `.RU_SAE`, `.RZ_SAE` (the EVEX b bit with the
|
||
rounding control in L'L), suppress-all-exceptions `.SAE`, and memory
|
||
broadcast `.BCST` (the b bit, the vector length preserved, disp8×N scaled
|
||
by the element size) — each combinable with the `.Z` zeroing suffix,
|
||
validated against the Go assembler's bytes, and rejected on instructions
|
||
that do not support them.
|
||
|
||
## [0.12.0] — 2026-07-17
|
||
|
||
GOOBJ emission: gasm-assembled functions drop into a `go build` without the
|
||
Go assembler.
|
||
|
||
### Added
|
||
|
||
- `asm`: **GOOBJ object output.** `gasm asm --format goobj -p <pkgpath>`
|
||
writes the Go toolchain's own object format — the one `cmd/link` consumes
|
||
directly: the functions as non-package symbols qualified with the package
|
||
path (exactly as `cmd/asm` records assembly symbols), the `GLOBL` data,
|
||
one serialized `FuncInfo` per function (argument/frame sizes, the asm
|
||
func flag, the start line, the file table) and the four pc-value tables
|
||
(`pcsp`, `pcfile`, `pcline`, `pcinline`). The `pcsp` table carries the
|
||
real stack deltas: the assembler now tracks every stack-adjustment
|
||
boundary through the prologue (`PUSHQ BP`, `SUBQ $frame, SP`) and each
|
||
`RET`'s epilogue, so frame-pointer functions unwind correctly. The
|
||
object preamble — the version-and-experiment header the linker compares
|
||
verbatim — is captured from the installed `go tool asm`, so the output is
|
||
always consistent with the toolchain that links it.
|
||
- `asm`: relocations against file-local `GLOBL` symbols become `R_PCREL`
|
||
entries in the GOOBJ output, with the instruction's displacement field
|
||
left zero for the linker to fill (as `cmd/asm` leaves it).
|
||
|
||
### Fixed
|
||
|
||
- `parser`: 64-bit `DATA` literals above `MaxInt64`
|
||
(`DATA mask<>+8(SB)/8, $0x800f…`) parse as unsigned and keep their bit
|
||
pattern, instead of being rejected as non-integer.
|
||
|
||
### Verified
|
||
|
||
- End-to-end: a gasm-emitted GOOBJ swapped into a `go build` in place of
|
||
the toolchain's assembly object links and runs with output identical to
|
||
the baseline binary (stack-argument calls and a `GLOBL` relocation
|
||
resolved by the Go linker). All 17 go-flac AVX2 kernel functions emit as
|
||
a GOOBJ that `go tool nm` reads back with every symbol intact.
|
||
|
||
## [0.11.0] — 2026-07-16
|
||
|
||
Linkable object output: external symbols and relocatable ELF / Mach-O
|
||
objects.
|
||
|
||
### Added
|
||
|
||
- `asm`: **object-file emission.** `gasm asm --format elf` writes an
|
||
ELF64 relocatable object and `--format macho` a Mach-O x86-64
|
||
`MH_OBJECT`: a code section (`.text` / `__TEXT,__text`) and a data
|
||
section (`.data` / `__DATA,__data`), a symbol table with one symbol per
|
||
`TEXT` and `GLOBL` (file-local `<>` symbols local, the rest global), and
|
||
one PC-relative relocation per static-symbol reference
|
||
(`R_X86_64_PC32` / `X86_64_RELOC_SIGNED`, the −4 addend the form needs).
|
||
The ELF output is verified end-to-end: a gasm-emitted object links with
|
||
a C driver and runs, resolving both a file-local constant and an
|
||
external symbol; the Mach-O output is verified structurally with
|
||
`debug/macho`.
|
||
- `asm`: **external symbol references.** A reference to a symbol no
|
||
`GLOBL` in the file defines no longer aborts assembly — it is recorded
|
||
as an external relocation (`Image.Externals`, `FuncLayout.Relocs`) and
|
||
becomes an undefined global symbol in the object output. The raw image
|
||
format (`--format raw`, the default) still reports them: only an object
|
||
file can represent a reference the linker must resolve.
|
||
|
||
### Changed
|
||
|
||
- `gasm asm` takes a `--format raw|elf|macho` flag selecting what `-o`
|
||
writes; without `--format` the behaviour is unchanged (the concatenated
|
||
image).
|
||
|
||
## [0.10.0] — 2026-07-15
|
||
|
||
The EVEX floating-point and conversion set: the packed-double arithmetic,
|
||
the scalar SD/SS forms, VMOVDDUP and the width-changing conversions, each
|
||
verified byte for byte against the Go assembler.
|
||
|
||
### Added
|
||
|
||
- `asm`: the rest of the common EVEX/VEX floating-point set — packed double
|
||
arithmetic (VSUBPD, VDIVPD, VMINPD, VMAXPD, VUNPCKLPD and the EVEX form of
|
||
VUNPCKHPD), the scalar double and single operations (VSUBSD, VDIVSD,
|
||
VMINSD, VMAXSD and the full VADDSS/VSUBSS/VMULSS/VDIVSS/VMINSS/VMAXSS
|
||
family in both VEX and EVEX — the EVEX scalar forms exist for masked and
|
||
zeroing use), and VMOVDDUP (lane duplication, VEX and EVEX).
|
||
- `asm`: the width-changing conversions — VCVTDQ2PS and VCVTPS2PD (VEX and
|
||
EVEX; the destination sets the length for PS→PD), the EVEX form of
|
||
VCVTDQ2PD, and the packed-double → dword family: VCVTPD2DQ/VCVTTPD2DQ
|
||
(EVEX-512 only, a ZMM source and an XMM destination) and their X/Y
|
||
spellings (VCVTPD2DQX/Y, VCVTTPD2DQX/Y), whose length follows the wider
|
||
source — a new operand form, since the destination is always XMM while
|
||
VEX.L / EVEX.L'L ride with the source (fixed by the spelling even for a
|
||
memory source).
|
||
- `asm`: masking and zeroing on every new form — the scalar SD/SS
|
||
arithmetic, the unpacks, VMOVDDUP and the conversions all accept the
|
||
explicit K1–K7 operand and the `.Z` suffix the way Go writes them.
|
||
|
||
### Documented
|
||
|
||
- VCVTPS2PD follows the Go assembler's encoding, which omits the F3
|
||
mandatory prefix (VEX.pp / EVEX.pp = 00) that Intel's maps prescribe; the
|
||
Go toolchain's machine code is the project's byte-for-byte oracle, and
|
||
gasm reproduces it exactly (and round-trips through the x86 decoder, which
|
||
shares the convention).
|
||
|
||
### Verified
|
||
|
||
- 58 new ground-truth cases — every instruction extracted from the Go
|
||
toolchain's own assembly (go build + an executable-segment dump), checked
|
||
byte for byte and round-tripped through the decoder, covering disp8×N for
|
||
the scalar (×8/×4), duplication (×8/×32/×64) and conversion (×8/×16/×32)
|
||
memory operands, the 5-bit register fields and the masked/zeroing P2
|
||
byte. All four go-flac/go-lz4 kernels still assemble byte-identically
|
||
and lint clean.
|
||
|
||
## [0.9.0] — 2026-07-14
|
||
|
||
AVX-512 masking and a wider EVEX integer set.
|
||
|
||
### Added
|
||
|
||
- `asm`: **EVEX masking** the way Go writes it — an explicit `K1`–`K7`
|
||
operand placed among the operands (merging mask), and a `.Z` mnemonic
|
||
suffix for zeroing (`VPADDD.Z Z1, Z2, K2, Z3`). Supported across the NDS,
|
||
reg/rm, immediate-shift, align, extract, convert and move forms, including
|
||
masked comparisons with a K destination (`VPCMPEQD Z0, Z3, K2, K1`). K0 is
|
||
rejected as an explicit mask, and `.Z` without a mask is an error, matching
|
||
the Go assembler.
|
||
- `asm`: the common AVX-512 F/BW integer set — VPADDB/W, VPSUBB/W, VPANDD/Q,
|
||
VPANDND/Q, VPMULLW, VPAVGB/W, the signed/unsigned min/max family for
|
||
B/W/D/Q elements, the variable shifts VPSLLVD/Q, VPSRLVD/Q, VPSRAVD/Q, the
|
||
EVEX forms of VPSHUFD/VPSHUFB, and the VMOVDQU8/VMOVDQU16 move aliases.
|
||
Register indices 16–31 encode correctly (the mod=11 quirk carries rm[4]
|
||
in X̄). All verified byte for byte against the Go assembler.
|
||
- `lint`: masked EVEX forms (`.Z` suffix, K operands) are recognised by
|
||
`unknown-instruction` and exempted from `operand-count`.
|
||
|
||
### Fixed
|
||
|
||
- `asm`: EVEX register–register operands with indices 16–31 encoded rm[4]
|
||
into B̄ instead of X̄ (the EVEX mod=11 extension quirk), producing wrong
|
||
prefix bytes for X16+/Y16+ r/m operands.
|
||
|
||
## [0.8.0] — 2026-07-13
|
||
|
||
Standard CLI ergonomics.
|
||
|
||
### Added
|
||
|
||
- `gasm --help` prints a proper top-level help (description, commands,
|
||
flags, examples), and every subcommand now answers `-h`/`--help` with its
|
||
own usage block (usage line, description, flag defaults), exiting 0. An
|
||
unknown command points at `gasm --help` instead of dumping the whole usage.
|
||
|
||
### Changed
|
||
|
||
- The version is primarily available as the standard `gasm --version` / `-V`
|
||
flag; the `gasm version` spelling remains as an alias.
|
||
|
||
## [0.7.0] — 2026-07-12
|
||
|
||
The formatter behaves like `go fmt` and canonicalises block separation.
|
||
|
||
### Added
|
||
|
||
- `gasm fmt` now works like `go fmt`: with no arguments — or with a directory
|
||
argument — it reformats every `.s` file below it in place and lists the
|
||
changed files, skipping `.` and `_` directories (`.git`, `_refs`, …).
|
||
Explicit file arguments keep the `-w` / standard-output behaviour.
|
||
|
||
### Changed
|
||
|
||
- `format`: canonical blank-line layout — a new block (a label, `TEXT` or
|
||
`GLOBL`) is preceded by exactly one blank line, neither more nor less.
|
||
Comments leading a block stay with it (the blank line goes before them),
|
||
stacked labels share their block, the function's first label keeps hugging
|
||
its `TEXT`, and runs of blank lines collapse to one. The output remains
|
||
idempotent and round-trips through the parser. All four go-flac/go-lz4
|
||
kernels were reformatted with this release and remain byte-identical when
|
||
assembled.
|
||
|
||
## [0.6.0] — 2026-07-11
|
||
|
||
Calibrated to the Go ABI: `register-clobber` stops reporting legal code, and
|
||
the encoder learns the legacy SSE moves.
|
||
|
||
### Changed
|
||
|
||
- `lint`: **`register-clobber` is now calibrated to the Go ABI**
|
||
(`cmd/compile/abi-internal.md`), not the platform ABI. Go's stack-based
|
||
ABI0 has no System V style callee-saved registers — amd64 `BX`, `R12`–`R15`
|
||
and the arm64/riscv64/loong64 scratch sets are caller-saved or permanent
|
||
scratch, and hand-written kernels may clobber them freely. The rule now
|
||
audits only the registers Go fixes across calls: the frame pointer and the
|
||
goroutine pointer (amd64 `BP`/`R14`, arm64 `R18`/`R28`/`R29`, riscv64
|
||
`X27`, loong64 `R22`), and the goroutine pointer is reported only when the
|
||
function can reach the runtime (is not `NOSPLIT` or makes a call) — the
|
||
ABI0 transition restores it on those paths, and NOSPLIT call-free leaves
|
||
may use it, exactly as the runtime's own assembly does. Both go-flac
|
||
kernels now lint with zero diagnostics.
|
||
|
||
### Fixed
|
||
|
||
- `lint`: the liveness analysis took the destination operand to be the
|
||
*first* operand on arm64, riscv64 and loong64; Plan 9 spelling puts it last
|
||
on every architecture Go supports. The def/use and save/restore
|
||
classification on those architectures was inverted.
|
||
- `format`: a comment that follows a `RET` (typically the next function's doc
|
||
comment) is no longer indented as if it were still inside the finished
|
||
function body.
|
||
|
||
### Added
|
||
|
||
- `asm`: the legacy (non-VEX) SSE moves — `MOVOU`/`MOVO` (the Plan 9 names
|
||
for MOVDQU/MOVDQA), `MOVUPS`/`MOVAPS`/`MOVUPD`/`MOVAPD` and the scalar
|
||
`MOVSD`/`MOVSS` — and `VMOVDQU64` in the EVEX set. All verified byte for
|
||
byte against the Go assembler.
|
||
|
||
## [0.5.0] — 2026-07-10
|
||
|
||
EVEX / AVX-512: the go-flac AVX-512 kernel now assembles, byte-identically to
|
||
the Go toolchain, completing the production-kernel coverage.
|
||
|
||
### Added
|
||
|
||
- `asm`: **EVEX (AVX-512) encoding** — the four-byte EVEX prefix with the
|
||
5-bit register fields (Z0–Z31, X/Y 16–31, with the reg-r/m X̄ quirk and
|
||
V'̄ shared between vvvv and the SIB index), opmask registers (K0–K7) as
|
||
operands and as mask destinations, and the compressed disp8×N displacement
|
||
(the multiplier follows the memory operand's size, as the Go assembler's
|
||
opcode tables prescribe). Covers every AVX-512 instruction the go-flac
|
||
kernels use: VPXORD/Q, VPADDD, VPSUBD/Q, VPUNPCK*DQ, VPMULLD/Q, VPERMD,
|
||
VPSLLD/VPSRAD/VPSRAQ, VALIGND, VPCMPEQD (K destination), VMOVDQU32,
|
||
VMOVUPD, VCVTQQ2PD, VPMOVSXDQ, the narrowing stores VPMOVDW/VPMOVQD, the
|
||
extracts VEXTRACTI64X4/VEXTRACTF64X4, VFMADD231PD, VADDPD, VMULPD, the
|
||
broadcasts VPBROADCASTD/Q (GPR and memory sources take different opcodes)
|
||
and the mask moves KMOVW/KTESTW. Masking/zeroing suffixes are out of scope
|
||
— the kernels use neither.
|
||
- `asm`: `AssembleFile` now accepts file-defined global (`non-<>`) symbols
|
||
too; a reference is external only when no `GLOBL` in the file defines it.
|
||
|
||
### Fixed
|
||
|
||
- `asm`: registers X16–Y31 force the EVEX encoding of dual-form mnemonics;
|
||
previously a `VPBROADCASTD AX, Y30` fell into the VEX encoder, which cannot
|
||
represent indices above 15 and silently truncated them.
|
||
- `asm`: the VEX encoder now rejects vector register indices 16–31 instead of
|
||
encoding a truncated (wrong) register.
|
||
|
||
### Verified
|
||
|
||
- All 10 functions of the go-flac `avx512_amd64.s` kernel assemble
|
||
byte-identically to the Go toolchain's machine code (the disp32 of the one
|
||
`VMOVDQU32 idx16(SB), Z13` load is linker-filled in Go and resolved within
|
||
gasm's own image — checked to reach the right constant bytes). The AVX2
|
||
kernel's 17 functions remain byte-identical.
|
||
|
||
## [0.4.0] — 2026-07-09
|
||
|
||
The standalone assembler reaches the whole go-flac AVX2 kernel: static
|
||
symbols assemble, and all 17 kernel functions now match the Go toolchain's
|
||
machine code byte for byte.
|
||
|
||
### Added
|
||
|
||
- `asm`: **file-level assembly** — `AssembleFile` turns a parsed file into an
|
||
`Image`: the function bodies in source order followed by a data section
|
||
built from the file's `GLOBL`/`DATA` directives (each symbol 16-aligned).
|
||
- `asm`: **static-symbol (`SB`) operands** — `mask<>(SB)` references encode as
|
||
RIP-relative loads with a patched disp32, resolved against the image layout
|
||
so the output is self-consistent and position-independent. External
|
||
(non-file-local) symbols are rejected with a clear error: they need
|
||
object-file emission.
|
||
- `gasm asm` prints the data section and symbol map alongside the functions
|
||
and writes the whole image (code + data) with `-o`.
|
||
|
||
### Verified
|
||
|
||
- All 17 functions of the go-flac `avx2_amd64.s` kernel assemble
|
||
byte-identically to the Go toolchain's machine code; the only differing
|
||
bytes are the displacements of the two `VMOVDQU mask24<>(SB), X15` loads,
|
||
which the Go linker fills at link time and gasm resolves within its own
|
||
image (checked to reach the right constant bytes).
|
||
|
||
## [0.3.0] — 2026-07-08
|
||
|
||
The assembler reaches byte-identical parity with the Go toolchain on the
|
||
production go-flac AVX2 kernels: every one of the 15 kernel functions that
|
||
avoid global symbols now assembles to exactly the Go assembler's bytes (the
|
||
two holdouts load a file-local constant through `SB` and wait on relocation
|
||
support).
|
||
|
||
### Added
|
||
|
||
- `asm`: the scalar instruction families the kernels use — `CMOVcc` and
|
||
`SETcc` (conditions spelled exactly like the jumps), `LZCNT`/`TZCNT`
|
||
(legacy `F3 0F BD/BC`), the sign/zero-extending moves (`MOVBLZX`, `MOVBQZX`,
|
||
`MOVWLZX`, `MOVWQZX`, `MOVWLSX`, `MOVLQSX`), `CVTSL2SD`/`CVTSQ2SD` (the
|
||
legacy SSE encoding, as the Go assembler emits it), the traditional
|
||
three-operand `IMUL3{W,L,Q}`, and the variable-count vector shifts
|
||
(`VPSRLQ X0, Y8, Y8` — the count in an XMM register or memory takes the
|
||
ordinary NDS form).
|
||
- `asm`: **jump relaxation** — jumps start in the short (rel8) form and
|
||
expand to rel32 when the settled displacement does not fit, iterating the
|
||
layout to a fixed point (CALL is always rel32).
|
||
- `asm`: **jump-to-jump folding** — a conditional jump to a label whose only
|
||
instruction is an unconditional jump is redirected to the ultimate
|
||
target, replicating the Go toolchain's linker, which chases such chains
|
||
before it encodes branches.
|
||
- `parser`: leading negative displacements with a base and index
|
||
(`LEAQ -4(DX)(R9*4), R9`) parse into a fully populated address.
|
||
|
||
### Fixed
|
||
|
||
- `asm`: `CMP` with a register or memory operand computed **second − first**
|
||
instead of first − second, silently inverting every condition that followed
|
||
(`CMPQ SI, R10; JGE` tested R10 ≥ SI). The encoding now always records
|
||
first − second — `CMP r/m, r` with the first operand in r/m, `CMP r, r/m`
|
||
with the first operand in reg — and is byte-identical to the Go assembler.
|
||
- `asm`: register-to-register `MOV` now uses the `r/m ← r` opcode (reg =
|
||
source), the Go assembler's choice; the output is byte-identical.
|
||
|
||
## [0.2.0] — 2026-07-07
|
||
|
||
The Phase 2 assembler grows the SIMD set: shuffles, extract/insert, permute
|
||
and the moves, on top of the Phase 1 VEX forms.
|
||
|
||
### Added
|
||
|
||
- `asm`: four new VEX (AVX/AVX2) operand forms, each validated by round-trip
|
||
decoding through `golang.org/x/arch` **and** byte-for-byte against the
|
||
machine code the real Go assembler emits:
|
||
- the immediate shuffle (`VPSHUFD`, `VPERMQ`),
|
||
- the three-operand-plus-immediate form (`VSHUFPD`, `VPERM2I128`,
|
||
`VINSERTI128`),
|
||
- the lane extract (`VEXTRACTI128`, `VEXTRACTF128` — the YMM source occupies
|
||
the ModRM.reg field, the XMM/memory destination the r/m field),
|
||
- the direction-sensitive moves (`VMOVDQU`, `VMOVUPD`, `VMOVD`, `VMOVQ`,
|
||
`VMOVSD` — each direction picks its own opcode and VEX.W; a vector→vector
|
||
move uses the store-form layout, matching the Go assembler),
|
||
- the no-operand `VZEROUPPER`, and `VPERMD` in the NDS form,
|
||
- the floating-point and FMA set (`VADDPD`, `VMULPD`, `VXORPD`,
|
||
`VUNPCKHPD`, the scalar `VADDSD`/`VMULSD`, `VCVTDQ2PD`, `VFMADD231PD`).
|
||
With the scalar set and the earlier NDS / reg-rm / immediate-shift forms,
|
||
the encoder now covers every integer, shuffle and FP instruction the
|
||
go-flac AVX2 kernels use.
|
||
- `asm`: `CMP` accepts the immediate in the second operand position
|
||
(`CMPL CX, $31`) — the spelling the Go assembler accepts — encoding it
|
||
identically to the immediate-first form.
|
||
|
||
### Fixed
|
||
|
||
- `asm`: an unused VEX.vvvv field is now stored as `1111` (v̄vvv = 1111), as
|
||
the hardware requires — the previous value (`0000`) made the two-operand
|
||
reg/rm forms (VPMOVSXWD, VPBROADCASTD, VMOVMSKPS, …) raise #UD on real CPUs
|
||
and differ from the Go assembler's bytes. The round-trip decoder ignores
|
||
the field on these instructions, which is why the byte-for-byte Go
|
||
comparison (added this release) is now part of the test suite.
|
||
|
||
## [0.1.0] — 2026-07-06
|
||
|
||
Initial release — the Phase 1 foundation.
|
||
|
||
### Added
|
||
|
||
- `token`, `lexer`, `ast`, `parser`: a hand-written, error-tolerant front end
|
||
for Plan 9 assembly. The lexer splices C-preprocessor line continuations
|
||
(`\` before a newline) so multi-line `#define` macros parse as one opaque
|
||
directive. Validated against the production AVX2/AVX-512 kernels in
|
||
`go-libraries/go-flac` and the Go runtime's `src/runtime/*.s` for all four
|
||
architectures, with zero parse errors.
|
||
- `arch`: register files and **complete** instruction tables for amd64,
|
||
arm64, riscv64 and loong64, with the middle-dot symbol separator and static
|
||
(`<>`) symbols. Instruction names are generated from the Go toolchain's own
|
||
assembler source (`just gen`) — the `anames` opcode lists plus the common
|
||
opcodes and the per-architecture front-end aliases (arm64 `B`/`BL`, the
|
||
`.P`/`.W` addressing suffixes, loong64 `JAL`, the x86 conditional-jump
|
||
spellings) — so every mnemonic the real assembler accepts is recognised.
|
||
- `lint`: conservative rules — `unknown-instruction`, `operand-count`,
|
||
`undefined-label`, `duplicate-label`, `missing-ret`,
|
||
`missing-textflag-include`, `abi-argsize` and `unreachable-code`. Macro
|
||
invocations are recognised (in-file `#define` names and underscore
|
||
identifiers) and the label/RET heuristics are suppressed in macro-using
|
||
files. `abi-argsize` parses the `// func` signature with the Go parser and
|
||
checks the declared TEXT argument size against Go's ABI0 layout;
|
||
`unreachable-code` flags dead code after `RET`, suppressed where reachability
|
||
is undecidable (PC-relative jumps, register-indirect branches, `#ifdef`).
|
||
Register liveness is computed by dataflow over the control-flow graph (basic
|
||
blocks, def/use, iterative backward iteration) and drives `register-clobber`,
|
||
an audit that flags a callee-saved register written but never saved/restored.
|
||
`funcdata-pcdata` validates the structure of `FUNCDATA`/`PCDATA` directives.
|
||
Zero error-severity diagnostics across the 90-file Go runtime corpus and the
|
||
production go-flac kernels (the `register-clobber` audit additionally reports
|
||
the go-flac kernels' unsaved callee-saved register use for review).
|
||
- `format`: an idempotent canonical formatter (operand spacing and per-function
|
||
mnemonic alignment) that preserves comments and round-trips through the
|
||
parser.
|
||
- `lsp`: a Language Server Protocol server over stdio providing completion,
|
||
hover documentation, document symbols, publish-diagnostics and semantic-token
|
||
highlighting.
|
||
- `asm`: a standalone amd64 (x86-64) assembler — an instruction encoder (REX/
|
||
ModR-M/SIB/displacement/immediate plus the scalar instruction set, and VEX/
|
||
AVX2 SIMD across three operand forms — NDS, reg/rm and immediate-shift —
|
||
covering the bulk of the integer SIMD set) validated by round-trip decoding
|
||
against `golang.org/x/arch`, and an assembler that drives the parser's AST
|
||
into the encoder with local-label resolution and `FP`/`SP` frame mapping
|
||
(plus Go prologue/epilogue generation), producing output byte-identical to the
|
||
Go assembler for the supported operand forms.
|
||
- `cmd/gasm`: the `gasm` binary with `tokens`, `parse`, `fmt`, `lint`, `asm`
|
||
and `lsp` subcommands.
|
||
- `_gen`: the generator that rebuilds the architecture instruction tables from
|
||
the Go toolchain source (`just gen`).
|