docs: add SECURITY.md and record the round in the CHANGELOG

Assisted-by: GLM 5.3 Flash
This commit is contained in:
2026-09-16 23:12:31 +02:00
parent 0078f7be5c
commit 3de043c494
2 changed files with 197 additions and 140 deletions
+157 -140
View File
@@ -7,11 +7,28 @@ and this project adheres to [Conventional Commits](https://www.conventionalcommi
## [development] ## [development]
### Added ### Changed
- - **Canonical just recipes.** `just gates` is the definition of done
(build, fmt-check, vet, test, race). `install` now builds and copies
the binary into `~/.local/bin` (`BINDIR` overrides) instead of
downloading module dependencies, and `install-bin` is gone. The test
gate sweeps the whole suite and computes the coverage floor over the
product packages with `-coverpkg`, `disasm` now included, so the
number is identical locally and in CI. `fuzz` requires its target
package.
- **The reported version comes from the build.** `gasm --version`
prints the version the toolchain recorded: the tag on a tagged
checkout, a pseudo-version naming the commit below one, `+dirty` on a
dirty tree and `(devel)` outside version control. Nothing is
injected with `-ldflags -X` any more.
- **CI realigned with the gate set.** The push pipeline runs the gates
minus race in one job, with a cached Go setup and the module as the
version source; the race detector moved to a hand-dispatched workflow
and into the release gates; the release builds without injection and
its smoke test requires the recorded tag and rejects `+dirty`.
## [0.33.0] — 2026-09-14 ## [0.33.0] - 2026-09-14
### Added ### Added
@@ -67,7 +84,7 @@ and this project adheres to [Conventional Commits](https://www.conventionalcommi
the toolchain picks, and the morestack block saves the link register the toolchain picks, and the morestack block saves the link register
with the toolchain's `OR` form on loong64. with the toolchain's `OR` form on loong64.
## [0.32.0] — 2026-08-31 ## [0.32.0] - 2026-08-31
### Added ### Added
@@ -218,7 +235,7 @@ and this project adheres to [Conventional Commits](https://www.conventionalcommi
outputs and operand strictness with `go tool asm`. outputs and operand strictness with `go tool asm`.
- **asm help text.** Updated to list arm64 as a supported architecture. - **asm help text.** Updated to list arm64 as a supported architecture.
## [0.31.1] — 2026-08-20 ## [0.31.1] - 2026-08-20
### Fixed ### Fixed
@@ -226,9 +243,9 @@ and this project adheres to [Conventional Commits](https://www.conventionalcommi
because the version variables in `justfile` and `cmd/gasm/main.go` were not because the version variables in `justfile` and `cmd/gasm/main.go` were not
bumped during the release commit. bumped during the release commit.
## [0.31.0] — 2026-08-20 ## [0.31.0] - 2026-08-20
The arm64 encoder (Phase 5 — complete) ships with ELF64 and GOOBJ emission, The arm64 encoder (Phase 5; complete) ships with ELF64 and GOOBJ emission,
verified byte-for-byte against `GOARCH=arm64 go tool asm` and linked into a verified byte-for-byte against `GOARCH=arm64 go tool asm` and linked into a
real `go build`. The encoder covers the full integer instruction set, FP real `go build`. The encoder covers the full integer instruction set, FP
arithmetic, conditional select, CRC32, and the MOV pseudo-instruction with arithmetic, conditional select, CRC32, and the MOV pseudo-instruction with
@@ -236,7 +253,7 @@ bitmask immediate encoding. The project now requires Go 1.27.
### Added ### Added
- **arm64 encoder (Phase 5 — complete).** `gasm asm` can now assemble `_arm64.s` - **arm64 encoder (Phase 5; complete).** `gasm asm` can now assemble `_arm64.s`
files: the AArch64 integer instruction set with the MOV pseudo-instruction and files: the AArch64 integer instruction set with the MOV pseudo-instruction and
its immediate-constant expansions (MOVZ/MOVN/MOVK for wide immediates, ORR with its immediate-constant expansions (MOVZ/MOVN/MOVK for wide immediates, ORR with
logical bitmask encoding for values like `$1`), data-processing (shifted logical bitmask encoding for values like `$1`), data-processing (shifted
@@ -245,7 +262,7 @@ bitmask immediate encoding. The project now requires Go 1.27.
SB/global symbol references (ADRP+ADD pairs with `R_ADDRARM64` relocations), SB/global symbol references (ADRP+ADD pairs with `R_ADDRARM64` relocations),
jump chain folding, and ELF64 emission (`gasm asm --format elf`). Ground-truth jump chain folding, and ELF64 emission (`gasm asm --format elf`). Ground-truth
verification against `GOARCH=arm64 go tool asm` matches byte-for-byte. Phase 5 verification against `GOARCH=arm64 go tool asm` matches byte-for-byte. Phase 5
(the other architectures — RISC-V, LoongArch, arm64) is now complete. (the other architectures; RISC-V, LoongArch, arm64) is now complete.
### Changed ### Changed
@@ -253,7 +270,7 @@ bitmask immediate encoding. The project now requires Go 1.27.
The `R_DWTXTADDR_U4` relocation type is detected at runtime for backward The `R_DWTXTADDR_U4` relocation type is detected at runtime for backward
compatibility. compatibility.
## [0.30.0] — 2026-08-13 ## [0.30.0] - 2026-08-13
The LoongArch encoder (Phase 5) ships with ELF64 and GOOBJ emission, verified The LoongArch encoder (Phase 5) ships with ELF64 and GOOBJ emission, verified
byte-for-byte against `GOARCH=loong64 go tool asm` and linked into a real byte-for-byte against `GOARCH=loong64 go tool asm` and linked into a real
@@ -278,7 +295,7 @@ tracks four hardware watchpoint slots, and the toolkit is Linux-only.
- **GOOBJ DWARF symbols.** The GOOBJ emitters now write the per-function - **GOOBJ DWARF symbols.** The GOOBJ emitters now write the per-function
DWARF symbols the linker requires (the subprogram DIE and the `.debug_line` DWARF symbols the linker requires (the subprogram DIE and the `.debug_line`
program, byte-identical to `cmd/asm`'s), and the pc-value table deltas are program, byte-identical to `cmd/asm`'s), and the pc-value table deltas are
in the architecture's MinLC units as the runtime expects — the amd64 link in the architecture's MinLC units as the runtime expects; the amd64 link
test now genuinely substitutes the gasm object, and the amd64/loong64 test now genuinely substitutes the gasm object, and the amd64/loong64
end-to-end GOOBJ link tests pass. end-to-end GOOBJ link tests pass.
- **RISC-V GOOBJ emission via the shared emitter.** RISC-V GOOBJ output is - **RISC-V GOOBJ emission via the shared emitter.** RISC-V GOOBJ output is
@@ -347,7 +364,7 @@ tracks four hardware watchpoint slots, and the toolkit is Linux-only.
`GOARCH=riscv64 go tool asm`. `GOARCH=riscv64 go tool asm`.
- **Debugger watchpoint slots.** `gasm debug`'s `watch` command always used - **Debugger watchpoint slots.** `gasm debug`'s `watch` command always used
hardware watchpoint slot 0, so a second `watch` call silently overwrote hardware watchpoint slot 0, so a second `watch` call silently overwrote
the first. Watchpoint slots are now tracked in the `Session` (DR0–DR3); the first. Watchpoint slots are now tracked in the `Session` (DR0-DR3);
`watch` picks the first free slot and reports an error if all four are in `watch` picks the first free slot and reports an error if all four are in
use, and `unwatch <slot>` clears one (no argument clears all). use, and `unwatch <slot>` clears one (no argument clears all).
@@ -359,7 +376,7 @@ tracks four hardware watchpoint slots, and the toolkit is Linux-only.
disassembly at PC, memory-write, watchpoints, and source-line mapping are disassembly at PC, memory-write, watchpoints, and source-line mapping are
all shipped. all shipped.
## [0.29.0] — 2026-08-07 ## [0.29.0] - 2026-08-07
RISC-V GOOBJ emission, YMM vector register display, named buffer allocation RISC-V GOOBJ emission, YMM vector register display, named buffer allocation
in the debugger, two new CLI commands (`diff`, `profile`), go-to-definition in in the debugger, two new CLI commands (`diff`, `profile`), go-to-definition in
@@ -369,30 +386,30 @@ new CLI commands. A signature-parser fix corrects grouped Go parameters.
### Added ### Added
- **RISC-V GOOBJ emission** — `gasm asm --format goobj` for RISC-V produces - **RISC-V GOOBJ emission**; `gasm asm --format goobj` for RISC-V produces
linkable Go objects with funcdata, pc-value tables, and RISC-V relocation linkable Go objects with funcdata, pc-value tables, and RISC-V relocation
types (same format as amd64 GOOBJ, with the RISC-V architecture marker). types (same format as amd64 GOOBJ, with the RISC-V architecture marker).
- **`gasm diff`** — compare the machine code of two assembly files byte-for-byte; - **`gasm diff`**; compare the machine code of two assembly files byte-for-byte;
shows which functions differ and the first few differing bytes. shows which functions differ and the first few differing bytes.
- **`gasm profile`** — show the basic-block structure of each function: labels, - **`gasm profile`**; show the basic-block structure of each function: labels,
offsets, frame size, and NOSPLIT flag. offsets, frame size, and NOSPLIT flag.
- **LSP go-to-definition** — `textDocument/definition` navigates from a label - **LSP go-to-definition**; `textDocument/definition` navigates from a label
reference to its definition. reference to its definition.
- **did-you-mean** — when the RISC-V assembler encounters an undefined label, it - **did-you-mean**; when the RISC-V assembler encounters an undefined label, it
suggests the closest existing label using Levenshtein distance. suggests the closest existing label using Levenshtein distance.
- **YMM vector register display** — `regs` in the debugger now shows YMM - **YMM vector register display**; `regs` in the debugger now shows YMM
registers via `PTRACE_GETFPREGS` (falls back to XMM when XSAVE is unavailable). registers via `PTRACE_GETFPREGS` (falls back to XMM when XSAVE is unavailable).
- **Named buffer allocation** — `gasm debug --buf name:size:pattern` allocates - **Named buffer allocation**; `gasm debug --buf name:size:pattern` allocates
buffers in the debuggee filled with `zero`, `ones`, `seq`, or a hex pattern; buffers in the debuggee filled with `zero`, `ones`, `seq`, or a hex pattern;
buffer pointers are placed into the argument block at the matching positions. buffer pointers are placed into the argument block at the matching positions.
- **Crash input storage** — `FuzzResult.CrashInput` stores the input that caused - **Crash input storage**; `FuzzResult.CrashInput` stores the input that caused
a crash or mismatch for reproducibility. a crash or mismatch for reproducibility.
- **ABI + fuzz combined** — `gasm verify --fuzz` now runs ABI checks (sentinel - **ABI + fuzz combined**; `gasm verify --fuzz` now runs ABI checks (sentinel
registers, canary, stack bounds) alongside differential fuzz testing. registers, canary, stack bounds) alongside differential fuzz testing.
- **`gasm diff --map`** — compare functions whose names differ between files - **`gasm diff --map`**; compare functions whose names differ between files
(e.g. `--map wideCopyAVX2=wideCopyAVX512` pairs two variants regardless (e.g. `--map wideCopyAVX2=wideCopyAVX512` pairs two variants regardless
of suffix). Unmapped functions fall back to the original name match. of suffix). Unmapped functions fall back to the original name match.
- **`gasm verify --call`** — invoke a single function with user-supplied buffers - **`gasm verify --call`**; invoke a single function with user-supplied buffers
(`--buf name:size:pattern`) instead of the smoke/abi/fuzz sweeps. Patterns: (`--buf name:size:pattern`) instead of the smoke/abi/fuzz sweeps. Patterns:
`zero`, `ones`, `seq`, or a hex blob. Useful for partial functions (e.g. `zero`, `ones`, `seq`, or a hex blob. Useful for partial functions (e.g.
decoders) that crash on random input but should succeed on valid data. decoders) that crash on random input but should succeed on valid data.
@@ -402,23 +419,23 @@ new CLI commands. A signature-parser fix corrects grouped Go parameters.
### Fixed ### Fixed
- **Signature parser** — grouped Go parameters like `dst, src []byte` are now - **Signature parser**; grouped Go parameters like `dst, src []byte` are now
parsed correctly (both get type `[]byte`). Previously the first name was parsed correctly (both get type `[]byte`). Previously the first name was
treated as its own type (`dst` with size 8), causing wrong ABI0 arg-block treated as its own type (`dst` with size 8), causing wrong ABI0 arg-block
layout in both `verify --call` and the fuzzer. layout in both `verify --call` and the fuzzer.
- **Flaky JIT tests** — `runtime.KeepAlive` guards and package-level buffers - **Flaky JIT tests**; `runtime.KeepAlive` guards and package-level buffers
prevent GC from collecting heap objects whose addresses were passed to JIT prevent GC from collecting heap objects whose addresses were passed to JIT
code via `unsafe.Pointer`; all verify tests pass 100/100 under `-race`. code via `unsafe.Pointer`; all verify tests pass 100/100 under `-race`.
### Changed ### Changed
- **Removed external kernel test dependencies** — the verify test suite no - **Removed external kernel test dependencies**; the verify test suite no
longer references production kernels from the separate go-libraries project. longer references production kernels from the separate go-libraries project.
The remaining test suite uses only `testdata/verify/*.s` kernels, which are The remaining test suite uses only `testdata/verify/*.s` kernels, which are
part of this repository. Coverage is identical locally and in CI (80.3 %). part of this repository. Coverage is identical locally and in CI (80.3 %).
## [0.28.0] — 2026-08-03 ## [0.28.0] - 2026-08-03
RISC-V encoder: full RV64IMAFDC instruction set with RVC compression, MOV RISC-V encoder: full RV64IMAFDC instruction set with RVC compression, MOV
pseudo-instruction, SB/global symbol references, ELF64 object emission, and pseudo-instruction, SB/global symbol references, ELF64 object emission, and
@@ -426,20 +443,20 @@ ground-truth verification against `GOARCH=riscv64 go tool asm`.
### Added ### Added
- **RISC-V encoder** — RV64I, RV64M, RV64A, RV64F/D, FMA, CSR, JALR. - **RISC-V encoder**; RV64I, RV64M, RV64A, RV64F/D, FMA, CSR, JALR.
- **MOV pseudo-instruction** — load, store, reg-to-reg, immediate, frame mapping. - **MOV pseudo-instruction**; load, store, reg-to-reg, immediate, frame mapping.
- **RVC compression** — 22 compressed instruction types (C.LDSP, C.SDSP, C.FLDSP, - **RVC compression**; 22 compressed instruction types (C.LDSP, C.SDSP, C.FLDSP,
C.FSDSP, C.ADDI, C.LI, C.LUI, C.ADDIW, C.MV, C.ADD, C.SUB, C.XOR, C.OR, C.AND, C.FSDSP, C.ADDI, C.LI, C.LUI, C.ADDIW, C.MV, C.ADD, C.SUB, C.XOR, C.OR, C.AND,
C.SLLI, C.SRLI, C.SRAI, C.ANDI, C.BEQZ, C.BNEZ, C.J, C.JR). C.SLLI, C.SRLI, C.SRAI, C.ANDI, C.BEQZ, C.BNEZ, C.J, C.JR).
- **SB/global symbols** — `MOV $sym(SB)`, `MOV sym(SB)`, `MOV rd, sym(SB)` - **SB/global symbols**; `MOV $sym(SB)`, `MOV sym(SB)`, `MOV rd, sym(SB)`
encoded as AUIPC pairs with R_RISCV_PCREL_HI20/LO12 relocations. encoded as AUIPC pairs with R_RISCV_PCREL_HI20/LO12 relocations.
- **GLOBL/DATA** — data section layout in `AssembleFileRISCV`. - **GLOBL/DATA**; data section layout in `AssembleFileRISCV`.
- **ELF64 emission** — `gasm asm --format elf` produces EM_RISCV objects - **ELF64 emission**; `gasm asm --format elf` produces EM_RISCV objects
(.text, .data, .symtab, .rela.text). (.text, .data, .symtab, .rela.text).
- **`gasm verify --ground-truth`** — byte-exact comparison against - **`gasm verify --ground-truth`**; byte-exact comparison against
`GOARCH=riscv64 go tool asm`. `GOARCH=riscv64 go tool asm`.
- **`gasm verify --profile`** — function layout listing for RISC-V. - **`gasm verify --profile`**; function layout listing for RISC-V.
- **CALL** — AUIPC + JALR pair encoding. - **CALL**; AUIPC + JALR pair encoding.
### Fixed ### Fixed
@@ -449,7 +466,7 @@ ground-truth verification against `GOARCH=riscv64 go tool asm`.
(bit-interleaved format). (bit-interleaved format).
## [0.27.0] — 2026-08-01 ## [0.27.0] - 2026-08-01
Subprocess isolation for `--fuzz`: each function is fuzzed in its own child Subprocess isolation for `--fuzz`: each function is fuzzed in its own child
process, so a partial function (decoder) that faults on random garbage is process, so a partial function (decoder) that faults on random garbage is
@@ -460,7 +477,7 @@ the parent. CRASH is informational (exit 0); only MISMATCH is an error.
- `gasm verify --fuzz` no longer crashes the process on partial functions. - `gasm verify --fuzz` no longer crashes the process on partial functions.
## [0.26.0] — 2026-07-31 ## [0.26.0] - 2026-07-31
Universal differential fuzzing: `gasm verify --fuzz` needs no hand-written Universal differential fuzzing: `gasm verify --fuzz` needs no hand-written
reference. It parses the `// func` signature from the assembly source, reference. It parses the `// func` signature from the assembly source,
@@ -471,7 +488,7 @@ area bit-for-bit.
### Added ### Added
- `verify`: `FuzzFunc` / `ExtractSignatures` / `parseFuncSig` — universal - `verify`: `FuzzFunc` / `ExtractSignatures` / `parseFuncSig`; universal
differential fuzz driven by the conventional `// func` comment. Each differential fuzz driven by the conventional `// func` comment. Each
version gets its own buffer set (deep copy) so functions that write to version gets its own buffer set (deep copy) so functions that write to
their arguments (histogram increments) don't corrupt the other's input. their arguments (histogram increments) don't corrupt the other's input.
@@ -486,16 +503,16 @@ area bit-for-bit.
over-copy paths read past the buffer on random garbage input. Subprocess over-copy paths read past the buffer on random garbage input. Subprocess
isolation (fork per function) is planned. Use `--ground-truth` for decoders. isolation (fork per function) is planned. Use `--ground-truth` for decoders.
## [0.25.0] — 2026-07-30 ## [0.25.0] - 2026-07-30
Universal ground-truth verification: `gasm verify --ground-truth` assembles Universal ground-truth verification: `gasm verify --ground-truth` assembles
any `.s` file with both gasm and `go tool asm`, then compares the machine any `.s` file with both gasm and `go tool asm`, then compares the machine
code byte-for-byte per function (relocation sites masked). No hand-written code byte-for-byte per function (relocation sites masked). No hand-written
reference needed — the Go toolchain IS the oracle. reference needed; the Go toolchain IS the oracle.
### Added ### Added
- `verify`: `GroundTruth` — shells out to `go tool asm`, parses the GOOBJ - `verify`: `GroundTruth`; shells out to `go tool asm`, parses the GOOBJ
output (minimal reader: block offsets, nonpkg symbol table, data index) output (minimal reader: block offsets, nonpkg symbol table, data index)
and returns per-function code bytes. and returns per-function code bytes.
- `gasm verify --ground-truth`: compares gasm's output against the Go - `gasm verify --ground-truth`: compares gasm's output against the Go
@@ -504,11 +521,11 @@ reference needed — the Go toolchain IS the oracle.
linker fills) are masked before comparison. linker fills) are masked before comparison.
- Verified: go-lz4 AVX2 2/2, go-flac AVX2 17/17 functions byte-identical. - Verified: go-lz4 AVX2 2/2, go-flac AVX2 17/17 functions byte-identical.
## [0.24.0] — 2026-07-29 ## [0.24.0] - 2026-07-29
The full analyze family and stereo PCM decode are now differentially tested. The full analyze family and stereo PCM decode are now differentially tested.
15 of 17 go-flac AVX2 kernels have bit-for-bit differential coverage; the 15 of 17 go-flac AVX2 kernels have bit-for-bit differential coverage; the
two remaining (autocorrAVX2 — FMA reassociation, lpcResidualAVX2 — complex two remaining (autocorrAVX2: FMA reassociation, lpcResidualAVX2: complex
multi-arg) are deferred. multi-arg) are deferred.
### Added ### Added
@@ -518,7 +535,7 @@ multi-arg) are deferred.
- `verify`: `decodeStereo16AVX2` differential test (500 random interleaved - `verify`: `decodeStereo16AVX2` differential test (500 random interleaved
stereo PCM buffers, both channels compared sample-by-sample). stereo PCM buffers, both channels compared sample-by-sample).
## [0.23.0] — 2026-07-28 ## [0.23.0] - 2026-07-28
The analyze family and 24-bit PCM decode join the differential suite. The analyze family and 24-bit PCM decode join the differential suite.
@@ -530,7 +547,7 @@ The analyze family and 24-bit PCM decode join the differential suite.
- `verify`: `decodeMono24AVX2` differential test (500 random 24-bit PCM - `verify`: `decodeMono24AVX2` differential test (500 random 24-bit PCM
buffers, sign-extension compared sample-by-sample). buffers, sign-extension compared sample-by-sample).
## [0.22.0] — 2026-07-27 ## [0.22.0] - 2026-07-27
The remaining go-flac encoder kernels join the differential suite. The remaining go-flac encoder kernels join the differential suite.
@@ -544,39 +561,39 @@ The remaining go-flac encoder kernels join the differential suite.
loop). loop).
## [0.21.0] — 2026-07-26 ## [0.21.0] - 2026-07-26
Differential testing extended to all four production kernels and the CLI Differential testing extended to all four production kernels and the CLI
exposes the full dynamic-analysis toolkit. exposes the full dynamic-analysis toolkit.
### Added ### Added
- `verify`: go-flac AVX2 differential tests — `decodeMono16AVX2` (500 - `verify`: go-flac AVX2 differential tests; `decodeMono16AVX2` (500
random PCM buffers), `pack16AVX2` (500 random int32→int16 packings) and random PCM buffers), `pack16AVX2` (500 random int32→int16 packings) and
all four decorrelation kernels (200 iterations each: left-side, side-right, all four decorrelation kernels (200 iterations each: left-side, side-right,
mid-side, interleave) compared bit-for-bit against the portable Go mid-side, interleave) compared bit-for-bit against the portable Go
references. references.
- `verify`: go-lz4 AVX-512 differential tests — `decodeBlockAVX512` (3 000 - `verify`: go-lz4 AVX-512 differential tests; `decodeBlockAVX512` (3 000
fuzzed LZ4 blocks + known answers) and `wideCopyAVX512` (0–1024 bytes) fuzzed LZ4 blocks + known answers) and `wideCopyAVX512` (0-1024 bytes)
against the same portable oracle as the AVX2 suite. against the same portable oracle as the AVX2 suite.
- `gasm verify --abi`: runs each NOSPLIT function with sentinel registers - `gasm verify --abi`: runs each NOSPLIT function with sentinel registers
and a red-zone canary, reporting violations. and a red-zone canary, reporting violations.
- `gasm verify --profile`: lists the static basic-block count per function. - `gasm verify --profile`: lists the static basic-block count per function.
## [0.20.0] — 2026-07-25 ## [0.20.0] - 2026-07-25
Coverage profiling: the third pillar of Phase 3. Static basic-block Coverage profiling: the third pillar of Phase 3. Static basic-block
enumeration from the assembler's label map, combined with multi-input path enumeration from the assembler's label map, combined with multi-input path
diversity measurement — how many observationally distinct execution paths a diversity measurement; how many observationally distinct execution paths a
test corpus exercises. test corpus exercises.
### Added ### Added
- `verify`: `Kernel.Blocks` / `Kernel.BlockCount` — enumerate basic blocks - `verify`: `Kernel.Blocks` / `Kernel.BlockCount`; enumerate basic blocks
from the assembler's local-label map (every jump target is a block from the assembler's local-label map (every jump target is a block
boundary; the function entry is always a block). `decodeBlockAVX2` has boundary; the function entry is always a block). `decodeBlockAVX2` has
27 blocks. 27 blocks.
- `verify`: `Kernel.ProfilePaths` — run the function with a corpus of - `verify`: `Kernel.ProfilePaths`; run the function with a corpus of
argument blocks and collect distinct output fingerprints (the result argument blocks and collect distinct output fingerprints (the result
words); reports path diversity as a lower bound on code coverage. words); reports path diversity as a lower bound on code coverage.
@@ -588,7 +605,7 @@ rt_sigaction handlers fragile in a Go process. The static + path-diversity
approach delivers the project's goal (proving the SIMD path and tail handling approach delivers the project's goal (proving the SIMD path and tail handling
execute) without fighting the runtime. execute) without fighting the runtime.
## [0.19.0] — 2026-07-24 ## [0.19.0] - 2026-07-24
Runtime ABI checks: the second pillar of Phase 3. The JIT trampoline now Runtime ABI checks: the second pillar of Phase 3. The JIT trampoline now
has an ABI-checking variant that sets sentinels in the callee-saved registers has an ABI-checking variant that sets sentinels in the callee-saved registers
@@ -598,41 +615,41 @@ detects any illegal write below the stack pointer.
### Added ### Added
- `verify`: `CallChecked` / `Kernel.CallFuncChecked` — ABI-checking JIT call - `verify`: `CallChecked` / `Kernel.CallFuncChecked`; ABI-checking JIT call
with sentinel registers and red-zone canary; returns an `ABIReport` with sentinel registers and red-zone canary; returns an `ABIReport`
(BPClobbered, R14Clobbered, RedZoneHit). (BPClobbered, R14Clobbered, RedZoneHit).
- `verify`: the raw `leaveJITCheckedRaw` trampoline — a TEXT symbol with no - `verify`: the raw `leaveJITCheckedRaw` trampoline; a TEXT symbol with no
ABIInternal wrapper (address obtained via GLOBL/DATA), so the JIT ABIInternal wrapper (address obtained via GLOBL/DATA), so the JIT
function's RET lands directly in the check code and sees the registers function's RET lands directly in the check code and sees the registers
exactly as the function left them. exactly as the function left them.
- Tests: deliberate BP/R14 clobberers detected; both go-lz4 kernels - Tests: deliberate BP/R14 clobberers detected; both go-lz4 kernels
confirmed ABI-clean (BP preserved, R14 preserved, red zone intact). confirmed ABI-clean (BP preserved, R14 preserved, red zone intact).
## [0.18.0] — 2026-07-23 ## [0.18.0] - 2026-07-23
Differential testing: the JIT-assembled go-lz4 `decodeBlockAVX2` kernel is Differential testing: the JIT-assembled go-lz4 `decodeBlockAVX2` kernel is
fuzzed against a portable Go reference — 5 000 valid LZ4 blocks compared fuzzed against a portable Go reference; 5 000 valid LZ4 blocks compared
bit-for-bit, plus 2 000 hostile (random garbage) inputs with matching error bit-for-bit, plus 2 000 hostile (random garbage) inputs with matching error
codes. This is the automated form of the project's bit-identical contract. codes. This is the automated form of the project's bit-identical contract.
### Added ### Added
- `verify`: differential fuzz tests — a random LZ4 block generator produces - `verify`: differential fuzz tests; a random LZ4 block generator produces
valid blocks (literals, overlapping matches, extension bytes) and the valid blocks (literals, overlapping matches, extension bytes) and the
JIT-assembled kernel's output is compared byte-for-byte against a portable JIT-assembled kernel's output is compared byte-for-byte against a portable
Go decoder; a hostile-input suite confirms error-code agreement on random Go decoder; a hostile-input suite confirms error-code agreement on random
garbage (no crashes, same classification). garbage (no crashes, same classification).
## [0.17.0] — 2026-07-22 ## [0.17.0] - 2026-07-22
Phase 3 begins: dynamic analysis. A JIT execution substrate that assembles Phase 3 begins: dynamic analysis. A JIT execution substrate that assembles
Plan 9 amd64 kernels into executable memory and calls them directly — pure Go Plan 9 amd64 kernels into executable memory and calls them directly; pure Go
(stdlib only, `syscall.Mmap` + an assembly trampoline), no cgo, no external (stdlib only, `syscall.Mmap` + an assembly trampoline), no cgo, no external
toolchain. toolchain.
### Added ### Added
- `verify` package: JIT infrastructure — `Map` copies machine code into a - `verify` package: JIT infrastructure; `Map` copies machine code into a
W^X memory mapping, `Call` invokes it through an ABI0 trampoline that W^X memory mapping, `Call` invokes it through an ABI0 trampoline that
switches to a prepared stack and back. `Load`/`LoadSource`/`LoadAST` switches to a prepared stack and back. `Load`/`LoadSource`/`LoadAST`
parse, assemble and map a `.s` file in one step; `Kernel.CallFunc` parse, assemble and map a `.s` file in one step; `Kernel.CallFunc`
@@ -641,34 +658,34 @@ toolchain.
available functions; with `-smoke`, calls each NOSPLIT function with available functions; with `-smoke`, calls each NOSPLIT function with
zeroed arguments to confirm the trampoline works end-to-end. zeroed arguments to confirm the trampoline works end-to-end.
- Integration tests: the go-lz4 `decodeBlockAVX2` and `wideCopyAVX2` - Integration tests: the go-lz4 `decodeBlockAVX2` and `wideCopyAVX2`
kernels (699 and 146 bytes) assemble, map and execute correctly — kernels (699 and 146 bytes) assemble, map and execute correctly;
known-answer LZ4 blocks decode bit-for-bit, wide copies of 0–1024 bytes known-answer LZ4 blocks decode bit-for-bit, wide copies of 0-1024 bytes
match, malformed input returns the correct error codes. match, malformed input returns the correct error codes.
## [0.16.0] — 2026-07-21 ## [0.16.0] - 2026-07-21
The scalar conversions between vector and general-purpose registers — the The scalar conversions between vector and general-purpose registers; the
last of the amd64 EVEX instruction set. last of the amd64 EVEX instruction set.
### Added ### Added
- `asm`: the GPR-interchanging conversions, byte for byte against the Go - `asm`: the GPR-interchanging conversions, byte for byte against the Go
assembler (28 ground-truth cases including memory sources and extended assembler (28 ground-truth cases including memory sources and extended
GPRs): vector to GPR — the signed and truncated VCVT{,T}S{D,S}2SI{,Q} GPRs): vector to GPR; the signed and truncated VCVT{,T}S{D,S}2SI{,Q}
in both VEX and EVEX, and the unsigned VCVT{,T}S{D,S}2USI{L,Q} in both VEX and EVEX, and the unsigned VCVT{,T}S{D,S}2USI{L,Q}
(EVEX only); GPR to vector — VCVTSI2SD{L,Q}/VCVTSI2SS{L,Q} (VEX and (EVEX only); GPR to vector; VCVTSI2SD{L,Q}/VCVTSI2SS{L,Q} (VEX and
EVEX) and VCVTUSI2SD{L,Q}/VCVTUSI2SS{L,Q} (EVEX only), whose preserved EVEX) and VCVTUSI2SD{L,Q}/VCVTUSI2SS{L,Q} (EVEX only), whose preserved
vector source sits in vvvv (three Plan 9 operands). vector source sits in vvvv (three Plan 9 operands).
## [0.15.0] — 2026-07-20 ## [0.15.0] - 2026-07-20
The last of the EVEX conversions and narrowing/extending moves — the EVEX The last of the EVEX conversions and narrowing/extending moves; the EVEX
instruction set is now complete save for the GPR-interchanging forms. instruction set is now complete save for the GPR-interchanging forms.
### Added ### Added
- `asm`: the unsigned and truncating conversions — VCVTPD2PS (and the X/Y - `asm`: the unsigned and truncating conversions; VCVTPD2PS (and the X/Y
spellings, whose length the spelling fixes), VCVTPD2UDQ (X/Y), spellings, whose length the spelling fixes), VCVTPD2UDQ (X/Y),
VCVTTPD2UDQ (X/Y), VCVTTPD2UQQ, VCVTPS2UDQ, VCVTTPS2UDQ, VCVTPS2UQQ, VCVTTPD2UDQ (X/Y), VCVTTPD2UQQ, VCVTPS2UDQ, VCVTTPS2UDQ, VCVTPS2UQQ,
VCVTTPS2UQQ, VCVTTPD2QQ, VCVTTPS2QQ, VCVTUQQ2PD, VCVTUQQ2PS (X/Y) and VCVTTPS2UQQ, VCVTTPD2QQ, VCVTTPS2QQ, VCVTUQQ2PD, VCVTUQQ2PS (X/Y) and
@@ -681,20 +698,20 @@ instruction set is now complete save for the GPR-interchanging forms.
D2M/Q2M), whose K register is a genuine operand rather than a mask and D2M/Q2M), whose K register is a genuine operand rather than a mask and
which therefore take no masking suffixes. which therefore take no masking suffixes.
## [0.14.0] — 2026-07-19 ## [0.14.0] - 2026-07-19
The floating-point helper and conversion tail of the AVX-512 set, plus The floating-point helper and conversion tail of the AVX-512 set, plus
gather and scatter with VSIB addressing — every encoding verified byte for gather and scatter with VSIB addressing; every encoding verified byte for
byte against the Go assembler. byte against the Go assembler.
### Added ### Added
- `asm`: the floating-point helpers — reciprocals and reciprocal square - `asm`: the floating-point helpers; reciprocals and reciprocal square
roots (VRCP14/VRSQRT14 PD/PS/SD/SS), exponents and mantissas (VGETEXP*, roots (VRCP14/VRSQRT14 PD/PS/SD/SS), exponents and mantissas (VGETEXP*,
VGETMANT*), scaling by powers of two (VSCALEF*), rounding (VRNDSCALE*), VGETMANT*), scaling by powers of two (VSCALEF*), rounding (VRNDSCALE*),
reduction (VREDUCE*), immediate fixup (VFIXUPIMM*) and range selection reduction (VREDUCE*), immediate fixup (VFIXUPIMM*) and range selection
(VRANGE*), and floating-point class tests (VFPCLASSPD/PS X/Y/Z and (VRANGE*), and floating-point class tests (VFPCLASSPD/PS X/Y/Z and
VFPCLASSSD/SS — a new immediate form whose reg field carries the opmask VFPCLASSSD/SS; a new immediate form whose reg field carries the opmask
destination). destination).
- `asm`: **gather and scatter with VSIB addressing.** The gathers take - `asm`: **gather and scatter with VSIB addressing.** The gathers take
both Go spellings: the VEX form with a vector mask register (OP mask, both Go spellings: the VEX form with a vector mask register (OP mask,
@@ -703,14 +720,14 @@ byte against the Go assembler.
data register (a ZMM index with an YMM destination encodes L'L = 10, as data register (a ZMM index with an YMM destination encodes L'L = 10, as
the Go assembler emits). The scatters (VSCATTER*/VPSCATTER*) are EVEX the Go assembler emits). The scatters (VSCATTER*/VPSCATTER*) are EVEX
only (OP src, K, vsib). All eight gather and eight scatter widths. only (OP src, K, vsib). All eight gather and eight scatter widths.
- `asm`: the remaining conversions — VCVTQQ2PS (the 512-bit source sets - `asm`: the remaining conversions; VCVTQQ2PS (the 512-bit source sets
the length), VCVTPD2QQ/UQQ, VCVTPS2QQ, VCVTUDQ2PD/PS, the half-precision the length), VCVTPD2QQ/UQQ, VCVTPS2QQ, VCVTUDQ2PD/PS, the half-precision
VCVTPH2PS and VCVTPS2PH (the extract layout with an immediate). VCVTPH2PS and VCVTPS2PH (the extract layout with an immediate).
## [0.13.0] — 2026-07-18 ## [0.13.0] - 2026-07-18
The wider AVX-512 set: ternary logic, permutes, compares, expand/compress, The wider AVX-512 set: ternary logic, permutes, compares, expand/compress,
the opmask instructions and the EVEX rounding/SAE/broadcast suffixes — every the opmask instructions and the EVEX rounding/SAE/broadcast suffixes; every
encoding verified byte for byte against the Go assembler. encoding verified byte for byte against the Go assembler.
### Added ### Added
@@ -718,7 +735,7 @@ encoding verified byte for byte against the Go assembler.
- `asm`: the wider EVEX/AVX-512 set, across roughly sixty new ground-truth - `asm`: the wider EVEX/AVX-512 set, across roughly sixty new ground-truth
cases: ternary logic (VPTERNLOGD/Q), the lane shuffles/inserts/extracts cases: ternary logic (VPTERNLOGD/Q), the lane shuffles/inserts/extracts
(VSHUF{F,I}{32,64}X{2,4}, the VINSERT*/VEXTRACT* {F,I}{32,64}X{2,4,8} (VSHUF{F,I}{32,64}X{2,4}, the VINSERT*/VEXTRACT* {F,I}{32,64}X{2,4,8}
family, VPALIGNR), compares with an opmask destination (VCMPPD/PS/SD/SS — family, VPALIGNR), compares with an opmask destination (VCMPPD/PS/SD/SS;
a new NDS-plus-immediate form with the K register in the reg field), the a new NDS-plus-immediate form with the K register in the reg field), the
permutes (VPERMB/W, VPERMI2/T2 D/Q/PD), the wider integer families permutes (VPERMB/W, VPERMI2/T2 D/Q/PD), the wider integer families
(VPMADDWD/UBSW, VPMULHUW, VPACKSSWB/USWB/SSDW/USDW, VPABS B/W/D/Q, the (VPMADDWD/UBSW, VPMULHUW, VPACKSSWB/USWB/SSDW/USDW, VPABS B/W/D/Q, the
@@ -732,15 +749,15 @@ encoding verified byte for byte against the Go assembler.
(VMOVSLDUP/VMOVSHDUP), the conversions (VCVTPS2DQ, VCVTTPS2DQ) and the (VMOVSLDUP/VMOVSHDUP), the conversions (VCVTPS2DQ, VCVTTPS2DQ) and the
remaining extending and narrowing moves (VPMOVSXBW, VPMOVZXBW, VPMOVWB, remaining extending and narrowing moves (VPMOVSXBW, VPMOVZXBW, VPMOVWB,
VPMOVQB). VPMOVQB).
- `asm`: the EVEX mnemonic suffixes the Go assembler accepts — the rounding - `asm`: the EVEX mnemonic suffixes the Go assembler accepts; the rounding
modes `.RN_SAE`, `.RD_SAE`, `.RU_SAE`, `.RZ_SAE` (the EVEX b bit with the modes `.RN_SAE`, `.RD_SAE`, `.RU_SAE`, `.RZ_SAE` (the EVEX b bit with the
rounding control in L'L), suppress-all-exceptions `.SAE`, and memory rounding control in L'L), suppress-all-exceptions `.SAE`, and memory
broadcast `.BCST` (the b bit, the vector length preserved, disp8×N scaled broadcast `.BCST` (the b bit, the vector length preserved, disp8×N scaled
by the element size) — each combinable with the `.Z` zeroing suffix, by the element size); each combinable with the `.Z` zeroing suffix,
validated against the Go assembler's bytes, and rejected on instructions validated against the Go assembler's bytes, and rejected on instructions
that do not support them. that do not support them.
## [0.12.0] — 2026-07-17 ## [0.12.0] - 2026-07-17
GOOBJ emission: gasm-assembled functions drop into a `go build` without the GOOBJ emission: gasm-assembled functions drop into a `go build` without the
Go assembler. Go assembler.
@@ -748,7 +765,7 @@ Go assembler.
### Added ### Added
- `asm`: **GOOBJ object output.** `gasm asm --format goobj -p <pkgpath>` - `asm`: **GOOBJ object output.** `gasm asm --format goobj -p <pkgpath>`
writes the Go toolchain's own object format — the one `cmd/link` consumes writes the Go toolchain's own object format; the one `cmd/link` consumes
directly: the functions as non-package symbols qualified with the package directly: the functions as non-package symbols qualified with the package
path (exactly as `cmd/asm` records assembly symbols), the `GLOBL` data, path (exactly as `cmd/asm` records assembly symbols), the `GLOBL` data,
one serialized `FuncInfo` per function (argument/frame sizes, the asm one serialized `FuncInfo` per function (argument/frame sizes, the asm
@@ -757,8 +774,8 @@ Go assembler.
real stack deltas: the assembler now tracks every stack-adjustment real stack deltas: the assembler now tracks every stack-adjustment
boundary through the prologue (`PUSHQ BP`, `SUBQ $frame, SP`) and each boundary through the prologue (`PUSHQ BP`, `SUBQ $frame, SP`) and each
`RET`'s epilogue, so frame-pointer functions unwind correctly. The `RET`'s epilogue, so frame-pointer functions unwind correctly. The
object preamble — the version-and-experiment header the linker compares object preamble; the version-and-experiment header the linker compares
verbatim — is captured from the installed `go tool asm`, so the output is verbatim; is captured from the installed `go tool asm`, so the output is
always consistent with the toolchain that links it. always consistent with the toolchain that links it.
- `asm`: relocations against file-local `GLOBL` symbols become `R_PCREL` - `asm`: relocations against file-local `GLOBL` symbols become `R_PCREL`
entries in the GOOBJ output, with the instruction's displacement field entries in the GOOBJ output, with the instruction's displacement field
@@ -771,7 +788,7 @@ Go assembler.
pattern, instead of being rejected as non-integer. pattern, instead of being rejected as non-integer.
## [0.11.0] — 2026-07-16 ## [0.11.0] - 2026-07-16
Linkable object output: external symbols and relocatable ELF / Mach-O Linkable object output: external symbols and relocatable ELF / Mach-O
objects. objects.
@@ -790,7 +807,7 @@ objects.
external symbol; the Mach-O output is verified structurally with external symbol; the Mach-O output is verified structurally with
`debug/macho`. `debug/macho`.
- `asm`: **external symbol references.** A reference to a symbol no - `asm`: **external symbol references.** A reference to a symbol no
`GLOBL` in the file defines no longer aborts assembly — it is recorded `GLOBL` in the file defines no longer aborts assembly; it is recorded
as an external relocation (`Image.Externals`, `FuncLayout.Relocs`) and as an external relocation (`Image.Externals`, `FuncLayout.Relocs`) and
becomes an undefined global symbol in the object output. The raw image becomes an undefined global symbol in the object output. The raw image
s (`--format raw`, the default) still reports them: only an object s (`--format raw`, the default) still reports them: only an object
@@ -802,7 +819,7 @@ objects.
writes; without `--format` the behaviour is unchanged (the concatenated writes; without `--format` the behaviour is unchanged (the concatenated
image). image).
## [0.10.0] — 2026-07-15 ## [0.10.0] - 2026-07-15
The EVEX floating-point and conversion set: the packed-double arithmetic, The EVEX floating-point and conversion set: the packed-double arithmetic,
the scalar SD/SS forms, VMOVDDUP and the width-changing conversions, each the scalar SD/SS forms, VMOVDDUP and the width-changing conversions, each
@@ -810,23 +827,23 @@ verified byte for byte against the Go assembler.
### Added ### Added
- `asm`: the rest of the common EVEX/VEX floating-point set — packed double - `asm`: the rest of the common EVEX/VEX floating-point set; packed double
arithmetic (VSUBPD, VDIVPD, VMINPD, VMAXPD, VUNPCKLPD and the EVEX form of arithmetic (VSUBPD, VDIVPD, VMINPD, VMAXPD, VUNPCKLPD and the EVEX form of
VUNPCKHPD), the scalar double and single operations (VSUBSD, VDIVSD, VUNPCKHPD), the scalar double and single operations (VSUBSD, VDIVSD,
VMINSD, VMAXSD and the full VADDSS/VSUBSS/VMULSS/VDIVSS/VMINSS/VMAXSS VMINSD, VMAXSD and the full VADDSS/VSUBSS/VMULSS/VDIVSS/VMINSS/VMAXSS
family in both VEX and EVEX — the EVEX scalar forms exist for masked and family in both VEX and EVEX; the EVEX scalar forms exist for masked and
zeroing use), and VMOVDDUP (lane duplication, VEX and EVEX). zeroing use), and VMOVDDUP (lane duplication, VEX and EVEX).
- `asm`: the width-changing conversions — VCVTDQ2PS and VCVTPS2PD (VEX and - `asm`: the width-changing conversions; VCVTDQ2PS and VCVTPS2PD (VEX and
EVEX; the destination sets the length for PS→PD), the EVEX form of EVEX; the destination sets the length for PS→PD), the EVEX form of
VCVTDQ2PD, and the packed-double → dword family: VCVTPD2DQ/VCVTTPD2DQ VCVTDQ2PD, and the packed-double → dword family: VCVTPD2DQ/VCVTTPD2DQ
(EVEX-512 only, a ZMM source and an XMM destination) and their X/Y (EVEX-512 only, a ZMM source and an XMM destination) and their X/Y
spellings (VCVTPD2DQX/Y, VCVTTPD2DQX/Y), whose length follows the wider spellings (VCVTPD2DQX/Y, VCVTTPD2DQX/Y), whose length follows the wider
source — a new operand form, since the destination is always XMM while source; a new operand form, since the destination is always XMM while
VEX.L / EVEX.L'L ride with the source (fixed by the spelling even for a VEX.L / EVEX.L'L ride with the source (fixed by the spelling even for a
memory source). memory source).
- `asm`: masking and zeroing on every new form — the scalar SD/SS - `asm`: masking and zeroing on every new form; the scalar SD/SS
arithmetic, the unpacks, VMOVDDUP and the conversions all accept the arithmetic, the unpacks, VMOVDDUP and the conversions all accept the
explicit K1–K7 operand and the `.Z` suffix the way Go writes them. explicit K1-K7 operand and the `.Z` suffix the way Go writes them.
### Changed ### Changed
@@ -837,35 +854,35 @@ verified byte for byte against the Go assembler.
shares the convention). shares the convention).
## [0.9.0] — 2026-07-14 ## [0.9.0] - 2026-07-14
AVX-512 masking and a wider EVEX integer set. AVX-512 masking and a wider EVEX integer set.
### Added ### Added
- `asm`: **EVEX masking** the way Go writes it — an explicit `K1`–`K7` - `asm`: **EVEX masking** the way Go writes it; an explicit `K1`-`K7`
operand placed among the operands (merging mask), and a `.Z` mnemonic operand placed among the operands (merging mask), and a `.Z` mnemonic
suffix for zeroing (`VPADDD.Z Z1, Z2, K2, Z3`). Supported across the NDS, suffix for zeroing (`VPADDD.Z Z1, Z2, K2, Z3`). Supported across the NDS,
reg/rm, immediate-shift, align, extract, convert and move forms, including reg/rm, immediate-shift, align, extract, convert and move forms, including
masked comparisons with a K destination (`VPCMPEQD Z0, Z3, K2, K1`). K0 is masked comparisons with a K destination (`VPCMPEQD Z0, Z3, K2, K1`). K0 is
rejected as an explicit mask, and `.Z` without a mask is an error, matching rejected as an explicit mask, and `.Z` without a mask is an error, matching
the Go assembler. the Go assembler.
- `asm`: the common AVX-512 F/BW integer set — VPADDB/W, VPSUBB/W, VPANDD/Q, - `asm`: the common AVX-512 F/BW integer set; VPADDB/W, VPSUBB/W, VPANDD/Q,
VPANDND/Q, VPMULLW, VPAVGB/W, the signed/unsigned min/max family for VPANDND/Q, VPMULLW, VPAVGB/W, the signed/unsigned min/max family for
B/W/D/Q elements, the variable shifts VPSLLVD/Q, VPSRLVD/Q, VPSRAVD/Q, the B/W/D/Q elements, the variable shifts VPSLLVD/Q, VPSRLVD/Q, VPSRAVD/Q, the
EVEX forms of VPSHUFD/VPSHUFB, and the VMOVDQU8/VMOVDQU16 move aliases. EVEX forms of VPSHUFD/VPSHUFB, and the VMOVDQU8/VMOVDQU16 move aliases.
Register indices 16–31 encode correctly (the mod=11 quirk carries rm[4] Register indices 16-31 encode correctly (the mod=11 quirk carries rm[4]
in X̄). All verified byte for byte against the Go assembler. in X̄). All verified byte for byte against the Go assembler.
- `lint`: masked EVEX forms (`.Z` suffix, K operands) are recognised by - `lint`: masked EVEX forms (`.Z` suffix, K operands) are recognised by
`unknown-instruction` and exempted from `operand-count`. `unknown-instruction` and exempted from `operand-count`.
### Fixed ### Fixed
- `asm`: EVEX register–register operands with indices 16–31 encoded rm[4] - `asm`: EVEX register-register operands with indices 16-31 encoded rm[4]
into B̄ instead of X̄ (the EVEX mod=11 extension quirk), producing wrong into B̄ instead of X̄ (the EVEX mod=11 extension quirk), producing wrong
prefix bytes for X16+/Y16+ r/m operands. prefix bytes for X16+/Y16+ r/m operands.
## [0.8.0] — 2026-07-13 ## [0.8.0] - 2026-07-13
Standard CLI ergonomics. Standard CLI ergonomics.
@@ -881,20 +898,20 @@ Standard CLI ergonomics.
- The version is primarily available as the standard `gasm --version` / `-V` - The version is primarily available as the standard `gasm --version` / `-V`
flag; the `gasm version` spelling remains as an alias. flag; the `gasm version` spelling remains as an alias.
## [0.7.0] — 2026-07-12 ## [0.7.0] - 2026-07-12
The formatter behaves like `go fmt` and canonicalises block separation. The formatter behaves like `go fmt` and canonicalises block separation.
### Added ### Added
- `gasm fmt` now works like `go fmt`: with no arguments — or with a directory - `gasm fmt` now works like `go fmt`: with no arguments; or with a directory
argument — it reformats every `.s` file below it in place and lists the argument; it reformats every `.s` file below it in place and lists the
changed files, skipping `.` and `_` directories (`.git`, `_refs`, …). changed files, skipping `.` and `_` directories (`.git`, `_refs`, …).
Explicit file arguments keep the `-w` / standard-output behaviour. Explicit file arguments keep the `-w` / standard-output behaviour.
### Changed ### Changed
- `s`: canonical blank-line layout — a new block (a label, `TEXT` or - `s`: canonical blank-line layout; a new block (a label, `TEXT` or
`GLOBL`) is preceded by exactly one blank line, neither more nor less. `GLOBL`) is preceded by exactly one blank line, neither more nor less.
Comments leading a block stay with it (the blank line goes before them), Comments leading a block stay with it (the blank line goes before them),
stacked labels share their block, the function's first label keeps hugging stacked labels share their block, the function's first label keeps hugging
@@ -903,7 +920,7 @@ The formatter behaves like `go fmt` and canonicalises block separation.
kernels were reformatted with this release and remain byte-identical when kernels were reformatted with this release and remain byte-identical when
assembled. assembled.
## [0.6.0] — 2026-07-11 ## [0.6.0] - 2026-07-11
Calibrated to the Go ABI: `register-clobber` stops reporting legal code, and Calibrated to the Go ABI: `register-clobber` stops reporting legal code, and
the encoder learns the legacy SSE moves. the encoder learns the legacy SSE moves.
@@ -912,13 +929,13 @@ the encoder learns the legacy SSE moves.
- `lint`: **`register-clobber` is now calibrated to the Go ABI** - `lint`: **`register-clobber` is now calibrated to the Go ABI**
(`cmd/compile/abi-internal.md`), not the platform ABI. Go's stack-based (`cmd/compile/abi-internal.md`), not the platform ABI. Go's stack-based
ABI0 has no System V style callee-saved registers — amd64 `BX`, `R12`–`R15` ABI0 has no System V style callee-saved registers; amd64 `BX`, `R12`-`R15`
and the arm64/riscv64/loong64 scratch sets are caller-saved or permanent and the arm64/riscv64/loong64 scratch sets are caller-saved or permanent
scratch, and hand-written kernels may clobber them freely. The rule now scratch, and hand-written kernels may clobber them freely. The rule now
audits only the registers Go fixes across calls: the frame pointer and the audits only the registers Go fixes across calls: the frame pointer and the
goroutine pointer (amd64 `BP`/`R14`, arm64 `R18`/`R28`/`R29`, riscv64 goroutine pointer (amd64 `BP`/`R14`, arm64 `R18`/`R28`/`R29`, riscv64
`X27`, loong64 `R22`), and the goroutine pointer is reported only when the `X27`, loong64 `R22`), and the goroutine pointer is reported only when the
function can reach the runtime (is not `NOSPLIT` or makes a call) — the function can reach the runtime (is not `NOSPLIT` or makes a call); the
ABI0 transition restores it on those paths, and NOSPLIT call-free leaves ABI0 transition restores it on those paths, and NOSPLIT call-free leaves
may use it, exactly as the runtime's own assembly does. Both go-flac may use it, exactly as the runtime's own assembly does. Both go-flac
kernels now lint with zero diagnostics. kernels now lint with zero diagnostics.
@@ -935,21 +952,21 @@ the encoder learns the legacy SSE moves.
### Added ### Added
- `asm`: the legacy (non-VEX) SSE moves — `MOVOU`/`MOVO` (the Plan 9 names - `asm`: the legacy (non-VEX) SSE moves; `MOVOU`/`MOVO` (the Plan 9 names
for MOVDQU/MOVDQA), `MOVUPS`/`MOVAPS`/`MOVUPD`/`MOVAPD` and the scalar for MOVDQU/MOVDQA), `MOVUPS`/`MOVAPS`/`MOVUPD`/`MOVAPD` and the scalar
`MOVSD`/`MOVSS` — and `VMOVDQU64` in the EVEX set. All verified byte for `MOVSD`/`MOVSS`; and `VMOVDQU64` in the EVEX set. All verified byte for
byte against the Go assembler. byte against the Go assembler.
## [0.5.0] — 2026-07-10 ## [0.5.0] - 2026-07-10
EVEX / AVX-512: the go-flac AVX-512 kernel now assembles, byte-identically to EVEX / AVX-512: the go-flac AVX-512 kernel now assembles, byte-identically to
the Go toolchain, completing the production-kernel coverage. the Go toolchain, completing the production-kernel coverage.
### Added ### Added
- `asm`: **EVEX (AVX-512) encoding** — the four-byte EVEX prefix with the - `asm`: **EVEX (AVX-512) encoding**; the four-byte EVEX prefix with the
5-bit register fields (Z0–Z31, X/Y 16–31, with the reg-r/m X̄ quirk and 5-bit register fields (Z0-Z31, X/Y 16-31, with the reg-r/m X̄ quirk and
V'̄ shared between vvvv and the SIB index), opmask registers (K0–K7) as V'̄ shared between vvvv and the SIB index), opmask registers (K0-K7) as
operands and as mask destinations, and the compressed disp8×N displacement operands and as mask destinations, and the compressed disp8×N displacement
(the multiplier follows the memory operand's size, as the Go assembler's (the multiplier follows the memory operand's size, as the Go assembler's
opcode tables prescribe). Covers every AVX-512 instruction the go-flac opcode tables prescribe). Covers every AVX-512 instruction the go-flac
@@ -959,20 +976,20 @@ the Go toolchain, completing the production-kernel coverage.
extracts VEXTRACTI64X4/VEXTRACTF64X4, VFMADD231PD, VADDPD, VMULPD, the extracts VEXTRACTI64X4/VEXTRACTF64X4, VFMADD231PD, VADDPD, VMULPD, the
broadcasts VPBROADCASTD/Q (GPR and memory sources take different opcodes) broadcasts VPBROADCASTD/Q (GPR and memory sources take different opcodes)
and the mask moves KMOVW/KTESTW. Masking/zeroing suffixes are out of scope and the mask moves KMOVW/KTESTW. Masking/zeroing suffixes are out of scope
— the kernels use neither. ; the kernels use neither.
- `asm`: `AssembleFile` now accepts file-defined global (`non-<>`) symbols - `asm`: `AssembleFile` now accepts file-defined global (`non-<>`) symbols
too; a reference is external only when no `GLOBL` in the file defines it. too; a reference is external only when no `GLOBL` in the file defines it.
### Fixed ### Fixed
- `asm`: registers X16–Y31 force the EVEX encoding of dual-form mnemonics; - `asm`: registers X16-Y31 force the EVEX encoding of dual-form mnemonics;
previously a `VPBROADCASTD AX, Y30` fell into the VEX encoder, which cannot previously a `VPBROADCASTD AX, Y30` fell into the VEX encoder, which cannot
represent indices above 15 and silently truncated them. represent indices above 15 and silently truncated them.
- `asm`: the VEX encoder now rejects vector register indices 16–31 instead of - `asm`: the VEX encoder now rejects vector register indices 16-31 instead of
encoding a truncated (wrong) register. encoding a truncated (wrong) register.
## [0.4.0] — 2026-07-09 ## [0.4.0] - 2026-07-09
The standalone assembler reaches the whole go-flac AVX2 kernel: static The standalone assembler reaches the whole go-flac AVX2 kernel: static
symbols assemble, and all 17 kernel functions now match the Go toolchain's symbols assemble, and all 17 kernel functions now match the Go toolchain's
@@ -980,10 +997,10 @@ machine code byte for byte.
### Added ### Added
- `asm`: **file-level assembly** — `AssembleFile` turns a parsed file into an - `asm`: **file-level assembly**; `AssembleFile` turns a parsed file into an
`Image`: the function bodies in source order followed by a data section `Image`: the function bodies in source order followed by a data section
built from the file's `GLOBL`/`DATA` directives (each symbol 16-aligned). built from the file's `GLOBL`/`DATA` directives (each symbol 16-aligned).
- `asm`: **static-symbol (`SB`) operands** — `mask<>(SB)` references encode as - `asm`: **static-symbol (`SB`) operands**; `mask<>(SB)` references encode as
RIP-relative loads with a patched disp32, resolved against the image layout RIP-relative loads with a patched disp32, resolved against the image layout
so the output is self-consistent and position-independent. External so the output is self-consistent and position-independent. External
(non-file-local) symbols are rejected with a clear error: they need (non-file-local) symbols are rejected with a clear error: they need
@@ -992,7 +1009,7 @@ machine code byte for byte.
and writes the whole image (code + data) with `-o`. and writes the whole image (code + data) with `-o`.
## [0.3.0] — 2026-07-08 ## [0.3.0] - 2026-07-08
The assembler reaches byte-identical parity with the Go toolchain on the The assembler reaches byte-identical parity with the Go toolchain on the
production go-flac AVX2 kernels: every one of the 15 kernel functions that production go-flac AVX2 kernels: every one of the 15 kernel functions that
@@ -1002,18 +1019,18 @@ support).
### Added ### Added
- `asm`: the scalar instruction families the kernels use — `CMOVcc` and - `asm`: the scalar instruction families the kernels use; `CMOVcc` and
`SETcc` (conditions spelled exactly like the jumps), `LZCNT`/`TZCNT` `SETcc` (conditions spelled exactly like the jumps), `LZCNT`/`TZCNT`
(legacy `F3 0F BD/BC`), the sign/zero-extending moves (`MOVBLZX`, `MOVBQZX`, (legacy `F3 0F BD/BC`), the sign/zero-extending moves (`MOVBLZX`, `MOVBQZX`,
`MOVWLZX`, `MOVWQZX`, `MOVWLSX`, `MOVLQSX`), `CVTSL2SD`/`CVTSQ2SD` (the `MOVWLZX`, `MOVWQZX`, `MOVWLSX`, `MOVLQSX`), `CVTSL2SD`/`CVTSQ2SD` (the
legacy SSE encoding, as the Go assembler emits it), the traditional legacy SSE encoding, as the Go assembler emits it), the traditional
three-operand `IMUL3{W,L,Q}`, and the variable-count vector shifts three-operand `IMUL3{W,L,Q}`, and the variable-count vector shifts
(`VPSRLQ X0, Y8, Y8` — the count in an XMM register or memory takes the (`VPSRLQ X0, Y8, Y8`; the count in an XMM register or memory takes the
ordinary NDS form). ordinary NDS form).
- `asm`: **jump relaxation** — jumps start in the short (rel8) form and - `asm`: **jump relaxation**; jumps start in the short (rel8) form and
expand to rel32 when the settled displacement does not fit, iterating the expand to rel32 when the settled displacement does not fit, iterating the
layout to a fixed point (CALL is always rel32). layout to a fixed point (CALL is always rel32).
- `asm`: **jump-to-jump folding** — a conditional jump to a label whose only - `asm`: **jump-to-jump folding**; a conditional jump to a label whose only
instruction is an unconditional jump is redirected to the ultimate instruction is an unconditional jump is redirected to the ultimate
target, replicating the Go toolchain's linker, which chases such chains target, replicating the Go toolchain's linker, which chases such chains
before it encodes branches. before it encodes branches.
@@ -1025,12 +1042,12 @@ support).
- `asm`: `CMP` with a register or memory operand computed **second − first** - `asm`: `CMP` with a register or memory operand computed **second − first**
instead of first − second, silently inverting every condition that followed instead of first − second, silently inverting every condition that followed
(`CMPQ SI, R10; JGE` tested R10 ≥ SI). The encoding now always records (`CMPQ SI, R10; JGE` tested R10 ≥ SI). The encoding now always records
first − second — `CMP r/m, r` with the first operand in r/m, `CMP r, r/m` first − second; `CMP r/m, r` with the first operand in r/m, `CMP r, r/m`
with the first operand in reg — and is byte-identical to the Go assembler. with the first operand in reg; and is byte-identical to the Go assembler.
- `asm`: register-to-register `MOV` now uses the `r/m ← r` opcode (reg = - `asm`: register-to-register `MOV` now uses the `r/m ← r` opcode (reg =
source), the Go assembler's choice; the output is byte-identical. source), the Go assembler's choice; the output is byte-identical.
## [0.2.0] — 2026-07-07 ## [0.2.0] - 2026-07-07
The Phase 2 assembler grows the SIMD set: shuffles, extract/insert, permute The Phase 2 assembler grows the SIMD set: shuffles, extract/insert, permute
and the moves, on top of the Phase 1 VEX forms. and the moves, on top of the Phase 1 VEX forms.
@@ -1043,10 +1060,10 @@ and the moves, on top of the Phase 1 VEX forms.
- the immediate shuffle (`VPSHUFD`, `VPERMQ`), - the immediate shuffle (`VPSHUFD`, `VPERMQ`),
- the three-operand-plus-immediate form (`VSHUFPD`, `VPERM2I128`, - the three-operand-plus-immediate form (`VSHUFPD`, `VPERM2I128`,
`VINSERTI128`), `VINSERTI128`),
- the lane extract (`VEXTRACTI128`, `VEXTRACTF128` — the YMM source occupies - the lane extract (`VEXTRACTI128`, `VEXTRACTF128`; the YMM source occupies
the ModRM.reg field, the XMM/memory destination the r/m field), the ModRM.reg field, the XMM/memory destination the r/m field),
- the direction-sensitive moves (`VMOVDQU`, `VMOVUPD`, `VMOVD`, `VMOVQ`, - the direction-sensitive moves (`VMOVDQU`, `VMOVUPD`, `VMOVD`, `VMOVQ`,
`VMOVSD` — each direction picks its own opcode and VEX.W; a vector→vector `VMOVSD`; each direction picks its own opcode and VEX.W; a vector→vector
move uses the store-form layout, matching the Go assembler), move uses the store-form layout, matching the Go assembler),
- the no-operand `VZEROUPPER`, and `VPERMD` in the NDS form, - the no-operand `VZEROUPPER`, and `VPERMD` in the NDS form,
- the floating-point and FMA set (`VADDPD`, `VMULPD`, `VXORPD`, - the floating-point and FMA set (`VADDPD`, `VMULPD`, `VXORPD`,
@@ -1055,21 +1072,21 @@ and the moves, on top of the Phase 1 VEX forms.
the encoder now covers every integer, shuffle and FP instruction the the encoder now covers every integer, shuffle and FP instruction the
go-flac AVX2 kernels use. go-flac AVX2 kernels use.
- `asm`: `CMP` accepts the immediate in the second operand position - `asm`: `CMP` accepts the immediate in the second operand position
(`CMPL CX, $31`) — the spelling the Go assembler accepts — encoding it (`CMPL CX, $31`), the spelling the Go assembler accepts, encoding it
identically to the immediate-first form. identically to the immediate-first form.
### Fixed ### Fixed
- `asm`: an unused VEX.vvvv field is now stored as `1111` (v̄vvv = 1111), as - `asm`: an unused VEX.vvvv field is now stored as `1111` (v̄vvv = 1111), as
the hardware requires — the previous value (`0000`) made the two-operand the hardware requires; the previous value (`0000`) made the two-operand
reg/rm forms (VPMOVSXWD, VPBROADCASTD, VMOVMSKPS, …) raise #UD on real CPUs reg/rm forms (VPMOVSXWD, VPBROADCASTD, VMOVMSKPS, …) raise #UD on real CPUs
and differ from the Go assembler's bytes. The round-trip decoder ignores and differ from the Go assembler's bytes. The round-trip decoder ignores
the field on these instructions, which is why the byte-for-byte Go the field on these instructions, which is why the byte-for-byte Go
comparison (added this release) is now part of the test suite. comparison (added this release) is now part of the test suite.
## [0.1.0] — 2026-07-06 ## [0.1.0] - 2026-07-06
Initial release — the Phase 1 foundation. Initial release; the Phase 1 foundation.
### Added ### Added
@@ -1082,11 +1099,11 @@ Initial release — the Phase 1 foundation.
- `arch`: register files and **complete** instruction tables for amd64, - `arch`: register files and **complete** instruction tables for amd64,
arm64, riscv64 and loong64, with the middle-dot symbol separator and static arm64, riscv64 and loong64, with the middle-dot symbol separator and static
(`<>`) symbols. Instruction names are generated from the Go toolchain's own (`<>`) symbols. Instruction names are generated from the Go toolchain's own
assembler source (`just gen`) — the `anames` opcode lists plus the common assembler source (`just gen`); the `anames` opcode lists plus the common
opcodes and the per-architecture front-end aliases (arm64 `B`/`BL`, the opcodes and the per-architecture front-end aliases (arm64 `B`/`BL`, the
`.P`/`.W` addressing suffixes, loong64 `JAL`, the x86 conditional-jump `.P`/`.W` addressing suffixes, loong64 `JAL`, the x86 conditional-jump
spellings) — so every mnemonic the real assembler accepts is recognised. spellings); so every mnemonic the real assembler accepts is recognised.
- `lint`: conservative rules — `unknown-instruction`, `operand-count`, - `lint`: conservative rules; `unknown-instruction`, `operand-count`,
`undefined-label`, `duplicate-label`, `missing-ret`, `undefined-label`, `duplicate-label`, `missing-ret`,
`missing-textflag-include`, `abi-argsize` and `unreachable-code`. Macro `missing-textflag-include`, `abi-argsize` and `unreachable-code`. Macro
invocations are recognised (in-file `#define` names and underscore invocations are recognised (in-file `#define` names and underscore
@@ -1108,9 +1125,9 @@ Initial release — the Phase 1 foundation.
- `lsp`: a Language Server Protocol server over stdio providing completion, - `lsp`: a Language Server Protocol server over stdio providing completion,
hover documentation, document symbols, publish-diagnostics and semantic-token hover documentation, document symbols, publish-diagnostics and semantic-token
highlighting. highlighting.
- `asm`: a standalone amd64 (x86-64) assembler — an instruction encoder (REX/ - `asm`: a standalone amd64 (x86-64) assembler; an instruction encoder (REX/
ModR-M/SIB/displacement/immediate plus the scalar instruction set, and VEX/ ModR-M/SIB/displacement/immediate plus the scalar instruction set, and VEX/
AVX2 SIMD across three operand forms — NDS, reg/rm and immediate-shift — AVX2 SIMD across the three operand forms NDS, reg/rm and immediate-shift,
covering the bulk of the integer SIMD set) validated by round-trip decoding covering the bulk of the integer SIMD set) validated by round-trip decoding
against `golang.org/x/arch`, and an assembler that drives the parser's AST against `golang.org/x/arch`, and an assembler that drives the parser's AST
into the encoder with local-label resolution and `FP`/`SP` frame mapping into the encoder with local-label resolution and `FP`/`SP` frame mapping
+40
View File
@@ -0,0 +1,40 @@
# Security policy
## Supported versions
Security fixes go to the newest release and to the `development` branch. Older
releases do not receive them.
| Version | Supported |
|---|---|
| 0.33.0 | yes |
| older releases | no |
## Reporting a vulnerability
**Do not open a public issue for a security problem.** A public report tells everyone
about the flaw before there is a fix. Report it privately to
**opensource@petrbalvin.org**.
Include:
- the version or commit you tested, and the platform
- what the problem is, and what an attacker gains from it
- the smallest reproducer you have, ideally a test or a single command
- a suggested fix, if you have one
## What to expect
- A human reads the report, and you get an acknowledgement.
- You are kept informed while the fix is being made, and told when it ships.
- The fix is released before the details are published, and the timing is agreed with
you.
- The reporter is credited in the release notes unless they ask otherwise.
## Out of scope
- Findings that require the attacker to already run code as the user, or to have local
access.
- Missing hardening with no demonstrated impact.
- Flaws in a third-party dependency: report them to that project, and to this one only
when this project's use of it makes them reachable.