docs: add SECURITY.md and record the round in the CHANGELOG

Assisted-by: GLM 5.3 Flash
This commit is contained in:
2026-09-16 23:12:31 +02:00
parent 0078f7be5c
commit 3de043c494
2 changed files with 197 additions and 140 deletions
+157 -140
View File
@@ -7,11 +7,28 @@ and this project adheres to [Conventional Commits](https://www.conventionalcommi
## [development]
### Added
### Changed
-
- **Canonical just recipes.** `just gates` is the definition of done
(build, fmt-check, vet, test, race). `install` now builds and copies
the binary into `~/.local/bin` (`BINDIR` overrides) instead of
downloading module dependencies, and `install-bin` is gone. The test
gate sweeps the whole suite and computes the coverage floor over the
product packages with `-coverpkg`, `disasm` now included, so the
number is identical locally and in CI. `fuzz` requires its target
package.
- **The reported version comes from the build.** `gasm --version`
prints the version the toolchain recorded: the tag on a tagged
checkout, a pseudo-version naming the commit below one, `+dirty` on a
dirty tree and `(devel)` outside version control. Nothing is
injected with `-ldflags -X` any more.
- **CI realigned with the gate set.** The push pipeline runs the gates
minus race in one job, with a cached Go setup and the module as the
version source; the race detector moved to a hand-dispatched workflow
and into the release gates; the release builds without injection and
its smoke test requires the recorded tag and rejects `+dirty`.
## [0.33.0] — 2026-09-14
## [0.33.0] - 2026-09-14
### Added
@@ -67,7 +84,7 @@ and this project adheres to [Conventional Commits](https://www.conventionalcommi
the toolchain picks, and the morestack block saves the link register
with the toolchain's `OR` form on loong64.
## [0.32.0] — 2026-08-31
## [0.32.0] - 2026-08-31
### Added
@@ -218,7 +235,7 @@ and this project adheres to [Conventional Commits](https://www.conventionalcommi
outputs and operand strictness with `go tool asm`.
- **asm help text.** Updated to list arm64 as a supported architecture.
## [0.31.1] — 2026-08-20
## [0.31.1] - 2026-08-20
### Fixed
@@ -226,9 +243,9 @@ and this project adheres to [Conventional Commits](https://www.conventionalcommi
because the version variables in `justfile` and `cmd/gasm/main.go` were not
bumped during the release commit.
## [0.31.0] — 2026-08-20
## [0.31.0] - 2026-08-20
The arm64 encoder (Phase 5 — complete) ships with ELF64 and GOOBJ emission,
The arm64 encoder (Phase 5; complete) ships with ELF64 and GOOBJ emission,
verified byte-for-byte against `GOARCH=arm64 go tool asm` and linked into a
real `go build`. The encoder covers the full integer instruction set, FP
arithmetic, conditional select, CRC32, and the MOV pseudo-instruction with
@@ -236,7 +253,7 @@ bitmask immediate encoding. The project now requires Go 1.27.
### Added
- **arm64 encoder (Phase 5 — complete).** `gasm asm` can now assemble `_arm64.s`
- **arm64 encoder (Phase 5; complete).** `gasm asm` can now assemble `_arm64.s`
files: the AArch64 integer instruction set with the MOV pseudo-instruction and
its immediate-constant expansions (MOVZ/MOVN/MOVK for wide immediates, ORR with
logical bitmask encoding for values like `$1`), data-processing (shifted
@@ -245,7 +262,7 @@ bitmask immediate encoding. The project now requires Go 1.27.
SB/global symbol references (ADRP+ADD pairs with `R_ADDRARM64` relocations),
jump chain folding, and ELF64 emission (`gasm asm --format elf`). Ground-truth
verification against `GOARCH=arm64 go tool asm` matches byte-for-byte. Phase 5
(the other architectures — RISC-V, LoongArch, arm64) is now complete.
(the other architectures; RISC-V, LoongArch, arm64) is now complete.
### Changed
@@ -253,7 +270,7 @@ bitmask immediate encoding. The project now requires Go 1.27.
The `R_DWTXTADDR_U4` relocation type is detected at runtime for backward
compatibility.
## [0.30.0] — 2026-08-13
## [0.30.0] - 2026-08-13
The LoongArch encoder (Phase 5) ships with ELF64 and GOOBJ emission, verified
byte-for-byte against `GOARCH=loong64 go tool asm` and linked into a real
@@ -278,7 +295,7 @@ tracks four hardware watchpoint slots, and the toolkit is Linux-only.
- **GOOBJ DWARF symbols.** The GOOBJ emitters now write the per-function
DWARF symbols the linker requires (the subprogram DIE and the `.debug_line`
program, byte-identical to `cmd/asm`'s), and the pc-value table deltas are
in the architecture's MinLC units as the runtime expects — the amd64 link
in the architecture's MinLC units as the runtime expects; the amd64 link
test now genuinely substitutes the gasm object, and the amd64/loong64
end-to-end GOOBJ link tests pass.
- **RISC-V GOOBJ emission via the shared emitter.** RISC-V GOOBJ output is
@@ -347,7 +364,7 @@ tracks four hardware watchpoint slots, and the toolkit is Linux-only.
`GOARCH=riscv64 go tool asm`.
- **Debugger watchpoint slots.** `gasm debug`'s `watch` command always used
hardware watchpoint slot 0, so a second `watch` call silently overwrote
the first. Watchpoint slots are now tracked in the `Session` (DR0–DR3);
the first. Watchpoint slots are now tracked in the `Session` (DR0-DR3);
`watch` picks the first free slot and reports an error if all four are in
use, and `unwatch <slot>` clears one (no argument clears all).
@@ -359,7 +376,7 @@ tracks four hardware watchpoint slots, and the toolkit is Linux-only.
disassembly at PC, memory-write, watchpoints, and source-line mapping are
all shipped.
## [0.29.0] — 2026-08-07
## [0.29.0] - 2026-08-07
RISC-V GOOBJ emission, YMM vector register display, named buffer allocation
in the debugger, two new CLI commands (`diff`, `profile`), go-to-definition in
@@ -369,30 +386,30 @@ new CLI commands. A signature-parser fix corrects grouped Go parameters.
### Added
- **RISC-V GOOBJ emission** — `gasm asm --format goobj` for RISC-V produces
- **RISC-V GOOBJ emission**; `gasm asm --format goobj` for RISC-V produces
linkable Go objects with funcdata, pc-value tables, and RISC-V relocation
types (same format as amd64 GOOBJ, with the RISC-V architecture marker).
- **`gasm diff`** — compare the machine code of two assembly files byte-for-byte;
- **`gasm diff`**; compare the machine code of two assembly files byte-for-byte;
shows which functions differ and the first few differing bytes.
- **`gasm profile`** — show the basic-block structure of each function: labels,
- **`gasm profile`**; show the basic-block structure of each function: labels,
offsets, frame size, and NOSPLIT flag.
- **LSP go-to-definition** — `textDocument/definition` navigates from a label
- **LSP go-to-definition**; `textDocument/definition` navigates from a label
reference to its definition.
- **did-you-mean** — when the RISC-V assembler encounters an undefined label, it
- **did-you-mean**; when the RISC-V assembler encounters an undefined label, it
suggests the closest existing label using Levenshtein distance.
- **YMM vector register display** — `regs` in the debugger now shows YMM
- **YMM vector register display**; `regs` in the debugger now shows YMM
registers via `PTRACE_GETFPREGS` (falls back to XMM when XSAVE is unavailable).
- **Named buffer allocation** — `gasm debug --buf name:size:pattern` allocates
- **Named buffer allocation**; `gasm debug --buf name:size:pattern` allocates
buffers in the debuggee filled with `zero`, `ones`, `seq`, or a hex pattern;
buffer pointers are placed into the argument block at the matching positions.
- **Crash input storage** — `FuzzResult.CrashInput` stores the input that caused
- **Crash input storage**; `FuzzResult.CrashInput` stores the input that caused
a crash or mismatch for reproducibility.
- **ABI + fuzz combined** — `gasm verify --fuzz` now runs ABI checks (sentinel
- **ABI + fuzz combined**; `gasm verify --fuzz` now runs ABI checks (sentinel
registers, canary, stack bounds) alongside differential fuzz testing.
- **`gasm diff --map`** — compare functions whose names differ between files
- **`gasm diff --map`**; compare functions whose names differ between files
(e.g. `--map wideCopyAVX2=wideCopyAVX512` pairs two variants regardless
of suffix). Unmapped functions fall back to the original name match.
- **`gasm verify --call`** — invoke a single function with user-supplied buffers
- **`gasm verify --call`**; invoke a single function with user-supplied buffers
(`--buf name:size:pattern`) instead of the smoke/abi/fuzz sweeps. Patterns:
`zero`, `ones`, `seq`, or a hex blob. Useful for partial functions (e.g.
decoders) that crash on random input but should succeed on valid data.
@@ -402,23 +419,23 @@ new CLI commands. A signature-parser fix corrects grouped Go parameters.
### Fixed
- **Signature parser** — grouped Go parameters like `dst, src []byte` are now
- **Signature parser**; grouped Go parameters like `dst, src []byte` are now
parsed correctly (both get type `[]byte`). Previously the first name was
treated as its own type (`dst` with size 8), causing wrong ABI0 arg-block
layout in both `verify --call` and the fuzzer.
- **Flaky JIT tests** — `runtime.KeepAlive` guards and package-level buffers
- **Flaky JIT tests**; `runtime.KeepAlive` guards and package-level buffers
prevent GC from collecting heap objects whose addresses were passed to JIT
code via `unsafe.Pointer`; all verify tests pass 100/100 under `-race`.
### Changed
- **Removed external kernel test dependencies** — the verify test suite no
- **Removed external kernel test dependencies**; the verify test suite no
longer references production kernels from the separate go-libraries project.
The remaining test suite uses only `testdata/verify/*.s` kernels, which are
part of this repository. Coverage is identical locally and in CI (80.3 %).
## [0.28.0] — 2026-08-03
## [0.28.0] - 2026-08-03
RISC-V encoder: full RV64IMAFDC instruction set with RVC compression, MOV
pseudo-instruction, SB/global symbol references, ELF64 object emission, and
@@ -426,20 +443,20 @@ ground-truth verification against `GOARCH=riscv64 go tool asm`.
### Added
- **RISC-V encoder** — RV64I, RV64M, RV64A, RV64F/D, FMA, CSR, JALR.
- **MOV pseudo-instruction** — load, store, reg-to-reg, immediate, frame mapping.
- **RVC compression** — 22 compressed instruction types (C.LDSP, C.SDSP, C.FLDSP,
- **RISC-V encoder**; RV64I, RV64M, RV64A, RV64F/D, FMA, CSR, JALR.
- **MOV pseudo-instruction**; load, store, reg-to-reg, immediate, frame mapping.
- **RVC compression**; 22 compressed instruction types (C.LDSP, C.SDSP, C.FLDSP,
C.FSDSP, C.ADDI, C.LI, C.LUI, C.ADDIW, C.MV, C.ADD, C.SUB, C.XOR, C.OR, C.AND,
C.SLLI, C.SRLI, C.SRAI, C.ANDI, C.BEQZ, C.BNEZ, C.J, C.JR).
- **SB/global symbols** — `MOV $sym(SB)`, `MOV sym(SB)`, `MOV rd, sym(SB)`
- **SB/global symbols**; `MOV $sym(SB)`, `MOV sym(SB)`, `MOV rd, sym(SB)`
encoded as AUIPC pairs with R_RISCV_PCREL_HI20/LO12 relocations.
- **GLOBL/DATA** — data section layout in `AssembleFileRISCV`.
- **ELF64 emission** — `gasm asm --format elf` produces EM_RISCV objects
- **GLOBL/DATA**; data section layout in `AssembleFileRISCV`.
- **ELF64 emission**; `gasm asm --format elf` produces EM_RISCV objects
(.text, .data, .symtab, .rela.text).
- **`gasm verify --ground-truth`** — byte-exact comparison against
- **`gasm verify --ground-truth`**; byte-exact comparison against
`GOARCH=riscv64 go tool asm`.
- **`gasm verify --profile`** — function layout listing for RISC-V.
- **CALL** — AUIPC + JALR pair encoding.
- **`gasm verify --profile`**; function layout listing for RISC-V.
- **CALL**; AUIPC + JALR pair encoding.
### Fixed
@@ -449,7 +466,7 @@ ground-truth verification against `GOARCH=riscv64 go tool asm`.
(bit-interleaved format).
## [0.27.0] — 2026-08-01
## [0.27.0] - 2026-08-01
Subprocess isolation for `--fuzz`: each function is fuzzed in its own child
process, so a partial function (decoder) that faults on random garbage is
@@ -460,7 +477,7 @@ the parent. CRASH is informational (exit 0); only MISMATCH is an error.
- `gasm verify --fuzz` no longer crashes the process on partial functions.
## [0.26.0] — 2026-07-31
## [0.26.0] - 2026-07-31
Universal differential fuzzing: `gasm verify --fuzz` needs no hand-written
reference. It parses the `// func` signature from the assembly source,
@@ -471,7 +488,7 @@ area bit-for-bit.
### Added
- `verify`: `FuzzFunc` / `ExtractSignatures` / `parseFuncSig` — universal
- `verify`: `FuzzFunc` / `ExtractSignatures` / `parseFuncSig`; universal
differential fuzz driven by the conventional `// func` comment. Each
version gets its own buffer set (deep copy) so functions that write to
their arguments (histogram increments) don't corrupt the other's input.
@@ -486,16 +503,16 @@ area bit-for-bit.
over-copy paths read past the buffer on random garbage input. Subprocess
isolation (fork per function) is planned. Use `--ground-truth` for decoders.
## [0.25.0] — 2026-07-30
## [0.25.0] - 2026-07-30
Universal ground-truth verification: `gasm verify --ground-truth` assembles
any `.s` file with both gasm and `go tool asm`, then compares the machine
code byte-for-byte per function (relocation sites masked). No hand-written
reference needed — the Go toolchain IS the oracle.
reference needed; the Go toolchain IS the oracle.
### Added
- `verify`: `GroundTruth` — shells out to `go tool asm`, parses the GOOBJ
- `verify`: `GroundTruth`; shells out to `go tool asm`, parses the GOOBJ
output (minimal reader: block offsets, nonpkg symbol table, data index)
and returns per-function code bytes.
- `gasm verify --ground-truth`: compares gasm's output against the Go
@@ -504,11 +521,11 @@ reference needed — the Go toolchain IS the oracle.
linker fills) are masked before comparison.
- Verified: go-lz4 AVX2 2/2, go-flac AVX2 17/17 functions byte-identical.
## [0.24.0] — 2026-07-29
## [0.24.0] - 2026-07-29
The full analyze family and stereo PCM decode are now differentially tested.
15 of 17 go-flac AVX2 kernels have bit-for-bit differential coverage; the
two remaining (autocorrAVX2 — FMA reassociation, lpcResidualAVX2 — complex
two remaining (autocorrAVX2: FMA reassociation, lpcResidualAVX2: complex
multi-arg) are deferred.
### Added
@@ -518,7 +535,7 @@ multi-arg) are deferred.
- `verify`: `decodeStereo16AVX2` differential test (500 random interleaved
stereo PCM buffers, both channels compared sample-by-sample).
## [0.23.0] — 2026-07-28
## [0.23.0] - 2026-07-28
The analyze family and 24-bit PCM decode join the differential suite.
@@ -530,7 +547,7 @@ The analyze family and 24-bit PCM decode join the differential suite.
- `verify`: `decodeMono24AVX2` differential test (500 random 24-bit PCM
buffers, sign-extension compared sample-by-sample).
## [0.22.0] — 2026-07-27
## [0.22.0] - 2026-07-27
The remaining go-flac encoder kernels join the differential suite.
@@ -544,39 +561,39 @@ The remaining go-flac encoder kernels join the differential suite.
loop).
## [0.21.0] — 2026-07-26
## [0.21.0] - 2026-07-26
Differential testing extended to all four production kernels and the CLI
exposes the full dynamic-analysis toolkit.
### Added
- `verify`: go-flac AVX2 differential tests — `decodeMono16AVX2` (500
- `verify`: go-flac AVX2 differential tests; `decodeMono16AVX2` (500
random PCM buffers), `pack16AVX2` (500 random int32→int16 packings) and
all four decorrelation kernels (200 iterations each: left-side, side-right,
mid-side, interleave) compared bit-for-bit against the portable Go
references.
- `verify`: go-lz4 AVX-512 differential tests — `decodeBlockAVX512` (3 000
fuzzed LZ4 blocks + known answers) and `wideCopyAVX512` (0–1024 bytes)
- `verify`: go-lz4 AVX-512 differential tests; `decodeBlockAVX512` (3 000
fuzzed LZ4 blocks + known answers) and `wideCopyAVX512` (0-1024 bytes)
against the same portable oracle as the AVX2 suite.
- `gasm verify --abi`: runs each NOSPLIT function with sentinel registers
and a red-zone canary, reporting violations.
- `gasm verify --profile`: lists the static basic-block count per function.
## [0.20.0] — 2026-07-25
## [0.20.0] - 2026-07-25
Coverage profiling: the third pillar of Phase 3. Static basic-block
enumeration from the assembler's label map, combined with multi-input path
diversity measurement — how many observationally distinct execution paths a
diversity measurement; how many observationally distinct execution paths a
test corpus exercises.
### Added
- `verify`: `Kernel.Blocks` / `Kernel.BlockCount` — enumerate basic blocks
- `verify`: `Kernel.Blocks` / `Kernel.BlockCount`; enumerate basic blocks
from the assembler's local-label map (every jump target is a block
boundary; the function entry is always a block). `decodeBlockAVX2` has
27 blocks.
- `verify`: `Kernel.ProfilePaths` — run the function with a corpus of
- `verify`: `Kernel.ProfilePaths`; run the function with a corpus of
argument blocks and collect distinct output fingerprints (the result
words); reports path diversity as a lower bound on code coverage.
@@ -588,7 +605,7 @@ rt_sigaction handlers fragile in a Go process. The static + path-diversity
approach delivers the project's goal (proving the SIMD path and tail handling
execute) without fighting the runtime.
## [0.19.0] — 2026-07-24
## [0.19.0] - 2026-07-24
Runtime ABI checks: the second pillar of Phase 3. The JIT trampoline now
has an ABI-checking variant that sets sentinels in the callee-saved registers
@@ -598,41 +615,41 @@ detects any illegal write below the stack pointer.
### Added
- `verify`: `CallChecked` / `Kernel.CallFuncChecked` — ABI-checking JIT call
- `verify`: `CallChecked` / `Kernel.CallFuncChecked`; ABI-checking JIT call
with sentinel registers and red-zone canary; returns an `ABIReport`
(BPClobbered, R14Clobbered, RedZoneHit).
- `verify`: the raw `leaveJITCheckedRaw` trampoline — a TEXT symbol with no
- `verify`: the raw `leaveJITCheckedRaw` trampoline; a TEXT symbol with no
ABIInternal wrapper (address obtained via GLOBL/DATA), so the JIT
function's RET lands directly in the check code and sees the registers
exactly as the function left them.
- Tests: deliberate BP/R14 clobberers detected; both go-lz4 kernels
confirmed ABI-clean (BP preserved, R14 preserved, red zone intact).
## [0.18.0] — 2026-07-23
## [0.18.0] - 2026-07-23
Differential testing: the JIT-assembled go-lz4 `decodeBlockAVX2` kernel is
fuzzed against a portable Go reference — 5 000 valid LZ4 blocks compared
fuzzed against a portable Go reference; 5 000 valid LZ4 blocks compared
bit-for-bit, plus 2 000 hostile (random garbage) inputs with matching error
codes. This is the automated form of the project's bit-identical contract.
### Added
- `verify`: differential fuzz tests — a random LZ4 block generator produces
- `verify`: differential fuzz tests; a random LZ4 block generator produces
valid blocks (literals, overlapping matches, extension bytes) and the
JIT-assembled kernel's output is compared byte-for-byte against a portable
Go decoder; a hostile-input suite confirms error-code agreement on random
garbage (no crashes, same classification).
## [0.17.0] — 2026-07-22
## [0.17.0] - 2026-07-22
Phase 3 begins: dynamic analysis. A JIT execution substrate that assembles
Plan 9 amd64 kernels into executable memory and calls them directly — pure Go
Plan 9 amd64 kernels into executable memory and calls them directly; pure Go
(stdlib only, `syscall.Mmap` + an assembly trampoline), no cgo, no external
toolchain.
### Added
- `verify` package: JIT infrastructure — `Map` copies machine code into a
- `verify` package: JIT infrastructure; `Map` copies machine code into a
W^X memory mapping, `Call` invokes it through an ABI0 trampoline that
switches to a prepared stack and back. `Load`/`LoadSource`/`LoadAST`
parse, assemble and map a `.s` file in one step; `Kernel.CallFunc`
@@ -641,34 +658,34 @@ toolchain.
available functions; with `-smoke`, calls each NOSPLIT function with
zeroed arguments to confirm the trampoline works end-to-end.
- Integration tests: the go-lz4 `decodeBlockAVX2` and `wideCopyAVX2`
kernels (699 and 146 bytes) assemble, map and execute correctly —
known-answer LZ4 blocks decode bit-for-bit, wide copies of 0–1024 bytes
kernels (699 and 146 bytes) assemble, map and execute correctly;
known-answer LZ4 blocks decode bit-for-bit, wide copies of 0-1024 bytes
match, malformed input returns the correct error codes.
## [0.16.0] — 2026-07-21
## [0.16.0] - 2026-07-21
The scalar conversions between vector and general-purpose registers — the
The scalar conversions between vector and general-purpose registers; the
last of the amd64 EVEX instruction set.
### Added
- `asm`: the GPR-interchanging conversions, byte for byte against the Go
assembler (28 ground-truth cases including memory sources and extended
GPRs): vector to GPR — the signed and truncated VCVT{,T}S{D,S}2SI{,Q}
GPRs): vector to GPR; the signed and truncated VCVT{,T}S{D,S}2SI{,Q}
in both VEX and EVEX, and the unsigned VCVT{,T}S{D,S}2USI{L,Q}
(EVEX only); GPR to vector — VCVTSI2SD{L,Q}/VCVTSI2SS{L,Q} (VEX and
(EVEX only); GPR to vector; VCVTSI2SD{L,Q}/VCVTSI2SS{L,Q} (VEX and
EVEX) and VCVTUSI2SD{L,Q}/VCVTUSI2SS{L,Q} (EVEX only), whose preserved
vector source sits in vvvv (three Plan 9 operands).
## [0.15.0] — 2026-07-20
## [0.15.0] - 2026-07-20
The last of the EVEX conversions and narrowing/extending moves — the EVEX
The last of the EVEX conversions and narrowing/extending moves; the EVEX
instruction set is now complete save for the GPR-interchanging forms.
### Added
- `asm`: the unsigned and truncating conversions — VCVTPD2PS (and the X/Y
- `asm`: the unsigned and truncating conversions; VCVTPD2PS (and the X/Y
spellings, whose length the spelling fixes), VCVTPD2UDQ (X/Y),
VCVTTPD2UDQ (X/Y), VCVTTPD2UQQ, VCVTPS2UDQ, VCVTTPS2UDQ, VCVTPS2UQQ,
VCVTTPS2UQQ, VCVTTPD2QQ, VCVTTPS2QQ, VCVTUQQ2PD, VCVTUQQ2PS (X/Y) and
@@ -681,20 +698,20 @@ instruction set is now complete save for the GPR-interchanging forms.
D2M/Q2M), whose K register is a genuine operand rather than a mask and
which therefore take no masking suffixes.
## [0.14.0] — 2026-07-19
## [0.14.0] - 2026-07-19
The floating-point helper and conversion tail of the AVX-512 set, plus
gather and scatter with VSIB addressing — every encoding verified byte for
gather and scatter with VSIB addressing; every encoding verified byte for
byte against the Go assembler.
### Added
- `asm`: the floating-point helpers — reciprocals and reciprocal square
- `asm`: the floating-point helpers; reciprocals and reciprocal square
roots (VRCP14/VRSQRT14 PD/PS/SD/SS), exponents and mantissas (VGETEXP*,
VGETMANT*), scaling by powers of two (VSCALEF*), rounding (VRNDSCALE*),
reduction (VREDUCE*), immediate fixup (VFIXUPIMM*) and range selection
(VRANGE*), and floating-point class tests (VFPCLASSPD/PS X/Y/Z and
VFPCLASSSD/SS — a new immediate form whose reg field carries the opmask
VFPCLASSSD/SS; a new immediate form whose reg field carries the opmask
destination).
- `asm`: **gather and scatter with VSIB addressing.** The gathers take
both Go spellings: the VEX form with a vector mask register (OP mask,
@@ -703,14 +720,14 @@ byte against the Go assembler.
data register (a ZMM index with an YMM destination encodes L'L = 10, as
the Go assembler emits). The scatters (VSCATTER*/VPSCATTER*) are EVEX
only (OP src, K, vsib). All eight gather and eight scatter widths.
- `asm`: the remaining conversions — VCVTQQ2PS (the 512-bit source sets
- `asm`: the remaining conversions; VCVTQQ2PS (the 512-bit source sets
the length), VCVTPD2QQ/UQQ, VCVTPS2QQ, VCVTUDQ2PD/PS, the half-precision
VCVTPH2PS and VCVTPS2PH (the extract layout with an immediate).
## [0.13.0] — 2026-07-18
## [0.13.0] - 2026-07-18
The wider AVX-512 set: ternary logic, permutes, compares, expand/compress,
the opmask instructions and the EVEX rounding/SAE/broadcast suffixes — every
the opmask instructions and the EVEX rounding/SAE/broadcast suffixes; every
encoding verified byte for byte against the Go assembler.
### Added
@@ -718,7 +735,7 @@ encoding verified byte for byte against the Go assembler.
- `asm`: the wider EVEX/AVX-512 set, across roughly sixty new ground-truth
cases: ternary logic (VPTERNLOGD/Q), the lane shuffles/inserts/extracts
(VSHUF{F,I}{32,64}X{2,4}, the VINSERT*/VEXTRACT* {F,I}{32,64}X{2,4,8}
family, VPALIGNR), compares with an opmask destination (VCMPPD/PS/SD/SS —
family, VPALIGNR), compares with an opmask destination (VCMPPD/PS/SD/SS;
a new NDS-plus-immediate form with the K register in the reg field), the
permutes (VPERMB/W, VPERMI2/T2 D/Q/PD), the wider integer families
(VPMADDWD/UBSW, VPMULHUW, VPACKSSWB/USWB/SSDW/USDW, VPABS B/W/D/Q, the
@@ -732,15 +749,15 @@ encoding verified byte for byte against the Go assembler.
(VMOVSLDUP/VMOVSHDUP), the conversions (VCVTPS2DQ, VCVTTPS2DQ) and the
remaining extending and narrowing moves (VPMOVSXBW, VPMOVZXBW, VPMOVWB,
VPMOVQB).
- `asm`: the EVEX mnemonic suffixes the Go assembler accepts — the rounding
- `asm`: the EVEX mnemonic suffixes the Go assembler accepts; the rounding
modes `.RN_SAE`, `.RD_SAE`, `.RU_SAE`, `.RZ_SAE` (the EVEX b bit with the
rounding control in L'L), suppress-all-exceptions `.SAE`, and memory
broadcast `.BCST` (the b bit, the vector length preserved, disp8×N scaled
by the element size) — each combinable with the `.Z` zeroing suffix,
by the element size); each combinable with the `.Z` zeroing suffix,
validated against the Go assembler's bytes, and rejected on instructions
that do not support them.
## [0.12.0] — 2026-07-17
## [0.12.0] - 2026-07-17
GOOBJ emission: gasm-assembled functions drop into a `go build` without the
Go assembler.
@@ -748,7 +765,7 @@ Go assembler.
### Added
- `asm`: **GOOBJ object output.** `gasm asm --format goobj -p <pkgpath>`
writes the Go toolchain's own object format — the one `cmd/link` consumes
writes the Go toolchain's own object format; the one `cmd/link` consumes
directly: the functions as non-package symbols qualified with the package
path (exactly as `cmd/asm` records assembly symbols), the `GLOBL` data,
one serialized `FuncInfo` per function (argument/frame sizes, the asm
@@ -757,8 +774,8 @@ Go assembler.
real stack deltas: the assembler now tracks every stack-adjustment
boundary through the prologue (`PUSHQ BP`, `SUBQ $frame, SP`) and each
`RET`'s epilogue, so frame-pointer functions unwind correctly. The
object preamble — the version-and-experiment header the linker compares
verbatim — is captured from the installed `go tool asm`, so the output is
object preamble; the version-and-experiment header the linker compares
verbatim; is captured from the installed `go tool asm`, so the output is
always consistent with the toolchain that links it.
- `asm`: relocations against file-local `GLOBL` symbols become `R_PCREL`
entries in the GOOBJ output, with the instruction's displacement field
@@ -771,7 +788,7 @@ Go assembler.
pattern, instead of being rejected as non-integer.
## [0.11.0] — 2026-07-16
## [0.11.0] - 2026-07-16
Linkable object output: external symbols and relocatable ELF / Mach-O
objects.
@@ -790,7 +807,7 @@ objects.
external symbol; the Mach-O output is verified structurally with
`debug/macho`.
- `asm`: **external symbol references.** A reference to a symbol no
`GLOBL` in the file defines no longer aborts assembly — it is recorded
`GLOBL` in the file defines no longer aborts assembly; it is recorded
as an external relocation (`Image.Externals`, `FuncLayout.Relocs`) and
becomes an undefined global symbol in the object output. The raw image
s (`--format raw`, the default) still reports them: only an object
@@ -802,7 +819,7 @@ objects.
writes; without `--format` the behaviour is unchanged (the concatenated
image).
## [0.10.0] — 2026-07-15
## [0.10.0] - 2026-07-15
The EVEX floating-point and conversion set: the packed-double arithmetic,
the scalar SD/SS forms, VMOVDDUP and the width-changing conversions, each
@@ -810,23 +827,23 @@ verified byte for byte against the Go assembler.
### Added
- `asm`: the rest of the common EVEX/VEX floating-point set — packed double
- `asm`: the rest of the common EVEX/VEX floating-point set; packed double
arithmetic (VSUBPD, VDIVPD, VMINPD, VMAXPD, VUNPCKLPD and the EVEX form of
VUNPCKHPD), the scalar double and single operations (VSUBSD, VDIVSD,
VMINSD, VMAXSD and the full VADDSS/VSUBSS/VMULSS/VDIVSS/VMINSS/VMAXSS
family in both VEX and EVEX — the EVEX scalar forms exist for masked and
family in both VEX and EVEX; the EVEX scalar forms exist for masked and
zeroing use), and VMOVDDUP (lane duplication, VEX and EVEX).
- `asm`: the width-changing conversions — VCVTDQ2PS and VCVTPS2PD (VEX and
- `asm`: the width-changing conversions; VCVTDQ2PS and VCVTPS2PD (VEX and
EVEX; the destination sets the length for PS→PD), the EVEX form of
VCVTDQ2PD, and the packed-double → dword family: VCVTPD2DQ/VCVTTPD2DQ
(EVEX-512 only, a ZMM source and an XMM destination) and their X/Y
spellings (VCVTPD2DQX/Y, VCVTTPD2DQX/Y), whose length follows the wider
source — a new operand form, since the destination is always XMM while
source; a new operand form, since the destination is always XMM while
VEX.L / EVEX.L'L ride with the source (fixed by the spelling even for a
memory source).
- `asm`: masking and zeroing on every new form — the scalar SD/SS
- `asm`: masking and zeroing on every new form; the scalar SD/SS
arithmetic, the unpacks, VMOVDDUP and the conversions all accept the
explicit K1–K7 operand and the `.Z` suffix the way Go writes them.
explicit K1-K7 operand and the `.Z` suffix the way Go writes them.
### Changed
@@ -837,35 +854,35 @@ verified byte for byte against the Go assembler.
shares the convention).
## [0.9.0] — 2026-07-14
## [0.9.0] - 2026-07-14
AVX-512 masking and a wider EVEX integer set.
### Added
- `asm`: **EVEX masking** the way Go writes it — an explicit `K1`–`K7`
- `asm`: **EVEX masking** the way Go writes it; an explicit `K1`-`K7`
operand placed among the operands (merging mask), and a `.Z` mnemonic
suffix for zeroing (`VPADDD.Z Z1, Z2, K2, Z3`). Supported across the NDS,
reg/rm, immediate-shift, align, extract, convert and move forms, including
masked comparisons with a K destination (`VPCMPEQD Z0, Z3, K2, K1`). K0 is
rejected as an explicit mask, and `.Z` without a mask is an error, matching
the Go assembler.
- `asm`: the common AVX-512 F/BW integer set — VPADDB/W, VPSUBB/W, VPANDD/Q,
- `asm`: the common AVX-512 F/BW integer set; VPADDB/W, VPSUBB/W, VPANDD/Q,
VPANDND/Q, VPMULLW, VPAVGB/W, the signed/unsigned min/max family for
B/W/D/Q elements, the variable shifts VPSLLVD/Q, VPSRLVD/Q, VPSRAVD/Q, the
EVEX forms of VPSHUFD/VPSHUFB, and the VMOVDQU8/VMOVDQU16 move aliases.
Register indices 16–31 encode correctly (the mod=11 quirk carries rm[4]
Register indices 16-31 encode correctly (the mod=11 quirk carries rm[4]
in X̄). All verified byte for byte against the Go assembler.
- `lint`: masked EVEX forms (`.Z` suffix, K operands) are recognised by
`unknown-instruction` and exempted from `operand-count`.
### Fixed
- `asm`: EVEX register–register operands with indices 16–31 encoded rm[4]
- `asm`: EVEX register-register operands with indices 16-31 encoded rm[4]
into B̄ instead of X̄ (the EVEX mod=11 extension quirk), producing wrong
prefix bytes for X16+/Y16+ r/m operands.
## [0.8.0] — 2026-07-13
## [0.8.0] - 2026-07-13
Standard CLI ergonomics.
@@ -881,20 +898,20 @@ Standard CLI ergonomics.
- The version is primarily available as the standard `gasm --version` / `-V`
flag; the `gasm version` spelling remains as an alias.
## [0.7.0] — 2026-07-12
## [0.7.0] - 2026-07-12
The formatter behaves like `go fmt` and canonicalises block separation.
### Added
- `gasm fmt` now works like `go fmt`: with no arguments — or with a directory
argument — it reformats every `.s` file below it in place and lists the
- `gasm fmt` now works like `go fmt`: with no arguments; or with a directory
argument; it reformats every `.s` file below it in place and lists the
changed files, skipping `.` and `_` directories (`.git`, `_refs`, …).
Explicit file arguments keep the `-w` / standard-output behaviour.
### Changed
- `s`: canonical blank-line layout — a new block (a label, `TEXT` or
- `s`: canonical blank-line layout; a new block (a label, `TEXT` or
`GLOBL`) is preceded by exactly one blank line, neither more nor less.
Comments leading a block stay with it (the blank line goes before them),
stacked labels share their block, the function's first label keeps hugging
@@ -903,7 +920,7 @@ The formatter behaves like `go fmt` and canonicalises block separation.
kernels were reformatted with this release and remain byte-identical when
assembled.
## [0.6.0] — 2026-07-11
## [0.6.0] - 2026-07-11
Calibrated to the Go ABI: `register-clobber` stops reporting legal code, and
the encoder learns the legacy SSE moves.
@@ -912,13 +929,13 @@ the encoder learns the legacy SSE moves.
- `lint`: **`register-clobber` is now calibrated to the Go ABI**
(`cmd/compile/abi-internal.md`), not the platform ABI. Go's stack-based
ABI0 has no System V style callee-saved registers — amd64 `BX`, `R12`–`R15`
ABI0 has no System V style callee-saved registers; amd64 `BX`, `R12`-`R15`
and the arm64/riscv64/loong64 scratch sets are caller-saved or permanent
scratch, and hand-written kernels may clobber them freely. The rule now
audits only the registers Go fixes across calls: the frame pointer and the
goroutine pointer (amd64 `BP`/`R14`, arm64 `R18`/`R28`/`R29`, riscv64
`X27`, loong64 `R22`), and the goroutine pointer is reported only when the
function can reach the runtime (is not `NOSPLIT` or makes a call) — the
function can reach the runtime (is not `NOSPLIT` or makes a call); the
ABI0 transition restores it on those paths, and NOSPLIT call-free leaves
may use it, exactly as the runtime's own assembly does. Both go-flac
kernels now lint with zero diagnostics.
@@ -935,21 +952,21 @@ the encoder learns the legacy SSE moves.
### Added
- `asm`: the legacy (non-VEX) SSE moves — `MOVOU`/`MOVO` (the Plan 9 names
- `asm`: the legacy (non-VEX) SSE moves; `MOVOU`/`MOVO` (the Plan 9 names
for MOVDQU/MOVDQA), `MOVUPS`/`MOVAPS`/`MOVUPD`/`MOVAPD` and the scalar
`MOVSD`/`MOVSS` — and `VMOVDQU64` in the EVEX set. All verified byte for
`MOVSD`/`MOVSS`; and `VMOVDQU64` in the EVEX set. All verified byte for
byte against the Go assembler.
## [0.5.0] — 2026-07-10
## [0.5.0] - 2026-07-10
EVEX / AVX-512: the go-flac AVX-512 kernel now assembles, byte-identically to
the Go toolchain, completing the production-kernel coverage.
### Added
- `asm`: **EVEX (AVX-512) encoding** — the four-byte EVEX prefix with the
5-bit register fields (Z0–Z31, X/Y 16–31, with the reg-r/m X̄ quirk and
V'̄ shared between vvvv and the SIB index), opmask registers (K0–K7) as
- `asm`: **EVEX (AVX-512) encoding**; the four-byte EVEX prefix with the
5-bit register fields (Z0-Z31, X/Y 16-31, with the reg-r/m X̄ quirk and
V'̄ shared between vvvv and the SIB index), opmask registers (K0-K7) as
operands and as mask destinations, and the compressed disp8×N displacement
(the multiplier follows the memory operand's size, as the Go assembler's
opcode tables prescribe). Covers every AVX-512 instruction the go-flac
@@ -959,20 +976,20 @@ the Go toolchain, completing the production-kernel coverage.
extracts VEXTRACTI64X4/VEXTRACTF64X4, VFMADD231PD, VADDPD, VMULPD, the
broadcasts VPBROADCASTD/Q (GPR and memory sources take different opcodes)
and the mask moves KMOVW/KTESTW. Masking/zeroing suffixes are out of scope
— the kernels use neither.
; the kernels use neither.
- `asm`: `AssembleFile` now accepts file-defined global (`non-<>`) symbols
too; a reference is external only when no `GLOBL` in the file defines it.
### Fixed
- `asm`: registers X16–Y31 force the EVEX encoding of dual-form mnemonics;
- `asm`: registers X16-Y31 force the EVEX encoding of dual-form mnemonics;
previously a `VPBROADCASTD AX, Y30` fell into the VEX encoder, which cannot
represent indices above 15 and silently truncated them.
- `asm`: the VEX encoder now rejects vector register indices 16–31 instead of
- `asm`: the VEX encoder now rejects vector register indices 16-31 instead of
encoding a truncated (wrong) register.
## [0.4.0] — 2026-07-09
## [0.4.0] - 2026-07-09
The standalone assembler reaches the whole go-flac AVX2 kernel: static
symbols assemble, and all 17 kernel functions now match the Go toolchain's
@@ -980,10 +997,10 @@ machine code byte for byte.
### Added
- `asm`: **file-level assembly** — `AssembleFile` turns a parsed file into an
- `asm`: **file-level assembly**; `AssembleFile` turns a parsed file into an
`Image`: the function bodies in source order followed by a data section
built from the file's `GLOBL`/`DATA` directives (each symbol 16-aligned).
- `asm`: **static-symbol (`SB`) operands** — `mask<>(SB)` references encode as
- `asm`: **static-symbol (`SB`) operands**; `mask<>(SB)` references encode as
RIP-relative loads with a patched disp32, resolved against the image layout
so the output is self-consistent and position-independent. External
(non-file-local) symbols are rejected with a clear error: they need
@@ -992,7 +1009,7 @@ machine code byte for byte.
and writes the whole image (code + data) with `-o`.
## [0.3.0] — 2026-07-08
## [0.3.0] - 2026-07-08
The assembler reaches byte-identical parity with the Go toolchain on the
production go-flac AVX2 kernels: every one of the 15 kernel functions that
@@ -1002,18 +1019,18 @@ support).
### Added
- `asm`: the scalar instruction families the kernels use — `CMOVcc` and
- `asm`: the scalar instruction families the kernels use; `CMOVcc` and
`SETcc` (conditions spelled exactly like the jumps), `LZCNT`/`TZCNT`
(legacy `F3 0F BD/BC`), the sign/zero-extending moves (`MOVBLZX`, `MOVBQZX`,
`MOVWLZX`, `MOVWQZX`, `MOVWLSX`, `MOVLQSX`), `CVTSL2SD`/`CVTSQ2SD` (the
legacy SSE encoding, as the Go assembler emits it), the traditional
three-operand `IMUL3{W,L,Q}`, and the variable-count vector shifts
(`VPSRLQ X0, Y8, Y8` — the count in an XMM register or memory takes the
(`VPSRLQ X0, Y8, Y8`; the count in an XMM register or memory takes the
ordinary NDS form).
- `asm`: **jump relaxation** — jumps start in the short (rel8) form and
- `asm`: **jump relaxation**; jumps start in the short (rel8) form and
expand to rel32 when the settled displacement does not fit, iterating the
layout to a fixed point (CALL is always rel32).
- `asm`: **jump-to-jump folding** — a conditional jump to a label whose only
- `asm`: **jump-to-jump folding**; a conditional jump to a label whose only
instruction is an unconditional jump is redirected to the ultimate
target, replicating the Go toolchain's linker, which chases such chains
before it encodes branches.
@@ -1025,12 +1042,12 @@ support).
- `asm`: `CMP` with a register or memory operand computed **second − first**
instead of first − second, silently inverting every condition that followed
(`CMPQ SI, R10; JGE` tested R10 ≥ SI). The encoding now always records
first − second — `CMP r/m, r` with the first operand in r/m, `CMP r, r/m`
with the first operand in reg — and is byte-identical to the Go assembler.
first − second; `CMP r/m, r` with the first operand in r/m, `CMP r, r/m`
with the first operand in reg; and is byte-identical to the Go assembler.
- `asm`: register-to-register `MOV` now uses the `r/m ← r` opcode (reg =
source), the Go assembler's choice; the output is byte-identical.
## [0.2.0] — 2026-07-07
## [0.2.0] - 2026-07-07
The Phase 2 assembler grows the SIMD set: shuffles, extract/insert, permute
and the moves, on top of the Phase 1 VEX forms.
@@ -1043,10 +1060,10 @@ and the moves, on top of the Phase 1 VEX forms.
- the immediate shuffle (`VPSHUFD`, `VPERMQ`),
- the three-operand-plus-immediate form (`VSHUFPD`, `VPERM2I128`,
`VINSERTI128`),
- the lane extract (`VEXTRACTI128`, `VEXTRACTF128` — the YMM source occupies
- the lane extract (`VEXTRACTI128`, `VEXTRACTF128`; the YMM source occupies
the ModRM.reg field, the XMM/memory destination the r/m field),
- the direction-sensitive moves (`VMOVDQU`, `VMOVUPD`, `VMOVD`, `VMOVQ`,
`VMOVSD` — each direction picks its own opcode and VEX.W; a vector→vector
`VMOVSD`; each direction picks its own opcode and VEX.W; a vector→vector
move uses the store-form layout, matching the Go assembler),
- the no-operand `VZEROUPPER`, and `VPERMD` in the NDS form,
- the floating-point and FMA set (`VADDPD`, `VMULPD`, `VXORPD`,
@@ -1055,21 +1072,21 @@ and the moves, on top of the Phase 1 VEX forms.
the encoder now covers every integer, shuffle and FP instruction the
go-flac AVX2 kernels use.
- `asm`: `CMP` accepts the immediate in the second operand position
(`CMPL CX, $31`) — the spelling the Go assembler accepts — encoding it
(`CMPL CX, $31`), the spelling the Go assembler accepts, encoding it
identically to the immediate-first form.
### Fixed
- `asm`: an unused VEX.vvvv field is now stored as `1111` (v̄vvv = 1111), as
the hardware requires — the previous value (`0000`) made the two-operand
the hardware requires; the previous value (`0000`) made the two-operand
reg/rm forms (VPMOVSXWD, VPBROADCASTD, VMOVMSKPS, …) raise #UD on real CPUs
and differ from the Go assembler's bytes. The round-trip decoder ignores
the field on these instructions, which is why the byte-for-byte Go
comparison (added this release) is now part of the test suite.
## [0.1.0] — 2026-07-06
## [0.1.0] - 2026-07-06
Initial release — the Phase 1 foundation.
Initial release; the Phase 1 foundation.
### Added
@@ -1082,11 +1099,11 @@ Initial release — the Phase 1 foundation.
- `arch`: register files and **complete** instruction tables for amd64,
arm64, riscv64 and loong64, with the middle-dot symbol separator and static
(`<>`) symbols. Instruction names are generated from the Go toolchain's own
assembler source (`just gen`) — the `anames` opcode lists plus the common
assembler source (`just gen`); the `anames` opcode lists plus the common
opcodes and the per-architecture front-end aliases (arm64 `B`/`BL`, the
`.P`/`.W` addressing suffixes, loong64 `JAL`, the x86 conditional-jump
spellings) — so every mnemonic the real assembler accepts is recognised.
- `lint`: conservative rules — `unknown-instruction`, `operand-count`,
spellings); so every mnemonic the real assembler accepts is recognised.
- `lint`: conservative rules; `unknown-instruction`, `operand-count`,
`undefined-label`, `duplicate-label`, `missing-ret`,
`missing-textflag-include`, `abi-argsize` and `unreachable-code`. Macro
invocations are recognised (in-file `#define` names and underscore
@@ -1108,9 +1125,9 @@ Initial release — the Phase 1 foundation.
- `lsp`: a Language Server Protocol server over stdio providing completion,
hover documentation, document symbols, publish-diagnostics and semantic-token
highlighting.
- `asm`: a standalone amd64 (x86-64) assembler — an instruction encoder (REX/
- `asm`: a standalone amd64 (x86-64) assembler; an instruction encoder (REX/
ModR-M/SIB/displacement/immediate plus the scalar instruction set, and VEX/
AVX2 SIMD across three operand forms — NDS, reg/rm and immediate-shift —
AVX2 SIMD across the three operand forms NDS, reg/rm and immediate-shift,
covering the bulk of the integer SIMD set) validated by round-trip decoding
against `golang.org/x/arch`, and an assembler that drives the parser's AST
into the encoder with local-label resolution and `FP`/`SP` frame mapping
+40
View File
@@ -0,0 +1,40 @@
# Security policy
## Supported versions
Security fixes go to the newest release and to the `development` branch. Older
releases do not receive them.
| Version | Supported |
|---|---|
| 0.33.0 | yes |
| older releases | no |
## Reporting a vulnerability
**Do not open a public issue for a security problem.** A public report tells everyone
about the flaw before there is a fix. Report it privately to
**opensource@petrbalvin.org**.
Include:
- the version or commit you tested, and the platform
- what the problem is, and what an attacker gains from it
- the smallest reproducer you have, ideally a test or a single command
- a suggested fix, if you have one
## What to expect
- A human reads the report, and you get an acknowledgement.
- You are kept informed while the fix is being made, and told when it ships.
- The fix is released before the details are published, and the timing is agreed with
you.
- The reporter is credited in the release notes unless they ask otherwise.
## Out of scope
- Findings that require the attacker to already run code as the user, or to have local
access.
- Missing hardening with no demonstrated impact.
- Flaws in a third-party dependency: report them to that project, and to this one only
when this project's use of it makes them reachable.