From 3de043c49468a05ffd4156b2daec91d438dcfea6 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Petr=20Balv=C3=ADn?= Date: Wed, 16 Sep 2026 23:12:31 +0200 Subject: [PATCH] docs: add SECURITY.md and record the round in the CHANGELOG Assisted-by: GLM 5.3 Flash --- CHANGELOG.md | 297 +++++++++++++++++++++++++++------------------------ SECURITY.md | 40 +++++++ 2 files changed, 197 insertions(+), 140 deletions(-) create mode 100644 SECURITY.md diff --git a/CHANGELOG.md b/CHANGELOG.md index 06ee596..bcd1c04 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -7,11 +7,28 @@ and this project adheres to [Conventional Commits](https://www.conventionalcommi ## [development] -### Added +### Changed -- +- **Canonical just recipes.** `just gates` is the definition of done + (build, fmt-check, vet, test, race). `install` now builds and copies + the binary into `~/.local/bin` (`BINDIR` overrides) instead of + downloading module dependencies, and `install-bin` is gone. The test + gate sweeps the whole suite and computes the coverage floor over the + product packages with `-coverpkg`, `disasm` now included, so the + number is identical locally and in CI. `fuzz` requires its target + package. +- **The reported version comes from the build.** `gasm --version` + prints the version the toolchain recorded: the tag on a tagged + checkout, a pseudo-version naming the commit below one, `+dirty` on a + dirty tree and `(devel)` outside version control. Nothing is + injected with `-ldflags -X` any more. +- **CI realigned with the gate set.** The push pipeline runs the gates + minus race in one job, with a cached Go setup and the module as the + version source; the race detector moved to a hand-dispatched workflow + and into the release gates; the release builds without injection and + its smoke test requires the recorded tag and rejects `+dirty`. -## [0.33.0] — 2026-09-14 +## [0.33.0] - 2026-09-14 ### Added @@ -67,7 +84,7 @@ and this project adheres to [Conventional Commits](https://www.conventionalcommi the toolchain picks, and the morestack block saves the link register with the toolchain's `OR` form on loong64. -## [0.32.0] — 2026-08-31 +## [0.32.0] - 2026-08-31 ### Added @@ -218,7 +235,7 @@ and this project adheres to [Conventional Commits](https://www.conventionalcommi outputs and operand strictness with `go tool asm`. - **asm help text.** Updated to list arm64 as a supported architecture. -## [0.31.1] — 2026-08-20 +## [0.31.1] - 2026-08-20 ### Fixed @@ -226,9 +243,9 @@ and this project adheres to [Conventional Commits](https://www.conventionalcommi because the version variables in `justfile` and `cmd/gasm/main.go` were not bumped during the release commit. -## [0.31.0] — 2026-08-20 +## [0.31.0] - 2026-08-20 -The arm64 encoder (Phase 5 — complete) ships with ELF64 and GOOBJ emission, +The arm64 encoder (Phase 5; complete) ships with ELF64 and GOOBJ emission, verified byte-for-byte against `GOARCH=arm64 go tool asm` and linked into a real `go build`. The encoder covers the full integer instruction set, FP arithmetic, conditional select, CRC32, and the MOV pseudo-instruction with @@ -236,7 +253,7 @@ bitmask immediate encoding. The project now requires Go 1.27. ### Added -- **arm64 encoder (Phase 5 — complete).** `gasm asm` can now assemble `_arm64.s` +- **arm64 encoder (Phase 5; complete).** `gasm asm` can now assemble `_arm64.s` files: the AArch64 integer instruction set with the MOV pseudo-instruction and its immediate-constant expansions (MOVZ/MOVN/MOVK for wide immediates, ORR with logical bitmask encoding for values like `$1`), data-processing (shifted @@ -245,7 +262,7 @@ bitmask immediate encoding. The project now requires Go 1.27. SB/global symbol references (ADRP+ADD pairs with `R_ADDRARM64` relocations), jump chain folding, and ELF64 emission (`gasm asm --format elf`). Ground-truth verification against `GOARCH=arm64 go tool asm` matches byte-for-byte. Phase 5 - (the other architectures — RISC-V, LoongArch, arm64) is now complete. + (the other architectures; RISC-V, LoongArch, arm64) is now complete. ### Changed @@ -253,7 +270,7 @@ bitmask immediate encoding. The project now requires Go 1.27. The `R_DWTXTADDR_U4` relocation type is detected at runtime for backward compatibility. -## [0.30.0] — 2026-08-13 +## [0.30.0] - 2026-08-13 The LoongArch encoder (Phase 5) ships with ELF64 and GOOBJ emission, verified byte-for-byte against `GOARCH=loong64 go tool asm` and linked into a real @@ -278,7 +295,7 @@ tracks four hardware watchpoint slots, and the toolkit is Linux-only. - **GOOBJ DWARF symbols.** The GOOBJ emitters now write the per-function DWARF symbols the linker requires (the subprogram DIE and the `.debug_line` program, byte-identical to `cmd/asm`'s), and the pc-value table deltas are - in the architecture's MinLC units as the runtime expects — the amd64 link + in the architecture's MinLC units as the runtime expects; the amd64 link test now genuinely substitutes the gasm object, and the amd64/loong64 end-to-end GOOBJ link tests pass. - **RISC-V GOOBJ emission via the shared emitter.** RISC-V GOOBJ output is @@ -347,7 +364,7 @@ tracks four hardware watchpoint slots, and the toolkit is Linux-only. `GOARCH=riscv64 go tool asm`. - **Debugger watchpoint slots.** `gasm debug`'s `watch` command always used hardware watchpoint slot 0, so a second `watch` call silently overwrote - the first. Watchpoint slots are now tracked in the `Session` (DR0–DR3); + the first. Watchpoint slots are now tracked in the `Session` (DR0-DR3); `watch` picks the first free slot and reports an error if all four are in use, and `unwatch ` clears one (no argument clears all). @@ -359,7 +376,7 @@ tracks four hardware watchpoint slots, and the toolkit is Linux-only. disassembly at PC, memory-write, watchpoints, and source-line mapping are all shipped. -## [0.29.0] — 2026-08-07 +## [0.29.0] - 2026-08-07 RISC-V GOOBJ emission, YMM vector register display, named buffer allocation in the debugger, two new CLI commands (`diff`, `profile`), go-to-definition in @@ -369,30 +386,30 @@ new CLI commands. A signature-parser fix corrects grouped Go parameters. ### Added -- **RISC-V GOOBJ emission** — `gasm asm --format goobj` for RISC-V produces +- **RISC-V GOOBJ emission**; `gasm asm --format goobj` for RISC-V produces linkable Go objects with funcdata, pc-value tables, and RISC-V relocation types (same format as amd64 GOOBJ, with the RISC-V architecture marker). -- **`gasm diff`** — compare the machine code of two assembly files byte-for-byte; +- **`gasm diff`**; compare the machine code of two assembly files byte-for-byte; shows which functions differ and the first few differing bytes. -- **`gasm profile`** — show the basic-block structure of each function: labels, +- **`gasm profile`**; show the basic-block structure of each function: labels, offsets, frame size, and NOSPLIT flag. -- **LSP go-to-definition** — `textDocument/definition` navigates from a label +- **LSP go-to-definition**; `textDocument/definition` navigates from a label reference to its definition. -- **did-you-mean** — when the RISC-V assembler encounters an undefined label, it +- **did-you-mean**; when the RISC-V assembler encounters an undefined label, it suggests the closest existing label using Levenshtein distance. -- **YMM vector register display** — `regs` in the debugger now shows YMM +- **YMM vector register display**; `regs` in the debugger now shows YMM registers via `PTRACE_GETFPREGS` (falls back to XMM when XSAVE is unavailable). -- **Named buffer allocation** — `gasm debug --buf name:size:pattern` allocates +- **Named buffer allocation**; `gasm debug --buf name:size:pattern` allocates buffers in the debuggee filled with `zero`, `ones`, `seq`, or a hex pattern; buffer pointers are placed into the argument block at the matching positions. -- **Crash input storage** — `FuzzResult.CrashInput` stores the input that caused +- **Crash input storage**; `FuzzResult.CrashInput` stores the input that caused a crash or mismatch for reproducibility. -- **ABI + fuzz combined** — `gasm verify --fuzz` now runs ABI checks (sentinel +- **ABI + fuzz combined**; `gasm verify --fuzz` now runs ABI checks (sentinel registers, canary, stack bounds) alongside differential fuzz testing. -- **`gasm diff --map`** — compare functions whose names differ between files +- **`gasm diff --map`**; compare functions whose names differ between files (e.g. `--map wideCopyAVX2=wideCopyAVX512` pairs two variants regardless of suffix). Unmapped functions fall back to the original name match. -- **`gasm verify --call`** — invoke a single function with user-supplied buffers +- **`gasm verify --call`**; invoke a single function with user-supplied buffers (`--buf name:size:pattern`) instead of the smoke/abi/fuzz sweeps. Patterns: `zero`, `ones`, `seq`, or a hex blob. Useful for partial functions (e.g. decoders) that crash on random input but should succeed on valid data. @@ -402,23 +419,23 @@ new CLI commands. A signature-parser fix corrects grouped Go parameters. ### Fixed -- **Signature parser** — grouped Go parameters like `dst, src []byte` are now +- **Signature parser**; grouped Go parameters like `dst, src []byte` are now parsed correctly (both get type `[]byte`). Previously the first name was treated as its own type (`dst` with size 8), causing wrong ABI0 arg-block layout in both `verify --call` and the fuzzer. -- **Flaky JIT tests** — `runtime.KeepAlive` guards and package-level buffers +- **Flaky JIT tests**; `runtime.KeepAlive` guards and package-level buffers prevent GC from collecting heap objects whose addresses were passed to JIT code via `unsafe.Pointer`; all verify tests pass 100/100 under `-race`. ### Changed -- **Removed external kernel test dependencies** — the verify test suite no +- **Removed external kernel test dependencies**; the verify test suite no longer references production kernels from the separate go-libraries project. The remaining test suite uses only `testdata/verify/*.s` kernels, which are part of this repository. Coverage is identical locally and in CI (80.3 %). -## [0.28.0] — 2026-08-03 +## [0.28.0] - 2026-08-03 RISC-V encoder: full RV64IMAFDC instruction set with RVC compression, MOV pseudo-instruction, SB/global symbol references, ELF64 object emission, and @@ -426,20 +443,20 @@ ground-truth verification against `GOARCH=riscv64 go tool asm`. ### Added -- **RISC-V encoder** — RV64I, RV64M, RV64A, RV64F/D, FMA, CSR, JALR. -- **MOV pseudo-instruction** — load, store, reg-to-reg, immediate, frame mapping. -- **RVC compression** — 22 compressed instruction types (C.LDSP, C.SDSP, C.FLDSP, +- **RISC-V encoder**; RV64I, RV64M, RV64A, RV64F/D, FMA, CSR, JALR. +- **MOV pseudo-instruction**; load, store, reg-to-reg, immediate, frame mapping. +- **RVC compression**; 22 compressed instruction types (C.LDSP, C.SDSP, C.FLDSP, C.FSDSP, C.ADDI, C.LI, C.LUI, C.ADDIW, C.MV, C.ADD, C.SUB, C.XOR, C.OR, C.AND, C.SLLI, C.SRLI, C.SRAI, C.ANDI, C.BEQZ, C.BNEZ, C.J, C.JR). -- **SB/global symbols** — `MOV $sym(SB)`, `MOV sym(SB)`, `MOV rd, sym(SB)` +- **SB/global symbols**; `MOV $sym(SB)`, `MOV sym(SB)`, `MOV rd, sym(SB)` encoded as AUIPC pairs with R_RISCV_PCREL_HI20/LO12 relocations. -- **GLOBL/DATA** — data section layout in `AssembleFileRISCV`. -- **ELF64 emission** — `gasm asm --format elf` produces EM_RISCV objects +- **GLOBL/DATA**; data section layout in `AssembleFileRISCV`. +- **ELF64 emission**; `gasm asm --format elf` produces EM_RISCV objects (.text, .data, .symtab, .rela.text). -- **`gasm verify --ground-truth`** — byte-exact comparison against +- **`gasm verify --ground-truth`**; byte-exact comparison against `GOARCH=riscv64 go tool asm`. -- **`gasm verify --profile`** — function layout listing for RISC-V. -- **CALL** — AUIPC + JALR pair encoding. +- **`gasm verify --profile`**; function layout listing for RISC-V. +- **CALL**; AUIPC + JALR pair encoding. ### Fixed @@ -449,7 +466,7 @@ ground-truth verification against `GOARCH=riscv64 go tool asm`. (bit-interleaved format). -## [0.27.0] — 2026-08-01 +## [0.27.0] - 2026-08-01 Subprocess isolation for `--fuzz`: each function is fuzzed in its own child process, so a partial function (decoder) that faults on random garbage is @@ -460,7 +477,7 @@ the parent. CRASH is informational (exit 0); only MISMATCH is an error. - `gasm verify --fuzz` no longer crashes the process on partial functions. -## [0.26.0] — 2026-07-31 +## [0.26.0] - 2026-07-31 Universal differential fuzzing: `gasm verify --fuzz` needs no hand-written reference. It parses the `// func` signature from the assembly source, @@ -471,7 +488,7 @@ area bit-for-bit. ### Added -- `verify`: `FuzzFunc` / `ExtractSignatures` / `parseFuncSig` — universal +- `verify`: `FuzzFunc` / `ExtractSignatures` / `parseFuncSig`; universal differential fuzz driven by the conventional `// func` comment. Each version gets its own buffer set (deep copy) so functions that write to their arguments (histogram increments) don't corrupt the other's input. @@ -486,16 +503,16 @@ area bit-for-bit. over-copy paths read past the buffer on random garbage input. Subprocess isolation (fork per function) is planned. Use `--ground-truth` for decoders. -## [0.25.0] — 2026-07-30 +## [0.25.0] - 2026-07-30 Universal ground-truth verification: `gasm verify --ground-truth` assembles any `.s` file with both gasm and `go tool asm`, then compares the machine code byte-for-byte per function (relocation sites masked). No hand-written -reference needed — the Go toolchain IS the oracle. +reference needed; the Go toolchain IS the oracle. ### Added -- `verify`: `GroundTruth` — shells out to `go tool asm`, parses the GOOBJ +- `verify`: `GroundTruth`; shells out to `go tool asm`, parses the GOOBJ output (minimal reader: block offsets, nonpkg symbol table, data index) and returns per-function code bytes. - `gasm verify --ground-truth`: compares gasm's output against the Go @@ -504,11 +521,11 @@ reference needed — the Go toolchain IS the oracle. linker fills) are masked before comparison. - Verified: go-lz4 AVX2 2/2, go-flac AVX2 17/17 functions byte-identical. -## [0.24.0] — 2026-07-29 +## [0.24.0] - 2026-07-29 The full analyze family and stereo PCM decode are now differentially tested. 15 of 17 go-flac AVX2 kernels have bit-for-bit differential coverage; the -two remaining (autocorrAVX2 — FMA reassociation, lpcResidualAVX2 — complex +two remaining (autocorrAVX2: FMA reassociation, lpcResidualAVX2: complex multi-arg) are deferred. ### Added @@ -518,7 +535,7 @@ multi-arg) are deferred. - `verify`: `decodeStereo16AVX2` differential test (500 random interleaved stereo PCM buffers, both channels compared sample-by-sample). -## [0.23.0] — 2026-07-28 +## [0.23.0] - 2026-07-28 The analyze family and 24-bit PCM decode join the differential suite. @@ -530,7 +547,7 @@ The analyze family and 24-bit PCM decode join the differential suite. - `verify`: `decodeMono24AVX2` differential test (500 random 24-bit PCM buffers, sign-extension compared sample-by-sample). -## [0.22.0] — 2026-07-27 +## [0.22.0] - 2026-07-27 The remaining go-flac encoder kernels join the differential suite. @@ -544,39 +561,39 @@ The remaining go-flac encoder kernels join the differential suite. loop). -## [0.21.0] — 2026-07-26 +## [0.21.0] - 2026-07-26 Differential testing extended to all four production kernels and the CLI exposes the full dynamic-analysis toolkit. ### Added -- `verify`: go-flac AVX2 differential tests — `decodeMono16AVX2` (500 +- `verify`: go-flac AVX2 differential tests; `decodeMono16AVX2` (500 random PCM buffers), `pack16AVX2` (500 random int32→int16 packings) and all four decorrelation kernels (200 iterations each: left-side, side-right, mid-side, interleave) compared bit-for-bit against the portable Go references. -- `verify`: go-lz4 AVX-512 differential tests — `decodeBlockAVX512` (3 000 - fuzzed LZ4 blocks + known answers) and `wideCopyAVX512` (0–1024 bytes) +- `verify`: go-lz4 AVX-512 differential tests; `decodeBlockAVX512` (3 000 + fuzzed LZ4 blocks + known answers) and `wideCopyAVX512` (0-1024 bytes) against the same portable oracle as the AVX2 suite. - `gasm verify --abi`: runs each NOSPLIT function with sentinel registers and a red-zone canary, reporting violations. - `gasm verify --profile`: lists the static basic-block count per function. -## [0.20.0] — 2026-07-25 +## [0.20.0] - 2026-07-25 Coverage profiling: the third pillar of Phase 3. Static basic-block enumeration from the assembler's label map, combined with multi-input path -diversity measurement — how many observationally distinct execution paths a +diversity measurement; how many observationally distinct execution paths a test corpus exercises. ### Added -- `verify`: `Kernel.Blocks` / `Kernel.BlockCount` — enumerate basic blocks +- `verify`: `Kernel.Blocks` / `Kernel.BlockCount`; enumerate basic blocks from the assembler's local-label map (every jump target is a block boundary; the function entry is always a block). `decodeBlockAVX2` has 27 blocks. -- `verify`: `Kernel.ProfilePaths` — run the function with a corpus of +- `verify`: `Kernel.ProfilePaths`; run the function with a corpus of argument blocks and collect distinct output fingerprints (the result words); reports path diversity as a lower bound on code coverage. @@ -588,7 +605,7 @@ rt_sigaction handlers fragile in a Go process. The static + path-diversity approach delivers the project's goal (proving the SIMD path and tail handling execute) without fighting the runtime. -## [0.19.0] — 2026-07-24 +## [0.19.0] - 2026-07-24 Runtime ABI checks: the second pillar of Phase 3. The JIT trampoline now has an ABI-checking variant that sets sentinels in the callee-saved registers @@ -598,41 +615,41 @@ detects any illegal write below the stack pointer. ### Added -- `verify`: `CallChecked` / `Kernel.CallFuncChecked` — ABI-checking JIT call +- `verify`: `CallChecked` / `Kernel.CallFuncChecked`; ABI-checking JIT call with sentinel registers and red-zone canary; returns an `ABIReport` (BPClobbered, R14Clobbered, RedZoneHit). -- `verify`: the raw `leaveJITCheckedRaw` trampoline — a TEXT symbol with no +- `verify`: the raw `leaveJITCheckedRaw` trampoline; a TEXT symbol with no ABIInternal wrapper (address obtained via GLOBL/DATA), so the JIT function's RET lands directly in the check code and sees the registers exactly as the function left them. - Tests: deliberate BP/R14 clobberers detected; both go-lz4 kernels confirmed ABI-clean (BP preserved, R14 preserved, red zone intact). -## [0.18.0] — 2026-07-23 +## [0.18.0] - 2026-07-23 Differential testing: the JIT-assembled go-lz4 `decodeBlockAVX2` kernel is -fuzzed against a portable Go reference — 5 000 valid LZ4 blocks compared +fuzzed against a portable Go reference; 5 000 valid LZ4 blocks compared bit-for-bit, plus 2 000 hostile (random garbage) inputs with matching error codes. This is the automated form of the project's bit-identical contract. ### Added -- `verify`: differential fuzz tests — a random LZ4 block generator produces +- `verify`: differential fuzz tests; a random LZ4 block generator produces valid blocks (literals, overlapping matches, extension bytes) and the JIT-assembled kernel's output is compared byte-for-byte against a portable Go decoder; a hostile-input suite confirms error-code agreement on random garbage (no crashes, same classification). -## [0.17.0] — 2026-07-22 +## [0.17.0] - 2026-07-22 Phase 3 begins: dynamic analysis. A JIT execution substrate that assembles -Plan 9 amd64 kernels into executable memory and calls them directly — pure Go +Plan 9 amd64 kernels into executable memory and calls them directly; pure Go (stdlib only, `syscall.Mmap` + an assembly trampoline), no cgo, no external toolchain. ### Added -- `verify` package: JIT infrastructure — `Map` copies machine code into a +- `verify` package: JIT infrastructure; `Map` copies machine code into a W^X memory mapping, `Call` invokes it through an ABI0 trampoline that switches to a prepared stack and back. `Load`/`LoadSource`/`LoadAST` parse, assemble and map a `.s` file in one step; `Kernel.CallFunc` @@ -641,34 +658,34 @@ toolchain. available functions; with `-smoke`, calls each NOSPLIT function with zeroed arguments to confirm the trampoline works end-to-end. - Integration tests: the go-lz4 `decodeBlockAVX2` and `wideCopyAVX2` - kernels (699 and 146 bytes) assemble, map and execute correctly — - known-answer LZ4 blocks decode bit-for-bit, wide copies of 0–1024 bytes + kernels (699 and 146 bytes) assemble, map and execute correctly; + known-answer LZ4 blocks decode bit-for-bit, wide copies of 0-1024 bytes match, malformed input returns the correct error codes. -## [0.16.0] — 2026-07-21 +## [0.16.0] - 2026-07-21 -The scalar conversions between vector and general-purpose registers — the +The scalar conversions between vector and general-purpose registers; the last of the amd64 EVEX instruction set. ### Added - `asm`: the GPR-interchanging conversions, byte for byte against the Go assembler (28 ground-truth cases including memory sources and extended - GPRs): vector to GPR — the signed and truncated VCVT{,T}S{D,S}2SI{,Q} + GPRs): vector to GPR; the signed and truncated VCVT{,T}S{D,S}2SI{,Q} in both VEX and EVEX, and the unsigned VCVT{,T}S{D,S}2USI{L,Q} - (EVEX only); GPR to vector — VCVTSI2SD{L,Q}/VCVTSI2SS{L,Q} (VEX and + (EVEX only); GPR to vector; VCVTSI2SD{L,Q}/VCVTSI2SS{L,Q} (VEX and EVEX) and VCVTUSI2SD{L,Q}/VCVTUSI2SS{L,Q} (EVEX only), whose preserved vector source sits in vvvv (three Plan 9 operands). -## [0.15.0] — 2026-07-20 +## [0.15.0] - 2026-07-20 -The last of the EVEX conversions and narrowing/extending moves — the EVEX +The last of the EVEX conversions and narrowing/extending moves; the EVEX instruction set is now complete save for the GPR-interchanging forms. ### Added -- `asm`: the unsigned and truncating conversions — VCVTPD2PS (and the X/Y +- `asm`: the unsigned and truncating conversions; VCVTPD2PS (and the X/Y spellings, whose length the spelling fixes), VCVTPD2UDQ (X/Y), VCVTTPD2UDQ (X/Y), VCVTTPD2UQQ, VCVTPS2UDQ, VCVTTPS2UDQ, VCVTPS2UQQ, VCVTTPS2UQQ, VCVTTPD2QQ, VCVTTPS2QQ, VCVTUQQ2PD, VCVTUQQ2PS (X/Y) and @@ -681,20 +698,20 @@ instruction set is now complete save for the GPR-interchanging forms. D2M/Q2M), whose K register is a genuine operand rather than a mask and which therefore take no masking suffixes. -## [0.14.0] — 2026-07-19 +## [0.14.0] - 2026-07-19 The floating-point helper and conversion tail of the AVX-512 set, plus -gather and scatter with VSIB addressing — every encoding verified byte for +gather and scatter with VSIB addressing; every encoding verified byte for byte against the Go assembler. ### Added -- `asm`: the floating-point helpers — reciprocals and reciprocal square +- `asm`: the floating-point helpers; reciprocals and reciprocal square roots (VRCP14/VRSQRT14 PD/PS/SD/SS), exponents and mantissas (VGETEXP*, VGETMANT*), scaling by powers of two (VSCALEF*), rounding (VRNDSCALE*), reduction (VREDUCE*), immediate fixup (VFIXUPIMM*) and range selection (VRANGE*), and floating-point class tests (VFPCLASSPD/PS X/Y/Z and - VFPCLASSSD/SS — a new immediate form whose reg field carries the opmask + VFPCLASSSD/SS; a new immediate form whose reg field carries the opmask destination). - `asm`: **gather and scatter with VSIB addressing.** The gathers take both Go spellings: the VEX form with a vector mask register (OP mask, @@ -703,14 +720,14 @@ byte against the Go assembler. data register (a ZMM index with an YMM destination encodes L'L = 10, as the Go assembler emits). The scatters (VSCATTER*/VPSCATTER*) are EVEX only (OP src, K, vsib). All eight gather and eight scatter widths. -- `asm`: the remaining conversions — VCVTQQ2PS (the 512-bit source sets +- `asm`: the remaining conversions; VCVTQQ2PS (the 512-bit source sets the length), VCVTPD2QQ/UQQ, VCVTPS2QQ, VCVTUDQ2PD/PS, the half-precision VCVTPH2PS and VCVTPS2PH (the extract layout with an immediate). -## [0.13.0] — 2026-07-18 +## [0.13.0] - 2026-07-18 The wider AVX-512 set: ternary logic, permutes, compares, expand/compress, -the opmask instructions and the EVEX rounding/SAE/broadcast suffixes — every +the opmask instructions and the EVEX rounding/SAE/broadcast suffixes; every encoding verified byte for byte against the Go assembler. ### Added @@ -718,7 +735,7 @@ encoding verified byte for byte against the Go assembler. - `asm`: the wider EVEX/AVX-512 set, across roughly sixty new ground-truth cases: ternary logic (VPTERNLOGD/Q), the lane shuffles/inserts/extracts (VSHUF{F,I}{32,64}X{2,4}, the VINSERT*/VEXTRACT* {F,I}{32,64}X{2,4,8} - family, VPALIGNR), compares with an opmask destination (VCMPPD/PS/SD/SS — + family, VPALIGNR), compares with an opmask destination (VCMPPD/PS/SD/SS; a new NDS-plus-immediate form with the K register in the reg field), the permutes (VPERMB/W, VPERMI2/T2 D/Q/PD), the wider integer families (VPMADDWD/UBSW, VPMULHUW, VPACKSSWB/USWB/SSDW/USDW, VPABS B/W/D/Q, the @@ -732,15 +749,15 @@ encoding verified byte for byte against the Go assembler. (VMOVSLDUP/VMOVSHDUP), the conversions (VCVTPS2DQ, VCVTTPS2DQ) and the remaining extending and narrowing moves (VPMOVSXBW, VPMOVZXBW, VPMOVWB, VPMOVQB). -- `asm`: the EVEX mnemonic suffixes the Go assembler accepts — the rounding +- `asm`: the EVEX mnemonic suffixes the Go assembler accepts; the rounding modes `.RN_SAE`, `.RD_SAE`, `.RU_SAE`, `.RZ_SAE` (the EVEX b bit with the rounding control in L'L), suppress-all-exceptions `.SAE`, and memory broadcast `.BCST` (the b bit, the vector length preserved, disp8×N scaled - by the element size) — each combinable with the `.Z` zeroing suffix, + by the element size); each combinable with the `.Z` zeroing suffix, validated against the Go assembler's bytes, and rejected on instructions that do not support them. -## [0.12.0] — 2026-07-17 +## [0.12.0] - 2026-07-17 GOOBJ emission: gasm-assembled functions drop into a `go build` without the Go assembler. @@ -748,7 +765,7 @@ Go assembler. ### Added - `asm`: **GOOBJ object output.** `gasm asm --format goobj -p ` - writes the Go toolchain's own object format — the one `cmd/link` consumes + writes the Go toolchain's own object format; the one `cmd/link` consumes directly: the functions as non-package symbols qualified with the package path (exactly as `cmd/asm` records assembly symbols), the `GLOBL` data, one serialized `FuncInfo` per function (argument/frame sizes, the asm @@ -757,8 +774,8 @@ Go assembler. real stack deltas: the assembler now tracks every stack-adjustment boundary through the prologue (`PUSHQ BP`, `SUBQ $frame, SP`) and each `RET`'s epilogue, so frame-pointer functions unwind correctly. The - object preamble — the version-and-experiment header the linker compares - verbatim — is captured from the installed `go tool asm`, so the output is + object preamble; the version-and-experiment header the linker compares + verbatim; is captured from the installed `go tool asm`, so the output is always consistent with the toolchain that links it. - `asm`: relocations against file-local `GLOBL` symbols become `R_PCREL` entries in the GOOBJ output, with the instruction's displacement field @@ -771,7 +788,7 @@ Go assembler. pattern, instead of being rejected as non-integer. -## [0.11.0] — 2026-07-16 +## [0.11.0] - 2026-07-16 Linkable object output: external symbols and relocatable ELF / Mach-O objects. @@ -790,7 +807,7 @@ objects. external symbol; the Mach-O output is verified structurally with `debug/macho`. - `asm`: **external symbol references.** A reference to a symbol no - `GLOBL` in the file defines no longer aborts assembly — it is recorded + `GLOBL` in the file defines no longer aborts assembly; it is recorded as an external relocation (`Image.Externals`, `FuncLayout.Relocs`) and becomes an undefined global symbol in the object output. The raw image s (`--format raw`, the default) still reports them: only an object @@ -802,7 +819,7 @@ objects. writes; without `--format` the behaviour is unchanged (the concatenated image). -## [0.10.0] — 2026-07-15 +## [0.10.0] - 2026-07-15 The EVEX floating-point and conversion set: the packed-double arithmetic, the scalar SD/SS forms, VMOVDDUP and the width-changing conversions, each @@ -810,23 +827,23 @@ verified byte for byte against the Go assembler. ### Added -- `asm`: the rest of the common EVEX/VEX floating-point set — packed double +- `asm`: the rest of the common EVEX/VEX floating-point set; packed double arithmetic (VSUBPD, VDIVPD, VMINPD, VMAXPD, VUNPCKLPD and the EVEX form of VUNPCKHPD), the scalar double and single operations (VSUBSD, VDIVSD, VMINSD, VMAXSD and the full VADDSS/VSUBSS/VMULSS/VDIVSS/VMINSS/VMAXSS - family in both VEX and EVEX — the EVEX scalar forms exist for masked and + family in both VEX and EVEX; the EVEX scalar forms exist for masked and zeroing use), and VMOVDDUP (lane duplication, VEX and EVEX). -- `asm`: the width-changing conversions — VCVTDQ2PS and VCVTPS2PD (VEX and +- `asm`: the width-changing conversions; VCVTDQ2PS and VCVTPS2PD (VEX and EVEX; the destination sets the length for PS→PD), the EVEX form of VCVTDQ2PD, and the packed-double → dword family: VCVTPD2DQ/VCVTTPD2DQ (EVEX-512 only, a ZMM source and an XMM destination) and their X/Y spellings (VCVTPD2DQX/Y, VCVTTPD2DQX/Y), whose length follows the wider - source — a new operand form, since the destination is always XMM while + source; a new operand form, since the destination is always XMM while VEX.L / EVEX.L'L ride with the source (fixed by the spelling even for a memory source). -- `asm`: masking and zeroing on every new form — the scalar SD/SS +- `asm`: masking and zeroing on every new form; the scalar SD/SS arithmetic, the unpacks, VMOVDDUP and the conversions all accept the - explicit K1–K7 operand and the `.Z` suffix the way Go writes them. + explicit K1-K7 operand and the `.Z` suffix the way Go writes them. ### Changed @@ -837,35 +854,35 @@ verified byte for byte against the Go assembler. shares the convention). -## [0.9.0] — 2026-07-14 +## [0.9.0] - 2026-07-14 AVX-512 masking and a wider EVEX integer set. ### Added -- `asm`: **EVEX masking** the way Go writes it — an explicit `K1`–`K7` +- `asm`: **EVEX masking** the way Go writes it; an explicit `K1`-`K7` operand placed among the operands (merging mask), and a `.Z` mnemonic suffix for zeroing (`VPADDD.Z Z1, Z2, K2, Z3`). Supported across the NDS, reg/rm, immediate-shift, align, extract, convert and move forms, including masked comparisons with a K destination (`VPCMPEQD Z0, Z3, K2, K1`). K0 is rejected as an explicit mask, and `.Z` without a mask is an error, matching the Go assembler. -- `asm`: the common AVX-512 F/BW integer set — VPADDB/W, VPSUBB/W, VPANDD/Q, +- `asm`: the common AVX-512 F/BW integer set; VPADDB/W, VPSUBB/W, VPANDD/Q, VPANDND/Q, VPMULLW, VPAVGB/W, the signed/unsigned min/max family for B/W/D/Q elements, the variable shifts VPSLLVD/Q, VPSRLVD/Q, VPSRAVD/Q, the EVEX forms of VPSHUFD/VPSHUFB, and the VMOVDQU8/VMOVDQU16 move aliases. - Register indices 16–31 encode correctly (the mod=11 quirk carries rm[4] + Register indices 16-31 encode correctly (the mod=11 quirk carries rm[4] in X̄). All verified byte for byte against the Go assembler. - `lint`: masked EVEX forms (`.Z` suffix, K operands) are recognised by `unknown-instruction` and exempted from `operand-count`. ### Fixed -- `asm`: EVEX register–register operands with indices 16–31 encoded rm[4] +- `asm`: EVEX register-register operands with indices 16-31 encoded rm[4] into B̄ instead of X̄ (the EVEX mod=11 extension quirk), producing wrong prefix bytes for X16+/Y16+ r/m operands. -## [0.8.0] — 2026-07-13 +## [0.8.0] - 2026-07-13 Standard CLI ergonomics. @@ -881,20 +898,20 @@ Standard CLI ergonomics. - The version is primarily available as the standard `gasm --version` / `-V` flag; the `gasm version` spelling remains as an alias. -## [0.7.0] — 2026-07-12 +## [0.7.0] - 2026-07-12 The formatter behaves like `go fmt` and canonicalises block separation. ### Added -- `gasm fmt` now works like `go fmt`: with no arguments — or with a directory - argument — it reformats every `.s` file below it in place and lists the +- `gasm fmt` now works like `go fmt`: with no arguments; or with a directory + argument; it reformats every `.s` file below it in place and lists the changed files, skipping `.` and `_` directories (`.git`, `_refs`, …). Explicit file arguments keep the `-w` / standard-output behaviour. ### Changed -- `s`: canonical blank-line layout — a new block (a label, `TEXT` or +- `s`: canonical blank-line layout; a new block (a label, `TEXT` or `GLOBL`) is preceded by exactly one blank line, neither more nor less. Comments leading a block stay with it (the blank line goes before them), stacked labels share their block, the function's first label keeps hugging @@ -903,7 +920,7 @@ The formatter behaves like `go fmt` and canonicalises block separation. kernels were reformatted with this release and remain byte-identical when assembled. -## [0.6.0] — 2026-07-11 +## [0.6.0] - 2026-07-11 Calibrated to the Go ABI: `register-clobber` stops reporting legal code, and the encoder learns the legacy SSE moves. @@ -912,13 +929,13 @@ the encoder learns the legacy SSE moves. - `lint`: **`register-clobber` is now calibrated to the Go ABI** (`cmd/compile/abi-internal.md`), not the platform ABI. Go's stack-based - ABI0 has no System V style callee-saved registers — amd64 `BX`, `R12`–`R15` + ABI0 has no System V style callee-saved registers; amd64 `BX`, `R12`-`R15` and the arm64/riscv64/loong64 scratch sets are caller-saved or permanent scratch, and hand-written kernels may clobber them freely. The rule now audits only the registers Go fixes across calls: the frame pointer and the goroutine pointer (amd64 `BP`/`R14`, arm64 `R18`/`R28`/`R29`, riscv64 `X27`, loong64 `R22`), and the goroutine pointer is reported only when the - function can reach the runtime (is not `NOSPLIT` or makes a call) — the + function can reach the runtime (is not `NOSPLIT` or makes a call); the ABI0 transition restores it on those paths, and NOSPLIT call-free leaves may use it, exactly as the runtime's own assembly does. Both go-flac kernels now lint with zero diagnostics. @@ -935,21 +952,21 @@ the encoder learns the legacy SSE moves. ### Added -- `asm`: the legacy (non-VEX) SSE moves — `MOVOU`/`MOVO` (the Plan 9 names +- `asm`: the legacy (non-VEX) SSE moves; `MOVOU`/`MOVO` (the Plan 9 names for MOVDQU/MOVDQA), `MOVUPS`/`MOVAPS`/`MOVUPD`/`MOVAPD` and the scalar - `MOVSD`/`MOVSS` — and `VMOVDQU64` in the EVEX set. All verified byte for + `MOVSD`/`MOVSS`; and `VMOVDQU64` in the EVEX set. All verified byte for byte against the Go assembler. -## [0.5.0] — 2026-07-10 +## [0.5.0] - 2026-07-10 EVEX / AVX-512: the go-flac AVX-512 kernel now assembles, byte-identically to the Go toolchain, completing the production-kernel coverage. ### Added -- `asm`: **EVEX (AVX-512) encoding** — the four-byte EVEX prefix with the - 5-bit register fields (Z0–Z31, X/Y 16–31, with the reg-r/m X̄ quirk and - V'̄ shared between vvvv and the SIB index), opmask registers (K0–K7) as +- `asm`: **EVEX (AVX-512) encoding**; the four-byte EVEX prefix with the + 5-bit register fields (Z0-Z31, X/Y 16-31, with the reg-r/m X̄ quirk and + V'̄ shared between vvvv and the SIB index), opmask registers (K0-K7) as operands and as mask destinations, and the compressed disp8×N displacement (the multiplier follows the memory operand's size, as the Go assembler's opcode tables prescribe). Covers every AVX-512 instruction the go-flac @@ -959,20 +976,20 @@ the Go toolchain, completing the production-kernel coverage. extracts VEXTRACTI64X4/VEXTRACTF64X4, VFMADD231PD, VADDPD, VMULPD, the broadcasts VPBROADCASTD/Q (GPR and memory sources take different opcodes) and the mask moves KMOVW/KTESTW. Masking/zeroing suffixes are out of scope - — the kernels use neither. + ; the kernels use neither. - `asm`: `AssembleFile` now accepts file-defined global (`non-<>`) symbols too; a reference is external only when no `GLOBL` in the file defines it. ### Fixed -- `asm`: registers X16–Y31 force the EVEX encoding of dual-form mnemonics; +- `asm`: registers X16-Y31 force the EVEX encoding of dual-form mnemonics; previously a `VPBROADCASTD AX, Y30` fell into the VEX encoder, which cannot represent indices above 15 and silently truncated them. -- `asm`: the VEX encoder now rejects vector register indices 16–31 instead of +- `asm`: the VEX encoder now rejects vector register indices 16-31 instead of encoding a truncated (wrong) register. -## [0.4.0] — 2026-07-09 +## [0.4.0] - 2026-07-09 The standalone assembler reaches the whole go-flac AVX2 kernel: static symbols assemble, and all 17 kernel functions now match the Go toolchain's @@ -980,10 +997,10 @@ machine code byte for byte. ### Added -- `asm`: **file-level assembly** — `AssembleFile` turns a parsed file into an +- `asm`: **file-level assembly**; `AssembleFile` turns a parsed file into an `Image`: the function bodies in source order followed by a data section built from the file's `GLOBL`/`DATA` directives (each symbol 16-aligned). -- `asm`: **static-symbol (`SB`) operands** — `mask<>(SB)` references encode as +- `asm`: **static-symbol (`SB`) operands**; `mask<>(SB)` references encode as RIP-relative loads with a patched disp32, resolved against the image layout so the output is self-consistent and position-independent. External (non-file-local) symbols are rejected with a clear error: they need @@ -992,7 +1009,7 @@ machine code byte for byte. and writes the whole image (code + data) with `-o`. -## [0.3.0] — 2026-07-08 +## [0.3.0] - 2026-07-08 The assembler reaches byte-identical parity with the Go toolchain on the production go-flac AVX2 kernels: every one of the 15 kernel functions that @@ -1002,18 +1019,18 @@ support). ### Added -- `asm`: the scalar instruction families the kernels use — `CMOVcc` and +- `asm`: the scalar instruction families the kernels use; `CMOVcc` and `SETcc` (conditions spelled exactly like the jumps), `LZCNT`/`TZCNT` (legacy `F3 0F BD/BC`), the sign/zero-extending moves (`MOVBLZX`, `MOVBQZX`, `MOVWLZX`, `MOVWQZX`, `MOVWLSX`, `MOVLQSX`), `CVTSL2SD`/`CVTSQ2SD` (the legacy SSE encoding, as the Go assembler emits it), the traditional three-operand `IMUL3{W,L,Q}`, and the variable-count vector shifts - (`VPSRLQ X0, Y8, Y8` — the count in an XMM register or memory takes the + (`VPSRLQ X0, Y8, Y8`; the count in an XMM register or memory takes the ordinary NDS form). -- `asm`: **jump relaxation** — jumps start in the short (rel8) form and +- `asm`: **jump relaxation**; jumps start in the short (rel8) form and expand to rel32 when the settled displacement does not fit, iterating the layout to a fixed point (CALL is always rel32). -- `asm`: **jump-to-jump folding** — a conditional jump to a label whose only +- `asm`: **jump-to-jump folding**; a conditional jump to a label whose only instruction is an unconditional jump is redirected to the ultimate target, replicating the Go toolchain's linker, which chases such chains before it encodes branches. @@ -1025,12 +1042,12 @@ support). - `asm`: `CMP` with a register or memory operand computed **second − first** instead of first − second, silently inverting every condition that followed (`CMPQ SI, R10; JGE` tested R10 ≥ SI). The encoding now always records - first − second — `CMP r/m, r` with the first operand in r/m, `CMP r, r/m` - with the first operand in reg — and is byte-identical to the Go assembler. + first − second; `CMP r/m, r` with the first operand in r/m, `CMP r, r/m` + with the first operand in reg; and is byte-identical to the Go assembler. - `asm`: register-to-register `MOV` now uses the `r/m ← r` opcode (reg = source), the Go assembler's choice; the output is byte-identical. -## [0.2.0] — 2026-07-07 +## [0.2.0] - 2026-07-07 The Phase 2 assembler grows the SIMD set: shuffles, extract/insert, permute and the moves, on top of the Phase 1 VEX forms. @@ -1043,10 +1060,10 @@ and the moves, on top of the Phase 1 VEX forms. - the immediate shuffle (`VPSHUFD`, `VPERMQ`), - the three-operand-plus-immediate form (`VSHUFPD`, `VPERM2I128`, `VINSERTI128`), - - the lane extract (`VEXTRACTI128`, `VEXTRACTF128` — the YMM source occupies + - the lane extract (`VEXTRACTI128`, `VEXTRACTF128`; the YMM source occupies the ModRM.reg field, the XMM/memory destination the r/m field), - the direction-sensitive moves (`VMOVDQU`, `VMOVUPD`, `VMOVD`, `VMOVQ`, - `VMOVSD` — each direction picks its own opcode and VEX.W; a vector→vector + `VMOVSD`; each direction picks its own opcode and VEX.W; a vector→vector move uses the store-form layout, matching the Go assembler), - the no-operand `VZEROUPPER`, and `VPERMD` in the NDS form, - the floating-point and FMA set (`VADDPD`, `VMULPD`, `VXORPD`, @@ -1055,21 +1072,21 @@ and the moves, on top of the Phase 1 VEX forms. the encoder now covers every integer, shuffle and FP instruction the go-flac AVX2 kernels use. - `asm`: `CMP` accepts the immediate in the second operand position - (`CMPL CX, $31`) — the spelling the Go assembler accepts — encoding it + (`CMPL CX, $31`), the spelling the Go assembler accepts, encoding it identically to the immediate-first form. ### Fixed - `asm`: an unused VEX.vvvv field is now stored as `1111` (v̄vvv = 1111), as - the hardware requires — the previous value (`0000`) made the two-operand + the hardware requires; the previous value (`0000`) made the two-operand reg/rm forms (VPMOVSXWD, VPBROADCASTD, VMOVMSKPS, …) raise #UD on real CPUs and differ from the Go assembler's bytes. The round-trip decoder ignores the field on these instructions, which is why the byte-for-byte Go comparison (added this release) is now part of the test suite. -## [0.1.0] — 2026-07-06 +## [0.1.0] - 2026-07-06 -Initial release — the Phase 1 foundation. +Initial release; the Phase 1 foundation. ### Added @@ -1082,11 +1099,11 @@ Initial release — the Phase 1 foundation. - `arch`: register files and **complete** instruction tables for amd64, arm64, riscv64 and loong64, with the middle-dot symbol separator and static (`<>`) symbols. Instruction names are generated from the Go toolchain's own - assembler source (`just gen`) — the `anames` opcode lists plus the common + assembler source (`just gen`); the `anames` opcode lists plus the common opcodes and the per-architecture front-end aliases (arm64 `B`/`BL`, the `.P`/`.W` addressing suffixes, loong64 `JAL`, the x86 conditional-jump - spellings) — so every mnemonic the real assembler accepts is recognised. -- `lint`: conservative rules — `unknown-instruction`, `operand-count`, + spellings); so every mnemonic the real assembler accepts is recognised. +- `lint`: conservative rules; `unknown-instruction`, `operand-count`, `undefined-label`, `duplicate-label`, `missing-ret`, `missing-textflag-include`, `abi-argsize` and `unreachable-code`. Macro invocations are recognised (in-file `#define` names and underscore @@ -1108,9 +1125,9 @@ Initial release — the Phase 1 foundation. - `lsp`: a Language Server Protocol server over stdio providing completion, hover documentation, document symbols, publish-diagnostics and semantic-token highlighting. -- `asm`: a standalone amd64 (x86-64) assembler — an instruction encoder (REX/ +- `asm`: a standalone amd64 (x86-64) assembler; an instruction encoder (REX/ ModR-M/SIB/displacement/immediate plus the scalar instruction set, and VEX/ - AVX2 SIMD across three operand forms — NDS, reg/rm and immediate-shift — + AVX2 SIMD across the three operand forms NDS, reg/rm and immediate-shift, covering the bulk of the integer SIMD set) validated by round-trip decoding against `golang.org/x/arch`, and an assembler that drives the parser's AST into the encoder with local-label resolution and `FP`/`SP` frame mapping diff --git a/SECURITY.md b/SECURITY.md new file mode 100644 index 0000000..f391217 --- /dev/null +++ b/SECURITY.md @@ -0,0 +1,40 @@ +# Security policy + +## Supported versions + +Security fixes go to the newest release and to the `development` branch. Older +releases do not receive them. + +| Version | Supported | +|---|---| +| 0.33.0 | yes | +| older releases | no | + +## Reporting a vulnerability + +**Do not open a public issue for a security problem.** A public report tells everyone +about the flaw before there is a fix. Report it privately to +**opensource@petrbalvin.org**. + +Include: + +- the version or commit you tested, and the platform +- what the problem is, and what an attacker gains from it +- the smallest reproducer you have, ideally a test or a single command +- a suggested fix, if you have one + +## What to expect + +- A human reads the report, and you get an acknowledgement. +- You are kept informed while the fix is being made, and told when it ships. +- The fix is released before the details are published, and the timing is agreed with + you. +- The reporter is credited in the release notes unless they ask otherwise. + +## Out of scope + +- Findings that require the attacker to already run code as the user, or to have local + access. +- Missing hardening with no demonstrated impact. +- Flaws in a third-party dependency: report them to that project, and to this one only + when this project's use of it makes them reachable.