chore: prepare release v0.36.0
Test / test (push) Successful in 32s
Release / build (amd64, freebsd) (push) Successful in 52s
Release / build (amd64, linux) (push) Successful in 13s
Release / build (arm64, freebsd) (push) Successful in 51s
Release / build (arm64, linux) (push) Successful in 29s
Release / build (loong64, linux) (push) Successful in 29s
Release / build (riscv64, linux) (push) Successful in 31s
Release / release (push) Successful in 10s

Assisted-by: GLM 5.3 Flash
This commit is contained in:
petrbalvin committed 2026-10-07 22:37:42 +02:00
1 parent e5df0a9d84
commit a77dd12d2d
13 files changed
+217 -77

No files matched your search

+109 -3
View File
@@ -7,17 +7,75 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
## [development]
## [0.36.0] - 2026-10-07
### Added
- **The SVE2 and SVE2.1 instruction sets on arm64.** Over 350 SVE
mnemonics across 537 encoded forms join the extension layer: the
narrowing two-to-one arithmetic family, the SVE2 cryptographic set
including ZADCLB, BFloat16 arithmetic, predicate counters and
reductions, the pairwise and quadword forms, multiple-structure loads
and stores (LD2/ST2 through LD4/ST4), shift by vector and the
shift-immediate scheme with its element-size encoding, CLASTA and
CLASTB, the last-active and compare predicate families, and the
vector-length arithmetic ADDVL, ADDPL and RDVL. Every form derives
from the toolchain's own encoder tables and is pinned against the
GOROOT corpus words.
- **The complete AVX512-FP16 set and AVX-VNNI-INT16 on amd64.** The
extension layer reaches 191 entries: the packed and scalar FMA
families (VFMADD, VFMSUB, VFMADDSUB and VFMSUBADD in the PH widths
with their SH mirrors), the complex multiply and complex FMA pairs
(VFMULC, VFCMULC, VFMADDC, VFCMADDC, packed and scalar, each
conjugating the source its prefix names), VMINMAXPH and VMINMAXSH
under their imm8 control, and the AVX-VNNI-INT16 dot products
(VPDPWSUD, VPDPWSUDS, VPDPWUSD, VPDPWUSDS) through a new VEX
encoding path beside the EVEX one. Every entry carries golden
vectors byte-checked against GNU as.
- **Extension-layer instructions assemble straight from `.s` source on
amd64.** The registered mnemonics dispatch from the assembly front
end with their operand grammar: `k0` through `k7` write masks with
merging and zeroing, `{1toN}` broadcast, embedded rounding and
`{sae}`, memory and scaled-index operands, and the imm8-control
forms, the VEX forms beside the EVEX ones. `gasm lint` surfaces the
layer's refusals as `extension-form` errors instead of letting a bad
shape reach the encoder, and every distinct mnemonic is pinned from
source text against the registry's bytes.
- **arm64 output byte-identical with the toolchain across the GOROOT
tree.** 31 of the 36 files the differential harness walks (90
functions, 27408 bytes) now assemble without a differing byte: the
missing shapes are in (GETCALLERPC, the REM, REMW, UREM and UREMW
family, DWORD, the FCVTHS, FCVTSH, FCVTDH and FCVTHD conversions, TLS
local-exec loads as a single MOVZ with the TLS relocation), the
literal pool drains mid-function when a function's literals would
overflow the ±512 KiB load-literal range, and frame-relative
addresses and flag-setting logicals materialise the way
cmd/internal/obj/arm64 writes them.
- **riscv64 END and GETCALLERPC.** END is accepted anywhere and emits
nothing, and GETCALLERPC reads the return address the toolchain's
rewrite describes: a leaf function reads LR, a framed body reads the
prologue's save slot, with the compressed spellings where they fit.
- **The loong64 register-pair spelling.** `MULV R4:R5, R6` parses the
way the toolchain's pair sugar does: the colon splits the operand and
the halves swap, any instruction shape takes the spelling, and the
malformed forms fail with the toolchain's own wording.
- **Fuzz targets across the tool.** The assembler carries
per-architecture FuzzAssemble targets and the front end carries
targets over the lexer, parser, formatter and the extension
registry; the seed corpora run as ordinary tests, and `just fuzz`
drives the campaigns locally. Hardening the targets found caps GLOBL
and DATA at 64 MiB per symbol and 128 MiB per section, so a crafted
input cannot materialise unbounded memory.
- **The FreeBSD port of the debugger.** `gasm debug` runs on FreeBSD on
amd64, arm64 and riscv64 with the same interactive surface as on Linux:
amd64 and arm64 with the same interactive surface as on Linux:
breakpoints, hardware watchpoints (x86 debug registers, the arm64 debug
register file), single-stepping, register and memory access, all behind
the kernel's own ptrace requests, with tracee memory through `PT_IO` and
stop reports through `PT_LWPINFO`. The JIT substrate maps executable
memory through `golang.org/x/sys/unix`, so `verify` builds on FreeBSD
too. The pipeline compile-gates all three architectures; live
validation awaits a FreeBSD machine.
too. The dispatched suite compile-gates both architectures, the
release pipeline ships the two binaries, and live validation awaits a
FreeBSD machine.
- **Workspace-wide navigation in the language server.** `gasm lsp` indexes
the `.s` files under the workspace root beyond the documents the editor
has open, so go-to-definition, find references and workspace symbol search
@@ -75,6 +133,23 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
### Changed
- **TEXT flag operands behave like the toolchain's.** A numeric or
parenthesised flag list (`TEXT ·f(SB), 4, $4096-0`,
`(NOSPLIT|NOFRAME)`) now suppresses the prologue and stack guard the
names suppress, where the numbers were silently ignored and the
prologue was emitted anyway; unknown flag names are rejected with the
toolchain's wording instead of passing unnoticed; and the toolchain's
own TEXT diagnostics fire, `ABIInternal requires NOSPLIT` and
`NOFRAME functions must have a frame size of 0`.
- **The disassembler names what it used to print raw.** The 195 amd64
encodings that decoded to anonymous renders now decode to their
instructions and round-trip, the BMI1 and BMI2 VEX set, RDSEED, RDPID,
the WAITPKG family, ENDBR, CLDEMOTE, UD1 and RORX among them, with the
arm64 HVC, SMC and SB forms and 74 loong64 rendering corrections
beside.
- **Variadic macro parameters are rejected like the toolchain rejects
them.** A macro declaration whose parameter list the toolchain
refuses no longer parses into a broken definition.
- **The corpus audit assembles like the build.** A file's `//go:build`
constraint decides which target architectures attempt it: cpu_x86.s is
an x86 build alone, and the msan and goexperiment.runtimesecret trees
@@ -94,6 +169,37 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
### Fixed
- **riscv64 memory offsets beyond the 12-bit immediate no longer
truncate silently.** `LD 4096(X6), X5` encoded as `LD X5, 0(X6)` and
read the wrong address; the expansion the toolchain performs (the
upper bits into its temporary register, then the access) is emitted
now, byte-identically, and offsets the toolchain refuses are refused
with its wording.
- **riscv64 operand acceptance matches the toolchain's.** 398 operand
shapes the toolchain rejects assembled without complaint (a float
register where an integer one is required, out-of-range CSRs, vector
shapes against the wrong bank, MOV family widths), and 488 more
produced the wrong diagnostic; all now fail with the toolchain's
texts, and a 1030-case error-parity catalogue pins them against the
toolchain's own negative files.
- **The debugger survives its failure paths.** Killing a debuggee that
had already died or sat ptrace-stopped could block the session
forever; Kill now wakes the tracee, signals and reaps it without
waiting on a parked child, so every launch-failure path returns
instead of hanging the tool.
- **Encoder defects found by fuzzing and the differential audits.** A
GLOBL with a negative size panicked the assembler and a small crafted
input materialised 3.92 GiB of DATA; certain SVE VTBL operands
panicked; an over-strict guard refused BIC into RSP where the
toolchain assembles it; byte-register operands now select the 8-bit
forms the toolchain picks (`XADDL DL, DL` and its companions); and
the loong64 shift compositions and indexed-offset forms that dropped
fields re-encode faithfully.
- **Formatter edges around labels and macros.** A stacked label that
names a macro stays attached to it, macro content behind a label
stays on its line, a selector-folded label resolves to its macro, an
empty TEXT body formats, and a file ending in a comment no longer
loses the comment.
- **Rename edits land in their own documents.** A rename collected the
ranges of every reference across the open documents but applied them all
to the document that started it, so renaming a symbol used in a second