chore: prepare release v0.36.0
Test / test (push) Successful in 32s
Release / build (amd64, freebsd) (push) Successful in 52s
Release / build (amd64, linux) (push) Successful in 13s
Release / build (arm64, freebsd) (push) Successful in 51s
Release / build (arm64, linux) (push) Successful in 29s
Release / build (loong64, linux) (push) Successful in 29s
Release / build (riscv64, linux) (push) Successful in 31s
Release / release (push) Successful in 10s

Assisted-by: GLM 5.3 Flash
This commit is contained in:
petrbalvin committed 2026-10-07 22:37:42 +02:00
1 parent e5df0a9d84
commit a77dd12d2d
13 files changed
+217 -77

No files matched your search

+9
View File
@@ -100,6 +100,15 @@ encoder backlog that `gasm audit-instructions` measures. The families:
whose register list rides the inverted V′VVV field. Mixing VEX and legacy
SSE in one loop pays the AVX-SSE transition penalty on every switch: keep
a loop in one dialect.
- **The extension layer.** The families the toolchain's own table carries
late or not at all assemble through the extension mechanism: BF16,
VP2INTERSECT, the complete AVX512-FP16 set (the packed and scalar FMA
families, the complex multiply and complex FMA pairs, VMINMAXPH and
VMINMAXSH), and the AVX-VNNI-INT16 dot products through a VEX path. The
layer takes `k0` to `k7` write masks with merging and zeroing, the
`{1toN}` broadcast, the embedded rounding and `{sae}` decorations and the
imm8 controls, and its refusals surface as `extension-form` lint errors
with the reason the encoder would give.
- **Cryptographic and counting extensions.** AES-NI, SHA-1 and SHA-256,
PCLMULQDQ, GFNI.
- **System.** `CPUID`, `RDTSC`, `SYSCALL`, the fences, `LDMXCSR` and
+29 -2
View File
@@ -12,8 +12,8 @@ toolchain's. The complete mnemonic inventory lives in the generated appendix
the stack pointer. There is no R31: thirty-one names and ZR.
- Floating-point and SIMD share one file written `Vn`; where an instruction
is scalar floating point the operand may be written `Fn` (`F0` to `F31`).
- SVE register names (`Z0` to `Z31`, `P0` to `P15`) exist in the assembler's
tables.
- SVE registers `Z0` to `Z31` and predicates `P0` to `P15`; the SVE2 and
SVE2.1 families encode, see SVE below.
- Roles the convention fixes: `RSP` is the stack pointer, `R29` the frame
pointer, `R30` the link register, `R26` the closure context and `R27` the
assembler's scratch register. The goroutine pointer lives in `R28` and is
@@ -106,6 +106,33 @@ scalar floating-point instructions being the exceptions. Operands carry an
arrangement suffix, `V5.H8`, and structure loads and stores use bracket
lists, `[V21.B16]`, with element selection as `V9.S[1]`.
## SVE
The scalable vector extension encodes through the `Z0` to `Z31` vector
registers and the `P0` to `P15` predicates, with arrangement suffixes on
the `Z` registers and the merging and zeroing qualifiers the toolchain's
SVE2 and SVE2.1 families carry: the narrowing two-to-one arithmetic, the
cryptographic set including ZADCLB, BFloat16 arithmetic, the predicate
counters and reductions, the pairwise and quadword forms, the
multiple-structure loads and stores, shift by vector and the
shift-immediate scheme, CLASTA and CLASTB, the compare and last-active
predicate families, and the vector-length arithmetic `ADDVL`, `ADDPL` and
`RDVL`, whose immediates count vectors or predicates rather than bytes.
The gather and scatter loads take their addresses through the five
addressing modes the toolchain defines. The source grammar is the
toolchain's own, and the byte output is pinned against it corpus-wide.
## Synthesised forms
- `GETCALLERPC` reads the return address the frame state describes: a leaf
function reads `R30`, a framed body reads the prologue's save slot.
- `REM`, `REMW`, `UREM` and `UREMW` synthesise a remainder from `SDIV` or
`UDIV` and the `MSUB` tail, with `RSP` refused as a destination.
- `DWORD $imm` lays eight little-endian bytes per immediate.
- `MOVD tls_g(SB), Rd` materialises a TLS local-exec load as a single
`MOVZ` carrying `R_ARM64_TLS_LE`, keyed off the file's own
`GLOBL ... TLSBSS` declaration.
## Alignment
`PCALIGN $n` pads to a power-of-two boundary between 8 and 2048 and also
+11 -6
View File
@@ -23,11 +23,16 @@ TEXT symbol(SB), [flags,] $framesize[-argsize]
- The symbol is an `·Name(SB)` reference into the current package, or a
fully qualified name.
- The optional flag argument is a constant expression, normally an OR of the
names from `textflag.h`, the table below. Without `#include "textflag.h"`
the names are not macros and the assembler reports the misleading error
`illegal or missing addressing mode for symbol NOSPLIT`: include the
header first.
- The optional flag argument is a constant expression: an OR of the names
from `textflag.h`, the table below, a bare number (`4`), or a
parenthesised combination (`(NOSPLIT|NOFRAME)`). Every spelling
suppresses exactly what the names suppress. Without
`#include "textflag.h"` the names are not macros and the assembler
reports the misleading error `illegal or missing addressing mode for
symbol NOSPLIT`: include the header first. A name outside the table is
rejected with the toolchain's wording, and two of the toolchain's own
TEXT checks fire here: `ABIInternal requires NOSPLIT`, and
`NOFRAME functions must have a frame size of 0` for a positive frame.
- `$framesize-argsize` is two constants, not a subtraction: the local frame
size in bytes, and the caller's argument area in bytes. The argument size
may be omitted entirely, `$16`, which marks the argument size unknown
@@ -59,7 +64,7 @@ Values from `textflag.h`, in agreement with `cmd/internal/obj/textflag.go`:
| WRAPPER | 32 | TEXT | a wrapper; must not disable `recover` |
| NEEDCTXT | 64 | TEXT | a closure consuming the context register |
| TLSBSS | 256 | data | a thread local word in BSS |
| NOFRAME | 512 | TEXT | no frame setup; the zero-frame spelling (the toolchain accepts a positive frame beside it and still allocates the frame, and the arm64 BSD syscall stubs pair it with $-8 deliberately) |
| NOFRAME | 512 | TEXT | no frame setup; the frame size must be zero or negative (the toolchain rejects a positive frame beside it, and the arm64 BSD syscall stubs pair it with $-8 deliberately) |
| REFLECTMETHOD | 1024 | TEXT | the function calls `reflect.Type.Method` or `MethodByName` |
| TOPFRAME | 2048 | TEXT | the outermost frame; unwinders stop here |
| ABIWRAPPER | 4096 | TEXT | an ABI transition wrapper |
+6
View File
@@ -41,6 +41,12 @@ R13` is `add.d R13, R12, R11`, and the two-operand form
- Jump and branch instructions keep the GNU order: `BEQ R0, R4, label1`.
- The bitfield family is `BSTRINSW`, `BSTRINSV`, `BSTRPICKW`, `BSTRPICKV`
`$<msb>, <Rj>, $<lsb>, <Rd>`.
- **The register-pair spelling** `R4:R5` names two registers in one
operand, the way the toolchain's pair syntax does: the colon splits the
operand and the halves swap, so `MULV R4:R5, R6` encodes exactly as
`MULV R5, R4, R6`. Any instruction shape accepts the spelling, decided
afterwards by ordinary operand matching, and a malformed pair fails with
the toolchain's wording.
## Addressing
+13
View File
@@ -55,6 +55,10 @@ the FCSR.
a literal pool in the binary.
- A 32-bit constant is accepted by `ADDI`, `ANDI`, `ORI` and `XORI`, and
the assembler synthesises values that exceed the 12-bit encoding window.
- A memory offset beyond the 12-bit immediate expands the way the
toolchain's loads and stores do: the upper bits go into the assembler's
temporary register and the access runs against it, byte-identically, and
an offset the toolchain refuses is refused with its wording.
- `MOVF` and `MOVD` materialise floating-point constants, encoding them as
`FLW` and `FLD` from a pool location unless the constant is exactly 0.0.
@@ -85,6 +89,15 @@ of the source, and register choice influences how much compresses.
Hand-writing compressed instructions in source is accepted but discouraged.
The debug flag `compressinstructions=0` turns the automatic conversion off.
## Pseudo-operations
- `END` is accepted anywhere and emits nothing, the function-end marker
the toolchain's front end also swallows.
- `GETCALLERPC` reads the return address the frame state describes: a leaf
function reads `X1`, a framed body reads the prologue's save slot at
`0(SP)`, and both spellings take the compressed encoding where the
register choice allows it.
## Vector extension
`VSETVLI` writes its vtype components in uppercase with the destination