chore: prepare release v0.36.0
Test / test (push) Successful in 32s
Release / build (amd64, freebsd) (push) Successful in 52s
Release / build (amd64, linux) (push) Successful in 13s
Release / build (arm64, freebsd) (push) Successful in 51s
Release / build (arm64, linux) (push) Successful in 29s
Release / build (loong64, linux) (push) Successful in 29s
Release / build (riscv64, linux) (push) Successful in 31s
Release / release (push) Successful in 10s

Assisted-by: GLM 5.3 Flash
This commit is contained in:
petrbalvin committed 2026-10-07 22:37:42 +02:00
1 parent e5df0a9d84
commit a77dd12d2d
13 files changed
+217 -77

No files matched your search

+29 -2
View File
@@ -12,8 +12,8 @@ toolchain's. The complete mnemonic inventory lives in the generated appendix
the stack pointer. There is no R31: thirty-one names and ZR.
- Floating-point and SIMD share one file written `Vn`; where an instruction
is scalar floating point the operand may be written `Fn` (`F0` to `F31`).
- SVE register names (`Z0` to `Z31`, `P0` to `P15`) exist in the assembler's
tables.
- SVE registers `Z0` to `Z31` and predicates `P0` to `P15`; the SVE2 and
SVE2.1 families encode, see SVE below.
- Roles the convention fixes: `RSP` is the stack pointer, `R29` the frame
pointer, `R30` the link register, `R26` the closure context and `R27` the
assembler's scratch register. The goroutine pointer lives in `R28` and is
@@ -106,6 +106,33 @@ scalar floating-point instructions being the exceptions. Operands carry an
arrangement suffix, `V5.H8`, and structure loads and stores use bracket
lists, `[V21.B16]`, with element selection as `V9.S[1]`.
## SVE
The scalable vector extension encodes through the `Z0` to `Z31` vector
registers and the `P0` to `P15` predicates, with arrangement suffixes on
the `Z` registers and the merging and zeroing qualifiers the toolchain's
SVE2 and SVE2.1 families carry: the narrowing two-to-one arithmetic, the
cryptographic set including ZADCLB, BFloat16 arithmetic, the predicate
counters and reductions, the pairwise and quadword forms, the
multiple-structure loads and stores, shift by vector and the
shift-immediate scheme, CLASTA and CLASTB, the compare and last-active
predicate families, and the vector-length arithmetic `ADDVL`, `ADDPL` and
`RDVL`, whose immediates count vectors or predicates rather than bytes.
The gather and scatter loads take their addresses through the five
addressing modes the toolchain defines. The source grammar is the
toolchain's own, and the byte output is pinned against it corpus-wide.
## Synthesised forms
- `GETCALLERPC` reads the return address the frame state describes: a leaf
function reads `R30`, a framed body reads the prologue's save slot.
- `REM`, `REMW`, `UREM` and `UREMW` synthesise a remainder from `SDIV` or
`UDIV` and the `MSUB` tail, with `RSP` refused as a destination.
- `DWORD $imm` lays eight little-endian bytes per immediate.
- `MOVD tls_g(SB), Rd` materialises a TLS local-exec load as a single
`MOVZ` carrying `R_ARM64_TLS_LE`, keyed off the file's own
`GLOBL ... TLSBSS` declaration.
## Alignment
`PCALIGN $n` pads to a power-of-two boundary between 8 and 2048 and also