From a77dd12d2d4fdd2abdfb6833ad599de060f1e29a Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Petr=20Balv=C3=ADn?= Date: Wed, 7 Oct 2026 22:37:42 +0200 Subject: [PATCH] chore: prepare release v0.36.0 Assisted-by: GLM 5.3 Flash --- .gitea/workflows/freebsd.yml | 41 ------------- .gitea/workflows/release.yml | 13 +++- .gitea/workflows/suite.yml | 18 ++++-- CHANGELOG.md | 112 ++++++++++++++++++++++++++++++++++- CONTRIBUTING.md | 5 +- README.md | 18 +++--- docs/CLI.md | 6 +- docs/DEVELOPMENT.md | 5 +- docs/asm/AMD64.md | 9 +++ docs/asm/ARM64.md | 31 +++++++++- docs/asm/DIRECTIVES.md | 17 ++++-- docs/asm/LOONG64.md | 6 ++ docs/asm/RISCV64.md | 13 ++++ 13 files changed, 217 insertions(+), 77 deletions(-) delete mode 100644 .gitea/workflows/freebsd.yml diff --git a/.gitea/workflows/freebsd.yml b/.gitea/workflows/freebsd.yml deleted file mode 100644 index 2d15865..0000000 --- a/.gitea/workflows/freebsd.yml +++ /dev/null @@ -1,41 +0,0 @@ -# FreeBSD compile gates. Dispatched by hand, never on a push. -# -# The debugger's ptrace surface and the JIT substrate are the two -# FreeBSD-portable layers the tree carries; the forge has no FreeBSD runner, -# so they can only be compile-gated, and three foreign-GOOS builds of the -# whole module are minutes of one-core work the push pipeline's budget cannot -# carry. The push pipeline stays fast and light; this workflow is the -# deliberate run, before a release or after touching the ported layers. -# Running the ptrace suite itself needs real FreeBSD hardware. -# -# A dispatched workflow takes no concurrency block: it is one deliberate run. -name: FreeBSD build - -on: - workflow_dispatch: - -env: - # One core: parallelism buys no speed here and costs memory the box does not have. - GOFLAGS: -p=1 - GOMAXPROCS: "2" - -jobs: - build: - runs-on: fedora - steps: - - uses: actions/checkout@v7 - - - uses: actions/setup-go@v6 - with: - # The module is the source of truth for the version, so it cannot drift. - go-version-file: go.mod - cache: true - - - name: FreeBSD build (amd64) - run: GOOS=freebsd GOARCH=amd64 go build ./... - - - name: FreeBSD build (arm64) - run: GOOS=freebsd GOARCH=arm64 go build ./... - - - name: FreeBSD build (riscv64) - run: GOOS=freebsd GOARCH=riscv64 go build ./... diff --git a/.gitea/workflows/release.yml b/.gitea/workflows/release.yml index ccb7df5..757601d 100644 --- a/.gitea/workflows/release.yml +++ b/.gitea/workflows/release.yml @@ -11,9 +11,12 @@ # (devel) even at its own /vX.Y.Z tag and this workflow's smoke test can never # pass for it. A Go repository is one module at the root. # -# The matrix carries the platforms the project ships: Linux on amd64, arm64, loong64 -# and riscv64. The FreeBSD port compiles in its own dispatched workflow and ships no -# binary. Nothing is installed: the fedora job image carries git, perl and node +# The matrix carries the platforms the project ships: Linux on amd64, arm64, +# loong64 and riscv64, and FreeBSD on amd64 and arm64. Go cross-compiles +# FreeBSD natively and the tree builds under CGO_ENABLED=0, which the +# suite's cross-build gates prove for the whole module; the ptrace suite +# itself still needs a FreeBSD machine, so nothing FreeBSD runs here. +# Nothing is installed: the fedora job image carries git, perl and node # (verified on the runner, 2026-10-04). # # Each job validates the tag for itself rather than passing a value between jobs, so @@ -46,6 +49,10 @@ jobs: goarch: loong64 - goos: linux goarch: riscv64 + - goos: freebsd + goarch: amd64 + - goos: freebsd + goarch: arm64 steps: - uses: actions/checkout@v7 diff --git a/.gitea/workflows/suite.yml b/.gitea/workflows/suite.yml index 9c39cff..e3cd273 100644 --- a/.gitea/workflows/suite.yml +++ b/.gitea/workflows/suite.yml @@ -1,9 +1,11 @@ # Suite, Go. Dispatched by hand, on development. Never a push gate. # -# The complete gate set minus race: the build, both static gates, the full suite with -# every short-layer skip unskipped, and the coverage floor. It is the pipeline form of -# the local `just test`, for the moments when the tree must be proven end to end and -# nobody is at the keyboard. +# The complete gate set minus race: the build, the FreeBSD cross-builds the +# port's compile proof needs (amd64 and arm64, the same two the release +# matrix ships), both static gates, the full suite with every short-layer +# skip unskipped, and the coverage floor. It is the pipeline form of the +# local `just test`, for the moments when the tree must be proven end to end +# and nobody is at the keyboard. # # Race never runs in CI. It roughly doubles the time and the memory on a box shared # with the forge, and the local `just gates` races the tree on the machine at the @@ -33,6 +35,14 @@ jobs: - name: Build run: go build ./... + - name: FreeBSD build (amd64) + # The port is pure Go: the cross-build needs no C toolchain and the + # release matrix ships the same two binaries. + run: CGO_ENABLED=0 GOOS=freebsd GOARCH=amd64 go build ./... + + - name: FreeBSD build (arm64) + run: CGO_ENABLED=0 GOOS=freebsd GOARCH=arm64 go build ./... + - name: Format run: | perl -e ' diff --git a/CHANGELOG.md b/CHANGELOG.md index d175f46..922b36f 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -7,17 +7,75 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ## [development] +## [0.36.0] - 2026-10-07 + ### Added +- **The SVE2 and SVE2.1 instruction sets on arm64.** Over 350 SVE + mnemonics across 537 encoded forms join the extension layer: the + narrowing two-to-one arithmetic family, the SVE2 cryptographic set + including ZADCLB, BFloat16 arithmetic, predicate counters and + reductions, the pairwise and quadword forms, multiple-structure loads + and stores (LD2/ST2 through LD4/ST4), shift by vector and the + shift-immediate scheme with its element-size encoding, CLASTA and + CLASTB, the last-active and compare predicate families, and the + vector-length arithmetic ADDVL, ADDPL and RDVL. Every form derives + from the toolchain's own encoder tables and is pinned against the + GOROOT corpus words. +- **The complete AVX512-FP16 set and AVX-VNNI-INT16 on amd64.** The + extension layer reaches 191 entries: the packed and scalar FMA + families (VFMADD, VFMSUB, VFMADDSUB and VFMSUBADD in the PH widths + with their SH mirrors), the complex multiply and complex FMA pairs + (VFMULC, VFCMULC, VFMADDC, VFCMADDC, packed and scalar, each + conjugating the source its prefix names), VMINMAXPH and VMINMAXSH + under their imm8 control, and the AVX-VNNI-INT16 dot products + (VPDPWSUD, VPDPWSUDS, VPDPWUSD, VPDPWUSDS) through a new VEX + encoding path beside the EVEX one. Every entry carries golden + vectors byte-checked against GNU as. +- **Extension-layer instructions assemble straight from `.s` source on + amd64.** The registered mnemonics dispatch from the assembly front + end with their operand grammar: `k0` through `k7` write masks with + merging and zeroing, `{1toN}` broadcast, embedded rounding and + `{sae}`, memory and scaled-index operands, and the imm8-control + forms, the VEX forms beside the EVEX ones. `gasm lint` surfaces the + layer's refusals as `extension-form` errors instead of letting a bad + shape reach the encoder, and every distinct mnemonic is pinned from + source text against the registry's bytes. +- **arm64 output byte-identical with the toolchain across the GOROOT + tree.** 31 of the 36 files the differential harness walks (90 + functions, 27408 bytes) now assemble without a differing byte: the + missing shapes are in (GETCALLERPC, the REM, REMW, UREM and UREMW + family, DWORD, the FCVTHS, FCVTSH, FCVTDH and FCVTHD conversions, TLS + local-exec loads as a single MOVZ with the TLS relocation), the + literal pool drains mid-function when a function's literals would + overflow the ±512 KiB load-literal range, and frame-relative + addresses and flag-setting logicals materialise the way + cmd/internal/obj/arm64 writes them. +- **riscv64 END and GETCALLERPC.** END is accepted anywhere and emits + nothing, and GETCALLERPC reads the return address the toolchain's + rewrite describes: a leaf function reads LR, a framed body reads the + prologue's save slot, with the compressed spellings where they fit. +- **The loong64 register-pair spelling.** `MULV R4:R5, R6` parses the + way the toolchain's pair sugar does: the colon splits the operand and + the halves swap, any instruction shape takes the spelling, and the + malformed forms fail with the toolchain's own wording. +- **Fuzz targets across the tool.** The assembler carries + per-architecture FuzzAssemble targets and the front end carries + targets over the lexer, parser, formatter and the extension + registry; the seed corpora run as ordinary tests, and `just fuzz` + drives the campaigns locally. Hardening the targets found caps GLOBL + and DATA at 64 MiB per symbol and 128 MiB per section, so a crafted + input cannot materialise unbounded memory. - **The FreeBSD port of the debugger.** `gasm debug` runs on FreeBSD on - amd64, arm64 and riscv64 with the same interactive surface as on Linux: + amd64 and arm64 with the same interactive surface as on Linux: breakpoints, hardware watchpoints (x86 debug registers, the arm64 debug register file), single-stepping, register and memory access, all behind the kernel's own ptrace requests, with tracee memory through `PT_IO` and stop reports through `PT_LWPINFO`. The JIT substrate maps executable memory through `golang.org/x/sys/unix`, so `verify` builds on FreeBSD - too. The pipeline compile-gates all three architectures; live - validation awaits a FreeBSD machine. + too. The dispatched suite compile-gates both architectures, the + release pipeline ships the two binaries, and live validation awaits a + FreeBSD machine. - **Workspace-wide navigation in the language server.** `gasm lsp` indexes the `.s` files under the workspace root beyond the documents the editor has open, so go-to-definition, find references and workspace symbol search @@ -75,6 +133,23 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ### Changed +- **TEXT flag operands behave like the toolchain's.** A numeric or + parenthesised flag list (`TEXT ·f(SB), 4, $4096-0`, + `(NOSPLIT|NOFRAME)`) now suppresses the prologue and stack guard the + names suppress, where the numbers were silently ignored and the + prologue was emitted anyway; unknown flag names are rejected with the + toolchain's wording instead of passing unnoticed; and the toolchain's + own TEXT diagnostics fire, `ABIInternal requires NOSPLIT` and + `NOFRAME functions must have a frame size of 0`. +- **The disassembler names what it used to print raw.** The 195 amd64 + encodings that decoded to anonymous renders now decode to their + instructions and round-trip, the BMI1 and BMI2 VEX set, RDSEED, RDPID, + the WAITPKG family, ENDBR, CLDEMOTE, UD1 and RORX among them, with the + arm64 HVC, SMC and SB forms and 74 loong64 rendering corrections + beside. +- **Variadic macro parameters are rejected like the toolchain rejects + them.** A macro declaration whose parameter list the toolchain + refuses no longer parses into a broken definition. - **The corpus audit assembles like the build.** A file's `//go:build` constraint decides which target architectures attempt it: cpu_x86.s is an x86 build alone, and the msan and goexperiment.runtimesecret trees @@ -94,6 +169,37 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ### Fixed +- **riscv64 memory offsets beyond the 12-bit immediate no longer + truncate silently.** `LD 4096(X6), X5` encoded as `LD X5, 0(X6)` and + read the wrong address; the expansion the toolchain performs (the + upper bits into its temporary register, then the access) is emitted + now, byte-identically, and offsets the toolchain refuses are refused + with its wording. +- **riscv64 operand acceptance matches the toolchain's.** 398 operand + shapes the toolchain rejects assembled without complaint (a float + register where an integer one is required, out-of-range CSRs, vector + shapes against the wrong bank, MOV family widths), and 488 more + produced the wrong diagnostic; all now fail with the toolchain's + texts, and a 1030-case error-parity catalogue pins them against the + toolchain's own negative files. +- **The debugger survives its failure paths.** Killing a debuggee that + had already died or sat ptrace-stopped could block the session + forever; Kill now wakes the tracee, signals and reaps it without + waiting on a parked child, so every launch-failure path returns + instead of hanging the tool. +- **Encoder defects found by fuzzing and the differential audits.** A + GLOBL with a negative size panicked the assembler and a small crafted + input materialised 3.92 GiB of DATA; certain SVE VTBL operands + panicked; an over-strict guard refused BIC into RSP where the + toolchain assembles it; byte-register operands now select the 8-bit + forms the toolchain picks (`XADDL DL, DL` and its companions); and + the loong64 shift compositions and indexed-offset forms that dropped + fields re-encode faithfully. +- **Formatter edges around labels and macros.** A stacked label that + names a macro stays attached to it, macro content behind a label + stays on its line, a selector-folded label resolves to its macro, an + empty TEXT body formats, and a file ending in a comment no longer + loses the comment. - **Rename edits land in their own documents.** A rename collected the ranges of every reference across the open documents but applied them all to the document that started it, so renaming a symbol used in a second diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 29badd9..85cfecf 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -118,9 +118,8 @@ Workflows live in `.gitea/workflows/` and run on the project's own runners: | Workflow | Trigger | What it does | |---|---|---| | Test | push or pull request to `development` | the format check, `go vet`, and the short layer of the test suite with the coverage profile and the 80 % floor, inside the two-minute budget | -| Suite | dispatched by hand | the complete gate set minus race: the build, both static gates, the full suite and the coverage floor | -| FreeBSD build | dispatched by hand | the three FreeBSD compile gates for amd64, arm64 and riscv64 | -| Release | a `v*` tag | the matrix build, the version smoke test and the release with its assets; no gate runs at the tag | +| Suite | dispatched by hand | the complete gate set minus race: the build, the two FreeBSD cross-build gates (amd64, arm64), both static gates, the full suite and the coverage floor | +| Release | a `v*` tag | the matrix build (Linux on four architectures, FreeBSD on two), the version smoke test and the release with its assets; no gate runs at the tag | The local equivalent is `just gates`, which is the same set plus the race detector. Race never runs in CI, on a push or a tag: it would double the time and the memory a shared diff --git a/README.md b/README.md index 7628be4..b69dfe0 100644 --- a/README.md +++ b/README.md @@ -95,7 +95,7 @@ to give that syntax the tooling it deserves. breakpoints (optionally conditional), hardware watchpoints, register and memory inspection, and headless script runs that report instruction and label coverage; it runs on Linux (all four architectures) and FreeBSD - (amd64, arm64, riscv64). + (amd64, arm64). - **Language server.** `gasm lsp` serves completion, hover, document symbols, push and pull diagnostics, semantic-token highlighting, go-to-definition, find references, rename, formatting, inlay hints, code actions, signature @@ -151,7 +151,7 @@ actually been executed. |---|---|---| | Encoding: byte-for-byte against `go tool asm` | native hardware | native hardware (the toolchain cross-assembles any GOARCH on any host) | | Execution: JIT calls, ABI checks, differential fuzzing | native hardware | qemu-user emulation | -| Debugger: ptrace tracing, breakpoints, watchpoints, coverage | native hardware | emulation cannot run ptrace; the layer compiles and its architecture-neutral units run under `go test ./...`, nothing more. FreeBSD (amd64, arm64, riscv64) is in the same position: the port compiles behind the cross-build gate and its integration test is ready, but no FreeBSD machine has executed it | +| Debugger: ptrace tracing, breakpoints, watchpoints, coverage | native hardware | emulation cannot run ptrace; the layer compiles and its architecture-neutral units run under `go test ./...`, nothing more. FreeBSD (amd64, arm64) is in the same position: the port compiles behind the cross-build gate and its integration test is ready, but no FreeBSD machine has executed it | Consequences, stated plainly. An emulator is a model of a CPU, not the CPU: instruction semantics are implemented in software and can differ @@ -222,18 +222,18 @@ The plan, in the order it is being worked: toolchain itself does not support; through ELF, Plan 9 assembly becomes usable outside Go entirely. - **Platforms: Linux and FreeBSD.** Linux is supported today on all four - architectures and is where the binary builds. FreeBSD follows on amd64, - arm64 and riscv64: the JIT's executable-memory mapping and the ptrace - debugger layer are ported (the debugger's live validation awaits a - FreeBSD machine, as the validation status states). Other unix systems - may follow those two. + architectures and is where the binary builds. FreeBSD follows on amd64 + and arm64: the JIT's executable-memory mapping and the ptrace debugger + layer are ported, the release carries the two FreeBSD binaries, and the + debugger's live validation awaits a FreeBSD machine (the validation + status states it). Other unix systems may follow those two. - **Four architectures, no more.** amd64, arm64, riscv64 and loong64. No others are planned. ## Install -Prebuilt binaries for linux/amd64, linux/arm64, linux/riscv64 and -linux/loong64 are on the +Prebuilt binaries for linux/amd64, linux/arm64, linux/riscv64, +linux/loong64, freebsd/amd64 and freebsd/arm64 are on the [releases page](https://sourcedock.dev/petrbalvin/gasm-sdk/releases). From source (Go 1.27.1): diff --git a/docs/CLI.md b/docs/CLI.md index cc55edc..aeb61e9 100644 --- a/docs/CLI.md +++ b/docs/CLI.md @@ -134,7 +134,7 @@ Rules: `unknown-instruction`, `operand-count`, `undefined-label`, `nonportable-register-name`, `unencodable-instruction`, `reserved-register-write`, `missing-argsize`, `noframe-frame-size`, `unnamed-fp-reference`, `hardware-sp-addressing`, `vex-sse-mixing`, -`unnamed-result`, `data-width`, `data-value-overflow`, +`unnamed-result`, `extension-form`, `data-width`, `data-value-overflow`, `data-string-width`, `data-without-globl` and `data-exceeds-globl`. ```sh @@ -290,8 +290,8 @@ Usage: gasm debug --func The debugger re-executes the binary it is running as (`os.Executable()`) for the traced child, so the child is the same `gasm`, whether it is installed on `$PATH` or run with `go run ./cmd/gasm`; nothing has to be installed first. Requires -Linux or FreeBSD (ptrace): all four architectures on Linux, amd64, arm64 and -riscv64 on FreeBSD. +Linux or FreeBSD (ptrace): all four architectures on Linux, amd64 and +arm64 on FreeBSD. REPL commands: diff --git a/docs/DEVELOPMENT.md b/docs/DEVELOPMENT.md index 02574e0..a89b43b 100644 --- a/docs/DEVELOPMENT.md +++ b/docs/DEVELOPMENT.md @@ -145,9 +145,8 @@ Workflows live in `.gitea/workflows/` and run on the project's own runners: | Workflow | Trigger | What it does | |---|---|---| | Test | push or pull request to `development` | the format check, `go vet`, the short layer of the suite with the coverage profile and the 80 % floor, inside the two-minute budget | -| Suite | dispatched by hand | the complete gate set minus race: the build, the format check, `go vet`, `go fix -diff`, the full suite and the coverage floor | -| FreeBSD build | dispatched by hand | the three FreeBSD compile gates: `GOOS=freebsd` for amd64, arm64 and riscv64 | -| Release | a `v*` tag | the matrix build, the version smoke test, and the release with its assets; no gate runs at the tag | +| Suite | dispatched by hand | the complete gate set minus race: the build, the two FreeBSD cross-build gates (`GOOS=freebsd` for amd64 and arm64), `go vet`, `go fix -diff`, the full suite and the coverage floor | +| Release | a `v*` tag | the matrix build (Linux on four architectures, FreeBSD on two), the version smoke test, and the release with its assets; no gate runs at the tag | The pipelines are written by hand rather than through `just`, but they enforce the same gates: the push path carries the affordable subset and the heavy tests diff --git a/docs/asm/AMD64.md b/docs/asm/AMD64.md index 87c996e..29df6de 100644 --- a/docs/asm/AMD64.md +++ b/docs/asm/AMD64.md @@ -100,6 +100,15 @@ encoder backlog that `gasm audit-instructions` measures. The families: whose register list rides the inverted V′VVV field. Mixing VEX and legacy SSE in one loop pays the AVX-SSE transition penalty on every switch: keep a loop in one dialect. +- **The extension layer.** The families the toolchain's own table carries + late or not at all assemble through the extension mechanism: BF16, + VP2INTERSECT, the complete AVX512-FP16 set (the packed and scalar FMA + families, the complex multiply and complex FMA pairs, VMINMAXPH and + VMINMAXSH), and the AVX-VNNI-INT16 dot products through a VEX path. The + layer takes `k0` to `k7` write masks with merging and zeroing, the + `{1toN}` broadcast, the embedded rounding and `{sae}` decorations and the + imm8 controls, and its refusals surface as `extension-form` lint errors + with the reason the encoder would give. - **Cryptographic and counting extensions.** AES-NI, SHA-1 and SHA-256, PCLMULQDQ, GFNI. - **System.** `CPUID`, `RDTSC`, `SYSCALL`, the fences, `LDMXCSR` and diff --git a/docs/asm/ARM64.md b/docs/asm/ARM64.md index eeecdc7..a4f57ee 100644 --- a/docs/asm/ARM64.md +++ b/docs/asm/ARM64.md @@ -12,8 +12,8 @@ toolchain's. The complete mnemonic inventory lives in the generated appendix the stack pointer. There is no R31: thirty-one names and ZR. - Floating-point and SIMD share one file written `Vn`; where an instruction is scalar floating point the operand may be written `Fn` (`F0` to `F31`). -- SVE register names (`Z0` to `Z31`, `P0` to `P15`) exist in the assembler's - tables. +- SVE registers `Z0` to `Z31` and predicates `P0` to `P15`; the SVE2 and + SVE2.1 families encode, see SVE below. - Roles the convention fixes: `RSP` is the stack pointer, `R29` the frame pointer, `R30` the link register, `R26` the closure context and `R27` the assembler's scratch register. The goroutine pointer lives in `R28` and is @@ -106,6 +106,33 @@ scalar floating-point instructions being the exceptions. Operands carry an arrangement suffix, `V5.H8`, and structure loads and stores use bracket lists, `[V21.B16]`, with element selection as `V9.S[1]`. +## SVE + +The scalable vector extension encodes through the `Z0` to `Z31` vector +registers and the `P0` to `P15` predicates, with arrangement suffixes on +the `Z` registers and the merging and zeroing qualifiers the toolchain's +SVE2 and SVE2.1 families carry: the narrowing two-to-one arithmetic, the +cryptographic set including ZADCLB, BFloat16 arithmetic, the predicate +counters and reductions, the pairwise and quadword forms, the +multiple-structure loads and stores, shift by vector and the +shift-immediate scheme, CLASTA and CLASTB, the compare and last-active +predicate families, and the vector-length arithmetic `ADDVL`, `ADDPL` and +`RDVL`, whose immediates count vectors or predicates rather than bytes. +The gather and scatter loads take their addresses through the five +addressing modes the toolchain defines. The source grammar is the +toolchain's own, and the byte output is pinned against it corpus-wide. + +## Synthesised forms + +- `GETCALLERPC` reads the return address the frame state describes: a leaf + function reads `R30`, a framed body reads the prologue's save slot. +- `REM`, `REMW`, `UREM` and `UREMW` synthesise a remainder from `SDIV` or + `UDIV` and the `MSUB` tail, with `RSP` refused as a destination. +- `DWORD $imm` lays eight little-endian bytes per immediate. +- `MOVD tls_g(SB), Rd` materialises a TLS local-exec load as a single + `MOVZ` carrying `R_ARM64_TLS_LE`, keyed off the file's own + `GLOBL ... TLSBSS` declaration. + ## Alignment `PCALIGN $n` pads to a power-of-two boundary between 8 and 2048 and also diff --git a/docs/asm/DIRECTIVES.md b/docs/asm/DIRECTIVES.md index fbc0eb2..b9534c6 100644 --- a/docs/asm/DIRECTIVES.md +++ b/docs/asm/DIRECTIVES.md @@ -23,11 +23,16 @@ TEXT symbol(SB), [flags,] $framesize[-argsize] - The symbol is an `·Name(SB)` reference into the current package, or a fully qualified name. -- The optional flag argument is a constant expression, normally an OR of the - names from `textflag.h`, the table below. Without `#include "textflag.h"` - the names are not macros and the assembler reports the misleading error - `illegal or missing addressing mode for symbol NOSPLIT`: include the - header first. +- The optional flag argument is a constant expression: an OR of the names + from `textflag.h`, the table below, a bare number (`4`), or a + parenthesised combination (`(NOSPLIT|NOFRAME)`). Every spelling + suppresses exactly what the names suppress. Without + `#include "textflag.h"` the names are not macros and the assembler + reports the misleading error `illegal or missing addressing mode for + symbol NOSPLIT`: include the header first. A name outside the table is + rejected with the toolchain's wording, and two of the toolchain's own + TEXT checks fire here: `ABIInternal requires NOSPLIT`, and + `NOFRAME functions must have a frame size of 0` for a positive frame. - `$framesize-argsize` is two constants, not a subtraction: the local frame size in bytes, and the caller's argument area in bytes. The argument size may be omitted entirely, `$16`, which marks the argument size unknown @@ -59,7 +64,7 @@ Values from `textflag.h`, in agreement with `cmd/internal/obj/textflag.go`: | WRAPPER | 32 | TEXT | a wrapper; must not disable `recover` | | NEEDCTXT | 64 | TEXT | a closure consuming the context register | | TLSBSS | 256 | data | a thread local word in BSS | -| NOFRAME | 512 | TEXT | no frame setup; the zero-frame spelling (the toolchain accepts a positive frame beside it and still allocates the frame, and the arm64 BSD syscall stubs pair it with $-8 deliberately) | +| NOFRAME | 512 | TEXT | no frame setup; the frame size must be zero or negative (the toolchain rejects a positive frame beside it, and the arm64 BSD syscall stubs pair it with $-8 deliberately) | | REFLECTMETHOD | 1024 | TEXT | the function calls `reflect.Type.Method` or `MethodByName` | | TOPFRAME | 2048 | TEXT | the outermost frame; unwinders stop here | | ABIWRAPPER | 4096 | TEXT | an ABI transition wrapper | diff --git a/docs/asm/LOONG64.md b/docs/asm/LOONG64.md index f157417..50be5ed 100644 --- a/docs/asm/LOONG64.md +++ b/docs/asm/LOONG64.md @@ -41,6 +41,12 @@ R13` is `add.d R13, R12, R11`, and the two-operand form - Jump and branch instructions keep the GNU order: `BEQ R0, R4, label1`. - The bitfield family is `BSTRINSW`, `BSTRINSV`, `BSTRPICKW`, `BSTRPICKV` `$, , $, `. +- **The register-pair spelling** `R4:R5` names two registers in one + operand, the way the toolchain's pair syntax does: the colon splits the + operand and the halves swap, so `MULV R4:R5, R6` encodes exactly as + `MULV R5, R4, R6`. Any instruction shape accepts the spelling, decided + afterwards by ordinary operand matching, and a malformed pair fails with + the toolchain's wording. ## Addressing diff --git a/docs/asm/RISCV64.md b/docs/asm/RISCV64.md index 95f8bf9..21a2162 100644 --- a/docs/asm/RISCV64.md +++ b/docs/asm/RISCV64.md @@ -55,6 +55,10 @@ the FCSR. a literal pool in the binary. - A 32-bit constant is accepted by `ADDI`, `ANDI`, `ORI` and `XORI`, and the assembler synthesises values that exceed the 12-bit encoding window. +- A memory offset beyond the 12-bit immediate expands the way the + toolchain's loads and stores do: the upper bits go into the assembler's + temporary register and the access runs against it, byte-identically, and + an offset the toolchain refuses is refused with its wording. - `MOVF` and `MOVD` materialise floating-point constants, encoding them as `FLW` and `FLD` from a pool location unless the constant is exactly 0.0. @@ -85,6 +89,15 @@ of the source, and register choice influences how much compresses. Hand-writing compressed instructions in source is accepted but discouraged. The debug flag `compressinstructions=0` turns the automatic conversion off. +## Pseudo-operations + +- `END` is accepted anywhere and emits nothing, the function-end marker + the toolchain's front end also swallows. +- `GETCALLERPC` reads the return address the frame state describes: a leaf + function reads `X1`, a framed body reads the prologue's save slot at + `0(SP)`, and both spellings take the compressed encoding where the + register choice allows it. + ## Vector extension `VSETVLI` writes its vtype components in uppercase with the destination