chore: prepare release v0.36.0
Test / test (push) Successful in 32s
Release / build (amd64, freebsd) (push) Successful in 52s
Release / build (amd64, linux) (push) Successful in 13s
Release / build (arm64, freebsd) (push) Successful in 51s
Release / build (arm64, linux) (push) Successful in 29s
Release / build (loong64, linux) (push) Successful in 29s
Release / build (riscv64, linux) (push) Successful in 31s
Release / release (push) Successful in 10s
Test / test (push) Successful in 32s
Release / build (amd64, freebsd) (push) Successful in 52s
Release / build (amd64, linux) (push) Successful in 13s
Release / build (arm64, freebsd) (push) Successful in 51s
Release / build (arm64, linux) (push) Successful in 29s
Release / build (loong64, linux) (push) Successful in 29s
Release / build (riscv64, linux) (push) Successful in 31s
Release / release (push) Successful in 10s
Assisted-by: GLM 5.3 Flash
This commit is contained in:
1 parent
e5df0a9d84
commit
a77dd12d2d
13 files changed
+217
-77
No files matched your search
@@ -1,41 +0,0 @@
|
||||
# FreeBSD compile gates. Dispatched by hand, never on a push.
|
||||
#
|
||||
# The debugger's ptrace surface and the JIT substrate are the two
|
||||
# FreeBSD-portable layers the tree carries; the forge has no FreeBSD runner,
|
||||
# so they can only be compile-gated, and three foreign-GOOS builds of the
|
||||
# whole module are minutes of one-core work the push pipeline's budget cannot
|
||||
# carry. The push pipeline stays fast and light; this workflow is the
|
||||
# deliberate run, before a release or after touching the ported layers.
|
||||
# Running the ptrace suite itself needs real FreeBSD hardware.
|
||||
#
|
||||
# A dispatched workflow takes no concurrency block: it is one deliberate run.
|
||||
name: FreeBSD build
|
||||
|
||||
on:
|
||||
workflow_dispatch:
|
||||
|
||||
env:
|
||||
# One core: parallelism buys no speed here and costs memory the box does not have.
|
||||
GOFLAGS: -p=1
|
||||
GOMAXPROCS: "2"
|
||||
|
||||
jobs:
|
||||
build:
|
||||
runs-on: fedora
|
||||
steps:
|
||||
- uses: actions/checkout@v7
|
||||
|
||||
- uses: actions/setup-go@v6
|
||||
with:
|
||||
# The module is the source of truth for the version, so it cannot drift.
|
||||
go-version-file: go.mod
|
||||
cache: true
|
||||
|
||||
- name: FreeBSD build (amd64)
|
||||
run: GOOS=freebsd GOARCH=amd64 go build ./...
|
||||
|
||||
- name: FreeBSD build (arm64)
|
||||
run: GOOS=freebsd GOARCH=arm64 go build ./...
|
||||
|
||||
- name: FreeBSD build (riscv64)
|
||||
run: GOOS=freebsd GOARCH=riscv64 go build ./...
|
||||
@@ -11,9 +11,12 @@
|
||||
# (devel) even at its own <module>/vX.Y.Z tag and this workflow's smoke test can never
|
||||
# pass for it. A Go repository is one module at the root.
|
||||
#
|
||||
# The matrix carries the platforms the project ships: Linux on amd64, arm64, loong64
|
||||
# and riscv64. The FreeBSD port compiles in its own dispatched workflow and ships no
|
||||
# binary. Nothing is installed: the fedora job image carries git, perl and node
|
||||
# The matrix carries the platforms the project ships: Linux on amd64, arm64,
|
||||
# loong64 and riscv64, and FreeBSD on amd64 and arm64. Go cross-compiles
|
||||
# FreeBSD natively and the tree builds under CGO_ENABLED=0, which the
|
||||
# suite's cross-build gates prove for the whole module; the ptrace suite
|
||||
# itself still needs a FreeBSD machine, so nothing FreeBSD runs here.
|
||||
# Nothing is installed: the fedora job image carries git, perl and node
|
||||
# (verified on the runner, 2026-10-04).
|
||||
#
|
||||
# Each job validates the tag for itself rather than passing a value between jobs, so
|
||||
@@ -46,6 +49,10 @@ jobs:
|
||||
goarch: loong64
|
||||
- goos: linux
|
||||
goarch: riscv64
|
||||
- goos: freebsd
|
||||
goarch: amd64
|
||||
- goos: freebsd
|
||||
goarch: arm64
|
||||
steps:
|
||||
- uses: actions/checkout@v7
|
||||
|
||||
|
||||
@@ -1,9 +1,11 @@
|
||||
# Suite, Go. Dispatched by hand, on development. Never a push gate.
|
||||
#
|
||||
# The complete gate set minus race: the build, both static gates, the full suite with
|
||||
# every short-layer skip unskipped, and the coverage floor. It is the pipeline form of
|
||||
# the local `just test`, for the moments when the tree must be proven end to end and
|
||||
# nobody is at the keyboard.
|
||||
# The complete gate set minus race: the build, the FreeBSD cross-builds the
|
||||
# port's compile proof needs (amd64 and arm64, the same two the release
|
||||
# matrix ships), both static gates, the full suite with every short-layer
|
||||
# skip unskipped, and the coverage floor. It is the pipeline form of the
|
||||
# local `just test`, for the moments when the tree must be proven end to end
|
||||
# and nobody is at the keyboard.
|
||||
#
|
||||
# Race never runs in CI. It roughly doubles the time and the memory on a box shared
|
||||
# with the forge, and the local `just gates` races the tree on the machine at the
|
||||
@@ -33,6 +35,14 @@ jobs:
|
||||
- name: Build
|
||||
run: go build ./...
|
||||
|
||||
- name: FreeBSD build (amd64)
|
||||
# The port is pure Go: the cross-build needs no C toolchain and the
|
||||
# release matrix ships the same two binaries.
|
||||
run: CGO_ENABLED=0 GOOS=freebsd GOARCH=amd64 go build ./...
|
||||
|
||||
- name: FreeBSD build (arm64)
|
||||
run: CGO_ENABLED=0 GOOS=freebsd GOARCH=arm64 go build ./...
|
||||
|
||||
- name: Format
|
||||
run: |
|
||||
perl -e '
|
||||
|
||||
+109
-3
@@ -7,17 +7,75 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
||||
|
||||
## [development]
|
||||
|
||||
## [0.36.0] - 2026-10-07
|
||||
|
||||
### Added
|
||||
|
||||
- **The SVE2 and SVE2.1 instruction sets on arm64.** Over 350 SVE
|
||||
mnemonics across 537 encoded forms join the extension layer: the
|
||||
narrowing two-to-one arithmetic family, the SVE2 cryptographic set
|
||||
including ZADCLB, BFloat16 arithmetic, predicate counters and
|
||||
reductions, the pairwise and quadword forms, multiple-structure loads
|
||||
and stores (LD2/ST2 through LD4/ST4), shift by vector and the
|
||||
shift-immediate scheme with its element-size encoding, CLASTA and
|
||||
CLASTB, the last-active and compare predicate families, and the
|
||||
vector-length arithmetic ADDVL, ADDPL and RDVL. Every form derives
|
||||
from the toolchain's own encoder tables and is pinned against the
|
||||
GOROOT corpus words.
|
||||
- **The complete AVX512-FP16 set and AVX-VNNI-INT16 on amd64.** The
|
||||
extension layer reaches 191 entries: the packed and scalar FMA
|
||||
families (VFMADD, VFMSUB, VFMADDSUB and VFMSUBADD in the PH widths
|
||||
with their SH mirrors), the complex multiply and complex FMA pairs
|
||||
(VFMULC, VFCMULC, VFMADDC, VFCMADDC, packed and scalar, each
|
||||
conjugating the source its prefix names), VMINMAXPH and VMINMAXSH
|
||||
under their imm8 control, and the AVX-VNNI-INT16 dot products
|
||||
(VPDPWSUD, VPDPWSUDS, VPDPWUSD, VPDPWUSDS) through a new VEX
|
||||
encoding path beside the EVEX one. Every entry carries golden
|
||||
vectors byte-checked against GNU as.
|
||||
- **Extension-layer instructions assemble straight from `.s` source on
|
||||
amd64.** The registered mnemonics dispatch from the assembly front
|
||||
end with their operand grammar: `k0` through `k7` write masks with
|
||||
merging and zeroing, `{1toN}` broadcast, embedded rounding and
|
||||
`{sae}`, memory and scaled-index operands, and the imm8-control
|
||||
forms, the VEX forms beside the EVEX ones. `gasm lint` surfaces the
|
||||
layer's refusals as `extension-form` errors instead of letting a bad
|
||||
shape reach the encoder, and every distinct mnemonic is pinned from
|
||||
source text against the registry's bytes.
|
||||
- **arm64 output byte-identical with the toolchain across the GOROOT
|
||||
tree.** 31 of the 36 files the differential harness walks (90
|
||||
functions, 27408 bytes) now assemble without a differing byte: the
|
||||
missing shapes are in (GETCALLERPC, the REM, REMW, UREM and UREMW
|
||||
family, DWORD, the FCVTHS, FCVTSH, FCVTDH and FCVTHD conversions, TLS
|
||||
local-exec loads as a single MOVZ with the TLS relocation), the
|
||||
literal pool drains mid-function when a function's literals would
|
||||
overflow the ±512 KiB load-literal range, and frame-relative
|
||||
addresses and flag-setting logicals materialise the way
|
||||
cmd/internal/obj/arm64 writes them.
|
||||
- **riscv64 END and GETCALLERPC.** END is accepted anywhere and emits
|
||||
nothing, and GETCALLERPC reads the return address the toolchain's
|
||||
rewrite describes: a leaf function reads LR, a framed body reads the
|
||||
prologue's save slot, with the compressed spellings where they fit.
|
||||
- **The loong64 register-pair spelling.** `MULV R4:R5, R6` parses the
|
||||
way the toolchain's pair sugar does: the colon splits the operand and
|
||||
the halves swap, any instruction shape takes the spelling, and the
|
||||
malformed forms fail with the toolchain's own wording.
|
||||
- **Fuzz targets across the tool.** The assembler carries
|
||||
per-architecture FuzzAssemble targets and the front end carries
|
||||
targets over the lexer, parser, formatter and the extension
|
||||
registry; the seed corpora run as ordinary tests, and `just fuzz`
|
||||
drives the campaigns locally. Hardening the targets found caps GLOBL
|
||||
and DATA at 64 MiB per symbol and 128 MiB per section, so a crafted
|
||||
input cannot materialise unbounded memory.
|
||||
- **The FreeBSD port of the debugger.** `gasm debug` runs on FreeBSD on
|
||||
amd64, arm64 and riscv64 with the same interactive surface as on Linux:
|
||||
amd64 and arm64 with the same interactive surface as on Linux:
|
||||
breakpoints, hardware watchpoints (x86 debug registers, the arm64 debug
|
||||
register file), single-stepping, register and memory access, all behind
|
||||
the kernel's own ptrace requests, with tracee memory through `PT_IO` and
|
||||
stop reports through `PT_LWPINFO`. The JIT substrate maps executable
|
||||
memory through `golang.org/x/sys/unix`, so `verify` builds on FreeBSD
|
||||
too. The pipeline compile-gates all three architectures; live
|
||||
validation awaits a FreeBSD machine.
|
||||
too. The dispatched suite compile-gates both architectures, the
|
||||
release pipeline ships the two binaries, and live validation awaits a
|
||||
FreeBSD machine.
|
||||
- **Workspace-wide navigation in the language server.** `gasm lsp` indexes
|
||||
the `.s` files under the workspace root beyond the documents the editor
|
||||
has open, so go-to-definition, find references and workspace symbol search
|
||||
@@ -75,6 +133,23 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
||||
|
||||
### Changed
|
||||
|
||||
- **TEXT flag operands behave like the toolchain's.** A numeric or
|
||||
parenthesised flag list (`TEXT ·f(SB), 4, $4096-0`,
|
||||
`(NOSPLIT|NOFRAME)`) now suppresses the prologue and stack guard the
|
||||
names suppress, where the numbers were silently ignored and the
|
||||
prologue was emitted anyway; unknown flag names are rejected with the
|
||||
toolchain's wording instead of passing unnoticed; and the toolchain's
|
||||
own TEXT diagnostics fire, `ABIInternal requires NOSPLIT` and
|
||||
`NOFRAME functions must have a frame size of 0`.
|
||||
- **The disassembler names what it used to print raw.** The 195 amd64
|
||||
encodings that decoded to anonymous renders now decode to their
|
||||
instructions and round-trip, the BMI1 and BMI2 VEX set, RDSEED, RDPID,
|
||||
the WAITPKG family, ENDBR, CLDEMOTE, UD1 and RORX among them, with the
|
||||
arm64 HVC, SMC and SB forms and 74 loong64 rendering corrections
|
||||
beside.
|
||||
- **Variadic macro parameters are rejected like the toolchain rejects
|
||||
them.** A macro declaration whose parameter list the toolchain
|
||||
refuses no longer parses into a broken definition.
|
||||
- **The corpus audit assembles like the build.** A file's `//go:build`
|
||||
constraint decides which target architectures attempt it: cpu_x86.s is
|
||||
an x86 build alone, and the msan and goexperiment.runtimesecret trees
|
||||
@@ -94,6 +169,37 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
||||
|
||||
### Fixed
|
||||
|
||||
- **riscv64 memory offsets beyond the 12-bit immediate no longer
|
||||
truncate silently.** `LD 4096(X6), X5` encoded as `LD X5, 0(X6)` and
|
||||
read the wrong address; the expansion the toolchain performs (the
|
||||
upper bits into its temporary register, then the access) is emitted
|
||||
now, byte-identically, and offsets the toolchain refuses are refused
|
||||
with its wording.
|
||||
- **riscv64 operand acceptance matches the toolchain's.** 398 operand
|
||||
shapes the toolchain rejects assembled without complaint (a float
|
||||
register where an integer one is required, out-of-range CSRs, vector
|
||||
shapes against the wrong bank, MOV family widths), and 488 more
|
||||
produced the wrong diagnostic; all now fail with the toolchain's
|
||||
texts, and a 1030-case error-parity catalogue pins them against the
|
||||
toolchain's own negative files.
|
||||
- **The debugger survives its failure paths.** Killing a debuggee that
|
||||
had already died or sat ptrace-stopped could block the session
|
||||
forever; Kill now wakes the tracee, signals and reaps it without
|
||||
waiting on a parked child, so every launch-failure path returns
|
||||
instead of hanging the tool.
|
||||
- **Encoder defects found by fuzzing and the differential audits.** A
|
||||
GLOBL with a negative size panicked the assembler and a small crafted
|
||||
input materialised 3.92 GiB of DATA; certain SVE VTBL operands
|
||||
panicked; an over-strict guard refused BIC into RSP where the
|
||||
toolchain assembles it; byte-register operands now select the 8-bit
|
||||
forms the toolchain picks (`XADDL DL, DL` and its companions); and
|
||||
the loong64 shift compositions and indexed-offset forms that dropped
|
||||
fields re-encode faithfully.
|
||||
- **Formatter edges around labels and macros.** A stacked label that
|
||||
names a macro stays attached to it, macro content behind a label
|
||||
stays on its line, a selector-folded label resolves to its macro, an
|
||||
empty TEXT body formats, and a file ending in a comment no longer
|
||||
loses the comment.
|
||||
- **Rename edits land in their own documents.** A rename collected the
|
||||
ranges of every reference across the open documents but applied them all
|
||||
to the document that started it, so renaming a symbol used in a second
|
||||
|
||||
+2
-3
@@ -118,9 +118,8 @@ Workflows live in `.gitea/workflows/` and run on the project's own runners:
|
||||
| Workflow | Trigger | What it does |
|
||||
|---|---|---|
|
||||
| Test | push or pull request to `development` | the format check, `go vet`, and the short layer of the test suite with the coverage profile and the 80 % floor, inside the two-minute budget |
|
||||
| Suite | dispatched by hand | the complete gate set minus race: the build, both static gates, the full suite and the coverage floor |
|
||||
| FreeBSD build | dispatched by hand | the three FreeBSD compile gates for amd64, arm64 and riscv64 |
|
||||
| Release | a `v*` tag | the matrix build, the version smoke test and the release with its assets; no gate runs at the tag |
|
||||
| Suite | dispatched by hand | the complete gate set minus race: the build, the two FreeBSD cross-build gates (amd64, arm64), both static gates, the full suite and the coverage floor |
|
||||
| Release | a `v*` tag | the matrix build (Linux on four architectures, FreeBSD on two), the version smoke test and the release with its assets; no gate runs at the tag |
|
||||
|
||||
The local equivalent is `just gates`, which is the same set plus the race detector. Race
|
||||
never runs in CI, on a push or a tag: it would double the time and the memory a shared
|
||||
|
||||
@@ -95,7 +95,7 @@ to give that syntax the tooling it deserves.
|
||||
breakpoints (optionally conditional), hardware watchpoints, register and
|
||||
memory inspection, and headless script runs that report instruction and
|
||||
label coverage; it runs on Linux (all four architectures) and FreeBSD
|
||||
(amd64, arm64, riscv64).
|
||||
(amd64, arm64).
|
||||
- **Language server.** `gasm lsp` serves completion, hover, document symbols,
|
||||
push and pull diagnostics, semantic-token highlighting, go-to-definition,
|
||||
find references, rename, formatting, inlay hints, code actions, signature
|
||||
@@ -151,7 +151,7 @@ actually been executed.
|
||||
|---|---|---|
|
||||
| Encoding: byte-for-byte against `go tool asm` | native hardware | native hardware (the toolchain cross-assembles any GOARCH on any host) |
|
||||
| Execution: JIT calls, ABI checks, differential fuzzing | native hardware | qemu-user emulation |
|
||||
| Debugger: ptrace tracing, breakpoints, watchpoints, coverage | native hardware | emulation cannot run ptrace; the layer compiles and its architecture-neutral units run under `go test ./...`, nothing more. FreeBSD (amd64, arm64, riscv64) is in the same position: the port compiles behind the cross-build gate and its integration test is ready, but no FreeBSD machine has executed it |
|
||||
| Debugger: ptrace tracing, breakpoints, watchpoints, coverage | native hardware | emulation cannot run ptrace; the layer compiles and its architecture-neutral units run under `go test ./...`, nothing more. FreeBSD (amd64, arm64) is in the same position: the port compiles behind the cross-build gate and its integration test is ready, but no FreeBSD machine has executed it |
|
||||
|
||||
Consequences, stated plainly. An emulator is a model of a CPU, not the
|
||||
CPU: instruction semantics are implemented in software and can differ
|
||||
@@ -222,18 +222,18 @@ The plan, in the order it is being worked:
|
||||
toolchain itself does not support; through ELF, Plan 9 assembly becomes
|
||||
usable outside Go entirely.
|
||||
- **Platforms: Linux and FreeBSD.** Linux is supported today on all four
|
||||
architectures and is where the binary builds. FreeBSD follows on amd64,
|
||||
arm64 and riscv64: the JIT's executable-memory mapping and the ptrace
|
||||
debugger layer are ported (the debugger's live validation awaits a
|
||||
FreeBSD machine, as the validation status states). Other unix systems
|
||||
may follow those two.
|
||||
architectures and is where the binary builds. FreeBSD follows on amd64
|
||||
and arm64: the JIT's executable-memory mapping and the ptrace debugger
|
||||
layer are ported, the release carries the two FreeBSD binaries, and the
|
||||
debugger's live validation awaits a FreeBSD machine (the validation
|
||||
status states it). Other unix systems may follow those two.
|
||||
- **Four architectures, no more.** amd64, arm64, riscv64 and loong64.
|
||||
No others are planned.
|
||||
|
||||
## Install
|
||||
|
||||
Prebuilt binaries for linux/amd64, linux/arm64, linux/riscv64 and
|
||||
linux/loong64 are on the
|
||||
Prebuilt binaries for linux/amd64, linux/arm64, linux/riscv64,
|
||||
linux/loong64, freebsd/amd64 and freebsd/arm64 are on the
|
||||
[releases page](https://sourcedock.dev/petrbalvin/gasm-sdk/releases).
|
||||
From source (Go 1.27.1):
|
||||
|
||||
|
||||
+3
-3
@@ -134,7 +134,7 @@ Rules: `unknown-instruction`, `operand-count`, `undefined-label`,
|
||||
`nonportable-register-name`, `unencodable-instruction`,
|
||||
`reserved-register-write`, `missing-argsize`, `noframe-frame-size`,
|
||||
`unnamed-fp-reference`, `hardware-sp-addressing`, `vex-sse-mixing`,
|
||||
`unnamed-result`, `data-width`, `data-value-overflow`,
|
||||
`unnamed-result`, `extension-form`, `data-width`, `data-value-overflow`,
|
||||
`data-string-width`, `data-without-globl` and `data-exceeds-globl`.
|
||||
|
||||
```sh
|
||||
@@ -290,8 +290,8 @@ Usage: gasm debug <file.s> --func <name>
|
||||
The debugger re-executes the binary it is running as (`os.Executable()`) for the
|
||||
traced child, so the child is the same `gasm`, whether it is installed on `$PATH`
|
||||
or run with `go run ./cmd/gasm`; nothing has to be installed first. Requires
|
||||
Linux or FreeBSD (ptrace): all four architectures on Linux, amd64, arm64 and
|
||||
riscv64 on FreeBSD.
|
||||
Linux or FreeBSD (ptrace): all four architectures on Linux, amd64 and
|
||||
arm64 on FreeBSD.
|
||||
|
||||
REPL commands:
|
||||
|
||||
|
||||
+2
-3
@@ -145,9 +145,8 @@ Workflows live in `.gitea/workflows/` and run on the project's own runners:
|
||||
| Workflow | Trigger | What it does |
|
||||
|---|---|---|
|
||||
| Test | push or pull request to `development` | the format check, `go vet`, the short layer of the suite with the coverage profile and the 80 % floor, inside the two-minute budget |
|
||||
| Suite | dispatched by hand | the complete gate set minus race: the build, the format check, `go vet`, `go fix -diff`, the full suite and the coverage floor |
|
||||
| FreeBSD build | dispatched by hand | the three FreeBSD compile gates: `GOOS=freebsd` for amd64, arm64 and riscv64 |
|
||||
| Release | a `v*` tag | the matrix build, the version smoke test, and the release with its assets; no gate runs at the tag |
|
||||
| Suite | dispatched by hand | the complete gate set minus race: the build, the two FreeBSD cross-build gates (`GOOS=freebsd` for amd64 and arm64), `go vet`, `go fix -diff`, the full suite and the coverage floor |
|
||||
| Release | a `v*` tag | the matrix build (Linux on four architectures, FreeBSD on two), the version smoke test, and the release with its assets; no gate runs at the tag |
|
||||
|
||||
The pipelines are written by hand rather than through `just`, but they enforce
|
||||
the same gates: the push path carries the affordable subset and the heavy tests
|
||||
|
||||
@@ -100,6 +100,15 @@ encoder backlog that `gasm audit-instructions` measures. The families:
|
||||
whose register list rides the inverted V′VVV field. Mixing VEX and legacy
|
||||
SSE in one loop pays the AVX-SSE transition penalty on every switch: keep
|
||||
a loop in one dialect.
|
||||
- **The extension layer.** The families the toolchain's own table carries
|
||||
late or not at all assemble through the extension mechanism: BF16,
|
||||
VP2INTERSECT, the complete AVX512-FP16 set (the packed and scalar FMA
|
||||
families, the complex multiply and complex FMA pairs, VMINMAXPH and
|
||||
VMINMAXSH), and the AVX-VNNI-INT16 dot products through a VEX path. The
|
||||
layer takes `k0` to `k7` write masks with merging and zeroing, the
|
||||
`{1toN}` broadcast, the embedded rounding and `{sae}` decorations and the
|
||||
imm8 controls, and its refusals surface as `extension-form` lint errors
|
||||
with the reason the encoder would give.
|
||||
- **Cryptographic and counting extensions.** AES-NI, SHA-1 and SHA-256,
|
||||
PCLMULQDQ, GFNI.
|
||||
- **System.** `CPUID`, `RDTSC`, `SYSCALL`, the fences, `LDMXCSR` and
|
||||
|
||||
+29
-2
@@ -12,8 +12,8 @@ toolchain's. The complete mnemonic inventory lives in the generated appendix
|
||||
the stack pointer. There is no R31: thirty-one names and ZR.
|
||||
- Floating-point and SIMD share one file written `Vn`; where an instruction
|
||||
is scalar floating point the operand may be written `Fn` (`F0` to `F31`).
|
||||
- SVE register names (`Z0` to `Z31`, `P0` to `P15`) exist in the assembler's
|
||||
tables.
|
||||
- SVE registers `Z0` to `Z31` and predicates `P0` to `P15`; the SVE2 and
|
||||
SVE2.1 families encode, see SVE below.
|
||||
- Roles the convention fixes: `RSP` is the stack pointer, `R29` the frame
|
||||
pointer, `R30` the link register, `R26` the closure context and `R27` the
|
||||
assembler's scratch register. The goroutine pointer lives in `R28` and is
|
||||
@@ -106,6 +106,33 @@ scalar floating-point instructions being the exceptions. Operands carry an
|
||||
arrangement suffix, `V5.H8`, and structure loads and stores use bracket
|
||||
lists, `[V21.B16]`, with element selection as `V9.S[1]`.
|
||||
|
||||
## SVE
|
||||
|
||||
The scalable vector extension encodes through the `Z0` to `Z31` vector
|
||||
registers and the `P0` to `P15` predicates, with arrangement suffixes on
|
||||
the `Z` registers and the merging and zeroing qualifiers the toolchain's
|
||||
SVE2 and SVE2.1 families carry: the narrowing two-to-one arithmetic, the
|
||||
cryptographic set including ZADCLB, BFloat16 arithmetic, the predicate
|
||||
counters and reductions, the pairwise and quadword forms, the
|
||||
multiple-structure loads and stores, shift by vector and the
|
||||
shift-immediate scheme, CLASTA and CLASTB, the compare and last-active
|
||||
predicate families, and the vector-length arithmetic `ADDVL`, `ADDPL` and
|
||||
`RDVL`, whose immediates count vectors or predicates rather than bytes.
|
||||
The gather and scatter loads take their addresses through the five
|
||||
addressing modes the toolchain defines. The source grammar is the
|
||||
toolchain's own, and the byte output is pinned against it corpus-wide.
|
||||
|
||||
## Synthesised forms
|
||||
|
||||
- `GETCALLERPC` reads the return address the frame state describes: a leaf
|
||||
function reads `R30`, a framed body reads the prologue's save slot.
|
||||
- `REM`, `REMW`, `UREM` and `UREMW` synthesise a remainder from `SDIV` or
|
||||
`UDIV` and the `MSUB` tail, with `RSP` refused as a destination.
|
||||
- `DWORD $imm` lays eight little-endian bytes per immediate.
|
||||
- `MOVD tls_g(SB), Rd` materialises a TLS local-exec load as a single
|
||||
`MOVZ` carrying `R_ARM64_TLS_LE`, keyed off the file's own
|
||||
`GLOBL ... TLSBSS` declaration.
|
||||
|
||||
## Alignment
|
||||
|
||||
`PCALIGN $n` pads to a power-of-two boundary between 8 and 2048 and also
|
||||
|
||||
+11
-6
@@ -23,11 +23,16 @@ TEXT symbol(SB), [flags,] $framesize[-argsize]
|
||||
|
||||
- The symbol is an `·Name(SB)` reference into the current package, or a
|
||||
fully qualified name.
|
||||
- The optional flag argument is a constant expression, normally an OR of the
|
||||
names from `textflag.h`, the table below. Without `#include "textflag.h"`
|
||||
the names are not macros and the assembler reports the misleading error
|
||||
`illegal or missing addressing mode for symbol NOSPLIT`: include the
|
||||
header first.
|
||||
- The optional flag argument is a constant expression: an OR of the names
|
||||
from `textflag.h`, the table below, a bare number (`4`), or a
|
||||
parenthesised combination (`(NOSPLIT|NOFRAME)`). Every spelling
|
||||
suppresses exactly what the names suppress. Without
|
||||
`#include "textflag.h"` the names are not macros and the assembler
|
||||
reports the misleading error `illegal or missing addressing mode for
|
||||
symbol NOSPLIT`: include the header first. A name outside the table is
|
||||
rejected with the toolchain's wording, and two of the toolchain's own
|
||||
TEXT checks fire here: `ABIInternal requires NOSPLIT`, and
|
||||
`NOFRAME functions must have a frame size of 0` for a positive frame.
|
||||
- `$framesize-argsize` is two constants, not a subtraction: the local frame
|
||||
size in bytes, and the caller's argument area in bytes. The argument size
|
||||
may be omitted entirely, `$16`, which marks the argument size unknown
|
||||
@@ -59,7 +64,7 @@ Values from `textflag.h`, in agreement with `cmd/internal/obj/textflag.go`:
|
||||
| WRAPPER | 32 | TEXT | a wrapper; must not disable `recover` |
|
||||
| NEEDCTXT | 64 | TEXT | a closure consuming the context register |
|
||||
| TLSBSS | 256 | data | a thread local word in BSS |
|
||||
| NOFRAME | 512 | TEXT | no frame setup; the zero-frame spelling (the toolchain accepts a positive frame beside it and still allocates the frame, and the arm64 BSD syscall stubs pair it with $-8 deliberately) |
|
||||
| NOFRAME | 512 | TEXT | no frame setup; the frame size must be zero or negative (the toolchain rejects a positive frame beside it, and the arm64 BSD syscall stubs pair it with $-8 deliberately) |
|
||||
| REFLECTMETHOD | 1024 | TEXT | the function calls `reflect.Type.Method` or `MethodByName` |
|
||||
| TOPFRAME | 2048 | TEXT | the outermost frame; unwinders stop here |
|
||||
| ABIWRAPPER | 4096 | TEXT | an ABI transition wrapper |
|
||||
|
||||
@@ -41,6 +41,12 @@ R13` is `add.d R13, R12, R11`, and the two-operand form
|
||||
- Jump and branch instructions keep the GNU order: `BEQ R0, R4, label1`.
|
||||
- The bitfield family is `BSTRINSW`, `BSTRINSV`, `BSTRPICKW`, `BSTRPICKV`
|
||||
`$<msb>, <Rj>, $<lsb>, <Rd>`.
|
||||
- **The register-pair spelling** `R4:R5` names two registers in one
|
||||
operand, the way the toolchain's pair syntax does: the colon splits the
|
||||
operand and the halves swap, so `MULV R4:R5, R6` encodes exactly as
|
||||
`MULV R5, R4, R6`. Any instruction shape accepts the spelling, decided
|
||||
afterwards by ordinary operand matching, and a malformed pair fails with
|
||||
the toolchain's wording.
|
||||
|
||||
## Addressing
|
||||
|
||||
|
||||
@@ -55,6 +55,10 @@ the FCSR.
|
||||
a literal pool in the binary.
|
||||
- A 32-bit constant is accepted by `ADDI`, `ANDI`, `ORI` and `XORI`, and
|
||||
the assembler synthesises values that exceed the 12-bit encoding window.
|
||||
- A memory offset beyond the 12-bit immediate expands the way the
|
||||
toolchain's loads and stores do: the upper bits go into the assembler's
|
||||
temporary register and the access runs against it, byte-identically, and
|
||||
an offset the toolchain refuses is refused with its wording.
|
||||
- `MOVF` and `MOVD` materialise floating-point constants, encoding them as
|
||||
`FLW` and `FLD` from a pool location unless the constant is exactly 0.0.
|
||||
|
||||
@@ -85,6 +89,15 @@ of the source, and register choice influences how much compresses.
|
||||
Hand-writing compressed instructions in source is accepted but discouraged.
|
||||
The debug flag `compressinstructions=0` turns the automatic conversion off.
|
||||
|
||||
## Pseudo-operations
|
||||
|
||||
- `END` is accepted anywhere and emits nothing, the function-end marker
|
||||
the toolchain's front end also swallows.
|
||||
- `GETCALLERPC` reads the return address the frame state describes: a leaf
|
||||
function reads `X1`, a framed body reads the prologue's save slot at
|
||||
`0(SP)`, and both spellings take the compressed encoding where the
|
||||
register choice allows it.
|
||||
|
||||
## Vector extension
|
||||
|
||||
`VSETVLI` writes its vtype components in uppercase with the destination
|
||||
|
||||
Reference in new issue
Block a user