chore: prepare release v0.36.0
Test / test (push) Successful in 32s
Release / build (amd64, freebsd) (push) Successful in 52s
Release / build (amd64, linux) (push) Successful in 13s
Release / build (arm64, freebsd) (push) Successful in 51s
Release / build (arm64, linux) (push) Successful in 29s
Release / build (loong64, linux) (push) Successful in 29s
Release / build (riscv64, linux) (push) Successful in 31s
Release / release (push) Successful in 10s

Assisted-by: GLM 5.3 Flash
This commit is contained in:
petrbalvin committed 2026-10-07 22:37:42 +02:00
1 parent e5df0a9d84
commit a77dd12d2d
13 files changed
+217 -77

No files matched your search

-41
View File
@@ -1,41 +0,0 @@
# FreeBSD compile gates. Dispatched by hand, never on a push.
#
# The debugger's ptrace surface and the JIT substrate are the two
# FreeBSD-portable layers the tree carries; the forge has no FreeBSD runner,
# so they can only be compile-gated, and three foreign-GOOS builds of the
# whole module are minutes of one-core work the push pipeline's budget cannot
# carry. The push pipeline stays fast and light; this workflow is the
# deliberate run, before a release or after touching the ported layers.
# Running the ptrace suite itself needs real FreeBSD hardware.
#
# A dispatched workflow takes no concurrency block: it is one deliberate run.
name: FreeBSD build
on:
workflow_dispatch:
env:
# One core: parallelism buys no speed here and costs memory the box does not have.
GOFLAGS: -p=1
GOMAXPROCS: "2"
jobs:
build:
runs-on: fedora
steps:
- uses: actions/checkout@v7
- uses: actions/setup-go@v6
with:
# The module is the source of truth for the version, so it cannot drift.
go-version-file: go.mod
cache: true
- name: FreeBSD build (amd64)
run: GOOS=freebsd GOARCH=amd64 go build ./...
- name: FreeBSD build (arm64)
run: GOOS=freebsd GOARCH=arm64 go build ./...
- name: FreeBSD build (riscv64)
run: GOOS=freebsd GOARCH=riscv64 go build ./...
+10 -3
View File
@@ -11,9 +11,12 @@
# (devel) even at its own <module>/vX.Y.Z tag and this workflow's smoke test can never
# pass for it. A Go repository is one module at the root.
#
# The matrix carries the platforms the project ships: Linux on amd64, arm64, loong64
# and riscv64. The FreeBSD port compiles in its own dispatched workflow and ships no
# binary. Nothing is installed: the fedora job image carries git, perl and node
# The matrix carries the platforms the project ships: Linux on amd64, arm64,
# loong64 and riscv64, and FreeBSD on amd64 and arm64. Go cross-compiles
# FreeBSD natively and the tree builds under CGO_ENABLED=0, which the
# suite's cross-build gates prove for the whole module; the ptrace suite
# itself still needs a FreeBSD machine, so nothing FreeBSD runs here.
# Nothing is installed: the fedora job image carries git, perl and node
# (verified on the runner, 2026-10-04).
#
# Each job validates the tag for itself rather than passing a value between jobs, so
@@ -46,6 +49,10 @@ jobs:
goarch: loong64
- goos: linux
goarch: riscv64
- goos: freebsd
goarch: amd64
- goos: freebsd
goarch: arm64
steps:
- uses: actions/checkout@v7
+14 -4
View File
@@ -1,9 +1,11 @@
# Suite, Go. Dispatched by hand, on development. Never a push gate.
#
# The complete gate set minus race: the build, both static gates, the full suite with
# every short-layer skip unskipped, and the coverage floor. It is the pipeline form of
# the local `just test`, for the moments when the tree must be proven end to end and
# nobody is at the keyboard.
# The complete gate set minus race: the build, the FreeBSD cross-builds the
# port's compile proof needs (amd64 and arm64, the same two the release
# matrix ships), both static gates, the full suite with every short-layer
# skip unskipped, and the coverage floor. It is the pipeline form of the
# local `just test`, for the moments when the tree must be proven end to end
# and nobody is at the keyboard.
#
# Race never runs in CI. It roughly doubles the time and the memory on a box shared
# with the forge, and the local `just gates` races the tree on the machine at the
@@ -33,6 +35,14 @@ jobs:
- name: Build
run: go build ./...
- name: FreeBSD build (amd64)
# The port is pure Go: the cross-build needs no C toolchain and the
# release matrix ships the same two binaries.
run: CGO_ENABLED=0 GOOS=freebsd GOARCH=amd64 go build ./...
- name: FreeBSD build (arm64)
run: CGO_ENABLED=0 GOOS=freebsd GOARCH=arm64 go build ./...
- name: Format
run: |
perl -e '
+109 -3
View File
@@ -7,17 +7,75 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
## [development]
## [0.36.0] - 2026-10-07
### Added
- **The SVE2 and SVE2.1 instruction sets on arm64.** Over 350 SVE
mnemonics across 537 encoded forms join the extension layer: the
narrowing two-to-one arithmetic family, the SVE2 cryptographic set
including ZADCLB, BFloat16 arithmetic, predicate counters and
reductions, the pairwise and quadword forms, multiple-structure loads
and stores (LD2/ST2 through LD4/ST4), shift by vector and the
shift-immediate scheme with its element-size encoding, CLASTA and
CLASTB, the last-active and compare predicate families, and the
vector-length arithmetic ADDVL, ADDPL and RDVL. Every form derives
from the toolchain's own encoder tables and is pinned against the
GOROOT corpus words.
- **The complete AVX512-FP16 set and AVX-VNNI-INT16 on amd64.** The
extension layer reaches 191 entries: the packed and scalar FMA
families (VFMADD, VFMSUB, VFMADDSUB and VFMSUBADD in the PH widths
with their SH mirrors), the complex multiply and complex FMA pairs
(VFMULC, VFCMULC, VFMADDC, VFCMADDC, packed and scalar, each
conjugating the source its prefix names), VMINMAXPH and VMINMAXSH
under their imm8 control, and the AVX-VNNI-INT16 dot products
(VPDPWSUD, VPDPWSUDS, VPDPWUSD, VPDPWUSDS) through a new VEX
encoding path beside the EVEX one. Every entry carries golden
vectors byte-checked against GNU as.
- **Extension-layer instructions assemble straight from `.s` source on
amd64.** The registered mnemonics dispatch from the assembly front
end with their operand grammar: `k0` through `k7` write masks with
merging and zeroing, `{1toN}` broadcast, embedded rounding and
`{sae}`, memory and scaled-index operands, and the imm8-control
forms, the VEX forms beside the EVEX ones. `gasm lint` surfaces the
layer's refusals as `extension-form` errors instead of letting a bad
shape reach the encoder, and every distinct mnemonic is pinned from
source text against the registry's bytes.
- **arm64 output byte-identical with the toolchain across the GOROOT
tree.** 31 of the 36 files the differential harness walks (90
functions, 27408 bytes) now assemble without a differing byte: the
missing shapes are in (GETCALLERPC, the REM, REMW, UREM and UREMW
family, DWORD, the FCVTHS, FCVTSH, FCVTDH and FCVTHD conversions, TLS
local-exec loads as a single MOVZ with the TLS relocation), the
literal pool drains mid-function when a function's literals would
overflow the ±512 KiB load-literal range, and frame-relative
addresses and flag-setting logicals materialise the way
cmd/internal/obj/arm64 writes them.
- **riscv64 END and GETCALLERPC.** END is accepted anywhere and emits
nothing, and GETCALLERPC reads the return address the toolchain's
rewrite describes: a leaf function reads LR, a framed body reads the
prologue's save slot, with the compressed spellings where they fit.
- **The loong64 register-pair spelling.** `MULV R4:R5, R6` parses the
way the toolchain's pair sugar does: the colon splits the operand and
the halves swap, any instruction shape takes the spelling, and the
malformed forms fail with the toolchain's own wording.
- **Fuzz targets across the tool.** The assembler carries
per-architecture FuzzAssemble targets and the front end carries
targets over the lexer, parser, formatter and the extension
registry; the seed corpora run as ordinary tests, and `just fuzz`
drives the campaigns locally. Hardening the targets found caps GLOBL
and DATA at 64 MiB per symbol and 128 MiB per section, so a crafted
input cannot materialise unbounded memory.
- **The FreeBSD port of the debugger.** `gasm debug` runs on FreeBSD on
amd64, arm64 and riscv64 with the same interactive surface as on Linux:
amd64 and arm64 with the same interactive surface as on Linux:
breakpoints, hardware watchpoints (x86 debug registers, the arm64 debug
register file), single-stepping, register and memory access, all behind
the kernel's own ptrace requests, with tracee memory through `PT_IO` and
stop reports through `PT_LWPINFO`. The JIT substrate maps executable
memory through `golang.org/x/sys/unix`, so `verify` builds on FreeBSD
too. The pipeline compile-gates all three architectures; live
validation awaits a FreeBSD machine.
too. The dispatched suite compile-gates both architectures, the
release pipeline ships the two binaries, and live validation awaits a
FreeBSD machine.
- **Workspace-wide navigation in the language server.** `gasm lsp` indexes
the `.s` files under the workspace root beyond the documents the editor
has open, so go-to-definition, find references and workspace symbol search
@@ -75,6 +133,23 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
### Changed
- **TEXT flag operands behave like the toolchain's.** A numeric or
parenthesised flag list (`TEXT ·f(SB), 4, $4096-0`,
`(NOSPLIT|NOFRAME)`) now suppresses the prologue and stack guard the
names suppress, where the numbers were silently ignored and the
prologue was emitted anyway; unknown flag names are rejected with the
toolchain's wording instead of passing unnoticed; and the toolchain's
own TEXT diagnostics fire, `ABIInternal requires NOSPLIT` and
`NOFRAME functions must have a frame size of 0`.
- **The disassembler names what it used to print raw.** The 195 amd64
encodings that decoded to anonymous renders now decode to their
instructions and round-trip, the BMI1 and BMI2 VEX set, RDSEED, RDPID,
the WAITPKG family, ENDBR, CLDEMOTE, UD1 and RORX among them, with the
arm64 HVC, SMC and SB forms and 74 loong64 rendering corrections
beside.
- **Variadic macro parameters are rejected like the toolchain rejects
them.** A macro declaration whose parameter list the toolchain
refuses no longer parses into a broken definition.
- **The corpus audit assembles like the build.** A file's `//go:build`
constraint decides which target architectures attempt it: cpu_x86.s is
an x86 build alone, and the msan and goexperiment.runtimesecret trees
@@ -94,6 +169,37 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
### Fixed
- **riscv64 memory offsets beyond the 12-bit immediate no longer
truncate silently.** `LD 4096(X6), X5` encoded as `LD X5, 0(X6)` and
read the wrong address; the expansion the toolchain performs (the
upper bits into its temporary register, then the access) is emitted
now, byte-identically, and offsets the toolchain refuses are refused
with its wording.
- **riscv64 operand acceptance matches the toolchain's.** 398 operand
shapes the toolchain rejects assembled without complaint (a float
register where an integer one is required, out-of-range CSRs, vector
shapes against the wrong bank, MOV family widths), and 488 more
produced the wrong diagnostic; all now fail with the toolchain's
texts, and a 1030-case error-parity catalogue pins them against the
toolchain's own negative files.
- **The debugger survives its failure paths.** Killing a debuggee that
had already died or sat ptrace-stopped could block the session
forever; Kill now wakes the tracee, signals and reaps it without
waiting on a parked child, so every launch-failure path returns
instead of hanging the tool.
- **Encoder defects found by fuzzing and the differential audits.** A
GLOBL with a negative size panicked the assembler and a small crafted
input materialised 3.92 GiB of DATA; certain SVE VTBL operands
panicked; an over-strict guard refused BIC into RSP where the
toolchain assembles it; byte-register operands now select the 8-bit
forms the toolchain picks (`XADDL DL, DL` and its companions); and
the loong64 shift compositions and indexed-offset forms that dropped
fields re-encode faithfully.
- **Formatter edges around labels and macros.** A stacked label that
names a macro stays attached to it, macro content behind a label
stays on its line, a selector-folded label resolves to its macro, an
empty TEXT body formats, and a file ending in a comment no longer
loses the comment.
- **Rename edits land in their own documents.** A rename collected the
ranges of every reference across the open documents but applied them all
to the document that started it, so renaming a symbol used in a second
+2 -3
View File
@@ -118,9 +118,8 @@ Workflows live in `.gitea/workflows/` and run on the project's own runners:
| Workflow | Trigger | What it does |
|---|---|---|
| Test | push or pull request to `development` | the format check, `go vet`, and the short layer of the test suite with the coverage profile and the 80 % floor, inside the two-minute budget |
| Suite | dispatched by hand | the complete gate set minus race: the build, both static gates, the full suite and the coverage floor |
| FreeBSD build | dispatched by hand | the three FreeBSD compile gates for amd64, arm64 and riscv64 |
| Release | a `v*` tag | the matrix build, the version smoke test and the release with its assets; no gate runs at the tag |
| Suite | dispatched by hand | the complete gate set minus race: the build, the two FreeBSD cross-build gates (amd64, arm64), both static gates, the full suite and the coverage floor |
| Release | a `v*` tag | the matrix build (Linux on four architectures, FreeBSD on two), the version smoke test and the release with its assets; no gate runs at the tag |
The local equivalent is `just gates`, which is the same set plus the race detector. Race
never runs in CI, on a push or a tag: it would double the time and the memory a shared
+9 -9
View File
@@ -95,7 +95,7 @@ to give that syntax the tooling it deserves.
breakpoints (optionally conditional), hardware watchpoints, register and
memory inspection, and headless script runs that report instruction and
label coverage; it runs on Linux (all four architectures) and FreeBSD
(amd64, arm64, riscv64).
(amd64, arm64).
- **Language server.** `gasm lsp` serves completion, hover, document symbols,
push and pull diagnostics, semantic-token highlighting, go-to-definition,
find references, rename, formatting, inlay hints, code actions, signature
@@ -151,7 +151,7 @@ actually been executed.
|---|---|---|
| Encoding: byte-for-byte against `go tool asm` | native hardware | native hardware (the toolchain cross-assembles any GOARCH on any host) |
| Execution: JIT calls, ABI checks, differential fuzzing | native hardware | qemu-user emulation |
| Debugger: ptrace tracing, breakpoints, watchpoints, coverage | native hardware | emulation cannot run ptrace; the layer compiles and its architecture-neutral units run under `go test ./...`, nothing more. FreeBSD (amd64, arm64, riscv64) is in the same position: the port compiles behind the cross-build gate and its integration test is ready, but no FreeBSD machine has executed it |
| Debugger: ptrace tracing, breakpoints, watchpoints, coverage | native hardware | emulation cannot run ptrace; the layer compiles and its architecture-neutral units run under `go test ./...`, nothing more. FreeBSD (amd64, arm64) is in the same position: the port compiles behind the cross-build gate and its integration test is ready, but no FreeBSD machine has executed it |
Consequences, stated plainly. An emulator is a model of a CPU, not the
CPU: instruction semantics are implemented in software and can differ
@@ -222,18 +222,18 @@ The plan, in the order it is being worked:
toolchain itself does not support; through ELF, Plan 9 assembly becomes
usable outside Go entirely.
- **Platforms: Linux and FreeBSD.** Linux is supported today on all four
architectures and is where the binary builds. FreeBSD follows on amd64,
arm64 and riscv64: the JIT's executable-memory mapping and the ptrace
debugger layer are ported (the debugger's live validation awaits a
FreeBSD machine, as the validation status states). Other unix systems
may follow those two.
architectures and is where the binary builds. FreeBSD follows on amd64
and arm64: the JIT's executable-memory mapping and the ptrace debugger
layer are ported, the release carries the two FreeBSD binaries, and the
debugger's live validation awaits a FreeBSD machine (the validation
status states it). Other unix systems may follow those two.
- **Four architectures, no more.** amd64, arm64, riscv64 and loong64.
No others are planned.
## Install
Prebuilt binaries for linux/amd64, linux/arm64, linux/riscv64 and
linux/loong64 are on the
Prebuilt binaries for linux/amd64, linux/arm64, linux/riscv64,
linux/loong64, freebsd/amd64 and freebsd/arm64 are on the
[releases page](https://sourcedock.dev/petrbalvin/gasm-sdk/releases).
From source (Go 1.27.1):
+3 -3
View File
@@ -134,7 +134,7 @@ Rules: `unknown-instruction`, `operand-count`, `undefined-label`,
`nonportable-register-name`, `unencodable-instruction`,
`reserved-register-write`, `missing-argsize`, `noframe-frame-size`,
`unnamed-fp-reference`, `hardware-sp-addressing`, `vex-sse-mixing`,
`unnamed-result`, `data-width`, `data-value-overflow`,
`unnamed-result`, `extension-form`, `data-width`, `data-value-overflow`,
`data-string-width`, `data-without-globl` and `data-exceeds-globl`.
```sh
@@ -290,8 +290,8 @@ Usage: gasm debug <file.s> --func <name>
The debugger re-executes the binary it is running as (`os.Executable()`) for the
traced child, so the child is the same `gasm`, whether it is installed on `$PATH`
or run with `go run ./cmd/gasm`; nothing has to be installed first. Requires
Linux or FreeBSD (ptrace): all four architectures on Linux, amd64, arm64 and
riscv64 on FreeBSD.
Linux or FreeBSD (ptrace): all four architectures on Linux, amd64 and
arm64 on FreeBSD.
REPL commands:
+2 -3
View File
@@ -145,9 +145,8 @@ Workflows live in `.gitea/workflows/` and run on the project's own runners:
| Workflow | Trigger | What it does |
|---|---|---|
| Test | push or pull request to `development` | the format check, `go vet`, the short layer of the suite with the coverage profile and the 80 % floor, inside the two-minute budget |
| Suite | dispatched by hand | the complete gate set minus race: the build, the format check, `go vet`, `go fix -diff`, the full suite and the coverage floor |
| FreeBSD build | dispatched by hand | the three FreeBSD compile gates: `GOOS=freebsd` for amd64, arm64 and riscv64 |
| Release | a `v*` tag | the matrix build, the version smoke test, and the release with its assets; no gate runs at the tag |
| Suite | dispatched by hand | the complete gate set minus race: the build, the two FreeBSD cross-build gates (`GOOS=freebsd` for amd64 and arm64), `go vet`, `go fix -diff`, the full suite and the coverage floor |
| Release | a `v*` tag | the matrix build (Linux on four architectures, FreeBSD on two), the version smoke test, and the release with its assets; no gate runs at the tag |
The pipelines are written by hand rather than through `just`, but they enforce
the same gates: the push path carries the affordable subset and the heavy tests
+9
View File
@@ -100,6 +100,15 @@ encoder backlog that `gasm audit-instructions` measures. The families:
whose register list rides the inverted V′VVV field. Mixing VEX and legacy
SSE in one loop pays the AVX-SSE transition penalty on every switch: keep
a loop in one dialect.
- **The extension layer.** The families the toolchain's own table carries
late or not at all assemble through the extension mechanism: BF16,
VP2INTERSECT, the complete AVX512-FP16 set (the packed and scalar FMA
families, the complex multiply and complex FMA pairs, VMINMAXPH and
VMINMAXSH), and the AVX-VNNI-INT16 dot products through a VEX path. The
layer takes `k0` to `k7` write masks with merging and zeroing, the
`{1toN}` broadcast, the embedded rounding and `{sae}` decorations and the
imm8 controls, and its refusals surface as `extension-form` lint errors
with the reason the encoder would give.
- **Cryptographic and counting extensions.** AES-NI, SHA-1 and SHA-256,
PCLMULQDQ, GFNI.
- **System.** `CPUID`, `RDTSC`, `SYSCALL`, the fences, `LDMXCSR` and
+29 -2
View File
@@ -12,8 +12,8 @@ toolchain's. The complete mnemonic inventory lives in the generated appendix
the stack pointer. There is no R31: thirty-one names and ZR.
- Floating-point and SIMD share one file written `Vn`; where an instruction
is scalar floating point the operand may be written `Fn` (`F0` to `F31`).
- SVE register names (`Z0` to `Z31`, `P0` to `P15`) exist in the assembler's
tables.
- SVE registers `Z0` to `Z31` and predicates `P0` to `P15`; the SVE2 and
SVE2.1 families encode, see SVE below.
- Roles the convention fixes: `RSP` is the stack pointer, `R29` the frame
pointer, `R30` the link register, `R26` the closure context and `R27` the
assembler's scratch register. The goroutine pointer lives in `R28` and is
@@ -106,6 +106,33 @@ scalar floating-point instructions being the exceptions. Operands carry an
arrangement suffix, `V5.H8`, and structure loads and stores use bracket
lists, `[V21.B16]`, with element selection as `V9.S[1]`.
## SVE
The scalable vector extension encodes through the `Z0` to `Z31` vector
registers and the `P0` to `P15` predicates, with arrangement suffixes on
the `Z` registers and the merging and zeroing qualifiers the toolchain's
SVE2 and SVE2.1 families carry: the narrowing two-to-one arithmetic, the
cryptographic set including ZADCLB, BFloat16 arithmetic, the predicate
counters and reductions, the pairwise and quadword forms, the
multiple-structure loads and stores, shift by vector and the
shift-immediate scheme, CLASTA and CLASTB, the compare and last-active
predicate families, and the vector-length arithmetic `ADDVL`, `ADDPL` and
`RDVL`, whose immediates count vectors or predicates rather than bytes.
The gather and scatter loads take their addresses through the five
addressing modes the toolchain defines. The source grammar is the
toolchain's own, and the byte output is pinned against it corpus-wide.
## Synthesised forms
- `GETCALLERPC` reads the return address the frame state describes: a leaf
function reads `R30`, a framed body reads the prologue's save slot.
- `REM`, `REMW`, `UREM` and `UREMW` synthesise a remainder from `SDIV` or
`UDIV` and the `MSUB` tail, with `RSP` refused as a destination.
- `DWORD $imm` lays eight little-endian bytes per immediate.
- `MOVD tls_g(SB), Rd` materialises a TLS local-exec load as a single
`MOVZ` carrying `R_ARM64_TLS_LE`, keyed off the file's own
`GLOBL ... TLSBSS` declaration.
## Alignment
`PCALIGN $n` pads to a power-of-two boundary between 8 and 2048 and also
+11 -6
View File
@@ -23,11 +23,16 @@ TEXT symbol(SB), [flags,] $framesize[-argsize]
- The symbol is an `·Name(SB)` reference into the current package, or a
fully qualified name.
- The optional flag argument is a constant expression, normally an OR of the
names from `textflag.h`, the table below. Without `#include "textflag.h"`
the names are not macros and the assembler reports the misleading error
`illegal or missing addressing mode for symbol NOSPLIT`: include the
header first.
- The optional flag argument is a constant expression: an OR of the names
from `textflag.h`, the table below, a bare number (`4`), or a
parenthesised combination (`(NOSPLIT|NOFRAME)`). Every spelling
suppresses exactly what the names suppress. Without
`#include "textflag.h"` the names are not macros and the assembler
reports the misleading error `illegal or missing addressing mode for
symbol NOSPLIT`: include the header first. A name outside the table is
rejected with the toolchain's wording, and two of the toolchain's own
TEXT checks fire here: `ABIInternal requires NOSPLIT`, and
`NOFRAME functions must have a frame size of 0` for a positive frame.
- `$framesize-argsize` is two constants, not a subtraction: the local frame
size in bytes, and the caller's argument area in bytes. The argument size
may be omitted entirely, `$16`, which marks the argument size unknown
@@ -59,7 +64,7 @@ Values from `textflag.h`, in agreement with `cmd/internal/obj/textflag.go`:
| WRAPPER | 32 | TEXT | a wrapper; must not disable `recover` |
| NEEDCTXT | 64 | TEXT | a closure consuming the context register |
| TLSBSS | 256 | data | a thread local word in BSS |
| NOFRAME | 512 | TEXT | no frame setup; the zero-frame spelling (the toolchain accepts a positive frame beside it and still allocates the frame, and the arm64 BSD syscall stubs pair it with $-8 deliberately) |
| NOFRAME | 512 | TEXT | no frame setup; the frame size must be zero or negative (the toolchain rejects a positive frame beside it, and the arm64 BSD syscall stubs pair it with $-8 deliberately) |
| REFLECTMETHOD | 1024 | TEXT | the function calls `reflect.Type.Method` or `MethodByName` |
| TOPFRAME | 2048 | TEXT | the outermost frame; unwinders stop here |
| ABIWRAPPER | 4096 | TEXT | an ABI transition wrapper |
+6
View File
@@ -41,6 +41,12 @@ R13` is `add.d R13, R12, R11`, and the two-operand form
- Jump and branch instructions keep the GNU order: `BEQ R0, R4, label1`.
- The bitfield family is `BSTRINSW`, `BSTRINSV`, `BSTRPICKW`, `BSTRPICKV`
`$<msb>, <Rj>, $<lsb>, <Rd>`.
- **The register-pair spelling** `R4:R5` names two registers in one
operand, the way the toolchain's pair syntax does: the colon splits the
operand and the halves swap, so `MULV R4:R5, R6` encodes exactly as
`MULV R5, R4, R6`. Any instruction shape accepts the spelling, decided
afterwards by ordinary operand matching, and a malformed pair fails with
the toolchain's wording.
## Addressing
+13
View File
@@ -55,6 +55,10 @@ the FCSR.
a literal pool in the binary.
- A 32-bit constant is accepted by `ADDI`, `ANDI`, `ORI` and `XORI`, and
the assembler synthesises values that exceed the 12-bit encoding window.
- A memory offset beyond the 12-bit immediate expands the way the
toolchain's loads and stores do: the upper bits go into the assembler's
temporary register and the access runs against it, byte-identically, and
an offset the toolchain refuses is refused with its wording.
- `MOVF` and `MOVD` materialise floating-point constants, encoding them as
`FLW` and `FLD` from a pool location unless the constant is exactly 0.0.
@@ -85,6 +89,15 @@ of the source, and register choice influences how much compresses.
Hand-writing compressed instructions in source is accepted but discouraged.
The debug flag `compressinstructions=0` turns the automatic conversion off.
## Pseudo-operations
- `END` is accepted anywhere and emits nothing, the function-end marker
the toolchain's front end also swallows.
- `GETCALLERPC` reads the return address the frame state describes: a leaf
function reads `X1`, a framed body reads the prologue's save slot at
`0(SP)`, and both spellings take the compressed encoding where the
register choice allows it.
## Vector extension
`VSETVLI` writes its vtype components in uppercase with the destination