docs(asm): describe the four target architectures
Test / test (push) Failing after 2m23s

Assisted-by: GLM 5.3 Flash
This commit is contained in:
2026-09-21 20:15:55 +02:00
parent 8a36af7c7d
commit 53de91b2df
6 changed files with 461 additions and 6 deletions
+124
View File
@@ -0,0 +1,124 @@
# AMD64
Layer 1, target page. Verified against `go tool asm` of Go 1.27.1 and against
gasm's encoder, whose output is compared byte for byte with the toolchain's
and executed on real hardware (`gasm verify`). The complete mnemonic
inventory lives in the generated appendix
[INSTRUCTIONS-AMD64.md](INSTRUCTIONS-AMD64.md); this page is the grammar and
the conventions.
## Registers
| Group | Names | Notes |
|---|---|---|
| General purpose, 64-bit | `AX` `BX` `CX` `DX` `SI` `DI` `BP` `SP` `R8` to `R15` | bare names, no prefix |
| Sub-registers | `AL` `CL` `DL` `BL` `AH` family; `R8B` `R8W` `R8D` for the byte, word and double word of `R8` | width rides the mnemonic as well |
| Vector | `X0` to `X15` (128-bit), `Y0` to `Y15` (256-bit), `Z0` to `Z31` (512-bit) | SSE, AVX and AVX-512 |
| Mask | `K0` to `K7` | AVX-512 opmask |
| System | `TLS` | the thread pointer, see below |
Roles the calling convention fixes, which assembly must respect and can rely
on:
- `SP` is the hardware stack pointer; the virtual frame pointer of the
common language is the pseudo-register SP of OPERANDS.md, a different
spelling with a different meaning.
- `BP` is callee-save. The assembler inserts the save and restore whenever
the function has a non-zero frame, so using BP as a general register
interferes with sampling profilers that walk the frame chain.
- `R14` holds `g`, the goroutine pointer, in the register ABI; `RDX` holds
the closure context; `R12` and `R13` are the register ABI's scratch pair
and `R15` its GOT temporary; `X15` is the zeroing register the compiler
uses. An ABI0 assembly function called from Go sees none of these live
across the call, but runtime assembly reads them directly.
- The legacy spellings for the goroutine pointer are the macros of
`runtime/go_tls.h`: `get_tls(r)` expands to `MOVQ TLS, r` and `g(r)` to
`0(r)(TLS*1)`, the segment base riding the index field.
## Addressing
The common forms of OPERANDS.md, with the amd64 specifics:
```text
offset(base) MOVQ 16(BX), AX
offset(base)(index*scale) MOVL foo+32(SP)(R9*8), CX
scale is 1, 2, 4 or 8
name±offset(SB) MOVQ ·table(SB), CX
```
- Global references assemble as absolute addresses and produce R_ADDR
relocations; branch targets produce R_PCREL.
- Vector indexed memory, the VSIB form with an X, Y or Z register in the
index position, exists for the gather and scatter families.
- There are no segment overrides in source; the one segment-flavoured form
is the TLS base in the index field shown above.
## The frame and the split check
The assembler manages the frame, not the programmer:
- It inserts the `BP` save and restore for any non-zero frame.
- It inserts the stack-split check for any function that is not NoSplit:
the check compares SP against the guard, and on exhaustion calls
`runtime.morestack_noctxt`. Frames at or below 128 bytes, StackSmall, use
the small compare; frames at or below 4096 bytes, StackBig, use the
adjusted form; larger frames compare in two steps.
- On amd64 the assembler marks a function NoSplit itself when the frame is
under StackSmall and the body calls nothing that needs stack: such a
function carries the NoSplit flag in the object without the source ever
writing NOSPLIT.
Results and arguments are stack-only in ABI0: the caller's frame carries
them at FP offsets, per the Go prototype.
## Instructions
The inventory counts 1654 recognised mnemonics today, of which the encoder
emits 1113; both numbers are generated in the appendix, and the gap is the
encoder backlog that `gasm audit-instructions` measures. The families:
- **Integer base.** The ALU and move set with width suffixes, `MOVB`,
`MOVW`, `MOVL`, `MOVQ`; the extension moves `MOVBLZX`, `MOVWLSX`,
`MOVLQSX` and their siblings, which the compiler's output leans on;
`LEA`; `PUSH` and `POP`; the shifts and rotates; the bit operations `BT`
through `BTC`, `BSF`, `BSR`, `LZCNT`, `TZCNT`, `POPCNT`, `BSWAP`; the
string primitives `MOVS` and `STOS`.
- **Exchange and atomics.** `XCHG`, `CMPXCHG`, `XADD`; the extended-carry
pair `ADCX` and `ADOX`; `CRC32`.
- **Scalar floating point.** The SSE2 scalar moves and arithmetic
(`MOVSD`, `MOVSS`, `ADDSD`, and the `CVT` family). Floating-point
immediates are not encodable on this target, so the assembler
materialises them: the constant lands in a synthesised read-only pool,
and a positive zero collapses to `XORPS` of the register with itself,
exactly as the toolchain does.
- **Legacy SIMD, SSE.** The `MOVO`, `MOVOU`, `MOVAPS` family and the packed
integer and floating operations, shuffles, lane extracts and inserts and
the imm8-controlled forms.
- **VEX and EVEX.** The `V`-prefixed forms for 256 and 512-bit work,
opmask operations on `K0` to `K7`, gathers and scatters, and the
quad-register families 4FMAPS, 4FNMADD, 4VNNIW, VP4DPWSSD and VP4DPWSSDS,
whose register list rides the inverted V′VVV field. Mixing VEX and legacy
SSE in one loop pays the AVX-SSE transition penalty on every switch: keep
a loop in one dialect.
- **Cryptographic and counting extensions.** AES-NI, SHA-1 and SHA-256,
PCLMULQDQ, GFNI.
- **System.** `CPUID`, `RDTSC`, `SYSCALL`, the fences, `LDMXCSR` and
`STMXCSR`, the prefetch family.
- **Pseudo-operations.** `BYTE`, `WORD`, `LONG`, `QUAD` lay raw bytes or
words into the stream for encodings the assembler does not know; `ADJSP`
adjusts the stack pointer; `DUFFCOPY` and `DUFFZERO` and `GETCALLERPC`
are compiler-side names the table recognises but an encoder need not
emit.
A mnemonic the appendix lists with `gasm encodes: no` assembles nowhere:
gasm reports it as an explicit error, never as wrong bytes, and the
`unencodable-instruction` lint flags it at edit time.
## Relocations
The relocations an amd64 object carries, all specified in
[GOOBJ.md](../GOOBJ.md): `R_ADDR` for absolute globals, `R_PCREL` for
relative addresses, `R_CALL` for direct calls, `R_TLS_LE` and `R_TLS_IE` for
thread local access and `R_GOTPCREL` for GOT relative sequences, plus
`R_DWTXTADDR_U4` inside the DWARF records, which the assembler always
emits in the four-byte flavour.
+122
View File
@@ -0,0 +1,122 @@
# ARM64
Layer 1, target page. Verified against `go tool asm` of Go 1.27.1, against the
toolchain's own arm64 assembler manual (`cmd/internal/obj/arm64/doc.go`) and
against gasm's encoder, whose output is compared byte for byte with the
toolchain's. The complete mnemonic inventory lives in the generated appendix
[INSTRUCTIONS-ARM64.md](INSTRUCTIONS-ARM64.md).
## Registers
- General purpose: `R0` to `R30`, plus `ZR`, the zero register, and `RSP`,
the stack pointer. There is no R31: thirty-one names and ZR.
- Floating-point and SIMD share one file written `Vn`; where an instruction
is scalar floating point the operand may be written `Fn` (`F0` to `F31`).
- SVE register names (`Z0` to `Z31`, `P0` to `P15`) exist in the assembler's
tables.
- Roles the convention fixes: `RSP` is the stack pointer, `R29` the frame
pointer, `R30` the link register, `R26` the closure context and `R27` the
assembler's scratch register. The goroutine pointer lives in `R28` and is
written `g` in source, its fields as `g_m(g)`, `g_sched(g)`; `R18` is the
platform-reserved register and the Go toolchain never addresses it.
## Loads, stores and the width suffixes
The MOV series is the load and store interface, with the width in the
mnemonic rather than the register name:
| Mnemonic | Machine instruction |
|---|---|
| `MOVD` | ldr, str, stur, 64-bit |
| `MOVW` | ldrsw, str, stur, 32-bit sign extending |
| `MOVWU` | ldr, 32-bit zero extending |
| `MOVH` | ldrsh, strh, sturh |
| `MOVHU` | ldrh |
| `MOVB` | ldrsb, strb, sturb |
| `MOVBU` | ldrb |
Post-index and pre-index addressing take the `.P` and `.W` suffixes on the
mnemonic: `MOVD.P -8(R10), R8` is `ldr x8, [x10],#-8`, and `MOVB.W
16(R16), R10` is `ldrsb x10, [x16,#16]!`.
## Addressing
```text
imm(Rn|RSP) 28(R17)
(Rn|RSP) (R22)
(Rn)(Rm) (R27)(R23)
(Rn)(Rm<<scale) (R4)(R12<<2)
(Rn)(Rm.UXTW<<3) extended and shifted index
(Rt1, Rt2) register pair for LDP, STP and the exclusive pair forms
```
Branch targets are labels, `(R3)` for indirect, `name(SB)` for static.
## Operand order and the special forms
Most instructions appear in left-to-right assignment order: `ADD R11,
RSP, R25` computes into R25. The exceptions the toolchain's manual lists,
each with its own order:
- stores and `CBZ`, `CBNZ` keep the GNU order: `MOVD R29, 384(R19)`.
- The multiply-accumulate family `MADD`, `MSUB`, `SMADDL` and friends are
`<Rm>, <Ra>, <Rn>, <Rd>`.
- The scalar FMA family `FMADDD` and friends are `<Fm>, <Fa>, <Fn>, <Fd>`.
- The bitfield family `BFI`, `BFXIL`, `SBFIZ`, `SBFX`, `UBFIZ`, `UBFX` is
`$<lsb>, <Rn>, $<width>, <Rd>`.
- The conditional compare and select families carry the condition as the
**first** operand: `CSEL GT, R0, R19, R1`, `CCMP MI, R22, $12, $13`,
`FCCMPD AL, F8, F26, $0`.
- The exclusive stores are `<Rf>, (<Rn>), <Rs>` with the status register
last: `STLXR ZR, (R15), R16`.
- `TBZ` and `TBNZ` are `$<imm>, <Rt>, <label>`.
Shifted and extended register operands ride the register: `R19>>30`,
`R26->24` for arithmetic right shift, `@>` for rotate, and the extend forms
`R19.UXTB<<4`, `R14.SXTX` with extend operators UXTB, UXTH, UXTW, UXTX,
SXTB, SXTH, SXTW, SXTX.
## Conditions, branches and names
- Conditions ride the branch mnemonic: `B.EQ`, or the canonical
per-condition names such as `BEQ`. Both spellings exist; the canonical
names are what the generated inventory lists.
- `br` is `JMP` and `blr` is `CALL` in this dialect; indirect branches are
`JMP (R3)` and `CALL (R17)`.
- `NOP` is a zero-width pseudo-instruction; the hardware nop is `NOOP`,
an alias of `HINT $0`.
- `umov` is written as `VMOV`.
## Constants
- A 16-bit immediate optionally shifted: `MOVK $(10<<32), R20`, with
`MOVZ`, `MOVN` and their W variants; a zero shift is rejected by the
assembler.
- Large integer constants: `MOV` materialises any 64-bit constant, the
closest-instruction way.
- Vector constants: `VMOVS`, `VMOVD` and `VMOVQ`, the last taking two
64-bit halves for a 128-bit value:
`VMOVQ $0x1122334455667788, $0x99aabbccddeeff00, V2`.
## SIMD
Floating-point and SIMD instructions mostly carry a `V` prefix
(`VADD`, `VFMLA`), the cryptographic extensions (`AESD`, `SHA256H`) and the
scalar floating-point instructions being the exceptions. Operands carry an
arrangement suffix, `V5.H8`, and structure loads and stores use bracket
lists, `[V21.B16]`, with element selection as `V9.S[1]`.
## Alignment
`PCALIGN $n` pads to a power-of-two boundary between 8 and 2048 and also
raises the function's alignment to the coarsest boundary any of its PCALIGN
directives asks for. Functions default to 16-byte alignment on this target.
## Relocations
`R_ADDRARM64` for the adrp-plus-add pair, `R_ARM64_PCREL` and the
`R_ARM64_PCREL_LDST` family for PC relative addressing, `R_ARM64_LDST` for
the load and store immediates, `R_ARM64_GOTPCREL` and `R_ARM64_GOT` for the
GOT, `R_ARM64_TLS_LE` and `R_ARM64_TLS_IE` for thread local storage and
`R_CALLARM64` for direct calls, all specified in
[GOOBJ.md](../GOOBJ.md).
+94
View File
@@ -0,0 +1,94 @@
# LoongArch 64
Layer 1, target page. Verified against `go tool asm` of Go 1.27.1, against
the toolchain's own loong64 assembler manual (`cmd/internal/obj/loong64/doc.go`)
and against gasm's encoder, whose output is compared byte for byte with the
toolchain's. The complete mnemonic inventory lives in the generated appendix
[INSTRUCTIONS-LOONG64.md](INSTRUCTIONS-LOONG64.md).
## Registers
- General purpose `R0` to `R31`, floating point `F0` to `F31`, LSX vectors
`V0` to `V31` and LASX vectors `X0` to `X31`.
- Fixed roles from the toolchain's table: `R0` is the constant zero, `R1`
the return address, `R3` the stack pointer, `R22` the goroutine pointer,
`R29` the closure context and `R30` the assembler's temporary. `R12`,
`R13`, `R14`, `R15` and `R20` serve the PLT and trampoline sequences:
usable in assembly, but saved before any call.
## Widths ride the mnemonic
| Suffix | Width |
|---|---|
| `B`, `BU` | 8-bit, 8-bit unsigned |
| `H`, `HU` | 16-bit, 16-bit unsigned |
| `W`, `WU` | 32-bit, 32-bit unsigned |
| `V` | 64-bit |
| `F`, `D` | 32-bit and 64-bit float |
| `V` prefix (LSX) | 128-bit vector |
| `XV` prefix (LASX) | 256-bit vector |
The MOV series is the load and store interface: `MOVB (R2), R3` loads a
byte, `MOVV (R2), R3` a double word, `VMOVQ (R2), V1` a 128-bit vector and
`XVMOVQ (R2), X1` a 256-bit one.
## Operand order
Most instructions appear in left-to-right assignment order: `ADDV R11, R12,
R13` is `add.d R13, R12, R11`, and the two-operand form
`OR R5, R6` assigns into R6. Exceptions:
- Jump and branch instructions keep the GNU order: `BEQ R0, R4, label1`.
- The bitfield family is `BSTRINSW`, `BSTRINSV`, `BSTRPICKW`, `BSTRPICKV`
`$<msb>, <Rj>, $<lsb>, <Rd>`.
## Addressing
- Plain: `offset(Rbase)`.
- Base plus offset **register**, no scale: `(R4)(R5)`, as in
`MOVB (R4)(R5), R6`, the `ldx` family.
- The pointer loads and stores `MOVWP` and `MOVVP` take a source-level
16-bit offset that the encoder halves into the 14-bit field, writing
`MOVWP 8(R4), R5` as `ldptr.w r5, r4, $2`.
## Vector element syntax
The `VMOVQ` and `XVMOVQ` transfer family covers register-to-vector moves
with arrangement and index suffixes: `VMOVQ Rj, Vd.B[index]` inserts a
general register into one lane, `VMOVQ Vj.B[index], Rd` extracts one,
`VMOVQ Rj, Vd.B16` broadcasts across all sixteen, and `VMOVQ Vj.B[index],
Vd.B16` replicates one lane. The broadcast-from-memory form takes the true
byte offset at source level, which the encoder rescales per arrangement.
The permute and extract families take their 8-bit control word first:
`VPERMIW ui8, Vj, Vd`, `VEXTRINSB ui8, Vj, Vd`.
## Alignment
`PCALIGN $n` pads with NOOP to a power-of-two boundary between 8 and 2048,
and this target additionally auto-aligns loop heads to 16 bytes.
## Atomics, barriers and prefetch
- The `AM` atomic family comes in plain and `_DB` flavours; the `_DB`
forms, such as `AMSWAPDBW`, complete the atomic sequence and act as a
full data barrier. Within the AM family the destination and base
registers may not coincide and the destination may not equal the operand
register: one is an exception, the other silently unspecified.
- `DBAR` carries the graded hint encoding documented for LA664 and later,
with hint 0x700 as the read-after-read lightweight barrier; older cores
treat every hint as the full barrier.
- `PRELD offset(Rbase), $hint` prefetches with the documented hints (0
load to L1, 2 load to L3, 8 store to L1); `PRELDX` adds the encoded
block descriptor.
- `ALSL`-family shift-and-add writes the desired shift amount in source and
encodes one less: `ALSLV $4, R4, R5, R6` shifts by 4.
- `ADDV16 si16<<16, Rj, Rd` is the high-immediate add paired with the
pointer loads for GOT relative access.
## Relocations
`R_CALLLOONG64` for the 28-bit BL, `R_LOONG64_CALL36` for the
PCADDU18I-plus-JIRL pair, the `R_LOONG64_ADDR`, `ADDR64`, `TLS_LE`, `TLS_IE`,
`GOT` and `GOT64` high and low pairs, the aligned conditional jump forms
`R_JMP16LOONG64` and `R_JMP21LOONG64`, and `R_LOONG64_ADD64` and `SUB64`
for in-place arithmetic, all specified in [GOOBJ.md](../GOOBJ.md).
+9 -6
View File
@@ -41,13 +41,16 @@ layers 1 or 2; it extends them.
| [DIRECTIVES.md](DIRECTIVES.md) | 1 | TEXT, DATA, GLOBL, FUNCDATA, PCDATA, PCALIGN and the function flags |
| [PREPROCESSOR.md](PREPROCESSOR.md) | 1 | `#include`, `#define`, `#ifdef` and friends, `-D`, `-I` |
| [RUNTIME.md](RUNTIME.md) | 2 | ABI0, prototypes, `go_asm.h`, `funcdata.h`, `go vet` |
| AMD64, ARM64, RISCV64, LOONG64 | 1, 3 | per architecture: registers, conventions, addressing, instruction families and the generated instruction appendices |
| [AMD64.md](AMD64.md) | 1 | registers, addressing, the frame and split check, families, relocations |
| [ARM64.md](ARM64.md) | 1 | registers, the MOV load and store series, special operand orders, SIMD |
| [RISCV64.md](RISCV64.md) | 1 | registers and their constrained names, per class operand order, profiles, vector extension |
| [LOONG64.md](LOONG64.md) | 1 | registers, width suffixes, vector element syntax, atomics and barriers |
| INSTRUCTIONS-AMD64.md and the other three | 1 | generated per architecture inventory of every accepted mnemonic |
| STANDALONE.md | 3 | the language outside Go |
## Status
The common-language core and the Go-embedded layer are written and verified.
The four per-architecture pages and their generated instruction appendices
follow, architecture by architecture; STANDALONE.md lands with the standalone
compilation phase. The object format these pages feed is specified in
[GOOBJ.md](../GOOBJ.md).
The common-language core, the Go-embedded layer, all four per-architecture
pages and the generated instruction appendices are written and verified.
STANDALONE.md lands with the standalone compilation phase. The object format
these pages feed is specified in [GOOBJ.md](../GOOBJ.md).
+104
View File
@@ -0,0 +1,104 @@
# RISC-V 64
Layer 1, target page. Verified against `go tool asm` of Go 1.27.1, against
the toolchain's own riscv64 assembler manual (`cmd/internal/obj/riscv/doc.go`)
and against gasm's encoder, whose output is compared byte for byte with the
toolchain's. The complete mnemonic inventory lives in the generated appendix
[INSTRUCTIONS-RISCV64.md](INSTRUCTIONS-RISCV64.md).
## Registers
- Integer: `X0` to `X31`. `X0` is hardwired zero. Three names the toolchain
constrains: `X4` must be written through its ABI name `TP`; `X27`, the
goroutine pointer, must be written `g` and may not be written `S11`; in
shared builds `X3` is off limits and must be written `GP`.
- The other integer registers may be written `Xn` or by their ABI names
(`A0`, `T0`, `S1`, and so on).
- Floating point: `F0` to `F31`. Vector: `V0` to `V31`.
- `X26` is the closure pointer and `X31` is the assembler's own scratch
register: its value may be clobbered by instruction sequences the
assembler inserts, so hand-written code must not rely on it.
- There is no reserved frame pointer register on this target.
## Operand order
The ordering differs from the ISA manual, and per instruction class:
- **R-type** is reversed: `ADD X10, X11, X12` is `add x12, x11, x10`.
- **I-type arithmetic** keeps that shape with the immediate first:
`ADDI $1, X11, X12`.
- **Loads and stores** are source first, like every Plan 9 dialect:
`MOV 16(X2), X10` loads and `MOV X10, (X2)` stores. The MOV series hides
the width; `MOVB` through `MOVD` spell it out.
- **Branches** keep the ISA order: `BLT X12, X23, loop1`, which jumps when
X12 < X23, the reverse of the SLT operand order.
- **FMA** is rotated one place left so the destination comes last:
`FMADDS F1, F2, F3, F4`.
- **AMO** is likewise rotated: `AMOSWAPW X5, (X6), X7`.
- **Ternary abbreviation** is supported and encouraged: `ADD X10, X12` means
`ADD X10, X12, X12`.
Where an R-type instruction has an I-type sibling, the assembler picks the
immediate form from the operand: `AND $3, X12, X13` assembles as `ANDI`.
## Names, suffixes and rounding
Dots are removed and suffixes are upper-cased: the ISA's `fmv.w.x` is
`FMVWX`. Floating-point rounding modes become suffixes, `FCVTLUS.RNE F0,
X5`, with RTZ assumed when the suffix is omitted; the toolchain never sets
the FCSR.
## Constants
- `MOV` materialises any 64-bit integer constant, synthesising it from a
few arithmetic instructions where possible and otherwise loading it from
a literal pool in the binary.
- A 32-bit constant is accepted by `ADDI`, `ANDI`, `ORI` and `XORI`, and
the assembler synthesises values that exceed the 12-bit encoding window.
- `MOVF` and `MOVD` materialise floating-point constants, encoding them as
`FLW` and `FLD` from a pool location unless the constant is exactly 0.0.
## Extensions and profiles
The default target profile is rva20u64, selected or raised with the
GORISCV64 environment variable. A short list of instructions outside the
default profile is synthesised by the assembler when the profile does not
provide them, so they are safe without guards: `ANDN`, `MAX`, `MAXU`, `MIN`,
`MINU`, `MOVB`, `MOVH`, `MOVHU`, `MOVWU`, `ORN`, `ROL`, `ROLW`, `ROR`,
`RORI`, `RORIW`, `RORW`, `XNOR`. The header `asm_riscv64.h` defines the
`hasZba`, `hasZbb`, `hasZbs` and `hasV` macros for guarding everything else.
## Fences and atomics
`FENCE` takes predecessor and successor sets in that order, uppercase
letters, `FENCE R, RW`; a bare `FENCE` is a full fence, as is
`FENCE IORW, IORW`. `FENCE.TSO` exists. The ordering bits of `LR`, `SC`
and the AMO instructions are not specifiable in source: the assembler sets
acquire and release on the AMO instructions, acquire on `LR` and release on
`SC`, always.
## Compressed instructions
The assembler converts 32-bit instructions to their compressed encodings
automatically; the conversion is a property of the emitted machine code, not
of the source, and register choice influences how much compresses.
Hand-writing compressed instructions in source is accepted but discouraged.
The debug flag `compressinstructions=0` turns the automatic conversion off.
## Vector extension
`VSETVLI` writes its vtype components in uppercase with the destination
last: `VSETVLI X10, E8, M1, TU, MU, X12`. Vector loads and stores are
source first like the scalar ones, with an optional stride or index register
second and the mask register, when present, always penultimate:
`VLE8V (X10), V3`, `VLE8V (X10), V0, V3` for the masked form. Vector
arithmetic reverses its operands, `VADDVV V1, V2, V3`, with the mask again
penultimate.
## Relocations
`R_RISCV_JAL`, `R_RISCV_CALL`, the `R_RISCV_PCREL_ITYPE` and `STYPE` pairs,
`R_RISCV_BRANCH`, the compressed branch and jump forms, the TLS and GOT
families and `R_RISCV_ADD32` and `SUB32`, all specified in
[GOOBJ.md](../GOOBJ.md). The assembler always emits the four-byte
`R_DWTXTADDR_U4` flavour inside its DWARF records.