Files
gasm-sdk/docs/asm/ARM64.md
T

149 lines
6.4 KiB
Markdown
Raw Normal View History

# ARM64
Layer 1, target page. Verified against `go tool asm` of Go 1.27.1, against the
toolchain's own arm64 assembler manual (`cmd/internal/obj/arm64/doc.go`) and
against gasm's encoder, whose output is compared byte for byte with the
toolchain's. The complete mnemonic inventory lives in the generated appendix
[INSTRUCTIONS-ARM64.md](INSTRUCTIONS-ARM64.md).
## Registers
- General purpose: `R0` to `R30`, plus `ZR`, the zero register, and `RSP`,
the stack pointer. There is no R31: thirty-one names and ZR.
- Floating-point and SIMD share one file written `Vn`; where an instruction
is scalar floating point the operand may be written `Fn` (`F0` to `F31`).
2026-10-07 22:37:42 +02:00
- SVE registers `Z0` to `Z31` and predicates `P0` to `P15`; the SVE2 and
SVE2.1 families encode, see SVE below.
- Roles the convention fixes: `RSP` is the stack pointer, `R29` the frame
pointer, `R30` the link register, `R26` the closure context and `R27` the
assembler's scratch register. The goroutine pointer lives in `R28` and is
written `g` in source, its fields as `g_m(g)`, `g_sched(g)`; `R18` is the
platform-reserved register and the Go toolchain never addresses it.
## Loads, stores and the width suffixes
The MOV series is the load and store interface, with the width in the
mnemonic rather than the register name:
| Mnemonic | Machine instruction |
|---|---|
| `MOVD` | ldr, str, stur, 64-bit |
| `MOVW` | ldrsw, str, stur, 32-bit sign extending |
| `MOVWU` | ldr, 32-bit zero extending |
| `MOVH` | ldrsh, strh, sturh |
| `MOVHU` | ldrh |
| `MOVB` | ldrsb, strb, sturb |
| `MOVBU` | ldrb |
Post-index and pre-index addressing take the `.P` and `.W` suffixes on the
mnemonic: `MOVD.P -8(R10), R8` is `ldr x8, [x10],#-8`, and `MOVB.W
16(R16), R10` is `ldrsb x10, [x16,#16]!`.
## Addressing
```text
imm(Rn|RSP) 28(R17)
(Rn|RSP) (R22)
(Rn)(Rm) (R27)(R23)
(Rn)(Rm<<scale) (R4)(R12<<2)
(Rn)(Rm.UXTW<<3) extended and shifted index
(Rt1, Rt2) register pair for LDP, STP and the exclusive pair forms
```
Branch targets are labels, `(R3)` for indirect, `name(SB)` for static.
## Operand order and the special forms
Most instructions appear in left-to-right assignment order: `ADD R11,
RSP, R25` computes into R25. The exceptions the toolchain's manual lists,
each with its own order:
- stores and `CBZ`, `CBNZ` keep the GNU order: `MOVD R29, 384(R19)`.
- The multiply-accumulate family `MADD`, `MSUB`, `SMADDL` and friends are
`<Rm>, <Ra>, <Rn>, <Rd>`.
- The scalar FMA family `FMADDD` and friends are `<Fm>, <Fa>, <Fn>, <Fd>`.
- The bitfield family `BFI`, `BFXIL`, `SBFIZ`, `SBFX`, `UBFIZ`, `UBFX` is
`$<lsb>, <Rn>, $<width>, <Rd>`.
- The conditional compare and select families carry the condition as the
**first** operand: `CSEL GT, R0, R19, R1`, `CCMP MI, R22, $12, $13`,
`FCCMPD AL, F8, F26, $0`.
- The exclusive stores are `<Rf>, (<Rn>), <Rs>` with the status register
last: `STLXR ZR, (R15), R16`.
- `TBZ` and `TBNZ` are `$<imm>, <Rt>, <label>`.
Shifted and extended register operands ride the register: `R19>>30`,
`R26->24` for arithmetic right shift, `@>` for rotate, and the extend forms
`R19.UXTB<<4`, `R14.SXTX` with extend operators UXTB, UXTH, UXTW, UXTX,
SXTB, SXTH, SXTW, SXTX.
## Conditions, branches and names
- Conditions ride the branch mnemonic: `B.EQ`, or the canonical
per-condition names such as `BEQ`. Both spellings exist; the canonical
names are what the generated inventory lists.
- `br` is `JMP` and `blr` is `CALL` in this dialect; indirect branches are
`JMP (R3)` and `CALL (R17)`.
- `NOP` is a zero-width pseudo-instruction; the hardware nop is `NOOP`,
an alias of `HINT $0`.
- `umov` is written as `VMOV`.
## Constants
- A 16-bit immediate optionally shifted: `MOVK $(10<<32), R20`, with
`MOVZ`, `MOVN` and their W variants; a zero shift is rejected by the
assembler.
- Large integer constants: `MOV` materialises any 64-bit constant, the
closest-instruction way.
- Vector constants: `VMOVS`, `VMOVD` and `VMOVQ`, the last taking two
64-bit halves for a 128-bit value:
`VMOVQ $0x1122334455667788, $0x99aabbccddeeff00, V2`.
## SIMD
Floating-point and SIMD instructions mostly carry a `V` prefix
(`VADD`, `VFMLA`), the cryptographic extensions (`AESD`, `SHA256H`) and the
scalar floating-point instructions being the exceptions. Operands carry an
arrangement suffix, `V5.H8`, and structure loads and stores use bracket
lists, `[V21.B16]`, with element selection as `V9.S[1]`.
2026-10-07 22:37:42 +02:00
## SVE
The scalable vector extension encodes through the `Z0` to `Z31` vector
registers and the `P0` to `P15` predicates, with arrangement suffixes on
the `Z` registers and the merging and zeroing qualifiers the toolchain's
SVE2 and SVE2.1 families carry: the narrowing two-to-one arithmetic, the
cryptographic set including ZADCLB, BFloat16 arithmetic, the predicate
counters and reductions, the pairwise and quadword forms, the
multiple-structure loads and stores, shift by vector and the
shift-immediate scheme, CLASTA and CLASTB, the compare and last-active
predicate families, and the vector-length arithmetic `ADDVL`, `ADDPL` and
`RDVL`, whose immediates count vectors or predicates rather than bytes.
The gather and scatter loads take their addresses through the five
addressing modes the toolchain defines. The source grammar is the
toolchain's own, and the byte output is pinned against it corpus-wide.
## Synthesised forms
- `GETCALLERPC` reads the return address the frame state describes: a leaf
function reads `R30`, a framed body reads the prologue's save slot.
- `REM`, `REMW`, `UREM` and `UREMW` synthesise a remainder from `SDIV` or
`UDIV` and the `MSUB` tail, with `RSP` refused as a destination.
- `DWORD $imm` lays eight little-endian bytes per immediate.
- `MOVD tls_g(SB), Rd` materialises a TLS local-exec load as a single
`MOVZ` carrying `R_ARM64_TLS_LE`, keyed off the file's own
`GLOBL ... TLSBSS` declaration.
## Alignment
`PCALIGN $n` pads to a power-of-two boundary between 8 and 2048 and also
raises the function's alignment to the coarsest boundary any of its PCALIGN
directives asks for. Functions default to 16-byte alignment on this target.
## Relocations
`R_ADDRARM64` for the adrp-plus-add pair, `R_ARM64_PCREL` and the
`R_ARM64_PCREL_LDST` family for PC relative addressing, `R_ARM64_LDST` for
the load and store immediates, `R_ARM64_GOTPCREL` and `R_ARM64_GOT` for the
GOT, `R_ARM64_TLS_LE` and `R_ARM64_TLS_IE` for thread local storage and
`R_CALLARM64` for direct calls, all specified in
[GOOBJ.md](../GOOBJ.md).