2026-09-21 20:11:39 +02:00
|
|
|
# ARM64
|
|
|
|
|
|
|
|
|
|
Layer 1, target page. Verified against `go tool asm` of Go 1.27.1, against the
|
|
|
|
|
toolchain's own arm64 assembler manual (`cmd/internal/obj/arm64/doc.go`) and
|
|
|
|
|
against gasm's encoder, whose output is compared byte for byte with the
|
|
|
|
|
toolchain's. The complete mnemonic inventory lives in the generated appendix
|
|
|
|
|
[INSTRUCTIONS-ARM64.md](INSTRUCTIONS-ARM64.md).
|
|
|
|
|
|
|
|
|
|
## Registers
|
|
|
|
|
|
|
|
|
|
- General purpose: `R0` to `R30`, plus `ZR`, the zero register, and `RSP`,
|
|
|
|
|
the stack pointer. There is no R31: thirty-one names and ZR.
|
|
|
|
|
- Floating-point and SIMD share one file written `Vn`; where an instruction
|
|
|
|
|
is scalar floating point the operand may be written `Fn` (`F0` to `F31`).
|
2026-10-07 22:37:42 +02:00
|
|
|
- SVE registers `Z0` to `Z31` and predicates `P0` to `P15`; the SVE2 and
|
|
|
|
|
SVE2.1 families encode, see SVE below.
|
2026-09-21 20:11:39 +02:00
|
|
|
- Roles the convention fixes: `RSP` is the stack pointer, `R29` the frame
|
|
|
|
|
pointer, `R30` the link register, `R26` the closure context and `R27` the
|
|
|
|
|
assembler's scratch register. The goroutine pointer lives in `R28` and is
|
|
|
|
|
written `g` in source, its fields as `g_m(g)`, `g_sched(g)`; `R18` is the
|
|
|
|
|
platform-reserved register and the Go toolchain never addresses it.
|
|
|
|
|
|
|
|
|
|
## Loads, stores and the width suffixes
|
|
|
|
|
|
|
|
|
|
The MOV series is the load and store interface, with the width in the
|
|
|
|
|
mnemonic rather than the register name:
|
|
|
|
|
|
|
|
|
|
| Mnemonic | Machine instruction |
|
|
|
|
|
|---|---|
|
|
|
|
|
| `MOVD` | ldr, str, stur, 64-bit |
|
|
|
|
|
| `MOVW` | ldrsw, str, stur, 32-bit sign extending |
|
|
|
|
|
| `MOVWU` | ldr, 32-bit zero extending |
|
|
|
|
|
| `MOVH` | ldrsh, strh, sturh |
|
|
|
|
|
| `MOVHU` | ldrh |
|
|
|
|
|
| `MOVB` | ldrsb, strb, sturb |
|
|
|
|
|
| `MOVBU` | ldrb |
|
|
|
|
|
|
|
|
|
|
Post-index and pre-index addressing take the `.P` and `.W` suffixes on the
|
|
|
|
|
mnemonic: `MOVD.P -8(R10), R8` is `ldr x8, [x10],#-8`, and `MOVB.W
|
|
|
|
|
16(R16), R10` is `ldrsb x10, [x16,#16]!`.
|
|
|
|
|
|
|
|
|
|
## Addressing
|
|
|
|
|
|
|
|
|
|
```text
|
|
|
|
|
imm(Rn|RSP) 28(R17)
|
|
|
|
|
(Rn|RSP) (R22)
|
|
|
|
|
(Rn)(Rm) (R27)(R23)
|
|
|
|
|
(Rn)(Rm<<scale) (R4)(R12<<2)
|
|
|
|
|
(Rn)(Rm.UXTW<<3) extended and shifted index
|
|
|
|
|
(Rt1, Rt2) register pair for LDP, STP and the exclusive pair forms
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
Branch targets are labels, `(R3)` for indirect, `name(SB)` for static.
|
|
|
|
|
|
|
|
|
|
## Operand order and the special forms
|
|
|
|
|
|
|
|
|
|
Most instructions appear in left-to-right assignment order: `ADD R11,
|
|
|
|
|
RSP, R25` computes into R25. The exceptions the toolchain's manual lists,
|
|
|
|
|
each with its own order:
|
|
|
|
|
|
|
|
|
|
- stores and `CBZ`, `CBNZ` keep the GNU order: `MOVD R29, 384(R19)`.
|
|
|
|
|
- The multiply-accumulate family `MADD`, `MSUB`, `SMADDL` and friends are
|
|
|
|
|
`<Rm>, <Ra>, <Rn>, <Rd>`.
|
|
|
|
|
- The scalar FMA family `FMADDD` and friends are `<Fm>, <Fa>, <Fn>, <Fd>`.
|
|
|
|
|
- The bitfield family `BFI`, `BFXIL`, `SBFIZ`, `SBFX`, `UBFIZ`, `UBFX` is
|
|
|
|
|
`$<lsb>, <Rn>, $<width>, <Rd>`.
|
|
|
|
|
- The conditional compare and select families carry the condition as the
|
|
|
|
|
**first** operand: `CSEL GT, R0, R19, R1`, `CCMP MI, R22, $12, $13`,
|
|
|
|
|
`FCCMPD AL, F8, F26, $0`.
|
|
|
|
|
- The exclusive stores are `<Rf>, (<Rn>), <Rs>` with the status register
|
|
|
|
|
last: `STLXR ZR, (R15), R16`.
|
|
|
|
|
- `TBZ` and `TBNZ` are `$<imm>, <Rt>, <label>`.
|
|
|
|
|
|
|
|
|
|
Shifted and extended register operands ride the register: `R19>>30`,
|
|
|
|
|
`R26->24` for arithmetic right shift, `@>` for rotate, and the extend forms
|
|
|
|
|
`R19.UXTB<<4`, `R14.SXTX` with extend operators UXTB, UXTH, UXTW, UXTX,
|
|
|
|
|
SXTB, SXTH, SXTW, SXTX.
|
|
|
|
|
|
|
|
|
|
## Conditions, branches and names
|
|
|
|
|
|
|
|
|
|
- Conditions ride the branch mnemonic: `B.EQ`, or the canonical
|
|
|
|
|
per-condition names such as `BEQ`. Both spellings exist; the canonical
|
|
|
|
|
names are what the generated inventory lists.
|
|
|
|
|
- `br` is `JMP` and `blr` is `CALL` in this dialect; indirect branches are
|
|
|
|
|
`JMP (R3)` and `CALL (R17)`.
|
|
|
|
|
- `NOP` is a zero-width pseudo-instruction; the hardware nop is `NOOP`,
|
|
|
|
|
an alias of `HINT $0`.
|
|
|
|
|
- `umov` is written as `VMOV`.
|
|
|
|
|
|
|
|
|
|
## Constants
|
|
|
|
|
|
|
|
|
|
- A 16-bit immediate optionally shifted: `MOVK $(10<<32), R20`, with
|
|
|
|
|
`MOVZ`, `MOVN` and their W variants; a zero shift is rejected by the
|
|
|
|
|
assembler.
|
|
|
|
|
- Large integer constants: `MOV` materialises any 64-bit constant, the
|
|
|
|
|
closest-instruction way.
|
|
|
|
|
- Vector constants: `VMOVS`, `VMOVD` and `VMOVQ`, the last taking two
|
|
|
|
|
64-bit halves for a 128-bit value:
|
|
|
|
|
`VMOVQ $0x1122334455667788, $0x99aabbccddeeff00, V2`.
|
|
|
|
|
|
|
|
|
|
## SIMD
|
|
|
|
|
|
|
|
|
|
Floating-point and SIMD instructions mostly carry a `V` prefix
|
|
|
|
|
(`VADD`, `VFMLA`), the cryptographic extensions (`AESD`, `SHA256H`) and the
|
|
|
|
|
scalar floating-point instructions being the exceptions. Operands carry an
|
|
|
|
|
arrangement suffix, `V5.H8`, and structure loads and stores use bracket
|
|
|
|
|
lists, `[V21.B16]`, with element selection as `V9.S[1]`.
|
|
|
|
|
|
2026-10-07 22:37:42 +02:00
|
|
|
## SVE
|
|
|
|
|
|
|
|
|
|
The scalable vector extension encodes through the `Z0` to `Z31` vector
|
|
|
|
|
registers and the `P0` to `P15` predicates, with arrangement suffixes on
|
|
|
|
|
the `Z` registers and the merging and zeroing qualifiers the toolchain's
|
|
|
|
|
SVE2 and SVE2.1 families carry: the narrowing two-to-one arithmetic, the
|
|
|
|
|
cryptographic set including ZADCLB, BFloat16 arithmetic, the predicate
|
|
|
|
|
counters and reductions, the pairwise and quadword forms, the
|
|
|
|
|
multiple-structure loads and stores, shift by vector and the
|
|
|
|
|
shift-immediate scheme, CLASTA and CLASTB, the compare and last-active
|
|
|
|
|
predicate families, and the vector-length arithmetic `ADDVL`, `ADDPL` and
|
|
|
|
|
`RDVL`, whose immediates count vectors or predicates rather than bytes.
|
|
|
|
|
The gather and scatter loads take their addresses through the five
|
|
|
|
|
addressing modes the toolchain defines. The source grammar is the
|
|
|
|
|
toolchain's own, and the byte output is pinned against it corpus-wide.
|
|
|
|
|
|
|
|
|
|
## Synthesised forms
|
|
|
|
|
|
|
|
|
|
- `GETCALLERPC` reads the return address the frame state describes: a leaf
|
|
|
|
|
function reads `R30`, a framed body reads the prologue's save slot.
|
|
|
|
|
- `REM`, `REMW`, `UREM` and `UREMW` synthesise a remainder from `SDIV` or
|
|
|
|
|
`UDIV` and the `MSUB` tail, with `RSP` refused as a destination.
|
|
|
|
|
- `DWORD $imm` lays eight little-endian bytes per immediate.
|
|
|
|
|
- `MOVD tls_g(SB), Rd` materialises a TLS local-exec load as a single
|
|
|
|
|
`MOVZ` carrying `R_ARM64_TLS_LE`, keyed off the file's own
|
|
|
|
|
`GLOBL ... TLSBSS` declaration.
|
|
|
|
|
|
2026-09-21 20:11:39 +02:00
|
|
|
## Alignment
|
|
|
|
|
|
|
|
|
|
`PCALIGN $n` pads to a power-of-two boundary between 8 and 2048 and also
|
|
|
|
|
raises the function's alignment to the coarsest boundary any of its PCALIGN
|
|
|
|
|
directives asks for. Functions default to 16-byte alignment on this target.
|
|
|
|
|
|
|
|
|
|
## Relocations
|
|
|
|
|
|
|
|
|
|
`R_ADDRARM64` for the adrp-plus-add pair, `R_ARM64_PCREL` and the
|
|
|
|
|
`R_ARM64_PCREL_LDST` family for PC relative addressing, `R_ARM64_LDST` for
|
|
|
|
|
the load and store immediates, `R_ARM64_GOTPCREL` and `R_ARM64_GOT` for the
|
|
|
|
|
GOT, `R_ARM64_TLS_LE` and `R_ARM64_TLS_IE` for thread local storage and
|
|
|
|
|
`R_CALLARM64` for direct calls, all specified in
|
|
|
|
|
[GOOBJ.md](../GOOBJ.md).
|