Assisted-by: GLM 5.3 Flash
This commit is contained in:
@@ -0,0 +1,122 @@
|
||||
# ARM64
|
||||
|
||||
Layer 1, target page. Verified against `go tool asm` of Go 1.27.1, against the
|
||||
toolchain's own arm64 assembler manual (`cmd/internal/obj/arm64/doc.go`) and
|
||||
against gasm's encoder, whose output is compared byte for byte with the
|
||||
toolchain's. The complete mnemonic inventory lives in the generated appendix
|
||||
[INSTRUCTIONS-ARM64.md](INSTRUCTIONS-ARM64.md).
|
||||
|
||||
## Registers
|
||||
|
||||
- General purpose: `R0` to `R30`, plus `ZR`, the zero register, and `RSP`,
|
||||
the stack pointer. There is no R31: thirty-one names and ZR.
|
||||
- Floating-point and SIMD share one file written `Vn`; where an instruction
|
||||
is scalar floating point the operand may be written `Fn` (`F0` to `F31`).
|
||||
- SVE register names (`Z0` to `Z31`, `P0` to `P15`) exist in the assembler's
|
||||
tables.
|
||||
- Roles the convention fixes: `RSP` is the stack pointer, `R29` the frame
|
||||
pointer, `R30` the link register, `R26` the closure context and `R27` the
|
||||
assembler's scratch register. The goroutine pointer lives in `R28` and is
|
||||
written `g` in source, its fields as `g_m(g)`, `g_sched(g)`; `R18` is the
|
||||
platform-reserved register and the Go toolchain never addresses it.
|
||||
|
||||
## Loads, stores and the width suffixes
|
||||
|
||||
The MOV series is the load and store interface, with the width in the
|
||||
mnemonic rather than the register name:
|
||||
|
||||
| Mnemonic | Machine instruction |
|
||||
|---|---|
|
||||
| `MOVD` | ldr, str, stur, 64-bit |
|
||||
| `MOVW` | ldrsw, str, stur, 32-bit sign extending |
|
||||
| `MOVWU` | ldr, 32-bit zero extending |
|
||||
| `MOVH` | ldrsh, strh, sturh |
|
||||
| `MOVHU` | ldrh |
|
||||
| `MOVB` | ldrsb, strb, sturb |
|
||||
| `MOVBU` | ldrb |
|
||||
|
||||
Post-index and pre-index addressing take the `.P` and `.W` suffixes on the
|
||||
mnemonic: `MOVD.P -8(R10), R8` is `ldr x8, [x10],#-8`, and `MOVB.W
|
||||
16(R16), R10` is `ldrsb x10, [x16,#16]!`.
|
||||
|
||||
## Addressing
|
||||
|
||||
```text
|
||||
imm(Rn|RSP) 28(R17)
|
||||
(Rn|RSP) (R22)
|
||||
(Rn)(Rm) (R27)(R23)
|
||||
(Rn)(Rm<<scale) (R4)(R12<<2)
|
||||
(Rn)(Rm.UXTW<<3) extended and shifted index
|
||||
(Rt1, Rt2) register pair for LDP, STP and the exclusive pair forms
|
||||
```
|
||||
|
||||
Branch targets are labels, `(R3)` for indirect, `name(SB)` for static.
|
||||
|
||||
## Operand order and the special forms
|
||||
|
||||
Most instructions appear in left-to-right assignment order: `ADD R11,
|
||||
RSP, R25` computes into R25. The exceptions the toolchain's manual lists,
|
||||
each with its own order:
|
||||
|
||||
- stores and `CBZ`, `CBNZ` keep the GNU order: `MOVD R29, 384(R19)`.
|
||||
- The multiply-accumulate family `MADD`, `MSUB`, `SMADDL` and friends are
|
||||
`<Rm>, <Ra>, <Rn>, <Rd>`.
|
||||
- The scalar FMA family `FMADDD` and friends are `<Fm>, <Fa>, <Fn>, <Fd>`.
|
||||
- The bitfield family `BFI`, `BFXIL`, `SBFIZ`, `SBFX`, `UBFIZ`, `UBFX` is
|
||||
`$<lsb>, <Rn>, $<width>, <Rd>`.
|
||||
- The conditional compare and select families carry the condition as the
|
||||
**first** operand: `CSEL GT, R0, R19, R1`, `CCMP MI, R22, $12, $13`,
|
||||
`FCCMPD AL, F8, F26, $0`.
|
||||
- The exclusive stores are `<Rf>, (<Rn>), <Rs>` with the status register
|
||||
last: `STLXR ZR, (R15), R16`.
|
||||
- `TBZ` and `TBNZ` are `$<imm>, <Rt>, <label>`.
|
||||
|
||||
Shifted and extended register operands ride the register: `R19>>30`,
|
||||
`R26->24` for arithmetic right shift, `@>` for rotate, and the extend forms
|
||||
`R19.UXTB<<4`, `R14.SXTX` with extend operators UXTB, UXTH, UXTW, UXTX,
|
||||
SXTB, SXTH, SXTW, SXTX.
|
||||
|
||||
## Conditions, branches and names
|
||||
|
||||
- Conditions ride the branch mnemonic: `B.EQ`, or the canonical
|
||||
per-condition names such as `BEQ`. Both spellings exist; the canonical
|
||||
names are what the generated inventory lists.
|
||||
- `br` is `JMP` and `blr` is `CALL` in this dialect; indirect branches are
|
||||
`JMP (R3)` and `CALL (R17)`.
|
||||
- `NOP` is a zero-width pseudo-instruction; the hardware nop is `NOOP`,
|
||||
an alias of `HINT $0`.
|
||||
- `umov` is written as `VMOV`.
|
||||
|
||||
## Constants
|
||||
|
||||
- A 16-bit immediate optionally shifted: `MOVK $(10<<32), R20`, with
|
||||
`MOVZ`, `MOVN` and their W variants; a zero shift is rejected by the
|
||||
assembler.
|
||||
- Large integer constants: `MOV` materialises any 64-bit constant, the
|
||||
closest-instruction way.
|
||||
- Vector constants: `VMOVS`, `VMOVD` and `VMOVQ`, the last taking two
|
||||
64-bit halves for a 128-bit value:
|
||||
`VMOVQ $0x1122334455667788, $0x99aabbccddeeff00, V2`.
|
||||
|
||||
## SIMD
|
||||
|
||||
Floating-point and SIMD instructions mostly carry a `V` prefix
|
||||
(`VADD`, `VFMLA`), the cryptographic extensions (`AESD`, `SHA256H`) and the
|
||||
scalar floating-point instructions being the exceptions. Operands carry an
|
||||
arrangement suffix, `V5.H8`, and structure loads and stores use bracket
|
||||
lists, `[V21.B16]`, with element selection as `V9.S[1]`.
|
||||
|
||||
## Alignment
|
||||
|
||||
`PCALIGN $n` pads to a power-of-two boundary between 8 and 2048 and also
|
||||
raises the function's alignment to the coarsest boundary any of its PCALIGN
|
||||
directives asks for. Functions default to 16-byte alignment on this target.
|
||||
|
||||
## Relocations
|
||||
|
||||
`R_ADDRARM64` for the adrp-plus-add pair, `R_ARM64_PCREL` and the
|
||||
`R_ARM64_PCREL_LDST` family for PC relative addressing, `R_ARM64_LDST` for
|
||||
the load and store immediates, `R_ARM64_GOTPCREL` and `R_ARM64_GOT` for the
|
||||
GOT, `R_ARM64_TLS_LE` and `R_ARM64_TLS_IE` for thread local storage and
|
||||
`R_CALLARM64` for direct calls, all specified in
|
||||
[GOOBJ.md](../GOOBJ.md).
|
||||
Reference in New Issue
Block a user