5.0 KiB
ARM64
Layer 1, target page. Verified against go tool asm of Go 1.27.1, against the
toolchain's own arm64 assembler manual (cmd/internal/obj/arm64/doc.go) and
against gasm's encoder, whose output is compared byte for byte with the
toolchain's. The complete mnemonic inventory lives in the generated appendix
INSTRUCTIONS-ARM64.md.
Registers
- General purpose:
R0toR30, plusZR, the zero register, andRSP, the stack pointer. There is no R31: thirty-one names and ZR. - Floating-point and SIMD share one file written
Vn; where an instruction is scalar floating point the operand may be writtenFn(F0toF31). - SVE register names (
Z0toZ31,P0toP15) exist in the assembler's tables. - Roles the convention fixes:
RSPis the stack pointer,R29the frame pointer,R30the link register,R26the closure context andR27the assembler's scratch register. The goroutine pointer lives inR28and is writtengin source, its fields asg_m(g),g_sched(g);R18is the platform-reserved register and the Go toolchain never addresses it.
Loads, stores and the width suffixes
The MOV series is the load and store interface, with the width in the mnemonic rather than the register name:
| Mnemonic | Machine instruction |
|---|---|
MOVD |
ldr, str, stur, 64-bit |
MOVW |
ldrsw, str, stur, 32-bit sign extending |
MOVWU |
ldr, 32-bit zero extending |
MOVH |
ldrsh, strh, sturh |
MOVHU |
ldrh |
MOVB |
ldrsb, strb, sturb |
MOVBU |
ldrb |
Post-index and pre-index addressing take the .P and .W suffixes on the
mnemonic: MOVD.P -8(R10), R8 is ldr x8, [x10],#-8, and MOVB.W 16(R16), R10 is ldrsb x10, [x16,#16]!.
Addressing
imm(Rn|RSP) 28(R17)
(Rn|RSP) (R22)
(Rn)(Rm) (R27)(R23)
(Rn)(Rm<<scale) (R4)(R12<<2)
(Rn)(Rm.UXTW<<3) extended and shifted index
(Rt1, Rt2) register pair for LDP, STP and the exclusive pair forms
Branch targets are labels, (R3) for indirect, name(SB) for static.
Operand order and the special forms
Most instructions appear in left-to-right assignment order: ADD R11, RSP, R25 computes into R25. The exceptions the toolchain's manual lists,
each with its own order:
- stores and
CBZ,CBNZkeep the GNU order:MOVD R29, 384(R19). - The multiply-accumulate family
MADD,MSUB,SMADDLand friends are<Rm>, <Ra>, <Rn>, <Rd>. - The scalar FMA family
FMADDDand friends are<Fm>, <Fa>, <Fn>, <Fd>. - The bitfield family
BFI,BFXIL,SBFIZ,SBFX,UBFIZ,UBFXis$<lsb>, <Rn>, $<width>, <Rd>. - The conditional compare and select families carry the condition as the
first operand:
CSEL GT, R0, R19, R1,CCMP MI, R22, $12, $13,FCCMPD AL, F8, F26, $0. - The exclusive stores are
<Rf>, (<Rn>), <Rs>with the status register last:STLXR ZR, (R15), R16. TBZandTBNZare$<imm>, <Rt>, <label>.
Shifted and extended register operands ride the register: R19>>30,
R26->24 for arithmetic right shift, @> for rotate, and the extend forms
R19.UXTB<<4, R14.SXTX with extend operators UXTB, UXTH, UXTW, UXTX,
SXTB, SXTH, SXTW, SXTX.
Conditions, branches and names
- Conditions ride the branch mnemonic:
B.EQ, or the canonical per-condition names such asBEQ. Both spellings exist; the canonical names are what the generated inventory lists. brisJMPandblrisCALLin this dialect; indirect branches areJMP (R3)andCALL (R17).NOPis a zero-width pseudo-instruction; the hardware nop isNOOP, an alias ofHINT $0.umovis written asVMOV.
Constants
- A 16-bit immediate optionally shifted:
MOVK $(10<<32), R20, withMOVZ,MOVNand their W variants; a zero shift is rejected by the assembler. - Large integer constants:
MOVmaterialises any 64-bit constant, the closest-instruction way. - Vector constants:
VMOVS,VMOVDandVMOVQ, the last taking two 64-bit halves for a 128-bit value:VMOVQ $0x1122334455667788, $0x99aabbccddeeff00, V2.
SIMD
Floating-point and SIMD instructions mostly carry a V prefix
(VADD, VFMLA), the cryptographic extensions (AESD, SHA256H) and the
scalar floating-point instructions being the exceptions. Operands carry an
arrangement suffix, V5.H8, and structure loads and stores use bracket
lists, [V21.B16], with element selection as V9.S[1].
Alignment
PCALIGN $n pads to a power-of-two boundary between 8 and 2048 and also
raises the function's alignment to the coarsest boundary any of its PCALIGN
directives asks for. Functions default to 16-byte alignment on this target.
Relocations
R_ADDRARM64 for the adrp-plus-add pair, R_ARM64_PCREL and the
R_ARM64_PCREL_LDST family for PC relative addressing, R_ARM64_LDST for
the load and store immediates, R_ARM64_GOTPCREL and R_ARM64_GOT for the
GOT, R_ARM64_TLS_LE and R_ARM64_TLS_IE for thread local storage and
R_CALLARM64 for direct calls, all specified in
GOOBJ.md.