Assisted-by: GLM 5.3 Flash
6.4 KiB
ARM64
Layer 1, target page. Verified against go tool asm of Go 1.27.1, against the
toolchain's own arm64 assembler manual (cmd/internal/obj/arm64/doc.go) and
against gasm's encoder, whose output is compared byte for byte with the
toolchain's. The complete mnemonic inventory lives in the generated appendix
INSTRUCTIONS-ARM64.md.
Registers
- General purpose:
R0toR30, plusZR, the zero register, andRSP, the stack pointer. There is no R31: thirty-one names and ZR. - Floating-point and SIMD share one file written
Vn; where an instruction is scalar floating point the operand may be writtenFn(F0toF31). - SVE registers
Z0toZ31and predicatesP0toP15; the SVE2 and SVE2.1 families encode, see SVE below. - Roles the convention fixes:
RSPis the stack pointer,R29the frame pointer,R30the link register,R26the closure context andR27the assembler's scratch register. The goroutine pointer lives inR28and is writtengin source, its fields asg_m(g),g_sched(g);R18is the platform-reserved register and the Go toolchain never addresses it.
Loads, stores and the width suffixes
The MOV series is the load and store interface, with the width in the mnemonic rather than the register name:
| Mnemonic | Machine instruction |
|---|---|
MOVD |
ldr, str, stur, 64-bit |
MOVW |
ldrsw, str, stur, 32-bit sign extending |
MOVWU |
ldr, 32-bit zero extending |
MOVH |
ldrsh, strh, sturh |
MOVHU |
ldrh |
MOVB |
ldrsb, strb, sturb |
MOVBU |
ldrb |
Post-index and pre-index addressing take the .P and .W suffixes on the
mnemonic: MOVD.P -8(R10), R8 is ldr x8, [x10],#-8, and MOVB.W 16(R16), R10 is ldrsb x10, [x16,#16]!.
Addressing
imm(Rn|RSP) 28(R17)
(Rn|RSP) (R22)
(Rn)(Rm) (R27)(R23)
(Rn)(Rm<<scale) (R4)(R12<<2)
(Rn)(Rm.UXTW<<3) extended and shifted index
(Rt1, Rt2) register pair for LDP, STP and the exclusive pair forms
Branch targets are labels, (R3) for indirect, name(SB) for static.
Operand order and the special forms
Most instructions appear in left-to-right assignment order: ADD R11, RSP, R25 computes into R25. The exceptions the toolchain's manual lists,
each with its own order:
- stores and
CBZ,CBNZkeep the GNU order:MOVD R29, 384(R19). - The multiply-accumulate family
MADD,MSUB,SMADDLand friends are<Rm>, <Ra>, <Rn>, <Rd>. - The scalar FMA family
FMADDDand friends are<Fm>, <Fa>, <Fn>, <Fd>. - The bitfield family
BFI,BFXIL,SBFIZ,SBFX,UBFIZ,UBFXis$<lsb>, <Rn>, $<width>, <Rd>. - The conditional compare and select families carry the condition as the
first operand:
CSEL GT, R0, R19, R1,CCMP MI, R22, $12, $13,FCCMPD AL, F8, F26, $0. - The exclusive stores are
<Rf>, (<Rn>), <Rs>with the status register last:STLXR ZR, (R15), R16. TBZandTBNZare$<imm>, <Rt>, <label>.
Shifted and extended register operands ride the register: R19>>30,
R26->24 for arithmetic right shift, @> for rotate, and the extend forms
R19.UXTB<<4, R14.SXTX with extend operators UXTB, UXTH, UXTW, UXTX,
SXTB, SXTH, SXTW, SXTX.
Conditions, branches and names
- Conditions ride the branch mnemonic:
B.EQ, or the canonical per-condition names such asBEQ. Both spellings exist; the canonical names are what the generated inventory lists. brisJMPandblrisCALLin this dialect; indirect branches areJMP (R3)andCALL (R17).NOPis a zero-width pseudo-instruction; the hardware nop isNOOP, an alias ofHINT $0.umovis written asVMOV.
Constants
- A 16-bit immediate optionally shifted:
MOVK $(10<<32), R20, withMOVZ,MOVNand their W variants; a zero shift is rejected by the assembler. - Large integer constants:
MOVmaterialises any 64-bit constant, the closest-instruction way. - Vector constants:
VMOVS,VMOVDandVMOVQ, the last taking two 64-bit halves for a 128-bit value:VMOVQ $0x1122334455667788, $0x99aabbccddeeff00, V2.
SIMD
Floating-point and SIMD instructions mostly carry a V prefix
(VADD, VFMLA), the cryptographic extensions (AESD, SHA256H) and the
scalar floating-point instructions being the exceptions. Operands carry an
arrangement suffix, V5.H8, and structure loads and stores use bracket
lists, [V21.B16], with element selection as V9.S[1].
SVE
The scalable vector extension encodes through the Z0 to Z31 vector
registers and the P0 to P15 predicates, with arrangement suffixes on
the Z registers and the merging and zeroing qualifiers the toolchain's
SVE2 and SVE2.1 families carry: the narrowing two-to-one arithmetic, the
cryptographic set including ZADCLB, BFloat16 arithmetic, the predicate
counters and reductions, the pairwise and quadword forms, the
multiple-structure loads and stores, shift by vector and the
shift-immediate scheme, CLASTA and CLASTB, the compare and last-active
predicate families, and the vector-length arithmetic ADDVL, ADDPL and
RDVL, whose immediates count vectors or predicates rather than bytes.
The gather and scatter loads take their addresses through the five
addressing modes the toolchain defines. The source grammar is the
toolchain's own, and the byte output is pinned against it corpus-wide.
Synthesised forms
GETCALLERPCreads the return address the frame state describes: a leaf function readsR30, a framed body reads the prologue's save slot.REM,REMW,UREMandUREMWsynthesise a remainder fromSDIVorUDIVand theMSUBtail, withRSPrefused as a destination.DWORD $immlays eight little-endian bytes per immediate.MOVD tls_g(SB), Rdmaterialises a TLS local-exec load as a singleMOVZcarryingR_ARM64_TLS_LE, keyed off the file's ownGLOBL ... TLSBSSdeclaration.
Alignment
PCALIGN $n pads to a power-of-two boundary between 8 and 2048 and also
raises the function's alignment to the coarsest boundary any of its PCALIGN
directives asks for. Functions default to 16-byte alignment on this target.
Relocations
R_ADDRARM64 for the adrp-plus-add pair, R_ARM64_PCREL and the
R_ARM64_PCREL_LDST family for PC relative addressing, R_ARM64_LDST for
the load and store immediates, R_ARM64_GOTPCREL and R_ARM64_GOT for the
GOT, R_ARM64_TLS_LE and R_ARM64_TLS_IE for thread local storage and
R_CALLARM64 for direct calls, all specified in
GOOBJ.md.