6.0 KiB
AMD64
Layer 1, target page. Verified against go tool asm of Go 1.27.1 and against
gasm's encoder, whose output is compared byte for byte with the toolchain's
and executed on real hardware (gasm verify). The complete mnemonic
inventory lives in the generated appendix
INSTRUCTIONS-AMD64.md; this page is the grammar and
the conventions.
Registers
| Group | Names | Notes |
|---|---|---|
| General purpose, 64-bit | AX BX CX DX SI DI BP SP R8 to R15 |
bare names, no prefix |
| Sub-registers | AL CL DL BL AH family; R8B R8W R8D for the byte, word and double word of R8 |
width rides the mnemonic as well |
| Vector | X0 to X15 (128-bit), Y0 to Y15 (256-bit), Z0 to Z31 (512-bit) |
SSE, AVX and AVX-512 |
| Mask | K0 to K7 |
AVX-512 opmask |
| System | TLS |
the thread pointer, see below |
Roles the calling convention fixes, which assembly must respect and can rely on:
SPis the hardware stack pointer; the virtual frame pointer of the common language is the pseudo-register SP of OPERANDS.md, a different spelling with a different meaning.BPis callee-save. The assembler inserts the save and restore whenever the function has a non-zero frame, so using BP as a general register interferes with sampling profilers that walk the frame chain.R14holdsg, the goroutine pointer, in the register ABI;RDXholds the closure context;R12andR13are the register ABI's scratch pair andR15its GOT temporary;X15is the zeroing register the compiler uses. An ABI0 assembly function called from Go sees none of these live across the call, but runtime assembly reads them directly.- The legacy spellings for the goroutine pointer are the macros of
runtime/go_tls.h:get_tls(r)expands toMOVQ TLS, randg(r)to0(r)(TLS*1), the segment base riding the index field.
Addressing
The common forms of OPERANDS.md, with the amd64 specifics:
offset(base) MOVQ 16(BX), AX
offset(base)(index*scale) MOVL foo+32(SP)(R9*8), CX
scale is 1, 2, 4 or 8
name±offset(SB) MOVQ ·table(SB), CX
- Global references assemble as absolute addresses and produce R_ADDR relocations; branch targets produce R_PCREL.
- Vector indexed memory, the VSIB form with an X, Y or Z register in the index position, exists for the gather and scatter families.
- There are no segment overrides in source; the one segment-flavoured form is the TLS base in the index field shown above.
The frame and the split check
The assembler manages the frame, not the programmer:
- It inserts the
BPsave and restore for any non-zero frame. - It inserts the stack-split check for any function that is not NoSplit:
the check compares SP against the guard, and on exhaustion calls
runtime.morestack_noctxt. Frames at or below 128 bytes, StackSmall, use the small compare; frames at or below 4096 bytes, StackBig, use the adjusted form; larger frames compare in two steps. - On amd64 the assembler marks a function NoSplit itself when the frame is under StackSmall and the body calls nothing that needs stack: such a function carries the NoSplit flag in the object without the source ever writing NOSPLIT.
Results and arguments are stack-only in ABI0: the caller's frame carries them at FP offsets, per the Go prototype.
Instructions
The inventory counts 1654 recognised mnemonics today, of which the encoder
emits 1113; both numbers are generated in the appendix, and the gap is the
encoder backlog that gasm audit-instructions measures. The families:
- Integer base. The ALU and move set with width suffixes,
MOVB,MOVW,MOVL,MOVQ; the extension movesMOVBLZX,MOVWLSX,MOVLQSXand their siblings, which the compiler's output leans on;LEA;PUSHandPOP; the shifts and rotates; the bit operationsBTthroughBTC,BSF,BSR,LZCNT,TZCNT,POPCNT,BSWAP; the string primitivesMOVSandSTOS. - Exchange and atomics.
XCHG,CMPXCHG,XADD; the extended-carry pairADCXandADOX;CRC32. - Scalar floating point. The SSE2 scalar moves and arithmetic
(
MOVSD,MOVSS,ADDSD, and theCVTfamily). Floating-point immediates are not encodable on this target, so the assembler materialises them: the constant lands in a synthesised read-only pool, and a positive zero collapses toXORPSof the register with itself, exactly as the toolchain does. - Legacy SIMD, SSE. The
MOVO,MOVOU,MOVAPSfamily and the packed integer and floating operations, shuffles, lane extracts and inserts and the imm8-controlled forms. - VEX and EVEX. The
V-prefixed forms for 256 and 512-bit work, opmask operations onK0toK7, gathers and scatters, and the quad-register families 4FMAPS, 4FNMADD, 4VNNIW, VP4DPWSSD and VP4DPWSSDS, whose register list rides the inverted V′VVV field. Mixing VEX and legacy SSE in one loop pays the AVX-SSE transition penalty on every switch: keep a loop in one dialect. - Cryptographic and counting extensions. AES-NI, SHA-1 and SHA-256, PCLMULQDQ, GFNI.
- System.
CPUID,RDTSC,SYSCALL, the fences,LDMXCSRandSTMXCSR, the prefetch family. - Pseudo-operations.
BYTE,WORD,LONG,QUADlay raw bytes or words into the stream for encodings the assembler does not know;ADJSPadjusts the stack pointer;DUFFCOPYandDUFFZEROandGETCALLERPCare compiler-side names the table recognises but an encoder need not emit.
A mnemonic the appendix lists with gasm encodes: no assembles nowhere:
gasm reports it as an explicit error, never as wrong bytes, and the
unencodable-instruction lint flags it at edit time.
Relocations
The relocations an amd64 object carries, all specified in
GOOBJ.md: R_ADDR for absolute globals, R_PCREL for
relative addresses, R_CALL for direct calls, R_TLS_LE and R_TLS_IE for
thread local access and R_GOTPCREL for GOT relative sequences, plus
R_DWTXTADDR_U4 inside the DWARF records, which the assembler always
emits in the four-byte flavour.