Files
gasm-sdk/docs/asm/AMD64.md
T
petrbalvin 53de91b2df
Test / test (push) Failing after 2m23s
docs(asm): describe the four target architectures
Assisted-by: GLM 5.3 Flash
2026-09-21 20:15:55 +02:00

6.0 KiB
Raw Permalink Blame History

AMD64

Layer 1, target page. Verified against go tool asm of Go 1.27.1 and against gasm's encoder, whose output is compared byte for byte with the toolchain's and executed on real hardware (gasm verify). The complete mnemonic inventory lives in the generated appendix INSTRUCTIONS-AMD64.md; this page is the grammar and the conventions.

Registers

Group Names Notes
General purpose, 64-bit AX BX CX DX SI DI BP SP R8 to R15 bare names, no prefix
Sub-registers AL CL DL BL AH family; R8B R8W R8D for the byte, word and double word of R8 width rides the mnemonic as well
Vector X0 to X15 (128-bit), Y0 to Y15 (256-bit), Z0 to Z31 (512-bit) SSE, AVX and AVX-512
Mask K0 to K7 AVX-512 opmask
System TLS the thread pointer, see below

Roles the calling convention fixes, which assembly must respect and can rely on:

  • SP is the hardware stack pointer; the virtual frame pointer of the common language is the pseudo-register SP of OPERANDS.md, a different spelling with a different meaning.
  • BP is callee-save. The assembler inserts the save and restore whenever the function has a non-zero frame, so using BP as a general register interferes with sampling profilers that walk the frame chain.
  • R14 holds g, the goroutine pointer, in the register ABI; RDX holds the closure context; R12 and R13 are the register ABI's scratch pair and R15 its GOT temporary; X15 is the zeroing register the compiler uses. An ABI0 assembly function called from Go sees none of these live across the call, but runtime assembly reads them directly.
  • The legacy spellings for the goroutine pointer are the macros of runtime/go_tls.h: get_tls(r) expands to MOVQ TLS, r and g(r) to 0(r)(TLS*1), the segment base riding the index field.

Addressing

The common forms of OPERANDS.md, with the amd64 specifics:

offset(base)               MOVQ 16(BX), AX
offset(base)(index*scale)  MOVL foo+32(SP)(R9*8), CX
scale is 1, 2, 4 or 8
name±offset(SB)            MOVQ ·table(SB), CX
  • Global references assemble as absolute addresses and produce R_ADDR relocations; branch targets produce R_PCREL.
  • Vector indexed memory, the VSIB form with an X, Y or Z register in the index position, exists for the gather and scatter families.
  • There are no segment overrides in source; the one segment-flavoured form is the TLS base in the index field shown above.

The frame and the split check

The assembler manages the frame, not the programmer:

  • It inserts the BP save and restore for any non-zero frame.
  • It inserts the stack-split check for any function that is not NoSplit: the check compares SP against the guard, and on exhaustion calls runtime.morestack_noctxt. Frames at or below 128 bytes, StackSmall, use the small compare; frames at or below 4096 bytes, StackBig, use the adjusted form; larger frames compare in two steps.
  • On amd64 the assembler marks a function NoSplit itself when the frame is under StackSmall and the body calls nothing that needs stack: such a function carries the NoSplit flag in the object without the source ever writing NOSPLIT.

Results and arguments are stack-only in ABI0: the caller's frame carries them at FP offsets, per the Go prototype.

Instructions

The inventory counts 1654 recognised mnemonics today, of which the encoder emits 1113; both numbers are generated in the appendix, and the gap is the encoder backlog that gasm audit-instructions measures. The families:

  • Integer base. The ALU and move set with width suffixes, MOVB, MOVW, MOVL, MOVQ; the extension moves MOVBLZX, MOVWLSX, MOVLQSX and their siblings, which the compiler's output leans on; LEA; PUSH and POP; the shifts and rotates; the bit operations BT through BTC, BSF, BSR, LZCNT, TZCNT, POPCNT, BSWAP; the string primitives MOVS and STOS.
  • Exchange and atomics. XCHG, CMPXCHG, XADD; the extended-carry pair ADCX and ADOX; CRC32.
  • Scalar floating point. The SSE2 scalar moves and arithmetic (MOVSD, MOVSS, ADDSD, and the CVT family). Floating-point immediates are not encodable on this target, so the assembler materialises them: the constant lands in a synthesised read-only pool, and a positive zero collapses to XORPS of the register with itself, exactly as the toolchain does.
  • Legacy SIMD, SSE. The MOVO, MOVOU, MOVAPS family and the packed integer and floating operations, shuffles, lane extracts and inserts and the imm8-controlled forms.
  • VEX and EVEX. The V-prefixed forms for 256 and 512-bit work, opmask operations on K0 to K7, gathers and scatters, and the quad-register families 4FMAPS, 4FNMADD, 4VNNIW, VP4DPWSSD and VP4DPWSSDS, whose register list rides the inverted V′VVV field. Mixing VEX and legacy SSE in one loop pays the AVX-SSE transition penalty on every switch: keep a loop in one dialect.
  • Cryptographic and counting extensions. AES-NI, SHA-1 and SHA-256, PCLMULQDQ, GFNI.
  • System. CPUID, RDTSC, SYSCALL, the fences, LDMXCSR and STMXCSR, the prefetch family.
  • Pseudo-operations. BYTE, WORD, LONG, QUAD lay raw bytes or words into the stream for encodings the assembler does not know; ADJSP adjusts the stack pointer; DUFFCOPY and DUFFZERO and GETCALLERPC are compiler-side names the table recognises but an encoder need not emit.

A mnemonic the appendix lists with gasm encodes: no assembles nowhere: gasm reports it as an explicit error, never as wrong bytes, and the unencodable-instruction lint flags it at edit time.

Relocations

The relocations an amd64 object carries, all specified in GOOBJ.md: R_ADDR for absolute globals, R_PCREL for relative addresses, R_CALL for direct calls, R_TLS_LE and R_TLS_IE for thread local access and R_GOTPCREL for GOT relative sequences, plus R_DWTXTADDR_U4 inside the DWARF records, which the assembler always emits in the four-byte flavour.