Files
gasm-sdk/docs/asm/AMD64.md
T
petrbalvin a77dd12d2d
Test / test (push) Successful in 32s
Release / build (amd64, freebsd) (push) Successful in 52s
Release / build (amd64, linux) (push) Successful in 13s
Release / build (arm64, freebsd) (push) Successful in 51s
Release / build (arm64, linux) (push) Successful in 29s
Release / build (loong64, linux) (push) Successful in 29s
Release / build (riscv64, linux) (push) Successful in 31s
Release / release (push) Successful in 10s
chore: prepare release v0.36.0
Assisted-by: GLM 5.3 Flash
2026-10-07 22:37:42 +02:00

6.6 KiB
Raw Blame History

AMD64

Layer 1, target page. Verified against go tool asm of Go 1.27.1 and against gasm's encoder, whose output is compared byte for byte with the toolchain's and executed on real hardware (gasm verify). The complete mnemonic inventory lives in the generated appendix INSTRUCTIONS-AMD64.md; this page is the grammar and the conventions.

Registers

Group Names Notes
General purpose, 64-bit AX BX CX DX SI DI BP SP R8 to R15 bare names, no prefix
Sub-registers AL CL DL BL AH family; R8B R8W R8D for the byte, word and double word of R8 width rides the mnemonic as well
Vector X0 to X15 (128-bit), Y0 to Y15 (256-bit), Z0 to Z31 (512-bit) SSE, AVX and AVX-512
Mask K0 to K7 AVX-512 opmask
System TLS the thread pointer, see below

Roles the calling convention fixes, which assembly must respect and can rely on:

  • SP is the hardware stack pointer; the virtual frame pointer of the common language is the pseudo-register SP of OPERANDS.md, a different spelling with a different meaning.
  • BP is callee-save. The assembler inserts the save and restore whenever the function has a non-zero frame, so using BP as a general register interferes with sampling profilers that walk the frame chain.
  • R14 holds g, the goroutine pointer, in the register ABI; RDX holds the closure context; R12 and R13 are the register ABI's scratch pair and R15 its GOT temporary; X15 is the zeroing register the compiler uses. An ABI0 assembly function called from Go sees none of these live across the call, but runtime assembly reads them directly.
  • The legacy spellings for the goroutine pointer are the macros of runtime/go_tls.h: get_tls(r) expands to MOVQ TLS, r and g(r) to 0(r)(TLS*1), the segment base riding the index field.

Addressing

The common forms of OPERANDS.md, with the amd64 specifics:

offset(base)               MOVQ 16(BX), AX
offset(base)(index*scale)  MOVL foo+32(SP)(R9*8), CX
scale is 1, 2, 4 or 8
name±offset(SB)            MOVQ ·table(SB), CX
  • Global references assemble as absolute addresses and produce R_ADDR relocations; branch targets produce R_PCREL.
  • Vector indexed memory, the VSIB form with an X, Y or Z register in the index position, exists for the gather and scatter families.
  • There are no segment overrides in source; the one segment-flavoured form is the TLS base in the index field shown above.

The frame and the split check

The assembler manages the frame, not the programmer:

  • It inserts the BP save and restore for any non-zero frame.
  • It inserts the stack-split check for any function that is not NoSplit: the check compares SP against the guard, and on exhaustion calls runtime.morestack_noctxt. Frames at or below 128 bytes, StackSmall, use the small compare; frames at or below 4096 bytes, StackBig, use the adjusted form; larger frames compare in two steps.
  • On amd64 the assembler marks a function NoSplit itself when the frame is under StackSmall and the body calls nothing that needs stack: such a function carries the NoSplit flag in the object without the source ever writing NOSPLIT.

Results and arguments are stack-only in ABI0: the caller's frame carries them at FP offsets, per the Go prototype.

Instructions

The inventory counts 1654 recognised mnemonics today, of which the encoder emits 1113; both numbers are generated in the appendix, and the gap is the encoder backlog that gasm audit-instructions measures. The families:

  • Integer base. The ALU and move set with width suffixes, MOVB, MOVW, MOVL, MOVQ; the extension moves MOVBLZX, MOVWLSX, MOVLQSX and their siblings, which the compiler's output leans on; LEA; PUSH and POP; the shifts and rotates; the bit operations BT through BTC, BSF, BSR, LZCNT, TZCNT, POPCNT, BSWAP; the string primitives MOVS and STOS.
  • Exchange and atomics. XCHG, CMPXCHG, XADD; the extended-carry pair ADCX and ADOX; CRC32.
  • Scalar floating point. The SSE2 scalar moves and arithmetic (MOVSD, MOVSS, ADDSD, and the CVT family). Floating-point immediates are not encodable on this target, so the assembler materialises them: the constant lands in a synthesised read-only pool, and a positive zero collapses to XORPS of the register with itself, exactly as the toolchain does.
  • Legacy SIMD, SSE. The MOVO, MOVOU, MOVAPS family and the packed integer and floating operations, shuffles, lane extracts and inserts and the imm8-controlled forms.
  • VEX and EVEX. The V-prefixed forms for 256 and 512-bit work, opmask operations on K0 to K7, gathers and scatters, and the quad-register families 4FMAPS, 4FNMADD, 4VNNIW, VP4DPWSSD and VP4DPWSSDS, whose register list rides the inverted V′VVV field. Mixing VEX and legacy SSE in one loop pays the AVX-SSE transition penalty on every switch: keep a loop in one dialect.
  • The extension layer. The families the toolchain's own table carries late or not at all assemble through the extension mechanism: BF16, VP2INTERSECT, the complete AVX512-FP16 set (the packed and scalar FMA families, the complex multiply and complex FMA pairs, VMINMAXPH and VMINMAXSH), and the AVX-VNNI-INT16 dot products through a VEX path. The layer takes k0 to k7 write masks with merging and zeroing, the {1toN} broadcast, the embedded rounding and {sae} decorations and the imm8 controls, and its refusals surface as extension-form lint errors with the reason the encoder would give.
  • Cryptographic and counting extensions. AES-NI, SHA-1 and SHA-256, PCLMULQDQ, GFNI.
  • System. CPUID, RDTSC, SYSCALL, the fences, LDMXCSR and STMXCSR, the prefetch family.
  • Pseudo-operations. BYTE, WORD, LONG, QUAD lay raw bytes or words into the stream for encodings the assembler does not know; ADJSP adjusts the stack pointer; DUFFCOPY and DUFFZERO and GETCALLERPC are compiler-side names the table recognises but an encoder need not emit.

A mnemonic the appendix lists with gasm encodes: no assembles nowhere: gasm reports it as an explicit error, never as wrong bytes, and the unencodable-instruction lint flags it at edit time.

Relocations

The relocations an amd64 object carries, all specified in GOOBJ.md: R_ADDR for absolute globals, R_PCREL for relative addresses, R_CALL for direct calls, R_TLS_LE and R_TLS_IE for thread local access and R_GOTPCREL for GOT relative sequences, plus R_DWTXTADDR_U4 inside the DWARF records, which the assembler always emits in the four-byte flavour.