125 lines
6.0 KiB
Markdown
125 lines
6.0 KiB
Markdown
# AMD64
|
||
|
||
Layer 1, target page. Verified against `go tool asm` of Go 1.27.1 and against
|
||
gasm's encoder, whose output is compared byte for byte with the toolchain's
|
||
and executed on real hardware (`gasm verify`). The complete mnemonic
|
||
inventory lives in the generated appendix
|
||
[INSTRUCTIONS-AMD64.md](INSTRUCTIONS-AMD64.md); this page is the grammar and
|
||
the conventions.
|
||
|
||
## Registers
|
||
|
||
| Group | Names | Notes |
|
||
|---|---|---|
|
||
| General purpose, 64-bit | `AX` `BX` `CX` `DX` `SI` `DI` `BP` `SP` `R8` to `R15` | bare names, no prefix |
|
||
| Sub-registers | `AL` `CL` `DL` `BL` `AH` family; `R8B` `R8W` `R8D` for the byte, word and double word of `R8` | width rides the mnemonic as well |
|
||
| Vector | `X0` to `X15` (128-bit), `Y0` to `Y15` (256-bit), `Z0` to `Z31` (512-bit) | SSE, AVX and AVX-512 |
|
||
| Mask | `K0` to `K7` | AVX-512 opmask |
|
||
| System | `TLS` | the thread pointer, see below |
|
||
|
||
Roles the calling convention fixes, which assembly must respect and can rely
|
||
on:
|
||
|
||
- `SP` is the hardware stack pointer; the virtual frame pointer of the
|
||
common language is the pseudo-register SP of OPERANDS.md, a different
|
||
spelling with a different meaning.
|
||
- `BP` is callee-save. The assembler inserts the save and restore whenever
|
||
the function has a non-zero frame, so using BP as a general register
|
||
interferes with sampling profilers that walk the frame chain.
|
||
- `R14` holds `g`, the goroutine pointer, in the register ABI; `RDX` holds
|
||
the closure context; `R12` and `R13` are the register ABI's scratch pair
|
||
and `R15` its GOT temporary; `X15` is the zeroing register the compiler
|
||
uses. An ABI0 assembly function called from Go sees none of these live
|
||
across the call, but runtime assembly reads them directly.
|
||
- The legacy spellings for the goroutine pointer are the macros of
|
||
`runtime/go_tls.h`: `get_tls(r)` expands to `MOVQ TLS, r` and `g(r)` to
|
||
`0(r)(TLS*1)`, the segment base riding the index field.
|
||
|
||
## Addressing
|
||
|
||
The common forms of OPERANDS.md, with the amd64 specifics:
|
||
|
||
```text
|
||
offset(base) MOVQ 16(BX), AX
|
||
offset(base)(index*scale) MOVL foo+32(SP)(R9*8), CX
|
||
scale is 1, 2, 4 or 8
|
||
name±offset(SB) MOVQ ·table(SB), CX
|
||
```
|
||
|
||
- Global references assemble as absolute addresses and produce R_ADDR
|
||
relocations; branch targets produce R_PCREL.
|
||
- Vector indexed memory, the VSIB form with an X, Y or Z register in the
|
||
index position, exists for the gather and scatter families.
|
||
- There are no segment overrides in source; the one segment-flavoured form
|
||
is the TLS base in the index field shown above.
|
||
|
||
## The frame and the split check
|
||
|
||
The assembler manages the frame, not the programmer:
|
||
|
||
- It inserts the `BP` save and restore for any non-zero frame.
|
||
- It inserts the stack-split check for any function that is not NoSplit:
|
||
the check compares SP against the guard, and on exhaustion calls
|
||
`runtime.morestack_noctxt`. Frames at or below 128 bytes, StackSmall, use
|
||
the small compare; frames at or below 4096 bytes, StackBig, use the
|
||
adjusted form; larger frames compare in two steps.
|
||
- On amd64 the assembler marks a function NoSplit itself when the frame is
|
||
under StackSmall and the body calls nothing that needs stack: such a
|
||
function carries the NoSplit flag in the object without the source ever
|
||
writing NOSPLIT.
|
||
|
||
Results and arguments are stack-only in ABI0: the caller's frame carries
|
||
them at FP offsets, per the Go prototype.
|
||
|
||
## Instructions
|
||
|
||
The inventory counts 1654 recognised mnemonics today, of which the encoder
|
||
emits 1113; both numbers are generated in the appendix, and the gap is the
|
||
encoder backlog that `gasm audit-instructions` measures. The families:
|
||
|
||
- **Integer base.** The ALU and move set with width suffixes, `MOVB`,
|
||
`MOVW`, `MOVL`, `MOVQ`; the extension moves `MOVBLZX`, `MOVWLSX`,
|
||
`MOVLQSX` and their siblings, which the compiler's output leans on;
|
||
`LEA`; `PUSH` and `POP`; the shifts and rotates; the bit operations `BT`
|
||
through `BTC`, `BSF`, `BSR`, `LZCNT`, `TZCNT`, `POPCNT`, `BSWAP`; the
|
||
string primitives `MOVS` and `STOS`.
|
||
- **Exchange and atomics.** `XCHG`, `CMPXCHG`, `XADD`; the extended-carry
|
||
pair `ADCX` and `ADOX`; `CRC32`.
|
||
- **Scalar floating point.** The SSE2 scalar moves and arithmetic
|
||
(`MOVSD`, `MOVSS`, `ADDSD`, and the `CVT` family). Floating-point
|
||
immediates are not encodable on this target, so the assembler
|
||
materialises them: the constant lands in a synthesised read-only pool,
|
||
and a positive zero collapses to `XORPS` of the register with itself,
|
||
exactly as the toolchain does.
|
||
- **Legacy SIMD, SSE.** The `MOVO`, `MOVOU`, `MOVAPS` family and the packed
|
||
integer and floating operations, shuffles, lane extracts and inserts and
|
||
the imm8-controlled forms.
|
||
- **VEX and EVEX.** The `V`-prefixed forms for 256 and 512-bit work,
|
||
opmask operations on `K0` to `K7`, gathers and scatters, and the
|
||
quad-register families 4FMAPS, 4FNMADD, 4VNNIW, VP4DPWSSD and VP4DPWSSDS,
|
||
whose register list rides the inverted V′VVV field. Mixing VEX and legacy
|
||
SSE in one loop pays the AVX-SSE transition penalty on every switch: keep
|
||
a loop in one dialect.
|
||
- **Cryptographic and counting extensions.** AES-NI, SHA-1 and SHA-256,
|
||
PCLMULQDQ, GFNI.
|
||
- **System.** `CPUID`, `RDTSC`, `SYSCALL`, the fences, `LDMXCSR` and
|
||
`STMXCSR`, the prefetch family.
|
||
- **Pseudo-operations.** `BYTE`, `WORD`, `LONG`, `QUAD` lay raw bytes or
|
||
words into the stream for encodings the assembler does not know; `ADJSP`
|
||
adjusts the stack pointer; `DUFFCOPY` and `DUFFZERO` and `GETCALLERPC`
|
||
are compiler-side names the table recognises but an encoder need not
|
||
emit.
|
||
|
||
A mnemonic the appendix lists with `gasm encodes: no` assembles nowhere:
|
||
gasm reports it as an explicit error, never as wrong bytes, and the
|
||
`unencodable-instruction` lint flags it at edit time.
|
||
|
||
## Relocations
|
||
|
||
The relocations an amd64 object carries, all specified in
|
||
[GOOBJ.md](../GOOBJ.md): `R_ADDR` for absolute globals, `R_PCREL` for
|
||
relative addresses, `R_CALL` for direct calls, `R_TLS_LE` and `R_TLS_IE` for
|
||
thread local access and `R_GOTPCREL` for GOT relative sequences, plus
|
||
`R_DWTXTADDR_U4` inside the DWARF records, which the assembler always
|
||
emits in the four-byte flavour.
|