# LoongArch 64 Layer 1, target page. Verified against `go tool asm` of Go 1.27.1, against the toolchain's own loong64 assembler manual (`cmd/internal/obj/loong64/doc.go`) and against gasm's encoder, whose output is compared byte for byte with the toolchain's. The complete mnemonic inventory lives in the generated appendix [INSTRUCTIONS-LOONG64.md](INSTRUCTIONS-LOONG64.md). ## Registers - General purpose `R0` to `R31`, floating point `F0` to `F31`, LSX vectors `V0` to `V31` and LASX vectors `X0` to `X31`. - Fixed roles from the toolchain's table: `R0` is the constant zero, `R1` the return address, `R3` the stack pointer, `R22` the goroutine pointer, `R29` the closure context and `R30` the assembler's temporary. `R12`, `R13`, `R14`, `R15` and `R20` serve the PLT and trampoline sequences: usable in assembly, but saved before any call. ## Widths ride the mnemonic | Suffix | Width | |---|---| | `B`, `BU` | 8-bit, 8-bit unsigned | | `H`, `HU` | 16-bit, 16-bit unsigned | | `W`, `WU` | 32-bit, 32-bit unsigned | | `V` | 64-bit | | `F`, `D` | 32-bit and 64-bit float | | `V` prefix (LSX) | 128-bit vector | | `XV` prefix (LASX) | 256-bit vector | The MOV series is the load and store interface: `MOVB (R2), R3` loads a byte, `MOVV (R2), R3` a double word, `VMOVQ (R2), V1` a 128-bit vector and `XVMOVQ (R2), X1` a 256-bit one. ## Operand order Most instructions appear in left-to-right assignment order: `ADDV R11, R12, R13` is `add.d R13, R12, R11`, and the two-operand form `OR R5, R6` assigns into R6. Exceptions: - Jump and branch instructions keep the GNU order: `BEQ R0, R4, label1`. - The bitfield family is `BSTRINSW`, `BSTRINSV`, `BSTRPICKW`, `BSTRPICKV` `$, , $, `. ## Addressing - Plain: `offset(Rbase)`. - Base plus offset **register**, no scale: `(R4)(R5)`, as in `MOVB (R4)(R5), R6`, the `ldx` family. - The pointer loads and stores `MOVWP` and `MOVVP` take a source-level 16-bit offset that the encoder halves into the 14-bit field, writing `MOVWP 8(R4), R5` as `ldptr.w r5, r4, $2`. ## Vector element syntax The `VMOVQ` and `XVMOVQ` transfer family covers register-to-vector moves with arrangement and index suffixes: `VMOVQ Rj, Vd.B[index]` inserts a general register into one lane, `VMOVQ Vj.B[index], Rd` extracts one, `VMOVQ Rj, Vd.B16` broadcasts across all sixteen, and `VMOVQ Vj.B[index], Vd.B16` replicates one lane. The broadcast-from-memory form takes the true byte offset at source level, which the encoder rescales per arrangement. The permute and extract families take their 8-bit control word first: `VPERMIW ui8, Vj, Vd`, `VEXTRINSB ui8, Vj, Vd`. ## Alignment `PCALIGN $n` pads with NOOP to a power-of-two boundary between 8 and 2048, and this target additionally auto-aligns loop heads to 16 bytes. ## Atomics, barriers and prefetch - The `AM` atomic family comes in plain and `_DB` flavours; the `_DB` forms, such as `AMSWAPDBW`, complete the atomic sequence and act as a full data barrier. Within the AM family the destination and base registers may not coincide and the destination may not equal the operand register: one is an exception, the other silently unspecified. - `DBAR` carries the graded hint encoding documented for LA664 and later, with hint 0x700 as the read-after-read lightweight barrier; older cores treat every hint as the full barrier. - `PRELD offset(Rbase), $hint` prefetches with the documented hints (0 load to L1, 2 load to L3, 8 store to L1); `PRELDX` adds the encoded block descriptor. - `ALSL`-family shift-and-add writes the desired shift amount in source and encodes one less: `ALSLV $4, R4, R5, R6` shifts by 4. - `ADDV16 si16<<16, Rj, Rd` is the high-immediate add paired with the pointer loads for GOT relative access. ## Relocations `R_CALLLOONG64` for the 28-bit BL, `R_LOONG64_CALL36` for the PCADDU18I-plus-JIRL pair, the `R_LOONG64_ADDR`, `ADDR64`, `TLS_LE`, `TLS_IE`, `GOT` and `GOT64` high and low pairs, the aligned conditional jump forms `R_JMP16LOONG64` and `R_JMP21LOONG64`, and `R_LOONG64_ADD64` and `SUB64` for in-place arithmetic, all specified in [GOOBJ.md](../GOOBJ.md).