Files
petrbalvin 53de91b2df
Test / test (push) Failing after 2m23s
docs(asm): describe the four target architectures
Assisted-by: GLM 5.3 Flash
2026-09-21 20:15:55 +02:00

4.8 KiB

RISC-V 64

Layer 1, target page. Verified against go tool asm of Go 1.27.1, against the toolchain's own riscv64 assembler manual (cmd/internal/obj/riscv/doc.go) and against gasm's encoder, whose output is compared byte for byte with the toolchain's. The complete mnemonic inventory lives in the generated appendix INSTRUCTIONS-RISCV64.md.

Registers

  • Integer: X0 to X31. X0 is hardwired zero. Three names the toolchain constrains: X4 must be written through its ABI name TP; X27, the goroutine pointer, must be written g and may not be written S11; in shared builds X3 is off limits and must be written GP.
  • The other integer registers may be written Xn or by their ABI names (A0, T0, S1, and so on).
  • Floating point: F0 to F31. Vector: V0 to V31.
  • X26 is the closure pointer and X31 is the assembler's own scratch register: its value may be clobbered by instruction sequences the assembler inserts, so hand-written code must not rely on it.
  • There is no reserved frame pointer register on this target.

Operand order

The ordering differs from the ISA manual, and per instruction class:

  • R-type is reversed: ADD X10, X11, X12 is add x12, x11, x10.
  • I-type arithmetic keeps that shape with the immediate first: ADDI $1, X11, X12.
  • Loads and stores are source first, like every Plan 9 dialect: MOV 16(X2), X10 loads and MOV X10, (X2) stores. The MOV series hides the width; MOVB through MOVD spell it out.
  • Branches keep the ISA order: BLT X12, X23, loop1, which jumps when X12 < X23, the reverse of the SLT operand order.
  • FMA is rotated one place left so the destination comes last: FMADDS F1, F2, F3, F4.
  • AMO is likewise rotated: AMOSWAPW X5, (X6), X7.
  • Ternary abbreviation is supported and encouraged: ADD X10, X12 means ADD X10, X12, X12.

Where an R-type instruction has an I-type sibling, the assembler picks the immediate form from the operand: AND $3, X12, X13 assembles as ANDI.

Names, suffixes and rounding

Dots are removed and suffixes are upper-cased: the ISA's fmv.w.x is FMVWX. Floating-point rounding modes become suffixes, FCVTLUS.RNE F0, X5, with RTZ assumed when the suffix is omitted; the toolchain never sets the FCSR.

Constants

  • MOV materialises any 64-bit integer constant, synthesising it from a few arithmetic instructions where possible and otherwise loading it from a literal pool in the binary.
  • A 32-bit constant is accepted by ADDI, ANDI, ORI and XORI, and the assembler synthesises values that exceed the 12-bit encoding window.
  • MOVF and MOVD materialise floating-point constants, encoding them as FLW and FLD from a pool location unless the constant is exactly 0.0.

Extensions and profiles

The default target profile is rva20u64, selected or raised with the GORISCV64 environment variable. A short list of instructions outside the default profile is synthesised by the assembler when the profile does not provide them, so they are safe without guards: ANDN, MAX, MAXU, MIN, MINU, MOVB, MOVH, MOVHU, MOVWU, ORN, ROL, ROLW, ROR, RORI, RORIW, RORW, XNOR. The header asm_riscv64.h defines the hasZba, hasZbb, hasZbs and hasV macros for guarding everything else.

Fences and atomics

FENCE takes predecessor and successor sets in that order, uppercase letters, FENCE R, RW; a bare FENCE is a full fence, as is FENCE IORW, IORW. FENCE.TSO exists. The ordering bits of LR, SC and the AMO instructions are not specifiable in source: the assembler sets acquire and release on the AMO instructions, acquire on LR and release on SC, always.

Compressed instructions

The assembler converts 32-bit instructions to their compressed encodings automatically; the conversion is a property of the emitted machine code, not of the source, and register choice influences how much compresses. Hand-writing compressed instructions in source is accepted but discouraged. The debug flag compressinstructions=0 turns the automatic conversion off.

Vector extension

VSETVLI writes its vtype components in uppercase with the destination last: VSETVLI X10, E8, M1, TU, MU, X12. Vector loads and stores are source first like the scalar ones, with an optional stride or index register second and the mask register, when present, always penultimate: VLE8V (X10), V3, VLE8V (X10), V0, V3 for the masked form. Vector arithmetic reverses its operands, VADDVV V1, V2, V3, with the mask again penultimate.

Relocations

R_RISCV_JAL, R_RISCV_CALL, the R_RISCV_PCREL_ITYPE and STYPE pairs, R_RISCV_BRANCH, the compressed branch and jump forms, the TLS and GOT families and R_RISCV_ADD32 and SUB32, all specified in GOOBJ.md. The assembler always emits the four-byte R_DWTXTADDR_U4 flavour inside its DWARF records.