105 lines
4.8 KiB
Markdown
105 lines
4.8 KiB
Markdown
# RISC-V 64
|
|||
|
|
|
||
|
|
Layer 1, target page. Verified against `go tool asm` of Go 1.27.1, against
|
||
|
|
the toolchain's own riscv64 assembler manual (`cmd/internal/obj/riscv/doc.go`)
|
||
|
|
and against gasm's encoder, whose output is compared byte for byte with the
|
||
|
|
toolchain's. The complete mnemonic inventory lives in the generated appendix
|
||
|
|
[INSTRUCTIONS-RISCV64.md](INSTRUCTIONS-RISCV64.md).
|
||
|
|
|
||
|
|
## Registers
|
||
|
|
|
||
|
|
- Integer: `X0` to `X31`. `X0` is hardwired zero. Three names the toolchain
|
||
|
|
constrains: `X4` must be written through its ABI name `TP`; `X27`, the
|
||
|
|
goroutine pointer, must be written `g` and may not be written `S11`; in
|
||
|
|
shared builds `X3` is off limits and must be written `GP`.
|
||
|
|
- The other integer registers may be written `Xn` or by their ABI names
|
||
|
|
(`A0`, `T0`, `S1`, and so on).
|
||
|
|
- Floating point: `F0` to `F31`. Vector: `V0` to `V31`.
|
||
|
|
- `X26` is the closure pointer and `X31` is the assembler's own scratch
|
||
|
|
register: its value may be clobbered by instruction sequences the
|
||
|
|
assembler inserts, so hand-written code must not rely on it.
|
||
|
|
- There is no reserved frame pointer register on this target.
|
||
|
|
|
||
|
|
## Operand order
|
||
|
|
|
||
|
|
The ordering differs from the ISA manual, and per instruction class:
|
||
|
|
|
||
|
|
- **R-type** is reversed: `ADD X10, X11, X12` is `add x12, x11, x10`.
|
||
|
|
- **I-type arithmetic** keeps that shape with the immediate first:
|
||
|
|
`ADDI $1, X11, X12`.
|
||
|
|
- **Loads and stores** are source first, like every Plan 9 dialect:
|
||
|
|
`MOV 16(X2), X10` loads and `MOV X10, (X2)` stores. The MOV series hides
|
||
|
|
the width; `MOVB` through `MOVD` spell it out.
|
||
|
|
- **Branches** keep the ISA order: `BLT X12, X23, loop1`, which jumps when
|
||
|
|
X12 < X23, the reverse of the SLT operand order.
|
||
|
|
- **FMA** is rotated one place left so the destination comes last:
|
||
|
|
`FMADDS F1, F2, F3, F4`.
|
||
|
|
- **AMO** is likewise rotated: `AMOSWAPW X5, (X6), X7`.
|
||
|
|
- **Ternary abbreviation** is supported and encouraged: `ADD X10, X12` means
|
||
|
|
`ADD X10, X12, X12`.
|
||
|
|
|
||
|
|
Where an R-type instruction has an I-type sibling, the assembler picks the
|
||
|
|
immediate form from the operand: `AND $3, X12, X13` assembles as `ANDI`.
|
||
|
|
|
||
|
|
## Names, suffixes and rounding
|
||
|
|
|
||
|
|
Dots are removed and suffixes are upper-cased: the ISA's `fmv.w.x` is
|
||
|
|
`FMVWX`. Floating-point rounding modes become suffixes, `FCVTLUS.RNE F0,
|
||
|
|
X5`, with RTZ assumed when the suffix is omitted; the toolchain never sets
|
||
|
|
the FCSR.
|
||
|
|
|
||
|
|
## Constants
|
||
|
|
|
||
|
|
- `MOV` materialises any 64-bit integer constant, synthesising it from a
|
||
|
|
few arithmetic instructions where possible and otherwise loading it from
|
||
|
|
a literal pool in the binary.
|
||
|
|
- A 32-bit constant is accepted by `ADDI`, `ANDI`, `ORI` and `XORI`, and
|
||
|
|
the assembler synthesises values that exceed the 12-bit encoding window.
|
||
|
|
- `MOVF` and `MOVD` materialise floating-point constants, encoding them as
|
||
|
|
`FLW` and `FLD` from a pool location unless the constant is exactly 0.0.
|
||
|
|
|
||
|
|
## Extensions and profiles
|
||
|
|
|
||
|
|
The default target profile is rva20u64, selected or raised with the
|
||
|
|
GORISCV64 environment variable. A short list of instructions outside the
|
||
|
|
default profile is synthesised by the assembler when the profile does not
|
||
|
|
provide them, so they are safe without guards: `ANDN`, `MAX`, `MAXU`, `MIN`,
|
||
|
|
`MINU`, `MOVB`, `MOVH`, `MOVHU`, `MOVWU`, `ORN`, `ROL`, `ROLW`, `ROR`,
|
||
|
|
`RORI`, `RORIW`, `RORW`, `XNOR`. The header `asm_riscv64.h` defines the
|
||
|
|
`hasZba`, `hasZbb`, `hasZbs` and `hasV` macros for guarding everything else.
|
||
|
|
|
||
|
|
## Fences and atomics
|
||
|
|
|
||
|
|
`FENCE` takes predecessor and successor sets in that order, uppercase
|
||
|
|
letters, `FENCE R, RW`; a bare `FENCE` is a full fence, as is
|
||
|
|
`FENCE IORW, IORW`. `FENCE.TSO` exists. The ordering bits of `LR`, `SC`
|
||
|
|
and the AMO instructions are not specifiable in source: the assembler sets
|
||
|
|
acquire and release on the AMO instructions, acquire on `LR` and release on
|
||
|
|
`SC`, always.
|
||
|
|
|
||
|
|
## Compressed instructions
|
||
|
|
|
||
|
|
The assembler converts 32-bit instructions to their compressed encodings
|
||
|
|
automatically; the conversion is a property of the emitted machine code, not
|
||
|
|
of the source, and register choice influences how much compresses.
|
||
|
|
Hand-writing compressed instructions in source is accepted but discouraged.
|
||
|
|
The debug flag `compressinstructions=0` turns the automatic conversion off.
|
||
|
|
|
||
|
|
## Vector extension
|
||
|
|
|
||
|
|
`VSETVLI` writes its vtype components in uppercase with the destination
|
||
|
|
last: `VSETVLI X10, E8, M1, TU, MU, X12`. Vector loads and stores are
|
||
|
|
source first like the scalar ones, with an optional stride or index register
|
||
|
|
second and the mask register, when present, always penultimate:
|
||
|
|
`VLE8V (X10), V3`, `VLE8V (X10), V0, V3` for the masked form. Vector
|
||
|
|
arithmetic reverses its operands, `VADDVV V1, V2, V3`, with the mask again
|
||
|
|
penultimate.
|
||
|
|
|
||
|
|
## Relocations
|
||
|
|
|
||
|
|
`R_RISCV_JAL`, `R_RISCV_CALL`, the `R_RISCV_PCREL_ITYPE` and `STYPE` pairs,
|
||
|
|
`R_RISCV_BRANCH`, the compressed branch and jump forms, the TLS and GOT
|
||
|
|
families and `R_RISCV_ADD32` and `SUB32`, all specified in
|
||
|
|
[GOOBJ.md](../GOOBJ.md). The assembler always emits the four-byte
|
||
|
|
`R_DWTXTADDR_U4` flavour inside its DWARF records.
|