Compare commits

..
31 Commits
Author SHA1 Message Date
petrbalvin cf05b9b384 feat(cmd): GOOS-aware headers, audit battery shapes and semicolon spacing
Test / test (push) Canceled after 1m8s
Assisted-by: GLM 5.3 Flash
2026-09-20 22:02:19 +02:00
petrbalvin a7744c24bd fix(parser): substitute macro parameters behind element selectors
Assisted-by: GLM 5.3 Flash
2026-09-20 22:02:19 +02:00
petrbalvin 522e6f2ae8 feat(parser): bracket register ranges, index-only VSIB and bare trailing immediates
Assisted-by: GLM 5.3 Flash
2026-09-20 22:02:19 +02:00
petrbalvin 81d4bd81e4 test(verify): register the wave kernels
Test / test (push) Failing after 2m20s
Assisted-by: GLM 5.3 Flash
2026-09-20 21:17:31 +02:00
petrbalvin 687678a2ea feat(elf): emit data relocations on arm64, riscv64 and loong64
Assisted-by: GLM 5.3 Flash
2026-09-20 21:17:20 +02:00
petrbalvin b0f9071bf5 feat(arm64): whole-vector moves, bookkeeping ops and truncating-move lowering
Assisted-by: GLM 5.3 Flash
2026-09-20 21:17:20 +02:00
petrbalvin 81e2673923 feat(amd64): encode the AVX-512 and BMI corpus families
Assisted-by: GLM 5.3 Flash
2026-09-20 21:17:20 +02:00
petrbalvin 75e9fd771b feat(parser): split plain statements on semicolons in the raw parse
Test / test (push) Failing after 2m30s
Assisted-by: GLM 5.3 Flash
2026-09-20 19:15:35 +02:00
petrbalvin 863926abd6 test(verify): register the loong64 vector kernels
Assisted-by: GLM 5.3 Flash
2026-09-20 19:15:05 +02:00
petrbalvin 241e7256f6 fix(arm64): reject bare BTI with a diagnostic and accept the full family
Assisted-by: GLM 5.3 Flash
2026-09-20 19:15:05 +02:00
petrbalvin 6556b85abf feat(asm): symbol-valued DATA, division slash in symbols and plain semicolons
Assisted-by: GLM 5.3 Flash
2026-09-20 19:15:05 +02:00
petrbalvin 289cabe993 feat(loong64): encode the full LSX and LASX table
Assisted-by: GLM 5.3 Flash
2026-09-20 19:14:43 +02:00
petrbalvin d6cf7cfa44 fix(format): keep statement separators and canonical macro bodies
Assisted-by: GLM 5.3 Flash
2026-09-20 19:14:43 +02:00
petrbalvin 4cc2f0eba5 feat(cmd): generate go_asm.h for package-context assembly
Assisted-by: GLM 5.3 Flash
2026-09-20 19:14:43 +02:00
petrbalvin 97dfaa7526 docs: changelog for macro expansion and the corrected corpus audit
Test / test (push) Successful in 2m14s
Assisted-by: GLM 5.3 Flash
2026-09-20 14:25:47 +02:00
petrbalvin 66aa4dbc8b test(verify): register the campaign kernels in the ground-truth suites
Assisted-by: GLM 5.3 Flash
2026-09-20 14:25:47 +02:00
petrbalvin dce5d31462 feat(amd64): LOCK and REP prefixes, literal data pseudo-ops and ADJSP
Assisted-by: GLM 5.3 Flash
2026-09-20 14:25:47 +02:00
petrbalvin 9dc3987e02 feat(riscv64,loong64): PCALIGN, branch relaxation and operand shapes
Assisted-by: GLM 5.3 Flash
2026-09-20 14:25:47 +02:00
petrbalvin 9b238a525a feat(arm64): wide immediates, SIMD compare and system operand forms
Assisted-by: GLM 5.3 Flash
2026-09-20 14:25:47 +02:00
petrbalvin ad82aac663 feat(parser): macro expansion, conditionals and include splicing with -I
Assisted-by: GLM 5.3 Flash
2026-09-20 14:25:47 +02:00
petrbalvin 0629f5e2df feat(arm64): assemble PCALIGN padding and BYTE literal bytes
Test / test (push) Successful in 2m16s
Assisted-by: GLM 5.3 Flash
2026-09-20 11:49:05 +02:00
petrbalvin ecb203dcf5 fix(lexer): treat trailing CR as line end so comment text is idempotent
Test / test (push) Successful in 2m13s
Assisted-by: GLM 5.3 Flash
2026-09-20 11:40:39 +02:00
petrbalvin 6c672567f3 feat(amd64): assemble the double-shift and static-SB operand shapes
Assisted-by: GLM 5.3 Flash
2026-09-20 11:40:39 +02:00
petrbalvin cc6e416c59 fix(lint): exempt shift counts, SETcc and ABIInternal from false positives
Assisted-by: GLM 5.3 Flash
2026-09-20 11:40:39 +02:00
petrbalvin c66a47973a fix(format): preserve square brackets in SIMD operands
Test / test (push) Successful in 2m15s
Assisted-by: GLM 5.3 Flash
2026-09-20 09:58:25 +02:00
petrbalvin 9629897202 docs: changelog and readme for the instruction wave and the honest corpus rate
Test / test (push) Successful in 2m17s
Assisted-by: GLM 5.3 Flash
2026-09-20 06:45:03 +02:00
petrbalvin 5399a8a724 feat(audit): probe the new operand shapes and measure attemptable files
Assisted-by: GLM 5.3 Flash
2026-09-20 06:45:03 +02:00
petrbalvin de5d9f358e feat(riscv64,loong64): encode AMO atomics, vector slices and bit ops
Assisted-by: GLM 5.3 Flash
2026-09-20 06:44:51 +02:00
petrbalvin ca3fdce0e0 feat(arm64): encode pairs, atomics, crypto, system and NEON slices
Assisted-by: GLM 5.3 Flash
2026-09-20 06:44:51 +02:00
petrbalvin fc2d92eabd feat(amd64): encode the GOROOT instruction families
Assisted-by: GLM 5.3 Flash
2026-09-20 06:44:51 +02:00
petrbalvin 39d2e80145 ci(release): refuse empty assets and verify what the release serves
Test / test (push) Successful in 2m10s
Assisted-by: DeepSeek V4.1 Flash
2026-09-20 02:01:36 +02:00
125 changed files with 22652 additions and 478 deletions
+37
View File
@@ -325,6 +325,17 @@ jobs:
chomp $id;
my @files = grep { -f $_ } glob(q{dist/*/*});
@files or die qq{ERROR: no assets under dist/\n};
# A file that arrived empty from the artifact step would be uploaded as an
# empty attachment, every status would still be 201, and the run would go
# green over a release nobody can install. Refuse it here, before the
# upload, and verify what was stored afterwards.
my %size;
for my $path (@files) {
my $n = -s $path // 0;
(my $name = $path) =~ s{.*/}{};
$n > 0 or die qq{ERROR: $path is empty, so there is nothing to upload\n};
$size{$name} = $n;
}
my $bad = 0;
for my $path (@files) {
(my $name = $path) =~ s{.*/}{};
@@ -346,5 +357,31 @@ jobs:
printf qq{%s: HTTP %s\n}, $name, $code;
$bad = 1 if $code ne q{201};
}
# Read every asset back through the release download route and require the
# served length to be the file that was sent: stored but empty is a broken
# release however green the run looks.
open(my $v, q{<}, q{version-no-v.txt}) or die qq{version-no-v.txt: $!};
my $v = <$v>;
close($v);
chomp $v;
for my $name (sort keys %size) {
my $url = qq{$ENV{GITEA_SERVER_URL}/$ENV{GITEA_REPOSITORY}/releases/download/v$v/$name};
my @head = (q{curl}, q{-sS}, q{-I}, q{-H}, qq{Authorization: token $ENV{GITEA_TOKEN}}, $url);
open(my $h, q{-|}, @head) or die qq{curl: $!};
my $len;
my $status;
while (my $l = <$h>) {
$status = $1 if $l =~ m{^HTTP/\S+\s+(\d+)};
$len = $1 if $l =~ m{^content-length:\s*(\d+)}i;
}
my $ok = close($h);
$len = defined $len ? $len : 0;
if (!$ok || $status != 200 || $len != $size{$name}) {
printf qq{ERROR: %s serves %s bytes, expected %d\n}, $name, $len, $size{$name};
$bad = 1;
next;
}
printf qq{%s: serves %d bytes\n}, $name, $len;
}
exit($bad ? 1 : 0);
'
+38
View File
@@ -9,6 +9,39 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
### Added
- **Macro expansion and include splicing.** `gasm asm`, `gasm diff` and
`gasm audit-instructions` now preprocess assembly the way the
toolchain does: object and parameterised `#define` macros expand at
the point of use, `#undef` and the `#ifdef`/`#ifndef`/`#else`/
`#endif` family select branches, `#include` splices headers resolved
through the source directory and the new repeatable `-I` flag, `;`
separates statements, and constant expressions left in operands
(`$(32-7)`, `$~63`, `(index*4)(base)`) fold at parse. Expansion
happens only on the assembly path: `gasm lint`, `gasm fmt` and the
language server keep reading the raw file.
- **The GOROOT instruction wave, part 1.** The encoder now covers the
instruction families GOROOT's real code uses that gasm lacked,
byte-verified against `go tool asm`: on amd64 the carry ALU, the
atomics (CMPXCHG, XADD, XCHG), AES-NI, SHA-1/256, PCLMULQDQ, CRC32,
GFNI, ADX, BMI, the string primitives, the system set (CPUID, RDTSC,
SYSCALL, fences, MXCSR) and the SSE/AVX/EVEX gaps; on arm64 the pair
loads and stores (LDP/STP), acquire/release and LSE atomics, AES and
SHA, the system operations, the bit ops and the NEON slice including
structure loads and the literal-pool moves; on riscv64 the RV64A AMO
family with aq/rl ordering, the Zbb pseudos with their RVC
compressions, the FMA forms and the RVV slice with `vsetvli`/
`vsetivli`; on loong64 the AM atomics with acquire/release forms, the
LSX/LASX slice, the `VMOVQ`/`XVMOVQ` transfer family and FSEL.
Also fixed on the way: arm64 `CASD`/`CASW` lacked an opcode bit, and
riscv64 `VSETVLI` with an immediate length now canonicalises to
`vsetivli` as the toolchain does.
- **The corpus audit measures honestly.** Files named for Go ports gasm
does not target (arm, 386, s390x, ...) are no longer attempted for the
four supported architectures (no supported build compiles them), and
the headline rate is reported over attemptable files: 136 of 433 on
the full corpus (31.4 %), 135 of 383 on real code (35.2 %), from the
127 that the previous release measured. The probe battery that
decides encodability gained the operand shapes the new families use.
-
## [0.34.0] - 2026-09-20
@@ -106,6 +139,11 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
### Fixed
- **The corpus audit attempts fewer files that no build would compile.**
Files named for Go ports gasm does not target (arm, 386, s390x, ...)
are reported as other-port and never attempted, the headline rate is
computed over attemptable files, and the audit searches the
toolchain's shipped headers (funcdata.h and friends) automatically.
- **riscv64 JALR silently jumped to the wrong register.** The trampoline
form `JALR X0, 0(X5)` read the memory operand's base as the destination,
encoding a jump to X0 with no diagnostic; the destination is the first
+3 -2
View File
@@ -124,8 +124,9 @@ can emit today is narrower, and a recognised but unencodable instruction is
reported as an explicit error, never as a wrong byte.
The same measurement runs over GOROOT's whole assembly corpus:
`gasm audit-instructions --corpus` reports 127 of 627 files (20.3 %)
assembling for every target architecture today, with the top failure
`gasm audit-instructions --corpus` reports 136 of 433 attemptable files
(31.4 %) assembling for every target architecture today (files named for
other Go ports are counted but never attempted), with the top failure
reasons per architecture; the number moves with every release.
### Validation status
+4
View File
@@ -70,6 +70,10 @@ func amd64Registers() []Register {
for i := 0; i <= 7; i++ {
add(fmt.Sprintf("K%d", i), Mask, "AVX-512 mask register")
}
// x87 stack registers (FMOVD and the other x87 moves).
for i := 0; i <= 7; i++ {
add(fmt.Sprintf("F%d", i), Float, "x87 stack register")
}
return regs
}
+28
View File
@@ -33,6 +33,7 @@ func arm64Registers() []Register {
for i := 0; i <= 30; i++ {
add(fmt.Sprintf("R%d", i), GPR, "64-bit general-purpose register")
}
add("R18_PLATFORM", GPR, "R18 under its toolchain-reserved Windows name (an alias of R18)")
add("ZR", Special, "zero register (reads as 0)")
add("SP", Special, "stack pointer")
add("LR", Special, "link register (alias of R30)")
@@ -150,6 +151,33 @@ func arm64Curated() []Instr {
t = append(t, i(op, "Atomic memory operation"))
}
// Register-pair loads and stores.
for _, op := range []string{"LDP", "STP", "LDPW", "STPW", "FLDPD", "FSTPD"} {
t = append(t, ic(op, "Register-pair load or store", 2, 2))
}
// Cache maintenance and prefetch.
t = append(t, i("DC", "Data cache maintenance"))
t = append(t, i("PRFM", "Memory prefetch"))
for _, op := range []string{"LDADDAL", "LDCLRAL", "LDORAL", "SWPAL"} {
t = append(t, i(op, "Atomic memory operation with acquire and release semantics"))
}
// Cryptographic extensions.
for _, op := range []string{"AESE", "AESD", "AESMC", "AESIMC"} {
t = append(t, i(op, "AES round"))
}
for _, op := range []string{
"SHA1C", "SHA1P", "SHA1M", "SHA1H", "SHA1SU0", "SHA1SU1",
"SHA256H", "SHA256H2", "SHA256SU0", "SHA256SU1",
"SHA512H", "SHA512H2", "SHA512SU0", "SHA512SU1",
} {
t = append(t, i(op, "SHA round"))
}
for _, op := range []string{"VEOR3", "VBCAX", "VXAR", "VRAX1"} {
t = append(t, i(op, "Three-way XOR / rotate crypto vector operation"))
}
// Floating-point scalar.
for _, op := range []string{
"FADD", "FSUB", "FMUL", "FDIV", "FNEG", "FABS", "FSQRT", "FMIN", "FMAX",
+101
View File
@@ -245,3 +245,104 @@ func main() {
t.Error("binary does not contain expected symbol")
}
}
// TestGOObjectAARCH64DataSymbolLink does for symbol-valued DATA fields what
// the rt0 files do ("DATA _rt0…lib+0(SB)/8, $_rt0…lib(SB)"): the gasm object
// carries an R_ADDR against the file's own TEXT symbol, the toolchain links
// it, and the binary is checked for the symbol (no arm64 host to run it).
func TestGOObjectAARCH64DataSymbolLink(t *testing.T) {
goBin, err := exec.LookPath("go")
if err != nil {
t.Skip("no Go toolchain available")
}
dir := t.TempDir()
asmSrc := `#include "textflag.h"
GLOBL entry(SB), NOPTR, $8
DATA entry+0(SB)/8, $·keepme(SB)
TEXT ·keepme(SB), NOSPLIT, $0-0
RET
TEXT ·entryptr(SB), NOSPLIT, $0-8
MOVD entry+0(SB), R4
MOVD R4, ret+0(FP)
RET
`
if err := os.WriteFile(filepath.Join(dir, "main_arm64.s"), []byte(asmSrc), 0o644); err != nil {
t.Fatal(err)
}
mainSrc := `package main
func keepme()
func entryptr() uintptr
func main() {
if entryptr() == 0 {
panic("the entry word is empty")
}
}
`
if err := os.WriteFile(filepath.Join(dir, "main.go"), []byte(mainSrc), 0o644); err != nil {
t.Fatal(err)
}
if err := os.WriteFile(filepath.Join(dir, "go.mod"), []byte("module a64dlink\n\ngo 1.21\n"), 0o644); err != nil {
t.Fatal(err)
}
build := exec.Command(goBin, "build", "-x", "-work", "-o", filepath.Join(dir, "prog"), ".")
build.Dir = dir
build.Env = append(os.Environ(), "GOARCH=arm64")
buildLog, err := build.CombinedOutput()
if err != nil {
t.Fatalf("baseline build: %v\n%s", err, buildLog)
}
var work, linkLine, asmObj string
for line := range strings.SplitSeq(string(buildLog), "\n") {
switch {
case strings.HasPrefix(line, "WORK="):
work = strings.TrimPrefix(line, "WORK=")
case strings.Contains(line, "/asm ") && strings.Contains(line, "main_arm64.s") && !strings.Contains(line, "-gensymabis"):
asmObj = fieldAfter(line, "-o")
case strings.Contains(line, "/link ") && strings.Contains(line, "-importcfg"):
linkLine = line
}
}
if work == "" || asmObj == "" || linkLine == "" {
t.Skipf("could not parse build log (work=%q asmObj=%q link=%q)", work, asmObj, linkLine)
}
defer os.RemoveAll(work)
asmObj = strings.ReplaceAll(asmObj, "$WORK", work)
linkLine = strings.ReplaceAll(linkLine, "$WORK", work)
src, err := os.ReadFile(filepath.Join(dir, "main_arm64.s"))
if err != nil {
t.Fatal(err)
}
f, errs := parser.Parse("main_arm64.s", string(src))
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileARM64(f)
if err != nil {
t.Fatalf("AssembleFileARM64: %v", err)
}
gasmObj, err := img.GOObjectAARCH64("a64dlink", "main_arm64.s")
if err != nil {
t.Fatalf("GOObjectAARCH64: %v", err)
}
if err := os.WriteFile(asmObj, gasmObj, 0o644); err != nil {
t.Fatalf("write gasm object: %v", err)
}
linkCmd := exec.Command("bash", "-c", "cd "+dir+" && "+linkLine)
linkCmd.Env = append(os.Environ(), "GOARCH=arm64")
if out, err := linkCmd.CombinedOutput(); err != nil {
t.Fatalf("re-link with gasm object: %v\n%s", err, out)
}
binData, err := os.ReadFile(filepath.Join(dir, "prog"))
if err != nil {
t.Fatal(err)
}
if !strings.Contains(string(binData), "keepme") {
t.Error("binary does not contain the keepme symbol")
}
}
+2763 -97
View File
File diff suppressed because it is too large Load Diff
+758 -27
View File
@@ -27,7 +27,14 @@ package asm
// Uncond-branch 0x6B<<25 | opc<<21 | Rn<<5 | Rd (BR/BLR/RET)
// ADR/ADRP p<<31 | 0x10<<24 | immlo<<29 | immhi<<5 | Rd
import "maps"
import (
"maps"
"math/bits"
"strconv"
"strings"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
)
// arm64RegNum returns the 5-bit register number for an AArch64 register name:
// R0-R30 (integer), F0-F31 (floating point), and the ABI aliases the
@@ -72,6 +79,11 @@ func arm64RegNum(name string) int {
return 17
case "R18":
return 18
case "R18_PLATFORM":
// The toolchain's Windows spelling: R18 is renamed R18_PLATFORM in
// cmd/asm/internal/arch so assembly cannot use it by accident, and
// sys_windows_arm64.s references it only through this name.
return 18
case "R19":
return 19
case "R20":
@@ -153,6 +165,67 @@ func a64MoveWide(sf, opc, hw, imm16, rd uint32) uint32 {
return sf<<31 | opc<<29 | 0x25<<23 | hw<<21 | imm16<<5 | rd
}
// ---- logical immediate ----
// a64LogicalImm encodes v as the AArch64 logical (bitmask) immediate for the
// given lane width (32 or 64): it returns the N, immr and imms fields of the
// imm13 encoding. The algorithm mirrors cmd/internal/obj/arm64's
// encodeLogicalImmArrEncoding: replicate the value, shrink it to the smallest
// repeating element, find the run of ones and its rotation. ok is false when
// v is not expressible (all zeros, all ones, or not a single cyclic run).
func a64LogicalImm(v int64, width int) (n, immr, imms uint32, ok bool) {
u := uint64(v)
if width == 32 {
u &= 0xFFFFFFFF
}
size := uint64(width)
mask := ^uint64(0)
if size < 64 {
mask = uint64(1)<<size - 1
}
u &= mask
// All zeros and all ones are MOV territory, not bitmask immediates.
if u == 0 || u == mask {
return 0, 0, 0, false
}
// Shrink to the smallest repeating element.
for size > 2 {
half := size / 2
hm := uint64(1)<<half - 1
if u&hm == u>>half&hm {
size = half
u &= hm
} else {
break
}
}
ones := bits.OnesCount64(u)
// Find the right-rotation that lays the ones out contiguously at the
// bottom of the element; the hardware applies the inverse rotation.
em := uint64(1)<<size - 1
expected := uint64(1)<<ones - 1
rot := -1
for r := 0; r < int(size); r++ {
rotated := u>>r | u<<(int(size)-r)
if size < 64 {
rotated &= em
}
if rotated == expected {
rot = r
break
}
}
if rot < 0 {
return 0, 0, 0, false
}
if size == 64 {
n = 1
}
immr = uint32((int(size) - rot) % int(size))
imms = ^uint32(uint32(size*2-1))&0x3F | uint32(ones-1)
return n, immr, imms, true
}
// ---- load/store (unsigned immediate, scaled) ----
// a64LSU encodes a load/store register (unsigned immediate, scaled):
@@ -240,6 +313,8 @@ const (
a64CondLT = 0xb
a64CondGT = 0xc
a64CondLE = 0xd
a64CondAL = 0xe
a64CondNV = 0xf
)
// arm64CondMap maps Go assembler condition mnemonics to AArch64 condition codes.
@@ -260,6 +335,8 @@ var arm64CondMap = map[string]uint32{
"LT": a64CondLT,
"GT": a64CondGT,
"LE": a64CondLE,
"AL": a64CondAL,
"NV": a64CondNV,
}
// ---- instruction format tags ----
@@ -267,28 +344,47 @@ var arm64CondMap = map[string]uint32{
type a64Format uint8
const (
a64FDPSR a64Format = iota // data-processing (shifted register): ADD, SUB, AND, ORR, EOR, etc.
a64FMovWide // move wide: MOVZ, MOVN, MOVK
a64FBranch // unconditional branch (B/BL)
a64FBranchCond // conditional branch (B.cond)
a64FUncondBranch // unconditional branch register (BR/BLR/RET)
a64FADR // ADR/ADRP
a64FEXTR // EXTR
a64FBitfield // bitfield: BFI/BFXIL/SBFM/UBFM/BFM
a64FShift // shifts: LSL/LSR/ASR alias SBFM/UBFM, ROR aliases EXTR; register forms are two-source
a64FDPR4 // data-processing 4-register: MADD/MSUB, Ra in bits 14:10
a64FFP3 // FP 3-operand (Rm, Rn, Rd): FADD, FSUB, FMUL, FDIV, etc.
a64FFPUnary // FP unary (Rn, Rd): FMOV, FABS, FNEG, FSQRT, FCVT, FRINT*
a64FFP4 // FP 4-operand FMA (Ra, Rm, Rn, Rd): FMADD, FMSUB, etc.
a64FFPCmp // FP compare (Rm, Rn): FCMP, FCMPE
a64FFPCCmp // FP conditional compare (Rm, Rn, nzcv, cond): FCCMP, FCCMPE
a64FFPCvt // FP↔integer conversion: FCVTZS, SCVTF, etc.
a64FFPSel // FP conditional select (Rm, Rn, Rd, cond): FCSEL
a64FCRC32 // CRC32
a64FCSEL // conditional select: CSEL, CSINC, CSINV, CSNEG
a64FExcl // exclusive load/store: LDXR, STXR, LDAXR, STLXR and pair forms LDXP, STXP
a64FLSE // LSE atomics: LDADD, CAS, SWP
a64FSIMD3 // SIMD 3-operand: VADD, VSUB, VMUL
a64FDPSR a64Format = iota // data-processing (shifted register): ADD, SUB, AND, ORR, EOR, etc.
a64FMovWide // move wide: MOVZ, MOVN, MOVK
a64FBranch // unconditional branch (B/BL)
a64FBranchCond // conditional branch (B.cond)
a64FUncondBranch // unconditional branch register (BR/BLR/RET)
a64FADR // ADR/ADRP
a64FEXTR // EXTR
a64FBitfield // bitfield: BFI/BFXIL/SBFM/UBFM/BFM
a64FBitfieldAlias // bitfield alias: BFI/BFXIL/SBFIZ/UBFIZ, ($lsb, Rn, $width, Rd)
a64FShift // shifts: LSL/LSR/ASR alias SBFM/UBFM, ROR aliases EXTR; register forms are two-source
a64FDPR4 // data-processing 4-register: MADD/MSUB, Ra in bits 14:10
a64FFP3 // FP 3-operand (Rm, Rn, Rd): FADD, FSUB, FMUL, FDIV, etc.
a64FFPUnary // FP unary (Rn, Rd): FMOV, FABS, FNEG, FSQRT, FCVT, FRINT*
a64FFP4 // FP 4-operand FMA (Ra, Rm, Rn, Rd): FMADD, FMSUB, etc.
a64FFPCmp // FP compare (Rm, Rn): FCMP, FCMPE
a64FFPCCmp // FP conditional compare (Rm, Rn, nzcv, cond): FCCMP, FCCMPE
a64FFPCvt // FP↔integer conversion: FCVTZS, SCVTF, etc.
a64FFPSel // FP conditional select (Rm, Rn, Rd, cond): FCSEL
a64FCRC32 // CRC32
a64FCSEL // conditional select: CSEL, CSINC, CSINV, CSNEG
a64FExcl // exclusive load/store: LDXR, STXR, LDAXR, STLXR and pair forms LDXP, STXP
a64FLSE // LSE atomics: LDADD, CAS, SWP
a64FDP1 // data-processing (1 source): RBIT, REV, CLZ, CLS
a64FBitfield2 // bitfield extract: UBFX, SBFX and the W forms
a64FCondCmp // conditional compare: CCMP, CCMN
a64FBranch19 // compare-and-branch: CBZ, CBNZ and the W forms
a64FTestBranch // test-and-branch: TBZ, TBNZ and the W forms
a64FPair // load/store pair: LDP, STP, LDPW, STPW, FLDPD, FSTPD
a64FAcqRel // acquire/release: LDAR family, STLR family
a64FSys // system: BRK, SVC, DMB, DSB, ISB, DC, MRS, MSR, PRFM
a64FCrypto2 // crypto 2-register: AESD, AESE, AESIMC, AESMC, SHA1H, ...
a64FCrypto3 // crypto 3-register: SHA1C, SHA256H, SHA512SU1, ...
a64FSIMDV // SIMD 3-register with arrangement: VADD, VAND, VCMEQ, VZIP1, ...
a64FSIMDVZero // SIMD compare against zero: VCMEQ $0, Vn, Vd
a64FSIMDV2 // SIMD 2-register with arrangement: VREV32, VREV64, VUADDLV, VMOV
a64FSIMDV4 // SIMD 4-register / imm 3-register: VEOR3, VBCAX, VXAR, VEXT
a64FVTBL // SIMD table lookup: VTBL
a64FDUP // SIMD element moves: VDUP, VMOV with element indices
a64FVLDST // SIMD structure loads/stores: VLD1, VST1, VLD1R, VLD4R
a64FShiftImm // SIMD shift by immediate: VSHL, VUSHR, VSRI
a64FMoviLit // VMOVS/VMOVD/VMOVQ with a large constant (literal pool)
)
// a64Enc is one instruction's encoding: its bit layout (format) and the
@@ -401,6 +497,16 @@ func init() {
a64InstrTable["MADDW"] = a64Enc{format: a64FDPR4, op: 0<<31 | 0x1b<<24}
a64InstrTable["MSUB"] = a64Enc{format: a64FDPR4, op: 1<<31 | 0x1b<<24 | 1<<15}
a64InstrTable["MSUBW"] = a64Enc{format: a64FDPR4, op: 0<<31 | 0x1b<<24 | 1<<15}
// The widening multiplies: a 64-bit result riding the same layout, the
// three-operand forms reading the accumulate register as ZR.
a64InstrTable["SMADDL"] = a64Enc{format: a64FDPR4, op: 1<<31 | 0x1b<<24 | 1<<21}
a64InstrTable["UMADDL"] = a64Enc{format: a64FDPR4, op: 1<<31 | 0x1b<<24 | 1<<21 | 1<<23}
a64InstrTable["SMSUBL"] = a64Enc{format: a64FDPR4, op: 1<<31 | 0x1b<<24 | 1<<21 | 1<<15}
a64InstrTable["UMSUBL"] = a64Enc{format: a64FDPR4, op: 1<<31 | 0x1b<<24 | 1<<21 | 1<<23 | 1<<15}
a64InstrTable["SMULL"] = a64Enc{format: a64FDPR4, op: 1<<31 | 0x1b<<24 | 1<<21 | 31<<10}
a64InstrTable["UMULL"] = a64Enc{format: a64FDPR4, op: 1<<31 | 0x1b<<24 | 1<<21 | 1<<23 | 31<<10}
a64InstrTable["SMNEGL"] = a64Enc{format: a64FDPR4, op: 1<<31 | 0x1b<<24 | 1<<21 | 1<<15 | 31<<10}
a64InstrTable["UMNEGL"] = a64Enc{format: a64FDPR4, op: 1<<31 | 0x1b<<24 | 1<<21 | 1<<23 | 1<<15 | 31<<10}
// ---- move wide ----
// MOVZ/MOVN/MOVK
@@ -450,6 +556,15 @@ func init() {
// ---- bitfield ----
a64InstrTable["BFM"] = a64Enc{format: a64FBitfield, op: 1<<31 | 1<<29 | 0x26<<23 | 1<<22}
a64InstrTable["BFMW"] = a64Enc{format: a64FBitfield, op: 0<<31 | 1<<29 | 0x26<<23 | 0<<22}
// The four-operand bitfield aliases: ($lsb, Rn, $width, Rd).
a64InstrTable["BFI"] = a64Enc{format: a64FBitfieldAlias, op: 1<<31 | 1<<29 | 0x26<<23 | 1<<22}
a64InstrTable["BFIW"] = a64Enc{format: a64FBitfieldAlias, op: 0<<31 | 1<<29 | 0x26<<23}
a64InstrTable["BFXIL"] = a64Enc{format: a64FBitfieldAlias, op: 1<<31 | 1<<29 | 0x26<<23 | 1<<22}
a64InstrTable["BFXILW"] = a64Enc{format: a64FBitfieldAlias, op: 0<<31 | 1<<29 | 0x26<<23}
a64InstrTable["SBFIZ"] = a64Enc{format: a64FBitfieldAlias, op: 0x93400000}
a64InstrTable["SBFIZW"] = a64Enc{format: a64FBitfieldAlias, op: 0x13000000}
a64InstrTable["UBFIZ"] = a64Enc{format: a64FBitfieldAlias, op: 0x53000000}
a64InstrTable["UBFIZW"] = a64Enc{format: a64FBitfieldAlias, op: 0x33000000}
a64InstrTable["SBFM"] = a64Enc{format: a64FBitfield, op: 1<<31 | 0<<29 | 0x26<<23 | 1<<22}
a64InstrTable["SBFMW"] = a64Enc{format: a64FBitfield, op: 0<<31 | 0<<29 | 0x26<<23 | 0<<22}
a64InstrTable["UBFM"] = a64Enc{format: a64FBitfield, op: 1<<31 | 2<<29 | 0x26<<23 | 1<<22}
@@ -622,10 +737,626 @@ func init() {
a64InstrTable["SWPD"] = a64Enc{format: a64FLSE, op: 3<<30 | 0x1c1<<21 | 0x20<<10}
a64InstrTable["SWPW"] = a64Enc{format: a64FLSE, op: 2<<30 | 0x1c1<<21 | 0x20<<10}
// ---- SIMD basics ----
a64InstrTable["VADD"] = a64Enc{format: a64FSIMD3, op: 0x0e208400}
a64InstrTable["VSUB"] = a64Enc{format: a64FSIMD3, op: 0x2e208400}
a64InstrTable["VMUL"] = a64Enc{format: a64FSIMD3, op: 0x0e209c00}
// ---- SIMD: the arrangement-aware tables in this file carry VADD,
// VSUB, VMUL and every other three-register vector op. ----
// ---- data-processing (1 source): sf 10 11010110 opcode 00000 Rn Rd ----
dp1 := map[string]uint32{
"RBIT": 0xdac00000, "REV16": 0xdac00400, "REV32": 0xdac00800,
"REV": 0xdac00c00, "CLZ": 0xdac01000, "CLS": 0xdac01400,
"RBITW": 0x5ac00000, "REVW": 0x5ac00800, "CLZW": 0x5ac01000, "CLSW": 0x5ac01400,
// Extend and byte-reverse: the UBFM/SBFM aliases with imms fixing
// the source width.
"SXTB": 0x93401c00, "SXTBW": 0x13001c00, "SXTH": 0x93403c00,
"SXTHW": 0x13003c00, "SXTW": 0x93407c00,
"UXTB": 0x53001c00, "UXTBW": 0x53001c00, "UXTH": 0x53403c00,
"UXTHW": 0x53003c00, "UXTW": 0x53407c00,
"REV16W": 0x5ac00400,
}
for m, op := range dp1 {
a64InstrTable[m] = a64Enc{format: a64FDP1, op: op}
}
// ---- bitfield extract: the UBFM/SBFM bases, immediate operands wrap ----
a64InstrTable["UBFX"] = a64Enc{format: a64FBitfield2, op: 0xd3400000}
a64InstrTable["SBFX"] = a64Enc{format: a64FBitfield2, op: 0x93400000}
a64InstrTable["UBFXW"] = a64Enc{format: a64FBitfield2, op: 0x53000000}
a64InstrTable["SBFXW"] = a64Enc{format: a64FBitfield2, op: 0x13000000}
// ---- conditional compare: sf 1 1 101001 0 imm5/Rm cond op2 Rn nzcv ----
a64InstrTable["CCMP"] = a64Enc{format: a64FCondCmp, op: 0xfa400000}
a64InstrTable["CCMN"] = a64Enc{format: a64FCondCmp, op: 0xba400000}
a64InstrTable["CCMPW"] = a64Enc{format: a64FCondCmp, op: 0x7a400000}
a64InstrTable["CCMNW"] = a64Enc{format: a64FCondCmp, op: 0x3a400000}
// ---- system operations ----
for _, m := range []string{"BRK", "SVC", "DMB", "DSB", "ISB", "CLREX", "HINT", "BTI", "HLT", "SMC", "HVC", "DCPS1", "DCPS2", "DCPS3", "DRPS", "ERET", "AUTIASP", "AUTIBSP", "AUTIA1716", "AUTIB1716", "SEVL", "SEV", "WFE", "WFI", "YIELD", "DC", "MRS", "MSR", "PRFM"} {
a64InstrTable[m] = a64Enc{format: a64FSys}
}
// ---- compare/test and branch ----
a64InstrTable["CBZ"] = a64Enc{format: a64FBranch19, op: 0xb4000000}
a64InstrTable["CBZW"] = a64Enc{format: a64FBranch19, op: 0x34000000}
a64InstrTable["CBNZ"] = a64Enc{format: a64FBranch19, op: 0xb5000000}
a64InstrTable["CBNZW"] = a64Enc{format: a64FBranch19, op: 0x35000000}
a64InstrTable["TBZ"] = a64Enc{format: a64FTestBranch, op: 0x36000000}
a64InstrTable["TBNZ"] = a64Enc{format: a64FTestBranch, op: 0x37000000}
// ---- load/store pair (signed offset) ----
a64InstrTable["LDP"] = a64Enc{format: a64FPair, op: 0xa9400000}
a64InstrTable["LDPW"] = a64Enc{format: a64FPair, op: 0x29400000}
a64InstrTable["STP"] = a64Enc{format: a64FPair, op: 0xa9000000}
a64InstrTable["STPW"] = a64Enc{format: a64FPair, op: 0x29000000}
a64InstrTable["FLDPD"] = a64Enc{format: a64FPair, op: 0x6d400000}
a64InstrTable["FSTPD"] = a64Enc{format: a64FPair, op: 0x6d000000}
// ---- acquire/release loads and stores ----
a64InstrTable["LDAR"] = a64Enc{format: a64FAcqRel, op: 0xc8dffc00}
a64InstrTable["LDARB"] = a64Enc{format: a64FAcqRel, op: 0x08dffc00}
a64InstrTable["LDARH"] = a64Enc{format: a64FAcqRel, op: 0x48dffc00}
a64InstrTable["LDARW"] = a64Enc{format: a64FAcqRel, op: 0x88dffc00}
a64InstrTable["STLR"] = a64Enc{format: a64FAcqRel, op: 0xc89ffc00}
a64InstrTable["STLRB"] = a64Enc{format: a64FAcqRel, op: 0x089ffc00}
a64InstrTable["STLRH"] = a64Enc{format: a64FAcqRel, op: 0x489ffc00}
a64InstrTable["STLRW"] = a64Enc{format: a64FAcqRel, op: 0x889ffc00}
// ---- LSE atomics with acquire and release semantics ----
// CAS carries a preset fixed op field and a real Rs; the LDADD/LDCLR/
// LDOR/SWP families leave Rs free for the returned value.
lse := map[string]uint32{
"CASALD": 0xc8e0fc00,
"CASALW": 0x88e0fc00,
"LDADDALD": 0xf8e00000,
"LDADDALW": 0xb8e00000,
"LDCLRALB": 0x38e01000,
"LDCLRALW": 0xb8e01000,
"LDCLRALD": 0xf8e01000,
"LDORALB": 0x38e03000,
"LDORALW": 0xb8e03000,
"LDORALD": 0xf8e03000,
"SWPALB": 0x38e08000,
"SWPALW": 0xb8e08000,
"SWPALD": 0xf8e08000,
}
for m, op := range lse {
a64InstrTable[m] = a64Enc{format: a64FLSE, op: op}
}
// The remaining width and ordering spellings of the same shapes, and the
// CAS compare-and-swap family, word-verified against go tool asm.
lseMore := map[string]uint32{
"LDADDAB": 0x38a00000,
"LDADDAH": 0x78a00000,
"LDADDALB": 0x38e00000,
"LDADDALH": 0x78e00000,
"LDADDLB": 0x38600000,
"LDADDLD": 0xf8600000,
"LDADDLH": 0x78600000,
"LDADDLW": 0xb8600000,
"LDCLRAB": 0x38a01000,
"LDCLRAH": 0x78a01000,
"LDCLRALH": 0x78e01000,
"LDCLRB": 0x38201000,
"LDCLRD": 0xf8201000,
"LDCLRH": 0x78201000,
"LDCLRLB": 0x38601000,
"LDCLRLD": 0xf8601000,
"LDCLRLH": 0x78601000,
"LDCLRLW": 0xb8601000,
"LDCLRW": 0xb8201000,
"LDEORAB": 0x38a02000,
"LDEORAD": 0xf8a02000,
"LDEORAH": 0x78a02000,
"LDEORALB": 0x38e02000,
"LDEORALH": 0x78e02000,
"LDEORAW": 0xb8a02000,
"LDEORB": 0x38202000,
"LDEORD": 0xf8202000,
"LDEORH": 0x78202000,
"LDEORLB": 0x38602000,
"LDEORLD": 0xf8602000,
"LDEORLH": 0x78602000,
"LDEORLW": 0xb8602000,
"LDEORW": 0xb8202000,
"LDORAB": 0x38a03000,
"LDORAD": 0xf8a03000,
"LDORAH": 0x78a03000,
"LDORALH": 0x78e03000,
"LDORAW": 0xb8a03000,
"LDORB": 0x38203000,
"LDORD": 0xf8203000,
"LDORH": 0x78203000,
"LDORLB": 0x38603000,
"LDORLD": 0xf8603000,
"LDORLH": 0x78603000,
"LDORLW": 0xb8603000,
"LDORW": 0xb8203000,
"SWPAB": 0x38a08000,
"SWPAD": 0xf8a08000,
"SWPAH": 0x78a08000,
"SWPALH": 0x78e08000,
"SWPAW": 0xb8a08000,
"SWPB": 0x38208000,
"SWPH": 0x78208000,
"SWPLB": 0x38608000,
"SWPLD": 0xf8608000,
"SWPLH": 0x78608000,
"SWPLW": 0xb8608000,
"CASAD": 0xc8e07c00,
"CASALB": 0x08e0fc00,
"CASLW": 0x88a0fc00,
}
for m, op := range lseMore {
a64InstrTable[m] = a64Enc{format: a64FLSE, op: op}
}
// ---- carry-setting/carry-using arithmetic and widening multiply ----
// MUL and SMULH/UMULH are the MADD/MSUB layout with the accumulate
// register preset to ZR (bits 14:10 = 11111).
dpsrExtra := map[string]uint32{
"ADC": 0x9a000000, "ADCW": 0x1a000000,
"ADCS": 0xba000000, "ADCSW": 0x3a000000,
"SBC": 0xda000000, "SBCW": 0x5a000000,
"SBCS": 0xfa000000, "SBCSW": 0x7a000000,
// MNEG/MSUB and NGC/SBC with the complementing register preset to ZR.
"MNEG": 0x9b00fc00, "MNEGW": 0x1b00fc00,
"NGC": 0xda000000, "NGCW": 0x5a000000,
"NGCS": 0xfa000000, "NGCSW": 0x7a000000,
"NEGSW": 0x6b000000,
"MUL": 0x9b007c00, "MULW": 0x1b007c00,
"SMULH": 0x9b407c00, "UMULH": 0x9bc07c00,
}
for m, op := range dpsrExtra {
a64InstrTable[m] = a64Enc{format: a64FDPSR, op: op}
}
// ---- crypto, 2-register (Rn, Rd) and 3-register (Rm, Rn, Rd) forms ----
crypto2 := map[string]uint32{
"AESD": 0x4e285800, "AESE": 0x4e284800,
"AESIMC": 0x4e287800, "AESMC": 0x4e286800,
"SHA1H": 0x5e280800, "SHA1SU1": 0x5e281800,
"SHA256SU0": 0x5e282800, "SHA512SU0": 0xcec08000,
}
for m, op := range crypto2 {
a64InstrTable[m] = a64Enc{format: a64FCrypto2, op: op}
}
crypto3 := map[string]uint32{
"SHA1C": 0x5e000000, "SHA1P": 0x5e001000,
"SHA1M": 0x5e002000, "SHA1SU0": 0x5e003000,
"SHA256H": 0x5e004000, "SHA256H2": 0x5e005000,
"SHA256SU1": 0x5e006000, "SHA512H": 0xce608000,
"SHA512H2": 0xce608400, "SHA512SU1": 0xce608800,
}
for m, op := range crypto3 {
a64InstrTable[m] = a64Enc{format: a64FCrypto3, op: op}
}
// ---- arrangement-aware SIMD, see a64SimdVTable and a64SimdV2Table ----
a64InstrTable["VEOR3"] = a64Enc{format: a64FSIMDV4, op: 0xce000000}
a64InstrTable["VBCAX"] = a64Enc{format: a64FSIMDV4, op: 0xce200000}
a64InstrTable["VXAR"] = a64Enc{format: a64FSIMDV4, op: 0xce800000}
a64InstrTable["VEXT"] = a64Enc{format: a64FSIMDV4, op: 0x2e000000}
a64InstrTable["VTBL"] = a64Enc{format: a64FVTBL}
a64InstrTable["VDUP"] = a64Enc{format: a64FDUP}
a64InstrTable["VMOVS"] = a64Enc{format: a64FMoviLit, op: 0xbd400000}
a64InstrTable["VMOVD"] = a64Enc{format: a64FMoviLit, op: 0xfd400000}
a64InstrTable["VMOVQ"] = a64Enc{format: a64FMoviLit, op: 0x3dc00000}
a64InstrTable["VSHL"] = a64Enc{format: a64FShiftImm, op: 0x0f000000 | 21<<10}
a64InstrTable["VUSHR"] = a64Enc{format: a64FShiftImm, op: 0x2f000000 | 1<<10}
a64InstrTable["VSRI"] = a64Enc{format: a64FShiftImm, op: 0x2f000000 | 17<<10}
a64InstrTable["VSSHR"] = a64Enc{format: a64FShiftImm, op: 0x0f000000 | 1<<10}
a64InstrTable["VSRA"] = a64Enc{format: a64FShiftImm, op: 0x0f000000 | 17<<10}
a64InstrTable["VSRSHR"] = a64Enc{format: a64FShiftImm, op: 0x0f000000 | 9<<10}
a64InstrTable["VSLI"] = a64Enc{format: a64FShiftImm, op: 0x2f000000 | 21<<10}
a64InstrTable["VSQSHL"] = a64Enc{format: a64FShiftImm, op: 0x0f000000 | 29<<10}
a64InstrTable["VUQSHL"] = a64Enc{format: a64FShiftImm, op: 0x2f000000 | 29<<10}
a64InstrTable["VLD1"] = a64Enc{format: a64FVLDST}
a64InstrTable["VLD1.P"] = a64Enc{format: a64FVLDST, op: 1}
a64InstrTable["VST1"] = a64Enc{format: a64FVLDST}
a64InstrTable["VST1.P"] = a64Enc{format: a64FVLDST, op: 1}
a64InstrTable["VLD1R"] = a64Enc{format: a64FVLDST}
a64InstrTable["VLD1R.P"] = a64Enc{format: a64FVLDST, op: 1}
a64InstrTable["VLD4R"] = a64Enc{format: a64FVLDST}
a64InstrTable["VLD4R.P"] = a64Enc{format: a64FVLDST, op: 1}
}
// a64SimdVSpec is one arrangement-aware SIMD instruction: the 8B base word,
// the set of arrangements it accepts as a bitmask over the a64Arr index and,
// for instructions that exist at a single arrangement and carry that
// arrangement's bits inside the base already, the fixed flag.
type a64SimdVSpec struct {
base uint32
arrs uint16
fixed bool
}
// a64Arr names the vector arrangements the encoders deal with, indexed by
// a64Arr. The source spellings put the element letter first: B8, H4, S2,
// D1 and the 128-bit halves B16, H8, S4, D2.
const (
a64Arr8B = iota
a64Arr16B
a64Arr4H
a64Arr8H
a64Arr2S
a64Arr4S
a64Arr2D
a64ArrD1
a64ArrQ1
a64ArrCount
)
// a64ArrNames maps an arrangement to its source spelling (element letter
// first, as the toolchain writes it).
var a64ArrNames = [a64ArrCount]string{
a64Arr8B: "B8", a64Arr16B: "B16", a64Arr4H: "H4", a64Arr8H: "H8",
a64Arr2S: "S2", a64Arr4S: "S4", a64Arr2D: "D2", a64ArrD1: "D1", a64ArrQ1: "Q1",
}
// a64ArrIndex resolves a source spelling to its a64Arr index, -1 when
// unknown.
func a64ArrIndex(s string) int {
for i, n := range a64ArrNames {
if n == s {
return i
}
}
return -1
}
// a64ElemLetter reports whether s is a bare element spelling (B, H, S, D, Q)
// as it appears in element operands such as V13.S[0].
func a64ElemLetter(s string) bool {
switch s {
case "B", "H", "S", "D", "Q":
return true
}
return false
}
// fpSimdArrs and fpAcrossArrs bound the arrangements the FP SIMD forms
// accept: H, S and D widths for the pairwise data-processing, H and S for
// the across-vector reductions.
var fpSimdArrs = uint16(1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D)
var fpAcrossArrs = uint16(1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S)
// a64SimdQOnly names the forms whose arrangement contributes the 128-bit
// flag alone, without the size bits: the FP converts, the FP round-to-integral
// and pairwise compares among them. Word-verified against go tool asm.
var a64SimdQOnly = map[string]bool{
"VSCVTF": true, "VUCVTF": true, "VFCVTZS": true, "VFCVTZU": true,
"VFABS": true, "VFNEG": true, "VFSQRT": true,
"VFRINTN": true, "VFRINTP": true, "VFRINTM": true, "VFRINTZ": true,
"VFADDP": true, "VFMAXP": true, "VFMAXNMP": true,
"VFMAXV": true, "VFMAXNMV": true,
}
// a64ArrBits carries the fixed bits an arrangement contributes to the
// three-same word shape: the element size at bits 23:22 and the 128-bit
// flag at bit 30. Bit 29 belongs to the instruction's own base.
var a64ArrBits = [a64ArrCount]uint32{
a64Arr8B: 0,
a64Arr16B: 1 << 30,
a64Arr4H: 1 << 22,
a64Arr8H: 1<<30 | 1<<22,
a64Arr2S: 1 << 23,
a64Arr4S: 1<<30 | 1<<23,
a64Arr2D: 1<<30 | 1<<23 | 1<<22,
a64ArrD1: 1<<23 | 1<<22,
a64ArrQ1: 0,
}
// a64SimdVTable holds the arrangement-aware three-register SIMD
// instructions (word = base | arrBits | Rm<<16 | Rn<<5 | Rd). Every base
// word and arrangement bit was read off go tool asm.
var a64SimdVTable = map[string]a64SimdVSpec{
"VADD": {0x0e208400, 0x7f, false},
"VSUB": {0x2e208400, 0x7f, false},
"VMUL": {0x0e209c00, 0x3f, false}, // no 2D: integer multiply stops at 4S
"VAND": {0x0e201c00, 0x03, false}, // logical ops accept 8B and 16B only
"VEOR": {0x2e201c00, 0x03, false},
"VORR": {0x0ea01c00, 0x03, false},
"VADDP": {0x0e20bc00, 0x7f, false},
"VZIP1": {0x0e003800, 0x7f, false},
"VZIP2": {0x0e007800, 0x7f, false},
"VCMEQ": {0x2e208c00, 0x7f, false},
"VCMGE": {0x0e203c00, 0x7f, false},
"VCMGT": {0x0e203400, 0x7f, false},
"VCMHI": {0x2e203400, 0x7f, false},
"VCMHS": {0x2e203c00, 0x7f, false},
// FP compares take H, S and D arrangements only (the toolchain rejects
// the byte forms), and VFCMLE/VFCMLT have no register form at all.
"VFCMEQ": {0x0e20e400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VFCMGE": {0x2e20e400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VFCMGT": {0x2ea0e400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
// FP arithmetic shares the same arrangement restriction.
"VFADD": {0x0e20d400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VFSUB": {0x0ea0d400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VFMUL": {0x2e20dc00, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VFDIV": {0x2e20fc00, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VFMAX": {0x0e20f400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VFMIN": {0x0ea0f400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VFMAXNM": {0x0e20c400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VFMINNM": {0x0ea0c400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VFMLA": {0x0e20cc00, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VFMLS": {0x0ea0cc00, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
// Saturating, halving, polynomial and pairwise arithmetic, the logical
// VBIT/VBSL family and the FP pairwise forms: word-verified against go
// tool asm.
"VBIC": {0x0e601c00, 0x7f, false},
"VBIF": {0x2ee01c00, 0x7f, false},
"VBIT": {0x6ea01c00, 0x7f, false},
"VBSL": {0x6e601c00, 0x7f, false},
"VCMTST": {0x0e208c00, 0x7f, false},
"VFADDP": {0x2e20d400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VFMAXP": {0x2e20f400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VFMINP": {0x6ea0f400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VFMAXNMP": {0x2e20c400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VFMINNMP": {0x6ea0c400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
"VMLA": {0x4ea09400, 0x7f, false},
"VMLS": {0x6ea09400, 0x7f, false},
"VORN": {0x4ee01c00, 0x7f, false},
"VSHADD": {0x4ea00400, 0x7f, false},
"VSRHADD": {0x4ea01400, 0x7f, false},
"VUHADD": {0x6ea00400, 0x7f, false},
"VURHADD": {0x6ea01400, 0x7f, false},
"VSMAX": {0x4ea06400, 0x7f, false},
"VSMIN": {0x4ea06c00, 0x7f, false},
"VSMAXP": {0x4ea0a400, 0x7f, false},
"VSMINP": {0x4ea0ac00, 0x7f, false},
"VUMAX": {0x2e206400, 0x7f, false},
"VUMIN": {0x2e206c00, 0x7f, false},
"VUMAXP": {0x6ea0a400, 0x7f, false},
"VUMINP": {0x6ea0ac00, 0x7f, false},
"VSQADD": {0x4ea00c00, 0x7f, false},
"VUQADD": {0x6ea00c00, 0x7f, false},
"VSQSUB": {0x4ea02c00, 0x7f, false},
"VUQSUB": {0x6ea02c00, 0x7f, false},
"VSSHL": {0x4ee04400, 0x7f, false},
"VUSHL": {0x6ee04400, 0x7f, false},
"VUZP1": {0x0e001800, 0x7f, false},
"VUZP2": {0x4ec05800, 0x7f, false},
"VTRN1": {0x4ec02800, 0x7f, false},
"VTRN2": {0x4ec06800, 0x7f, false},
"VRAX1": {0xce608c00, 1 << a64Arr2D, true}, // SHA3 group, D2 only
"VPMULL": {0x0e20e000, 1<<a64Arr8B | 1<<a64ArrD1, false},
"VPMULL2": {0x0e20e000, 1<<a64Arr16B | 1<<a64Arr2D, false},
}
// a64SimdVZero holds the compare-against-zero words of the SIMD compares
// spelled with a $0 first operand (word = base | arrBits | Rn<<5 | Rd).
// VCMHI and VCMHS have no zero form: the toolchain reports an illegal
// combination for them, so they stay out and the encoder rejects the shape.
var a64SimdVZero = map[string]uint32{
"VCMEQ": 0x0e209800,
"VCMGT": 0x0e208800,
"VCMGE": 0x2e208800,
"VCMLT": 0x0e20a800,
"VCMLE": 0x2e209800,
// FP compares against (0.0): the register forms above carry the U and op
// bits; the zero forms reshape them.
"VFCMEQ": 0x0ea0d800,
"VFCMGE": 0x2ea0c800,
"VFCMGT": 0x0ea0c800,
"VFCMLE": 0x2ea0d800,
"VFCMLT": 0x0ea0e800,
}
// a64SimdV2Table holds the arrangement-aware two-register SIMD instructions
// (word = base | arrBits | Rn<<5 | Rd). VMOV is served from here too, with
// the register pair spelling ORR Vd, Vn, Vm.
var a64SimdV2Table = map[string]a64SimdVSpec{
"VREV32": {0x2e200800, 1<<a64Arr8B | 1<<a64Arr16B | 1<<a64Arr4H | 1<<a64Arr8H, false},
"VREV64": {0x0e200800, 0x3f, false},
"VREV16": {0x0e201800, 1<<a64Arr8B | 1<<a64Arr16B, false},
"VUADDLV": {0x2e303800, 0x3f, false},
"VMOV": {0x0ea01c00, 1<<a64Arr8B | 1<<a64Arr16B, false},
// Two-register data-processing across one arrangement.
"VABS": {0x0e20b800, 0x7f, false},
"VNEG": {0x2e20b800, 0x7f, false},
"VCLS": {0x0e204800, 0x7f, false},
"VCLZ": {0x2e204800, 0x7f, false},
"VCNT": {0x0e205800, 0x7f, false},
"VNOT": {0x2e205800, 0x7f, false},
"VSQABS": {0x0e207800, 0x7f, false},
"VSQNEG": {0x2e207800, 0x7f, false},
"VRBIT": {0x6e605800, 0x7f, false},
"VSCVTF": {0x4e21d800, fpSimdArrs, false},
"VUCVTF": {0x6e21d800, fpSimdArrs, false},
"VFCVTZS": {0x4ea1b800, fpSimdArrs, false},
"VFCVTZU": {0x6ea1b800, fpSimdArrs, false},
"VFABS": {0x0ea0f800, fpSimdArrs, false},
"VFNEG": {0x2ea0f800, fpSimdArrs, false},
"VFSQRT": {0x2ea1f800, fpSimdArrs, false},
"VFRINTN": {0x0e218800, fpSimdArrs, false},
"VFRINTP": {0x0ea18800, fpSimdArrs, false},
"VFRINTM": {0x0e219800, fpSimdArrs, false},
"VFRINTZ": {0x0ea19800, fpSimdArrs, false},
// Across-vector reductions: the operand arrangement rides as usual and
// the destination stays a bare V register.
"VADDV": {0x0e31b800, 0x3f, false},
"VSMAXV": {0x0e30a800, 0x3f, false},
"VSMINV": {0x0e31a800, 0x3f, false},
"VUMAXV": {0x2e30a800, 0x3f, false},
"VUMINV": {0x2e31a800, 0x3f, false},
"VFMAXV": {0x2e30f800, fpAcrossArrs, false},
"VFMINV": {0x2eb0f800, fpAcrossArrs, false},
"VFMAXNMV": {0x2e30c800, fpAcrossArrs, false},
"VFMINNMV": {0x2eb0c800, fpAcrossArrs, false},
}
// a64CryptoArr is the arrangement each crypto instruction's operands must
// carry when they spell one at all; a bare V/F spelling is accepted as is.
var a64CryptoArr = map[string]int{
"AESD": a64Arr16B, "AESE": a64Arr16B, "AESIMC": a64Arr16B, "AESMC": a64Arr16B,
"SHA1H": a64Arr4S, "SHA1SU1": a64Arr4S, "SHA256SU0": a64Arr4S, "SHA512SU0": a64Arr2D,
"SHA1C": a64Arr4S, "SHA1P": a64Arr4S, "SHA1M": a64Arr4S, "SHA1SU0": a64Arr4S,
"SHA256H": a64Arr4S, "SHA256H2": a64Arr4S, "SHA256SU1": a64Arr4S,
"SHA512H": a64Arr2D, "SHA512H2": a64Arr2D, "SHA512SU1": a64Arr2D,
}
// a64DCOps maps the data-cache maintenance operation names to their fixed
// word (the register rides bits 4:0).
var a64DCOps = map[string]uint32{
"IVAC": 0xd5087620, "ZVA": 0xd50b7420,
"CVAC": 0xd50b7a20, "CVAU": 0xd50b7b20, "CIVAC": 0xd50b7e20,
}
// a64MRSOps maps the system register names GOROOT reads to their fixed word
// (the destination register rides bits 4:0).
var a64MRSOps = map[string]uint32{
"ELR_EL1": 0xd5384020, "MIDR_EL1": 0xd5380000,
"ID_AA64PFR0_EL1": 0xd5380400, "ID_AA64ISAR0_EL1": 0xd5380600,
"ID_AA64ISAR1_EL1": 0xd5380620, "CNTFRQ_EL0": 0xd53be000,
"CNTPCT_EL0": 0xd53be020, "CNTVCT_EL0": 0xd53be040,
"DCZID_EL0": 0xd53b00e0, "DIT": 0xd53b42a0, "ID_AA64ZFR0_EL1": 0xd5380480,
"NZCV": 0xd53b4200, "FPCR": 0xd53b4400, "FPSR": 0xd53b4420,
}
// a64MSRRegOps maps the system register names GOROOT writes through the
// MSR (register) form, spelled in Go assembly as MOVD Rn, <sysreg> or
// MSR Rn, <sysreg>; the source register rides bits 4:0.
var a64MSRRegOps = map[string]uint32{
"NZCV": 0xd51b4200, "FPCR": 0xd51b4400, "FPSR": 0xd51b4420,
"ELR_EL1": 0xd5184020,
}
// a64MSROps maps the system register names GOROOT writes to their fixed
// word; the immediate rides CRm at bits 11:8 and Rt is the fixed 11111.
var a64MSROps = map[string]uint32{
"SPSel": 0xd50040a0, "DAIFSet": 0xd50340c0, "DAIFClr": 0xd50340e0, "DIT": 0xd5034040,
}
// a64PRFOps maps the prefetch operation names to their prfop immediate
// (word = 0xf9800000 | Rn<<5 | prfop).
var a64PRFOps = map[string]int{
"PLDL1KEEP": 0x00, "PLDL1STRM": 0x01, "PLDL2KEEP": 0x02, "PLDL2STRM": 0x03,
"PLDL3KEEP": 0x04, "PLDL3STRM": 0x05,
"PLIL1KEEP": 0x08, "PLIL1STRM": 0x09, "PLIL2KEEP": 0x0a, "PLIL2STRM": 0x0b,
"PLIL3KEEP": 0x0c, "PLIL3STRM": 0x0d,
"PSTL1KEEP": 0x10, "PSTL1STRM": 0x11, "PSTL2KEEP": 0x12, "PSTL2STRM": 0x13,
"PSTL3KEEP": 0x14, "PSTL3STRM": 0x15,
}
// a64VLD1Base holds the fixed words of the multi-register structure
// accesses, indexed by register count 1..4, before the Q and size bits.
// Post-index spellings add 0x9f0000 (post bit and Rm = 11111).
var a64VLD1Base = [5]uint32{0, 0x0c407000, 0x0c40a000, 0x0c406000, 0x0c402000}
var a64VST1Base = [5]uint32{0, 0x0c007000, 0x0c00a000, 0x0c006000, 0x0c002000}
// a64Vec is a parsed vector operand: the register number, the arrangement
// ("" when the operand spells none) and, for element forms, the lane index.
type a64Vec struct {
reg int
arr string
idx int
hasIdx bool
}
// a64VecReg parses a vector register operand: V0..V31 (F0..F31 as an alias,
// the same architectural registers the scalar floating-point spellings use),
// optionally with an arrangement suffix such as V0.B16 and, for element
// forms, a lane index such as V13.S[0]. It reports ok=false for anything
// else, including X/W and R spellings, which the toolchain's vector
// operands reject as well.
func a64VecReg(name string) (v a64Vec, ok bool) {
s := strings.TrimSpace(name)
if i := strings.IndexByte(s, '.'); i >= 0 {
v.arr = strings.TrimSpace(s[i+1:])
s = s[:i]
}
if v.arr != "" {
// Element form: B[3], S[2] and friends.
if j := strings.IndexByte(v.arr, '['); j >= 0 {
k := strings.LastIndexByte(v.arr, ']')
if k < j {
return v, false
}
n, err := strconv.Atoi(strings.TrimSpace(v.arr[j+1 : k]))
if err != nil || n < 0 {
return v, false
}
v.idx, v.hasIdx = n, true
v.arr = strings.TrimSpace(v.arr[:j])
}
if a64ArrIndex(v.arr) < 0 && !a64ElemLetter(v.arr) {
return v, false
}
}
if len(s) < 2 || (s[0] != 'V' && s[0] != 'F') {
return v, false
}
n := 0
for i := 1; i < len(s); i++ {
if s[i] < '0' || s[i] > '9' {
return v, false
}
n = n*10 + int(s[i]-'0')
}
if n > 31 {
return v, false
}
v.reg = n
return v, true
}
// a64ElemField encodes a lane index for the copy/insert group: imm5 = the
// index shifted by the element scale, with the scale's own bit set. B gets
// shift 1 (the Q bit rides elsewhere), H shift 2, S shift 3 and D shift 4.
func a64ElemField(arr string, idx int) (uint32, bool) {
var shift, low uint32
switch arr {
case "B8", "B16", "B":
shift, low = 1, 1
case "H4", "H8", "H":
shift, low = 2, 2
case "S2", "S4", "S":
shift, low = 3, 4
case "D1", "D2", "D":
shift, low = 4, 8
default:
return 0, false
}
if idx < 0 || idx >= 1<<(5-shift) {
return 0, false
}
return uint32(idx)<<shift | low, true
}
// a64VecListOf recovers the register list of a VLD1/VST1/VTBL operand run.
// The parser keeps parenthesised groups whole but splits bracketed lists on
// the commas, so a list arrives as one operand run whose first Raw starts
// with "[" and whose last Raw ends with "]". It returns the parsed
// registers with the brackets and spaces removed.
func a64VecListOf(ops []*ast.Operand, start int) (vs []a64Vec, end int, ok bool) {
if start >= len(ops) || !strings.HasPrefix(strings.TrimSpace(ops[start].Raw), "[") {
return nil, 0, false
}
end = start
for end < len(ops) {
if strings.HasSuffix(strings.TrimSpace(ops[end].Raw), "]") {
break
}
end++
}
if end >= len(ops) {
return nil, 0, false
}
for i := start; i <= end; i++ {
s := strings.TrimSpace(ops[i].Raw)
s = strings.TrimPrefix(s, "[")
s = strings.TrimSuffix(s, "]")
if s == "" && len(ops) > start+1 {
return nil, 0, false
}
for part := range strings.SplitSeq(s, ",") {
v, ok := a64VecReg(part)
if !ok {
return nil, 0, false
}
vs = append(vs, v)
}
}
return vs, end, true
}
// ---- load/store helper tables ----
+737 -18
View File
@@ -4,6 +4,7 @@
package asm
import (
"strings"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
@@ -104,6 +105,7 @@ func TestArm64RegNum(t *testing.T) {
}{
{"R0", 0}, {"R4", 4}, {"R29", 29}, {"R30", 30}, {"R31", 31},
{"FP", 29}, {"LR", 30}, {"LINK", 30}, {"SP", 31}, {"ZR", 31},
{"R18_PLATFORM", 18},
{"F0", 0}, {"F4", 4}, {"F31", 31},
{"INVALID", -1}, {"X0", -1}, {"", -1},
}
@@ -472,12 +474,579 @@ TEXT ·f(SB), NOSPLIT, $0-0
}
}
// TestArm64SIMD tests SIMD encoding (via the instruction table).
// TestArm64SIMD tests SIMD encoding (via the arrangement-aware table).
func TestArm64SIMD(t *testing.T) {
// Verify SIMD instructions are in the table.
for _, mnem := range []string{"VADD", "VSUB", "VMUL"} {
if _, ok := a64InstrTable[mnem]; !ok {
t.Errorf("%s not in instruction table", mnem)
// Verify SIMD instructions are in the arrangement table.
for _, mnem := range []string{"VADD", "VSUB", "VMUL", "VAND", "VEOR", "VORR", "VCMEQ", "VZIP1", "VZIP2"} {
if _, ok := a64SimdVTable[mnem]; !ok {
t.Errorf("%s not in the SIMD arrangement table", mnem)
}
}
}
// TestArm64CarryAndBitOps pins the carry-setting arithmetic, the widening
// multiplies and the data-processing (1 source) group against go tool asm.
func TestArm64CarryAndBitOps(t *testing.T) {
got := arm64Words(t, "\tADC R0, R2, R12\n\tADCS $0, R1\n\tSBCS R5, R9, R5\n\tSBC R25, R10, R26\n"+
"\tMUL R4, R3, R0\n\tUMULH R24, R20, R24\n\tSMULH R1, R2, R3\n\tMSUB R19, R16, R26, R2\n"+
"\tRBIT R11, R4\n\tREV R1, R2\n\tCLZ R21, R9\n\tREVW R1, R2\n\tCLSW R1, R2\n")
want := []uint32{
0x9a00004c, // ADC R12, R2, R0
0xba1f0021, // ADCS R1, R1, ZR
0xfa050125, // SBCS R5, R9, R5
0xda19015a, // SBC R26, R10, R25
0x9b047c60, // MUL R0, R3, R4
0x9bd87e98, // UMULH R24, R20, R24
0x9b417c43, // SMULH R3, R2, R1
0x9b13c342, // MSUB R2, R26, R19, R16
0xdac00164, // RBIT R4, R11
0xdac00c22, // REV R2, R1
0xdac012a9, // CLZ R9, R21
0x5ac00822, // REVW R2, R1
0x5ac01422, // CLSW R2, R1
0xd65f03c0, // RET
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
}
}
}
// TestArm64BitfieldExtract pins UBFX/SBFX: immr wraps to the register
// width, an out-of-range imms is an error.
func TestArm64BitfieldExtract(t *testing.T) {
got := arm64Words(t, "\tUBFX $33, R17, $25, R5\n\tUBFXW $4, R1, $9, R2\n")
want := []uint32{
0xd361e625, // UBFX immr=1 (33 wrapped), imms=25
0x53043022, // UBFXW immr=4, imms=9
0xd65f03c0,
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
}
}
for _, body := range []string{"\tUBFX $33, R17, $70, R5\n", "\tUBFX $-1, R17, $3, R5\n"} {
f, errs := parser.Parse("test_arm64.s", "#include \"textflag.h\"\n\nTEXT ·f(SB), NOSPLIT, $0-0\n"+body+"\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
if _, err := AssembleFileARM64(f); err == nil {
t.Errorf("%s: expected an error, got none", body)
}
}
}
// TestArm64CondCompare pins CCMP/CCMN.
func TestArm64CondCompare(t *testing.T) {
got := arm64Words(t, "\tCCMP LE, R7, $19, $3\n\tCCMP LT, R30, R6, $7\n\tCCMN EQ, R1, R2, $3\n\tCCMPW LE, R7, $19, $3\n")
want := []uint32{
0xfa53d8e3, // CCMP imm form
0xfa46b3c7, // CCMP register form
0xba420023, // CCMN register form
0x7a53d8e3, // CCMPW
0xd65f03c0,
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
}
}
}
// TestArm64CompareBranch pins CBZ/CBNZ/TBZ/TBNZ against a label five and
// six words ahead, matching go tool asm's own offsets.
func TestArm64CompareBranch(t *testing.T) {
// Layout: CBZ(0) TBZ(4) TBNZ(8) CBNZ(12) NOP(16) NOP(17th word...) done.
src := "#include \"textflag.h\"\n\nTEXT ·f(SB), NOSPLIT, $0-0\n" +
"\tCBZ R1, done\n\tTBZ $4, R7, done\n\tTBNZ $33, R7, done\n\tCBNZW R2, done\n" +
"\tNOP\n\tNOP\n\tdone:\tNOP\n\tRET\n"
f, errs := parser.Parse("test_arm64.s", src)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileARM64(f)
if err != nil {
t.Fatalf("AssembleFileARM64: %v", err)
}
got := leWords(img.Code)
// done sits at word 6 from each branch's own pc: CBZ rel 6, TBZ rel 5,
// TBNZ rel 4, CBNZW rel 3.
want := []uint32{
0xb40000c1, // CBZ R1, +6
0x362000a7, // TBZ $4, R7, +5
0xb7080087, // TBNZ $33, R7, +4
0x35000062, // CBNZW R2, +3
0xd503201f, 0xd503201f, 0xd503201f,
0xd65f03c0,
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
}
}
}
// TestArm64ADR pins ADR against a forward label.
func TestArm64ADR(t *testing.T) {
src := "#include \"textflag.h\"\n\nTEXT ·f(SB), NOSPLIT, $0-0\n" +
"\tADR done, R10\n\tNOP\n\tNOP\n\tdone:\tNOP\n\tRET\n"
f, errs := parser.Parse("test_arm64.s", src)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileARM64(f)
if err != nil {
t.Fatalf("AssembleFileARM64: %v", err)
}
got := leWords(img.Code)
// rel = 12 bytes: immlo 0, immhi 3.
want := []uint32{0x1000006a, 0xd503201f, 0xd503201f, 0xd503201f, 0xd65f03c0}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
}
}
}
// TestArm64PairLoadStore pins LDP/STP/LDPW/FLDPD/FSTPD.
func TestArm64PairLoadStore(t *testing.T) {
got := arm64Words(t, "\tSTP (R2, R3), 8(R5)\n\tLDP -8(R5), (R2, R3)\n\tLDPW 4(R0), (R1, R2)\n\tSTPW (R1, R2), 4(R0)\n"+
"\tFLDPD 8(R0), (F1, F2)\n\tFSTPD (F3, F4), -8(R5)\n")
want := []uint32{
0xa9008ca2, // STP (R2, R3), 8(R5)
0xa97f8ca2, // LDP -8(R5), (R2, R3)
0x29408801, // LDPW 4(R0), (R1, R2)
0x29008801, // STPW (R1, R2), 4(R0)
0x6d408801, // FLDPD 8(R0), (F1, F2)
0x6d3f90a3, // FSTPD (F3, F4), -8(R5)
0xd65f03c0,
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
}
}
}
// TestArm64AcquireRelease pins LDAR/STLR and the acquire/release LSE
// families.
func TestArm64AcquireRelease(t *testing.T) {
got := arm64Words(t, "\tLDAR (R27), R22\n\tLDARB (R25), R2\n\tLDARW (R12), R29\n\tSTLR R3, (R24)\n\tSTLRB R11, (R22)\n"+
"\tCASALD R5, (R6), R7\n\tLDADDALD R5, (R6), R7\n\tLDCLRALB R5, (R6), R7\n\tLDORALD R5, (RSP), R7\n\tSWPALW R5, (R6), R7\n")
want := []uint32{
0xc8dfff76, // LDAR R22, (R27)
0x08dfff22, // LDARB R2, (R25)
0x88dffd9d, // LDARW R29, (R12)
0xc89fff03, // STLR R3, (R24)
0x089ffecb, // STLRB R11, (R22)
0xc8e5fcc7, // CASALD R7, (R6), R5
0xf8e500c7, // LDADDALD R7, (R6), R5
0x38e510c7, // LDCLRALB R7, (R6), R5
0xf8e533e7, // LDORALD R7, (RSP), R5
0xb8e580c7, // SWPALW R7, (R6), R5
0xd65f03c0,
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
}
}
}
// TestArm64BTI pins the landing-pad family against the toolchain words:
// only the uppercase C/J/JC spellings assemble, and bare BTI is a
// diagnostic, never a panic.
func TestArm64BTI(t *testing.T) {
got := arm64Words(t, "\tBTI C\n\tBTI J\n\tBTI JC\n")
want := []uint32{
0xd503245f, // BTI C
0xd503249f, // BTI J
0xd50324df, // BTI JC
0xd65f03c0, // RET
}
if len(got) != len(want) {
t.Fatalf("got %d words, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %#x, want %#x", i, got[i], want[i])
}
}
for _, src := range []string{"\tBTI\n", "\tBTI c\n", "\tBTI B\n"} {
f, errs := parser.Parse("test_arm64.s", "#include \"textflag.h\"\n\nTEXT ·f(SB), NOSPLIT, $0-0\n"+src+"\tRET\n")
if len(errs) > 0 {
continue
}
if _, err := AssembleFileARM64(f); err == nil {
t.Errorf("BTI spelling %q should be rejected, as go tool asm rejects it", src)
}
}
}
// TestArm64System pins BRK, SVC, the barriers, cache maintenance and the
// system register accesses.
func TestArm64System(t *testing.T) {
got := arm64Words(t, "\tBRK $35943\n\tBRK\n\tSVC $7165\n\tDMB $1\n\tDSB $1\n\tISB $15\n"+
"\tDC ZVA, R4\n\tDC IVAC, R1\n\tMRS DCZID_EL0, R3\n\tMRS CNTVCT_EL0, R0\n\tMSR $9, DAIFSet\n\tMSR $3, SPSel\n"+
"\tPRFM (R0), PLDL1KEEP\n\tPRFM (R3), PLDL3KEEP\n\tPRFM (R2), $25\n")
want := []uint32{
0xd4318ce0, // BRK $35943
0xd4200000, // BRK
0xd4037fa1, // SVC $7165
0xd50331bf, // DMB $1
0xd503319f, // DSB $1
0xd5033fdf, // ISB $15
0xd50b7424, // DC ZVA, R4
0xd5087621, // DC IVAC, R1
0xd53b00e3, // MRS DCZID_EL0, R3
0xd53be040, // MRS CNTVCT_EL0, R0
0xd50349df, // MSR $9, DAIFSet
0xd50043bf, // MSR $3, SPSel
0xf9800000, // PRFM (R0), PLDL1KEEP
0xf9800064, // PRFM (R3), PLDL3KEEP
0xf9800059, // PRFM (R2), $25
0xd65f03c0,
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
}
}
}
// TestArm64Crypto pins the AES and SHA families.
func TestArm64Crypto(t *testing.T) {
got := arm64Words(t, "\tAESE V31.B16, V29.B16\n\tAESD V22.B16, V19.B16\n\tAESIMC V12.B16, V27.B16\n\tAESMC V14.B16, V28.B16\n"+
"\tSHA1C V8.S4, V8, V2\n\tSHA1H V17, V25\n\tSHA1P V3.S4, V20, V27\n\tSHA1SU0 V17.S4, V13.S4, V16.S4\n\tSHA1SU1 V24.S4, V23.S4\n"+
"\tSHA256H V4.S4, V2, V11\n\tSHA256H2 V6.S4, V16, V11\n\tSHA256SU0 V0.S4, V16.S4\n\tSHA256SU1 V31.S4, V3.S4, V15.S4\n"+
"\tSHA512H V2.D2, V1, V0\n\tSHA512H2 V4.D2, V3, V2\n\tSHA512SU0 V9.D2, V8.D2\n\tSHA512SU1 V7.D2, V6.D2, V5.D2\n")
want := []uint32{
0x4e284bfd, // AESE
0x4e285ad3, // AESD
0x4e28799b, // AESIMC
0x4e2869dc, // AESMC
0x5e080102, // SHA1C
0x5e280a39, // SHA1H
0x5e03129b, // SHA1P
0x5e1131b0, // SHA1SU0
0x5e281b17, // SHA1SU1
0x5e04404b, // SHA256H
0x5e06520b, // SHA256H2
0x5e282810, // SHA256SU0
0x5e1f606f, // SHA256SU1
0xce628020, // SHA512H
0xce648462, // SHA512H2
0xcec08128, // SHA512SU0
0xce6788c5, // SHA512SU1
0xd65f03c0,
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
}
}
}
// TestArm64SIMDLogical pins the arrangement-aware three- and two-register
// SIMD paths.
func TestArm64SIMDLogical(t *testing.T) {
got := arm64Words(t, "\tVADD V1.B16, V2.B16, V3.B16\n\tVAND V4.B16, V4.B16, V9.B16\n\tVEOR V0.B16, V1.B16, V0.B16\n"+
"\tVORR V5.B16, V4.B16, V3.B16\n\tVADDP V1.H8, V2.H8, V3.H8\n\tVZIP1 V16.H8, V3.H8, V19.H8\n\tVZIP2 V22.D2, V25.D2, V21.D2\n"+
"\tVCMEQ V24.S4, V13.S4, V12.S4\n\tVCMEQ $0, V2.H4, V3.H4\n\tVREV32 V2.H8, V1.H8\n\tVREV64 V2.S4, V3.S4\n\tVUADDLV V31.S4, V11\n"+
"\tVPMULL V2.D1, V1.D1, V3.Q1\n\tVPMULL2 V2.B16, V1.B16, V4.H8\n\tVRAX1 V26.D2, V29.D2, V30.D2\n\tVMOV V2.B16, V4.B16\n")
want := []uint32{
0x4e218443, // VADD 16B
0x4e241c89, // VAND
0x6e201c20, // VEOR
0x4ea51c83, // VORR
0x4e61bc43, // VADDP 8H
0x4e503873, // VZIP1 8H
0x4ed67b35, // VZIP2 2D
0x6eb88dac, // VCMEQ 4S
0x0e609843, // VCMEQ $0, 4H
0x6e600841, // VREV32 8H
0x4ea00843, // VREV64 4S
0x6eb03beb, // VUADDLV 4S
0x0ee2e023, // VPMULL D1
0x4e22e024, // VPMULL2 16B
0xce7a8fbe, // VRAX1 2D
0x4ea21c44, // VMOV 16B pair
0xd65f03c0,
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
}
}
}
// TestArm64SIMDWide pins the four-register crypto group, VXAR, VEXT and the
// shift-by-immediate encodings.
func TestArm64SIMDWide(t *testing.T) {
got := arm64Words(t, "\tVEOR3 V2.B16, V7.B16, V12.B16, V25.B16\n\tVBCAX V1.B16, V2.B16, V26.B16, V31.B16\n"+
"\tVXAR $63, V27.D2, V21.D2, V26.D2\n\tVEXT $4, V2.B8, V1.B8, V3.B8\n\tVEXT $8, V2.B16, V1.B16, V3.B16\n"+
"\tVSHL $7, V22.D2, V25.D2\n\tVUSHR $6, V22.H8, V23.H8\n\tVSRI $24, V1.S4, V2.S4\n")
want := []uint32{
0xce070999, // VEOR3
0xce22075f, // VBCAX
0xce9bfeba, // VXAR
0x2e022023, // VEXT B8
0x6e024023, // VEXT B16
0x4f4756d9, // VSHL D2 $7
0x6f1a06d7, // VUSHR H8 $6
0x6f284422, // VSRI S4 $24
0xd65f03c0,
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
}
}
}
// TestArm64SIMDElement pins VDUP and the VMOV element forms.
func TestArm64SIMDElement(t *testing.T) {
got := arm64Words(t, "\tVDUP V31.B[15], V18\n\tVDUP V19.S[3], V18.S4\n\tVDUP V1.D[1], V2.D2\n"+
"\tVMOV V13.S[0], R20\n\tVMOV V11.B[11], V16.B[12]\n\tVMOV R20, V21.B[2]\n")
want := []uint32{
0x5e1f07f2, // VDUP element to register
0x4e1c0672, // VDUP element across S4
0x4e180422, // VDUP element across D2
0x0e043db4, // VMOV element to register
0x6e195d70, // VMOV element to element
0x4e051e95, // VMOV register into element
0xd65f03c0,
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
}
}
}
// TestArm64GPIntoVector pins the whole-vector moves VMOV/VDUP Rs, Vd.<T>
// against `go tool asm -S` output (Go 1.27, arm64): word = Q | 7<<25 |
// imm5<<16 | 3<<10 | rs<<5 | rd, shared by both mnemonics, the form
// sys_windows_arm64.s and the bytealg loops use. The D1 destination is
// rejected, as the toolchain rejects it.
func TestArm64GPIntoVector(t *testing.T) {
got := arm64Words(t, "\tVMOV R5, V5.B16\n\tVMOV R1, V2.B8\n\tVMOV R3, V4.H4\n"+
"\tVMOV R9, V10.S4\n\tVMOV R7, V31.H8\n\tVMOV R11, V12.D2\n"+
"\tVDUP R5, V5.B16\n\tVDUP R9, V10.H8\n\tVMOV V4.B16, V20.B16\n")
want := []uint32{
0x4e010ca5, // VMOV R5, V5.B16
0x0e010c22, // VMOV R1, V2.B8
0x0e020c64, // VMOV R3, V4.H4
0x4e040d2a, // VMOV R9, V10.S4
0x4e020cff, // VMOV R7, V31.H8
0x4e080d6c, // VMOV R11, V12.D2
0x4e010ca5, // VDUP R5, V5.B16 (same word as VMOV)
0x4e020d2a, // VDUP R9, V10.H8
0x4ea41c94, // VMOV V4.B16, V20.B16 (vector to vector stays ORR)
0xd65f03c0,
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
}
}
f, errs := parser.Parse("test_arm64.s", "#include \"textflag.h\"\n\nTEXT ·f(SB), NOSPLIT, $0-0\n\tVMOV R7, V8.D1\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
if _, err := AssembleFileARM64(f); err == nil {
t.Errorf("VMOV R7, V8.D1 assembled, want an arrangement error")
}
}
// TestArm64SimdTwoOperand pins the two-operand accumulate spellings
// VADD/VSUB Vm, Vn against `go tool asm -S` output (Go 1.27, arm64):
// word = 5<<28|7<<25|7<<21|1<<15|1<<10 for VADD (7<<28 for VSUB) with
// rf<<16 | rn<<5 | rn, bare V registers only (asm7.go case 89).
func TestArm64SimdTwoOperand(t *testing.T) {
got := arm64Words(t, "\tVADD V7, V8\n\tVSUB V7, V8\n\tVADD V1, V2\n\tVADD V0.B16, V1.B16, V2.B16\n")
want := []uint32{
0x5ee78508, // VADD V7, V8
0x7ee78508, // VSUB V7, V8
0x5ee18442, // VADD V1, V2
0x4e208422, // VADD arranged: the ordinary three-register path
0xd65f03c0,
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
}
}
}
// TestArm64TruncMove pins the truncating register moves against
// `go tool asm -S` output (Go 1.27, arm64): the signed forms lower to SXTB,
// SXTH and SXTW (SBFM), the unsigned byte and halfword forms to UXTB and
// UXTH (UBFM), MOVWU to a W ORR, and a narrow move out of the zero register
// drops to the W ORR too (asm7.go case 45).
func TestArm64TruncMove(t *testing.T) {
got := arm64Words(t, "\tMOVB R3, R4\n\tMOVH R5, R6\n\tMOVW R9, R10\n"+
"\tMOVBU R3, R4\n\tMOVHU R3, R4\n\tMOVWU R3, R4\n\tMOVD R3, R4\n"+
"\tMOVD ZR, R4\n\tMOVB ZR, R4\n\tMOVWU ZR, R5\n")
want := []uint32{
0x93401c64, // MOVB = SXTB
0x93403ca6, // MOVH = SXTH
0x93407d2a, // MOVW = SXTW
0xd3401c64, // MOVBU = UXTB
0xd3403c64, // MOVHU = UXTH
0x2a0303e4, // MOVWU = ORR W
0xaa0303e4, // MOVD = ORR X
0xaa1f03e4, // MOVD ZR, R4 keeps the X form
0x2a1f03e4, // MOVB ZR, R4 drops to the W form
0x2a1f03e5, // MOVWU ZR, R5
0xd65f03c0,
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
}
}
}
// TestArm64SIMDLoadStore pins the structure loads and stores.
func TestArm64SIMDLoadStore(t *testing.T) {
got := arm64Words(t, "\tVLD1 (R2), [V21.B16]\n\tVLD1 (R1), [V2.B16, V3.B16]\n\tVLD1 (R29), [V14.D1, V15.D1, V16.D1, V17.D1]\n"+
"\tVLD1.P 32(R1), [V2.B16, V3.B16]\n\tVST1 [V2.S4, V3.S4, V4.S4, V5.S4], (R14)\n\tVST1.P [V2.B16], (R1)\n"+
"\tVLD1R (R1), [V9.B8]\n\tVLD4R (R0), [V0.B8, V1.B8, V2.B8, V3.B8]\n")
want := []uint32{
0x4c407055, // VLD1 one register
0x4c40a022, // VLD1 two registers
0x0c402fae, // VLD1 four registers D1
0x4cdfa022, // VLD1.P two registers
0x4c0029c2, // VST1 four registers S4
0x4c9f7022, // VST1.P one register
0x0d40c029, // VLD1R
0x0d60e000, // VLD4R
0xd65f03c0,
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
}
}
}
// TestArm64MoviLiteral pins the VMOVS/VMOVD/VMOVQ constant loads: three
// words each (ADRP, ADD, wide load) plus the pooled literal in the data
// section.
func TestArm64MoviLiteral(t *testing.T) {
src := "#include \"textflag.h\"\n\nTEXT ·f(SB), NOSPLIT, $0-0\n" +
"\tVMOVS $0x80402010, V11\n\tVMOVD $0x8040201008040201, V20\n" +
"\tVMOVQ $0x7040201008040201, $0x8040201008040201, V10\n\tRET\n"
f, errs := parser.Parse("test_arm64.s", src)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileARM64(f)
if err != nil {
t.Fatalf("AssembleFileARM64: %v", err)
}
if img.Funcs[0].Size != 12*3+4 {
t.Errorf("func size = %d, want %d", img.Funcs[0].Size, 12*3+4)
}
want := []uint32{
0x9000001b, 0x9100037b, 0xbd40036b, // VMOVS: ADRP, ADD, LDR S
0x9000001b, 0x9100037b, 0xfd400374, // VMOVD: ADRP, ADD, LDR D
0x9000001b, 0x9100037b, 0x3dc0036a, // VMOVQ: ADRP, ADD, LDR Q
0xd65f03c0,
}
got := leWords(img.Code)
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
}
}
// The literals sit in the data section.
var found32, found64, found128 bool
for _, d := range img.DataSyms {
switch d.Name {
case "$i32.80402010":
found32 = d.Size == 4
case "$i64.8040201008040201":
found64 = d.Size == 8
case "$i128.80402010080402017040201008040201":
found128 = d.Size == 16
}
}
if !found32 || !found64 || !found128 {
t.Errorf("literals missing: i32=%v i64=%v i128=%v", found32, found64, found128)
}
}
// TestArm64MOVK pins standalone MOVK with the hw field derived from the
// chunk position.
func TestArm64MOVK(t *testing.T) {
got := arm64Words(t, "\tMOVK $1234, R5\n\tMOVK $305397760, R5\n\tMOVKW $1234, R5\n")
want := []uint32{
0xf2809a45, // MOVK hw=0
0xf2a24685, // MOVK hw=1
0x72809a45, // MOVKW hw=0
0xd65f03c0,
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
}
}
}
@@ -851,20 +1420,170 @@ func TestArm64ExclNoOffset(t *testing.T) {
}
}
// TestArm64AddSubImmRange: immediates that cannot ride the imm12 field are
// rejected instead of wrapping through int32.
func TestArm64AddSubImmRange(t *testing.T) {
for _, body := range []string{
"\tADD $0x100000000, R0, R1\n",
"\tSUB $-0x100000000, R0, R1\n",
"\tCMP $0x100000000, R0\n",
} {
f, errs := parser.Parse("test_arm64.s", "#include \"textflag.h\"\n\nTEXT ·f(SB), NOSPLIT, $0-0\n"+body+"\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
// TestArm64AddSubImmWide pins the wide-immediate classification the toolchain
// applies to the ADD/SUB family (asm7.go cases 48, 62, 13): the ADDCON2 split
// into two imm12 instructions for plain ADD/SUB, the bitmask ORR into REGTMP,
// and the MOVZ/MOVN/MOVK materialisations followed by the register form.
// Comparisons never split, and the W forms classify the 32-bit value. Every
// word is go tool asm's own for the same source.
func TestArm64AddSubImmWide(t *testing.T) {
got := arm64Words(t, strings.Join([]string{
"\tADD $0xaaaaaa, R2, R3",
"\tSUB $0xaaaaaa, R2",
"\tADD $0x186a0, R2, R5",
"\tADD $0x1ffe00, R2, R3",
"\tADD $0x3fffffffc000, R5",
"\tADD $-100000, R2, R3",
"\tADD $-2048, R2, R3",
"\tCMP $0xaaaaaa, R2",
"\tCMP $0xffffffffffa0, R3",
"\tCMPW $27745, R2",
"\tCMPW $0x60060, R2",
"\tADDS $0xaaaaaa, R2, R3",
"\tADD $0x12345678, R2, R3",
"\tADDW $0x60060, R2",
"\tSUB $0xe7791f700, R3, R1",
"\tADDW $0x12345678, R2, R3",
"\tCMN $0x1000000, R2",
}, "\n")+"\n")
want := []uint32{
0x912aa843, 0x916aa863, // ADD $0xaaaaaa, R2, R3: ADDCON2 split
0xd12aa842, 0xd16aa842, // SUB $0xaaaaaa, R2: split with Rd = Rn
0x911a8045, 0x914060a5, // ADD $0x186a0, R2, R5: split
0xb2772ffb, 0x8b1b0043, // ADD $0x1ffe00: bitmask beats the split
0xb2727ffb, 0x8b1b00a5, // ADD $0x3fffffffc000: bitmask into REGTMP
0x9290d3fb, 0xf2bfffdb, 0x8b1b0043, // ADD $-100000: MOVN + MOVK
0x9280fffb, 0x8b1b0043, // ADD $-2048: single MOVN + ADD
0xd295555b, 0xf2a0155b, 0xeb1b005f, // CMP: never split, MOVZ + MOVK
0x92800bfb, 0xf2e0001b, 0xeb1b007f, // CMP $0xffffffffffa0: MOVN + fixup
0x528d8c3b, 0x6b1b005f, // CMPW $27745: W movcon, single MOVZW
0x52800c1b, 0x72a000db, 0x6b1b005f, // CMPW $0x60060: S form skips the split
0xd295555b, 0xf2a0155b, 0xab1b0043, // ADDS $0xaaaaaa: MOVZ + MOVK + ADDS
0xd28acf1b, 0xf2a2469b, 0x8b1b0043, // ADD $0x12345678: MOVZ + MOVK
0x11018042, 0x11418042, // ADDW $0x60060: W split
0xd29ee01b, 0xf2aef23b, 0xf2c001db, 0xcb1b0061, // SUB $0xe7791f700
0x528acf1b, 0x72a2469b, 0x0b1b0043, // ADDW $0x12345678: MOVZW + MOVKW
0xd2a0201b, 0xab1b005f, // CMN $0x1000000: single MOVZ + CMN
0xd65f03c0, // RET
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("wide word %d = %08x, want %08x", i, got[i], want[i])
}
if _, err := AssembleFileARM64(f); err == nil {
t.Errorf("%s: expected an error, got none", body)
}
}
// TestArm64CarryImmWide pins the carry family's $0 spellings in two and
// three operands, the ROR shift on the logical group (and its rejection for
// the arithmetic forms), the NGC/MNEG zero-register aliases and the vector
// alias with an element selector. Words are go tool asm's own.
func TestArm64CarryShiftAlias(t *testing.T) {
got := arm64Words(t, "\tADC $0, R20\n\tADC $0, R20, R4\n\tSBCS $0, R4, R12\n"+
"\tSBCS R15, R4, R12\n\tANDW R9@>7, R19, R26\n\tAND R1@>33, R2, R3\n"+
"\tNEGSW R23<<1, R30\n\tNGC R2, R7\n\tMNEG R14, R27, R23\n")
want := []uint32{
0x9a1f0294, // ADC ZR, R20, R20
0x9a1f0284, // ADC ZR, R20, R4
0xfa1f008c, // SBCS ZR, R4, R12
0xfa0f008c, // SBCS R15, R4, R12
0x0ac91e7a, // ANDW R9 ROR 7, R19, R26
0x8ac18443, // AND R1 ROR 33, R2, R3
0x6b1707fe, // SUBSW ZR, R30, R23 LSL 1
0xda0203e7, // SBC ZR, R7, R2
0x9b0eff77, // MSUB ZR, R27, R14, R23
0xd65f03c0, // RET
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("carry word %d = %08x, want %08x", i, got[i], want[i])
}
}
// ROR on an arithmetic form is unallocated: the toolchain reports an
// unsupported shift operator.
f, errs := parser.Parse("test_arm64.s", "#include \"textflag.h\"\n\nTEXT ·f(SB), NOSPLIT, $0-0\n\tADD R1@>33, R2, R3\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
if _, err := AssembleFileARM64(f); err == nil {
t.Error("ADD R1@>33: expected an error, got none")
}
}
// TestArm64VecAliasElement pins the register-alias rewrite inside a vector
// operand with an element selector and inside a split register list: the
// aliases resolve textually where the parser carries the selector apart from
// the name. Words are go tool asm's own.
func TestArm64VecAliasElement(t *testing.T) {
src := `#include "textflag.h"
#define POLY V15
#define ACC0 V8
#define ACC1 V9
TEXT ·f(SB), NOSPLIT, $0-0
VMOV R1, POLY.D[0]
VEOR POLY.B16, POLY.B16, POLY.B16
VLD1 (R0), [ACC0.B16]
VLD1.P (R0), [ACC0.B16, ACC1.B16]
VST1.P [ACC0.B16, ACC1.B16], 32(R1)
RET
`
f, errs := parser.Parse("test_arm64.s", src)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileARM64(f)
if err != nil {
t.Fatalf("AssembleFileARM64: %v", err)
}
got := leWords(img.Code)
want := []uint32{
0x4e081c2f, // INS V15.D[0], R1
0x6e2f1def, // VEOR V15.B16, V15.B16, V15.B16
0x4c407008, // VLD1 (R0), [V8.B16]
0x4cdfa008, // VLD1.P (R0), [V8.B16, V9.B16]
0x4c9fa028, // VST1.P [V8.B16, V9.B16], 32(R1)
0xd65f03c0, // RET
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("vecalias word %d = %08x, want %08x", i, got[i], want[i])
}
}
}
// TestArm64AddSubImmBeyond32 pins the materialisation the toolchain applies
// once the value leaves every imm12 form: a constant sequence into REGTMP
// (R27) followed by the register form. SUB $-0x100000000 is a bitmask
// immediate, so it rides the ORR form; the others take MOVZ. Words are go
// tool asm's own.
func TestArm64AddSubImmBeyond32(t *testing.T) {
got := arm64Words(t, "\tADD $0x100000000, R0, R1\n\tSUB $-0x100000000, R0, R1\n\tCMP $0x100000000, R0\n")
want := []uint32{
0xd2c0003b, // MOVZ $(1<<32>>16), R27 (hw=2)
0x8b1b0001, // ADD R27, R0, R1
0xb2607ffb, // ORR $-4294967296, ZR, R27 (bitmask)
0xcb1b0001, // SUB R27, R0, R1
0xd2c0003b, // MOVZ $(1<<32>>16), R27 (hw=2)
0xeb1b001f, // CMP R27, R0
0xd65f03c0, // RET
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
}
}
}
+73
View File
@@ -67,6 +67,9 @@ type spadjStep struct {
// patch sites (for the file-level layout to resolve), the label table and the
// stack-adjustment boundaries.
func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, []spadjStep, []LineEntry, error) {
if err := checkAdjspBalance(t); err != nil {
return nil, nil, nil, nil, nil, err
}
fi := computeFrame(t)
chain := jumpChain(t)
resolve := func(name string) string {
@@ -203,6 +206,14 @@ func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, [
spadjStep{guardLen + len(fi.prologue), 8 + fi.size},
)
}
// frameBase is the SP delta the prologue leaves: 8 for the saved base
// pointer plus the frame, 0 frameless. bodyDelta tracks the ADJSP
// statements' straight-line sum, so a mid-body step's value is the
// frame base plus what the body has opened so far.
frameBase, bodyDelta := 0, 0
if fi.useFP {
frameBase = 8 + fi.size
}
pos := guardLen + len(fi.prologue)
for i, stmt := range t.Body {
s, ok := stmt.(*ast.Instr)
@@ -230,6 +241,16 @@ func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, [
ps[k].kind = RelCall
}
}
if strings.ToUpper(s.Mnemonic.Text) == "ADJSP" && len(s.Operands) == 1 && s.Operands[0].Imm.HasVal {
// The statement shifted SP mid-body: record the new running
// delta as the value in effect from just past the instruction.
v := s.Operands[0].Imm.Val
if s.Operands[0].Imm.Neg {
v = -v
}
bodyDelta += int(v)
steps = append(steps, spadjStep{pos + len(code), frameBase + bodyDelta})
}
patches = append(patches, ps...)
lines = append(lines, LineEntry{Offset: pos, Line: s.Pos().Line})
out = append(out, code...)
@@ -412,6 +433,40 @@ func hasCall(t *ast.Text) bool {
return false
}
// checkAdjspBalance mirrors the toolchain's push/pop walk: every ADJSP
// shifts SP away from the entry state and every RET must see the shifts
// closed. The assembler's own prologue and epilogue contribute matching
// deltas on both sides, so the statements' straight-line sum must be zero
// at each RET; branches do not reset the walk, which runs over the program
// list in source order. go tool asm reports an offender as "unbalanced
// PUSH/POP" (verified against ADJSP $16 before a RET, accepted as a
// $16/$-16 pair, per-RET rather than per-function).
func checkAdjspBalance(t *ast.Text) error {
delta := 0
for _, stmt := range t.Body {
in, ok := stmt.(*ast.Instr)
if !ok {
continue
}
switch strings.ToUpper(in.Mnemonic.Text) {
case "ADJSP":
if len(in.Operands) != 1 || !in.Operands[0].Imm.HasVal {
continue // reported during emission
}
v := in.Operands[0].Imm.Val
if in.Operands[0].Imm.Neg {
v = -v
}
delta += int(v)
case "RET":
if delta != 0 {
return fmt.Errorf("unbalanced PUSH/POP")
}
}
}
return nil
}
// guardLen returns the byte length of the stack-split guard prefix. The
// final conditional branch (JBE, and JB in the big class) is 2 bytes in the
// short form and 6 in the long form.
@@ -791,6 +846,14 @@ func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Op
case ast.OpAddr:
a := op.Addr
// A bracketed register range, [Z0-Z3]: the four-register source of
// the 4FMAPS/4VNNIW families. The EVEX quad-register emit path
// needs an encoder operand of its own, so the shape stays a named
// gap rather than an encoding.
if a.Range != nil {
return nil, fmt.Errorf("register range %q needs quad-register encoder support", op.Raw)
}
// FP-relative: x+N(FP) → (N + fpAdjust)(SP). The offset N lives in the
// symbol, not the address displacement.
if a.Sym != nil && a.Sym.Pseudo == "FP" {
@@ -838,6 +901,16 @@ func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Op
}
return m, nil
}
// Index-only memory: the VSIB form the gather/scatter families
// read, 8(X4*1). A scaled vector index addresses memory with no
// base register; the mod=00 SIB with base field 101 carries it.
if a.Index != "" {
idx, ok := ParseReg(a.Index)
if !ok {
return nil, fmt.Errorf("unknown index register %q", a.Index)
}
return Mem{Index: idx, Scale: a.Scale, Disp: a.Offset, HasIndex: true, Size: size}, nil
}
// Bare register.
if a.Sym != nil && a.Sym.Pseudo == "" && a.Sym.Name != "" {
if r, ok := ParseReg(a.Sym.Name); ok {
+129
View File
@@ -439,3 +439,132 @@ func TestSubSPEncodings(t *testing.T) {
}
}
}
// TestAssemblePseudoStatements runs LOCK/REP, BYTE/WORD and END through the
// full statement pipeline, pinned against go tool asm (Go 1.27, amd64). It
// asserts the three behaviours the toolchain shows: each prefix statement is
// a standalone byte with a PC of its own (so a label placed on the LOCK
// points at the F0), the data pseudo-ops write their literal bytes inline,
// and END terminates nothing (the statements after it still belong to the
// function and carry no trace of it).
func TestAssemblePseudoStatements(t *testing.T) {
fn := firstText(t, `
#include "textflag.h"
TEXT ·pseudo(SB), NOSPLIT, $0-0
pfx:
LOCK
CMPXCHGQ AX, (BX)
REP
MOVSQ
BYTE $0x0f
BYTE $0x1f
WORD $0x1234
END
BYTE $0x02
RET
`)
code, labels, err := Assemble(fn)
if err != nil {
t.Fatalf("Assemble: %v", err)
}
// go tool asm: f0 480fb103 f3 48a5 0f 1f 3412 02 c3
want := []byte{
0xf0,
0x48, 0x0f, 0xb1, 0x03,
0xf3, 0x48, 0xa5,
0x0f, 0x1f, 0x34, 0x12,
0x02, 0xc3,
}
if hexBytes(code) != hexBytes(want) {
t.Errorf("pseudo statements:\n got: %s\n want: %s", hexBytes(code), hexBytes(want))
}
// The label sits on the LOCK byte, exactly where the toolchain's PC
// listing puts it.
if off := labels["pfx"]; off != 0 {
t.Errorf("label pfx = %d, want 0 (the LOCK's own byte)", off)
}
// The trailing BYTE lands where the layout says: after the 8 bytes of
// LOCK, CMPXCHGQ, REP and MOVSQ plus the 4 data bytes, END contributing
// none.
if code[12] != 0x02 {
t.Errorf("byte at 12 = %02x, want 02 (the BYTE after END)", code[12])
}
}
// TestAssembleAdjspBalance pins the toolchain's push/pop balance rule over
// ADJSP: the straight-line sum of the adjustments must be zero at each
// RET, branches in between counting for nothing (verified against go tool
// asm: ADJSP $16 before a RET is reported as "unbalanced PUSH/POP", a
// $16/$-16 pair with a JMP in between assembles).
func TestAssembleAdjspBalance(t *testing.T) {
// Balanced pair with a branch in between, bytes pinned from go tool asm.
fn := firstText(t, `
#include "textflag.h"
TEXT ·adjsp(SB), NOSPLIT, $0-0
ADJSP $16
JMP body
body:
ADJSP $-16
RET
`)
code, _, err := Assemble(fn)
if err != nil {
t.Fatalf("Assemble: %v", err)
}
want := []byte{0x48, 0x83, 0xEC, 0x10, 0xEB, 0x00, 0x48, 0x83, 0xC4, 0x10, 0xC3}
if hexBytes(code) != hexBytes(want) {
t.Errorf("adjsp pair:\n got: %s\n want: %s", hexBytes(code), hexBytes(want))
}
// Unbalanced at the RET: the toolchain diagnoses, so must we.
_, _, err = Assemble(firstText(t, `
#include "textflag.h"
TEXT ·unbalanced(SB), NOSPLIT, $0-0
ADJSP $16
RET
`))
if err == nil || !strings.Contains(err.Error(), "unbalanced PUSH/POP") {
t.Errorf("unbalanced ADJSP: err = %v, want unbalanced PUSH/POP", err)
}
// The check runs per RET: a closed pair before the first RET does not
// excuse an open adjustment before the second.
_, _, err = Assemble(firstText(t, `
#include "textflag.h"
TEXT ·tworet(SB), NOSPLIT, $0-0
ADJSP $8
ADJSP $-8
RET
mid:
ADJSP $8
RET
`))
if err == nil || !strings.Contains(err.Error(), "unbalanced PUSH/POP") {
t.Errorf("second RET with open ADJSP: err = %v, want unbalanced PUSH/POP", err)
}
// A framed function: the assembler's own prologue and epilogue
// contribute matching deltas, so the pair in the body still balances,
// and the bytes match go tool asm end to end.
fn = firstText(t, `
#include "textflag.h"
TEXT ·framed(SB), $16-8
ADJSP $8
ADJSP $-8
RET
`)
code, _, err = Assemble(fn)
if err != nil {
t.Fatalf("Assemble framed: %v", err)
}
want = []byte{
0x55, 0x48, 0x89, 0xE5, 0x48, 0x83, 0xEC, 0x10, // prologue
0x48, 0x83, 0xEC, 0x08, // ADJSP $8
0x48, 0x83, 0xC4, 0x08, // ADJSP $-8
0x48, 0x83, 0xC4, 0x10, 0x5D, // epilogue
0xC3,
}
if hexBytes(code) != hexBytes(want) {
t.Errorf("framed adjsp:\n got: %s\n want: %s", hexBytes(code), hexBytes(want))
}
}
+65 -10
View File
@@ -43,6 +43,9 @@ const (
stInfoShift = 4
rX8664PC32 = 2
// R_X86_64_32 (debug/elf): the absolute 32-bit address of a symbol, the
// R_ADDR shape a 4-byte DATA field carries.
rX8664Abs32 = 10
// R_X86_64_TPOFF32 (debug/elf): the local-exec TLS offset the stack
// guard loads from FS. 20 is R_X86_64_TLSLD, a different relocation.
rX8664TPOFF32 = 23
@@ -158,6 +161,50 @@ func (img *Image) ELFObject() ([]byte, error) {
}
}
// The data symbols' symbol-valued DATA fields ("DATA s+0(SB)/8,
// $other(SB)") become .rela.data entries: an absolute relocation of the
// DATA line's width at the field's data-section offset, S + A with no
// PC term. Widths 4 and 8 have ELF relocation shapes; narrower fields
// cannot hold an address, so they are refused rather than truncated.
var dataRelas []elfRela
for _, d := range img.DataSyms {
for _, r := range d.Relocs {
idx, ok := symIdx[r.Name]
if !ok {
return nil, fmt.Errorf("data relocation references unknown symbol %q", r.Name)
}
var typ uint32
switch r.Siz {
case 8:
typ = rX8664Abs64
case 4:
typ = rX8664Abs32
default:
return nil, fmt.Errorf("DATA %q: a symbol value of width %d has no ELF relocation", d.Name, r.Siz)
}
dataRelas = append(dataRelas, elfRela{
off: uint64(d.Offset + r.Off),
sym: idx,
typ: typ,
addend: r.Addend,
})
}
}
// Section presence: .rela.text only when there are code relocations,
// .rela.data only when a DATA line holds a symbol value.
hasRela := len(relas) > 0
hasDataRela := len(dataRelas) > 0
nSections := 6 // NULL, .text, .data, .symtab, .strtab, .shstrtab
if hasRela {
nSections++
}
if hasDataRela {
nSections++
}
secSymtab, secStrtab := 3, 4
secShstr := nSections - 1
// Serialise the string tables.
stNames := newElfStrtab()
for _, s := range syms {
@@ -167,19 +214,13 @@ func (img *Image) ELFObject() ([]byte, error) {
for _, n := range []string{".text", ".data", ".symtab", ".strtab", ".rela.text", ".shstrtab"} {
stSections.add(n)
}
if hasDataRela {
stSections.add(".rela.data")
}
for _, n := range dwarfSectionNames {
stSections.add(n)
}
// Section presence: .rela.text only when there are relocations.
hasRela := len(relas) > 0
nSections := 6 // NULL, .text, .data, .symtab, .strtab, .shstrtab
if hasRela {
nSections = 7
}
secSymtab, secStrtab := 3, 4
secShstr := nSections - 1
// Lay the file out: header, section data, section headers.
var out []byte
out = append(out, make([]byte, 64)...) // ELF header, filled last
@@ -214,7 +255,7 @@ func (img *Image) ELFObject() ([]byte, error) {
strtabOff := len(out)
out = append(out, stNames.bytes()...)
var relaOff int
var relaOff, relaDataOff int
if hasRela {
align(8)
relaOff = len(out)
@@ -226,6 +267,17 @@ func (img *Image) ELFObject() ([]byte, error) {
out = append(out, b[:]...)
}
}
if hasDataRela {
align(8)
relaDataOff = len(out)
for _, r := range dataRelas {
var b [24]byte
le.PutUint64(b[0:], r.off)
le.PutUint64(b[8:], uint64(r.sym)<<32|uint64(r.typ))
le.PutUint64(b[16:], uint64(r.addend))
out = append(out, b[:]...)
}
}
shstrOff := len(out)
out = append(out, stSections.bytes()...)
@@ -284,6 +336,9 @@ func (img *Image) ELFObject() ([]byte, error) {
if hasRela {
putSh(".rela.text", shtRela, 0, relaOff, 24*len(relas), secSymtab, secText, 8, 24)
}
if hasDataRela {
putSh(".rela.data", shtRela, 0, relaDataOff, 24*len(dataRelas), secSymtab, secData, 8, 24)
}
putSh(".shstrtab", shtStrtab, 0, shstrOff, len(stSections.bytes()), 0, 0, 1, 0)
// DWARF section headers; their indices follow the write order.
+108
View File
@@ -750,3 +750,111 @@ func readFormSkip(t *testing.T, r *ulebIter, form uint64) {
t.Fatalf("unsupported form %#x", form)
}
}
// TestELFObjectDataRelocation checks that a symbol-valued DATA field ("DATA
// s+0(SB)/8, $other(SB)") reaches the ELF object as a .rela.data entry: an
// absolute 64-bit relocation at the field's offset within .data, against
// the named symbol, external targets included.
func TestELFObjectDataRelocation(t *testing.T) {
f, errs := parser.Parse("t_amd64.s", `#include "textflag.h"
TEXT ·Keep(SB), NOSPLIT, $0-8
RET
GLOBL holder(SB), NOPTR, $24
DATA holder+0(SB)/8, $·Keep+5(SB)
DATA holder+8(SB)/8, $holder(SB)
DATA holder+16(SB)/8, $extvar(SB)
`)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFile(f)
if err != nil {
t.Fatalf("assemble: %v", err)
}
obj, err := img.ELFObject()
if err != nil {
t.Fatalf("ELFObject: %v", err)
}
ef, err := elf.NewFile(bytes.NewReader(obj))
if err != nil {
t.Fatalf("parse emitted object: %v", err)
}
defer ef.Close()
relaData := ef.Section(".rela.data")
if relaData == nil {
t.Fatal("missing .rela.data section")
}
if relaData.Link == 0 || ef.Sections[relaData.Link].Name != ".symtab" {
t.Errorf(".rela.data sh_link = %d, want the .symtab index", relaData.Link)
}
if ef.Sections[relaData.Info].Name != ".data" {
t.Errorf(".rela.data sh_info = %d, want the .data index", relaData.Info)
}
relas, err := relaData.Data()
if err != nil {
t.Fatal(err)
}
var got []struct {
off uint64
sym uint32
typ uint32
addend int64
}
for i := 0; i+24 <= len(relas); i += 24 {
got = append(got, struct {
off uint64
sym uint32
typ uint32
addend int64
}{
off: binary.LittleEndian.Uint64(relas[i:]),
// r_info packs the type in the low dword and the symbol index
// in the high dword.
typ: binary.LittleEndian.Uint32(relas[i+8:]),
sym: binary.LittleEndian.Uint32(relas[i+12:]),
addend: int64(binary.LittleEndian.Uint64(relas[i+16:])),
})
}
// debug/elf hides the table's null entry, so raw index s names syms[s-1].
syms, err := ef.Symbols()
if err != nil {
t.Fatal(err)
}
name := func(idx uint32) string {
if idx >= 1 && int(idx) <= len(syms) {
return syms[idx-1].Name
}
return ""
}
// The offsets are data-section-relative: the field's DATA offset plus
// the symbol's position in .data (the layout aligns each symbol to 16).
base := uint64(0)
for _, d := range img.DataSyms {
if d.Name == "holder" {
base = uint64(d.Offset)
}
}
want := []struct {
off uint64
typ uint32
addend int64
target string
}{
{off: base + 0, typ: uint32(elf.R_X86_64_64), addend: 5, target: "Keep"},
{off: base + 8, typ: uint32(elf.R_X86_64_64), addend: 0, target: "holder"},
{off: base + 16, typ: uint32(elf.R_X86_64_64), addend: 0, target: "extvar"},
}
if len(got) != len(want) {
t.Fatalf(".rela.data entries = %d, want %d", len(got), len(want))
}
for i, w := range want {
g := got[i]
if g.off != w.off || g.typ != w.typ || g.addend != w.addend {
t.Errorf("entry %d = {off %d typ %d addend %d}, want {off %d typ %d addend %d}",
i, g.off, g.typ, g.addend, w.off, w.typ, w.addend)
}
if n := name(g.sym); n != w.target {
t.Errorf("entry %d names %q, want %q", i, n, w.target)
}
}
}
+68 -11
View File
@@ -18,12 +18,16 @@ const (
rArm64AddAbsLo12NC = 277 // R_AARCH64_ADD_ABS_LO12_NC (ADD page offset)
rArm64Call26 = 283 // R_AARCH64_CALL26 (BL instruction)
rArm64Ldst64Lo12NC = 286 // R_AARCH64_LDST64_ABS_LO12_NC (64-bit LDR/STR page offset)
// R_AARCH64_ABS32 (debug/elf 258): the absolute 32-bit address of a
// symbol, the R_ADDR shape a 4-byte DATA field carries. ABS64 (257)
// lives with the DWARF fixup constants as rAARCH64Abs64.
rArm64Abs32 = 258
)
// ELFAARCH64Object returns the image as an ELF64 relocatable object file for
// AArch64 (EM_AARCH64, 64-bit, little-endian). The structure mirrors the
// amd64 and RISC-V ELF emitters: .text, .data, .symtab, .strtab and an
// optional .rela.text.
// amd64 and RISC-V ELF emitters: .text, .data, .symtab, .strtab, an
// optional .rela.text and an optional .rela.data.
func (img *Image) ELFAARCH64Object() ([]byte, error) {
le := binary.LittleEndian
@@ -133,6 +137,50 @@ func (img *Image) ELFAARCH64Object() ([]byte, error) {
}
}
// The data symbols' symbol-valued DATA fields ("DATA s+0(SB)/8,
// $other(SB)") become .rela.data entries: an absolute relocation of the
// DATA line's width at the field's data-section offset, S + A with no
// PC term. Widths 4 and 8 have ELF relocation shapes; narrower fields
// cannot hold an address, so they are refused rather than truncated.
var dataRelas []elfRela
for _, d := range img.DataSyms {
for _, r := range d.Relocs {
idx, ok := symIdx[r.Name]
if !ok {
return nil, fmt.Errorf("data relocation references unknown symbol %q", r.Name)
}
var typ uint32
switch r.Siz {
case 8:
typ = rAARCH64Abs64
case 4:
typ = rArm64Abs32
default:
return nil, fmt.Errorf("DATA %q: a symbol value of width %d has no ELF relocation", d.Name, r.Siz)
}
dataRelas = append(dataRelas, elfRela{
off: uint64(d.Offset + r.Off),
sym: idx,
typ: typ,
addend: r.Addend,
})
}
}
// Section presence: .rela.text only when there are code relocations,
// .rela.data only when a DATA line holds a symbol value.
hasRela := len(relas) > 0
hasDataRela := len(dataRelas) > 0
nSections := 6
if hasRela {
nSections++
}
if hasDataRela {
nSections++
}
secSymtab, secStrtab := 3, 4
secShstr := nSections - 1
// String tables.
stNames := newElfStrtab()
for _, s := range syms {
@@ -142,18 +190,13 @@ func (img *Image) ELFAARCH64Object() ([]byte, error) {
for _, n := range []string{".text", ".data", ".symtab", ".strtab", ".rela.text", ".shstrtab"} {
stSections.add(n)
}
if hasDataRela {
stSections.add(".rela.data")
}
for _, n := range dwarfSectionNames {
stSections.add(n)
}
hasRela := len(relas) > 0
nSections := 6
if hasRela {
nSections = 7
}
secSymtab, secStrtab := 3, 4
secShstr := nSections - 1
// Layout.
var out []byte
out = append(out, make([]byte, 64)...)
@@ -188,7 +231,7 @@ func (img *Image) ELFAARCH64Object() ([]byte, error) {
strtabOff := len(out)
out = append(out, stNames.bytes()...)
var relaOff int
var relaOff, relaDataOff int
if hasRela {
align(8)
relaOff = len(out)
@@ -200,6 +243,17 @@ func (img *Image) ELFAARCH64Object() ([]byte, error) {
out = append(out, b[:]...)
}
}
if hasDataRela {
align(8)
relaDataOff = len(out)
for _, r := range dataRelas {
var b [24]byte
le.PutUint64(b[0:], r.off)
le.PutUint64(b[8:], uint64(r.sym)<<32|uint64(r.typ))
le.PutUint64(b[16:], uint64(r.addend))
out = append(out, b[:]...)
}
}
shstrOff := len(out)
out = append(out, stSections.bytes()...)
@@ -257,6 +311,9 @@ func (img *Image) ELFAARCH64Object() ([]byte, error) {
if hasRela {
putSh(".rela.text", shtRela, 0, relaOff, 24*len(relas), secSymtab, secText, 8, 24)
}
if hasDataRela {
putSh(".rela.data", shtRela, 0, relaDataOff, 24*len(dataRelas), secSymtab, secData, 8, 24)
}
putSh(".shstrtab", shtStrtab, 0, shstrOff, len(stSections.bytes()), 0, 0, 1, 0)
// DWARF section headers; their indices follow the write order.
if dw != nil {
+115
View File
@@ -197,3 +197,118 @@ TEXT ·add(SB), NOSPLIT, $0-24
t.Error("unexpected .rela.text section when there are no relocations")
}
}
// TestELFAARCH64ObjectDataRelocation checks that a symbol-valued DATA field
// ("DATA s+0(SB)/8, $other(SB)") reaches the AArch64 ELF object as a
// .rela.data entry: an R_AARCH64_ABS64 (ABS32 for a width-4 field) at the
// field's offset within .data, against the named symbol, external targets
// included.
func TestELFAARCH64ObjectDataRelocation(t *testing.T) {
f, errs := parser.Parse("t_arm64.s", `#include "textflag.h"
TEXT ·Keep(SB), NOSPLIT, $0-0
RET
GLOBL holder(SB), NOPTR, $32
DATA holder+0(SB)/8, $·Keep+5(SB)
DATA holder+8(SB)/8, $holder(SB)
DATA holder+16(SB)/8, $extvar(SB)
DATA holder+24(SB)/4, $Keep(SB)
`)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileARM64(f)
if err != nil {
t.Fatalf("AssembleFileARM64: %v", err)
}
obj, err := img.ELFAARCH64Object()
if err != nil {
t.Fatalf("ELFAARCH64Object: %v", err)
}
checkELFSectionAccounting(t, obj)
ef, err := elf.NewFile(bytes.NewReader(obj))
if err != nil {
t.Fatalf("parse emitted object: %v", err)
}
defer ef.Close()
relaData := ef.Section(".rela.data")
if relaData == nil {
t.Fatal("missing .rela.data section")
}
if relaData.Type != elf.SHT_RELA {
t.Errorf(".rela.data type = %v, want SHT_RELA", relaData.Type)
}
if relaData.Link == 0 || ef.Sections[relaData.Link].Name != ".symtab" {
t.Errorf(".rela.data sh_link = %d, want the .symtab index", relaData.Link)
}
if ef.Sections[relaData.Info].Name != ".data" {
t.Errorf(".rela.data sh_info = %d, want the .data index", relaData.Info)
}
relas, err := relaData.Data()
if err != nil {
t.Fatal(err)
}
var got []struct {
off uint64
sym uint32
typ uint32
addend int64
}
for i := 0; i+24 <= len(relas); i += 24 {
got = append(got, struct {
off uint64
sym uint32
typ uint32
addend int64
}{
off: binary.LittleEndian.Uint64(relas[i:]),
// r_info packs the type in the low dword and the symbol index
// in the high dword.
typ: binary.LittleEndian.Uint32(relas[i+8:]),
sym: binary.LittleEndian.Uint32(relas[i+12:]),
addend: int64(binary.LittleEndian.Uint64(relas[i+16:])),
})
}
// debug/elf hides the table's null entry, so raw index s names syms[s-1].
syms, err := ef.Symbols()
if err != nil {
t.Fatal(err)
}
name := func(idx uint32) string {
if idx >= 1 && int(idx) <= len(syms) {
return syms[idx-1].Name
}
return ""
}
// The offsets are data-section-relative: the field's DATA offset plus
// the symbol's position in .data (the layout aligns each symbol to 16).
base := uint64(0)
for _, d := range img.DataSyms {
if d.Name == "holder" {
base = uint64(d.Offset)
}
}
want := []struct {
off uint64
typ uint32
addend int64
target string
}{
{off: base + 0, typ: uint32(elf.R_AARCH64_ABS64), addend: 5, target: "Keep"},
{off: base + 8, typ: uint32(elf.R_AARCH64_ABS64), addend: 0, target: "holder"},
{off: base + 16, typ: uint32(elf.R_AARCH64_ABS64), addend: 0, target: "extvar"},
{off: base + 24, typ: uint32(elf.R_AARCH64_ABS32), addend: 0, target: "Keep"},
}
if len(got) != len(want) {
t.Fatalf(".rela.data entries = %d, want %d", len(got), len(want))
}
for i, w := range want {
g := got[i]
if g.off != w.off || g.typ != w.typ || g.addend != w.addend {
t.Errorf("entry %d = {off %d typ %d addend %d}, want {off %d typ %d addend %d}",
i, g.off, g.typ, g.addend, w.off, w.typ, w.addend)
}
if n := name(g.sym); n != w.target {
t.Errorf("entry %d names %q, want %q", i, n, w.target)
}
}
}
+68 -11
View File
@@ -23,12 +23,16 @@ const (
rLarchPCALAHI20 = 71 // R_LARCH_PCALA_HI20 (pcalau12i)
rLarchPCALALO12 = 72 // R_LARCH_PCALA_LO12 (addi.d/ld/st)
rLarchB26 = 66 // R_LARCH_B26 (b/bl, matches the Go linker's mapping)
// R_LARCH_32 (debug/elf 1): the absolute 32-bit address of a symbol,
// the R_ADDR shape a 4-byte DATA field carries. R_LARCH_64 (2) lives
// with the DWARF fixup constants as rLarchAbs64.
rLarchAbs32 = 1
)
// ELFLOONG64Object returns the image as an ELF64 relocatable object file for
// LoongArch (EM_LOONGARCH, 64-bit, little-endian). The structure mirrors the
// amd64 and RISC-V ELF emitters: .text, .data, .symtab, .strtab and an
// optional .rela.text.
// amd64 and RISC-V ELF emitters: .text, .data, .symtab, .strtab, an
// optional .rela.text and an optional .rela.data.
func (img *Image) ELFLOONG64Object() ([]byte, error) {
le := binary.LittleEndian
@@ -117,6 +121,50 @@ func (img *Image) ELFLOONG64Object() ([]byte, error) {
}
}
// The data symbols' symbol-valued DATA fields ("DATA s+0(SB)/8,
// $other(SB)") become .rela.data entries: an absolute relocation of the
// DATA line's width at the field's data-section offset, S + A with no
// PC term. Widths 4 and 8 have ELF relocation shapes; narrower fields
// cannot hold an address, so they are refused rather than truncated.
var dataRelas []elfRela
for _, d := range img.DataSyms {
for _, r := range d.Relocs {
idx, ok := symIdx[r.Name]
if !ok {
return nil, fmt.Errorf("data relocation references unknown symbol %q", r.Name)
}
var typ uint32
switch r.Siz {
case 8:
typ = rLarchAbs64
case 4:
typ = rLarchAbs32
default:
return nil, fmt.Errorf("DATA %q: a symbol value of width %d has no ELF relocation", d.Name, r.Siz)
}
dataRelas = append(dataRelas, elfRela{
off: uint64(d.Offset + r.Off),
sym: idx,
typ: typ,
addend: r.Addend,
})
}
}
// Section presence: .rela.text only when there are code relocations,
// .rela.data only when a DATA line holds a symbol value.
hasRela := len(relas) > 0
hasDataRela := len(dataRelas) > 0
nSections := 6
if hasRela {
nSections++
}
if hasDataRela {
nSections++
}
secSymtab, secStrtab := 3, 4
secShstr := nSections - 1
// String tables.
stNames := newElfStrtab()
for _, s := range syms {
@@ -126,18 +174,13 @@ func (img *Image) ELFLOONG64Object() ([]byte, error) {
for _, n := range []string{".text", ".data", ".symtab", ".strtab", ".rela.text", ".shstrtab"} {
stSections.add(n)
}
if hasDataRela {
stSections.add(".rela.data")
}
for _, n := range dwarfSectionNames {
stSections.add(n)
}
hasRela := len(relas) > 0
nSections := 6
if hasRela {
nSections = 7
}
secSymtab, secStrtab := 3, 4
secShstr := nSections - 1
// Layout.
var out []byte
out = append(out, make([]byte, 64)...)
@@ -172,7 +215,7 @@ func (img *Image) ELFLOONG64Object() ([]byte, error) {
strtabOff := len(out)
out = append(out, stNames.bytes()...)
var relaOff int
var relaOff, relaDataOff int
if hasRela {
align(8)
relaOff = len(out)
@@ -184,6 +227,17 @@ func (img *Image) ELFLOONG64Object() ([]byte, error) {
out = append(out, b[:]...)
}
}
if hasDataRela {
align(8)
relaDataOff = len(out)
for _, r := range dataRelas {
var b [24]byte
le.PutUint64(b[0:], r.off)
le.PutUint64(b[8:], uint64(r.sym)<<32|uint64(r.typ))
le.PutUint64(b[16:], uint64(r.addend))
out = append(out, b[:]...)
}
}
shstrOff := len(out)
out = append(out, stSections.bytes()...)
@@ -239,6 +293,9 @@ func (img *Image) ELFLOONG64Object() ([]byte, error) {
if hasRela {
putSh(".rela.text", shtRela, 0, relaOff, 24*len(relas), secSymtab, secText, 8, 24)
}
if hasDataRela {
putSh(".rela.data", shtRela, 0, relaDataOff, 24*len(dataRelas), secSymtab, secData, 8, 24)
}
putSh(".shstrtab", shtStrtab, 0, shstrOff, len(stSections.bytes()), 0, 0, 1, 0)
// DWARF section headers; their indices follow the write order.
if dw != nil {
+115
View File
@@ -245,3 +245,118 @@ func TestELFLOONG64BranchRelocation(t *testing.T) {
}
}
}
// TestELFLOONG64ObjectDataRelocation checks that a symbol-valued DATA field
// ("DATA s+0(SB)/8, $other(SB)") reaches the LoongArch ELF object as a
// .rela.data entry: an R_LARCH_64 (R_LARCH_32 for a width-4 field) at the
// field's offset within .data, against the named symbol, external targets
// included.
func TestELFLOONG64ObjectDataRelocation(t *testing.T) {
f, errs := parser.Parse("t_loong64.s", `#include "textflag.h"
TEXT ·Keep(SB), NOSPLIT, $0-0
RET
GLOBL holder(SB), NOPTR, $32
DATA holder+0(SB)/8, $·Keep+5(SB)
DATA holder+8(SB)/8, $holder(SB)
DATA holder+16(SB)/8, $extvar(SB)
DATA holder+24(SB)/4, $Keep(SB)
`)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileLOONG64(f)
if err != nil {
t.Fatalf("AssembleFileLOONG64: %v", err)
}
obj, err := img.ELFLOONG64Object()
if err != nil {
t.Fatalf("ELFLOONG64Object: %v", err)
}
checkELFSectionAccounting(t, obj)
ef, err := elf.NewFile(bytes.NewReader(obj))
if err != nil {
t.Fatalf("parse emitted object: %v", err)
}
defer ef.Close()
relaData := ef.Section(".rela.data")
if relaData == nil {
t.Fatal("missing .rela.data section")
}
if relaData.Type != elf.SHT_RELA {
t.Errorf(".rela.data type = %v, want SHT_RELA", relaData.Type)
}
if relaData.Link == 0 || ef.Sections[relaData.Link].Name != ".symtab" {
t.Errorf(".rela.data sh_link = %d, want the .symtab index", relaData.Link)
}
if ef.Sections[relaData.Info].Name != ".data" {
t.Errorf(".rela.data sh_info = %d, want the .data index", relaData.Info)
}
relas, err := relaData.Data()
if err != nil {
t.Fatal(err)
}
var got []struct {
off uint64
sym uint32
typ uint32
addend int64
}
for i := 0; i+24 <= len(relas); i += 24 {
got = append(got, struct {
off uint64
sym uint32
typ uint32
addend int64
}{
off: binary.LittleEndian.Uint64(relas[i:]),
// r_info packs the type in the low dword and the symbol index
// in the high dword.
typ: binary.LittleEndian.Uint32(relas[i+8:]),
sym: binary.LittleEndian.Uint32(relas[i+12:]),
addend: int64(binary.LittleEndian.Uint64(relas[i+16:])),
})
}
// debug/elf hides the table's null entry, so raw index s names syms[s-1].
syms, err := ef.Symbols()
if err != nil {
t.Fatal(err)
}
name := func(idx uint32) string {
if idx >= 1 && int(idx) <= len(syms) {
return syms[idx-1].Name
}
return ""
}
// The offsets are data-section-relative: the field's DATA offset plus
// the symbol's position in .data (the layout aligns each symbol to 16).
base := uint64(0)
for _, d := range img.DataSyms {
if d.Name == "holder" {
base = uint64(d.Offset)
}
}
want := []struct {
off uint64
typ uint32
addend int64
target string
}{
{off: base + 0, typ: uint32(elf.R_LARCH_64), addend: 5, target: "Keep"},
{off: base + 8, typ: uint32(elf.R_LARCH_64), addend: 0, target: "holder"},
{off: base + 16, typ: uint32(elf.R_LARCH_64), addend: 0, target: "extvar"},
{off: base + 24, typ: uint32(elf.R_LARCH_32), addend: 0, target: "Keep"},
}
if len(got) != len(want) {
t.Fatalf(".rela.data entries = %d, want %d", len(got), len(want))
}
for i, w := range want {
g := got[i]
if g.off != w.off || g.typ != w.typ || g.addend != w.addend {
t.Errorf("entry %d = {off %d typ %d addend %d}, want {off %d typ %d addend %d}",
i, g.off, g.typ, g.addend, w.off, w.typ, w.addend)
}
if n := name(g.sym); n != w.target {
t.Errorf("entry %d names %q, want %q", i, n, w.target)
}
}
}
+68 -10
View File
@@ -24,11 +24,16 @@ const (
rRISCVPCRELHI20 = 23 // R_RISCV_PCREL_HI20
rRISCVPCRELLO12I = 24 // R_RISCV_PCREL_LO12_I
rRISCVPCRELLO12S = 25 // R_RISCV_PCREL_LO12_S
// R_RISCV_32 (debug/elf 1): the absolute 32-bit address of a symbol,
// the R_ADDR shape a 4-byte DATA field carries. R_RISCV_64 (2) lives
// with the DWARF fixup constants as rRISCVAbs64.
rRISVCAbs32 = 1
)
// ELFRISCVObject returns the image as an ELF64 relocatable object file for
// RISC-V (EM_RISCV, 64-bit, little-endian). The structure mirrors the amd64
// ELF emission: .text, .data, .symtab, .strtab and optional .rela.text.
// ELF emission: .text, .data, .symtab, .strtab, an optional .rela.text and
// an optional .rela.data.
func (img *Image) ELFRISCVObject() ([]byte, error) {
le := binary.LittleEndian
@@ -129,6 +134,50 @@ func (img *Image) ELFRISCVObject() ([]byte, error) {
}
}
// The data symbols' symbol-valued DATA fields ("DATA s+0(SB)/8,
// $other(SB)") become .rela.data entries: an absolute relocation of the
// DATA line's width at the field's data-section offset, S + A with no
// PC term. Widths 4 and 8 have ELF relocation shapes; narrower fields
// cannot hold an address, so they are refused rather than truncated.
var dataRelas []elfRela
for _, d := range img.DataSyms {
for _, r := range d.Relocs {
idx, ok := symIdx[r.Name]
if !ok {
return nil, fmt.Errorf("data relocation references unknown symbol %q", r.Name)
}
var typ uint32
switch r.Siz {
case 8:
typ = rRISCVAbs64
case 4:
typ = rRISVCAbs32
default:
return nil, fmt.Errorf("DATA %q: a symbol value of width %d has no ELF relocation", d.Name, r.Siz)
}
dataRelas = append(dataRelas, elfRela{
off: uint64(d.Offset + r.Off),
sym: idx,
typ: typ,
addend: r.Addend,
})
}
}
// Section presence: .rela.text only when there are code relocations,
// .rela.data only when a DATA line holds a symbol value.
hasRela := len(relas) > 0
hasDataRela := len(dataRelas) > 0
nSections := 6
if hasRela {
nSections++
}
if hasDataRela {
nSections++
}
secSymtab, secStrtab := 3, 4
secShstr := nSections - 1
// String tables.
stNames := newElfStrtab()
for _, s := range syms {
@@ -138,18 +187,13 @@ func (img *Image) ELFRISCVObject() ([]byte, error) {
for _, n := range []string{".text", ".data", ".symtab", ".strtab", ".rela.text", ".shstrtab"} {
stSections.add(n)
}
if hasDataRela {
stSections.add(".rela.data")
}
for _, n := range dwarfSectionNames {
stSections.add(n)
}
hasRela := len(relas) > 0
nSections := 6
if hasRela {
nSections = 7
}
secSymtab, secStrtab := 3, 4
secShstr := nSections - 1
// Layout.
var out []byte
out = append(out, make([]byte, 64)...)
@@ -184,7 +228,7 @@ func (img *Image) ELFRISCVObject() ([]byte, error) {
strtabOff := len(out)
out = append(out, stNames.bytes()...)
var relaOff int
var relaOff, relaDataOff int
if hasRela {
align(8)
relaOff = len(out)
@@ -196,6 +240,17 @@ func (img *Image) ELFRISCVObject() ([]byte, error) {
out = append(out, b[:]...)
}
}
if hasDataRela {
align(8)
relaDataOff = len(out)
for _, r := range dataRelas {
var b [24]byte
le.PutUint64(b[0:], r.off)
le.PutUint64(b[8:], uint64(r.sym)<<32|uint64(r.typ))
le.PutUint64(b[16:], uint64(r.addend))
out = append(out, b[:]...)
}
}
shstrOff := len(out)
out = append(out, stSections.bytes()...)
@@ -251,6 +306,9 @@ func (img *Image) ELFRISCVObject() ([]byte, error) {
if hasRela {
putSh(".rela.text", shtRela, 0, relaOff, 24*len(relas), secSymtab, secText, 8, 24)
}
if hasDataRela {
putSh(".rela.data", shtRela, 0, relaDataOff, 24*len(dataRelas), secSymtab, secData, 8, 24)
}
putSh(".shstrtab", shtStrtab, 0, shstrOff, len(stSections.bytes()), 0, 0, 1, 0)
// DWARF section headers; their indices follow the write order.
if dw != nil {
+128
View File
@@ -0,0 +1,128 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
package asm
import (
"bytes"
"debug/elf"
"encoding/binary"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
)
// TestELFRISCVObjectDataRelocation checks that a symbol-valued DATA field
// ("DATA s+0(SB)/8, $other(SB)") reaches the RISC-V ELF object as a
// .rela.data entry: an R_RISCV_64 (R_RISCV_32 for a width-4 field) at the
// field's offset within .data, against the named symbol, external targets
// included.
func TestELFRISCVObjectDataRelocation(t *testing.T) {
f, errs := parser.Parse("t_riscv64.s", `#include "textflag.h"
TEXT ·Keep(SB), NOSPLIT, $0-0
RET
GLOBL holder(SB), NOPTR, $32
DATA holder+0(SB)/8, $·Keep+5(SB)
DATA holder+8(SB)/8, $holder(SB)
DATA holder+16(SB)/8, $extvar(SB)
DATA holder+24(SB)/4, $Keep(SB)
`)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileRISCV(f)
if err != nil {
t.Fatalf("AssembleFileRISCV: %v", err)
}
obj, err := img.ELFRISCVObject()
if err != nil {
t.Fatalf("ELFRISCVObject: %v", err)
}
checkELFSectionAccounting(t, obj)
ef, err := elf.NewFile(bytes.NewReader(obj))
if err != nil {
t.Fatalf("parse emitted object: %v", err)
}
defer ef.Close()
relaData := ef.Section(".rela.data")
if relaData == nil {
t.Fatal("missing .rela.data section")
}
if relaData.Type != elf.SHT_RELA {
t.Errorf(".rela.data type = %v, want SHT_RELA", relaData.Type)
}
if relaData.Link == 0 || ef.Sections[relaData.Link].Name != ".symtab" {
t.Errorf(".rela.data sh_link = %d, want the .symtab index", relaData.Link)
}
if ef.Sections[relaData.Info].Name != ".data" {
t.Errorf(".rela.data sh_info = %d, want the .data index", relaData.Info)
}
relas, err := relaData.Data()
if err != nil {
t.Fatal(err)
}
var got []struct {
off uint64
sym uint32
typ uint32
addend int64
}
for i := 0; i+24 <= len(relas); i += 24 {
got = append(got, struct {
off uint64
sym uint32
typ uint32
addend int64
}{
off: binary.LittleEndian.Uint64(relas[i:]),
// r_info packs the type in the low dword and the symbol index
// in the high dword.
typ: binary.LittleEndian.Uint32(relas[i+8:]),
sym: binary.LittleEndian.Uint32(relas[i+12:]),
addend: int64(binary.LittleEndian.Uint64(relas[i+16:])),
})
}
// debug/elf hides the table's null entry, so raw index s names syms[s-1].
syms, err := ef.Symbols()
if err != nil {
t.Fatal(err)
}
name := func(idx uint32) string {
if idx >= 1 && int(idx) <= len(syms) {
return syms[idx-1].Name
}
return ""
}
// The offsets are data-section-relative: the field's DATA offset plus
// the symbol's position in .data (the layout aligns each symbol to 16).
base := uint64(0)
for _, d := range img.DataSyms {
if d.Name == "holder" {
base = uint64(d.Offset)
}
}
want := []struct {
off uint64
typ uint32
addend int64
target string
}{
{off: base + 0, typ: uint32(elf.R_RISCV_64), addend: 5, target: "Keep"},
{off: base + 8, typ: uint32(elf.R_RISCV_64), addend: 0, target: "holder"},
{off: base + 16, typ: uint32(elf.R_RISCV_64), addend: 0, target: "extvar"},
{off: base + 24, typ: uint32(elf.R_RISCV_32), addend: 0, target: "Keep"},
}
if len(got) != len(want) {
t.Fatalf(".rela.data entries = %d, want %d", len(got), len(want))
}
for i, w := range want {
g := got[i]
if g.off != w.off || g.typ != w.typ || g.addend != w.addend {
t.Errorf("entry %d = {off %d typ %d addend %d}, want {off %d typ %d addend %d}",
i, g.off, g.typ, g.addend, w.off, w.typ, w.addend)
}
if n := name(g.sym); n != w.target {
t.Errorf("entry %d names %q, want %q", i, n, w.target)
}
}
}
+35 -8
View File
@@ -18,7 +18,14 @@ func Encodable(mnemonic string) bool {
// Fixed-name instructions (no size suffix).
switch upper {
case "RET", "NOP", "CALL", "JMP":
case "RET", "NOP", "CALL", "JMP",
"POPFQ", "PUSHFQ", "INT", "LDMXCSR", "STMXCSR", "CMPSD", "SHA256RNDS2",
// The literal-data pseudo-ops, the accepted-and-ignored END and the
// SP adjust.
"BYTE", "WORD", "LONG", "QUAD", "END", "ADJSP":
return true
}
if _, ok := noOperandTable[upper]; ok {
return true
}
if _, ok := condCode(upper); ok {
@@ -31,7 +38,7 @@ func Encodable(mnemonic string) bool {
return false
}
if isVex(base) || isEvex(base) || isKOp(base) || isGather(base) || isScatter(base) ||
base == "KMOVW" || base == "KMOVQ" {
base == "KMOVW" || base == "KMOVQ" || base == "KMOVB" || base == "KMOVD" {
return true
}
@@ -53,13 +60,28 @@ func Encodable(mnemonic string) bool {
}
}
// Legacy SSE shuffles and packed binaries dispatch on the full name.
// Legacy SSE shuffles and packed binaries dispatch on the full name; so
// do the imm8-controlled instructions, the lane extracts and inserts and
// the packed integer shifts (their trailing width letters belong to the
// mnemonic).
if _, ok := sseShufTable[upper]; ok {
return true
}
if _, ok := sseBinTable[upper]; ok {
return true
}
if _, ok := sseImm3Table[upper]; ok {
return true
}
if _, ok := sseExtractTable[upper]; ok {
return true
}
if _, ok := sseInsertTable[upper]; ok {
return true
}
if _, ok := sseShiftImm[upper]; ok {
return true
}
// The size-suffix split: retry the tables and the scalar switch on the
// base.
@@ -74,12 +96,15 @@ func Encodable(mnemonic string) bool {
}
}
switch base2 {
case "MOV",
"ADD", "SUB", "AND", "OR", "XOR", "CMP",
case "MOV", "MOVD",
"ADD", "SUB", "AND", "OR", "XOR", "CMP", "ADC", "SBB",
"TEST",
"LEA",
"INC", "DEC", "NEG", "NOT",
"SHL", "SHR", "SAR",
"INC", "DEC", "NEG", "NOT", "MUL", "DIV", "IDIV",
"SHL", "SHR", "SAR", "SAL", "ROL", "ROR", "RCL", "RCR",
"BT", "BTS", "BTR", "BTC",
"XCHG", "CMPXCHG", "XADD", "CRC32", "ADCX", "ADOX",
"MOVS", "STOS",
"IMUL", "IMUL3",
"PUSH", "POP",
"BSF", "BSR", "LZCNT", "TZCNT", "POPCNT",
@@ -88,7 +113,9 @@ func Encodable(mnemonic string) bool {
"MOVBLZX", "MOVBQZX", "MOVWLZX", "MOVWQZX", "MOVWLSX", "MOVLQSX",
"MOVBWZX", "MOVBWSX", "MOVBLSX", "MOVBQSX", "MOVWQSX", "MOVLQZX",
"CVTSL2SD", "CVTSQ2SD",
"MOVOU", "MOVO", "MOVUPS", "MOVAPS", "MOVUPD", "MOVAPD", "MOVSD", "MOVSS":
"CVTSD2S", "CVTTSD2S", "CVTSS2S", "CVTTSS2S",
"FMOVD",
"MOVOU", "MOVO", "MOVOA", "MOVUPS", "MOVAPS", "MOVUPD", "MOVAPD", "MOVSD", "MOVSS":
return true
}
// Full-name dispatches the size split would eat (a trailing width
+160 -7
View File
@@ -58,6 +58,51 @@ func (e *enc) encode(mnem string, ops []Operand) error {
if cc, ok := condCode(upper); ok {
return e.encodeJcc(cc, ops)
}
// No-operand system and string-control instructions (CPUID, RDTSC,
// SYSCALL, the fences, UNDEF, …).
if op, ok := noOperandTable[upper]; ok {
if len(ops) != 0 {
return fmt.Errorf("%s takes no operands, got %d", upper, len(ops))
}
return e.emit(&instr{opcode: op, modrm: -1, sib: -1})
}
// POPFQ/PUSHFQ are exact names: the bare POPF/PUSHF and the L spellings
// are rejected by go tool asm in 64-bit mode, so they stay unsupported.
switch upper {
case "POPFQ":
if len(ops) != 0 {
return fmt.Errorf("POPFQ takes no operands, got %d", len(ops))
}
return e.emit(&instr{opcode: []byte{0x9D}, modrm: -1, sib: -1})
case "PUSHFQ":
if len(ops) != 0 {
return fmt.Errorf("PUSHFQ takes no operands, got %d", len(ops))
}
return e.emit(&instr{opcode: []byte{0x9C}, modrm: -1, sib: -1})
case "INT":
return e.encodeInt(ops)
case "LDMXCSR":
return e.encodeMxcsr(2, ops)
case "STMXCSR":
return e.encodeMxcsr(3, ops)
// CMPSD is the scalar double compare, whose predicate immediate comes
// LAST in Plan 9 order (src, dst, $imm).
case "CMPSD":
return e.encodeCmpsd(ops)
// SHA256RNDS2 carries the round constant in a literal X0 first operand.
case "SHA256RNDS2":
return e.encodeSha256rnds2(ops)
// BYTE, WORD, LONG and QUAD write the immediate into the text stream
// itself: 1, 2, 4 or 8 literal bytes, little-endian. END is accepted
// and ignored. ADJSP adjusts SP by the immediate, sign-chosen between
// the SUBQ and ADDQ forms.
case "BYTE", "WORD", "LONG", "QUAD":
return e.encodeData(upper, ops)
case "END":
return e.encodeEnd(ops)
case "ADJSP":
return e.encodeAdjsp(ops)
}
// VEX (AVX/AVX2) and EVEX (AVX-512) instructions: the trailing
// B/W/L/Q/D is part of the mnemonic, not a size suffix, so dispatch
@@ -67,7 +112,9 @@ func (e *enc) encode(mnem string, ops []Operand) error {
if err != nil {
return err
}
if isVex(base) || isEvex(base) || isKOp(base) || isGather(base) || isScatter(base) || base == "KMOVW" || base == "KMOVQ" {
if isVex(base) || isEvex(base) || isKOp(base) || isGather(base) || isScatter(base) ||
isEvexPrefGather(base) ||
base == "KMOVW" || base == "KMOVQ" || base == "KMOVB" || base == "KMOVD" {
return e.encodeVec(base, ops, sfx)
}
if sfx.any() {
@@ -101,6 +148,21 @@ func (e *enc) encode(mnem string, ops []Operand) error {
if m, ok := sseBinTable[base]; ok {
return e.encodeSSEBin(m, ops)
}
// The imm8-controlled legacy instructions, the lane extracts and inserts
// and the packed integer shifts all dispatch on the full name: a trailing
// width letter here belongs to the mnemonic, not to the size split.
if m, ok := sseImm3Table[upper]; ok {
return e.encodeSSEImm3(m, ops)
}
if m, ok := sseExtractTable[upper]; ok {
return e.encodeSSEExtract(m, ops)
}
if m, ok := sseInsertTable[upper]; ok {
return e.encodeSSEInsert(m, ops)
}
if _, ok := sseShiftImm[upper]; ok {
return e.encodeSSEShift(upper, ops)
}
// PMOVMSKB ends in a width letter the size split would eat, so it
// dispatches on the full name like the packed binaries above.
if upper == "PMOVMSKB" {
@@ -109,16 +171,36 @@ func (e *enc) encode(mnem string, ops []Operand) error {
switch base {
case "MOV":
return e.encodeMov(ops, size)
case "ADD", "SUB", "AND", "OR", "XOR", "CMP":
// MOVD is the Go assembler's alias of MOVQ: the same byte forms, 64-bit
// REX.W and all.
case "MOVD":
return e.encodeMov(ops, 8)
case "ADD", "SUB", "AND", "OR", "XOR", "CMP", "ADC", "SBB":
return e.encodeALU(aluOp[base], ops, size)
case "TEST":
return e.encodeTest(ops, size)
case "LEA":
return e.encodeLea(ops, size)
case "INC", "DEC", "NEG", "NOT":
case "INC", "DEC", "NEG", "NOT", "MUL", "DIV", "IDIV":
return e.encodeUnary(unaryOp[base], ops, size)
case "SHL", "SHR", "SAR":
return e.encodeShift(shiftOp[base], ops, size)
case "SHL", "SHR", "SAR", "SAL", "ROL", "ROR", "RCL", "RCR":
return e.encodeShift(base, ops, size)
case "BT", "BTS", "BTR", "BTC":
return e.encodeBitTest(base, ops, size)
case "XCHG":
return e.encodeExchange(ops, size)
case "CMPXCHG":
return e.encodeRegRegOp(0xB0, 0xB1, base, ops, size)
case "XADD":
return e.encodeRegRegOp(0xC0, 0xC1, base, ops, size)
case "CRC32":
return e.encodeCrc32(ops, size)
case "ADCX":
return e.encodeCarryExt(0x66, ops, size)
case "ADOX":
return e.encodeCarryExt(0xF3, ops, size)
case "MOVS", "STOS":
return e.encodeStringOp(base, ops, size)
case "IMUL", "IMUL3":
return e.encodeImul(ops, size)
case "PUSH":
@@ -136,7 +218,11 @@ func (e *enc) encode(mnem string, ops []Operand) error {
return e.encodeMovExtend(base, ops)
case "CVTSL2SD", "CVTSQ2SD":
return e.encodeCvtsi2sd(base == "CVTSQ2SD", ops)
case "MOVOU", "MOVO", "MOVUPS", "MOVAPS", "MOVUPD", "MOVAPD", "MOVSD", "MOVSS":
case "CVTSD2S", "CVTTSD2S", "CVTSS2S", "CVTTSS2S":
return e.encodeCvtInt(base, ops, size)
case "FMOVD":
return e.encodeFmov(ops)
case "MOVOU", "MOVO", "MOVOA", "MOVUPS", "MOVAPS", "MOVUPD", "MOVAPD", "MOVSD", "MOVSS":
return e.encodeSSEMove(sseMoveTable[base], ops)
}
return fmt.Errorf("unsupported instruction %q", mnem)
@@ -166,6 +252,73 @@ var prefetchVariant = map[string]int{
"PREFETCHT2": 3,
}
// dataWidth is the literal byte count of each data-emission pseudo-op.
var dataWidth = map[string]int{
"BYTE": 1,
"WORD": 2,
"LONG": 4,
"QUAD": 8,
}
// encodeData emits the literal-data pseudo-ops: BYTE, WORD, LONG and QUAD
// write the immediate into the text stream as 1, 2, 4 or 8 bytes,
// little-endian, with no opcode lookup. The value is truncated to the
// width rather than range-checked, exactly as go tool asm behaves (BYTE
// $0x1FF emits FF, WORD $0x12345 emits 45 23, both without an error), and
// exactly one immediate is accepted: the toolchain rejects a list such as
// BYTE $1, $2, $3.
func (e *enc) encodeData(mnem string, ops []Operand) error {
if len(ops) != 1 {
return fmt.Errorf("%s expects 1 immediate operand, got %d", mnem, len(ops))
}
imm, ok := ops[0].(Imm)
if !ok {
return fmt.Errorf("%s requires an integer immediate", mnem)
}
width := dataWidth[mnem]
out := make([]byte, width)
u := uint64(imm)
for i := range width {
out[i] = byte(u >> (8 * i))
}
e.out = append(e.out, out...)
return nil
}
// encodeEnd accepts-and-ignores END. go tool asm drops the statement
// entirely: the AEND Prog is skipped when the program list is flushed, so
// the statements after an END still belong to the same function and the
// encoded body carries no trace of it, whatever operands follow the name
// (the toolchain takes END $0 and END AX alike). Zero bytes, no effect.
func (e *enc) encodeEnd(ops []Operand) error {
return nil
}
// encodeAdjsp emits ADJSP $imm: a positive value is SUBQ $imm, SP, a
// negative one ADDQ $-imm, SP, in the imm8 or imm32 form the magnitude
// picks (the same selection subSP and addSP make for the frame). go tool
// asm refuses ADJSP $0 outright, so a zero value is an error here too; the
// statement's effect on the SP balance is checked by the function-level
// assembly (checkAdjspBalance), as the toolchain's push/pop walk does.
func (e *enc) encodeAdjsp(ops []Operand) error {
if len(ops) != 1 {
return fmt.Errorf("ADJSP expects 1 immediate operand, got %d", len(ops))
}
imm, ok := ops[0].(Imm)
if !ok {
return fmt.Errorf("ADJSP requires an integer immediate")
}
switch v := int(imm); {
case v > 0:
e.out = append(e.out, subSP(v)...)
case v < 0:
e.out = append(e.out, addSP(-v)...)
default:
return fmt.Errorf("ADJSP $0 has no encoding")
}
return nil
}
// splitSize separates a trailing B/W/L/Q size suffix from the mnemonic.
func splitSize(upper string) (base string, size int) {
if upper == "" {
@@ -195,7 +348,7 @@ func (e *enc) encodeVec(upper string, ops []Operand, sfx evexSuffix) error {
if ss, ok := scatterTable[upper]; ok {
return e.encodeScatter(upper, ss, ops, sfx)
}
if upper == "KMOVW" || upper == "KMOVQ" {
if upper == "KMOVW" || upper == "KMOVQ" || upper == "KMOVB" || upper == "KMOVD" {
if sfx.any() {
return fmt.Errorf("%s takes no EVEX suffixes", upper)
}
+437
View File
@@ -200,10 +200,60 @@ func TestUnary(t *testing.T) {
func TestShift(t *testing.T) {
checkSyntax(t, "shl rdx, 0x2", "SHLQ", Imm(2), DX)
checkSyntax(t, "shl rdx, cl", "SHLQ", CL, DX)
checkSyntax(t, "shl rdx, cl", "SHLQ", CX, DX)
checkSyntax(t, "shl rdx, 0x1", "SHLQ", Imm(1), DX)
checkSyntax(t, "sar rcx, 0x1f", "SARQ", Imm(31), CX)
}
// TestDoubleShift pins the three-operand SHL/SHR form, which encodes as
// SHLD/SHRD: go tool asm accepts it for SHL/SHR at W/L/Q widths and rejects
// it for SAR, SAL, the rotates and the B width. The byte pins mirror the
// oracle's objdump output (48 0f a4 fe 0d for the first case, and so on).
func TestDoubleShift(t *testing.T) {
cases := []struct {
name string
mnem string
ops []Operand
want string // hex encoding
}{
{"SHLQ imm", "SHLQ", []Operand{Imm(0x0d), DI, SI}, "480fa4fe0d"},
{"SHLQ CX high regs", "SHLQ", []Operand{CX, Reg{idx: 8, size: 8}, Reg{idx: 9, size: 8}}, "4d0fa5c1"},
{"SHRQ imm", "SHRQ", []Operand{Imm(1), AX, CX}, "480facc101"},
{"SHLW imm", "SHLW", []Operand{Imm(1), AX, CX}, "660fa4c101"},
{"SHRD CL", "SHRQ", []Operand{CL, AX, CX}, "480fadc1"},
{"SHLD imm high regs", "SHLQ", []Operand{Imm(2), Reg{idx: 10, size: 8}, Reg{idx: 11, size: 8}}, "4d0fa4d302"},
{"SHRD imm max", "SHRQ", []Operand{Imm(63), Reg{idx: 9, size: 8}, Reg{idx: 15, size: 8}}, "4d0faccf3f"},
}
for _, c := range cases {
code, err := Encode(c.mnem, c.ops...)
if err != nil {
t.Errorf("%s: Encode: %v", c.name, err)
continue
}
if got := hexCompact(code); got != c.want {
t.Errorf("%s: bytes %s, want %s", c.name, got, c.want)
}
}
// Rejected forms: the oracle rejects every one of these.
rejected := []struct {
name string
mnem string
ops []Operand
}{
{"SARQ three operands", "SARQ", []Operand{Imm(1), AX, CX}},
{"SALQ three operands", "SALQ", []Operand{Imm(1), AX, CX}},
{"ROLQ three operands", "ROLQ", []Operand{Imm(1), AX, CX}},
{"SHLB three operands", "SHLB", []Operand{Imm(1), AL, CL}},
{"SHRQ memory source", "SHRQ", []Operand{Imm(1), Ptr(AX, 0, 8), CX}},
{"SHRQ ECX count", "SHRQ", []Operand{Reg{idx: 1, size: 4}, AX, CX}},
}
for _, c := range rejected {
if _, err := Encode(c.mnem, c.ops...); err == nil {
t.Errorf("%s: Encode succeeded, want rejection", c.name)
}
}
}
func TestImul(t *testing.T) {
checkSyntax(t, "imul rdx, rcx", "IMULQ", CX, DX)
checkSyntax(t, "imul edx, edx, 0x3", "IMULL", Imm(3), DX, DX)
@@ -277,6 +327,12 @@ func TestSSEMoveGroundTruth(t *testing.T) {
{"MOVSD (SI),X1", "MOVSD", []Operand{Ptr(SI, 0, 8), vreg(t, "X1")}, "f20f100e", "MOVSD_XMM"},
{"MOVSD X1,X2", "MOVSD", []Operand{vreg(t, "X1"), vreg(t, "X2")}, "f20f10d1", "MOVSD_XMM"},
{"MOVSS X3,(DI)", "MOVSS", []Operand{vreg(t, "X3"), Ptr(DI, 0, 4)}, "f30f111f", "MOVSS"},
// Static-symbol (SB) references: the GOROOT crypto kernels load and
// store octa constants by name (MOVOU bswapMask<>+0(SB), X0).
{"MOVOU sym,X0", "MOVOU", []Operand{sbMem{size: 16, name: "bswapMask"}, vreg(t, "X0")}, "f30f6f0500000000", "MOVDQU"},
{"MOVOU X0,sym+8", "MOVOU", []Operand{vreg(t, "X0"), sbMem{size: 16, name: "bswapMask", addend: 8}}, "f30f7f0500000000", "MOVDQU"},
{"MOVO sym,X1", "MOVO", []Operand{sbMem{size: 16, name: "gcmPoly"}, vreg(t, "X1")}, "660f6f0d00000000", "MOVDQA"},
{"MOVO X2,sym", "MOVO", []Operand{vreg(t, "X2"), sbMem{size: 16, name: "gcmPoly"}}, "660f7f1500000000", "MOVDQA"},
}
for _, c := range cases {
code, err := Encode(c.mnem, c.ops...)
@@ -546,6 +602,248 @@ func TestEncodableCmovSize(t *testing.T) {
}
}
// TestCarryShiftMulGroundTruth pins the carry-flag ALU family (ADC/SBB with
// their accumulator immediate forms), the rotate family, MUL/DIV/IDIV and the
// bit-test family byte for byte against go tool asm (see
// testdata/verify/scalar_amd64.s).
func TestCarryShiftMulGroundTruth(t *testing.T) {
cases := []struct {
name string
mnem string
ops []Operand
want string
}{
{"ADCQ AX,BX", "ADCQ", []Operand{AX, BX}, "4811c3"},
{"ADCL AX,BX", "ADCL", []Operand{AX, BX}, "11c3"},
{"ADCB AL,BL", "ADCB", []Operand{AL, BL}, "10c3"},
{"ADCW AX,BX", "ADCW", []Operand{AX, BX}, "6611c3"},
{"SBBQ AX,BX", "SBBQ", []Operand{AX, BX}, "4819c3"},
{"ADCQ $5,BX", "ADCQ", []Operand{Imm(5), BX}, "4883d305"},
{"ADCQ $300,BX", "ADCQ", []Operand{Imm(300), BX}, "4881d32c010000"},
{"ADCQ $300,AX", "ADCQ", []Operand{Imm(300), AX}, "48152c010000"},
{"ADCB $5,AL", "ADCB", []Operand{Imm(5), AL}, "1405"},
{"SBBQ $300,AX", "SBBQ", []Operand{Imm(300), AX}, "481d2c010000"},
{"ADCQ AX,(BX)", "ADCQ", []Operand{AX, Ptr(BX, 0, 8)}, "481103"},
{"ROLQ $3,AX", "ROLQ", []Operand{Imm(3), AX}, "48c1c003"},
{"ROLL CX,BX", "ROLL", []Operand{CL, BX}, "d3c3"},
{"RORQ CL,AX", "RORQ", []Operand{CL, AX}, "48d3c8"},
{"RCRQ $1,BX", "RCRQ", []Operand{Imm(1), BX}, "48d1db"},
{"RCLQ $3,AX", "RCLQ", []Operand{Imm(3), AX}, "48c1d003"},
{"RORB CL,BL", "RORB", []Operand{CL, BL}, "d2cb"},
{"SALQ $2,AX", "SALQ", []Operand{Imm(2), AX}, "48c1e002"},
{"ROLW $1,AX", "ROLW", []Operand{Imm(1), AX}, "66d1c0"},
{"MULQ CX", "MULQ", []Operand{CX}, "48f7e1"},
{"MULL CX", "MULL", []Operand{CX}, "f7e1"},
{"MULB CL", "MULB", []Operand{CL}, "f6e1"},
{"DIVL CX", "DIVL", []Operand{CX}, "f7f1"},
{"IDIVQ CX", "IDIVQ", []Operand{CX}, "48f7f9"},
{"MULW CX", "MULW", []Operand{CX}, "66f7e1"},
{"BTQ AX,DX", "BTQ", []Operand{AX, DX}, "480fa3c2"},
{"BTL AX,DX", "BTL", []Operand{AX, DX}, "0fa3c2"},
{"BTW AX,DX", "BTW", []Operand{AX, DX}, "660fa3c2"},
{"BTQ $3,BX", "BTQ", []Operand{Imm(3), BX}, "480fbae303"},
{"BTQ $3,(AX)", "BTQ", []Operand{Imm(3), Ptr(AX, 0, 8)}, "480fba2003"},
{"BTSQ $5,BX", "BTSQ", []Operand{Imm(5), BX}, "480fbaeb05"},
{"BTCQ AX,BX", "BTCQ", []Operand{AX, BX}, "480fbbc3"},
{"BTRQ $7,BX", "BTRQ", []Operand{Imm(7), BX}, "480fbaf307"},
}
for _, c := range cases {
code, err := Encode(c.mnem, c.ops...)
if err != nil {
t.Errorf("%s: Encode: %v", c.name, err)
continue
}
if got := fmt.Sprintf("%x", code); got != c.want {
t.Errorf("%s = %s, want %s", c.name, got, c.want)
}
}
// The bit-test immediate is an unsigned bit index with the negative
// spelling accepted, the shuffle convention: BTQ $300 must be rejected.
if _, err := Encode("BTQ", Imm(300), AX); err == nil {
t.Errorf("BTQ $300: expected an error, got none")
}
}
// TestAtomicSystemGroundTruth pins the exchange/compare-exchange/accumulate
// family, the string primitives, the flag and system instructions, the MXCSR
// pair, the scalar float-to-int conversions and the x87 FMOVD byte for byte
// against go tool asm (see testdata/verify/atomics_amd64.s and
// testdata/verify/system_amd64.s).
func TestAtomicSystemGroundTruth(t *testing.T) {
r8 := Reg{idx: 8, size: 8}
cases := []struct {
name string
mnem string
ops []Operand
want string
}{
{"XCHGQ AX,BX", "XCHGQ", []Operand{AX, BX}, "4893"},
{"XCHGQ BX,AX", "XCHGQ", []Operand{BX, AX}, "4893"},
{"XCHGL AX,BX", "XCHGL", []Operand{AX, BX}, "93"},
{"XCHGB AL,BL", "XCHGB", []Operand{AL, BL}, "86c3"},
{"XCHGW AX,BX", "XCHGW", []Operand{AX, BX}, "6693"},
{"XCHGQ R8,R9", "XCHGQ", []Operand{r8, Reg{idx: 9, size: 8}}, "4d87c1"},
{"XCHGQ BX,(AX)", "XCHGQ", []Operand{BX, Ptr(AX, 0, 8)}, "488718"},
{"XCHGQ (AX),BX", "XCHGQ", []Operand{Ptr(AX, 0, 8), BX}, "488718"},
{"XCHGQ AX,(BX)", "XCHGQ", []Operand{AX, Ptr(BX, 0, 8)}, "488703"},
{"CMPXCHGL AX,BX", "CMPXCHGL", []Operand{AX, BX}, "0fb1c3"},
{"CMPXCHGQ AX,(BX)", "CMPXCHGQ", []Operand{AX, Ptr(BX, 0, 8)}, "480fb103"},
{"CMPXCHGB AL,(BX)", "CMPXCHGB", []Operand{AL, Ptr(BX, 0, 1)}, "0fb003"},
{"CMPXCHGW AX,BX", "CMPXCHGW", []Operand{AX, BX}, "660fb1c3"},
{"XADDL AX,BX", "XADDL", []Operand{AX, BX}, "0fc1c3"},
{"XADDQ AX,(BX)", "XADDQ", []Operand{AX, Ptr(BX, 0, 8)}, "480fc103"},
{"XADDB AL,(BX)", "XADDB", []Operand{AL, Ptr(BX, 0, 1)}, "0fc003"},
{"XADDW AX,BX", "XADDW", []Operand{AX, BX}, "660fc1c3"},
{"ADCXL AX,CX", "ADCXL", []Operand{AX, CX}, "660f38f6c8"},
{"ADCXQ AX,CX", "ADCXQ", []Operand{AX, CX}, "66480f38f6c8"},
{"ADOXL AX,CX", "ADOXL", []Operand{AX, CX}, "f30f38f6c8"},
{"ADOXQ AX,CX", "ADOXQ", []Operand{AX, CX}, "f3480f38f6c8"},
{"CRC32B AX,CX", "CRC32B", []Operand{AX, CX}, "f20f38f0c8"},
{"CRC32W AX,CX", "CRC32W", []Operand{AX, CX}, "66f20f38f1c8"},
{"CRC32L AX,CX", "CRC32L", []Operand{AX, CX}, "f20f38f1c8"},
{"CRC32Q AX,CX", "CRC32Q", []Operand{AX, CX}, "f2480f38f1c8"},
{"CRC32L (AX),CX", "CRC32L", []Operand{Ptr(AX, 0, 4), CX}, "f20f38f108"},
{"MOVSQ", "MOVSQ", []Operand{}, "48a5"},
{"MOVSL", "MOVSL", []Operand{}, "a5"},
{"MOVSB", "MOVSB", []Operand{}, "a4"},
{"MOVSW", "MOVSW", []Operand{}, "66a5"},
{"STOSB", "STOSB", []Operand{}, "aa"},
{"STOSQ", "STOSQ", []Operand{}, "48ab"},
{"STOSL", "STOSL", []Operand{}, "ab"},
{"STOSW", "STOSW", []Operand{}, "66ab"},
{"CLD", "CLD", []Operand{}, "fc"},
{"STD", "STD", []Operand{}, "fd"},
{"POPFQ", "POPFQ", []Operand{}, "9d"},
{"PUSHFQ", "PUSHFQ", []Operand{}, "9c"},
{"CPUID", "CPUID", []Operand{}, "0fa2"},
{"RDTSC", "RDTSC", []Operand{}, "0f31"},
{"RDTSCP", "RDTSCP", []Operand{}, "0f01f9"},
{"SYSCALL", "SYSCALL", []Operand{}, "0f05"},
{"XGETBV", "XGETBV", []Operand{}, "0f01d0"},
{"PAUSE", "PAUSE", []Operand{}, "f390"},
{"LFENCE", "LFENCE", []Operand{}, "0faee8"},
{"MFENCE", "MFENCE", []Operand{}, "0faef0"},
{"SFENCE", "SFENCE", []Operand{}, "0faef8"},
{"UNDEF", "UNDEF", []Operand{}, "0f0b"},
{"INT $3", "INT", []Operand{Imm(3)}, "cd03"},
{"LDMXCSR (AX)", "LDMXCSR", []Operand{Ptr(AX, 0, 4)}, "0fae10"},
{"STMXCSR (AX)", "STMXCSR", []Operand{Ptr(AX, 0, 4)}, "0fae18"},
{"CVTSD2SL X0,AX", "CVTSD2SL", []Operand{vreg(t, "X0"), AX}, "f20f2dc0"},
{"CVTTSD2SQ X0,AX", "CVTTSD2SQ", []Operand{vreg(t, "X0"), AX}, "f2480f2cc0"},
{"CVTTSD2SL X0,AX", "CVTTSD2SL", []Operand{vreg(t, "X0"), AX}, "f20f2cc0"},
{"CVTSS2SQ X0,AX", "CVTSS2SQ", []Operand{vreg(t, "X0"), AX}, "f3480f2dc0"},
{"FMOVD (AX),F0", "FMOVD", []Operand{Ptr(AX, 0, 8), vreg(t, "F0")}, "dd00"},
{"FMOVD F0,(AX)", "FMOVD", []Operand{vreg(t, "F0"), Ptr(AX, 0, 8)}, "dd10"},
{"FMOVD F0,F1", "FMOVD", []Operand{vreg(t, "F0"), vreg(t, "F1")}, "ddd1"},
{"MOVD AX,X0", "MOVD", []Operand{AX, vreg(t, "X0")}, "66480f6ec0"},
{"MOVD X0,AX", "MOVD", []Operand{vreg(t, "X0"), AX}, "66480f7ec0"},
{"MOVD X0,X1", "MOVD", []Operand{vreg(t, "X0"), vreg(t, "X1")}, "f30f7ec8"},
{"MOVD (AX),X0", "MOVD", []Operand{Ptr(AX, 0, 8), vreg(t, "X0")}, "f30f7e00"},
{"MOVD X0,(AX)", "MOVD", []Operand{vreg(t, "X0"), Ptr(AX, 0, 8)}, "660fd600"},
}
for _, c := range cases {
code, err := Encode(c.mnem, c.ops...)
if err != nil {
t.Errorf("%s: Encode: %v", c.name, err)
continue
}
if got := fmt.Sprintf("%x", code); got != c.want {
t.Errorf("%s = %s, want %s", c.name, got, c.want)
}
}
// LDMXCSR/STMXCSR take a memory operand only.
if _, err := Encode("LDMXCSR", AX); err == nil {
t.Errorf("LDMXCSR AX: expected an error, got none")
}
}
// TestSSEGapsGroundTruth pins the legacy SSE gap families: the scalar
// compare and square root, the Plan 9 packed spellings, the imm8-controlled
// shuffles, the lane extracts and inserts, the packed integer shifts and the
// AES/SHA round instructions, byte for byte against go tool asm (see
// testdata/verify/crypto_amd64.s and testdata/verify/sse_amd64.s).
func TestSSEGapsGroundTruth(t *testing.T) {
cases := []struct {
name string
mnem string
ops []Operand
want string
}{
{"ANDNPD X0,X1", "ANDNPD", []Operand{vreg(t, "X0"), vreg(t, "X1")}, "660f55c8"},
{"ANDNPS X0,X1", "ANDNPS", []Operand{vreg(t, "X0"), vreg(t, "X1")}, "0f55c8"},
{"COMISD X0,X1", "COMISD", []Operand{vreg(t, "X0"), vreg(t, "X1")}, "660f2fc8"},
{"SQRTSD X0,X1", "SQRTSD", []Operand{vreg(t, "X0"), vreg(t, "X1")}, "f20f51c8"},
{"PSHUFL $3,X0,X1", "PSHUFL", []Operand{Imm(3), vreg(t, "X0"), vreg(t, "X1")}, "660f70c803"},
{"PALIGNR $2,X0,X1", "PALIGNR", []Operand{Imm(2), vreg(t, "X0"), vreg(t, "X1")}, "660f3a0fc802"},
{"PBLENDW $3,X0,X1", "PBLENDW", []Operand{Imm(3), vreg(t, "X0"), vreg(t, "X1")}, "660f3a0ec803"},
{"PCMPESTRI $1,X0,X1", "PCMPESTRI", []Operand{Imm(1), vreg(t, "X0"), vreg(t, "X1")}, "660f3a61c801"},
{"PCLMULQDQ $0,X0,X1", "PCLMULQDQ", []Operand{Imm(0), vreg(t, "X0"), vreg(t, "X1")}, "660f3a44c800"},
{"PCLMULQDQ $0,(AX),X1", "PCLMULQDQ", []Operand{Imm(0), Ptr(AX, 0, 16), vreg(t, "X1")}, "660f3a440800"},
{"PEXTRB $1,X0,AX", "PEXTRB", []Operand{Imm(1), vreg(t, "X0"), AX}, "660f3a14c001"},
{"PEXTRD $1,X0,AX", "PEXTRD", []Operand{Imm(1), vreg(t, "X0"), AX}, "660f3a16c001"},
{"PEXTRQ $1,X0,AX", "PEXTRQ", []Operand{Imm(1), vreg(t, "X0"), AX}, "66480f3a16c001"},
{"PEXTRW $1,X0,AX", "PEXTRW", []Operand{Imm(1), vreg(t, "X0"), AX}, "660fc5c001"},
{"PEXTRW $1,X0,(AX)", "PEXTRW", []Operand{Imm(1), vreg(t, "X0"), Ptr(AX, 0, 2)}, "660f3a150001"},
{"PINSRB $1,AX,X0", "PINSRB", []Operand{Imm(1), AX, vreg(t, "X0")}, "660f3a20c001"},
{"PINSRD $1,AX,X0", "PINSRD", []Operand{Imm(1), AX, vreg(t, "X0")}, "660f3a22c001"},
{"PINSRQ $1,AX,X0", "PINSRQ", []Operand{Imm(1), AX, vreg(t, "X0")}, "66480f3a22c001"},
{"PINSRW $1,AX,X0", "PINSRW", []Operand{Imm(1), AX, vreg(t, "X0")}, "660fc4c001"},
{"PINSRW $1,(AX),X0", "PINSRW", []Operand{Imm(1), Ptr(AX, 0, 2), vreg(t, "X0")}, "660fc40001"},
{"PSLLL $2,X0", "PSLLL", []Operand{Imm(2), vreg(t, "X0")}, "660f72f002"},
{"PSRAL $2,X0", "PSRAL", []Operand{Imm(2), vreg(t, "X0")}, "660f72e002"},
{"PSRLL $2,X0", "PSRLL", []Operand{Imm(2), vreg(t, "X0")}, "660f72d002"},
{"PSRLQ $2,X0", "PSRLQ", []Operand{Imm(2), vreg(t, "X0")}, "660f73d002"},
{"PSLLQ $2,X0", "PSLLQ", []Operand{Imm(2), vreg(t, "X0")}, "660f73f002"},
{"PSLLW $2,X0", "PSLLW", []Operand{Imm(2), vreg(t, "X0")}, "660f71f002"},
{"PSRLW $2,X0", "PSRLW", []Operand{Imm(2), vreg(t, "X0")}, "660f71d002"},
{"PSRAW $2,X0", "PSRAW", []Operand{Imm(2), vreg(t, "X0")}, "660f71e002"},
{"PSLLDQ $2,X0", "PSLLDQ", []Operand{Imm(2), vreg(t, "X0")}, "660f73f802"},
{"PSRLDQ $2,X0", "PSRLDQ", []Operand{Imm(2), vreg(t, "X0")}, "660f73d802"},
{"PSLLL X0,X1", "PSLLL", []Operand{vreg(t, "X0"), vreg(t, "X1")}, "660ff2c8"},
{"PSRLQ X0,X1", "PSRLQ", []Operand{vreg(t, "X0"), vreg(t, "X1")}, "660fd3c8"},
{"PSLLL (AX),X1", "PSLLL", []Operand{Ptr(AX, 0, 16), vreg(t, "X1")}, "660ff208"},
{"PSUBL X0,X1", "PSUBL", []Operand{vreg(t, "X0"), vreg(t, "X1")}, "660ffac8"},
{"PADDL X0,X1", "PADDL", []Operand{vreg(t, "X0"), vreg(t, "X1")}, "660ffec8"},
{"PCMPEQL X0,X1", "PCMPEQL", []Operand{vreg(t, "X0"), vreg(t, "X1")}, "660f76c8"},
{"PUNPCKLBW X0,X1", "PUNPCKLBW", []Operand{vreg(t, "X0"), vreg(t, "X1")}, "660f60c8"},
{"MOVOA X0,X1", "MOVOA", []Operand{vreg(t, "X0"), vreg(t, "X1")}, "660f6fc8"},
{"MOVOA (AX),X1", "MOVOA", []Operand{Ptr(AX, 0, 16), vreg(t, "X1")}, "660f6f08"},
{"MOVOA X0,(AX)", "MOVOA", []Operand{vreg(t, "X0"), Ptr(AX, 0, 16)}, "660f7f00"},
{"AESIMC X0,X1", "AESIMC", []Operand{vreg(t, "X0"), vreg(t, "X1")}, "660f38dbc8"},
{"AESIMC (AX),X1", "AESIMC", []Operand{Ptr(AX, 0, 16), vreg(t, "X1")}, "660f38db08"},
{"AESENC X0,X1", "AESENC", []Operand{vreg(t, "X0"), vreg(t, "X1")}, "660f38dcc8"},
{"AESENCLAST X0,X1", "AESENCLAST", []Operand{vreg(t, "X0"), vreg(t, "X1")}, "660f38ddc8"},
{"AESDEC X0,X1", "AESDEC", []Operand{vreg(t, "X0"), vreg(t, "X1")}, "660f38dec8"},
{"AESDECLAST X0,X1", "AESDECLAST", []Operand{vreg(t, "X0"), vreg(t, "X1")}, "660f38dfc8"},
{"AESKEYGENASSIST $0,X0,X1", "AESKEYGENASSIST", []Operand{Imm(0), vreg(t, "X0"), vreg(t, "X1")}, "660f3adfc800"},
{"SHA1MSG1 X0,X1", "SHA1MSG1", []Operand{vreg(t, "X0"), vreg(t, "X1")}, "0f38c9c8"},
{"SHA1MSG2 X0,X1", "SHA1MSG2", []Operand{vreg(t, "X0"), vreg(t, "X1")}, "0f38cac8"},
{"SHA1NEXTE X0,X1", "SHA1NEXTE", []Operand{vreg(t, "X0"), vreg(t, "X1")}, "0f38c8c8"},
{"SHA1RNDS4 $0,X0,X1", "SHA1RNDS4", []Operand{Imm(0), vreg(t, "X0"), vreg(t, "X1")}, "0f3accc800"},
{"SHA256MSG1 X0,X1", "SHA256MSG1", []Operand{vreg(t, "X0"), vreg(t, "X1")}, "0f38ccc8"},
{"SHA256MSG2 X0,X1", "SHA256MSG2", []Operand{vreg(t, "X0"), vreg(t, "X1")}, "0f38cdc8"},
{"SHA256RNDS2 X0,X1,X2", "SHA256RNDS2", []Operand{vreg(t, "X0"), vreg(t, "X1"), vreg(t, "X2")}, "0f38cbd1"},
}
for _, c := range cases {
code, err := Encode(c.mnem, c.ops...)
if err != nil {
t.Errorf("%s: Encode: %v", c.name, err)
continue
}
if got := fmt.Sprintf("%x", code); got != c.want {
t.Errorf("%s = %s, want %s", c.name, got, c.want)
}
}
// SHA256RNDS2's first operand must be the literal X0.
if _, err := Encode("SHA256RNDS2", vreg(t, "X1"), vreg(t, "X2"), vreg(t, "X3")); err == nil {
t.Errorf("SHA256RNDS2 X1,...: expected an error, got none")
}
// PSLLDQ has no variable-count form.
if _, err := Encode("PSLLDQ", vreg(t, "X0"), vreg(t, "X1")); err == nil {
t.Errorf("PSLLDQ X0,X1: expected an error, got none")
}
}
// TestSSEBinGroundTruth checks the legacy packed/scalar binary family
// byte for byte (no prefix / 66 / F2 / F3 variants).
func TestSSEBinGroundTruth(t *testing.T) {
@@ -627,3 +925,142 @@ func TestMOVQXMMGroundTruth(t *testing.T) {
}
}
}
// TestPrefixStatements pins LOCK, REP and REPN. go tool asm encodes each as
// a standalone one-byte instruction with a PC of its own (F0, F3, F2), not a
// prefix field merged into the following instruction, and it validates
// nothing about the pairing (LOCK before NOP assembles). The prefixed
// atomic and string shapes are the bytes the runtime's own kernels need.
func TestPrefixStatements(t *testing.T) {
cases := []struct {
name string
mnem string
ops []Operand
want string
}{
{"LOCK", "LOCK", nil, "f0"},
{"REP", "REP", nil, "f3"},
{"REPN", "REPN", nil, "f2"},
// LOCK; CMPXCHGQ AX, (BX)
{"LOCK CMPXCHGQ", "CMPXCHGQ", []Operand{AX, Ptr(BX, 0, 8)}, "480fb103"},
// REP; MOVSQ
{"REP MOVSQ", "MOVSQ", nil, "48a5"},
// REPN; MOVSB
{"REPN MOVSB", "MOVSB", nil, "a4"},
}
for _, c := range cases {
code, err := Encode(c.mnem, c.ops...)
if err != nil {
t.Errorf("%s: %v", c.name, err)
continue
}
if got := fmt.Sprintf("%x", code); got != c.want {
t.Errorf("%s = %s, want %s", c.name, got, c.want)
}
}
// The prefix statements take no operands, as the toolchain reports for
// LOCK AX.
if _, err := Encode("LOCK", AX); err == nil {
t.Error("LOCK AX assembled, want an error")
}
if _, err := Encode("REP", Imm(1)); err == nil {
t.Error("REP $1 assembled, want an error")
}
}
// TestDataEmission pins BYTE, WORD, LONG and QUAD: the immediate lands in
// the text stream as 1, 2, 4 or 8 little-endian bytes with no opcode
// lookup, truncated to the width rather than range-checked (go tool asm
// emits FF for BYTE $0x1FF and 45 23 for WORD $0x12345, both silently).
func TestDataEmission(t *testing.T) {
cases := []struct {
name string
mnem string
imm Imm
want string
}{
{"BYTE", "BYTE", 0x0f, "0f"},
{"BYTE negative", "BYTE", -1, "ff"},
{"BYTE truncated", "BYTE", 0x1ff, "ff"},
{"WORD", "WORD", 0x1234, "3412"},
{"WORD negative", "WORD", -1, "ffff"},
{"WORD truncated", "WORD", 0x12345, "4523"},
{"LONG", "LONG", 0x11223344, "44332211"},
{"LONG negative", "LONG", -1, "ffffffff"},
{"QUAD", "QUAD", 0x1122334455667788, "8877665544332211"},
{"QUAD negative", "QUAD", -2, "feffffffffffffff"},
}
for _, c := range cases {
code, err := Encode(c.mnem, c.imm)
if err != nil {
t.Errorf("%s: %v", c.name, err)
continue
}
if got := fmt.Sprintf("%x", code); got != c.want {
t.Errorf("%s = %s, want %s", c.name, got, c.want)
}
}
// Exactly one immediate: the toolchain rejects BYTE $1, $2, $3, and a
// register or a missing operand is no immediate at all.
if _, err := Encode("BYTE"); err == nil {
t.Error("BYTE with no operand assembled, want an error")
}
if _, err := Encode("BYTE", Imm(1), Imm(2)); err == nil {
t.Error("BYTE $1, $2 assembled, want an error")
}
if _, err := Encode("WORD", AX); err == nil {
t.Error("WORD AX assembled, want an error")
}
}
// TestEndIgnored pins END: go tool asm drops the statement entirely, so it
// encodes to zero bytes and takes any operands without complaint (the
// toolchain accepts END $0 and END AX alike).
func TestEndIgnored(t *testing.T) {
for _, ops := range [][]Operand{nil, {Imm(0)}, {AX}} {
code, err := Encode("END", ops...)
if err != nil {
t.Errorf("END: %v", err)
continue
}
if len(code) != 0 {
t.Errorf("END = %x, want no bytes", code)
}
}
}
// TestAdjsp pins ADJSP: a positive immediate is SUBQ $imm, SP, a negative
// one ADDQ $-imm, SP, in the imm8 or imm32 form the magnitude picks; $0
// has no encoding (go tool asm refuses ADJSP $0 outright).
func TestAdjsp(t *testing.T) {
cases := []struct {
name string
imm Imm
want string
}{
{"imm8", 112, "4883ec70"},
{"imm8 negative", -112, "4883c470"},
{"imm32", 200, "4881ecc8000000"},
{"imm32 negative", -200, "4881c4c8000000"},
{"small", 8, "4883ec08"},
}
for _, c := range cases {
code, err := Encode("ADJSP", c.imm)
if err != nil {
t.Errorf("%s: %v", c.name, err)
continue
}
if got := fmt.Sprintf("%x", code); got != c.want {
t.Errorf("ADJSP %d = %s, want %s", int64(c.imm), got, c.want)
}
}
if _, err := Encode("ADJSP", Imm(0)); err == nil {
t.Error("ADJSP $0 assembled, want an error")
}
if _, err := Encode("ADJSP"); err == nil {
t.Error("ADJSP with no operand assembled, want an error")
}
if _, err := Encode("ADJSP", AX); err == nil {
t.Error("ADJSP AX assembled, want an error")
}
}
+531 -40
View File
@@ -5,6 +5,7 @@ package asm
import (
"fmt"
"slices"
"strings"
)
@@ -91,7 +92,7 @@ var evexTable = map[string]evexSpec{
"VPSRAD": {1, 0x72, 0, 1, 4, vexShiftImm, [3]int{16, 32, 64}},
// EVEX.128/256/512.66.0F.W1, variable shift with an XMM count (VPSRAQ;
// the W bit distinguishes it from VPSRAD's E2 form).
"VPSRAQ": {1, 0xE2, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSRAQ": {1, 0x72, 1, 1, 4, vexShiftImm, [3]int{16, 32, 64}},
// EVEX.128/256/512.F3.0F.W1, signed qword to packed double (reg=dst,
// rm=src, no vvvv).
@@ -138,7 +139,7 @@ var evexTable = map[string]evexSpec{
// EVEX.66.0F, the EVEX forms of the VEX two-source shuffle.
"VSHUFPD": {1, 0xC6, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VSHUFPS": {1, 0xC6, 0, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VSHUFPS": {1, 0xC6, 0, 0, -1, vexNDS3Imm, [3]int{16, 32, 64}},
// EVEX.66.0F3A, lane insert ($imm, xsrc, zsrc1, zdst).
"VINSERTF32X4": {3, 0x18, 0, 1, -1, vexNDS3Imm, [3]int{0, 16, 32}},
@@ -180,10 +181,19 @@ var evexTable = map[string]evexSpec{
"VPCMPUQ": {3, 0x1E, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
// EVEX.66.0F38, permutes (NDS form).
"VPERMB": {2, 0x8D, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMW": {2, 0x8D, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMI2D": {2, 0x76, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMI2Q": {2, 0x76, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMB": {2, 0x8D, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMW": {2, 0x8D, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMI2B": {2, 0x75, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMI2D": {2, 0x76, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMI2Q": {2, 0x76, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
// EVEX.66.0F38, population count (reg=dst, rm=src; W selects byte/word
// against dword/qword).
"VPOPCNTB": {2, 0x54, 0, 1, -1, vexRM, [3]int{16, 32, 64}},
"VPOPCNTD": {2, 0x55, 0, 1, -1, vexRM, [3]int{16, 32, 64}},
"VPOPCNTQ": {2, 0x55, 1, 1, -1, vexRM, [3]int{16, 32, 64}},
// EVEX.66.0F.W1, the qword spelling of the packed OR (VPORQ has no VEX
// form in the Go assembler: it always encodes through EVEX).
"VPORQ": {1, 0xEB, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMT2D": {2, 0x7E, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMT2Q": {2, 0x7E, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMT2PD": {2, 0x7F, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
@@ -193,7 +203,7 @@ var evexTable = map[string]evexSpec{
"VPMULHUW": {1, 0xE4, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPMADDUBSW": {2, 0x04, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSLLVW": {2, 0x12, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSRLVW": {2, 0x11, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSRLVW": {2, 0x10, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPACKSSWB": {1, 0x63, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPACKUSWB": {1, 0x67, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPACKSSDW": {1, 0x6B, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
@@ -315,7 +325,7 @@ var evexTable = map[string]evexSpec{
"VCVTPD2UQQ": {1, 0x79, 1, 1, -1, vexRM, [3]int{16, 32, 64}},
"VCVTPS2QQ": {1, 0x7B, 0, 1, -1, vexRM, [3]int{8, 16, 32}},
"VCVTUDQ2PD": {1, 0x7A, 0, 2, -1, vexRM, [3]int{8, 16, 32}},
"VCVTUDQ2PS": {1, 0x7A, 0, 0, -1, vexRM, [3]int{8, 16, 32}},
"VCVTUDQ2PS": {1, 0x7A, 0, 3, -1, vexRM, [3]int{8, 16, 32}},
// EVEX.66.0F38, half-precision convert (half-width source).
"VCVTPH2PS": {2, 0x13, 0, 1, -1, vexRM, [3]int{8, 16, 32}},
// EVEX.66.0F3A, half-precision convert back ($imm, src, dst: reg=src,
@@ -485,6 +495,246 @@ var evexTable = map[string]evexSpec{
// destination (VPMOVDW dword→word, VPMOVQD qword→dword).
"VPMOVDW": {2, 0x33, 0, 2, -1, vexRMRev, [3]int{8, 16, 32}},
"VPMOVQD": {2, 0x35, 0, 2, -1, vexRMRev, [3]int{8, 16, 32}},
// --- the AVX-512 families the avx512enc corpus exercises, read off
// the toolchain opcodetables ---
"VAESDEC": {2, 0xDE, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VAESDECLAST": {2, 0xDF, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VAESENC": {2, 0xDC, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VAESENCLAST": {2, 0xDD, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VALIGNQ": {3, 0x03, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VANDNPD": {1, 0x55, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VANDPD": {1, 0x54, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VBLENDMPD": {2, 0x65, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VBLENDMPS": {2, 0x65, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VBROADCASTF32X2": {2, 0x19, 0, 1, -1, vexRM, [3]int{0, 8, 8}},
"VBROADCASTF32X4": {2, 0x1A, 0, 1, -1, vexRM, [3]int{0, 16, 16}},
"VBROADCASTF32X8": {2, 0x1B, 0, 1, -1, vexRM, [3]int{0, 0, 32}},
"VBROADCASTF64X2": {2, 0x1A, 1, 1, -1, vexRM, [3]int{0, 16, 16}},
"VBROADCASTF64X4": {2, 0x1B, 1, 1, -1, vexRM, [3]int{0, 0, 32}},
"VBROADCASTI32X2": {2, 0x59, 0, 1, -1, vexRM, [3]int{8, 8, 8}},
"VBROADCASTI32X4": {2, 0x5A, 0, 1, -1, vexRM, [3]int{0, 16, 16}},
"VBROADCASTI32X8": {2, 0x5B, 0, 1, -1, vexRM, [3]int{0, 0, 32}},
"VBROADCASTI64X2": {2, 0x5A, 1, 1, -1, vexRM, [3]int{0, 16, 16}},
"VBROADCASTI64X4": {2, 0x5B, 1, 1, -1, vexRM, [3]int{0, 0, 32}},
"VCOMISD": {1, 0x2F, 1, 1, -1, vexRM, [3]int{8, 0, 0}},
"VCVTSD2SS": {1, 0x5A, 1, 3, -1, vexNDS3, [3]int{8, 0, 0}},
"VCVTSS2SD": {1, 0x5A, 0, 2, -1, vexNDS3, [3]int{4, 0, 0}},
"VDBPSADBW": {3, 0x42, 0, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VEXP2PD": {2, 0xC8, 1, 1, -1, vexRM, [3]int{0, 0, 64}},
"VEXP2PS": {2, 0xC8, 0, 1, -1, vexRM, [3]int{0, 0, 64}},
"VFMADD132PD": {2, 0x98, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMADD132PS": {2, 0x98, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMADD132SD": {2, 0x99, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFMADD132SS": {2, 0x99, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFMADD213PD": {2, 0xA8, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMADD213PS": {2, 0xA8, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMADD213SD": {2, 0xA9, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFMADD213SS": {2, 0xA9, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFMADD231PS": {2, 0xB8, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMADD231SD": {2, 0xB9, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFMADD231SS": {2, 0xB9, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFMADDSUB132PD": {2, 0x96, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMADDSUB132PS": {2, 0x96, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMADDSUB213PD": {2, 0xA6, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMADDSUB213PS": {2, 0xA6, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMADDSUB231PD": {2, 0xB6, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMADDSUB231PS": {2, 0xB6, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUB132PD": {2, 0x9A, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUB132PS": {2, 0x9A, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUB132SD": {2, 0x9B, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFMSUB132SS": {2, 0x9B, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFMSUB213PD": {2, 0xAA, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUB213PS": {2, 0xAA, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUB213SD": {2, 0xAB, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFMSUB213SS": {2, 0xAB, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFMSUB231PD": {2, 0xBA, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUB231PS": {2, 0xBA, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUB231SD": {2, 0xBB, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFMSUB231SS": {2, 0xBB, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFMSUBADD132PD": {2, 0x97, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUBADD132PS": {2, 0x97, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUBADD213PD": {2, 0xA7, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUBADD213PS": {2, 0xA7, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUBADD231PD": {2, 0xB7, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUBADD231PS": {2, 0xB7, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMADD132PD": {2, 0x9C, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMADD132PS": {2, 0x9C, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMADD132SD": {2, 0x9D, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFNMADD132SS": {2, 0x9D, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFNMADD213PD": {2, 0xAC, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMADD213PS": {2, 0xAC, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMADD213SD": {2, 0xAD, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFNMADD213SS": {2, 0xAD, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFNMADD231PD": {2, 0xBC, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMADD231PS": {2, 0xBC, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMADD231SD": {2, 0xBD, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFNMADD231SS": {2, 0xBD, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFNMSUB132PD": {2, 0x9E, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMSUB132PS": {2, 0x9E, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMSUB132SD": {2, 0x9F, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFNMSUB132SS": {2, 0x9F, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFNMSUB213PD": {2, 0xAE, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMSUB213PS": {2, 0xAE, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMSUB213SD": {2, 0xAF, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFNMSUB213SS": {2, 0xAF, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFNMSUB231PD": {2, 0xBE, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMSUB231PS": {2, 0xBE, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMSUB231SD": {2, 0xBF, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFNMSUB231SS": {2, 0xBF, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VGF2P8AFFINEINVQB": {3, 0xCF, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VGF2P8AFFINEQB": {3, 0xCE, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VGF2P8MULB": {2, 0xCF, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VMOVNTDQ": {1, 0xE7, 0, 1, -1, vexRMRev, [3]int{16, 32, 64}},
"VMOVNTDQA": {2, 0x2A, 0, 1, -1, vexRM, [3]int{16, 32, 64}},
"VMOVNTPD": {1, 0x2B, 1, 1, -1, vexRMRev, [3]int{16, 32, 64}},
"VORPD": {1, 0x56, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPADDSB": {1, 0xEC, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPADDSW": {1, 0xED, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPADDUSB": {1, 0xDC, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPADDUSW": {1, 0xDD, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPBLENDMB": {2, 0x66, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPBLENDMD": {2, 0x64, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPBLENDMQ": {2, 0x64, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPBLENDMW": {2, 0x66, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPBROADCASTMB2Q": {2, 0x2A, 1, 2, -1, vexRM, [3]int{0, 0, 0}},
"VPBROADCASTMW2D": {2, 0x3A, 0, 2, -1, vexRM, [3]int{0, 0, 0}},
"VPCLMULQDQ": {3, 0x44, 0, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VPCMPEQB": {1, 0x74, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPCMPEQQ": {2, 0x29, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPCMPEQW": {1, 0x75, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPCMPGTB": {1, 0x64, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPCMPGTD": {1, 0x66, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPCMPGTQ": {2, 0x37, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPCMPGTW": {1, 0x65, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPCOMPRESSB": {2, 0x63, 0, 1, -1, vexRMRev, [3]int{1, 1, 1}},
"VPCOMPRESSW": {2, 0x63, 1, 1, -1, vexRMRev, [3]int{2, 2, 2}},
"VPCONFLICTD": {2, 0xC4, 0, 1, -1, vexRM, [3]int{16, 32, 64}},
"VPCONFLICTQ": {2, 0xC4, 1, 1, -1, vexRM, [3]int{16, 32, 64}},
"VPDPBUSD": {2, 0x50, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPDPBUSDS": {2, 0x51, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPDPWSSD": {2, 0x52, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPDPWSSDS": {2, 0x53, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMI2PD": {2, 0x77, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMI2PS": {2, 0x77, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMI2W": {2, 0x75, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMPS": {2, 0x16, 0, 1, -1, vexNDS3, [3]int{0, 32, 64}},
"VPERMT2B": {2, 0x7D, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMT2PS": {2, 0x7F, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMT2W": {2, 0x7D, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPEXPANDB": {2, 0x62, 0, 1, -1, vexRM, [3]int{1, 1, 1}},
"VPEXPANDW": {2, 0x62, 1, 1, -1, vexRM, [3]int{2, 2, 2}},
"VPINSRD": {3, 0x22, 0, 1, -1, vexNDS3Imm, [3]int{4, 0, 0}},
"VPINSRQ": {3, 0x22, 1, 1, -1, vexNDS3Imm, [3]int{8, 0, 0}},
"VPLZCNTD": {2, 0x44, 0, 1, -1, vexRM, [3]int{16, 32, 64}},
"VPLZCNTQ": {2, 0x44, 1, 1, -1, vexRM, [3]int{16, 32, 64}},
"VPMADD52HUQ": {2, 0xB5, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPMADD52LUQ": {2, 0xB4, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPMULDQ": {2, 0x28, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPMULHRSW": {2, 0x0B, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPMULHW": {1, 0xE5, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPMULTISHIFTQB": {2, 0x83, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPMULUDQ": {1, 0xF4, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPOPCNTW": {2, 0x54, 1, 1, -1, vexRM, [3]int{16, 32, 64}},
"VPORD": {1, 0xEB, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPROLVD": {2, 0x15, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPROLVQ": {2, 0x15, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPRORVD": {2, 0x14, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPRORVQ": {2, 0x14, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSADBW": {1, 0xF6, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSHLDD": {3, 0x71, 0, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VPSHLDQ": {3, 0x71, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VPSHLDVD": {2, 0x71, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSHLDVQ": {2, 0x71, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSHLDVW": {2, 0x70, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSHLDW": {3, 0x70, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VPSHRDD": {3, 0x73, 0, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VPSHRDQ": {3, 0x73, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VPSHRDVD": {2, 0x73, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSHRDVQ": {2, 0x73, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSHRDVW": {2, 0x72, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSHRDW": {3, 0x72, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VPSHUFBITQMB": {2, 0x8F, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSRAVW": {2, 0x11, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSRLD": {1, 0x72, 0, 1, 2, vexShiftImm, [3]int{16, 32, 64}},
"VPSRLDQ": {1, 0x73, 0, 1, 3, vexShiftImm, [3]int{16, 32, 64}},
// EVEX.66.0F73 /7, the byte-quad shift left (the count is always an
// immediate; there is no register-count twin).
"VPSLLDQ": {1, 0x73, 0, 1, 7, vexShiftImm, [3]int{16, 32, 64}},
// EVEX.128/256/512.0F.W0, the plain-prefix (no 66) packed spellings
// whose EVEX form drops the legacy prefix entirely.
"VANDNPS": {1, 0x55, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}},
"VANDPS": {1, 0x54, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}},
"VORPS": {1, 0x56, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}},
"VXORPS": {1, 0x57, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}},
"VUNPCKLPS": {1, 0x14, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}},
"VUNPCKHPS": {1, 0x15, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}},
"VSQRTPS": {1, 0x51, 0, 0, -1, vexRM, [3]int{16, 32, 64}},
"VCOMISS": {1, 0x2F, 0, 0, -1, vexRM, [3]int{4, 0, 0}},
"VUCOMISS": {1, 0x2E, 0, 0, -1, vexRM, [3]int{4, 0, 0}},
"VMOVNTPS": {1, 0x2B, 0, 0, -1, vexRMRev, [3]int{16, 32, 64}},
"VPSUBSB": {1, 0xE8, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSUBSW": {1, 0xE9, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSUBUSB": {1, 0xD8, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSUBUSW": {1, 0xD9, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPTESTMB": {2, 0x26, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPTESTMD": {2, 0x27, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPTESTMQ": {2, 0x27, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPTESTMW": {2, 0x26, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPTESTNMB": {2, 0x26, 0, 2, -1, vexNDS3, [3]int{16, 32, 64}},
"VPTESTNMD": {2, 0x27, 0, 2, -1, vexNDS3, [3]int{16, 32, 64}},
"VPTESTNMQ": {2, 0x27, 1, 2, -1, vexNDS3, [3]int{16, 32, 64}},
"VPTESTNMW": {2, 0x26, 1, 2, -1, vexNDS3, [3]int{16, 32, 64}},
"VPUNPCKHBW": {1, 0x68, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPUNPCKHQDQ": {1, 0x6D, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPUNPCKHWD": {1, 0x69, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPUNPCKLBW": {1, 0x60, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPUNPCKLQDQ": {1, 0x6C, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPUNPCKLWD": {1, 0x61, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VRCP28PD": {2, 0xCA, 1, 1, -1, vexRM, [3]int{0, 0, 64}},
"VRCP28PS": {2, 0xCA, 0, 1, -1, vexRM, [3]int{0, 0, 64}},
"VRCP28SD": {2, 0xCB, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VRCP28SS": {2, 0xCB, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VRSQRT28PD": {2, 0xCC, 1, 1, -1, vexRM, [3]int{0, 0, 64}},
"VRSQRT28PS": {2, 0xCC, 0, 1, -1, vexRM, [3]int{0, 0, 64}},
"VRSQRT28SD": {2, 0xCD, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VRSQRT28SS": {2, 0xCD, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VSQRTPD": {1, 0x51, 1, 1, -1, vexRM, [3]int{16, 32, 64}},
"VSQRTSD": {1, 0x51, 1, 3, -1, vexNDS3, [3]int{8, 0, 0}},
"VSQRTSS": {1, 0x51, 0, 2, -1, vexNDS3, [3]int{4, 0, 0}},
"VUCOMISD": {1, 0x2E, 1, 1, -1, vexRM, [3]int{8, 0, 0}},
"VXORPD": {1, 0x57, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
// EVEX.128/256/512.0F.F3/F2.W0, word shuffles with an immediate
// ($imm, src, dst: reg = dst, rm = src, imm8). The F3/F2 prefixes
// split the high/low lane spellings.
"VPSHUFHW": {1, 0x70, 0, 2, -1, vexImmRM, [3]int{16, 32, 64}},
"VPSHUFLW": {1, 0x70, 0, 3, -1, vexImmRM, [3]int{16, 32, 64}},
// EVEX.128.66.0F3A, lane extract to a general-purpose register or
// memory ($imm, xsrc, GPR/mem dst: reg = source, rm = destination).
"VPEXTRB": {3, 0x14, 0, 1, -1, vexExtractGPR, [3]int{1, 1, 1}},
"VPEXTRW": {3, 0x15, 0, 1, -1, vexExtractGPR, [3]int{2, 2, 2}},
"VPEXTRD": {3, 0x16, 0, 1, -1, vexExtractGPR, [3]int{4, 4, 4}},
"VPEXTRQ": {3, 0x16, 1, 1, -1, vexExtractGPR, [3]int{8, 8, 8}},
// EVEX.66.0F3A.W1, the qword permutes with an immediate control
// ($imm, src, dst: reg = dst, rm = src, imm8); the register-count
// forms live in evexRegFormTable.
"VPERMQ": {3, 0x00, 1, 1, -1, vexImmRM, [3]int{16, 32, 64}},
"VPERMPD": {3, 0x01, 1, 1, -1, vexImmRM, [3]int{16, 32, 64}},
// EVEX.66.0F3A, the packed permute shuffles with an immediate control.
"VPERMILPS": {3, 0x04, 0, 1, -1, vexImmRM, [3]int{16, 32, 64}},
"VPERMILPD": {3, 0x05, 1, 1, -1, vexImmRM, [3]int{16, 32, 64}},
// EVEX.128.0F.W0, high/low half moves. VMOVHPS carries the
// three-operand insert form (rm = m64 source, vvvv = preserved,
// reg = dst) and the two-operand store (reg = source, rm = m64);
// the encoder splits on the operand count. VMOVLHPS is the
// three-operand form alone.
"VMOVHPS": {1, 0x16, 0, 0, -1, vexNDS3, [3]int{8, 0, 0}},
"VMOVLHPS": {1, 0x16, 0, 0, -1, vexNDS3, [3]int{8, 0, 0}},
}
// evexBcastSpec describes an EVEX broadcast (VPBROADCASTD/Q): the opcode
@@ -520,32 +770,38 @@ type evexMoveSpec struct {
n [3]int
vecOK bool // the non-memory operand may be a vector register
xmmOnly bool // wider than XMM registers are rejected
nds3 bool // a three-operand register form exists (VMOVSD/VMOVSS)
}
// evexMoveTable maps an upper-case EVEX move mnemonic to its encoding.
var evexMoveTable = map[string]evexMoveSpec{
// EVEX.128/256/512.F3.0F.W0, unaligned integer move.
"VMOVDQU32": {1, 2, 0x6F, 0x7F, 0, [3]int{16, 32, 64}, true, false},
"VMOVDQU32": {1, 2, 0x6F, 0x7F, 0, [3]int{16, 32, 64}, true, false, false},
// EVEX.128/256/512.F3.0F.W1, unaligned qword move.
"VMOVDQU64": {1, 2, 0x6F, 0x7F, 1, [3]int{16, 32, 64}, true, false},
"VMOVDQU64": {1, 2, 0x6F, 0x7F, 1, [3]int{16, 32, 64}, true, false, false},
// EVEX.128/256/512.F2.0F.W0, unaligned byte move (byte/word moves use the
// F2 prefix, dword/qword moves F3; the element size only changes the tuple
// semantics).
"VMOVDQU8": {1, 3, 0x6F, 0x7F, 0, [3]int{16, 32, 64}, true, false},
"VMOVDQU8": {1, 3, 0x6F, 0x7F, 0, [3]int{16, 32, 64}, true, false, false},
// EVEX.128/256/512.F2.0F.W1, unaligned word move (shares the qword
// encoding).
"VMOVDQU16": {1, 3, 0x6F, 0x7F, 1, [3]int{16, 32, 64}, true, false},
"VMOVDQU16": {1, 3, 0x6F, 0x7F, 1, [3]int{16, 32, 64}, true, false, false},
// EVEX.128/256/512.66.0F.W1, unaligned packed double move.
"VMOVUPD": {1, 1, 0x10, 0x11, 1, [3]int{16, 32, 64}, true, false},
"VMOVUPD": {1, 1, 0x10, 0x11, 1, [3]int{16, 32, 64}, true, false, false},
// EVEX.128/256/512, aligned packed moves.
"VMOVAPS": {1, 0, 0x28, 0x29, 0, [3]int{16, 32, 64}, true, false},
"VMOVAPD": {1, 1, 0x28, 0x29, 1, [3]int{16, 32, 64}, true, false},
"VMOVAPS": {1, 0, 0x28, 0x29, 0, [3]int{16, 32, 64}, true, false, false},
"VMOVAPD": {1, 1, 0x28, 0x29, 1, [3]int{16, 32, 64}, true, false, false},
// EVEX.128/256/512.66.0F, aligned integer moves.
"VMOVDQA32": {1, 1, 0x6F, 0x7F, 0, [3]int{16, 32, 64}, true, false},
"VMOVDQA64": {1, 1, 0x6F, 0x7F, 1, [3]int{16, 32, 64}, true, false},
"VMOVDQA32": {1, 1, 0x6F, 0x7F, 0, [3]int{16, 32, 64}, true, false, false},
"VMOVDQA64": {1, 1, 0x6F, 0x7F, 1, [3]int{16, 32, 64}, true, false, false},
// EVEX.128.F3.0F.W0, scalar single move, memory operands (the
// three-operand register form is not supported).
"VMOVSS": {1, 2, 0x10, 0x11, 0, [3]int{4, 4, 4}, false, true},
"VMOVSS": {1, 2, 0x10, 0x11, 0, [3]int{4, 4, 4}, false, true, true},
// EVEX.128.F2.0F.W1, scalar double move: memory operands and the
// three-operand register form (VMOVSD dst, src1, src2).
"VMOVSD": {1, 3, 0x10, 0x11, 1, [3]int{8, 8, 8}, false, true, true},
// EVEX.128/256/512.0F.W0, unaligned packed single move.
"VMOVUPS": {1, 0, 0x10, 0x11, 0, [3]int{16, 32, 64}, true, false, false},
}
// isEvex reports whether the mnemonic has an EVEX encoding we handle.
@@ -570,6 +826,13 @@ func evexRequired(upper string, ops []Operand) bool {
if !inVex && !inVexMove {
return true // EVEX-only mnemonic
}
// The byte-quad shifts have VEX register forms but EVEX-only memory
// forms: a memory count source forces the EVEX encoding.
if upper == "VPSLLDQ" || upper == "VPSRLDQ" {
if slices.ContainsFunc(ops, memOperand) {
return true
}
}
for _, op := range ops {
if r, ok := op.(Reg); ok && (r.size == 64 || r.mask || (r.isVec() && r.idx >= 16)) {
return true
@@ -743,6 +1006,27 @@ func (e *enc) encodeEvex(mnemUpper string, ops []Operand, sfx evexSuffix) error
}
spec.n = [3]int{n, n, n}
}
// A mnemonic with an immediate and a register spelling (the
// variable-count shifts, the permutes) encodes the register one
// when the first operand is not an immediate.
if len(ops) > 0 {
if _, isImm := ops[0].(Imm); !isImm {
if alt, ok := evexRegFormTable[mnemUpper]; ok {
spec, inTable = alt, true
}
}
}
// The high/low half moves split by operand count: three operands
// insert, two store (VMOVHPS m64, X1).
if hs, ok := evexHptrTable[mnemUpper]; ok {
if len(ops) == 2 {
if hs.store.opcode == 0 {
return fmt.Errorf("%s has no two-operand form", mnemUpper)
}
return e.encodeEvexRMRev(hs.store, ops, 0, sfx)
}
spec = hs.insert
}
} else if sfx.evexOnly() {
return fmt.Errorf("%s: the instruction does not take rounding/SAE/broadcast suffixes", mnemUpper)
}
@@ -802,6 +1086,12 @@ func (e *enc) encodeEvex(mnemUpper string, ops []Operand, sfx evexSuffix) error
}
return e.encodeEvexMove(mnemUpper, ms, ops, mask, sfx)
}
if ps, ok := evexPrefGatherTable[mnemUpper]; ok {
if sfx.any() {
return fmt.Errorf("%s takes no EVEX suffixes", mnemUpper)
}
return e.encodeEvexPrefGather(mnemUpper, ps, ops, mask, sfx)
}
if !inTable {
return fmt.Errorf("unsupported instruction %q for ZMM/K operands", mnemUpper)
}
@@ -820,6 +1110,8 @@ func (e *enc) encodeEvex(mnemUpper string, ops []Operand, sfx evexSuffix) error
return e.encodeEvexNDS3Imm(spec, ops, mask, sfx)
case vexExtract:
return e.encodeEvexExtract(spec, ops, mask, sfx)
case vexExtractGPR:
return e.encodeEvexExtractGPR(spec, ops, mask, sfx)
case vexRMSrcLen:
return e.encodeEvexRMSrcLen(spec, ops, mask, sfx)
}
@@ -899,6 +1191,11 @@ func (e *enc) encodeEvexImmRM(spec evexSpec, ops []Operand, mask int, sfx evexSu
if dstReg.mask {
if r, ok := src.(Reg); ok && r.isVec() {
ll = r.vecLenBit()
} else if l, err := soleLen(spec.n); err == nil {
// A memory source with a length-fixed mnemonic
// (VFPCLASSPDX/Y/Z): the length comes from the table's
// single valid slot, not from the operand.
ll = l
}
} else if r, ok := src.(Reg); ok && r.isVec() {
ll = r.vecLenBit()
@@ -925,9 +1222,11 @@ func (e *enc) encodeEvexShiftImm(spec evexSpec, ops []Operand, mask int, sfx eve
if !ok {
return fmt.Errorf("shift count must be an immediate")
}
srcReg, ok := src.(Reg)
if !ok || !srcReg.isVec() {
return fmt.Errorf("shift source must be a vector register")
// The count source is a vector register or memory; the length the L'L
// field and the disp8×N multiplier follow is the destination's either
// way.
if !vecOrMem(src) {
return fmt.Errorf("shift source must be a vector register or memory")
}
dstReg, ok := dst.(Reg)
if !ok || !dstReg.isVec() {
@@ -937,7 +1236,7 @@ func (e *enc) encodeEvexShiftImm(spec evexSpec, ops []Operand, mask int, sfx eve
if err != nil {
return err
}
if err := e.emitEvexFields(spec, dstReg.vecLenBit(), spec.opdigit, dstReg.idx, srcReg, mask, sfx); err != nil {
if err := e.emitEvexFields(spec, dstReg.vecLenBit(), spec.opdigit, dstReg.idx, src, mask, sfx); err != nil {
return err
}
e.out = append(e.out, immByte)
@@ -1009,10 +1308,77 @@ func (e *enc) encodeEvexExtract(spec evexSpec, ops []Operand, mask int, sfx evex
return nil
}
// encodeEvexExtractGPR encodes the lane extract to a general-purpose
// register or memory: OP $imm, xsrc, dst (reg = the XMM source, rm = the
// destination, imm8). The encoding is 128-bit regardless of register
// numbers, so L'L is fixed at 0 and the disp8×N multiplier is the extracted
// element size the table carries.
func (e *enc) encodeEvexExtractGPR(spec evexSpec, ops []Operand, mask int, sfx evexSuffix) error {
if len(ops) != 3 {
return fmt.Errorf("extract expects 3 operands ($imm, xsrc, dst), got %d", len(ops))
}
imm, src, dst := ops[0], ops[1], ops[2]
immVal, ok := imm.(Imm)
if !ok {
return fmt.Errorf("extract lane must be an immediate")
}
srcReg, ok := src.(Reg)
if !ok || !srcReg.isVec() {
return fmt.Errorf("extract source must be a vector register")
}
switch dst.(type) {
case Reg:
if dst.(Reg).isVec() {
return fmt.Errorf("extract destination must be a general-purpose register or memory")
}
case Mem, sbMem:
default:
return fmt.Errorf("extract destination must be a general-purpose register or memory")
}
immByte, err := imm8(int64(immVal))
if err != nil {
return err
}
if err := e.emitEvexFields(spec, 0, srcReg.idx, -1, dst, mask, sfx); err != nil {
return err
}
e.out = append(e.out, immByte)
return nil
}
// encodeEvexMove encodes a two-operand EVEX move; a vector→vector move uses
// the store-form opcode (reg = source, rm = destination), matching the Go
// assembler.
// assembler. The scalar moves also carry a three-operand register form
// (VMOVSD dst, src1, src2: the load opcode with vvvv = src1), which ms.nds3
// opens.
func (e *enc) encodeEvexMove(mnem string, ms evexMoveSpec, ops []Operand, mask int, sfx evexSuffix) error {
if len(ops) == 3 {
if !ms.nds3 {
return fmt.Errorf("EVEX move expects 2 operands, got %d", len(ops))
}
// The masked scalar register form keeps the Go assembler's own
// layout: the store opcode with reg = op0, vvvv = op1 and the
// destination in r/m (op2) — the bytes go tool asm emits, not
// the manual's NDS reading.
src, src1, dst := ops[0], ops[1], ops[2]
reg, ok := src.(Reg)
if !ok || !reg.isVec() {
return fmt.Errorf("%s: first operand must be a vector register", mnem)
}
vvvvReg, ok := src1.(Reg)
if !ok || !vvvvReg.isVec() {
return fmt.Errorf("%s: second operand must be a vector register", mnem)
}
dstReg, ok := dst.(Reg)
if !ok || !dstReg.isVec() {
return fmt.Errorf("%s: destination must be a vector register", mnem)
}
if ms.xmmOnly && (reg.size != 16 || vvvvReg.size != 16 || dstReg.size != 16) {
return fmt.Errorf("%s operates on XMM registers only", mnem)
}
spec := evexSpec{mapSel: ms.mapSel, opcode: ms.store, w: ms.w, pp: ms.pp, opdigit: -1, n: ms.n}
return e.emitEvexFields(spec, dstReg.vecLenBit(), reg.idx, vvvvReg.idx, dst, mask, sfx)
}
if len(ops) != 2 {
return fmt.Errorf("EVEX move expects 2 operands, got %d", len(ops))
}
@@ -1132,12 +1498,20 @@ func (e *enc) encodeEvexBcast(bs evexBcastSpec, ops []Operand, mask int, sfx eve
return fmt.Errorf("broadcast destination must be a vector register")
}
spec := evexSpec{mapSel: bs.mapSel, w: bs.w, pp: 1, opdigit: -1}
switch src.(type) {
switch r := src.(type) {
case Mem, sbMem:
spec.opcode = bs.opMem
spec.n = [3]int{bs.n, bs.n, bs.n}
case Reg:
spec.opcode = bs.opReg
// A GPR source uses the register broadcast opcode; a vector
// source shares the xmm/mem one (the low byte is copied from
// the lane or from the memory operand).
if r.isVec() {
spec.opcode = bs.opMem
spec.n = [3]int{bs.n, bs.n, bs.n}
} else {
spec.opcode = bs.opReg
}
default:
return fmt.Errorf("broadcast source must be a register or memory")
}
@@ -1340,6 +1714,102 @@ func isScatter(upper string) bool {
return ok
}
// isEvexPrefGather reports whether the mnemonic is a gather/scatter
// prefetch hint.
func isEvexPrefGather(upper string) bool {
_, ok := evexPrefGatherTable[upper]
return ok
}
// evexRegFormTable holds the register-count twin of the immediate-form
// entries in evexTable. Several mnemonics name two encodings: an immediate
// count or control ($imm, src, dst …) and a register-count one whose second
// operand is a vector register or memory (count, src2, src1, dst). The
// immediate spelling lives in evexTable, this table carries the register
// spelling, and encodeEvex picks by whether the first operand is an
// immediate, the way vexVarShift does on the VEX side.
var evexRegFormTable = map[string]evexSpec{
"VPSLLD": {1, 0xF2, 0, 1, -1, vexNDS3, [3]int{16, 16, 16}},
"VPSLLQ": {1, 0xF3, 1, 1, -1, vexNDS3, [3]int{16, 16, 16}},
"VPSLLW": {1, 0xF1, 0, 1, -1, vexNDS3, [3]int{16, 16, 16}},
"VPSRAD": {1, 0xE2, 0, 1, -1, vexNDS3, [3]int{16, 16, 16}},
"VPSRAQ": {1, 0xE2, 1, 1, -1, vexNDS3, [3]int{16, 16, 16}},
"VPSRAW": {1, 0xE1, 0, 1, -1, vexNDS3, [3]int{16, 16, 16}},
"VPSRLD": {1, 0xD2, 0, 1, -1, vexNDS3, [3]int{16, 16, 16}},
"VPSRLQ": {1, 0xD3, 1, 1, -1, vexNDS3, [3]int{16, 16, 16}},
"VPSRLW": {1, 0xD1, 0, 1, -1, vexNDS3, [3]int{16, 16, 16}},
// EVEX.NDS.0F38.W1, the register-count permutes (the immediate
// controls live in evexTable under 0F3A).
"VPERMQ": {2, 0x36, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMPD": {2, 0x16, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
// EVEX.NDS.0F38, the register-count permil shuffles.
"VPERMILPS": {2, 0x0C, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMILPD": {2, 0x0D, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
}
// evexPrefGatherSpec describes a gather/scatter prefetch hint: one memory
// operand with a VSIB index and an opmask register, no destination. The
// ModRM.reg field carries a fixed /digit, the L'L field is fixed at 512, and
// the mask register is the instruction's only register operand.
type evexPrefGatherSpec struct {
mapSel int
opcode byte
w int
pp int
opdigit int
n int
}
var evexPrefGatherTable = map[string]evexPrefGatherSpec{
"VGATHERPF0DPD": {2, 0xC6, 1, 1, 1, 8},
"VGATHERPF0DPS": {2, 0xC6, 0, 1, 1, 4},
"VGATHERPF0QPD": {2, 0xC7, 1, 1, 1, 8},
"VGATHERPF0QPS": {2, 0xC7, 0, 1, 1, 4},
"VGATHERPF1DPD": {2, 0xC6, 1, 1, 2, 8},
"VGATHERPF1DPS": {2, 0xC6, 0, 1, 2, 4},
"VGATHERPF1QPD": {2, 0xC7, 1, 1, 2, 8},
"VGATHERPF1QPS": {2, 0xC7, 0, 1, 2, 4},
"VSCATTERPF0DPD": {2, 0xC6, 1, 1, 5, 8},
"VSCATTERPF0DPS": {2, 0xC6, 0, 1, 5, 4},
"VSCATTERPF0QPD": {2, 0xC7, 1, 1, 5, 8},
"VSCATTERPF0QPS": {2, 0xC7, 0, 1, 5, 4},
"VSCATTERPF1DPD": {2, 0xC6, 1, 1, 6, 8},
"VSCATTERPF1DPS": {2, 0xC6, 0, 1, 6, 4},
"VSCATTERPF1QPD": {2, 0xC7, 1, 1, 6, 8},
"VSCATTERPF1QPS": {2, 0xC7, 0, 1, 6, 4},
}
// evexHptrSpec describes the high/low half moves (VMOVHPS family): the
// three-operand insert shares an opcode with a two-operand store whose
// source is the vector register and whose destination is m64.
type evexHptrSpec struct {
insert evexSpec
store evexSpec // store.opcode == 0 when the mnemonic has no store form
}
var evexHptrTable = map[string]evexHptrSpec{
"VMOVHPS": {
insert: evexSpec{mapSel: 1, opcode: 0x16, w: 0, pp: 0, opdigit: -1, form: vexNDS3, n: [3]int{8, 0, 0}},
store: evexSpec{mapSel: 1, opcode: 0x17, w: 0, pp: 0, opdigit: -1, form: vexRMRev, n: [3]int{8, 0, 0}},
},
"VMOVLHPS": {
insert: evexSpec{mapSel: 1, opcode: 0x16, w: 0, pp: 0, opdigit: -1, form: vexNDS3, n: [3]int{8, 0, 0}},
},
}
// encodeEvexPrefGather encodes a gather/scatter prefetch hint: OP K, vsib.
func (e *enc) encodeEvexPrefGather(upper string, ps evexPrefGatherSpec, ops []Operand, mask int, sfx evexSuffix) error {
if len(ops) != 1 {
return fmt.Errorf("%s expects 2 operands (K, vsib memory), got %d", upper, len(ops)+1)
}
m, ok := ops[0].(Mem)
if !ok || !m.HasIndex || !m.Index.isVec() {
return fmt.Errorf("%s: operand must be a VSIB memory reference with a vector index", upper)
}
spec := evexSpec{mapSel: ps.mapSel, opcode: ps.opcode, w: ps.w, pp: ps.pp, opdigit: ps.opdigit, n: [3]int{ps.n, ps.n, ps.n}}
return e.emitEvexFields(spec, 2, ps.opdigit, -1, m, mask, sfx)
}
// vsibLen validates a VSIB memory operand (the index must be a vector
// register) and returns it with the vector length the index selects, the
// EVEX L'L field follows the index register, not the data register.
@@ -1361,7 +1831,9 @@ func (e *enc) encodeGather(upper string, gs gatherSpec, ops []Operand, sfx evexS
return err
}
if mask != 0 || sfx.any() {
// EVEX form: OP vsib, K, dst.
// EVEX form: OP vsib, K, dst. The L'L field is the wider of the
// index and the data register lengths (the Go assembler's
// layout); the disp8×N multiplier stays the index element size.
if len(rest) != 2 {
return fmt.Errorf("%s expects 3 operands (vsib, K, dst), got %d", upper, len(ops))
}
@@ -1373,6 +1845,9 @@ func (e *enc) encodeGather(upper string, gs gatherSpec, ops []Operand, sfx evexS
if !ok || !dst.isVec() {
return fmt.Errorf("%s: destination must be a vector register", upper)
}
if d := dst.vecLenBit(); d > ll {
ll = d
}
evex := evexSpec{mapSel: 2, opcode: gs.opcode, w: gs.w, pp: 1, opdigit: -1, n: [3]int{gs.n, gs.n, gs.n}}
return e.emitEvexFields(evex, ll, dst.idx, -1, vsib, mask, sfx)
}
@@ -1422,6 +1897,11 @@ func (e *enc) encodeScatter(upper string, ss gatherSpec, ops []Operand, sfx evex
if err != nil {
return err
}
// The L'L field is the wider of the data register and the VSIB index
// lengths, the bytes go tool asm emits.
if d := src.vecLenBit(); d > ll {
ll = d
}
evex := evexSpec{mapSel: 2, opcode: ss.opcode, w: ss.w, pp: 1, opdigit: -1, n: [3]int{ss.n, ss.n, ss.n}}
return e.emitEvexFields(evex, ll, src.idx, -1, vsib, mask, sfx)
}
@@ -1432,21 +1912,25 @@ func (e *enc) encodeScatter(upper string, ss gatherSpec, ops []Operand, sfx evex
var evexKOperand = map[string]bool{
"VPMOVM2B": true, "VPMOVM2W": true, "VPMOVM2D": true, "VPMOVM2Q": true,
"VPMOVB2M": true, "VPMOVW2M": true, "VPMOVD2M": true, "VPMOVQ2M": true,
// The K-to-vector broadcast reads its opmask source from r/m.
"VPBROADCASTMB2Q": true, "VPBROADCASTMW2D": true,
}
// kmovSpec describes a KMOV width: the opcode depends on the operand
// direction, kk (k/mem → K is 90, k → k uses the same), kmem (K → mem),
// gprk (GPR/mem → K), kgpr (K → GPR), and the GPR forms carry a mandatory
// prefix and W for the wider widths.
// direction, kk (k → k), kmem (k → mem), gprk (GPR/mem → k) and kgpr
// (k → GPR). Each direction group carries its own mandatory prefix and W:
// the k-destination/source forms share one pair, the GPR forms another.
type kmovSpec struct {
kk, kmem, gprk, kgpr byte
gprPP int
w int
kPP, kW int // prefix and VEX.W for the k forms
gprPP, gprW int // prefix and VEX.W for the GPR forms
}
var kmovTable = map[string]kmovSpec{
"KMOVW": {0x90, 0x91, 0x92, 0x93, 0, 0},
"KMOVQ": {0x90, 0x91, 0x92, 0x93, 3, 1},
"KMOVW": {0x90, 0x91, 0x92, 0x93, 0, 0, 0, 0},
"KMOVB": {0x90, 0x91, 0x92, 0x93, 1, 0, 1, 0},
"KMOVD": {0x90, 0x91, 0x92, 0x93, 1, 1, 3, 0},
"KMOVQ": {0x90, 0x91, 0x92, 0x93, 0, 1, 3, 1},
}
// encodeKmov encodes a KMOV width, selecting the opcode by direction.
@@ -1460,14 +1944,14 @@ func (e *enc) encodeKmov(upper string, ops []Operand) error {
dstReg, dstIsReg := dst.(Reg)
srcK := srcIsReg && srcReg.mask
dstK := dstIsReg && dstReg.mask
spec := vexSpec{mapSel: 1, w: ks.w, pp: 0, opdigit: -1}
switch {
case srcK && dstK:
spec.opcode = ks.kk // k ← k: reg = dst, rm = src
// k ← k: reg = dst, rm = src.
spec := vexSpec{mapSel: 1, opcode: ks.kk, w: ks.kW, pp: ks.kPP, opdigit: -1}
return e.emitVexFields(spec, 0, dstReg.idx&7, 0, 15, src)
case srcK && dstIsReg:
spec.opcode = ks.kgpr // GPR ← k: reg = dst, rm = src
spec.pp = ks.gprPP
// GPR ← k: reg = dst, rm = src.
spec := vexSpec{mapSel: 1, opcode: ks.kgpr, w: ks.gprW, pp: ks.gprPP, opdigit: -1}
rBit := 0
if dstReg.idx >= 8 {
rBit = 1
@@ -1477,11 +1961,17 @@ func (e *enc) encodeKmov(upper string, ops []Operand) error {
if _, ok := dst.(Mem); !ok {
return fmt.Errorf("%s: invalid destination operand", upper)
}
spec.opcode = ks.kmem // mem ← k: reg = src, rm = dst
// mem ← k: reg = src, rm = dst.
spec := vexSpec{mapSel: 1, opcode: ks.kmem, w: ks.kW, pp: ks.kPP, opdigit: -1}
return e.emitVexFields(spec, 0, srcReg.idx&7, 0, 15, dst)
case dstK:
spec.opcode = ks.gprk // k ← GPR/mem: reg = dst, rm = src
spec.pp = ks.gprPP
// k ← GPR: reg = dst, rm = src. A memory source shares the k ← k
// opcode and prefix group (the ykmovb layout the Go assembler uses).
opcode, w, pp := ks.gprk, ks.gprW, ks.gprPP
if memOperand(src) {
opcode, w, pp = ks.kk, ks.kW, ks.kPP
}
spec := vexSpec{mapSel: 1, opcode: opcode, w: w, pp: pp, opdigit: -1}
return e.emitVexFields(spec, 0, dstReg.idx&7, 0, 15, src)
}
return fmt.Errorf("%s requires a K register operand", upper)
@@ -1524,6 +2014,7 @@ var kOpsTable = map[string]kOpSpec{
"KXORD": {1, 0x47, 1, 1, 1, vexNDS3},
"KXORQ": {1, 0x47, 1, 0, 1, vexNDS3},
"KUNPCKBW": {1, 0x4B, 0, 1, 1, vexNDS3},
"KUNPCKWD": {1, 0x4B, 0, 0, 1, vexNDS3},
"KUNPCKDQ": {1, 0x4B, 1, 0, 1, vexNDS3},
"KADDB": {1, 0x4A, 0, 1, 1, vexNDS3},
"KADDW": {1, 0x4A, 0, 0, 1, vexNDS3},
+108
View File
@@ -38,6 +38,15 @@ func TestEvexGroundTruth(t *testing.T) {
{"VADDPD Z11,Z10,Z10", "VADDPD", []Operand{vreg(t, "Z11"), vreg(t, "Z10"), vreg(t, "Z10")}, "6251ad4858d3"},
{"VMULPD Z13,Z12,Z12", "VMULPD", []Operand{vreg(t, "Z13"), vreg(t, "Z12"), vreg(t, "Z12")}, "62519d4859e5"},
{"VFMADD231PD Z14,Z12,Z10", "VFMADD231PD", []Operand{vreg(t, "Z14"), vreg(t, "Z12"), vreg(t, "Z10")}, "62529d48b8d6"},
// The qword OR spelling always encodes through EVEX.
{"VPORQ Y0,Y1,Y2", "VPORQ", []Operand{vreg(t, "Y0"), vreg(t, "Y1"), vreg(t, "Y2")}, "62f1f528ebd0"},
{"VPORQ X0,X1,X2", "VPORQ", []Operand{vreg(t, "X0"), vreg(t, "X1"), vreg(t, "X2")}, "62f1f508ebd0"},
// Byte permute and population count.
{"VPERMI2B X0,X1,X2", "VPERMI2B", []Operand{vreg(t, "X0"), vreg(t, "X1"), vreg(t, "X2")}, "62f2750875d0"},
{"VPOPCNTB X0,X1", "VPOPCNTB", []Operand{vreg(t, "X0"), vreg(t, "X1")}, "62f27d0854c8"},
{"VPOPCNTD X0,X1", "VPOPCNTD", []Operand{vreg(t, "X0"), vreg(t, "X1")}, "62f27d0855c8"},
{"VPOPCNTD Y0,Y1", "VPOPCNTD", []Operand{vreg(t, "Y0"), vreg(t, "Y1")}, "62f27d2855c8"},
{"VPOPCNTQ X0,X1", "VPOPCNTQ", []Operand{vreg(t, "X0"), vreg(t, "X1")}, "62f2fd0855c8"},
// Align (NDS + imm8).
{"VALIGND $12,Z12,Z0,Z1", "VALIGND", []Operand{Imm(12), vreg(t, "Z12"), vreg(t, "Z0"), vreg(t, "Z1")}, "62d37d4803cc0c"},
{"VALIGND $15,Z9,Z0,Z1", "VALIGND", []Operand{Imm(15), vreg(t, "Z9"), vreg(t, "Z0"), vreg(t, "Z1")}, "62d37d4803c90f"},
@@ -52,6 +61,16 @@ func TestEvexGroundTruth(t *testing.T) {
{"KMOVW K1,CX", "KMOVW", []Operand{vreg(t, "K1"), CX}, "c5f893c9"},
{"KMOVW K1,R12", "KMOVW", []Operand{vreg(t, "K1"), vreg(t, "R12")}, "c57893e1"},
{"KTESTW K1,K1", "KTESTW", []Operand{vreg(t, "K1"), vreg(t, "K1")}, "c5f899c9"},
{"KMOVB K1,K2", "KMOVB", []Operand{vreg(t, "K1"), vreg(t, "K2")}, "c5f990d1"},
{"KMOVB AX,K1", "KMOVB", []Operand{AX, vreg(t, "K1")}, "c5f992c8"},
{"KMOVB K1,AX", "KMOVB", []Operand{vreg(t, "K1"), AX}, "c5f993c1"},
{"KMOVB K1,(AX)", "KMOVB", []Operand{vreg(t, "K1"), Ptr(AX, 0, 1)}, "c5f99108"},
{"KMOVD K1,K2", "KMOVD", []Operand{vreg(t, "K1"), vreg(t, "K2")}, "c4e1f990d1"},
{"KMOVD AX,K1", "KMOVD", []Operand{AX, vreg(t, "K1")}, "c5fb92c8"},
{"KMOVD K1,AX", "KMOVD", []Operand{vreg(t, "K1"), AX}, "c5fb93c1"},
{"KMOVD K1,(AX)", "KMOVD", []Operand{vreg(t, "K1"), Ptr(AX, 0, 4)}, "c4e1f99108"},
{"KMOVB (AX),K1", "KMOVB", []Operand{Ptr(AX, 0, 1), vreg(t, "K1")}, "c5f99008"},
{"KMOVQ (AX),K1", "KMOVQ", []Operand{Ptr(AX, 0, 8), vreg(t, "K1")}, "c4e1f89008"},
// Moves, incl. disp8×N (64 for a 512-bit operand).
{"VMOVDQU32 (SI)(R15*4),Z3", "VMOVDQU32", []Operand{Idx(SI, vreg(t, "R15"), 4, 0, 64), vreg(t, "Z3")}, "62b17e486f1cbe"},
{"VMOVDQU32 4(SI)(AX*1),Z4", "VMOVDQU32", []Operand{Idx(SI, AX, 1, 4, 64), vreg(t, "Z4")}, "62f17e486fa40604000000"},
@@ -702,3 +721,92 @@ func hexCompact(b []byte) string {
}
return string(out)
}
// TestAvx512CorpusFamilies pins representative encodings of the AVX-512
// families the toolchain's avx512enc corpus exercises: the bytes are the
// go tool asm output for exactly these operands, and the same families are
// covered end to end by the avx512_amd64.s differential kernel.
func TestAvx512CorpusFamilies(t *testing.T) {
vsib := func(base, idx string, scale int) Operand {
return Idx(vreg(t, base), vreg(t, idx), scale, 0, 0)
}
cases := []struct {
name string
mnem string
ops []Operand
want string
}{
// AES rounds (EVEX NDS, VEX twin routed by operand width).
{"VAESDEC Z", "VAESDEC", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "Z3")}, "62f26d48ded9"},
// Integer VNNI and the bit algorithm group.
{"VPDPBUSD", "VPDPBUSD", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "K2"), vreg(t, "Z3")}, "62f26d4a50d9"},
{"VPOPCNTW", "VPOPCNTW", []Operand{vreg(t, "Z1"), vreg(t, "K3"), vreg(t, "Z2")}, "62f2fd4b54d1"},
{"VPCONFLICTD", "VPCONFLICTD", []Operand{vreg(t, "Z1"), vreg(t, "K1"), vreg(t, "Z2")}, "62f27d49c4d1"},
{"VPLZCNTQ masked", "VPLZCNTQ", []Operand{vreg(t, "Z7"), vreg(t, "K1"), vreg(t, "Z8")}, "6272fd4944c7"},
{"VPERMT2B", "VPERMT2B", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "K1"), vreg(t, "Z3")}, "62f26d497dd9"},
{"VPMULTISHIFTQB", "VPMULTISHIFTQB", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "K3"), vreg(t, "Z4")}, "62f2ed4b83e1"},
{"VDBPSADBW", "VDBPSADBW", []Operand{Imm(3), vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "K3"), vreg(t, "Z3")}, "62f36d4b42d903"},
{"VPSHUFBITQMB", "VPSHUFBITQMB", []Operand{vreg(t, "Z9"), vreg(t, "Z10"), vreg(t, "K3")}, "62d22d488fd9"},
{"VPTESTNMQ", "VPTESTNMQ", []Operand{vreg(t, "Z13"), vreg(t, "Z14"), vreg(t, "K5")}, "62d28e4827ed"},
// Permutations: immediate and register counts.
{"VALIGNQ", "VALIGNQ", []Operand{Imm(3), vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "K1"), vreg(t, "Z3")}, "62f3ed4903d903"},
{"VPERMQ imm", "VPERMQ", []Operand{Imm(1), vreg(t, "Z1"), vreg(t, "K2"), vreg(t, "Z2")}, "62f3fd4a00d101"},
{"VPERMQ reg", "VPERMQ", []Operand{vreg(t, "Z3"), vreg(t, "Z4"), vreg(t, "K2"), vreg(t, "Z5")}, "62f2dd4a36eb"},
{"VPERMPD reg", "VPERMPD", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "Z3")}, "62f2ed4816d9"},
{"VPERMILPS imm", "VPERMILPS", []Operand{Imm(5), vreg(t, "Z9"), vreg(t, "K2"), vreg(t, "Z10")}, "62537d4a04d105"},
{"VPERMILPS reg", "VPERMILPS", []Operand{vreg(t, "Z11"), vreg(t, "Z12"), vreg(t, "K2"), vreg(t, "Z13")}, "62521d4a0ceb"},
// Shifts: immediate, register-count and memory-count forms; the
// count source carries its own XMM tuple width.
{"VPSLLW imm mask", "VPSLLW", []Operand{Imm(3), vreg(t, "Z1"), vreg(t, "K2"), vreg(t, "Z2")}, "62f16d4a71f103"},
{"VPSLLD reg count", "VPSLLD", []Operand{vreg(t, "X1"), vreg(t, "Z2"), vreg(t, "K1"), vreg(t, "Z3")}, "62f16d49f2d9"},
{"VPSLLDQ", "VPSLLDQ", []Operand{Imm(9), vreg(t, "Z7"), vreg(t, "Z8")}, "62f13d4873ff09"},
{"VPSRLDQ mem", "VPSRLDQ", []Operand{Imm(11), Ptr(SI, 16, 16), vreg(t, "Z4")}, "62f15d48739e100000000b"},
{"VPSRLVW", "VPSRLVW", []Operand{vreg(t, "Z3"), vreg(t, "Z4"), vreg(t, "K1"), vreg(t, "Z5")}, "62f2dd4910eb"},
// Conversions and shuffles with the F2 prefix and no prefix.
{"VCVTUDQ2PS", "VCVTUDQ2PS", []Operand{vreg(t, "Z1"), vreg(t, "K1"), vreg(t, "Z2")}, "62f17f497ad1"},
{"VSHUFPS", "VSHUFPS", []Operand{Imm(2), vreg(t, "Z4"), vreg(t, "Z5"), vreg(t, "K1"), vreg(t, "Z6")}, "62f15449c6f402"},
// Gather and scatter prefetch hints (memory-only, /digit in reg).
{"VGATHERPF0DPD", "VGATHERPF0DPD", []Operand{vreg(t, "K5"), vsib("R10", "Y29", 8)}, "6292fd45c60cea"},
{"VSCATTERPF1DPS", "VSCATTERPF1DPS", []Operand{vreg(t, "K2"), vsib("R10", "Z28", 4)}, "62927d42c634a2"},
// Opmask broadcasts and the K logic.
{"VPBROADCASTMB2Q", "VPBROADCASTMB2Q", []Operand{vreg(t, "K1"), vreg(t, "Z2")}, "62f2fe482ad1"},
{"VPBROADCASTMW2D", "VPBROADCASTMW2D", []Operand{vreg(t, "K3"), vreg(t, "Z4")}, "62f27e483ae3"},
{"KUNPCKWD", "KUNPCKWD", []Operand{vreg(t, "K6"), vreg(t, "K4"), vreg(t, "K1")}, "c5dc4bce"},
{"KADDB", "KADDB", []Operand{vreg(t, "K2"), vreg(t, "K3"), vreg(t, "K5")}, "c5e54aea"},
// Lane extracts to general registers (EVEX and VEX routes).
{"VPEXTRB", "VPEXTRB", []Operand{Imm(3), vreg(t, "X26"), AX}, "62637d0814d003"},
{"VPEXTRD", "VPEXTRD", []Operand{Imm(1), vreg(t, "X26"), vreg(t, "R9")}, "62437d0816d101"},
{"VPEXTRD vex", "VPEXTRD", []Operand{Imm(1), vreg(t, "X2"), DI}, "c4e37916d701"},
{"VPINSRQ", "VPINSRQ", []Operand{Imm(1), DI, vreg(t, "X3"), vreg(t, "X4")}, "c4e3e122e701"},
// Moves: masked unaligned, masked scalar register form, half moves
// and non-temporal stores.
{"VMOVUPS mask", "VMOVUPS", []Operand{vreg(t, "Z1"), vreg(t, "K2"), vreg(t, "Z3")}, "62f17c4a11cb"},
{"VMOVSD 3op", "VMOVSD", []Operand{vreg(t, "X14"), vreg(t, "X5"), vreg(t, "K3"), vreg(t, "X22")}, "6231d70b11f6"},
{"VMOVSS 3op", "VMOVSS", []Operand{vreg(t, "X18"), vreg(t, "X3"), vreg(t, "K2"), vreg(t, "X25")}, "6281660a11d1"},
{"VMOVHPS insert", "VMOVHPS", []Operand{Ptr(SI, 0, 8), vreg(t, "X18"), vreg(t, "X19")}, "62e16c00161e"},
{"VMOVHPS store", "VMOVHPS", []Operand{vreg(t, "X20"), Ptr(SI, 8, 8)}, "62e17c08176601"},
{"VMOVLHPS", "VMOVLHPS", []Operand{vreg(t, "X16"), vreg(t, "X5"), vreg(t, "X17")}, "62a1540816c8"},
{"VMOVNTDQ", "VMOVNTDQ", []Operand{vreg(t, "Z7"), Ptr(SI, 0, 64)}, "62f17d48e73e"},
{"VMOVNTDQA", "VMOVNTDQA", []Operand{Ptr(SI, 64, 64), vreg(t, "Z8")}, "62727d482a4601"},
{"VMOVNTPS", "VMOVNTPS", []Operand{vreg(t, "Z9"), Ptr(SI, 0, 64)}, "62717c482b0e"},
// Scalar compares with and without the 66 prefix.
{"VCOMISD", "VCOMISD", []Operand{vreg(t, "X5"), vreg(t, "X6")}, "c5f92ff5"},
{"VUCOMISS", "VUCOMISS", []Operand{vreg(t, "X7"), vreg(t, "X8")}, "c5782ec7"},
// Floating point helpers.
{"VSQRTSD", "VSQRTSD", []Operand{vreg(t, "X1"), vreg(t, "X2"), vreg(t, "K1"), vreg(t, "X3")}, "62f1ef0951d9"},
{"VEXP2PD", "VEXP2PD", []Operand{vreg(t, "Z5"), vreg(t, "K1"), vreg(t, "Z6")}, "62f2fd49c8f5"},
{"VRCP28SD", "VRCP28SD", []Operand{vreg(t, "X9"), vreg(t, "X8"), vreg(t, "K1"), vreg(t, "X10")}, "6252bd09cbd1"},
{"VBROADCASTF32X2", "VBROADCASTF32X2", []Operand{vreg(t, "X1"), vreg(t, "K1"), vreg(t, "Z2")}, "62f27d4919d1"},
{"VPCOMPRESSB", "VPCOMPRESSB", []Operand{vreg(t, "Z1"), vreg(t, "K1"), Ptr(SI, 0, 64)}, "62f27d49630e"},
}
for _, c := range cases {
code, err := Encode(c.mnem, c.ops...)
if err != nil {
t.Errorf("%s: Encode: %v", c.name, err)
continue
}
if got := hexCompact(code); got != c.want {
t.Errorf("%s: got %s, want %s", c.name, got, c.want)
}
}
}
+55
View File
@@ -463,6 +463,61 @@ func (img *Image) emitGOObject(pkgPath, srcPath string, pre []byte, minLC int, r
symRelocs[si] = append(symRelocs[si], rec[:]...)
}
}
// The data symbols' own relocations: the symbol-valued DATA fields
// ("DATA s+0(SB)/8, $other(SB)"). The toolchain patches each field
// with the target's absolute address through an R_ADDR of the DATA
// line's width, on every architecture (the code relocations are
// per-architecture PC-relative shapes; a data pointer word is not), so
// this mapping bypasses relocField. The definitions were appended in
// DataSyms order, so data symbol i is definition index i.
for i, d := range img.DataSyms {
for _, r := range d.Relocs {
if r.Kind != RelAddr {
return nil, fmt.Errorf("GOOBJ emission: data symbol %q carries a non-data relocation", d.Name)
}
var rec [23]byte
binary.LittleEndian.PutUint32(rec[0:], uint32(int32(r.Off)))
rec[4] = r.Siz
binary.LittleEndian.PutUint16(rec[5:], relocAddr)
binary.LittleEndian.PutUint64(rec[7:], uint64(r.Addend))
switch {
case r.External && r.Name == goobjBuiltinMorestack:
binary.LittleEndian.PutUint32(rec[15:], pkgIdxBuiltin)
binary.LittleEndian.PutUint32(rec[19:], goobjBuiltinMorestackNoctxt)
case r.External:
pkg, name := splitQualified(r.Name)
if pkg == "" {
return nil, fmt.Errorf("GOOBJ emission: external symbol %q has no package prefix", r.Name)
}
pIdx, ok := extPkgIdx[pkg]
if !ok {
return nil, fmt.Errorf("GOOBJ emission: package %q not resolved", pkg)
}
sIdx, ok := extSymIdx[pkg+"·"+name]
if !ok {
return nil, fmt.Errorf("GOOBJ emission: symbol %s·%s not resolved", pkg, name)
}
binary.LittleEndian.PutUint32(rec[15:], uint32(pIdx))
binary.LittleEndian.PutUint32(rec[19:], uint32(sIdx))
default:
if di, ok := defIdx[r.Name]; ok {
binary.LittleEndian.PutUint32(rec[15:], pkgIdxSelf)
binary.LittleEndian.PutUint32(rec[19:], uint32(di))
break
}
// A DATA field may hold the address of a TEXT function of
// the same file (the rt0 lib entry spelling), which is a
// non-package definition.
ni, isText := textNpIdx[r.Name]
if !isText {
return nil, fmt.Errorf("GOOBJ emission: reference to unknown symbol %q", r.Name)
}
binary.LittleEndian.PutUint32(rec[15:], pkgIdxNone)
binary.LittleEndian.PutUint32(rec[19:], uint32(ni))
}
symRelocs[i] = append(symRelocs[i], rec[:]...)
}
}
// The DWARF symbols' own relocations (the function address references).
for _, ds := range dwarfRelocs {
for _, r := range ds.relocs {
+706 -14
View File
@@ -14,30 +14,82 @@ var aluOp = map[string]struct {
}{
"ADD": {0x01, 0},
"OR": {0x09, 1},
"ADC": {0x11, 2},
"SBB": {0x19, 3},
"AND": {0x21, 4},
"SUB": {0x29, 5},
"XOR": {0x31, 6},
"CMP": {0x39, 7},
}
// unaryOp maps INC/DEC/NEG/NOT to their /digit and base opcode. INC/DEC use
// the 0xFE/0xFF group (the short 0x40-0x4F forms are REX prefixes in 64-bit
// mode); NEG/NOT use the 0xF6/0xF7 group.
// unaryOp maps INC/DEC/NEG/NOT/MUL/DIV/IDIV to their /digit and base opcode.
// INC/DEC use the 0xFE/0xFF group (the short 0x40-0x4F forms are REX prefixes
// in 64-bit mode); NEG/NOT/MUL/DIV/IDIV use the 0xF6/0xF7 group (MUL /4,
// DIV /6, IDIV /7; the accumulator is the implicit other operand).
var unaryOp = map[string]struct {
digit int
op byte
}{
"INC": {0, 0xFF},
"DEC": {1, 0xFF},
"NOT": {2, 0xF7},
"NEG": {3, 0xF7},
"INC": {0, 0xFF},
"DEC": {1, 0xFF},
"NOT": {2, 0xF7},
"NEG": {3, 0xF7},
"MUL": {4, 0xF7},
"DIV": {6, 0xF7},
"IDIV": {7, 0xF7},
}
// shiftOp maps SHL/SHR/SAR to their /digit in the 0xC0/0xC1/0xD0-0xD3 group.
// shiftOp maps SHL/SAL/SHR/SAR/ROL/ROR/RCL/RCR to their /digit in the
// 0xC0/0xC1/0xD0-0xD3 group. SAL is the same encoding as SHL (/4).
var shiftOp = map[string]int{
"SHL": 4,
"SAL": 4,
"SHR": 5,
"SAR": 7,
"ROL": 0,
"ROR": 1,
"RCL": 2,
"RCR": 3,
}
// bitTestOp maps BT/BTS/BTR/BTC to their /digit in the 0F BA immediate form;
// the register form is 0F A3/AB/B3/BB, the same digit in the low nibble's
// opcode row.
var bitTestOp = map[string]int{
"BT": 4,
"BTS": 5,
"BTR": 6,
"BTC": 7,
}
// noOperandTable maps a fixed no-operand mnemonic to its opcode bytes. The
// fence names carry their opcode inside the 0F AE /digit group spelled out in
// full (E8/F0/F8), and PAUSE is F3 90.
//
// LOCK, REP and REPN are the prefix statements. go tool asm encodes each as
// a standalone one-byte instruction with a PC of its own (F0, F3 and F2
// respectively), not as a prefix field merged into the next instruction: the
// statement that follows is encoded unaware of it, and nothing validates
// that the pairing is a legal one (LOCK before NOP assembles without
// complaint, each byte pinned against the toolchain). Because the bytes
// land in the stream before the following statement anyway, a LOCKed
// CMPXCHGQ encodes identically to a prefixed form.
var noOperandTable = map[string][]byte{
"CPUID": {0x0F, 0xA2},
"RDTSC": {0x0F, 0x31},
"RDTSCP": {0x0F, 0x01, 0xF9},
"SYSCALL": {0x0F, 0x05},
"XGETBV": {0x0F, 0x01, 0xD0},
"CLD": {0xFC},
"STD": {0xFD},
"PAUSE": {0xF3, 0x90},
"LFENCE": {0x0F, 0xAE, 0xE8},
"MFENCE": {0x0F, 0xAE, 0xF0},
"SFENCE": {0x0F, 0xAE, 0xF8},
"UNDEF": {0x0F, 0x0B},
"LOCK": {0xF0},
"REP": {0xF3},
"REPN": {0xF2},
}
// --- MOV --------------------------------------------------------------------
@@ -309,6 +361,13 @@ func (e *enc) encodeALUImm(digit int, dst Operand, imm int64, size int) error {
if err != nil {
return err
}
// The byte accumulator short form (0x04+digit*8, no ModR/M) when
// the destination is AL, the form the Go assembler prefers here.
if r, ok := dst.(Reg); ok && r.idx == 0 {
i := &instr{opcode: []byte{byte(0x04 + digit*8)}, modrm: -1, sib: -1}
i.imm = immBytes
return e.emit(i)
}
i := newInstr(1, []byte{0x80})
if err := setRMDigit(i, digit, dst, 1); err != nil {
return err
@@ -451,13 +510,34 @@ func (e *enc) encodeUnary(op struct {
// --- SHL/SHR/SAR ------------------------------------------------------------
func (e *enc) encodeShift(digit int, ops []Operand, size int) error {
// doubleShiftOp maps the two mnemonics whose three-operand form go tool asm
// accepts to the SHLD/SHRD opcode pair (imm8 form, CL form). SAR, SAL and
// the rotates have no such form: the oracle rejects SARQ/ROLQ with three
// operands, and so do we.
var doubleShiftOp = map[string][2]byte{
"SHL": {0xA4, 0xA5}, // SHLD
"SHR": {0xAC, 0xAD}, // SHRD
}
// isShiftCountCL reports whether a count operand is the CL register or its
// CX spelling: go tool asm accepts both (CX names the same low byte) and
// rejects ECX/RCX.
func isShiftCountCL(o Operand) bool {
reg, ok := o.(Reg)
return ok && reg.idx == 1 && (reg.size == 1 || reg.size == 2)
}
func (e *enc) encodeShift(base string, ops []Operand, size int) error {
digit := shiftOp[base]
if len(ops) == 3 {
return e.encodeDoubleShift(base, ops, size)
}
if len(ops) != 2 {
return fmt.Errorf("shift expects 2 operands, got %d", len(ops))
}
count, dst := ops[0], ops[1]
// Count is $1, %CL, or an imm8.
if reg, ok := count.(Reg); ok && reg.idx == 1 && reg.size <= 1 {
// Count is $1, CL (or its CX spelling), or an imm8.
if isShiftCountCL(count) {
// CL: 0xD2 (8-bit) / 0xD3.
op := byte(0xD3)
if size == 1 {
@@ -504,6 +584,44 @@ func (e *enc) encodeShift(digit int, ops []Operand, size int) error {
return e.emit(i)
}
// encodeDoubleShift emits the three-operand SHL/SHR form, which the Go
// assembler spells as a shift but encodes as SHLD/SHRD (0F A4/A5, 0F AC/AD):
// the first operand is the count ($imm or CL), the second feeds the vacated
// bits (the reg field) and the third is the shifted value (the r/m field),
// matching go tool asm byte for byte. The W/L/Q widths exist; the oracle
// rejects the three-operand B form and every SAR/rotate one.
func (e *enc) encodeDoubleShift(base string, ops []Operand, size int) error {
opc, ok := doubleShiftOp[base]
if !ok || size == 1 {
return fmt.Errorf("%s: shift expects 2 operands, got %d", base, len(ops))
}
count, src, dst := ops[0], ops[1], ops[2]
srcReg, ok := src.(Reg)
if !ok {
return fmt.Errorf("%s: middle operand must be a register, like go tool asm", base)
}
i := newInstr(size, []byte{0x0F, opc[0]})
if isShiftCountCL(count) {
// CL (or CX) form: 0F A5/AD.
i.opcode[1] = opc[1]
} else {
imm, ok := count.(Imm)
if !ok {
return fmt.Errorf("shift count must be $1, CL or an immediate")
}
// The count is an unsigned imm8: the same range convention as the
// two-operand shift above.
if imm < 0 || imm > 255 {
return fmt.Errorf("shift count $%d is out of the 0..255 range", int64(imm))
}
i.imm = []byte{byte(imm)}
}
if err := setRMReg(i, srcReg.idx, srcReg.idx >= 8, false, dst, size); err != nil {
return err
}
return e.emit(i)
}
// --- IMUL -------------------------------------------------------------------
func (e *enc) encodeImul(ops []Operand, size int) error {
@@ -913,6 +1031,7 @@ type sseMove struct {
var sseMoveTable = map[string]sseMove{
"MOVOU": {0xF3, 0x6F, 0x7F}, // MOVDQU, unaligned octa
"MOVO": {0x66, 0x6F, 0x7F}, // MOVDQA, aligned octa
"MOVOA": {0x66, 0x6F, 0x7F}, // MOVDQA, the aligned octa alias
"MOVUPS": {0x00, 0x10, 0x11}, // unaligned packed single
"MOVAPS": {0x00, 0x28, 0x29}, // aligned packed single
"MOVUPD": {0x66, 0x10, 0x11}, // unaligned packed double
@@ -938,12 +1057,12 @@ func (e *enc) encodeSSEMove(m sseMove, ops []Operand) error {
op = m.load
reg, rm = dstReg, src
case srcVec:
if _, ok := dst.(Mem); !ok {
if !isX86Mem(dst) {
return fmt.Errorf("SSE move: invalid destination operand")
}
reg, rm = srcReg, dst
case dstVec:
if _, ok := src.(Mem); !ok {
if !isX86Mem(src) {
return fmt.Errorf("SSE move: invalid source operand")
}
op = m.load
@@ -1000,10 +1119,120 @@ var sseBinTable = map[string]sseBin{
"PSUBB": {0x66, 0xF8, false}, "PSUBW": {0x66, 0xF9, false},
"PSUBD": {0x66, 0xFA, false}, "PSUBQ": {0x66, 0xFB, false},
"PCMPEQB": {0x66, 0x74, false}, "PCMPEQW": {0x66, 0x75, false},
"PCMPEQD": {0x66, 0x76, false},
"PCMPEQD": {0x66, 0x76, false}, "PCMPEQL": {0x66, 0x76, false},
"PCMPGTB": {0x66, 0x64, false}, "PCMPGTW": {0x66, 0x65, false},
"PCMPGTD": {0x66, 0x66, false},
"PSHUFB": {0x66, 0x00, true},
// Scalar compares and square root, packed adds/subtracts and the byte
// unpack, the spellings the Plan 9 table uses (COMISD orders the
// operands like every other two-operand form).
"ANDNPD": {0x66, 0x55, false},
"ANDNPS": {0x00, 0x55, false},
"COMISD": {0x66, 0x2F, false},
"SQRTSD": {0xF2, 0x51, false},
"PADDL": {0x66, 0xFE, false},
"PSUBL": {0x66, 0xFA, false},
"PUNPCKLBW": {0x66, 0x60, false},
// AES round functions (66 0F38) and the SHA message schedule helpers
// (no prefix, 0F38).
"AESENC": {0x66, 0xDC, true},
"AESENCLAST": {0x66, 0xDD, true},
"AESDEC": {0x66, 0xDE, true},
"AESDECLAST": {0x66, 0xDF, true},
"AESIMC": {0x66, 0xDB, true},
"SHA1MSG1": {0x00, 0xC9, true},
"SHA1MSG2": {0x00, 0xCA, true},
"SHA1NEXTE": {0x00, 0xC8, true},
"SHA256MSG1": {0x00, 0xCC, true},
"SHA256MSG2": {0x00, 0xCD, true},
}
// sseImm3 describes a legacy SSE instruction taking a leading imm8 and two
// further operands: OP $imm, src, dst with reg = dst, rm = src. map38 and
// map3A select the opcode map the same way as sseBin's.
type sseImm3 struct {
prefix byte
op byte
map3A bool // opcode lives under 0F3A instead of 0F38
}
// sseImm3Table covers the imm8-controlled legacy instructions: the SSSE3
// align/blend shuffles, the string compare, carry-less multiply and the AES
// key assistant. SHA1RNDS4 carries no prefix, unlike its 0F3A siblings.
var sseImm3Table = map[string]sseImm3{
"PALIGNR": {0x66, 0x0F, true},
"PBLENDW": {0x66, 0x0E, true},
"PCMPESTRI": {0x66, 0x61, true},
"PCLMULQDQ": {0x66, 0x44, true},
"AESKEYGENASSIST": {0x66, 0xDF, true},
"SHA1RNDS4": {0x00, 0xCC, true},
}
// sseExtract describes a lane extract: OP $imm, xsrc, dst with reg = the XMM
// source and rm = the destination (GPR or memory). PEXTRW's GPR destination
// uses the older 0F C5 form; its memory destination the SSE4.1 0F3A 15 one,
// so it carries both opcodes.
type sseExtract struct {
op []byte
opMem []byte // used when the destination is memory; nil shares op
rexW bool // PEXTRQ's REX.W
}
var sseExtractTable = map[string]sseExtract{
"PEXTRB": {[]byte{0x0F, 0x3A, 0x14}, nil, false},
"PEXTRD": {[]byte{0x0F, 0x3A, 0x16}, nil, false},
"PEXTRQ": {[]byte{0x0F, 0x3A, 0x16}, nil, true},
"PEXTRW": {[]byte{0x0F, 0xC5}, []byte{0x0F, 0x3A, 0x15}, false},
}
// sseInsert describes a lane insert: OP $imm, src, xdst with reg = the XMM
// destination and rm = the source (GPR or memory).
type sseInsert struct {
op []byte
rexW bool // PINSRQ's REX.W
}
var sseInsertTable = map[string]sseInsert{
"PINSRB": {[]byte{0x0F, 0x3A, 0x20}, false},
"PINSRD": {[]byte{0x0F, 0x3A, 0x22}, false},
"PINSRQ": {[]byte{0x0F, 0x3A, 0x22}, true},
"PINSRW": {[]byte{0x0F, 0xC4}, false},
}
// sseShiftImm maps the legacy packed integer shifts' immediate form:
// OP $imm, dst (66 0F 71/72/73 /digit). The Plan 9 dword spellings end in L
// (PSLLL/PSRAL/PSRLL) and the octa byte shifts are PSLLDQ/PSRLDQ.
var sseShiftImm = map[string]sseShift{
"PSLLW": {0x71, 6},
"PSRLW": {0x71, 2},
"PSRAW": {0x71, 4},
"PSLLL": {0x72, 6},
"PSRLL": {0x72, 2},
"PSRAL": {0x72, 4},
"PSLLQ": {0x73, 6},
"PSRLQ": {0x73, 2},
"PSLLDQ": {0x73, 7},
"PSRLDQ": {0x73, 3},
}
// sseShiftVar maps the variable-count forms (the count comes from an XMM
// register or memory): OP count, dst (66 0F D1-F3). PSLLDQ/PSRLDQ have no
// variable form.
var sseShiftVar = map[string]byte{
"PSLLW": 0xF1,
"PSRLW": 0xD1,
"PSRAW": 0xE1,
"PSLLL": 0xF2,
"PSRLL": 0xD2,
"PSRAL": 0xE2,
"PSLLQ": 0xF3,
"PSRLQ": 0xD3,
}
// sseShift is one /digit selector in the 0F 71/72/73 immediate group.
type sseShift struct {
op byte
digit int
}
// sseShuf describes a legacy SSE shuffle taking a trailing imm8
@@ -1016,6 +1245,7 @@ type sseShuf struct {
var sseShufTable = map[string]sseShuf{
"SHUFPS": {0, 0xC6}, "SHUFPD": {0x66, 0xC6},
"PSHUFD": {0x66, 0x70}, "PSHUFHW": {0xF3, 0x70}, "PSHUFLW": {0xF2, 0x70},
"PSHUFL": {0x66, 0x70},
}
// encodeSSEBin encodes reg = reg op rm (memory allowed for rm).
@@ -1090,3 +1320,465 @@ func (e *enc) encodeCvtsi2sd(quad bool, ops []Operand) error {
}
return e.emit(i)
}
// --- carry, bit test, exchange and accumulate -------------------------------
// encodeBitTest encodes BT/BTS/BTR/BTC. The bit index goes first in Plan 9
// order (BTQ AX, BX tests BX at the offset in AX, encoding 0F A3 with
// reg = index, rm = target); an immediate index uses 0F BA /digit with imm8.
func (e *enc) encodeBitTest(name string, ops []Operand, size int) error {
if len(ops) != 2 {
return fmt.Errorf("%s expects 2 operands, got %d", name, len(ops))
}
digit := bitTestOp[name]
index, target := ops[0], ops[1]
if reg, ok := index.(Reg); ok {
// Register index: 0F A3 (BT) / 0F AB (BTS) / 0F B3 (BTR) / 0F BB (BTC),
// the /digit base plus eight per step.
i := newInstr(size, []byte{0x0F, 0xA3 + byte(digit-4)<<3})
if err := setRM(i, reg, target, size); err != nil {
return err
}
return e.emit(i)
}
imm, ok := index.(Imm)
if !ok {
return fmt.Errorf("%s index must be a register or an immediate", name)
}
immByte, err := imm8(int64(imm))
if err != nil {
return err
}
i := newInstr(size, []byte{0x0F, 0xBA})
if err := setRMDigit(i, digit, target, size); err != nil {
return err
}
i.imm = []byte{immByte}
return e.emit(i)
}
// encodeExchange encodes XCHG. A register-to-register exchange where either
// operand is AX uses the 0x90+r accumulator form (with REX.W for the quad
// form, as the Go assembler emits it); everything else uses 0x86/0x87 with
// the register operand in ModRM.reg, the memory (or second register) in r/m.
func (e *enc) encodeExchange(ops []Operand, size int) error {
if len(ops) != 2 {
return fmt.Errorf("XCHG expects 2 operands, got %d", len(ops))
}
src, dst := ops[0], ops[1]
srcReg, srcIsReg := src.(Reg)
dstReg, dstIsReg := dst.(Reg)
if srcIsReg && dstIsReg && size > 1 && (srcReg.idx == 0 || dstReg.idx == 0) {
// 0x90+r: r is the non-AX register, whichever side it sits on.
r := dstReg
if srcReg.idx == 0 {
r = dstReg
} else {
r = srcReg
}
i := newInstr(size, []byte{0x90 + byte(r.idx&7)})
i.rexB = r.idx >= 8
return e.emit(i)
}
op := byte(0x87)
if size == 1 {
op = 0x86
}
switch {
case srcIsReg:
i := newInstr(size, []byte{op})
if err := setRM(i, srcReg, dst, size); err != nil {
return err
}
return e.emit(i)
case dstIsReg:
i := newInstr(size, []byte{op})
if err := setRM(i, dstReg, src, size); err != nil {
return err
}
return e.emit(i)
}
return fmt.Errorf("XCHG: at least one operand must be a register")
}
// encodeRegRegOp encodes the two-operand read-modify-write pair CMPXCHG
// (0F B0/B1) and XADD (0F C0/C1): reg = source, rm = destination, with the
// destination writable (register or memory).
func (e *enc) encodeRegRegOp(op8, op byte, name string, ops []Operand, size int) error {
if len(ops) != 2 {
return fmt.Errorf("%s expects 2 operands, got %d", name, len(ops))
}
srcReg, ok := ops[0].(Reg)
if !ok {
return fmt.Errorf("%s source must be a register", name)
}
opc := op
if size == 1 {
opc = op8
}
i := newInstr(size, []byte{0x0F, opc})
if err := setRM(i, srcReg, ops[1], size); err != nil {
return err
}
return e.emit(i)
}
// encodeCrc32 encodes the CRC32 family: F2 0F38 F0 for the byte form, F1 for
// the rest; the word form carries a 0x66 operand-size prefix (66 F2, the
// prefix order the Go assembler emits) and the quad form REX.W. reg = GPR
// accumulator, rm = the data source.
func (e *enc) encodeCrc32(ops []Operand, size int) error {
if len(ops) != 2 {
return fmt.Errorf("CRC32 expects 2 operands, got %d", len(ops))
}
dstReg, ok := ops[1].(Reg)
if !ok || dstReg.isVec() {
return fmt.Errorf("CRC32 destination must be a general register")
}
i := &instr{opSize16: size == 2, prefix: 0xF2, opcode: []byte{0x0F, 0x38, 0xF0}, modrm: -1, sib: -1}
if size > 1 {
i.opcode[2] = 0xF1
}
i.rexW = size == 8
if err := setRM(i, dstReg, ops[0], size); err != nil {
return err
}
return e.emit(i)
}
// encodeCarryExt encodes ADCX (66 0F38 F6) and ADOX (F3 0F38 F6): reg =
// destination, rm = source, the carry/overflow flag as the carry-in.
func (e *enc) encodeCarryExt(prefix byte, ops []Operand, size int) error {
if len(ops) != 2 {
return fmt.Errorf("ADCX/ADOX expects 2 operands, got %d", len(ops))
}
dstReg, ok := ops[1].(Reg)
if !ok || dstReg.isVec() {
return fmt.Errorf("ADCX/ADOX destination must be a general register")
}
i := &instr{prefix: prefix, opcode: []byte{0x0F, 0x38, 0xF6}, modrm: -1, sib: -1, rexW: size == 8}
if err := setRM(i, dstReg, ops[0], size); err != nil {
return err
}
return e.emit(i)
}
// --- string primitives, flags and INT ----------------------------------------
// encodeStringOp encodes the no-operand string primitives MOVS (A4/A5) and
// STOS (AA/AB); the size suffix picks the byte form and supplies the 0x66 or
// REX.W prefix.
func (e *enc) encodeStringOp(base string, ops []Operand, size int) error {
if len(ops) != 0 {
return fmt.Errorf("%s takes no operands, got %d", base, len(ops))
}
var op byte
switch base {
case "MOVS":
op = 0xA5
if size == 1 {
op = 0xA4
}
case "STOS":
op = 0xAB
if size == 1 {
op = 0xAA
}
default:
return fmt.Errorf("unsupported string instruction %q", base)
}
return e.emit(newInstr(size, []byte{op}))
}
// encodeInt encodes INT with its single imm8 operand. The field takes the
// low byte silently inside the 32-bit span, matching the scalar convention
// (go tool asm encodes INT $256 as CD 00).
func (e *enc) encodeInt(ops []Operand) error {
if len(ops) != 1 {
return fmt.Errorf("INT expects 1 operand, got %d", len(ops))
}
imm, ok := ops[0].(Imm)
if !ok {
return fmt.Errorf("INT operand must be an immediate")
}
if imm < -(1<<31) || imm > (1<<32)-1 {
return fmt.Errorf("immediate $%d does not fit in 32 bits", int64(imm))
}
return e.emit(&instr{opcode: []byte{0xCD}, modrm: -1, sib: -1, imm: []byte{byte(imm)}})
}
// encodeMxcsr encodes LDMXCSR (0F AE /2) and STMXCSR (0F AE /3); both take a
// single 32-bit memory operand.
func (e *enc) encodeMxcsr(digit int, ops []Operand) error {
if len(ops) != 1 {
return fmt.Errorf("MXCSR instruction expects 1 operand, got %d", len(ops))
}
m, ok := ops[0].(Mem)
if !ok {
return fmt.Errorf("MXCSR instruction requires a memory operand")
}
i := &instr{opcode: []byte{0x0F, 0xAE}, modrm: -1, sib: -1}
if err := setMem(i, digit, m); err != nil {
return err
}
return e.emit(i)
}
// cvtIntOp maps the scalar float-to-integer conversions to their mandatory
// prefix and opcode: 0F 2D (CVTSD2S, CVTSS2S) and 0F 2C (their truncating
// CVTT forms). The mnemonic's Q/L suffix fixes the GPR destination width.
var cvtIntOp = map[string]struct {
prefix byte
op byte
}{
"CVTSD2S": {0xF2, 0x2D},
"CVTTSD2S": {0xF2, 0x2C},
"CVTSS2S": {0xF3, 0x2D},
"CVTTSS2S": {0xF3, 0x2C},
}
// encodeCvtInt encodes a scalar float-to-integer conversion: F2/F3 0F 2D/2C
// with reg = GPR destination, rm = XMM (or memory) source; REX.W follows the
// quad spellings.
func (e *enc) encodeCvtInt(base string, ops []Operand, size int) error {
if len(ops) != 2 {
return fmt.Errorf("%s expects 2 operands, got %d", base, len(ops))
}
spec := cvtIntOp[base]
src, dst := ops[0], ops[1]
dstReg, ok := dst.(Reg)
if !ok || dstReg.isVec() {
return fmt.Errorf("%s destination must be a general register", base)
}
i := newInstr(size, []byte{0x0F, spec.op})
i.prefix = spec.prefix
if err := setRM(i, dstReg, src, size); err != nil {
return err
}
return e.emit(i)
}
// encodeFmov encodes the x87 double move. The memory forms are DD /0
// (FMOVD mem, F: load) and DD /2 (FMOVD F, mem: store); a register-to-register
// move is DD C0+dst (FLD st(dst)), the form the Go assembler emits.
func (e *enc) encodeFmov(ops []Operand) error {
if len(ops) != 2 {
return fmt.Errorf("FMOVD expects 2 operands, got %d", len(ops))
}
src, dst := ops[0], ops[1]
srcReg, srcIsF := src.(Reg)
dstReg, dstIsF := dst.(Reg)
srcF := srcIsF && srcReg.fp
dstF := dstIsF && dstReg.fp
switch {
case srcF && dstF:
// The register form is DD /2 with rm = the destination (FST st(dst)).
i := &instr{opcode: []byte{0xDD}, modrm: -1, sib: -1}
if err := setRMDigit(i, 2, dstReg, 8); err != nil {
return err
}
return e.emit(i)
case dstF:
m, ok := src.(Mem)
if !ok {
return fmt.Errorf("FMOVD: invalid source operand")
}
i := &instr{opcode: []byte{0xDD}, modrm: -1, sib: -1}
if err := setMem(i, 0, m); err != nil {
return err
}
return e.emit(i)
case srcF:
m, ok := dst.(Mem)
if !ok {
return fmt.Errorf("FMOVD: invalid destination operand")
}
i := &instr{opcode: []byte{0xDD}, modrm: -1, sib: -1}
if err := setMem(i, 2, m); err != nil {
return err
}
return e.emit(i)
}
return fmt.Errorf("FMOVD needs an x87 register operand")
}
// --- legacy SSE imm8, extract, insert and packed shift families --------------
// encodeSSEImm3 encodes an imm8-controlled three-operand form: OP $imm, src,
// dst with reg = dst, rm = src and the immediate appended last (PALIGNR,
// PBLENDW, PCMPESTRI, PCLMULQDQ, AESKEYGENASSIST, SHA1RNDS4).
func (e *enc) encodeSSEImm3(m sseImm3, ops []Operand) error {
if len(ops) != 3 {
return fmt.Errorf("SSE imm8 instruction expects 3 operands ($imm, src, dst), got %d", len(ops))
}
imm, ok := ops[0].(Imm)
if !ok {
return fmt.Errorf("SSE imm8 instruction needs an immediate first operand")
}
immByte, err := imm8(int64(imm))
if err != nil {
return err
}
src, dst := ops[1], ops[2]
dstReg, ok2 := dst.(Reg)
if !ok2 || !dstReg.isVec() {
return fmt.Errorf("SSE imm8 instruction destination must be a vector register")
}
opcode := []byte{0x0F, 0x38, m.op}
if m.map3A {
opcode = []byte{0x0F, 0x3A, m.op}
}
i := &instr{prefix: m.prefix, opcode: opcode, modrm: -1, sib: -1}
if err := setRM(i, dstReg, src, 8); err != nil {
return err
}
i.imm = []byte{immByte}
return e.emit(i)
}
// encodeSSEExtract encodes a lane extract: OP $imm, xsrc, dst with reg = the
// XMM source, rm = the GPR or memory destination (PEXTRB/PEXTRD/PEXTRQ and
// PEXTRW, whose GPR form is the older 0F C5 opcode and whose memory form the
// SSE4.1 0F3A 15 one).
func (e *enc) encodeSSEExtract(m sseExtract, ops []Operand) error {
if len(ops) != 3 {
return fmt.Errorf("extract expects 3 operands ($imm, src, dst), got %d", len(ops))
}
imm, ok := ops[0].(Imm)
if !ok {
return fmt.Errorf("extract needs an immediate first operand")
}
immByte, err := imm8(int64(imm))
if err != nil {
return err
}
srcReg, srcVec := vecReg(ops[1])
if !srcVec {
return fmt.Errorf("extract source must be an XMM register")
}
opcode := m.op
if m.opMem != nil && memOperand(ops[2]) {
opcode = m.opMem
}
i := &instr{prefix: 0x66, opcode: opcode, modrm: -1, sib: -1, rexW: m.rexW}
if err := setRM(i, srcReg, ops[2], 8); err != nil {
return err
}
i.imm = []byte{immByte}
return e.emit(i)
}
// encodeSSEInsert encodes a lane insert: OP $imm, src, xdst with reg = the
// XMM destination and rm = the GPR or memory source (PINSRB/PINSRD/PINSRQ and
// PINSRW).
func (e *enc) encodeSSEInsert(m sseInsert, ops []Operand) error {
if len(ops) != 3 {
return fmt.Errorf("insert expects 3 operands ($imm, src, dst), got %d", len(ops))
}
imm, ok := ops[0].(Imm)
if !ok {
return fmt.Errorf("insert needs an immediate first operand")
}
immByte, err := imm8(int64(imm))
if err != nil {
return err
}
dstReg, dstVec := vecReg(ops[2])
if !dstVec {
return fmt.Errorf("insert destination must be an XMM register")
}
i := &instr{prefix: 0x66, opcode: m.op, modrm: -1, sib: -1, rexW: m.rexW}
if err := setRM(i, dstReg, ops[1], 8); err != nil {
return err
}
i.imm = []byte{immByte}
return e.emit(i)
}
// encodeSSEShift encodes the legacy packed integer shifts. The immediate
// form is OP $imm, dst (66 0F 71/72/73 /digit); the variable form
// OP count, dst carries the count in an XMM register (or memory) on the
// 66 0F D1-F3 opcodes. The destination is always the register written.
func (e *enc) encodeSSEShift(name string, ops []Operand) error {
if len(ops) != 2 {
return fmt.Errorf("%s expects 2 operands, got %d", name, len(ops))
}
dstReg, ok := ops[1].(Reg)
if !ok || !dstReg.isVec() {
return fmt.Errorf("%s destination must be the second, vector operand", name)
}
if imm, isImm := ops[0].(Imm); isImm {
spec := sseShiftImm[name]
immByte, err := imm8(int64(imm))
if err != nil {
return err
}
i := &instr{prefix: 0x66, opcode: []byte{0x0F, spec.op}, modrm: -1, sib: -1}
if err := setRMDigit(i, spec.digit, dstReg, 8); err != nil {
return err
}
i.imm = []byte{immByte}
return e.emit(i)
}
if !vecOrMem(ops[0]) {
return fmt.Errorf("%s count must be an immediate, a vector register or memory", name)
}
op, ok := sseShiftVar[name]
if !ok {
return fmt.Errorf("%s has no variable-count form", name)
}
i := &instr{prefix: 0x66, opcode: []byte{0x0F, op}, modrm: -1, sib: -1}
if err := setRM(i, dstReg, ops[0], 8); err != nil {
return err
}
return e.emit(i)
}
// encodeCmpsd encodes CMPSD, the scalar double compare with its predicate
// immediate LAST in Plan 9 order (src, dst, $imm), unlike the shuffle family:
// F2 0F C2 with reg = dst, rm = src.
func (e *enc) encodeCmpsd(ops []Operand) error {
if len(ops) != 3 {
return fmt.Errorf("CMPSD expects 3 operands (src, dst, $imm), got %d", len(ops))
}
imm, ok := ops[2].(Imm)
if !ok {
return fmt.Errorf("CMPSD predicate must be an immediate")
}
immByte, err := imm8(int64(imm))
if err != nil {
return err
}
dstReg, ok2 := ops[1].(Reg)
if !ok2 || !dstReg.isVec() {
return fmt.Errorf("CMPSD destination must be a vector register")
}
i := &instr{prefix: 0xF2, opcode: []byte{0x0F, 0xC2}, modrm: -1, sib: -1}
if err := setRM(i, dstReg, ops[0], 8); err != nil {
return err
}
i.imm = []byte{immByte}
return e.emit(i)
}
// encodeSha256rnds2 encodes SHA256RNDS2, whose first operand must be the
// literal X0 carrying the round constant: OP X0, src, dst (0F38 CB, no
// prefix, reg = dst, rm = src; X0 is implicit on the wire).
func (e *enc) encodeSha256rnds2(ops []Operand) error {
if len(ops) != 3 {
return fmt.Errorf("SHA256RNDS2 expects 3 operands (X0, src, dst), got %d", len(ops))
}
x0, ok := ops[0].(Reg)
if !ok || !x0.isVec() || x0.idx != 0 || x0.size != 16 {
return fmt.Errorf("SHA256RNDS2 first operand must be X0")
}
dstReg, ok2 := ops[2].(Reg)
if !ok2 || !dstReg.isVec() {
return fmt.Errorf("SHA256RNDS2 destination must be a vector register")
}
i := &instr{opcode: []byte{0x0F, 0x38, 0xCB}, modrm: -1, sib: -1}
if err := setRM(i, dstReg, ops[1], 8); err != nil {
return err
}
return e.emit(i)
}
+217
View File
@@ -0,0 +1,217 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
package asm
import (
"bytes"
"encoding/binary"
"os"
"os/exec"
"path/filepath"
"runtime"
"strings"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
)
// The differential kernels for the DATA-path and front-end gaps are kept in
// testdata/verify beside the campaign's other kernels; the verify package's
// suites are not open to the asm package, so this test is their runner: each
// kernel assembles through gasm and through go tool asm, and the functions'
// bytes must agree with the relocation sites masked on both sides.
// toolAsmObject assembles path with the installed toolchain's assembler for
// goarch ("" = the host) and returns the object bytes.
func toolAsmObject(t *testing.T, path, goarch string) []byte {
t.Helper()
goBin, err := exec.LookPath("go")
if err != nil {
t.Skip("no Go toolchain available")
}
out, err := exec.Command(goBin, "env", "GOROOT").Output()
if err != nil {
t.Fatalf("go env GOROOT: %v", err)
}
includeDir := filepath.Join(strings.TrimSpace(string(out)), "pkg", "include")
pkg := strings.TrimSuffix(filepath.Base(path), ".s")
pkg = strings.TrimSuffix(pkg, "_amd64")
pkg = strings.TrimSuffix(pkg, "_arm64")
objPath := filepath.Join(t.TempDir(), "oracle.o")
cmd := exec.Command(goBin, "tool", "asm", "-I", includeDir, "-p", pkg, "-o", objPath, path)
if goarch != "" {
environ := os.Environ()
env := make([]string, 0, len(environ)+1)
for _, e := range environ {
if !strings.HasPrefix(e, "GOARCH=") {
env = append(env, e)
}
}
cmd.Env = append(env, "GOARCH="+goarch)
}
if out, err := cmd.CombinedOutput(); err != nil {
t.Fatalf("go tool asm %s: %v\n%s", filepath.Base(path), err, out)
}
obj, err := os.ReadFile(objPath)
if err != nil {
t.Fatal(err)
}
return obj
}
// oracleFuncCode extracts the non-package TEXT functions' code bytes from a
// toolchain object, keyed by the name the object records (pkg.name).
func oracleFuncCode(t *testing.T, obj []byte) map[string][]byte {
t.Helper()
v := openGoobj(t, obj)
le := binary.LittleEndian
const symSize = 21
nps := v.syms(blkNonpkgdef)
data := v.blk(blkData)
didx := v.blk(blkDataIdx)
preceding := 0
for _, bi := range []int{blkSymdef, blkHashed64def, blkHasheddef} {
preceding += len(v.blk(bi)) / symSize
}
total := preceding + len(nps)
out := make(map[string][]byte, len(nps))
for i, s := range nps {
if s.typ != kindSTEXT {
continue
}
start := le.Uint32(didx[4*(preceding+i):])
end := uint32(len(data))
if preceding+i+1 < total {
end = le.Uint32(didx[4*(preceding+i+1):])
}
out[s.name] = data[start:end]
}
return out
}
// maskCode zeroes every relocation field, the way the toolchain's object
// leaves them for the linker.
func maskCode(code []byte, relocs []Reloc) []byte {
for _, r := range relocs {
for j := r.Off; j < r.Off+4 && j < len(code); j++ {
code[j] = 0
}
}
return code
}
// code assembles src for amd64 and returns the image's code bytes.
func code(path, src string) []byte {
f, errs := parser.Parse(path, src)
if len(errs) > 0 {
return nil
}
img, err := AssembleFile(f)
if err != nil {
return nil
}
return img.Code
}
// TestDifferentialKernels pins the new kernels against the oracle.
func TestDifferentialKernels(t *testing.T) {
if runtime.GOARCH != "amd64" {
t.Skip("the amd64 kernels assume an amd64 host assembler default")
}
for _, k := range []struct {
path string
goarch string
arm64 bool
}{
{filepath.Join("..", "testdata", "verify", "datarel_amd64.s"), "", false},
{filepath.Join("..", "testdata", "verify", "divslash_amd64.s"), "", false},
{filepath.Join("..", "testdata", "verify", "semicolons_amd64.s"), "", false},
{filepath.Join("..", "testdata", "verify", "datarel_arm64.s"), "arm64", true},
{filepath.Join("..", "testdata", "verify", "divslash_arm64.s"), "arm64", true},
} {
t.Run(filepath.Base(k.path), func(t *testing.T) {
src, err := os.ReadFile(k.path)
if err != nil {
t.Fatalf("read: %v", err)
}
f, errs := parser.Parse(k.path, string(src))
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
var img *Image
if k.arm64 {
img, err = AssembleFileARM64(f)
} else {
img, err = AssembleFile(f)
}
if err != nil {
t.Fatalf("assemble: %v", err)
}
gt := oracleFuncCode(t, toolAsmObject(t, k.path, k.goarch))
// The oracle keys its functions by the qualified object name
// (pkg.name); match on the local part.
byLocal := make(map[string][]byte, len(gt))
for name, code := range gt {
if _, after, ok := strings.Cut(name, "."); ok {
name = after
}
byLocal[name] = code
}
matched := 0
for _, fn := range img.Funcs {
gasmCode := maskCode(append([]byte(nil), img.Code[fn.Offset:fn.Offset+fn.Size]...), fn.Relocs)
goCode, ok := byLocal[fn.Name]
if !ok {
t.Errorf("%s: not in ground truth (%d functions: %v)", fn.Name, len(gt), keysOf(byLocal))
continue
}
goCode = maskCode(append([]byte(nil), goCode...), fn.Relocs)
cmpLen := min(len(goCode), len(gasmCode))
if !bytes.Equal(gasmCode[:cmpLen], goCode[:cmpLen]) {
t.Errorf("%s: MISMATCH gasm=%d go=%d bytes\ngasm %x\ngo %x", fn.Name, len(gasmCode), len(goCode), gasmCode, goCode)
continue
}
for _, b := range goCode[len(gasmCode):] {
if b != 0 {
t.Errorf("%s: non-zero trailing bytes in go tool asm output", fn.Name)
break
}
}
matched++
t.Logf("%s: MATCH (%d bytes)", fn.Name, len(gasmCode))
}
if matched == 0 {
t.Fatal("no functions matched")
}
})
}
}
func keysOf(m map[string][]byte) []string {
out := make([]string, 0, len(m))
for k := range m {
out = append(out, k)
}
return out
}
// TestSemicolonSpellingParity pins that the ';' statement separator changes
// nothing about the encoding: the one-line spelling assembles to exactly the
// bytes of the same statements written one per line.
func TestSemicolonSpellingParity(t *testing.T) {
for _, tt := range []struct{ one, two string }{
{"\tROLQ $3, DI; ROLQ $13, DI\n", "\tROLQ $3, DI\n\tROLQ $13, DI\n"},
{"\tREP; MOVSQ\n", "\tREP\n\tMOVSQ\n"},
{"\tXORQ AX, AX; XORQ CX, CX\n", "\tXORQ AX, AX\n\tXORQ CX, CX\n"},
} {
one := code("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n"+tt.one+"\tRET\n")
two := code("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n"+tt.two+"\tRET\n")
if !bytes.Equal(one, two) {
t.Errorf("semicolon spelling %q: %x, want the two-line bytes %x", tt.one, one, two)
}
}
}
+68 -3
View File
@@ -103,6 +103,7 @@ const (
RelArm64Branch // R_CALLARM64 (BL instruction)
RelArm64LDST64 // R_ARM64_PCREL_LDST64 (ADRP + 64-bit LDR/STR pair)
RelLoong64Branch // R_CALLLOONG64 (BL instruction)
RelAddr // R_ADDR: the absolute address of a symbol held in a DATA field
)
type Reloc struct {
@@ -112,13 +113,17 @@ type Reloc struct {
// Addend select the target: the symbol plus the byte offset. An
// External relocation names a symbol no GLOBL in the file defines;
// the object-file emitters carry it into the output's relocation
// table.
// table. Siz is the width of the patched field and is set only for
// data-field relocations (RelAddr, Off relative to the data symbol),
// whose width is the DATA line's; code relocations take their width
// from the architecture's instruction encoding.
Off int
After int
Name string
Addend int64
External bool
Kind RelocKind
Siz uint8
}
// DataSymbol describes one GLOBL symbol laid out in the data section.
@@ -130,6 +135,11 @@ type DataSymbol struct {
Static bool // the <> marker: file-local, not exported
Rodata bool // the RODATA flag: read-only data
Dupok bool // the DUPOK flag: duplicate-OK
// Relocs carries the symbol-valued DATA initialisers ("DATA s+0(SB)/8,
// $other(SB)"): fields of this symbol's data that hold another symbol's
// address, resolved by the linker. Off is relative to the symbol's
// data start.
Relocs []Reloc
}
// Bytes returns the whole image: code, then data.
@@ -257,6 +267,22 @@ func AssembleFile(f *ast.File) (*Image, error) {
img.Funcs[i].Relocs = append(img.Funcs[i].Relocs, reloc)
}
}
// The data symbols' symbol-valued DATA fields resolve the same way the
// code references do: a name the file defines (GLOBL or TEXT) stays an
// internal reference the emitters resolve, anything else is external.
// img.DataSyms was laid out in dataSyms order, so the indexes line up.
for i := range img.DataSyms {
for _, r := range dataSyms[i].relocs {
reloc := r
if _, ok := img.Symbols[reloc.Name]; !ok {
if _, ok := textOff[reloc.Name]; !ok {
reloc.External = true
externals[reloc.Name] = true
}
}
img.DataSyms[i].Relocs = append(img.DataSyms[i].Relocs, reloc)
}
}
for name := range externals {
img.Externals = append(img.Externals, name)
}
@@ -429,6 +455,23 @@ func markExternals(img *Image, dataSyms []dataSym) {
}
}
}
// The declared data symbols carry the file's own relocations (the
// symbol-valued DATA fields); the layouts appended img.DataSyms in
// dataSyms order, so the indexes line up. The trailing entries (the
// pooled arm64 literals) have no source relocations.
for i := range img.DataSyms {
if i >= len(dataSyms) {
break
}
for _, r := range dataSyms[i].relocs {
reloc := r
if !known[reloc.Name] {
reloc.External = true
externals[reloc.Name] = true
}
img.DataSyms[i].Relocs = append(img.DataSyms[i].Relocs, reloc)
}
}
for name := range externals {
img.Externals = append(img.Externals, name)
}
@@ -444,6 +487,9 @@ type dataSym struct {
static bool
rodata bool
dupok bool
// relocs are the symbol-valued DATA fields, in declaration order; Off
// is relative to the symbol's data start.
relocs []Reloc
}
// collectData gathers the file's static symbols (GLOBL) and their initial
@@ -511,8 +557,8 @@ func collectData(f *ast.File) ([]dataSym, error) {
if !ok {
return nil, fmt.Errorf("DATA %q: no matching GLOBL", dd.Name.Name)
}
if dd.Value == nil || !dd.Value.Imm.HasVal {
return nil, fmt.Errorf("DATA %q: value must be an integer immediate", dd.Name.Name)
if dd.Value == nil {
return nil, fmt.Errorf("DATA %q: missing value", dd.Name.Name)
}
w := dd.Width
switch w {
@@ -525,6 +571,25 @@ func collectData(f *ast.File) ([]dataSym, error) {
if off < 0 || off+int64(w) > int64(len(buf)) {
return nil, fmt.Errorf("DATA %q+%d/%d exceeds GLOBL size %d", dd.Name.Name, off, w, len(buf))
}
// A symbol value ("DATA s+0(SB)/8, $other(SB)", the rt0 spelling)
// leaves the field zero and records a relocation against the named
// symbol: the linker patches the absolute address at this data
// offset. The toolchain emits the same shape, an R_ADDR of the
// DATA width with the value's offset as the addend, on every
// architecture.
if sym := dd.Value.Imm.Sym; !dd.Value.Imm.HasVal && sym != nil {
syms[i].relocs = append(syms[i].relocs, Reloc{
Off: int(off),
Name: sym.Name,
Addend: sym.Offset,
Kind: RelAddr,
Siz: uint8(w),
})
continue
}
if !dd.Value.Imm.HasVal {
return nil, fmt.Errorf("DATA %q: value must be an integer immediate or a symbol address", dd.Name.Name)
}
v := dd.Value.Imm.Val
if dd.Value.Imm.Neg {
v = -v
+252
View File
@@ -4,6 +4,10 @@
package asm
import (
"encoding/binary"
"os"
"os/exec"
"path/filepath"
"strings"
"testing"
@@ -166,3 +170,251 @@ func TestCollectDataNumericFlags(t *testing.T) {
}
}
}
// TestCollectDataSymbolValue covers the symbol-valued DATA field ("DATA
// s+0(SB)/8, $other(SB)", the rt0 spelling): the field stays zero in the
// image and the relocation is recorded against the named symbol, whatever
// the file defines (a TEXT function, a GLOBL) or leaves external.
func TestCollectDataSymbolValue(t *testing.T) {
src := `#include "textflag.h"
TEXT ·Keep(SB), NOSPLIT, $0-8
MOVQ target+0(FP), AX
RET
GLOBL holder(SB), NOPTR, $32
DATA holder+0(SB)/8, $·Keep(SB)
DATA holder+8(SB)/8, $·Keep+5(SB)
DATA holder+16(SB)/8, $holder(SB)
GLOBL spare(SB), NOPTR, $8
DATA spare+0(SB)/8, $extvar(SB)
`
f, errs := parser.Parse("f_amd64.s", src)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFile(f)
if err != nil {
t.Fatalf("assemble: %v", err)
}
byName := map[string]DataSymbol{}
for _, d := range img.DataSyms {
byName[d.Name] = d
}
want := []struct {
sym string
off int
name string
addend int64
ext bool
}{
{"holder", 0, "Keep", 0, false},
{"holder", 8, "Keep", 5, false},
{"holder", 16, "holder", 0, false},
{"spare", 0, "extvar", 0, true},
}
var flat []struct {
sym string
r Reloc
}
for _, d := range img.DataSyms {
for _, r := range d.Relocs {
flat = append(flat, struct {
sym string
r Reloc
}{d.Name, r})
}
}
if len(flat) != len(want) {
t.Fatalf("data relocations = %d, want %d", len(flat), len(want))
}
for i, w := range want {
g := flat[i]
r := g.r
if g.sym != w.sym {
t.Errorf("relocation %d sits on %q, want %q", i, g.sym, w.sym)
continue
}
if r.Off != w.off || r.Name != w.name || r.Addend != w.addend || r.External != w.ext {
t.Errorf("relocation %d = {+%d %q addend %d ext %v}, want {+%d %q addend %d ext %v}",
i, r.Off, r.Name, r.Addend, r.External, w.off, w.name, w.addend, w.ext)
}
if r.Kind != RelAddr {
t.Errorf("relocation %d kind = %v, want RelAddr", i, r.Kind)
}
if r.Siz != 8 {
t.Errorf("relocation %d siz = %d, want 8", i, r.Siz)
}
}
// The fields themselves stay zero: only the linker fills them.
for _, b := range img.Data {
if b != 0 {
t.Fatal("data section is not all zero before relocation")
}
}
if len(img.Externals) != 1 || img.Externals[0] != "extvar" {
t.Errorf("Externals = %v, want [extvar]", img.Externals)
}
}
// TestGOObjectDataSymbolReloc pins the GOOBJ record a symbol-valued DATA
// field produces, against the shape the toolchain emits for the same
// source: an R_ADDR of the DATA width at the field offset, pkgIdxNone plus
// the non-package definition index when the target is the file's own TEXT
// function (the rt0 lib entry spelling).
func TestGOObjectDataSymbolReloc(t *testing.T) {
f, errs := parser.Parse("f_amd64.s", `#include "textflag.h"
TEXT ·Keep(SB), NOSPLIT, $0-8
RET
GLOBL holder(SB), NOPTR, $16
DATA holder+0(SB)/8, $·Keep+5(SB)
`)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFile(f)
if err != nil {
t.Fatalf("assemble: %v", err)
}
obj, err := img.GOObject("main", "f_amd64.s")
if err != nil {
t.Fatalf("GOObject: %v", err)
}
v := openGoobj(t, obj)
// Walk every relocation record; the data record is the one of Siz 8
// and type R_ADDR.
var off, add int64
var pkg, sym uint32
found := false
for data := v.blk(blkReloc); len(data) >= 23; data = data[23:] {
if data[4] != 8 || binary.LittleEndian.Uint16(data[5:]) != relocAddr {
continue
}
found = true
off = int64(int32(binary.LittleEndian.Uint32(data[0:])))
add = int64(binary.LittleEndian.Uint64(data[7:]))
pkg = binary.LittleEndian.Uint32(data[15:])
sym = binary.LittleEndian.Uint32(data[19:])
break
}
if !found {
t.Fatal("no data relocation record in the object")
}
if off != 0 || add != 5 {
t.Errorf("data reloc = {off %d addend %d}, want {off 0 addend 5}", off, add)
}
if pkg != pkgIdxNone {
t.Errorf("data reloc pkg = %#x, want pkgIdxNone (the TEXT function)", pkg)
}
// The function's non-package definition index: the four pc tables
// precede it, so index 4.
if sym != 4 {
t.Errorf("data reloc sym = %d, want 4", sym)
}
}
// TestGOObjectDataSymbolLink is the end-to-end proof for symbol-valued DATA
// fields: the gasm object is substituted for the toolchain's and re-linked,
// then executed, and the linked data word must hold the real address of the
// function the DATA line named (runtime.FuncForPC identifies it).
func TestGOObjectDataSymbolLink(t *testing.T) {
goBin, err := exec.LookPath("go")
if err != nil {
t.Skip("no Go toolchain available")
}
dir := t.TempDir()
asmSrc := `#include "textflag.h"
GLOBL entry(SB), NOPTR, $8
DATA entry+0(SB)/8, $·keepme(SB)
TEXT ·keepme(SB), NOSPLIT, $0-0
RET
TEXT ·entryptr(SB), NOSPLIT, $0-8
MOVQ entry+0(SB), AX
MOVQ AX, ret+0(FP)
RET
`
if err := os.WriteFile(filepath.Join(dir, "main_amd64.s"), []byte(asmSrc), 0o644); err != nil {
t.Fatal(err)
}
mainSrc := `package main
import "runtime"
func keepme()
func entryptr() uintptr
func main() {
pc := entryptr()
fn := runtime.FuncForPC(pc)
if fn == nil {
panic("the entry word does not point at a function")
}
if fn.Name() != "main.keepme" {
panic("the entry word points at " + fn.Name())
}
}
`
if err := os.WriteFile(filepath.Join(dir, "main.go"), []byte(mainSrc), 0o644); err != nil {
t.Fatal(err)
}
if err := os.WriteFile(filepath.Join(dir, "go.mod"), []byte("module dlink\n\ngo 1.21\n"), 0o644); err != nil {
t.Fatal(err)
}
// Capture the build: the package archive's asm object and the link line.
build := exec.Command(goBin, "build", "-x", "-work", "-o", filepath.Join(dir, "prog"), ".")
build.Dir = dir
buildLog, err := build.CombinedOutput()
if err != nil {
t.Fatalf("baseline build: %v\n%s", err, buildLog)
}
var work, linkLine, asmObj string
for line := range strings.SplitSeq(string(buildLog), "\n") {
switch {
case strings.HasPrefix(line, "WORK="):
work = strings.TrimPrefix(line, "WORK=")
case strings.Contains(line, "/asm ") && strings.Contains(line, "main_amd64.s") && !strings.Contains(line, "-gensymabis"):
asmObj = fieldAfter(line, "-o")
case strings.Contains(line, "/link ") && strings.Contains(line, "-importcfg"):
linkLine = line
}
}
if work == "" || asmObj == "" || linkLine == "" {
t.Skipf("could not parse build log (work=%q asmObj=%q link=%q)", work, asmObj, linkLine)
}
defer os.RemoveAll(work)
asmObj = strings.ReplaceAll(asmObj, "$WORK", work)
linkLine = strings.ReplaceAll(linkLine, "$WORK", work)
// Assemble the same source with gasm and substitute the object.
src, err := os.ReadFile(filepath.Join(dir, "main_amd64.s"))
if err != nil {
t.Fatal(err)
}
f, errs := parser.Parse("main_amd64.s", string(src))
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFile(f)
if err != nil {
t.Fatalf("AssembleFile: %v", err)
}
gasmObj, err := img.GOObject("dlink", "main_amd64.s")
if err != nil {
t.Fatalf("GOObject: %v", err)
}
if err := os.WriteFile(asmObj, gasmObj, 0o644); err != nil {
t.Fatalf("write gasm object: %v", err)
}
linkCmd := exec.Command("bash", "-c", "cd "+dir+" && "+linkLine)
if out, err := linkCmd.CombinedOutput(); err != nil {
t.Fatalf("re-link with gasm object: %v\n%s", err, out)
}
// The linked program must run and find the right function behind the
// data word.
out, err := exec.Command(filepath.Join(dir, "prog")).CombinedOutput()
if err != nil {
t.Fatalf("linked program failed: %v\n%s", err, out)
}
}
+872 -45
View File
File diff suppressed because it is too large Load Diff
+608 -6
View File
@@ -30,7 +30,10 @@ package asm
// of the immediate and register fields), mirroring the toolchain's OP_*
// helpers, so each l64* function only ORs its fields in.
import "maps"
import (
"maps"
"strings"
)
// loong64RegNum returns the 5-bit register number for a LoongArch register
// name: R0-R31 (integer), F0-F31 (floating point), FCC0-FCC7 (condition
@@ -103,7 +106,12 @@ func loong64RegNum(name string) int {
case "R31", "S8":
return 31
}
// F0-F31, FCC0-FCC7, FCSR0-FCSR31.
// F0-F31, FCC0-FCC7, FCSR0-FCSR31. The LSX/LASX vector banks (V0-V31,
// X0-X31) are deliberately NOT accepted here: they are a separate
// register class, and the toolchain rejects V/X names wherever an
// integer or FP register is expected (GOARCH=loong64 go tool asm reports
// "unrecognized instruction" for `BEQZ X0`). Vector operands are
// resolved only through loong64VecRegNum.
if len(name) >= 4 && name[:4] == "FCSR" {
return loong64RegSpecial(name[4:], 31)
}
@@ -148,6 +156,19 @@ func loong64RegSpecial(digits string, max int) int {
return -1
}
// loong64VecRegNum resolves an LSX/LASX vector register name (V0-V31 or
// X0-X31) to its 5-bit number, or -1. The vector banks are a register class
// of their own: the toolchain accepts them only in the vector operands of the
// LSX/LASX instructions (GOARCH=loong64 go tool asm assembles `VADDV V0, V1,
// V2` and `XVADDV X0, X1, X2`, and rejects `VADDV R4, R5, R6`), so the V/X
// spellings never reach the integer/FP resolver.
func loong64VecRegNum(name string) int {
if len(name) < 2 || (name[0] != 'V' && name[0] != 'X') {
return -1
}
return loong64RegSpecial(name[1:], 31)
}
// ---- format helpers ----
// l64rrr encodes a 3R instruction: op | rk<<10 | rj<<5 | rd.
@@ -247,7 +268,7 @@ const (
l64Firr14 // 2RI14 (ldptr/stptr)
l64Firr16 // 2RI16 (addu16i.d)
l64Fir20 // 2RI20 (lu12i.w, lu32i.d, pcalau12i, pcaddu12i)
l64Frrrr // 4R (fmadd/fmsub/fnmadd/fnmsub)
l64Frrrr // 4R (fmadd/fmsub/fnmadd/fnmsub, fsel)
l64Firir // bstrins/bstrpick
l64Firrr // alsl
l64Fi15 // syscall/break/dbar
@@ -255,6 +276,9 @@ const (
l64Frdtime // rdtime (rd at bits [9:5], rj at bits [4:0])
l64Fshift // 2RI12 with a 5/6-bit shift immediate
l64Fpreld // preld (2RI12 + 5-bit hint)
l64Fvvv // 3R vector (LSX/LASX): op | vk<<10 | vj<<5 | vd
l64Fvcf // vector-to-condition: op | subop<<10 | vj<<5 | fcc
l64Fvvvv // 4R vector shuffle: op | va<<15 | vk<<10 | vj<<5 | vd
)
// l64Enc is one instruction's encoding: its bit layout (format) and the
@@ -277,10 +301,72 @@ type l64DualEnc struct {
var l64DualTable = map[string]l64DualEnc{}
// l64InstrTable maps LoongArch mnemonics (as the Go assembler spells them)
// to their encoding. SIMD (LSX/LASX: V*/XV*) instructions are not covered
// yet; the base integer, memory and floating-point ISA is complete.
// to their encoding.
var l64InstrTable = map[string]l64Enc{}
// l64Vec3Enc pairs a vector opcode with its register bank: false = LSX
// (V0-V31), true = LASX (X0-X31). The toolchain accepts one bank per
// spelling: GOARCH=loong64 go tool asm assembles `VADDV V1, V2, V3` and
// `XVADDV X1, X2, X3`, and rejects the crossed spellings.
type l64Vec3Enc struct {
op uint32
lasx bool
}
// l64VecImmEnc carries the immediate-form encoding of a vector mnemonic:
// the opcode, the bank, the accepted immediate range, the bias the toolchain
// adds (vsrai.b encodes imm+8) and the mask of the encoded field (vseqi.b
// keeps a 5-bit two's-complement value, vseqi.d a 7-bit one).
type l64VecImmEnc struct {
op uint32
lasx bool
min, max int
bias int
mask int
}
// l64VecBank marks the LSX/LASX mnemonics and records which register bank
// each accepts; presence in the map routes the mnemonic through the vector
// dispatcher rather than the integer/FP formats.
var l64VecBank = map[string]bool{}
// l64VecImmInfo mirrors l64VecImmTable for the dispatcher.
var l64VecImmInfo = map[string]l64VecImmEnc{}
// l64Vec2R marks the two-operand vector mnemonics (INSTR vj, vd, such as
// vpcnt.v).
var l64Vec2R = map[string]bool{}
// l64Vec4R marks the four-operand vector mnemonics (INSTR va, vk, vj, vd,
// such as vshuf.b).
var l64Vec4R = map[string]bool{}
// l64VmovqOps holds the VMOVQ/XVMOVQ opcode constants (pre-shifted to bit
// 15), read off `go tool objdump` of GOARCH=loong64 `go tool asm` kernels.
type l64VmovqEnc struct {
ld, st, ldx, stx uint32 // plain and indexed load/store
replB, replH, replW, replD uint32 // vldrepl: load and replicate element
pickS, pickU uint32 // vpickve2gr.{,u} element extract
ins uint32 // vinsgr2vr element insert
dup uint32 // vreplgr2vr duplicate (width in [11:10])
move uint32 // vori.b/xvori.b $0 register move
}
var l64VmovqTable = map[bool]l64VmovqEnc{
false: { // VMOVQ, the LSX (V) bank
ld: 0x5800 << 15, st: 0x5880 << 15, ldx: 0x7080 << 15, stx: 0x7088 << 15,
replB: 0x6100 << 15, replH: 0x6080 << 15, replW: 0x6040 << 15, replD: 0x6020 << 15,
pickS: 0xE5DF << 15, pickU: 0xE5E7 << 15,
ins: 0xE5D7 << 15, dup: 0xE53E << 15, move: 0xE65A << 15,
},
true: { // XVMOVQ, the LASX (X) bank
ld: 0x5900 << 15, st: 0x5980 << 15, ldx: 0x7090 << 15, stx: 0x7098 << 15,
replB: 0x6500 << 15, replH: 0x6480 << 15, replW: 0x6440 << 15, replD: 0x6420 << 15,
pickS: 0xEDDF << 15, pickU: 0xEDE7 << 15,
ins: 0xEDD7 << 15, dup: 0xED3E << 15, move: 0xEE5A << 15,
},
}
func init() {
// 3R, integer.
rrr := map[string]uint32{
@@ -360,6 +446,18 @@ func init() {
"FTINTRZVF": 0x46a9 << 10, "FTINTRZVD": 0x46aa << 10,
"FTINTRNEWF": 0x46b1 << 10, "FTINTRNEWD": 0x46b2 << 10,
"FTINTRNEVF": 0x46b9 << 10, "FTINTRNEVD": 0x46ba << 10,
// LSX: convert a 64-bit integer lane to a double float. The operand
// bank is the FP registers (the toolchain spells it `FFINTDV F0, F1`),
// so the entry stays on the 2R integer/FP format.
"FFINTDV": 0x474a << 10,
// The rest of the scalar conversions (all F-bank, 2R).
"FFINTFW": 0x4744 << 10, // ffint.s.w
"FFINTFV": 0x4746 << 10, // ffint.s.l
"FFINTDW": 0x4748 << 10, // ffint.d.w
"FTINTWF": 0x46c1 << 10, // ftint.w.s
"FTINTWD": 0x46c2 << 10, // ftint.w.d
"FTINTVF": 0x46c9 << 10, // ftint.l.s
"FTINTVD": 0x46ca << 10, // ftint.l.d
}
for m, op := range rr {
l64InstrTable[m] = l64Enc{format: l64Frr, op: op}
@@ -416,12 +514,14 @@ func init() {
// LUI is the Plan 9 spelling of lu12i.w.
l64InstrTable["LUI"] = l64Enc{format: l64Fir20, op: 0x0a << 25}
// 4R, fused multiply-add.
// 4R, fused multiply-add, and FSEL (fsel.d: the first operand is a FCC
// condition flag, the layout matches the 4R shape).
rrrr := map[string]uint32{
"FMADDF": 0x81 << 20, "FMADDD": 0x82 << 20,
"FMSUBF": 0x85 << 20, "FMSUBD": 0x86 << 20,
"FNMADDF": 0x89 << 20, "FNMADDD": 0x8a << 20,
"FNMSUBF": 0x8d << 20, "FNMSUBD": 0x8e << 20,
"FSEL": 0x340 << 18,
}
for m, op := range rrrr {
l64InstrTable[m] = l64Enc{format: l64Frrrr, op: op}
@@ -455,6 +555,10 @@ func init() {
l64InstrTable["PRELD"] = l64Enc{format: l64Fpreld, op: 0x0ab << 22}
// Atomics, 3R with the AM field order (rk=value, rj=address, rd=result).
// The toolchain's form is three operands, `AMADDW rk, (rj), rd`
// (cmd/asm/internal/asm/testdata/loong64enc1.s and
// internal/runtime/atomic/atomic_loong64.s); the two-register spelling
// is rejected by the oracle.
am := map[string]uint32{
"AMSWAPB": 0x070B8 << 15, "AMSWAPH": 0x070B9 << 15,
"AMSWAPW": 0x070C0 << 15, "AMSWAPV": 0x070C1 << 15,
@@ -472,10 +576,508 @@ func init() {
"AMSWAPDBW": 0x070D2 << 15, "AMSWAPDBV": 0x070D3 << 15,
"AMCASDBB": 0x070B4 << 15, "AMCASDBH": 0x070B5 << 15,
"AMCASDBW": 0x070B6 << 15, "AMCASDBV": 0x070B7 << 15,
// The _dbar (acquire/release) add, and, or variants: opcodes read off
// `go tool objdump` of `AMADDDBW R14, (R13), R12` and friends.
"AMADDDBW": 0x070D4 << 15, "AMADDDBV": 0x070D5 << 15,
"AMANDDBW": 0x070D6 << 15, "AMANDDBV": 0x070D7 << 15,
"AMORDBW": 0x070D8 << 15, "AMORDBV": 0x070D9 << 15,
// The remaining _dbar exchange variants (loong64enc1.s).
"AMXORDBW": 0x070DA << 15, "AMXORDBV": 0x070DB << 15,
"AMMAXDBW": 0x070DC << 15, "AMMAXDBV": 0x070DD << 15,
"AMMINDBW": 0x070DE << 15, "AMMINDBV": 0x070DF << 15,
"AMMAXDBWU": 0x070E0 << 15, "AMMAXDBVU": 0x070E1 << 15,
"AMMINDBWU": 0x070E2 << 15, "AMMINDBVU": 0x070E3 << 15,
}
for m, op := range am {
l64InstrTable[m] = l64Enc{format: l64Fam, op: op}
}
// ---- LSX/LASX (V*/XV*) ----
// Every opcode below was read off `go tool objdump` of a GOARCH=loong64
// `go tool asm` kernel (the toolchain's own loong64enc1.s cross-checks
// most of them), not assumed from the LoongArch manual.
// Three vector registers: INSTR vk, vj, vd (or INSTR vk, vd with
// vj = vd). l64Vec3Enc.lasx selects the register bank the toolchain
// accepts: LSX spellings take V0-V31, LASX spellings X0-X31.
vec3 := map[string]l64Vec3Enc{
"VADDW": {0xE016 << 15, false}, "VADDV": {0xE017 << 15, false},
"VANDV": {0xE24C << 15, false}, "VXORV": {0xE24E << 15, false},
"VSEQB": {0xE000 << 15, false}, "VSEQV": {0xE003 << 15, false},
"VSRAB": {0xE1D8 << 15, false}, "VROTRW": {0xE1DE << 15, false},
"XVADDV": {0xE817 << 15, true},
"XVANDV": {0xEA4C << 15, true}, "XVXORV": {0xEA4E << 15, true},
"XVSEQB": {0xE800 << 15, true}, "XVSEQV": {0xE803 << 15, true},
}
// The integer and FP add/subtract families: [X]VADD and [X]VSUB by lane
// width, plus the [X]VSADD/[X]VSSUB saturating pairs.
// Opcodes transcribed from the toolchain's loong64enc1.s.
addsub := map[string]l64Vec3Enc{
"VADDB": {0xE014 << 15, false}, "VADDH": {0xE015 << 15, false},
"VADDD": {0xE262 << 15, false}, "VADDF": {0xE261 << 15, false},
"VADDQ": {0xE25A << 15, false},
"VSUBB": {0xE018 << 15, false}, "VSUBH": {0xE019 << 15, false},
"VSUBW": {0xE01A << 15, false}, "VSUBV": {0xE01B << 15, false},
"VSUBQ": {0xE25B << 15, false},
"VSUBF": {0xE265 << 15, false}, "VSUBD": {0xE266 << 15, false},
"VSADDB": {0xE08C << 15, false}, "VSADDH": {0xE08D << 15, false},
"VSADDW": {0xE08E << 15, false}, "VSADDV": {0xE08F << 15, false},
"VSADDBU": {0xE094 << 15, false}, "VSADDHU": {0xE095 << 15, false},
"VSADDWU": {0xE096 << 15, false}, "VSADDVU": {0xE097 << 15, false},
"VSSUBB": {0xE090 << 15, false}, "VSSUBH": {0xE091 << 15, false},
"VSSUBW": {0xE092 << 15, false}, "VSSUBV": {0xE093 << 15, false},
"VSSUBBU": {0xE098 << 15, false}, "VSSUBHU": {0xE099 << 15, false},
"VSSUBWU": {0xE09A << 15, false}, "VSSUBVU": {0xE09B << 15, false},
"XVADDB": {0xE814 << 15, true}, "XVADDH": {0xE815 << 15, true},
"XVADDW": {0xE816 << 15, true},
"XVADDD": {0xEA62 << 15, true}, "XVADDF": {0xEA61 << 15, true},
"XVADDQ": {0xEA5A << 15, true},
"XVSUBB": {0xE818 << 15, true}, "XVSUBH": {0xE819 << 15, true},
"XVSUBW": {0xE81A << 15, true}, "XVSUBV": {0xE81B << 15, true},
"XVSUBQ": {0xEA5B << 15, true},
"XVSUBF": {0xEA65 << 15, true}, "XVSUBD": {0xEA66 << 15, true},
"XVSADDB": {0xE88C << 15, true}, "XVSADDH": {0xE88D << 15, true},
"XVSADDW": {0xE88E << 15, true}, "XVSADDV": {0xE88F << 15, true},
"XVSADDBU": {0xE894 << 15, true}, "XVSADDHU": {0xE895 << 15, true},
"XVSADDWU": {0xE896 << 15, true}, "XVSADDVU": {0xE897 << 15, true},
"XVSSUBB": {0xE890 << 15, true}, "XVSSUBH": {0xE891 << 15, true},
"XVSSUBW": {0xE892 << 15, true}, "XVSSUBV": {0xE893 << 15, true},
"XVSSUBBU": {0xE898 << 15, true}, "XVSSUBHU": {0xE899 << 15, true},
"XVSSUBWU": {0xE89A << 15, true}, "XVSSUBVU": {0xE89B << 15, true},
}
// The multiply families: plain and high-half [X]VMUL/[X]VMUH, the
// widening [X]VMULW{EV,OD} ladder and its accumulating [X]VMADDW twins,
// plus the [X]VMADD/[X]VMSUB fused multiply-add and the [X]VDIV/[X]VMOD
// divide and modulo pairs.
muldiv := map[string]l64Vec3Enc{
"VMULB": {0xE108 << 15, false}, "VMULH": {0xE109 << 15, false},
"VMULW": {0xE10A << 15, false}, "VMULV": {0xE10B << 15, false},
"VMUHB": {0xE10C << 15, false}, "VMUHH": {0xE10D << 15, false},
"VMUHW": {0xE10E << 15, false}, "VMUHV": {0xE10F << 15, false},
"VMUHBU": {0xE110 << 15, false}, "VMUHHU": {0xE111 << 15, false},
"VMUHWU": {0xE112 << 15, false}, "VMUHVU": {0xE113 << 15, false},
"VMULWEVHB": {0xE120 << 15, false}, "VMULWEVWH": {0xE121 << 15, false},
"VMULWEVVW": {0xE122 << 15, false}, "VMULWEVQV": {0xE123 << 15, false},
"VMULWODHB": {0xE124 << 15, false}, "VMULWODWH": {0xE125 << 15, false},
"VMULWODVW": {0xE126 << 15, false}, "VMULWODQV": {0xE127 << 15, false},
"VMULWEVHBU": {0xE130 << 15, false}, "VMULWEVWHU": {0xE131 << 15, false},
"VMULWEVVWU": {0xE132 << 15, false}, "VMULWEVQVU": {0xE133 << 15, false},
"VMULWODHBU": {0xE134 << 15, false}, "VMULWODWHU": {0xE135 << 15, false},
"VMULWODVWU": {0xE136 << 15, false}, "VMULWODQVU": {0xE137 << 15, false},
"VMULWEVHBUB": {0xE140 << 15, false}, "VMULWEVWHUH": {0xE141 << 15, false},
"VMULWEVVWUW": {0xE142 << 15, false}, "VMULWEVQVUV": {0xE143 << 15, false},
"VMULWODHBUB": {0xE144 << 15, false}, "VMULWODWHUH": {0xE145 << 15, false},
"VMULWODVWUW": {0xE146 << 15, false}, "VMULWODQVUV": {0xE147 << 15, false},
"VMADDB": {0xE150 << 15, false}, "VMADDH": {0xE151 << 15, false},
"VMADDW": {0xE152 << 15, false}, "VMADDV": {0xE153 << 15, false},
"VMSUBB": {0xE154 << 15, false}, "VMSUBH": {0xE155 << 15, false},
"VMSUBW": {0xE156 << 15, false}, "VMSUBV": {0xE157 << 15, false},
"VMADDWEVHB": {0xE158 << 15, false}, "VMADDWEVWH": {0xE159 << 15, false},
"VMADDWEVVW": {0xE15A << 15, false}, "VMADDWEVQV": {0xE15B << 15, false},
"VMADDWODHB": {0xE15C << 15, false}, "VMADDWODWH": {0xE15D << 15, false},
"VMADDWODVW": {0xE15E << 15, false}, "VMADDWODQV": {0xE15F << 15, false},
"VMADDWEVHBU": {0xE168 << 15, false}, "VMADDWEVWHU": {0xE169 << 15, false},
"VMADDWEVVWU": {0xE16A << 15, false}, "VMADDWEVQVU": {0xE16B << 15, false},
"VMADDWODHBU": {0xE16C << 15, false}, "VMADDWODWHU": {0xE16D << 15, false},
"VMADDWODVWU": {0xE16E << 15, false}, "VMADDWODQVU": {0xE16F << 15, false},
"VMADDWEVHBUB": {0xE178 << 15, false}, "VMADDWEVWHUH": {0xE179 << 15, false},
"VMADDWEVVWUW": {0xE17A << 15, false}, "VMADDWEVQVUV": {0xE17B << 15, false},
"VMADDWODHBUB": {0xE17C << 15, false}, "VMADDWODWHUH": {0xE17D << 15, false},
"VMADDWODVWUW": {0xE17E << 15, false}, "VMADDWODQVUV": {0xE17F << 15, false},
"VDIVB": {0xE1C0 << 15, false}, "VDIVH": {0xE1C1 << 15, false},
"VDIVW": {0xE1C2 << 15, false}, "VDIVV": {0xE1C3 << 15, false},
"VMODB": {0xE1C4 << 15, false}, "VMODH": {0xE1C5 << 15, false},
"VMODW": {0xE1C6 << 15, false}, "VMODV": {0xE1C7 << 15, false},
"VDIVBU": {0xE1C8 << 15, false}, "VDIVHU": {0xE1C9 << 15, false},
"VDIVWU": {0xE1CA << 15, false}, "VDIVVU": {0xE1CB << 15, false},
"VMODBU": {0xE1CC << 15, false}, "VMODHU": {0xE1CD << 15, false},
"VMODWU": {0xE1CE << 15, false}, "VMODVU": {0xE1CF << 15, false},
"VMULF": {0xE271 << 15, false}, "VMULD": {0xE272 << 15, false},
"VDIVF": {0xE275 << 15, false}, "VDIVD": {0xE276 << 15, false},
"XVMULB": {0xE908 << 15, true}, "XVMULH": {0xE909 << 15, true},
"XVMULW": {0xE90A << 15, true}, "XVMULV": {0xE90B << 15, true},
"XVMUHB": {0xE90C << 15, true}, "XVMUHH": {0xE90D << 15, true},
"XVMUHW": {0xE90E << 15, true}, "XVMUHV": {0xE90F << 15, true},
"XVMUHBU": {0xE910 << 15, true}, "XVMUHHU": {0xE911 << 15, true},
"XVMUHWU": {0xE912 << 15, true}, "XVMUHVU": {0xE913 << 15, true},
"XVMULWEVHB": {0xE920 << 15, true}, "XVMULWEVWH": {0xE921 << 15, true},
"XVMULWEVVW": {0xE922 << 15, true}, "XVMULWEVQV": {0xE923 << 15, true},
"XVMULWODHB": {0xE924 << 15, true}, "XVMULWODWH": {0xE925 << 15, true},
"XVMULWODVW": {0xE926 << 15, true}, "XVMULWODQV": {0xE927 << 15, true},
"XVMULWEVHBU": {0xE930 << 15, true}, "XVMULWEVWHU": {0xE931 << 15, true},
"XVMULWEVVWU": {0xE932 << 15, true}, "XVMULWEVQVU": {0xE933 << 15, true},
"XVMULWODHBU": {0xE934 << 15, true}, "XVMULWODWHU": {0xE935 << 15, true},
"XVMULWODVWU": {0xE936 << 15, true}, "XVMULWODQVU": {0xE937 << 15, true},
"XVMULWEVHBUB": {0xE940 << 15, true}, "XVMULWEVWHUH": {0xE941 << 15, true},
"XVMULWEVVWUW": {0xE942 << 15, true}, "XVMULWEVQVUV": {0xE943 << 15, true},
"XVMULWODHBUB": {0xE944 << 15, true}, "XVMULWODWHUH": {0xE945 << 15, true},
"XVMULWODVWUW": {0xE946 << 15, true}, "XVMULWODQVUV": {0xE947 << 15, true},
"XVMADDB": {0xE950 << 15, true}, "XVMADDH": {0xE951 << 15, true},
"XVMADDW": {0xE952 << 15, true}, "XVMADDV": {0xE953 << 15, true},
"XVMSUBB": {0xE954 << 15, true}, "XVMSUBH": {0xE955 << 15, true},
"XVMSUBW": {0xE956 << 15, true}, "XVMSUBV": {0xE957 << 15, true},
"XVMADDWEVHB": {0xE958 << 15, true}, "XVMADDWEVWH": {0xE959 << 15, true},
"XVMADDWEVVW": {0xE95A << 15, true}, "XVMADDWEVQV": {0xE95B << 15, true},
"XVMADDWODHB": {0xE95C << 15, true}, "XVMADDWODWH": {0xE95D << 15, true},
"XVMADDWODVW": {0xE95E << 15, true}, "XVMADDWODQV": {0xE95F << 15, true},
"XVMADDWEVHBU": {0xE968 << 15, true}, "XVMADDWEVWHU": {0xE969 << 15, true},
"XVMADDWEVVWU": {0xE96A << 15, true}, "XVMADDWEVQVU": {0xE96B << 15, true},
"XVMADDWODHBU": {0xE96C << 15, true}, "XVMADDWODWHU": {0xE96D << 15, true},
"XVMADDWODVWU": {0xE96E << 15, true}, "XVMADDWODQVU": {0xE96F << 15, true},
"XVMADDWEVHBUB": {0xE978 << 15, true}, "XVMADDWEVWHUH": {0xE979 << 15, true},
"XVMADDWEVVWUW": {0xE97A << 15, true}, "XVMADDWEVQVUV": {0xE97B << 15, true},
"XVMADDWODHBUB": {0xE97C << 15, true}, "XVMADDWODWHUH": {0xE97D << 15, true},
"XVMADDWODVWUW": {0xE97E << 15, true}, "XVMADDWODQVUV": {0xE97F << 15, true},
"XVDIVB": {0xE9C0 << 15, true}, "XVDIVH": {0xE9C1 << 15, true},
"XVDIVW": {0xE9C2 << 15, true}, "XVDIVV": {0xE9C3 << 15, true},
"XVMODB": {0xE9C4 << 15, true}, "XVMODH": {0xE9C5 << 15, true},
"XVMODW": {0xE9C6 << 15, true}, "XVMODV": {0xE9C7 << 15, true},
"XVDIVBU": {0xE9C8 << 15, true}, "XVDIVHU": {0xE9C9 << 15, true},
"XVDIVWU": {0xE9CA << 15, true}, "XVDIVVU": {0xE9CB << 15, true},
"XVMODBU": {0xE9CC << 15, true}, "XVMODHU": {0xE9CD << 15, true},
"XVMODWU": {0xE9CE << 15, true}, "XVMODVU": {0xE9CF << 15, true},
"XVMULF": {0xEA71 << 15, true}, "XVMULD": {0xEA72 << 15, true},
"XVDIVF": {0xEA75 << 15, true}, "XVDIVD": {0xEA76 << 15, true},
}
// The lane-wise shifts and rotates (three-register forms; the immediate
// forms live in l64VecImmInfo), the interleave families, the bit
// clear/set/rev register forms, the remaining logic and compare
// spellings, the widening add/subtract ladder and the vector FP
// arithmetic.
vecmisc := map[string]l64Vec3Enc{
"VSLLB": {0xE1D0 << 15, false}, "VSLLH": {0xE1D1 << 15, false},
"VSLLW": {0xE1D2 << 15, false}, "VSLLV": {0xE1D3 << 15, false},
"VSRLB": {0xE1D4 << 15, false}, "VSRLH": {0xE1D5 << 15, false},
"VSRLW": {0xE1D6 << 15, false}, "VSRLV": {0xE1D7 << 15, false},
"VSRAH": {0xE1D9 << 15, false}, "VSRAW": {0xE1DA << 15, false},
"VSRAV": {0xE1DB << 15, false},
"VROTRB": {0xE1DC << 15, false}, "VROTRH": {0xE1DD << 15, false},
"VROTRV": {0xE1DF << 15, false},
"VILVLB": {0xE234 << 15, false}, "VILVLH": {0xE235 << 15, false},
"VILVLW": {0xE236 << 15, false}, "VILVLV": {0xE237 << 15, false},
"VILVHB": {0xE238 << 15, false}, "VILVHH": {0xE239 << 15, false},
"VILVHW": {0xE23A << 15, false}, "VILVHV": {0xE23B << 15, false},
"VBITCLRB": {0xE218 << 15, false}, "VBITCLRH": {0xE219 << 15, false},
"VBITCLRW": {0xE21A << 15, false}, "VBITCLRV": {0xE21B << 15, false},
"VBITSETB": {0xE21C << 15, false}, "VBITSETH": {0xE21D << 15, false},
"VBITSETW": {0xE21E << 15, false}, "VBITSETV": {0xE21F << 15, false},
"VBITREVB": {0xE220 << 15, false}, "VBITREVH": {0xE221 << 15, false},
"VBITREVW": {0xE222 << 15, false}, "VBITREVV": {0xE223 << 15, false},
"VORV": {0xE24D << 15, false}, "VNORV": {0xE24F << 15, false},
"VANDNV": {0xE250 << 15, false}, "VORNV": {0xE251 << 15, false},
"VSEQH": {0xE001 << 15, false}, "VSEQW": {0xE002 << 15, false},
"VSLTB": {0xE00C << 15, false}, "VSLTH": {0xE00D << 15, false},
"VSLTW": {0xE00E << 15, false}, "VSLTV": {0xE00F << 15, false},
"VSLTBU": {0xE010 << 15, false}, "VSLTHU": {0xE011 << 15, false},
"VSLTWU": {0xE012 << 15, false}, "VSLTVU": {0xE013 << 15, false},
"VADDWEVHB": {0xE03C << 15, false}, "VADDWEVWH": {0xE03D << 15, false},
"VADDWEVVW": {0xE03E << 15, false}, "VADDWEVQV": {0xE03F << 15, false},
"VSUBWEVHB": {0xE040 << 15, false}, "VSUBWEVWH": {0xE041 << 15, false},
"VSUBWEVVW": {0xE042 << 15, false}, "VSUBWEVQV": {0xE043 << 15, false},
"VADDWODHB": {0xE044 << 15, false}, "VADDWODWH": {0xE045 << 15, false},
"VADDWODVW": {0xE046 << 15, false}, "VADDWODQV": {0xE047 << 15, false},
"VSUBWODHB": {0xE048 << 15, false}, "VSUBWODWH": {0xE049 << 15, false},
"VSUBWODVW": {0xE04A << 15, false}, "VSUBWODQV": {0xE04B << 15, false},
"VSUBWEVHBU": {0xE060 << 15, false}, "VSUBWEVWHU": {0xE061 << 15, false},
"VSUBWEVVWU": {0xE062 << 15, false}, "VSUBWEVQVU": {0xE063 << 15, false},
"VADDWEVHBU": {0xE05C << 15, false}, "VADDWEVWHU": {0xE05D << 15, false},
"VADDWEVVWU": {0xE05E << 15, false}, "VADDWEVQVU": {0xE05F << 15, false},
"VADDWODHBU": {0xE064 << 15, false}, "VADDWODWHU": {0xE065 << 15, false},
"VADDWODVWU": {0xE066 << 15, false}, "VADDWODQVU": {0xE067 << 15, false},
"VSUBWODHBU": {0xE068 << 15, false}, "VSUBWODWHU": {0xE069 << 15, false},
"VSUBWODVWU": {0xE06A << 15, false}, "VSUBWODQVU": {0xE06B << 15, false},
"VSHUFH": {0xE2F5 << 15, false}, "VSHUFW": {0xE2F6 << 15, false},
"VSHUFV": {0xE2F7 << 15, false},
"XVSLLB": {0xE9D0 << 15, true}, "XVSLLH": {0xE9D1 << 15, true},
"XVSLLW": {0xE9D2 << 15, true}, "XVSLLV": {0xE9D3 << 15, true},
"XVSRLB": {0xE9D4 << 15, true}, "XVSRLH": {0xE9D5 << 15, true},
"XVSRLW": {0xE9D6 << 15, true}, "XVSRLV": {0xE9D7 << 15, true},
"XVSRAB": {0xE9D8 << 15, true}, "XVSRAH": {0xE9D9 << 15, true},
"XVSRAW": {0xE9DA << 15, true}, "XVSRAV": {0xE9DB << 15, true},
"XVROTRB": {0xE9DC << 15, true}, "XVROTRH": {0xE9DD << 15, true},
"XVROTRW": {0xE9DE << 15, true}, "XVROTRV": {0xE9DF << 15, true},
"XVILVLB": {0xEA34 << 15, true}, "XVILVLH": {0xEA35 << 15, true},
"XVILVLW": {0xEA36 << 15, true}, "XVILVLV": {0xEA37 << 15, true},
"XVILVHB": {0xEA38 << 15, true}, "XVILVHH": {0xEA39 << 15, true},
"XVILVHW": {0xEA3A << 15, true}, "XVILVHV": {0xEA3B << 15, true},
"XVBITCLRB": {0xEA18 << 15, true}, "XVBITCLRH": {0xEA19 << 15, true},
"XVBITCLRW": {0xEA1A << 15, true}, "XVBITCLRV": {0xEA1B << 15, true},
"XVBITSETB": {0xEA1C << 15, true}, "XVBITSETH": {0xEA1D << 15, true},
"XVBITSETW": {0xEA1E << 15, true}, "XVBITSETV": {0xEA1F << 15, true},
"XVBITREVB": {0xEA20 << 15, true}, "XVBITREVH": {0xEA21 << 15, true},
"XVBITREVW": {0xEA22 << 15, true}, "XVBITREVV": {0xEA23 << 15, true},
"XVORV": {0xEA4D << 15, true}, "XVNORV": {0xEA4F << 15, true},
"XVANDNV": {0xEA50 << 15, true}, "XVORNV": {0xEA51 << 15, true},
"XVSEQH": {0xE801 << 15, true}, "XVSEQW": {0xE802 << 15, true},
"XVSLTB": {0xE80C << 15, true}, "XVSLTH": {0xE80D << 15, true},
"XVSLTW": {0xE80E << 15, true}, "XVSLTV": {0xE80F << 15, true},
"XVSLTBU": {0xE810 << 15, true}, "XVSLTHU": {0xE811 << 15, true},
"XVSLTWU": {0xE812 << 15, true}, "XVSLTVU": {0xE813 << 15, true},
"XVADDWEVHB": {0xE83C << 15, true}, "XVADDWEVWH": {0xE83D << 15, true},
"XVADDWEVVW": {0xE83E << 15, true}, "XVADDWEVQV": {0xE83F << 15, true},
"XVSUBWEVHB": {0xE840 << 15, true}, "XVSUBWEVWH": {0xE841 << 15, true},
"XVSUBWEVVW": {0xE842 << 15, true}, "XVSUBWEVQV": {0xE843 << 15, true},
"XVADDWODHB": {0xE844 << 15, true}, "XVADDWODWH": {0xE845 << 15, true},
"XVADDWODVW": {0xE846 << 15, true}, "XVADDWODQV": {0xE847 << 15, true},
"XVSUBWODHB": {0xE848 << 15, true}, "XVSUBWODWH": {0xE849 << 15, true},
"XVSUBWODVW": {0xE84A << 15, true}, "XVSUBWODQV": {0xE84B << 15, true},
"XVADDWEVHBU": {0xE85C << 15, true}, "XVADDWEVWHU": {0xE85D << 15, true},
"XVADDWEVVWU": {0xE85E << 15, true}, "XVADDWEVQVU": {0xE85F << 15, true},
"XVSUBWEVHBU": {0xE860 << 15, true}, "XVSUBWEVWHU": {0xE861 << 15, true},
"XVSUBWEVVWU": {0xE862 << 15, true}, "XVSUBWEVQVU": {0xE863 << 15, true},
"XVADDWODHBU": {0xE864 << 15, true}, "XVADDWODWHU": {0xE865 << 15, true},
"XVADDWODVWU": {0xE866 << 15, true}, "XVADDWODQVU": {0xE867 << 15, true},
"XVSUBWODHBU": {0xE868 << 15, true}, "XVSUBWODWHU": {0xE869 << 15, true},
"XVSUBWODVWU": {0xE86A << 15, true}, "XVSUBWODQVU": {0xE86B << 15, true},
"XVSHUFH": {0xEAF5 << 15, true}, "XVSHUFW": {0xEAF6 << 15, true},
"XVSHUFV": {0xEAF7 << 15, true},
}
for _, tab := range []map[string]l64Vec3Enc{addsub, muldiv, vecmisc} {
for m, e := range tab {
if _, dup := vec3[m]; dup {
panic("loong64: duplicate vector mnemonic " + m)
}
vec3[m] = e
}
}
for m, e := range vec3 {
l64InstrTable[m] = l64Enc{format: l64Fvvv, op: e.op}
l64VecBank[m] = e.lasx
}
// Immediate forms: INSTR $imm, vj, vd (or INSTR $imm, vd). The immediate
// range, bias and field mask are the ones the toolchain encodes: vandi.b
// stores the raw 8-bit constant, vsrari.b stores imm+8 (lane-width
// bias), the si5 compares store 5-bit two's-complement values and vseqi.d
// a 7-bit field the toolchain range-checks down to si5.
// The mnemonics that also have a register form (the shifts, the bit
// clear/set/rev families, VSEQ and the logic immediates) keep their
// three-register entry in l64InstrTable; the dispatcher picks the
// immediate opcode from l64VecImmInfo by operand kind, so the immediate
// entries must not overwrite the table.
vecImm := map[string]l64VecImmEnc{
"VANDB": {0xE7A0 << 15, false, 0, 255, 0, 0xFF},
"XVANDB": {0xEFA0 << 15, true, 0, 255, 0, 0xFF},
"VORB": {0xE7A8 << 15, false, 0, 255, 0, 0xFF},
"XVORB": {0xEFA8 << 15, true, 0, 255, 0, 0xFF},
"VXORB": {0xE7B0 << 15, false, 0, 255, 0, 0xFF},
"XVXORB": {0xEFB0 << 15, true, 0, 255, 0, 0xFF},
"VNORB": {0xE7B8 << 15, false, 0, 255, 0, 0xFF},
"XVNORB": {0xEFB8 << 15, true, 0, 255, 0, 0xFF},
"VSEQB": {0xE500 << 15, false, -16, 15, 0, 0x1F},
"XVSEQB": {0xE900 << 15, true, -16, 15, 0, 0x1F},
// vseqi.h/w accept the same si5 window as vseqi.b; vseqi.d carries a
// 7-bit field, but the toolchain range-checks it down to si5 as well
// (GOARCH=loong64 go tool asm rejects VSEQV $32 and VSEQV $-64).
"VSEQH": {0xE501 << 15, false, -16, 15, 0, 0x1F},
"XVSEQH": {0xED01 << 15, true, -16, 15, 0, 0x1F},
"VSEQW": {0xE502 << 15, false, -16, 15, 0, 0x1F},
"XVSEQW": {0xED02 << 15, true, -16, 15, 0, 0x1F},
"VSEQV": {0xE503 << 15, false, -16, 15, 0, 0x7F},
"XVSEQV": {0xE903 << 15, true, -16, 15, 0, 0x7F},
// vslti compares against a signed (or, in the U spellings, unsigned)
// si5/ui5 constant.
"VSLTB": {0xE50C << 15, false, -16, 15, 0, 0x1F},
"XVSLTB": {0xED0C << 15, true, -16, 15, 0, 0x1F},
"VSLTH": {0xE50D << 15, false, -16, 15, 0, 0x1F},
"XVSLTH": {0xED0D << 15, true, -16, 15, 0, 0x1F},
"VSLTW": {0xE50E << 15, false, -16, 15, 0, 0x1F},
"XVSLTW": {0xED0E << 15, true, -16, 15, 0, 0x1F},
"VSLTV": {0xE50F << 15, false, -16, 15, 0, 0x1F},
"XVSLTV": {0xED0F << 15, true, -16, 15, 0, 0x1F},
"VSLTBU": {0xE510 << 15, false, 0, 31, 0, 0x1F},
"XVSLTBU": {0xED10 << 15, true, 0, 31, 0, 0x1F},
"VSLTHU": {0xE511 << 15, false, 0, 31, 0, 0x1F},
"XVSLTHU": {0xED11 << 15, true, 0, 31, 0, 0x1F},
"VSLTWU": {0xE512 << 15, false, 0, 31, 0, 0x1F},
"XVSLTWU": {0xED12 << 15, true, 0, 31, 0, 0x1F},
"VSLTVU": {0xE513 << 15, false, 0, 31, 0, 0x1F},
"XVSLTVU": {0xED13 << 15, true, 0, 31, 0, 0x1F},
// vaddi/vsubi take ui5 constants for every width on this toolchain
// (VADDVU $32 is rejected by the oracle although the field is ui8).
"VADDBU": {0xE514 << 15, false, 0, 31, 0, 0x1F},
"XVADDBU": {0xED14 << 15, true, 0, 31, 0, 0x1F},
"VADDHU": {0xE515 << 15, false, 0, 31, 0, 0x1F},
"XVADDHU": {0xED15 << 15, true, 0, 31, 0, 0x1F},
"VADDWU": {0xE516 << 15, false, 0, 31, 0, 0x1F},
"XVADDWU": {0xED16 << 15, true, 0, 31, 0, 0x1F},
"VADDVU": {0xE517 << 15, false, 0, 31, 0, 0x1F},
"XVADDVU": {0xED17 << 15, true, 0, 31, 0, 0x1F},
"VSUBBU": {0xE518 << 15, false, 0, 31, 0, 0x1F},
"XVSUBBU": {0xED18 << 15, true, 0, 31, 0, 0x1F},
"VSUBHU": {0xE519 << 15, false, 0, 31, 0, 0x1F},
"XVSUBHU": {0xED19 << 15, true, 0, 31, 0, 0x1F},
"VSUBWU": {0xE51A << 15, false, 0, 31, 0, 0x1F},
"XVSUBWU": {0xED1A << 15, true, 0, 31, 0, 0x1F},
"VSUBVU": {0xE51B << 15, false, 0, 31, 0, 0x1F},
"XVSUBVU": {0xED1B << 15, true, 0, 31, 0, 0x1F},
// The shift/rotate immediates ride in a width-sized field whose upper
// bits carry the lane-width code: vslli.b stores ui3 at [12:0] with
// bits [14:13] inside the opcode, vslli.h ui4 under a 4 bit mask, and
// the .w/.d spellings a raw ui5/ui6.
"VSLLB": {0x732C2000, false, 0, 7, 0, 0x7},
"XVSLLB": {0x772C2000, true, 0, 7, 0, 0x7},
"VSLLH": {0x732C4000, false, 0, 15, 0, 0xF},
"XVSLLH": {0x772C4000, true, 0, 15, 0, 0xF},
"VSLLW": {0xE659 << 15, false, 0, 31, 0, 0x1F},
"XVSLLW": {0xEE59 << 15, true, 0, 31, 0, 0x1F},
"VSLLV": {0xE65A << 15, false, 0, 63, 0, 0x3F},
"XVSLLV": {0xEE5A << 15, true, 0, 63, 0, 0x3F},
"VSRLB": {0x73302000, false, 0, 7, 0, 0x7},
"XVSRLB": {0x77302000, true, 0, 7, 0, 0x7},
"VSRLH": {0x73304000, false, 0, 15, 0, 0xF},
"XVSRLH": {0x77304000, true, 0, 15, 0, 0xF},
"VSRLW": {0xE661 << 15, false, 0, 31, 0, 0x1F},
"XVSRLW": {0xEE61 << 15, true, 0, 31, 0, 0x1F},
"VSRLV": {0xE662 << 15, false, 0, 63, 0, 0x3F},
"XVSRLV": {0xEE62 << 15, true, 0, 63, 0, 0x3F},
// vsrari/vrotri bias the field so the lane-width code rides above the
// shift amount (.b adds 8, .h 16, .w 32; .d is a raw ui6).
"VSRAB": {0xE668 << 15, false, 0, 7, 8, 0x1F},
"XVSRAB": {0xEE68 << 15, true, 0, 7, 8, 0x1F},
"VSRAH": {0x73344000, false, 0, 15, 0, 0xF},
"XVSRAH": {0x77344000, true, 0, 15, 0, 0xF},
"VSRAW": {0xE669 << 15, false, 0, 31, 0, 0x1F},
"XVSRAW": {0xEE69 << 15, true, 0, 31, 0, 0x1F},
"VSRAV": {0xE66A << 15, false, 0, 63, 0, 0x3F},
"XVSRAV": {0xEE6A << 15, true, 0, 63, 0, 0x3F},
"VROTRB": {0x72A02000, false, 0, 7, 0, 0x7},
"XVROTRB": {0x76A02000, true, 0, 7, 0, 0x7},
"VROTRH": {0x72A04000, false, 0, 15, 0, 0xF},
"XVROTRH": {0x76A04000, true, 0, 15, 0, 0xF},
"VROTRW": {0xE541 << 15, false, 0, 31, 0, 0x1F},
"XVROTRW": {0xED41 << 15, true, 0, 31, 0, 0x1F},
"VROTRV": {0xE542 << 15, false, 0, 63, 0, 0x3F},
"XVROTRV": {0xED42 << 15, true, 0, 63, 0, 0x3F},
// vbitclri/vbitseti/vbitrevi follow the same width-coded layout.
"VBITCLRB": {0x73102000, false, 0, 7, 0, 0x7},
"XVBITCLRB": {0x77102000, true, 0, 7, 0, 0x7},
"VBITCLRH": {0x73104000, false, 0, 15, 0, 0xF},
"XVBITCLRH": {0x77104000, true, 0, 15, 0, 0xF},
"VBITCLRW": {0xE621 << 15, false, 0, 31, 0, 0x1F},
"XVBITCLRW": {0xEE21 << 15, true, 0, 31, 0, 0x1F},
"VBITCLRV": {0xE622 << 15, false, 0, 63, 0, 0x3F},
"XVBITCLRV": {0xEE22 << 15, true, 0, 63, 0, 0x3F},
"VBITSETB": {0x73142000, false, 0, 7, 0, 0x7},
"XVBITSETB": {0x77142000, true, 0, 7, 0, 0x7},
"VBITSETH": {0x73144000, false, 0, 15, 0, 0xF},
"XVBITSETH": {0x77144000, true, 0, 15, 0, 0xF},
"VBITSETW": {0xE629 << 15, false, 0, 31, 0, 0x1F},
"XVBITSETW": {0xEE29 << 15, true, 0, 31, 0, 0x1F},
"VBITSETV": {0xE62A << 15, false, 0, 63, 0, 0x3F},
"XVBITSETV": {0xEE2A << 15, true, 0, 63, 0, 0x3F},
"VBITREVB": {0x73182000, false, 0, 7, 0, 0x7},
"XVBITREVB": {0x77182000, true, 0, 7, 0, 0x7},
"VBITREVH": {0x73184000, false, 0, 15, 0, 0xF},
"XVBITREVH": {0x77184000, true, 0, 15, 0, 0xF},
"VBITREVW": {0xE631 << 15, false, 0, 31, 0, 0x1F},
"XVBITREVW": {0xEE31 << 15, true, 0, 31, 0, 0x1F},
"VBITREVV": {0xE632 << 15, false, 0, 63, 0, 0x3F},
"XVBITREVV": {0xEE32 << 15, true, 0, 63, 0, 0x3F},
// The 4-bit-select shuffles and the byte-extract/insert permutations
// take ui8 (the .d shuffle ui4 range-checked to 0..15 by the
// toolchain) packing both position nibbles.
"VSHUF4IB": {0xE720 << 15, false, 0, 255, 0, 0xFF},
"XVSHUF4IB": {0xEF20 << 15, true, 0, 255, 0, 0xFF},
"VSHUF4IH": {0xE728 << 15, false, 0, 255, 0, 0xFF},
"XVSHUF4IH": {0xEF28 << 15, true, 0, 255, 0, 0xFF},
"VSHUF4IW": {0xE730 << 15, false, 0, 255, 0, 0xFF},
"XVSHUF4IW": {0xEF30 << 15, true, 0, 255, 0, 0xFF},
"VSHUF4IV": {0xE738 << 15, false, 0, 15, 0, 0xFF},
"XVSHUF4IV": {0xEF38 << 15, true, 0, 15, 0, 0xFF},
"VPERMIW": {0xE7C8 << 15, false, 0, 255, 0, 0xFF},
"XVPERMIW": {0xEFC8 << 15, true, 0, 255, 0, 0xFF},
"XVPERMIV": {0xEFD0 << 15, true, 0, 255, 0, 0xFF},
"XVPERMIQ": {0xEFD8 << 15, true, 0, 255, 0, 0xFF},
"VEXTRINSB": {0xE718 << 15, false, 0, 255, 0, 0xFF},
"XVEXTRINSB": {0xEF18 << 15, true, 0, 255, 0, 0xFF},
"VEXTRINSH": {0xE710 << 15, false, 0, 255, 0, 0xFF},
"XVEXTRINSH": {0xEF10 << 15, true, 0, 255, 0, 0xFF},
"VEXTRINSW": {0xE708 << 15, false, 0, 255, 0, 0xFF},
"XVEXTRINSW": {0xEF08 << 15, true, 0, 255, 0, 0xFF},
"VEXTRINSV": {0xE700 << 15, false, 0, 255, 0, 0xFF},
"XVEXTRINSV": {0xEF00 << 15, true, 0, 255, 0, 0xFF},
}
for m, e := range vecImm {
l64VecImmInfo[m] = e
l64VecBank[m] = e.lasx
}
// Vector-to-condition flag: INSTR vj, FCCn (vsetnez.v, vsetanyeqz.*,
// vsetallnez.*): the sub-op rides in the rk field.
vecCf := map[string]uint32{
"VSETNEV": 0xE539<<15 | 7<<10, "XVSETNEV": 0xED39<<15 | 7<<10,
"VSETANYEQB": 0xE539<<15 | 8<<10, "XVSETANYEQB": 0xED39<<15 | 8<<10,
"VSETANYEQV": 0xE539<<15 | 11<<10, "XVSETANYEQV": 0xED39<<15 | 11<<10,
"VSETALLNEV": 0xE539<<15 | 15<<10, "XVSETALLNEV": 0xED39<<15 | 15<<10,
"VSETEQV": 0xE539<<15 | 6<<10, "XVSETEQV": 0xED39<<15 | 6<<10,
"VSETANYEQH": 0xE539<<15 | 9<<10, "XVSETANYEQH": 0xED39<<15 | 9<<10,
"VSETANYEQW": 0xE539<<15 | 10<<10, "XVSETANYEQW": 0xED39<<15 | 10<<10,
"VSETALLNEB": 0xE539<<15 | 12<<10, "XVSETALLNEB": 0xED39<<15 | 12<<10,
"VSETALLNEH": 0xE539<<15 | 13<<10, "XVSETALLNEH": 0xED39<<15 | 13<<10,
"VSETALLNEW": 0xE539<<15 | 14<<10, "XVSETALLNEW": 0xED39<<15 | 14<<10,
}
for m, op := range vecCf {
l64InstrTable[m] = l64Enc{format: l64Fvcf, op: op}
l64VecBank[m] = strings.HasPrefix(m, "XV")
}
// Lane popcount and the two-operand vector FP/unary spellings: INSTR vj,
// vd (the 2R layout with the opcode extending over the unused vk field;
// the low byte of each constant is the instruction's own sub-op).
vec2r := map[string]l64Vec3Enc{
"VPCNTV": {0x1CA70B << 10, false}, "XVPCNTV": {0x1DA70B << 10, true},
}
// The rest of the lane popcounts, the vector negations and the vector FP
// unary conversions (loong64enc1.s).
vec2rMore := map[string]l64Vec3Enc{
"VPCNTB": {0x1CA708 << 10, false}, "VPCNTH": {0x1CA709 << 10, false},
"VPCNTW": {0x1CA70A << 10, false},
"VNEGB": {0x1CA70C << 10, false}, "VNEGH": {0x1CA70D << 10, false},
"VNEGW": {0x1CA70E << 10, false}, "VNEGV": {0x1CA70F << 10, false},
"VFCLASSF": {0x1CA735 << 10, false}, "VFCLASSD": {0x1CA736 << 10, false},
"VFSQRTF": {0x1CA739 << 10, false}, "VFSQRTD": {0x1CA73A << 10, false},
"VFRECIPF": {0x1CA73D << 10, false}, "VFRECIPD": {0x1CA73E << 10, false},
"VFRSQRTF": {0x1CA741 << 10, false}, "VFRSQRTD": {0x1CA742 << 10, false},
"VFRINTF": {0x1CA74D << 10, false}, "VFRINTD": {0x1CA74E << 10, false},
"VFRINTRMF": {0x1CA751 << 10, false}, "VFRINTRMD": {0x1CA752 << 10, false},
"VFRINTRPF": {0x1CA755 << 10, false}, "VFRINTRPD": {0x1CA756 << 10, false},
"VFRINTRZF": {0x1CA759 << 10, false}, "VFRINTRZD": {0x1CA75A << 10, false},
"VFRINTRNEF": {0x1CA75D << 10, false}, "VFRINTRNED": {0x1CA75E << 10, false},
"XVPCNTB": {0x1DA708 << 10, true}, "XVPCNTH": {0x1DA709 << 10, true},
"XVPCNTW": {0x1DA70A << 10, true},
"XVNEGB": {0x1DA70C << 10, true}, "XVNEGH": {0x1DA70D << 10, true},
"XVNEGW": {0x1DA70E << 10, true}, "XVNEGV": {0x1DA70F << 10, true},
"XVFCLASSF": {0x1DA735 << 10, true}, "XVFCLASSD": {0x1DA736 << 10, true},
"XVFSQRTF": {0x1DA739 << 10, true}, "XVFSQRTD": {0x1DA73A << 10, true},
"XVFRECIPF": {0x1DA73D << 10, true}, "XVFRECIPD": {0x1DA73E << 10, true},
"XVFRSQRTF": {0x1DA741 << 10, true}, "XVFRSQRTD": {0x1DA742 << 10, true},
"XVFRINTF": {0x1DA74D << 10, true}, "XVFRINTD": {0x1DA74E << 10, true},
"XVFRINTRMF": {0x1DA751 << 10, true}, "XVFRINTRMD": {0x1DA752 << 10, true},
"XVFRINTRPF": {0x1DA755 << 10, true}, "XVFRINTRPD": {0x1DA756 << 10, true},
"XVFRINTRZF": {0x1DA759 << 10, true}, "XVFRINTRZD": {0x1DA75A << 10, true},
"XVFRINTRNEF": {0x1DA75D << 10, true}, "XVFRINTRNED": {0x1DA75E << 10, true},
}
maps.Copy(vec2r, vec2rMore)
for m, e := range vec2r {
l64InstrTable[m] = l64Enc{format: l64Frr, op: e.op}
l64VecBank[m] = e.lasx
l64Vec2R[m] = true
}
// The four-register byte shuffle: INSTR va, vk, vj, vd (the operand the
// table reads in each field position, va at bits [19:15]).
vec4r := map[string]l64Vec3Enc{
"VSHUFB": {0x0D50 << 16, false}, "XVSHUFB": {0x0D60 << 16, true},
}
for m, e := range vec4r {
l64InstrTable[m] = l64Enc{format: l64Fvvvv, op: e.op}
l64VecBank[m] = e.lasx
l64Vec4R[m] = true
}
}
// l64FpMovTable maps (mnemonic, from-class, to-class) to the 2R opcode of the
+503
View File
@@ -263,6 +263,20 @@ func TestLOONG64_regNames(t *testing.T) {
t.Errorf("loong64RegNum(%q) = %d, want %d", name, got, want)
}
}
// The X/V spellings name the LSX/LASX vector banks, a register class of
// their own: the oracle (GOARCH=loong64 go tool asm) rejects `BEQZ X0`
// with "unrecognized instruction" while assembling `VADDV V0, V1, V2`
// and `XVADDV X0, X1, X2`, so loong64RegNum stays strict and the vector
// operands resolve through loong64VecRegNum only.
vecCases := map[string]int{
"V0": 0, "V31": 31, "X0": 0, "X31": 31,
"R4": -1, "F0": -1, "FCC0": -1, "V32": -1, "X32": -1, "V": -1, "X": -1,
}
for name, want := range vecCases {
if got := loong64VecRegNum(name); got != want {
t.Errorf("loong64VecRegNum(%q) = %d, want %d", name, got, want)
}
}
}
func TestLOONG64_bytesEqualGroundTruth(t *testing.T) {
@@ -328,3 +342,492 @@ TEXT ·f(SB), NOSPLIT, $0-0
0x4C000020, // jirl r0, r1, 0 (RET)
)
}
// TestLOONG64_vector pins the LSX/LASX slice against words read off
// GOARCH=loong64 go tool asm (cross-checked against the toolchain's own
// loong64enc1.s): the three-register forms, the immediate forms with their
// biases, the vector-to-condition forms, lane popcount, the FP conversion,
// FSEL and the VMOVQ move family.
func TestLOONG64_vector(t *testing.T) {
t.Run("three-register and immediate forms", func(t *testing.T) {
fn := firstTextLOONG64(t, `#include "textflag.h"
TEXT ·v(SB), NOSPLIT, $0
VADDV V1, V2, V3
VADDW V1, V2, V3
VADDV V2, V1
VANDV V1, V2
VXORV V1, V2, V3
VSEQB V1, V2, V3
VSEQV V1, V2, V3
VSRAB V1, V2, V3
VROTRW V1, V2, V3
VANDB $0, V2, V3
VANDB $255, V2
VSEQB $3, V2, V3
VSEQV $15, V2, V3
VSEQV $-15, V2, V3
VSRAB $7, V1, V2
VROTRW $16, V1, V2
VPCNTV V1, V2
XVADDV X1, X2, X3
XVXORV X1, X2, X3
XVSEQB X1, X2, X3
XVPCNTV X1, X2
RET
`)
code := assembleLOONG64Helper(t, fn)
wantWords(t, code,
0x700B8443, // vadd.v v3, v2, v1
0x700B0443, // vadd.w
0x700B8821, // vadd.v v1, v1, v2 (two-operand form)
0x71260442, // vand.v v2, v2, v1
0x71270443, // vxor.v
0x70000443, // vseq.b
0x70018443, // vseq.d
0x70EC0443, // vsra.b
0x70EF0443, // vrotr.w
0x73D00043, // vandi.b v3, v2, 0
0x73D3FC42, // vandi.b v2, v2, 255 (two-operand form)
0x72800C43, // vseqi.b v3, v2, 3
0x7281BC43, // vseqi.d v3, v2, 15
0x7281C443, // vseqi.d v3, v2, -15 (7-bit two's complement)
0x73343C22, // vsrai.b v2, v1, 7 (encoded as 7+8)
0x72A0C022, // vrotri.w v2, v1, 16
0x729C2C22, // vpcnt.d v2, v1
0x740B8443, // xvadd.d x3, x2, x1
0x75270443, // xvxor.d
0x74000443, // xvseq.b
0x769C2C22, // xvpcnt.d x2, x1
0x4C000020,
)
})
t.Run("vector-to-condition", func(t *testing.T) {
fn := firstTextLOONG64(t, `#include "textflag.h"
TEXT ·v(SB), NOSPLIT, $0
VSETNEV V1, FCC0
VSETANYEQB V1, FCC0
VSETANYEQV V2, FCC0
VSETALLNEV V0, FCC0
XVSETNEV X1, FCC0
XVSETALLNEV X1, FCC0
RET
`)
code := assembleLOONG64Helper(t, fn)
wantWords(t, code,
0x729C9C20, // vsetnez.d fcc0, v1
0x729CA020, // vsetanyeqz.b
0x729CAC40, // vsetanyeqz.d
0x729CBC00, // vsetallnez.d
0x769C9C20, // xvsetnez.d
0x769CBC20, // xvsetallnez.d
0x4C000020,
)
})
t.Run("FP convert and FSEL", func(t *testing.T) {
fn := firstTextLOONG64(t, `#include "textflag.h"
TEXT ·v(SB), NOSPLIT, $0
FFINTDV F0, F1
FSEL FCC0, F3, F4, F3
FSEL FCC1, F1, F2
RET
`)
code := assembleLOONG64Helper(t, fn)
wantWords(t, code,
0x011D2801, // ffint.d.v f1, f0
0x0D000C83, // fsel f3, f4, f3, fcc0
0x0D008442, // fsel f2, f2, f1, fcc1
0x4C000020,
)
})
t.Run("VMOVQ move family", func(t *testing.T) {
fn := firstTextLOONG64(t, `#include "textflag.h"
TEXT ·v(SB), NOSPLIT, $0
VMOVQ V1, V9
VMOVQ (R4), V2
VMOVQ 16(R4), V2
VMOVQ V0, (R4)
VMOVQ V0, 32(R4)
VMOVQ (R4)(R7), V3
VMOVQ V3, (R4)(R7)
VMOVQ R6, V0.B16
VMOVQ R6, V12.W4
VMOVQ (R4), V4.W4
XVMOVQ X3, X7
XVMOVQ (R4), X2
XVMOVQ X0, (R4)
XVMOVQ (R4)(R7), X4
XVMOVQ X0, (R4)(R7)
XVMOVQ R6, X0.B32
RET
`)
code := assembleLOONG64Helper(t, fn)
wantWords(t, code,
0x732D0029, // vori.b v9, v1, 0 (register move)
0x2C000082, // vld v2, r4, 0
0x2C004082, // vld v2, r4, 16
0x2C400080, // vst v0, r4, 0
0x2C408080, // vst v0, r4, 32
0x38401C83, // vldx v3, r4, r7
0x38441C83, // vstx v3, r4, r7
0x729F00C0, // vreplgr2vr.b v0, r6
0x729F08CC, // vreplgr2vr.w v12, r6
0x30200084, // vldrepl.w v4, r4, 0
0x772D0067, // xvori.b x7, x3, 0
0x2C800082, // xvld x2, r4, 0
0x2CC00080, // xvst x0, r4, 0
0x38481C84, // xvldx x4, r4, r7
0x384C1C80, // xvstx x0, r4, r7
0x769F00C0, // xvreplgr2vr.b x0, r6
0x4C000020,
)
})
t.Run("element extract and insert", func(t *testing.T) {
fn := firstTextLOONG64(t, `#include "textflag.h"
TEXT ·v(SB), NOSPLIT, $0
VMOVQ V0.V[0], R10
VMOVQ V6.V[1], R8
VMOVQ R9, V1.V[0]
XVMOVQ X0.V[0], R10
XVMOVQ X5.W[7], R7
XVMOVQ R4, X7.V[3]
RET
`)
code := assembleLOONG64Helper(t, fn)
wantWords(t, code,
0x72EFF00A, // vpickve2gr.d r10, v0, 0
0x72EFF4C8, // vpickve2gr.d r8, v6, 1
0x72EBF121, // vinsgr2vr.d v1, r9, 0
0x76EFE00A, // xvpickve2gr.d r10, x0, 0
0x76EFDCA7, // xvpickve2gr.w r7, x5, 7
0x76EBEC87, // xvinsgr2vr.d x7, r4, 3
0x4C000020,
)
})
// The integer and FP add/subtract families with their saturating pairs
// and immediate spellings (loong64enc1.s words).
t.Run("add and subtract families", func(t *testing.T) {
fn := firstTextLOONG64(t, `#include "textflag.h"
TEXT ·v(SB), NOSPLIT, $0
VADDB V1, V2, V3
VADDF V1, V2, V3
VADDD V1, V2, V3
VSUBD V1, V2, V3
VSADDV V1, V2, V3
VSSUBVU V1, V2, V3
VADDBU $1, V2, V1
VADDBU $1, V2
VSUBVU $31, V2
XVSADDV X3, X2, X1
XVSUBD X1, X2, X3
RET
`)
code := assembleLOONG64Helper(t, fn)
wantWords(t, code,
0x700A0443, // vadd.b
0x71308443, // vadd.f
0x71310443, // vadd.d
0x71330443, // vsub.d
0x70478443, // vsadd.v
0x704D8443, // vssub.u.d
0x728A0441, // vaddi.bu v1, v2, 1
0x728A0442, // vaddi.bu v2, v2, 1 (two-operand form)
0x728DFC42, // vsubi.du v2, v2, 31 (two-operand form)
0x74478C41, // xvsadd.d x1, x2, x3
0x75330443, // xvsub.d x3, x2, x1
0x4C000020,
)
})
// The multiply, divide and accumulate families.
t.Run("multiply and divide families", func(t *testing.T) {
fn := firstTextLOONG64(t, `#include "textflag.h"
TEXT ·v(SB), NOSPLIT, $0
VMULV V1, V2, V3
VMUHHU V1, V2, V3
VDIVBU V1, V2, V3
VMODV V1, V2, V3
VMADDB V1, V2, V3
VMSUBV V1, V2, V3
VMULWEVHB V1, V2, V3
VMULWODQV V1, V2, V3
VMADDWEVHBUB V1, V2, V3
XVDIVD X1, X2, X3
RET
`)
code := assembleLOONG64Helper(t, fn)
wantWords(t, code,
0x70858443, // vmul.v
0x70888443, // vmuh.u.d
0x70E40443, // vdiv.u.b
0x70E38443, // vmod.d
0x70A80443, // vmadd.b
0x70AB8443, // vmsub.d
0x70900443, // vmulwev.h.b
0x70938443, // vmulwod.q.d
0x70BC0443, // vmaddwev.h.bu.b
0x753B0443, // xvdiv.d
0x4C000020,
)
})
// The shift, bit and interleave families in register and immediate
// spellings, with the width-coded shift immediates.
t.Run("shift, bit and interleave families", func(t *testing.T) {
fn := firstTextLOONG64(t, `#include "textflag.h"
TEXT ·v(SB), NOSPLIT, $0
VSLLV V1, V2, V3
VROTRB V1, V2, V3
VBITCLRV V1, V2, V3
VBITSETW V1, V2, V3
VBITREVV V1, V2, V3
VILVLB V1, V2, V3
VILVHV V1, V2, V3
VSLLB $7, V1, V2
VSLLB $5, V1
VSRLH $15, V1, V2
VSRAW $31, V1, V2
VSRAV $63, V1, V2
VROTRV $63, V1, V2
VBITCLRB $7, V2, V3
VBITREVV $63, V2, V3
VSEQH $-16, V2, V3
VSLTB $1, V2, V3
VSLTHU $31, V2, V3
XVILVLV X3, X2, X1
XVSLLB $7, X2, X1
XVSRAV $63, X2, X1
XVBITREVV $63, X2, X1
RET
`)
code := assembleLOONG64Helper(t, fn)
wantWords(t, code,
0x70E98443, // vsll.d
0x70EE0443, // vrotr.b
0x710D8443, // vbitclr.d
0x710F0443, // vbitset.w
0x71118443, // vbitrev.d
0x711A0443, // vilvl.b
0x711D8443, // vilvh.d
0x732C3C22, // vslli.b v2, v1, 7
0x732C3421, // vslli.b v1, v1, 5 (two-operand form)
0x73307C22, // vsrli.h v2, v1, 15
0x7334FC22, // vsrai.w v2, v1, 31
0x7335FC22, // vsrai.d v2, v1, 63
0x72A1FC22, // vrotri.d v2, v1, 63
0x73103C43, // vbitclri.b v3, v2, 7
0x7319FC43, // vbitrevi.d v3, v2, 63
0x7280C043, // vseqi.h v3, v2, -16
0x72860443, // vslti.b v3, v2, 1
0x7288FC43, // vslti.hu v3, v2, 31
0x751B8C41, // xvilvl.d x1, x2, x3
0x772C3C41, // xvslli.b x1, x2, 7
0x7735FC41, // xvsrai.d x1, x2, 63
0x7719FC41, // xvbitrevi.d x1, x2, 63
0x4C000020,
)
})
// The shuffle, select and permutation families, including the
// four-register byte shuffle.
t.Run("shuffle and permutation families", func(t *testing.T) {
fn := firstTextLOONG64(t, `#include "textflag.h"
TEXT ·v(SB), NOSPLIT, $0
VSHUFH V1, V2, V3
VSHUFW V1, V2, V3
VSHUFV V1, V2, V3
VSHUFB V1, V2, V3, V4
XVSHUFB X1, X2, X3, X4
VSHUF4IB $255, V2, V1
VSHUF4IV $15, V2, V1
XVSHUF4IV $15, X1, X2
VEXTRINSB $0x18, V1, V2
XVEXTRINSV $0x81, X1, X2
VPERMIW $0x1B, V1, V2
XVPERMIQ $0x4B, X1, X2
RET
`)
code := assembleLOONG64Helper(t, fn)
wantWords(t, code,
0x717A8443, // vshuf.h
0x717B0443, // vshuf.w
0x717B8443, // vshuf.d
0x0D508864, // vshuf.b v4, v3, v2, v1
0x0D608864, // xvshuf.b
0x7393FC41, // vshuf4i.b v1, v2, 255
0x739C3C41, // vshuf4i.d v1, v2, 15
0x779C3C22, // xvshuf4i.d x2, x1, 15
0x738C6022, // vextrins.b v2, v1, 0x18
0x77820422, // xvextrins.d x2, x1, 0x81
0x73E46C22, // vpermi.w v2, v1, 0x1b
0x77ED2C22, // xvpermi.q x2, x1, 0x4b
0x4C000020,
)
})
// The vector FP families, the unary spellings, the compare-to-flag
// additions and the scalar int/float conversions.
t.Run("FP and conversion families", func(t *testing.T) {
fn := firstTextLOONG64(t, `#include "textflag.h"
TEXT ·v(SB), NOSPLIT, $0
VADDF V1, V2, V3
VMULF V1, V2, V3
VFCLASSD V1, V2
VFSQRTF V1, V2
VFRECIPD V1, V2
VFRSQRTF V1, V2
VFRINTF V1, V2
VFRINTRNED V1, V2
VNEGB V1, V2
VPCNTB V1, V2
XVNEGV X2, X1
XVPCNTW X3, X2
XVFRINTRNEF X1, X2
VSETEQV V1, FCC0
VSETANYEQH V1, FCC0
VSETALLNEB V1, FCC0
XVSETALLNEW X1, FCC0
FFINTFW F0, F1
FTINTVD F0, F1
RET
`)
code := assembleLOONG64Helper(t, fn)
wantWords(t, code,
0x71308443, // vfadd.s
0x71388443, // vfmul.s
0x729CD822, // vfclass.d
0x729CE422, // vfsqrt.s
0x729CF822, // vfrecip.d
0x729D0422, // vfrsqrt.s
0x729D3422, // vfrint.s
0x729D7822, // vfrintne.s
0x729C3022, // vneg.b
0x729C2022, // vpcnt.b
0x769C3C41, // xvneg.d x1, x2
0x769C2862, // xvpcnt.w x2, x3
0x769D7422, // xvfrintne.s x2, x1
0x729C9820, // vseteqz.d fcc0, v1
0x729CA420, // vsetanyeqz.h
0x729CB020, // vsetallnez.b
0x769CB820, // xvsetallnez.w
0x011D1001, // ffint.s.w f1, f0
0x011B2801, // ftint.l.d f1, f0
0x4C000020,
)
})
}
// TestLOONG64_vectorErrors pins the register-class and range diagnostics of
// the vector slice; each shape is rejected by the oracle as well
// (GOARCH=loong64 go tool asm).
func TestLOONG64_vectorErrors(t *testing.T) {
cases := []string{
// Integer registers in vector positions.
`TEXT ·e(SB), NOSPLIT, $0
VADDV R4, R5, R6
RET
`,
// Crossed banks: LSX spellings take V, LASX spellings X.
`TEXT ·e(SB), NOSPLIT, $0
VADDV X1, X2, X3
RET
`,
`TEXT ·e(SB), NOSPLIT, $0
XVADDV V1, V2, V3
RET
`,
// The LASX bank has no .b/.h element forms.
`TEXT ·e(SB), NOSPLIT, $0
XVMOVQ R4, X2.B[0]
RET
`,
// Immediate ranges.
`TEXT ·e(SB), NOSPLIT, $0
VANDB $256, V2
RET
`,
`TEXT ·e(SB), NOSPLIT, $0
VSEQB $16, V2, V3
RET
`,
`TEXT ·e(SB), NOSPLIT, $0
VROTRW $32, V1, V2
RET
`,
`TEXT ·e(SB), NOSPLIT, $0
VADDVU $32, V2
RET
`,
`TEXT ·e(SB), NOSPLIT, $0
VSEQV $32, V2, V3
RET
`,
`TEXT ·e(SB), NOSPLIT, $0
VSHUF4IV $16, V2, V1
RET
`,
`TEXT ·e(SB), NOSPLIT, $0
VEXTRINSB $256, V1, V2
RET
`,
`TEXT ·e(SB), NOSPLIT, $0
VSLTV $-17, V2, V3
RET
`,
// VSHUFB wants four vector registers.
`TEXT ·e(SB), NOSPLIT, $0
VSHUFB V1, V2, V3
RET
`,
// The FCC forms still refuse vector registers.
`TEXT ·e(SB), NOSPLIT, $0
VSETEQV V1, V2
RET
`,
// VSET* wants an FCC flag, not a vector register.
`TEXT ·e(SB), NOSPLIT, $0
VSETNEV V1, V2
RET
`,
}
for i, src := range cases {
fn := firstTextLOONG64(t, src)
if _, _, _, _, _, err := assembleLOONG64(fn); err == nil {
t.Errorf("case %d: expected an error, got none", i)
}
}
}
// TestLOONG64_dbarAtomics pins the _dbar (acquire/release) AMO variants.
// The oracle words come from GOARCH=loong64 go tool objdump of kernels
// assembled with go tool asm, and match the toolchain's loong64enc1.s.
func TestLOONG64_dbarAtomics(t *testing.T) {
fn := firstTextLOONG64(t, `#include "textflag.h"
TEXT ·atoms(SB), NOSPLIT, $0
AMADDDBW R14, (R13), R12
AMADDDBV R14, (R13), R12
AMANDDBW R5, (R4), R6
AMANDDBV R5, (R4), R6
AMORDBW R5, (R4), R0
AMORDBV R5, (R4), R6
AMSWAPDBW R5, (R4), R6
AMCASDBV R6, (R4), R5
RET
`)
code := assembleLOONG64Helper(t, fn)
wantWords(t, code,
0x386A39AC, // amadd_db.w r12, r13, r14
0x386AB9AC, // amadd_db.d
0x386B1486, // amand_db.w r6, r4, r5
0x386B9486, // amand_db.d
0x386C1480, // amor_db.w r0, r4, r5
0x386C9486, // amor_db.d
0x38691486, // amswap_db.w
0x385B9885, // amcas_db.w
0x4C000020,
)
}
+4 -1
View File
@@ -232,7 +232,10 @@ DATA ·table+0(SB)/8, $42
}
// TestLOONG64_errors checks the encoder's error paths: undefined labels,
// invalid register operands and operand-count mismatches.
// invalid register operands and operand-count mismatches. The X0 and
// AMADDW cases follow the oracle: GOARCH=loong64 go tool asm rejects
// `BEQZ X0` (the X bank is not an integer register) and the two-register
// `AMADDW R4, R5` (the AM* family is strictly `val, (addr), result`).
func TestLOONG64_errors(t *testing.T) {
cases := []string{
`TEXT ·e(SB), NOSPLIT, $0
+13
View File
@@ -48,3 +48,16 @@ type sbMem struct {
}
func (sbMem) isOperand() {}
// isX86Mem reports whether the operand is an amd64 memory reference: a base
// or indexed Mem, or an SB-relative sbMem. Encoders that gate on "memory in
// this position" must accept both; the r/m emitters distinguish the two
// themselves.
func isX86Mem(o Operand) bool {
switch o.(type) {
case Mem, sbMem:
return true
default:
return false
}
}
+6 -1
View File
@@ -17,12 +17,13 @@ import "strings"
// size. The high flag marks the legacy high-byte registers AH/CH/DH/BH, which
// occupy indices 4-7 yet take no REX prefix, unlike SPL/BPL/SIL/DIL that share
// those indices but require one. The mask flag marks the AVX-512 opmask
// registers K0-K7.
// registers K0-K7, the fp flag the x87 stack registers F0-F7.
type Reg struct {
idx int
size int // informational width implied by the name; the mnemonic decides
high bool // AH/CH/DH/BH
mask bool // K0-K7 opmask register
fp bool // F0-F7 x87 stack register
}
// Index returns the register number (0-15 for GPRs, 0-31 for vectors).
@@ -144,6 +145,10 @@ func buildRegByName() map[string]Reg {
for i := 0; i <= 7; i++ {
m["K"+itoa(i)] = Reg{idx: i, size: 8, mask: true}
}
// x87 stack: F0..F7.
for i := 0; i <= 7; i++ {
m["F"+itoa(i)] = Reg{idx: i, size: 8, fp: true}
}
return m
}
+1130 -47
View File
File diff suppressed because it is too large Load Diff
+150 -30
View File
@@ -141,10 +141,36 @@ func riscvRegNum(name string) int {
case "F31", "FT11":
return 31
default:
// Vector registers V0-V31 (the "V" extension). They share the
// register numbering with the integer file: a bare number 0-31.
if len(name) >= 2 && name[0] == 'V' {
if n, ok := parseRegDigits(name[1:], 31); ok {
return n
}
}
return -1
}
}
// parseRegDigits parses a decimal register suffix and reports whether it is
// within [0, max].
func parseRegDigits(digits string, max int) (int, bool) {
if digits == "" {
return 0, false
}
n := 0
for i := 0; i < len(digits); i++ {
if digits[i] < '0' || digits[i] > '9' {
return 0, false
}
n = n*10 + int(digits[i]-'0')
if n > max {
return 0, false
}
}
return n, true
}
// RISC-V instruction encoding parameters.
type riscvEnc struct {
opcode uint32 // bits [6:0]
@@ -193,6 +219,9 @@ var riscvInstrTable = map[string]riscvEnc{
"DIVUW": {0x3B, 0x5, 0x01},
"REMW": {0x3B, 0x6, 0x01},
"REMUW": {0x3B, 0x7, 0x01},
// Zicond conditional zeroing.
"CZEROEQZ": {0x33, 0x5, 0x07},
"CZERONEZ": {0x33, 0x7, 0x07},
// RV64I, I-type arithmetic.
"ADDI": {0x13, 0x0, 0x00},
"ADDIW": {0x1B, 0x0, 0x00},
@@ -221,36 +250,47 @@ var riscvInstrTable = map[string]riscvEnc{
"BGE": {0x63, 0x5, 0x00},
"BLTU": {0x63, 0x6, 0x00},
"BGEU": {0x63, 0x7, 0x00},
// The swapped-spelling comparison forms: encoded as BLT/BGE/BLTU/BGEU
// with the register operands swapped.
"BGT": {0x63, 0x4, 0x00},
"BLE": {0x63, 0x5, 0x00},
"BGTU": {0x63, 0x6, 0x00},
"BLEU": {0x63, 0x7, 0x00},
// U-type.
"LUI": {0x37, 0x0, 0x00},
"AUIPC": {0x17, 0x0, 0x00},
// System.
"ECALL": {0x73, 0x0, 0x00},
"EBREAK": {0x73, 0x0, 0x00},
"FENCE": {0x0F, 0x0, 0x00},
"ECALL": {0x73, 0x0, 0x00},
"EBREAK": {0x73, 0x0, 0x00},
"FENCE": {0x0F, 0x0, 0x00},
"FENCE.TSO": {0x0F, 0x0, 0x00},
"PAUSE": {0x0F, 0x0, 0x00},
// JALR, indirect jump/call (I-type).
"JALR": {0x67, 0x0, 0x00},
// RV64A, atomics (AMO opcode 0x2F).
// funct3: 0x2 = word, 0x3 = doubleword. funct5 in bits [31:27].
"AMOSWAPW": {0x2F, 0x2, 0x01 << 2},
"AMOSWAPD": {0x2F, 0x3, 0x01 << 2},
"AMOADDW": {0x2F, 0x2, 0x00 << 2},
"AMOADDD": {0x2F, 0x3, 0x00 << 2},
"AMOANDW": {0x2F, 0x2, 0x0C << 2},
"AMOANDD": {0x2F, 0x3, 0x0C << 2},
"AMOORW": {0x2F, 0x2, 0x06 << 2},
"AMOORD": {0x2F, 0x3, 0x06 << 2},
"AMOXORW": {0x2F, 0x2, 0x04 << 2},
"AMOXORD": {0x2F, 0x3, 0x04 << 2},
"AMOMAXW": {0x2F, 0x2, 0x14 << 2},
"AMOMAXD": {0x2F, 0x3, 0x14 << 2},
"AMOMINW": {0x2F, 0x2, 0x10 << 2},
"AMOMIND": {0x2F, 0x3, 0x10 << 2},
"AMOMAXUW": {0x2F, 0x2, 0x1C << 2},
"AMOMAXUD": {0x2F, 0x3, 0x1C << 2},
"AMOMINUW": {0x2F, 0x2, 0x18 << 2},
"AMOMINUD": {0x2F, 0x3, 0x18 << 2},
// funct3: 0x2 = word, 0x3 = doubleword. The stored funct7 is the full
// 7-bit field: funct5 in the upper five bits and the aq/rl ordering bits in
// the lower two, exactly as the toolchain writes them: every AMO sets both
// aq and rl (funct7 |= 3).
"AMOSWAPW": {0x2F, 0x2, 0x01<<2 | 0x3},
"AMOSWAPD": {0x2F, 0x3, 0x01<<2 | 0x3},
"AMOADDW": {0x2F, 0x2, 0x00<<2 | 0x3},
"AMOADDD": {0x2F, 0x3, 0x00<<2 | 0x3},
"AMOANDW": {0x2F, 0x2, 0x0C<<2 | 0x3},
"AMOANDD": {0x2F, 0x3, 0x0C<<2 | 0x3},
"AMOORW": {0x2F, 0x2, 0x08<<2 | 0x3},
"AMOORD": {0x2F, 0x3, 0x08<<2 | 0x3},
"AMOXORW": {0x2F, 0x2, 0x04<<2 | 0x3},
"AMOXORD": {0x2F, 0x3, 0x04<<2 | 0x3},
"AMOMAXW": {0x2F, 0x2, 0x14<<2 | 0x3},
"AMOMAXD": {0x2F, 0x3, 0x14<<2 | 0x3},
"AMOMINW": {0x2F, 0x2, 0x10<<2 | 0x3},
"AMOMIND": {0x2F, 0x3, 0x10<<2 | 0x3},
"AMOMAXUW": {0x2F, 0x2, 0x1C<<2 | 0x3},
"AMOMAXUD": {0x2F, 0x3, 0x1C<<2 | 0x3},
"AMOMINUW": {0x2F, 0x2, 0x18<<2 | 0x3},
"AMOMINUD": {0x2F, 0x3, 0x18<<2 | 0x3},
// RV64F/D, floating-point arithmetic.
"FADDS": {0x53, 0x0, 0x00},
@@ -273,12 +313,23 @@ var riscvInstrTable = map[string]riscvEnc{
"FMAXS": {0x53, 0x1, 0x14},
"FMIND": {0x53, 0x0, 0x15},
"FMAXD": {0x53, 0x1, 0x15},
// FP sign injection (double): rs2 carries the sign source.
"FSGNJD": {0x53, 0x0, 0x11},
"FSGNJS": {0x53, 0x0, 0x10},
"FSGNJX": {0x53, 0x0, 0x14},
"FSGNJXD": {0x53, 0x0, 0x15},
"FSGNJXS": {0x53, 0x0, 0x14},
"FSGNJND": {0x53, 0x1, 0x11},
"FSGNJNS": {0x53, 0x1, 0x10},
"FSGNJNX": {0x53, 0x1, 0x14},
// RV64A, load-reserved / store-conditional (funct5 0x02 / 0x03).
"LRW": {0x2F, 0x2, 0x02 << 2},
"LRD": {0x2F, 0x3, 0x02 << 2},
"SCW": {0x2F, 0x2, 0x03 << 2},
"SCD": {0x2F, 0x3, 0x03 << 2},
// The toolchain gives LR acquire ordering (aq = 1) and SC release
// ordering (rl = 1).
"LRW": {0x2F, 0x2, 0x02<<2 | 0x2},
"LRD": {0x2F, 0x3, 0x02<<2 | 0x2},
"SCW": {0x2F, 0x2, 0x03<<2 | 0x1},
"SCD": {0x2F, 0x3, 0x03<<2 | 0x1},
// FP compare, result in integer register (funct7 0x50/0x51).
"FEQS": {0x53, 0x2, 0x50},
@@ -296,11 +347,11 @@ func riscvRType(enc riscvEnc, rd, rs1, rs2 int) uint32 {
}
// riscvAMOType encodes an atomic (AMO) instruction.
// Layout: funct5 | aq | rl | rs2 | rs1 | funct3 | rd | opcode.
// The funct5 is stored in the upper bits of enc.funct7 (shifted left by 2).
// Layout: funct7 | rs2 | rs1 | funct3 | rd | opcode, where funct7 carries the
// funct5 in its upper five bits and the aq/rl ordering bits in the lower two
// (the table stores the full field, so the word needs no reassembly).
func riscvAMOType(enc riscvEnc, rd, rs1, rs2 int) uint32 {
funct5 := enc.funct7 >> 2 // extract funct5 from the stored value
return (funct5 << 27) | (uint32(rs2) << 20) | (uint32(rs1) << 15) |
return (enc.funct7 << 25) | (uint32(rs2) << 20) | (uint32(rs1) << 15) |
(enc.funct3 << 12) | (uint32(rd) << 7) | enc.opcode
}
@@ -342,6 +393,10 @@ var riscvCvtTable = map[string]riscvCvtEnc{
"FMVDX": {0x79, 0x0, 0x53}, // int64 → float64 (bit move)
"FMVXW": {0x70, 0x0, 0x53}, // float32 → int32 (bit move)
"FMVWX": {0x78, 0x0, 0x53}, // int32 → float32 (bit move)
// The toolchain's W/D suffix spellings of the same moves.
"FMVXS": {0x70, 0x0, 0x53},
"FMVFS": {0x78, 0x0, 0x53},
"FMVSX": {0x79, 0x0, 0x53},
}
// riscvCvtType encodes an FP conversion instruction.
@@ -441,6 +496,71 @@ func riscvJType(rd int, offset int32) uint32 {
0x6F // JAL opcode
}
// ---- RVV ("V" extension) encoding helpers ----
// The OP-V major opcode and its funct3 subclasses.
const (
riscvOpV = 0x57 // the vector operation opcode (also OPcfg for vset*)
// funct3 values: 0 OPIVV, 1 OPFVV, 2 OPMVV, 3 OPIVI, 4 OPIVX,
// 5 OPFVF, 6 OPMVX, 7 vsetvli.
riscvVf3VV = 0x0 // vector-vector
riscvVf3MV = 0x2 // vector mask
riscvVf3VI = 0x3 // vector-immediate
riscvVf3VX = 0x4 // vector-scalar
riscvVf3Cfg = 0x7 // vsetvli
)
// riscvVType composes the vsetvli/vsetivli vtype immediate: the register
// group multiplier in [2:0], the selected element width in [5:3] and the
// tail-agnostic and mask-agnostic policies in bits 6 and 7.
func riscvVType(vsew, vlmul, vta, vma int) int {
return vlmul | vsew<<3 | vta<<6 | vma<<7
}
// riscvVSetEnc encodes VSETVLI and VSETIVLI: imm[31:20] = vtype, rs1 = the
// avl register or 5-bit uimm, rd = the destination. Both carry funct3 7; a
// vsetivli is distinguished by bits [31:30] set in the immediate (the 0xC00
// the toolchain writes above its 10-bit vtype).
func riscvVSetEnc(vsetivli bool, avl, vtype, rd int) uint32 {
imm := vtype & 0x3FF
if vsetivli {
imm |= 0xC00
}
return uint32(imm)<<20 | uint32(avl&0x1F)<<15 | uint32(riscvVf3Cfg)<<12 |
uint32(rd)<<7 | riscvOpV
}
// riscvVLSType encodes a vector load or store: the full 32-bit word with the
// segment count in bits [31:29], the addressing mode in bits [28:26], the
// unmasked bit at 25 and the width in funct3. width follows the load
// convention (0 = 8-bit, 5 = 16-bit, 6 = 32-bit, 7 = 64-bit).
func riscvVLSType(op uint32, nf, mop, width int, rs2 int32, rs1, rd int) uint32 {
return uint32(nf&0x7)<<29 | uint32(mop&0x7)<<26 | 1<<25 |
uint32(rs2)<<20 | uint32(rs1)<<15 | uint32(width&0x7)<<12 |
uint32(rd)<<7 | op
}
// riscvVVInstr encodes an OP-V instruction with the six-bit operation code in
// funct7's upper bits, bit 25 as the unmasked flag and the three registers in
// the standard positions. vs1 may name an integer register for the *VX forms
// (the scalar sits in the rs1 field) or an immediate for the *VI forms.
func riscvVVInstr(funct6, funct3 int, vs1 int32, vs2, vd int) uint32 {
return uint32(funct6&0x3F)<<26 | 1<<25 | uint32(vs1)<<15 |
uint32(funct3)<<12 | uint32(vs2)<<20 | uint32(vd)<<7 | riscvOpV
}
// riscvVUnaryInstr encodes a one-vector-operand OP-V instruction whose fixed
// fields live where the second source register would be: rs1Field and vs2 are
// written verbatim (the oracle writes fixed non-zero constants there for some
// instructions, such as 0x11 in the rs1 field of vmfirst.m and vid.v).
func riscvVUnaryInstr(funct6, funct3 int, rs1Field int32, vs2, vd int) uint32 {
return uint32(funct6&0x3F)<<26 | 1<<25 | uint32(vs2&0x1F)<<20 |
uint32(rs1Field&0x1F)<<15 | uint32(funct3&0x7)<<12 | uint32(vd&0x1F)<<7 | riscvOpV
}
// riscvSegNF maps a segment count to the 3-bit nf field (count - 1).
func riscvSegNF(n int) int32 { return int32(n - 1) }
// ---- RVC (compressed) encoding helpers ----
// isRVCIntReg reports whether a register number can be encoded in the 3-bit
+242 -6
View File
@@ -5,6 +5,8 @@ package asm
import (
"bytes"
"encoding/binary"
"encoding/hex"
"strings"
"testing"
@@ -866,7 +868,7 @@ func encodeOneInstrRISCV(t *testing.T, src string, pc int, offsets map[string]in
t.Helper()
fn := firstTextRISCV(t, "#include \"textflag.h\"\n"+src)
instr := fn.Body[0].(*ast.Instr)
return encodeRISCVInstr(instr, pc, offsets, riscvFrameInfo{}, nil)
return encodeRISCVInstr(instr, pc, offsets, riscvFrameInfo{}, nil, nil)
}
// TestRISCVBranchJumpRange checks that displacements beyond the B-type span
@@ -903,9 +905,10 @@ func TestRISCVBranchJumpRange(t *testing.T) {
}
}
// TestRISCVBranchFarBody drives the range check through the full two-pass
// assembler: a forward branch over a body larger than the B-type span must
// error rather than wrap.
// TestRISCVBranchFarBody drives the relaxation pass through the full
// assembler: a forward branch over a body larger than the B-type span is
// rewritten as an inverted branch over an inserted JMP, the same layout the
// toolchain produces, instead of wrapping to a wrong target.
func TestRISCVBranchFarBody(t *testing.T) {
var sb strings.Builder
sb.WriteString("#include \"textflag.h\"\nTEXT ·far(SB), NOSPLIT, $0\n\tBEQ X10, X11, done\n")
@@ -914,8 +917,20 @@ func TestRISCVBranchFarBody(t *testing.T) {
}
sb.WriteString("done:\n\tRET\n")
fn := firstTextRISCV(t, sb.String())
if _, _, _, _, _, err := assembleRISCV(fn); err == nil {
t.Error("expected a branch-out-of-range error, got none")
out, _, _, _, _, err := assembleRISCV(fn)
if err != nil {
t.Fatalf("unexpected error: %v", err)
}
// The relaxed branch at offset 0 targets the inserted JMP at 4 (bne
// x10, x11, +4); the JMP at 4 carries the far forward displacement.
wantBranch := wordLE(riscvBType(riscvEnc{0x63, 0x1, 0x00}, 10, 11, 4))
if !bytes.Equal(out[0:4], wantBranch) {
t.Errorf("relaxed branch = %x, want %x", out[0:4], wantBranch)
}
// done sits after 1100 ADDs: 4 + 4400, i.e. offset 4404 from the JMP at 4.
wantJmp := wordLE(riscvJType(0, 4404))
if !bytes.Equal(out[4:8], wantJmp) {
t.Errorf("inserted JMP = %x, want %x", out[4:8], wantJmp)
}
}
@@ -971,3 +986,224 @@ TEXT ·edge(SB), NOSPLIT, $0
t.Errorf("int32-span immediates must assemble: %v", err)
}
}
// riscvWants decodes code as little-endian words and pins each one; the
// expected values below were read off GOARCH=riscv64 go tool objdump of
// kernels assembled with go tool asm (the toolchain's riscv64.s testdata
// cross-checks the same words).
func riscvWants(t *testing.T, code []byte, want ...uint32) {
t.Helper()
got := make([]uint32, 0, len(code)/4)
for i := 0; i+4 <= len(code); i += 4 {
got = append(got, binary.LittleEndian.Uint32(code[i:]))
}
if len(got) < len(want) {
t.Fatalf("word count = %d, want %d\ncode: % x", len(got), len(want), code)
}
// The RET (JALR) ends the sequence; only the pinned prefix is compared.
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
}
}
}
// riscvWantsHex pins the exact hex encoding of a function's instruction
// bytes, including any 2-byte compressed instructions in the stream; the
// expected strings were read off GOARCH=riscv64 go tool objdump of kernels
// assembled with go tool asm (the toolchain's riscv64.s testdata
// cross-checks the same words).
func riscvWantsHex(t *testing.T, code []byte, wantHex string) {
t.Helper()
got := hex.EncodeToString(code)
if got != wantHex {
t.Errorf("code = %s, want %s", got, wantHex)
}
}
// TestRISCV_extendedPseudos pins the toolchain-synthesised instructions:
// ANDN/ORN (XORI + AND/OR through the destination or TMP), the five-word
// MIN/MAX expansion, the four-word rotate, ROR's compressed reverse shift
// (C.SLLI when rd == rs1, both non-zero, 1 <= sll <= 63), the identical-
// input MIN/MAX fold to C.MV, FABSD (FSGNJX.D), SEQZ and RDTIME (csrrs with
// the time CSR).
func TestRISCV_extendedPseudos(t *testing.T) {
t.Run("logic and minmax", func(t *testing.T) {
fn := firstTextRISCV(t, `#include "textflag.h"
TEXT ·l(SB), NOSPLIT, $0
ANDN X19, X20, X21
ANDN X19, X20
ORN X20, X19
MAX X26, X28, X29
MIN X29, X30, X5
MAX X5, X5
MAX X5, X5, X6
SEQZ X5, X6
NEG X5, X6
NOT X5
RDTIME X5
RET
`)
code := assembleRISCVHelper(t, fn)
// Words 0-10 up to the folded C.MV pair (halfwords 96 82 and 16 83),
// then SEQZ, NEG, NOT and RDTIME.
riscvWantsHex(t, code,
"93caf9ffb37a5a01"+"93cff9ff337afa01"+"934ffaffb3e9f901"+
"b32fae01b30ff041b34eae01b3fedf01b34ede01"+
"b3afee01b30ff041b342df01b3f25f00b3425f00"+
"9682"+"1683"+
"13b31200"+"33035040"+"93c2f2ff"+"f32210c0"+"67800000")
})
t.Run("rotate", func(t *testing.T) {
fn := firstTextRISCV(t, `#include "textflag.h"
TEXT ·r(SB), NOSPLIT, $0
ROR X10, X11, X12
ROR X10, X11
ROR $63, X11
RORIW $31, X13, X14
RORIW $1, X14, X15
RORIW $3, X14
RORW X15, X16, X17
RORW $31, X13
RET
`)
code := assembleRISCVHelper(t, fn)
// The third ROR carries the compressed C.SLLI (05 86) in mid-stream.
riscvWantsHex(t, code,
"b30fa040b39ff50133d6a50033e6cf00"+
"b30fa040b39ff501b3d5a500b3e5bf00"+
"93dff5038605b3e5bf00"+
"9bdff6011b97160033e7ef00"+
"9b5f17009b17f701b3e7ff00"+
"9b5f37001b17d70133e7ef00"+
"b30ff040bb1ff801bb58f800b3e81f01"+
"9bdff6019b961600b3e6df00"+"67800000")
})
t.Run("fp and branches", func(t *testing.T) {
fn := firstTextRISCV(t, `#include "textflag.h"
TEXT ·f(SB), NOSPLIT, $0
FABSD F1, F2
FSGNJD F1, F0, F2
FMADDD F1, F2, F3, F4
FMSUBD F1, F2, F3, F4
FNMSUBD F1, F2, F3, F4
BGT X5, X6, tgt
BLE X5, X6, tgt
BGTU X5, X6, tgt
BLEU X5, X6, tgt
tgt:
RDTIME X5
RET
`)
code := assembleRISCVHelper(t, fn)
riscvWantsHex(t, code,
"53a11022"+"53011022"+"4382201a4782201a4b82201a"+
"63485300635653006364530063725300"+ // blt/bge/bltu/bgeu x6, x5
"f32210c0"+"67800000")
})
}
// TestRISCV_amoWords pins the full AMO family: every AMO carries aq and rl
// (funct7 |= 3), LR is acquire (funct7 |= 2) and SC release (funct7 |= 1),
// exactly as GOARCH=riscv64 go tool asm encodes them.
func TestRISCV_amoWords(t *testing.T) {
fn := firstTextRISCV(t, `#include "textflag.h"
TEXT ·amo(SB), NOSPLIT, $0
AMOSWAPW X5, (X6), X7
AMOSWAPD X5, (X6), X7
AMOADDW X5, (X6), X7
AMOADDD X5, (X6), X7
AMOANDW X5, (X6), X7
AMOANDD X5, (X6), X7
AMOORW X5, (X6), X7
AMOORD X5, (X6), X7
AMOXORW X5, (X6), X7
AMOXORD X5, (X6), X7
AMOMAXW X5, (X6), X7
AMOMAXD X5, (X6), X7
AMOMAXUW X5, (X6), X7
AMOMAXUD X5, (X6), X7
AMOMINUW X5, (X6), X7
AMOMINUD X5, (X6), X7
LRW (X5), X6
LRD (X5), X6
SCW X5, (X6), X7
SCD X5, (X6), X7
RET
`)
code := assembleRISCVHelper(t, fn)
riscvWants(t, code,
0x0E5323AF, // amoswap.w
0x0E5333AF, // amoswap.d
0x065323AF, // amoaddd.w
0x065333AF, // amoadd.d
0x665323AF, // amoand.w
0x665333AF, // amoand.d
0x465323AF, // amoor.w
0x465333AF, // amoor.d
0x265323AF, // amoxor.w
0x265333AF, // amoxor.d
0xA65323AF, // amomax.w
0xA65333AF, // amomax.d
0xE65323AF, // amomaxu.w
0xE65333AF, // amomaxu.d
0xC65323AF, // amominu.w
0xC65333AF, // amominu.d
0x1402A32F, // lr.w (aq)
0x1402B32F, // lr.d
0x1A5323AF, // sc.w (rl)
0x1A5333AF, // sc.d
)
}
// TestRISCV_vectorWords pins the RVV slice and the VSET* encodings. The
// toolchain canonicalises an immediate avl to vsetivli even under the
// VSETVLI spelling (`VSETVLI $15` and `VSETIVLI $15` come out byte-
// identical), which is what the 0xC00 bit of the first word carries.
func TestRISCV_vectorWords(t *testing.T) {
fn := firstTextRISCV(t, `#include "textflag.h"
TEXT ·v(SB), NOSPLIT, $0
VSETVLI X5, E8, M8, TA, MA, X6
VSETIVLI $4, E32, M1, TA, MA, X0
VSETVLI $15, E32, M1, TA, MA, X12
VADDVV V1, V2, V3
VADDVX X12, V12, V12
VXORVV V8, V16, V24
VMSEQVX X12, V8, V0
VMSNEVV V8, V16, V0
VSLLVI $8, V28, V30
VSRLVI $25, V29, V29
VFIRSTM V0, X6
VIDV V12
VMV4RV V8, V24
VLE8V (X10), V8
VSE8V V24, (X10)
VSE32V V9, (X11)
VLSSEG4E32V (X14), X0, V0
VLSSEG8E32V (X10), X0, V4
RET
`)
code := assembleRISCVHelper(t, fn)
riscvWants(t, code,
0x0C32F357, // vsetvli x6, x5, vtype 0xc3 (E8, M8, TA, MA)
0xCD027057, // vsetivli x0, 4
0xCD07F657, // vsetivli x12, 15: VSETVLI $15 canonicalises to the same word
0x022081D7, // vadd.vv v3, v2, v1
0x02C64657, // vadd.vx v12, v12, x12
0x2F040C57, // vxor.vv v24, v16, v8
0x62864057, // vmseq.vx v0, v8, x12
0x67040057, // vmsne.vv v0, v16, v8
0x97C43F57, // vsll.vi v30, v28, 8
0xA3DCBED7, // vsrl.vi v29, v29, 25
0x4208A357, // vmfirst.m x6, v0
0x5208A657, // vid.v v12
0x9E81BC57, // vmv4r.v v24, v8
0x02050407, // vle8.v v8, (x10)
0x02050C27, // vse8.v v24, (x10)
0x0205E4A7, // vse32.v v9, (x11)
0x6A076007, // vlsseg4e32.v v0, (x14), x0
0xEA056207, // vlsseg8e32.v v4, (x10), x0
)
}
+362 -5
View File
@@ -41,7 +41,7 @@ const (
vexExtract
// vexRMRev is the reversed two-operand form `OP src, dst` with the source
// in ModRM.reg and the destination in r/m, the layout of the EVEX
// narrowing stores (VPMOVDW, VPMOVQD).
// narrowing stores (VPMOVDW, VPMOVQD) and of the non-temporal VMOVNTDQ.
vexRMRev
// vexRMSrcLen is the two-operand conversion form `OP src, dst` whose
// vector length follows the source: the packed-double → dword
@@ -52,6 +52,30 @@ const (
vexRMSrcLen
// vexZero is the no-operand form (VZEROUPPER).
vexZero
// vexZeroAll is the no-operand form that zeroes the full upper state
// (VZEROALL, the L = 1 twin of VZEROUPPER).
vexZeroAll
// vexNDS3GPR is the three-operand NDS form over general-purpose
// registers (ANDN, MULX): reg = dst, vvvv = src1, rm = src2, L = 0.
vexNDS3GPR
// vexImmRMGPR is the immediate form over general-purpose registers
// (RORX): reg = dst, rm = src, imm8 = op0, L = 0.
vexImmRMGPR
// vexRMOpGPR is the two-operand /digit form over general-purpose
// registers (BLSI, BLSMSK, BLSR): ModRM.reg = /digit, ModRM.rm = src
// (op0), VEX.vvvv = dst (op1), L = 0.
vexRMOpGPR
// vexCountGPR is the three-operand count form over general-purpose
// registers (SHLX, SHRX, SARX, BEXTR, BZHI): the first operand rides
// VEX.vvvv and the second is r/m, the opposite pairing of the ANDN
// family, with reg = dst (op2), L = 0.
vexCountGPR
// vexExtractGPR is the lane-extract-to-GPR form `OP $imm, xsrc, GPR/mem
// dst`: ModRM.reg = xsrc (op1), ModRM.rm = destination (op2), imm8 =
// op0, the VPEXTRB/W/D/Q layout. EVEX only; the destination never
// carries a vector length, so the register the L'L field follows is the
// XMM source.
vexExtractGPR
)
// vexSpec describes one VEX instruction's encoding parameters.
@@ -125,6 +149,12 @@ var vexTable = map[string]vexSpec{
"VMAXSS": {1, 0x5F, 0, 2, -1, vexNDS3},
// VEX.128/256.66.0F38.W1, fused multiply-add (NDS form).
"VFMADD231PD": {2, 0xB8, 1, 1, -1, vexNDS3},
// Scalar fused multiply-add (NDS form). The Go assembler carries the
// same 66 prefix as the packed forms on every FMA row, and W1 on the
// double-precision spellings, so SD shares PD's prefix/W pair and the
// scalar width rides on the W bit.
"VFMADD213SD": {2, 0xA9, 1, 1, -1, vexNDS3},
"VFNMADD231SD": {2, 0xBD, 1, 1, -1, vexNDS3},
// VEX.128/256.66.0F38.WIG, sign/zero extend and broadcast (reg=dst, rm=src,
// no vvvv).
@@ -192,6 +222,57 @@ var vexTable = map[string]vexSpec{
// VEX.128.0F.W0, no operands.
"VZEROUPPER": {1, 0x77, 0, 0, -1, vexZero},
// VEX.256.0F.W0, zero all vector registers (the L = 1 twin).
"VZEROALL": {1, 0x77, 0, 0, -1, vexZeroAll},
// VEX.128/256.66.0F38, byte shuffle shifts and the packed byte compare.
"VPSLLDQ": {1, 0x73, 0, 1, 7, vexShiftImm},
"VPSRLDQ": {1, 0x73, 0, 1, 3, vexShiftImm},
"VPCMPEQB": {1, 0x74, 0, 1, -1, vexNDS3},
// VEX.128/256.0F.WIG, packed single XOR (NDS form).
"VXORPS": {1, 0x57, 0, 0, -1, vexNDS3},
// VEX.256.66.0F3A.W0, two-source permutes and blends with an imm8 control.
"VPERM2F128": {3, 0x06, 0, 1, -1, vexNDS3Imm},
"VPBLENDD": {3, 0x02, 0, 1, -1, vexNDS3Imm},
// VEX.128/256.66.0F3A.WIG, byte align (NDS + imm8); the ZMM spelling
// falls through to the EVEX table.
"VPALIGNR": {3, 0x0F, 0, 1, -1, vexNDS3Imm},
// VEX.128/256.66.0F3A.W0, carry-less multiply ($imm, src2, src1, dst).
"VPCLMULQDQ": {3, 0x44, 0, 1, -1, vexNDS3Imm},
// VEX.128/256.66.0F3A.W1, GF(2^8) affine transform (NDS + imm8).
"VGF2P8AFFINEQB": {3, 0xCE, 1, 1, -1, vexNDS3Imm},
// BMI1/BMI2 general-register VEX forms (see vexNDS3GPR/vexImmRMGPR).
"ANDNL": {2, 0xF2, 0, 0, -1, vexNDS3GPR},
"ANDNQ": {2, 0xF2, 1, 0, -1, vexNDS3GPR},
"MULXL": {2, 0xF6, 0, 3, -1, vexNDS3GPR},
"MULXQ": {2, 0xF6, 1, 3, -1, vexNDS3GPR},
// VEX.NDS.LZ.0F38, the BMI2 three-operand bit ops: BEXTR and BZHI
// share the F7/F5 opcodes across W, the variable shifts carry their
// direction in the prefix (SHLX 66, SHRX F2, SARX F3) and PDEP/PEXT
// in F2/F3.
"BEXTRL": {2, 0xF7, 0, 0, -1, vexCountGPR},
"BEXTRQ": {2, 0xF7, 1, 0, -1, vexCountGPR},
"BZHIL": {2, 0xF5, 0, 0, -1, vexCountGPR},
"BZHIQ": {2, 0xF5, 1, 0, -1, vexCountGPR},
"SARXL": {2, 0xF7, 0, 2, -1, vexCountGPR},
"SARXQ": {2, 0xF7, 1, 2, -1, vexCountGPR},
"SHLXL": {2, 0xF7, 0, 1, -1, vexCountGPR},
"SHLXQ": {2, 0xF7, 1, 1, -1, vexCountGPR},
"SHRXL": {2, 0xF7, 0, 3, -1, vexCountGPR},
"SHRXQ": {2, 0xF7, 1, 3, -1, vexCountGPR},
"PDEPL": {2, 0xF5, 0, 3, -1, vexNDS3GPR},
"PDEPQ": {2, 0xF5, 1, 3, -1, vexNDS3GPR},
"PEXTL": {2, 0xF5, 0, 2, -1, vexNDS3GPR},
"PEXTQ": {2, 0xF5, 1, 2, -1, vexNDS3GPR},
// VEX.LZ.0F38.W, the BMI1 unary bit ops (src, dst: ModRM.reg = /digit,
// rm = src, vvvv = dst).
"BLSIL": {2, 0xF3, 0, 0, 3, vexRMOpGPR},
"BLSIQ": {2, 0xF3, 1, 0, 3, vexRMOpGPR},
"BLSMSKL": {2, 0xF3, 0, 0, 2, vexRMOpGPR},
"BLSMSKQ": {2, 0xF3, 1, 0, 2, vexRMOpGPR},
"BLSRL": {2, 0xF3, 0, 0, 1, vexRMOpGPR},
"BLSRQ": {2, 0xF3, 1, 0, 1, vexRMOpGPR},
"RORXL": {3, 0xF0, 0, 3, -1, vexImmRMGPR},
"RORXQ": {3, 0xF0, 1, 3, -1, vexImmRMGPR},
// VEX.128.0F.W0, mask-register test (KTESTW k1, k2: reg = dst, rm = src).
"KTESTW": {1, 0x99, 0, 0, -1, vexRM},
@@ -200,6 +281,14 @@ var vexTable = map[string]vexSpec{
// rm=scalar memory; SD is 256-bit only).
"VBROADCASTSS": {2, 0x18, 0, 1, -1, vexRM},
"VBROADCASTSD": {2, 0x19, 0, 1, -1, vexRM},
// VEX.256.66.0F38.W0, broadcast a 128-bit lane into both halves of a
// YMM (the encoder rejects an XMM destination, as go tool asm does).
"VBROADCASTI128": {2, 0x5A, 0, 1, -1, vexRM},
// VEX.128/256.66.0F.WIG, non-temporal store (vector source in reg,
// memory destination in rm).
"VMOVNTDQ": {1, 0xE7, 0, 1, -1, vexRMRev},
// VEX.128/256.66.0F38.W0, test (reg=dst, rm=src, no vvvv).
"VPTEST": {2, 0x17, 0, 1, -1, vexRM},
// VEX.66.0F38.W0, half-precision convert (reg=dst, rm=half-width
// source).
"VCVTPH2PS": {2, 0x13, 0, 1, -1, vexRM},
@@ -242,6 +331,128 @@ var vexTable = map[string]vexSpec{
"VCVTPD2DQY": {1, 0xE6, 0, 3, -1, vexRMSrcLen},
"VCVTTPD2DQX": {1, 0xE6, 0, 1, -1, vexRMSrcLen},
"VCVTTPD2DQY": {1, 0xE6, 0, 1, -1, vexRMSrcLen},
// --- the VEX forms the avx512enc corpus exercises alongside the EVEX
// spellings, read off the toolchain opcode tables ---
"VAESDEC": {2, 0xDE, 0, 1, -1, vexNDS3},
"VAESDECLAST": {2, 0xDF, 0, 1, -1, vexNDS3},
"VAESENC": {2, 0xDC, 0, 1, -1, vexNDS3},
"VAESENCLAST": {2, 0xDD, 0, 1, -1, vexNDS3},
"VANDNPD": {1, 0x55, 0, 1, -1, vexNDS3},
"VANDPD": {1, 0x54, 0, 1, -1, vexNDS3},
"VCOMISD": {1, 0x2F, 0, 1, -1, vexRM},
"VCVTSD2SS": {1, 0x5A, 0, 3, -1, vexNDS3},
"VCVTSS2SD": {1, 0x5A, 0, 2, -1, vexNDS3},
"VFMADD132PD": {2, 0x98, 1, 1, -1, vexNDS3},
"VFMADD132PS": {2, 0x98, 0, 1, -1, vexNDS3},
"VFMADD132SD": {2, 0x99, 1, 1, -1, vexNDS3},
"VFMADD132SS": {2, 0x99, 0, 1, -1, vexNDS3},
"VFMADD213PD": {2, 0xA8, 1, 1, -1, vexNDS3},
"VFMADD213PS": {2, 0xA8, 0, 1, -1, vexNDS3},
"VFMADD213SS": {2, 0xA9, 0, 1, -1, vexNDS3},
"VFMADD231PS": {2, 0xB8, 0, 1, -1, vexNDS3},
"VFMADD231SD": {2, 0xB9, 1, 1, -1, vexNDS3},
"VFMADD231SS": {2, 0xB9, 0, 1, -1, vexNDS3},
"VFMADDSUB132PD": {2, 0x96, 1, 1, -1, vexNDS3},
"VFMADDSUB132PS": {2, 0x96, 0, 1, -1, vexNDS3},
"VFMADDSUB213PD": {2, 0xA6, 1, 1, -1, vexNDS3},
"VFMADDSUB213PS": {2, 0xA6, 0, 1, -1, vexNDS3},
"VFMADDSUB231PD": {2, 0xB6, 1, 1, -1, vexNDS3},
"VFMADDSUB231PS": {2, 0xB6, 0, 1, -1, vexNDS3},
"VFMSUB132PD": {2, 0x9A, 1, 1, -1, vexNDS3},
"VFMSUB132PS": {2, 0x9A, 0, 1, -1, vexNDS3},
"VFMSUB132SD": {2, 0x9B, 1, 1, -1, vexNDS3},
"VFMSUB132SS": {2, 0x9B, 0, 1, -1, vexNDS3},
"VFMSUB213PD": {2, 0xAA, 1, 1, -1, vexNDS3},
"VFMSUB213PS": {2, 0xAA, 0, 1, -1, vexNDS3},
"VFMSUB213SD": {2, 0xAB, 1, 1, -1, vexNDS3},
"VFMSUB213SS": {2, 0xAB, 0, 1, -1, vexNDS3},
"VFMSUB231PD": {2, 0xBA, 1, 1, -1, vexNDS3},
"VFMSUB231PS": {2, 0xBA, 0, 1, -1, vexNDS3},
"VFMSUB231SD": {2, 0xBB, 1, 1, -1, vexNDS3},
"VFMSUB231SS": {2, 0xBB, 0, 1, -1, vexNDS3},
"VFMSUBADD132PD": {2, 0x97, 1, 1, -1, vexNDS3},
"VFMSUBADD132PS": {2, 0x97, 0, 1, -1, vexNDS3},
"VFMSUBADD213PD": {2, 0xA7, 1, 1, -1, vexNDS3},
"VFMSUBADD213PS": {2, 0xA7, 0, 1, -1, vexNDS3},
"VFMSUBADD231PD": {2, 0xB7, 1, 1, -1, vexNDS3},
"VFMSUBADD231PS": {2, 0xB7, 0, 1, -1, vexNDS3},
"VFNMADD132PD": {2, 0x9C, 1, 1, -1, vexNDS3},
"VFNMADD132PS": {2, 0x9C, 0, 1, -1, vexNDS3},
"VFNMADD132SD": {2, 0x9D, 1, 1, -1, vexNDS3},
"VFNMADD132SS": {2, 0x9D, 0, 1, -1, vexNDS3},
"VFNMADD213PD": {2, 0xAC, 1, 1, -1, vexNDS3},
"VFNMADD213PS": {2, 0xAC, 0, 1, -1, vexNDS3},
"VFNMADD213SD": {2, 0xAD, 1, 1, -1, vexNDS3},
"VFNMADD213SS": {2, 0xAD, 0, 1, -1, vexNDS3},
"VFNMADD231PD": {2, 0xBC, 1, 1, -1, vexNDS3},
"VFNMADD231PS": {2, 0xBC, 0, 1, -1, vexNDS3},
"VFNMADD231SS": {2, 0xBD, 0, 1, -1, vexNDS3},
"VFNMSUB132PD": {2, 0x9E, 1, 1, -1, vexNDS3},
"VFNMSUB132PS": {2, 0x9E, 0, 1, -1, vexNDS3},
"VFNMSUB132SD": {2, 0x9F, 1, 1, -1, vexNDS3},
"VFNMSUB132SS": {2, 0x9F, 0, 1, -1, vexNDS3},
"VFNMSUB213PD": {2, 0xAE, 1, 1, -1, vexNDS3},
"VFNMSUB213PS": {2, 0xAE, 0, 1, -1, vexNDS3},
"VFNMSUB213SD": {2, 0xAF, 1, 1, -1, vexNDS3},
"VFNMSUB213SS": {2, 0xAF, 0, 1, -1, vexNDS3},
"VFNMSUB231PD": {2, 0xBE, 1, 1, -1, vexNDS3},
"VFNMSUB231PS": {2, 0xBE, 0, 1, -1, vexNDS3},
"VFNMSUB231SD": {2, 0xBF, 1, 1, -1, vexNDS3},
"VFNMSUB231SS": {2, 0xBF, 0, 1, -1, vexNDS3},
"VGF2P8AFFINEINVQB": {3, 0xCF, 1, 1, -1, vexNDS3Imm},
"VGF2P8MULB": {2, 0xCF, 0, 1, -1, vexNDS3},
"VMOVNTDQA": {2, 0x2A, 0, 1, -1, vexRM},
"VMOVNTPD": {1, 0x2B, 0, 1, -1, vexRMRev},
"VORPD": {1, 0x56, 0, 1, -1, vexNDS3},
"VPADDSB": {1, 0xEC, 0, 1, -1, vexNDS3},
"VPADDSW": {1, 0xED, 0, 1, -1, vexNDS3},
"VPADDUSB": {1, 0xDC, 0, 1, -1, vexNDS3},
"VPADDUSW": {1, 0xDD, 0, 1, -1, vexNDS3},
"VPCMPEQQ": {2, 0x29, 0, 1, -1, vexNDS3},
"VPCMPEQW": {1, 0x75, 0, 1, -1, vexNDS3},
"VPCMPGTB": {1, 0x64, 0, 1, -1, vexNDS3},
"VPCMPGTD": {1, 0x66, 0, 1, -1, vexNDS3},
"VPCMPGTW": {1, 0x65, 0, 1, -1, vexNDS3},
"VPERMPS": {2, 0x16, 0, 1, -1, vexNDS3},
"VPEXTRB": {3, 0x14, 0, 1, -1, vexExtract},
"VPEXTRD": {3, 0x16, 0, 1, -1, vexExtract},
"VPEXTRQ": {3, 0x16, 1, 1, -1, vexExtract},
"VPINSRD": {3, 0x22, 0, 1, -1, vexNDS3Imm},
"VPINSRQ": {3, 0x22, 1, 1, -1, vexNDS3Imm},
"VPMULHRSW": {2, 0x0B, 0, 1, -1, vexNDS3},
"VPMULHW": {1, 0xE5, 0, 1, -1, vexNDS3},
"VPMULUDQ": {1, 0xF4, 0, 1, -1, vexNDS3},
"VPSADBW": {1, 0xF6, 0, 1, -1, vexNDS3},
"VPSUBSB": {1, 0xE8, 0, 1, -1, vexNDS3},
"VPSUBSW": {1, 0xE9, 0, 1, -1, vexNDS3},
"VPSUBUSB": {1, 0xD8, 0, 1, -1, vexNDS3},
"VPSUBUSW": {1, 0xD9, 0, 1, -1, vexNDS3},
"VPUNPCKHBW": {1, 0x68, 0, 1, -1, vexNDS3},
"VPUNPCKHQDQ": {1, 0x6D, 0, 1, -1, vexNDS3},
"VPUNPCKHWD": {1, 0x69, 0, 1, -1, vexNDS3},
"VPUNPCKLBW": {1, 0x60, 0, 1, -1, vexNDS3},
"VPUNPCKLWD": {1, 0x61, 0, 1, -1, vexNDS3},
"VSQRTPD": {1, 0x51, 0, 1, -1, vexRM},
"VSQRTSD": {1, 0x51, 0, 3, -1, vexNDS3},
"VSQRTSS": {1, 0x51, 0, 2, -1, vexNDS3},
"VUCOMISD": {1, 0x2E, 0, 1, -1, vexRM},
// VEX.0F.WIG, the plain-prefix single/double arithmetic and unpack
// spellings (no 66 prefix; WIG, so W = 0).
"VANDNPS": {1, 0x55, 0, 0, -1, vexNDS3},
"VANDPS": {1, 0x54, 0, 0, -1, vexNDS3},
"VORPS": {1, 0x56, 0, 0, -1, vexNDS3},
"VUNPCKLPS": {1, 0x14, 0, 0, -1, vexNDS3},
"VUNPCKHPS": {1, 0x15, 0, 0, -1, vexNDS3},
"VSQRTPS": {1, 0x51, 0, 0, -1, vexRM},
"VMOVNTPS": {1, 0x2B, 0, 0, -1, vexRMRev},
// VEX.128.66.0F, the scalar and packed compare forms.
"VCOMISS": {1, 0x2F, 0, 1, -1, vexRM},
"VUCOMISS": {1, 0x2E, 0, 0, -1, vexRM},
// VEX.128.0F.F3/F2.W0, the high/low word shuffles ($imm, src, dst).
"VPSHUFHW": {1, 0x70, 0, 2, -1, vexImmRM},
"VPSHUFLW": {1, 0x70, 0, 3, -1, vexImmRM},
}
// vexSrcLen maps a source-length conversion mnemonic (the X/Y spellings of
@@ -290,6 +501,8 @@ type vexMoveSpec struct {
var vexMoveTable = map[string]vexMoveSpec{
// VEX.128/256.F3.0F.WIG, unaligned integer move.
"VMOVDQU": {1, 2, 0x6F, 0x7F, 0, 0, 0, 0, true, false, false},
// VEX.128/256.66.0F.WIG, aligned integer move.
"VMOVDQA": {1, 1, 0x6F, 0x7F, 0, 0, 0, 0, true, false, false},
// VEX.128/256.66.0F.WIG, unaligned packed double move.
"VMOVUPD": {1, 1, 0x10, 0x11, 0, 0, 0, 0, true, false, false},
// VEX.128.66.0F.W0, 32-bit GPR/memory ↔ XMM.
@@ -324,6 +537,14 @@ func (e *enc) encodeVex(mnemUpper string, ops []Operand) error {
return fmt.Errorf("%s: vector register index %d needs an EVEX (AVX-512) instruction", mnemUpper, r.idx)
}
}
// VBROADCASTI128 broadcasts a 128-bit lane into a 256-bit destination
// only; an XMM destination is rejected exactly as go tool asm does.
if mnemUpper == "VBROADCASTI128" {
dstReg, ok := ops[len(ops)-1].(Reg)
if len(ops) != 2 || !ok || dstReg.size != 32 {
return fmt.Errorf("VBROADCASTI128 requires a YMM destination")
}
}
if ms, ok := vexMoveTable[mnemUpper]; ok {
return e.encodeVexMove(mnemUpper, ms, ops)
}
@@ -356,6 +577,18 @@ func (e *enc) encodeVex(mnemUpper string, ops []Operand) error {
return e.encodeVexRMSrcLen(mnemUpper, spec, ops)
case vexZero:
return e.encodeVexZero(mnemUpper, spec, ops)
case vexZeroAll:
return e.encodeVexZeroAll(mnemUpper, spec, ops)
case vexNDS3GPR:
return e.encodeVexNDS3GPR(spec, ops)
case vexImmRMGPR:
return e.encodeVexImmRMGPR(spec, ops)
case vexRMOpGPR:
return e.encodeVexRMOpGPR(spec, ops)
case vexCountGPR:
return e.encodeVexCountGPR(spec, ops)
case vexRMRev:
return e.encodeVexRMRev(spec, ops)
}
return fmt.Errorf("unhandled VEX form for %s", mnemUpper)
}
@@ -457,9 +690,10 @@ func (e *enc) encodeVexShiftImm(spec vexSpec, ops []Operand) error {
if !ok {
return fmt.Errorf("shift count must be an immediate")
}
srcReg, ok := src.(Reg)
if !ok || !srcReg.isVec() {
return fmt.Errorf("shift source must be a vector register")
// The count source is a vector register or memory; the VEX length
// follows the destination register either way.
if !vecOrMem(src) {
return fmt.Errorf("shift source must be a vector register or memory")
}
dstReg, ok := dst.(Reg)
if !ok || !dstReg.isVec() {
@@ -467,7 +701,7 @@ func (e *enc) encodeVexShiftImm(spec vexSpec, ops []Operand) error {
}
vvvvBar := 15 - (dstReg.idx & 15)
if err := e.emitVexFields(spec, dstReg.vecLenBit(), spec.opdigit, 0, vvvvBar, srcReg); err != nil {
if err := e.emitVexFields(spec, dstReg.vecLenBit(), spec.opdigit, 0, vvvvBar, src); err != nil {
return err
}
immByte, err := imm8(int64(immVal))
@@ -607,6 +841,129 @@ func (e *enc) encodeVexZero(mnem string, spec vexSpec, ops []Operand) error {
return nil
}
// encodeVexZeroAll encodes a no-operand instruction (VZEROALL), the L = 1
// twin of VZEROUPPER.
func (e *enc) encodeVexZeroAll(mnem string, spec vexSpec, ops []Operand) error {
if len(ops) != 0 {
return fmt.Errorf("%s expects no operands, got %d", mnem, len(ops))
}
// 2-byte VEX: R̄ = 1, v̄vvv = 1111 (unused), L = 1.
e.out = append(e.out, 0xC5, byte(1<<7|15<<3|1<<2|spec.pp), spec.opcode)
return nil
}
// encodeVexNDS3GPR encodes the three-operand NDS form over general-purpose
// registers (ANDN, MULX): OP src2, src1, dst with reg = dst, vvvv = src1,
// rm = src2 and L = 0.
func (e *enc) encodeVexNDS3GPR(spec vexSpec, ops []Operand) error {
if len(ops) != 3 {
return fmt.Errorf("VEX NDS instruction expects 3 operands, got %d", len(ops))
}
src2, src1, dst := ops[0], ops[1], ops[2]
dstReg, ok := dst.(Reg)
if !ok || dstReg.isVec() {
return fmt.Errorf("VEX destination must be a general-purpose register")
}
vvvvReg, ok := src1.(Reg)
if !ok || vvvvReg.isVec() {
return fmt.Errorf("VEX vvvv operand must be a general-purpose register")
}
rBit := 0
if dstReg.idx >= 8 {
rBit = 1
}
return e.emitVexFields(spec, 0, dstReg.idx&7, rBit, 15-(vvvvReg.idx&15), src2)
}
// encodeVexImmRMGPR encodes the immediate form over general-purpose
// registers (RORX): OP $imm, src, dst with reg = dst, rm = src, L = 0.
func (e *enc) encodeVexImmRMGPR(spec vexSpec, ops []Operand) error {
if len(ops) != 3 {
return fmt.Errorf("instruction expects 3 operands ($imm, src, dst), got %d", len(ops))
}
imm, src, dst := ops[0], ops[1], ops[2]
immVal, ok := imm.(Imm)
if !ok {
return fmt.Errorf("shift control must be an immediate")
}
dstReg, ok := dst.(Reg)
if !ok || dstReg.isVec() {
return fmt.Errorf("VEX destination must be a general-purpose register")
}
immByte, err := imm8(int64(immVal))
if err != nil {
return err
}
if err := e.emitVexFields(spec, 0, dstReg.idx&7, 0, 15, src); err != nil {
return err
}
e.out = append(e.out, immByte)
return nil
}
// encodeVexRMOpGPR encodes the two-operand /digit form over general-purpose
// registers (BLSI, BLSMSK, BLSR): OP src, dst with ModRM.reg = /digit,
// ModRM.rm = src and VEX.vvvv = dst.
func (e *enc) encodeVexRMOpGPR(spec vexSpec, ops []Operand) error {
if len(ops) != 2 {
return fmt.Errorf("instruction expects 2 operands (src, dst), got %d", len(ops))
}
src, dst := ops[0], ops[1]
dstReg, ok := dst.(Reg)
if !ok || dstReg.isVec() {
return fmt.Errorf("VEX destination must be a general-purpose register")
}
return e.emitVexFields(spec, 0, spec.opdigit, 0, 15-(dstReg.idx&15), src)
}
// encodeVexCountGPR encodes the three-operand count form over general-purpose
// registers (SHLX, SHRX, SARX, BEXTR, BZHI): OP src, count, dst with
// VEX.vvvv = src (op0), ModRM.rm = count (op1), ModRM.reg = dst (op2).
func (e *enc) encodeVexCountGPR(spec vexSpec, ops []Operand) error {
if len(ops) != 3 {
return fmt.Errorf("VEX count instruction expects 3 operands, got %d", len(ops))
}
src, count, dst := ops[0], ops[1], ops[2]
dstReg, ok := dst.(Reg)
if !ok || dstReg.isVec() {
return fmt.Errorf("VEX destination must be a general-purpose register")
}
countReg, ok := count.(Reg)
if !ok || countReg.isVec() {
return fmt.Errorf("VEX count operand must be a general-purpose register")
}
srcReg, ok := src.(Reg)
if !ok || srcReg.isVec() {
return fmt.Errorf("VEX count source must be a general-purpose register")
}
rBit := 0
if dstReg.idx >= 8 {
rBit = 1
}
return e.emitVexFields(spec, 0, dstReg.idx&7, rBit, 15-(srcReg.idx&15), count)
}
// encodeVexRMRev encodes the reversed two-operand form: OP src, dst with the
// vector source in ModRM.reg and the memory destination in r/m (VMOVNTDQ,
// a store with no register-destination form).
func (e *enc) encodeVexRMRev(spec vexSpec, ops []Operand) error {
if len(ops) != 2 {
return fmt.Errorf("store expects 2 operands, got %d", len(ops))
}
srcReg, ok := ops[0].(Reg)
if !ok || !srcReg.isVec() {
return fmt.Errorf("store source must be a vector register")
}
if !memOperand(ops[1]) {
return fmt.Errorf("store destination must be memory")
}
rBit := 0
if srcReg.idx >= 8 {
rBit = 1
}
return e.emitVexFields(spec, srcReg.vecLenBit(), srcReg.idx&7, rBit, 15, ops[1])
}
// encodeVexMove encodes a two-operand move (VMOVDQU, VMOVUPD, VMOVD, VMOVQ,
// VMOVSD), picking the direction-specific opcode and VEX.W. A vector→vector
// move uses the store-form layout (reg = source, rm = destination), matching
+116
View File
@@ -19,6 +19,65 @@ func vreg(t *testing.T, name string) Reg {
return r
}
// x86asmUnrecognised lists the VEX mnemonics whose machine code the
// golang.org/x/arch decoder cannot resolve; their bytes are verified against
// go tool asm in the ground-truth tests instead.
var x86asmUnrecognised = map[string]bool{
"ANDNL": true,
"ANDNQ": true,
"MULXL": true,
"MULXQ": true,
"RORXL": true,
"RORXQ": true,
"VFMADD213SD": true,
"VFNMADD231SD": true,
// The scalar FMA spellings the decoder's tables lack entirely.
"VFMADD132SD": true,
"VFMADD132SS": true,
"VFMADD213SS": true,
"VFMADD231SD": true,
"VFMADD231SS": true,
"VFMSUB132SD": true,
"VFMSUB132SS": true,
"VFMSUB213SD": true,
"VFMSUB213SS": true,
"VFMSUB231SD": true,
"VFMSUB231SS": true,
"VFNMADD132SD": true,
"VFNMADD132SS": true,
"VFNMADD213SD": true,
"VFNMADD213SS": true,
"VFNMADD231SS": true,
"VFNMSUB132SD": true,
"VFNMSUB132SS": true,
"VFNMSUB213SD": true,
"VFNMSUB213SS": true,
"VFNMSUB231SD": true,
"VFNMSUB231SS": true,
// The BMI1 unary bit ops the decoder's AVX tables lack.
"BLSIL": true,
"BLSIQ": true,
"BLSMSKL": true,
"BLSMSKQ": true,
"BLSRL": true,
"BLSRQ": true,
// The BMI2 bit ops whose W1/LZ rows the decoder misses.
"BEXTRL": true,
"BEXTRQ": true,
"BZHIL": true,
"BZHIQ": true,
"PDEPL": true,
"PDEPQ": true,
"PEXTL": true,
"PEXTQ": true,
"SARXL": true,
"SARXQ": true,
"SHLXL": true,
"SHLXQ": true,
"SHRXL": true,
"SHRXQ": true,
}
// TestVexNDS3 encodes `mnem Y0, Y1, Y2` for every three-operand NDS
// instruction and verifies it round-trips through the x86 decoder to the same
// mnemonic. A wrong opcode/map/pp surfaces as a different decoded instruction.
@@ -37,8 +96,15 @@ func TestVexNDS3(t *testing.T) {
t.Errorf("%s: Encode: %v", mnem, err)
continue
}
// The x86 decoder's table lacks a handful of rows the Go assembler
// emits (the scalar 213/231 FMA spellings among them); those are
// pinned byte for byte against go tool asm in TestVexGroundTruth
// instead of round-tripped here.
inst, err := x86asm.Decode(code, 64)
if err != nil {
if strings.Contains(err.Error(), "unrecognized instruction") && x86asmUnrecognised[mnem] {
continue
}
t.Errorf("%s: Decode(% x): %v", mnem, err, code)
continue
}
@@ -184,6 +250,50 @@ func TestVexGroundTruth(t *testing.T) {
{"VMULSD X0,X1,X1", "VMULSD", []Operand{vreg(t, "X0"), vreg(t, "X1"), vreg(t, "X1")}, "c5f359c8", ""},
{"VFMADD231PD Y14,Y12,Y8", "VFMADD231PD", []Operand{vreg(t, "Y14"), vreg(t, "Y12"), vreg(t, "Y8")}, "c4429db8c6", ""},
{"VFMADD231PD (DI),Y12,Y8", "VFMADD231PD", []Operand{Ptr(DI, 0, 32), vreg(t, "Y12"), vreg(t, "Y8")}, "c4629db807", ""},
{"VFMADD213SD X0,X1,X2", "VFMADD213SD", []Operand{vreg(t, "X0"), vreg(t, "X1"), vreg(t, "X2")}, "c4e2f1a9d0", ""},
{"VFNMADD231SD X0,X1,X2", "VFNMADD231SD", []Operand{vreg(t, "X0"), vreg(t, "X1"), vreg(t, "X2")}, "c4e2f1bdd0", ""},
// Packed single XOR and byte compare (NDS form).
{"VXORPS Y0,Y1,Y2", "VXORPS", []Operand{vreg(t, "Y0"), vreg(t, "Y1"), vreg(t, "Y2")}, "c5f457d0", ""},
{"VPCMPEQB Y0,Y1,Y2", "VPCMPEQB", []Operand{vreg(t, "Y0"), vreg(t, "Y1"), vreg(t, "Y2")}, "c5f574d0", ""},
// Octa byte shifts (vvvv carries the destination).
{"VPSLLDQ $2,X0,X1", "VPSLLDQ", []Operand{Imm(2), vreg(t, "X0"), vreg(t, "X1")}, "c5f173f802", ""},
{"VPSRLDQ $2,Y0,Y1", "VPSRLDQ", []Operand{Imm(2), vreg(t, "Y0"), vreg(t, "Y1")}, "c5f573d802", ""},
// Two-source shuffle, blend and carry-less multiply (NDS + imm8).
{"VPERM2F128 $3,Y0,Y1,Y2", "VPERM2F128", []Operand{Imm(3), vreg(t, "Y0"), vreg(t, "Y1"), vreg(t, "Y2")}, "c4e37506d003", ""},
{"VPBLENDD $3,X0,X1,X2", "VPBLENDD", []Operand{Imm(3), vreg(t, "X0"), vreg(t, "X1"), vreg(t, "X2")}, "c4e37102d003", ""},
{"VPBLENDD $3,Y0,Y1,Y2", "VPBLENDD", []Operand{Imm(3), vreg(t, "Y0"), vreg(t, "Y1"), vreg(t, "Y2")}, "c4e37502d003", ""},
{"VPCLMULQDQ $0,X0,X1,X2", "VPCLMULQDQ", []Operand{Imm(0), vreg(t, "X0"), vreg(t, "X1"), vreg(t, "X2")}, "c4e37144d000", ""},
{"VGF2P8AFFINEQB $0,X0,X1,X2", "VGF2P8AFFINEQB", []Operand{Imm(0), vreg(t, "X0"), vreg(t, "X1"), vreg(t, "X2")}, "c4e3f1ced000", ""},
// Two-operand test and the non-temporal and broadcast stores.
{"VPTEST X0,X1", "VPTEST", []Operand{vreg(t, "X0"), vreg(t, "X1")}, "c4e27917c8", ""},
{"VPTEST Y0,Y1", "VPTEST", []Operand{vreg(t, "Y0"), vreg(t, "Y1")}, "c4e27d17c8", ""},
{"VMOVNTDQ Y0,(AX)", "VMOVNTDQ", []Operand{vreg(t, "Y0"), Ptr(AX, 0, 32)}, "c5fde700", ""},
{"VMOVNTDQ X0,(AX)", "VMOVNTDQ", []Operand{vreg(t, "X0"), Ptr(AX, 0, 16)}, "c5f9e700", ""},
{"VBROADCASTI128 (AX),Y1", "VBROADCASTI128", []Operand{Ptr(AX, 0, 16), vreg(t, "Y1")}, "c4e27d5a08", ""},
// Aligned integer move and the full zeroing form.
{"VMOVDQA X0,X1", "VMOVDQA", []Operand{vreg(t, "X0"), vreg(t, "X1")}, "c5f97fc1", ""},
{"VMOVDQA (AX),X1", "VMOVDQA", []Operand{Ptr(AX, 0, 16), vreg(t, "X1")}, "c5f96f08", ""},
{"VMOVDQA Y0,Y1", "VMOVDQA", []Operand{vreg(t, "Y0"), vreg(t, "Y1")}, "c5fd7fc1", ""},
{"VZEROALL", "VZEROALL", []Operand{}, "c5fc77", ""},
// BMI1/BMI2 general-register VEX forms.
{"ANDNL AX,BX,CX", "ANDNL", []Operand{AX, BX, CX}, "c4e260f2c8", ""},
{"ANDNQ AX,BX,CX", "ANDNQ", []Operand{AX, BX, CX}, "c4e2e0f2c8", ""},
{"MULXL AX,BX,CX", "MULXL", []Operand{AX, BX, CX}, "c4e263f6c8", ""},
{"MULXQ AX,BX,CX", "MULXQ", []Operand{AX, BX, CX}, "c4e2e3f6c8", ""},
{"RORXL $3,AX,CX", "RORXL", []Operand{Imm(3), AX, CX}, "c4e37bf0c803", ""},
{"RORXQ $3,AX,CX", "RORXQ", []Operand{Imm(3), AX, CX}, "c4e3fbf0c803", ""},
// BMI2 variable shifts and bit ops (three general registers).
{"SHLXL AX,CX,R15", "SHLXL", []Operand{AX, CX, vreg(t, "R15")}, "c46279f7f9", ""},
{"SHRXQ R8,DX,AX", "SHRXQ", []Operand{vreg(t, "R8"), DX, AX}, "c4e2bbf7c2", ""},
{"SARXQ AX,DX,R9", "SARXQ", []Operand{AX, DX, vreg(t, "R9")}, "c462faf7ca", ""},
{"BEXTRL AX,CX,R15", "BEXTRL", []Operand{AX, CX, vreg(t, "R15")}, "c46278f7f9", ""},
{"BZHIQ AX,CX,R15", "BZHIQ", []Operand{AX, CX, vreg(t, "R15")}, "c462f8f5f9", ""},
{"PDEPQ AX,CX,R15", "PDEPQ", []Operand{AX, CX, vreg(t, "R15")}, "c462f3f5f8", ""},
{"PEXTQ AX,CX,R15", "PEXTQ", []Operand{AX, CX, vreg(t, "R15")}, "c462f2f5f8", ""},
// BMI1 unary bit ops (src, dst: /digit in ModRM.reg, dst in vvvv).
{"BLSIL AX,CX", "BLSIL", []Operand{AX, CX}, "c4e270f3d8", ""},
{"BLSRQ AX,CX", "BLSRQ", []Operand{AX, CX}, "c4e2f0f3c8", ""},
{"BLSMSKQ AX,CX", "BLSMSKQ", []Operand{AX, CX}, "c4e2f0f3d0", ""},
// Two-operand reg/rm form (v̄vvv must be 1111).
{"VPMOVSXDQ X0,Y4", "VPMOVSXDQ", []Operand{vreg(t, "X0"), vreg(t, "Y4")}, "c4e27d25e0", ""},
{"VPMOVSXWD (SI),Y0", "VPMOVSXWD", []Operand{Ptr(SI, 0, 8), vreg(t, "Y0")}, "c4e27d2306", ""},
@@ -287,6 +397,12 @@ func TestVexGroundTruth(t *testing.T) {
}
inst, err := x86asm.Decode(code, 64)
if err != nil {
// The decoder's AVX/BMI table lacks a few rows the Go
// assembler emits (the GPR VEX forms and the scalar FMA
// spellings); their bytes are the ground truth here.
if x86asmUnrecognised[c.mnem] {
continue
}
t.Errorf("%s: Decode(% x): %v", c.name, code, err)
continue
}
+17 -7
View File
@@ -151,11 +151,21 @@ type Immediate struct {
// Address is a non-immediate operand: a register, a memory reference, a symbol
// reference or a label. Fields are populated best-effort from the syntax.
type Address struct {
Sym *Symbol // name reference (bare ident, or name+off(pseudo))
Base string // base register, from (base)
Index string // index register, from (index*scale)
Scale int // index scale; 0 when absent
Offset int64 // leading displacement, from off(base)
HasOff bool // a leading displacement is present
Shift string // verbatim arm64 shift suffix, e.g. "<< 2"
Sym *Symbol // name reference (bare ident, or name+off(pseudo))
Base string // base register, from (base)
Index string // index register, from (index*scale)
Scale int // index scale; 0 when absent
Offset int64 // leading displacement, from off(base)
HasOff bool // a leading displacement is present
Shift string // verbatim arm64 shift suffix, e.g. "<< 2"
Range *RegRange // bracketed register range; nil for every other form
}
// RegRange is a bracketed register range, [Z0-Z3]: the amd64 spelling of
// the four-register source of the 4FMAPS/4VNNIW families. Lo and Hi carry
// the verbatim register spellings; the range is inclusive at both ends.
type RegRange struct {
Lo string
Hi string
Pos token.Position
}
+367
View File
@@ -0,0 +1,367 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
package main
import (
"errors"
"fmt"
"go/ast"
"go/build"
"go/constant"
"go/parser"
"go/token"
"go/types"
"os"
"path/filepath"
"regexp"
"strings"
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
)
// go_asm.h is the header the Go compiler writes for every package that
// carries assembly (the compiler's -asmhdr output): "#define const_NAME
// value" for each package constant, and for each named struct type
// "#define TYPE__size size" plus one "#define TYPE_field offset" per field.
// GOROOT assembly includes it, and a standalone assembler has no compiler
// to have produced it, so gasm generates the equivalent itself: the package
// the .s file lives in is parsed and type-checked here, with the target
// architecture's own sizes, and the same defines are written out. The
// type-checking GOOS is selected by the caller: a GOOS-specific file
// (sys_darwin_arm64.s) needs its platform's defines, which a header from
// the ambient GOOS silently omits.
//
// The emitter mirrors cmd/compile's dumpasmhdr exactly: constants come out
// as "const_NAME", struct entries as "NAME__size" followed by the fields in
// declaration order, blank names are skipped, and float and complex
// constants are omitted (the assembler carries integers, bools and strings
// only). Aliases to structs are emitted, generic types are not: they have
// no fixed size. A define the assembly references but this header does not
// carry surfaces later as the assembler's own "undefined" diagnostic naming
// the define, which is the honest failure.
// goAsmInclude matches the #include "go_asm.h" directive, tolerant of
// whitespace, so the wiring knows which files need a generated header
// before the preprocessor runs and would report the header as missing.
var goAsmInclude = regexp.MustCompile(`(?m)^\s*#\s*include\s+"go_asm\.h"`)
// needsGoAsmHeader reports whether src includes go_asm.h.
func needsGoAsmHeader(src string) bool {
return goAsmInclude.MatchString(src)
}
// goAsmHeaderResolved reports whether the include of go_asm.h from a file in
// asmDir already resolves: to a header in the package directory itself, or
// in one of the -I directories, the way the preprocessor searches. Only an
// unresolved include is generated for; a header someone placed by hand is
// the tool the author chose, and it also wins the preprocessor's own search
// order, so generating a second copy would be dead weight at best.
func goAsmHeaderResolved(asmDir string, dirs []string) bool {
candidates := []string{filepath.Join(asmDir, "go_asm.h")}
for _, d := range dirs {
candidates = append(candidates, filepath.Join(d, "go_asm.h"))
}
for _, candidate := range candidates {
if st, err := os.Stat(candidate); err == nil && !st.IsDir() {
return true
}
}
return false
}
// generateGoAsmHeader type-checks the Go package in pkgDir for goos and
// goarch, writes its go_asm.h equivalent into dir, and returns dir. An
// empty goos means the ambient one. The caller owns the directory and its
// removal.
func generateGoAsmHeader(pkgDir, goos, goarch, dir string) (string, error) {
if goos == "" {
goos = build.Default.GOOS
}
imp := newSourceImporter(goos, goarch)
if imp.sizes == nil {
return "", fmt.Errorf("go_asm.h: unknown GOARCH %q", goarch)
}
bp, err := imp.ctxt.ImportDir(pkgDir, 0)
if err != nil {
return "", fmt.Errorf("go_asm.h for GOARCH %s in %s: %w", goarch, pkgDir, err)
}
files, errs := imp.parse(bp)
if len(errs) > 0 {
return "", fmt.Errorf("go_asm.h for GOARCH %s in %s: %s", goarch, pkgDir, errorList(errs))
}
_, info, errs := imp.checkPackage(bp, files)
if len(errs) > 0 {
return "", fmt.Errorf("go_asm.h for GOARCH %s in %s: package does not type-check: %s", goarch, pkgDir, errorList(errs))
}
var b strings.Builder
fmt.Fprintf(&b, "// generated by gasm from package %s (GOOS %s, GOARCH %s)\n\n", bp.Name, goos, goarch)
// Files in the build's own order and declarations in source order: the
// same walk the compiler's reader makes, so the header reads the same
// way the toolchain's does. Order carries no meaning to the assembler
// (defines form a table), only to a human diffing against one.
for _, f := range files {
for _, decl := range f.Decls {
gd, ok := decl.(*ast.GenDecl)
if !ok {
continue
}
for _, spec := range gd.Specs {
switch gd.Tok {
case token.CONST:
vs, ok := spec.(*ast.ValueSpec)
if !ok {
continue
}
for _, name := range vs.Names {
emitConst(&b, info.Defs[name], name.Name)
}
case token.TYPE:
ts, ok := spec.(*ast.TypeSpec)
if !ok {
continue
}
emitStruct(&b, imp.sizes, info.Defs[ts.Name], ts.Name.Name)
}
}
}
}
if err := os.MkdirAll(dir, 0o755); err != nil {
return "", fmt.Errorf("go_asm.h for GOARCH %s in %s: %w", goarch, pkgDir, err)
}
out := filepath.Join(dir, "go_asm.h")
if err := os.WriteFile(out, []byte(b.String()), 0o644); err != nil {
return "", fmt.Errorf("go_asm.h for GOARCH %s in %s: %w", goarch, pkgDir, err)
}
return dir, nil
}
// emitConst writes one const define, skipping what the toolchain skips:
// blank names, and float and complex values the assembler has no syntax for.
func emitConst(b *strings.Builder, obj types.Object, name string) {
c, ok := obj.(*types.Const)
if !ok || name == "_" {
return
}
switch c.Val().Kind() {
case constant.Float, constant.Complex, constant.Unknown:
return
}
fmt.Fprintf(b, "#define const_%s %s\n", name, c.Val().ExactString())
}
// emitStruct writes one named struct type's size and field offsets,
// skipping what the toolchain skips: blank names, non-struct types, and
// generic types, whose size depends on their instantiation.
func emitStruct(b *strings.Builder, sizes types.Sizes, obj types.Object, name string) {
tn, ok := obj.(*types.TypeName)
if !ok || name == "_" {
return
}
t := types.Unalias(tn.Type())
// Generic types are spelled *types.Named with a type-parameter list;
// a plain struct type or an instantiated one carries none.
if named, ok := t.(*types.Named); ok && named.TypeParams().Len() > 0 {
return
}
st, ok := t.Underlying().(*types.Struct)
if !ok {
return
}
fmt.Fprintf(b, "#define %s__size %d\n", name, sizes.Sizeof(t))
fields := make([]*types.Var, st.NumFields())
for i := range st.NumFields() {
fields[i] = st.Field(i)
}
for i, off := range sizes.Offsetsof(fields) {
fld := fields[i]
if fld.Name() == "_" {
continue
}
fmt.Fprintf(b, "#define %s_%s %d\n", name, fld.Name(), off)
}
}
// errorList renders at most three errors, enough to say what is wrong
// without burying the diagnostic the caller actually reads.
func errorList(errs []error) string {
if len(errs) > 3 {
errs = errs[:3]
}
msgs := make([]string, len(errs))
for i, err := range errs {
msgs[i] = err.Error()
}
return strings.Join(msgs, "; ")
}
// sourceImporter type-checks imported packages from source with the target
// architecture's sizes. go/importer's "source" importer pins the host
// GOARCH, which would lay out imported types (internal/cpu, internal/abi)
// for the wrong target on a cross-architecture header, so the recursion is
// carried here with one build context and one sizes instance per
// architecture.
type sourceImporter struct {
fset *token.FileSet
ctxt *build.Context
sizes types.Sizes
pkgs map[string]*types.Package
}
// newSourceImporter returns the importer for one target GOOS and GOARCH.
// Cgo is disabled so the file set is deterministic and independent of the
// host's C toolchain: cgo-tagged files drop out of the build exactly as
// they do from a CGO_ENABLED=0 build, whose assembly is what gasm targets.
func newSourceImporter(goos, goarch string) *sourceImporter {
ctxt := new(build.Context)
*ctxt = build.Default
ctxt.GOOS = goos
ctxt.GOARCH = goarch
ctxt.CgoEnabled = false
return &sourceImporter{
fset: token.NewFileSet(),
ctxt: ctxt,
sizes: types.SizesFor("gc", goarch),
pkgs: map[string]*types.Package{},
}
}
// Import type-checks one imported package and memoises it. "unsafe" must
// resolve to go/types' own package, never to the source in GOROOT/src/unsafe:
// the source declares Sizeof and Offsetof as ordinary functions over
// ArbitraryType, and checking against that signature rejects half the
// unsafe arithmetic the gc compiler accepts, which is exactly the divergence
// srcimporter guards against the same way.
func (im *sourceImporter) Import(path string) (*types.Package, error) {
if path == "unsafe" {
return types.Unsafe, nil
}
if p, ok := im.pkgs[path]; ok {
return p, nil
}
bp, err := im.ctxt.Import(path, "", 0)
if err != nil {
return nil, err
}
files, errs := im.parse(bp)
if len(errs) > 0 {
return nil, errors.New(errorList(errs))
}
pkg, _, _ := im.checkPackage(bp, files)
im.pkgs[path] = pkg
return pkg, nil
}
// parse reads the build package's Go files. Import-level failures (no Go
// files for the target, unreadable files) come back as errors, and the
// type-check decides the rest.
func (im *sourceImporter) parse(bp *build.Package) ([]*ast.File, []error) {
if len(bp.GoFiles) == 0 {
return nil, []error{fmt.Errorf("no Go source files for GOOS=%s GOARCH=%s", im.ctxt.GOOS, im.ctxt.GOARCH)}
}
var (
files []*ast.File
errs []error
)
for _, name := range bp.GoFiles {
f, err := parser.ParseFile(im.fset, filepath.Join(bp.Dir, name), nil, parser.SkipObjectResolution)
if err != nil {
errs = append(errs, err)
continue
}
files = append(files, f)
}
return files, errs
}
// checkPackage type-checks one package's files with the importer's sizes,
// recording every error: a header from a package that does not type-check
// could silently mis-state an offset, so the caller refuses the header
// rather than trusting it. The returned Defs map backs the root package's
// emission walk; imports only need the checked package itself.
func (im *sourceImporter) checkPackage(bp *build.Package, files []*ast.File) (*types.Package, *types.Info, []error) {
var errs []error
conf := &types.Config{
Importer: im,
Sizes: im.sizes,
Error: func(err error) { errs = append(errs, err) },
}
info := &types.Info{Defs: map[*ast.Ident]types.Object{}}
pkg, _ := conf.Check(bp.ImportPath, im.fset, files, info)
return pkg, info, errs
}
// asmhdrCache generates one go_asm.h per package directory and target
// architecture under one temp root, for callers that assemble many files
// (the corpus audit). Failures are cached too: a package that does not
// type-check must not be re-checked once per file.
type asmhdrCache struct {
root string
dirs map[string]string // "pkgDir\x00goos\x00goarch" -> directory holding go_asm.h
errs map[string]error
}
func newAsmhdrCache() (*asmhdrCache, error) {
root, err := os.MkdirTemp("", "gasm-asmhdr")
if err != nil {
return nil, err
}
return &asmhdrCache{root: root, dirs: map[string]string{}, errs: map[string]error{}}, nil
}
// dirFor returns the directory holding the generated go_asm.h for pkgDir
// under goos and goarch, generating it on first use. An empty goos means
// the ambient one, resolved here so that one package cannot generate twice
// under an explicit and an implicit spelling of the same GOOS.
func (c *asmhdrCache) dirFor(pkgDir, goos, goarch string) (string, error) {
if goos == "" {
goos = build.Default.GOOS
}
key := pkgDir + "\x00" + goos + "\x00" + goarch
if dir, ok := c.dirs[key]; ok {
return dir, nil
}
if err, ok := c.errs[key]; ok {
return "", err
}
dir := filepath.Join(c.root, fmt.Sprintf("h%d_%s_%s", len(c.dirs), goos, goarch))
if _, err := generateGoAsmHeader(pkgDir, goos, goarch, dir); err != nil {
c.errs[key] = err
return "", err
}
c.dirs[key] = dir
return dir, nil
}
// close removes the temp root.
func (c *asmhdrCache) close() { os.RemoveAll(c.root) }
// ensureGoAsmHeader prepares the include directory a file that includes
// go_asm.h needs: the generated header for the package in path's directory,
// for the file's target GOOS and architecture. It reports a usage error
// when the architecture cannot be determined, and passes through the
// generator's diagnostics, which name the package.
func ensureGoAsmHeader(path string, target arch.Arch, goos string, cache *asmhdrCache) (string, func(), error) {
if path == "-" {
return "", nil, errors.New("cannot generate go_asm.h for standard input (no package directory)")
}
if target == arch.Unknown {
return "", nil, errors.New("a file that includes go_asm.h needs a target architecture: name the file _<arch>.s or pass -GOARCH")
}
if cache != nil {
dir, err := cache.dirFor(filepath.Dir(path), goos, goarchName(target))
return dir, func() {}, err
}
root, err := os.MkdirTemp("", "gasm-asmhdr")
if err != nil {
return "", nil, err
}
dir, err := generateGoAsmHeader(filepath.Dir(path), goos, goarchName(target), root)
if err != nil {
os.RemoveAll(root)
return "", nil, err
}
return dir, func() { os.RemoveAll(root) }, nil
}
+311
View File
@@ -0,0 +1,311 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
package main
import (
"os"
"path/filepath"
"strings"
"testing"
)
// writePkg lays out a minimal Go package in a temp directory.
func writePkg(t *testing.T, files map[string]string) string {
t.Helper()
dir := t.TempDir()
for name, src := range files {
if err := os.WriteFile(filepath.Join(dir, name), []byte(src), 0o644); err != nil {
t.Fatal(err)
}
}
return dir
}
// generateFor generates the header for dir and returns its text.
func generateFor(t *testing.T, dir, goarch string) string {
t.Helper()
hdrDir, err := generateGoAsmHeader(dir, goarch, t.TempDir())
if err != nil {
t.Fatalf("generateGoAsmHeader(%q, %s): %v", dir, goarch, err)
}
b, err := os.ReadFile(filepath.Join(hdrDir, "go_asm.h"))
if err != nil {
t.Fatal(err)
}
return string(b)
}
func TestGenerateGoAsmHeaderShape(t *testing.T) {
dir := writePkg(t, map[string]string{"sample.go": `package sample
const bufSize = 1024
const (
a = iota * 8
b
c
)
const (
strConst = "hello"
boolConst = true
floatConst = 1.5
_ = "the blank identifier is skipped"
)
const shift = 1 << 20
type reader struct {
r int64
w int64
_ [4]byte
name string
}
type scalar int
type aliased struct {
k uint32
v uint32
}
type alias = aliased
`})
hdr := generateFor(t, dir, "amd64")
want := []string{
"#define const_bufSize 1024",
// iota resolves through go/types, one define per name.
"#define const_a 0",
"#define const_b 8",
"#define const_c 16",
`#define const_strConst "hello"`,
"#define const_boolConst true",
// Floats are the toolchain's own skip, as are blank names.
"#define const_shift 1048576",
// The blank field still occupies its bytes: the pad after w runs to
// the string's 8-byte alignment.
"#define reader__size 40",
"#define reader_r 0",
"#define reader_w 8",
"#define reader_name 24",
// Non-struct named types carry no defines; aliases to structs do.
"#define aliased__size 8",
"#define aliased_k 0",
"#define aliased_v 4",
"#define alias__size 8",
"#define alias_k 0",
"#define alias_v 4",
}
for _, w := range want {
if !strings.Contains(hdr, w+"\n") {
t.Errorf("header misses %q\ngot:\n%s", w, hdr)
}
}
for _, banned := range []string{"#define const_floatConst", "#define _ ", "#define scalar"} {
if strings.Contains(hdr, banned) {
t.Errorf("header must not carry %s\ngot:\n%s", banned, hdr)
}
}
}
func TestGenerateGoAsmHeaderPerArch(t *testing.T) {
dir := writePkg(t, map[string]string{
"common.go": `package perarch
type layout struct {
a int32
p uintptr
}
`,
// The build-tagged file set is part of the contract: a per-arch
// package is exactly how internal/cpu declares its layouts.
"const_amd64.go": `//go:build amd64
package perarch
const flavour = 1
`,
"const_arm64.go": `//go:build arm64
package perarch
const flavour = 2
`,
})
amd64 := generateFor(t, dir, "amd64")
arm64 := generateFor(t, dir, "arm64")
if !strings.Contains(amd64, "#define const_flavour 1\n") {
t.Errorf("amd64 header misses const_flavour 1:\n%s", amd64)
}
if !strings.Contains(arm64, "#define const_flavour 2\n") {
t.Errorf("arm64 header misses const_flavour 2:\n%s", arm64)
}
if strings.Contains(arm64, "#define const_flavour 1\n") {
t.Errorf("arm64 header must not carry the amd64 file's value")
}
// SizesFor makes the layout the target's: uintptr is 4 bytes wide on
// 386 and 8 on amd64, which must move p and grow the struct.
if !strings.Contains(amd64, "#define layout__size 16\n") || !strings.Contains(amd64, "#define layout_p 8\n") {
t.Errorf("amd64 layout wrong:\n%s", amd64)
}
w386 := generateFor(t, dir, "386")
if !strings.Contains(w386, "#define layout__size 8\n") || !strings.Contains(w386, "#define layout_p 4\n") {
t.Errorf("386 layout wrong:\n%s", w386)
}
}
func TestGenerateGoAsmHeaderErrors(t *testing.T) {
t.Run("type error", func(t *testing.T) {
dir := writePkg(t, map[string]string{"bad.go": `package bad
const x = undefinedIdent
`})
_, err := generateGoAsmHeader(dir, "amd64", t.TempDir())
if err == nil {
t.Fatal("generation must fail for a package that does not type-check")
}
if !strings.Contains(err.Error(), dir) {
t.Errorf("error must name the package directory: %v", err)
}
if !strings.Contains(err.Error(), "type-check") {
t.Errorf("error must say the package does not type-check: %v", err)
}
})
t.Run("no go files", func(t *testing.T) {
dir := t.TempDir()
_, err := generateGoAsmHeader(dir, "amd64", t.TempDir())
if err == nil {
t.Fatal("generation must fail without Go files")
}
if !strings.Contains(err.Error(), dir) {
t.Errorf("error must name the package directory: %v", err)
}
})
}
func TestNeedsGoAsmHeader(t *testing.T) {
yes := "#include \"go_asm.h\"\n#include \"textflag.h\"\n"
no := "#include \"textflag.h\"\n#include \"funcdata.h\"\n"
if !needsGoAsmHeader(yes) {
t.Error("needsGoAsmHeader(missing on a go_asm.h include)")
}
if needsGoAsmHeader(no) {
t.Error("needsGoAsmHeader claims other headers need generation")
}
}
func TestGoAsmHeaderResolved(t *testing.T) {
dir := t.TempDir()
if goAsmHeaderResolved(dir, nil) {
t.Error("resolved with no header anywhere")
}
other := t.TempDir()
if goAsmHeaderResolved(dir, []string{other}) {
t.Error("resolved with an empty -I directory")
}
if err := os.WriteFile(filepath.Join(dir, "go_asm.h"), nil, 0o644); err != nil {
t.Fatal(err)
}
if !goAsmHeaderResolved(dir, nil) {
t.Error("not resolved with the header in the package directory")
}
}
func TestOtherGOOSFile(t *testing.T) {
for path, want := range map[string]bool{
"/x/sys_windows_amd64.s": true,
"/x/rt0_js_wasm.s": true,
"/x/sys_darwin_arm64.s": true,
"/x/sys_linux_amd64.s": false,
"/x/time_linux_amd64.s": false,
"/x/memmove_amd64.s": false,
"/x/generic.s": false,
} {
if got := otherGOOSFile(path); got != want {
t.Errorf("otherGOOSFile(%q) = %v, want %v", path, got, want)
}
}
}
// TestRunCorpusAuditGoAsm covers the audit wiring end to end: a package
// beside its kernel, the kernel living off the generated defines, and the
// histogram recording a generation failure as its own reason.
func TestRunCorpusAuditGoAsm(t *testing.T) {
dir := t.TempDir()
write := func(name, src string) {
t.Helper()
if err := os.WriteFile(filepath.Join(dir, name), []byte(src), 0o644); err != nil {
t.Fatal(err)
}
}
write("pkg.go", `package corpus
const pageSize = 4096
type header struct {
magic uint64
flags uint64
}
`)
write("kern_amd64.s", "#include \"go_asm.h\"\nTEXT \xc2\xb7f(SB), NOSPLIT, $0-16\n\tMOVQ\t$const_pageSize, AX\n\tMOVQ\t$header__size, BX\n\tRET\n")
// The defines live in the file's own package; a kernel in a directory
// without Go files has no package to generate from.
if err := os.MkdirAll(filepath.Join(dir, "sub"), 0o755); err != nil {
t.Fatal(err)
}
write(filepath.Join("sub", "lonely_arm64.s"), "#include \"go_asm.h\"\nTEXT \xc2\xb7g(SB), NOSPLIT, $0-0\n\tRET\n")
stats, err := runCorpusAudit(dir, nil)
if err != nil {
t.Fatalf("runCorpusAudit: %v", err)
}
get := func(name string) *corpusTally {
for i, tg := range stats.targets {
if tg.name == name {
return stats.tallies[i]
}
}
t.Fatalf("no tally for %s", name)
return nil
}
if a := get("amd64"); a.attempted != 1 || a.assembled != 1 {
t.Errorf("amd64 = %d/%d, want 1/1", a.assembled, a.attempted)
}
// lonely_arm64.s is an arm64 file whose package cannot be generated.
if a := get("arm64"); a.attempted != 1 || a.assembled != 0 {
t.Errorf("arm64 = %d/%d, want 0/1", a.assembled, a.attempted)
}
if r := get("arm64").reasons["go_asm.h generation failed"]; r != 1 {
t.Errorf("arm64 go_asm.h failure count = %d, want 1", r)
}
}
// TestGenerateGoAsmHeaderRuntime pins the generator against the real thing:
// the runtime package, whose header the toolchain's own -asmhdr output was
// sampled from. Skipped in short mode: it type-checks the whole package.
func TestGenerateGoAsmHeaderRuntime(t *testing.T) {
if testing.Short() {
t.Skip("type-checks the whole runtime package")
}
dir, err := generateGoAsmHeader("/usr/local/go/src/runtime", "amd64", t.TempDir())
if err != nil {
t.Fatalf("generateGoAsmHeader(runtime): %v", err)
}
b, err := os.ReadFile(dir + "/go_asm.h")
if err != nil {
t.Fatal(err)
}
hdr := string(b)
for _, want := range []string{
"#define const_hashSize 8\n",
"#define const_avxSupported 1\n",
"#define const_pageSize 8192\n",
"#define g_stackguard0 16\n",
"#define m__size ",
} {
if !strings.Contains(hdr, want) {
t.Errorf("runtime header misses %q", want)
}
}
}
+276 -22
View File
@@ -5,6 +5,7 @@ package main
import (
"fmt"
"maps"
"os"
"os/exec"
"path/filepath"
@@ -37,7 +38,7 @@ import (
// construction and are excluded from the diff; the other architectures list
// their conditional branches outright.
func cmdAuditInstructions(args []string) error {
fs := newCommand("audit-instructions", "gasm audit-instructions [--corpus [dir]] [amd64|arm64|riscv64|loong64]", `
fs := newCommand("audit-instructions", "gasm audit-instructions [--corpus [dir]] [-I dir] [amd64|arm64|riscv64|loong64]", `
Compare the gasm encoder for the given architecture (default amd64) against
go tool asm and print the diff: superset encodings (gasm-only, shippable via
gasm asm --format goobj) and known-but-unencodable names (the backlog). The
@@ -57,11 +58,13 @@ per-architecture pass rates and the most common failure reasons, which drive
the encodability backlog by frequency rather than by table order.
`)
corpus := fs.Bool("corpus", false, "assemble a corpus of .s files and report pass rates and failure reasons")
var dirs includeDirs
fs.Var(&dirs, "I", "directory to search for #include files (may be repeated)")
if err := fs.Parse(args); err != nil {
return err
}
if *corpus {
return cmdAuditCorpus(fs.Args())
return cmdAuditCorpus(fs.Args(), dirs)
}
archName := "amd64"
switch n := len(fs.Args()); {
@@ -266,18 +269,76 @@ func probeShapes(a arch.Arch) []string {
// and takes R register spellings.
"EQ, R0, R1, R2", "EQ, R0, R1", "EQ, R0",
"GE, F0, F1, F2", "NE, F0, F1, $0",
// Pairs, acquire/release and exclusive atomics, LSE-AL forms.
"(R0), R1", "R0, (R1)", "R1, (R2), R3", "(R2, R3), 8(R1)",
"8(R1), (R2, R3)", "R1, R2, (R3)", "(R0)",
// System operations and their register/operand names.
"$4, R1, p2", "$35943", "$1", "$1, SPSel", "SPSel, R0",
"IVAC, R0", "(R0), PLDL1KEEP", "R1, R2, R3, R4",
// SIMD element, structure and literal-pool forms.
"(R0), [V1.B16]", "[V1.B16], (R0)", "V13.S[0], R1",
"R1, V2.B[3]", "$4, V1.B16, V2.B16", "V1.B16, (R0)",
"(R0), V1.B16", "",
// The spellings GOROOT's own kernels use, from the
// differential kernels this table was proven against.
"R0, p2", "R0, R1", "F0, F1, F2, F3", "$4, V1.B16, V2.B16, V3.B16, V4.B16",
"(R0), [V0.B8, V1.B8, V2.B8, V3.B8]", "$1, $2, V1",
"R0, R1, p2", "p2, R1", "$1234, R1", "DCZID_EL0, R1",
"$0", "R1, $4, EQ", "$33, R1, $25, R2", "$4, R1, p2",
"$4, V1.B8, V2.B8, V3.B8", "$63, V1.D2, V2.D2, V3.D2",
"V1.B16, [V2.B16], V3.B16", "V1.B8, [V2.B16, V3.B16], V4.B8",
"$4, V1.B16, V2.B16, V3.B16", "$15, V1", "V1, V2, p2",
"R0, R1, $1, $4, p2",
// The landing-pad kind, the compiler's PCDATA
// bookkeeping and the four-operand bitfield
// insert/extract family, as the toolchain's own
// testdata spells them.
"C", "$1, $0", "$0, R1, $1, R2",
}
case arch.RISCV:
return []string{
"X5, X6, X7", "X5, X6", "X5", "$1, X5", "X5, (X6)", "$1, X5, X6",
"(X5), X6", "F0, F1, F2", "F0, F1", "p2", "X1, p2", "X0, p2",
"X5, X6, p2", "p2(SB)",
// AMO atomics: destination, base, source.
"R5, (R4), R6", "X5, (X4), X6",
// Segment stores take the first vector register aligned
// to the segment count, as the toolchain requires.
"(X5), X6, V0, V8", "(X5), X6, V0", "(X5), X0, V4",
// The FP multiply-add family takes four registers.
"F0, F1, F2, F3",
// The RVV slice: register, vector-register and vtype forms.
"V1, V2, V3", "V1, X5, V2", "V1", "V1, (X5)", "(X5), V1",
"$15, V1", "$15", "V1, V2", "V1, X5",
"X5, X6, p2", "R5, R6, p2",
"X5, E8, M8, TA, MA, X6", "$4, E32, M1, TA, MA, X1",
"(X5), X6, V1, V2",
// The CSR immediate forms the toolchain's testdata spells:
// immediate, CSR name, destination.
"$2, TIME, X5",
"",
}
case arch.LOONG64:
return []string{
"R4, R5, R6", "R4, R5", "R4", "$1, R4", "R4, (R5)", "(R4), R5",
"F0, F1, F2", "F0, F1", "p2", "R1, p2", "R4, p2",
"$1, R4, R5, R6", "$65536, R4", "R4, R5, p2", "p2(SB)",
// AMO atomics: destination, base, source.
"R5, (R4), R6", "X5, (X4), X6",
// Segment stores take the first vector register aligned
// to the segment count, as the toolchain requires.
"(X5), X6, V0, V8", "(X5), X6, V0", "(X5), X0, V4",
// The LSX and LASX banks share the 5-bit numbering with F.
"V1, V2, V3", "X1, X2, X3", "V1, V2", "X1, X2", "V1", "X1",
// The vector compare-to-flag forms land in an FCC register.
"V1, FCC0", "X1, FCC0",
// The compiler's bookkeeping pair and the raw spellings the
// toolchain's own testdata carries: JIRL rd, rj, offset (the
// form RET lowers to), the prefetch with a 32-bit address and
// hint, and the byte-shuffle quads.
"$1, $0", "R1, R5, 0", "0(R7), $5, $0", "(R7), $5, $0",
"V1, V2, V3, V4", "X1, X2, X3, X4",
"",
}
}
return nil
@@ -351,8 +412,11 @@ func (t *corpusTally) fail(path, reason string) {
}
}
// cmdAuditCorpus implements audit-instructions --corpus.
func cmdAuditCorpus(args []string) error {
// cmdAuditCorpus implements audit-instructions --corpus. The include
// directories carry #include resolution over a corpus whose files refer to
// headers such as GOROOT/pkg/include, the same -I a toolchain comparison
// needs.
func cmdAuditCorpus(args []string, dirs includeDirs) error {
if len(args) > 1 {
return &usageError{fmt.Errorf("audit-instructions --corpus takes at most one directory argument")}
}
@@ -366,7 +430,27 @@ func cmdAuditCorpus(args []string) error {
}
root = filepath.Join(strings.TrimSpace(string(out)), "src")
}
stats, err := runCorpusAudit(root)
// The toolchain's shipped headers (funcdata.h and friends) define the
// macros GOROOT files include; a corpus audit measures those files, so
// the header directory joins the search path automatically. go_asm.h
// is compiler-generated per package, so it is not resolved from here:
// files that include it get one generated per target architecture,
// which runCorpusAudit arranges.
if out, err := exec.Command("go", "env", "GOROOT").Output(); err == nil {
pkgInclude := filepath.Join(strings.TrimSpace(string(out)), "pkg", "include")
if fi, err := os.Stat(pkgInclude); err == nil && fi.IsDir() {
seen := false
for _, d := range dirs {
if d == pkgInclude {
seen = true
}
}
if !seen {
dirs = append(dirs, pkgInclude)
}
}
}
stats, err := runCorpusAudit(root, dirs)
if err != nil {
return err
}
@@ -376,16 +460,107 @@ func cmdAuditCorpus(args []string) error {
// corpusStats is the outcome of one corpus audit run.
type corpusStats struct {
root string
files int
generic int // files attempted for all four architectures
full int // files that assembled for every target architecture
targets []corpusTarget
tallies []*corpusTally
root string
files int
generic int // files attempted for all four architectures
otherPort int // files named for another Go port: never attempted
full int // files that assembled for every target architecture
targets []corpusTarget
tallies []*corpusTally
}
// runCorpusAudit assembles every .s file under root and returns the stats.
func runCorpusAudit(root string) (*corpusStats, error) {
// goPortSuffixes lists every architecture the Go project ports to. A file
// named for one of them belongs to that port's build, not to the generic
// set, even when gasm does not support the architecture.
var goPortSuffixes = []string{
"386", "amd64", "arm", "arm64", "loong64", "mips", "mips64",
"mips64le", "mipsle", "ppc64", "ppc64le", "riscv", "riscv64",
"s390x", "wasm",
}
// otherPortFile reports whether the file belongs to a build no supported
// target ever compiles: either its name carries a Go-architecture suffix
// gasm does not support, or, for a file with no architecture suffix at all,
// it names another GOOS, which go/build drops from the file set
// (rt0_js_wasm.s is a javascript build, not a generic one).
func otherPortFile(path string) bool {
if otherGOOSFile(path) {
return true
}
base := path
if i := strings.LastIndexByte(base, '/'); i >= 0 {
base = base[i+1:]
}
for _, sfx := range goPortSuffixes {
if strings.HasSuffix(base, "_"+sfx+".s") {
return true
}
}
return false
}
// goOSNames are the GOOS values go/build recognises in file names.
var goOSNames = map[string]bool{
"aix": true, "android": true, "darwin": true, "dragonfly": true,
"freebsd": true, "hurd": true, "illumos": true, "ios": true,
"js": true, "linux": true, "nacl": true, "netbsd": true,
"openbsd": true, "plan9": true, "solaris": true, "wasip1": true,
"windows": true, "zos": true,
}
// resolveGOOS validates a -GOOS flag value, mirroring the architecture
// check's surface: a usage error naming what the tool accepts.
func resolveGOOS(name string) (string, error) {
lower := strings.ToLower(name)
if goOSNames[lower] {
return lower, nil
}
return "", &usageError{fmt.Errorf("unknown GOOS %q: want one of %s", name, strings.Join(slices.Sorted(maps.Keys(goOSNames)), ", "))}
}
// goosFromFilename returns the GOOS the file's name carries, by go/build's
// goodOSArchFile rule: the GOOS segment sits last, or last before the
// architecture segment (sys_darwin_arm64.s, vlop_arm.s carries none). An
// empty result means the name names no GOOS and the ambient one applies.
func goosFromFilename(path string) string {
base := path
if i := strings.LastIndexByte(base, '/'); i >= 0 {
base = base[i+1:]
}
base = strings.TrimSuffix(base, ".s")
// go/build ignores everything before the first underscore, so a GOOS
// segment is only ever looked for from there on.
i := strings.IndexByte(base, '_')
if i < 0 {
return ""
}
segs := strings.Split(base[i:], "_")
if n := len(segs); n >= 2 && goOSNames[segs[n-2]] && slices.Contains(goPortSuffixes, segs[n-1]) {
return segs[n-2]
}
if goOSNames[segs[len(segs)-1]] {
return segs[len(segs)-1]
}
return ""
}
// otherGOOSFile reports whether the file's name names a GOOS other than the
// host's, by go/build's file-name rules.
func otherGOOSFile(path string) bool {
base := path
if i := strings.LastIndexByte(base, '/'); i >= 0 {
base = base[i+1:]
}
for seg := range strings.SplitSeq(strings.TrimSuffix(base, ".s"), "_") {
if goOSNames[seg] && seg != runtime.GOOS {
return true
}
}
return false
}
func runCorpusAudit(root string, dirs includeDirs) (*corpusStats, error) {
files, err := asmFiles(root)
if err != nil {
return nil, err
@@ -403,14 +578,28 @@ func runCorpusAudit(root string) (*corpusStats, error) {
}
// full is the north-star number: a file counts when every architecture
// its name allows assembles it.
full, generic := 0, 0
full, generic, otherPort := 0, 0, 0
// Header generation is created on first use, so a corpus with no
// go_asm.h includes never pays for a temp directory.
var hdr *asmhdrCache
defer func() {
if hdr != nil {
hdr.close()
}
}()
for _, path := range files {
src, err := readSource(path)
if err != nil {
return nil, err
}
f, errs := parser.Parse(path, src)
// The GOOS the header generation type-checks under follows the
// file's name when the name carries one; the ambient GOOS is the
// honest guess otherwise (a build tag naming another GOOS is
// invisible to a file-name rule).
goos := goosFromFilename(path)
var wanted []int // indexes into targets
if a := arch.FromFilename(path); a != arch.Unknown {
@@ -419,6 +608,14 @@ func runCorpusAudit(root string) (*corpusStats, error) {
wanted = append(wanted, i)
}
}
} else if otherPortFile(path) {
// A file named for a Go port gasm does not support (arm,
// 386, s390x, ...) or for another GOOS is compiled by no
// supported-arch build, so it is neither generic nor a
// per-arch attempt: counting it as generic would make the
// headline unreachably low for reasons no supported target
// can fix.
otherPort++
} else {
generic++
for i := range targets {
@@ -426,6 +623,55 @@ func runCorpusAudit(root string) (*corpusStats, error) {
}
}
// A file that includes go_asm.h parses against a per-target header:
// the defines differ per architecture (internal/cpu's layout, for
// one) and per GOOS (sys_darwin_arm64.s's trampoline constants,
// for another), so the parse cannot be shared the way a
// header-free file's can. A generation failure is a failure for
// every target, named for the package rather than a bare "include
// not found". A header already resolvable in the package
// directory or the -I list is left alone.
if len(wanted) > 0 && needsGoAsmHeader(src) && !goAsmHeaderResolved(filepath.Dir(path), dirs) {
if hdr == nil {
if hdr, err = newAsmhdrCache(); err != nil {
return nil, err
}
}
pkgDir := filepath.Dir(path)
ok := true
for _, i := range wanted {
tg, t := targets[i], tallies[i]
t.attempted++
hdrDir, err := hdr.dirFor(pkgDir, goos, goarchName(tg.a))
if err != nil {
ok = false
t.fail(path, corpusReason(err))
continue
}
f, errs := parser.ParseWithOptions(path, src, parser.Options{
Expand: true,
IncludeDirs: append(slices.Clone(dirs), hdrDir),
})
if len(errs) > 0 {
ok = false
t.fail(path, corpusReason(errs[0]))
continue
}
if _, err := assembleFile(tg.a, f); err != nil {
ok = false
t.fail(path, corpusReason(err))
continue
}
t.assembled++
}
if ok && len(wanted) > 0 {
full++
}
continue
}
f, errs := parser.ParseWithOptions(path, src, parser.Options{Expand: true, IncludeDirs: dirs})
ok := true
for _, i := range wanted {
tg, t := targets[i], tallies[i]
@@ -449,19 +695,25 @@ func runCorpusAudit(root string) (*corpusStats, error) {
}
return &corpusStats{
root: root,
files: len(files),
generic: generic,
full: full,
targets: targets,
tallies: tallies,
root: root,
files: len(files),
generic: generic,
otherPort: otherPort,
full: full,
targets: targets,
tallies: tallies,
}, nil
}
// printCorpusStats renders the corpus audit report.
func printCorpusStats(s *corpusStats) {
fmt.Printf("corpus %s: %d files (%d generic, attempted for all architectures)\n", s.root, s.files, s.generic)
fmt.Printf(" assemble for every target architecture: %d (%.1f%%)\n", s.full, 100*float64(s.full)/float64(max(s.files, 1)))
fmt.Printf("corpus %s: %d files (%d generic, attempted for all architectures; %d named for other Go ports, never attempted)\n", s.root, s.files, s.generic, s.otherPort)
// The rate is over the files a supported build would attempt: the
// other ports' files sit in the count for completeness but can never
// assemble, so counting them in the denominator would report the gap
// of architectures gasm deliberately does not target.
attemptable := max(s.files-s.otherPort, 1)
fmt.Printf(" assemble for every target architecture: %d of %d attemptable (%.1f%%)\n", s.full, attemptable, 100*float64(s.full)/float64(attemptable))
for i, tg := range s.targets {
t := s.tallies[i]
fmt.Printf(" %s: %d/%d attempted\n", tg.name, t.assembled, t.attempted)
@@ -476,6 +728,8 @@ func printCorpusStats(s *corpusStats) {
func corpusReason(err error) string {
msg := err.Error()
switch {
case strings.Contains(msg, "go_asm.h for GOARCH"):
return "go_asm.h generation failed"
case strings.Contains(msg, "unsupported"), strings.Contains(msg, "cannot encode"):
return "instruction not encodable"
case strings.Contains(msg, "undefined label"):
+62 -11
View File
@@ -240,6 +240,16 @@ func readSource(path string) (string, error) {
return string(b), err
}
// includeDirs collects repeatable -I flags: the directories searched for
// #include files during macro expansion and include splicing.
type includeDirs []string
func (d *includeDirs) String() string { return strings.Join(*d, ",") }
func (d *includeDirs) Set(v string) error {
*d = append(*d, v)
return nil
}
func cmdTokens(args []string) int {
fs := newCommand("tokens", "gasm tokens <file>", `
Print the lexical token stream of FILE: position, token kind and text, one
@@ -476,7 +486,7 @@ hover, document symbols, diagnostics and semantic-token highlighting.
}
func cmdAsm(args []string) int {
fs := newCommand("asm", "gasm asm [--format raw|elf|goobj] [-p pkg] [-GOARCH arch] [-o out] <file>", `
fs := newCommand("asm", "gasm asm [--format raw|elf|goobj] [-I dir] [-p pkg] [-GOARCH arch] [-GOOS os] [-o out] <file>", `
Assemble FILE without the Go toolchain: every TEXT function is encoded to
machine code and printed as a hex dump. Supported architectures: amd64
(including VEX/AVX2 and EVEX/AVX-512), arm64 (AArch64 integer, FP,
@@ -493,14 +503,26 @@ system toolchain; goobj emits the Go toolchain's own object format, which
cmd/link consumes directly (it requires -p, the package path, and the
installed Go toolchain: the object preamble is captured from go tool asm
and the format version from go version).
A file that includes go_asm.h gets that header generated automatically from
the package it lives in (the .go files beside it, type-checked for the
target architecture, the toolchain's own defines), so GOROOT assembly
assembles without a compiler. -GOOS selects the type-checking GOOS for
that header: a GOOS-specific file (sys_darwin_arm64.s) needs its platform's
defines, which a header from the ambient GOOS silently omits. A package
that has no Go files for the target or does not type-check is a hard error
naming the package.
`)
out := fs.String("o", "", "write the output to this file")
format := fs.String("format", "raw", "output format: raw (concatenated image), elf or goobj (Go object)")
pkg := fs.String("p", "", "package path for --format goobj (qualifies the exported symbols)")
archName := fs.String("GOARCH", "", "target architecture: amd64, arm64, riscv64 or loong64 (overrides the file-name suffix)")
goosName := fs.String("GOOS", "", "operating system for go_asm.h generation: a GOOS go/build recognises (default: the host's)")
var dirs includeDirs
fs.Var(&dirs, "I", "directory to search for #include files (may be repeated)")
fs.Parse(args)
if fs.NArg() != 1 {
fmt.Fprintln(os.Stderr, "usage: gasm asm [--format raw|elf|goobj] [-p pkg] [-GOARCH arch] [-o out] <file>")
fmt.Fprintln(os.Stderr, "usage: gasm asm [--format raw|elf|goobj] [-I dir] [-p pkg] [-GOARCH arch] [-GOOS os] [-o out] <file>")
return 2
}
// The format is validated before anything else, so a bogus value exits 2
@@ -521,12 +543,38 @@ and the format version from go version).
}
targetArch = a
}
goos := ""
if *goosName != "" {
g, err := resolveGOOS(*goosName)
if err != nil {
fmt.Fprintf(os.Stderr, "gasm asm: %v\n", err)
return 2
}
goos = g
}
src, err := readSource(path)
if err != nil {
fmt.Fprintln(os.Stderr, "gasm:", err)
return 1
}
f, errs := parser.Parse(path, src)
// A file that includes go_asm.h cannot assemble without the package's
// defines, and without a compiler nothing else has generated them, so
// gasm produces the equivalent itself: automatic, because the compiler
// behaves the same way and a flag would only ever be forgotten. A
// generation failure is fatal and names the package: assembling against
// a missing header would fail later with a bare "undefined" instead.
// A go_asm.h that already resolves (placed by hand, or passed with -I)
// is left alone.
if needsGoAsmHeader(src) && !goAsmHeaderResolved(filepath.Dir(path), dirs) {
hdrDir, cleanup, err := ensureGoAsmHeader(path, targetArch, goos, nil)
if err != nil {
fmt.Fprintln(os.Stderr, "gasm asm:", err)
return 1
}
defer cleanup()
dirs = append(dirs, hdrDir)
}
f, errs := parser.ParseWithOptions(path, src, parser.Options{Expand: true, IncludeDirs: dirs})
for _, e := range errs {
fmt.Fprintf(os.Stderr, "%s: %v\n", path, e)
}
@@ -634,7 +682,7 @@ and the format version from go version).
// cmdDiff compares the machine code of two assembly files.
func cmdDiff(args []string) int {
set := newCommand("diff", "gasm diff [-GOARCH arch] <file1.s> <file2.s>", `
set := newCommand("diff", "gasm diff [-GOARCH arch] [-I dir] <file1.s> <file2.s>", `
Compare the machine code produced by assembling two files.
Shows which functions differ and the byte-level differences.
Useful for verifying that two implementations produce identical code,
@@ -645,9 +693,11 @@ e.g. --map wideCopyAVX2=wideCopyAVX512 pairs the two regardless of suffix.
`)
mapSpec := set.String("map", "", "comma-separated old=new pairs to match functions with different names")
archName := set.String("GOARCH", "", "target architecture for both files: amd64, arm64, riscv64 or loong64")
var dirs includeDirs
set.Var(&dirs, "I", "directory to search for #include files (may be repeated)")
set.Parse(args)
if set.NArg() != 2 {
fmt.Fprintln(os.Stderr, "usage: gasm diff [-GOARCH arch] <file1.s> <file2.s>")
fmt.Fprintln(os.Stderr, "usage: gasm diff [-GOARCH arch] [-I dir] <file1.s> <file2.s>")
return 2
}
path1, path2 := set.Arg(0), set.Arg(1)
@@ -675,12 +725,12 @@ e.g. --map wideCopyAVX2=wideCopyAVX512 pairs the two regardless of suffix.
}
// Assemble both files.
img1, err := assemblePath(path1, forced)
img1, err := assemblePath(path1, forced, dirs)
if err != nil {
fmt.Fprintf(os.Stderr, "gasm diff: %s: %v\n", path1, err)
return 1
}
img2, err := assemblePath(path2, forced)
img2, err := assemblePath(path2, forced, dirs)
if err != nil {
fmt.Fprintf(os.Stderr, "gasm diff: %s: %v\n", path2, err)
return 1
@@ -755,14 +805,15 @@ func assembleFile(targetArch arch.Arch, f *ast.File) (*asm.Image, error) {
}
}
// assemblePath reads, parses and assembles a file (used by cmdDiff). A
// non-Unknown forced architecture overrides the file-name suffix.
func assemblePath(path string, forced arch.Arch) (*asm.Image, error) {
// assemblePath reads, preprocesses, parses and assembles a file (used by
// cmdDiff). A non-Unknown forced architecture overrides the file-name
// suffix.
func assemblePath(path string, forced arch.Arch, dirs includeDirs) (*asm.Image, error) {
src, err := readSource(path)
if err != nil {
return nil, err
}
f, errs := parser.Parse(path, src)
f, errs := parser.ParseWithOptions(path, src, parser.Options{Expand: true, IncludeDirs: dirs})
for _, e := range errs {
fmt.Fprintf(os.Stderr, "%s: %v\n", path, e)
}
+1 -1
View File
@@ -403,7 +403,7 @@ func TestRunCorpusAudit(t *testing.T) {
write("generic.s", "#include \"textflag.h\"\nTEXT ·g(SB), NOSPLIT, $0-0\n\tRET\n")
write("broken.s", "#include \"textflag.h\"\nTEXT ·b(SB), NOSPLIT, $0-0\n\tJMP nowhere\n\tRET\n")
stats, err := runCorpusAudit(dir)
stats, err := runCorpusAudit(dir, nil)
if err != nil {
t.Fatalf("runCorpusAudit: %v", err)
}
+109
View File
@@ -0,0 +1,109 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
package main
import (
"os"
"path/filepath"
"strings"
"testing"
)
// writeTree writes a directory of files and returns its root.
func writeTree(t *testing.T, files map[string]string) string {
t.Helper()
dir := t.TempDir()
for name, content := range files {
path := filepath.Join(dir, name)
if err := os.MkdirAll(filepath.Dir(path), 0o755); err != nil {
t.Fatal(err)
}
if err := os.WriteFile(path, []byte(content), 0o644); err != nil {
t.Fatal(err)
}
}
return dir
}
// TestAsmMacroAndIncludeEndToEnd drives `gasm asm` over a source with an
// in-file parameterised macro and an include resolved through -I, and checks
// the assembled bytes came from the expansion (the loop body counts six
// increments, two per expanded iteration).
func TestAsmMacroAndIncludeEndToEnd(t *testing.T) {
if testing.Short() {
t.Skip("runs the assembler end to end")
}
dir := writeTree(t, map[string]string{
"inc/consts.h": "#define NITER 3\n",
"main_amd64.s": "#include \"textflag.h\"\n" +
"#include \"consts.h\"\n" +
"#define STEP(r) ADDQ $1, r; ADDQ $1, r\n" +
"TEXT ·f(SB), NOSPLIT, $0-8\n" +
"\tXORQ AX, AX\n" +
"\tMOVQ $NITER, CX\n" +
"loop:\n" +
"\tSTEP(AX)\n" +
"\tDECQ CX\n" +
"\tJNZ loop\n" +
"\tMOVQ AX, ret+0(FP)\n" +
"\tRET\n",
})
stdout, stderr, code := capture(func() int {
return cmdAsm([]string{"-I", filepath.Join(dir, "inc"), "-GOARCH", "amd64", filepath.Join(dir, "main_amd64.s")})
})
if code != 0 {
t.Fatalf("gasm asm exited %d: %s%s", code, stdout, stderr)
}
// The macro expanded to two ADDQ $1 encodings in the static body; the
// iteration count lives in the runtime loop.
if n := strings.Count(stdout, "83 c0 01"); n != 2 {
t.Errorf("found %d ADDQ $1 encodings in the image, want 2:\n%s", n, stdout)
}
}
// TestAsmIncludeResolutionOrder pins the -I search order end to end: the
// including file's directory wins over the -I directories.
func TestAsmIncludeResolutionOrder(t *testing.T) {
if testing.Short() {
t.Skip("runs the assembler end to end")
}
dir := writeTree(t, map[string]string{
"src/main_amd64.s": "#include \"textflag.h\"\n" +
"#include \"vals.h\"\n" +
"TEXT ·f(SB), NOSPLIT, $0\n" +
"\tMOVQ $VAL, AX\n" +
"\tRET\n",
"src/vals.h": "#define VAL 1\n",
"late/vals.h": "#define VAL 2\n",
"early/vals.h": "#define VAL 3\n",
})
stdout, stderr, code := capture(func() int {
return cmdAsm([]string{"-I", filepath.Join(dir, "early"), "-I", filepath.Join(dir, "late"),
"-GOARCH", "amd64", filepath.Join(dir, "src", "main_amd64.s")})
})
if code != 0 {
t.Fatalf("gasm asm exited %d: %s%s", code, stdout, stderr)
}
// VAL came from src/vals.h, not from either -I directory: the image
// loads the immediate 1.
if !strings.Contains(stdout, "b8 01 00 00 00") {
t.Errorf("expected the source-directory VAL (immediate 1) in:\n%s", stdout)
}
}
// TestAsmMissingIncludeIsAnError pins the diagnostic for an include that
// resolves nowhere on the assembly path.
func TestAsmMissingIncludeIsAnError(t *testing.T) {
if testing.Short() {
t.Skip("runs the assembler end to end")
}
path := writeTemp(t, "main_amd64.s", "#include \"textflag.h\"\n#include \"nothere.h\"\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n")
_, stderr, code := capture(func() int { return cmdAsm([]string{"-GOARCH", "amd64", path}) })
if code == 0 {
t.Fatal("gasm asm accepted a file whose include resolves nowhere")
}
if !strings.Contains(stderr, `#include "nothere.h"`) {
t.Errorf("stderr does not name the failing include: %s", stderr)
}
}
+23 -4
View File
@@ -141,14 +141,16 @@ gasm lint kernel_amd64.s
## asm
```text
Usage: gasm asm [--format raw|elf|goobj] [-p pkg] [-GOARCH arch] [-o out] <file>
Usage: gasm asm [--format raw|elf|goobj] [-I dir] [-p pkg] [-GOARCH arch] [-GOOS os] [-o out] <file>
```
| Flag | Default | Effect |
|---|---|---|
| `-format` | `raw` | output format: `raw` (concatenated image), `elf` or `goobj` (Go object) |
| `-I` | empty | directory to search for `#include` files; may be repeated, searched in order after the source directory |
| `-p` | empty | package path for `--format goobj`, qualifying the exported symbols |
| `-GOARCH` | empty | target architecture: `amd64`, `arm64`, `riscv64` or `loong64`; overrides the file-name suffix |
| `-GOOS` | empty | operating system for the generated `go_asm.h`: any GOOS `go/build` recognises in file names; default is the host's |
| `-o` | empty | write the output to this file instead of a hex dump on stdout |
Supported architectures: amd64 (VEX/AVX2 and EVEX/AVX-512 included), arm64,
@@ -162,6 +164,20 @@ system toolchain; `goobj` emits the Go toolchain's own object format, which
installed: the object preamble is captured from `go tool asm` and the format
version from `go version`. `raw` and `elf` need no toolchain at all.
A file that includes `go_asm.h` gets that header generated from the Go
files beside it, type-checked for the target. `-GOOS` selects the
type-checking GOOS for that header, because a GOOS-specific file needs its
platform's defines: `sys_darwin_arm64.s` fails against the ambient GOOS
(`machTimebaseInfo_numer` is missing from a linux type-check) and assembles
with `-GOOS darwin`.
Assembly preprocessing matches the toolchain's: `#define` macros (object and
parameterised) expand at the point of use, `#undef`, `#ifdef`, `#ifndef`,
`#else` and `#endif` behave as in `go tool asm`, `;` separates statements,
and `#include "file"` splices the named file in, resolved against the source
directory and then each `-I` directory in order. `textflag.h` is the one
header that is not spliced: gasm consumes its flag names natively.
```sh
gasm asm hello_amd64.s
```
@@ -305,12 +321,13 @@ gasm debug --func add --cover hello_amd64.s
## diff
```text
Usage: gasm diff [-GOARCH arch] <file1.s> <file2.s>
Usage: gasm diff [-GOARCH arch] [-I dir] <file1.s> <file2.s>
```
| Flag | Default | Effect |
|---|---|---|
| `-GOARCH` | empty | target architecture for both files, overriding the file-name suffixes |
| `-I` | empty | directory to search for `#include` files; may be repeated, searched in order after the source directory |
| `-map` | empty | comma-separated `old=new` pairs to match functions with different names |
Functions are paired by exact name unless `--map` says otherwise, so
@@ -348,7 +365,7 @@ add: 16 bytes, args=24, frame=0 NOSPLIT
## audit-instructions
```text
Usage: gasm audit-instructions [--corpus [dir]] [amd64|arm64|riscv64|loong64]
Usage: gasm audit-instructions [--corpus [dir]] [-I dir] [amd64|arm64|riscv64|loong64]
```
Compare the gasm encoder for the given architecture (default amd64) against the
@@ -381,7 +398,9 @@ With `--corpus` the audit changes shape: it assembles every `.s` file under
DIR (default `GOROOT/src`) with the gasm encoder only, no toolchain probing.
A file whose name carries a recognisable `_arch` suffix is attempted for that
architecture; a file without one is attempted for all four, exactly as a
`GOARCH` build would compile it. The report gives the headline number (files
`GOARCH` build would compile it, and a name that names a GOOS
(`sys_darwin_arm64.s`) type-checks its generated `go_asm.h` for that GOOS.
The report gives the headline number (files
that assemble for every target architecture), the per-architecture pass rates
and the most common failure reasons with one representative file each, which
drive the encodability backlog by frequency rather than by table order. A run
+18 -1
View File
@@ -2,7 +2,7 @@
.SH NAME
gasm-asm \- assemble Plan 9 assembly without the Go toolchain
.SH SYNOPSIS
.B gasm asm [\-\-format raw|elf|goobj] [\-p pkg] [\-GOARCH arch] [\-o out] <file>
.B gasm asm [\-\-format raw|elf|goobj] [\-I dir] [\-p pkg] [\-GOARCH arch] [\-GOOS os] [\-o out] <file>
.SH DESCRIPTION
Assemble FILE without the Go toolchain: every TEXT function is encoded
to machine code and printed as a hex dump. Supported architectures:
@@ -42,11 +42,24 @@ need no toolchain at all.
Framed functions receive the stack-split guard and the trailing
morestack block, byte-identical to the toolchain's output, so split
functions link too.
.PP
A file that includes go_asm.h gets that header generated from the Go
files beside it, type-checked for the target.
.B \-GOOS
selects the type-checking GOOS for that header, because a GOOS-specific
file needs its platform's defines: sys_darwin_arm64.s fails against the
ambient GOOS (machTimebaseInfo_numer is missing from a linux type-check)
and assembles with
.BR "\-GOOS darwin" .
.SH OPTIONS
.TP
.B \-\-format \fIraw|elf|goobj\fR
Output format; the default is raw.
.TP
.B \-I \fIdir\fR
Directory to search for #include files; may be repeated, searched in
order after the source directory.
.TP
.B \-p \fIpkg\fR
Package path for --format goobj, qualifying the exported symbols.
.TP
@@ -55,6 +68,10 @@ Target architecture: amd64, arm64, riscv64 or loong64; overrides the
file-name suffix, which is how the suffix-less majority of GOROOT's
files (cpu_x86.s, stub.s, ...) become assemblable.
.TP
.B \-GOOS \fIos\fR
Operating system for the generated go_asm.h: any GOOS go/build
recognises in file names; the default is the host's.
.TP
.B \-o \fIfile\fR
Write the output to this file instead of a hex dump on stdout.
.SH EXIT STATUS
+7 -1
View File
@@ -2,7 +2,7 @@
.SH NAME
gasm-audit-instructions \- diff the encoder against the Go toolchain, or measure a corpus
.SH SYNOPSIS
.B gasm audit\-instructions [\-\-corpus [\fIdir\fR]] [amd64|arm64|riscv64|loong64]
.B gasm audit\-instructions [\-\-corpus [\fIdir\fR]] [\-I dir] [amd64|arm64|riscv64|loong64]
.SH DESCRIPTION
Compare the gasm encoder for the given architecture (default amd64)
against
@@ -38,6 +38,12 @@ second.
.B \-\-corpus [\fIdir\fR]
Assemble a corpus of .s files and report pass rates and failure
reasons.
.TP
.B \-I \fIdir\fR
Directory to search for #include files; may be repeated, searched in
order after the source directory. A corpus run whose files include
toolchain headers (such as GOROOT/pkg/include) needs it, the same -I a
toolchain comparison takes.
.SH EXIT STATUS
The mnemonic-diff mode reports through its output and exits 0; a failed
probe or an unknown architecture exits non-zero.
+5 -1
View File
@@ -2,7 +2,7 @@
.SH NAME
gasm-diff \- compare the machine code of two assembly files
.SH SYNOPSIS
.B gasm diff [\-GOARCH arch] <file1.s> <file2.s>
.B gasm diff [\-GOARCH arch] [\-I dir] <file1.s> <file2.s>
.SH DESCRIPTION
Compare the machine code produced by assembling two files. Shows which
functions differ and the byte-level differences. Useful for verifying
@@ -20,6 +20,10 @@ pairs two variants regardless of suffix.
Target architecture for both files: amd64, arm64, riscv64 or loong64;
overrides the file-name suffixes.
.TP
.B \-I \fIdir\fR
Directory to search for #include files; may be repeated, searched in
order after the source directory.
.TP
.B \-\-map \fIspec\fR
Comma-separated old=new pairs to match functions with different names.
.SH EXIT STATUS
+72 -8
View File
@@ -241,6 +241,13 @@ func renderInstr(line []token.Token, width int) string {
if line[0].Kind != token.Ident {
return "\t" + mnem + " " + ops
}
// A statement separator belongs to the statement it ends: when the
// operands open with a ';', the alignment padding would land between
// the mnemonic and its own separator (REP ; MOVSQ), so such a line
// renders with a single space whatever the function's width.
if strings.HasPrefix(ops, ";") {
return "\t" + mnem + " " + ops
}
if width < len(mnem) {
width = len(mnem)
}
@@ -254,11 +261,12 @@ func renderPreproc(line []token.Token) string {
line[2].Kind == token.String {
return "#include " + line[2].Text
}
parts := make([]string, 0, len(line)-1)
for _, t := range line[1:] {
parts = append(parts, t.Text)
}
return "#" + strings.Join(parts, " ")
// The body of a directive, a macro definition included, is an ordinary
// token run: rendering it through renderOps applies the same punctuation
// rules as everywhere else, so a macro body keeps its canonical spelling
// ($v, (a, b), the ';' separators between statements) instead of being
// spread with a space between every token.
return "#" + renderOps(line[1:])
}
// renderOps re-spaces a run of operand tokens into canonical form. It never
@@ -300,6 +308,26 @@ func wouldMerge(prev, cur token.Token) bool {
return len(kinds) != 2 || kinds[0] != prev.Kind || kinds[1] != cur.Kind
}
// isOperandBracket reports whether t is one of the square-bracket tokens the
// lexer emits, as Illegal tokens carrying their spelling, for the arm64 and
// loong64 register lists and element selectors that valid GAsm source
// contains.
func isOperandBracket(t token.Token) bool {
return t.Kind == token.Illegal && (t.Text == "[" || t.Text == "]")
}
// isOpenBracket reports whether t is the '[' of a register list or element
// selector.
func isOpenBracket(t token.Token) bool {
return t.Kind == token.Illegal && t.Text == "["
}
// isCloseBracket reports whether t is the ']' that closes a register list or
// element selector.
func isCloseBracket(t token.Token) bool {
return t.Kind == token.Illegal && t.Text == "]"
}
// spaceBetween decides whether a single space separates prev and cur.
func spaceBetween(prev, cur token.Token) bool {
// '/' beside '/' or '*' would form a comment opener in the output and
@@ -308,10 +336,33 @@ func spaceBetween(prev, cur token.Token) bool {
return true
}
switch cur.Kind {
case token.Illegal:
// A closing bracket always glues to the text it closes. An opening
// bracket glues to the operand it extends (V31.B[15]) but takes its
// own space after a comma, a mnemonic or an operator, exactly like
// the parenthesis rule below. Any other Illegal spelling is stray.
if isCloseBracket(cur) {
return false
}
if isOpenBracket(cur) {
switch prev.Kind {
case token.Ident, token.Number, token.RParen, token.RAngle:
return false
}
return true
}
return true
case token.RParen:
return false
case token.Comma:
return false
case token.Semicolon:
// A ';' is a statement separator on the assembly path, not an
// operand: dropping it would fuse two statements into a line the
// assembler rejects, so it must survive as punctuation. It glues
// to the statement it ends and the next statement takes one space,
// matching the toolchain's listing style.
return false
case token.Star, token.Plus, token.Minus, token.Slash, token.Pipe:
return false
case token.LShift, token.RShift, token.Arrow, token.At:
@@ -328,6 +379,10 @@ func spaceBetween(prev, cur token.Token) bool {
}
}
switch prev.Kind {
case token.Illegal:
// '[' opens a bracket group and glues to what follows; ']' closes
// one, and what comes next takes its own space.
return !isOpenBracket(prev)
case token.LParen, token.Star, token.Plus, token.Minus, token.Slash, token.Pipe:
return false
case token.Dollar:
@@ -336,7 +391,9 @@ func spaceBetween(prev, cur token.Token) bool {
return false
case token.LAngle, token.RAngle:
return false
case token.Comma:
case token.Comma, token.Semicolon:
// The statement after a ';' separator takes its own space, exactly
// like the operand after a comma.
return true
}
return true
@@ -358,8 +415,15 @@ func splitLines(toks []token.Token) [][]token.Token {
// Illegal tokens carry no canonical spelling: the parser
// reports them as errors where they matter, and the formatter
// drops them so that a stray character cannot survive into the
// output and make the next pass render a different file.
continue
// output and make the next pass render a different file. The
// square brackets of the arm64 and loong64 vector syntaxes are
// the one exception: the lexer gives them no dedicated kind,
// but a register list [V0.B16, V1.B16] and an element selector
// V0.B[3] are valid, load-bearing source, so their tokens stay
// in the stream and renderOps glues them back where they were.
if !isOperandBracket(t) {
continue
}
}
if t.Kind == token.Newline {
lines = append(lines, cur)
+307
View File
@@ -5,9 +5,11 @@ package format
import (
"os"
"slices"
"strings"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
"sourcedock.dev/petrbalvin/gasm-devkit/lexer"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-devkit/token"
@@ -153,6 +155,121 @@ func TestOperandSpacing(t *testing.T) {
}
}
// TestVectorBracketSpacing pins the square-bracket operand forms of the
// arm64 and loong64 vector syntaxes. The lexer emits '[' and ']' as Illegal
// tokens carrying their spelling, and renderOps must glue them back exactly
// where they were: a register list and an element selector are load-bearing
// operands the assembler reads out of the operand text, so no bracket may be
// dropped, and the canonical spelling inside the brackets is tight.
func TestVectorBracketSpacing(t *testing.T) {
cases := map[string]string{
// Register lists of one to four registers.
"[V21.B16]": "[V21.B16]",
"[V17.B16, V18.B16]": "[V17.B16, V18.B16]",
"[V18.D1, V19.D1, V20.D1]": "[V18.D1, V19.D1, V20.D1]",
"[V14.B16, V15.B16, V16.B16, V17.B16]": "[V14.B16, V15.B16, V16.B16, V17.B16]",
// Element selectors.
"V31.B[15]": "V31.B[15]",
"V19.S[0]": "V19.S[0]",
"V1.D[1]": "V1.D[1]",
"V11.B[11], V16.B[12]": "V11.B[11], V16.B[12]",
// Lists beside address operands, on either side.
"32(R1), [V2.B16, V3.B16]": "32(R1), [V2.B16, V3.B16]",
"[V2.S4, V3.S4], (R14)": "[V2.S4, V3.S4], (R14)",
"(R24), [V18.D1, V19.D1]": "(R24), [V18.D1, V19.D1]",
// A spaced spelling canonicalises to the tight one.
"[ V21.B16 ]": "[V21.B16]",
"V31.B [15]": "V31.B[15]",
}
for in, want := range cases {
toks := lexOperands(in)
if got := renderOps(toks); got != want {
t.Errorf("renderOps(%q) = %q, want %q", in, got, want)
}
}
}
// TestSIMDBracketRoundTrip formats whole functions carrying the bracket
// shapes of the arm64 vector kernels and pins the output byte for byte. The
// brackets are load-bearing: formatting must not change what the file
// assembles to, so the formatted text keeps every bracket, re-formats to
// itself and still parses cleanly.
func TestSIMDBracketRoundTrip(t *testing.T) {
in := "#include \"textflag.h\"\n" +
"\n" +
"TEXT ·f(SB), NOSPLIT, $0\n" +
"VDUP V31.B[15], R3\n" +
"VTBL V22.B16, [V28.B16], V11.B16\n" +
"VLD1 (R2), [V21.B16]\n" +
"VMOVQ $0x70, $0x80, V10\n" +
"RET\n"
want := "#include \"textflag.h\"\n" +
"\n" +
"TEXT ·f(SB), NOSPLIT, $0\n" +
"\tVDUP V31.B[15], R3\n" +
"\tVTBL V22.B16, [V28.B16], V11.B16\n" +
"\tVLD1 (R2), [V21.B16]\n" +
"\tVMOVQ $0x70, $0x80, V10\n" +
"\tRET\n"
got := Source(in)
if got != want {
t.Fatalf("formatting mismatch:\n--- got ---\n%q\n--- want ---\n%q", got, want)
}
if again := Source(got); again != got {
t.Fatalf("not idempotent:\n%q", again)
}
if n := strings.Count(got, "["); n != 3 {
t.Errorf("output carries %d '[', want 3:\n%s", n, got)
}
if _, errs := parser.Parse("in.s", got); len(errs) > 0 {
t.Errorf("formatted output no longer parses: %v", errs)
}
}
// TestBracketFormsRoundTrip runs every bracket shape of the vector kernels
// through a full format pass as its own single-instruction function, where
// the canonical form is the line itself indented: formatting must be a no-op
// on each, so no bracket moves, vanishes or gains a space.
func TestBracketFormsRoundTrip(t *testing.T) {
for _, instr := range []string{
"VDUP V31.B[15], V18",
"VDUP V19.S[3], V18.S4",
"VDUP V1.D[1], V2.D2",
"VMOV V13.S[0], R20",
"VMOV V11.B[11], V16.B[12]",
"VMOV R20, V21.B[2]",
"VTBL V22.B16, [V28.B16], V11.B16",
"VTBL V18.B8, [V17.B16, V18.B16], V22.B8",
"VTBL V31.B8, [V14.B16, V15.B16, V16.B16, V17.B16], V15.B8",
"VLD1 (R2), [V21.B16]",
"VLD1 (R24), [V18.D1, V19.D1, V20.D1]",
"VLD1 (R29), [V14.D1, V15.D1, V16.D1, V17.D1]",
"VLD1.P 32(R1), [V2.B16, V3.B16]",
"VLD1R (R1), [V9.B8]",
"VLD4R (R0), [V0.B8, V1.B8, V2.B8, V3.B8]",
"VST1 [V2.S4, V3.S4, V4.S4, V5.S4], (R14)",
"VST1.P [V2.B16], (R1)",
"VST1.P [V2.B16, V3.B16], 32(R1)",
"VMOVQ $0x70, $0x80, V10",
} {
src := "TEXT ·f(SB), NOSPLIT, $0\n" + instr + "\nRET\n"
want := "TEXT ·f(SB), NOSPLIT, $0\n\t" + instr + "\n\tRET\n"
got := Source(src)
if got != want {
t.Errorf("formatting %q:\n got %q\n want %q", instr, got, want)
continue
}
if again := Source(got); again != got {
t.Errorf("not idempotent for %q:\n%q", instr, again)
}
if _, errs := parser.Parse("in.s", got); len(errs) > 0 {
t.Errorf("formatted output of %q no longer parses: %v", instr, errs)
}
}
}
// TestFlagListRoundTrip pins the '|' flag separator and the <ABIInternal>
// marker through a full format pass: the bars the Go toolchain requires and
// the ABI bracket must survive byte for byte, on TEXT and GLOBL alike.
@@ -183,6 +300,196 @@ func TestCRLFInputIsNormalisedToLF(t *testing.T) {
}
}
// TestSemicolonSeparators pins the treatment of ';' statement separators.
// The separator is load-bearing on the assembly path, where the parser reads
// semicolon-separated statements: a formatter that drops it fuses two
// statements into a line the assembler rejects, which is data corruption.
// Each row pins the canonical spelling, one space after the ';', tight
// before it, the way the toolchain's own sources and listings write it.
func TestSemicolonSeparators(t *testing.T) {
cases := []struct {
name string
in string
want string
}{
{
name: "between instructions, tight",
in: "TEXT ·f(SB), $0\nBYTE $0x48;BYTE $0xc7\nRET\n",
want: "TEXT ·f(SB), $0\n\tBYTE $0x48; BYTE $0xc7\n\tRET\n",
},
{
name: "between instructions, spaced",
in: "TEXT ·f(SB), $0\nBYTE $0x48 ; BYTE $0xc7\nRET\n",
want: "TEXT ·f(SB), $0\n\tBYTE $0x48; BYTE $0xc7\n\tRET\n",
},
{
name: "after a label",
in: "TEXT ·f(SB), $0\nlabel: BYTE $1; BYTE $2\nRET\n",
want: "TEXT ·f(SB), $0\nlabel:\n\tBYTE $1; BYTE $2\n\tRET\n",
},
{
// The continuation-spliced macro shape of the runtime sources:
// the lexer makes one logical line of the backslash continuations.
name: "inside a macro body, continued",
in: "#define MOVLTOREG(v, off) \\\n\tMOVL $v, AX; \\\n\tMOVL AX, ret+off(FP)\n",
want: "#define MOVLTOREG(v, off) MOVL $v, AX; MOVL AX, ret+off(FP)\n",
},
{
name: "inside a macro body, one line",
in: "#define PEAS BYTE $0x0a; BYTE $0x0b\n",
want: "#define PEAS BYTE $0x0a; BYTE $0x0b\n",
},
{
name: "several separators in one line",
in: "TEXT ·f(SB), $0\nBYTE $1; BYTE $2; BYTE $3\nRET\n",
want: "TEXT ·f(SB), $0\n\tBYTE $1; BYTE $2; BYTE $3\n\tRET\n",
},
{
name: "two separators back to back",
in: "TEXT ·f(SB), $0\nBYTE $1;; BYTE $2\nRET\n",
want: "TEXT ·f(SB), $0\n\tBYTE $1;; BYTE $2\n\tRET\n",
},
{
name: "inside a line comment, untouched",
in: "TEXT ·f(SB), $0\n// keep; the; separators\nBYTE $1\nRET\n",
want: "TEXT ·f(SB), $0\n\t// keep; the; separators\n\tBYTE $1\n\tRET\n",
},
{
name: "after a statement, before a comment",
in: "TEXT ·f(SB), $0\nMOVQ AX, BX; // tail\nRET\n",
want: "TEXT ·f(SB), $0\n\tMOVQ AX, BX; // tail\n\tRET\n",
},
{
name: "last character on a line",
in: "TEXT ·f(SB), $0\nBYTE $1;\nRET\n",
want: "TEXT ·f(SB), $0\n\tBYTE $1;\n\tRET\n",
},
{
// The REP shape: a prefix-style zero-operand statement
// followed by the instruction it prefixes. The separator
// belongs to the statement it ends, so the function's
// alignment width (MOVSQ is the widest mnemonic here) must
// not open a gap before it: one space after the mnemonic
// whatever the neighbours' lengths.
name: "after a prefix-style statement",
in: "TEXT ·f(SB), $0\nMOVQ AX, BX\nREP; MOVSQ\nRET\n",
want: "TEXT ·f(SB), $0\n\tMOVQ AX, BX\n\tREP ; MOVSQ\n\tRET\n",
},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
got := Source(tc.in)
if got != tc.want {
t.Fatalf("formatting mismatch:\n--- got ---\n%q\n--- want ---\n%q", got, tc.want)
}
if again := Source(got); again != got {
t.Fatalf("not idempotent:\n%q", again)
}
if in, out := strings.Count(tc.in, ";"), strings.Count(got, ";"); in != out {
t.Fatalf("semicolon count changed: %d -> %d\n%s", in, out, got)
}
if _, errs := parser.Parse("in.s", got); len(errs) > 0 {
t.Fatalf("formatted output no longer parses: %v", errs)
}
})
}
}
// TestSemicolonStatementRoundTrip proves the formatter's contract on the
// path where ';' separates statements: parse the source the way the
// assembler does, format it, re-parse the formatted text and compare the
// statement sequence. Raw operand texts are token-joined, so they are
// insensitive to the whitespace a format pass chooses, and the comparison
// can only fail when a token is lost: dropping a ';' fuses two statements
// into one, exactly the corruption the released formatter committed.
func TestSemicolonStatementRoundTrip(t *testing.T) {
src := "#define MOVLTOREG(v, off) \\\n" +
"\tMOVL $v, AX; \\\n" +
"\tMOVL AX, ret+off(FP)\n" +
"\n" +
"TEXT ·f(SB), NOSPLIT, $0\n" +
"BYTE $0x48; BYTE $0xc7\n" +
"first: BYTE $1; BYTE $2\n" +
"MOVLTOREG($42, 0)\n" +
"RET\n"
before, errs := parser.ParseWithOptions("in.s", src, parser.Options{Expand: true})
if len(errs) > 0 {
t.Fatalf("source does not parse: %v", errs)
}
formatted := Source(src)
after, errs := parser.ParseWithOptions("in.s", formatted, parser.Options{Expand: true})
if len(errs) > 0 {
t.Fatalf("formatted source does not parse: %v", errs)
}
want, got := stmtSignature(before), stmtSignature(after)
if !slices.Equal(got, want) {
t.Fatalf("statement sequence changed:\n--- before ---\n%q\n--- after ---\n%q", want, got)
}
if again := Source(formatted); again != formatted {
t.Fatalf("not idempotent:\n%q", again)
}
// The two BYTE statements on the first line must stay two: one fused
// statement here is the exact defect this package once shipped.
var bytes []string
for _, stmt := range stmtSignature(after) {
if rest, ok := strings.CutPrefix(stmt, "instr BYTE "); ok {
bytes = append(bytes, rest)
}
}
if want := []string{"$ 0x48", "$ 0xc7", "$ 1", "$ 2"}; !slices.Equal(bytes, want) {
t.Fatalf("BYTE statements after expansion = %q, want %q", bytes, want)
}
}
// stmtSignature flattens a parsed file into one string per declaration and
// statement, in source order. Every component is token-derived, so the
// signature is stable across format passes and moves only when a token is
// lost or gained.
func stmtSignature(f *ast.File) []string {
var out []string
for _, d := range f.Decls {
switch d := d.(type) {
case *ast.Text:
out = append(out, "text "+d.Name.Raw)
for _, s := range d.Body {
out = append(out, stmtText(s))
}
case *ast.Globl:
out = append(out, "globl "+d.Name.Raw)
case *ast.Data:
out = append(out, "data "+d.Name.Raw)
case *ast.Include:
out = append(out, "include "+d.Header.Text)
case *ast.Preproc:
out = append(out, "preproc "+d.Raw)
}
}
for _, s := range f.Orphans {
out = append(out, stmtText(s))
}
return out
}
// stmtText renders one statement for stmtSignature.
func stmtText(s ast.Stmt) string {
switch s := s.(type) {
case *ast.Label:
return "label " + s.Name.Text
case *ast.Instr:
parts := make([]string, 0, len(s.Operands)+1)
parts = append(parts, s.Mnemonic.Text)
for _, op := range s.Operands {
parts = append(parts, op.Raw)
}
return "instr " + strings.Join(parts, " ")
default:
return "stmt"
}
}
// lexOperands lexes a single operand string and drops the EOF token.
func lexOperands(s string) []token.Token {
toks := lexer.Tokenize(s)
+6
View File
@@ -31,6 +31,12 @@ func FuzzFormatIdempotency(f *testing.F) {
f.Add("TEXT ·f(SB), NOSPLIT, $0\n\tMOVQ AX, BX\n\tRET\n")
f.Add("TEXT ·f(SB),NOSPLIT,$0\n\tMOVQ AX,BX\n\n\n\tRET\n")
f.Add("garbage ### ???\n")
// Line-ending whitespace at the edge of a comment: a CR followed by more
// trailing whitespace once survived the first pass and disappeared on
// re-lexing, so formatting was not idempotent.
f.Add("//\r ")
f.Add("// loop \r\t\nMOVQ AX, BX\n")
f.Add("TEXT ·f(SB), NOSPLIT, $0 // tail\r\n\tMOVQ AX, BX\r\n\tRET\r\n")
f.Fuzz(func(t *testing.T, src string) {
once := Source(src)
@@ -0,0 +1,2 @@
go test fuzz v1
string("//\r ")
+62 -7
View File
@@ -16,8 +16,16 @@ import (
"sourcedock.dev/petrbalvin/gasm-devkit/token"
)
// middleDot is the Plan 9 symbol separator (U+00B7), used in ·funcName(SB).
const middleDot = '\u00B7'
const (
// middleDot is the Plan 9 symbol separator (U+00B7), used in
// ·funcName(SB): it stands for the period between package path and name.
middleDot = '\u00B7'
// divisionSlash is the Plan 9 path separator (U+2215), used inside the
// package path of a symbol: internal∕runtime∕atomic·Xchg. Like the
// middle dot it is an identifier character, so a package path containing
// it lexes as one name; the ordinary slash (U+002F) stays punctuation.
divisionSlash = '\u2215'
)
// Lexer scans a source string one token at a time.
type Lexer struct {
@@ -119,7 +127,9 @@ func (l *Lexer) Next() token.Token {
// is a C-preprocessor line continuation (used by #define macros in the
// runtime .s files): splice the lines together by consuming both, so
// the whole macro becomes one logical line that the parser treats as an
// opaque preprocessor directive.
// opaque preprocessor directive. The backslash may also reach its
// newline across whitespace and a trailing comment ("…; \ // note\n"),
// which the toolchain's scanner skips the same way.
for {
c := l.cur()
if c == ' ' || c == '\t' || c == '\r' {
@@ -136,6 +146,16 @@ func (l *Lexer) Next() token.Token {
}
continue
}
if c == '\\' && l.continuationAhead() {
l.advance() // backslash, then the runes the scan saw
for !l.atEnd() && l.cur() != '\n' {
l.advance()
}
if !l.atEnd() {
l.advance() // the newline that closes the continuation
}
continue
}
break
}
@@ -185,16 +205,42 @@ func (l *Lexer) Next() token.Token {
}
}
// continuationAhead reports, without consuming anything, whether the
// backslash at the current position closes onto a newline through nothing
// but horizontal whitespace and one line comment. Positions after the
// backslash are inspected directly on the rune slice so a non-match leaves
// the scanner state untouched.
func (l *Lexer) continuationAhead() bool {
i := l.i + 1
for i < len(l.src) {
switch r := l.src[i]; {
case r == ' ' || r == '\t' || r == '\r':
i++
case r == '/' && i+1 < len(l.src) && l.src[i+1] == '/':
for i < len(l.src) && l.src[i] != '\n' {
i++
}
default:
return r == '\n'
}
}
return false
}
// lineComment consumes a // comment up to, but not including, the newline. A
// trailing \r is part of a CRLF line ending rather than comment content:
// dropping it keeps the formatter's output uniformly LF-terminated.
// trailing run of \r, spaces and tabs is line-ending whitespace rather than
// comment content, so it never enters the token text. Trimming only a \r
// directly before the token's end would make the text depend on what follows
// the comment (a newline or the end of the input): "//x\r " would carry the
// "\r " while "//x\r\n" would not, and a formatter that terminates the line
// with \n would then re-lex its own output to a shorter comment.
func (l *Lexer) lineComment(start token.Position) token.Token {
var b strings.Builder
for !l.atEnd() && l.cur() != '\n' {
b.WriteRune(l.cur())
l.advance()
}
return l.make(token.Comment, start, strings.TrimSuffix(b.String(), "\r"))
return l.make(token.Comment, start, strings.TrimRight(b.String(), " \t\r"))
}
// blockComment consumes a /* ... */ comment, tolerating an unterminated one.
@@ -399,6 +445,15 @@ func (l *Lexer) punct(start token.Position) token.Token {
case '|':
l.advance()
return l.make(token.Pipe, start, "|")
case ';':
l.advance()
return l.make(token.Semicolon, start, ";")
case '&':
l.advance()
return l.make(token.Ampersand, start, "&")
case '~':
l.advance()
return l.make(token.Tilde, start, "~")
default:
// Unknown rune: emit it as Illegal and move on.
l.advance()
@@ -413,7 +468,7 @@ func isHexDigit(r rune) bool {
}
func isIdentStart(r rune) bool {
return r == '_' || r == middleDot || unicode.IsLetter(r)
return r == '_' || r == middleDot || r == divisionSlash || unicode.IsLetter(r)
}
func isIdentChar(r rune) bool {
+26
View File
@@ -85,6 +85,19 @@ func TestLabelAndComment(t *testing.T) {
[]token.Kind{token.Ident, token.Colon, token.Ident, token.Ident, token.Comment})
}
func TestLineCommentTrailingWhitespace(t *testing.T) {
// A trailing run of CR, spaces and tabs is line-ending whitespace, not
// comment content. The token text must not depend on what follows the
// comment: before the trim covered only a CR directly before the token's
// end, "// loop\r " kept the CR while "// loop\r\n" dropped it, and the
// formatter re-lexed its own output to a shorter comment.
eq(t, texts("// loop\r"), []string{"// loop"})
eq(t, texts("// loop\r "), []string{"// loop"})
eq(t, texts("// loop \r\t\nMOVQ AX, BX"), []string{"// loop", "MOVQ", "AX", ",", "BX"})
// A CR inside the comment is content and stays.
eq(t, texts("// loops\rall"), []string{"// loops\rall"})
}
func TestAVX512Mnemonics(t *testing.T) {
eq(t, texts("VFMADD231PD Z14, Z12, Z10"),
[]string{"VFMADD231PD", "Z14", ",", "Z12", ",", "Z10"})
@@ -169,6 +182,19 @@ func TestNulIsIllegal(t *testing.T) {
eq(t, texts("MOVQ \x00 AX"), []string{"MOVQ", "\x00", "AX"})
}
func TestDivisionSlashInIdentifiers(t *testing.T) {
// U+2215 DIVISION SLASH is an identifier character, the way the
// toolchain's tokenizer treats it: the package path of a symbol is
// written with it (internal∕runtime∕atomic·Xchg) and must lex as one
// name. The ordinary slash (U+002F) stays punctuation.
eq(t, texts("CALL internal∕runtime∕atomic·Xchg(SB)"),
[]string{"CALL", "internal∕runtime∕atomic·Xchg", "(", "SB", ")"})
eq(t, texts("MOVQ sync∕atomic·Align(SB), AX"),
[]string{"MOVQ", "sync∕atomic·Align", "(", "SB", ")", ",", "AX"})
// It may also begin a name, like any letter of the toolchain's rule.
eq(t, kinds("∕x"), []token.Kind{token.Ident})
}
// TestOffsetsAroundInvalidByte pins Position.Offset against the original
// bytes: an invalid UTF-8 byte decodes to RuneError but advances the offset
// table by exactly one byte, so every later position stays a true byte
+6
View File
@@ -21,6 +21,12 @@ import (
// The check requires a parseable signature; functions without one, and
// functions whose parameters are all covered by frame reads, stay silent.
func checkABI0Args(t *ast.Text) []Diagnostic {
// An explicit <ABIInternal> TEXT reads its arguments from the register
// file by declaration (runtime·memmove<ABIInternal> is the canonical
// example), so the ABI0 frame contract does not apply to it.
if t.Name != nil && t.Name.ABI != "" {
return nil
}
params, ok := abiParamNames(t.Doc)
if !ok || len(params) == 0 {
return nil
+17
View File
@@ -82,6 +82,23 @@ func TestABIArgSizeSkipsRegisterABI(t *testing.T) {
}
}
// TestABI0ArgsSkipsABIInternal verifies the frame-read check does not fire for
// a TEXT declared <ABIInternal>: runtime·memmove<ABIInternal> and friends read
// their arguments from the register file by declaration, which is the correct
// spelling there, not the register-args port bug the rule hunts.
func TestABI0ArgsSkipsABIInternal(t *testing.T) {
diags := lintSrc(t, "#include \"textflag.h\"\n"+
"// func memmove(to, from unsafe.Pointer, n uintptr)\n"+
"TEXT ·memmove<ABIInternal>(SB), NOSPLIT, $0-24\n"+
"\tMOVQ AX, DI\n"+
"\tMOVQ BX, SI\n"+
"\tMOVQ CX, BX\n"+
"\tRET\n")
if codes(diags)[CodeABI0RegisterArgs] != 0 {
t.Fatalf("ABIInternal TEXT must not be checked against the FP frame: %+v", diags)
}
}
// TestUnreachableCode exercises the dead-code detection and its guard rails.
func TestUnreachableCode(t *testing.T) {
// Code after a RET is unreachable.
+111 -8
View File
@@ -338,10 +338,8 @@ func lintText(t *ast.Text, tab *arch.Table, archKnown bool, cfg Config, macros m
}
if isJump(cfg.Arch, upper) {
for _, op := range st.Operands {
if name, pos, ok := localLabelRef(op); ok && !tab.IsRegister(name) && !arch.IsPseudoReg(name) {
referenced[name] = pos
}
if name, pos, ok := branchTargetRef(cfg.Arch, upper, st.Operands, tab); ok {
referenced[name] = pos
}
}
}
@@ -649,17 +647,65 @@ func localLabelRef(op *ast.Operand) (string, token.Position, bool) {
return sym.Name, op.Pos, true
}
// branchTargetRef returns the local label a branch transfers control to: the
// bare symbol in the destination position, the last operand, since that is
// where the Plan 9 branch target sits. A register-named target is a
// register-indirect branch (JMP AX, arm64 BR R5, riscv64 JALR X6, loong64
// JIRL R1) and yields no reference, unless the encoder reads the target
// positionally (positionalBranchTarget): there a label may legitimately
// collide with a register alias, riscv64 ZERO being the ABI name of X0, and
// a label named zero is ordinary code.
func branchTargetRef(a arch.Arch, upper string, ops []*ast.Operand, tab *arch.Table) (string, token.Position, bool) {
if len(ops) == 0 {
return "", token.Position{}, false
}
name, pos, ok := localLabelRef(ops[len(ops)-1])
if !ok {
return "", token.Position{}, false
}
if !positionalBranchTarget(a, upper) && (tab.IsRegister(name) || arch.IsPseudoReg(name)) {
return "", token.Position{}, false
}
return name, pos, true
}
// positionalBranchTarget reports whether the encoder reads a bare-symbol
// operand of the branch as its label target from a fixed position, without
// consulting the register file. The riscv64 branch, JMP and JAL encoders do
// (labelFromOperand in asm/riscv_assemble.go), as do the loong64 branch,
// BFPT/BFPF and jump encoders (l64Label in asm/loong64_assemble.go). amd64
// never does, because a bare register operand to JMP/CALL/Jcc is a
// register-indirect branch; nor do the register-indirect forms of the RISC
// families (arm64 BR/BLR, riscv64 JALR/JR, loong64 JIRL).
func positionalBranchTarget(a arch.Arch, upper string) bool {
switch a {
case arch.RISCV:
return riscvBranches[upper] || upper == "JMP" || upper == "JAL"
case arch.LOONG64:
return loong64Branches[upper] || upper == "JMP" || upper == "B" ||
upper == "JAL" || upper == "BL"
}
return false
}
// riscvBranches and loong64Branches are the conditional-branch mnemonics; they
// are listed explicitly rather than matched by a "B" prefix so that bit-manip
// instructions (BCLR, BSET, …) are never mistaken for branches.
// instructions (BCLR, BSET, …) are never mistaken for branches. The sets
// mirror the encoder's own branch cases: the B-type table entries
// (riscv_encode.go), the branch-zero pseudos and the reversed branches
// BGT/BGTU/BLE/BLEU (riscv_assemble.go), and for loong64 the 16-bit branch
// table plus the single-register forms of l64branch21Table (BEQZ/BNEZ and the
// floating-point branches BFPT/BFPF).
var riscvBranches = map[string]bool{
"BEQ": true, "BNE": true, "BLT": true, "BGE": true, "BLTU": true, "BGEU": true,
"BEQZ": true, "BNEZ": true, "BLEZ": true, "BGEZ": true, "BLTZ": true, "BGTZ": true,
"BGT": true, "BGTU": true, "BLE": true, "BLEU": true,
}
var loong64Branches = map[string]bool{
"BEQ": true, "BNE": true, "BLT": true, "BGE": true, "BLTU": true, "BGEU": true,
"BLEZ": true, "BLTZ": true, "BGEZ": true, "BGTZ": true,
"BEQZ": true, "BNEZ": true, "BFPT": true, "BFPF": true,
}
// isJump reports whether the mnemonic is any branch.
@@ -676,7 +722,8 @@ func isJump(a arch.Arch, upper string) bool {
upper == "JR" || upper == "BR"
case arch.LOONG64:
return upper == "CALL" || loong64Branches[upper] ||
upper == "JIRL" || upper == "JMP" || upper == "BR"
upper == "JIRL" || upper == "JMP" || upper == "BR" ||
upper == "B" || upper == "JAL" || upper == "BL"
default: // amd64
return upper == "CALL" || strings.HasPrefix(upper, "J")
}
@@ -692,7 +739,8 @@ func isUnconditionalJump(a arch.Arch, upper string) bool {
return upper == "JMP" || upper == "J" || upper == "JAL" ||
upper == "JALR" || upper == "JR" || upper == "BR"
case arch.LOONG64:
return upper == "JMP" || upper == "JIRL" || upper == "BR"
return upper == "JMP" || upper == "JIRL" || upper == "BR" || upper == "B" ||
upper == "JAL" || upper == "BL"
default:
return upper == "JMP"
}
@@ -792,6 +840,43 @@ func isSPReg(op *ast.Operand, a arch.Arch) bool {
return false
}
// shiftRotateBases are the shift and rotate mnemonics without their width
// suffix. These are the instructions whose encoder path (encodeShift) reads
// the count from the first operand.
var shiftRotateBases = map[string]bool{
"SHL": true, "SHR": true, "SAR": true, "SAL": true,
"ROL": true, "ROR": true, "RCL": true, "RCR": true,
}
// isShiftCountOperand reports whether operand i of mnem is the shift count.
// The ISA fixes the shift/rotate count register at CL: the D2/D3 group (and
// C0/C1 for immediates) encode the count outside the ModRM register field,
// so the count operand is 8-bit by definition no matter how wide the data is.
// The count arrives as the first of the two operands; the one-operand form
// does not exist.
func isShiftCountOperand(mnem string, i, nops int) bool {
if nops != 2 || i != 0 {
return false
}
if shiftRotateBases[mnem] {
return true
}
if len(mnem) > 1 {
switch mnem[len(mnem)-1] {
case 'Q', 'L', 'W', 'B':
return shiftRotateBases[mnem[:len(mnem)-1]]
}
}
return false
}
// isSetcc reports whether the mnemonic is a SETcc: SET plus a condition code.
// The membership test is the encoder's own SET dispatch, which asm.Encodable
// mirrors.
func isSetcc(mnem string) bool {
return strings.HasPrefix(mnem, "SET") && asm.Encodable(mnem)
}
// checkRegisterWidth detects amd64 register-width mismatches. The naming
// truth of the Go assembler governs: AX, BX, CX, DX, SI, DI, BP, SP and
// R8-R15 ARE the 64-bit register names (there are no separate EAX/RAX
@@ -802,6 +887,14 @@ func isSPReg(op *ast.Operand, a arch.Arch) bool {
// register (EAX under the gasm alias extension, or a byte form), and byte
// registers in L/W operations.
func checkRegisterWidth(mnem string, ops []*ast.Operand) string {
// A SETcc stores one byte: the destination is an 8-bit register or an
// 8-bit memory location by definition (0F 90+cc), whichever condition it
// tests. The trailing letter of spellings like SETPL or SETEQ is part of
// the condition code, not an operand width, so the whole family is
// exempt from the suffix logic.
if isSetcc(mnem) {
return ""
}
// Determine expected width from mnemonic suffix.
var expected int // 0=unknown, 8/4/2/1=bytes
switch {
@@ -816,10 +909,20 @@ func checkRegisterWidth(mnem string, ops []*ast.Operand) string {
default:
return "" // no suffix, can't determine width
}
for _, op := range ops {
for i, op := range ops {
if op.Kind != ast.OpAddr || op.Addr.Sym == nil {
continue
}
// Only a bare register carries a width to compare: frame and static
// symbol references (ch+8(FP), foo(SB)) and memory operands are not
// registers even when their name collides with one.
if op.Addr.Sym.Pseudo != "" || op.Addr.Base != "" || op.Addr.Index != "" {
continue
}
// The shift/rotate count is exempt: fixed at 8 bits by the ISA.
if isShiftCountOperand(mnem, i, len(ops)) {
continue
}
name := strings.ToLower(op.Addr.Sym.Name)
regWidth := amd64RegWidth(name)
if regWidth == 0 {
+214
View File
@@ -8,6 +8,7 @@ import (
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
"sourcedock.dev/petrbalvin/gasm-devkit/asm"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
)
@@ -21,6 +22,21 @@ func lintSrc(t *testing.T, src string) []Diagnostic {
return File(f, Config{Arch: arch.AMD64})
}
// lintArchFile parses and lints src under a, then hands the same file to
// assemble so the assertion is pinned against the encoder: a kernel the
// linter reasons about must also be one the encoder accepts.
func lintArchFile(t *testing.T, filename, src string, a arch.Arch, assemble func(*ast.File) (*asm.Image, error)) []Diagnostic {
t.Helper()
f, errs := parser.Parse(filename, src)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
if _, err := assemble(f); err != nil {
t.Fatalf("encoder rejects the kernel: %v", err)
}
return File(f, Config{Arch: a})
}
// lintSrcArch lints src under the architecture inferred from filename.
func lintSrcArch(t *testing.T, filename, src string) []Diagnostic {
t.Helper()
@@ -262,6 +278,138 @@ loop:
}
}
func TestRiscvBranchFamilyRegistersLabels(t *testing.T) {
// Every riscv64 pseudo-branch that references a label must register that
// reference: the reversed branches BGT/BGTU/BLE/BLEU (GOROOT's
// memmove_riscv64 branches with BGTU) and a label named like the ZERO
// register alias (GOROOT's memclr_riscv64 carries a label named zero;
// ZERO is the ABI name of X0) must not be reported unused.
diags := lintSrcArch(t, "f_riscv64.s", `
#include "textflag.h"
TEXT ·f(SB), NOSPLIT, $0
BGTU X10, X11, backward
BGT X10, X11, zero
BLE X10, X11, one
BLEU X10, X11, two
BEQZ X10, zero
BNEZ X10, one
JMP two
backward:
RET
zero:
RET
one:
RET
two:
RET
`)
if codes(diags)[CodeUnusedLabel] != 0 {
t.Fatalf("branch-referenced labels must not be flagged unused: %+v", diags)
}
if codes(diags)[CodeUndefinedLabel] != 0 {
t.Fatalf("defined labels must resolve: %+v", diags)
}
// A branch to a truly undefined label still reports.
diags = lintSrcArch(t, "f_riscv64.s", `
#include "textflag.h"
TEXT ·f(SB), NOSPLIT, $0
BGT X10, X11, nowhere
RET
`)
if codes(diags)[CodeUndefinedLabel] != 1 {
t.Fatalf("undefined branch target must be flagged: %+v", diags)
}
// A register-indirect JALR is not a label reference.
diags = lintSrcArch(t, "f_riscv64.s", `
#include "textflag.h"
TEXT ·f(SB), NOSPLIT, $0
JALR X1
RET
`)
if codes(diags)[CodeUndefinedLabel] != 0 {
t.Fatalf("register operand of JALR is not a label: %+v", diags)
}
}
func TestLoong64BranchFamilyRegistersLabels(t *testing.T) {
// The loong64 jumps and single-register branches (JAL, B, BL, BEQZ/BNEZ,
// BFPT/BFPF) all reference their label from the last operand; GOROOT's
// own basic kernels tail-call with JAL, so the reference must register.
diags := lintSrcArch(t, "f_loong64.s", `
#include "textflag.h"
TEXT ·f(SB), NOSPLIT, $0
BEQZ R4, fin
BNEZ R4, fin
BLTZ R4, fin
JAL fin
BL fin
B fin
RET
fin:
RET
`)
if codes(diags)[CodeUnusedLabel] != 0 {
t.Fatalf("branch-referenced labels must not be flagged unused: %+v", diags)
}
if codes(diags)[CodeUndefinedLabel] != 0 {
t.Fatalf("defined labels must resolve: %+v", diags)
}
}
// TestBranchFamiliesAssemble pins the lint branch sets to the encoder: every
// mnemonic the linter classifies as a riscv64 or loong64 label branch must be
// a branch the encoder actually assembles, with the label in the last
// operand. If the encoder gains or renames a branch, this test fails and the
// set follows it.
func TestBranchFamiliesAssemble(t *testing.T) {
riscvForms := map[string]string{}
for m := range riscvBranches {
riscvForms[m] = m + " X10, X11, tgt"
}
for _, m := range []string{"BEQZ", "BNEZ", "BLTZ", "BGEZ", "BLEZ", "BGTZ"} {
riscvForms[m] = m + " X10, tgt"
}
riscvForms["JMP"] = "JMP tgt"
riscvForms["JAL"] = "JAL tgt"
loongForms := map[string]string{}
for _, m := range []string{"BEQ", "BNE", "BLT", "BGE", "BLTU", "BGEU"} {
loongForms[m] = m + " R4, R5, tgt"
}
for _, m := range []string{"BEQZ", "BNEZ", "BLTZ", "BGEZ", "BLEZ", "BGTZ", "BFPT", "BFPF"} {
loongForms[m] = m + " R4, tgt"
}
loongForms["JMP"] = "JMP tgt"
loongForms["B"] = "B tgt"
loongForms["JAL"] = "JAL tgt"
loongForms["BL"] = "BL tgt"
for m, form := range riscvForms {
src := "#include \"textflag.h\"\n" +
"TEXT ·f(SB), NOSPLIT, $0\n" +
"\t" + form + "\n" +
"tgt:\n" +
"\tRET\n"
diags := lintArchFile(t, "f_riscv64.s", src, arch.RISCV, asm.AssembleFileRISCV)
if codes(diags)[CodeUnusedLabel] != 0 || codes(diags)[CodeUndefinedLabel] != 0 {
t.Errorf("riscv64 %s: label reference not registered: %+v", m, diags)
}
}
for m, form := range loongForms {
src := "#include \"textflag.h\"\n" +
"TEXT ·f(SB), NOSPLIT, $0\n" +
"\t" + form + "\n" +
"tgt:\n" +
"\tRET\n"
diags := lintArchFile(t, "f_loong64.s", src, arch.LOONG64, asm.AssembleFileLOONG64)
if codes(diags)[CodeUnusedLabel] != 0 || codes(diags)[CodeUndefinedLabel] != 0 {
t.Errorf("loong64 %s: label reference not registered: %+v", m, diags)
}
}
}
func TestInvalidTextflag(t *testing.T) {
diags := lintSrc(t, `
#include "textflag.h"
@@ -413,6 +561,72 @@ TEXT ·f(SB), NOSPLIT, $0
}
}
func TestRegisterWidthShiftCount(t *testing.T) {
// The shift and rotate count lives in CL by ISA definition (the D2/D3
// group encodes the count outside the ModRM register field), so the count
// operand is 8-bit no matter how wide the data is: SHLQ CL, AX is the
// normal spelling of a 64-bit shift. The data operand keeps its check.
diags := lintSrc(t, `
#include "textflag.h"
TEXT ·f(SB), NOSPLIT, $0
SHLQ CL, AX
SHRL CL, BX
SARQ CL, CX
ROLL CL, DX
RORQ CL, R8
RCLL CL, R9
RCRQ CL, R10
MOVQ CL, R10
RET
`)
if codes(diags)[CodeRegisterWidthMismatch] != 1 {
t.Fatalf("only the MOVQ CL data move must be flagged, got %+v", diags)
}
}
func TestRegisterWidthSetcc(t *testing.T) {
// A SETcc stores one byte whichever condition it tests (0F 90+cc), so
// SETNE AL is always right and the trailing letters of SETEQ, SETPL and
// SETLS are condition codes, not width suffixes.
diags := lintSrc(t, `
#include "textflag.h"
TEXT ·f(SB), NOSPLIT, $0
CMPQ AX, BX
SETNE AL
SETEQ AL
SETPL AL
SETLS AL
SETCC (BX)
SETGE (R8)
RET
`)
if codes(diags)[CodeRegisterWidthMismatch] != 0 {
t.Fatalf("SETcc destinations are 8-bit by definition: %+v", diags)
}
if codes(diags)[CodeUnknownInstr] != 0 {
t.Fatalf("every SETcc spelling must be known: %+v", diags)
}
}
func TestRegisterWidthFrameNames(t *testing.T) {
// GOROOT's BSD syscall stubs carry frame parameters whose names collide
// with byte register names (kevent's ch and nch): MOVQ ch+8(FP), SI is a
// frame reference, not the CH register.
diags := lintSrc(t, `
#include "textflag.h"
TEXT ·kevent(SB), NOSPLIT, $0-36
MOVL kq+0(FP), DI
MOVQ ch+8(FP), SI
MOVL nch+16(FP), DX
MOVQ ev+24(FP), R10
MOVQ AX, ret+32(FP)
RET
`)
if codes(diags)[CodeRegisterWidthMismatch] != 0 {
t.Fatalf("frame and static symbol names are not registers: %+v", diags)
}
}
func TestNonportableRegisterName(t *testing.T) {
diags := lintSrc(t, `
#include "textflag.h"
+146
View File
@@ -0,0 +1,146 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// Constant-expression folding for operands. The toolchain's assembler
// evaluates arithmetic in every operand position, and macro-heavy GOROOT
// sources lean on it: parameterised bodies carry offsets like
// ((index*4)+0)(base), immediates like $(32-shift) and masks like
// $~63 or $(1<<0|1<<9). Substituting the parameters textually therefore
// leaves constant arithmetic behind, and the parser folds it here, keeping
// the operand AST identical to what the same literals written out would
// produce. Anything that is not a closed integer expression fails to fold
// and falls through to the ordinary operand paths.
package parser
import (
"sourcedock.dev/petrbalvin/gasm-devkit/token"
)
// foldExpr evaluates the constant integer expression at the head of ts and
// returns its value together with the unconsumed tokens. ok is false when
// the tokens do not form an expression, which is the callers' signal to use
// the ordinary parsing paths.
func foldExpr(ts []token.Token) (val int64, rest []token.Token, ok bool) {
v, rest, ok := foldAdd(ts)
if !ok {
return 0, ts, false
}
return v, rest, true
}
// foldAdd parses addition-level expressions: +, - and | bind loosest, the
// Plan 9 convention that makes x<<1|3 read as (x<<1)|3.
func foldAdd(ts []token.Token) (int64, []token.Token, bool) {
v, rest, ok := foldMul(ts)
if !ok {
return 0, ts, false
}
for len(rest) > 0 {
kind := rest[0].Kind
if kind != token.Plus && kind != token.Minus && kind != token.Pipe {
return v, rest, true
}
w, r2, ok := foldMul(rest[1:])
if !ok {
return v, rest, true
}
switch kind {
case token.Plus:
v += w
case token.Minus:
v -= w
case token.Pipe:
v |= w
}
rest = r2
}
return v, rest, true
}
// foldMul parses multiplication-level expressions: *, / and the bit
// operators &, << and >>.
func foldMul(ts []token.Token) (int64, []token.Token, bool) {
v, rest, ok := foldFactor(ts)
if !ok {
return 0, ts, false
}
for len(rest) > 0 {
switch rest[0].Kind {
case token.Star:
w, r2, ok := foldFactor(rest[1:])
if !ok {
return v, rest, true
}
v *= w
rest = r2
case token.Slash:
w, r2, ok := foldFactor(rest[1:])
if !ok || w == 0 {
return v, rest, true
}
v /= w
rest = r2
case token.Ampersand:
w, r2, ok := foldFactor(rest[1:])
if !ok {
return v, rest, true
}
v &= w
rest = r2
case token.LShift:
w, r2, ok := foldFactor(rest[1:])
if !ok || w < 0 || w >= 64 {
return v, rest, true
}
v <<= uint(w)
rest = r2
case token.RShift:
w, r2, ok := foldFactor(rest[1:])
if !ok || w < 0 || w >= 64 {
return v, rest, true
}
v >>= uint(w)
rest = r2
default:
return v, rest, true
}
}
return v, rest, true
}
// foldFactor parses a number, a parenthesised expression, or a unary sign
// or complement.
func foldFactor(ts []token.Token) (int64, []token.Token, bool) {
if len(ts) == 0 {
return 0, ts, false
}
switch ts[0].Kind {
case token.Number:
v, ok := tryInt(ts[0].Text)
if !ok {
return 0, ts, false
}
return v, ts[1:], true
case token.LParen:
v, rest, ok := foldAdd(ts[1:])
if !ok || len(rest) == 0 || rest[0].Kind != token.RParen {
return 0, ts, false
}
return v, rest[1:], true
case token.Minus:
v, rest, ok := foldFactor(ts[1:])
if !ok {
return 0, ts, false
}
return -v, rest, true
case token.Plus:
return foldFactor(ts[1:])
case token.Tilde:
v, rest, ok := foldFactor(ts[1:])
if !ok {
return 0, ts, false
}
return ^v, rest, true
}
return 0, ts, false
}
+124 -8
View File
@@ -32,9 +32,8 @@ func (e Error) Error() string {
// returned file is usable even when errors is non-empty.
func Parse(path, src string) (*ast.File, []error) {
tokens := lexer.Tokenize(src)
lines := splitLines(tokens)
p := &state{path: path}
p.parse(lines)
p.parse(statementLines(tokens))
return p.file, p.errs
}
@@ -73,6 +72,46 @@ func splitLines(tokens []token.Token) [][]token.Token {
return lines
}
// statementLines turns the token stream into the logical lines the parser
// reads: physical lines split at the ';' statement separators, exactly the
// way the expansion path treats the expanded bodies. The runtime writes
// "ROLQ $3, DI; ROLQ $13, DI" and "REP; MOVSB" in plain files, and the
// separator carries no meaning beyond the break. Comments are statement
// text, not structure: the lexer delivers a whole comment as one token, so
// a ';' inside a comment is never a separator; a comment after a statement
// stays on that statement's line; and a comment that sits between
// statements (the runtime's "NO_LOCAL_POINTERS; /* … */" style) stands as
// its own logical line, like a whole-line comment.
func statementLines(tokens []token.Token) [][]token.Token {
var out [][]token.Token
var cur []token.Token
flush := func() {
if len(cur) > 0 {
out = append(out, cur)
cur = nil
}
}
for _, t := range tokens {
switch t.Kind {
case token.EOF:
// The stream's terminator is not statement content.
case token.Newline, token.Semicolon:
flush()
case token.Comment:
if len(cur) > 0 {
cur = append(cur, t)
} else {
out = append(out, []token.Token{t})
}
flush()
default:
cur = append(cur, t)
}
}
flush()
return out
}
func (p *state) parse(lines [][]token.Token) {
p.file = &ast.File{Path: p.path, Macros: map[string]bool{}}
for _, line := range lines {
@@ -265,7 +304,7 @@ func (p *state) parseGlobl(line []token.Token) *ast.Globl {
rest = rest[1:]
}
if len(rest) > 0 && rest[0].Kind == token.Dollar {
g.Size = parseOperand(rest)
g.Size = parseOperand(rest, false)
}
return g
}
@@ -282,7 +321,7 @@ func (p *state) parseData(line []token.Token) *ast.Data {
d.Name = sym
d.Width = width
if len(valuePart) > 0 {
d.Value = parseOperand(stripComment(valuePart))
d.Value = parseOperand(stripComment(valuePart), false)
}
return d
}
@@ -293,8 +332,13 @@ func (p *state) parseInstr(line []token.Token) {
return
}
instr := &ast.Instr{Mnemonic: body[0], Comment: comment}
for _, grp := range splitOperands(body[1:]) {
if op := parseOperand(grp); op != nil {
grps := splitOperands(body[1:])
for i, grp := range grps {
// Only the final operand slot may carry a bare constant: the
// toolchain reads the trailing 1 of CMPSD X1, X0, 1 as $1
// (math/floor_amd64.s), while an earlier bare number names an
// absolute address, a form this parser keeps out of the tree.
if op := parseOperand(grp, i == len(grps)-1); op != nil {
instr.Operands = append(instr.Operands, op)
}
}
@@ -378,8 +422,10 @@ func setName(raw string, sym *ast.Symbol) {
// --- operand parsing --------------------------------------------------------
// parseOperand parses one operand group into an Operand.
func parseOperand(g []token.Token) *ast.Operand {
// parseOperand parses one operand group into an Operand. allowBare marks
// the final operand slot of an instruction, where the toolchain reads a
// bare constant expression as an immediate.
func parseOperand(g []token.Token, allowBare bool) *ast.Operand {
g = stripComment(g)
if len(g) == 0 {
return nil
@@ -392,9 +438,25 @@ func parseOperand(g []token.Token) *ast.Operand {
}
op.Kind = ast.OpAddr
op.Addr = parseAddress(g)
// A trailing bare constant leaves every address field empty: the
// grammar sees no register, memory reference or symbol, and the closed
// constant expression is the whole group. Read it as the immediate it
// names, exactly what the $ spelling would produce.
if allowBare && isEmptyAddress(op.Addr) {
if v, rest, ok := foldExpr(g); ok && len(rest) == 0 {
op.Kind = ast.OpImmediate
op.Imm = ast.Immediate{Val: v, HasVal: true}
}
}
return op
}
// isEmptyAddress reports whether parseAddress populated nothing, its sign
// that the group is no register, memory reference, symbol or register range.
func isEmptyAddress(a ast.Address) bool {
return a.Sym == nil && a.Base == "" && a.Index == "" && a.Range == nil && a.Shift == ""
}
// parseImmediate parses the tokens following a '$'.
func parseImmediate(g []token.Token) ast.Immediate {
var imm ast.Immediate
@@ -408,6 +470,18 @@ func parseImmediate(g []token.Token) ast.Immediate {
return imm
}
}
// A constant expression introduced by '(' or '~'. Textual macro
// substitution leaves arithmetic such as $(32-shift) and $~63 behind,
// and the toolchain evaluates it in place; only shapes the ordinary
// paths below cannot read reach the folder, so every existing form
// keeps its exact parse.
if g[0].Kind == token.LParen || g[0].Kind == token.Tilde {
if v, rest, ok := foldExpr(g); ok && len(rest) == 0 {
imm.Val = v
imm.HasVal = true
return imm
}
}
i := 0
if g[i].Kind == token.Minus {
imm.Neg = true
@@ -446,6 +520,14 @@ func parseAddress(g []token.Token) ast.Address {
if len(g) == 0 {
return addr
}
// A bracketed register range, [Z0-Z3]: the amd64 4FMAPS/4VNNIW
// multi-source operand. The bracket runes arrive as Illegal tokens
// (the lexer has no bracket kind), so the shape matches on their text.
if isBracket(g[0], "[") && len(g) == 5 && g[1].Kind == token.Ident &&
g[2].Kind == token.Minus && g[3].Kind == token.Ident && isBracket(g[4], "]") {
addr.Range = &ast.RegRange{Lo: g[1].Text, Hi: g[3].Text, Pos: g[0].Pos}
return addr
}
// Symbol-with-pseudo form: name[<>][+off](PSEUDO).
// When the prefix is not a valid symbol name (e.g. a bare number like
// 0(SP) in RISC-V), sym is nil, and we fall through to regular memory
@@ -459,6 +541,17 @@ func parseAddress(g []token.Token) ast.Address {
}
i := 0
// A parenthesised constant expression as the displacement: substituted
// macro bodies carry ((index*4)+0)(base) shapes. As with the signed
// number path below, the value is committed only when a base group
// follows.
if i < len(g) && g[i].Kind == token.LParen {
if v, rest, ok := foldExpr(g[i:]); ok && len(rest) > 0 && rest[0].Kind == token.LParen {
addr.Offset = v
addr.HasOff = true
i = len(g) - len(rest)
}
}
// Optional leading displacement before a '(' base group. A sign pushes
// the parenthesis one token further out: -4(DX) has it at i+2.
if isSignedNumber(g, i) {
@@ -515,6 +608,16 @@ func parseAddress(g []token.Token) ast.Address {
}
}
}
// A lone (index*scale) group is the VSIB index-only form: the
// gather/scatter families address memory through a scaled vector index
// with no base register, 8(X4*1). The two-group grammar below reads
// (base)(index*scale), so a first group whose member carries a scale
// factor can only be an index.
if isIndexGroup(g[i:]) {
addr.Index = g[i+1].Text
addr.Scale = int(parseInt(g[i+3].Text))
i += 5
}
// First parenthesised group: the base register.
if i < len(g) && g[i].Kind == token.LParen {
i++
@@ -571,6 +674,19 @@ func findPseudoParen(g []token.Token) int {
return -1
}
// isBracket reports whether t is a square bracket. The lexer has no bracket
// kind, so '[' and ']' arrive as Illegal tokens.
func isBracket(t token.Token, text string) bool {
return t.Kind == token.Illegal && t.Text == text
}
// isIndexGroup reports whether g begins with a complete (index*scale) group:
// one identifier followed by a scale factor, all inside a single parenthesis.
func isIndexGroup(g []token.Token) bool {
return len(g) >= 5 && g[0].Kind == token.LParen && g[1].Kind == token.Ident &&
g[2].Kind == token.Star && g[3].Kind == token.Number && g[4].Kind == token.RParen
}
// --- token helpers ----------------------------------------------------------
// splitOperands splits a token slice on top-level commas (commas outside any
+233
View File
@@ -401,3 +401,236 @@ func TestInt64MinimumImmediate(t *testing.T) {
t.Errorf("imm.Float = %q, want empty", imm.Float)
}
}
// TestDivisionSlashPackagePath covers the runtime's package-path spelling:
// U+2215 DIVISION SLASH separates the elements of an import path inside a
// symbol (internal∕runtime∕atomic·Xchg), and the middle dot still separates
// the package from the name. The whole spelling must reach the symbol, not
// stop at the first slash.
func TestDivisionSlashPackagePath(t *testing.T) {
file, errs := Parse("t.s", "TEXT \u00b7f(SB), $0\n\tCALL internal∕runtime∕atomic·Xchg(SB)\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse errors: %v", errs)
}
txt := file.Decls[0].(*ast.Text)
instr := txt.Body[0].(*ast.Instr)
sym := instr.Operands[0].Addr.Sym
if sym == nil {
t.Fatal("operand carries no symbol")
}
if sym.Pkg != "internal∕runtime∕atomic" {
t.Errorf("pkg = %q, want internal∕runtime∕atomic", sym.Pkg)
}
if sym.Name != "Xchg" {
t.Errorf("name = %q, want Xchg", sym.Name)
}
if sym.Raw != "internal∕runtime∕atomic·Xchg(SB)" {
t.Errorf("raw = %q", sym.Raw)
}
}
// TestSemicolonStatements covers the plain parse path: ';' separates
// statements on one line exactly as it does inside macro expansion, and a
// ';' inside a comment is comment text.
func TestSemicolonStatements(t *testing.T) {
file, errs := Parse("t.s", "TEXT \u00b7f(SB), $0\n\tROLQ $3, DI; ROLQ $13, DI\n\tMOVQ AX, BX // note; still comment\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse errors: %v", errs)
}
txt := file.Decls[0].(*ast.Text)
if len(txt.Body) != 4 {
t.Fatalf("body = %d statements, want 4", len(txt.Body))
}
first := txt.Body[0].(*ast.Instr)
if first.Mnemonic.Text != "ROLQ" || len(first.Operands) != 2 {
t.Errorf("first statement = %+v, want ROLQ with two operands", first.Mnemonic)
}
second := txt.Body[1].(*ast.Instr)
if second.Mnemonic.Text != "ROLQ" || len(second.Operands) != 2 {
t.Errorf("second statement = %s, want ROLQ with two operands", second.Mnemonic.Text)
}
// The trailing comment belongs to the second MOVQ, semicolon included.
third := txt.Body[2].(*ast.Instr)
if third.Mnemonic.Text != "MOVQ" || third.Comment != "note; still comment" {
t.Errorf("third = %s, comment %q", third.Mnemonic.Text, third.Comment)
}
}
// TestSemicolonAfterLabel covers a label sharing its line with two
// statements.
func TestSemicolonAfterLabel(t *testing.T) {
file, errs := Parse("t.s", "TEXT \u00b7f(SB), $0\nloop: NOP; NOP\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse errors: %v", errs)
}
txt := file.Decls[0].(*ast.Text)
if len(txt.Body) != 4 {
t.Fatalf("body = %d statements, want 4 (label, two instructions, RET)", len(txt.Body))
}
if _, ok := txt.Body[0].(*ast.Label); !ok {
t.Errorf("first statement = %T, want *ast.Label", txt.Body[0])
}
for i, want := range []string{"NOP", "NOP", "RET"} {
in, ok := txt.Body[i+1].(*ast.Instr)
if !ok || in.Mnemonic.Text != want {
t.Errorf("statement %d = %v, want %s", i+1, txt.Body[i+1], want)
}
}
}
// TestParseEqualsZeroOptions pins the contract that ParseWithOptions with
// the zero Options reproduces Parse, here for the semicolon split.
func TestParseEqualsZeroOptions(t *testing.T) {
src := "TEXT \u00b7f(SB), $0\n\tNOP; NOP\n\tRET\n"
a, errsA := Parse("t.s", src)
b, errsB := ParseWithOptions("t.s", src, Options{})
if len(errsA) > 0 || len(errsB) > 0 {
t.Fatalf("errors: %v / %v", errsA, errsB)
}
ta, tb := texts(a), texts(b)
if len(ta) != len(tb) {
t.Fatalf("decl counts differ: %d vs %d", len(ta), len(tb))
}
for i := range ta {
if len(ta[i].Body) != len(tb[i].Body) {
t.Fatalf("TEXT %d: body lengths differ: %d vs %d", i, len(ta[i].Body), len(tb[i].Body))
}
}
}
// TestBracketRegisterRange pins the amd64 multi-source operand of the
// 4FMAPS/4VNNIW families: the bracket group [Z0-Z3] names four consecutive
// source registers and must reach the AST as a register range instead of an
// empty address.
func TestBracketRegisterRange(t *testing.T) {
file, errs := Parse("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tV4FMADDPS 17(SP), [Z0-Z3], K2, Z0\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse errors: %v", errs)
}
fn := file.Decls[0].(*ast.Text)
in := fn.Body[0].(*ast.Instr)
if len(in.Operands) != 4 {
t.Fatalf("operands = %d, want 4", len(in.Operands))
}
rng := in.Operands[1]
if rng.Kind != ast.OpAddr {
t.Errorf("range operand kind = %v, want OpAddr", rng.Kind)
}
if rng.Addr.Range == nil {
t.Fatalf("range operand = %+v, want a register range", rng.Addr)
}
if rng.Addr.Range.Lo != "Z0" || rng.Addr.Range.Hi != "Z3" {
t.Errorf("range = %s-%s, want Z0-Z3", rng.Addr.Range.Lo, rng.Addr.Range.Hi)
}
if rng.Addr.Sym != nil || rng.Addr.Base != "" || rng.Addr.Index != "" || rng.Addr.Shift != "" {
t.Errorf("range operand carries stray address fields: %+v", rng.Addr)
}
if rng.Raw != "[ Z0 - Z3 ]" {
t.Errorf("range raw = %q, want the verbatim spelling", rng.Raw)
}
}
// TestBracketRegisterRangeNotList pins that arm64-style register lists, whose
// members carry arrangements, stay out of the simple range shape: they remain
// plain bracketed groups the arm64 encoder reads from Raw. A comma inside
// brackets is a top-level comma, so a multi-member list spans several
// operands, exactly the shape the arm64 encoder's list scan stitches back.
func TestBracketRegisterRangeNotList(t *testing.T) {
file, errs := Parse("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tVLD1 (R2), [V21.B16]\n\tVLD1 (R1), [V2.B16, V3.B16]\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse errors: %v", errs)
}
fn := file.Decls[0].(*ast.Text)
for i, want := range []string{"[ V21.B16 ]", "V3.B16 ]"} {
in := fn.Body[i].(*ast.Instr)
op := in.Operands[len(in.Operands)-1]
if op.Addr.Range != nil {
t.Errorf("%s: range = %v, want nil", in.Mnemonic.Text, op.Addr.Range)
}
if op.Raw != want {
t.Errorf("operand %d raw = %q, want %q", i, op.Raw, want)
}
}
}
// TestVSIBIndexOnly pins the gather/scatter memory operand with a scaled
// vector index and no base register: 8(X4*1) must carry index and scale and
// leave the base empty, not strand the scale in the shift suffix.
func TestVSIBIndexOnly(t *testing.T) {
file, errs := Parse("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tVPGATHERDQ Y0, 8(X4*1), Y6\n\tVPGATHERDQ Y0, (X4*2), Y6\n\tVPGATHERDQ Y0, -8(X4*1), Y6\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse errors: %v", errs)
}
fn := file.Decls[0].(*ast.Text)
want := []ast.Address{
{Index: "X4", Scale: 1, Offset: 8, HasOff: true},
{Index: "X4", Scale: 2},
{Index: "X4", Scale: 1, Offset: -8, HasOff: true},
}
for i, w := range want {
in := fn.Body[i].(*ast.Instr)
a := in.Operands[1].Addr
if a.Base != "" || a.Index != w.Index || a.Scale != w.Scale || a.Offset != w.Offset || a.HasOff != w.HasOff || a.Shift != "" {
t.Errorf("operand %d = %+v, want %+v", i, a, w)
}
}
}
// TestVSIBTwoGroupKeepsBase pins that the ordinary (base)(index*scale)
// grammar is untouched by the index-only recognition.
func TestVSIBTwoGroupKeepsBase(t *testing.T) {
file, errs := Parse("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tVP4DPWSSD 7(SI)(DI*1), [Z2-Z5], K4, Z17\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse errors: %v", errs)
}
fn := file.Decls[0].(*ast.Text)
in := fn.Body[0].(*ast.Instr)
a := in.Operands[0].Addr
if a.Base != "SI" || a.Index != "DI" || a.Scale != 1 || a.Offset != 7 || !a.HasOff {
t.Errorf("address = %+v, want base SI index DI scale 1 offset 7", a)
}
if in.Operands[1].Addr.Range == nil || in.Operands[1].Addr.Range.Lo != "Z2" || in.Operands[1].Addr.Range.Hi != "Z5" {
t.Errorf("second operand = %+v, want range Z2-Z5", in.Operands[1].Addr)
}
}
// TestBareTrailingImmediate pins the toolchain's bare constant spelling in
// the final operand slot: CMPSD X1, X0, 1 reads as $1 (math/floor_amd64.s).
// Earlier slots keep the strict grammar, so a bare number there stays an
// address rather than becoming an immediate.
func TestBareTrailingImmediate(t *testing.T) {
file, errs := Parse("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tCMPSD X1, X0, 1\n\tCMPSD X1, X0, -1\n\tADDQ AX, 1+2\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse errors: %v", errs)
}
fn := file.Decls[0].(*ast.Text)
for i, want := range []int64{1, -1, 3} {
in := fn.Body[i].(*ast.Instr)
last := in.Operands[len(in.Operands)-1]
if last.Kind != ast.OpImmediate || !last.Imm.HasVal || last.Imm.Val != want {
t.Errorf("operand %d = %+v, want immediate %d", i, last, want)
}
}
// A bare number outside the final slot is not an immediate.
file2, errs2 := Parse("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tADDQ 1, AX\n\tRET\n")
if len(errs2) > 0 {
t.Fatalf("parse errors: %v", errs2)
}
fn2 := file2.Decls[0].(*ast.Text)
first := fn2.Body[0].(*ast.Instr).Operands[0]
if first.Kind != ast.OpAddr {
t.Errorf("non-final bare number kind = %v, want OpAddr", first.Kind)
}
// A bare name in the final slot stays a symbol: labels are names, not
// constants, and jump targets depend on the distinction.
file3, errs3 := Parse("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tJMP loop\nloop: NOP\n\tRET\n")
if len(errs3) > 0 {
t.Fatalf("parse errors: %v", errs3)
}
fn3 := file3.Decls[0].(*ast.Text)
jmp := fn3.Body[0].(*ast.Instr)
if jmp.Operands[0].Kind != ast.OpAddr || jmp.Operands[0].Addr.Sym == nil || jmp.Operands[0].Addr.Sym.Name != "loop" {
t.Errorf("jump target = %+v, want label loop", jmp.Operands[0])
}
}
+562
View File
@@ -0,0 +1,562 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// The preprocessor turns #define and #include directives into the token
// stream the parser really sees, the way the Go toolchain's assembler does:
// object and parameterised macros expand at the point of use, and an
// #include splices the named file's lines in place of the directive. The
// pass runs only on the assembly path (gasm asm, diff, the corpus audit),
// where the result is machine code; parsing for the linter, formatter and
// language server keeps the raw file so their view of #define lines, and
// therefore their macro-aware behaviour, is unchanged.
package parser
import (
"fmt"
"os"
"path/filepath"
"slices"
"strconv"
"strings"
"unicode/utf8"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
"sourcedock.dev/petrbalvin/gasm-devkit/lexer"
"sourcedock.dev/petrbalvin/gasm-devkit/token"
)
// Options controls the optional preprocessing applied before a file is
// parsed. The zero value reproduces Parse exactly.
type Options struct {
// IncludeDirs lists the -I directories searched for #include files,
// in order, after the including file's own directory.
IncludeDirs []string
// Expand enables macro expansion, include splicing and the
// statement-separator reading of ';' that the expanded bodies rely on.
Expand bool
}
// ParseWithOptions parses src like Parse, optionally preprocessing it first.
// The returned file is usable even when errors is non-empty.
func ParseWithOptions(path, src string, opts Options) (*ast.File, []error) {
tokens := lexer.Tokenize(src)
var lines [][]token.Token
var errs []error
if opts.Expand {
pp := &preproc{opts: opts, macros: map[string]*macroDef{}}
lines = pp.fileLines(path, tokens, token.Position{})
errs = pp.errs
} else {
lines = statementLines(tokens)
}
p := &state{path: path}
p.parse(lines)
return p.file, append(errs, p.errs...)
}
// maxExpansionDepth bounds recursive macro expansion; the toolchain's
// assembler gives up after 100 nested invocations without producing a token.
const maxExpansionDepth = 100
// textflagHeader names the one header gasm does not splice: its flag macros
// (NOSPLIT, RODATA, …) are consumed by name throughout gasm's parser,
// encoders and linter, and expanding them to their numeric constants would
// leave every consumer blind to them.
const textflagHeader = "textflag.h"
// macroDef is one #define. A nil args slice is an object macro; a non-nil
// (possibly empty) one is parameterised, the C distinction between
// "#define A(x)" and "#define A (x)".
type macroDef struct {
name string
args []string
body []token.Token
}
// preproc carries the state of one expansion pass: the live macro table, the
// chain of files currently being read, for cycle detection, and the
// conditional-inclusion stack of #ifdef regions.
type preproc struct {
opts Options
macros map[string]*macroDef
errs []error
stack []string // absolute paths of files being read, innermost last
ifdefStack []bool // one entry per open #ifdef/#ifndef, its truth
}
// enabled reports whether the position being read is inside a live
// conditional branch. Directives inside a disabled branch contribute
// nothing, and its content lines are dropped, exactly as the toolchain's
// input stack does.
func (pp *preproc) enabled() bool {
return len(pp.ifdefStack) == 0 || pp.ifdefStack[len(pp.ifdefStack)-1]
}
func (pp *preproc) errorf(pos token.Position, format string, args ...any) {
pp.errs = append(pp.errs, Error{Pos: pos, Msg: fmt.Sprintf(format, args...)})
}
// fileLines tokenizes and preprocesses one file into logical lines.
// Directive lines are kept (the parser records them for the tooling);
// #include lines are replaced by the included file's lines. includePos is
// the position of the #include that pulled this file in, zero for the
// top-level file, and only serves cycle diagnostics.
func (pp *preproc) fileLines(path string, tokens []token.Token, includePos token.Position) [][]token.Token {
abs, err := filepath.Abs(path)
if err != nil {
abs = filepath.Clean(path)
}
if slices.Contains(pp.stack, abs) {
if includePos.IsValid() {
pp.errorf(includePos, "#include %q: include cycle (%s is already being read)", path, filepath.Base(path))
}
return nil
}
pp.stack = append(pp.stack, abs)
var out [][]token.Token
for _, line := range splitLines(tokens) {
if len(line) == 0 {
out = append(out, line)
continue
}
if line[0].Kind == token.Hash {
out = append(out, pp.directive(line, filepath.Dir(path))...)
continue
}
if !pp.enabled() {
continue
}
out = append(out, splitOnSemicolons(pp.expandTokens(line))...)
}
pp.stack = pp.stack[:len(pp.stack)-1]
if len(pp.stack) == 0 && len(pp.ifdefStack) > 0 {
// The stack is per-input, shared across includes, so only the
// top-level file's end can decide the input was left unclosed.
pp.errorf(token.Position{Line: 1, Column: 1}, "unclosed #ifdef or #ifndef")
}
return out
}
// directive processes one '#' line and returns the lines to keep in the
// stream: every directive line is kept as-is for the parser (which records
// it), except #include, which is replaced by the spliced content.
// Conditionals are tracked on every line; every other directive is inert
// inside a disabled branch.
func (pp *preproc) directive(line []token.Token, dir string) [][]token.Token {
if len(line) < 2 || line[1].Kind != token.Ident {
return [][]token.Token{line}
}
switch line[1].Text {
case "ifdef", "ifndef":
pp.ifdef(line, line[1].Text == "ifndef")
case "else":
pp.elseBranch(line)
case "endif":
pp.endif(line)
case "define":
if pp.enabled() {
pp.define(line)
}
case "undef":
if pp.enabled() {
pp.undef(line)
}
case "include":
if pp.enabled() {
return pp.include(line, dir)
}
default:
// #line and unknown directives are recorded but not interpreted:
// conservative support keeps the parser's view intact and files
// using them fail on their content, not silently.
}
return [][]token.Token{line}
}
// ifdef handles "#ifdef NAME" and "#ifndef NAME", pushing the branch's truth
// onto the conditional stack. A branch opened inside a disabled region is
// itself disabled, however the name resolves.
func (pp *preproc) ifdef(line []token.Token, inverted bool) {
truth := false
if len(line) >= 3 && line[2].Kind == token.Ident {
_, defined := pp.macros[line[2].Text]
truth = defined != inverted
} else {
pp.errorf(line[0].Pos, "expected identifier after #%s", line[1].Text)
}
if !pp.enabled() {
truth = false
}
pp.ifdefStack = append(pp.ifdefStack, truth)
}
// elseBranch flips the innermost conditional's truth, but only when the
// region enclosing it is itself live: the toolchain keeps outer overrides.
func (pp *preproc) elseBranch(line []token.Token) {
if len(pp.ifdefStack) == 0 {
pp.errorf(line[0].Pos, "unmatched #else")
return
}
if len(pp.ifdefStack) == 1 || pp.ifdefStack[len(pp.ifdefStack)-2] {
pp.ifdefStack[len(pp.ifdefStack)-1] = !pp.ifdefStack[len(pp.ifdefStack)-1]
}
}
// endif closes the innermost conditional.
func (pp *preproc) endif(line []token.Token) {
if len(pp.ifdefStack) == 0 {
pp.errorf(line[0].Pos, "unmatched #endif")
return
}
pp.ifdefStack = pp.ifdefStack[:len(pp.ifdefStack)-1]
}
// define parses "#define NAME[(formals)] body" into the macro table. The
// body runs to the end of the logical line (the lexer has already spliced
// backslash continuations) and stops at a comment, which never expands.
func (pp *preproc) define(line []token.Token) {
if len(line) < 3 || line[2].Kind != token.Ident {
return
}
name := line[2]
args := []string(nil)
body := line[3:]
// The definition is parameterised only when '(' follows the name
// directly; the toolchain separates "#define A(x)" from
// "#define A (x)" by adjacency, and so does the column check here.
if len(body) > 0 && body[0].Kind == token.LParen &&
body[0].Pos.Column == name.Pos.Column+utf8.RuneCountInString(name.Text) {
args = []string{}
i := 1
for i < len(body) && body[i].Kind != token.RParen {
if body[i].Kind == token.Ident {
args = append(args, body[i].Text)
}
i++
}
if i < len(body) {
body = body[i+1:]
} else {
body = nil
}
}
if i := slices.IndexFunc(body, func(t token.Token) bool { return t.Kind == token.Comment }); i >= 0 {
body = body[:i]
}
if _, exists := pp.macros[name.Text]; exists {
// The toolchain refuses redefinition, so a file the oracle accepts
// never redefines; failing here keeps that contract visible.
pp.errorf(name.Pos, "redefinition of macro %s", name.Text)
}
pp.macros[name.Text] = &macroDef{name: name.Text, args: args, body: pp.bodyWithBreaks(body)}
}
// bodyWithBreaks records the statement boundaries the continuations carry.
// The lexer splices backslash-continued lines into one logical line, but the
// toolchain keeps the newline as a token in the stored body, which is how a
// multi-instruction body without semicolons (the arm64 style) still splits
// into statements on expansion. A line change inside the logical line is
// exactly a continuation, so the boundary is restored from the positions.
func (pp *preproc) bodyWithBreaks(body []token.Token) []token.Token {
out := make([]token.Token, 0, len(body))
for i, t := range body {
if i > 0 && t.Pos.Line != body[i-1].Pos.Line {
out = append(out, token.Token{Kind: token.Newline, Text: "\n", Pos: t.Pos, End: t.Pos})
}
out = append(out, t)
}
return out
}
// undef handles "#undef NAME", which the toolchain honours and requires to
// name a defined macro.
func (pp *preproc) undef(line []token.Token) {
if len(line) < 3 || line[2].Kind != token.Ident {
return
}
if _, ok := pp.macros[line[2].Text]; !ok {
pp.errorf(line[2].Pos, "#undef for undefined macro %s", line[2].Text)
return
}
delete(pp.macros, line[2].Text)
}
// include resolves and splices "#include \"file\"". A header that cannot be
// read keeps the directive line in the stream, with a diagnostic.
func (pp *preproc) include(line []token.Token, dir string) [][]token.Token {
if len(line) < 3 || line[2].Kind != token.String {
return [][]token.Token{line}
}
header := line[2]
name, err := strconv.Unquote(header.Text)
if err != nil {
pp.errorf(header.Pos, "unquoting include file name: %v", err)
return [][]token.Token{line}
}
if filepath.Base(name) == textflagHeader {
// Flag macros are handled natively (see textflagHeader); the
// directive stays so tools still see the include.
return [][]token.Token{line}
}
resolved, ok := pp.resolve(name, dir)
if !ok {
searched := append([]string{dir}, pp.opts.IncludeDirs...)
pp.errorf(header.Pos, "#include %q: file not found (searched %s)", name, strings.Join(searched, ", "))
return [][]token.Token{line}
}
src, err := os.ReadFile(resolved)
if err != nil {
pp.errorf(header.Pos, "#include %q: %v", name, err)
return [][]token.Token{line}
}
return pp.fileLines(resolved, lexer.Tokenize(string(src)), header.Pos)
}
// resolve looks an include name up the way the toolchain does: as written
// (relative to the working directory), then relative to the including
// file's directory, then in each -I directory in order.
func (pp *preproc) resolve(name, dir string) (string, bool) {
candidates := []string{name}
if !filepath.IsAbs(name) {
candidates = append(candidates, filepath.Join(dir, name))
for _, d := range pp.opts.IncludeDirs {
candidates = append(candidates, filepath.Join(d, name))
}
}
for _, c := range candidates {
if st, err := os.Stat(c); err == nil && !st.IsDir() {
return c, true
}
}
return "", false
}
// expandTokens expands every macro invocation in a token sequence,
// recursively, with a depth guard. A body is spliced into the sequence in
// place and rescanned, the way the toolchain's input stack re-reads pushed
// tokens: an object macro may name a parameterised one, and the argument
// list of the expansion may then come from the tokens that follow.
func (pp *preproc) expandTokens(in []token.Token) []token.Token {
s := in
i := 0
consecutive := 0
for i < len(s) {
t := s[i]
if t.Kind != token.Ident {
i++
consecutive = 0
continue
}
def, suffix := pp.macroFor(t.Text)
if def == nil {
i++
consecutive = 0
continue
}
// The guard mirrors the toolchain's: 100 nested invocations in a
// row without a plain token between them means recursion.
consecutive++
if consecutive > maxExpansionDepth {
pp.errorf(t.Pos, "recursive macro invocation (deeper than %d levels)", maxExpansionDepth)
return nil
}
if def.args == nil {
body := restamp(def.body, t.Pos)
if suffix != "" {
// The macro was reached only through a compound spelling
// (ACC0.B16 over "#define ACC0 V8"), so the selector has
// to travel with the expansion.
body = appendSelector(body, suffix, t.Pos)
}
s = append(s[:i], append(body, s[i+1:]...)...)
continue
}
// A parameterised macro invoked without its parentheses stands
// unexpanded, naming itself, as in the toolchain.
if i+1 >= len(s) || s[i+1].Kind != token.LParen {
i++
consecutive = 0
continue
}
args, next := pp.collectArgs(s, i+1, t)
if args == nil {
return nil
}
// A zero-argument macro may be invoked as NAME().
if len(def.args) == 0 && len(args) == 1 && len(args[0]) == 0 {
args = nil
}
if len(args) != len(def.args) {
pp.errorf(t.Pos, "wrong arg count for macro %s: got %d, want %d", t.Text, len(args), len(def.args))
i = next
consecutive = 0
continue
}
sub := make([]token.Token, 0, len(def.body))
for _, bt := range def.body {
if bt.Kind == token.Ident {
if k := slices.Index(def.args, bt.Text); k >= 0 {
sub = append(sub, restamp(args[k], t.Pos)...)
continue
}
// A parameter used with an element or lane selector: the
// lexer folds A.S4 into one identifier, so the whole-token
// match above cannot see the parameter. The toolchain
// lexes the period separately and substitutes the name
// alone; splitting at the FIRST period and pasting the
// argument back in front of the selector is the equivalent
// for this lexer.
if k, sel := parameterSelector(bt.Text, def.args); k >= 0 {
sub = append(sub, restamp(pasteSelector(args[k], sel), t.Pos)...)
continue
}
}
sub = append(sub, bt)
}
s = append(s[:i], append(sub, s[next:]...)...)
}
return s
}
// macroFor finds the macro a use names. The lexer folds NAME.selector into
// one identifier token, so a macro written behind a selector suffix
// (ACC0.B16 over "#define ACC0 V8") never matches a whole-token table
// lookup; the toolchain splits on the period and reads the two halves, so
// the prefix before the FIRST period is tried here as well and the caller
// re-attaches the suffix to whatever the macro expands to. Only a whole
// name counts: AB.S4 does not reach a macro named A, and a parameterised
// macro is not hidden behind a selector, because its invocation would need
// the parentheses to follow the bare name.
func (pp *preproc) macroFor(text string) (*macroDef, string) {
if def := pp.macros[text]; def != nil {
return def, ""
}
if j := strings.IndexByte(text, '.'); j > 0 {
if def := pp.macros[text[:j]]; def != nil && def.args == nil {
return def, text[j:]
}
}
return nil, ""
}
// appendSelector glues a selector suffix onto an object macro's expansion:
// the selector binds to the identifier the expansion ends with, the way the
// toolchain's operand parser reads V0 and .B16 back as one register
// spelling. An expansion that does not end in an identifier carries the
// selector as its own token, which the parser then reports where it cannot
// parse it.
func appendSelector(body []token.Token, suffix string, pos token.Position) []token.Token {
if n := len(body); n > 0 && body[n-1].Kind == token.Ident {
body[n-1].Text += suffix
return body
}
return append(body, token.Token{Kind: token.Ident, Text: suffix, Pos: pos, End: pos})
}
// parameterSelector reports the argument a compound body token names: the
// parameter whose whole name occupies the text before the token's FIRST
// period, with the selector that follows. k is negative when no parameter
// matches, which leaves tokens like AB.S4 untouched even though a parameter
// A is bound.
func parameterSelector(text string, args []string) (int, string) {
j := strings.IndexByte(text, '.')
if j <= 0 {
return -1, ""
}
if k := slices.Index(args, text[:j]); k >= 0 {
return k, text[j:]
}
return -1, ""
}
// pasteSelector joins an argument with the selector a compound body token
// carries, textually: the selector binds to the identifier the argument
// ends with, so A.S4 over the argument V0.B16 spells V0.B16.S4, exactly the
// operand the toolchain's split-then-substitute leaves behind. An argument
// with no trailing identifier carries the selector as a separate token,
// which the parser then reports where it cannot parse it.
func pasteSelector(val []token.Token, suffix string) []token.Token {
if len(val) == 0 {
return []token.Token{{Kind: token.Ident, Text: suffix}}
}
out := slices.Clone(val)
if n := len(out); out[n-1].Kind == token.Ident {
out[n-1].Text += suffix
return out
}
return append(out, token.Token{Kind: token.Ident, Text: suffix})
}
// collectArgs reads the actual argument tokens of an invocation; the opening
// parenthesis is at start. Commas separate arguments except inside nested
// parentheses. A nil result means the list was unterminated, which is a
// diagnostic.
func (pp *preproc) collectArgs(in []token.Token, start int, name token.Token) ([][]token.Token, int) {
var args [][]token.Token
var cur []token.Token
nesting := 0
for i := start + 1; i < len(in); i++ {
t := in[i]
switch t.Kind {
case token.LParen:
nesting++
cur = append(cur, t)
case token.RParen:
if nesting == 0 {
return append(args, cur), i + 1
}
nesting--
cur = append(cur, t)
case token.Comma:
if nesting == 0 {
args = append(args, cur)
cur = nil
continue
}
cur = append(cur, t)
case token.Comment:
pp.errorf(name.Pos, "unterminated arg list invoking macro %s", name.Text)
return nil, i
default:
cur = append(cur, t)
}
}
pp.errorf(name.Pos, "unterminated arg list invoking macro %s", name.Text)
return nil, len(in)
}
// restamp copies body tokens to the invocation's position, so diagnostics
// and the line table point where the macro was used, as the toolchain's
// input stack does.
func restamp(body []token.Token, pos token.Position) []token.Token {
out := make([]token.Token, len(body))
for i, t := range body {
t.Pos, t.End = pos, pos
out[i] = t
}
return out
}
// splitOnSemicolons breaks a token sequence at ';' statement separators and
// at the Newline markers that record continuation boundaries inside macro
// bodies, producing the logical lines the parser expects. The separators
// carry no meaning beyond the break, so the pieces are exactly what the same
// statements on separate lines would produce.
func splitOnSemicolons(ts []token.Token) [][]token.Token {
var out [][]token.Token
start := 0
for i, t := range ts {
if t.Kind == token.Semicolon || t.Kind == token.Newline {
if i > start {
out = append(out, ts[start:i])
}
start = i + 1
}
}
if start < len(ts) {
out = append(out, ts[start:])
}
return out
}
+675
View File
@@ -0,0 +1,675 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
package parser
import (
"os"
"path/filepath"
"strings"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
)
// expand parses src with preprocessing enabled and returns the first TEXT's
// body instructions as "MNEMONIC operand|operand" strings, the shape the
// expansion assertions below compare against. Runs of spaces are
// collapsed: Raw renders a token group as its tokens joined with single
// spaces, so "$(32-7)" arrives as "$ ( 32 - 7 )" and the comparison must
// not depend on that spelling.
func expand(t *testing.T, src string) (*ast.File, []string) {
t.Helper()
f, errs := ParseWithOptions("t_amd64.s", src, Options{Expand: true})
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
ts := texts(f)
if len(ts) == 0 {
t.Fatalf("no TEXT in:\n%s", src)
}
var got []string
for _, s := range ts[0].Body {
in, ok := s.(*ast.Instr)
if !ok {
continue
}
var ops []string
for _, op := range in.Operands {
ops = append(ops, op.Raw)
}
line := in.Mnemonic.Text + " " + strings.Join(ops, ", ")
got = append(got, strings.ReplaceAll(line, " ", ""))
}
return f, got
}
func wantLines(t *testing.T, got []string, want ...string) {
t.Helper()
strip := func(lines []string) string {
var out []string
for _, l := range lines {
out = append(out, strings.ReplaceAll(l, " ", ""))
}
return strings.Join(out, "\n")
}
if strip(got) != strip(want) {
t.Errorf("expanded body:\n %s\nwant:\n %s", strings.Join(got, "\n "), strings.Join(want, "\n "))
}
}
func TestObjectMacroExpandsAtUse(t *testing.T) {
_, got := expand(t, `
#define REGTMP CX
#define TWICE ADDQ CX, AX; ADDQ CX, AX
TEXT ·f(SB), NOSPLIT, $0
MOVQ 8(SP), REGTMP
TWICE
RET
`)
wantLines(t, got,
"MOVQ 8(SP), CX",
"ADDQ CX, AX",
"ADDQ CX, AX",
"RET",
)
}
func TestParameterisedMacroSubstitutesArguments(t *testing.T) {
f, errs := ParseWithOptions("t_amd64.s", `
#define ROUND1(a, index, const, shift) \
ADDQ $const, a; \
MOVW (index*4)(SP), a; \
RORQ $(32-shift), a
TEXT ·f(SB), NOSPLIT, $0
ROUND1(AX, 3, 0xd76aa478, 7)
RET
`, Options{Expand: true})
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
body := texts(f)[0].Body
add := body[0].(*ast.Instr)
if add.Mnemonic.Text != "ADDQ" || !add.Operands[0].Imm.HasVal ||
add.Operands[0].Imm.Val != 0xd76aa478 || add.Operands[1].Addr.Sym == nil ||
add.Operands[1].Addr.Sym.Name != "AX" {
t.Errorf("ADDQ operands substituted wrong: %+v %+v", add.Operands[0].Imm, add.Operands[1].Addr)
}
mov := body[1].(*ast.Instr)
if addr := mov.Operands[0].Addr; !addr.HasOff || addr.Offset != 12 {
t.Errorf("MOVW offset = %+v, want 12 from 3*4", addr)
}
ror := body[2].(*ast.Instr)
if !ror.Operands[0].Imm.HasVal || ror.Operands[0].Imm.Val != 25 {
t.Errorf("RORQ immediate = %+v, want 25 from (32-7)", ror.Operands[0].Imm)
}
}
func TestMacroArgumentsKeepCommasInParens(t *testing.T) {
// An argument may itself be an unparenthesised expression: the tokens
// substitute verbatim and the parser folds the result, as the
// toolchain's parser does.
f, errs := ParseWithOptions("t_amd64.s", `
#define LOAD(dst, off) MOVQ off(SP), dst
TEXT ·f(SB), NOSPLIT, $0
LOAD(AX, 1*8)
RET
`, Options{Expand: true})
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
in := texts(f)[0].Body[0].(*ast.Instr)
addr := in.Operands[0].Addr
if !addr.HasOff || addr.Offset != 8 {
t.Errorf("offset = %+v, want 8", addr)
}
if sym := in.Operands[1].Addr.Sym; sym == nil || sym.Name != "AX" {
t.Errorf("destination = %+v, want AX", in.Operands[1].Addr)
}
}
func TestNestedMacroInvocations(t *testing.T) {
// An object macro naming a parameterised one, and a parameterised body
// invoking another parameterised macro: the toolchain's input stack
// rescans substituted tokens, and so does expansion here.
_, got := expand(t, `
#define DOUBLE(x) ADDQ x, x
#define TWICE2 DOUBLE
#define FOUR(a, b) DOUBLE(a); DOUBLE(b)
TEXT ·f(SB), NOSPLIT, $0
TWICE2(AX)
FOUR(AX, CX)
RET
`)
wantLines(t, got,
"ADDQ AX, AX",
"ADDQ AX, AX",
"ADDQ CX, CX",
"RET",
)
}
func TestMultiLineBodySplitsWithoutSemicolons(t *testing.T) {
// The arm64 style: backslash-continued lines with no semicolons. The
// continuation newline is a statement boundary, as in the toolchain.
_, got := expand(t, `
#define PAIR \
ADDQ AX, AX \
MOVQ AX, CX
TEXT ·f(SB), NOSPLIT, $0
PAIR
RET
`)
wantLines(t, got,
"ADDQ AX, AX",
"MOVQ AX, CX",
"RET",
)
}
func TestZeroArgumentMacro(t *testing.T) {
_, got := expand(t, `
#define BARRIER()
TEXT ·f(SB), NOSPLIT, $0
BARRIER()
RET
`)
wantLines(t, got, "RET")
}
func TestParameterisedWithoutParensStandsAsName(t *testing.T) {
// A parameterised macro invoked without its parentheses names itself,
// which the parser then reports as an unknown instruction rather than
// silently expanding nothing.
f, errs := ParseWithOptions("t_amd64.s", `
#define M(x) ADDQ x, x
TEXT ·f(SB), NOSPLIT, $0
M
RET
`, Options{Expand: true})
if len(errs) != 0 {
t.Fatalf("parse: %v", errs)
}
fn := texts(f)[0]
if len(fn.Body) == 0 {
t.Fatal("body empty")
}
in, ok := fn.Body[0].(*ast.Instr)
if !ok || in.Mnemonic.Text != "M" {
t.Fatalf("bare parameterised macro did not stand as its name: %+v", fn.Body[0])
}
}
func TestDefinitionScoping(t *testing.T) {
// A definition applies from its point onward: the use before the
// #define stays untouched.
_, got := expand(t, `
TEXT ·f(SB), NOSPLIT, $0
SPECIAL
#define SPECIAL ADDQ AX, AX
SPECIAL
RET
`)
wantLines(t, got,
"SPECIAL",
"ADDQ AX, AX",
"RET",
)
}
func TestUndefRemovesMacro(t *testing.T) {
_, got := expand(t, `
#define TEMP AX
TEXT ·f(SB), NOSPLIT, $0
TEMP
#undef TEMP
TEMP
RET
`)
wantLines(t, got,
"AX",
"TEMP",
"RET",
)
}
func TestUndefUndefinedMacroIsAnError(t *testing.T) {
_, errs := ParseWithOptions("t_amd64.s", "#undef NOSUCH\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n", Options{Expand: true})
if len(errs) == 0 || !strings.Contains(errs[0].Error(), "undefined macro NOSUCH") {
t.Fatalf("#undef of an undefined macro: got %v, want an error naming it", errs)
}
}
func TestRedefinitionIsAnError(t *testing.T) {
_, errs := ParseWithOptions("t_amd64.s", "#define A X\n#define A Y\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n", Options{Expand: true})
if len(errs) == 0 || !strings.Contains(errs[0].Error(), "redefinition of macro A") {
t.Fatalf("redefinition: got %v, want an error", errs)
}
}
func TestRecursiveMacroIsAnError(t *testing.T) {
_, errs := ParseWithOptions("t_amd64.s", "#define A B\n#define B A\nTEXT ·f(SB), NOSPLIT, $0\n\tA\n\tRET\n", Options{Expand: true})
if len(errs) == 0 || !strings.Contains(errs[0].Error(), "recursive macro invocation") {
t.Fatalf("recursion: got %v, want a recursive-macro error, not a hang", errs)
}
}
func TestWrongArgumentCountIsAnError(t *testing.T) {
_, errs := ParseWithOptions("t_amd64.s", "#define M(a, b) ADDQ a, b\nTEXT ·f(SB), NOSPLIT, $0\n\tM(AX)\n\tRET\n", Options{Expand: true})
if len(errs) == 0 || !strings.Contains(errs[0].Error(), "wrong arg count for macro M") {
t.Fatalf("arg count: got %v, want an error", errs)
}
}
func TestConditionalsSelectOneBranch(t *testing.T) {
_, got := expand(t, `
#define MODE2
TEXT ·f(SB), NOSPLIT, $0
#ifdef MODE2
ADDQ AX, AX
#else
SUBQ AX, AX
#endif
#ifndef MODE2
SUBQ CX, CX
#else
ADDQ CX, CX
#endif
RET
`)
wantLines(t, got,
"ADDQ AX, AX",
"ADDQ CX, CX",
"RET",
)
}
func TestConditionalsHideDefinitionsAndIncludes(t *testing.T) {
// A definition inside a disabled branch must not exist, and an
// unresolvable include there must not be followed.
_, got := expand(t, `
TEXT ·f(SB), NOSPLIT, $0
#ifdef NOTDEFINED
#define HIDEN ADDQ AX, AX
#include "nowhere.h"
#endif
HIDEN
RET
`)
wantLines(t, got, "HIDEN", "RET")
}
func TestUnclosedConditionalIsAnError(t *testing.T) {
_, errs := ParseWithOptions("t_amd64.s", "#ifdef X\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n", Options{Expand: true})
if len(errs) == 0 || !strings.Contains(errs[0].Error(), "unclosed #ifdef") {
t.Fatalf("unclosed conditional: got %v, want an error", errs)
}
}
func TestUnmatchedConditionalDelimitersAreErrors(t *testing.T) {
_, errs := ParseWithOptions("t_amd64.s", "#endif\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n", Options{Expand: true})
if len(errs) == 0 || !strings.Contains(errs[0].Error(), "unmatched #endif") {
t.Fatalf("unmatched #endif: got %v, want an error", errs)
}
_, errs = ParseWithOptions("t_amd64.s", "#else\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n", Options{Expand: true})
if len(errs) == 0 || !strings.Contains(errs[0].Error(), "unmatched #else") {
t.Fatalf("unmatched #else: got %v, want an error", errs)
}
}
// includeTree writes a directory of include files and returns its path.
func includeTree(t *testing.T, files map[string]string) string {
t.Helper()
dir := t.TempDir()
for name, content := range files {
path := filepath.Join(dir, name)
if err := os.MkdirAll(filepath.Dir(path), 0o755); err != nil {
t.Fatal(err)
}
if err := os.WriteFile(path, []byte(content), 0o644); err != nil {
t.Fatal(err)
}
}
return dir
}
func TestIncludeSplicesAndDefinesAreShared(t *testing.T) {
dir := includeTree(t, map[string]string{
"consts.h": "#define KONST $42\n",
})
f, errs := ParseWithOptions("t_amd64.s", `
#include "consts.h"
TEXT ·f(SB), NOSPLIT, $0
MOVQ KONST, AX
RET
`, Options{Expand: true, IncludeDirs: []string{dir}})
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
in := texts(f)[0].Body[0].(*ast.Instr)
if in.Mnemonic.Text != "MOVQ" || strings.ReplaceAll(in.Operands[0].Raw, " ", "") != "$42" {
t.Fatalf("include splicing failed: %+v", in)
}
}
func TestIncludeResolutionOrder(t *testing.T) {
// The including file's directory wins over the -I list, and the -I list
// is searched in order.
src := includeTree(t, map[string]string{
"inc/main.s": "#include \"which.h\"\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n",
"inc/which.h": "#define WHO ONE\n",
"first/which.h": "#define WHO TWO\n",
"second/which.h": "#define WHO THREE\n",
})
main := filepath.Join(src, "inc", "main.s")
body, err := os.ReadFile(main)
if err != nil {
t.Fatal(err)
}
// The header exists in the including file's directory and in two -I
// directories; the source-directory copy must win.
f, errs := ParseWithOptions(main, string(body), Options{Expand: true, IncludeDirs: []string{
filepath.Join(src, "first"), filepath.Join(src, "second"),
}})
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
found := false
for _, d := range f.Decls {
if pp, ok := d.(*ast.Preproc); ok && strings.Contains(pp.Raw, "define WHO ONE") {
found = true
}
}
if !found {
t.Error("the including file's directory did not win include resolution")
}
}
func TestIncludeSearchesIncludeDirsInOrder(t *testing.T) {
src := includeTree(t, map[string]string{
"inc/main.s": "#include \"which.h\"\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n",
"first/which.h": "#define WHO TWO\n",
"second/which.h": "#define WHO THREE\n",
})
main := filepath.Join(src, "inc", "main.s")
body, err := os.ReadFile(main)
if err != nil {
t.Fatal(err)
}
f, errs := ParseWithOptions(main, string(body), Options{Expand: true, IncludeDirs: []string{
filepath.Join(src, "first"), filepath.Join(src, "second"),
}})
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
for _, d := range f.Decls {
if pp, ok := d.(*ast.Preproc); ok && strings.Contains(pp.Raw, "define WHO THREE") {
t.Error("the second -I directory was searched before the first")
}
}
}
func TestIncludeCycleIsDetected(t *testing.T) {
src := includeTree(t, map[string]string{
"a.s": "#include \"b.s\"\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n",
"b.s": "#include \"a.s\"\n",
})
_, errs := ParseWithOptions(filepath.Join(src, "a.s"), "#include \"b.s\"\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n",
Options{Expand: true})
if len(errs) == 0 || !strings.Contains(errs[0].Error(), "include cycle") {
t.Fatalf("include cycle: got %v, want a cycle diagnostic, not a hang", errs)
}
}
func TestUnresolvableIncludeIsAnError(t *testing.T) {
_, errs := ParseWithOptions("t_amd64.s", "#include \"nothere.h\"\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n",
Options{Expand: true, IncludeDirs: []string{t.TempDir()}})
if len(errs) == 0 || !strings.Contains(errs[0].Error(), `#include "nothere.h"`) {
t.Fatalf("missing include: got %v, want a clear diagnostic", errs)
}
}
func TestTextflagHeaderIsNeverSpliced(t *testing.T) {
// textflag.h resolves nowhere here, yet the file must parse: the flag
// names are consumed natively and the include stays in the tree.
f, errs := ParseWithOptions("t_amd64.s", `
#include "textflag.h"
TEXT ·f(SB), NOSPLIT, $0
RET
`, Options{Expand: true})
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
hasInclude := false
for _, d := range f.Decls {
if _, ok := d.(*ast.Include); ok {
hasInclude = true
}
}
if !hasInclude {
t.Error("textflag.h include was dropped from the tree")
}
}
func TestSemicolonSplitsRawLinesToo(t *testing.T) {
_, got := expand(t, `
TEXT ·f(SB), NOSPLIT, $0
BYTE $0x0f; BYTE $0x1f
RET
`)
wantLines(t, got, "BYTE $0x0f", "BYTE $0x1f", "RET")
}
func TestParseUnchangedWithoutExpand(t *testing.T) {
// Without Expand the preprocessor must not exist: a macro invocation
// stays an unexpanded instruction line. The ';' statement separator is
// not part of the preprocessor: the plain parse path splits on it the
// same way the expansion path does, so both spellings agree.
f, errs := Parse("t_amd64.s", `
#define TWICE ADDQ AX, AX
TEXT ·f(SB), NOSPLIT, $0
TWICE
BYTE $0x0f; BYTE $0x1f
RET
`)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
fn := texts(f)[0]
var mnemonics []string
for _, s := range fn.Body {
if in, ok := s.(*ast.Instr); ok {
mnemonics = append(mnemonics, in.Mnemonic.Text)
}
}
if strings.Join(mnemonics, " ") != "TWICE BYTE BYTE RET" {
t.Errorf("non-expanding parse changed: %v", mnemonics)
}
}
func TestConstantExpressionFolding(t *testing.T) {
// The shapes substituted macro bodies leave behind: parenthesised
// arithmetic in immediates and displacements, tilde complements. The
// assertions read the semantic fields; Raw keeps the operand's tokens
// in the canonicalised rendering, not the folded values.
f, errs := ParseWithOptions("t_amd64.s", `
TEXT ·f(SB), NOSPLIT, $0
RORQ $(32-7), AX
ANDQ $~63, AX
MOVQ ((2*4)+0)(SP), AX
MOVQ $((1<<3)|(1<<1)), AX
RET
`, Options{Expand: true})
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
body := texts(f)[0].Body
ror := body[0].(*ast.Instr)
if !ror.Operands[0].Imm.HasVal || ror.Operands[0].Imm.Val != 25 {
t.Errorf("RORQ immediate = %+v, want 25", ror.Operands[0].Imm)
}
and := body[1].(*ast.Instr)
if !and.Operands[0].Imm.HasVal || and.Operands[0].Imm.Val != -64 {
t.Errorf("ANDQ immediate = %+v, want -64", and.Operands[0].Imm)
}
mov := body[2].(*ast.Instr)
addr := mov.Operands[0].Addr
if !addr.HasOff || addr.Offset != 8 || addr.Base != "SP" {
t.Errorf("MOVQ address = %+v, want 8(SP)", addr)
}
mov2 := body[3].(*ast.Instr)
if !mov2.Operands[0].Imm.HasVal || mov2.Operands[0].Imm.Val != 10 {
t.Errorf("MOVQ immediate = %+v, want 10", mov2.Operands[0].Imm)
}
}
func TestConstantExpressionFoldsWithoutExpand(t *testing.T) {
// Folding is a parser capability, not a preprocessing one: a
// hand-written $(32-7) folds the same way with expansion off.
f, errs := ParseWithOptions("t_amd64.s", "TEXT ·f(SB), NOSPLIT, $0\n\tRORQ $(32-7), AX\n\tRET\n", Options{})
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
in := texts(f)[0].Body[0].(*ast.Instr)
if !in.Operands[0].Imm.HasVal || in.Operands[0].Imm.Val != 25 {
t.Errorf("Imm = %+v, want 25", in.Operands[0].Imm)
}
}
func TestParameterWithSelectorSubstitutes(t *testing.T) {
// The lexer folds A.S4 into one identifier token, so a parameter used
// with an element or lane selector never matched the whole-token
// substitution; the toolchain's lexer splits on the period and its
// substitution sees the name alone. Several parameters carry selectors
// in one body here, which is the chacha8_arm64.s QR shape in miniature.
_, got := expand(t, `
#define QR(A, B, C, D) VADD A.S4, B.S4, C.S4; VEOR D.B16, A.B16, D.B16
TEXT ·f(SB), NOSPLIT, $0
QR(V0, V1, V2, V3)
RET
`)
wantLines(t, got,
"VADD V0.S4, V1.S4, V2.S4",
"VEOR V3.B16, V0.B16, V3.B16",
"RET",
)
}
func TestSelectorWithCompoundArgumentPastesTextually(t *testing.T) {
// An argument that is itself one compound identifier pastes verbatim:
// A.S4 over V0.B16 spells V0.B16.S4, the operand the toolchain's
// split-then-substitute leaves behind.
_, got := expand(t, `
#define M(A) VADD A.S4, A.S4, A.S4
TEXT ·f(SB), NOSPLIT, $0
M(V0.B16)
RET
`)
wantLines(t, got, "VADD V0.B16.S4, V0.B16.S4, V0.B16.S4", "RET")
}
func TestSelectorAlongsideBareParameter(t *testing.T) {
// A body may use the parameter bare and suffixed, and the argument may
// itself end in a selector; neither disturbs the other.
_, got := expand(t, `
#define M(A) VADD A, A.S4, A
TEXT ·f(SB), NOSPLIT, $0
M(V0)
M(V1.B16)
RET
`)
wantLines(t, got,
"VADD V0, V0.S4, V0",
"VADD V1.B16, V1.B16.S4, V1.B16",
"RET",
)
}
func TestSelectorKeepsNonParameterPrefixes(t *testing.T) {
// The prefix before the period must be the whole parameter name:
// AB.S4 never reaches a parameter A.
_, got := expand(t, `
#define M(A) VADD AB.S4, A.S4, AB.S4
TEXT ·f(SB), NOSPLIT, $0
M(V0)
RET
`)
wantLines(t, got, "VADD AB.S4, V0.S4, AB.S4", "RET")
}
func TestSelectorExpandsMacroValuedArgument(t *testing.T) {
// gcm_arm64.s invokes mulRound(B1) where B1 is itself an object macro:
// the paste stays rescannable, so B1.D1 still expands to V1.D1 the way
// the toolchain's rescan of substituted tokens does.
_, got := expand(t, `
#define B1 V1
#define mulRound(X) VPMULL X.D1, T1.D1, T3.Q1
TEXT ·f(SB), NOSPLIT, $0
mulRound(B1)
RET
`)
wantLines(t, got, "VPMULL V1.D1, T1.D1, T3.Q1", "RET")
}
func TestObjectMacroBehindSelectorExpands(t *testing.T) {
// Ordinary code writes ACC0.B16 where ACC0 is an object macro; the
// toolchain expands the alias because its lexer reads the selector as
// its own token, and the lookup here must reach the macro through the
// compound spelling the same way.
_, got := expand(t, `
#define ACC0 V8
TEXT ·f(SB), NOSPLIT, $0
VEOR ACC0.B16, ACC0.B16, ACC0.B16
RET
`)
wantLines(t, got, "VEOR V8.B16, V8.B16, V8.B16", "RET")
}
func TestChacha8QRMacroExpands(t *testing.T) {
// The real QR round of chacha8_arm64.s end to end: every parameter
// carries a selector somewhere, and the round is sixteen instructions.
_, got := expand(t, `
#define QR(A, B, C, D) \
VADD A.S4, B.S4, A.S4; VEOR D.B16, A.B16, D.B16; VREV32 D.H8, D.H8; \
VADD C.S4, D.S4, C.S4; VEOR B.B16, C.B16, V30.B16; VSHL $12, V30.S4, B.S4; VSRI $20, V30.S4, B.S4; \
VADD A.S4, B.S4, A.S4; VEOR D.B16, A.B16, D.B16; VTBL V31.B16, [D.B16], D.B16; \
VADD C.S4, D.S4, C.S4; VEOR B.B16, C.B16, V30.B16; VSHL $7, V30.S4, B.S4; VSRI $25, V30.S4, B.S4
TEXT ·f(SB), NOSPLIT, $0
QR(V0, V1, V2, V3)
RET
`)
wantLines(t, got,
"VADD V0.S4, V1.S4, V0.S4",
"VEOR V3.B16, V0.B16, V3.B16",
"VREV32 V3.H8, V3.H8",
"VADD V2.S4, V3.S4, V2.S4",
"VEOR V1.B16, V2.B16, V30.B16",
"VSHL $12, V30.S4, V1.S4",
"VSRI $20, V30.S4, V1.S4",
"VADD V0.S4, V1.S4, V0.S4",
"VEOR V3.B16, V0.B16, V3.B16",
"VTBL V31.B16, [V3.B16], V3.B16",
"VADD V2.S4, V3.S4, V2.S4",
"VEOR V1.B16, V2.B16, V30.B16",
"VSHL $7, V30.S4, V1.S4",
"VSRI $25, V30.S4, V1.S4",
"RET",
)
}
func TestNotAnExpressionFallsBack(t *testing.T) {
// Symbol immediates and floats must keep their ordinary parse.
f, errs := ParseWithOptions("t_amd64.s", "TEXT ·f(SB), NOSPLIT, $0\n\tMOVQ $1.5, AX\n\tMOVQ $·sym(SB), AX\n\tRET\n", Options{})
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
fn := texts(f)[0]
mov1 := fn.Body[0].(*ast.Instr)
if mov1.Operands[0].Imm.HasVal || mov1.Operands[0].Imm.Float != "1.5" {
t.Errorf("float immediate parsed as %+v", mov1.Operands[0].Imm)
}
mov2 := fn.Body[1].(*ast.Instr)
if mov2.Operands[0].Imm.Sym == nil {
t.Errorf("symbol immediate parsed as %+v", mov2.Operands[0].Imm)
}
}
+44
View File
@@ -0,0 +1,44 @@
// Differential kernel: the _dbar (acquire/release) atomic exchange
// variants against the Go toolchain's loong64enc1.s rows.
#include "textflag.h"
TEXT ·AMXORDBW(SB), NOSPLIT, $0
AMXORDBW R14, (R13), R12
RET
TEXT ·AMXORDBV(SB), NOSPLIT, $0
AMXORDBV R14, (R13), R12
RET
TEXT ·AMMAXDBW(SB), NOSPLIT, $0
AMMAXDBW R14, (R13), R12
RET
TEXT ·AMMAXDBV(SB), NOSPLIT, $0
AMMAXDBV R14, (R13), R12
RET
TEXT ·AMMINDBW(SB), NOSPLIT, $0
AMMINDBW R14, (R13), R12
RET
TEXT ·AMMINDBV(SB), NOSPLIT, $0
AMMINDBV R14, (R13), R12
RET
TEXT ·AMMAXDBWU(SB), NOSPLIT, $0
AMMAXDBWU R14, (R13), R12
RET
TEXT ·AMMAXDBVU(SB), NOSPLIT, $0
AMMAXDBVU R14, (R13), R12
RET
TEXT ·AMMINDBWU(SB), NOSPLIT, $0
AMMINDBWU R14, (R13), R12
RET
TEXT ·AMMINDBVU(SB), NOSPLIT, $0
AMMINDBVU R14, (R13), R12
RET
+69
View File
@@ -0,0 +1,69 @@
// Atomics and carry-extending multi-word arithmetic: exchange,
// compare-exchange, exchange-add, ADCX/ADOX and the CRC-32 accumulator
// family. Every result is folded back so no instruction is dead.
#include "textflag.h"
// func xchg(p *uint64, v uint64) uint64
TEXT ·xchg(SB), NOSPLIT, $0-24
MOVQ p+0(FP), AX
MOVQ v+8(FP), BX
XCHGQ BX, (AX)
XCHGQ BX, CX
XCHGL BX, CX
XCHGW BX, CX
XCHGB BL, CL
MOVQ AX, ret+16(FP)
RET
// func cmpxchg(p *uint64, old, new uint64) uint8
TEXT ·cmpxchg(SB), NOSPLIT, $0-25
MOVQ p+0(FP), AX
MOVQ old+8(FP), BX
MOVQ new+16(FP), CX
CMPXCHGQ CX, (AX)
CMPXCHGL CX, BX
CMPXCHGW CX, BX
CMPXCHGB CL, BL
SETEQ AL
MOVB AL, ret+24(FP)
RET
// func xadd(p *uint64, v uint64) uint64
TEXT ·xadd(SB), NOSPLIT, $0-24
MOVQ p+0(FP), AX
MOVQ v+8(FP), BX
XADDQ BX, (AX)
XADDL BX, CX
XADDW BX, CX
XADDB BL, CL
MOVQ AX, ret+16(FP)
RET
// func adcx_adox(lo, hi, x, y uint64) uint64
TEXT ·adcx_adox(SB), NOSPLIT, $0-40
MOVQ lo+0(FP), AX
MOVQ hi+8(FP), DX
MOVQ x+16(FP), BX
MOVQ y+24(FP), CX
ADCXQ BX, AX
ADOXQ CX, DX
ADCXL BX, AX
ADOXL CX, DX
XORQ BX, BX
ADCXQ BX, AX
MOVQ AX, ret+32(FP)
RET
// func crc32(crc uint32, p *byte, n int) uint32
TEXT ·crc32(SB), NOSPLIT, $0-28
MOVL crc+0(FP), AX
MOVQ p+8(FP), SI
MOVQ n+16(FP), CX
CRC32B (SI), AX
CRC32Q (SI), CX
CRC32L (SI), AX
MOVW (SI), DX
CRC32W DX, AX
MOVL AX, ret+24(FP)
RET
+72
View File
@@ -0,0 +1,72 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// Differential kernel for the arm64 synchronisation instructions: the
// acquire/release loads and stores, the exclusive family and the LSE
// atomics with acquire and release semantics, plus the register-pair
// loads and stores. Every function is byte-compared against go tool asm.
#include "textflag.h"
// func acquireRelease()
TEXT ·acquireRelease(SB), NOSPLIT, $0-0
LDAR (R1), R2
LDARB (R3), R4
LDARH (R5), R6
LDARW (R7), R8
STLR R2, (R1)
STLRB R4, (R3)
STLRH R6, (R5)
STLRW R8, (R7)
RET
// func exclusive()
TEXT ·exclusive(SB), NOSPLIT, $0-0
LDAXR (R1), R2
LDAXRB (R3), R4
LDAXRW (R5), R6
STLXR R2, (R1), R8
STLXRB R4, (R3), R8
STLXRW R6, (R5), R8
RET
// func lseAcquireRelease()
TEXT ·lseAcquireRelease(SB), NOSPLIT, $0-0
CASALD R1, (R3), R2
CASALW R4, (R6), R5
LDADDALD R1, (R3), R2
LDADDALW R4, (R6), R5
LDCLRALB R1, (R3), R2
LDCLRALW R4, (R6), R5
LDCLRALD R1, (R3), R2
LDORALB R1, (R3), R2
LDORALW R4, (R6), R5
LDORALD R1, (R3), R2
SWPALB R1, (R3), R2
SWPALW R4, (R6), R5
SWPALD R1, (R3), R2
RET
// func lseBase()
TEXT ·lseBase(SB), NOSPLIT, $0-0
LDADDD R1, (R3), R2
LDADDW R4, (R6), R5
CASD R1, (R3), R2
CASW R4, (R6), R5
SWPD R1, (R3), R2
SWPW R4, (R6), R5
RET
// func pairs()
TEXT ·pairs(SB), NOSPLIT, $0-0
LDP (R1), (R2, R3)
LDP 8(R4), (R5, R6)
LDP -16(R1), (R2, R3)
LDPW 4(R4), (R5, R6)
STP (R2, R3), 24(R7)
STP (R2, R3),-8(R7)
STPW (R1, R2), 4(R0)
FLDPD (R8), (F1, F2)
FLDPD 8(R8), (F3, F4)
FSTPD (F3, F4),-8(R9)
RET
+49
View File
@@ -0,0 +1,49 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// Differential kernel for the loong64 atomics: the AM* family in its plain
// and _dbar (acquire/release) forms, spelled as the runtime's
// atomic_loong64.s spells them. Every AM* takes three operands:
// value, (address), result.
#include "textflag.h"
TEXT ·plain(SB), NOSPLIT, $0-0
AMSWAPB R14, (R13), R12
AMSWAPH R14, (R13), R12
AMSWAPW R5, (R4), R6
AMSWAPV R5, (R4), R0
AMCASB R14, (R13), R12
AMCASH R6, (R4), R5
AMCASW R6, (R4), R5
AMCASV R6, (R4), R5
AMADDW R5, (R4), R0
AMADDV R14, (R13), R12
AMANDW R5, (R4), R6
AMANDV R5, (R4), R6
AMORW R5, (R4), R0
AMORV R5, (R4), R6
AMXORW R5, (R4), R6
AMXORV R5, (R4), R6
AMMAXW R5, (R4), R6
AMMAXV R5, (R4), R6
AMMINW R5, (R4), R6
AMMINV R5, (R4), R6
AMMAXWU R5, (R4), R6
AMMAXVU R5, (R4), R6
AMMINWU R5, (R4), R6
AMMINVU R5, (R4), R6
RET
TEXT ·dbar(SB), NOSPLIT, $0-0
AMADDDBW R5, (R4), R6
AMADDDBV R5, (R4), R6
AMANDDBW R5, (R6), R0
AMANDDBV R5, (R4), R6
AMORDBW R5, (R6), R0
AMORDBV R5, (R4), R6
AMSWAPDBW R5, (R4), R6
AMSWAPDBV R5, (R4), R0
AMCASDBW R6, (R4), R5
AMCASDBV R6, (R4), R5
RET
+35
View File
@@ -0,0 +1,35 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// Differential kernel for the riscv64 atomics: the RV64A AMO family and the
// load-reserved / store-conditional pair, in the toolchain's spelling
// (value, (address), result). Both orderings sit in the encodings: the
// table gives every AMO aq and rl, LR acquire and SC release.
#include "textflag.h"
TEXT ·amo(SB), NOSPLIT, $0-0
AMOSWAPW X5, (X6), X7
AMOSWAPD X5, (X6), X7
AMOADDW X5, (X6), X7
AMOADDD X5, (X6), X7
AMOANDW X5, (X6), X7
AMOANDD X5, (X6), X7
AMOORW X5, (X6), X7
AMOORD X5, (X6), X7
AMOXORW X5, (X6), X7
AMOXORD X5, (X6), X7
AMOMAXW X5, (X6), X7
AMOMAXD X5, (X6), X7
AMOMAXUW X5, (X6), X7
AMOMAXUD X5, (X6), X7
AMOMINUW X5, (X6), X7
AMOMINUD X5, (X6), X7
RET
TEXT ·lrsc(SB), NOSPLIT, $0-0
LRW (X5), X6
LRD (X5), X6
SCW X5, (X6), X7
SCD X5, (X6), X7
RET
+191
View File
@@ -0,0 +1,191 @@
// The AVX-512 families behind the avx512enc gap: AES round ops, integer
// VNNI and bit algorithms, word shifts and permutes with an immediate or a
// register count, lane broadcasts and extracts, gather and scatter prefetch
// hints, opmask broadcasts, the high/low half moves and the non-temporal
// stores. Every result is folded back so no instruction is dead.
#include "textflag.h"
// func avx512int(p *byte, n int) uint64
TEXT ·avx512int(SB), NOSPLIT, $0-24
MOVQ p+0(FP), SI
MOVQ n+16(FP), CX
// AES rounds through the EVEX spellings, masks included.
VAESENC Z20, Z21, Z22
VAESENCLAST Z23, Z24, Z25
VAESDEC (SI), Z26, Z27
VAESDECLAST Z28, Z29, Z30
// Integer VNNI and the bit algorithm group.
VPDPBUSD Z1, Z2, K2, Z3
VPDPBUSDS Z4, Z5, K2, Z6
VPDPWSSD Z7, Z8, Z9
VPDPWSSDS Z10, Z11, K2, Z12
VPOPCNTW Z12, K3, Z13
VPOPCNTB Z14, Z15
VGF2P8MULB Z16, Z17, K4, Z18
VGF2P8AFFINEQB $7, Z18, Z19, K5, Z20
// Byte/word arithmetic with saturation and masks.
VPADDSB Z1, Z2, K1, Z3
VPADDUSW Z3, Z4, K1, Z5
VPSUBSW Z5, Z6, K1, Z7
VPSUBUSB Z7, Z8, K1, Z9
VPSADBW Z9, Z10, Z11
VPMULHRSW Z11, Z12, Z13
VPMULHW Z13, Z14, Z15
VPUNPCKLBW Z15, Z16, K2, Z17
VPUNPCKHBW Z17, Z18, K2, Z19
VPUNPCKLWD Z19, Z20, K2, Z21
VPUNPCKHWD Z21, Z22, K2, Z23
VPCMPEQB Z23, Z24, K2, K3
VPCMPGTW Z25, Z26, K2, K3
VPCMPEQQ Z27, Z28, K2
VPMULTISHIFTQB Z29, Z30, K3, Z31
VDBPSADBW $3, Z1, Z2, K3, Z3
MOVQ CX, ret+16(FP)
RET
// func avx512perm(p *byte) uint64
TEXT ·avx512perm(SB), NOSPLIT, $0-16
MOVQ p+0(FP), SI
// Permutations: immediate and register counts, ternary logic.
VALIGNQ $3, Z1, Z2, K1, Z3
VPERMT2B Z3, Z4, K1, Z5
VPERMT2W Z5, Z6, K1, Z7
VPERMT2PS Z7, Z8, K1, Z9
VPERMI2W Z9, Z10, K1, Z11
VPERMI2PS Z11, Z12, K1, Z13
VPERMI2PD Z13, Z14, K1, Z15
VPERMB Z15, Z16, K1, Z17
VPERMW Z17, Z18, K1, Z19
VPERMPS Z19, Z20, Z21
VPERMD Z20, Z21, Z22
VPERMQ $1, Z1, K2, Z2
VPERMQ Z3, Z4, K2, Z5
VPERMPD $1, Z5, K2, Z6
VPERMPD Z7, Z8, K2, Z9
VPERMILPS $5, Z9, K2, Z10
VPERMILPS Z11, Z12, K2, Z13
VPERMILPD $1, Z13, K2, Z14
VPERMILPD Z15, Z16, K2, Z17
VPTERNLOGD $6, Z17, Z18, K2, Z19
VPTERNLOGQ $9, Z19, Z20, K2, Z21
// Lane shuffle and blend families.
VSHUFPD $1, Z1, Z2, K1, Z3
VSHUFPS $2, Z4, Z5, K1, Z6
VBLENDMPD Z7, Z8, K1, Z9
VBLENDMPS Z9, Z10, K1, Z11
VPBLENDMB Z11, Z12, K1, Z13
VPBLENDMW Z13, Z14, K1, Z15
VPBLENDMD Z15, Z16, K1, Z17
VPBLENDMQ Z17, Z18, K1, Z19
// Conflicts and leading zero counts.
VPCONFLICTD Z1, K1, Z2
VPCONFLICTQ Z3, K1, Z4
VPLZCNTD Z5, K1, Z6
VPLZCNTQ Z7, K1, Z8
// Compress and expand, byte and word widths.
VPCOMPRESSB Z1, K1, (SI)
VPCOMPRESSW Z2, K1, (SI)
VPEXPANDB (SI), K1, Z3
VPEXPANDW (SI), K1, Z4
MOVQ SI, ret+8(FP)
RET
// func avx512shift(p *byte) uint64
TEXT ·avx512shift(SB), NOSPLIT, $0-16
MOVQ p+0(FP), SI
// Variable shifts and shuffles with masks.
VPSLLVW Z1, Z2, K1, Z3
VPSRLVW Z3, Z4, K1, Z5
VPSRAVW Z5, Z6, K1, Z7
VPSHLDVW Z7, Z8, K1, Z9
VPSHRDVW Z9, Z10, K1, Z11
VPSHLDVD Z11, Z12, K1, Z13
VPSHLDVQ Z13, Z14, K1, Z15
VPSHRDVD Z15, Z16, K1, Z17
VPSHRDVQ Z17, Z18, K1, Z19
// Immediate shifts, the word/byte-quad widths and masks.
VPSLLW $3, Z1, K2, Z2
VPSRLW $5, Z3, K2, Z4
VPSRAW $7, Z5, K2, Z6
VPSLLDQ $9, Z7, Z8
VPSRLDQ $11, Z9, Z10
// Register-count shifts and their memory-count forms.
VPSLLD X1, Z2, K1, Z3
VPSRLD 16(SI), Z4, K1, Z5
VPSLLQ X6, Z7, K1, Z8
VPSRLQ X9, Z10, K1, Z11
VPSLLW X12, Z13, K1, Z14
VPSRAW X15, Z16, K1, Z17
VPSRAQ $13, Z12, K1, Z13
VPSRAD X14, Z15, K1, Z16
// Lane shuffles in and out.
VPSHLDW $2, Z1, Z2, K1, Z3
VPSHLDQ $4, Z3, Z4, K1, Z5
VPSHRDW $6, Z5, Z6, K1, Z7
VPSHRDQ $8, Z7, Z8, K1, Z9
VPSHUFBITQMB Z9, Z10, K3
VPTESTMB Z11, Z12, K4
VPTESTNMQ Z13, Z14, K5
MOVQ SI, ret+8(FP)
RET
// func avx512float(x float64) float64
TEXT ·avx512float(SB), NOSPLIT, $0-16
// Square roots, compares and the EXP2/RCP28 helpers.
MOVQ x+0(FP), AX
VSQRTPD Z1, K1, Z2
VSQRTPS Z3, K1, Z4
VSQRTSD X1, X2, K1, X3
VSQRTSS X3, X4, X5
VCOMISD X5, X6
VUCOMISS X7, X8
VEXP2PD Z5, K1, Z6
VRCP28PD Z7, K1, Z8
VRCP28SD X9, X8, K1, X10
VRSQRT28PS Z11, K1, Z12
VRSQRT28SS X11, X10, K1, X12
VCVTSD2SS X1, X2, X3
VCVTSS2SD X3, X2, K1, X4
VFMADD132PD Z1, Z2, K1, Z3
VFMADD231SD X1, X2, K1, X3
VFMSUBADD213PS Z3, Z4, K1, Z5
VFNMSUB231PD Z5, Z6, K1, Z7
// Broadcasts and masked moves.
VBROADCASTF32X2 X1, K1, Z2
VBROADCASTI64X2 (SI), K1, Z3
VMOVUPS Z1, K2, Z3
VMOVSD X14, X5, K3, X22
VMOVSS X18, X3, K2, X25
VMOVHPS (SI), X18, X19
VMOVHPS X20, 8(SI)
VMOVLHPS X16, X5, X17
VMOVNTDQ Z7, (SI)
VMOVNTDQA 64(SI), Z8
VMOVNTPD Z9, (SI)
MOVQ SI, ret+8(FP)
RET
// func avx512mask(p *byte) uint64
TEXT ·avx512mask(SB), NOSPLIT, $0-16
MOVQ p+0(FP), SI
// Omask broadcasts and the K register logic.
VPBROADCASTMB2Q K1, Z2
VPBROADCASTMW2D K3, Z4
KUNPCKWD K6, K4, K1
KADDB K2, K3, K5
KORW K1, K2, K7
// Gather and scatter prefetch hints.
VGATHERPF0DPD K5, (SI)(Y29*8)
VSCATTERPF1DPS K2, (SI)(Z28*4)
// Masked gathers ride the EVEX spelling; the data length wins L'L.
VGATHERDPD (SI)(X10*4), K7, Y22
VPSCATTERDQ Y6, K2, (SI)(X4*1)
// Lane extracts to general registers.
VPEXTRB $3, X1, AX
VPEXTRD $1, X2, DI
VPINSRQ $1, SI, X3, X4
VEXTRACTI32X4 $1, Z1, X5
VINSERTI64X2 $1, X6, Z7, K2, Z8
MOVQ SI, ret+8(FP)
RET
+102
View File
@@ -0,0 +1,102 @@
// The AVX/AVX-512 gap families: fused scalar multiply-add, carries through
// GF(2^8) affine transforms, population counts, non-temporal stores, mask
// moves and the KMOV widths. Every result is folded back so no instruction
// is dead.
#include "textflag.h"
// func avxblend(a, b []float64) float64
TEXT ·avxblend(SB), NOSPLIT, $0-56
MOVQ a_base+0(FP), SI
MOVQ b_base+24(FP), DI
VMOVUPD (SI), Y0
VMOVUPD (DI), Y1
VXORPS Y2, Y2, Y2
VSHUFPD $5, Y0, Y1, Y3
VMOVUPD Y3, (SI)
VPBLENDD $3, Y0, Y1, Y4
VPERM2F128 $1, Y4, Y0, Y0
VEXTRACTF128 $1, Y0, X1
VZEROALL
VMOVSD X1, ret+48(FP)
RET
// func avxint(p *byte, n int) uint64
TEXT ·avxint(SB), NOSPLIT, $0-24
MOVQ p+0(FP), SI
VMOVDQU (SI), Y0
VPCMPEQB Y0, Y0, Y1
VPSLLDQ $2, X0, X0
VPSRLDQ $4, Y0, Y0
VPALIGNR $3, X0, X1, X1
VPCLMULQDQ $0, X0, X1, X2
VGF2P8AFFINEQB $7, X2, X0, X3
VPOPCNTB X3, X4
VPOPCNTD Y0, Y5
VPERMI2B X0, X1, X2
VPTEST X0, X0
VPMOVMSKB X1, AX
VZEROUPPER
MOVQ AX, ret+16(FP)
RET
// func avxnt(p *float64)
TEXT ·avxnt(SB), NOSPLIT, $0-8
MOVQ p+0(FP), DI
VMOVUPD (DI), Y0
VADDPD Y0, Y0, Y0
VMOVNTDQ Y0, (DI)
VMOVNTDQ X0, 16(DI)
VZEROALL
RET
// func avxmas(a, b []float64) float64
TEXT ·avxmas(SB), NOSPLIT, $0-56
MOVQ a_base+0(FP), SI
MOVQ b_base+24(FP), DI
VMOVSD (SI), X0
VMOVSD (DI), X1
VFMADD213SD X1, X0, X0
VFNMADD231SD X1, X0, X0
VADDSD X1, X0, X0
VMOVSD X0, ret+48(FP)
RET
// func avxgpr(x, y uint64) uint64
TEXT ·avxgpr(SB), NOSPLIT, $0-24
MOVQ x+0(FP), AX
MOVQ y+8(FP), BX
ANDNL BX, AX, CX
MULXQ BX, DX, SI
RORXL $3, AX, CX
RORXQ $7, BX, SI
MOVQ CX, ret+16(FP)
RET
// func avxmask(kin uint8, p *byte) uint8
TEXT ·avxmask(SB), NOSPLIT, $0-17
MOVQ p+8(FP), SI
KMOVB kin+0(FP), K1
KMOVB K1, K2
KMOVW K2, K1
KMOVD K1, K3
KMOVQ K3, K4
KMOVB K4, K1
KMOVB K1, AX
KMOVD K1, (SI)
MOVB AL, ret+8(FP)
RET
// func avx512(p *uint64, n int) uint64
TEXT ·avx512(SB), NOSPLIT, $0-24
MOVQ p+0(FP), SI
VMOVDQU64 (SI), Z0
VPORQ Z0, Z0, Z1
VPOPCNTQ Z1, Z2
VPERMB Z1, Z0, Z2
VPXORD Z2, Z1, Z0
VMOVDQA64 Z0, (SI)
VZEROUPPER
XORQ AX, AX
MOVQ AX, ret+16(FP)
RET
+64
View File
@@ -0,0 +1,64 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// Differential kernel for the riscv64 toolchain-synthesised instructions:
// the Zbb-style pseudos the assembler expands instruction-for-instruction
// (ANDN/ORN, MIN/MAX, ROR and friends, the reversed branches, FABSD), the
// CSR read RDTIME and the FP sign-injection and fused-multiply-add forms.
#include "textflag.h"
TEXT ·logic(SB), NOSPLIT, $0-0
ANDN X19, X20, X21
ANDN X19, X20
ANDN X21, X19, X21
ORN X20, X19
ORN X20, X19, X21
MAX X26, X28, X29
MAX X26, X28
MAXU X28, X29, X30
MAXU X28, X29
MIN X29, X30, X5
MIN X29, X30
MINU X30, X5, X6
MINU X30, X5
MAX X5, X5
MAX X5, X5, X6
SEQZ X5, X6
NEG X5, X6
NEG X5
NOT X5
NOT X5, X6
NOP
RET
TEXT ·rotate(SB), NOSPLIT, $0-0
ROR X10, X11, X12
ROR X10, X11
ROR $63, X11
RORIW $31, X13, X14
RORIW $1, X14, X15
RORIW $3, X14
RORW X15, X16, X17
RORW $31, X13
RET
TEXT ·fp(SB), NOSPLIT, $0-0
FABSD F1, F2
FSGNJD F1, F0, F2
FMADDD F1, F2, F3, F4
FMSUBD F1, F2, F3, F4
FNMSUBD F1, F2, F3, F4
FMADDS F1, F2, F3, F4
FNMADDS F1, F2, F3, F4
RET
TEXT ·branches(SB), NOSPLIT, $0-0
BGT X5, X6, tgt
BLE X5, X6, tgt
BGTU X5, X6, tgt
BLEU X5, X6, tgt
tgt:
RDTIME X5
RET
+31
View File
@@ -0,0 +1,31 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// Differential kernel for the arm64 bookkeeping statements: the funcdata.h
// pseudo-directives (GO_ARGS, NO_LOCAL_POINTERS, FUNCDATA, PCDATA) contribute
// no instruction bytes, and every function is byte-compared against
// go tool asm.
#include "textflag.h"
#include "funcdata.h"
// func bookkeep()
TEXT ·bookkeep(SB), NOSPLIT, $8-0
GO_ARGS
FUNCDATA $3, inline_tree(SB)
PCDATA $1, $2
MOVD R1, 0(RSP)
RET
// func bookkeepNoLocals()
TEXT ·bookkeepNoLocals(SB), NOSPLIT, $16-0
NO_LOCAL_POINTERS
PCDATA $0, $0
PCDATA $1, $1
MOVD R2, 8(RSP)
RET
// func bookkeepPlain()
TEXT ·bookkeepPlain(SB), NOSPLIT, $0-0
MOVD R3, R4
RET
File diff suppressed because it is too large Load Diff
+48
View File
@@ -0,0 +1,48 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// Carry arithmetic, logical shifts, register aliases with element selectors
// and the ADC/SBC immediate spellings: the shapes nat_arm64.s, p256 and
// gcm_arm64.s exercise. Byte-for-byte against go tool asm.
#include "textflag.h"
#define acc0 V8
#define acc1 V9
#define const0 R15
#define POLY V15
// carry pins the ADC/SBC family: the $0 spellings in two and three
// operands, and the register-carry forms.
TEXT ·carry(SB), NOSPLIT, $0-0
ADC $0, R20
ADC $0, R20, R4
SBCS $0, R4
SBCS $0, R4, R12
SBCS R15, R4, R12
SBC $0, R1
ADCSW $0, R2, R3
RET
// shift pins the shifted-register forms including ROR, which only the
// logical family accepts.
TEXT ·shift(SB), NOSPLIT, $0-0
ANDW R9@>7, R19, R26
AND R1@>33, R2, R3
ADD R1<<11, R2, R3
SUB R1->33, R2
ORR R5<<2, R6, R7
RET
// vecalias pins the vector aliases with element selectors and the
// structure loads with aliased members.
TEXT ·vecalias(SB), NOSPLIT, $0-0
MOVD $0xC2, R1
VMOV R1, POLY.D[0]
VMOV R0, POLY.D[1]
VEOR POLY.B16, POLY.B16, POLY.B16
VLD1 (R0), [acc0.B16]
VLD1.P (R0), [acc0.B16, acc1.B16]
VST1 [acc0.B16, acc1.B16], (R1)
VST1.P [acc0.B16, acc1.B16], 32(R1)
RET
+57
View File
@@ -0,0 +1,57 @@
// The AES-NI, SHA and carry-less multiply round instructions as GOROOT's
// crypto kernels spell them. Every result is folded back so no instruction
// is dead.
#include "textflag.h"
// func aesround(blk, rk *byte)
TEXT ·aesround(SB), NOSPLIT, $0-16
MOVQ blk+0(FP), SI
MOVQ rk+8(FP), DI
MOVOU (SI), X0
MOVOU (DI), X1
AESENC X1, X0
AESENCLAST X1, X0
AESDEC X1, X0
AESDECLAST X1, X0
AESIMC X1, X2
AESKEYGENASSIST $1, X1, X3
MOVOU X0, (SI)
MOVOU X2, (DI)
RET
// func sha1block(p *byte, n int, h *[5]uint32)
TEXT ·sha1block(SB), NOSPLIT, $0-24
MOVQ p+0(FP), SI
MOVQ h+16(FP), DI
MOVOU (SI), X0
MOVOU 16(SI), X1
SHA1RNDS4 $0, X1, X0
SHA1NEXTE X1, X0
SHA1MSG1 X1, X2
SHA1MSG2 X1, X2
MOVOU X0, (DI)
RET
// func sha256block(p *byte, n int, h *[8]uint32)
TEXT ·sha256block(SB), NOSPLIT, $0-24
MOVQ p+0(FP), SI
MOVQ h+16(FP), DI
MOVOU (SI), X0
MOVOU 16(SI), X1
SHA256RNDS2 X0, X1, X0
SHA256MSG1 X1, X2
SHA256MSG2 X1, X2
MOVOU X0, (DI)
RET
// func pclmul(a, b *byte)
TEXT ·pclmul(SB), NOSPLIT, $0-16
MOVQ a+0(FP), SI
MOVQ b+8(FP), DI
MOVOU (SI), X0
MOVOU (DI), X1
PCLMULQDQ $0, X1, X0
PCLMULQDQ $17, (DI), X0
MOVOU X0, (SI)
RET
+42
View File
@@ -0,0 +1,42 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// Differential kernel for the arm64 cryptographic extension: the AES round
// instructions and the SHA1, SHA256 and SHA512 families. Every function is
// byte-compared against go tool asm.
#include "textflag.h"
// func aesRound()
TEXT ·aesRound(SB), NOSPLIT, $0-0
AESE V31.B16, V29.B16
AESD V22.B16, V19.B16
AESMC V14.B16, V28.B16
AESIMC V12.B16, V27.B16
RET
// func sha1Round()
TEXT ·sha1Round(SB), NOSPLIT, $0-0
SHA1C V8.S4, V8, V2
SHA1P V3.S4, V20, V27
SHA1M V0.S4, V27, V27
SHA1H V17, V25
SHA1SU0 V17.S4, V13.S4, V16.S4
SHA1SU1 V24.S4, V23.S4
RET
// func sha256Round()
TEXT ·sha256Round(SB), NOSPLIT, $0-0
SHA256H V4.S4, V2, V11
SHA256H2 V6.S4, V16, V11
SHA256SU0 V0.S4, V16.S4
SHA256SU1 V31.S4, V3.S4, V15.S4
RET
// func sha512Round()
TEXT ·sha512Round(SB), NOSPLIT, $0-0
SHA512H V2.D2, V1, V0
SHA512H2 V4.D2, V3, V2
SHA512SU0 V9.D2, V8.D2
SHA512SU1 V7.D2, V6.D2, V5.D2
RET
+25
View File
@@ -0,0 +1,25 @@
#include "textflag.h"
// The kernel exercises the symbol-valued DATA spelling the runtime's rt0
// files use: a data word holding the address of a symbol, resolved by the
// linker through a relocation at the field.
// func lookup() ptr
TEXT ·lookup(SB), NOSPLIT, $0-8
MOVQ handlers+8(SB), AX
MOVQ AX, ret+0(FP)
RET
// func handler() int64
TEXT ·handler(SB), NOSPLIT, $0-8
MOVQ $42, AX
MOVQ AX, ret+0(FP)
RET
GLOBL handlers(SB), NOPTR, $24
DATA handlers+0(SB)/8, $·handler(SB)
DATA handlers+8(SB)/8, $table(SB)
DATA handlers+16(SB)/8, $·handler+5(SB)
GLOBL table(SB), RODATA, $8
DATA table+0(SB)/8, $0x123456789abcdef0
+25
View File
@@ -0,0 +1,25 @@
#include "textflag.h"
// The kernel exercises the symbol-valued DATA spelling the runtime's rt0
// files use: a data word holding the address of a symbol, resolved by the
// linker through a relocation at the field.
// func lookup() ptr
TEXT ·lookup(SB), NOSPLIT, $0-8
MOVD handlers+8(SB), R4
MOVD R4, ret+0(FP)
RET
// func handler() int64
TEXT ·handler(SB), NOSPLIT, $0-8
MOVZ $42, R4
MOVD R4, ret+0(FP)
RET
GLOBL handlers(SB), NOPTR, $24
DATA handlers+0(SB)/8, $·handler(SB)
DATA handlers+8(SB)/8, $table(SB)
DATA handlers+16(SB)/8, $extentry(SB)
GLOBL table(SB), RODATA, $8
DATA table+0(SB)/8, $0x123456789abcdef0
+19
View File
@@ -0,0 +1,19 @@
#include "textflag.h"
// The kernel exercises the U+2215 DIVISION SLASH inside a symbol's package
// path: internal∕runtime∕atomic·Xchg, the spelling sync/atomic/asm.s uses.
// The middle dot (U+00B7) still separates the package path from the name.
// func swap(a, b int64) int64
TEXT ·swap(SB), NOSPLIT, $0-24
MOVQ a+0(FP), DI
MOVQ b+8(FP), SI
CALL internal∕runtime∕atomic·Xchg(SB)
MOVQ AX, ret+16(FP)
RET
// func note() int64
TEXT ·note(SB), NOSPLIT, $0-8
CALL runtime∕debug·SetGCPercent(SB)
MOVQ AX, ret+0(FP)
RET
+19
View File
@@ -0,0 +1,19 @@
#include "textflag.h"
// The kernel exercises the U+2215 DIVISION SLASH inside a symbol's package
// path: internal∕runtime∕atomic·Xchg, the spelling sync/atomic/asm.s uses.
// The middle dot (U+00B7) still separates the package path from the name.
// func swap(a, b int64) int64
TEXT ·swap(SB), NOSPLIT, $0-24
MOVD a+0(FP), R4
MOVD b+8(FP), R5
CALL internal∕runtime∕atomic·Xchg(SB)
MOVD R4, ret+16(FP)
RET
// func note() int64
TEXT ·note(SB), NOSPLIT, $0-8
CALL runtime∕debug·SetGCPercent(SB)
MOVD R0, ret+0(FP)
RET
+33
View File
@@ -0,0 +1,33 @@
// The three-operand SHL/SHR forms, which go tool asm encodes as SHLD/SHRD:
// immediate and CL (or its CX spelling) counts at the Q and W widths, next
// to the two-operand CX-count spelling GOROOT's bignum kernels use. Every
// result is folded back so no instruction is dead.
#include "textflag.h"
// func dblshift(x, y uint64) uint64
TEXT ·dblshift(SB), NOSPLIT, $0-24
MOVQ x+0(FP), SI
MOVQ y+8(FP), DI
MOVQ $12, CX
SHLQ $13, SI, DI
SHRQ $7, DI, SI
SHLQ CX, SI, DI
SHRQ CX, DI, SI
SHLQ CX, SI
SHLQ $9, DI
SHLW $1, SI, DI
SHRW $3, DI, SI
XORQ DI, SI
MOVQ SI, ret+16(FP)
RET
// func dblshift32(a, b uint32) uint32
TEXT ·dblshift32(SB), NOSPLIT, $0-12
MOVL a+0(FP), SI
MOVL b+4(FP), DI
SHLL $5, SI, DI
SHRL $2, DI, SI
XORL SI, DI
MOVL DI, ret+8(FP)
RET
+66
View File
@@ -0,0 +1,66 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// Differential kernel for the arm64 integer slice: carry-setting arithmetic,
// widening multiplies, bit manipulation, conditional compares, the compare
// and test branches, ADR and the wide-constant moves. Every function is
// byte-compared against go tool asm.
#include "textflag.h"
// func carryArith()
TEXT ·carryArith(SB), NOSPLIT, $0-0
ADC R0, R2, R12
ADCS R23, R22, R22
ADC $0, R1
SBC R25, R10, R26
SBCS R5, R9, R5
SBCS $0, R1
RET
// func wideningMul()
TEXT ·wideningMul(SB), NOSPLIT, $0-0
MUL R4, R3, R0
MSUB R19, R16, R26, R2
SMULH R24, R20, R24
UMULH R24, R20, R24
RET
// func bitManip()
TEXT ·bitManip(SB), NOSPLIT, $0-0
RBIT R11, R4
REV R1, R2
CLZ R21, R9
REVW R1, R2
CLSW R1, R2
UBFX $33, R17, $25, R5
UBFXW $4, R1, $9, R2
RET
// func condCompare()
TEXT ·condCompare(SB), NOSPLIT, $0-0
CCMP LE, R7, $19, $3
CCMP LT, R30, R6, $7
CCMN EQ, R1, R2, $3
CCMPW LE, R7, $19, $3
RET
// func branchForms()
TEXT ·branchForms(SB), NOSPLIT, $0-0
CBZ R1, target
CBNZ R7, target
CBNZW R2, target
TBZ $4, R7, target
TBNZ $33, R7, target
ADR target, R10
target:
RET
// func wideMoves()
TEXT ·wideMoves(SB), NOSPLIT, $0-0
MOVK $1234, R5
MOVK $305397760, R5
MOVKW $1234, R5
MOVK $16771847290880, R21
RET
+51
View File
@@ -0,0 +1,51 @@
// The subtract-immediate fold, the TEQ/TNE trap pseudos, PRELDX, the FP
// condition branches and the N(PC) branch spellings, against the toolchain.
#include "textflag.h"
// func SubFold(x int64) int64
TEXT ·SubFold(SB), NOSPLIT, $0-16
MOVV x+0(FP), R8
SUBV $0, R8
SUBV $4, R9, R10
SUBV $4096, R11
SUBV $-4, R12
SUB $1, R13
SUBVU $4, R14
SUBV $1048576, R15
MOVV R8, ret+8(FP)
RET
// func Traps(x int64) int64
TEXT ·Traps(SB), NOSPLIT, $0-16
MOVV x+0(FP), R4
TEQ $4, R4, R5
TEQ $4, R4
TNE $6, R5, R6
MOVV R4, ret+8(FP)
RET
// func Prefetch(x int64) int64
TEXT ·Prefetch(SB), NOSPLIT, $0-16
MOVV x+0(FP), R7
PRELDX 0(R7), $0x80001021, $0
PRELDX -1(R7), $0x1021, $2
MOVV R7, ret+8(FP)
RET
// func BranchForms(x int64) int64
TEXT ·BranchForms(SB), NOSPLIT, $0-16
MOVV x+0(FP), R4
l1:
BFPT l1
BFPT FCC3, l1
BFPF l1
JMP -4(PC)
JAL 1(PC)
JAL (R4)
loop:
ADDV $1, R4
BEQ R4, R5, loop
BNE R4, l1
RET
+33
View File
@@ -0,0 +1,33 @@
// PCALIGN padding on loong64: andi $0, $0, 0 (the architecture's NOP), plus
// the automatic loop-head alignment to a 16-byte boundary.
#include "textflag.h"
// func Pad16(x int64) int64
TEXT ·Pad16(SB), NOSPLIT, $0-16
MOVV x+0(FP), R4
PCALIGN $16
ADDV $1, R4
MOVV R4, ret+8(FP)
RET
// func Pad32(x int64) int64
TEXT ·Pad32(SB), NOSPLIT, $0-16
MOVV x+0(FP), R4
PCALIGN $32
ADDV $1, R4
MOVV R4, ret+8(FP)
RET
// func LoopAlign(x int64) int64
TEXT ·LoopAlign(SB), NOSPLIT, $0-16
MOVV x+0(FP), R4
MOVV $10, R5
loop:
BEQ R4, R5, done
ADDV $1, R4
JMP loop
done:
MOVV R4, ret+8(FP)
RET
+35
View File
@@ -0,0 +1,35 @@
// PCALIGN padding on riscv64: 4-byte NOPs with a 2-byte compressed NOP when
// the pad is 2 mod 4, exactly as the toolchain lays the bytes down.
#include "textflag.h"
// func Pad8(x int64) int64
TEXT ·Pad8(SB), NOSPLIT, $0-16
MOV x+0(FP), X5
PCALIGN $8
ADD $1, X5
MOV X5, ret+8(FP)
RET
// func Pad16(x int64) int64
TEXT ·Pad16(SB), NOSPLIT, $0-16
MOV x+0(FP), X5
PCALIGN $16
ADD $1, X5
MOV X5, ret+8(FP)
RET
// func Pad32(x int64) int64
TEXT ·Pad32(SB), NOSPLIT, $0-16
MOV x+0(FP), X5
PCALIGN $32
ADD $1, X5
MOV X5, ret+8(FP)
RET
// func PadAfterOdd(x int64) int64
TEXT ·PadAfterOdd(SB), NOSPLIT, $0-16
MOV x+0(FP), X5
PCALIGN $8
ADD $1, X5
MOV X5, ret+8(FP)
RET
+103
View File
@@ -0,0 +1,103 @@
// Instruction prefixes: LOCK, REP and REPN. go tool asm encodes each
// statement as a standalone one-byte instruction with a PC of its own (F0,
// F3 and F2 respectively); the statement that follows is encoded unaware of
// it, and nothing validates the pairing. The shapes are the runtime's
// atomic read-modify-write family and the string moves, every result folded
// back.
#include "textflag.h"
// func cas64(ptr *uint64, old, new uint64) bool
TEXT ·cas64(SB), NOSPLIT, $0-25
MOVQ ptr+0(FP), BX
MOVQ old+8(FP), AX
MOVQ new+16(FP), CX
LOCK
CMPXCHGQ CX, 0(BX)
SETEQ ret+24(FP)
RET
// func casloop(addr *uint64, v uint64) uint64
// The runtime's Or64 shape: a LOCK inside a branch loop, the backward jump
// measuring over the prefix statement's own byte.
TEXT ·casloop(SB), NOSPLIT, $0-24
MOVQ addr+0(FP), BX
MOVQ v+8(FP), CX
loop:
MOVQ CX, DX
MOVQ (BX), AX
ORQ AX, DX
LOCK
CMPXCHGQ DX, (BX)
JNZ loop
MOVQ AX, ret+16(FP)
RET
// func xadd64(p *uint64, v uint64) uint64
TEXT ·xadd64(SB), NOSPLIT, $0-24
MOVQ p+0(FP), AX
MOVQ v+8(FP), BX
LOCK
XADDQ BX, (AX)
MOVQ AX, ret+16(FP)
RET
// func xaddw(p *uint16, v uint16) uint16
TEXT ·xaddw(SB), NOSPLIT, $0-12
MOVQ p+0(FP), AX
MOVW v+8(FP), BX
LOCK
XADDW BX, (AX)
MOVW AX, ret+8(FP)
RET
// func lockarith(p *uint64)
TEXT ·lockarith(SB), NOSPLIT, $0-8
MOVQ p+0(FP), AX
LOCK
ORQ CX, (AX)
LOCK
ANDL CX, (AX)
LOCK
INCQ (AX)
LOCK
DECQ (AX)
LOCK
ORB BX, (AX)
RET
// func repstring(dst, src *byte, n int)
// The memmove shapes: forward copy by quadwords, backward tails.
TEXT ·repstring(SB), NOSPLIT, $0-24
MOVQ dst+0(FP), DI
MOVQ src+8(FP), SI
REP
MOVSQ
REP
MOVSB
REPN
MOVSB
REP
STOSQ
REP
STOSB
RET
// func pfxlabel()
// Labels pinned on prefix statements' own bytes: pfx: sits on the LOCK,
// mid: on the REPN.
TEXT ·pfxlabel(SB), NOSPLIT, $0-0
pfx:
LOCK
XCHGL BX, (AX)
JMP done
mid:
REPN
MOVSB
done:
REP
STOSB
RET
+31
View File
@@ -0,0 +1,31 @@
// Differential kernel: the bookkeeping statements the assembler accepts and
// encodes to nothing (gasm v. go tool asm, byte for byte).
#include "textflag.h"
TEXT ·end(SB), NOSPLIT, $0
END
RET
TEXT ·funcdata(SB), NOSPLIT, $0
FUNCDATA $0, ref(SB)
RET
TEXT ·pcdata(SB), NOSPLIT, $0
PCDATA $0, $1
PCDATA $1, $-2
RET
TEXT ·getcallerpc(SB), NOSPLIT, $0
GETCALLERPC R4
RET
TEXT ·mixed(SB), NOSPLIT, $0
PCDATA $0, $1
ADDV R4, R5, R6
FUNCDATA $1, ref(SB)
GETCALLERPC R7
RET
ref:
RET
+62
View File
@@ -0,0 +1,62 @@
// Literal data emission: BYTE, WORD, LONG and QUAD write the immediate
// into the text stream as 1, 2, 4 or 8 little-endian bytes with no opcode
// lookup, truncated to the width rather than range-checked; END is
// accepted and ignored, contributing no bytes and ending nothing. The
// shapes mirror the runtime's hand-laid markers
// (crypto/internal/boring/sig/sig_amd64.s) and its syscall stubs
// (runtime/sys_linux_amd64.s).
#include "textflag.h"
// func marker()
// A boring/crypto-style marker: a hand-laid forward branch whose skip
// distance is patched at runtime. One BYTE per statement, as the
// runtime's own file spells it: the semicolon-separated one-liner the
// sys_linux_amd64.s stub uses does not survive gasm fmt, which drops the
// statement separators.
TEXT ·marker(SB), NOSPLIT, $0-0
BYTE $0xEB
BYTE $0x1D
BYTE $0xF4
BYTE $0x48
BYTE $0xF4
BYTE $0x4B
BYTE $0xC3
RET
// func stub()
// The sys_linux_amd64.s stub bytes: the sign-extended
// "48 c7 c0 0f 00 00 00" form of MOVQ $rt_sigreturn, AX.
TEXT ·stub(SB), NOSPLIT, $0-0
BYTE $0x48
BYTE $0xc7
BYTE $0xc0
BYTE $0x0f
BYTE $0x00
BYTE $0x00
BYTE $0x00
RET
// func words()
// The wider literals, and an END that ends nothing: the WORD after it
// still lands in this function.
TEXT ·words(SB), NOSPLIT, $0-0
WORD $0x1234
WORD $-1
LONG $0x11223344
LONG $-1
QUAD $0x1122334455667788
QUAD $-2
END
WORD $0xBEEF
RET
// func trunc()
// Truncation, not a range check: each literal keeps its low bytes, exactly
// as go tool asm emits them.
TEXT ·trunc(SB), NOSPLIT, $0-0
BYTE $0x1FF
WORD $0x12345
LONG $0x123456789
QUAD $-2
RET
+76
View File
@@ -0,0 +1,76 @@
// Carry arithmetic, rotates, unsigned/signed division and bit tests: the
// scalar families GOROOT's big-number and crypto kernels use. Every result
// is folded back so no instruction is dead.
#include "textflag.h"
// func carry(a, b uint64) uint64
TEXT ·carry(SB), NOSPLIT, $0-24
MOVQ a+0(FP), AX
MOVQ b+8(FP), BX
ADDQ BX, AX
ADCQ $0, AX
MOVQ BX, CX
SBBQ $1, CX
ADCL BX, AX
ADCB AL, BL
ADCW $7, CX
MOVQ AX, ret+16(FP)
RET
// func borrow(a, b uint64) uint64
TEXT ·borrow(SB), NOSPLIT, $0-24
MOVQ a+0(FP), AX
MOVQ b+8(FP), BX
SUBQ BX, AX
SBBQ $0, AX
SBBQ BX, CX
MOVQ AX, ret+16(FP)
RET
// func rot(x uint64, n uint32) uint64
TEXT ·rot(SB), NOSPLIT, $0-24
MOVQ x+0(FP), AX
MOVL n+8(FP), CX
ROLQ CL, AX
RORQ $7, AX
ROLL $1, AX
RORL CL, AX
RCLQ $1, AX
RCRQ CL, AX
ROLW $3, AX
SALQ $2, AX
SALB $1, AX
MOVQ AX, ret+8(FP)
RET
// func muldiv(a, b uint64) uint64
TEXT ·muldiv(SB), NOSPLIT, $0-24
MOVQ a+0(FP), AX
MOVQ b+8(FP), BX
MULQ BX
MULQ (BX)
MOVL (BX), CX
MULL CX
DIVQ BX
IDIVQ BX
MOVL a+0(FP), AX
DIVL CX
IDIVL CX
MOVQ AX, ret+16(FP)
RET
// func bitfield(w *uint64) uint64
TEXT ·bitfield(SB), NOSPLIT, $0-16
MOVQ (DI), AX
MOVQ (DI), CX
BTQ AX, CX
BTQ $3, (DI)
BTL AX, CX
BTW $1, CX
BTSQ $5, AX
BTRQ AX, CX
BTCQ $7, (DI)
SETCS AL
MOVQ AX, ret+8(FP)
RET
+21
View File
@@ -0,0 +1,21 @@
#include "textflag.h"
// The kernel exercises the ';' statement separator in a plain file, the way
// the runtime writes it ("ROLQ $3, DI; ROLQ $13, DI", "REP; MOVSQ"). Each
// statement assembles exactly as it would on a line of its own.
// func rol(x int64) int64
TEXT ·rol(SB), NOSPLIT, $0-16
ROLQ $3, DI; ROLQ $13, DI
MOVQ DI, ret+0(FP)
RET
// func move(dst, src unsafe.Pointer)
TEXT ·move(SB), NOSPLIT, $0-16
REP ; MOVSQ
RET
TEXT ·paired(SB), NOSPLIT, $0-8
XORQ AX, AX; XORQ CX, CX
MOVQ AX, ret+0(FP)
RET
+98
View File
@@ -0,0 +1,98 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// Differential kernel for the arm64 NEON slice: the logical and arithmetic
// three-register operations, permutations, comparisons, shifts, the crypto
// four-register group, element moves, table lookups and the structure
// loads and stores. Every function is byte-compared against go tool asm.
#include "textflag.h"
// func simdLogic()
TEXT ·simdLogic(SB), NOSPLIT, $0-0
VADD V1.B16, V2.B16, V3.B16
VADD V1.B8, V2.B8, V3.B8
VSUB V1.S4, V2.S4, V3.S4
VMUL V1.H8, V2.H8, V3.H8
VAND V4.B16, V4.B16, V9.B16
VORR V5.B16, V4.B16, V3.B16
VEOR V0.B16, V1.B16, V0.B16
VADDP V1.H8, V2.H8, V3.H8
VCMEQ V24.S4, V13.S4, V12.S4
VCMEQ $0, V2.H4, V3.H4
RET
// func simdPerm()
TEXT ·simdPerm(SB), NOSPLIT, $0-0
VZIP1 V16.H8, V3.H8, V19.H8
VZIP1 V6.D2, V9.D2, V11.D2
VZIP2 V22.D2, V25.D2, V21.D2
VREV32 V2.H8, V1.H8
VREV64 V2.S4, V3.S4
VUADDLV V31.S4, V11
VEXT $4, V2.B8, V1.B8, V3.B8
VEXT $8, V2.B16, V1.B16, V3.B16
RET
// func simdShift()
TEXT ·simdShift(SB), NOSPLIT, $0-0
VSHL $7, V22.D2, V25.D2
VSHL $24, V1.S4, V2.S4
VUSHR $6, V22.H8, V23.H8
VUSHR $56, V1.D2, V2.D2
VSRI $24, V1.S4, V2.S4
VSRI $56, V1.D2, V2.D2
RET
// func simdCrypto4()
TEXT ·simdCrypto4(SB), NOSPLIT, $0-0
VEOR3 V2.B16, V7.B16, V12.B16, V25.B16
VBCAX V1.B16, V2.B16, V26.B16, V31.B16
VXAR $63, V27.D2, V21.D2, V26.D2
VRAX1 V26.D2, V29.D2, V30.D2
VPMULL V2.D1, V1.D1, V3.Q1
VPMULL V2.B8, V1.B8, V3.H8
VPMULL2 V2.D2, V1.D2, V4.Q1
VPMULL2 V2.B16, V1.B16, V4.H8
RET
// func simdElement()
TEXT ·simdElement(SB), NOSPLIT, $0-0
VDUP V31.B[15], V18
VDUP V19.S[3], V18.S4
VDUP V1.D[1], V2.D2
VMOV V13.S[0], R20
VMOV V11.B[11], V16.B[12]
VMOV R20, V21.B[2]
VMOV V2.B16, V4.B16
RET
// func simdTable()
TEXT ·simdTable(SB), NOSPLIT, $0-0
VTBL V22.B16, [V28.B16], V11.B16
VTBL V18.B8, [V17.B16, V18.B16], V22.B8
VTBL V31.B8, [V14.B16, V15.B16, V16.B16, V17.B16], V15.B8
RET
// func simdLoadStore()
TEXT ·simdLoadStore(SB), NOSPLIT, $0-0
VLD1 (R2), [V21.B16]
VLD1 (R24), [V18.D1, V19.D1, V20.D1]
VLD1 (R29), [V14.D1, V15.D1, V16.D1, V17.D1]
VLD1.P 32(R1), [V2.B16, V3.B16]
VLD1.P 64(R4), [V5.B16, V6.B16, V7.B16, V8.B16]
VLD1R (R1), [V9.B8]
VLD1R (R0), [V0.B16]
VLD4R (R0), [V0.B8, V1.B8, V2.B8, V3.B8]
VST1 [V2.S4, V3.S4, V4.S4, V5.S4], (R14)
VST1 [V14.H4, V15.H4, V16.H4], (R27)
VST1.P [V2.B16], (R1)
VST1.P [V2.B16, V3.B16], 32(R1)
RET
// func simdLiteral()
TEXT ·simdLiteral(SB), NOSPLIT, $0-0
VMOVS $0x80402010, V11
VMOVD $0x8040201008040201, V20
VMOVQ $0x7040201008040201, $0x8040201008040201, V10
RET
+57
View File
@@ -0,0 +1,57 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// Differential kernel for the arm64 whole-vector moves between a general
// register and an arranged vector (VMOV/VDUP Rs, Vd.<T>), the two-operand
// accumulate spellings VADD/VSUB Vm, Vn, and the toolchain-reserved
// R18_PLATFORM register name. Every function is byte-compared against
// go tool asm.
#include "textflag.h"
// func gpIntoVector()
TEXT ·gpIntoVector(SB), NOSPLIT, $0-0
VMOV R1, V2.B8
VMOV R3, V4.B16
VMOV R5, V6.H4
VMOV R7, V8.H8
VMOV R9, V10.S2
VMOV R11, V12.S4
VMOV R13, V14.D2
VDUP R15, V16.B8
VDUP R17, V18.B16
VDUP R19, V20.H8
VDUP R21, V22.S4
VDUP R23, V24.D2
RET
// func simdAccumulate()
TEXT ·simdAccumulate(SB), NOSPLIT, $0-0
VADD V7, V8
VSUB V7, V8
VADD V1, V2
VSUB V30, V31
VADD V0.B16, V1.B16, V2.B16
VSUB V0.S4, V1.S4, V2.S4
RET
// func truncMove()
TEXT ·truncMove(SB), NOSPLIT, $0-0
MOVB R3, R4
MOVH R5, R6
MOVW R9, R10
MOVBU R3, R4
MOVHU R3, R4
MOVWU R3, R4
MOVD R3, R4
RET
// func platformRegister()
TEXT ·platformRegister(SB), NOSPLIT, $0-0
MOVD R18_PLATFORM, R3
MOVW R18_PLATFORM, R4
MOVD R3, R18_PLATFORM
MOVD 0x68(R18_PLATFORM), R5
MOVD R5, 0x68(R18_PLATFORM)
MOVW 8(R18_PLATFORM), R6
RET
+77
View File
@@ -0,0 +1,77 @@
// The legacy SSE gap families: scalar compares and square roots, the Plan 9
// packed spellings, shuffles, lane extracts and inserts, packed integer
// shifts and the octa moves. Every result is folded back so no instruction
// is dead.
#include "textflag.h"
// func cmporder(a, b *float64) int
TEXT ·cmporder(SB), NOSPLIT, $0-24
MOVQ a+0(FP), SI
MOVQ b+8(FP), DI
MOVSD (SI), X0
MOVSD (DI), X1
ANDNPD X0, X2
ANDNPS X0, X3
COMISD X0, X1
SQRTSD X0, X2
CMPSD X0, X1, $5
MOVL SI, CX
SETPL CL
MOVL CX, ret+16(FP)
RET
// func packed(w *uint64) uint64
TEXT ·packed(SB), NOSPLIT, $0-16
MOVQ w+0(FP), SI
MOVO (SI), X0
MOVOA (SI), X1
PADDL X0, X1
PSUBL X0, X1
PCMPEQL X0, X1
PUNPCKLBW X0, X1
PSHUFL $27, X0, X2
MOVOU X2, (SI)
MOVQ (SI), AX
MOVQ AX, ret+8(FP)
RET
// func lanes(p *byte, buf *byte)
TEXT ·lanes(SB), NOSPLIT, $0-16
MOVQ p+0(FP), SI
MOVQ buf+8(FP), DI
MOVO (SI), X0
MOVQ SI, AX
PINSRB $1, AX, X0
PINSRW $2, AX, X0
PINSRD $3, AX, X0
PINSRQ $1, AX, X0
PEXTRB $1, X0, AX
PEXTRW $2, X0, AX
PEXTRD $3, X0, AX
PEXTRQ $1, X0, CX
PCMPESTRI $4, X0, X0
MOVB AL, (DI)
MOVOU X0, (SI)
RET
// func shifts(p *uint64)
TEXT ·shifts(SB), NOSPLIT, $0-8
MOVQ p+0(FP), SI
MOVO (SI), X0
MOVO X0, X1
PSLLW $3, X0
PSRLW $1, X1
PSRAW $2, X0
PSLLL $4, X0
PSRLL $5, X1
PSRAL $1, X0
PSLLQ $7, X0
PSRLQ $9, X1
PSLLL X1, X0
PSRLQ X0, X1
PSLLDQ $2, X0
PSRLDQ $4, X1
MOVOU X0, (SI)
MOVOU X1, 16(SI)
RET
+27
View File
@@ -0,0 +1,27 @@
// Legacy SSE octa moves against static (SB) symbols: the load and store
// shapes GOROOT's AES-CTR, AES-GCM and P-256 kernels spell (MOVOU
// bswapMask<>+0(SB), X0 and the reverse), including offsets into the symbol
// and the aligned MOVO pair. Every result is folded back so no instruction
// is dead.
#include "textflag.h"
// func ssestatic() uint64
TEXT ·ssestatic(SB), NOSPLIT, $0-8
MOVOU bswapMask<>+0(SB), X0
MOVOU bswapMask<>+8(SB), X1
MOVO rodataMask<>+0(SB), X2
PXOR X1, X0
PXOR X2, X0
MOVOU X0, sink<>+0(SB)
MOVOU sink<>+0(SB), X3
PXOR X3, X0
MOVQ X0, AX
MOVQ AX, ret+0(FP)
RET
GLOBL bswapMask<>(SB), RODATA|NOPTR, $16
GLOBL rodataMask<>(SB), RODATA|NOPTR, $16
GLOBL sink<>(SB), NOPTR, $16
+76
View File
@@ -0,0 +1,76 @@
// System, string-primitive and x87 families: flag register moves, the
// serialising instructions, MOVS/STOS, the MXCSR pair, scalar float-to-int
// conversions and FMOVD. Every result is folded back so no instruction is
// dead.
#include "textflag.h"
// func system(x uint64) uint64
TEXT ·system(SB), NOSPLIT, $0-16
MOVQ x+0(FP), AX
PUSHFQ
POPFQ
CPUID
RDTSC
RDTSCP
SYSCALL
XGETBV
PAUSE
LFENCE
MFENCE
SFENCE
UNDEF
XORQ AX, BX
MOVQ BX, ret+8(FP)
RET
// func stringprim(p *byte, n int) uint64
TEXT ·stringprim(SB), NOSPLIT, $0-24
MOVQ p+0(FP), DI
MOVQ n+8(FP), CX
LEAQ buf<>(SB), AX
MOVQ AX, SI
CLD
MOVSB
MOVSW
MOVSL
MOVSQ
STOSB
STOSQ
STOSL
STOSW
MOVQ DI, ret+16(FP)
RET
DATA buf<>+0x00(SB)/8, $0
GLOBL buf<>(SB), NOPTR, $8
// func intgate(x uint64) uint64
TEXT ·intgate(SB), NOSPLIT, $0-16
MOVQ x+0(FP), AX
INT $3
MOVQ AX, ret+8(FP)
RET
// func fpmxcsr(x float64, csr *uint32) int64
TEXT ·fpmxcsr(SB), NOSPLIT, $0-24
MOVQ x+0(FP), X0
MOVQ csr+8(FP), AX
STMXCSR (AX)
LDMXCSR (AX)
CVTSD2SL X0, CX
CVTTSD2SQ X0, DX
MOVL (AX), SI
MOVQ SI, ret+8(FP)
RET
// func fmove(p *float64) float64
TEXT ·fmove(SB), NOSPLIT, $0-16
MOVQ p+0(FP), AX
FMOVD (AX), F0
FMOVD F0, F1
FMOVD F0, (AX)
MOVQ (AX), AX
MOVQ AX, ret+8(FP)
RET
+61
View File
@@ -0,0 +1,61 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// Differential kernel for the arm64 system instructions: barriers,
// cache maintenance, the system register accesses, supervisor calls,
// breakpoints and prefetches. Every function is byte-compared against
// go tool asm.
#include "textflag.h"
// func barriers()
TEXT ·barriers(SB), NOSPLIT, $0-0
DMB $15
DMB $1
DSB $15
DSB $4
ISB $15
ISB $1
RET
// func cacheOps()
TEXT ·cacheOps(SB), NOSPLIT, $0-0
DC ZVA, R4
DC IVAC, R1
DC CVAC, R2
DC CVAU, R3
DC CIVAC, R7
RET
// func sysRegs()
TEXT ·sysRegs(SB), NOSPLIT, $0-0
MRS DCZID_EL0, R3
MRS CNTVCT_EL0, R0
MRS CNTPCT_EL0, R1
MRS CNTFRQ_EL0, R2
MRS MIDR_EL1, R0
MRS ID_AA64PFR0_EL1, R0
MRS ID_AA64ISAR0_EL1, R0
MRS ID_AA64ISAR1_EL1, R0
MRS DIT, R0
MSR $3, SPSel
MSR $9, DAIFSet
MSR $6, DAIFClr
MSR $1, DIT
RET
// func exceptions()
TEXT ·exceptions(SB), NOSPLIT, $0-0
SVC $0
SVC $7165
BRK
BRK $35943
RET
// func prefetch()
TEXT ·prefetch(SB), NOSPLIT, $0-0
PRFM (R0), PLDL1KEEP
PRFM (R3), PLDL3KEEP
PRFM (R4), PSTL1KEEP
PRFM (R2), $25
RET
+209
View File
@@ -0,0 +1,209 @@
// Differential kernel: the arith_add vector slice against the Go
// toolchain's loong64enc1.s rows (gasm v. go tool asm, byte for byte).
#include "textflag.h"
TEXT ·VADDB(SB), NOSPLIT, $0
VADDB V1, V2, V3
RET
TEXT ·VADDH(SB), NOSPLIT, $0
VADDH V1, V2, V3
RET
TEXT ·VADDQ(SB), NOSPLIT, $0
VADDQ V1, V2, V3
RET
TEXT ·XVADDB(SB), NOSPLIT, $0
XVADDB X3, X2, X1
RET
TEXT ·XVADDH(SB), NOSPLIT, $0
XVADDH X3, X2, X1
RET
TEXT ·XVADDW(SB), NOSPLIT, $0
XVADDW X3, X2, X1
RET
TEXT ·XVADDQ(SB), NOSPLIT, $0
XVADDQ X3, X2, X1
RET
TEXT ·VADDBU(SB), NOSPLIT, $0
VADDBU $1, V2
VADDBU $1, V2, V1
RET
TEXT ·VADDHU(SB), NOSPLIT, $0
VADDHU $2, V2, V1
RET
TEXT ·VADDWU(SB), NOSPLIT, $0
VADDWU $3, V2, V1
RET
TEXT ·VADDVU(SB), NOSPLIT, $0
VADDVU $4, V2, V1
RET
TEXT ·XVADDBU(SB), NOSPLIT, $0
XVADDBU $9, X1, X2
RET
TEXT ·XVADDHU(SB), NOSPLIT, $0
XVADDHU $10, X1, X2
RET
TEXT ·XVADDWU(SB), NOSPLIT, $0
XVADDWU $11, X1, X2
RET
TEXT ·XVADDVU(SB), NOSPLIT, $0
XVADDVU $12, X1, X2
RET
TEXT ·VADDWEVHB(SB), NOSPLIT, $0
VADDWEVHB V1, V2, V3
RET
TEXT ·VADDWEVWH(SB), NOSPLIT, $0
VADDWEVWH V1, V2, V3
RET
TEXT ·VADDWEVVW(SB), NOSPLIT, $0
VADDWEVVW V1, V2, V3
RET
TEXT ·VADDWEVQV(SB), NOSPLIT, $0
VADDWEVQV V1, V2, V3
RET
TEXT ·VADDWODHB(SB), NOSPLIT, $0
VADDWODHB V1, V2, V3
RET
TEXT ·VADDWODWH(SB), NOSPLIT, $0
VADDWODWH V1, V2, V3
RET
TEXT ·VADDWODVW(SB), NOSPLIT, $0
VADDWODVW V1, V2, V3
RET
TEXT ·VADDWODQV(SB), NOSPLIT, $0
VADDWODQV V1, V2, V3
RET
TEXT ·XVADDWEVHB(SB), NOSPLIT, $0
XVADDWEVHB X1, X2, X3
RET
TEXT ·XVADDWEVWH(SB), NOSPLIT, $0
XVADDWEVWH X1, X2, X3
RET
TEXT ·XVADDWEVVW(SB), NOSPLIT, $0
XVADDWEVVW X1, X2, X3
RET
TEXT ·XVADDWEVQV(SB), NOSPLIT, $0
XVADDWEVQV X1, X2, X3
RET
TEXT ·XVADDWODHB(SB), NOSPLIT, $0
XVADDWODHB X1, X2, X3
RET
TEXT ·XVADDWODWH(SB), NOSPLIT, $0
XVADDWODWH X1, X2, X3
RET
TEXT ·XVADDWODVW(SB), NOSPLIT, $0
XVADDWODVW X1, X2, X3
RET
TEXT ·XVADDWODQV(SB), NOSPLIT, $0
XVADDWODQV X1, X2, X3
RET
TEXT ·VADDWEVHBU(SB), NOSPLIT, $0
VADDWEVHBU V1, V2, V3
RET
TEXT ·VADDWEVWHU(SB), NOSPLIT, $0
VADDWEVWHU V1, V2, V3
RET
TEXT ·VADDWEVVWU(SB), NOSPLIT, $0
VADDWEVVWU V1, V2, V3
RET
TEXT ·VADDWEVQVU(SB), NOSPLIT, $0
VADDWEVQVU V1, V2, V3
RET
TEXT ·VADDWODHBU(SB), NOSPLIT, $0
VADDWODHBU V1, V2, V3
RET
TEXT ·VADDWODWHU(SB), NOSPLIT, $0
VADDWODWHU V1, V2, V3
RET
TEXT ·VADDWODVWU(SB), NOSPLIT, $0
VADDWODVWU V1, V2, V3
RET
TEXT ·VADDWODQVU(SB), NOSPLIT, $0
VADDWODQVU V1, V2, V3
RET
TEXT ·XVADDWEVHBU(SB), NOSPLIT, $0
XVADDWEVHBU X1, X2, X3
RET
TEXT ·XVADDWEVWHU(SB), NOSPLIT, $0
XVADDWEVWHU X1, X2, X3
RET
TEXT ·XVADDWEVVWU(SB), NOSPLIT, $0
XVADDWEVVWU X1, X2, X3
RET
TEXT ·XVADDWEVQVU(SB), NOSPLIT, $0
XVADDWEVQVU X1, X2, X3
RET
TEXT ·XVADDWODHBU(SB), NOSPLIT, $0
XVADDWODHBU X1, X2, X3
RET
TEXT ·XVADDWODWHU(SB), NOSPLIT, $0
XVADDWODWHU X1, X2, X3
RET
TEXT ·XVADDWODVWU(SB), NOSPLIT, $0
XVADDWODVWU X1, X2, X3
RET
TEXT ·XVADDWODQVU(SB), NOSPLIT, $0
XVADDWODQVU X1, X2, X3
RET
TEXT ·VADDF(SB), NOSPLIT, $0
VADDF V1, V2, V3
RET
TEXT ·VADDD(SB), NOSPLIT, $0
VADDD V1, V2, V3
RET
TEXT ·XVADDF(SB), NOSPLIT, $0
XVADDF X1, X2, X3
RET
TEXT ·XVADDD(SB), NOSPLIT, $0
XVADDD X1, X2, X3
RET

Some files were not shown because too many files have changed in this diff Show More