Compare commits

..
37 Commits
Author SHA1 Message Date
petrbalvin cf6bc6987e fix(ci): pass the upload file to curl, not its interpolation
Test / test (push) Successful in 2m26s
Release / gates (push) Successful in 2m25s
Release / build (amd64, linux) (push) Successful in 1m15s
Release / build (arm64, linux) (push) Successful in 1m16s
Release / build (loong64, linux) (push) Successful in 1m16s
Release / build (riscv64, linux) (push) Successful in 1m16s
Release / release (push) Successful in 34s
2026-09-22 01:31:49 +02:00
petrbalvin ff7b1452b1 docs: name 0.35.0 as the supported release
Test / test (push) Successful in 2m33s
Release / gates (push) Successful in 2m29s
Release / build (amd64, linux) (push) Successful in 1m18s
Release / build (arm64, linux) (push) Successful in 1m20s
Release / build (loong64, linux) (push) Successful in 1m17s
Release / build (riscv64, linux) (push) Successful in 1m26s
Release / release (push) Failing after 35s
2026-09-22 00:52:56 +02:00
petrbalvin 517c1cea25 chore: prepare release v0.35.0
Test / test (push) Successful in 2m33s
Release / gates (push) Failing after 46s
Release / build (amd64, linux) (push) Skipped
Release / build (arm64, linux) (push) Skipped
Release / build (loong64, linux) (push) Skipped
Release / build (riscv64, linux) (push) Skipped
Release / release (push) Skipped
2026-09-22 00:44:10 +02:00
petrbalvin a3e3010e0f fix(cmd): resolve the runtime header test GOROOT from the go command
Test / test (push) Successful in 2m39s
2026-09-21 22:46:07 +02:00
petrbalvin 057c4eb545 docs: complete the release delta in the changelog and readme 2026-09-21 22:45:56 +02:00
petrbalvin f720381d43 feat(asm): the segment-absolute and crash-store forms GOROOT writes
Test / test (push) Failing after 2m28s
Assisted-by: GLM 5.3 Flash
2026-09-21 22:19:53 +02:00
petrbalvin 2c9042d62c feat(asm): PCALIGN alignment on amd64
Assisted-by: GLM 5.3 Flash
2026-09-21 22:00:30 +02:00
petrbalvin 82ef289d3a feat(asm): the immediate multiply and arm64 indirect branches GOROOT writes
Assisted-by: GLM 5.3 Flash
2026-09-21 21:50:11 +02:00
petrbalvin 7246b0e002 feat(asm): the TLS access pair in the toolchain's one-instruction form
Assisted-by: GLM 5.3 Flash
2026-09-21 21:35:15 +02:00
petrbalvin 8cfd40aac8 feat(asm): the operand forms and defines GOROOT writes
Assisted-by: GLM 5.3 Flash
2026-09-21 21:17:34 +02:00
petrbalvin 5382c9a8e4 feat(audit): list every corpus failure per architecture 2026-09-21 21:17:34 +02:00
petrbalvin 53de91b2df docs(asm): describe the four target architectures
Test / test (push) Failing after 2m23s
Assisted-by: GLM 5.3 Flash
2026-09-21 20:15:55 +02:00
petrbalvin 8a36af7c7d docs(asm): generate the instruction appendices
Assisted-by: GLM 5.3 Flash
2026-09-21 20:15:55 +02:00
petrbalvin e9789ce3f4 chore(arch): regenerate the instruction tables 2026-09-21 20:15:55 +02:00
petrbalvin 837231c068 docs(asm): open the assembly language reference
Assisted-by: GLM 5.3 Flash
2026-09-21 19:49:04 +02:00
petrbalvin 95025be1bc docs(changelog): describe the encoder entries by content
Test / test (push) Failing after 2m33s
2026-09-21 19:20:01 +02:00
petrbalvin 03a964bb2d docs(goobj): document the GOOBJ object file format 2026-09-21 19:19:53 +02:00
petrbalvin 123a16e346 docs(readme): state the documentation goal 2026-09-21 18:35:27 +02:00
petrbalvin 9701812bee docs: changelog for the completeness waves
Test / test (push) Failing after 3m6s
Assisted-by: GLM 5.3 Flash
2026-09-21 02:04:44 +02:00
petrbalvin 29ac03468e feat(amd64): floating-point immediates through a synthesised pool
Assisted-by: GLM 5.3 Flash
2026-09-21 02:04:44 +02:00
petrbalvin bfb7701db1 feat(amd64): emit the quad-register EVEX families
Assisted-by: GLM 5.3 Flash
2026-09-21 02:02:19 +02:00
petrbalvin e8b6ff5d7c fix(parser): fold a signed parenthesised displacement expression
Test / test (push) Failing after 2m21s
Assisted-by: GLM 5.3 Flash
2026-09-21 00:45:33 +02:00
petrbalvin 1456907000 feat(riscv64,loong64): operand tail, float DATA and honest port classification
Assisted-by: GLM 5.3 Flash
2026-09-21 00:44:47 +02:00
petrbalvin ec1c521187 feat(cmd): GOOS-aware headers, audit battery shapes and semicolon spacing
Test / test (push) Failing after 2m21s
Assisted-by: GLM 5.3 Flash
2026-09-20 22:02:46 +02:00
petrbalvin a7744c24bd fix(parser): substitute macro parameters behind element selectors
Assisted-by: GLM 5.3 Flash
2026-09-20 22:02:19 +02:00
petrbalvin 522e6f2ae8 feat(parser): bracket register ranges, index-only VSIB and bare trailing immediates
Assisted-by: GLM 5.3 Flash
2026-09-20 22:02:19 +02:00
petrbalvin 81d4bd81e4 test(verify): register the wave kernels
Test / test (push) Failing after 2m20s
Assisted-by: GLM 5.3 Flash
2026-09-20 21:17:31 +02:00
petrbalvin 687678a2ea feat(elf): emit data relocations on arm64, riscv64 and loong64
Assisted-by: GLM 5.3 Flash
2026-09-20 21:17:20 +02:00
petrbalvin b0f9071bf5 feat(arm64): whole-vector moves, bookkeeping ops and truncating-move lowering
Assisted-by: GLM 5.3 Flash
2026-09-20 21:17:20 +02:00
petrbalvin 81e2673923 feat(amd64): encode the AVX-512 and BMI corpus families
Assisted-by: GLM 5.3 Flash
2026-09-20 21:17:20 +02:00
petrbalvin 75e9fd771b feat(parser): split plain statements on semicolons in the raw parse
Test / test (push) Failing after 2m30s
Assisted-by: GLM 5.3 Flash
2026-09-20 19:15:35 +02:00
petrbalvin 863926abd6 test(verify): register the loong64 vector kernels
Assisted-by: GLM 5.3 Flash
2026-09-20 19:15:05 +02:00
petrbalvin 241e7256f6 fix(arm64): reject bare BTI with a diagnostic and accept the full family
Assisted-by: GLM 5.3 Flash
2026-09-20 19:15:05 +02:00
petrbalvin 6556b85abf feat(asm): symbol-valued DATA, division slash in symbols and plain semicolons
Assisted-by: GLM 5.3 Flash
2026-09-20 19:15:05 +02:00
petrbalvin 289cabe993 feat(loong64): encode the full LSX and LASX table
Assisted-by: GLM 5.3 Flash
2026-09-20 19:14:43 +02:00
petrbalvin d6cf7cfa44 fix(format): keep statement separators and canonical macro bodies
Assisted-by: GLM 5.3 Flash
2026-09-20 19:14:43 +02:00
petrbalvin 4cc2f0eba5 feat(cmd): generate go_asm.h for package-context assembly
Assisted-by: GLM 5.3 Flash
2026-09-20 19:14:43 +02:00
116 changed files with 16970 additions and 340 deletions
+4 -1
View File
@@ -342,7 +342,10 @@ jobs:
my @cmd = (q{curl}, q{-sS}, q{-o}, q{/dev/null}, q{-w}, q{%{http_code}},
q{-H}, qq{Authorization: token $ENV{GITEA_TOKEN}},
q{-H}, q{Content-Type: application/octet-stream},
q{-X}, q{POST}, q{--data-binary}, qq{@$path},
# The @ must not sit inside a qq{} string: there it starts an
# array interpolation and the upload body collapses to empty,
# which Gitea stores as a 201-created zero-byte attachment.
q{-X}, q{POST}, q{--data-binary}, q{@} . $path,
qq{$ENV{GITEA_SERVER_URL}/api/v1/repos/$ENV{GITEA_REPOSITORY}/releases/$id/assets?name=$name});
open(my $curl, q{-|}, @cmd) or die qq{curl: $!};
my $code = <$curl>;
+111 -9
View File
@@ -9,6 +9,31 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
### Added
-
## [0.35.0] - 2026-09-22
### Added
- **The go_asm.h generator.** `gasm asm` generates the package's go_asm.h
itself when an assembly file includes it: the Go files beside the source
are type-checked for the target architecture and the constants and field
offsets become assembler defines, so package-context files assemble with
no compiler and no `go build` in the loop. `-GOOS` selects the
type-checking GOOS for GOOS-specific files, and the corpus audit derives
the GOOS from the file name.
- **ELF data relocations on arm64, riscv64 and loong64.** `gasm asm
--format elf` emits `.rela.data` for symbol-valued DATA initialisers on
every architecture (amd64 carried them already), so standalone ELF
objects link on all four targets.
- **Corpus failure listing.** `gasm audit-instructions --corpus --list`
prints every failing file with its failure reason, per architecture,
instead of one representative file per reason.
- **DATA with symbol values and relaxed symbol spellings.** DATA
initialisers accept `$symbol(SB)` values, laid down as an absolute
relocation at the data field (GOOBJ on all four architectures and ELF
on all four as of this release), and U+2215 is accepted inside symbol
package paths.
- **Macro expansion and include splicing.** `gasm asm`, `gasm diff` and
`gasm audit-instructions` now preprocess assembly the way the
toolchain does: object and parameterised `#define` macros expand at
@@ -19,7 +44,8 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
(`$(32-7)`, `$~63`, `(index*4)(base)`) fold at parse. Expansion
happens only on the assembly path: `gasm lint`, `gasm fmt` and the
language server keep reading the raw file.
- **The GOROOT instruction wave, part 1.** The encoder now covers the
- **Encoder coverage: the instruction families GOROOT's real code
uses.** The encoder now covers the
instruction families GOROOT's real code uses that gasm lacked,
byte-verified against `go tool asm`: on amd64 the carry ALU, the
atomics (CMPXCHG, XADD, XCHG), AES-NI, SHA-1/256, PCLMULQDQ, CRC32,
@@ -35,14 +61,90 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
Also fixed on the way: arm64 `CASD`/`CASW` lacked an opcode bit, and
riscv64 `VSETVLI` with an immediate length now canonicalises to
`vsetivli` as the toolchain does.
- **The corpus audit measures honestly.** Files named for Go ports gasm
does not target (arm, 386, s390x, ...) are no longer attempted for the
four supported architectures (no supported build compiles them), and
the headline rate is reported over attemptable files: 136 of 433 on
the full corpus (31.4 %), 135 of 383 on real code (35.2 %), from the
127 that the previous release measured. The probe battery that
decides encodability gained the operand shapes the new families use.
-
- **Encoder coverage: quad-register AVX-512 and floating-point
immediates.** The encoder gains the
quad-register AVX-512 families (4FMAPS, 4FNMADD, 4VNNIW, VP4DPWSSD,
VP4DPWSSDS) with the register list riding the inverted V'VVVV field,
floating-point immediates on the SSE scalar moves and arithmetic
(the constant lands in a synthesised read-only pool, a positive zero
collapses to XORPS exactly as the toolchain does), accept-and-ignore
FUNCDATA and PCDATA, three-operand double shifts, static-symbol
operands for the legacy SSE moves, and the pooled 64-bit immediate
materialisation on riscv64. The parser carries bracketed register
ranges, index-only VSIB memory operands and bare trailing immediates;
macro substitution reaches parameters used with element suffixes
(`A.S4`), and `;` separates statements in plain files.
- **Per-architecture reference pages.** [docs/asm/](docs/asm/README.md)
gains AMD64, ARM64, RISCV64 and LOONG64: the register files and the
roles the ABI fixes, addressing, operand order with every special form,
constants and materialisation, alignment, fences and the relocations
each target emits. An instruction inventory appendix per architecture
is generated from the toolchain's own tables by `just gen`, and the
regenerated tables recognise 147 more mnemonics than the previous
release carried (arm64 107, riscv64 31, loong64 9).
- **The Plan 9 assembly language reference.** [docs/asm/](docs/asm/README.md)
opens the complete language reference with its common core: the lexicon,
statement structure and constant expressions, the operand grammar with
the pseudo-registers and symbol naming, the directives and the function
flag vocabulary, preprocessing with `#define` and `#include`, and the
Go-embedded layer (ABI0, prototypes, `go_asm.h`, `funcdata.h` and the
runtime contract). Every claim is verified against `go tool asm` of
Go 1.27.1 and gasm's differential tests; the per-architecture pages and
generated instruction appendices follow.
- **GOOBJ format specification.** [docs/GOOBJ.md](docs/GOOBJ.md)
documents the Go object file format in full: both containers, the 96
byte header and all 19 blocks, every structure with its byte
offsets, symbol kinds and flag bits, all 106 relocation types with
the weak variants, aux symbols, the FuncInfo payload, the pc-value
table encoding, the content hashes and the builtin table, all
verified byte for byte against objects produced by Go 1.27.1's own
tools.
### Changed
- **The corpus audit measures like a build.** Files named for a Go port
gasm does not target (arm, 386, s390x, ...) are never attempted, because
no supported build compiles them; the GOOS comes from the file name; and
each target's go_asm.h is generated on the fly. The headline is reported
over attemptable files: 291 of 353 on the full corpus (82.4 %) assemble
for every target architecture and 295 of 303 on real code (97.4 %),
against 108 of 627 over all files (17.2 %) that the previous release
measured.
### Fixed
- **The operand forms GOROOT writes.** Numeric PC-relative jumps
(`JEQ 2(PC)`, the park loop `JMP 0(PC)`) resolve with the toolchain's
own instruction counting and fold jump-to-jump chains exactly as its
branch optimiser does; symbol immediates (`MOVQ $sym(SB), AX`)
assemble to the toolchain's RIP-relative LEA with an R_PCREL
relocation; negated constant expressions in operands (`ADJSP
$-(REGS - 8)`, the shape the cgo ABI macros write) fold; the immediate
multiply (`IMULQ $1000000000, AX`) encodes with the toolchain's
0x69/0x6B selection; the TLS access pair assembles as the toolchain's
one-instruction form (the bare `MOVQ TLS, r` load nops out and
`off(r)(TLS*1)` folds to the segment-prefixed absolute whose disp32
carries the R_TLSLE relocation, per-GOOS); arm64 accepts the
bare-register indirect branch (`BL R9` beside `BL (R9)`, both BLR) and
the zero-immediate store (`MOVD $0, mem` through the zero register,
rejecting non-zero immediates as the toolchain does); `PCALIGN` now
aligns on amd64, padding with the toolchain's greedy
single-instruction NOPs; the segment-absolute forms (`MOVQ 0x30(GS),
AX` and the store direction) and the absolute crash-store
(`MOVL $0xf1, 0xf1`) encode; and `gasm asm` predefines the
`GOARCH_<arch>` and `GOOS_<goos>` macros the go command passes to
`go tool asm`, so GOROOT headers' `#ifdef GOARCH_amd64` platform
blocks (`go_tls.h`'s `get_tls` and friends) select as intended. The
GOROOT corpus measure moves to 291 of 353 files assembling for every
target architecture (82.4 %), 97.4 % of the real-code corpus, from
70.8 % and 82.2 %.
- **Tool corrections across the pipeline.** The formatter keeps square
brackets in SIMD operands, statement separators and canonical macro
bodies; the linter drops false positives on shift counts, SETcc
spellings and ABIInternal references; the lexer treats a trailing
carriage return as a line end so comment text stays idempotent; and
arm64 rejects bare BTI with a diagnostic while accepting the full
family.
## [0.34.0] - 2026-09-20
+44 -6
View File
@@ -81,7 +81,10 @@ to give that syntax the tooling it deserves.
GOOBJ format, which needs the installed toolchain and which `go build`
consumes in place of the toolchain's output. Framed functions get the
stack-split guard and the morestack block, byte-identical to the
toolchain's, so split functions link too.
toolchain's, so split functions link too. The assembler preprocesses
like the toolchain (`#define`, `#include` with `-I`, `#ifdef`), generates
`go_asm.h` from the package's Go files, and carries `PCALIGN`, the
`LOCK`/`REP` prefixes and the literal-data pseudo-ops.
- **Disassembler.** `gasm dis` lists a `.s` file's functions at their real
offsets after assembling, or disassembles raw bytes from a file or stdin.
- **Dynamic verification.** `gasm verify` JIT-loads assembled functions into
@@ -110,9 +113,9 @@ Four architectures, the four that matter in practice:
| Architecture | GOARCH | File suffix | Instructions recognised |
|--------------|-------------|--------------|---------------------------------------------|
| AMD64 | `amd64` | `_amd64.s` | 1600 + common opcodes + traditional aliases |
| ARM64 | `arm64` | `_arm64.s` | 538 + common opcodes |
| RISC-V | `riscv64` | `_riscv64.s` | 961 + common opcodes |
| LoongArch | `loong64` | `_loong64.s` | 799 + common opcodes |
| ARM64 | `arm64` | `_arm64.s` | 645 + common opcodes |
| RISC-V | `riscv64` | `_riscv64.s` | 992 + common opcodes |
| LoongArch | `loong64` | `_loong64.s` | 808 + common opcodes |
"Common opcodes" are the instructions shared by every architecture (`RET`,
`JMP`, `NOP`, `CALL`, `TEXT`, `FUNCDATA`, `PCDATA`, ...). AMD64 additionally
@@ -124,8 +127,8 @@ can emit today is narrower, and a recognised but unencodable instruction is
reported as an explicit error, never as a wrong byte.
The same measurement runs over GOROOT's whole assembly corpus:
`gasm audit-instructions --corpus` reports 136 of 433 attemptable files
(31.4 %) assembling for every target architecture today (files named for
`gasm audit-instructions --corpus` reports 291 of 353 attemptable files
(82.4 %) assembling for every target architecture today (files named for
other Go ports are counted but never attempted), with the top failure
reasons per architecture; the number moves with every release.
@@ -157,6 +160,39 @@ been compiled and read, never executed. Its architecture-neutral units
run under `go test ./...`, which the race workflow and a manual run
perform; the default `just test` gate does not sweep `./debug/...`.
## The documentation goal
The toolkit is the primary goal. The secondary one is documentation: a
specification of the Plan 9 assembly language and of the GOOBJ object
format that is 100 % complete, detailed enough to implement against,
and written to a professional standard. These are the two subjects this
project works with every day, and they are the two for which no usable
documentation exists.
Go documents the language on a single page, "A Quick Guide to Go's
Assembler", which carries no section for loong64, one of the four
architectures gasm supports, and covers a fraction of what each
assembler accepts. What exists beyond it lives as comments inside the
toolchain's internal source: per-architecture reference manuals for
arm64, ppc64, riscv64 and loong64, written for the toolchain's own
maintainers rather than for an outside reader, and none at all for
amd64. GOOBJ fares worst of all. The format that `go build` consumes
has no specification anywhere: it is described by a comment in an
internal package, it is not a stable interface, and it can change with
any toolchain release.
The gap is therefore filled the only way it can be filled: by reverse
engineering the toolchain itself, the same work the encoders already
perform. Most of the documentation can come from nowhere else, and it
is written as that knowledge is produced during development. It is
verified the way the code is verified: an encoding documented here is
one that differential tests against `go tool asm` confirm
byte-for-byte, and a format field documented here is one the linker
demonstrably reads. The work has begun: [docs/GOOBJ.md](docs/GOOBJ.md)
specifies the object file format completely, and
[docs/asm/README.md](docs/asm/README.md) opens the language reference
with its common core. The per-architecture pages follow.
## Direction
The plan, in the order it is being worked:
@@ -282,6 +318,8 @@ recipe.
~/.local/share/man (MANDIR overrides); `just uninstall-man` removes
them
- [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md): components and data flow
- [docs/GOOBJ.md](docs/GOOBJ.md): the GOOBJ object file format specification
- [docs/asm/](docs/asm/README.md): the Plan 9 assembly language reference
- [docs/DEVELOPMENT.md](docs/DEVELOPMENT.md): development setup and recipes
- [CHANGELOG.md](CHANGELOG.md): release history
+1 -1
View File
@@ -7,7 +7,7 @@ releases do not receive them.
| Version | Supported |
|---|---|
| 0.34.0 | yes |
| 0.35.0 | yes |
| older releases | no |
## Reporting a vulnerability
+101 -3
View File
@@ -8,6 +8,10 @@
// names so gasm-devkit supports every instruction the real assembler does,
// with no hand-maintained (and therefore inevitably incomplete) lists.
//
// The same data feeds the generated instruction appendices of the assembly
// language reference, docs/asm/INSTRUCTIONS-<ARCH>.md, so that the reference
// cannot drift from the tables it documents.
//
// Usage (via the justfile):
//
// just gen
@@ -26,6 +30,9 @@ import (
"path/filepath"
"sort"
"strings"
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
"sourcedock.dev/petrbalvin/gasm-devkit/asm"
)
// archDirs maps a gasm-devkit architecture name to its obj sub-directory.
@@ -39,11 +46,30 @@ var archDirs = []struct {
{"loong64", "loong64"},
}
// docPages maps an architecture to its generated appendix in the language
// reference. The amd64 page carries a per-mnemonic encodability column,
// decided by asm.Encodable, which mirrors the encoder's own dispatch; the
// other targets have no single cheap predicate, so their pages carry the
// inventory and point at the live measurement instead.
var docPages = []struct {
arch arch.Arch
title string
file string
anames string
encodable bool
}{
{arch.AMD64, "AMD64", "INSTRUCTIONS-AMD64.md", "cmd/internal/obj/x86/anames.go", true},
{arch.ARM64, "ARM64", "INSTRUCTIONS-ARM64.md", "cmd/internal/obj/arm64/anames.go", false},
{arch.RISCV, "RISC-V 64", "INSTRUCTIONS-RISCV64.md", "cmd/internal/obj/riscv/anames.go", false},
{arch.LOONG64, "LoongArch 64", "INSTRUCTIONS-LOONG64.md", "cmd/internal/obj/loong64/anames.go", false},
}
func main() {
goroot := strings.TrimSpace(runGoEnvGOROOT())
if goroot == "" {
fatal("could not determine GOROOT")
}
version := strings.TrimSpace(runGoEnv("GOVERSION"))
// The common opcodes shared by every architecture (RET, JMP, NOP, CALL,
// TEXT, FUNCDATA, …) live in cmd/internal/obj/util.go.
commonPath := filepath.Join(goroot, "src", "cmd", "internal", "obj", "util.go")
@@ -57,16 +83,24 @@ func main() {
}
fmt.Printf("%-8s %4d instructions -> arch/common_gen.go\n", "common", len(common))
names := map[string][]string{}
for _, a := range archDirs {
path := filepath.Join(goroot, "src", "cmd", "internal", "obj", a.sub, "anames.go")
names, err := extractInstrs(path)
names[a.arch], err = extractInstrs(path)
if err != nil {
fatal("extract %s: %v", a.arch, err)
}
if err := writeGen(a.arch, a.sub, names); err != nil {
if err := writeGen(a.arch, a.sub, names[a.arch]); err != nil {
fatal("write %s: %v", a.arch, err)
}
fmt.Printf("%-8s %4d instructions -> arch/%s_gen.go\n", a.arch, len(names), a.arch)
fmt.Printf("%-8s %4d instructions -> arch/%s_gen.go\n", a.arch, len(names[a.arch]), a.arch)
}
for _, p := range docPages {
if err := writeDocPage(p.arch, p.title, p.file, p.anames, version, p.encodable); err != nil {
fatal("write %s: %v", p.file, err)
}
fmt.Printf("%-8s -> docs/asm/%s\n", p.arch, p.file)
}
}
@@ -172,6 +206,61 @@ func writeGen(arch, sub string, names []string) error {
return os.WriteFile(filepath.Join("arch", arch+"_gen.go"), []byte(b.String()), 0o644)
}
// writeDocPage emits docs/asm/<file>, the generated instruction appendix of
// the language reference for one architecture: every mnemonic the toolchain
// accepts, with the curated summary where the architecture table carries one
// and, on amd64, a per-mnemonic encodability column.
func writeDocPage(a arch.Arch, title, file, anames, version string, encodable bool) error {
table := arch.ForArch(a)
instrs := table.Instructions()
var b strings.Builder
b.WriteString("# " + title + ": instruction inventory\n\n")
b.WriteString("Generated by gasm-devkit's `_gen` from the Go toolchain's instruction table\n")
b.WriteString("(`" + anames + "`, " + version + "); DO NOT EDIT. This page lists every mnemonic\n")
b.WriteString("`go tool asm` accepts on this target, which is the upper bound of the\n")
b.WriteString("language on it: a name absent here is not an instruction of the target,\n")
b.WriteString("and a name present here may still be one gasm's encoder cannot emit yet.\n\n")
encodableCount := 0
if encodable {
b.WriteString("The `gasm encodes` column reports whether gasm's encoder can emit the\n")
b.WriteString("mnemonic today; the gap is the encoder backlog, measured live by\n")
b.WriteString("`gasm audit-instructions`.\n\n")
b.WriteString("| Mnemonic | gasm encodes | Notes |\n")
b.WriteString("|---|---|---|\n")
for _, in := range instrs {
ok := asm.Encodable(in.Name)
if ok {
encodableCount++
}
b.WriteString("| `" + in.Name + "` | " + yesNo(ok) + " | " + in.Summary + " |\n")
}
b.WriteString("\n")
fmt.Fprintf(&b, "Recognised: %d mnemonics. gasm encodes: %d.\n", len(instrs), encodableCount)
} else {
b.WriteString("The inventory carries no per-mnemonic encoder column: on this target\n")
b.WriteString("encodability is decided per operand shape, and the live measured\n")
b.WriteString("coverage is reported by `gasm audit-instructions`.\n\n")
b.WriteString("| Mnemonic | Notes |\n")
b.WriteString("|---|---|\n")
for _, in := range instrs {
b.WriteString("| `" + in.Name + "` | " + in.Summary + " |\n")
}
b.WriteString("\n")
fmt.Fprintf(&b, "Recognised: %d mnemonics.\n", len(instrs))
}
return os.WriteFile(filepath.Join("docs", "asm", file), []byte(b.String()), 0o644)
}
// yesNo renders a boolean as the word the appendix tables use.
func yesNo(v bool) string {
if v {
return "yes"
}
return "no"
}
func runGoEnvGOROOT() string {
out, err := exec.Command("go", "env", "GOROOT").Output()
if err != nil {
@@ -180,6 +269,15 @@ func runGoEnvGOROOT() string {
return string(out)
}
// runGoEnv runs `go env` for a single variable.
func runGoEnv(name string) string {
out, err := exec.Command("go", "env", name).Output()
if err != nil {
return ""
}
return string(out)
}
func fatal(format string, args ...any) {
fmt.Fprintf(os.Stderr, "gen: "+format+"\n", args...)
os.Exit(1)
+1
View File
@@ -33,6 +33,7 @@ func arm64Registers() []Register {
for i := 0; i <= 30; i++ {
add(fmt.Sprintf("R%d", i), GPR, "64-bit general-purpose register")
}
add("R18_PLATFORM", GPR, "R18 under its toolchain-reserved Windows name (an alias of R18)")
add("ZR", Special, "zero register (reads as 0)")
add("SP", Special, "stack pointer")
add("LR", Special, "link register (alias of R30)")
+107
View File
@@ -364,6 +364,8 @@ var arm64GeneratedInstrs = []string{
"REVW",
"ROR",
"RORW",
"RPRFM",
"SB",
"SBC",
"SBCS",
"SBCSW",
@@ -477,23 +479,68 @@ var arm64GeneratedInstrs = []string{
"UXTH",
"UXTHW",
"UXTW",
"VABS",
"VADD",
"VADDP",
"VADDV",
"VAND",
"VBCAX",
"VBIC",
"VBIF",
"VBIT",
"VBSL",
"VCLS",
"VCLZ",
"VCMEQ",
"VCMGE",
"VCMGT",
"VCMHI",
"VCMHS",
"VCMLE",
"VCMLT",
"VCMTST",
"VCNT",
"VDUP",
"VEOR",
"VEOR3",
"VEXT",
"VFABS",
"VFADD",
"VFADDP",
"VFCMEQ",
"VFCMGE",
"VFCMGT",
"VFCMLE",
"VFCMLT",
"VFCVTL",
"VFCVTL2",
"VFCVTN",
"VFCVTN2",
"VFCVTZS",
"VFCVTZU",
"VFDIV",
"VFMAX",
"VFMAXNM",
"VFMAXNMP",
"VFMAXNMV",
"VFMAXP",
"VFMAXV",
"VFMIN",
"VFMINNM",
"VFMINNMP",
"VFMINNMV",
"VFMINP",
"VFMINV",
"VFMLA",
"VFMLS",
"VFMUL",
"VFNEG",
"VFRINTM",
"VFRINTN",
"VFRINTP",
"VFRINTZ",
"VFSQRT",
"VFSUB",
"VLD1",
"VLD1R",
"VLD2",
@@ -502,11 +549,17 @@ var arm64GeneratedInstrs = []string{
"VLD3R",
"VLD4",
"VLD4R",
"VMLA",
"VMLS",
"VMOV",
"VMOVD",
"VMOVI",
"VMOVQ",
"VMOVS",
"VMUL",
"VNEG",
"VNOT",
"VORN",
"VORR",
"VPMULL",
"VPMULL2",
@@ -515,14 +568,47 @@ var arm64GeneratedInstrs = []string{
"VREV16",
"VREV32",
"VREV64",
"VSCVTF",
"VSHADD",
"VSHL",
"VSHRN",
"VSHRN2",
"VSLI",
"VSMAX",
"VSMAXP",
"VSMAXV",
"VSMIN",
"VSMINP",
"VSMINV",
"VSMLAL",
"VSMLAL2",
"VSMLSL",
"VSMLSL2",
"VSMULL",
"VSMULL2",
"VSQABS",
"VSQADD",
"VSQNEG",
"VSQSHL",
"VSQSUB",
"VSQXTN",
"VSQXTN2",
"VSQXTUN",
"VSQXTUN2",
"VSRHADD",
"VSRI",
"VSRSHR",
"VSSHL",
"VSSHLL",
"VSSHLL2",
"VSSHR",
"VST1",
"VST2",
"VST3",
"VST4",
"VSUB",
"VSXTL",
"VSXTL2",
"VTBL",
"VTBX",
"VTRN1",
@@ -530,8 +616,27 @@ var arm64GeneratedInstrs = []string{
"VUADDLV",
"VUADDW",
"VUADDW2",
"VUCVTF",
"VUHADD",
"VUMAX",
"VUMAXP",
"VUMAXV",
"VUMIN",
"VUMINP",
"VUMINV",
"VUMLAL",
"VUMLAL2",
"VUMLSL",
"VUMLSL2",
"VUMULL",
"VUMULL2",
"VUQADD",
"VUQSHL",
"VUQSUB",
"VUQXTN",
"VUQXTN2",
"VURHADD",
"VUSHL",
"VUSHLL",
"VUSHLL2",
"VUSHR",
@@ -541,6 +646,8 @@ var arm64GeneratedInstrs = []string{
"VUZP1",
"VUZP2",
"VXAR",
"VXTN",
"VXTN2",
"VZIP1",
"VZIP2",
"WFE",
+9
View File
@@ -152,6 +152,8 @@ var loong64GeneratedInstrs = []string{
"FNMADDF",
"FNMSUBD",
"FNMSUBF",
"FRINTD",
"FRINTF",
"FSCALEBD",
"FSCALEBF",
"FSEL",
@@ -177,7 +179,10 @@ var loong64GeneratedInstrs = []string{
"FTINTWF",
"JIRL",
"LL",
"LLACQV",
"LLACQW",
"LLV",
"LLW",
"LU12IW",
"LU32ID",
"LU52ID",
@@ -248,7 +253,11 @@ var loong64GeneratedInstrs = []string{
"ROTR",
"ROTRV",
"SC",
"SCQ",
"SCRELV",
"SCRELW",
"SCV",
"SCW",
"SGT",
"SGTU",
"SLL",
+31
View File
@@ -81,6 +81,9 @@ var riscvGeneratedInstrs = []string{
"CLD",
"CLDSP",
"CLI",
"CLMUL",
"CLMULH",
"CLMULR",
"CLUI",
"CLW",
"CLWSP",
@@ -95,13 +98,20 @@ var riscvGeneratedInstrs = []string{
"CSDSP",
"CSLLI",
"CSRAI",
"CSRC",
"CSRCI",
"CSRLI",
"CSRR",
"CSRRC",
"CSRRCI",
"CSRRS",
"CSRRSI",
"CSRRW",
"CSRRWI",
"CSRS",
"CSRSI",
"CSRW",
"CSRWI",
"CSUB",
"CSUBW",
"CSW",
@@ -259,6 +269,7 @@ var riscvGeneratedInstrs = []string{
"ORCB",
"ORI",
"ORN",
"PAUSE",
"RDCYCLE",
"RDINSTRET",
"RDTIME",
@@ -322,6 +333,8 @@ var riscvGeneratedInstrs = []string{
"VADDVI",
"VADDVV",
"VADDVX",
"VANDNVV",
"VANDNVX",
"VANDVI",
"VANDVV",
"VANDVX",
@@ -329,8 +342,17 @@ var riscvGeneratedInstrs = []string{
"VASUBUVX",
"VASUBVV",
"VASUBVX",
"VBREV8V",
"VBREVV",
"VCLMULHVV",
"VCLMULHVX",
"VCLMULVV",
"VCLMULVX",
"VCLZV",
"VCOMPRESSVM",
"VCPOPM",
"VCPOPV",
"VCTZV",
"VDIVUVV",
"VDIVUVX",
"VDIVVV",
@@ -743,10 +765,16 @@ var riscvGeneratedInstrs = []string{
"VREMUVX",
"VREMVV",
"VREMVX",
"VREV8V",
"VRGATHEREI16VV",
"VRGATHERVI",
"VRGATHERVV",
"VRGATHERVX",
"VROLVV",
"VROLVX",
"VRORVI",
"VRORVV",
"VRORVX",
"VRSUBVI",
"VRSUBVX",
"VS1RV",
@@ -950,6 +978,9 @@ var riscvGeneratedInstrs = []string{
"VWMULVX",
"VWREDSUMUVS",
"VWREDSUMVS",
"VWSLLVI",
"VWSLLVV",
"VWSLLVX",
"VWSUBUVV",
"VWSUBUVX",
"VWSUBUWV",
+101
View File
@@ -245,3 +245,104 @@ func main() {
t.Error("binary does not contain expected symbol")
}
}
// TestGOObjectAARCH64DataSymbolLink does for symbol-valued DATA fields what
// the rt0 files do ("DATA _rt0…lib+0(SB)/8, $_rt0…lib(SB)"): the gasm object
// carries an R_ADDR against the file's own TEXT symbol, the toolchain links
// it, and the binary is checked for the symbol (no arm64 host to run it).
func TestGOObjectAARCH64DataSymbolLink(t *testing.T) {
goBin, err := exec.LookPath("go")
if err != nil {
t.Skip("no Go toolchain available")
}
dir := t.TempDir()
asmSrc := `#include "textflag.h"
GLOBL entry(SB), NOPTR, $8
DATA entry+0(SB)/8, $·keepme(SB)
TEXT ·keepme(SB), NOSPLIT, $0-0
RET
TEXT ·entryptr(SB), NOSPLIT, $0-8
MOVD entry+0(SB), R4
MOVD R4, ret+0(FP)
RET
`
if err := os.WriteFile(filepath.Join(dir, "main_arm64.s"), []byte(asmSrc), 0o644); err != nil {
t.Fatal(err)
}
mainSrc := `package main
func keepme()
func entryptr() uintptr
func main() {
if entryptr() == 0 {
panic("the entry word is empty")
}
}
`
if err := os.WriteFile(filepath.Join(dir, "main.go"), []byte(mainSrc), 0o644); err != nil {
t.Fatal(err)
}
if err := os.WriteFile(filepath.Join(dir, "go.mod"), []byte("module a64dlink\n\ngo 1.21\n"), 0o644); err != nil {
t.Fatal(err)
}
build := exec.Command(goBin, "build", "-x", "-work", "-o", filepath.Join(dir, "prog"), ".")
build.Dir = dir
build.Env = append(os.Environ(), "GOARCH=arm64")
buildLog, err := build.CombinedOutput()
if err != nil {
t.Fatalf("baseline build: %v\n%s", err, buildLog)
}
var work, linkLine, asmObj string
for line := range strings.SplitSeq(string(buildLog), "\n") {
switch {
case strings.HasPrefix(line, "WORK="):
work = strings.TrimPrefix(line, "WORK=")
case strings.Contains(line, "/asm ") && strings.Contains(line, "main_arm64.s") && !strings.Contains(line, "-gensymabis"):
asmObj = fieldAfter(line, "-o")
case strings.Contains(line, "/link ") && strings.Contains(line, "-importcfg"):
linkLine = line
}
}
if work == "" || asmObj == "" || linkLine == "" {
t.Skipf("could not parse build log (work=%q asmObj=%q link=%q)", work, asmObj, linkLine)
}
defer os.RemoveAll(work)
asmObj = strings.ReplaceAll(asmObj, "$WORK", work)
linkLine = strings.ReplaceAll(linkLine, "$WORK", work)
src, err := os.ReadFile(filepath.Join(dir, "main_arm64.s"))
if err != nil {
t.Fatal(err)
}
f, errs := parser.Parse("main_arm64.s", string(src))
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileARM64(f)
if err != nil {
t.Fatalf("AssembleFileARM64: %v", err)
}
gasmObj, err := img.GOObjectAARCH64("a64dlink", "main_arm64.s")
if err != nil {
t.Fatalf("GOObjectAARCH64: %v", err)
}
if err := os.WriteFile(asmObj, gasmObj, 0o644); err != nil {
t.Fatalf("write gasm object: %v", err)
}
linkCmd := exec.Command("bash", "-c", "cd "+dir+" && "+linkLine)
linkCmd.Env = append(os.Environ(), "GOARCH=arm64")
if out, err := linkCmd.CombinedOutput(); err != nil {
t.Fatalf("re-link with gasm object: %v\n%s", err, out)
}
binData, err := os.ReadFile(filepath.Join(dir, "prog"))
if err != nil {
t.Fatal(err)
}
if !strings.Contains(string(binData), "keepme") {
t.Error("binary does not contain the keepme symbol")
}
}
+160 -31
View File
@@ -263,9 +263,10 @@ func arm64InstrSize(instr *ast.Instr, fi arm64FrameInfo, pos int) int {
return 4
}
}
// The funcdata pseudo-statements contribute no bytes.
// The funcdata pseudo-statements contribute no bytes, the expanded
// FUNCDATA/PCDATA forms included.
switch mnem {
case "NO_LOCAL_POINTERS", "GO_ARGS", "GO_RESULTS_INITIALIZED":
case "NO_LOCAL_POINTERS", "GO_ARGS", "GO_RESULTS_INITIALIZED", "END", "FUNCDATA", "PCDATA":
return 0
}
switch mnem {
@@ -348,10 +349,27 @@ func encodeARM64Instr(instr *ast.Instr, pc int, offsets map[string]int, fi arm64
// The funcdata.h pseudo-statements (NO_LOCAL_POINTERS, GO_ARGS,
// GO_RESULTS_INITIALIZED) carry metadata for the linker, not machine
// code: the toolchain emits zero instruction bytes for them, and so does
// the encoder here.
// the encoder here. Files that include funcdata.h spell them after
// macro expansion as FUNCDATA $n, sym(SB), so the expanded forms are
// bookkeeping too (the same treatment the loong64 encoder applies).
switch mnem {
case "NO_LOCAL_POINTERS", "GO_ARGS", "GO_RESULTS_INITIALIZED":
return nil, nil
case "END":
if len(ops) != 0 {
return nil, fmt.Errorf("END expects no operands, got %d", len(ops))
}
return nil, nil
case "FUNCDATA":
if len(ops) != 2 || !isImmOperand(ops[0]) {
return nil, fmt.Errorf("FUNCDATA expects $n, sym(SB)")
}
return nil, nil
case "PCDATA":
if len(ops) != 2 || !isImmOperand(ops[0]) || !isImmOperand(ops[1]) {
return nil, fmt.Errorf("PCDATA expects $n, $n")
}
return nil, nil
}
// Conditional branches (BEQ, BNE, BGE, BLT, BGT, BLE, etc.).
@@ -520,8 +538,17 @@ func encodeARM64Instr(instr *ast.Instr, pc int, offsets map[string]int, fi arm64
// SIMD element moves (VDUP, VMOV with lane indices) take precedence
// over the plain arrangement paths, which carry no index.
if (mnem == "VDUP" || mnem == "VMOV") && arm64SimdHasElement(ops) {
return encodeARM64Dup(mnem, ops)
if mnem == "VDUP" || mnem == "VMOV" {
if arm64SimdHasElement(ops) {
return encodeARM64Dup(mnem, ops)
}
// VMOV/VDUP Rn, Vd.<T>: a general register into an arranged whole
// vector (asm7.go case 82, shared by both mnemonics). The element
// paths above only run when a lane index is spelled, so this is the
// whole-vector shape's only route.
if b, ok, err := encodeARM64GPToVec(mnem, ops); ok {
return b, err
}
}
// Arrangement-aware SIMD three-register (VADD, VAND, VCMEQ, VZIP1,
@@ -610,6 +637,20 @@ func encodeARM64Branch(mnem string, ops []*ast.Operand, pc int, offsets map[stri
return a64wordLE(a64UncondBranch(opc, uint32(rn), 0)), nil
}
// The bare spelling BL R9 is the same indirect branch: the parser reads
// a bare identifier as a symbol, and one named for a register is an
// indirect branch through it, which the toolchain accepts alongside the
// parenthesised form (BL (R3) and BL R3 both encode BLR R3).
if op.Addr.Sym != nil && op.Addr.Sym.Pseudo == "" && op.Addr.Base == "" && op.Addr.Index == "" {
if rn := arm64RegNum(op.Addr.Sym.Name); rn >= 0 {
opc := uint32(0) // BR
if link {
opc = 1 // BLR
}
return a64wordLE(a64UncondBranch(opc, uint32(rn), 0)), nil
}
}
// Symbol reference: BL sym(SB), or B sym(SB) for a tail call, against a
// relocation (R_CALLARM64 either way).
if op.Addr.Sym != nil && op.Addr.Sym.Pseudo == "SB" {
@@ -1427,6 +1468,15 @@ func encodeARM64Mov(instr *ast.Instr, mnem string, wb string, fi arm64FrameInfo,
}
return encodeARM64SBAddr(src.Imm.Sym, rd, relocs), nil
}
// Immediate → memory: only storing zero is encodable (the ZR
// register); the toolchain rejects any other immediate-to-memory
// combination ("illegal combination").
if isMemOperand(dst) {
if arm64Imm64(src) != 0 {
return nil, fmt.Errorf("%s: illegal combination: an immediate store must be zero", mnem)
}
return encodeARM64MemOp(mnem, dst, 31, false, fi, "")
}
rd := arm64RegNum(operandRegName(dst))
if rd < 0 {
return nil, fmt.Errorf("%s $imm: invalid destination register", mnem)
@@ -1733,10 +1783,33 @@ func encodeARM64RegMove(mnem string, src, dst *ast.Operand) ([]byte, error) {
return a64wordLE(sf<<31 | 0x1E<<24 | typ<<22 | 1<<21 | 7<<16 | uint32(rs)<<5 | uint32(rd)), nil
}
// Integer → integer: ORR Rd, ZR, Rs.
// Integer → integer. Every truncating register move lowers to an
// extend in the toolchain (asm7.go case 45): the signed forms to SBFM
// (SXTB, SXTH, SXTW), the unsigned byte and halfword forms to UBFM
// (UXTB, UXTH), and only MOVWU to an ORR against WZR. MOVD stays
// ORR Xd, XZR, Xm.
if rs != 31 {
switch mnem {
case "MOVB":
return a64wordLE(0x93400000 | 7<<10 | uint32(rs)<<5 | uint32(rd)), nil
case "MOVH":
return a64wordLE(0x93400000 | 15<<10 | uint32(rs)<<5 | uint32(rd)), nil
case "MOVW":
return a64wordLE(0x93400000 | 31<<10 | uint32(rs)<<5 | uint32(rd)), nil
case "MOVBU":
return a64wordLE(0xd3400000 | 7<<10 | uint32(rs)<<5 | uint32(rd)), nil
case "MOVHU":
return a64wordLE(0xd3400000 | 15<<10 | uint32(rs)<<5 | uint32(rd)), nil
}
}
sf := uint32(1) // 64-bit
if mnem == "MOVW" || mnem == "MOVWU" || mnem == "MOVB" || mnem == "MOVBU" ||
mnem == "MOVH" || mnem == "MOVHU" {
if mnem == "MOVWU" {
sf = 0
}
// A narrow move out of the zero register loses its width: the
// toolchain rewrites it as MOVWU (asm7.go case 45), an ORR against
// WZR. MOVD and MOV keep the 64-bit form.
if rs == 31 && mnem != "MOVD" && mnem != "MOV" {
sf = 0
}
op := uint32(1<<29 | 0x0a<<24) // ORR
@@ -2856,8 +2929,14 @@ func encodeARM64Sys(mnem string, ops []*ast.Operand) ([]byte, error) {
}
return a64wordLE(0xd503201f | uint32(v)<<5), nil
case "BTI":
// The toolchain requires the landing-pad kind: bare BTI is
// rejected ("missing operand"), and only the uppercase C/J/JC
// spellings assemble (0xd503245f/49f/4df).
if len(ops) != 1 {
return nil, fmt.Errorf("%s expects C, J, or JC", mnem)
}
op := operandRegName(ops[0])
base, ok := map[string]uint32{"C": 0xd503245f}[op]
base, ok := map[string]uint32{"C": 0xd503245f, "J": 0xd503249f, "JC": 0xd50324df}[op]
if !ok {
return nil, fmt.Errorf("%s: unknown kind %q", mnem, op)
}
@@ -3179,6 +3258,22 @@ func encodeARM64SimdV(mnem string, spec a64SimdVSpec, ops []*ast.Operand) ([]byt
}
return a64wordLE(base | a64ArrBits[arr] | uint32(vn.reg)<<5 | uint32(vd.reg)), nil
}
// The toolchain's two-operand spellings VADD/VSUB Vm, Vn accumulate Vn
// with Vm in place (asm7.go case 89, r defaulting to rt). They exist
// for bare V registers alone: the arranged forms and every other
// three-register mnemonic are rejected outright.
if len(ops) == 2 && (mnem == "VADD" || mnem == "VSUB") {
vm, ok1 := arm64VecOf(ops[0])
vn, ok2 := arm64VecOf(ops[1])
if !ok1 || !ok2 || vm.hasIdx || vn.hasIdx || vm.arr != "" || vn.arr != "" {
return nil, fmt.Errorf("%s: two-operand form takes bare V registers", mnem)
}
base := uint32(0x5ee08400) // VADD
if mnem == "VSUB" {
base = 0x7ee08400
}
return a64wordLE(base | uint32(vm.reg)<<16 | uint32(vn.reg)<<5 | uint32(vn.reg)), nil
}
if len(ops) != 3 {
return nil, fmt.Errorf("%s expects 3 operands, got %d", mnem, len(ops))
}
@@ -3372,6 +3467,60 @@ func encodeARM64VTBL(mnem string, ops []*ast.Operand) ([]byte, error) {
return a64wordLE(base | q | uint32(len(ts)-1)<<13 | uint32(vi.reg)<<16 | uint32(ts[0].reg)<<5 | uint32(vd.reg)), nil
}
// encodeARM64GPToVec encodes the whole-vector move VMOV/VDUP Rs, Vd.<T>: a
// general register into an arranged vector, the spelling asm7.go's case 82
// calls vmov/vdup Rn, Vd.<T>. ok is false for anything that is not that
// shape, so the caller falls through to the arrangement and element paths;
// the toolchain rejects the bare spellings outright, and the reverse
// Vd.<T>, Rs with them.
func encodeARM64GPToVec(mnem string, ops []*ast.Operand) ([]byte, bool, error) {
if len(ops) != 2 {
return nil, false, nil
}
if ops[0].Addr.Base != "" || isImmOperand(ops[0]) {
return nil, false, nil
}
rs := arm64RegNum(operandRegName(ops[0]))
if rs < 0 {
return nil, false, nil
}
dst, ok := arm64VecOf(ops[1])
if !ok || dst.hasIdx || dst.arr == "" {
return nil, false, nil
}
b, err := a64GPVecWhole(mnem, rs, dst)
return b, true, err
}
// a64GPVecWhole lays down the general-register-into-a-whole-vector move:
// word = Q | 7<<25 | imm5<<16 | 3<<10 | rs<<5 | rd, with imm5 naming the
// lane width and Q the vector length. Both VMOV and VDUP take this form
// (asm7.go case 82); INS-into-one-lane is encoded elsewhere.
func a64GPVecWhole(mnem string, rs int, dst a64Vec) ([]byte, error) {
var imm5, q uint32
switch dst.arr {
case "B8":
imm5, q = 1, 0
case "B16":
imm5, q = 1, 1<<30
case "H4":
imm5, q = 2, 0
case "H8":
imm5, q = 2, 1<<30
case "S2":
imm5, q = 4, 0
case "S4":
imm5, q = 4, 1<<30
case "D2":
imm5, q = 8, 1<<30
default:
// D1 rides no case-82 row: the toolchain rejects the one-doubleword
// spelling for this form, so the encoder refuses it too.
return nil, fmt.Errorf("%s: invalid destination arrangement %q", mnem, dst.arr)
}
return a64wordLE(q | 0x0e000c00 | imm5<<16 | uint32(rs)<<5 | uint32(dst.reg)), nil
}
// encodeARM64Dup encodes the SIMD element moves VDUP and VMOV spell with
// lane indices:
//
@@ -3407,28 +3556,8 @@ func encodeARM64Dup(mnem string, ops []*ast.Operand) ([]byte, error) {
return nil, fmt.Errorf("%s: source must be a general register", mnem)
}
if !dst.hasIdx {
var imm5, q uint32
switch dst.arr {
case "B8":
imm5, q = 1, 0
case "B16":
imm5, q = 1, 1<<30
case "H4":
imm5, q = 2, 0
case "H8":
imm5, q = 2, 1<<30
case "S2":
imm5, q = 4, 0
case "S4":
imm5, q = 4, 1<<30
case "D1":
imm5, q = 8, 0
case "D2":
imm5, q = 8, 1<<30
default:
return nil, fmt.Errorf("%s: invalid destination arrangement %q", mnem, dst.arr)
}
return a64wordLE(q | 0x0e000c00 | imm5<<16 | uint32(rs)<<5 | uint32(dst.reg)), nil
// Duplicates the register across every lane (DUP Vd.T, Rn).
return a64GPVecWhole(mnem, rs, dst)
}
f, ok := a64ElemField(dst.arr, dst.idx)
if !ok {
+5
View File
@@ -79,6 +79,11 @@ func arm64RegNum(name string) int {
return 17
case "R18":
return 18
case "R18_PLATFORM":
// The toolchain's Windows spelling: R18 is renamed R18_PLATFORM in
// cmd/asm/internal/arch so assembly cannot use it by accident, and
// sys_windows_arm64.s references it only through this name.
return 18
case "R19":
return 19
case "R20":
+124
View File
@@ -105,6 +105,7 @@ func TestArm64RegNum(t *testing.T) {
}{
{"R0", 0}, {"R4", 4}, {"R29", 29}, {"R30", 30}, {"R31", 31},
{"FP", 29}, {"LR", 30}, {"LINK", 30}, {"SP", 31}, {"ZR", 31},
{"R18_PLATFORM", 18},
{"F0", 0}, {"F4", 4}, {"F31", 31},
{"INVALID", -1}, {"X0", -1}, {"", -1},
}
@@ -675,6 +676,36 @@ func TestArm64AcquireRelease(t *testing.T) {
}
}
// TestArm64BTI pins the landing-pad family against the toolchain words:
// only the uppercase C/J/JC spellings assemble, and bare BTI is a
// diagnostic, never a panic.
func TestArm64BTI(t *testing.T) {
got := arm64Words(t, "\tBTI C\n\tBTI J\n\tBTI JC\n")
want := []uint32{
0xd503245f, // BTI C
0xd503249f, // BTI J
0xd50324df, // BTI JC
0xd65f03c0, // RET
}
if len(got) != len(want) {
t.Fatalf("got %d words, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %#x, want %#x", i, got[i], want[i])
}
}
for _, src := range []string{"\tBTI\n", "\tBTI c\n", "\tBTI B\n"} {
f, errs := parser.Parse("test_arm64.s", "#include \"textflag.h\"\n\nTEXT ·f(SB), NOSPLIT, $0-0\n"+src+"\tRET\n")
if len(errs) > 0 {
continue
}
if _, err := AssembleFileARM64(f); err == nil {
t.Errorf("BTI spelling %q should be rejected, as go tool asm rejects it", src)
}
}
}
// TestArm64System pins BRK, SVC, the barriers, cache maintenance and the
// system register accesses.
func TestArm64System(t *testing.T) {
@@ -831,6 +862,99 @@ func TestArm64SIMDElement(t *testing.T) {
}
}
// TestArm64GPIntoVector pins the whole-vector moves VMOV/VDUP Rs, Vd.<T>
// against `go tool asm -S` output (Go 1.27, arm64): word = Q | 7<<25 |
// imm5<<16 | 3<<10 | rs<<5 | rd, shared by both mnemonics, the form
// sys_windows_arm64.s and the bytealg loops use. The D1 destination is
// rejected, as the toolchain rejects it.
func TestArm64GPIntoVector(t *testing.T) {
got := arm64Words(t, "\tVMOV R5, V5.B16\n\tVMOV R1, V2.B8\n\tVMOV R3, V4.H4\n"+
"\tVMOV R9, V10.S4\n\tVMOV R7, V31.H8\n\tVMOV R11, V12.D2\n"+
"\tVDUP R5, V5.B16\n\tVDUP R9, V10.H8\n\tVMOV V4.B16, V20.B16\n")
want := []uint32{
0x4e010ca5, // VMOV R5, V5.B16
0x0e010c22, // VMOV R1, V2.B8
0x0e020c64, // VMOV R3, V4.H4
0x4e040d2a, // VMOV R9, V10.S4
0x4e020cff, // VMOV R7, V31.H8
0x4e080d6c, // VMOV R11, V12.D2
0x4e010ca5, // VDUP R5, V5.B16 (same word as VMOV)
0x4e020d2a, // VDUP R9, V10.H8
0x4ea41c94, // VMOV V4.B16, V20.B16 (vector to vector stays ORR)
0xd65f03c0,
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
}
}
f, errs := parser.Parse("test_arm64.s", "#include \"textflag.h\"\n\nTEXT ·f(SB), NOSPLIT, $0-0\n\tVMOV R7, V8.D1\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
if _, err := AssembleFileARM64(f); err == nil {
t.Errorf("VMOV R7, V8.D1 assembled, want an arrangement error")
}
}
// TestArm64SimdTwoOperand pins the two-operand accumulate spellings
// VADD/VSUB Vm, Vn against `go tool asm -S` output (Go 1.27, arm64):
// word = 5<<28|7<<25|7<<21|1<<15|1<<10 for VADD (7<<28 for VSUB) with
// rf<<16 | rn<<5 | rn, bare V registers only (asm7.go case 89).
func TestArm64SimdTwoOperand(t *testing.T) {
got := arm64Words(t, "\tVADD V7, V8\n\tVSUB V7, V8\n\tVADD V1, V2\n\tVADD V0.B16, V1.B16, V2.B16\n")
want := []uint32{
0x5ee78508, // VADD V7, V8
0x7ee78508, // VSUB V7, V8
0x5ee18442, // VADD V1, V2
0x4e208422, // VADD arranged: the ordinary three-register path
0xd65f03c0,
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
}
}
}
// TestArm64TruncMove pins the truncating register moves against
// `go tool asm -S` output (Go 1.27, arm64): the signed forms lower to SXTB,
// SXTH and SXTW (SBFM), the unsigned byte and halfword forms to UXTB and
// UXTH (UBFM), MOVWU to a W ORR, and a narrow move out of the zero register
// drops to the W ORR too (asm7.go case 45).
func TestArm64TruncMove(t *testing.T) {
got := arm64Words(t, "\tMOVB R3, R4\n\tMOVH R5, R6\n\tMOVW R9, R10\n"+
"\tMOVBU R3, R4\n\tMOVHU R3, R4\n\tMOVWU R3, R4\n\tMOVD R3, R4\n"+
"\tMOVD ZR, R4\n\tMOVB ZR, R4\n\tMOVWU ZR, R5\n")
want := []uint32{
0x93401c64, // MOVB = SXTB
0x93403ca6, // MOVH = SXTH
0x93407d2a, // MOVW = SXTW
0xd3401c64, // MOVBU = UXTB
0xd3403c64, // MOVHU = UXTH
0x2a0303e4, // MOVWU = ORR W
0xaa0303e4, // MOVD = ORR X
0xaa1f03e4, // MOVD ZR, R4 keeps the X form
0x2a1f03e4, // MOVB ZR, R4 drops to the W form
0x2a1f03e5, // MOVWU ZR, R5
0xd65f03c0,
}
if len(got) != len(want) {
t.Fatalf("word count = %d, want %d", len(got), len(want))
}
for i := range want {
if got[i] != want[i] {
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
}
}
}
// TestArm64SIMDLoadStore pins the structure loads and stores.
func TestArm64SIMDLoadStore(t *testing.T) {
got := arm64Words(t, "\tVLD1 (R2), [V21.B16]\n\tVLD1 (R1), [V2.B16, V3.B16]\n\tVLD1 (R29), [V14.D1, V15.D1, V16.D1, V17.D1]\n"+
+497 -39
View File
@@ -5,6 +5,7 @@ package asm
import (
"fmt"
"strconv"
"strings"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
@@ -28,7 +29,7 @@ import (
// emitted: the bytes match go tool asm only for NOSPLIT functions or
// zero-frame leaves, where the toolchain emits no guard either.
func Assemble(t *ast.Text) ([]byte, map[string]int, error) {
code, _, labels, _, _, err := assemble(t, nil)
code, _, labels, _, _, _, err := assemble(t, nil)
return code, labels, err
}
@@ -37,10 +38,25 @@ func Assemble(t *ast.Text) ([]byte, map[string]int, error) {
// rejects SB operands outright (single-function assembly cannot resolve
// them). When allowExternal is set, a reference to a symbol no GLOBL in the
// file defines is recorded as an external relocation instead of failing
// the object-file emitters resolve it at link time.
// the object-file emitters resolve it at link time. goos selects the TLS
// access form: the empty default behaves as linux.
type linkInfo struct {
symbols map[string]bool
allowExternal bool
goos string
}
// tlsOneInsn reports the one-instruction TLS form, obj6.go's
// CanUse1InsnTLS for the GOOS gasm supports: the bare TLS load nops out and
// the (TLS*1) index folds to a segment-absolute access. Windows and plan9
// keep the two-instruction form; shared linux does too, which gasm's raw
// path does not model and therefore does not select.
func (l *linkInfo) tlsOneInsn() bool {
switch l.goos {
case "", "linux", "freebsd":
return true
}
return false
}
// sbPatch is a function-relative static-symbol relocation: the disp32 field
@@ -66,9 +82,9 @@ type spadjStep struct {
// assemble encodes a TEXT body, returning the machine code, the static-symbol
// patch sites (for the file-level layout to resolve), the label table and the
// stack-adjustment boundaries.
func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, []spadjStep, []LineEntry, error) {
func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, []spadjStep, []LineEntry, []floatPoolEntry, error) {
if err := checkAdjspBalance(t); err != nil {
return nil, nil, nil, nil, nil, err
return nil, nil, nil, nil, nil, nil, err
}
fi := computeFrame(t)
chain := jumpChain(t)
@@ -85,29 +101,118 @@ func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, [
// outgrows the short form.
long := make([]bool, len(t.Body))
sizes := make([]int, len(t.Body))
numTargets := make([]int, len(t.Body))
for i := range numTargets {
numTargets[i] = -1
}
offsets := map[string]int{}
pcs := make([]int, len(t.Body))
var guardJBlong, guardJBElong, moreJMPlong bool
poolSeen := map[string]bool{}
var poolList []floatPoolEntry
for {
guard := fi.guardLen(guardJBlong, guardJBElong)
pos := guard + len(fi.prologue)
for i := range numTargets {
numTargets[i] = -1
}
idxAtPc := map[int]int{}
for i, stmt := range t.Body {
switch s := stmt.(type) {
case *ast.Label:
offsets[s.Name.Text] = pos
case *ast.Instr:
if strings.ToUpper(s.Mnemonic.Text) == "PCALIGN" {
// The alignment pseudo-statement: its size is the
// padding to the next boundary at this very position,
// filled with NOPs at emission.
pad, err := pcAlignPad(pcAlignValue(s), pos)
if err != nil {
return nil, nil, nil, nil, nil, nil, fmt.Errorf("PCALIGN: %w", err)
}
sizes[i] = pad
pcs[i] = pos
pos += pad
continue
}
sz, err := instrSize(s, fi, long[i], link)
if err != nil {
return nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", s.Mnemonic.Text, err)
return nil, nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", s.Mnemonic.Text, err)
}
sizes[i] = sz
pcs[i] = pos
idxAtPc[pos] = i
pos += sz
}
}
bodyLen := pos - (guard + len(fi.prologue))
// Expand any short jump whose displacement no longer fits rel8.
changed := false
// Numeric ±N(PC) jumps resolve against this iteration's layout; the
// emission pass reads the same table after the loop converges. A
// target that is itself an unconditional local JMP is chased to the
// ultimate target: the toolchain's brloop pass collapses branch-to-
// branch chains before it encodes, so matching its bytes requires
// the same redirection.
for i := range numTargets {
numTargets[i] = -1
}
for i, stmt := range t.Body {
s, ok := stmt.(*ast.Instr)
if !ok {
continue
}
if len(s.Operands) == 1 {
if n, isNum := pcJumpOffset(s.Operands[0]); isNum {
if target, okT := pcJumpTarget(t, i, n, pcs); okT {
numTargets[i] = target
}
}
}
}
for i := range numTargets {
if numTargets[i] < 0 {
continue
}
tgt := numTargets[i]
for hop := 0; hop < len(t.Body); hop++ {
idx, ok := idxAtPc[tgt]
if !ok {
break
}
in, ok := t.Body[idx].(*ast.Instr)
if !ok || strings.ToUpper(in.Mnemonic.Text) != "JMP" || len(in.Operands) != 1 {
break
}
if name, isLabel := labelName(in.Operands[0]); isLabel {
tgt = offsets[resolve(name)]
continue
}
if n, isNum := pcJumpOffset(in.Operands[0]); isNum {
next, okT := pcJumpTarget(t, idx, n, pcs)
if !okT {
break
}
tgt = next
continue
}
break // JMP through a register or memory: the chain ends
}
numTargets[i] = tgt
}
for i, stmt := range t.Body {
s, ok := stmt.(*ast.Instr)
if !ok {
continue
}
if numTargets[i] >= 0 && !long[i] {
rel := int64(numTargets[i] - (pcs[i] + jumpSize(strings.ToUpper(s.Mnemonic.Text), false)))
if !fits8(rel) {
long[i] = true
changed = true
}
}
}
for i, stmt := range t.Body {
s, ok := stmt.(*ast.Instr)
if !ok {
@@ -229,12 +334,18 @@ func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, [
spadjStep{pos + epi, 0},
)
}
code, ps, err := encodeInstr(s, pos, offsets, fi, long[i], resolve, link)
code, ps, pool, err := encodeInstr(s, pos, offsets, fi, long[i], resolve, link, numTargets[i])
if err != nil {
return nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", s.Mnemonic.Text, err)
return nil, nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", s.Mnemonic.Text, err)
}
for _, entry := range pool {
if !poolSeen[entry.name] {
poolSeen[entry.name] = true
poolList = append(poolList, entry)
}
}
if len(code) != sizes[i] {
return nil, nil, nil, nil, nil, fmt.Errorf("%s: size mismatch (%d vs %d)", s.Mnemonic.Text, len(code), sizes[i])
return nil, nil, nil, nil, nil, nil, fmt.Errorf("%s: size mismatch (%d vs %d)", s.Mnemonic.Text, len(code), sizes[i])
}
if strings.ToUpper(s.Mnemonic.Text) == "CALL" {
for k := range ps {
@@ -272,7 +383,7 @@ func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, [
pos += len(suffix)
}
_ = pos
return out, patches, offsets, steps, lines, nil
return out, patches, offsets, steps, lines, poolList, nil
}
// jumpChain precomputes jump-to-jump folding: a label whose first instruction
@@ -419,6 +530,83 @@ func computeFrame(t *ast.Text) frameInfo {
return fi
}
// pcJumpOffset recognises the numeric relative jump operand ±N(PC) and
// returns N: the toolchain counts instructions, not bytes, so +2(PC) targets
// the second instruction boundary after the branch.
func pcJumpOffset(op *ast.Operand) (int, bool) {
if op.Kind != ast.OpAddr || op.Addr.Base != "PC" {
return 0, false
}
return int(op.Addr.Offset), true
}
// pcJumpTarget resolves a numeric jump at statement index j: N counts the
// instruction statements after the jump itself (N = 0 is the jump's own
// address, the classic park loop), and the target is the start of the Nth
// one. It reports false when the count runs past the end of the function.
func pcJumpTarget(t *ast.Text, j, n int, pcs []int) (int, bool) {
if n == 0 {
return pcs[j], true
}
seen := 0
for k := j + 1; k < len(t.Body); k++ {
if _, ok := t.Body[k].(*ast.Instr); !ok {
continue
}
seen++
if seen == n {
return pcs[k], true
}
}
return 0, false
}
// x86 NOP encodings, single-instruction no-ops of lengths 1 to 9 (the
// toolchain's asm6.go nop table); longer padding repeats the largest that
// fits, greedy from the end.
var x86Nops = [][]byte{
{0x90},
{0x66, 0x90},
{0x0F, 0x1F, 0x00},
{0x0F, 0x1F, 0x40, 0x00},
{0x0F, 0x1F, 0x44, 0x00, 0x00},
{0x66, 0x0F, 0x1F, 0x44, 0x00, 0x00},
{0x0F, 0x1F, 0x80, 0x00, 0x00, 0x00, 0x00},
{0x0F, 0x1F, 0x84, 0x00, 0x00, 0x00, 0x00, 0x00},
{0x66, 0x0F, 0x1F, 0x84, 0x00, 0x00, 0x00, 0x00, 0x00},
}
// fillNOPs fills p with the greedy largest single-instruction NOPs, exactly
// the toolchain's fillnop.
func fillNOPs(p []byte) {
for len(p) > 0 {
m := min(len(p), len(x86Nops))
copy(p[:m], x86Nops[m-1])
p = p[m:]
}
}
// pcAlignPad computes the padding PCALIGN $align inserts at pos: the
// alignment must be a power of two in [8, 2048] and the padding runs to the
// next boundary (zero when the position is already aligned).
func pcAlignPad(align, pos int) (int, error) {
if align <= 0 || align&(align-1) != 0 || align < 8 || align > 2048 {
return 0, fmt.Errorf("alignment value of an instruction must be a power of two and in the range [8, 2048], got %d", align)
}
if lob := pos & (align - 1); lob != 0 {
return align - lob, nil
}
return 0, nil
}
// pcAlignValue reads a PCALIGN statement's alignment operand.
func pcAlignValue(s *ast.Instr) int {
if len(s.Operands) == 1 && s.Operands[0].Kind == ast.OpImmediate && s.Operands[0].Imm.HasVal {
return int(s.Operands[0].Imm.Val)
}
return 0 // rejected by pcAlignPad's range check
}
// hasCall reports whether the function body contains a CALL instruction.
func hasCall(t *ast.Text) bool {
for _, stmt := range t.Body {
@@ -593,7 +781,7 @@ func instrSize(s *ast.Instr, fi frameInfo, long bool, link *linkInfo) (int, erro
}
return jumpSize(mnem, long), nil
}
code, _, err := encodeInstr(s, 0, nil, fi, false, nil, link)
code, _, _, err := encodeInstr(s, 0, nil, fi, false, nil, link, -1)
if err != nil {
return 0, err
}
@@ -628,9 +816,21 @@ func jumpSize(mnem string, long bool) int {
// (relative to pc, the instruction's own offset). A RET in a frame-pointer
// function is prefixed with the epilogue. resolve, when non-nil, redirects a
// jump label through the jump-to-jump chain before the offset lookup.
func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, long bool, resolve func(string) string, link *linkInfo) ([]byte, []sbPatch, error) {
func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, long bool, resolve func(string) string, link *linkInfo, numTarget int) ([]byte, []sbPatch, []floatPoolEntry, error) {
mnem := strings.ToUpper(s.Mnemonic.Text)
if mnem == "PCALIGN" {
// The layout pass already accounted the padding; emit the same
// amount of NOP bytes for the statement's own position.
pad, err := pcAlignPad(pcAlignValue(s), pc)
if err != nil {
return nil, nil, nil, err
}
out := make([]byte, pad)
fillNOPs(out)
return out, nil, nil, nil
}
var prefix []byte
if mnem == "RET" && fi.useFP {
prefix = fi.epilogue
@@ -638,6 +838,7 @@ func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, lon
var code []byte
var ps []sbPatch
var pool []floatPoolEntry
var err error
if isJumpMnemonic(mnem) {
if (mnem == "CALL" || mnem == "JMP") && isSBCall(s) {
@@ -646,7 +847,7 @@ func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, lon
// or the linker.
code, ps, err = encodeSBCall(s, link)
if err != nil {
return nil, nil, err
return nil, nil, nil, err
}
for i := range ps {
ps[i].kind = RelCall
@@ -656,23 +857,23 @@ func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, lon
ps[i].off += body
ps[i].after = body + len(code)
}
return append(prefix, code...), ps, nil
return append(prefix, code...), ps, nil, nil
}
if (mnem == "CALL" || mnem == "JMP") && indirectJumpTarget(s) {
// JMP/CALL through a register or memory: no relocation and no
// label to resolve, the operand fully determines the bytes.
code, err = encodeIndirectJump(s, mnem)
if err != nil {
return nil, nil, err
return nil, nil, nil, err
}
return append(prefix, code...), nil, nil
return append(prefix, code...), nil, nil, nil
}
code, err = encodeJump(s, mnem, pc+len(prefix), offsets, long, resolve)
code, err = encodeJump(s, mnem, pc+len(prefix), offsets, long, resolve, numTarget)
} else {
code, ps, err = encodeNormal(s, fi, link)
code, ps, pool, err = encodeNormal(s, fi, link)
}
if err != nil {
return nil, nil, err
return nil, nil, nil, err
}
// Anchor the patch fields at function-relative positions: off indexes the
// disp32 field, after is the address just past the instruction.
@@ -681,49 +882,183 @@ func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, lon
ps[i].off += body
ps[i].after = body + len(code)
}
return append(prefix, code...), ps, nil
return append(prefix, code...), ps, pool, nil
}
func encodeNormal(s *ast.Instr, fi frameInfo, link *linkInfo) ([]byte, []sbPatch, error) {
_, size := splitSize(strings.ToUpper(s.Mnemonic.Text))
func encodeNormal(s *ast.Instr, fi frameInfo, link *linkInfo) ([]byte, []sbPatch, []floatPoolEntry, error) {
mnemUpper := strings.ToUpper(s.Mnemonic.Text)
if mnemUpper == "FUNCDATA" || mnemUpper == "PCDATA" {
code, err := encodeBookkeeping(mnemUpper, s)
if err != nil {
return nil, nil, nil, err
}
return code, nil, nil, nil
}
// MOVQ $sym±off(SB), r64: the toolchain assembles a symbol immediate as
// LEAQ disp32(RIP), r64 with an R_PCREL relocation at the disp32 field,
// never as a 64-bit absolute immediate (verified against go tool asm).
// MOVD is the MOVQ alias; the narrower widths reject the form outright.
if (mnemUpper == "MOVQ" || mnemUpper == "MOVD") && len(s.Operands) == 2 &&
s.Operands[0].Kind == ast.OpImmediate && s.Operands[0].Imm.Sym != nil &&
s.Operands[0].Imm.Sym.Pseudo == "SB" {
mem := &ast.Operand{Kind: ast.OpAddr, Addr: ast.Address{Sym: s.Operands[0].Imm.Sym}}
src, err := operandFromAST(mnemUpper, mem, 8, fi, link)
if err != nil {
return nil, nil, nil, err
}
dst, err := operandFromAST(mnemUpper, s.Operands[1], 8, fi, link)
if err != nil {
return nil, nil, nil, err
}
e := &enc{}
if err := e.encodeLea([]Operand{src, dst}, 8); err != nil {
return nil, nil, nil, err
}
ps := make([]sbPatch, len(e.patches))
for i, p := range e.patches {
ps[i] = sbPatch{off: p.off, name: p.name, addend: p.addend}
}
return e.out, ps, nil, nil
}
// MOVQ/MOVL TLS, r: the bare TLS load. The toolchain's progedit nops
// it out on the one-instruction TLS systems (linux and freebsd, not
// shared) and encodes the segment-prefixed load elsewhere; get_tls(r),
// the macro GOROOT's go_tls.h defines, expands to exactly this
// statement, and the toolchain's pairing pass removes it whenever the
// following instruction's (TLS*1) index folds.
if (mnemUpper == "MOVQ" || mnemUpper == "MOVL") && len(s.Operands) == 2 && isBareTLS(s.Operands[0]) {
return encodeTLSBaseLoad(s, fi, link)
}
_, size := splitSize(mnemUpper)
if size == 0 {
size = 8
}
ops := make([]Operand, len(s.Operands))
for i, op := range s.Operands {
o, err := operandFromAST(op, size, fi, link)
o, err := operandFromAST(mnemUpper, op, size, fi, link)
if err != nil {
return nil, nil, err
return nil, nil, nil, err
}
ops[i] = o
}
e := &enc{}
if err := e.encode(s.Mnemonic.Text, ops); err != nil {
return nil, nil, err
return nil, nil, nil, err
}
ps := make([]sbPatch, len(e.patches))
for i, p := range e.patches {
ps[i] = sbPatch{off: p.off, name: p.name, addend: p.addend}
if p.tls {
ps[i].kind = RelTLSLE
}
}
return e.out, ps, nil
return e.out, ps, e.floatPoolList(), nil
}
// isBareTLS reports whether the operand is the bare TLS pseudo-register
// load source, the expansion of go_tls.h's get_tls(r) macro.
func isBareTLS(op *ast.Operand) bool {
return op.Kind == ast.OpAddr && op.Addr.Sym != nil &&
op.Addr.Sym.Pseudo == "" && op.Addr.Sym.Name == "TLS" &&
op.Addr.Base == "" && op.Addr.Index == ""
}
// encodeTLSBaseLoad assembles MOVQ/MOVL TLS, r. On the one-instruction TLS
// systems (linux and freebsd outside -shared, obj6.go's CanUse1InsnTLS) the
// statement nops out: the following (TLS*1) access folds to a direct
// segment-absolute load. The two-instruction systems keep the segment load,
// nine bytes with the R_TLSLE patch site at the disp32.
func encodeTLSBaseLoad(s *ast.Instr, fi frameInfo, link *linkInfo) ([]byte, []sbPatch, []floatPoolEntry, error) {
_, size := splitSize(strings.ToUpper(s.Mnemonic.Text))
if size == 0 {
size = 8
}
dst, err := operandFromAST("MOVQ", s.Operands[1], 8, fi, link)
if err != nil {
return nil, nil, nil, err
}
reg, ok := dst.(Reg)
if !ok || reg.isVec() {
return nil, nil, nil, fmt.Errorf("TLS: destination must be a general register")
}
if link == nil || link.tlsOneInsn() {
return nil, nil, nil, nil // noped out
}
seg := byte(0x64) // FS
if link.goos == "windows" {
seg = 0x65 // GS
}
e := &enc{}
i := &instr{
prefix: seg,
rexW: size == 8,
rexR: reg.idx >= 8,
opcode: []byte{0x8B},
modrm: 0x04 | (reg.idx&7)<<3,
sib: 0x25,
disp: le32(0),
tls: true,
}
if err := e.emit(i); err != nil {
return nil, nil, nil, err
}
ps := make([]sbPatch, len(e.patches))
for i, p := range e.patches {
ps[i] = sbPatch{off: p.off, name: p.name, addend: p.addend, kind: RelTLSLE}
}
return e.out, ps, nil, nil
}
// encodeBookkeeping accepts-and-ignores FUNCDATA and PCDATA at the statement
// level, before operand conversion: the toolchain's shapes are FUNCDATA
// $n, sym(SB) and PCDATA $n, $m, and neither contributes a byte to the
// function body. The symbol reference must not run through the SB-operand
// path, which demands file-level resolution the statement never needs.
func encodeBookkeeping(upper string, s *ast.Instr) ([]byte, error) {
if len(s.Operands) != 2 {
return nil, fmt.Errorf("%s expects 2 operands, got %d", upper, len(s.Operands))
}
a, b := s.Operands[0], s.Operands[1]
if a.Kind != ast.OpImmediate || !a.Imm.HasVal {
return nil, fmt.Errorf("%s: first operand must be an integer immediate", upper)
}
switch upper {
case "FUNCDATA":
if b.Kind != ast.OpAddr || b.Addr.Sym == nil || b.Addr.Sym.Pseudo != "SB" {
return nil, fmt.Errorf("FUNCDATA: second operand must be a symbol reference")
}
case "PCDATA":
if b.Kind != ast.OpImmediate || !b.Imm.HasVal {
return nil, fmt.Errorf("PCDATA: second operand must be an integer immediate")
}
}
return nil, nil
}
// encodeJump encodes a JMP/CALL/Jcc with a relative offset resolved from the
// target label, in the short (rel8) or long (rel32) form.
func encodeJump(s *ast.Instr, mnem string, pc int, offsets map[string]int, long bool, resolve func(string) string) ([]byte, error) {
// target label or from a numeric ±N(PC) instruction count, in the short
// (rel8) or long (rel32) form. numTarget is the resolved byte offset of a
// numeric operand, negative when the operand is not one.
func encodeJump(s *ast.Instr, mnem string, pc int, offsets map[string]int, long bool, resolve func(string) string, numTarget int) ([]byte, error) {
if len(s.Operands) != 1 {
return nil, fmt.Errorf("jump expects 1 operand, got %d", len(s.Operands))
}
name, ok := labelName(s.Operands[0])
if !ok {
name, isLabel := labelName(s.Operands[0])
if !isLabel && numTarget < 0 {
return nil, fmt.Errorf("jump target must be a local label")
}
if resolve != nil && mnem != "CALL" {
name = resolve(name)
}
target, ok := offsets[name]
if !ok {
return nil, fmt.Errorf("undefined label %q", name)
var target int
if isLabel {
if resolve != nil && mnem != "CALL" {
name = resolve(name)
}
t, ok := offsets[name]
if !ok {
return nil, fmt.Errorf("undefined label %q", name)
}
target = t
} else {
target = numTarget
}
rel := int64(target - (pc + jumpSize(mnem, long)))
@@ -756,7 +1091,7 @@ func isSBCall(s *ast.Instr) bool {
// encodeSBCall encodes CALL sym(SB) as E8 rel32 with a patch site.
func encodeSBCall(s *ast.Instr, link *linkInfo) ([]byte, []sbPatch, error) {
o, err := operandFromAST(s.Operands[0], 8, frameInfo{}, link)
o, err := operandFromAST(strings.ToUpper(s.Mnemonic.Text), s.Operands[0], 8, frameInfo{}, link)
if err != nil {
return nil, nil, err
}
@@ -797,6 +1132,11 @@ func indirectJumpTarget(s *ast.Instr) bool {
return false
}
a := s.Operands[0].Addr
// ±N(PC) is the numeric relative form, the PC counts instructions from
// the branch: relative, not indirect.
if a.Base == "PC" || a.Index == "PC" {
return false
}
if a.Base != "" || a.Index != "" {
return true
}
@@ -813,7 +1153,7 @@ func indirectJumpTarget(s *ast.Instr) bool {
func encodeIndirectJump(s *ast.Instr, mnem string) ([]byte, error) {
ops := make([]Operand, len(s.Operands))
for i, op := range s.Operands {
o, err := operandFromAST(op, 8, frameInfo{}, nil)
o, err := operandFromAST(mnem, op, 8, frameInfo{}, nil)
if err != nil {
return nil, err
}
@@ -830,8 +1170,11 @@ func encodeIndirectJump(s *ast.Instr, mnem string) ([]byte, error) {
var spReg = Reg{idx: 4, size: 8}
// operandFromAST converts a parsed operand into an encoder Operand, applying
// the frame translation to FP/SP pseudo-register operands.
func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Operand, error) {
// the frame translation to FP/SP pseudo-register operands. mnemUpper is the
// instruction's upper-case mnemonic, which the floating-point immediate gate
// needs: only the SSE mnemonics whose encoding takes an XMM/memory source
// accept one.
func operandFromAST(mnemUpper string, op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Operand, error) {
switch op.Kind {
case ast.OpImmediate:
if op.Imm.HasVal {
@@ -841,11 +1184,45 @@ func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Op
}
return Imm(v), nil
}
// A floating-point immediate: $1.5, $-1.0 or the parenthesised
// $(-1.0) spelling (the constant-expression folder only folds
// integers, so that shape arrives with an empty Immediate and only
// the raw spelling carries the value). The toolchain rewrites it
// into a pooled-constant read on the SSE scalar paths and rejects
// it everywhere else.
if text, neg, ok := floatImmText(op); ok {
if !sseFloatImm[mnemUpper] {
return nil, fmt.Errorf("%s does not take a floating-point immediate", mnemUpper)
}
return FloatImm{Text: text, Neg: neg}, nil
}
return nil, fmt.Errorf("non-integer immediate not supported")
case ast.OpAddr:
a := op.Addr
// A bracketed register range, [Z0-Z3]: the four-register source of
// the 4FMAPS/4VNNIW families. The range must span four consecutive
// same-width vector registers, exactly what the toolchain's parser
// takes; the EVEX quad-register emit path reads the low end.
if a.Range != nil {
lo, ok := ParseReg(a.Range.Lo)
if !ok {
return nil, fmt.Errorf("unknown register %q in range", a.Range.Lo)
}
hi, ok := ParseReg(a.Range.Hi)
if !ok {
return nil, fmt.Errorf("unknown register %q in range", a.Range.Hi)
}
if !lo.isVec() || lo.size != hi.size {
return nil, fmt.Errorf("register range %q must span four same-width vector registers", op.Raw)
}
if hi.idx != lo.idx+3 {
return nil, fmt.Errorf("register range %q must span four consecutive registers", op.Raw)
}
return RegList{Lo: lo, Hi: hi}, nil
}
// FP-relative: x+N(FP) → (N + fpAdjust)(SP). The offset N lives in the
// symbol, not the address displacement.
if a.Sym != nil && a.Sym.Pseudo == "FP" {
@@ -877,12 +1254,44 @@ func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Op
// Memory with a real base register: (base), off(base), (base)(index*scale).
if a.Base != "" {
// Segment-absolute: 0x30(GS) and 0x28(FS), the windows TLS
// spellings. The segment override prefixes a disp32 absolute
// reference with no relocation.
if a.Base == "GS" || a.Base == "FS" {
seg := byte(0x64)
if a.Base == "GS" {
seg = 0x65
}
return SegAbs{Disp: a.Offset, Size: size, Seg: seg}, nil
}
base, ok := ParseReg(a.Base)
if !ok {
return nil, fmt.Errorf("unknown base register %q", a.Base)
}
m := Mem{Base: base, Disp: a.Offset, HasBase: true, Size: size}
if a.Index != "" {
if a.Index == "TLS" {
// off(base)(TLS*1): the thread-local annotation. The
// one-instruction TLS form folds it to off(TLS), the
// segment-prefixed absolute whose disp32 carries an
// R_TLS_LE patch site; the base register disappears
// from the encoding, exactly as the toolchain's
// progedit rewrites the address.
seg := byte(0x64) // FS on linux, freebsd, plan9
if link != nil && link.goos == "windows" {
seg = 0x65 // GS
}
return TLSMem{Disp: a.Offset, Size: size, Seg: seg}, nil
}
if a.Index == "GS" || a.Index == "FS" {
// 0(CX)(GS): the segment annotation rides the base
// access as the override prefix.
m.Seg = 0x64
if a.Index == "GS" {
m.Seg = 0x65
}
return m, nil
}
idx, ok := ParseReg(a.Index)
if !ok {
return nil, fmt.Errorf("unknown index register %q", a.Index)
@@ -893,6 +1302,21 @@ func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Op
}
return m, nil
}
// Index-only memory: the VSIB form the gather/scatter families
// read, 8(X4*1). A scaled vector index addresses memory with no
// base register; the mod=00 SIB with base field 101 carries it.
if a.Index != "" {
idx, ok := ParseReg(a.Index)
if !ok {
return nil, fmt.Errorf("unknown index register %q", a.Index)
}
return Mem{Index: idx, Scale: a.Scale, Disp: a.Offset, HasIndex: true, Size: size}, nil
}
// A bare displacement with no base: the absolute address form,
// MOVL $0xf1, 0xf1. No segment and no relocation.
if a.Sym == nil && a.Base == "" && a.Index == "" && a.HasOff {
return SegAbs{Disp: a.Offset, Size: size}, nil
}
// Bare register.
if a.Sym != nil && a.Sym.Pseudo == "" && a.Sym.Name != "" {
if r, ok := ParseReg(a.Sym.Name); ok {
@@ -903,3 +1327,37 @@ func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Op
}
return nil, fmt.Errorf("unsupported operand")
}
// floatImmText recovers a floating-point immediate's magnitude and sign from
// the parsed operand. The ordinary spellings arrive in Imm.Float; the
// parenthesised $(-1.0) leaves the Immediate empty, because the integer
// folder cannot read it, and only the verbatim operand text still carries
// the value. Anything that is not a number a float parser accepts reports
// not-ok, so every other shape keeps its existing diagnostic.
func floatImmText(op *ast.Operand) (text string, neg bool, ok bool) {
if op.Imm.Float != "" {
return op.Imm.Float, op.Imm.Neg, true
}
if op.Imm.HasVal || op.Imm.Str != "" || op.Imm.Sym != nil {
return "", false, false
}
// joinRaw spaced the token texts; the compact spelling is what matters.
compact := strings.ReplaceAll(op.Raw, " ", "")
inner, ok := strings.CutPrefix(compact, "$(")
if !ok || !strings.HasSuffix(inner, ")") {
return "", false, false
}
inner = strings.TrimSuffix(inner, ")")
inner = strings.TrimPrefix(inner, "+")
if s, ok := strings.CutPrefix(inner, "-"); ok {
neg = true
inner = s
}
if inner == "" || !strings.ContainsAny(inner, "0123456789") {
return "", false, false
}
if _, err := strconv.ParseFloat(inner, 64); err != nil {
return "", false, false
}
return inner, neg, true
}
+25
View File
@@ -568,3 +568,28 @@ TEXT ·framed(SB), $16-8
t.Errorf("framed adjsp:\n got: %s\n want: %s", hexBytes(code), hexBytes(want))
}
}
// TestAssembleRegRange pins the bracketed register range at the statement
// level: exactly four consecutive same-width vector registers assemble, the
// toolchain's rejected shapes all report an error.
func TestAssembleRegRange(t *testing.T) {
asm := func(t *testing.T, op string) ([]byte, error) {
t.Helper()
f, errs := parser.Parse("f_amd64.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tV4FMADDPS 17(SP), "+op+", K2, Z0\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse %s: %v", op, errs)
}
code, _, err := Assemble(f.Decls[0].(*ast.Text))
return code, err
}
for _, op := range []string{"[Z0-Z3]", "[Z4-Z7]", "[Z28-Z31]"} {
if _, err := asm(t, op); err != nil {
t.Errorf("%s: %v", op, err)
}
}
for _, op := range []string{"[Z0-Z4]", "[Z0-Z2]", "[Z0-Z0]", "[Z4-Z0]", "[Z1-Z0]", "[AX-Z3]", "[Z0-AX]"} {
if _, err := asm(t, op); err == nil {
t.Errorf("%s: assembled, want an error", op)
}
}
}
+65 -10
View File
@@ -43,6 +43,9 @@ const (
stInfoShift = 4
rX8664PC32 = 2
// R_X86_64_32 (debug/elf): the absolute 32-bit address of a symbol, the
// R_ADDR shape a 4-byte DATA field carries.
rX8664Abs32 = 10
// R_X86_64_TPOFF32 (debug/elf): the local-exec TLS offset the stack
// guard loads from FS. 20 is R_X86_64_TLSLD, a different relocation.
rX8664TPOFF32 = 23
@@ -158,6 +161,50 @@ func (img *Image) ELFObject() ([]byte, error) {
}
}
// The data symbols' symbol-valued DATA fields ("DATA s+0(SB)/8,
// $other(SB)") become .rela.data entries: an absolute relocation of the
// DATA line's width at the field's data-section offset, S + A with no
// PC term. Widths 4 and 8 have ELF relocation shapes; narrower fields
// cannot hold an address, so they are refused rather than truncated.
var dataRelas []elfRela
for _, d := range img.DataSyms {
for _, r := range d.Relocs {
idx, ok := symIdx[r.Name]
if !ok {
return nil, fmt.Errorf("data relocation references unknown symbol %q", r.Name)
}
var typ uint32
switch r.Siz {
case 8:
typ = rX8664Abs64
case 4:
typ = rX8664Abs32
default:
return nil, fmt.Errorf("DATA %q: a symbol value of width %d has no ELF relocation", d.Name, r.Siz)
}
dataRelas = append(dataRelas, elfRela{
off: uint64(d.Offset + r.Off),
sym: idx,
typ: typ,
addend: r.Addend,
})
}
}
// Section presence: .rela.text only when there are code relocations,
// .rela.data only when a DATA line holds a symbol value.
hasRela := len(relas) > 0
hasDataRela := len(dataRelas) > 0
nSections := 6 // NULL, .text, .data, .symtab, .strtab, .shstrtab
if hasRela {
nSections++
}
if hasDataRela {
nSections++
}
secSymtab, secStrtab := 3, 4
secShstr := nSections - 1
// Serialise the string tables.
stNames := newElfStrtab()
for _, s := range syms {
@@ -167,19 +214,13 @@ func (img *Image) ELFObject() ([]byte, error) {
for _, n := range []string{".text", ".data", ".symtab", ".strtab", ".rela.text", ".shstrtab"} {
stSections.add(n)
}
if hasDataRela {
stSections.add(".rela.data")
}
for _, n := range dwarfSectionNames {
stSections.add(n)
}
// Section presence: .rela.text only when there are relocations.
hasRela := len(relas) > 0
nSections := 6 // NULL, .text, .data, .symtab, .strtab, .shstrtab
if hasRela {
nSections = 7
}
secSymtab, secStrtab := 3, 4
secShstr := nSections - 1
// Lay the file out: header, section data, section headers.
var out []byte
out = append(out, make([]byte, 64)...) // ELF header, filled last
@@ -214,7 +255,7 @@ func (img *Image) ELFObject() ([]byte, error) {
strtabOff := len(out)
out = append(out, stNames.bytes()...)
var relaOff int
var relaOff, relaDataOff int
if hasRela {
align(8)
relaOff = len(out)
@@ -226,6 +267,17 @@ func (img *Image) ELFObject() ([]byte, error) {
out = append(out, b[:]...)
}
}
if hasDataRela {
align(8)
relaDataOff = len(out)
for _, r := range dataRelas {
var b [24]byte
le.PutUint64(b[0:], r.off)
le.PutUint64(b[8:], uint64(r.sym)<<32|uint64(r.typ))
le.PutUint64(b[16:], uint64(r.addend))
out = append(out, b[:]...)
}
}
shstrOff := len(out)
out = append(out, stSections.bytes()...)
@@ -284,6 +336,9 @@ func (img *Image) ELFObject() ([]byte, error) {
if hasRela {
putSh(".rela.text", shtRela, 0, relaOff, 24*len(relas), secSymtab, secText, 8, 24)
}
if hasDataRela {
putSh(".rela.data", shtRela, 0, relaDataOff, 24*len(dataRelas), secSymtab, secData, 8, 24)
}
putSh(".shstrtab", shtStrtab, 0, shstrOff, len(stSections.bytes()), 0, 0, 1, 0)
// DWARF section headers; their indices follow the write order.
+108
View File
@@ -750,3 +750,111 @@ func readFormSkip(t *testing.T, r *ulebIter, form uint64) {
t.Fatalf("unsupported form %#x", form)
}
}
// TestELFObjectDataRelocation checks that a symbol-valued DATA field ("DATA
// s+0(SB)/8, $other(SB)") reaches the ELF object as a .rela.data entry: an
// absolute 64-bit relocation at the field's offset within .data, against
// the named symbol, external targets included.
func TestELFObjectDataRelocation(t *testing.T) {
f, errs := parser.Parse("t_amd64.s", `#include "textflag.h"
TEXT ·Keep(SB), NOSPLIT, $0-8
RET
GLOBL holder(SB), NOPTR, $24
DATA holder+0(SB)/8, $·Keep+5(SB)
DATA holder+8(SB)/8, $holder(SB)
DATA holder+16(SB)/8, $extvar(SB)
`)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFile(f)
if err != nil {
t.Fatalf("assemble: %v", err)
}
obj, err := img.ELFObject()
if err != nil {
t.Fatalf("ELFObject: %v", err)
}
ef, err := elf.NewFile(bytes.NewReader(obj))
if err != nil {
t.Fatalf("parse emitted object: %v", err)
}
defer ef.Close()
relaData := ef.Section(".rela.data")
if relaData == nil {
t.Fatal("missing .rela.data section")
}
if relaData.Link == 0 || ef.Sections[relaData.Link].Name != ".symtab" {
t.Errorf(".rela.data sh_link = %d, want the .symtab index", relaData.Link)
}
if ef.Sections[relaData.Info].Name != ".data" {
t.Errorf(".rela.data sh_info = %d, want the .data index", relaData.Info)
}
relas, err := relaData.Data()
if err != nil {
t.Fatal(err)
}
var got []struct {
off uint64
sym uint32
typ uint32
addend int64
}
for i := 0; i+24 <= len(relas); i += 24 {
got = append(got, struct {
off uint64
sym uint32
typ uint32
addend int64
}{
off: binary.LittleEndian.Uint64(relas[i:]),
// r_info packs the type in the low dword and the symbol index
// in the high dword.
typ: binary.LittleEndian.Uint32(relas[i+8:]),
sym: binary.LittleEndian.Uint32(relas[i+12:]),
addend: int64(binary.LittleEndian.Uint64(relas[i+16:])),
})
}
// debug/elf hides the table's null entry, so raw index s names syms[s-1].
syms, err := ef.Symbols()
if err != nil {
t.Fatal(err)
}
name := func(idx uint32) string {
if idx >= 1 && int(idx) <= len(syms) {
return syms[idx-1].Name
}
return ""
}
// The offsets are data-section-relative: the field's DATA offset plus
// the symbol's position in .data (the layout aligns each symbol to 16).
base := uint64(0)
for _, d := range img.DataSyms {
if d.Name == "holder" {
base = uint64(d.Offset)
}
}
want := []struct {
off uint64
typ uint32
addend int64
target string
}{
{off: base + 0, typ: uint32(elf.R_X86_64_64), addend: 5, target: "Keep"},
{off: base + 8, typ: uint32(elf.R_X86_64_64), addend: 0, target: "holder"},
{off: base + 16, typ: uint32(elf.R_X86_64_64), addend: 0, target: "extvar"},
}
if len(got) != len(want) {
t.Fatalf(".rela.data entries = %d, want %d", len(got), len(want))
}
for i, w := range want {
g := got[i]
if g.off != w.off || g.typ != w.typ || g.addend != w.addend {
t.Errorf("entry %d = {off %d typ %d addend %d}, want {off %d typ %d addend %d}",
i, g.off, g.typ, g.addend, w.off, w.typ, w.addend)
}
if n := name(g.sym); n != w.target {
t.Errorf("entry %d names %q, want %q", i, n, w.target)
}
}
}
+68 -11
View File
@@ -18,12 +18,16 @@ const (
rArm64AddAbsLo12NC = 277 // R_AARCH64_ADD_ABS_LO12_NC (ADD page offset)
rArm64Call26 = 283 // R_AARCH64_CALL26 (BL instruction)
rArm64Ldst64Lo12NC = 286 // R_AARCH64_LDST64_ABS_LO12_NC (64-bit LDR/STR page offset)
// R_AARCH64_ABS32 (debug/elf 258): the absolute 32-bit address of a
// symbol, the R_ADDR shape a 4-byte DATA field carries. ABS64 (257)
// lives with the DWARF fixup constants as rAARCH64Abs64.
rArm64Abs32 = 258
)
// ELFAARCH64Object returns the image as an ELF64 relocatable object file for
// AArch64 (EM_AARCH64, 64-bit, little-endian). The structure mirrors the
// amd64 and RISC-V ELF emitters: .text, .data, .symtab, .strtab and an
// optional .rela.text.
// amd64 and RISC-V ELF emitters: .text, .data, .symtab, .strtab, an
// optional .rela.text and an optional .rela.data.
func (img *Image) ELFAARCH64Object() ([]byte, error) {
le := binary.LittleEndian
@@ -133,6 +137,50 @@ func (img *Image) ELFAARCH64Object() ([]byte, error) {
}
}
// The data symbols' symbol-valued DATA fields ("DATA s+0(SB)/8,
// $other(SB)") become .rela.data entries: an absolute relocation of the
// DATA line's width at the field's data-section offset, S + A with no
// PC term. Widths 4 and 8 have ELF relocation shapes; narrower fields
// cannot hold an address, so they are refused rather than truncated.
var dataRelas []elfRela
for _, d := range img.DataSyms {
for _, r := range d.Relocs {
idx, ok := symIdx[r.Name]
if !ok {
return nil, fmt.Errorf("data relocation references unknown symbol %q", r.Name)
}
var typ uint32
switch r.Siz {
case 8:
typ = rAARCH64Abs64
case 4:
typ = rArm64Abs32
default:
return nil, fmt.Errorf("DATA %q: a symbol value of width %d has no ELF relocation", d.Name, r.Siz)
}
dataRelas = append(dataRelas, elfRela{
off: uint64(d.Offset + r.Off),
sym: idx,
typ: typ,
addend: r.Addend,
})
}
}
// Section presence: .rela.text only when there are code relocations,
// .rela.data only when a DATA line holds a symbol value.
hasRela := len(relas) > 0
hasDataRela := len(dataRelas) > 0
nSections := 6
if hasRela {
nSections++
}
if hasDataRela {
nSections++
}
secSymtab, secStrtab := 3, 4
secShstr := nSections - 1
// String tables.
stNames := newElfStrtab()
for _, s := range syms {
@@ -142,18 +190,13 @@ func (img *Image) ELFAARCH64Object() ([]byte, error) {
for _, n := range []string{".text", ".data", ".symtab", ".strtab", ".rela.text", ".shstrtab"} {
stSections.add(n)
}
if hasDataRela {
stSections.add(".rela.data")
}
for _, n := range dwarfSectionNames {
stSections.add(n)
}
hasRela := len(relas) > 0
nSections := 6
if hasRela {
nSections = 7
}
secSymtab, secStrtab := 3, 4
secShstr := nSections - 1
// Layout.
var out []byte
out = append(out, make([]byte, 64)...)
@@ -188,7 +231,7 @@ func (img *Image) ELFAARCH64Object() ([]byte, error) {
strtabOff := len(out)
out = append(out, stNames.bytes()...)
var relaOff int
var relaOff, relaDataOff int
if hasRela {
align(8)
relaOff = len(out)
@@ -200,6 +243,17 @@ func (img *Image) ELFAARCH64Object() ([]byte, error) {
out = append(out, b[:]...)
}
}
if hasDataRela {
align(8)
relaDataOff = len(out)
for _, r := range dataRelas {
var b [24]byte
le.PutUint64(b[0:], r.off)
le.PutUint64(b[8:], uint64(r.sym)<<32|uint64(r.typ))
le.PutUint64(b[16:], uint64(r.addend))
out = append(out, b[:]...)
}
}
shstrOff := len(out)
out = append(out, stSections.bytes()...)
@@ -257,6 +311,9 @@ func (img *Image) ELFAARCH64Object() ([]byte, error) {
if hasRela {
putSh(".rela.text", shtRela, 0, relaOff, 24*len(relas), secSymtab, secText, 8, 24)
}
if hasDataRela {
putSh(".rela.data", shtRela, 0, relaDataOff, 24*len(dataRelas), secSymtab, secData, 8, 24)
}
putSh(".shstrtab", shtStrtab, 0, shstrOff, len(stSections.bytes()), 0, 0, 1, 0)
// DWARF section headers; their indices follow the write order.
if dw != nil {
+115
View File
@@ -197,3 +197,118 @@ TEXT ·add(SB), NOSPLIT, $0-24
t.Error("unexpected .rela.text section when there are no relocations")
}
}
// TestELFAARCH64ObjectDataRelocation checks that a symbol-valued DATA field
// ("DATA s+0(SB)/8, $other(SB)") reaches the AArch64 ELF object as a
// .rela.data entry: an R_AARCH64_ABS64 (ABS32 for a width-4 field) at the
// field's offset within .data, against the named symbol, external targets
// included.
func TestELFAARCH64ObjectDataRelocation(t *testing.T) {
f, errs := parser.Parse("t_arm64.s", `#include "textflag.h"
TEXT ·Keep(SB), NOSPLIT, $0-0
RET
GLOBL holder(SB), NOPTR, $32
DATA holder+0(SB)/8, $·Keep+5(SB)
DATA holder+8(SB)/8, $holder(SB)
DATA holder+16(SB)/8, $extvar(SB)
DATA holder+24(SB)/4, $Keep(SB)
`)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileARM64(f)
if err != nil {
t.Fatalf("AssembleFileARM64: %v", err)
}
obj, err := img.ELFAARCH64Object()
if err != nil {
t.Fatalf("ELFAARCH64Object: %v", err)
}
checkELFSectionAccounting(t, obj)
ef, err := elf.NewFile(bytes.NewReader(obj))
if err != nil {
t.Fatalf("parse emitted object: %v", err)
}
defer ef.Close()
relaData := ef.Section(".rela.data")
if relaData == nil {
t.Fatal("missing .rela.data section")
}
if relaData.Type != elf.SHT_RELA {
t.Errorf(".rela.data type = %v, want SHT_RELA", relaData.Type)
}
if relaData.Link == 0 || ef.Sections[relaData.Link].Name != ".symtab" {
t.Errorf(".rela.data sh_link = %d, want the .symtab index", relaData.Link)
}
if ef.Sections[relaData.Info].Name != ".data" {
t.Errorf(".rela.data sh_info = %d, want the .data index", relaData.Info)
}
relas, err := relaData.Data()
if err != nil {
t.Fatal(err)
}
var got []struct {
off uint64
sym uint32
typ uint32
addend int64
}
for i := 0; i+24 <= len(relas); i += 24 {
got = append(got, struct {
off uint64
sym uint32
typ uint32
addend int64
}{
off: binary.LittleEndian.Uint64(relas[i:]),
// r_info packs the type in the low dword and the symbol index
// in the high dword.
typ: binary.LittleEndian.Uint32(relas[i+8:]),
sym: binary.LittleEndian.Uint32(relas[i+12:]),
addend: int64(binary.LittleEndian.Uint64(relas[i+16:])),
})
}
// debug/elf hides the table's null entry, so raw index s names syms[s-1].
syms, err := ef.Symbols()
if err != nil {
t.Fatal(err)
}
name := func(idx uint32) string {
if idx >= 1 && int(idx) <= len(syms) {
return syms[idx-1].Name
}
return ""
}
// The offsets are data-section-relative: the field's DATA offset plus
// the symbol's position in .data (the layout aligns each symbol to 16).
base := uint64(0)
for _, d := range img.DataSyms {
if d.Name == "holder" {
base = uint64(d.Offset)
}
}
want := []struct {
off uint64
typ uint32
addend int64
target string
}{
{off: base + 0, typ: uint32(elf.R_AARCH64_ABS64), addend: 5, target: "Keep"},
{off: base + 8, typ: uint32(elf.R_AARCH64_ABS64), addend: 0, target: "holder"},
{off: base + 16, typ: uint32(elf.R_AARCH64_ABS64), addend: 0, target: "extvar"},
{off: base + 24, typ: uint32(elf.R_AARCH64_ABS32), addend: 0, target: "Keep"},
}
if len(got) != len(want) {
t.Fatalf(".rela.data entries = %d, want %d", len(got), len(want))
}
for i, w := range want {
g := got[i]
if g.off != w.off || g.typ != w.typ || g.addend != w.addend {
t.Errorf("entry %d = {off %d typ %d addend %d}, want {off %d typ %d addend %d}",
i, g.off, g.typ, g.addend, w.off, w.typ, w.addend)
}
if n := name(g.sym); n != w.target {
t.Errorf("entry %d names %q, want %q", i, n, w.target)
}
}
}
+68 -11
View File
@@ -23,12 +23,16 @@ const (
rLarchPCALAHI20 = 71 // R_LARCH_PCALA_HI20 (pcalau12i)
rLarchPCALALO12 = 72 // R_LARCH_PCALA_LO12 (addi.d/ld/st)
rLarchB26 = 66 // R_LARCH_B26 (b/bl, matches the Go linker's mapping)
// R_LARCH_32 (debug/elf 1): the absolute 32-bit address of a symbol,
// the R_ADDR shape a 4-byte DATA field carries. R_LARCH_64 (2) lives
// with the DWARF fixup constants as rLarchAbs64.
rLarchAbs32 = 1
)
// ELFLOONG64Object returns the image as an ELF64 relocatable object file for
// LoongArch (EM_LOONGARCH, 64-bit, little-endian). The structure mirrors the
// amd64 and RISC-V ELF emitters: .text, .data, .symtab, .strtab and an
// optional .rela.text.
// amd64 and RISC-V ELF emitters: .text, .data, .symtab, .strtab, an
// optional .rela.text and an optional .rela.data.
func (img *Image) ELFLOONG64Object() ([]byte, error) {
le := binary.LittleEndian
@@ -117,6 +121,50 @@ func (img *Image) ELFLOONG64Object() ([]byte, error) {
}
}
// The data symbols' symbol-valued DATA fields ("DATA s+0(SB)/8,
// $other(SB)") become .rela.data entries: an absolute relocation of the
// DATA line's width at the field's data-section offset, S + A with no
// PC term. Widths 4 and 8 have ELF relocation shapes; narrower fields
// cannot hold an address, so they are refused rather than truncated.
var dataRelas []elfRela
for _, d := range img.DataSyms {
for _, r := range d.Relocs {
idx, ok := symIdx[r.Name]
if !ok {
return nil, fmt.Errorf("data relocation references unknown symbol %q", r.Name)
}
var typ uint32
switch r.Siz {
case 8:
typ = rLarchAbs64
case 4:
typ = rLarchAbs32
default:
return nil, fmt.Errorf("DATA %q: a symbol value of width %d has no ELF relocation", d.Name, r.Siz)
}
dataRelas = append(dataRelas, elfRela{
off: uint64(d.Offset + r.Off),
sym: idx,
typ: typ,
addend: r.Addend,
})
}
}
// Section presence: .rela.text only when there are code relocations,
// .rela.data only when a DATA line holds a symbol value.
hasRela := len(relas) > 0
hasDataRela := len(dataRelas) > 0
nSections := 6
if hasRela {
nSections++
}
if hasDataRela {
nSections++
}
secSymtab, secStrtab := 3, 4
secShstr := nSections - 1
// String tables.
stNames := newElfStrtab()
for _, s := range syms {
@@ -126,18 +174,13 @@ func (img *Image) ELFLOONG64Object() ([]byte, error) {
for _, n := range []string{".text", ".data", ".symtab", ".strtab", ".rela.text", ".shstrtab"} {
stSections.add(n)
}
if hasDataRela {
stSections.add(".rela.data")
}
for _, n := range dwarfSectionNames {
stSections.add(n)
}
hasRela := len(relas) > 0
nSections := 6
if hasRela {
nSections = 7
}
secSymtab, secStrtab := 3, 4
secShstr := nSections - 1
// Layout.
var out []byte
out = append(out, make([]byte, 64)...)
@@ -172,7 +215,7 @@ func (img *Image) ELFLOONG64Object() ([]byte, error) {
strtabOff := len(out)
out = append(out, stNames.bytes()...)
var relaOff int
var relaOff, relaDataOff int
if hasRela {
align(8)
relaOff = len(out)
@@ -184,6 +227,17 @@ func (img *Image) ELFLOONG64Object() ([]byte, error) {
out = append(out, b[:]...)
}
}
if hasDataRela {
align(8)
relaDataOff = len(out)
for _, r := range dataRelas {
var b [24]byte
le.PutUint64(b[0:], r.off)
le.PutUint64(b[8:], uint64(r.sym)<<32|uint64(r.typ))
le.PutUint64(b[16:], uint64(r.addend))
out = append(out, b[:]...)
}
}
shstrOff := len(out)
out = append(out, stSections.bytes()...)
@@ -239,6 +293,9 @@ func (img *Image) ELFLOONG64Object() ([]byte, error) {
if hasRela {
putSh(".rela.text", shtRela, 0, relaOff, 24*len(relas), secSymtab, secText, 8, 24)
}
if hasDataRela {
putSh(".rela.data", shtRela, 0, relaDataOff, 24*len(dataRelas), secSymtab, secData, 8, 24)
}
putSh(".shstrtab", shtStrtab, 0, shstrOff, len(stSections.bytes()), 0, 0, 1, 0)
// DWARF section headers; their indices follow the write order.
if dw != nil {
+115
View File
@@ -245,3 +245,118 @@ func TestELFLOONG64BranchRelocation(t *testing.T) {
}
}
}
// TestELFLOONG64ObjectDataRelocation checks that a symbol-valued DATA field
// ("DATA s+0(SB)/8, $other(SB)") reaches the LoongArch ELF object as a
// .rela.data entry: an R_LARCH_64 (R_LARCH_32 for a width-4 field) at the
// field's offset within .data, against the named symbol, external targets
// included.
func TestELFLOONG64ObjectDataRelocation(t *testing.T) {
f, errs := parser.Parse("t_loong64.s", `#include "textflag.h"
TEXT ·Keep(SB), NOSPLIT, $0-0
RET
GLOBL holder(SB), NOPTR, $32
DATA holder+0(SB)/8, $·Keep+5(SB)
DATA holder+8(SB)/8, $holder(SB)
DATA holder+16(SB)/8, $extvar(SB)
DATA holder+24(SB)/4, $Keep(SB)
`)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileLOONG64(f)
if err != nil {
t.Fatalf("AssembleFileLOONG64: %v", err)
}
obj, err := img.ELFLOONG64Object()
if err != nil {
t.Fatalf("ELFLOONG64Object: %v", err)
}
checkELFSectionAccounting(t, obj)
ef, err := elf.NewFile(bytes.NewReader(obj))
if err != nil {
t.Fatalf("parse emitted object: %v", err)
}
defer ef.Close()
relaData := ef.Section(".rela.data")
if relaData == nil {
t.Fatal("missing .rela.data section")
}
if relaData.Type != elf.SHT_RELA {
t.Errorf(".rela.data type = %v, want SHT_RELA", relaData.Type)
}
if relaData.Link == 0 || ef.Sections[relaData.Link].Name != ".symtab" {
t.Errorf(".rela.data sh_link = %d, want the .symtab index", relaData.Link)
}
if ef.Sections[relaData.Info].Name != ".data" {
t.Errorf(".rela.data sh_info = %d, want the .data index", relaData.Info)
}
relas, err := relaData.Data()
if err != nil {
t.Fatal(err)
}
var got []struct {
off uint64
sym uint32
typ uint32
addend int64
}
for i := 0; i+24 <= len(relas); i += 24 {
got = append(got, struct {
off uint64
sym uint32
typ uint32
addend int64
}{
off: binary.LittleEndian.Uint64(relas[i:]),
// r_info packs the type in the low dword and the symbol index
// in the high dword.
typ: binary.LittleEndian.Uint32(relas[i+8:]),
sym: binary.LittleEndian.Uint32(relas[i+12:]),
addend: int64(binary.LittleEndian.Uint64(relas[i+16:])),
})
}
// debug/elf hides the table's null entry, so raw index s names syms[s-1].
syms, err := ef.Symbols()
if err != nil {
t.Fatal(err)
}
name := func(idx uint32) string {
if idx >= 1 && int(idx) <= len(syms) {
return syms[idx-1].Name
}
return ""
}
// The offsets are data-section-relative: the field's DATA offset plus
// the symbol's position in .data (the layout aligns each symbol to 16).
base := uint64(0)
for _, d := range img.DataSyms {
if d.Name == "holder" {
base = uint64(d.Offset)
}
}
want := []struct {
off uint64
typ uint32
addend int64
target string
}{
{off: base + 0, typ: uint32(elf.R_LARCH_64), addend: 5, target: "Keep"},
{off: base + 8, typ: uint32(elf.R_LARCH_64), addend: 0, target: "holder"},
{off: base + 16, typ: uint32(elf.R_LARCH_64), addend: 0, target: "extvar"},
{off: base + 24, typ: uint32(elf.R_LARCH_32), addend: 0, target: "Keep"},
}
if len(got) != len(want) {
t.Fatalf(".rela.data entries = %d, want %d", len(got), len(want))
}
for i, w := range want {
g := got[i]
if g.off != w.off || g.typ != w.typ || g.addend != w.addend {
t.Errorf("entry %d = {off %d typ %d addend %d}, want {off %d typ %d addend %d}",
i, g.off, g.typ, g.addend, w.off, w.typ, w.addend)
}
if n := name(g.sym); n != w.target {
t.Errorf("entry %d names %q, want %q", i, n, w.target)
}
}
}
+68 -10
View File
@@ -24,11 +24,16 @@ const (
rRISCVPCRELHI20 = 23 // R_RISCV_PCREL_HI20
rRISCVPCRELLO12I = 24 // R_RISCV_PCREL_LO12_I
rRISCVPCRELLO12S = 25 // R_RISCV_PCREL_LO12_S
// R_RISCV_32 (debug/elf 1): the absolute 32-bit address of a symbol,
// the R_ADDR shape a 4-byte DATA field carries. R_RISCV_64 (2) lives
// with the DWARF fixup constants as rRISCVAbs64.
rRISVCAbs32 = 1
)
// ELFRISCVObject returns the image as an ELF64 relocatable object file for
// RISC-V (EM_RISCV, 64-bit, little-endian). The structure mirrors the amd64
// ELF emission: .text, .data, .symtab, .strtab and optional .rela.text.
// ELF emission: .text, .data, .symtab, .strtab, an optional .rela.text and
// an optional .rela.data.
func (img *Image) ELFRISCVObject() ([]byte, error) {
le := binary.LittleEndian
@@ -129,6 +134,50 @@ func (img *Image) ELFRISCVObject() ([]byte, error) {
}
}
// The data symbols' symbol-valued DATA fields ("DATA s+0(SB)/8,
// $other(SB)") become .rela.data entries: an absolute relocation of the
// DATA line's width at the field's data-section offset, S + A with no
// PC term. Widths 4 and 8 have ELF relocation shapes; narrower fields
// cannot hold an address, so they are refused rather than truncated.
var dataRelas []elfRela
for _, d := range img.DataSyms {
for _, r := range d.Relocs {
idx, ok := symIdx[r.Name]
if !ok {
return nil, fmt.Errorf("data relocation references unknown symbol %q", r.Name)
}
var typ uint32
switch r.Siz {
case 8:
typ = rRISCVAbs64
case 4:
typ = rRISVCAbs32
default:
return nil, fmt.Errorf("DATA %q: a symbol value of width %d has no ELF relocation", d.Name, r.Siz)
}
dataRelas = append(dataRelas, elfRela{
off: uint64(d.Offset + r.Off),
sym: idx,
typ: typ,
addend: r.Addend,
})
}
}
// Section presence: .rela.text only when there are code relocations,
// .rela.data only when a DATA line holds a symbol value.
hasRela := len(relas) > 0
hasDataRela := len(dataRelas) > 0
nSections := 6
if hasRela {
nSections++
}
if hasDataRela {
nSections++
}
secSymtab, secStrtab := 3, 4
secShstr := nSections - 1
// String tables.
stNames := newElfStrtab()
for _, s := range syms {
@@ -138,18 +187,13 @@ func (img *Image) ELFRISCVObject() ([]byte, error) {
for _, n := range []string{".text", ".data", ".symtab", ".strtab", ".rela.text", ".shstrtab"} {
stSections.add(n)
}
if hasDataRela {
stSections.add(".rela.data")
}
for _, n := range dwarfSectionNames {
stSections.add(n)
}
hasRela := len(relas) > 0
nSections := 6
if hasRela {
nSections = 7
}
secSymtab, secStrtab := 3, 4
secShstr := nSections - 1
// Layout.
var out []byte
out = append(out, make([]byte, 64)...)
@@ -184,7 +228,7 @@ func (img *Image) ELFRISCVObject() ([]byte, error) {
strtabOff := len(out)
out = append(out, stNames.bytes()...)
var relaOff int
var relaOff, relaDataOff int
if hasRela {
align(8)
relaOff = len(out)
@@ -196,6 +240,17 @@ func (img *Image) ELFRISCVObject() ([]byte, error) {
out = append(out, b[:]...)
}
}
if hasDataRela {
align(8)
relaDataOff = len(out)
for _, r := range dataRelas {
var b [24]byte
le.PutUint64(b[0:], r.off)
le.PutUint64(b[8:], uint64(r.sym)<<32|uint64(r.typ))
le.PutUint64(b[16:], uint64(r.addend))
out = append(out, b[:]...)
}
}
shstrOff := len(out)
out = append(out, stSections.bytes()...)
@@ -251,6 +306,9 @@ func (img *Image) ELFRISCVObject() ([]byte, error) {
if hasRela {
putSh(".rela.text", shtRela, 0, relaOff, 24*len(relas), secSymtab, secText, 8, 24)
}
if hasDataRela {
putSh(".rela.data", shtRela, 0, relaDataOff, 24*len(dataRelas), secSymtab, secData, 8, 24)
}
putSh(".shstrtab", shtStrtab, 0, shstrOff, len(stSections.bytes()), 0, 0, 1, 0)
// DWARF section headers; their indices follow the write order.
if dw != nil {
+128
View File
@@ -0,0 +1,128 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
package asm
import (
"bytes"
"debug/elf"
"encoding/binary"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
)
// TestELFRISCVObjectDataRelocation checks that a symbol-valued DATA field
// ("DATA s+0(SB)/8, $other(SB)") reaches the RISC-V ELF object as a
// .rela.data entry: an R_RISCV_64 (R_RISCV_32 for a width-4 field) at the
// field's offset within .data, against the named symbol, external targets
// included.
func TestELFRISCVObjectDataRelocation(t *testing.T) {
f, errs := parser.Parse("t_riscv64.s", `#include "textflag.h"
TEXT ·Keep(SB), NOSPLIT, $0-0
RET
GLOBL holder(SB), NOPTR, $32
DATA holder+0(SB)/8, $·Keep+5(SB)
DATA holder+8(SB)/8, $holder(SB)
DATA holder+16(SB)/8, $extvar(SB)
DATA holder+24(SB)/4, $Keep(SB)
`)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileRISCV(f)
if err != nil {
t.Fatalf("AssembleFileRISCV: %v", err)
}
obj, err := img.ELFRISCVObject()
if err != nil {
t.Fatalf("ELFRISCVObject: %v", err)
}
checkELFSectionAccounting(t, obj)
ef, err := elf.NewFile(bytes.NewReader(obj))
if err != nil {
t.Fatalf("parse emitted object: %v", err)
}
defer ef.Close()
relaData := ef.Section(".rela.data")
if relaData == nil {
t.Fatal("missing .rela.data section")
}
if relaData.Type != elf.SHT_RELA {
t.Errorf(".rela.data type = %v, want SHT_RELA", relaData.Type)
}
if relaData.Link == 0 || ef.Sections[relaData.Link].Name != ".symtab" {
t.Errorf(".rela.data sh_link = %d, want the .symtab index", relaData.Link)
}
if ef.Sections[relaData.Info].Name != ".data" {
t.Errorf(".rela.data sh_info = %d, want the .data index", relaData.Info)
}
relas, err := relaData.Data()
if err != nil {
t.Fatal(err)
}
var got []struct {
off uint64
sym uint32
typ uint32
addend int64
}
for i := 0; i+24 <= len(relas); i += 24 {
got = append(got, struct {
off uint64
sym uint32
typ uint32
addend int64
}{
off: binary.LittleEndian.Uint64(relas[i:]),
// r_info packs the type in the low dword and the symbol index
// in the high dword.
typ: binary.LittleEndian.Uint32(relas[i+8:]),
sym: binary.LittleEndian.Uint32(relas[i+12:]),
addend: int64(binary.LittleEndian.Uint64(relas[i+16:])),
})
}
// debug/elf hides the table's null entry, so raw index s names syms[s-1].
syms, err := ef.Symbols()
if err != nil {
t.Fatal(err)
}
name := func(idx uint32) string {
if idx >= 1 && int(idx) <= len(syms) {
return syms[idx-1].Name
}
return ""
}
// The offsets are data-section-relative: the field's DATA offset plus
// the symbol's position in .data (the layout aligns each symbol to 16).
base := uint64(0)
for _, d := range img.DataSyms {
if d.Name == "holder" {
base = uint64(d.Offset)
}
}
want := []struct {
off uint64
typ uint32
addend int64
target string
}{
{off: base + 0, typ: uint32(elf.R_RISCV_64), addend: 5, target: "Keep"},
{off: base + 8, typ: uint32(elf.R_RISCV_64), addend: 0, target: "holder"},
{off: base + 16, typ: uint32(elf.R_RISCV_64), addend: 0, target: "extvar"},
{off: base + 24, typ: uint32(elf.R_RISCV_32), addend: 0, target: "Keep"},
}
if len(got) != len(want) {
t.Fatalf(".rela.data entries = %d, want %d", len(got), len(want))
}
for i, w := range want {
g := got[i]
if g.off != w.off || g.typ != w.typ || g.addend != w.addend {
t.Errorf("entry %d = {off %d typ %d addend %d}, want {off %d typ %d addend %d}",
i, g.off, g.typ, g.addend, w.off, w.typ, w.addend)
}
if n := name(g.sym); n != w.target {
t.Errorf("entry %d names %q, want %q", i, n, w.target)
}
}
}
+3 -3
View File
@@ -20,9 +20,9 @@ func Encodable(mnemonic string) bool {
switch upper {
case "RET", "NOP", "CALL", "JMP",
"POPFQ", "PUSHFQ", "INT", "LDMXCSR", "STMXCSR", "CMPSD", "SHA256RNDS2",
// The literal-data pseudo-ops, the accepted-and-ignored END and the
// SP adjust.
"BYTE", "WORD", "LONG", "QUAD", "END", "ADJSP":
// The literal-data pseudo-ops, the accepted-and-ignored END and
// bookkeeping statements, and the SP adjust.
"BYTE", "WORD", "LONG", "QUAD", "END", "ADJSP", "FUNCDATA", "PCDATA":
return true
}
if _, ok := noOperandTable[upper]; ok {
+229 -2
View File
@@ -5,6 +5,8 @@ package asm
import (
"fmt"
"math"
"strconv"
"strings"
)
@@ -21,6 +23,39 @@ func Encode(mnemonic string, ops ...Operand) ([]byte, error) {
type enc struct {
out []byte
patches []encPatch // disp32 fields awaiting static-symbol resolution
// FloatPool collects the pooled constants the floating-point
// immediates reference, in first-use order.
floatPool []floatPoolEntry
floatPoolSeen map[string]bool
}
// floatPoolEntry is one pooled floating-point constant: the symbol name
// the emitted RIP-relative load refers to and its IEEE-754 bytes.
type floatPoolEntry struct {
name string
data []byte
}
// addFloatPool records a pooled constant, deduplicated by symbol name.
func (e *enc) addFloatPool(name string, bits uint64, width int) {
if e.floatPoolSeen == nil {
e.floatPoolSeen = map[string]bool{}
}
if e.floatPoolSeen[name] {
return
}
e.floatPoolSeen[name] = true
data := make([]byte, width)
for i := range width {
data[i] = byte(bits >> (8 * i))
}
e.floatPool = append(e.floatPool, floatPoolEntry{name: name, data: data})
}
// floatPoolList returns the pooled constants in first-use order.
func (e *enc) floatPoolList() []floatPoolEntry {
return e.floatPool
}
// encPatch marks a 4-byte displacement field in enc.out that must receive the
@@ -29,6 +64,7 @@ type encPatch struct {
off int
name string
addend int64
tls bool // a TLS slot offset: the patch is R_TLSLE with no symbol
}
func (e *enc) encode(mnem string, ops []Operand) error {
@@ -102,6 +138,11 @@ func (e *enc) encode(mnem string, ops []Operand) error {
return e.encodeEnd(ops)
case "ADJSP":
return e.encodeAdjsp(ops)
// The runtime's bookkeeping statements carry no text bytes: go tool asm
// records FUNCDATA and PCDATA in the program list only, so the encoded
// body shows nothing, on every architecture.
case "FUNCDATA", "PCDATA":
return e.encodeFuncdata(upper, ops)
}
// VEX (AVX/AVX2) and EVEX (AVX-512) instructions: the trailing
@@ -113,6 +154,7 @@ func (e *enc) encode(mnem string, ops []Operand) error {
return err
}
if isVex(base) || isEvex(base) || isKOp(base) || isGather(base) || isScatter(base) ||
isEvexPrefGather(base) ||
base == "KMOVW" || base == "KMOVQ" || base == "KMOVB" || base == "KMOVD" {
return e.encodeVec(base, ops, sfx)
}
@@ -140,11 +182,18 @@ func (e *enc) encode(mnem string, ops []Operand) error {
}
// Legacy SSE packed binaries dispatch on the full name: the packed
// integer mnemonics carry real width suffixes (PADDB/PCMPGTW/...),
// which the size split must not eat.
// which the size split must not eat. A floating-point immediate
// rewrites into a pooled-constant read on the scalar members.
if m, ok := sseBinTable[upper]; ok {
if f, isFloat := floatImmOperand(ops); isFloat {
return e.encodeSSEFloatBin(upper, m, f, ops)
}
return e.encodeSSEBin(m, ops)
}
if m, ok := sseBinTable[base]; ok {
if f, isFloat := floatImmOperand(ops); isFloat {
return e.encodeSSEFloatBin(upper, m, f, ops)
}
return e.encodeSSEBin(m, ops)
}
// The imm8-controlled legacy instructions, the lane extracts and inserts
@@ -221,7 +270,12 @@ func (e *enc) encode(mnem string, ops []Operand) error {
return e.encodeCvtInt(base, ops, size)
case "FMOVD":
return e.encodeFmov(ops)
case "MOVOU", "MOVO", "MOVOA", "MOVUPS", "MOVAPS", "MOVUPD", "MOVAPD", "MOVSD", "MOVSS":
case "MOVSD", "MOVSS":
if f, isFloat := floatImmOperand(ops); isFloat {
return e.encodeSSEFloatMove(upper, f, ops)
}
return e.encodeSSEMove(sseMoveTable[base], ops)
case "MOVOU", "MOVO", "MOVOA", "MOVUPS", "MOVAPS", "MOVUPD", "MOVAPD":
return e.encodeSSEMove(sseMoveTable[base], ops)
}
return fmt.Errorf("unsupported instruction %q", mnem)
@@ -284,6 +338,33 @@ func (e *enc) encodeData(mnem string, ops []Operand) error {
return nil
}
// encodeFuncdata accepts-and-ignores the runtime bookkeeping statements:
// FUNCDATA $n, sym(SB) and PCDATA $n, $m. go tool asm emits no text bytes
// for either (the entries live in the object's ancillary tables, not the
// function body), and the operand shapes it takes are exactly these: an
// integer count first, then a symbol reference for FUNCDATA and an integer
// value for PCDATA. The other architectures accept-and-ignore the same
// statements; amd64 now matches.
func (e *enc) encodeFuncdata(upper string, ops []Operand) error {
if len(ops) != 2 {
return fmt.Errorf("%s expects 2 operands, got %d", upper, len(ops))
}
if _, ok := ops[0].(Imm); !ok {
return fmt.Errorf("%s: first operand must be an integer immediate", upper)
}
switch upper {
case "FUNCDATA":
if _, ok := ops[1].(sbMem); !ok {
return fmt.Errorf("FUNCDATA: second operand must be a symbol reference")
}
case "PCDATA":
if _, ok := ops[1].(Imm); !ok {
return fmt.Errorf("PCDATA: second operand must be an integer immediate")
}
}
return nil
}
// encodeEnd accepts-and-ignores END. go tool asm drops the statement
// entirely: the AEND Prog is skipped when the program list is flushed, so
// the statements after an END still belong to the same function and the
@@ -318,6 +399,120 @@ func (e *enc) encodeAdjsp(ops []Operand) error {
return nil
}
// --- floating-point immediates ----------------------------------------------
// sseFloatImm lists the mnemonics whose first operand may be a floating-point
// immediate, the set go tool asm rewrites into a pooled-constant read: the
// scalar moves, the four scalar arithmetic pairs and the scalar compares.
// The packed members and the uniform forms (MAXSD, MINSD, SQRTSD, CMPSD)
// reject the immediate in the toolchain and are absent here on purpose.
var sseFloatImm = map[string]bool{
"MOVSD": true, "MOVSS": true,
"ADDSD": true, "ADDSS": true,
"SUBSD": true, "SUBSS": true,
"MULSD": true, "MULSS": true,
"DIVSD": true, "DIVSS": true,
"COMISD": true, "COMISS": true,
"UCOMISD": true, "UCOMISS": true,
}
// floatImmOperand reports whether the operand list opens with a
// floating-point immediate in the two-operand spelling (imm, dst).
func floatImmOperand(ops []Operand) (FloatImm, bool) {
if len(ops) != 2 {
return FloatImm{}, false
}
f, ok := ops[0].(FloatImm)
return f, ok
}
// floatPoolValue evaluates a floating-point immediate at the width its
// mnemonic encodes and names the pool constant the toolchain synthesises:
// $f64.<16 hex> for the doubles, $f32.<8 hex> for the singles (the float32
// rounding of the parsed value). The name carries the IEEE-754 bits; the
// section holds them little-endian.
func floatPoolValue(mnem string, f FloatImm) (bits uint64, name string, err error) {
v, err := strconv.ParseFloat(f.Text, 64)
if err != nil {
return 0, "", fmt.Errorf("invalid floating-point immediate %q", f.Text)
}
if f.Neg {
v = -v
}
if strings.HasSuffix(mnem, "D") {
bits = math.Float64bits(v)
return bits, fmt.Sprintf("$f64.%016x", bits), nil
}
bits = uint64(math.Float32bits(float32(v)))
return bits, fmt.Sprintf("$f32.%08x", bits), nil
}
// encodeSSEFloatMove encodes MOVSD/MOVSS with a floating-point immediate
// source. A positive zero needs no memory read: the toolchain emits
// XORPS dst, dst. Anything else loads the pooled constant RIP-relative
// ($f64.<hex>(SB) / $f32.<hex>(SB)), the displacement a patch site the
// file-level layout or the linker resolves.
func (e *enc) encodeSSEFloatMove(mnem string, f FloatImm, ops []Operand) error {
if !sseFloatImm[mnem] {
return fmt.Errorf("%s does not take a floating-point immediate", mnem)
}
dst, ok := ops[1].(Reg)
if !ok || !dst.isVec() {
return fmt.Errorf("%s: destination must be a vector register", mnem)
}
bits, name, err := floatPoolValue(mnem, f)
if err != nil {
return err
}
e.addFloatPool(name, bits, mwidth(mnem))
if bits == 0 {
i := &instr{opcode: []byte{0x0F, 0x57}, modrm: -1, sib: -1} // XORPS
if err := setRM(i, dst, dst, 8); err != nil {
return err
}
return e.emit(i)
}
m := sseMoveTable[mnem]
i := &instr{prefix: m.prefix, opcode: []byte{0x0F, m.load}, modrm: -1, sib: -1}
if err := setRM(i, dst, sbMem{size: mwidth(mnem), name: name}, 8); err != nil {
return err
}
return e.emit(i)
}
// encodeSSEFloatBin encodes the scalar arithmetic and compare mnemonics with
// a floating-point immediate source: the constant is read from the pool into
// the instruction's r/m side (reg = destination), the rewrite go tool asm
// performs at the source level.
func (e *enc) encodeSSEFloatBin(mnem string, m sseBin, f FloatImm, ops []Operand) error {
if !sseFloatImm[mnem] {
return fmt.Errorf("%s does not take a floating-point immediate", mnem)
}
dst, ok := ops[1].(Reg)
if !ok || !dst.isVec() {
return fmt.Errorf("%s: destination must be a vector register", mnem)
}
bits, name, err := floatPoolValue(mnem, f)
if err != nil {
return err
}
e.addFloatPool(name, bits, mwidth(mnem))
i := &instr{prefix: m.prefix, opcode: []byte{0x0F, m.op}, modrm: -1, sib: -1}
if err := setRM(i, dst, sbMem{size: mwidth(mnem), name: name}, 8); err != nil {
return err
}
return e.emit(i)
}
// mwidth returns the operand width a scalar SSE mnemonic encodes: the double
// spellings end in D, the single spellings in S.
func mwidth(mnem string) int {
if strings.HasSuffix(mnem, "D") {
return 8
}
return 4
}
// splitSize separates a trailing B/W/L/Q size suffix from the mnemonic.
func splitSize(upper string) (base string, size int) {
if upper == "" {
@@ -384,6 +579,7 @@ type instr struct {
disp []byte
imm []byte
sb *sbRef // static-symbol displacement in disp, awaiting resolution
tls bool // the displacement is a TLS slot offset, patched R_TLSLE
}
// sbRef records that an instruction's displacement refers to a static symbol
@@ -426,6 +622,9 @@ func (e *enc) emit(i *instr) error {
if i.sb != nil {
e.patches = append(e.patches, encPatch{off: len(e.out), name: i.sb.name, addend: i.sb.addend})
}
if i.tls {
e.patches = append(e.patches, encPatch{off: len(e.out), tls: true})
}
e.out = append(e.out, i.disp...)
e.out = append(e.out, i.imm...)
return nil
@@ -480,12 +679,30 @@ func setRMReg(i *instr, regField int, rexR, regForced bool, rm Operand, opSize i
i.disp = le32(0)
i.sb = &sbRef{name: r.name, addend: r.addend}
return nil
case TLSMem:
// off(TLS): the segment-prefixed absolute access, mod=00 with the
// SIB escape's disp32 absolute form. The displacement is the TLS
// slot offset, patched by the linker's TLS relocation.
i.prefix = r.Seg
i.modrm = 0x04 | regField<<3
i.sib = 0x25
i.disp = le32(r.Disp)
i.tls = true
return nil
case SegAbs:
// 0x30(GS): the segment override with the SIB escape's disp32
// absolute form, no relocation.
setSegAbs(i, regField, r)
return nil
default:
return fmt.Errorf("invalid r/m operand %T", rm)
}
}
func setMem(i *instr, regField int, m Mem) error {
if m.Seg != 0 {
i.prefix = m.Seg
}
modrm, sib, disp, xBit, bBit, err := memComponents(regField, m)
if err != nil {
return err
@@ -498,6 +715,16 @@ func setMem(i *instr, regField int, m Mem) error {
return nil
}
// setSegAbs assembles a segment-absolute operand, 0x30(GS): the segment
// override with the mod=00 SIB escape's disp32 absolute form and no
// relocation.
func setSegAbs(i *instr, regField int, m SegAbs) {
i.prefix = m.Seg
i.modrm = 0x04 | regField<<3
i.sib = 0x25
i.disp = le32(m.Disp)
}
// memComponents computes the ModR/M byte (with the given reg field), the SIB
// byte (-1 if none), the displacement bytes, and the high index/base bits, for
// a memory operand. It is shared by the REX (scalar) and VEX (vector) paths.
+157
View File
@@ -9,6 +9,9 @@ import (
"testing"
"golang.org/x/arch/x86/x86asm"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
)
// decode encodes an instruction and decodes it back, returning the decoded
@@ -1064,3 +1067,157 @@ func TestAdjsp(t *testing.T) {
t.Error("ADJSP AX assembled, want an error")
}
}
// TestFloatImmediateGroundTruth pins the floating-point immediate rewrite
// byte for byte against go tool asm: the scalar moves and the scalar
// arithmetic read the constant from a synthesised read-only pool symbol
// ($f64.<hex>, $f32.<hex>) RIP-relative with the displacement left to the
// relocation, and a positive zero on the moves collapses to XORPS dst, dst.
func TestFloatImmediateGroundTruth(t *testing.T) {
cases := []struct {
name string
mnem string
ops []Operand
want string
}{
{"MOVSD -1.0", "MOVSD", []Operand{FloatImm{Text: "1.0", Neg: true}, vreg(t, "X2")}, "f20f101500000000"},
{"MOVSD 1.5", "MOVSD", []Operand{FloatImm{Text: "1.5"}, vreg(t, "X3")}, "f20f101d00000000"},
{"MOVSS 2.5", "MOVSS", []Operand{FloatImm{Text: "2.5"}, vreg(t, "X4")}, "f30f102500000000"},
{"MOVSS -0.5", "MOVSS", []Operand{FloatImm{Text: "0.5", Neg: true}, vreg(t, "X5")}, "f30f102d00000000"},
{"MOVSS +0.0 is XORPS", "MOVSS", []Operand{FloatImm{Text: "0.0"}, vreg(t, "X10")}, "450f57d2"},
{"MOVSD +0.0 is XORPS", "MOVSD", []Operand{FloatImm{Text: "0.0"}, vreg(t, "X6")}, "0f57f6"},
{"ADDSD 1.0", "ADDSD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}, "f20f580500000000"},
{"ADDSS 0.5", "ADDSS", []Operand{FloatImm{Text: "0.5"}, vreg(t, "X1")}, "f30f580d00000000"},
{"SUBSD 2.0", "SUBSD", []Operand{FloatImm{Text: "2.0"}, vreg(t, "X3")}, "f20f5c1d00000000"},
{"MULSD -2.5", "MULSD", []Operand{FloatImm{Text: "2.5", Neg: true}, vreg(t, "X3")}, "f20f591d00000000"},
{"DIVSD 1.0", "DIVSD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}, "f20f5e0500000000"},
{"COMISD 1.0", "COMISD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}, "660f2f0500000000"},
}
for _, c := range cases {
code, err := Encode(c.mnem, c.ops...)
if err != nil {
t.Errorf("%s: Encode: %v", c.name, err)
continue
}
if got := hexCompact(code); got != c.want {
t.Errorf("%s: got %s, want %s", c.name, got, c.want)
}
}
// The pool names carry the IEEE-754 bits, the float32 narrowing for the
// single spellings; negative zero keeps its sign bit and never takes the
// XORPS shortcut.
for _, c := range []struct {
mnem string
imm FloatImm
want string
}{
{"MOVSD", FloatImm{Text: "1.0", Neg: true}, "$f64.bff0000000000000"},
{"MOVSD", FloatImm{Text: "0.5"}, "$f64.3fe0000000000000"},
{"MOVSS", FloatImm{Text: "2.5"}, "$f32.40200000"},
{"MOVSS", FloatImm{Text: "0.5", Neg: true}, "$f32.bf000000"},
{"MOVSD", FloatImm{Text: "0.0", Neg: true}, "$f64.8000000000000000"},
} {
_, name, err := floatPoolValue(c.mnem, c.imm)
if err != nil {
t.Errorf("%s %s: %v", c.mnem, c.imm.Text, err)
continue
}
if name != c.want {
t.Errorf("%s $%s: pool name %s, want %s", c.mnem, c.imm.Text, name, c.want)
}
}
// The shapes the toolchain's parser rejects: the packed and uniform
// forms, a non-vector destination, and the integer spellings.
for _, c := range []struct {
name string
mnem string
ops []Operand
}{
{"MAXSD rejects the immediate", "MAXSD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}},
{"MINSD rejects the immediate", "MINSD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}},
{"SQRTSD rejects the immediate", "SQRTSD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}},
{"integer destination", "MOVSD", []Operand{FloatImm{Text: "1.0"}, AX}},
} {
if _, err := Encode(c.mnem, c.ops...); err == nil {
t.Errorf("%s: expected an error, got none", c.name)
}
}
}
// TestBookkeepingGroundTruth pins FUNCDATA and PCDATA as accept-and-ignore:
// go tool asm emits no text bytes for either, on every architecture.
func TestBookkeepingGroundTruth(t *testing.T) {
for _, c := range []struct {
name string
mnem string
ops []Operand
}{
{"FUNCDATA", "FUNCDATA", []Operand{Imm(3), sbMem{name: "\u00b7f.arginfo0"}}},
{"PCDATA", "PCDATA", []Operand{Imm(1), Imm(-1)}},
} {
code, err := Encode(c.mnem, c.ops...)
if err != nil {
t.Errorf("%s: Encode: %v", c.name, err)
continue
}
if len(code) != 0 {
t.Errorf("%s: emitted %x, want no bytes", c.name, code)
}
}
for _, c := range []struct {
name string
mnem string
ops []Operand
}{
{"FUNCDATA arity", "FUNCDATA", []Operand{Imm(3)}},
{"FUNCDATA missing the count", "FUNCDATA", []Operand{sbMem{name: "x"}}},
{"FUNCDATA integer value", "FUNCDATA", []Operand{Imm(3), Imm(4)}},
{"PCDATA arity", "PCDATA", []Operand{Imm(1)}},
{"PCDATA register value", "PCDATA", []Operand{Imm(1), AX}},
} {
if _, err := Encode(c.mnem, c.ops...); err == nil {
t.Errorf("%s: expected an error, got none", c.name)
}
}
// At the statement level the bookkeeping lines sit between real
// instructions and contribute nothing to the body, symbol reference
// included: the FUNCDATA operand never needs file-level resolution.
f, errs := parser.Parse("t_amd64.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tNOP\n\tFUNCDATA $3, \u00b7f.arginfo0(SB)\n\tPCDATA $1, $-1\n\tFUNCDATA $0, x<>(SB)\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFile(f)
if err != nil {
t.Fatalf("assemble: %v", err)
}
want := "90c3"
if got := hexCompact(img.Code); got != want {
t.Errorf("body %s, want %s (the bookkeeping lines contribute nothing)", got, want)
}
if _, err := AssembleFile(mustParse(t, "TEXT \u00b7f(SB), NOSPLIT, $0\n\tFUNCDATA $1, X0\n\tRET\n")); err == nil {
t.Error("FUNCDATA $1, X0 assembled, want an error")
}
if _, err := AssembleFile(mustParse(t, "TEXT \u00b7f(SB), NOSPLIT, $0\n\tPCDATA $1, X0\n\tRET\n")); err == nil {
t.Error("PCDATA $1, X0 assembled, want an error")
}
// Encodable mirrors Encode for the names this work touched.
for _, mnem := range []string{"FUNCDATA", "PCDATA", "V4FMADDPS", "V4FMADDSS", "V4FNMADDPS", "V4FNMADDSS", "VP4DPWSSD", "VP4DPWSSDS"} {
if !Encodable(mnem) {
t.Errorf("Encodable(%s) = false, want true", mnem)
}
}
}
// mustParse parses src or fails the test.
func mustParse(t *testing.T, src string) *ast.File {
t.Helper()
f, errs := parser.Parse("t_amd64.s", src)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
return f
}
+593 -24
View File
@@ -5,6 +5,7 @@ package asm
import (
"fmt"
"slices"
"strings"
)
@@ -91,7 +92,7 @@ var evexTable = map[string]evexSpec{
"VPSRAD": {1, 0x72, 0, 1, 4, vexShiftImm, [3]int{16, 32, 64}},
// EVEX.128/256/512.66.0F.W1, variable shift with an XMM count (VPSRAQ;
// the W bit distinguishes it from VPSRAD's E2 form).
"VPSRAQ": {1, 0xE2, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSRAQ": {1, 0x72, 1, 1, 4, vexShiftImm, [3]int{16, 32, 64}},
// EVEX.128/256/512.F3.0F.W1, signed qword to packed double (reg=dst,
// rm=src, no vvvv).
@@ -138,7 +139,7 @@ var evexTable = map[string]evexSpec{
// EVEX.66.0F, the EVEX forms of the VEX two-source shuffle.
"VSHUFPD": {1, 0xC6, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VSHUFPS": {1, 0xC6, 0, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VSHUFPS": {1, 0xC6, 0, 0, -1, vexNDS3Imm, [3]int{16, 32, 64}},
// EVEX.66.0F3A, lane insert ($imm, xsrc, zsrc1, zdst).
"VINSERTF32X4": {3, 0x18, 0, 1, -1, vexNDS3Imm, [3]int{0, 16, 32}},
@@ -202,7 +203,7 @@ var evexTable = map[string]evexSpec{
"VPMULHUW": {1, 0xE4, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPMADDUBSW": {2, 0x04, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSLLVW": {2, 0x12, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSRLVW": {2, 0x11, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSRLVW": {2, 0x10, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPACKSSWB": {1, 0x63, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPACKUSWB": {1, 0x67, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPACKSSDW": {1, 0x6B, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
@@ -324,7 +325,7 @@ var evexTable = map[string]evexSpec{
"VCVTPD2UQQ": {1, 0x79, 1, 1, -1, vexRM, [3]int{16, 32, 64}},
"VCVTPS2QQ": {1, 0x7B, 0, 1, -1, vexRM, [3]int{8, 16, 32}},
"VCVTUDQ2PD": {1, 0x7A, 0, 2, -1, vexRM, [3]int{8, 16, 32}},
"VCVTUDQ2PS": {1, 0x7A, 0, 0, -1, vexRM, [3]int{8, 16, 32}},
"VCVTUDQ2PS": {1, 0x7A, 0, 3, -1, vexRM, [3]int{8, 16, 32}},
// EVEX.66.0F38, half-precision convert (half-width source).
"VCVTPH2PS": {2, 0x13, 0, 1, -1, vexRM, [3]int{8, 16, 32}},
// EVEX.66.0F3A, half-precision convert back ($imm, src, dst: reg=src,
@@ -494,6 +495,331 @@ var evexTable = map[string]evexSpec{
// destination (VPMOVDW dword→word, VPMOVQD qword→dword).
"VPMOVDW": {2, 0x33, 0, 2, -1, vexRMRev, [3]int{8, 16, 32}},
"VPMOVQD": {2, 0x35, 0, 2, -1, vexRMRev, [3]int{8, 16, 32}},
// --- the AVX-512 families the avx512enc corpus exercises, read off
// the toolchain opcodetables ---
"VAESDEC": {2, 0xDE, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VAESDECLAST": {2, 0xDF, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VAESENC": {2, 0xDC, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VAESENCLAST": {2, 0xDD, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VALIGNQ": {3, 0x03, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VANDNPD": {1, 0x55, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VANDPD": {1, 0x54, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VBLENDMPD": {2, 0x65, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VBLENDMPS": {2, 0x65, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VBROADCASTF32X2": {2, 0x19, 0, 1, -1, vexRM, [3]int{0, 8, 8}},
"VBROADCASTF32X4": {2, 0x1A, 0, 1, -1, vexRM, [3]int{0, 16, 16}},
"VBROADCASTF32X8": {2, 0x1B, 0, 1, -1, vexRM, [3]int{0, 0, 32}},
"VBROADCASTF64X2": {2, 0x1A, 1, 1, -1, vexRM, [3]int{0, 16, 16}},
"VBROADCASTF64X4": {2, 0x1B, 1, 1, -1, vexRM, [3]int{0, 0, 32}},
"VBROADCASTI32X2": {2, 0x59, 0, 1, -1, vexRM, [3]int{8, 8, 8}},
"VBROADCASTI32X4": {2, 0x5A, 0, 1, -1, vexRM, [3]int{0, 16, 16}},
"VBROADCASTI32X8": {2, 0x5B, 0, 1, -1, vexRM, [3]int{0, 0, 32}},
"VBROADCASTI64X2": {2, 0x5A, 1, 1, -1, vexRM, [3]int{0, 16, 16}},
"VBROADCASTI64X4": {2, 0x5B, 1, 1, -1, vexRM, [3]int{0, 0, 32}},
"VCOMISD": {1, 0x2F, 1, 1, -1, vexRM, [3]int{8, 0, 0}},
"VCVTSD2SS": {1, 0x5A, 1, 3, -1, vexNDS3, [3]int{8, 0, 0}},
"VCVTSS2SD": {1, 0x5A, 0, 2, -1, vexNDS3, [3]int{4, 0, 0}},
"VDBPSADBW": {3, 0x42, 0, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VEXP2PD": {2, 0xC8, 1, 1, -1, vexRM, [3]int{0, 0, 64}},
"VEXP2PS": {2, 0xC8, 0, 1, -1, vexRM, [3]int{0, 0, 64}},
"VFMADD132PD": {2, 0x98, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMADD132PS": {2, 0x98, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMADD132SD": {2, 0x99, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFMADD132SS": {2, 0x99, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFMADD213PD": {2, 0xA8, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMADD213PS": {2, 0xA8, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMADD213SD": {2, 0xA9, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFMADD213SS": {2, 0xA9, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFMADD231PS": {2, 0xB8, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMADD231SD": {2, 0xB9, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFMADD231SS": {2, 0xB9, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFMADDSUB132PD": {2, 0x96, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMADDSUB132PS": {2, 0x96, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMADDSUB213PD": {2, 0xA6, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMADDSUB213PS": {2, 0xA6, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMADDSUB231PD": {2, 0xB6, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMADDSUB231PS": {2, 0xB6, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUB132PD": {2, 0x9A, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUB132PS": {2, 0x9A, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUB132SD": {2, 0x9B, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFMSUB132SS": {2, 0x9B, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFMSUB213PD": {2, 0xAA, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUB213PS": {2, 0xAA, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUB213SD": {2, 0xAB, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFMSUB213SS": {2, 0xAB, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFMSUB231PD": {2, 0xBA, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUB231PS": {2, 0xBA, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUB231SD": {2, 0xBB, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFMSUB231SS": {2, 0xBB, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFMSUBADD132PD": {2, 0x97, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUBADD132PS": {2, 0x97, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUBADD213PD": {2, 0xA7, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUBADD213PS": {2, 0xA7, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUBADD231PD": {2, 0xB7, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFMSUBADD231PS": {2, 0xB7, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMADD132PD": {2, 0x9C, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMADD132PS": {2, 0x9C, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMADD132SD": {2, 0x9D, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFNMADD132SS": {2, 0x9D, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFNMADD213PD": {2, 0xAC, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMADD213PS": {2, 0xAC, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMADD213SD": {2, 0xAD, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFNMADD213SS": {2, 0xAD, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFNMADD231PD": {2, 0xBC, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMADD231PS": {2, 0xBC, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMADD231SD": {2, 0xBD, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFNMADD231SS": {2, 0xBD, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFNMSUB132PD": {2, 0x9E, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMSUB132PS": {2, 0x9E, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMSUB132SD": {2, 0x9F, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFNMSUB132SS": {2, 0x9F, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFNMSUB213PD": {2, 0xAE, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMSUB213PS": {2, 0xAE, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMSUB213SD": {2, 0xAF, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFNMSUB213SS": {2, 0xAF, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VFNMSUB231PD": {2, 0xBE, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMSUB231PS": {2, 0xBE, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VFNMSUB231SD": {2, 0xBF, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VFNMSUB231SS": {2, 0xBF, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VGF2P8AFFINEINVQB": {3, 0xCF, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VGF2P8AFFINEQB": {3, 0xCE, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VGF2P8MULB": {2, 0xCF, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VMOVNTDQ": {1, 0xE7, 0, 1, -1, vexRMRev, [3]int{16, 32, 64}},
"VMOVNTDQA": {2, 0x2A, 0, 1, -1, vexRM, [3]int{16, 32, 64}},
"VMOVNTPD": {1, 0x2B, 1, 1, -1, vexRMRev, [3]int{16, 32, 64}},
"VORPD": {1, 0x56, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPADDSB": {1, 0xEC, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPADDSW": {1, 0xED, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPADDUSB": {1, 0xDC, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPADDUSW": {1, 0xDD, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPBLENDMB": {2, 0x66, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPBLENDMD": {2, 0x64, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPBLENDMQ": {2, 0x64, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPBLENDMW": {2, 0x66, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPBROADCASTMB2Q": {2, 0x2A, 1, 2, -1, vexRM, [3]int{0, 0, 0}},
"VPBROADCASTMW2D": {2, 0x3A, 0, 2, -1, vexRM, [3]int{0, 0, 0}},
"VPCLMULQDQ": {3, 0x44, 0, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VPCMPEQB": {1, 0x74, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPCMPEQQ": {2, 0x29, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPCMPEQW": {1, 0x75, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPCMPGTB": {1, 0x64, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPCMPGTD": {1, 0x66, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPCMPGTQ": {2, 0x37, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPCMPGTW": {1, 0x65, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPCOMPRESSB": {2, 0x63, 0, 1, -1, vexRMRev, [3]int{1, 1, 1}},
"VPCOMPRESSW": {2, 0x63, 1, 1, -1, vexRMRev, [3]int{2, 2, 2}},
"VPCONFLICTD": {2, 0xC4, 0, 1, -1, vexRM, [3]int{16, 32, 64}},
"VPCONFLICTQ": {2, 0xC4, 1, 1, -1, vexRM, [3]int{16, 32, 64}},
"VPDPBUSD": {2, 0x50, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPDPBUSDS": {2, 0x51, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPDPWSSD": {2, 0x52, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPDPWSSDS": {2, 0x53, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMI2PD": {2, 0x77, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMI2PS": {2, 0x77, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMI2W": {2, 0x75, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMPS": {2, 0x16, 0, 1, -1, vexNDS3, [3]int{0, 32, 64}},
"VPERMT2B": {2, 0x7D, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMT2PS": {2, 0x7F, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMT2W": {2, 0x7D, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPEXPANDB": {2, 0x62, 0, 1, -1, vexRM, [3]int{1, 1, 1}},
"VPEXPANDW": {2, 0x62, 1, 1, -1, vexRM, [3]int{2, 2, 2}},
"VPINSRD": {3, 0x22, 0, 1, -1, vexNDS3Imm, [3]int{4, 0, 0}},
"VPINSRQ": {3, 0x22, 1, 1, -1, vexNDS3Imm, [3]int{8, 0, 0}},
"VPLZCNTD": {2, 0x44, 0, 1, -1, vexRM, [3]int{16, 32, 64}},
"VPLZCNTQ": {2, 0x44, 1, 1, -1, vexRM, [3]int{16, 32, 64}},
"VPMADD52HUQ": {2, 0xB5, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPMADD52LUQ": {2, 0xB4, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPMULDQ": {2, 0x28, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPMULHRSW": {2, 0x0B, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPMULHW": {1, 0xE5, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPMULTISHIFTQB": {2, 0x83, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPMULUDQ": {1, 0xF4, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPOPCNTW": {2, 0x54, 1, 1, -1, vexRM, [3]int{16, 32, 64}},
"VPORD": {1, 0xEB, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPROLVD": {2, 0x15, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPROLVQ": {2, 0x15, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPRORVD": {2, 0x14, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPRORVQ": {2, 0x14, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSADBW": {1, 0xF6, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSHLDD": {3, 0x71, 0, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VPSHLDQ": {3, 0x71, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VPSHLDVD": {2, 0x71, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSHLDVQ": {2, 0x71, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSHLDVW": {2, 0x70, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSHLDW": {3, 0x70, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VPSHRDD": {3, 0x73, 0, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VPSHRDQ": {3, 0x73, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VPSHRDVD": {2, 0x73, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSHRDVQ": {2, 0x73, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSHRDVW": {2, 0x72, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSHRDW": {3, 0x72, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
"VPSHUFBITQMB": {2, 0x8F, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSRAVW": {2, 0x11, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSRLD": {1, 0x72, 0, 1, 2, vexShiftImm, [3]int{16, 32, 64}},
"VPSRLDQ": {1, 0x73, 0, 1, 3, vexShiftImm, [3]int{16, 32, 64}},
// EVEX.66.0F73 /7, the byte-quad shift left (the count is always an
// immediate; there is no register-count twin).
"VPSLLDQ": {1, 0x73, 0, 1, 7, vexShiftImm, [3]int{16, 32, 64}},
// EVEX.128/256/512.0F.W0, the plain-prefix (no 66) packed spellings
// whose EVEX form drops the legacy prefix entirely.
"VANDNPS": {1, 0x55, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}},
"VANDPS": {1, 0x54, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}},
"VORPS": {1, 0x56, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}},
"VXORPS": {1, 0x57, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}},
"VUNPCKLPS": {1, 0x14, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}},
"VUNPCKHPS": {1, 0x15, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}},
"VSQRTPS": {1, 0x51, 0, 0, -1, vexRM, [3]int{16, 32, 64}},
"VCOMISS": {1, 0x2F, 0, 0, -1, vexRM, [3]int{4, 0, 0}},
"VUCOMISS": {1, 0x2E, 0, 0, -1, vexRM, [3]int{4, 0, 0}},
"VMOVNTPS": {1, 0x2B, 0, 0, -1, vexRMRev, [3]int{16, 32, 64}},
"VPSUBSB": {1, 0xE8, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSUBSW": {1, 0xE9, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSUBUSB": {1, 0xD8, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPSUBUSW": {1, 0xD9, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPTESTMB": {2, 0x26, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPTESTMD": {2, 0x27, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPTESTMQ": {2, 0x27, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPTESTMW": {2, 0x26, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPTESTNMB": {2, 0x26, 0, 2, -1, vexNDS3, [3]int{16, 32, 64}},
"VPTESTNMD": {2, 0x27, 0, 2, -1, vexNDS3, [3]int{16, 32, 64}},
"VPTESTNMQ": {2, 0x27, 1, 2, -1, vexNDS3, [3]int{16, 32, 64}},
"VPTESTNMW": {2, 0x26, 1, 2, -1, vexNDS3, [3]int{16, 32, 64}},
"VPUNPCKHBW": {1, 0x68, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPUNPCKHQDQ": {1, 0x6D, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPUNPCKHWD": {1, 0x69, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPUNPCKLBW": {1, 0x60, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPUNPCKLQDQ": {1, 0x6C, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPUNPCKLWD": {1, 0x61, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VRCP28PD": {2, 0xCA, 1, 1, -1, vexRM, [3]int{0, 0, 64}},
"VRCP28PS": {2, 0xCA, 0, 1, -1, vexRM, [3]int{0, 0, 64}},
"VRCP28SD": {2, 0xCB, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VRCP28SS": {2, 0xCB, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VRSQRT28PD": {2, 0xCC, 1, 1, -1, vexRM, [3]int{0, 0, 64}},
"VRSQRT28PS": {2, 0xCC, 0, 1, -1, vexRM, [3]int{0, 0, 64}},
"VRSQRT28SD": {2, 0xCD, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
"VRSQRT28SS": {2, 0xCD, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
"VSQRTPD": {1, 0x51, 1, 1, -1, vexRM, [3]int{16, 32, 64}},
"VSQRTSD": {1, 0x51, 1, 3, -1, vexNDS3, [3]int{8, 0, 0}},
"VSQRTSS": {1, 0x51, 0, 2, -1, vexNDS3, [3]int{4, 0, 0}},
"VUCOMISD": {1, 0x2E, 1, 1, -1, vexRM, [3]int{8, 0, 0}},
"VXORPD": {1, 0x57, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
// EVEX.128/256/512.0F.F3/F2.W0, word shuffles with an immediate
// ($imm, src, dst: reg = dst, rm = src, imm8). The F3/F2 prefixes
// split the high/low lane spellings.
"VPSHUFHW": {1, 0x70, 0, 2, -1, vexImmRM, [3]int{16, 32, 64}},
"VPSHUFLW": {1, 0x70, 0, 3, -1, vexImmRM, [3]int{16, 32, 64}},
// EVEX.128.66.0F3A, lane extract to a general-purpose register or
// memory ($imm, xsrc, GPR/mem dst: reg = source, rm = destination).
"VPEXTRB": {3, 0x14, 0, 1, -1, vexExtractGPR, [3]int{1, 1, 1}},
"VPEXTRW": {3, 0x15, 0, 1, -1, vexExtractGPR, [3]int{2, 2, 2}},
"VPEXTRD": {3, 0x16, 0, 1, -1, vexExtractGPR, [3]int{4, 4, 4}},
"VPEXTRQ": {3, 0x16, 1, 1, -1, vexExtractGPR, [3]int{8, 8, 8}},
// EVEX.66.0F3A.W1, the qword permutes with an immediate control
// ($imm, src, dst: reg = dst, rm = src, imm8); the register-count
// forms live in evexRegFormTable.
"VPERMQ": {3, 0x00, 1, 1, -1, vexImmRM, [3]int{16, 32, 64}},
"VPERMPD": {3, 0x01, 1, 1, -1, vexImmRM, [3]int{16, 32, 64}},
// EVEX.66.0F3A, the packed permute shuffles with an immediate control.
"VPERMILPS": {3, 0x04, 0, 1, -1, vexImmRM, [3]int{16, 32, 64}},
"VPERMILPD": {3, 0x05, 1, 1, -1, vexImmRM, [3]int{16, 32, 64}},
// EVEX.128.0F.W0, high/low half moves. VMOVHPS carries the
// three-operand insert form (rm = m64 source, vvvv = preserved,
// reg = dst) and the two-operand store (reg = source, rm = m64);
// the encoder splits on the operand count. VMOVLHPS is the
// three-operand form alone.
"VMOVHPS": {1, 0x16, 0, 0, -1, vexNDS3, [3]int{8, 0, 0}},
"VMOVLHPS": {1, 0x16, 0, 0, -1, vexNDS3, [3]int{8, 0, 0}},
}
// evexQuad describes one quad-register instruction: the opcode under
// EVEX.0F38.W0 with the F2 mandatory prefix, and the width of the vector
// registers the bracketed list and the destination take (512-bit ZMM for
// the packed forms, 128-bit XMM for the scalar ones).
type evexQuad struct {
opcode byte
width int // register width in bytes: 64 (ZMM) or 16 (XMM)
}
// evexQuadTable maps the quad-register instructions (the 4FMAPS and 4VNNIW
// families) to their encoding. The operand shape is fixed: a single memory
// source in r/m, the bracketed register list whose LOW register travels the
// inverted 5-bit V'VVVV field, an optional opmask in aaa and the vector
// destination in reg. The vector length follows the destination (512-bit
// for the ZMM list forms, 128-bit for the scalar ones) while the disp8×N
// multiplier stays 16 for every member, the toolchain's own tuple choice.
var evexQuadTable = map[string]evexQuad{
"V4FMADDPS": {0x9A, 64},
"V4FMADDSS": {0x9B, 16},
"V4FNMADDPS": {0xAA, 64},
"V4FNMADDSS": {0xAB, 16},
"VP4DPWSSD": {0x52, 64},
"VP4DPWSSDS": {0x53, 64},
}
// isEvexQuad reports whether the mnemonic is a quad-register instruction.
func isEvexQuad(upper string) bool {
_, ok := evexQuadTable[upper]
return ok
}
// encodeEvexQuad encodes the quad-register form: OP mem, [Zn-Zn+3], (K), dst.
// The register list is the VVVV-side source: its low register fills the
// inverted V'VVVV bits, which is why an indexed memory source above Z15 (no
// spare EVEX.X bit once V' is taken) is refused. Masking rides the standard
// aaa field, zeroing keeps the usual requires-a-mask rule, and no other
// suffix applies.
func (e *enc) encodeEvexQuad(mnem string, q evexQuad, ops []Operand, sfx evexSuffix) error {
if len(ops) != 3 && len(ops) != 4 {
return fmt.Errorf("%s expects 3 or 4 operands (mem, [Zn-Zn+3], (K), dst), got %d", mnem, len(ops))
}
mem, lst := ops[0], ops[1]
dst := ops[len(ops)-1]
mask := 0
if len(ops) == 4 {
k, ok := ops[2].(Reg)
if !ok || !k.mask {
return fmt.Errorf("%s: third operand must be an opmask register", mnem)
}
if k.idx == 0 {
return fmt.Errorf("k0 is not a usable mask register")
}
mask = k.idx
}
list, ok := lst.(RegList)
if !ok {
return fmt.Errorf("%s: second operand must be a four-register list", mnem)
}
if list.Lo.size != q.width {
return fmt.Errorf("%s: the register list must hold %d-bit vector registers", mnem, q.width*8)
}
dstReg, ok := dst.(Reg)
if !ok || !dstReg.isVec() {
return fmt.Errorf("%s: destination must be a vector register", mnem)
}
if dstReg.size != q.width {
return fmt.Errorf("%s: the destination must be a %d-bit vector register", mnem, q.width*8)
}
if !memOperand(mem) {
return fmt.Errorf("%s: the source must be a memory operand", mnem)
}
// The list owns V'VVVV; a scaled index in the EVEX-only half would fold
// its fifth bit into the same field the list's low register occupies.
if m, ok := mem.(Mem); ok && m.HasIndex && m.Index.idx >= 16 {
return fmt.Errorf("%s: an index register above Z15 has no EVEX bit free", mnem)
}
if sfx.zeroing && mask == 0 {
return fmt.Errorf("%s: zeroing (.Z) requires a mask register", mnem)
}
spec := evexSpec{mapSel: 2, opcode: q.opcode, w: 0, pp: 3, opdigit: -1, n: [3]int{16, 16, 16}}
// The vector length follows the destination (512-bit for the ZMM forms,
// 128-bit for the scalar ones), exactly as the oracle encodes it.
return e.emitEvexFields(spec, dstReg.vecLenBit(), dstReg.idx, list.Lo.idx, mem, mask, sfx)
}
// evexBcastSpec describes an EVEX broadcast (VPBROADCASTD/Q): the opcode
@@ -529,32 +855,38 @@ type evexMoveSpec struct {
n [3]int
vecOK bool // the non-memory operand may be a vector register
xmmOnly bool // wider than XMM registers are rejected
nds3 bool // a three-operand register form exists (VMOVSD/VMOVSS)
}
// evexMoveTable maps an upper-case EVEX move mnemonic to its encoding.
var evexMoveTable = map[string]evexMoveSpec{
// EVEX.128/256/512.F3.0F.W0, unaligned integer move.
"VMOVDQU32": {1, 2, 0x6F, 0x7F, 0, [3]int{16, 32, 64}, true, false},
"VMOVDQU32": {1, 2, 0x6F, 0x7F, 0, [3]int{16, 32, 64}, true, false, false},
// EVEX.128/256/512.F3.0F.W1, unaligned qword move.
"VMOVDQU64": {1, 2, 0x6F, 0x7F, 1, [3]int{16, 32, 64}, true, false},
"VMOVDQU64": {1, 2, 0x6F, 0x7F, 1, [3]int{16, 32, 64}, true, false, false},
// EVEX.128/256/512.F2.0F.W0, unaligned byte move (byte/word moves use the
// F2 prefix, dword/qword moves F3; the element size only changes the tuple
// semantics).
"VMOVDQU8": {1, 3, 0x6F, 0x7F, 0, [3]int{16, 32, 64}, true, false},
"VMOVDQU8": {1, 3, 0x6F, 0x7F, 0, [3]int{16, 32, 64}, true, false, false},
// EVEX.128/256/512.F2.0F.W1, unaligned word move (shares the qword
// encoding).
"VMOVDQU16": {1, 3, 0x6F, 0x7F, 1, [3]int{16, 32, 64}, true, false},
"VMOVDQU16": {1, 3, 0x6F, 0x7F, 1, [3]int{16, 32, 64}, true, false, false},
// EVEX.128/256/512.66.0F.W1, unaligned packed double move.
"VMOVUPD": {1, 1, 0x10, 0x11, 1, [3]int{16, 32, 64}, true, false},
"VMOVUPD": {1, 1, 0x10, 0x11, 1, [3]int{16, 32, 64}, true, false, false},
// EVEX.128/256/512, aligned packed moves.
"VMOVAPS": {1, 0, 0x28, 0x29, 0, [3]int{16, 32, 64}, true, false},
"VMOVAPD": {1, 1, 0x28, 0x29, 1, [3]int{16, 32, 64}, true, false},
"VMOVAPS": {1, 0, 0x28, 0x29, 0, [3]int{16, 32, 64}, true, false, false},
"VMOVAPD": {1, 1, 0x28, 0x29, 1, [3]int{16, 32, 64}, true, false, false},
// EVEX.128/256/512.66.0F, aligned integer moves.
"VMOVDQA32": {1, 1, 0x6F, 0x7F, 0, [3]int{16, 32, 64}, true, false},
"VMOVDQA64": {1, 1, 0x6F, 0x7F, 1, [3]int{16, 32, 64}, true, false},
"VMOVDQA32": {1, 1, 0x6F, 0x7F, 0, [3]int{16, 32, 64}, true, false, false},
"VMOVDQA64": {1, 1, 0x6F, 0x7F, 1, [3]int{16, 32, 64}, true, false, false},
// EVEX.128.F3.0F.W0, scalar single move, memory operands (the
// three-operand register form is not supported).
"VMOVSS": {1, 2, 0x10, 0x11, 0, [3]int{4, 4, 4}, false, true},
"VMOVSS": {1, 2, 0x10, 0x11, 0, [3]int{4, 4, 4}, false, true, true},
// EVEX.128.F2.0F.W1, scalar double move: memory operands and the
// three-operand register form (VMOVSD dst, src1, src2).
"VMOVSD": {1, 3, 0x10, 0x11, 1, [3]int{8, 8, 8}, false, true, true},
// EVEX.128/256/512.0F.W0, unaligned packed single move.
"VMOVUPS": {1, 0, 0x10, 0x11, 0, [3]int{16, 32, 64}, true, false, false},
}
// isEvex reports whether the mnemonic has an EVEX encoding we handle.
@@ -565,8 +897,10 @@ func isEvex(mnemUpper string) bool {
if _, ok := evexBcastTable[mnemUpper]; ok {
return true
}
_, ok := evexMoveTable[mnemUpper]
return ok
if _, ok := evexMoveTable[mnemUpper]; ok {
return true
}
return isEvexQuad(mnemUpper)
}
// evexRequired reports whether the operands force the EVEX encoding of a
@@ -579,6 +913,13 @@ func evexRequired(upper string, ops []Operand) bool {
if !inVex && !inVexMove {
return true // EVEX-only mnemonic
}
// The byte-quad shifts have VEX register forms but EVEX-only memory
// forms: a memory count source forces the EVEX encoding.
if upper == "VPSLLDQ" || upper == "VPSRLDQ" {
if slices.ContainsFunc(ops, memOperand) {
return true
}
}
for _, op := range ops {
if r, ok := op.(Reg); ok && (r.size == 64 || r.mask || (r.isVec() && r.idx >= 16)) {
return true
@@ -741,6 +1082,14 @@ func (e *enc) encodeEvex(mnemUpper string, ops []Operand, sfx evexSuffix) error
return e.encodeEvexRM(spec, ops, 0, sfx)
}
spec, inTable := evexTable[mnemUpper]
if q, ok := evexQuadTable[mnemUpper]; ok {
// The quad-register family carries no rounding, SAE or broadcast;
// only masking and zeroing apply.
if sfx.sae || sfx.bcst || sfx.rounding >= 0 {
return fmt.Errorf("%s takes no rounding/SAE/broadcast suffix", mnemUpper)
}
return e.encodeEvexQuad(mnemUpper, q, ops, sfx)
}
if inTable {
if (sfx.rounding >= 0 || sfx.sae) && !evexRound[mnemUpper] {
return fmt.Errorf("%s: rounding/SAE is not supported for this instruction", mnemUpper)
@@ -752,6 +1101,27 @@ func (e *enc) encodeEvex(mnemUpper string, ops []Operand, sfx evexSuffix) error
}
spec.n = [3]int{n, n, n}
}
// A mnemonic with an immediate and a register spelling (the
// variable-count shifts, the permutes) encodes the register one
// when the first operand is not an immediate.
if len(ops) > 0 {
if _, isImm := ops[0].(Imm); !isImm {
if alt, ok := evexRegFormTable[mnemUpper]; ok {
spec, inTable = alt, true
}
}
}
// The high/low half moves split by operand count: three operands
// insert, two store (VMOVHPS m64, X1).
if hs, ok := evexHptrTable[mnemUpper]; ok {
if len(ops) == 2 {
if hs.store.opcode == 0 {
return fmt.Errorf("%s has no two-operand form", mnemUpper)
}
return e.encodeEvexRMRev(hs.store, ops, 0, sfx)
}
spec = hs.insert
}
} else if sfx.evexOnly() {
return fmt.Errorf("%s: the instruction does not take rounding/SAE/broadcast suffixes", mnemUpper)
}
@@ -811,6 +1181,12 @@ func (e *enc) encodeEvex(mnemUpper string, ops []Operand, sfx evexSuffix) error
}
return e.encodeEvexMove(mnemUpper, ms, ops, mask, sfx)
}
if ps, ok := evexPrefGatherTable[mnemUpper]; ok {
if sfx.any() {
return fmt.Errorf("%s takes no EVEX suffixes", mnemUpper)
}
return e.encodeEvexPrefGather(mnemUpper, ps, ops, mask, sfx)
}
if !inTable {
return fmt.Errorf("unsupported instruction %q for ZMM/K operands", mnemUpper)
}
@@ -829,6 +1205,8 @@ func (e *enc) encodeEvex(mnemUpper string, ops []Operand, sfx evexSuffix) error
return e.encodeEvexNDS3Imm(spec, ops, mask, sfx)
case vexExtract:
return e.encodeEvexExtract(spec, ops, mask, sfx)
case vexExtractGPR:
return e.encodeEvexExtractGPR(spec, ops, mask, sfx)
case vexRMSrcLen:
return e.encodeEvexRMSrcLen(spec, ops, mask, sfx)
}
@@ -908,6 +1286,11 @@ func (e *enc) encodeEvexImmRM(spec evexSpec, ops []Operand, mask int, sfx evexSu
if dstReg.mask {
if r, ok := src.(Reg); ok && r.isVec() {
ll = r.vecLenBit()
} else if l, err := soleLen(spec.n); err == nil {
// A memory source with a length-fixed mnemonic
// (VFPCLASSPDX/Y/Z): the length comes from the table's
// single valid slot, not from the operand.
ll = l
}
} else if r, ok := src.(Reg); ok && r.isVec() {
ll = r.vecLenBit()
@@ -934,9 +1317,11 @@ func (e *enc) encodeEvexShiftImm(spec evexSpec, ops []Operand, mask int, sfx eve
if !ok {
return fmt.Errorf("shift count must be an immediate")
}
srcReg, ok := src.(Reg)
if !ok || !srcReg.isVec() {
return fmt.Errorf("shift source must be a vector register")
// The count source is a vector register or memory; the length the L'L
// field and the disp8×N multiplier follow is the destination's either
// way.
if !vecOrMem(src) {
return fmt.Errorf("shift source must be a vector register or memory")
}
dstReg, ok := dst.(Reg)
if !ok || !dstReg.isVec() {
@@ -946,7 +1331,7 @@ func (e *enc) encodeEvexShiftImm(spec evexSpec, ops []Operand, mask int, sfx eve
if err != nil {
return err
}
if err := e.emitEvexFields(spec, dstReg.vecLenBit(), spec.opdigit, dstReg.idx, srcReg, mask, sfx); err != nil {
if err := e.emitEvexFields(spec, dstReg.vecLenBit(), spec.opdigit, dstReg.idx, src, mask, sfx); err != nil {
return err
}
e.out = append(e.out, immByte)
@@ -1018,10 +1403,77 @@ func (e *enc) encodeEvexExtract(spec evexSpec, ops []Operand, mask int, sfx evex
return nil
}
// encodeEvexExtractGPR encodes the lane extract to a general-purpose
// register or memory: OP $imm, xsrc, dst (reg = the XMM source, rm = the
// destination, imm8). The encoding is 128-bit regardless of register
// numbers, so L'L is fixed at 0 and the disp8×N multiplier is the extracted
// element size the table carries.
func (e *enc) encodeEvexExtractGPR(spec evexSpec, ops []Operand, mask int, sfx evexSuffix) error {
if len(ops) != 3 {
return fmt.Errorf("extract expects 3 operands ($imm, xsrc, dst), got %d", len(ops))
}
imm, src, dst := ops[0], ops[1], ops[2]
immVal, ok := imm.(Imm)
if !ok {
return fmt.Errorf("extract lane must be an immediate")
}
srcReg, ok := src.(Reg)
if !ok || !srcReg.isVec() {
return fmt.Errorf("extract source must be a vector register")
}
switch dst.(type) {
case Reg:
if dst.(Reg).isVec() {
return fmt.Errorf("extract destination must be a general-purpose register or memory")
}
case Mem, sbMem:
default:
return fmt.Errorf("extract destination must be a general-purpose register or memory")
}
immByte, err := imm8(int64(immVal))
if err != nil {
return err
}
if err := e.emitEvexFields(spec, 0, srcReg.idx, -1, dst, mask, sfx); err != nil {
return err
}
e.out = append(e.out, immByte)
return nil
}
// encodeEvexMove encodes a two-operand EVEX move; a vector→vector move uses
// the store-form opcode (reg = source, rm = destination), matching the Go
// assembler.
// assembler. The scalar moves also carry a three-operand register form
// (VMOVSD dst, src1, src2: the load opcode with vvvv = src1), which ms.nds3
// opens.
func (e *enc) encodeEvexMove(mnem string, ms evexMoveSpec, ops []Operand, mask int, sfx evexSuffix) error {
if len(ops) == 3 {
if !ms.nds3 {
return fmt.Errorf("EVEX move expects 2 operands, got %d", len(ops))
}
// The masked scalar register form keeps the Go assembler's own
// layout: the store opcode with reg = op0, vvvv = op1 and the
// destination in r/m (op2) — the bytes go tool asm emits, not
// the manual's NDS reading.
src, src1, dst := ops[0], ops[1], ops[2]
reg, ok := src.(Reg)
if !ok || !reg.isVec() {
return fmt.Errorf("%s: first operand must be a vector register", mnem)
}
vvvvReg, ok := src1.(Reg)
if !ok || !vvvvReg.isVec() {
return fmt.Errorf("%s: second operand must be a vector register", mnem)
}
dstReg, ok := dst.(Reg)
if !ok || !dstReg.isVec() {
return fmt.Errorf("%s: destination must be a vector register", mnem)
}
if ms.xmmOnly && (reg.size != 16 || vvvvReg.size != 16 || dstReg.size != 16) {
return fmt.Errorf("%s operates on XMM registers only", mnem)
}
spec := evexSpec{mapSel: ms.mapSel, opcode: ms.store, w: ms.w, pp: ms.pp, opdigit: -1, n: ms.n}
return e.emitEvexFields(spec, dstReg.vecLenBit(), reg.idx, vvvvReg.idx, dst, mask, sfx)
}
if len(ops) != 2 {
return fmt.Errorf("EVEX move expects 2 operands, got %d", len(ops))
}
@@ -1141,12 +1593,20 @@ func (e *enc) encodeEvexBcast(bs evexBcastSpec, ops []Operand, mask int, sfx eve
return fmt.Errorf("broadcast destination must be a vector register")
}
spec := evexSpec{mapSel: bs.mapSel, w: bs.w, pp: 1, opdigit: -1}
switch src.(type) {
switch r := src.(type) {
case Mem, sbMem:
spec.opcode = bs.opMem
spec.n = [3]int{bs.n, bs.n, bs.n}
case Reg:
spec.opcode = bs.opReg
// A GPR source uses the register broadcast opcode; a vector
// source shares the xmm/mem one (the low byte is copied from
// the lane or from the memory operand).
if r.isVec() {
spec.opcode = bs.opMem
spec.n = [3]int{bs.n, bs.n, bs.n}
} else {
spec.opcode = bs.opReg
}
default:
return fmt.Errorf("broadcast source must be a register or memory")
}
@@ -1349,6 +1809,102 @@ func isScatter(upper string) bool {
return ok
}
// isEvexPrefGather reports whether the mnemonic is a gather/scatter
// prefetch hint.
func isEvexPrefGather(upper string) bool {
_, ok := evexPrefGatherTable[upper]
return ok
}
// evexRegFormTable holds the register-count twin of the immediate-form
// entries in evexTable. Several mnemonics name two encodings: an immediate
// count or control ($imm, src, dst …) and a register-count one whose second
// operand is a vector register or memory (count, src2, src1, dst). The
// immediate spelling lives in evexTable, this table carries the register
// spelling, and encodeEvex picks by whether the first operand is an
// immediate, the way vexVarShift does on the VEX side.
var evexRegFormTable = map[string]evexSpec{
"VPSLLD": {1, 0xF2, 0, 1, -1, vexNDS3, [3]int{16, 16, 16}},
"VPSLLQ": {1, 0xF3, 1, 1, -1, vexNDS3, [3]int{16, 16, 16}},
"VPSLLW": {1, 0xF1, 0, 1, -1, vexNDS3, [3]int{16, 16, 16}},
"VPSRAD": {1, 0xE2, 0, 1, -1, vexNDS3, [3]int{16, 16, 16}},
"VPSRAQ": {1, 0xE2, 1, 1, -1, vexNDS3, [3]int{16, 16, 16}},
"VPSRAW": {1, 0xE1, 0, 1, -1, vexNDS3, [3]int{16, 16, 16}},
"VPSRLD": {1, 0xD2, 0, 1, -1, vexNDS3, [3]int{16, 16, 16}},
"VPSRLQ": {1, 0xD3, 1, 1, -1, vexNDS3, [3]int{16, 16, 16}},
"VPSRLW": {1, 0xD1, 0, 1, -1, vexNDS3, [3]int{16, 16, 16}},
// EVEX.NDS.0F38.W1, the register-count permutes (the immediate
// controls live in evexTable under 0F3A).
"VPERMQ": {2, 0x36, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMPD": {2, 0x16, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
// EVEX.NDS.0F38, the register-count permil shuffles.
"VPERMILPS": {2, 0x0C, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
"VPERMILPD": {2, 0x0D, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
}
// evexPrefGatherSpec describes a gather/scatter prefetch hint: one memory
// operand with a VSIB index and an opmask register, no destination. The
// ModRM.reg field carries a fixed /digit, the L'L field is fixed at 512, and
// the mask register is the instruction's only register operand.
type evexPrefGatherSpec struct {
mapSel int
opcode byte
w int
pp int
opdigit int
n int
}
var evexPrefGatherTable = map[string]evexPrefGatherSpec{
"VGATHERPF0DPD": {2, 0xC6, 1, 1, 1, 8},
"VGATHERPF0DPS": {2, 0xC6, 0, 1, 1, 4},
"VGATHERPF0QPD": {2, 0xC7, 1, 1, 1, 8},
"VGATHERPF0QPS": {2, 0xC7, 0, 1, 1, 4},
"VGATHERPF1DPD": {2, 0xC6, 1, 1, 2, 8},
"VGATHERPF1DPS": {2, 0xC6, 0, 1, 2, 4},
"VGATHERPF1QPD": {2, 0xC7, 1, 1, 2, 8},
"VGATHERPF1QPS": {2, 0xC7, 0, 1, 2, 4},
"VSCATTERPF0DPD": {2, 0xC6, 1, 1, 5, 8},
"VSCATTERPF0DPS": {2, 0xC6, 0, 1, 5, 4},
"VSCATTERPF0QPD": {2, 0xC7, 1, 1, 5, 8},
"VSCATTERPF0QPS": {2, 0xC7, 0, 1, 5, 4},
"VSCATTERPF1DPD": {2, 0xC6, 1, 1, 6, 8},
"VSCATTERPF1DPS": {2, 0xC6, 0, 1, 6, 4},
"VSCATTERPF1QPD": {2, 0xC7, 1, 1, 6, 8},
"VSCATTERPF1QPS": {2, 0xC7, 0, 1, 6, 4},
}
// evexHptrSpec describes the high/low half moves (VMOVHPS family): the
// three-operand insert shares an opcode with a two-operand store whose
// source is the vector register and whose destination is m64.
type evexHptrSpec struct {
insert evexSpec
store evexSpec // store.opcode == 0 when the mnemonic has no store form
}
var evexHptrTable = map[string]evexHptrSpec{
"VMOVHPS": {
insert: evexSpec{mapSel: 1, opcode: 0x16, w: 0, pp: 0, opdigit: -1, form: vexNDS3, n: [3]int{8, 0, 0}},
store: evexSpec{mapSel: 1, opcode: 0x17, w: 0, pp: 0, opdigit: -1, form: vexRMRev, n: [3]int{8, 0, 0}},
},
"VMOVLHPS": {
insert: evexSpec{mapSel: 1, opcode: 0x16, w: 0, pp: 0, opdigit: -1, form: vexNDS3, n: [3]int{8, 0, 0}},
},
}
// encodeEvexPrefGather encodes a gather/scatter prefetch hint: OP K, vsib.
func (e *enc) encodeEvexPrefGather(upper string, ps evexPrefGatherSpec, ops []Operand, mask int, sfx evexSuffix) error {
if len(ops) != 1 {
return fmt.Errorf("%s expects 2 operands (K, vsib memory), got %d", upper, len(ops)+1)
}
m, ok := ops[0].(Mem)
if !ok || !m.HasIndex || !m.Index.isVec() {
return fmt.Errorf("%s: operand must be a VSIB memory reference with a vector index", upper)
}
spec := evexSpec{mapSel: ps.mapSel, opcode: ps.opcode, w: ps.w, pp: ps.pp, opdigit: ps.opdigit, n: [3]int{ps.n, ps.n, ps.n}}
return e.emitEvexFields(spec, 2, ps.opdigit, -1, m, mask, sfx)
}
// vsibLen validates a VSIB memory operand (the index must be a vector
// register) and returns it with the vector length the index selects, the
// EVEX L'L field follows the index register, not the data register.
@@ -1370,7 +1926,9 @@ func (e *enc) encodeGather(upper string, gs gatherSpec, ops []Operand, sfx evexS
return err
}
if mask != 0 || sfx.any() {
// EVEX form: OP vsib, K, dst.
// EVEX form: OP vsib, K, dst. The L'L field is the wider of the
// index and the data register lengths (the Go assembler's
// layout); the disp8×N multiplier stays the index element size.
if len(rest) != 2 {
return fmt.Errorf("%s expects 3 operands (vsib, K, dst), got %d", upper, len(ops))
}
@@ -1382,6 +1940,9 @@ func (e *enc) encodeGather(upper string, gs gatherSpec, ops []Operand, sfx evexS
if !ok || !dst.isVec() {
return fmt.Errorf("%s: destination must be a vector register", upper)
}
if d := dst.vecLenBit(); d > ll {
ll = d
}
evex := evexSpec{mapSel: 2, opcode: gs.opcode, w: gs.w, pp: 1, opdigit: -1, n: [3]int{gs.n, gs.n, gs.n}}
return e.emitEvexFields(evex, ll, dst.idx, -1, vsib, mask, sfx)
}
@@ -1431,6 +1992,11 @@ func (e *enc) encodeScatter(upper string, ss gatherSpec, ops []Operand, sfx evex
if err != nil {
return err
}
// The L'L field is the wider of the data register and the VSIB index
// lengths, the bytes go tool asm emits.
if d := src.vecLenBit(); d > ll {
ll = d
}
evex := evexSpec{mapSel: 2, opcode: ss.opcode, w: ss.w, pp: 1, opdigit: -1, n: [3]int{ss.n, ss.n, ss.n}}
return e.emitEvexFields(evex, ll, src.idx, -1, vsib, mask, sfx)
}
@@ -1441,6 +2007,8 @@ func (e *enc) encodeScatter(upper string, ss gatherSpec, ops []Operand, sfx evex
var evexKOperand = map[string]bool{
"VPMOVM2B": true, "VPMOVM2W": true, "VPMOVM2D": true, "VPMOVM2Q": true,
"VPMOVB2M": true, "VPMOVW2M": true, "VPMOVD2M": true, "VPMOVQ2M": true,
// The K-to-vector broadcast reads its opmask source from r/m.
"VPBROADCASTMB2Q": true, "VPBROADCASTMW2D": true,
}
// kmovSpec describes a KMOV width: the opcode depends on the operand
@@ -1541,6 +2109,7 @@ var kOpsTable = map[string]kOpSpec{
"KXORD": {1, 0x47, 1, 1, 1, vexNDS3},
"KXORQ": {1, 0x47, 1, 0, 1, vexNDS3},
"KUNPCKBW": {1, 0x4B, 0, 1, 1, vexNDS3},
"KUNPCKWD": {1, 0x4B, 0, 0, 1, vexNDS3},
"KUNPCKDQ": {1, 0x4B, 1, 0, 1, vexNDS3},
"KADDB": {1, 0x4A, 0, 1, 1, vexNDS3},
"KADDW": {1, 0x4A, 0, 0, 1, vexNDS3},
+213
View File
@@ -721,3 +721,216 @@ func hexCompact(b []byte) string {
}
return string(out)
}
// TestAvx512CorpusFamilies pins representative encodings of the AVX-512
// families the toolchain's avx512enc corpus exercises: the bytes are the
// go tool asm output for exactly these operands, and the same families are
// covered end to end by the avx512_amd64.s differential kernel.
func TestAvx512CorpusFamilies(t *testing.T) {
vsib := func(base, idx string, scale int) Operand {
return Idx(vreg(t, base), vreg(t, idx), scale, 0, 0)
}
cases := []struct {
name string
mnem string
ops []Operand
want string
}{
// AES rounds (EVEX NDS, VEX twin routed by operand width).
{"VAESDEC Z", "VAESDEC", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "Z3")}, "62f26d48ded9"},
// Integer VNNI and the bit algorithm group.
{"VPDPBUSD", "VPDPBUSD", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "K2"), vreg(t, "Z3")}, "62f26d4a50d9"},
{"VPOPCNTW", "VPOPCNTW", []Operand{vreg(t, "Z1"), vreg(t, "K3"), vreg(t, "Z2")}, "62f2fd4b54d1"},
{"VPCONFLICTD", "VPCONFLICTD", []Operand{vreg(t, "Z1"), vreg(t, "K1"), vreg(t, "Z2")}, "62f27d49c4d1"},
{"VPLZCNTQ masked", "VPLZCNTQ", []Operand{vreg(t, "Z7"), vreg(t, "K1"), vreg(t, "Z8")}, "6272fd4944c7"},
{"VPERMT2B", "VPERMT2B", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "K1"), vreg(t, "Z3")}, "62f26d497dd9"},
{"VPMULTISHIFTQB", "VPMULTISHIFTQB", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "K3"), vreg(t, "Z4")}, "62f2ed4b83e1"},
{"VDBPSADBW", "VDBPSADBW", []Operand{Imm(3), vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "K3"), vreg(t, "Z3")}, "62f36d4b42d903"},
{"VPSHUFBITQMB", "VPSHUFBITQMB", []Operand{vreg(t, "Z9"), vreg(t, "Z10"), vreg(t, "K3")}, "62d22d488fd9"},
{"VPTESTNMQ", "VPTESTNMQ", []Operand{vreg(t, "Z13"), vreg(t, "Z14"), vreg(t, "K5")}, "62d28e4827ed"},
// Permutations: immediate and register counts.
{"VALIGNQ", "VALIGNQ", []Operand{Imm(3), vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "K1"), vreg(t, "Z3")}, "62f3ed4903d903"},
{"VPERMQ imm", "VPERMQ", []Operand{Imm(1), vreg(t, "Z1"), vreg(t, "K2"), vreg(t, "Z2")}, "62f3fd4a00d101"},
{"VPERMQ reg", "VPERMQ", []Operand{vreg(t, "Z3"), vreg(t, "Z4"), vreg(t, "K2"), vreg(t, "Z5")}, "62f2dd4a36eb"},
{"VPERMPD reg", "VPERMPD", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "Z3")}, "62f2ed4816d9"},
{"VPERMILPS imm", "VPERMILPS", []Operand{Imm(5), vreg(t, "Z9"), vreg(t, "K2"), vreg(t, "Z10")}, "62537d4a04d105"},
{"VPERMILPS reg", "VPERMILPS", []Operand{vreg(t, "Z11"), vreg(t, "Z12"), vreg(t, "K2"), vreg(t, "Z13")}, "62521d4a0ceb"},
// Shifts: immediate, register-count and memory-count forms; the
// count source carries its own XMM tuple width.
{"VPSLLW imm mask", "VPSLLW", []Operand{Imm(3), vreg(t, "Z1"), vreg(t, "K2"), vreg(t, "Z2")}, "62f16d4a71f103"},
{"VPSLLD reg count", "VPSLLD", []Operand{vreg(t, "X1"), vreg(t, "Z2"), vreg(t, "K1"), vreg(t, "Z3")}, "62f16d49f2d9"},
{"VPSLLDQ", "VPSLLDQ", []Operand{Imm(9), vreg(t, "Z7"), vreg(t, "Z8")}, "62f13d4873ff09"},
{"VPSRLDQ mem", "VPSRLDQ", []Operand{Imm(11), Ptr(SI, 16, 16), vreg(t, "Z4")}, "62f15d48739e100000000b"},
{"VPSRLVW", "VPSRLVW", []Operand{vreg(t, "Z3"), vreg(t, "Z4"), vreg(t, "K1"), vreg(t, "Z5")}, "62f2dd4910eb"},
// Conversions and shuffles with the F2 prefix and no prefix.
{"VCVTUDQ2PS", "VCVTUDQ2PS", []Operand{vreg(t, "Z1"), vreg(t, "K1"), vreg(t, "Z2")}, "62f17f497ad1"},
{"VSHUFPS", "VSHUFPS", []Operand{Imm(2), vreg(t, "Z4"), vreg(t, "Z5"), vreg(t, "K1"), vreg(t, "Z6")}, "62f15449c6f402"},
// Gather and scatter prefetch hints (memory-only, /digit in reg).
{"VGATHERPF0DPD", "VGATHERPF0DPD", []Operand{vreg(t, "K5"), vsib("R10", "Y29", 8)}, "6292fd45c60cea"},
{"VSCATTERPF1DPS", "VSCATTERPF1DPS", []Operand{vreg(t, "K2"), vsib("R10", "Z28", 4)}, "62927d42c634a2"},
// Opmask broadcasts and the K logic.
{"VPBROADCASTMB2Q", "VPBROADCASTMB2Q", []Operand{vreg(t, "K1"), vreg(t, "Z2")}, "62f2fe482ad1"},
{"VPBROADCASTMW2D", "VPBROADCASTMW2D", []Operand{vreg(t, "K3"), vreg(t, "Z4")}, "62f27e483ae3"},
{"KUNPCKWD", "KUNPCKWD", []Operand{vreg(t, "K6"), vreg(t, "K4"), vreg(t, "K1")}, "c5dc4bce"},
{"KADDB", "KADDB", []Operand{vreg(t, "K2"), vreg(t, "K3"), vreg(t, "K5")}, "c5e54aea"},
// Lane extracts to general registers (EVEX and VEX routes).
{"VPEXTRB", "VPEXTRB", []Operand{Imm(3), vreg(t, "X26"), AX}, "62637d0814d003"},
{"VPEXTRD", "VPEXTRD", []Operand{Imm(1), vreg(t, "X26"), vreg(t, "R9")}, "62437d0816d101"},
{"VPEXTRD vex", "VPEXTRD", []Operand{Imm(1), vreg(t, "X2"), DI}, "c4e37916d701"},
{"VPINSRQ", "VPINSRQ", []Operand{Imm(1), DI, vreg(t, "X3"), vreg(t, "X4")}, "c4e3e122e701"},
// Moves: masked unaligned, masked scalar register form, half moves
// and non-temporal stores.
{"VMOVUPS mask", "VMOVUPS", []Operand{vreg(t, "Z1"), vreg(t, "K2"), vreg(t, "Z3")}, "62f17c4a11cb"},
{"VMOVSD 3op", "VMOVSD", []Operand{vreg(t, "X14"), vreg(t, "X5"), vreg(t, "K3"), vreg(t, "X22")}, "6231d70b11f6"},
{"VMOVSS 3op", "VMOVSS", []Operand{vreg(t, "X18"), vreg(t, "X3"), vreg(t, "K2"), vreg(t, "X25")}, "6281660a11d1"},
{"VMOVHPS insert", "VMOVHPS", []Operand{Ptr(SI, 0, 8), vreg(t, "X18"), vreg(t, "X19")}, "62e16c00161e"},
{"VMOVHPS store", "VMOVHPS", []Operand{vreg(t, "X20"), Ptr(SI, 8, 8)}, "62e17c08176601"},
{"VMOVLHPS", "VMOVLHPS", []Operand{vreg(t, "X16"), vreg(t, "X5"), vreg(t, "X17")}, "62a1540816c8"},
{"VMOVNTDQ", "VMOVNTDQ", []Operand{vreg(t, "Z7"), Ptr(SI, 0, 64)}, "62f17d48e73e"},
{"VMOVNTDQA", "VMOVNTDQA", []Operand{Ptr(SI, 64, 64), vreg(t, "Z8")}, "62727d482a4601"},
{"VMOVNTPS", "VMOVNTPS", []Operand{vreg(t, "Z9"), Ptr(SI, 0, 64)}, "62717c482b0e"},
// Scalar compares with and without the 66 prefix.
{"VCOMISD", "VCOMISD", []Operand{vreg(t, "X5"), vreg(t, "X6")}, "c5f92ff5"},
{"VUCOMISS", "VUCOMISS", []Operand{vreg(t, "X7"), vreg(t, "X8")}, "c5782ec7"},
// Floating point helpers.
{"VSQRTSD", "VSQRTSD", []Operand{vreg(t, "X1"), vreg(t, "X2"), vreg(t, "K1"), vreg(t, "X3")}, "62f1ef0951d9"},
{"VEXP2PD", "VEXP2PD", []Operand{vreg(t, "Z5"), vreg(t, "K1"), vreg(t, "Z6")}, "62f2fd49c8f5"},
{"VRCP28SD", "VRCP28SD", []Operand{vreg(t, "X9"), vreg(t, "X8"), vreg(t, "K1"), vreg(t, "X10")}, "6252bd09cbd1"},
{"VBROADCASTF32X2", "VBROADCASTF32X2", []Operand{vreg(t, "X1"), vreg(t, "K1"), vreg(t, "Z2")}, "62f27d4919d1"},
{"VPCOMPRESSB", "VPCOMPRESSB", []Operand{vreg(t, "Z1"), vreg(t, "K1"), Ptr(SI, 0, 64)}, "62f27d49630e"},
}
for _, c := range cases {
code, err := Encode(c.mnem, c.ops...)
if err != nil {
t.Errorf("%s: Encode: %v", c.name, err)
continue
}
if got := hexCompact(code); got != c.want {
t.Errorf("%s: got %s, want %s", c.name, got, c.want)
}
}
}
// TestEvexQuadRegisterGroundTruth pins the quad-register instructions (the
// 4FMAPS and 4VNNIW families) byte for byte against go tool asm: the memory
// source keeps r/m, the bracketed list's LOW register travels the inverted
// 5-bit V'VVVV field, the destination sits in reg, the opmask rides aaa and
// the vector length follows the destination (L'L=512 for the ZMM forms,
// 128 for the scalar ones) while the disp8×N multiplier stays 16 for every
// member. The x86 decoder has no view of these forms, so no decode check
// runs.
func TestEvexQuadRegisterGroundTruth(t *testing.T) {
sp := vreg(t, "RSP")
cases := []struct {
name string
mnem string
ops []Operand
want string
}{
{"V4FMADDPS 17(SP) [Z0-Z3] K2 Z0", "V4FMADDPS",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "K2"), vreg(t, "Z0")},
"62f27f4a9a842411000000"},
{"V4FMADDPS [Z10-Z13]", "V4FMADDPS",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z10"), vreg(t, "Z13")}, vreg(t, "K2"), vreg(t, "Z0")},
"62f22f4a9a842411000000"},
{"V4FMADDPS [Z20-Z23]", "V4FMADDPS",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z20"), vreg(t, "Z23")}, vreg(t, "K2"), vreg(t, "Z0")},
"62f25f429a842411000000"},
{"V4FMADDPS Z8 dst", "V4FMADDPS",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "K2"), vreg(t, "Z8")},
"62727f4a9a842411000000"},
{"V4FMADDPS disp8x16", "V4FMADDPS",
[]Operand{Ptr(sp, 64, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "K2"), vreg(t, "Z0")},
"62f27f4a9a442404"},
{"V4FMADDPS unmasked", "V4FMADDPS",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "Z0")},
"62f27f489a842411000000"},
{"V4FMADDSS 7(AX) [X0-X3] K5 X22", "V4FMADDSS",
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X0"), vreg(t, "X3")}, vreg(t, "K5"), vreg(t, "X22")},
"62e27f0d9bb007000000"},
{"V4FMADDSS (DI)", "V4FMADDSS",
[]Operand{Ptr(DI, 0, 8), RegList{vreg(t, "X0"), vreg(t, "X3")}, vreg(t, "K5"), vreg(t, "X22")},
"62e27f0d9b37"},
{"V4FMADDSS [X10-X13]", "V4FMADDSS",
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X10"), vreg(t, "X13")}, vreg(t, "K5"), vreg(t, "X22")},
"62e22f0d9bb007000000"},
{"V4FMADDSS [X20-X23]", "V4FMADDSS",
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X20"), vreg(t, "X23")}, vreg(t, "K5"), vreg(t, "X22")},
"62e25f059bb007000000"},
{"V4FMADDSS X30 dst", "V4FMADDSS",
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X0"), vreg(t, "X3")}, vreg(t, "K5"), vreg(t, "X30")},
"62627f0d9bb007000000"},
{"V4FMADDSS X3 dst", "V4FMADDSS",
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X0"), vreg(t, "X3")}, vreg(t, "K5"), vreg(t, "X3")},
"62f27f0d9b9807000000"},
{"V4FMADDSS disp8x16", "V4FMADDSS",
[]Operand{Ptr(AX, 16, 8), RegList{vreg(t, "X20"), vreg(t, "X23")}, vreg(t, "K5"), vreg(t, "X30")},
"62625f059b7001"},
{"V4FNMADDPS", "V4FNMADDPS",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "K2"), vreg(t, "Z0")},
"62f27f4aaa842411000000"},
{"V4FNMADDSS", "V4FNMADDSS",
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X0"), vreg(t, "X3")}, vreg(t, "K5"), vreg(t, "X22")},
"62e27f0dabb007000000"},
{"VP4DPWSSD", "VP4DPWSSD",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "K2"), vreg(t, "Z0")},
"62f27f4a52842411000000"},
{"VP4DPWSSDS unmasked", "VP4DPWSSDS",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "Z0")},
"62f27f4853842411000000"},
}
for _, c := range cases {
code, err := Encode(c.mnem, c.ops...)
if err != nil {
t.Errorf("%s: Encode: %v", c.name, err)
continue
}
if got := hexCompact(code); got != c.want {
t.Errorf("%s: got %s, want %s", c.name, got, c.want)
}
}
}
// TestEvexQuadRegisterErrors pins the operand shapes the toolchain rejects:
// the register class the list and the destination take is fixed per
// instruction, the source is memory only, the opmask slot is positional and
// the list's low register owns V'VVVV.
func TestEvexQuadRegisterErrors(t *testing.T) {
sp := vreg(t, "RSP")
list := func(lo, hi string) RegList {
return RegList{vreg(t, lo), vreg(t, hi)}
}
cases := []struct {
name string
mnem string
ops []Operand
}{
{"X list on the PS form", "V4FMADDPS",
[]Operand{Ptr(sp, 0, 8), list("X0", "X3"), vreg(t, "K2"), vreg(t, "Z0")}},
{"Z list on the SS form", "V4FMADDSS",
[]Operand{Ptr(AX, 0, 8), list("Z0", "Z3"), vreg(t, "K5"), vreg(t, "X22")}},
{"Y destination", "V4FMADDPS",
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "K2"), vreg(t, "Y0")}},
{"register source", "V4FMADDPS",
[]Operand{vreg(t, "Z1"), list("Z0", "Z3"), vreg(t, "K2"), vreg(t, "Z0")}},
{"non-mask third operand", "V4FMADDPS",
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "Z4"), vreg(t, "Z0")}},
{"k0 mask", "V4FMADDPS",
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "K0"), vreg(t, "Z0")}},
{"K after the destination", "V4FMADDPS",
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "Z0"), vreg(t, "K2")}},
{"zeroing without a mask", "V4FMADDPS.Z",
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "Z0")}},
{"SAE suffix", "V4FMADDPS.SAE",
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "K2"), vreg(t, "Z0")}},
{"high index source", "VP4DPWSSD",
[]Operand{Idx(DI, vreg(t, "X16"), 1, 0, 8), list("Z0", "Z3"), vreg(t, "K2"), vreg(t, "Z0")}},
{"short operand list", "V4FMADDPS",
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3")}},
}
for _, c := range cases {
if _, err := Encode(c.mnem, c.ops...); err == nil {
t.Errorf("%s: expected an error, got none", c.name)
}
}
}
+55
View File
@@ -463,6 +463,61 @@ func (img *Image) emitGOObject(pkgPath, srcPath string, pre []byte, minLC int, r
symRelocs[si] = append(symRelocs[si], rec[:]...)
}
}
// The data symbols' own relocations: the symbol-valued DATA fields
// ("DATA s+0(SB)/8, $other(SB)"). The toolchain patches each field
// with the target's absolute address through an R_ADDR of the DATA
// line's width, on every architecture (the code relocations are
// per-architecture PC-relative shapes; a data pointer word is not), so
// this mapping bypasses relocField. The definitions were appended in
// DataSyms order, so data symbol i is definition index i.
for i, d := range img.DataSyms {
for _, r := range d.Relocs {
if r.Kind != RelAddr {
return nil, fmt.Errorf("GOOBJ emission: data symbol %q carries a non-data relocation", d.Name)
}
var rec [23]byte
binary.LittleEndian.PutUint32(rec[0:], uint32(int32(r.Off)))
rec[4] = r.Siz
binary.LittleEndian.PutUint16(rec[5:], relocAddr)
binary.LittleEndian.PutUint64(rec[7:], uint64(r.Addend))
switch {
case r.External && r.Name == goobjBuiltinMorestack:
binary.LittleEndian.PutUint32(rec[15:], pkgIdxBuiltin)
binary.LittleEndian.PutUint32(rec[19:], goobjBuiltinMorestackNoctxt)
case r.External:
pkg, name := splitQualified(r.Name)
if pkg == "" {
return nil, fmt.Errorf("GOOBJ emission: external symbol %q has no package prefix", r.Name)
}
pIdx, ok := extPkgIdx[pkg]
if !ok {
return nil, fmt.Errorf("GOOBJ emission: package %q not resolved", pkg)
}
sIdx, ok := extSymIdx[pkg+"·"+name]
if !ok {
return nil, fmt.Errorf("GOOBJ emission: symbol %s·%s not resolved", pkg, name)
}
binary.LittleEndian.PutUint32(rec[15:], uint32(pIdx))
binary.LittleEndian.PutUint32(rec[19:], uint32(sIdx))
default:
if di, ok := defIdx[r.Name]; ok {
binary.LittleEndian.PutUint32(rec[15:], pkgIdxSelf)
binary.LittleEndian.PutUint32(rec[19:], uint32(di))
break
}
// A DATA field may hold the address of a TEXT function of
// the same file (the rt0 lib entry spelling), which is a
// non-package definition.
ni, isText := textNpIdx[r.Name]
if !isText {
return nil, fmt.Errorf("GOOBJ emission: reference to unknown symbol %q", r.Name)
}
binary.LittleEndian.PutUint32(rec[15:], pkgIdxNone)
binary.LittleEndian.PutUint32(rec[19:], uint32(ni))
}
symRelocs[i] = append(symRelocs[i], rec[:]...)
}
}
// The DWARF symbols' own relocations (the function address references).
for _, ds := range dwarfRelocs {
for _, r := range ds.relocs {
+4 -4
View File
@@ -307,12 +307,12 @@ func TestStackGuardBytesLOONG64(t *testing.T) {
func TestStackGuardGOObjInternalCall(t *testing.T) {
for _, tt := range []struct {
src string
assemble func(*ast.File) (*Image, error)
assemble func(*ast.File, ...AssembleOption) (*Image, error)
}{
{"g_amd64.s", AssembleFile},
{"g_arm64.s", AssembleFileARM64},
{"g_riscv64.s", AssembleFileRISCV},
{"g_loong64.s", AssembleFileLOONG64},
{"g_arm64.s", func(f *ast.File, _ ...AssembleOption) (*Image, error) { return AssembleFileARM64(f) }},
{"g_riscv64.s", func(f *ast.File, _ ...AssembleOption) (*Image, error) { return AssembleFileRISCV(f) }},
{"g_loong64.s", func(f *ast.File, _ ...AssembleOption) (*Image, error) { return AssembleFileLOONG64(f) }},
} {
f, errs := parser.Parse(tt.src, "TEXT \u00b7callsmall(SB), $16-0\n\tCALL \u00b7other(SB)\n\tRET\nTEXT \u00b7other(SB), NOSPLIT, $0\n\tRET\n")
if len(errs) > 0 {
+74 -21
View File
@@ -192,6 +192,30 @@ func (e *enc) encodeMov(ops []Operand, size int) error {
}
return e.emit(i)
case TLSMem:
if !dstIsReg {
return fmt.Errorf("MOV: two memory operands")
}
// MOV r, off(TLS): the segment-prefixed absolute load, reg=dst,
// rm=src(tlsMem) through the SIB escape; the disp32 is the TLS slot
// offset with its R_TLSLE patch site.
i := newInstr(size, []byte{movRR(size)})
if err := setRM(i, dstReg, src, size); err != nil {
return err
}
return e.emit(i)
case SegAbs:
if !dstIsReg {
return fmt.Errorf("MOV: two memory operands")
}
// MOV r, 0x30(GS): the segment-absolute load.
i := newInstr(size, []byte{movRR(size)})
if err := setRM(i, dstReg, src, size); err != nil {
return err
}
return e.emit(i)
case Imm:
if dstIsReg {
v := int64(src)
@@ -232,11 +256,24 @@ func (e *enc) encodeMov(ops []Operand, size int) error {
i.imm = imm
return e.emit(i)
}
// MOV r/m, imm: 0xC6 (8-bit) / 0xC7 /0.
// MOV r/m, imm: 0xC6 (8-bit) / 0xC7 /0. An immediate in the
// destination slot is the absolute-address crash-store spelling,
// MOVL $0xf1, 0xf1: the parser reads the trailing bare constant
// as an immediate, and the store's disp32 carries the address.
op := byte(0xC7)
if size == 1 {
op = 0xC6
}
if d, ok := dst.(Imm); ok {
i := newInstr(size, []byte{op})
setSegAbs(i, 0, SegAbs{Disp: int64(d)})
immBytes, err := immediate(int64(src), size, false)
if err != nil {
return err
}
i.imm = immBytes
return e.emit(i)
}
i := newInstr(size, []byte{op})
if err := setRMDigit(i, 0, dst, size); err != nil {
return err
@@ -627,7 +664,16 @@ func (e *enc) encodeDoubleShift(base string, ops []Operand, size int) error {
func (e *enc) encodeImul(ops []Operand, size int) error {
switch len(ops) {
case 2:
// IMUL r, r/m: 0x0F 0xAF.
// Two shapes. The leading-immediate spelling IMUL $imm, r multiplies
// r in place (dst = rm = r): the shape GOROOT's clock code writes.
// Otherwise IMUL r, r/m: 0x0F 0xAF.
if imm, ok := ops[0].(Imm); ok {
dstReg, isReg := ops[1].(Reg)
if !isReg {
return fmt.Errorf("IMUL: destination must be a register")
}
return e.encodeImulImm(imm, dstReg, dstReg, size)
}
dstReg, ok := ops[1].(Reg)
if !ok {
return fmt.Errorf("IMUL: destination must be a register")
@@ -647,29 +693,36 @@ func (e *enc) encodeImul(ops []Operand, size int) error {
if !ok {
return fmt.Errorf("IMUL: immediate operand expected first")
}
// Plan 9 order: IMUL $imm, src, dst.
if fits8(int64(imm)) {
i := newInstr(size, []byte{0x6B})
if err := setRM(i, dstReg, ops[1], size); err != nil {
return err
}
i.imm = []byte{byte(int8(imm))}
return e.emit(i)
}
i := newInstr(size, []byte{0x69})
if err := setRM(i, dstReg, ops[1], size); err != nil {
return err
}
immBytes, err := immediate(int64(imm), size, false)
if err != nil {
return err
}
i.imm = immBytes
return e.emit(i)
// Plan 9 order: IMUL $imm, src, dst; the source stays a general
// r/m operand (setRM takes registers and memory alike).
return e.encodeImulImm(imm, ops[1], dstReg, size)
}
return fmt.Errorf("IMUL expects 2 or 3 operands, got %d", len(ops))
}
// encodeImulImm emits the immediate multiply: 0x6B with a sign-extended imm8
// when the value fits, 0x69 with a 32-bit immediate otherwise.
func (e *enc) encodeImulImm(imm Imm, rm Operand, dst Reg, size int) error {
if fits8(int64(imm)) {
i := newInstr(size, []byte{0x6B})
if err := setRM(i, dst, rm, size); err != nil {
return err
}
i.imm = []byte{byte(int8(imm))}
return e.emit(i)
}
i := newInstr(size, []byte{0x69})
if err := setRM(i, dst, rm, size); err != nil {
return err
}
immBytes, err := immediate(int64(imm), size, false)
if err != nil {
return err
}
i.imm = immBytes
return e.emit(i)
}
// --- PUSH / POP -------------------------------------------------------------
func (e *enc) encodePushPop(ops []Operand, size int, push bool) error {
+219
View File
@@ -0,0 +1,219 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
package asm
import (
"bytes"
"encoding/binary"
"os"
"os/exec"
"path/filepath"
"runtime"
"strings"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
)
// The differential kernels for the DATA-path and front-end gaps are kept in
// testdata/verify beside the campaign's other kernels; the verify package's
// suites are not open to the asm package, so this test is their runner: each
// kernel assembles through gasm and through go tool asm, and the functions'
// bytes must agree with the relocation sites masked on both sides.
// toolAsmObject assembles path with the installed toolchain's assembler for
// goarch ("" = the host) and returns the object bytes.
func toolAsmObject(t *testing.T, path, goarch string) []byte {
t.Helper()
goBin, err := exec.LookPath("go")
if err != nil {
t.Skip("no Go toolchain available")
}
out, err := exec.Command(goBin, "env", "GOROOT").Output()
if err != nil {
t.Fatalf("go env GOROOT: %v", err)
}
includeDir := filepath.Join(strings.TrimSpace(string(out)), "pkg", "include")
pkg := strings.TrimSuffix(filepath.Base(path), ".s")
pkg = strings.TrimSuffix(pkg, "_amd64")
pkg = strings.TrimSuffix(pkg, "_arm64")
objPath := filepath.Join(t.TempDir(), "oracle.o")
cmd := exec.Command(goBin, "tool", "asm", "-I", includeDir, "-p", pkg, "-o", objPath, path)
if goarch != "" {
environ := os.Environ()
env := make([]string, 0, len(environ)+1)
for _, e := range environ {
if !strings.HasPrefix(e, "GOARCH=") {
env = append(env, e)
}
}
cmd.Env = append(env, "GOARCH="+goarch)
}
if out, err := cmd.CombinedOutput(); err != nil {
t.Fatalf("go tool asm %s: %v\n%s", filepath.Base(path), err, out)
}
obj, err := os.ReadFile(objPath)
if err != nil {
t.Fatal(err)
}
return obj
}
// oracleFuncCode extracts the non-package TEXT functions' code bytes from a
// toolchain object, keyed by the name the object records (pkg.name). Each
// function's span is its own symbol size: a toolchain object that follows
// the text with data symbols (the synthesised float-constant pool) would
// otherwise fold them into the last function's bytes.
func oracleFuncCode(t *testing.T, obj []byte) map[string][]byte {
t.Helper()
v := openGoobj(t, obj)
le := binary.LittleEndian
const symSize = 21
nps := v.syms(blkNonpkgdef)
data := v.blk(blkData)
didx := v.blk(blkDataIdx)
preceding := 0
for _, bi := range []int{blkSymdef, blkHashed64def, blkHasheddef} {
preceding += len(v.blk(bi)) / symSize
}
out := make(map[string][]byte, len(nps))
for i, s := range nps {
if s.typ != kindSTEXT {
continue
}
start := le.Uint32(didx[4*(preceding+i):])
out[s.name] = data[start : start+s.size]
}
return out
}
// maskCode zeroes every relocation field, the way the toolchain's object
// leaves them for the linker.
func maskCode(code []byte, relocs []Reloc) []byte {
for _, r := range relocs {
for j := r.Off; j < r.Off+4 && j < len(code); j++ {
code[j] = 0
}
}
return code
}
// code assembles src for amd64 and returns the image's code bytes.
func code(path, src string) []byte {
f, errs := parser.Parse(path, src)
if len(errs) > 0 {
return nil
}
img, err := AssembleFile(f)
if err != nil {
return nil
}
return img.Code
}
// TestDifferentialKernels pins the new kernels against the oracle.
func TestDifferentialKernels(t *testing.T) {
if runtime.GOARCH != "amd64" {
t.Skip("the amd64 kernels assume an amd64 host assembler default")
}
for _, k := range []struct {
path string
goarch string
arm64 bool
}{
{filepath.Join("..", "testdata", "verify", "datarel_amd64.s"), "", false},
{filepath.Join("..", "testdata", "verify", "divslash_amd64.s"), "", false},
{filepath.Join("..", "testdata", "verify", "semicolons_amd64.s"), "", false},
{filepath.Join("..", "testdata", "verify", "quadreg_amd64.s"), "", false},
{filepath.Join("..", "testdata", "verify", "floatimm_amd64.s"), "", false},
{filepath.Join("..", "testdata", "verify", "bookkeep_amd64.s"), "", false},
{filepath.Join("..", "testdata", "verify", "forms_amd64.s"), "", false},
{filepath.Join("..", "testdata", "verify", "datarel_arm64.s"), "arm64", true},
{filepath.Join("..", "testdata", "verify", "divslash_arm64.s"), "arm64", true},
} {
t.Run(filepath.Base(k.path), func(t *testing.T) {
src, err := os.ReadFile(k.path)
if err != nil {
t.Fatalf("read: %v", err)
}
f, errs := parser.Parse(k.path, string(src))
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
var img *Image
if k.arm64 {
img, err = AssembleFileARM64(f)
} else {
img, err = AssembleFile(f)
}
if err != nil {
t.Fatalf("assemble: %v", err)
}
gt := oracleFuncCode(t, toolAsmObject(t, k.path, k.goarch))
// The oracle keys its functions by the qualified object name
// (pkg.name); match on the local part.
byLocal := make(map[string][]byte, len(gt))
for name, code := range gt {
if _, after, ok := strings.Cut(name, "."); ok {
name = after
}
byLocal[name] = code
}
matched := 0
for _, fn := range img.Funcs {
gasmCode := maskCode(append([]byte(nil), img.Code[fn.Offset:fn.Offset+fn.Size]...), fn.Relocs)
goCode, ok := byLocal[fn.Name]
if !ok {
t.Errorf("%s: not in ground truth (%d functions: %v)", fn.Name, len(gt), keysOf(byLocal))
continue
}
goCode = maskCode(append([]byte(nil), goCode...), fn.Relocs)
cmpLen := min(len(goCode), len(gasmCode))
if !bytes.Equal(gasmCode[:cmpLen], goCode[:cmpLen]) {
t.Errorf("%s: MISMATCH gasm=%d go=%d bytes\ngasm %x\ngo %x", fn.Name, len(gasmCode), len(goCode), gasmCode, goCode)
continue
}
for _, b := range goCode[len(gasmCode):] {
if b != 0 {
t.Errorf("%s: non-zero trailing bytes in go tool asm output", fn.Name)
break
}
}
matched++
t.Logf("%s: MATCH (%d bytes)", fn.Name, len(gasmCode))
}
if matched == 0 {
t.Fatal("no functions matched")
}
})
}
}
func keysOf(m map[string][]byte) []string {
out := make([]string, 0, len(m))
for k := range m {
out = append(out, k)
}
return out
}
// TestSemicolonSpellingParity pins that the ';' statement separator changes
// nothing about the encoding: the one-line spelling assembles to exactly the
// bytes of the same statements written one per line.
func TestSemicolonSpellingParity(t *testing.T) {
for _, tt := range []struct{ one, two string }{
{"\tROLQ $3, DI; ROLQ $13, DI\n", "\tROLQ $3, DI\n\tROLQ $13, DI\n"},
{"\tREP; MOVSQ\n", "\tREP\n\tMOVSQ\n"},
{"\tXORQ AX, AX; XORQ CX, CX\n", "\tXORQ AX, AX\n\tXORQ CX, CX\n"},
} {
one := code("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n"+tt.one+"\tRET\n")
two := code("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n"+tt.two+"\tRET\n")
if !bytes.Equal(one, two) {
t.Errorf("semicolon spelling %q: %x, want the two-line bytes %x", tt.one, one, two)
}
}
}
+176 -11
View File
@@ -5,6 +5,7 @@ package asm
import (
"fmt"
"math"
"sort"
"strconv"
@@ -103,6 +104,7 @@ const (
RelArm64Branch // R_CALLARM64 (BL instruction)
RelArm64LDST64 // R_ARM64_PCREL_LDST64 (ADRP + 64-bit LDR/STR pair)
RelLoong64Branch // R_CALLLOONG64 (BL instruction)
RelAddr // R_ADDR: the absolute address of a symbol held in a DATA field
)
type Reloc struct {
@@ -112,13 +114,17 @@ type Reloc struct {
// Addend select the target: the symbol plus the byte offset. An
// External relocation names a symbol no GLOBL in the file defines;
// the object-file emitters carry it into the output's relocation
// table.
// table. Siz is the width of the patched field and is set only for
// data-field relocations (RelAddr, Off relative to the data symbol),
// whose width is the DATA line's; code relocations take their width
// from the architecture's instruction encoding.
Off int
After int
Name string
Addend int64
External bool
Kind RelocKind
Siz uint8
}
// DataSymbol describes one GLOBL symbol laid out in the data section.
@@ -130,6 +136,11 @@ type DataSymbol struct {
Static bool // the <> marker: file-local, not exported
Rodata bool // the RODATA flag: read-only data
Dupok bool // the DUPOK flag: duplicate-OK
// Relocs carries the symbol-valued DATA initialisers ("DATA s+0(SB)/8,
// $other(SB)"): fields of this symbol's data that hold another symbol's
// address, resolved by the linker. Off is relative to the symbol's
// data start.
Relocs []Reloc
}
// Bytes returns the whole image: code, then data.
@@ -139,6 +150,18 @@ func (img *Image) Bytes() []byte {
return append(out, img.Data...)
}
// AssembleOption adjusts the file-level assembly context.
type AssembleOption func(*linkInfo)
// WithGOOS selects the target operating system for the forms that depend on
// it, the TLS access shape above all: linux and freebsd take the
// one-instruction form, windows and plan9 keep the two-instruction load.
func WithGOOS(goos string) AssembleOption {
return func(l *linkInfo) {
l.goos = goos
}
}
// AssembleFile assembles every TEXT function of a parsed file and lays out
// its static symbols (GLOBL/DATA) in a data section behind the code. Each
// reference to a file-local static symbol becomes a RIP-relative load whose
@@ -146,7 +169,7 @@ func (img *Image) Bytes() []byte {
// GLOBL defines is recorded as an external relocation (Externals) with its
// displacement left zero, the object-file emitters resolve it at link
// time, while the raw image (Bytes) cannot represent it.
func AssembleFile(f *ast.File) (*Image, error) {
func AssembleFile(f *ast.File, opts ...AssembleOption) (*Image, error) {
dataSyms, err := collectData(f)
if err != nil {
return nil, err
@@ -155,7 +178,18 @@ func AssembleFile(f *ast.File) (*Image, error) {
for _, d := range dataSyms {
known[d.name] = true
}
// TEXT symbols are file-level definitions too: a symbol immediate
// ($fn(SB)) may name one, exactly as a data reference names a GLOBL.
for _, d := range f.Decls {
if t, ok := d.(*ast.Text); ok {
known[t.Name.Name] = true
}
}
link := &linkInfo{symbols: known, allowExternal: true}
for _, o := range opts {
o(link)
}
poolSeen := map[string]bool{}
img := &Image{Symbols: map[string]int{}, SourcePath: f.Path}
textOff := map[string]int{}
@@ -169,7 +203,26 @@ func AssembleFile(f *ast.File) (*Image, error) {
if !ok {
continue
}
code, patches, labels, steps, lines, err := assemble(t, link)
code, patches, labels, steps, lines, pool, err := assemble(t, link)
if err != nil {
return nil, fmt.Errorf("%s: %w", t.Name.Name, err)
}
// The pooled floating-point constants join the declared data as
// read-only symbols, deduplicated across the file (the toolchain
// synthesises the same symbols into its rodata).
for _, entry := range pool {
if poolSeen[entry.name] {
continue
}
poolSeen[entry.name] = true
dataSyms = append(dataSyms, dataSym{
name: entry.name,
buf: entry.data,
size: len(entry.data),
rodata: true,
dupok: true,
})
}
if err != nil {
return nil, fmt.Errorf("%s: %w", t.Name.Name, err)
}
@@ -257,6 +310,22 @@ func AssembleFile(f *ast.File) (*Image, error) {
img.Funcs[i].Relocs = append(img.Funcs[i].Relocs, reloc)
}
}
// The data symbols' symbol-valued DATA fields resolve the same way the
// code references do: a name the file defines (GLOBL or TEXT) stays an
// internal reference the emitters resolve, anything else is external.
// img.DataSyms was laid out in dataSyms order, so the indexes line up.
for i := range img.DataSyms {
for _, r := range dataSyms[i].relocs {
reloc := r
if _, ok := img.Symbols[reloc.Name]; !ok {
if _, ok := textOff[reloc.Name]; !ok {
reloc.External = true
externals[reloc.Name] = true
}
}
img.DataSyms[i].Relocs = append(img.DataSyms[i].Relocs, reloc)
}
}
for name := range externals {
img.Externals = append(img.Externals, name)
}
@@ -273,6 +342,10 @@ func AssembleFileRISCV(f *ast.File) (*Image, error) {
if err != nil {
return nil, err
}
// The pooled $i64 constants the wide MOV immediate loads refer to join
// the declared data as read-only symbols, deduplicated across the file
// (the toolchain synthesises the same symbols into its rodata).
litSeen := map[string]bool{}
img := &Image{Symbols: map[string]int{}, SourcePath: f.Path}
for _, d := range f.Decls {
@@ -280,10 +353,23 @@ func AssembleFileRISCV(f *ast.File) (*Image, error) {
if !ok {
continue
}
code, labels, relocs, lines, spadj, err := assembleRISCV(t)
code, labels, relocs, lines, spadj, lits, err := assembleRISCV(t)
if err != nil {
return nil, fmt.Errorf("%s: %w", t.Name.Name, err)
}
for _, lit := range lits {
if litSeen[lit.Name] {
continue
}
litSeen[lit.Name] = true
dataSyms = append(dataSyms, dataSym{
name: lit.Name,
buf: lit.Data,
size: len(lit.Data),
rodata: true,
dupok: true,
})
}
fl := FuncLayout{
Name: t.Name.Name,
Pkg: t.Name.Pkg,
@@ -429,6 +515,23 @@ func markExternals(img *Image, dataSyms []dataSym) {
}
}
}
// The declared data symbols carry the file's own relocations (the
// symbol-valued DATA fields); the layouts appended img.DataSyms in
// dataSyms order, so the indexes line up. The trailing entries (the
// pooled arm64 literals) have no source relocations.
for i := range img.DataSyms {
if i >= len(dataSyms) {
break
}
for _, r := range dataSyms[i].relocs {
reloc := r
if !known[reloc.Name] {
reloc.External = true
externals[reloc.Name] = true
}
img.DataSyms[i].Relocs = append(img.DataSyms[i].Relocs, reloc)
}
}
for name := range externals {
img.Externals = append(img.Externals, name)
}
@@ -444,6 +547,9 @@ type dataSym struct {
static bool
rodata bool
dupok bool
// relocs are the symbol-valued DATA fields, in declaration order; Off
// is relative to the symbol's data start.
relocs []Reloc
}
// collectData gathers the file's static symbols (GLOBL) and their initial
@@ -511,20 +617,79 @@ func collectData(f *ast.File) ([]dataSym, error) {
if !ok {
return nil, fmt.Errorf("DATA %q: no matching GLOBL", dd.Name.Name)
}
if dd.Value == nil || !dd.Value.Imm.HasVal {
return nil, fmt.Errorf("DATA %q: value must be an integer immediate", dd.Name.Name)
if dd.Value == nil {
return nil, fmt.Errorf("DATA %q: missing value", dd.Name.Name)
}
w := dd.Width
switch w {
case 1, 2, 4, 8:
default:
return nil, fmt.Errorf("DATA %q: invalid width %d (want 1, 2, 4 or 8)", dd.Name.Name, w)
}
off := dd.Name.Offset
buf := syms[i].buf
if off < 0 || off+int64(w) > int64(len(buf)) {
return nil, fmt.Errorf("DATA %q+%d/%d exceeds GLOBL size %d", dd.Name.Name, off, w, len(buf))
}
// A symbol value ("DATA s+0(SB)/8, $other(SB)", the rt0 spelling)
// leaves the field zero and records a relocation against the named
// symbol: the linker patches the absolute address at this data
// offset. The toolchain emits the same shape, an R_ADDR of the
// DATA width with the value's offset as the addend, on every
// architecture.
if sym := dd.Value.Imm.Sym; !dd.Value.Imm.HasVal && sym != nil {
syms[i].relocs = append(syms[i].relocs, Reloc{
Off: int(off),
Name: sym.Name,
Addend: sym.Offset,
Kind: RelAddr,
Siz: uint8(w),
})
continue
}
// A string or rune value ("DATA s+0(SB)/20, $"text"") writes its
// bytes into the field and leaves the rest zero, the toolchain's
// WriteString: the declared width must hold every byte, and any
// width is legal.
if s := dd.Value.Imm.Str; s != "" && !dd.Value.Imm.HasVal {
text, err := strconv.Unquote(s)
if err != nil {
return nil, fmt.Errorf("DATA %q: invalid string value %s", dd.Name.Name, s)
}
if len(text) > w {
return nil, fmt.Errorf("DATA %q: string of %d bytes does not fit width %d", dd.Name.Name, len(text), w)
}
copy(buf[off:], text)
continue
}
// A floating-point value stores its IEEE-754 bits: /4 the float32
// rounding of the parsed double, /8 the full 64 bits, the
// toolchain's WriteFloat32 and WriteFloat64.
if f := dd.Value.Imm.Float; f != "" && !dd.Value.Imm.HasVal {
num, err := strconv.ParseFloat(f, 64)
if err != nil {
return nil, fmt.Errorf("DATA %q: invalid floating-point value %q", dd.Name.Name, f)
}
if dd.Value.Imm.Neg {
num = -num
}
var v uint64
switch w {
case 4:
v = uint64(math.Float32bits(float32(num)))
case 8:
v = math.Float64bits(num)
default:
return nil, fmt.Errorf("DATA %q: invalid width %d for a float (want 4 or 8)", dd.Name.Name, w)
}
for j := range w {
buf[off+int64(j)] = byte(v >> (8 * j))
}
continue
}
if !dd.Value.Imm.HasVal {
return nil, fmt.Errorf("DATA %q: value must be an integer immediate or a symbol address", dd.Name.Name)
}
switch w {
case 1, 2, 4, 8:
default:
return nil, fmt.Errorf("DATA %q: invalid width %d (want 1, 2, 4 or 8)", dd.Name.Name, w)
}
v := dd.Value.Imm.Val
if dd.Value.Imm.Neg {
v = -v
+325
View File
@@ -4,6 +4,11 @@
package asm
import (
"encoding/binary"
"fmt"
"os"
"os/exec"
"path/filepath"
"strings"
"testing"
@@ -166,3 +171,323 @@ func TestCollectDataNumericFlags(t *testing.T) {
}
}
}
// TestCollectDataSymbolValue covers the symbol-valued DATA field ("DATA
// s+0(SB)/8, $other(SB)", the rt0 spelling): the field stays zero in the
// image and the relocation is recorded against the named symbol, whatever
// the file defines (a TEXT function, a GLOBL) or leaves external.
func TestCollectDataSymbolValue(t *testing.T) {
src := `#include "textflag.h"
TEXT ·Keep(SB), NOSPLIT, $0-8
MOVQ target+0(FP), AX
RET
GLOBL holder(SB), NOPTR, $32
DATA holder+0(SB)/8, $·Keep(SB)
DATA holder+8(SB)/8, $·Keep+5(SB)
DATA holder+16(SB)/8, $holder(SB)
GLOBL spare(SB), NOPTR, $8
DATA spare+0(SB)/8, $extvar(SB)
`
f, errs := parser.Parse("f_amd64.s", src)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFile(f)
if err != nil {
t.Fatalf("assemble: %v", err)
}
byName := map[string]DataSymbol{}
for _, d := range img.DataSyms {
byName[d.Name] = d
}
want := []struct {
sym string
off int
name string
addend int64
ext bool
}{
{"holder", 0, "Keep", 0, false},
{"holder", 8, "Keep", 5, false},
{"holder", 16, "holder", 0, false},
{"spare", 0, "extvar", 0, true},
}
var flat []struct {
sym string
r Reloc
}
for _, d := range img.DataSyms {
for _, r := range d.Relocs {
flat = append(flat, struct {
sym string
r Reloc
}{d.Name, r})
}
}
if len(flat) != len(want) {
t.Fatalf("data relocations = %d, want %d", len(flat), len(want))
}
for i, w := range want {
g := flat[i]
r := g.r
if g.sym != w.sym {
t.Errorf("relocation %d sits on %q, want %q", i, g.sym, w.sym)
continue
}
if r.Off != w.off || r.Name != w.name || r.Addend != w.addend || r.External != w.ext {
t.Errorf("relocation %d = {+%d %q addend %d ext %v}, want {+%d %q addend %d ext %v}",
i, r.Off, r.Name, r.Addend, r.External, w.off, w.name, w.addend, w.ext)
}
if r.Kind != RelAddr {
t.Errorf("relocation %d kind = %v, want RelAddr", i, r.Kind)
}
if r.Siz != 8 {
t.Errorf("relocation %d siz = %d, want 8", i, r.Siz)
}
}
// The fields themselves stay zero: only the linker fills them.
for _, b := range img.Data {
if b != 0 {
t.Fatal("data section is not all zero before relocation")
}
}
if len(img.Externals) != 1 || img.Externals[0] != "extvar" {
t.Errorf("Externals = %v, want [extvar]", img.Externals)
}
}
// TestGOObjectDataSymbolReloc pins the GOOBJ record a symbol-valued DATA
// field produces, against the shape the toolchain emits for the same
// source: an R_ADDR of the DATA width at the field offset, pkgIdxNone plus
// the non-package definition index when the target is the file's own TEXT
// function (the rt0 lib entry spelling).
func TestGOObjectDataSymbolReloc(t *testing.T) {
f, errs := parser.Parse("f_amd64.s", `#include "textflag.h"
TEXT ·Keep(SB), NOSPLIT, $0-8
RET
GLOBL holder(SB), NOPTR, $16
DATA holder+0(SB)/8, $·Keep+5(SB)
`)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFile(f)
if err != nil {
t.Fatalf("assemble: %v", err)
}
obj, err := img.GOObject("main", "f_amd64.s")
if err != nil {
t.Fatalf("GOObject: %v", err)
}
v := openGoobj(t, obj)
// Walk every relocation record; the data record is the one of Siz 8
// and type R_ADDR.
var off, add int64
var pkg, sym uint32
found := false
for data := v.blk(blkReloc); len(data) >= 23; data = data[23:] {
if data[4] != 8 || binary.LittleEndian.Uint16(data[5:]) != relocAddr {
continue
}
found = true
off = int64(int32(binary.LittleEndian.Uint32(data[0:])))
add = int64(binary.LittleEndian.Uint64(data[7:]))
pkg = binary.LittleEndian.Uint32(data[15:])
sym = binary.LittleEndian.Uint32(data[19:])
break
}
if !found {
t.Fatal("no data relocation record in the object")
}
if off != 0 || add != 5 {
t.Errorf("data reloc = {off %d addend %d}, want {off 0 addend 5}", off, add)
}
if pkg != pkgIdxNone {
t.Errorf("data reloc pkg = %#x, want pkgIdxNone (the TEXT function)", pkg)
}
// The function's non-package definition index: the four pc tables
// precede it, so index 4.
if sym != 4 {
t.Errorf("data reloc sym = %d, want 4", sym)
}
}
// TestGOObjectDataSymbolLink is the end-to-end proof for symbol-valued DATA
// fields: the gasm object is substituted for the toolchain's and re-linked,
// then executed, and the linked data word must hold the real address of the
// function the DATA line named (runtime.FuncForPC identifies it).
func TestGOObjectDataSymbolLink(t *testing.T) {
goBin, err := exec.LookPath("go")
if err != nil {
t.Skip("no Go toolchain available")
}
dir := t.TempDir()
asmSrc := `#include "textflag.h"
GLOBL entry(SB), NOPTR, $8
DATA entry+0(SB)/8, $·keepme(SB)
TEXT ·keepme(SB), NOSPLIT, $0-0
RET
TEXT ·entryptr(SB), NOSPLIT, $0-8
MOVQ entry+0(SB), AX
MOVQ AX, ret+0(FP)
RET
`
if err := os.WriteFile(filepath.Join(dir, "main_amd64.s"), []byte(asmSrc), 0o644); err != nil {
t.Fatal(err)
}
mainSrc := `package main
import "runtime"
func keepme()
func entryptr() uintptr
func main() {
pc := entryptr()
fn := runtime.FuncForPC(pc)
if fn == nil {
panic("the entry word does not point at a function")
}
if fn.Name() != "main.keepme" {
panic("the entry word points at " + fn.Name())
}
}
`
if err := os.WriteFile(filepath.Join(dir, "main.go"), []byte(mainSrc), 0o644); err != nil {
t.Fatal(err)
}
if err := os.WriteFile(filepath.Join(dir, "go.mod"), []byte("module dlink\n\ngo 1.21\n"), 0o644); err != nil {
t.Fatal(err)
}
// Capture the build: the package archive's asm object and the link line.
build := exec.Command(goBin, "build", "-x", "-work", "-o", filepath.Join(dir, "prog"), ".")
build.Dir = dir
buildLog, err := build.CombinedOutput()
if err != nil {
t.Fatalf("baseline build: %v\n%s", err, buildLog)
}
var work, linkLine, asmObj string
for line := range strings.SplitSeq(string(buildLog), "\n") {
switch {
case strings.HasPrefix(line, "WORK="):
work = strings.TrimPrefix(line, "WORK=")
case strings.Contains(line, "/asm ") && strings.Contains(line, "main_amd64.s") && !strings.Contains(line, "-gensymabis"):
asmObj = fieldAfter(line, "-o")
case strings.Contains(line, "/link ") && strings.Contains(line, "-importcfg"):
linkLine = line
}
}
if work == "" || asmObj == "" || linkLine == "" {
t.Skipf("could not parse build log (work=%q asmObj=%q link=%q)", work, asmObj, linkLine)
}
defer os.RemoveAll(work)
asmObj = strings.ReplaceAll(asmObj, "$WORK", work)
linkLine = strings.ReplaceAll(linkLine, "$WORK", work)
// Assemble the same source with gasm and substitute the object.
src, err := os.ReadFile(filepath.Join(dir, "main_amd64.s"))
if err != nil {
t.Fatal(err)
}
f, errs := parser.Parse("main_amd64.s", string(src))
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFile(f)
if err != nil {
t.Fatalf("AssembleFile: %v", err)
}
gasmObj, err := img.GOObject("dlink", "main_amd64.s")
if err != nil {
t.Fatalf("GOObject: %v", err)
}
if err := os.WriteFile(asmObj, gasmObj, 0o644); err != nil {
t.Fatalf("write gasm object: %v", err)
}
linkCmd := exec.Command("bash", "-c", "cd "+dir+" && "+linkLine)
if out, err := linkCmd.CombinedOutput(); err != nil {
t.Fatalf("re-link with gasm object: %v\n%s", err, out)
}
// The linked program must run and find the right function behind the
// data word.
out, err := exec.Command(filepath.Join(dir, "prog")).CombinedOutput()
if err != nil {
t.Fatalf("linked program failed: %v\n%s", err, out)
}
}
// TestCollectDataFloatAndStringValues covers the non-integer DATA values the
// runtime's math and asm files use: floating-point initialisers store their
// IEEE-754 bits (/4 the float32 rounding, /8 the full double) and string
// initialisers write their bytes zero-padded within the declared width.
func TestCollectDataFloatAndStringValues(t *testing.T) {
src := `#include "textflag.h"
TEXT ·Keep(SB), NOSPLIT, $0-8
RET
GLOBL vals<>(SB), RODATA, $44
DATA vals<>+0(SB)/8, $0.5
DATA vals<>+8(SB)/8, $-1.0
DATA vals<>+16(SB)/4, $1.5
DATA vals<>+20(SB)/16, $"call frame too "
DATA vals<>+36(SB)/4, $"hi"
`
f, errs := parser.Parse("fvals_amd64.s", src)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFile(f)
if err != nil {
t.Fatalf("AssembleFile: %v", err)
}
byName := map[string]DataSymbol{}
for _, d := range img.DataSyms {
byName[d.Name] = d
}
d := byName["vals"]
if d.Size != 44 {
t.Fatalf("vals size = %d, want 44", d.Size)
}
buf := img.Data[d.Offset : d.Offset+44]
// 0.5 = 0x3FE0000000000000, -1.0 = 0xBFF0000000000000 (float64);
// 1.5 = 0x3FC00000 (float32).
for _, c := range []struct {
off int
want []byte
}{
{0, []byte{0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0xE0, 0x3F}},
{8, []byte{0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0xF0, 0xBF}},
{16, []byte{0x00, 0x00, 0xC0, 0x3F}},
{20, []byte("call frame too ")},
{36, []byte{'h', 'i', 0x00, 0x00}},
} {
if string(buf[c.off:c.off+len(c.want)]) != string(c.want) {
t.Errorf("vals+%d: got % x, want % x", c.off, buf[c.off:c.off+len(c.want)], c.want)
}
}
}
// TestCollectDataValueErrors pins the value-kind width rules: a float needs
// width 4 or 8, a string must fit its declared width, and a bad float
// literal is diagnosed rather than stored.
func TestCollectDataValueErrors(t *testing.T) {
cases := []string{
`GLOBL v<>(SB), RODATA, $4
DATA v<>+0(SB)/1, $0.5`,
`GLOBL v<>(SB), RODATA, $2
DATA v<>+0(SB)/2, $"toolarge"`,
}
for i, src := range cases {
full := "#include \"textflag.h\"\nTEXT ·Keep(SB), NOSPLIT, $0-8\n\tRET\n" + src
f, errs := parser.Parse(fmt.Sprintf("verr%d_amd64.s", i), full)
if len(errs) > 0 {
t.Fatalf("case %d parse: %v", i, errs)
}
if _, err := AssembleFile(f); err == nil {
t.Errorf("case %d: expected an error, got none", i)
}
}
}
+58
View File
@@ -372,6 +372,10 @@ func loong64InstrSize(instr *ast.Instr, fi loong64FrameInfo) int {
return len(loong64Return(fi))
}
switch mnem {
case "END", "FUNCDATA", "PCDATA":
return 0 // bookkeeping statements contribute no bytes
case "GETCALLERPC":
return 4 // or rd, r1, r0
case "TEQ", "TNE":
return 8 // bne/beq over the BREAK, then BREAK
case "PRELDX":
@@ -480,6 +484,36 @@ func encodeLOONG64Instr(instr *ast.Instr, pc int, offsets map[string]int, fi loo
return nil, fmt.Errorf("WORD expects 1 operand, got %d", len(ops))
}
return l64wordLE(uint32(immFromOperand(ops[0]))), nil
case "END", "FUNCDATA", "PCDATA", "GETCALLERPC":
// The assembler's bookkeeping statements. END, FUNCDATA and PCDATA
// contribute no bytes, the same shapes GOARCH=loong64 go tool asm
// accepts and emits nothing for; GETCALLERPC reads the caller's
// address out of R1 (RA) as or rd, r1, r0.
switch mnem {
case "END":
if len(ops) != 0 {
return nil, fmt.Errorf("END expects no operands, got %d", len(ops))
}
return nil, nil
case "FUNCDATA":
if len(ops) != 2 || !isImmOperand(ops[0]) {
return nil, fmt.Errorf("FUNCDATA expects $n, sym(SB)")
}
return nil, nil
case "PCDATA":
if len(ops) != 2 || !isImmOperand(ops[0]) || !isImmOperand(ops[1]) {
return nil, fmt.Errorf("PCDATA expects $n, $n")
}
return nil, nil
}
if len(ops) != 1 || isMemOperand(ops[0]) || loong64RegClass(operandRegName(ops[0])) != l64ClsGR {
return nil, fmt.Errorf("GETCALLERPC expects a general register")
}
rd := loong64RegNum(operandRegName(ops[0]))
if rd < 0 {
return nil, fmt.Errorf("GETCALLERPC: invalid register operand")
}
return l64wordLE(l64rrr(l64movRegTable["MOVV"].op, 0, 1, rd)), nil
case "NEGW", "NEGV":
// The integer negation pseudo is a subtract from zero:
// NEGW src, dst → sub.w r0, src, dst.
@@ -2084,6 +2118,30 @@ func encodeLOONG64Vector(instr *ast.Instr, mnem string, fi loong64FrameInfo) ([]
return l64wordLE(l64rr(l64InstrTable[mnem].op, vj, fcc)), true, nil
}
// Four-register forms (vshuf.b): INSTR va, vk, vj, vd.
if l64Vec4R[mnem] {
if len(ops) != 4 {
return nil, true, fmt.Errorf("%s expects 4 operands, got %d", mnem, len(ops))
}
va, err := vec(ops[0])
if err != nil {
return nil, true, err
}
vk, err := vec(ops[1])
if err != nil {
return nil, true, err
}
vj, err := vec(ops[2])
if err != nil {
return nil, true, err
}
vd, err := vec(ops[3])
if err != nil {
return nil, true, err
}
return l64wordLE(l64InstrTable[mnem].op | uint32(va&0x1f)<<15 | uint32(vk&0x1f)<<10 | uint32(vj&0x1f)<<5 | uint32(vd&0x1f)), true, nil
}
// Three-register forms: INSTR vk, vj, vd or INSTR vk, vd (vj = vd).
if len(ops) != 2 && len(ops) != 3 {
return nil, true, fmt.Errorf("%s expects 2 or 3 operands, got %d", mnem, len(ops))
+449 -12
View File
@@ -278,6 +278,7 @@ const (
l64Fpreld // preld (2RI12 + 5-bit hint)
l64Fvvv // 3R vector (LSX/LASX): op | vk<<10 | vj<<5 | vd
l64Fvcf // vector-to-condition: op | subop<<10 | vj<<5 | fcc
l64Fvvvv // 4R vector shuffle: op | va<<15 | vk<<10 | vj<<5 | vd
)
// l64Enc is one instruction's encoding: its bit layout (format) and the
@@ -336,6 +337,10 @@ var l64VecImmInfo = map[string]l64VecImmEnc{}
// vpcnt.v).
var l64Vec2R = map[string]bool{}
// l64Vec4R marks the four-operand vector mnemonics (INSTR va, vk, vj, vd,
// such as vshuf.b).
var l64Vec4R = map[string]bool{}
// l64VmovqOps holds the VMOVQ/XVMOVQ opcode constants (pre-shifted to bit
// 15), read off `go tool objdump` of GOARCH=loong64 `go tool asm` kernels.
type l64VmovqEnc struct {
@@ -445,6 +450,14 @@ func init() {
// bank is the FP registers (the toolchain spells it `FFINTDV F0, F1`),
// so the entry stays on the 2R integer/FP format.
"FFINTDV": 0x474a << 10,
// The rest of the scalar conversions (all F-bank, 2R).
"FFINTFW": 0x4744 << 10, // ffint.s.w
"FFINTFV": 0x4746 << 10, // ffint.s.l
"FFINTDW": 0x4748 << 10, // ffint.d.w
"FTINTWF": 0x46c1 << 10, // ftint.w.s
"FTINTWD": 0x46c2 << 10, // ftint.w.d
"FTINTVF": 0x46c9 << 10, // ftint.l.s
"FTINTVD": 0x46ca << 10, // ftint.l.d
}
for m, op := range rr {
l64InstrTable[m] = l64Enc{format: l64Frr, op: op}
@@ -568,6 +581,12 @@ func init() {
"AMADDDBW": 0x070D4 << 15, "AMADDDBV": 0x070D5 << 15,
"AMANDDBW": 0x070D6 << 15, "AMANDDBV": 0x070D7 << 15,
"AMORDBW": 0x070D8 << 15, "AMORDBV": 0x070D9 << 15,
// The remaining _dbar exchange variants (loong64enc1.s).
"AMXORDBW": 0x070DA << 15, "AMXORDBV": 0x070DB << 15,
"AMMAXDBW": 0x070DC << 15, "AMMAXDBV": 0x070DD << 15,
"AMMINDBW": 0x070DE << 15, "AMMINDBV": 0x070DF << 15,
"AMMAXDBWU": 0x070E0 << 15, "AMMAXDBVU": 0x070E1 << 15,
"AMMINDBWU": 0x070E2 << 15, "AMMINDBVU": 0x070E3 << 15,
}
for m, op := range am {
l64InstrTable[m] = l64Enc{format: l64Fam, op: op}
@@ -590,6 +609,240 @@ func init() {
"XVANDV": {0xEA4C << 15, true}, "XVXORV": {0xEA4E << 15, true},
"XVSEQB": {0xE800 << 15, true}, "XVSEQV": {0xE803 << 15, true},
}
// The integer and FP add/subtract families: [X]VADD and [X]VSUB by lane
// width, plus the [X]VSADD/[X]VSSUB saturating pairs.
// Opcodes transcribed from the toolchain's loong64enc1.s.
addsub := map[string]l64Vec3Enc{
"VADDB": {0xE014 << 15, false}, "VADDH": {0xE015 << 15, false},
"VADDD": {0xE262 << 15, false}, "VADDF": {0xE261 << 15, false},
"VADDQ": {0xE25A << 15, false},
"VSUBB": {0xE018 << 15, false}, "VSUBH": {0xE019 << 15, false},
"VSUBW": {0xE01A << 15, false}, "VSUBV": {0xE01B << 15, false},
"VSUBQ": {0xE25B << 15, false},
"VSUBF": {0xE265 << 15, false}, "VSUBD": {0xE266 << 15, false},
"VSADDB": {0xE08C << 15, false}, "VSADDH": {0xE08D << 15, false},
"VSADDW": {0xE08E << 15, false}, "VSADDV": {0xE08F << 15, false},
"VSADDBU": {0xE094 << 15, false}, "VSADDHU": {0xE095 << 15, false},
"VSADDWU": {0xE096 << 15, false}, "VSADDVU": {0xE097 << 15, false},
"VSSUBB": {0xE090 << 15, false}, "VSSUBH": {0xE091 << 15, false},
"VSSUBW": {0xE092 << 15, false}, "VSSUBV": {0xE093 << 15, false},
"VSSUBBU": {0xE098 << 15, false}, "VSSUBHU": {0xE099 << 15, false},
"VSSUBWU": {0xE09A << 15, false}, "VSSUBVU": {0xE09B << 15, false},
"XVADDB": {0xE814 << 15, true}, "XVADDH": {0xE815 << 15, true},
"XVADDW": {0xE816 << 15, true},
"XVADDD": {0xEA62 << 15, true}, "XVADDF": {0xEA61 << 15, true},
"XVADDQ": {0xEA5A << 15, true},
"XVSUBB": {0xE818 << 15, true}, "XVSUBH": {0xE819 << 15, true},
"XVSUBW": {0xE81A << 15, true}, "XVSUBV": {0xE81B << 15, true},
"XVSUBQ": {0xEA5B << 15, true},
"XVSUBF": {0xEA65 << 15, true}, "XVSUBD": {0xEA66 << 15, true},
"XVSADDB": {0xE88C << 15, true}, "XVSADDH": {0xE88D << 15, true},
"XVSADDW": {0xE88E << 15, true}, "XVSADDV": {0xE88F << 15, true},
"XVSADDBU": {0xE894 << 15, true}, "XVSADDHU": {0xE895 << 15, true},
"XVSADDWU": {0xE896 << 15, true}, "XVSADDVU": {0xE897 << 15, true},
"XVSSUBB": {0xE890 << 15, true}, "XVSSUBH": {0xE891 << 15, true},
"XVSSUBW": {0xE892 << 15, true}, "XVSSUBV": {0xE893 << 15, true},
"XVSSUBBU": {0xE898 << 15, true}, "XVSSUBHU": {0xE899 << 15, true},
"XVSSUBWU": {0xE89A << 15, true}, "XVSSUBVU": {0xE89B << 15, true},
}
// The multiply families: plain and high-half [X]VMUL/[X]VMUH, the
// widening [X]VMULW{EV,OD} ladder and its accumulating [X]VMADDW twins,
// plus the [X]VMADD/[X]VMSUB fused multiply-add and the [X]VDIV/[X]VMOD
// divide and modulo pairs.
muldiv := map[string]l64Vec3Enc{
"VMULB": {0xE108 << 15, false}, "VMULH": {0xE109 << 15, false},
"VMULW": {0xE10A << 15, false}, "VMULV": {0xE10B << 15, false},
"VMUHB": {0xE10C << 15, false}, "VMUHH": {0xE10D << 15, false},
"VMUHW": {0xE10E << 15, false}, "VMUHV": {0xE10F << 15, false},
"VMUHBU": {0xE110 << 15, false}, "VMUHHU": {0xE111 << 15, false},
"VMUHWU": {0xE112 << 15, false}, "VMUHVU": {0xE113 << 15, false},
"VMULWEVHB": {0xE120 << 15, false}, "VMULWEVWH": {0xE121 << 15, false},
"VMULWEVVW": {0xE122 << 15, false}, "VMULWEVQV": {0xE123 << 15, false},
"VMULWODHB": {0xE124 << 15, false}, "VMULWODWH": {0xE125 << 15, false},
"VMULWODVW": {0xE126 << 15, false}, "VMULWODQV": {0xE127 << 15, false},
"VMULWEVHBU": {0xE130 << 15, false}, "VMULWEVWHU": {0xE131 << 15, false},
"VMULWEVVWU": {0xE132 << 15, false}, "VMULWEVQVU": {0xE133 << 15, false},
"VMULWODHBU": {0xE134 << 15, false}, "VMULWODWHU": {0xE135 << 15, false},
"VMULWODVWU": {0xE136 << 15, false}, "VMULWODQVU": {0xE137 << 15, false},
"VMULWEVHBUB": {0xE140 << 15, false}, "VMULWEVWHUH": {0xE141 << 15, false},
"VMULWEVVWUW": {0xE142 << 15, false}, "VMULWEVQVUV": {0xE143 << 15, false},
"VMULWODHBUB": {0xE144 << 15, false}, "VMULWODWHUH": {0xE145 << 15, false},
"VMULWODVWUW": {0xE146 << 15, false}, "VMULWODQVUV": {0xE147 << 15, false},
"VMADDB": {0xE150 << 15, false}, "VMADDH": {0xE151 << 15, false},
"VMADDW": {0xE152 << 15, false}, "VMADDV": {0xE153 << 15, false},
"VMSUBB": {0xE154 << 15, false}, "VMSUBH": {0xE155 << 15, false},
"VMSUBW": {0xE156 << 15, false}, "VMSUBV": {0xE157 << 15, false},
"VMADDWEVHB": {0xE158 << 15, false}, "VMADDWEVWH": {0xE159 << 15, false},
"VMADDWEVVW": {0xE15A << 15, false}, "VMADDWEVQV": {0xE15B << 15, false},
"VMADDWODHB": {0xE15C << 15, false}, "VMADDWODWH": {0xE15D << 15, false},
"VMADDWODVW": {0xE15E << 15, false}, "VMADDWODQV": {0xE15F << 15, false},
"VMADDWEVHBU": {0xE168 << 15, false}, "VMADDWEVWHU": {0xE169 << 15, false},
"VMADDWEVVWU": {0xE16A << 15, false}, "VMADDWEVQVU": {0xE16B << 15, false},
"VMADDWODHBU": {0xE16C << 15, false}, "VMADDWODWHU": {0xE16D << 15, false},
"VMADDWODVWU": {0xE16E << 15, false}, "VMADDWODQVU": {0xE16F << 15, false},
"VMADDWEVHBUB": {0xE178 << 15, false}, "VMADDWEVWHUH": {0xE179 << 15, false},
"VMADDWEVVWUW": {0xE17A << 15, false}, "VMADDWEVQVUV": {0xE17B << 15, false},
"VMADDWODHBUB": {0xE17C << 15, false}, "VMADDWODWHUH": {0xE17D << 15, false},
"VMADDWODVWUW": {0xE17E << 15, false}, "VMADDWODQVUV": {0xE17F << 15, false},
"VDIVB": {0xE1C0 << 15, false}, "VDIVH": {0xE1C1 << 15, false},
"VDIVW": {0xE1C2 << 15, false}, "VDIVV": {0xE1C3 << 15, false},
"VMODB": {0xE1C4 << 15, false}, "VMODH": {0xE1C5 << 15, false},
"VMODW": {0xE1C6 << 15, false}, "VMODV": {0xE1C7 << 15, false},
"VDIVBU": {0xE1C8 << 15, false}, "VDIVHU": {0xE1C9 << 15, false},
"VDIVWU": {0xE1CA << 15, false}, "VDIVVU": {0xE1CB << 15, false},
"VMODBU": {0xE1CC << 15, false}, "VMODHU": {0xE1CD << 15, false},
"VMODWU": {0xE1CE << 15, false}, "VMODVU": {0xE1CF << 15, false},
"VMULF": {0xE271 << 15, false}, "VMULD": {0xE272 << 15, false},
"VDIVF": {0xE275 << 15, false}, "VDIVD": {0xE276 << 15, false},
"XVMULB": {0xE908 << 15, true}, "XVMULH": {0xE909 << 15, true},
"XVMULW": {0xE90A << 15, true}, "XVMULV": {0xE90B << 15, true},
"XVMUHB": {0xE90C << 15, true}, "XVMUHH": {0xE90D << 15, true},
"XVMUHW": {0xE90E << 15, true}, "XVMUHV": {0xE90F << 15, true},
"XVMUHBU": {0xE910 << 15, true}, "XVMUHHU": {0xE911 << 15, true},
"XVMUHWU": {0xE912 << 15, true}, "XVMUHVU": {0xE913 << 15, true},
"XVMULWEVHB": {0xE920 << 15, true}, "XVMULWEVWH": {0xE921 << 15, true},
"XVMULWEVVW": {0xE922 << 15, true}, "XVMULWEVQV": {0xE923 << 15, true},
"XVMULWODHB": {0xE924 << 15, true}, "XVMULWODWH": {0xE925 << 15, true},
"XVMULWODVW": {0xE926 << 15, true}, "XVMULWODQV": {0xE927 << 15, true},
"XVMULWEVHBU": {0xE930 << 15, true}, "XVMULWEVWHU": {0xE931 << 15, true},
"XVMULWEVVWU": {0xE932 << 15, true}, "XVMULWEVQVU": {0xE933 << 15, true},
"XVMULWODHBU": {0xE934 << 15, true}, "XVMULWODWHU": {0xE935 << 15, true},
"XVMULWODVWU": {0xE936 << 15, true}, "XVMULWODQVU": {0xE937 << 15, true},
"XVMULWEVHBUB": {0xE940 << 15, true}, "XVMULWEVWHUH": {0xE941 << 15, true},
"XVMULWEVVWUW": {0xE942 << 15, true}, "XVMULWEVQVUV": {0xE943 << 15, true},
"XVMULWODHBUB": {0xE944 << 15, true}, "XVMULWODWHUH": {0xE945 << 15, true},
"XVMULWODVWUW": {0xE946 << 15, true}, "XVMULWODQVUV": {0xE947 << 15, true},
"XVMADDB": {0xE950 << 15, true}, "XVMADDH": {0xE951 << 15, true},
"XVMADDW": {0xE952 << 15, true}, "XVMADDV": {0xE953 << 15, true},
"XVMSUBB": {0xE954 << 15, true}, "XVMSUBH": {0xE955 << 15, true},
"XVMSUBW": {0xE956 << 15, true}, "XVMSUBV": {0xE957 << 15, true},
"XVMADDWEVHB": {0xE958 << 15, true}, "XVMADDWEVWH": {0xE959 << 15, true},
"XVMADDWEVVW": {0xE95A << 15, true}, "XVMADDWEVQV": {0xE95B << 15, true},
"XVMADDWODHB": {0xE95C << 15, true}, "XVMADDWODWH": {0xE95D << 15, true},
"XVMADDWODVW": {0xE95E << 15, true}, "XVMADDWODQV": {0xE95F << 15, true},
"XVMADDWEVHBU": {0xE968 << 15, true}, "XVMADDWEVWHU": {0xE969 << 15, true},
"XVMADDWEVVWU": {0xE96A << 15, true}, "XVMADDWEVQVU": {0xE96B << 15, true},
"XVMADDWODHBU": {0xE96C << 15, true}, "XVMADDWODWHU": {0xE96D << 15, true},
"XVMADDWODVWU": {0xE96E << 15, true}, "XVMADDWODQVU": {0xE96F << 15, true},
"XVMADDWEVHBUB": {0xE978 << 15, true}, "XVMADDWEVWHUH": {0xE979 << 15, true},
"XVMADDWEVVWUW": {0xE97A << 15, true}, "XVMADDWEVQVUV": {0xE97B << 15, true},
"XVMADDWODHBUB": {0xE97C << 15, true}, "XVMADDWODWHUH": {0xE97D << 15, true},
"XVMADDWODVWUW": {0xE97E << 15, true}, "XVMADDWODQVUV": {0xE97F << 15, true},
"XVDIVB": {0xE9C0 << 15, true}, "XVDIVH": {0xE9C1 << 15, true},
"XVDIVW": {0xE9C2 << 15, true}, "XVDIVV": {0xE9C3 << 15, true},
"XVMODB": {0xE9C4 << 15, true}, "XVMODH": {0xE9C5 << 15, true},
"XVMODW": {0xE9C6 << 15, true}, "XVMODV": {0xE9C7 << 15, true},
"XVDIVBU": {0xE9C8 << 15, true}, "XVDIVHU": {0xE9C9 << 15, true},
"XVDIVWU": {0xE9CA << 15, true}, "XVDIVVU": {0xE9CB << 15, true},
"XVMODBU": {0xE9CC << 15, true}, "XVMODHU": {0xE9CD << 15, true},
"XVMODWU": {0xE9CE << 15, true}, "XVMODVU": {0xE9CF << 15, true},
"XVMULF": {0xEA71 << 15, true}, "XVMULD": {0xEA72 << 15, true},
"XVDIVF": {0xEA75 << 15, true}, "XVDIVD": {0xEA76 << 15, true},
}
// The lane-wise shifts and rotates (three-register forms; the immediate
// forms live in l64VecImmInfo), the interleave families, the bit
// clear/set/rev register forms, the remaining logic and compare
// spellings, the widening add/subtract ladder and the vector FP
// arithmetic.
vecmisc := map[string]l64Vec3Enc{
"VSLLB": {0xE1D0 << 15, false}, "VSLLH": {0xE1D1 << 15, false},
"VSLLW": {0xE1D2 << 15, false}, "VSLLV": {0xE1D3 << 15, false},
"VSRLB": {0xE1D4 << 15, false}, "VSRLH": {0xE1D5 << 15, false},
"VSRLW": {0xE1D6 << 15, false}, "VSRLV": {0xE1D7 << 15, false},
"VSRAH": {0xE1D9 << 15, false}, "VSRAW": {0xE1DA << 15, false},
"VSRAV": {0xE1DB << 15, false},
"VROTRB": {0xE1DC << 15, false}, "VROTRH": {0xE1DD << 15, false},
"VROTRV": {0xE1DF << 15, false},
"VILVLB": {0xE234 << 15, false}, "VILVLH": {0xE235 << 15, false},
"VILVLW": {0xE236 << 15, false}, "VILVLV": {0xE237 << 15, false},
"VILVHB": {0xE238 << 15, false}, "VILVHH": {0xE239 << 15, false},
"VILVHW": {0xE23A << 15, false}, "VILVHV": {0xE23B << 15, false},
"VBITCLRB": {0xE218 << 15, false}, "VBITCLRH": {0xE219 << 15, false},
"VBITCLRW": {0xE21A << 15, false}, "VBITCLRV": {0xE21B << 15, false},
"VBITSETB": {0xE21C << 15, false}, "VBITSETH": {0xE21D << 15, false},
"VBITSETW": {0xE21E << 15, false}, "VBITSETV": {0xE21F << 15, false},
"VBITREVB": {0xE220 << 15, false}, "VBITREVH": {0xE221 << 15, false},
"VBITREVW": {0xE222 << 15, false}, "VBITREVV": {0xE223 << 15, false},
"VORV": {0xE24D << 15, false}, "VNORV": {0xE24F << 15, false},
"VANDNV": {0xE250 << 15, false}, "VORNV": {0xE251 << 15, false},
"VSEQH": {0xE001 << 15, false}, "VSEQW": {0xE002 << 15, false},
"VSLTB": {0xE00C << 15, false}, "VSLTH": {0xE00D << 15, false},
"VSLTW": {0xE00E << 15, false}, "VSLTV": {0xE00F << 15, false},
"VSLTBU": {0xE010 << 15, false}, "VSLTHU": {0xE011 << 15, false},
"VSLTWU": {0xE012 << 15, false}, "VSLTVU": {0xE013 << 15, false},
"VADDWEVHB": {0xE03C << 15, false}, "VADDWEVWH": {0xE03D << 15, false},
"VADDWEVVW": {0xE03E << 15, false}, "VADDWEVQV": {0xE03F << 15, false},
"VSUBWEVHB": {0xE040 << 15, false}, "VSUBWEVWH": {0xE041 << 15, false},
"VSUBWEVVW": {0xE042 << 15, false}, "VSUBWEVQV": {0xE043 << 15, false},
"VADDWODHB": {0xE044 << 15, false}, "VADDWODWH": {0xE045 << 15, false},
"VADDWODVW": {0xE046 << 15, false}, "VADDWODQV": {0xE047 << 15, false},
"VSUBWODHB": {0xE048 << 15, false}, "VSUBWODWH": {0xE049 << 15, false},
"VSUBWODVW": {0xE04A << 15, false}, "VSUBWODQV": {0xE04B << 15, false},
"VSUBWEVHBU": {0xE060 << 15, false}, "VSUBWEVWHU": {0xE061 << 15, false},
"VSUBWEVVWU": {0xE062 << 15, false}, "VSUBWEVQVU": {0xE063 << 15, false},
"VADDWEVHBU": {0xE05C << 15, false}, "VADDWEVWHU": {0xE05D << 15, false},
"VADDWEVVWU": {0xE05E << 15, false}, "VADDWEVQVU": {0xE05F << 15, false},
"VADDWODHBU": {0xE064 << 15, false}, "VADDWODWHU": {0xE065 << 15, false},
"VADDWODVWU": {0xE066 << 15, false}, "VADDWODQVU": {0xE067 << 15, false},
"VSUBWODHBU": {0xE068 << 15, false}, "VSUBWODWHU": {0xE069 << 15, false},
"VSUBWODVWU": {0xE06A << 15, false}, "VSUBWODQVU": {0xE06B << 15, false},
"VSHUFH": {0xE2F5 << 15, false}, "VSHUFW": {0xE2F6 << 15, false},
"VSHUFV": {0xE2F7 << 15, false},
"XVSLLB": {0xE9D0 << 15, true}, "XVSLLH": {0xE9D1 << 15, true},
"XVSLLW": {0xE9D2 << 15, true}, "XVSLLV": {0xE9D3 << 15, true},
"XVSRLB": {0xE9D4 << 15, true}, "XVSRLH": {0xE9D5 << 15, true},
"XVSRLW": {0xE9D6 << 15, true}, "XVSRLV": {0xE9D7 << 15, true},
"XVSRAB": {0xE9D8 << 15, true}, "XVSRAH": {0xE9D9 << 15, true},
"XVSRAW": {0xE9DA << 15, true}, "XVSRAV": {0xE9DB << 15, true},
"XVROTRB": {0xE9DC << 15, true}, "XVROTRH": {0xE9DD << 15, true},
"XVROTRW": {0xE9DE << 15, true}, "XVROTRV": {0xE9DF << 15, true},
"XVILVLB": {0xEA34 << 15, true}, "XVILVLH": {0xEA35 << 15, true},
"XVILVLW": {0xEA36 << 15, true}, "XVILVLV": {0xEA37 << 15, true},
"XVILVHB": {0xEA38 << 15, true}, "XVILVHH": {0xEA39 << 15, true},
"XVILVHW": {0xEA3A << 15, true}, "XVILVHV": {0xEA3B << 15, true},
"XVBITCLRB": {0xEA18 << 15, true}, "XVBITCLRH": {0xEA19 << 15, true},
"XVBITCLRW": {0xEA1A << 15, true}, "XVBITCLRV": {0xEA1B << 15, true},
"XVBITSETB": {0xEA1C << 15, true}, "XVBITSETH": {0xEA1D << 15, true},
"XVBITSETW": {0xEA1E << 15, true}, "XVBITSETV": {0xEA1F << 15, true},
"XVBITREVB": {0xEA20 << 15, true}, "XVBITREVH": {0xEA21 << 15, true},
"XVBITREVW": {0xEA22 << 15, true}, "XVBITREVV": {0xEA23 << 15, true},
"XVORV": {0xEA4D << 15, true}, "XVNORV": {0xEA4F << 15, true},
"XVANDNV": {0xEA50 << 15, true}, "XVORNV": {0xEA51 << 15, true},
"XVSEQH": {0xE801 << 15, true}, "XVSEQW": {0xE802 << 15, true},
"XVSLTB": {0xE80C << 15, true}, "XVSLTH": {0xE80D << 15, true},
"XVSLTW": {0xE80E << 15, true}, "XVSLTV": {0xE80F << 15, true},
"XVSLTBU": {0xE810 << 15, true}, "XVSLTHU": {0xE811 << 15, true},
"XVSLTWU": {0xE812 << 15, true}, "XVSLTVU": {0xE813 << 15, true},
"XVADDWEVHB": {0xE83C << 15, true}, "XVADDWEVWH": {0xE83D << 15, true},
"XVADDWEVVW": {0xE83E << 15, true}, "XVADDWEVQV": {0xE83F << 15, true},
"XVSUBWEVHB": {0xE840 << 15, true}, "XVSUBWEVWH": {0xE841 << 15, true},
"XVSUBWEVVW": {0xE842 << 15, true}, "XVSUBWEVQV": {0xE843 << 15, true},
"XVADDWODHB": {0xE844 << 15, true}, "XVADDWODWH": {0xE845 << 15, true},
"XVADDWODVW": {0xE846 << 15, true}, "XVADDWODQV": {0xE847 << 15, true},
"XVSUBWODHB": {0xE848 << 15, true}, "XVSUBWODWH": {0xE849 << 15, true},
"XVSUBWODVW": {0xE84A << 15, true}, "XVSUBWODQV": {0xE84B << 15, true},
"XVADDWEVHBU": {0xE85C << 15, true}, "XVADDWEVWHU": {0xE85D << 15, true},
"XVADDWEVVWU": {0xE85E << 15, true}, "XVADDWEVQVU": {0xE85F << 15, true},
"XVSUBWEVHBU": {0xE860 << 15, true}, "XVSUBWEVWHU": {0xE861 << 15, true},
"XVSUBWEVVWU": {0xE862 << 15, true}, "XVSUBWEVQVU": {0xE863 << 15, true},
"XVADDWODHBU": {0xE864 << 15, true}, "XVADDWODWHU": {0xE865 << 15, true},
"XVADDWODVWU": {0xE866 << 15, true}, "XVADDWODQVU": {0xE867 << 15, true},
"XVSUBWODHBU": {0xE868 << 15, true}, "XVSUBWODWHU": {0xE869 << 15, true},
"XVSUBWODVWU": {0xE86A << 15, true}, "XVSUBWODQVU": {0xE86B << 15, true},
"XVSHUFH": {0xEAF5 << 15, true}, "XVSHUFW": {0xEAF6 << 15, true},
"XVSHUFV": {0xEAF7 << 15, true},
}
for _, tab := range []map[string]l64Vec3Enc{addsub, muldiv, vecmisc} {
for m, e := range tab {
if _, dup := vec3[m]; dup {
panic("loong64: duplicate vector mnemonic " + m)
}
vec3[m] = e
}
}
for m, e := range vec3 {
l64InstrTable[m] = l64Enc{format: l64Fvvv, op: e.op}
l64VecBank[m] = e.lasx
@@ -597,21 +850,156 @@ func init() {
// Immediate forms: INSTR $imm, vj, vd (or INSTR $imm, vd). The immediate
// range, bias and field mask are the ones the toolchain encodes: vandi.b
// stores the raw 8-bit constant, vsrai.b stores imm+8 (byte-lane bias),
// vseqi.b and vseqi.d store 5-bit and 7-bit two's-complement values.
// The mnemonics that also have a register form (VSEQB, VSEQV, VSRAB,
// VROTRW) keep their three-register entry in l64InstrTable; the
// dispatcher picks the immediate opcode from l64VecImmInfo by operand
// kind, so the immediate entries must not overwrite the table.
// stores the raw 8-bit constant, vsrari.b stores imm+8 (lane-width
// bias), the si5 compares store 5-bit two's-complement values and vseqi.d
// a 7-bit field the toolchain range-checks down to si5.
// The mnemonics that also have a register form (the shifts, the bit
// clear/set/rev families, VSEQ and the logic immediates) keep their
// three-register entry in l64InstrTable; the dispatcher picks the
// immediate opcode from l64VecImmInfo by operand kind, so the immediate
// entries must not overwrite the table.
vecImm := map[string]l64VecImmEnc{
"VANDB": {0xE7A0 << 15, false, 0, 255, 0, 0xFF},
"XVANDB": {0xEFA0 << 15, true, 0, 255, 0, 0xFF},
"VORB": {0xE7A8 << 15, false, 0, 255, 0, 0xFF},
"XVORB": {0xEFA8 << 15, true, 0, 255, 0, 0xFF},
"VXORB": {0xE7B0 << 15, false, 0, 255, 0, 0xFF},
"XVXORB": {0xEFB0 << 15, true, 0, 255, 0, 0xFF},
"VNORB": {0xE7B8 << 15, false, 0, 255, 0, 0xFF},
"XVNORB": {0xEFB8 << 15, true, 0, 255, 0, 0xFF},
"VSEQB": {0xE500 << 15, false, -16, 15, 0, 0x1F},
"XVSEQB": {0xE900 << 15, true, -16, 15, 0, 0x1F},
"VSEQV": {0xE503 << 15, false, -64, 63, 0, 0x7F},
"XVSEQV": {0xE903 << 15, true, -64, 63, 0, 0x7F},
"VSRAB": {0xE668 << 15, false, 0, 7, 8, 0x1F},
"VROTRW": {0xE541 << 15, false, 0, 31, 0, 0x1F},
// vseqi.h/w accept the same si5 window as vseqi.b; vseqi.d carries a
// 7-bit field, but the toolchain range-checks it down to si5 as well
// (GOARCH=loong64 go tool asm rejects VSEQV $32 and VSEQV $-64).
"VSEQH": {0xE501 << 15, false, -16, 15, 0, 0x1F},
"XVSEQH": {0xED01 << 15, true, -16, 15, 0, 0x1F},
"VSEQW": {0xE502 << 15, false, -16, 15, 0, 0x1F},
"XVSEQW": {0xED02 << 15, true, -16, 15, 0, 0x1F},
"VSEQV": {0xE503 << 15, false, -16, 15, 0, 0x7F},
"XVSEQV": {0xE903 << 15, true, -16, 15, 0, 0x7F},
// vslti compares against a signed (or, in the U spellings, unsigned)
// si5/ui5 constant.
"VSLTB": {0xE50C << 15, false, -16, 15, 0, 0x1F},
"XVSLTB": {0xED0C << 15, true, -16, 15, 0, 0x1F},
"VSLTH": {0xE50D << 15, false, -16, 15, 0, 0x1F},
"XVSLTH": {0xED0D << 15, true, -16, 15, 0, 0x1F},
"VSLTW": {0xE50E << 15, false, -16, 15, 0, 0x1F},
"XVSLTW": {0xED0E << 15, true, -16, 15, 0, 0x1F},
"VSLTV": {0xE50F << 15, false, -16, 15, 0, 0x1F},
"XVSLTV": {0xED0F << 15, true, -16, 15, 0, 0x1F},
"VSLTBU": {0xE510 << 15, false, 0, 31, 0, 0x1F},
"XVSLTBU": {0xED10 << 15, true, 0, 31, 0, 0x1F},
"VSLTHU": {0xE511 << 15, false, 0, 31, 0, 0x1F},
"XVSLTHU": {0xED11 << 15, true, 0, 31, 0, 0x1F},
"VSLTWU": {0xE512 << 15, false, 0, 31, 0, 0x1F},
"XVSLTWU": {0xED12 << 15, true, 0, 31, 0, 0x1F},
"VSLTVU": {0xE513 << 15, false, 0, 31, 0, 0x1F},
"XVSLTVU": {0xED13 << 15, true, 0, 31, 0, 0x1F},
// vaddi/vsubi take ui5 constants for every width on this toolchain
// (VADDVU $32 is rejected by the oracle although the field is ui8).
"VADDBU": {0xE514 << 15, false, 0, 31, 0, 0x1F},
"XVADDBU": {0xED14 << 15, true, 0, 31, 0, 0x1F},
"VADDHU": {0xE515 << 15, false, 0, 31, 0, 0x1F},
"XVADDHU": {0xED15 << 15, true, 0, 31, 0, 0x1F},
"VADDWU": {0xE516 << 15, false, 0, 31, 0, 0x1F},
"XVADDWU": {0xED16 << 15, true, 0, 31, 0, 0x1F},
"VADDVU": {0xE517 << 15, false, 0, 31, 0, 0x1F},
"XVADDVU": {0xED17 << 15, true, 0, 31, 0, 0x1F},
"VSUBBU": {0xE518 << 15, false, 0, 31, 0, 0x1F},
"XVSUBBU": {0xED18 << 15, true, 0, 31, 0, 0x1F},
"VSUBHU": {0xE519 << 15, false, 0, 31, 0, 0x1F},
"XVSUBHU": {0xED19 << 15, true, 0, 31, 0, 0x1F},
"VSUBWU": {0xE51A << 15, false, 0, 31, 0, 0x1F},
"XVSUBWU": {0xED1A << 15, true, 0, 31, 0, 0x1F},
"VSUBVU": {0xE51B << 15, false, 0, 31, 0, 0x1F},
"XVSUBVU": {0xED1B << 15, true, 0, 31, 0, 0x1F},
// The shift/rotate immediates ride in a width-sized field whose upper
// bits carry the lane-width code: vslli.b stores ui3 at [12:0] with
// bits [14:13] inside the opcode, vslli.h ui4 under a 4 bit mask, and
// the .w/.d spellings a raw ui5/ui6.
"VSLLB": {0x732C2000, false, 0, 7, 0, 0x7},
"XVSLLB": {0x772C2000, true, 0, 7, 0, 0x7},
"VSLLH": {0x732C4000, false, 0, 15, 0, 0xF},
"XVSLLH": {0x772C4000, true, 0, 15, 0, 0xF},
"VSLLW": {0xE659 << 15, false, 0, 31, 0, 0x1F},
"XVSLLW": {0xEE59 << 15, true, 0, 31, 0, 0x1F},
"VSLLV": {0xE65A << 15, false, 0, 63, 0, 0x3F},
"XVSLLV": {0xEE5A << 15, true, 0, 63, 0, 0x3F},
"VSRLB": {0x73302000, false, 0, 7, 0, 0x7},
"XVSRLB": {0x77302000, true, 0, 7, 0, 0x7},
"VSRLH": {0x73304000, false, 0, 15, 0, 0xF},
"XVSRLH": {0x77304000, true, 0, 15, 0, 0xF},
"VSRLW": {0xE661 << 15, false, 0, 31, 0, 0x1F},
"XVSRLW": {0xEE61 << 15, true, 0, 31, 0, 0x1F},
"VSRLV": {0xE662 << 15, false, 0, 63, 0, 0x3F},
"XVSRLV": {0xEE62 << 15, true, 0, 63, 0, 0x3F},
// vsrari/vrotri bias the field so the lane-width code rides above the
// shift amount (.b adds 8, .h 16, .w 32; .d is a raw ui6).
"VSRAB": {0xE668 << 15, false, 0, 7, 8, 0x1F},
"XVSRAB": {0xEE68 << 15, true, 0, 7, 8, 0x1F},
"VSRAH": {0x73344000, false, 0, 15, 0, 0xF},
"XVSRAH": {0x77344000, true, 0, 15, 0, 0xF},
"VSRAW": {0xE669 << 15, false, 0, 31, 0, 0x1F},
"XVSRAW": {0xEE69 << 15, true, 0, 31, 0, 0x1F},
"VSRAV": {0xE66A << 15, false, 0, 63, 0, 0x3F},
"XVSRAV": {0xEE6A << 15, true, 0, 63, 0, 0x3F},
"VROTRB": {0x72A02000, false, 0, 7, 0, 0x7},
"XVROTRB": {0x76A02000, true, 0, 7, 0, 0x7},
"VROTRH": {0x72A04000, false, 0, 15, 0, 0xF},
"XVROTRH": {0x76A04000, true, 0, 15, 0, 0xF},
"VROTRW": {0xE541 << 15, false, 0, 31, 0, 0x1F},
"XVROTRW": {0xED41 << 15, true, 0, 31, 0, 0x1F},
"VROTRV": {0xE542 << 15, false, 0, 63, 0, 0x3F},
"XVROTRV": {0xED42 << 15, true, 0, 63, 0, 0x3F},
// vbitclri/vbitseti/vbitrevi follow the same width-coded layout.
"VBITCLRB": {0x73102000, false, 0, 7, 0, 0x7},
"XVBITCLRB": {0x77102000, true, 0, 7, 0, 0x7},
"VBITCLRH": {0x73104000, false, 0, 15, 0, 0xF},
"XVBITCLRH": {0x77104000, true, 0, 15, 0, 0xF},
"VBITCLRW": {0xE621 << 15, false, 0, 31, 0, 0x1F},
"XVBITCLRW": {0xEE21 << 15, true, 0, 31, 0, 0x1F},
"VBITCLRV": {0xE622 << 15, false, 0, 63, 0, 0x3F},
"XVBITCLRV": {0xEE22 << 15, true, 0, 63, 0, 0x3F},
"VBITSETB": {0x73142000, false, 0, 7, 0, 0x7},
"XVBITSETB": {0x77142000, true, 0, 7, 0, 0x7},
"VBITSETH": {0x73144000, false, 0, 15, 0, 0xF},
"XVBITSETH": {0x77144000, true, 0, 15, 0, 0xF},
"VBITSETW": {0xE629 << 15, false, 0, 31, 0, 0x1F},
"XVBITSETW": {0xEE29 << 15, true, 0, 31, 0, 0x1F},
"VBITSETV": {0xE62A << 15, false, 0, 63, 0, 0x3F},
"XVBITSETV": {0xEE2A << 15, true, 0, 63, 0, 0x3F},
"VBITREVB": {0x73182000, false, 0, 7, 0, 0x7},
"XVBITREVB": {0x77182000, true, 0, 7, 0, 0x7},
"VBITREVH": {0x73184000, false, 0, 15, 0, 0xF},
"XVBITREVH": {0x77184000, true, 0, 15, 0, 0xF},
"VBITREVW": {0xE631 << 15, false, 0, 31, 0, 0x1F},
"XVBITREVW": {0xEE31 << 15, true, 0, 31, 0, 0x1F},
"VBITREVV": {0xE632 << 15, false, 0, 63, 0, 0x3F},
"XVBITREVV": {0xEE32 << 15, true, 0, 63, 0, 0x3F},
// The 4-bit-select shuffles and the byte-extract/insert permutations
// take ui8 (the .d shuffle ui4 range-checked to 0..15 by the
// toolchain) packing both position nibbles.
"VSHUF4IB": {0xE720 << 15, false, 0, 255, 0, 0xFF},
"XVSHUF4IB": {0xEF20 << 15, true, 0, 255, 0, 0xFF},
"VSHUF4IH": {0xE728 << 15, false, 0, 255, 0, 0xFF},
"XVSHUF4IH": {0xEF28 << 15, true, 0, 255, 0, 0xFF},
"VSHUF4IW": {0xE730 << 15, false, 0, 255, 0, 0xFF},
"XVSHUF4IW": {0xEF30 << 15, true, 0, 255, 0, 0xFF},
"VSHUF4IV": {0xE738 << 15, false, 0, 15, 0, 0xFF},
"XVSHUF4IV": {0xEF38 << 15, true, 0, 15, 0, 0xFF},
"VPERMIW": {0xE7C8 << 15, false, 0, 255, 0, 0xFF},
"XVPERMIW": {0xEFC8 << 15, true, 0, 255, 0, 0xFF},
"XVPERMIV": {0xEFD0 << 15, true, 0, 255, 0, 0xFF},
"XVPERMIQ": {0xEFD8 << 15, true, 0, 255, 0, 0xFF},
"VEXTRINSB": {0xE718 << 15, false, 0, 255, 0, 0xFF},
"XVEXTRINSB": {0xEF18 << 15, true, 0, 255, 0, 0xFF},
"VEXTRINSH": {0xE710 << 15, false, 0, 255, 0, 0xFF},
"XVEXTRINSH": {0xEF10 << 15, true, 0, 255, 0, 0xFF},
"VEXTRINSW": {0xE708 << 15, false, 0, 255, 0, 0xFF},
"XVEXTRINSW": {0xEF08 << 15, true, 0, 255, 0, 0xFF},
"VEXTRINSV": {0xE700 << 15, false, 0, 255, 0, 0xFF},
"XVEXTRINSV": {0xEF00 << 15, true, 0, 255, 0, 0xFF},
}
for m, e := range vecImm {
l64VecImmInfo[m] = e
@@ -625,22 +1013,71 @@ func init() {
"VSETANYEQB": 0xE539<<15 | 8<<10, "XVSETANYEQB": 0xED39<<15 | 8<<10,
"VSETANYEQV": 0xE539<<15 | 11<<10, "XVSETANYEQV": 0xED39<<15 | 11<<10,
"VSETALLNEV": 0xE539<<15 | 15<<10, "XVSETALLNEV": 0xED39<<15 | 15<<10,
"VSETEQV": 0xE539<<15 | 6<<10, "XVSETEQV": 0xED39<<15 | 6<<10,
"VSETANYEQH": 0xE539<<15 | 9<<10, "XVSETANYEQH": 0xED39<<15 | 9<<10,
"VSETANYEQW": 0xE539<<15 | 10<<10, "XVSETANYEQW": 0xED39<<15 | 10<<10,
"VSETALLNEB": 0xE539<<15 | 12<<10, "XVSETALLNEB": 0xED39<<15 | 12<<10,
"VSETALLNEH": 0xE539<<15 | 13<<10, "XVSETALLNEH": 0xED39<<15 | 13<<10,
"VSETALLNEW": 0xE539<<15 | 14<<10, "XVSETALLNEW": 0xED39<<15 | 14<<10,
}
for m, op := range vecCf {
l64InstrTable[m] = l64Enc{format: l64Fvcf, op: op}
l64VecBank[m] = strings.HasPrefix(m, "XV")
}
// Lane popcount: INSTR vj, vd (the 2R layout with the opcode extending
// over the unused vk field).
// Lane popcount and the two-operand vector FP/unary spellings: INSTR vj,
// vd (the 2R layout with the opcode extending over the unused vk field;
// the low byte of each constant is the instruction's own sub-op).
vec2r := map[string]l64Vec3Enc{
"VPCNTV": {0x1CA70B << 10, false}, "XVPCNTV": {0x1DA70B << 10, true},
}
// The rest of the lane popcounts, the vector negations and the vector FP
// unary conversions (loong64enc1.s).
vec2rMore := map[string]l64Vec3Enc{
"VPCNTB": {0x1CA708 << 10, false}, "VPCNTH": {0x1CA709 << 10, false},
"VPCNTW": {0x1CA70A << 10, false},
"VNEGB": {0x1CA70C << 10, false}, "VNEGH": {0x1CA70D << 10, false},
"VNEGW": {0x1CA70E << 10, false}, "VNEGV": {0x1CA70F << 10, false},
"VFCLASSF": {0x1CA735 << 10, false}, "VFCLASSD": {0x1CA736 << 10, false},
"VFSQRTF": {0x1CA739 << 10, false}, "VFSQRTD": {0x1CA73A << 10, false},
"VFRECIPF": {0x1CA73D << 10, false}, "VFRECIPD": {0x1CA73E << 10, false},
"VFRSQRTF": {0x1CA741 << 10, false}, "VFRSQRTD": {0x1CA742 << 10, false},
"VFRINTF": {0x1CA74D << 10, false}, "VFRINTD": {0x1CA74E << 10, false},
"VFRINTRMF": {0x1CA751 << 10, false}, "VFRINTRMD": {0x1CA752 << 10, false},
"VFRINTRPF": {0x1CA755 << 10, false}, "VFRINTRPD": {0x1CA756 << 10, false},
"VFRINTRZF": {0x1CA759 << 10, false}, "VFRINTRZD": {0x1CA75A << 10, false},
"VFRINTRNEF": {0x1CA75D << 10, false}, "VFRINTRNED": {0x1CA75E << 10, false},
"XVPCNTB": {0x1DA708 << 10, true}, "XVPCNTH": {0x1DA709 << 10, true},
"XVPCNTW": {0x1DA70A << 10, true},
"XVNEGB": {0x1DA70C << 10, true}, "XVNEGH": {0x1DA70D << 10, true},
"XVNEGW": {0x1DA70E << 10, true}, "XVNEGV": {0x1DA70F << 10, true},
"XVFCLASSF": {0x1DA735 << 10, true}, "XVFCLASSD": {0x1DA736 << 10, true},
"XVFSQRTF": {0x1DA739 << 10, true}, "XVFSQRTD": {0x1DA73A << 10, true},
"XVFRECIPF": {0x1DA73D << 10, true}, "XVFRECIPD": {0x1DA73E << 10, true},
"XVFRSQRTF": {0x1DA741 << 10, true}, "XVFRSQRTD": {0x1DA742 << 10, true},
"XVFRINTF": {0x1DA74D << 10, true}, "XVFRINTD": {0x1DA74E << 10, true},
"XVFRINTRMF": {0x1DA751 << 10, true}, "XVFRINTRMD": {0x1DA752 << 10, true},
"XVFRINTRPF": {0x1DA755 << 10, true}, "XVFRINTRPD": {0x1DA756 << 10, true},
"XVFRINTRZF": {0x1DA759 << 10, true}, "XVFRINTRZD": {0x1DA75A << 10, true},
"XVFRINTRNEF": {0x1DA75D << 10, true}, "XVFRINTRNED": {0x1DA75E << 10, true},
}
maps.Copy(vec2r, vec2rMore)
for m, e := range vec2r {
l64InstrTable[m] = l64Enc{format: l64Frr, op: e.op}
l64VecBank[m] = e.lasx
l64Vec2R[m] = true
}
// The four-register byte shuffle: INSTR va, vk, vj, vd (the operand the
// table reads in each field position, va at bits [19:15]).
vec4r := map[string]l64Vec3Enc{
"VSHUFB": {0x0D50 << 16, false}, "XVSHUFB": {0x0D60 << 16, true},
}
for m, e := range vec4r {
l64InstrTable[m] = l64Enc{format: l64Fvvvv, op: e.op}
l64VecBank[m] = e.lasx
l64Vec4R[m] = true
}
}
// l64FpMovTable maps (mnemonic, from-class, to-class) to the 2R opcode of the
+242
View File
@@ -507,6 +507,218 @@ TEXT ·v(SB), NOSPLIT, $0
0x4C000020,
)
})
// The integer and FP add/subtract families with their saturating pairs
// and immediate spellings (loong64enc1.s words).
t.Run("add and subtract families", func(t *testing.T) {
fn := firstTextLOONG64(t, `#include "textflag.h"
TEXT ·v(SB), NOSPLIT, $0
VADDB V1, V2, V3
VADDF V1, V2, V3
VADDD V1, V2, V3
VSUBD V1, V2, V3
VSADDV V1, V2, V3
VSSUBVU V1, V2, V3
VADDBU $1, V2, V1
VADDBU $1, V2
VSUBVU $31, V2
XVSADDV X3, X2, X1
XVSUBD X1, X2, X3
RET
`)
code := assembleLOONG64Helper(t, fn)
wantWords(t, code,
0x700A0443, // vadd.b
0x71308443, // vadd.f
0x71310443, // vadd.d
0x71330443, // vsub.d
0x70478443, // vsadd.v
0x704D8443, // vssub.u.d
0x728A0441, // vaddi.bu v1, v2, 1
0x728A0442, // vaddi.bu v2, v2, 1 (two-operand form)
0x728DFC42, // vsubi.du v2, v2, 31 (two-operand form)
0x74478C41, // xvsadd.d x1, x2, x3
0x75330443, // xvsub.d x3, x2, x1
0x4C000020,
)
})
// The multiply, divide and accumulate families.
t.Run("multiply and divide families", func(t *testing.T) {
fn := firstTextLOONG64(t, `#include "textflag.h"
TEXT ·v(SB), NOSPLIT, $0
VMULV V1, V2, V3
VMUHHU V1, V2, V3
VDIVBU V1, V2, V3
VMODV V1, V2, V3
VMADDB V1, V2, V3
VMSUBV V1, V2, V3
VMULWEVHB V1, V2, V3
VMULWODQV V1, V2, V3
VMADDWEVHBUB V1, V2, V3
XVDIVD X1, X2, X3
RET
`)
code := assembleLOONG64Helper(t, fn)
wantWords(t, code,
0x70858443, // vmul.v
0x70888443, // vmuh.u.d
0x70E40443, // vdiv.u.b
0x70E38443, // vmod.d
0x70A80443, // vmadd.b
0x70AB8443, // vmsub.d
0x70900443, // vmulwev.h.b
0x70938443, // vmulwod.q.d
0x70BC0443, // vmaddwev.h.bu.b
0x753B0443, // xvdiv.d
0x4C000020,
)
})
// The shift, bit and interleave families in register and immediate
// spellings, with the width-coded shift immediates.
t.Run("shift, bit and interleave families", func(t *testing.T) {
fn := firstTextLOONG64(t, `#include "textflag.h"
TEXT ·v(SB), NOSPLIT, $0
VSLLV V1, V2, V3
VROTRB V1, V2, V3
VBITCLRV V1, V2, V3
VBITSETW V1, V2, V3
VBITREVV V1, V2, V3
VILVLB V1, V2, V3
VILVHV V1, V2, V3
VSLLB $7, V1, V2
VSLLB $5, V1
VSRLH $15, V1, V2
VSRAW $31, V1, V2
VSRAV $63, V1, V2
VROTRV $63, V1, V2
VBITCLRB $7, V2, V3
VBITREVV $63, V2, V3
VSEQH $-16, V2, V3
VSLTB $1, V2, V3
VSLTHU $31, V2, V3
XVILVLV X3, X2, X1
XVSLLB $7, X2, X1
XVSRAV $63, X2, X1
XVBITREVV $63, X2, X1
RET
`)
code := assembleLOONG64Helper(t, fn)
wantWords(t, code,
0x70E98443, // vsll.d
0x70EE0443, // vrotr.b
0x710D8443, // vbitclr.d
0x710F0443, // vbitset.w
0x71118443, // vbitrev.d
0x711A0443, // vilvl.b
0x711D8443, // vilvh.d
0x732C3C22, // vslli.b v2, v1, 7
0x732C3421, // vslli.b v1, v1, 5 (two-operand form)
0x73307C22, // vsrli.h v2, v1, 15
0x7334FC22, // vsrai.w v2, v1, 31
0x7335FC22, // vsrai.d v2, v1, 63
0x72A1FC22, // vrotri.d v2, v1, 63
0x73103C43, // vbitclri.b v3, v2, 7
0x7319FC43, // vbitrevi.d v3, v2, 63
0x7280C043, // vseqi.h v3, v2, -16
0x72860443, // vslti.b v3, v2, 1
0x7288FC43, // vslti.hu v3, v2, 31
0x751B8C41, // xvilvl.d x1, x2, x3
0x772C3C41, // xvslli.b x1, x2, 7
0x7735FC41, // xvsrai.d x1, x2, 63
0x7719FC41, // xvbitrevi.d x1, x2, 63
0x4C000020,
)
})
// The shuffle, select and permutation families, including the
// four-register byte shuffle.
t.Run("shuffle and permutation families", func(t *testing.T) {
fn := firstTextLOONG64(t, `#include "textflag.h"
TEXT ·v(SB), NOSPLIT, $0
VSHUFH V1, V2, V3
VSHUFW V1, V2, V3
VSHUFV V1, V2, V3
VSHUFB V1, V2, V3, V4
XVSHUFB X1, X2, X3, X4
VSHUF4IB $255, V2, V1
VSHUF4IV $15, V2, V1
XVSHUF4IV $15, X1, X2
VEXTRINSB $0x18, V1, V2
XVEXTRINSV $0x81, X1, X2
VPERMIW $0x1B, V1, V2
XVPERMIQ $0x4B, X1, X2
RET
`)
code := assembleLOONG64Helper(t, fn)
wantWords(t, code,
0x717A8443, // vshuf.h
0x717B0443, // vshuf.w
0x717B8443, // vshuf.d
0x0D508864, // vshuf.b v4, v3, v2, v1
0x0D608864, // xvshuf.b
0x7393FC41, // vshuf4i.b v1, v2, 255
0x739C3C41, // vshuf4i.d v1, v2, 15
0x779C3C22, // xvshuf4i.d x2, x1, 15
0x738C6022, // vextrins.b v2, v1, 0x18
0x77820422, // xvextrins.d x2, x1, 0x81
0x73E46C22, // vpermi.w v2, v1, 0x1b
0x77ED2C22, // xvpermi.q x2, x1, 0x4b
0x4C000020,
)
})
// The vector FP families, the unary spellings, the compare-to-flag
// additions and the scalar int/float conversions.
t.Run("FP and conversion families", func(t *testing.T) {
fn := firstTextLOONG64(t, `#include "textflag.h"
TEXT ·v(SB), NOSPLIT, $0
VADDF V1, V2, V3
VMULF V1, V2, V3
VFCLASSD V1, V2
VFSQRTF V1, V2
VFRECIPD V1, V2
VFRSQRTF V1, V2
VFRINTF V1, V2
VFRINTRNED V1, V2
VNEGB V1, V2
VPCNTB V1, V2
XVNEGV X2, X1
XVPCNTW X3, X2
XVFRINTRNEF X1, X2
VSETEQV V1, FCC0
VSETANYEQH V1, FCC0
VSETALLNEB V1, FCC0
XVSETALLNEW X1, FCC0
FFINTFW F0, F1
FTINTVD F0, F1
RET
`)
code := assembleLOONG64Helper(t, fn)
wantWords(t, code,
0x71308443, // vfadd.s
0x71388443, // vfmul.s
0x729CD822, // vfclass.d
0x729CE422, // vfsqrt.s
0x729CF822, // vfrecip.d
0x729D0422, // vfrsqrt.s
0x729D3422, // vfrint.s
0x729D7822, // vfrintne.s
0x729C3022, // vneg.b
0x729C2022, // vpcnt.b
0x769C3C41, // xvneg.d x1, x2
0x769C2862, // xvpcnt.w x2, x3
0x769D7422, // xvfrintne.s x2, x1
0x729C9820, // vseteqz.d fcc0, v1
0x729CA420, // vsetanyeqz.h
0x729CB020, // vsetallnez.b
0x769CB820, // xvsetallnez.w
0x011D1001, // ffint.s.w f1, f0
0x011B2801, // ftint.l.d f1, f0
0x4C000020,
)
})
}
// TestLOONG64_vectorErrors pins the register-class and range diagnostics of
@@ -545,6 +757,36 @@ func TestLOONG64_vectorErrors(t *testing.T) {
`TEXT ·e(SB), NOSPLIT, $0
VROTRW $32, V1, V2
RET
`,
`TEXT ·e(SB), NOSPLIT, $0
VADDVU $32, V2
RET
`,
`TEXT ·e(SB), NOSPLIT, $0
VSEQV $32, V2, V3
RET
`,
`TEXT ·e(SB), NOSPLIT, $0
VSHUF4IV $16, V2, V1
RET
`,
`TEXT ·e(SB), NOSPLIT, $0
VEXTRINSB $256, V1, V2
RET
`,
`TEXT ·e(SB), NOSPLIT, $0
VSLTV $-17, V2, V3
RET
`,
// VSHUFB wants four vector registers.
`TEXT ·e(SB), NOSPLIT, $0
VSHUFB V1, V2, V3
RET
`,
// The FCC forms still refuse vector registers.
`TEXT ·e(SB), NOSPLIT, $0
VSETEQV V1, V2
RET
`,
// VSET* wants an FCC flag, not a vector register.
`TEXT ·e(SB), NOSPLIT, $0
+46
View File
@@ -14,6 +14,51 @@ type Imm int64
func (Imm) isOperand() {}
// RegList is a bracketed register range, [Z0-Z3]: the four-register source
// of the 4FMAPS and 4VNNIW families. The EVEX emit path carries the list's
// low register through the inverted 5-bit V'VVVV field; the three higher
// registers are implied by the instruction, so only the pair travels here.
type RegList struct {
Lo Reg
Hi Reg // implied by the encoding; Lo.idx+3 by construction
}
func (RegList) isOperand() {}
// FloatImm is a floating-point immediate ($-1.0). The SSE mnemonics whose
// encoding takes an XMM/memory source at that position rewrite it as a read
// from a read-only pool constant ($f64.<hex> or $f32.<hex>), the toolchain's
// own behaviour; every other instruction rejects it.
type FloatImm struct {
Text string // the numeric text as written, sign excluded
Neg bool // a leading minus
}
func (FloatImm) isOperand() {}
// TLSMem is a thread-local access, the source form off(base)(TLS*1) with the
// base dropped: the toolchain's one-instruction TLS rewrite assembles it as
// the segment-prefixed absolute whose disp32 carries an R_TLS_LE patch site
// (the linker fills the TLS slot offset).
type TLSMem struct {
Disp int64
Size int
Seg byte // the segment override: FS (0x64) or GS (0x65) on windows
}
func (TLSMem) isOperand() {}
// SegAbs is a segment-absolute access, 0x30(GS): the segment override
// prefixes a disp32 absolute reference with no relocation. The base
// register spellings GS and FS produce it.
type SegAbs struct {
Disp int64
Size int
Seg byte // 0x64 FS, 0x65 GS
}
func (SegAbs) isOperand() {}
// Mem is a memory operand of the form disp(base)(index*scale).
type Mem struct {
Base Reg
@@ -23,6 +68,7 @@ type Mem struct {
Size int // operand width in bytes
HasBase bool
HasIndex bool
Seg byte // segment override prefix (0x64 FS, 0x65 GS); 0 = none
}
func (Mem) isOperand() {}
+286 -35
View File
@@ -6,6 +6,7 @@ package asm
import (
"errors"
"fmt"
"math/bits"
"slices"
"strings"
@@ -14,13 +15,14 @@ import (
// assembleRISCV assembles a RISC-V TEXT function body into machine code.
// It handles the full RV64IMAFDC instruction set including RVC compression.
func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, []SpadjStep, error) {
func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, []SpadjStep, []RiscvLiteral, error) {
fi := riscvComputeFrame(t)
prologue := riscvPrologue(fi)
guardLen, err := riscvGuardLen(fi)
if err != nil {
return nil, nil, nil, nil, nil, err
return nil, nil, nil, nil, nil, nil, err
}
lits := &riscvLiterals{}
var relocs []Reloc
var spadj []SpadjStep
@@ -74,9 +76,9 @@ func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, [
pc := len(prologue)
for i := range recs {
branchLike := isBranchLike(recs[i].instr.Mnemonic.Text) || riscvIsCondBranch(recs[i].instr.Mnemonic.Text)
code, err := encodeRISCVInstr(recs[i].instr, pc, offsets, fi, nil, nil) // no relocs in Pass 2
code, err := encodeRISCVInstr(recs[i].instr, pc, offsets, fi, nil, nil, lits) // no relocs in Pass 2
if err != nil && !(branchLike && riscvIsRangeError(err)) {
return nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", recs[i].instr.Mnemonic.Text, err)
return nil, nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", recs[i].instr.Mnemonic.Text, err)
}
if err != nil {
code = make([]byte, 4)
@@ -178,13 +180,20 @@ func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, [
}
}
if !changed {
// Capture the final pcs for the N(PC) branch forms: their target
// is the instruction N source slots away, resolved by index.
// Capture the final pcs for the N(PC) branch and jump forms: the
// target is the instruction N source slots away (N=0 the branch
// itself, N negative backwards), resolved by index against the
// final layout.
pcRelPcs = map[*ast.Instr]int{}
for i := range recs {
if _, ok := riscvPCRelOffset(recs[i].instr); ok {
pcRelPcs[recs[i].instr] = pcs[i]
n, ok := riscvPCRelOffset(recs[i].instr)
if !ok {
continue
}
if i+n < 0 || i+n >= len(recs) {
continue
}
pcRelPcs[recs[i].instr] = pcs[i+n]
}
break
}
@@ -198,7 +207,7 @@ func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, [
var out []byte
guardBytes, guardReloc, err := riscvGuard(fi)
if err != nil {
return nil, nil, nil, nil, nil, err
return nil, nil, nil, nil, nil, nil, err
}
if fi.needSplit {
out = append(out, guardBytes...)
@@ -220,11 +229,11 @@ func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, [
// The JMP a relaxation inserted: JAL X0 to the original target.
targetOff, ok := offsets[r.jmpTo]
if !ok {
return nil, nil, nil, nil, nil, fmt.Errorf("undefined label %q", r.jmpTo)
return nil, nil, nil, nil, nil, nil, fmt.Errorf("undefined label %q", r.jmpTo)
}
offset := int32(targetOff - pc)
if err := riscvCheckJumpOffset(r.jmpTo, offset); err != nil {
return nil, nil, nil, nil, nil, err
return nil, nil, nil, nil, nil, nil, err
}
word := riscvJType(0, offset)
code = []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}
@@ -233,7 +242,7 @@ func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, [
// JMP, always the very next instruction (offset 4).
enc, rs1, rs2, ok := riscvInvertedBranchEnc(strings.ToUpper(r.instr.Mnemonic.Text), r.instr.Operands)
if !ok {
return nil, nil, nil, nil, nil, fmt.Errorf("%s: cannot relax branch", r.instr.Mnemonic.Text)
return nil, nil, nil, nil, nil, nil, fmt.Errorf("%s: cannot relax branch", r.instr.Mnemonic.Text)
}
word := riscvBType(enc, rs1, rs2, 4)
code = []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}
@@ -241,9 +250,9 @@ func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, [
code = r.code
default:
var err error
code, err = encodeRISCVInstr(r.instr, pc, offsets, fi, &relocs, pcRelPcs)
code, err = encodeRISCVInstr(r.instr, pc, offsets, fi, &relocs, pcRelPcs, lits)
if err != nil {
return nil, nil, nil, nil, nil, err
return nil, nil, nil, nil, nil, nil, err
}
if c16, ok := tryCompressRVC(r.instr, fi); ok {
code = []byte{byte(c16), byte(c16 >> 8)}
@@ -270,7 +279,7 @@ func assembleRISCV(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry, [
if fi.needSplit {
relocs = append(relocs, guardReloc)
}
return out, offsets, relocs, lines, spadj, nil
return out, offsets, relocs, lines, spadj, lits.list(), nil
}
// riscvImmAlias maps the R-type ALU mnemonics onto their I-type immediate
@@ -348,6 +357,11 @@ func riscvPadBytes(pad int) []byte {
func riscvInstrSize(instr *ast.Instr, fi riscvFrameInfo) int {
mnem := instr.Mnemonic.Text
ops := instr.Operands
mnem = riscvNormalisePseudo(mnem)
if mnem == "FUNCDATA" || mnem == "PCDATA" {
// The bookkeeping statements contribute no bytes.
return 0
}
var immNeg bool
mnem, immNeg = riscvNormaliseImmAlias(mnem, ops)
if mnem == "RET" {
@@ -368,7 +382,25 @@ func riscvInstrSize(instr *ast.Instr, fi riscvFrameInfo) int {
}
// MOV $imm, rd → size depends on the immediate and RVC compression.
if isImmOperand(ops[0]) && ops[0].Imm.Sym == nil {
return riscvMovImmSize(regFromOperand(ops[1]), immFromOperand(ops[0]))
imm := riscvOperandImm64(ops[0])
if int64(int32(imm)) != imm {
return riscvMovImm64Size(regFromOperand(ops[1]), imm)
}
return riscvMovImmSize(regFromOperand(ops[1]), int32(imm))
}
// MOV $sym+off(FP|SP), rd → the frame-adjusted offset as an ADDI,
// compressed like riscvSPAddiBytes encodes it.
if isImmOperand(ops[0]) && ops[0].Imm.Sym != nil &&
(ops[0].Imm.Sym.Pseudo == "FP" || ops[0].Imm.Sym.Pseudo == "SP") {
rd := regFromOperand(ops[1])
_, off := riscvResolvePseudo(ops[0].Imm.Sym, fi)
if rd > 0 && off == 0 {
return 2 // C.MV rd, SP
}
if isRVCIntReg(rd) && off > 0 && off < 1024 && off%4 == 0 {
return 2 // C.ADDI4SPN
}
return riscvItypeImmediateSize("ADDI", off)
}
// Frame-relative loads and stores: a frame offset beyond the signed
// 12-bit range materialises the address in X31 first.
@@ -664,9 +696,10 @@ func riscvCheckJumpOffset(target string, off int32) error {
}
// encodeRISCVInstr encodes a single RISC-V instruction.
func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscvFrameInfo, relocs *[]Reloc, pcRelPcs map[*ast.Instr]int) ([]byte, error) {
func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscvFrameInfo, relocs *[]Reloc, pcRelPcs map[*ast.Instr]int, lits *riscvLiterals) ([]byte, error) {
mnem := instr.Mnemonic.Text
ops := instr.Operands
mnem = riscvNormalisePseudo(mnem)
var immNeg bool
mnem, immNeg = riscvNormaliseImmAlias(mnem, ops)
var word uint32
@@ -677,6 +710,22 @@ func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscv
// RET = epilogue (restore LR and close the frame when present) +
// uncompressed JALR X0, 0(X1) (the toolchain never compresses RET).
return riscvReturn(fi), nil
case "FUNCDATA":
// The assembler's bookkeeping statement, the expanded form of the
// GO_ARGS and NO_LOCAL_POINTERS macros: FUNCDATA $n, sym(SB)
// contributes no bytes, exactly as the toolchain's listing shows
// (the FUNCDATA entries and the instruction after them share a PC).
if len(ops) != 2 || !isImmOperand(ops[0]) {
return nil, fmt.Errorf("FUNCDATA expects $n, sym(SB)")
}
return nil, nil
case "PCDATA":
// The other bookkeeping statement, the expanded form of
// GO_RESULTS_INITIALIZED: PCDATA $n, $m contributes no bytes too.
if len(ops) != 2 || !isImmOperand(ops[0]) || !isImmOperand(ops[1]) {
return nil, fmt.Errorf("PCDATA expects $n, $m")
}
return nil, nil
case "WORD":
// WORD $w lays down a raw 32-bit little-endian word.
if len(ops) != 1 {
@@ -739,6 +788,22 @@ func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscv
return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil
}
target = labelFromOperand(ops[0])
// JMP N(PC): the PC-relative slot form, resolved like the
// branches (the toolchain counts source instructions at a
// uniform 4 bytes, so JMP 0(PC) is a self-loop and JMP -3(PC)
// reaches twelve bytes back). It must be recognised before the
// indirect-register form, whose operand it resembles.
if off, isPCRel, err := riscvPCRelTargetOff(instr, pc, pcRelPcs); isPCRel {
if err != nil {
return nil, err
}
offset := int32(off - pc)
if err := riscvCheckJumpOffset("", offset); err != nil {
return nil, err
}
word = riscvJType(0, offset)
return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil
}
// JMP (X5): an indirect branch, the toolchain's JALR X0, 0(X5).
if ops[0].Addr.Sym == nil && ops[0].Addr.Base != "" {
if ops[0].Addr.Offset != 0 || ops[0].Addr.Index != "" {
@@ -751,17 +816,6 @@ func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscv
word = riscvIType(riscvEnc{0x67, 0x0, 0x00}, 0, rs1, 0)
return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil
}
if off, isPCRel, err := riscvPCRelTargetOff(instr, pc, pcRelPcs); isPCRel {
if err != nil {
return nil, err
}
offset := int32(off - pc)
if err := riscvCheckJumpOffset("", offset); err != nil {
return nil, err
}
word = riscvJType(0, offset)
return []byte{byte(word), byte(word >> 8), byte(word >> 16), byte(word >> 24)}, nil
}
}
targetOff, ok := offsets[target]
if !ok {
@@ -810,7 +864,7 @@ func encodeRISCVInstr(instr *ast.Instr, pc int, offsets map[string]int, fi riscv
// (MOVB/MOVH/MOVW and unsigned forms) select the access width, and
// MOVD/MOVF address the FP registers.
case "MOV", "MOVB", "MOVBU", "MOVH", "MOVHU", "MOVW", "MOVWU", "MOVF", "MOVD":
return encodeRISCVMov(instr, fi, relocs)
return encodeRISCVMov(instr, fi, relocs, lits)
// JALR: indirect jump/call. Plan 9: JALR rs1, rd or JALR offset(rs1).
case "JALR":
@@ -1316,7 +1370,7 @@ func isImmOperand(op *ast.Operand) bool {
// - MOV Rs, (Rd) register-relative store
// - MOV Rs, Rd register-to-register move (ADDI $0)
// - MOV $imm, Rd load immediate (ADDI or LUI+ADDIW)
func encodeRISCVMov(instr *ast.Instr, fi riscvFrameInfo, relocs *[]Reloc) ([]byte, error) {
func encodeRISCVMov(instr *ast.Instr, fi riscvFrameInfo, relocs *[]Reloc, lits *riscvLiterals) ([]byte, error) {
ops := instr.Operands
if len(ops) != 2 {
return nil, fmt.Errorf("MOV expects 2 operands, got %d", len(ops))
@@ -1335,8 +1389,21 @@ func encodeRISCVMov(instr *ast.Instr, fi riscvFrameInfo, relocs *[]Reloc) ([]byt
}
return encodeRISCVSBAddr(src.Imm.Sym, rd, relocs), nil
}
// MOV $sym+off(FP|SP), rd: the address of a frame slot as an
// immediate is the frame-adjusted offset against the hardware SP,
// the toolchain's ADDI $adj, SP, rd (argframe+0(FP) in the runtime's
// reflect trampolines is the spelling).
if src.Imm.Sym != nil && (src.Imm.Sym.Pseudo == "FP" || src.Imm.Sym.Pseudo == "SP") {
rd := regFromOperand(dst)
if rd < 0 {
return nil, fmt.Errorf("MOV $%s(%s): invalid destination register", src.Imm.Sym.Name, src.Imm.Sym.Pseudo)
}
_, off := riscvResolvePseudo(src.Imm.Sym, fi)
return riscvSPAddiBytes(rd, off), nil
}
// MOV $sym(FP/SP), rd, not supported: immediate symbol references
// other than SB cannot be encoded as a simple immediate.
// other than the frame pseudos cannot be encoded as a simple
// immediate.
if src.Imm.Sym != nil && src.Imm.Sym.Pseudo != "" {
return nil, fmt.Errorf("MOV $%s(%s): unsupported immediate symbol reference (only SB is supported)", src.Imm.Sym.Name, src.Imm.Sym.Pseudo)
}
@@ -1344,11 +1411,14 @@ func encodeRISCVMov(instr *ast.Instr, fi riscvFrameInfo, relocs *[]Reloc) ([]byt
if rd < 0 {
return nil, fmt.Errorf("MOV $imm: invalid destination register")
}
imm, err := riscvImm32FromOperand(src, false)
if err != nil {
return nil, err
imm := riscvOperandImm64(src)
if int64(int32(imm)) != imm {
// Beyond the signed 32-bit span the toolchain either builds the
// value from a shifted 32-bit part or loads it from the pooled
// $i64 constant it synthesises for the purpose.
return riscvLoadImm64(rd, imm, lits, relocs), nil
}
return encodeRISCVLoadImm(rd, imm), nil
return encodeRISCVLoadImm(rd, int32(imm)), nil
}
// Memory → register (load).
@@ -1557,6 +1627,186 @@ func splitRISCV32Imm(imm int32) (low, high int32) {
return low, high
}
// riscvNormalisePseudo rewrites the toolchain's UNDEF spelling onto EBREAK:
// the assembler accepts UNDEF where the hardware wants the trap instruction
// and emits ebreak (compressed to C.EBREAK under RVC), so every pass sees the
// canonical name.
func riscvNormalisePseudo(mnem string) string {
if strings.EqualFold(mnem, "UNDEF") {
return "EBREAK"
}
return mnem
}
// riscvOperandImm64 reads an immediate operand as a full signed 64-bit value,
// where immFromOperand would truncate to int32; the MOV immediate path uses
// it to classify the wide constants.
func riscvOperandImm64(op *ast.Operand) int64 {
if !op.Imm.HasVal {
return 0
}
v := op.Imm.Val
if op.Imm.Neg {
v = -v
}
return v
}
// riscvSplitShiftConst mirrors cmd/internal/obj/riscv's splitShiftConst: it
// looks for the signed 32-bit integer a constant can be rebuilt from with a
// left shift, a left-and-right shift pair (a run of ones), or a zero-extended
// 32-bit pattern. A constant that fits none of the shapes is materialised
// from the pooled $i64 data symbol instead.
func riscvSplitShiftConst(v int64) (imm int64, lsh int, rsh int, ok bool) {
// Rebuild from a signed 32-bit integer shifted left.
lsh = bits.TrailingZeros64(uint64(v))
c := v >> lsh
if int64(int32(c)) == c {
return c, lsh, 0, true
}
// Rebuild from a small negative constant: shift left into place, then
// shift the sign-extended ones run right.
rsh = bits.LeadingZeros64(uint64(v))
ones := bits.OnesCount64((uint64(v) >> lsh) >> 11)
if rsh+ones+lsh+11 == 64 {
c = (1<<11 | ((v >> lsh) & 0x7ff)) << 52 >> 52 // sign extend 12 bits
if lsh > 0 || c != -1 {
lsh += rsh
}
return c, lsh, rsh, true
}
// Rebuild from a zero-extended signed 32-bit integer.
if int64(uint32(c)) == c {
c = int64(int32(c))
lsh, rsh = 32, 32-lsh
return c, lsh, rsh, true
}
return 0, 0, 0, false
}
// riscvSPAddiBytes encodes ADDI rd, SP, imm for the frame-address immediates
// (the MOV $sym+off(FP|SP) form), using the compressed forms the toolchain
// picks under RVC: C.ADDI4SPN for a positive 4-byte multiple that fits, C.MV
// for the zero offset, the plain ADDI otherwise.
func riscvSPAddiBytes(rd int, imm int32) []byte {
if rd != 0 && imm == 0 {
return word16(rvcCR(0x8, uint32(rd), 2)) // C.MV rd, SP
}
if isRVCIntReg(rd) && imm > 0 && imm < 1024 && imm%4 == 0 {
return word16(rvcCIW(0x0, rvcReg3(rd), uint32(imm)))
}
return wordLE(riscvIType(riscvInstrTable["ADDI"], rd, 2, imm))
}
// riscvMovImm64Size returns the encoded byte length of MOV $imm, rd when the
// immediate sits outside the signed 32-bit span: the shifted-part sequences
// of riscvLoadImm64, or the 8-byte AUIPC+LD pool load.
func riscvMovImm64Size(rd int, imm int64) int {
c, lsh, rsh, ok := riscvSplitShiftConst(imm)
if !ok {
return 8 // AUIPC + LD against the $i64 pool symbol
}
size := riscvMovImmSize(rd, int32(c))
if lsh > 0 {
size += riscvShiftImmSize(rd, true)
}
if rsh > 0 {
size += riscvShiftImmSize(rd, false)
}
return size
}
// riscvShiftImmSize returns the encoded size of one SLLI/SRLI expansion
// part: two bytes under RVC when the destination can carry a compressed
// shift (C.SLLI admits every register but X0, C.SRLI only X8 to X15), four
// otherwise.
func riscvShiftImmSize(rd int, left bool) int {
if rd != 0 && (left || isRVCIntReg(rd)) {
return 2
}
return 4
}
// riscvLoadImm64 encodes MOV $imm, rd for an immediate beyond the signed
// 32-bit span, mirroring the toolchain's instructionsForMOVConst: when a
// shifted 32-bit part rebuilds the value it emits that part (compressed like
// any written MOV) followed by the SLLI and SRLI shifts; otherwise it loads
// the constant from the pooled read-only $i64.<hex> symbol via AUIPC + LD
// and registers the literal so the data section carries its bytes.
func riscvLoadImm64(rd int, imm int64, lits *riscvLiterals, relocs *[]Reloc) []byte {
c, lsh, rsh, ok := riscvSplitShiftConst(imm)
if !ok {
name := fmt.Sprintf("$i64.%016x", uint64(imm))
if lits != nil {
lits.add(name, riscvLiteralBytes(imm))
}
return encodeRISCVSBLoad(&ast.Symbol{Name: name}, rd, relocs)
}
out := encodeRISCVLoadImm(rd, int32(c))
if lsh > 0 {
out = append(out, riscvShiftImmBytes(rd, lsh, true)...)
}
if rsh > 0 {
out = append(out, riscvShiftImmBytes(rd, rsh, false)...)
}
return out
}
// riscvShiftImmBytes encodes one SLLI (left) or SRLI expansion part, using
// the compressed form the toolchain picks under RVC: C.SLLI admits every
// register but X0, C.SRLI only X8 to X15.
func riscvShiftImmBytes(rd, shamt int, left bool) []byte {
if rd != 0 && shamt >= 1 && shamt <= 63 && (left || isRVCIntReg(rd)) {
if left {
return word16(rvcSLLI(uint32(rd), uint32(shamt)&0x3F))
}
return word16(rvcCBShift(0x0, rvcReg3(rd), uint32(shamt)&0x3F))
}
enc := riscvEnc{0x13, 0x1, 0x00} // SLLI
imm := int32(shamt)
if !left {
enc = riscvEnc{0x13, 0x5, 0x00} // SRLI: funct6 000000, funct3 101
}
return wordLE(riscvIType(enc, rd, rd, imm))
}
// riscvLiteralBytes renders a 64-bit constant as the little-endian bytes the
// $i64 pool symbol holds.
func riscvLiteralBytes(v int64) []byte {
return []byte{byte(v), byte(v >> 8), byte(v >> 16), byte(v >> 24),
byte(v >> 32), byte(v >> 40), byte(v >> 48), byte(v >> 56)}
}
// RiscvLiteral is one pooled 64-bit constant: a MOV whose immediate sits
// beyond both the 32-bit span and the shift sequences loads its bits from a
// read-only data symbol named like the toolchain's $i64 pool.
type RiscvLiteral struct {
Name string
Data []byte
}
// riscvLiterals collects the pooled constants the MOV expansions refer to,
// deduplicated by name, in first-use order.
type riscvLiterals struct {
order []RiscvLiteral
seen map[string]bool
}
func (l *riscvLiterals) add(name string, data []byte) {
if l.seen == nil {
l.seen = map[string]bool{}
}
if !l.seen[name] {
l.seen[name] = true
l.order = append(l.order, RiscvLiteral{Name: name, Data: data})
}
}
func (l *riscvLiterals) list() []RiscvLiteral { return l.order }
// encodeRISCVItypeImmediate encodes an I-type arithmetic instruction, expanding
// large immediates for ADDI/ANDI/ORI/XORI into LUI+ADDIW+op (or two ADDIs for
// ADDI), matching the Go assembler.
@@ -1734,6 +1984,7 @@ func encodeRISCVJALR(instr *ast.Instr, fi riscvFrameInfo) ([]byte, error) {
// RVC form. It returns the compressed instruction word and true on success.
func tryCompressRVC(instr *ast.Instr, fi riscvFrameInfo) (uint16, bool) {
mnem := riscvCompressMnem(instr)
mnem = riscvNormalisePseudo(mnem)
ops := instr.Operands
// The immediate aliases fold onto their I-type mnemonics before
// compression: the toolchain compresses ADD $imm, rd as c.addi, exactly
+1 -1
View File
@@ -63,7 +63,7 @@ func riscvRegNum(name string) int {
return 24
case "X25", "S9":
return 25
case "X26", "S10":
case "X26", "S10", "CTXT":
return 26
case "X27", "S11", "g":
return 27
+152 -19
View File
@@ -33,7 +33,7 @@ func firstTextRISCV(t *testing.T, src string) *ast.Text {
// assembleRISCVHelper assembles one TEXT function and returns its code bytes.
func assembleRISCVHelper(t *testing.T, fn *ast.Text) []byte {
t.Helper()
code, _, _, _, _, err := assembleRISCV(fn)
code, _, _, _, _, _, err := assembleRISCV(fn)
if err != nil {
t.Fatalf("assemble: %v", err)
}
@@ -785,16 +785,140 @@ TEXT ·sys(SB), NOSPLIT, $0
}
}
func TestRISCV_MOV_sym_FP_error(t *testing.T) {
// MOV $sym(FP), rd should return an error (unsupported).
func TestRISCV_MOV_sym_FP(t *testing.T) {
// MOV $sym(FP), rd lowers to the frame-adjusted ADDI against SP: the
// toolchain's argframe spelling. A zero frame leaves the offset at the
// 8-byte link slot, compressed to C.ADDI4SPN.
fn := firstTextRISCV(t, `#include "textflag.h"
TEXT ·badfp(SB), NOSPLIT, $0
TEXT ·argfp(SB), NOSPLIT, $0
MOV $arg(FP), X10
RET
`)
_, _, _, _, _, err := assembleRISCV(fn)
if err == nil {
t.Error("expected error for MOV $arg(FP), got nil")
code, _, _, _, _, _, err := assembleRISCV(fn)
if err != nil {
t.Fatalf("assemble: %v", err)
}
// prologue (0: leaf, zero frame) + C.ADDI4SPN (2) + RET (4) = 6
want := []byte{0x28, 0x00, 0x67, 0x80, 0x00, 0x00}
if string(code) != string(want) {
t.Errorf("got % x, want % x", code, want)
}
}
func TestRISCV_Bookkeeping(t *testing.T) {
// FUNCDATA and PCDATA contribute no bytes; UNDEF is the toolchain's
// ebreak, compressed to C.EBREAK under RVC.
fn := firstTextRISCV(t, `#include "textflag.h"
TEXT ·book(SB), NOSPLIT, $0-8
FUNCDATA $0, marks<>(SB)
PCDATA $1, $1
UNDEF
MOV $1, X10
MOV X10, ret+0(FP)
RET
`)
code, _, _, _, _, _, err := assembleRISCV(fn)
if err != nil {
t.Fatalf("assemble: %v", err)
}
// C.EBREAK (2) + C.LI X10, 1 (2) + C.SWSP (2) + RET (4) = 10: the
// FUNCDATA and PCDATA statements contribute nothing.
want := []byte{0x02, 0x90, 0x05, 0x45, 0x2a, 0xe4, 0x67, 0x80, 0x00, 0x00}
if string(code) != string(want) {
t.Errorf("got % x, want % x", code, want)
}
}
func TestRISCV_JMPPCRel(t *testing.T) {
// JMP N(PC): the displacement tracks the instruction N source slots
// away in the final layout (0 the jump itself, negative backwards).
fn := firstTextRISCV(t, `#include "textflag.h"
TEXT ·slots(SB), NOSPLIT, $0-0
JMP 2(PC)
MOV $1, X11
MOV $2, X12
MOV X12, X11
JMP -3(PC)
RET
`)
code, _, _, _, _, _, err := assembleRISCV(fn)
if err != nil {
t.Fatalf("assemble: %v", err)
}
// JMP 2(PC) lands on the C.MV six bytes ahead; JMP -3(PC) lands back on
// the first C.LI, six bytes behind.
want := []byte{
0x6f, 0x00, 0x60, 0x00, // JAL X0, 6
0x85, 0x45, // C.LI X11, 1
0x09, 0x46, // C.LI X12, 2
0xb2, 0x85, // C.MV X11, X12
0x6f, 0xf0, 0xbf, 0xff, // JAL X0, -6
0x67, 0x80, 0x00, 0x00, // RET
}
if string(code) != string(want) {
t.Errorf("got % x, want % x", code, want)
}
}
func TestRISCV_MOVWideImm(t *testing.T) {
// Shift-sequence constants compress like the toolchain's expansion.
fn := firstTextRISCV(t, `#include "textflag.h"
TEXT ·wide(SB), NOSPLIT, $0-0
MOV $0x8000000000000000, X5
MOV $0x100000000, X5
MOV $0x000fffffffffffda, X5
RET
`)
code, _, _, _, _, _, err := assembleRISCV(fn)
if err != nil {
t.Fatalf("assemble: %v", err)
}
// C.LI -1, C.SLLI 63; C.LI 1, C.SLLI 32; C.LI -19, C.SLLI 13, SRLI 12.
want := []byte{
0xfd, 0x52, 0xfe, 0x12,
0x85, 0x42, 0x82, 0x12,
0xb5, 0x52, 0xb6, 0x02, 0x93, 0xd2, 0xc2, 0x00,
0x67, 0x80, 0x00, 0x00,
}
if string(code) != string(want) {
t.Errorf("got % x, want % x", code, want)
}
}
func TestRISCV_MOVImmPool(t *testing.T) {
// A constant outside the shift shapes loads from the pooled $i64 data
// symbol via AUIPC+LD, named like the toolchain's pool.
src := `#include "textflag.h"
TEXT ·pool(SB), NOSPLIT, $0-8
MOV $0x0101010101010101, X16
MOV X16, ret+0(FP)
RET
`
f, errs := parser.Parse("pool_riscv64.s", src)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFileRISCV(f)
if err != nil {
t.Fatalf("AssembleFileRISCV: %v", err)
}
// AUIPC X16, 0 + LD X16, 0(X16): the relocation pair carries the symbol.
wantCode := []byte{0x17, 0x08, 0x00, 0x00, 0x03, 0x38, 0x08, 0x00}
if string(img.Code[0:8]) != string(wantCode) {
t.Errorf("pool load: got % x", img.Code[0:8])
}
var lit *DataSymbol
for i := range img.DataSyms {
if img.DataSyms[i].Name == "$i64.0101010101010101" {
lit = &img.DataSyms[i]
}
}
if lit == nil {
t.Fatalf("pool symbol missing: %v", img.DataSyms)
}
wantData := []byte{0x01, 0x01, 0x01, 0x01, 0x01, 0x01, 0x01, 0x01}
if string(img.Data[lit.Offset:lit.Offset+8]) != string(wantData) {
t.Errorf("pool bytes: got % x", img.Data[lit.Offset:lit.Offset+8])
}
}
@@ -805,7 +929,7 @@ TEXT ·calltest(SB), NOSPLIT, $0
CALL ext(SB)
RET
`)
code, _, relocs, _, _, err := assembleRISCV(fn)
code, _, relocs, _, _, _, err := assembleRISCV(fn)
if err != nil {
t.Fatalf("assemble: %v", err)
}
@@ -834,7 +958,7 @@ TEXT ·calllocal(SB), NOSPLIT, $0
sub:
RET
`)
_, _, _, _, _, err := assembleRISCV(fn)
_, _, _, _, _, _, err := assembleRISCV(fn)
if err == nil {
t.Error("expected error for CALL to local label, got nil")
}
@@ -868,7 +992,7 @@ func encodeOneInstrRISCV(t *testing.T, src string, pc int, offsets map[string]in
t.Helper()
fn := firstTextRISCV(t, "#include \"textflag.h\"\n"+src)
instr := fn.Body[0].(*ast.Instr)
return encodeRISCVInstr(instr, pc, offsets, riscvFrameInfo{}, nil, nil)
return encodeRISCVInstr(instr, pc, offsets, riscvFrameInfo{}, nil, nil, nil)
}
// TestRISCVBranchJumpRange checks that displacements beyond the B-type span
@@ -917,7 +1041,7 @@ func TestRISCVBranchFarBody(t *testing.T) {
}
sb.WriteString("done:\n\tRET\n")
fn := firstTextRISCV(t, sb.String())
out, _, _, _, _, err := assembleRISCV(fn)
out, _, _, _, _, _, err := assembleRISCV(fn)
if err != nil {
t.Fatalf("unexpected error: %v", err)
}
@@ -943,7 +1067,7 @@ TEXT ·csrhi(SB), NOSPLIT, $0
CSRRW $4096, X10, X11
RET
`)
if _, _, _, _, _, err := assembleRISCV(fn); err == nil {
if _, _, _, _, _, _, err := assembleRISCV(fn); err == nil {
t.Error("expected an out-of-range error for CSR $4096, got none")
}
fn = firstTextRISCV(t, `#include "textflag.h"
@@ -951,25 +1075,24 @@ TEXT ·csrmax(SB), NOSPLIT, $0
CSRRW $4095, X10, X11
RET
`)
if _, _, _, _, _, err := assembleRISCV(fn); err != nil {
if _, _, _, _, _, _, err := assembleRISCV(fn); err != nil {
t.Errorf("CSR $4095 must assemble: %v", err)
}
}
// TestRISCV_Imm64Rejected checks that immediates outside the signed 32-bit
// span are diagnosed instead of silently truncated to their low 32 bits (the
// toolchain materialises such constants via SLLI expansion, which this
// assembler does not implement).
// span are diagnosed instead of silently truncated to their low 32 bits for
// the I-type arithmetic; the MOV forms materialise the wide constant instead
// (shift sequence or pooled load), like the toolchain.
func TestRISCV_Imm64Rejected(t *testing.T) {
cases := []string{
"MOV $0x123456789, X10",
"ADDI $0x100000000, X10, X11",
"ANDI $-0x800000001, X10, X11",
"SUB $0x100000000, X10, X11",
}
for _, src := range cases {
fn := firstTextRISCV(t, "#include \"textflag.h\"\nTEXT ·wide(SB), NOSPLIT, $0\n\t"+src+"\n\tRET\n")
if _, _, _, _, _, err := assembleRISCV(fn); err == nil {
if _, _, _, _, _, _, err := assembleRISCV(fn); err == nil {
t.Errorf("%s: expected an out-of-range error, got none", src)
}
}
@@ -982,9 +1105,19 @@ TEXT ·edge(SB), NOSPLIT, $0
SUB $0x80000000, X12, X13
RET
`)
if _, _, _, _, _, err := assembleRISCV(fn); err != nil {
if _, _, _, _, _, _, err := assembleRISCV(fn); err != nil {
t.Errorf("int32-span immediates must assemble: %v", err)
}
// Beyond the span the MOV forms materialise the constant like the
// toolchain instead of diagnosing it.
fn = firstTextRISCV(t, `#include "textflag.h"
TEXT ·pool(SB), NOSPLIT, $0
MOV $0x123456789, X10
RET
`)
if _, _, _, _, _, _, err := assembleRISCV(fn); err != nil {
t.Errorf("MOV with a 64-bit immediate must assemble: %v", err)
}
}
// riscvWants decodes code as little-endian words and pins each one; the
+221 -7
View File
@@ -61,6 +61,21 @@ const (
// vexImmRMGPR is the immediate form over general-purpose registers
// (RORX): reg = dst, rm = src, imm8 = op0, L = 0.
vexImmRMGPR
// vexRMOpGPR is the two-operand /digit form over general-purpose
// registers (BLSI, BLSMSK, BLSR): ModRM.reg = /digit, ModRM.rm = src
// (op0), VEX.vvvv = dst (op1), L = 0.
vexRMOpGPR
// vexCountGPR is the three-operand count form over general-purpose
// registers (SHLX, SHRX, SARX, BEXTR, BZHI): the first operand rides
// VEX.vvvv and the second is r/m, the opposite pairing of the ANDN
// family, with reg = dst (op2), L = 0.
vexCountGPR
// vexExtractGPR is the lane-extract-to-GPR form `OP $imm, xsrc, GPR/mem
// dst`: ModRM.reg = xsrc (op1), ModRM.rm = destination (op2), imm8 =
// op0, the VPEXTRB/W/D/Q layout. EVEX only; the destination never
// carries a vector length, so the register the L'L field follows is the
// XMM source.
vexExtractGPR
)
// vexSpec describes one VEX instruction's encoding parameters.
@@ -230,8 +245,34 @@ var vexTable = map[string]vexSpec{
"ANDNQ": {2, 0xF2, 1, 0, -1, vexNDS3GPR},
"MULXL": {2, 0xF6, 0, 3, -1, vexNDS3GPR},
"MULXQ": {2, 0xF6, 1, 3, -1, vexNDS3GPR},
"RORXL": {3, 0xF0, 0, 3, -1, vexImmRMGPR},
"RORXQ": {3, 0xF0, 1, 3, -1, vexImmRMGPR},
// VEX.NDS.LZ.0F38, the BMI2 three-operand bit ops: BEXTR and BZHI
// share the F7/F5 opcodes across W, the variable shifts carry their
// direction in the prefix (SHLX 66, SHRX F2, SARX F3) and PDEP/PEXT
// in F2/F3.
"BEXTRL": {2, 0xF7, 0, 0, -1, vexCountGPR},
"BEXTRQ": {2, 0xF7, 1, 0, -1, vexCountGPR},
"BZHIL": {2, 0xF5, 0, 0, -1, vexCountGPR},
"BZHIQ": {2, 0xF5, 1, 0, -1, vexCountGPR},
"SARXL": {2, 0xF7, 0, 2, -1, vexCountGPR},
"SARXQ": {2, 0xF7, 1, 2, -1, vexCountGPR},
"SHLXL": {2, 0xF7, 0, 1, -1, vexCountGPR},
"SHLXQ": {2, 0xF7, 1, 1, -1, vexCountGPR},
"SHRXL": {2, 0xF7, 0, 3, -1, vexCountGPR},
"SHRXQ": {2, 0xF7, 1, 3, -1, vexCountGPR},
"PDEPL": {2, 0xF5, 0, 3, -1, vexNDS3GPR},
"PDEPQ": {2, 0xF5, 1, 3, -1, vexNDS3GPR},
"PEXTL": {2, 0xF5, 0, 2, -1, vexNDS3GPR},
"PEXTQ": {2, 0xF5, 1, 2, -1, vexNDS3GPR},
// VEX.LZ.0F38.W, the BMI1 unary bit ops (src, dst: ModRM.reg = /digit,
// rm = src, vvvv = dst).
"BLSIL": {2, 0xF3, 0, 0, 3, vexRMOpGPR},
"BLSIQ": {2, 0xF3, 1, 0, 3, vexRMOpGPR},
"BLSMSKL": {2, 0xF3, 0, 0, 2, vexRMOpGPR},
"BLSMSKQ": {2, 0xF3, 1, 0, 2, vexRMOpGPR},
"BLSRL": {2, 0xF3, 0, 0, 1, vexRMOpGPR},
"BLSRQ": {2, 0xF3, 1, 0, 1, vexRMOpGPR},
"RORXL": {3, 0xF0, 0, 3, -1, vexImmRMGPR},
"RORXQ": {3, 0xF0, 1, 3, -1, vexImmRMGPR},
// VEX.128.0F.W0, mask-register test (KTESTW k1, k2: reg = dst, rm = src).
"KTESTW": {1, 0x99, 0, 0, -1, vexRM},
@@ -290,6 +331,128 @@ var vexTable = map[string]vexSpec{
"VCVTPD2DQY": {1, 0xE6, 0, 3, -1, vexRMSrcLen},
"VCVTTPD2DQX": {1, 0xE6, 0, 1, -1, vexRMSrcLen},
"VCVTTPD2DQY": {1, 0xE6, 0, 1, -1, vexRMSrcLen},
// --- the VEX forms the avx512enc corpus exercises alongside the EVEX
// spellings, read off the toolchain opcode tables ---
"VAESDEC": {2, 0xDE, 0, 1, -1, vexNDS3},
"VAESDECLAST": {2, 0xDF, 0, 1, -1, vexNDS3},
"VAESENC": {2, 0xDC, 0, 1, -1, vexNDS3},
"VAESENCLAST": {2, 0xDD, 0, 1, -1, vexNDS3},
"VANDNPD": {1, 0x55, 0, 1, -1, vexNDS3},
"VANDPD": {1, 0x54, 0, 1, -1, vexNDS3},
"VCOMISD": {1, 0x2F, 0, 1, -1, vexRM},
"VCVTSD2SS": {1, 0x5A, 0, 3, -1, vexNDS3},
"VCVTSS2SD": {1, 0x5A, 0, 2, -1, vexNDS3},
"VFMADD132PD": {2, 0x98, 1, 1, -1, vexNDS3},
"VFMADD132PS": {2, 0x98, 0, 1, -1, vexNDS3},
"VFMADD132SD": {2, 0x99, 1, 1, -1, vexNDS3},
"VFMADD132SS": {2, 0x99, 0, 1, -1, vexNDS3},
"VFMADD213PD": {2, 0xA8, 1, 1, -1, vexNDS3},
"VFMADD213PS": {2, 0xA8, 0, 1, -1, vexNDS3},
"VFMADD213SS": {2, 0xA9, 0, 1, -1, vexNDS3},
"VFMADD231PS": {2, 0xB8, 0, 1, -1, vexNDS3},
"VFMADD231SD": {2, 0xB9, 1, 1, -1, vexNDS3},
"VFMADD231SS": {2, 0xB9, 0, 1, -1, vexNDS3},
"VFMADDSUB132PD": {2, 0x96, 1, 1, -1, vexNDS3},
"VFMADDSUB132PS": {2, 0x96, 0, 1, -1, vexNDS3},
"VFMADDSUB213PD": {2, 0xA6, 1, 1, -1, vexNDS3},
"VFMADDSUB213PS": {2, 0xA6, 0, 1, -1, vexNDS3},
"VFMADDSUB231PD": {2, 0xB6, 1, 1, -1, vexNDS3},
"VFMADDSUB231PS": {2, 0xB6, 0, 1, -1, vexNDS3},
"VFMSUB132PD": {2, 0x9A, 1, 1, -1, vexNDS3},
"VFMSUB132PS": {2, 0x9A, 0, 1, -1, vexNDS3},
"VFMSUB132SD": {2, 0x9B, 1, 1, -1, vexNDS3},
"VFMSUB132SS": {2, 0x9B, 0, 1, -1, vexNDS3},
"VFMSUB213PD": {2, 0xAA, 1, 1, -1, vexNDS3},
"VFMSUB213PS": {2, 0xAA, 0, 1, -1, vexNDS3},
"VFMSUB213SD": {2, 0xAB, 1, 1, -1, vexNDS3},
"VFMSUB213SS": {2, 0xAB, 0, 1, -1, vexNDS3},
"VFMSUB231PD": {2, 0xBA, 1, 1, -1, vexNDS3},
"VFMSUB231PS": {2, 0xBA, 0, 1, -1, vexNDS3},
"VFMSUB231SD": {2, 0xBB, 1, 1, -1, vexNDS3},
"VFMSUB231SS": {2, 0xBB, 0, 1, -1, vexNDS3},
"VFMSUBADD132PD": {2, 0x97, 1, 1, -1, vexNDS3},
"VFMSUBADD132PS": {2, 0x97, 0, 1, -1, vexNDS3},
"VFMSUBADD213PD": {2, 0xA7, 1, 1, -1, vexNDS3},
"VFMSUBADD213PS": {2, 0xA7, 0, 1, -1, vexNDS3},
"VFMSUBADD231PD": {2, 0xB7, 1, 1, -1, vexNDS3},
"VFMSUBADD231PS": {2, 0xB7, 0, 1, -1, vexNDS3},
"VFNMADD132PD": {2, 0x9C, 1, 1, -1, vexNDS3},
"VFNMADD132PS": {2, 0x9C, 0, 1, -1, vexNDS3},
"VFNMADD132SD": {2, 0x9D, 1, 1, -1, vexNDS3},
"VFNMADD132SS": {2, 0x9D, 0, 1, -1, vexNDS3},
"VFNMADD213PD": {2, 0xAC, 1, 1, -1, vexNDS3},
"VFNMADD213PS": {2, 0xAC, 0, 1, -1, vexNDS3},
"VFNMADD213SD": {2, 0xAD, 1, 1, -1, vexNDS3},
"VFNMADD213SS": {2, 0xAD, 0, 1, -1, vexNDS3},
"VFNMADD231PD": {2, 0xBC, 1, 1, -1, vexNDS3},
"VFNMADD231PS": {2, 0xBC, 0, 1, -1, vexNDS3},
"VFNMADD231SS": {2, 0xBD, 0, 1, -1, vexNDS3},
"VFNMSUB132PD": {2, 0x9E, 1, 1, -1, vexNDS3},
"VFNMSUB132PS": {2, 0x9E, 0, 1, -1, vexNDS3},
"VFNMSUB132SD": {2, 0x9F, 1, 1, -1, vexNDS3},
"VFNMSUB132SS": {2, 0x9F, 0, 1, -1, vexNDS3},
"VFNMSUB213PD": {2, 0xAE, 1, 1, -1, vexNDS3},
"VFNMSUB213PS": {2, 0xAE, 0, 1, -1, vexNDS3},
"VFNMSUB213SD": {2, 0xAF, 1, 1, -1, vexNDS3},
"VFNMSUB213SS": {2, 0xAF, 0, 1, -1, vexNDS3},
"VFNMSUB231PD": {2, 0xBE, 1, 1, -1, vexNDS3},
"VFNMSUB231PS": {2, 0xBE, 0, 1, -1, vexNDS3},
"VFNMSUB231SD": {2, 0xBF, 1, 1, -1, vexNDS3},
"VFNMSUB231SS": {2, 0xBF, 0, 1, -1, vexNDS3},
"VGF2P8AFFINEINVQB": {3, 0xCF, 1, 1, -1, vexNDS3Imm},
"VGF2P8MULB": {2, 0xCF, 0, 1, -1, vexNDS3},
"VMOVNTDQA": {2, 0x2A, 0, 1, -1, vexRM},
"VMOVNTPD": {1, 0x2B, 0, 1, -1, vexRMRev},
"VORPD": {1, 0x56, 0, 1, -1, vexNDS3},
"VPADDSB": {1, 0xEC, 0, 1, -1, vexNDS3},
"VPADDSW": {1, 0xED, 0, 1, -1, vexNDS3},
"VPADDUSB": {1, 0xDC, 0, 1, -1, vexNDS3},
"VPADDUSW": {1, 0xDD, 0, 1, -1, vexNDS3},
"VPCMPEQQ": {2, 0x29, 0, 1, -1, vexNDS3},
"VPCMPEQW": {1, 0x75, 0, 1, -1, vexNDS3},
"VPCMPGTB": {1, 0x64, 0, 1, -1, vexNDS3},
"VPCMPGTD": {1, 0x66, 0, 1, -1, vexNDS3},
"VPCMPGTW": {1, 0x65, 0, 1, -1, vexNDS3},
"VPERMPS": {2, 0x16, 0, 1, -1, vexNDS3},
"VPEXTRB": {3, 0x14, 0, 1, -1, vexExtract},
"VPEXTRD": {3, 0x16, 0, 1, -1, vexExtract},
"VPEXTRQ": {3, 0x16, 1, 1, -1, vexExtract},
"VPINSRD": {3, 0x22, 0, 1, -1, vexNDS3Imm},
"VPINSRQ": {3, 0x22, 1, 1, -1, vexNDS3Imm},
"VPMULHRSW": {2, 0x0B, 0, 1, -1, vexNDS3},
"VPMULHW": {1, 0xE5, 0, 1, -1, vexNDS3},
"VPMULUDQ": {1, 0xF4, 0, 1, -1, vexNDS3},
"VPSADBW": {1, 0xF6, 0, 1, -1, vexNDS3},
"VPSUBSB": {1, 0xE8, 0, 1, -1, vexNDS3},
"VPSUBSW": {1, 0xE9, 0, 1, -1, vexNDS3},
"VPSUBUSB": {1, 0xD8, 0, 1, -1, vexNDS3},
"VPSUBUSW": {1, 0xD9, 0, 1, -1, vexNDS3},
"VPUNPCKHBW": {1, 0x68, 0, 1, -1, vexNDS3},
"VPUNPCKHQDQ": {1, 0x6D, 0, 1, -1, vexNDS3},
"VPUNPCKHWD": {1, 0x69, 0, 1, -1, vexNDS3},
"VPUNPCKLBW": {1, 0x60, 0, 1, -1, vexNDS3},
"VPUNPCKLWD": {1, 0x61, 0, 1, -1, vexNDS3},
"VSQRTPD": {1, 0x51, 0, 1, -1, vexRM},
"VSQRTSD": {1, 0x51, 0, 3, -1, vexNDS3},
"VSQRTSS": {1, 0x51, 0, 2, -1, vexNDS3},
"VUCOMISD": {1, 0x2E, 0, 1, -1, vexRM},
// VEX.0F.WIG, the plain-prefix single/double arithmetic and unpack
// spellings (no 66 prefix; WIG, so W = 0).
"VANDNPS": {1, 0x55, 0, 0, -1, vexNDS3},
"VANDPS": {1, 0x54, 0, 0, -1, vexNDS3},
"VORPS": {1, 0x56, 0, 0, -1, vexNDS3},
"VUNPCKLPS": {1, 0x14, 0, 0, -1, vexNDS3},
"VUNPCKHPS": {1, 0x15, 0, 0, -1, vexNDS3},
"VSQRTPS": {1, 0x51, 0, 0, -1, vexRM},
"VMOVNTPS": {1, 0x2B, 0, 0, -1, vexRMRev},
// VEX.128.66.0F, the scalar and packed compare forms.
"VCOMISS": {1, 0x2F, 0, 1, -1, vexRM},
"VUCOMISS": {1, 0x2E, 0, 0, -1, vexRM},
// VEX.128.0F.F3/F2.W0, the high/low word shuffles ($imm, src, dst).
"VPSHUFHW": {1, 0x70, 0, 2, -1, vexImmRM},
"VPSHUFLW": {1, 0x70, 0, 3, -1, vexImmRM},
}
// vexSrcLen maps a source-length conversion mnemonic (the X/Y spellings of
@@ -420,6 +583,10 @@ func (e *enc) encodeVex(mnemUpper string, ops []Operand) error {
return e.encodeVexNDS3GPR(spec, ops)
case vexImmRMGPR:
return e.encodeVexImmRMGPR(spec, ops)
case vexRMOpGPR:
return e.encodeVexRMOpGPR(spec, ops)
case vexCountGPR:
return e.encodeVexCountGPR(spec, ops)
case vexRMRev:
return e.encodeVexRMRev(spec, ops)
}
@@ -523,9 +690,10 @@ func (e *enc) encodeVexShiftImm(spec vexSpec, ops []Operand) error {
if !ok {
return fmt.Errorf("shift count must be an immediate")
}
srcReg, ok := src.(Reg)
if !ok || !srcReg.isVec() {
return fmt.Errorf("shift source must be a vector register")
// The count source is a vector register or memory; the VEX length
// follows the destination register either way.
if !vecOrMem(src) {
return fmt.Errorf("shift source must be a vector register or memory")
}
dstReg, ok := dst.(Reg)
if !ok || !dstReg.isVec() {
@@ -533,7 +701,7 @@ func (e *enc) encodeVexShiftImm(spec vexSpec, ops []Operand) error {
}
vvvvBar := 15 - (dstReg.idx & 15)
if err := e.emitVexFields(spec, dstReg.vecLenBit(), spec.opdigit, 0, vvvvBar, srcReg); err != nil {
if err := e.emitVexFields(spec, dstReg.vecLenBit(), spec.opdigit, 0, vvvvBar, src); err != nil {
return err
}
immByte, err := imm8(int64(immVal))
@@ -700,7 +868,11 @@ func (e *enc) encodeVexNDS3GPR(spec vexSpec, ops []Operand) error {
if !ok || vvvvReg.isVec() {
return fmt.Errorf("VEX vvvv operand must be a general-purpose register")
}
return e.emitVexFields(spec, 0, dstReg.idx&7, 0, 15-(vvvvReg.idx&15), src2)
rBit := 0
if dstReg.idx >= 8 {
rBit = 1
}
return e.emitVexFields(spec, 0, dstReg.idx&7, rBit, 15-(vvvvReg.idx&15), src2)
}
// encodeVexImmRMGPR encodes the immediate form over general-purpose
@@ -729,6 +901,48 @@ func (e *enc) encodeVexImmRMGPR(spec vexSpec, ops []Operand) error {
return nil
}
// encodeVexRMOpGPR encodes the two-operand /digit form over general-purpose
// registers (BLSI, BLSMSK, BLSR): OP src, dst with ModRM.reg = /digit,
// ModRM.rm = src and VEX.vvvv = dst.
func (e *enc) encodeVexRMOpGPR(spec vexSpec, ops []Operand) error {
if len(ops) != 2 {
return fmt.Errorf("instruction expects 2 operands (src, dst), got %d", len(ops))
}
src, dst := ops[0], ops[1]
dstReg, ok := dst.(Reg)
if !ok || dstReg.isVec() {
return fmt.Errorf("VEX destination must be a general-purpose register")
}
return e.emitVexFields(spec, 0, spec.opdigit, 0, 15-(dstReg.idx&15), src)
}
// encodeVexCountGPR encodes the three-operand count form over general-purpose
// registers (SHLX, SHRX, SARX, BEXTR, BZHI): OP src, count, dst with
// VEX.vvvv = src (op0), ModRM.rm = count (op1), ModRM.reg = dst (op2).
func (e *enc) encodeVexCountGPR(spec vexSpec, ops []Operand) error {
if len(ops) != 3 {
return fmt.Errorf("VEX count instruction expects 3 operands, got %d", len(ops))
}
src, count, dst := ops[0], ops[1], ops[2]
dstReg, ok := dst.(Reg)
if !ok || dstReg.isVec() {
return fmt.Errorf("VEX destination must be a general-purpose register")
}
countReg, ok := count.(Reg)
if !ok || countReg.isVec() {
return fmt.Errorf("VEX count operand must be a general-purpose register")
}
srcReg, ok := src.(Reg)
if !ok || srcReg.isVec() {
return fmt.Errorf("VEX count source must be a general-purpose register")
}
rBit := 0
if dstReg.idx >= 8 {
rBit = 1
}
return e.emitVexFields(spec, 0, dstReg.idx&7, rBit, 15-(srcReg.idx&15), count)
}
// encodeVexRMRev encodes the reversed two-operand form: OP src, dst with the
// vector source in ModRM.reg and the memory destination in r/m (VMOVNTDQ,
// a store with no register-destination form).
+57
View File
@@ -31,6 +31,51 @@ var x86asmUnrecognised = map[string]bool{
"RORXQ": true,
"VFMADD213SD": true,
"VFNMADD231SD": true,
// The scalar FMA spellings the decoder's tables lack entirely.
"VFMADD132SD": true,
"VFMADD132SS": true,
"VFMADD213SS": true,
"VFMADD231SD": true,
"VFMADD231SS": true,
"VFMSUB132SD": true,
"VFMSUB132SS": true,
"VFMSUB213SD": true,
"VFMSUB213SS": true,
"VFMSUB231SD": true,
"VFMSUB231SS": true,
"VFNMADD132SD": true,
"VFNMADD132SS": true,
"VFNMADD213SD": true,
"VFNMADD213SS": true,
"VFNMADD231SS": true,
"VFNMSUB132SD": true,
"VFNMSUB132SS": true,
"VFNMSUB213SD": true,
"VFNMSUB213SS": true,
"VFNMSUB231SD": true,
"VFNMSUB231SS": true,
// The BMI1 unary bit ops the decoder's AVX tables lack.
"BLSIL": true,
"BLSIQ": true,
"BLSMSKL": true,
"BLSMSKQ": true,
"BLSRL": true,
"BLSRQ": true,
// The BMI2 bit ops whose W1/LZ rows the decoder misses.
"BEXTRL": true,
"BEXTRQ": true,
"BZHIL": true,
"BZHIQ": true,
"PDEPL": true,
"PDEPQ": true,
"PEXTL": true,
"PEXTQ": true,
"SARXL": true,
"SARXQ": true,
"SHLXL": true,
"SHLXQ": true,
"SHRXL": true,
"SHRXQ": true,
}
// TestVexNDS3 encodes `mnem Y0, Y1, Y2` for every three-operand NDS
@@ -237,6 +282,18 @@ func TestVexGroundTruth(t *testing.T) {
{"MULXQ AX,BX,CX", "MULXQ", []Operand{AX, BX, CX}, "c4e2e3f6c8", ""},
{"RORXL $3,AX,CX", "RORXL", []Operand{Imm(3), AX, CX}, "c4e37bf0c803", ""},
{"RORXQ $3,AX,CX", "RORXQ", []Operand{Imm(3), AX, CX}, "c4e3fbf0c803", ""},
// BMI2 variable shifts and bit ops (three general registers).
{"SHLXL AX,CX,R15", "SHLXL", []Operand{AX, CX, vreg(t, "R15")}, "c46279f7f9", ""},
{"SHRXQ R8,DX,AX", "SHRXQ", []Operand{vreg(t, "R8"), DX, AX}, "c4e2bbf7c2", ""},
{"SARXQ AX,DX,R9", "SARXQ", []Operand{AX, DX, vreg(t, "R9")}, "c462faf7ca", ""},
{"BEXTRL AX,CX,R15", "BEXTRL", []Operand{AX, CX, vreg(t, "R15")}, "c46278f7f9", ""},
{"BZHIQ AX,CX,R15", "BZHIQ", []Operand{AX, CX, vreg(t, "R15")}, "c462f8f5f9", ""},
{"PDEPQ AX,CX,R15", "PDEPQ", []Operand{AX, CX, vreg(t, "R15")}, "c462f3f5f8", ""},
{"PEXTQ AX,CX,R15", "PEXTQ", []Operand{AX, CX, vreg(t, "R15")}, "c462f2f5f8", ""},
// BMI1 unary bit ops (src, dst: /digit in ModRM.reg, dst in vvvv).
{"BLSIL AX,CX", "BLSIL", []Operand{AX, CX}, "c4e270f3d8", ""},
{"BLSRQ AX,CX", "BLSRQ", []Operand{AX, CX}, "c4e2f0f3c8", ""},
{"BLSMSKQ AX,CX", "BLSMSKQ", []Operand{AX, CX}, "c4e2f0f3d0", ""},
// Two-operand reg/rm form (v̄vvv must be 1111).
{"VPMOVSXDQ X0,Y4", "VPMOVSXDQ", []Operand{vreg(t, "X0"), vreg(t, "Y4")}, "c4e27d25e0", ""},
{"VPMOVSXWD (SI),Y0", "VPMOVSXWD", []Operand{Ptr(SI, 0, 8), vreg(t, "Y0")}, "c4e27d2306", ""},
+17 -7
View File
@@ -151,11 +151,21 @@ type Immediate struct {
// Address is a non-immediate operand: a register, a memory reference, a symbol
// reference or a label. Fields are populated best-effort from the syntax.
type Address struct {
Sym *Symbol // name reference (bare ident, or name+off(pseudo))
Base string // base register, from (base)
Index string // index register, from (index*scale)
Scale int // index scale; 0 when absent
Offset int64 // leading displacement, from off(base)
HasOff bool // a leading displacement is present
Shift string // verbatim arm64 shift suffix, e.g. "<< 2"
Sym *Symbol // name reference (bare ident, or name+off(pseudo))
Base string // base register, from (base)
Index string // index register, from (index*scale)
Scale int // index scale; 0 when absent
Offset int64 // leading displacement, from off(base)
HasOff bool // a leading displacement is present
Shift string // verbatim arm64 shift suffix, e.g. "<< 2"
Range *RegRange // bracketed register range; nil for every other form
}
// RegRange is a bracketed register range, [Z0-Z3]: the amd64 spelling of
// the four-register source of the 4FMAPS/4VNNIW families. Lo and Hi carry
// the verbatim register spellings; the range is inclusive at both ends.
type RegRange struct {
Lo string
Hi string
Pos token.Position
}
+367
View File
@@ -0,0 +1,367 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
package main
import (
"errors"
"fmt"
"go/ast"
"go/build"
"go/constant"
"go/parser"
"go/token"
"go/types"
"os"
"path/filepath"
"regexp"
"strings"
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
)
// go_asm.h is the header the Go compiler writes for every package that
// carries assembly (the compiler's -asmhdr output): "#define const_NAME
// value" for each package constant, and for each named struct type
// "#define TYPE__size size" plus one "#define TYPE_field offset" per field.
// GOROOT assembly includes it, and a standalone assembler has no compiler
// to have produced it, so gasm generates the equivalent itself: the package
// the .s file lives in is parsed and type-checked here, with the target
// architecture's own sizes, and the same defines are written out. The
// type-checking GOOS is selected by the caller: a GOOS-specific file
// (sys_darwin_arm64.s) needs its platform's defines, which a header from
// the ambient GOOS silently omits.
//
// The emitter mirrors cmd/compile's dumpasmhdr exactly: constants come out
// as "const_NAME", struct entries as "NAME__size" followed by the fields in
// declaration order, blank names are skipped, and float and complex
// constants are omitted (the assembler carries integers, bools and strings
// only). Aliases to structs are emitted, generic types are not: they have
// no fixed size. A define the assembly references but this header does not
// carry surfaces later as the assembler's own "undefined" diagnostic naming
// the define, which is the honest failure.
// goAsmInclude matches the #include "go_asm.h" directive, tolerant of
// whitespace, so the wiring knows which files need a generated header
// before the preprocessor runs and would report the header as missing.
var goAsmInclude = regexp.MustCompile(`(?m)^\s*#\s*include\s+"go_asm\.h"`)
// needsGoAsmHeader reports whether src includes go_asm.h.
func needsGoAsmHeader(src string) bool {
return goAsmInclude.MatchString(src)
}
// goAsmHeaderResolved reports whether the include of go_asm.h from a file in
// asmDir already resolves: to a header in the package directory itself, or
// in one of the -I directories, the way the preprocessor searches. Only an
// unresolved include is generated for; a header someone placed by hand is
// the tool the author chose, and it also wins the preprocessor's own search
// order, so generating a second copy would be dead weight at best.
func goAsmHeaderResolved(asmDir string, dirs []string) bool {
candidates := []string{filepath.Join(asmDir, "go_asm.h")}
for _, d := range dirs {
candidates = append(candidates, filepath.Join(d, "go_asm.h"))
}
for _, candidate := range candidates {
if st, err := os.Stat(candidate); err == nil && !st.IsDir() {
return true
}
}
return false
}
// generateGoAsmHeader type-checks the Go package in pkgDir for goos and
// goarch, writes its go_asm.h equivalent into dir, and returns dir. An
// empty goos means the ambient one. The caller owns the directory and its
// removal.
func generateGoAsmHeader(pkgDir, goos, goarch, dir string) (string, error) {
if goos == "" {
goos = build.Default.GOOS
}
imp := newSourceImporter(goos, goarch)
if imp.sizes == nil {
return "", fmt.Errorf("go_asm.h: unknown GOARCH %q", goarch)
}
bp, err := imp.ctxt.ImportDir(pkgDir, 0)
if err != nil {
return "", fmt.Errorf("go_asm.h for GOARCH %s in %s: %w", goarch, pkgDir, err)
}
files, errs := imp.parse(bp)
if len(errs) > 0 {
return "", fmt.Errorf("go_asm.h for GOARCH %s in %s: %s", goarch, pkgDir, errorList(errs))
}
_, info, errs := imp.checkPackage(bp, files)
if len(errs) > 0 {
return "", fmt.Errorf("go_asm.h for GOARCH %s in %s: package does not type-check: %s", goarch, pkgDir, errorList(errs))
}
var b strings.Builder
fmt.Fprintf(&b, "// generated by gasm from package %s (GOOS %s, GOARCH %s)\n\n", bp.Name, goos, goarch)
// Files in the build's own order and declarations in source order: the
// same walk the compiler's reader makes, so the header reads the same
// way the toolchain's does. Order carries no meaning to the assembler
// (defines form a table), only to a human diffing against one.
for _, f := range files {
for _, decl := range f.Decls {
gd, ok := decl.(*ast.GenDecl)
if !ok {
continue
}
for _, spec := range gd.Specs {
switch gd.Tok {
case token.CONST:
vs, ok := spec.(*ast.ValueSpec)
if !ok {
continue
}
for _, name := range vs.Names {
emitConst(&b, info.Defs[name], name.Name)
}
case token.TYPE:
ts, ok := spec.(*ast.TypeSpec)
if !ok {
continue
}
emitStruct(&b, imp.sizes, info.Defs[ts.Name], ts.Name.Name)
}
}
}
}
if err := os.MkdirAll(dir, 0o755); err != nil {
return "", fmt.Errorf("go_asm.h for GOARCH %s in %s: %w", goarch, pkgDir, err)
}
out := filepath.Join(dir, "go_asm.h")
if err := os.WriteFile(out, []byte(b.String()), 0o644); err != nil {
return "", fmt.Errorf("go_asm.h for GOARCH %s in %s: %w", goarch, pkgDir, err)
}
return dir, nil
}
// emitConst writes one const define, skipping what the toolchain skips:
// blank names, and float and complex values the assembler has no syntax for.
func emitConst(b *strings.Builder, obj types.Object, name string) {
c, ok := obj.(*types.Const)
if !ok || name == "_" {
return
}
switch c.Val().Kind() {
case constant.Float, constant.Complex, constant.Unknown:
return
}
fmt.Fprintf(b, "#define const_%s %s\n", name, c.Val().ExactString())
}
// emitStruct writes one named struct type's size and field offsets,
// skipping what the toolchain skips: blank names, non-struct types, and
// generic types, whose size depends on their instantiation.
func emitStruct(b *strings.Builder, sizes types.Sizes, obj types.Object, name string) {
tn, ok := obj.(*types.TypeName)
if !ok || name == "_" {
return
}
t := types.Unalias(tn.Type())
// Generic types are spelled *types.Named with a type-parameter list;
// a plain struct type or an instantiated one carries none.
if named, ok := t.(*types.Named); ok && named.TypeParams().Len() > 0 {
return
}
st, ok := t.Underlying().(*types.Struct)
if !ok {
return
}
fmt.Fprintf(b, "#define %s__size %d\n", name, sizes.Sizeof(t))
fields := make([]*types.Var, st.NumFields())
for i := range st.NumFields() {
fields[i] = st.Field(i)
}
for i, off := range sizes.Offsetsof(fields) {
fld := fields[i]
if fld.Name() == "_" {
continue
}
fmt.Fprintf(b, "#define %s_%s %d\n", name, fld.Name(), off)
}
}
// errorList renders at most three errors, enough to say what is wrong
// without burying the diagnostic the caller actually reads.
func errorList(errs []error) string {
if len(errs) > 3 {
errs = errs[:3]
}
msgs := make([]string, len(errs))
for i, err := range errs {
msgs[i] = err.Error()
}
return strings.Join(msgs, "; ")
}
// sourceImporter type-checks imported packages from source with the target
// architecture's sizes. go/importer's "source" importer pins the host
// GOARCH, which would lay out imported types (internal/cpu, internal/abi)
// for the wrong target on a cross-architecture header, so the recursion is
// carried here with one build context and one sizes instance per
// architecture.
type sourceImporter struct {
fset *token.FileSet
ctxt *build.Context
sizes types.Sizes
pkgs map[string]*types.Package
}
// newSourceImporter returns the importer for one target GOOS and GOARCH.
// Cgo is disabled so the file set is deterministic and independent of the
// host's C toolchain: cgo-tagged files drop out of the build exactly as
// they do from a CGO_ENABLED=0 build, whose assembly is what gasm targets.
func newSourceImporter(goos, goarch string) *sourceImporter {
ctxt := new(build.Context)
*ctxt = build.Default
ctxt.GOOS = goos
ctxt.GOARCH = goarch
ctxt.CgoEnabled = false
return &sourceImporter{
fset: token.NewFileSet(),
ctxt: ctxt,
sizes: types.SizesFor("gc", goarch),
pkgs: map[string]*types.Package{},
}
}
// Import type-checks one imported package and memoises it. "unsafe" must
// resolve to go/types' own package, never to the source in GOROOT/src/unsafe:
// the source declares Sizeof and Offsetof as ordinary functions over
// ArbitraryType, and checking against that signature rejects half the
// unsafe arithmetic the gc compiler accepts, which is exactly the divergence
// srcimporter guards against the same way.
func (im *sourceImporter) Import(path string) (*types.Package, error) {
if path == "unsafe" {
return types.Unsafe, nil
}
if p, ok := im.pkgs[path]; ok {
return p, nil
}
bp, err := im.ctxt.Import(path, "", 0)
if err != nil {
return nil, err
}
files, errs := im.parse(bp)
if len(errs) > 0 {
return nil, errors.New(errorList(errs))
}
pkg, _, _ := im.checkPackage(bp, files)
im.pkgs[path] = pkg
return pkg, nil
}
// parse reads the build package's Go files. Import-level failures (no Go
// files for the target, unreadable files) come back as errors, and the
// type-check decides the rest.
func (im *sourceImporter) parse(bp *build.Package) ([]*ast.File, []error) {
if len(bp.GoFiles) == 0 {
return nil, []error{fmt.Errorf("no Go source files for GOOS=%s GOARCH=%s", im.ctxt.GOOS, im.ctxt.GOARCH)}
}
var (
files []*ast.File
errs []error
)
for _, name := range bp.GoFiles {
f, err := parser.ParseFile(im.fset, filepath.Join(bp.Dir, name), nil, parser.SkipObjectResolution)
if err != nil {
errs = append(errs, err)
continue
}
files = append(files, f)
}
return files, errs
}
// checkPackage type-checks one package's files with the importer's sizes,
// recording every error: a header from a package that does not type-check
// could silently mis-state an offset, so the caller refuses the header
// rather than trusting it. The returned Defs map backs the root package's
// emission walk; imports only need the checked package itself.
func (im *sourceImporter) checkPackage(bp *build.Package, files []*ast.File) (*types.Package, *types.Info, []error) {
var errs []error
conf := &types.Config{
Importer: im,
Sizes: im.sizes,
Error: func(err error) { errs = append(errs, err) },
}
info := &types.Info{Defs: map[*ast.Ident]types.Object{}}
pkg, _ := conf.Check(bp.ImportPath, im.fset, files, info)
return pkg, info, errs
}
// asmhdrCache generates one go_asm.h per package directory and target
// architecture under one temp root, for callers that assemble many files
// (the corpus audit). Failures are cached too: a package that does not
// type-check must not be re-checked once per file.
type asmhdrCache struct {
root string
dirs map[string]string // "pkgDir\x00goos\x00goarch" -> directory holding go_asm.h
errs map[string]error
}
func newAsmhdrCache() (*asmhdrCache, error) {
root, err := os.MkdirTemp("", "gasm-asmhdr")
if err != nil {
return nil, err
}
return &asmhdrCache{root: root, dirs: map[string]string{}, errs: map[string]error{}}, nil
}
// dirFor returns the directory holding the generated go_asm.h for pkgDir
// under goos and goarch, generating it on first use. An empty goos means
// the ambient one, resolved here so that one package cannot generate twice
// under an explicit and an implicit spelling of the same GOOS.
func (c *asmhdrCache) dirFor(pkgDir, goos, goarch string) (string, error) {
if goos == "" {
goos = build.Default.GOOS
}
key := pkgDir + "\x00" + goos + "\x00" + goarch
if dir, ok := c.dirs[key]; ok {
return dir, nil
}
if err, ok := c.errs[key]; ok {
return "", err
}
dir := filepath.Join(c.root, fmt.Sprintf("h%d_%s_%s", len(c.dirs), goos, goarch))
if _, err := generateGoAsmHeader(pkgDir, goos, goarch, dir); err != nil {
c.errs[key] = err
return "", err
}
c.dirs[key] = dir
return dir, nil
}
// close removes the temp root.
func (c *asmhdrCache) close() { os.RemoveAll(c.root) }
// ensureGoAsmHeader prepares the include directory a file that includes
// go_asm.h needs: the generated header for the package in path's directory,
// for the file's target GOOS and architecture. It reports a usage error
// when the architecture cannot be determined, and passes through the
// generator's diagnostics, which name the package.
func ensureGoAsmHeader(path string, target arch.Arch, goos string, cache *asmhdrCache) (string, func(), error) {
if path == "-" {
return "", nil, errors.New("cannot generate go_asm.h for standard input (no package directory)")
}
if target == arch.Unknown {
return "", nil, errors.New("a file that includes go_asm.h needs a target architecture: name the file _<arch>.s or pass -GOARCH")
}
if cache != nil {
dir, err := cache.dirFor(filepath.Dir(path), goos, goarchName(target))
return dir, func() {}, err
}
root, err := os.MkdirTemp("", "gasm-asmhdr")
if err != nil {
return "", nil, err
}
dir, err := generateGoAsmHeader(filepath.Dir(path), goos, goarchName(target), root)
if err != nil {
os.RemoveAll(root)
return "", nil, err
}
return dir, func() { os.RemoveAll(root) }, nil
}
+430
View File
@@ -0,0 +1,430 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
package main
import (
"os"
"os/exec"
"path/filepath"
"strings"
"testing"
)
// writePkg lays out a minimal Go package in a temp directory.
func writePkg(t *testing.T, files map[string]string) string {
t.Helper()
dir := t.TempDir()
for name, src := range files {
if err := os.WriteFile(filepath.Join(dir, name), []byte(src), 0o644); err != nil {
t.Fatal(err)
}
}
return dir
}
// generateFor generates the header for dir and returns its text. An empty
// goos means the ambient one.
func generateFor(t *testing.T, dir, goos, goarch string) string {
t.Helper()
hdrDir, err := generateGoAsmHeader(dir, goos, goarch, t.TempDir())
if err != nil {
t.Fatalf("generateGoAsmHeader(%q, %s, %s): %v", dir, goos, goarch, err)
}
b, err := os.ReadFile(filepath.Join(hdrDir, "go_asm.h"))
if err != nil {
t.Fatal(err)
}
return string(b)
}
func TestGenerateGoAsmHeaderShape(t *testing.T) {
dir := writePkg(t, map[string]string{"sample.go": `package sample
const bufSize = 1024
const (
a = iota * 8
b
c
)
const (
strConst = "hello"
boolConst = true
floatConst = 1.5
_ = "the blank identifier is skipped"
)
const shift = 1 << 20
type reader struct {
r int64
w int64
_ [4]byte
name string
}
type scalar int
type aliased struct {
k uint32
v uint32
}
type alias = aliased
`})
hdr := generateFor(t, dir, "", "amd64")
want := []string{
"#define const_bufSize 1024",
// iota resolves through go/types, one define per name.
"#define const_a 0",
"#define const_b 8",
"#define const_c 16",
`#define const_strConst "hello"`,
"#define const_boolConst true",
// Floats are the toolchain's own skip, as are blank names.
"#define const_shift 1048576",
// The blank field still occupies its bytes: the pad after w runs to
// the string's 8-byte alignment.
"#define reader__size 40",
"#define reader_r 0",
"#define reader_w 8",
"#define reader_name 24",
// Non-struct named types carry no defines; aliases to structs do.
"#define aliased__size 8",
"#define aliased_k 0",
"#define aliased_v 4",
"#define alias__size 8",
"#define alias_k 0",
"#define alias_v 4",
}
for _, w := range want {
if !strings.Contains(hdr, w+"\n") {
t.Errorf("header misses %q\ngot:\n%s", w, hdr)
}
}
for _, banned := range []string{"#define const_floatConst", "#define _ ", "#define scalar"} {
if strings.Contains(hdr, banned) {
t.Errorf("header must not carry %s\ngot:\n%s", banned, hdr)
}
}
}
func TestGenerateGoAsmHeaderPerArch(t *testing.T) {
dir := writePkg(t, map[string]string{
"common.go": `package perarch
type layout struct {
a int32
p uintptr
}
`,
// The build-tagged file set is part of the contract: a per-arch
// package is exactly how internal/cpu declares its layouts.
"const_amd64.go": `//go:build amd64
package perarch
const flavour = 1
`,
"const_arm64.go": `//go:build arm64
package perarch
const flavour = 2
`,
})
amd64 := generateFor(t, dir, "", "amd64")
arm64 := generateFor(t, dir, "", "arm64")
if !strings.Contains(amd64, "#define const_flavour 1\n") {
t.Errorf("amd64 header misses const_flavour 1:\n%s", amd64)
}
if !strings.Contains(arm64, "#define const_flavour 2\n") {
t.Errorf("arm64 header misses const_flavour 2:\n%s", arm64)
}
if strings.Contains(arm64, "#define const_flavour 1\n") {
t.Errorf("arm64 header must not carry the amd64 file's value")
}
// SizesFor makes the layout the target's: uintptr is 4 bytes wide on
// 386 and 8 on amd64, which must move p and grow the struct.
if !strings.Contains(amd64, "#define layout__size 16\n") || !strings.Contains(amd64, "#define layout_p 8\n") {
t.Errorf("amd64 layout wrong:\n%s", amd64)
}
w386 := generateFor(t, dir, "", "386")
if !strings.Contains(w386, "#define layout__size 8\n") || !strings.Contains(w386, "#define layout_p 4\n") {
t.Errorf("386 layout wrong:\n%s", w386)
}
}
// TestGenerateGoAsmHeaderGOOS pins the GOOS half of the target: only the
// platform's own files type-check into the header, which is why
// sys_darwin_arm64.s cannot assemble against a linux-generated one.
func TestGenerateGoAsmHeaderGOOS(t *testing.T) {
dir := writePkg(t, map[string]string{
"common.go": `package goosaware
type shared struct {
a int32
}
`,
"plat_darwin.go": `//go:build darwin
package goosaware
type platform struct {
trampoline_numer int64
}
`,
"plat_windows.go": `//go:build windows
package goosaware
type platform struct {
callbackArgs__size int32
}
`,
})
darwin := generateFor(t, dir, "darwin", "arm64")
if !strings.Contains(darwin, "#define platform__size 8\n") || !strings.Contains(darwin, "#define platform_trampoline_numer 0\n") {
t.Errorf("darwin header misses the darwin layout:\n%s", darwin)
}
if strings.Contains(darwin, "callbackArgs") {
t.Errorf("darwin header must not carry the windows layout:\n%s", darwin)
}
windows := generateFor(t, dir, "windows", "arm64")
if !strings.Contains(windows, "#define platform_callbackArgs__size 0\n") {
t.Errorf("windows header misses the windows layout:\n%s", windows)
}
if strings.Contains(windows, "trampoline_numer") {
t.Errorf("windows header must not carry the darwin layout:\n%s", windows)
}
// The ambient GOOS is neither of the two, so only shared's defines are
// emitted; the shared type keeps its layout there.
ambient := generateFor(t, dir, "", "arm64")
if !strings.Contains(ambient, "#define shared__size 4\n") {
t.Errorf("ambient header misses the shared layout:\n%s", ambient)
}
if strings.Contains(ambient, "#define platform_") {
t.Errorf("ambient header must not carry either platform layout:\n%s", ambient)
}
}
func TestGoosFromFilename(t *testing.T) {
for path, want := range map[string]string{
"/x/sys_darwin_arm64.s": "darwin",
"/x/sys_windows_arm64.s": "windows",
"/x/asm_linux_amd64.s": "linux",
"/x/rt0_darwin_arm64.s": "darwin",
"/x/vgetrandom_zos_s390x.s": "zos",
"/x/rt0_js_wasm.s": "js",
"/x/memmove_amd64.s": "",
"/x/vlop_arm.s": "",
"/x/stubs.s": "",
} {
if got := goosFromFilename(path); got != want {
t.Errorf("goosFromFilename(%q) = %q, want %q", path, got, want)
}
}
}
func TestGenerateGoAsmHeaderErrors(t *testing.T) {
t.Run("type error", func(t *testing.T) {
dir := writePkg(t, map[string]string{"bad.go": `package bad
const x = undefinedIdent
`})
_, err := generateGoAsmHeader(dir, "", "amd64", t.TempDir())
if err == nil {
t.Fatal("generation must fail for a package that does not type-check")
}
if !strings.Contains(err.Error(), dir) {
t.Errorf("error must name the package directory: %v", err)
}
if !strings.Contains(err.Error(), "type-check") {
t.Errorf("error must say the package does not type-check: %v", err)
}
})
t.Run("no go files", func(t *testing.T) {
dir := t.TempDir()
_, err := generateGoAsmHeader(dir, "", "amd64", t.TempDir())
if err == nil {
t.Fatal("generation must fail without Go files")
}
if !strings.Contains(err.Error(), dir) {
t.Errorf("error must name the package directory: %v", err)
}
})
}
func TestNeedsGoAsmHeader(t *testing.T) {
yes := "#include \"go_asm.h\"\n#include \"textflag.h\"\n"
no := "#include \"textflag.h\"\n#include \"funcdata.h\"\n"
if !needsGoAsmHeader(yes) {
t.Error("needsGoAsmHeader(missing on a go_asm.h include)")
}
if needsGoAsmHeader(no) {
t.Error("needsGoAsmHeader claims other headers need generation")
}
}
func TestGoAsmHeaderResolved(t *testing.T) {
dir := t.TempDir()
if goAsmHeaderResolved(dir, nil) {
t.Error("resolved with no header anywhere")
}
other := t.TempDir()
if goAsmHeaderResolved(dir, []string{other}) {
t.Error("resolved with an empty -I directory")
}
if err := os.WriteFile(filepath.Join(dir, "go_asm.h"), nil, 0o644); err != nil {
t.Fatal(err)
}
if !goAsmHeaderResolved(dir, nil) {
t.Error("not resolved with the header in the package directory")
}
}
func TestOtherGOOSFile(t *testing.T) {
for path, want := range map[string]bool{
"/x/sys_windows_amd64.s": true,
"/x/rt0_js_wasm.s": true,
"/x/sys_darwin_arm64.s": true,
"/x/sys_linux_amd64.s": false,
"/x/time_linux_amd64.s": false,
"/x/memmove_amd64.s": false,
"/x/generic.s": false,
} {
if got := otherGOOSFile(path); got != want {
t.Errorf("otherGOOSFile(%q) = %v, want %v", path, got, want)
}
}
}
// TestRunCorpusAuditGoAsm covers the audit wiring end to end: a package
// beside its kernel, the kernel living off the generated defines, and the
// histogram recording a generation failure as its own reason.
func TestRunCorpusAuditGoAsm(t *testing.T) {
dir := t.TempDir()
write := func(name, src string) {
t.Helper()
if err := os.WriteFile(filepath.Join(dir, name), []byte(src), 0o644); err != nil {
t.Fatal(err)
}
}
write("pkg.go", `package corpus
const pageSize = 4096
type header struct {
magic uint64
flags uint64
}
`)
write("kern_amd64.s", "#include \"go_asm.h\"\nTEXT \xc2\xb7f(SB), NOSPLIT, $0-16\n\tMOVQ\t$const_pageSize, AX\n\tMOVQ\t$header__size, BX\n\tRET\n")
// The defines live in the file's own package; a kernel in a directory
// without Go files has no package to generate from.
if err := os.MkdirAll(filepath.Join(dir, "sub"), 0o755); err != nil {
t.Fatal(err)
}
write(filepath.Join("sub", "lonely_arm64.s"), "#include \"go_asm.h\"\nTEXT \xc2\xb7g(SB), NOSPLIT, $0-0\n\tRET\n")
stats, err := runCorpusAudit(dir, nil)
if err != nil {
t.Fatalf("runCorpusAudit: %v", err)
}
get := func(name string) *corpusTally {
for i, tg := range stats.targets {
if tg.name == name {
return stats.tallies[i]
}
}
t.Fatalf("no tally for %s", name)
return nil
}
if a := get("amd64"); a.attempted != 1 || a.assembled != 1 {
t.Errorf("amd64 = %d/%d, want 1/1", a.assembled, a.attempted)
}
// lonely_arm64.s is an arm64 file whose package cannot be generated.
if a := get("arm64"); a.attempted != 1 || a.assembled != 0 {
t.Errorf("arm64 = %d/%d, want 0/1", a.assembled, a.attempted)
}
if r := get("arm64").reasons["go_asm.h generation failed"]; r != 1 {
t.Errorf("arm64 go_asm.h failure count = %d, want 1", r)
}
}
// TestRunCorpusAuditGOOS covers the filename-derived GOOS end to end: a
// kernel whose name names darwin must have its header type-checked with
// GOOS=darwin, so the darwin-only constant it offsets with is defined. The
// operand mirrors sys_darwin_arm64.s's trampoline, where a missing define
// leaves an unexpanded symbol in the offset and fails.
func TestRunCorpusAuditGOOS(t *testing.T) {
dir := t.TempDir()
write := func(name, src string) {
t.Helper()
if err := os.WriteFile(filepath.Join(dir, name), []byte(src), 0o644); err != nil {
t.Fatal(err)
}
}
write("pkg.go", "package corpus\n")
write("plat_darwin.go", "//go:build darwin\n\npackage corpus\n\nconst trampolineNumer = 8\n")
write("kern_darwin_arm64.s", "#include \"go_asm.h\"\n"+
"GLOBL timebase<>(SB), NOPTR, $16\n"+
"TEXT \xc2\xb7g(SB), NOSPLIT, $0-0\n"+
"\tMOVD\ttimebase<>+const_trampolineNumer(SB), R0\n"+
"\tRET\n")
stats, err := runCorpusAudit(dir, nil)
if err != nil {
t.Fatalf("runCorpusAudit: %v", err)
}
var arm *corpusTally
for i, tg := range stats.targets {
if tg.name == "arm64" {
arm = stats.tallies[i]
}
}
if arm == nil {
t.Fatal("no arm64 tally")
}
if arm.attempted != 1 || arm.assembled != 1 {
t.Errorf("arm64 = %d/%d, want 1/1; reasons: %v", arm.assembled, arm.attempted, arm.reasons)
}
}
// TestGenerateGoAsmHeaderRuntime pins the generator against the real thing:
// the runtime package of the ambient toolchain, whose header the toolchain's
// own -asmhdr output was sampled from. Skipped in short mode: it type-checks
// the whole package. The GOROOT comes from the go command itself, so the
// test follows whatever toolchain the host provides.
func TestGenerateGoAsmHeaderRuntime(t *testing.T) {
if testing.Short() {
t.Skip("type-checks the whole runtime package")
}
out, err := exec.Command("go", "env", "GOROOT").Output()
if err != nil {
t.Skipf("no Go toolchain: %v", err)
}
runtimeDir := filepath.Join(strings.TrimSpace(string(out)), "src", "runtime")
dir, err := generateGoAsmHeader(runtimeDir, "", "amd64", t.TempDir())
if err != nil {
t.Fatalf("generateGoAsmHeader(runtime): %v", err)
}
b, err := os.ReadFile(dir + "/go_asm.h")
if err != nil {
t.Fatal(err)
}
hdr := string(b)
for _, want := range []string{
"#define const_hashSize 8\n",
"#define const_avxSupported 1\n",
"#define const_pageSize 8192\n",
"#define g_stackguard0 16\n",
"#define m__size ",
} {
if !strings.Contains(hdr, want) {
t.Errorf("runtime header misses %q", want)
}
}
}
+196 -19
View File
@@ -5,6 +5,7 @@ package main
import (
"fmt"
"maps"
"os"
"os/exec"
"path/filepath"
@@ -37,7 +38,7 @@ import (
// construction and are excluded from the diff; the other architectures list
// their conditional branches outright.
func cmdAuditInstructions(args []string) error {
fs := newCommand("audit-instructions", "gasm audit-instructions [--corpus [dir]] [-I dir] [amd64|arm64|riscv64|loong64]", `
fs := newCommand("audit-instructions", "gasm audit-instructions [--corpus [dir]] [--list] [-I dir] [amd64|arm64|riscv64|loong64]", `
Compare the gasm encoder for the given architecture (default amd64) against
go tool asm and print the diff: superset encodings (gasm-only, shippable via
gasm asm --format goobj) and known-but-unencodable names (the backlog). The
@@ -54,16 +55,19 @@ toolchain probing. A file whose name carries a recognisable _arch suffix is
attempted for that architecture; a file without one is attempted for all
four, exactly as a GOARCH build would compile it. The report gives the
per-architecture pass rates and the most common failure reasons, which drive
the encodability backlog by frequency rather than by table order.
the encodability backlog by frequency rather than by table order. With
-list the report also prints every failing file with its reason, per
architecture.
`)
corpus := fs.Bool("corpus", false, "assemble a corpus of .s files and report pass rates and failure reasons")
list := fs.Bool("list", false, "with --corpus, list every failing file with its reason, per architecture")
var dirs includeDirs
fs.Var(&dirs, "I", "directory to search for #include files (may be repeated)")
if err := fs.Parse(args); err != nil {
return err
}
if *corpus {
return cmdAuditCorpus(fs.Args(), dirs)
return cmdAuditCorpus(fs.Args(), dirs, *list)
}
archName := "amd64"
switch n := len(fs.Args()); {
@@ -288,6 +292,11 @@ func probeShapes(a arch.Arch) []string {
"V1.B16, [V2.B16], V3.B16", "V1.B8, [V2.B16, V3.B16], V4.B8",
"$4, V1.B16, V2.B16, V3.B16", "$15, V1", "V1, V2, p2",
"R0, R1, $1, $4, p2",
// The landing-pad kind, the compiler's PCDATA
// bookkeeping and the four-operand bitfield
// insert/extract family, as the toolchain's own
// testdata spells them.
"C", "$1, $0", "$0, R1, $1, R2",
}
case arch.RISCV:
return []string{
@@ -307,6 +316,9 @@ func probeShapes(a arch.Arch) []string {
"X5, X6, p2", "R5, R6, p2",
"X5, E8, M8, TA, MA, X6", "$4, E32, M1, TA, MA, X1",
"(X5), X6, V1, V2",
// The CSR immediate forms the toolchain's testdata spells:
// immediate, CSR name, destination.
"$2, TIME, X5",
"",
}
case arch.LOONG64:
@@ -323,6 +335,12 @@ func probeShapes(a arch.Arch) []string {
"V1, V2, V3", "X1, X2, X3", "V1, V2", "X1, X2", "V1", "X1",
// The vector compare-to-flag forms land in an FCC register.
"V1, FCC0", "X1, FCC0",
// The compiler's bookkeeping pair and the raw spellings the
// toolchain's own testdata carries: JIRL rd, rj, offset (the
// form RET lowers to), the prefetch with a 32-bit address and
// hint, and the byte-shuffle quads.
"$1, $0", "R1, R5, 0", "0(R7), $5, $0", "(R7), $5, $0",
"V1, V2, V3, V4", "X1, X2, X3, X4",
"",
}
}
@@ -388,20 +406,30 @@ type corpusTally struct {
assembled int
reasons map[string]int // failure reason → count
example map[string]string // failure reason → one representative file
fails []corpusFailure // every failure, in file order, for --list
}
func (t *corpusTally) fail(path, reason string) {
// corpusFailure is one failed attempt, recorded for the --list report.
type corpusFailure struct {
path string
reason string
detail string
}
func (t *corpusTally) fail(path string, err error) {
reason := corpusReason(err)
t.reasons[reason]++
if t.example[reason] == "" {
t.example[reason] = path
}
t.fails = append(t.fails, corpusFailure{path: path, reason: reason, detail: firstLine(err.Error())})
}
// cmdAuditCorpus implements audit-instructions --corpus. The include
// directories carry #include resolution over a corpus whose files refer to
// headers such as GOROOT/pkg/include, the same -I a toolchain comparison
// needs.
func cmdAuditCorpus(args []string, dirs includeDirs) error {
func cmdAuditCorpus(args []string, dirs includeDirs, list bool) error {
if len(args) > 1 {
return &usageError{fmt.Errorf("audit-instructions --corpus takes at most one directory argument")}
}
@@ -418,7 +446,9 @@ func cmdAuditCorpus(args []string, dirs includeDirs) error {
// The toolchain's shipped headers (funcdata.h and friends) define the
// macros GOROOT files include; a corpus audit measures those files, so
// the header directory joins the search path automatically. go_asm.h
// is compiler-generated per package and stays unresolvable on purpose.
// is compiler-generated per package, so it is not resolved from here:
// files that include it get one generated per target architecture,
// which runCorpusAudit arranges.
if out, err := exec.Command("go", "env", "GOROOT").Output(); err == nil {
pkgInclude := filepath.Join(strings.TrimSpace(string(out)), "pkg", "include")
if fi, err := os.Stat(pkgInclude); err == nil && fi.IsDir() {
@@ -437,7 +467,7 @@ func cmdAuditCorpus(args []string, dirs includeDirs) error {
if err != nil {
return err
}
printCorpusStats(stats)
printCorpusStats(stats, list)
return nil
}
@@ -458,13 +488,19 @@ type corpusStats struct {
// set, even when gasm does not support the architecture.
var goPortSuffixes = []string{
"386", "amd64", "arm", "arm64", "loong64", "mips", "mips64",
"mips64le", "mipsle", "ppc64", "ppc64le", "riscv", "riscv64",
"s390x", "wasm",
"mips64le", "mipsle", "mips64x", "mipsx", "ppc64", "ppc64le",
"ppc64x", "riscv", "riscv64", "s390x", "wasm",
}
// otherPortFile reports whether the file's name carries a Go-architecture
// suffix gasm does not support.
// otherPortFile reports whether the file belongs to a build no supported
// target ever compiles: either its name carries a Go-architecture suffix
// gasm does not support, or, for a file with no architecture suffix at all,
// it names another GOOS, which go/build drops from the file set
// (rt0_js_wasm.s is a javascript build, not a generic one).
func otherPortFile(path string) bool {
if otherGOOSFile(path) {
return true
}
base := path
if i := strings.LastIndexByte(base, '/'); i >= 0 {
base = base[i+1:]
@@ -477,6 +513,66 @@ func otherPortFile(path string) bool {
return false
}
// goOSNames are the GOOS values go/build recognises in file names.
var goOSNames = map[string]bool{
"aix": true, "android": true, "darwin": true, "dragonfly": true,
"freebsd": true, "hurd": true, "illumos": true, "ios": true,
"js": true, "linux": true, "nacl": true, "netbsd": true,
"openbsd": true, "plan9": true, "solaris": true, "wasip1": true,
"windows": true, "zos": true,
}
// resolveGOOS validates a -GOOS flag value, mirroring the architecture
// check's surface: a usage error naming what the tool accepts.
func resolveGOOS(name string) (string, error) {
lower := strings.ToLower(name)
if goOSNames[lower] {
return lower, nil
}
return "", &usageError{fmt.Errorf("unknown GOOS %q: want one of %s", name, strings.Join(slices.Sorted(maps.Keys(goOSNames)), ", "))}
}
// goosFromFilename returns the GOOS the file's name carries, by go/build's
// goodOSArchFile rule: the GOOS segment sits last, or last before the
// architecture segment (sys_darwin_arm64.s, vlop_arm.s carries none). An
// empty result means the name names no GOOS and the ambient one applies.
func goosFromFilename(path string) string {
base := path
if i := strings.LastIndexByte(base, '/'); i >= 0 {
base = base[i+1:]
}
base = strings.TrimSuffix(base, ".s")
// go/build ignores everything before the first underscore, so a GOOS
// segment is only ever looked for from there on.
i := strings.IndexByte(base, '_')
if i < 0 {
return ""
}
segs := strings.Split(base[i:], "_")
if n := len(segs); n >= 2 && goOSNames[segs[n-2]] && slices.Contains(goPortSuffixes, segs[n-1]) {
return segs[n-2]
}
if goOSNames[segs[len(segs)-1]] {
return segs[len(segs)-1]
}
return ""
}
// otherGOOSFile reports whether the file's name names a GOOS other than the
// host's, by go/build's file-name rules.
func otherGOOSFile(path string) bool {
base := path
if i := strings.LastIndexByte(base, '/'); i >= 0 {
base = base[i+1:]
}
for seg := range strings.SplitSeq(strings.TrimSuffix(base, ".s"), "_") {
if goOSNames[seg] && seg != runtime.GOOS {
return true
}
}
return false
}
func runCorpusAudit(root string, dirs includeDirs) (*corpusStats, error) {
files, err := asmFiles(root)
if err != nil {
@@ -497,12 +593,26 @@ func runCorpusAudit(root string, dirs includeDirs) (*corpusStats, error) {
// its name allows assembles it.
full, generic, otherPort := 0, 0, 0
// Header generation is created on first use, so a corpus with no
// go_asm.h includes never pays for a temp directory.
var hdr *asmhdrCache
defer func() {
if hdr != nil {
hdr.close()
}
}()
for _, path := range files {
src, err := readSource(path)
if err != nil {
return nil, err
}
f, errs := parser.ParseWithOptions(path, src, parser.Options{Expand: true, IncludeDirs: dirs})
// The GOOS the header generation type-checks under follows the
// file's name when the name carries one; the ambient GOOS is the
// honest guess otherwise (a build tag naming another GOOS is
// invisible to a file-name rule).
goos := goosFromFilename(path)
var wanted []int // indexes into targets
if a := arch.FromFilename(path); a != arch.Unknown {
@@ -513,10 +623,11 @@ func runCorpusAudit(root string, dirs includeDirs) (*corpusStats, error) {
}
} else if otherPortFile(path) {
// A file named for a Go port gasm does not support (arm,
// 386, s390x, ...) is compiled by no supported-arch build,
// so it is neither generic nor a per-arch attempt: counting
// it as generic would make the headline unreachably low
// for reasons no supported target can fix.
// 386, s390x, ...) or for another GOOS is compiled by no
// supported-arch build, so it is neither generic nor a
// per-arch attempt: counting it as generic would make the
// headline unreachably low for reasons no supported target
// can fix.
otherPort++
} else {
generic++
@@ -525,19 +636,76 @@ func runCorpusAudit(root string, dirs includeDirs) (*corpusStats, error) {
}
}
// A file that includes go_asm.h parses against a per-target header:
// the defines differ per architecture (internal/cpu's layout, for
// one) and per GOOS (sys_darwin_arm64.s's trampoline constants,
// for another), so the parse cannot be shared the way a
// header-free file's can. A generation failure is a failure for
// every target, named for the package rather than a bare "include
// not found". A header already resolvable in the package
// directory or the -I list is left alone.
if len(wanted) > 0 && needsGoAsmHeader(src) && !goAsmHeaderResolved(filepath.Dir(path), dirs) {
if hdr == nil {
if hdr, err = newAsmhdrCache(); err != nil {
return nil, err
}
}
pkgDir := filepath.Dir(path)
ok := true
for _, i := range wanted {
tg, t := targets[i], tallies[i]
t.attempted++
hdrDir, err := hdr.dirFor(pkgDir, goos, goarchName(tg.a))
if err != nil {
ok = false
t.fail(path, err)
continue
}
f, errs := parser.ParseWithOptions(path, src, parser.Options{
Expand: true,
IncludeDirs: append(slices.Clone(dirs), hdrDir),
Predefines: platformPredefinesFor(goarchName(tg.a), goos),
})
if len(errs) > 0 {
ok = false
t.fail(path, errs[0])
continue
}
if _, err := assembleFile(tg.a, f, goos); err != nil {
ok = false
t.fail(path, err)
continue
}
t.assembled++
}
if ok && len(wanted) > 0 {
full++
}
continue
}
ok := true
for _, i := range wanted {
tg, t := targets[i], tallies[i]
t.attempted++
// The parse carries the target's platform predefines, so it
// cannot be shared across targets the way a header-free file's
// could: a #ifdef GOARCH_arm block must be live on arm64 and
// dead everywhere else.
f, errs := parser.ParseWithOptions(path, src, parser.Options{
Expand: true,
IncludeDirs: dirs,
Predefines: platformPredefinesFor(goarchName(tg.a), goos),
})
var err error
if len(errs) > 0 {
err = errs[0] // a parse failure is a failure for every target
} else {
_, err = assembleFile(tg.a, f)
_, err = assembleFile(tg.a, f, goos)
}
if err != nil {
ok = false
t.fail(path, corpusReason(err))
t.fail(path, err)
continue
}
t.assembled++
@@ -559,7 +727,7 @@ func runCorpusAudit(root string, dirs includeDirs) (*corpusStats, error) {
}
// printCorpusStats renders the corpus audit report.
func printCorpusStats(s *corpusStats) {
func printCorpusStats(s *corpusStats, list bool) {
fmt.Printf("corpus %s: %d files (%d generic, attempted for all architectures; %d named for other Go ports, never attempted)\n", s.root, s.files, s.generic, s.otherPort)
// The rate is over the files a supported build would attempt: the
// other ports' files sit in the count for completeness but can never
@@ -574,6 +742,13 @@ func printCorpusStats(s *corpusStats) {
fmt.Printf(" %4d %s\n", t.reasons[r], r)
fmt.Printf(" e.g. %s\n", t.example[r])
}
if !list {
continue
}
for _, f := range t.fails {
fmt.Printf(" FAIL %s\n", f.path)
fmt.Printf(" %s: %s\n", f.reason, f.detail)
}
}
}
@@ -581,6 +756,8 @@ func printCorpusStats(s *corpusStats) {
func corpusReason(err error) string {
msg := err.Error()
switch {
case strings.Contains(msg, "go_asm.h for GOARCH"):
return "go_asm.h generation failed"
case strings.Contains(msg, "unsupported"), strings.Contains(msg, "cannot encode"):
return "instruction not encodable"
case strings.Contains(msg, "undefined label"):
+1 -1
View File
@@ -84,7 +84,7 @@ func disSource(path string, target arch.Arch) int {
if len(errs) > 0 {
return 1
}
img, err := assembleFile(target, f)
img, err := assembleFile(target, f, "")
if err != nil {
fmt.Fprintf(os.Stderr, "gasm dis: %v\n", err)
return 1
+72 -13
View File
@@ -486,7 +486,7 @@ hover, document symbols, diagnostics and semantic-token highlighting.
}
func cmdAsm(args []string) int {
fs := newCommand("asm", "gasm asm [--format raw|elf|goobj] [-I dir] [-p pkg] [-GOARCH arch] [-o out] <file>", `
fs := newCommand("asm", "gasm asm [--format raw|elf|goobj] [-I dir] [-p pkg] [-GOARCH arch] [-GOOS os] [-o out] <file>", `
Assemble FILE without the Go toolchain: every TEXT function is encoded to
machine code and printed as a hex dump. Supported architectures: amd64
(including VEX/AVX2 and EVEX/AVX-512), arm64 (AArch64 integer, FP,
@@ -503,16 +503,26 @@ system toolchain; goobj emits the Go toolchain's own object format, which
cmd/link consumes directly (it requires -p, the package path, and the
installed Go toolchain: the object preamble is captured from go tool asm
and the format version from go version).
A file that includes go_asm.h gets that header generated automatically from
the package it lives in (the .go files beside it, type-checked for the
target architecture, the toolchain's own defines), so GOROOT assembly
assembles without a compiler. -GOOS selects the type-checking GOOS for
that header: a GOOS-specific file (sys_darwin_arm64.s) needs its platform's
defines, which a header from the ambient GOOS silently omits. A package
that has no Go files for the target or does not type-check is a hard error
naming the package.
`)
out := fs.String("o", "", "write the output to this file")
format := fs.String("format", "raw", "output format: raw (concatenated image), elf or goobj (Go object)")
pkg := fs.String("p", "", "package path for --format goobj (qualifies the exported symbols)")
archName := fs.String("GOARCH", "", "target architecture: amd64, arm64, riscv64 or loong64 (overrides the file-name suffix)")
goosName := fs.String("GOOS", "", "operating system for go_asm.h generation: a GOOS go/build recognises (default: the host's)")
var dirs includeDirs
fs.Var(&dirs, "I", "directory to search for #include files (may be repeated)")
fs.Parse(args)
if fs.NArg() != 1 {
fmt.Fprintln(os.Stderr, "usage: gasm asm [--format raw|elf|goobj] [-I dir] [-p pkg] [-GOARCH arch] [-o out] <file>")
fmt.Fprintln(os.Stderr, "usage: gasm asm [--format raw|elf|goobj] [-I dir] [-p pkg] [-GOARCH arch] [-GOOS os] [-o out] <file>")
return 2
}
// The format is validated before anything else, so a bogus value exits 2
@@ -533,12 +543,38 @@ and the format version from go version).
}
targetArch = a
}
goos := ""
if *goosName != "" {
g, err := resolveGOOS(*goosName)
if err != nil {
fmt.Fprintf(os.Stderr, "gasm asm: %v\n", err)
return 2
}
goos = g
}
src, err := readSource(path)
if err != nil {
fmt.Fprintln(os.Stderr, "gasm:", err)
return 1
}
f, errs := parser.ParseWithOptions(path, src, parser.Options{Expand: true, IncludeDirs: dirs})
// A file that includes go_asm.h cannot assemble without the package's
// defines, and without a compiler nothing else has generated them, so
// gasm produces the equivalent itself: automatic, because the compiler
// behaves the same way and a flag would only ever be forgotten. A
// generation failure is fatal and names the package: assembling against
// a missing header would fail later with a bare "undefined" instead.
// A go_asm.h that already resolves (placed by hand, or passed with -I)
// is left alone.
if needsGoAsmHeader(src) && !goAsmHeaderResolved(filepath.Dir(path), dirs) {
hdrDir, cleanup, err := ensureGoAsmHeader(path, targetArch, goos, nil)
if err != nil {
fmt.Fprintln(os.Stderr, "gasm asm:", err)
return 1
}
defer cleanup()
dirs = append(dirs, hdrDir)
}
f, errs := parser.ParseWithOptions(path, src, parser.Options{Expand: true, IncludeDirs: dirs, Predefines: platformPredefinesFor(string(targetArch), goos)})
for _, e := range errs {
fmt.Fprintf(os.Stderr, "%s: %v\n", path, e)
}
@@ -546,7 +582,7 @@ and the format version from go version).
return 1
}
img, err := assembleFile(targetArch, f)
img, err := assembleFile(targetArch, f, goos)
if err != nil {
fmt.Fprintf(os.Stderr, "%s: %v\n", path, err)
return 1
@@ -753,11 +789,34 @@ e.g. --map wideCopyAVX2=wideCopyAVX512 pairs the two regardless of suffix.
return 1
}
// platformPredefines mirrors the go command's assembler invocation, which
// defines GOOS_<goos> and GOARCH_<arch> as -D macros: GOROOT headers
// (go_tls.h, asm_riscv64.h) select their platform blocks with #ifdef on
// exactly those names, so an assembler without them cannot see the platform
// definitions at all.
func platformPredefines(goarch, goos string) map[string]string {
return map[string]string{
"GOARCH_" + goarch: "1",
"GOOS_" + goos: "1",
}
}
// platformPredefinesFor resolves the ambient GOOS the way a build would: a
// file whose name carries one (sys_darwin_arm64.s) is compiled for that GOOS
// and nothing else.
func platformPredefinesFor(goarch string, fileGoos string) map[string]string {
goos := fileGoos
if goos == "" {
goos = runtime.GOOS
}
return platformPredefines(goarch, goos)
}
// assembleFile assembles a parsed file for the given architecture and returns the image.
func assembleFile(targetArch arch.Arch, f *ast.File) (*asm.Image, error) {
func assembleFile(targetArch arch.Arch, f *ast.File, goos string) (*asm.Image, error) {
switch targetArch {
case arch.AMD64:
return asm.AssembleFile(f)
return asm.AssembleFile(f, asm.WithGOOS(goos))
case arch.RISCV:
return asm.AssembleFileRISCV(f)
case arch.ARM64:
@@ -777,18 +836,18 @@ func assemblePath(path string, forced arch.Arch, dirs includeDirs) (*asm.Image,
if err != nil {
return nil, err
}
f, errs := parser.ParseWithOptions(path, src, parser.Options{Expand: true, IncludeDirs: dirs})
target := forced
if target == arch.Unknown {
target = arch.FromFilename(path)
}
f, errs := parser.ParseWithOptions(path, src, parser.Options{Expand: true, IncludeDirs: dirs, Predefines: platformPredefinesFor(string(target), "")})
for _, e := range errs {
fmt.Fprintf(os.Stderr, "%s: %v\n", path, e)
}
if len(errs) > 0 {
return nil, fmt.Errorf("parse errors")
}
target := forced
if target == arch.Unknown {
target = arch.FromFilename(path)
}
return assembleFile(target, f)
return assembleFile(target, f, "")
}
// printByteDiff shows the first few byte differences between two code blocks.
@@ -884,7 +943,7 @@ func cmdVerifyNonJIT(path string, targetArch arch.Arch, groundTruth, profile boo
if len(errs) > 0 {
return 1
}
img, err := assembleFile(targetArch, f)
img, err := assembleFile(targetArch, f, "")
if err != nil {
fmt.Fprintf(os.Stderr, "gasm verify: %v\n", err)
return 1
+14
View File
@@ -119,6 +119,20 @@ identifier is a register or a label is an *architecture* question, so it is
left to `arch` and resolved in the lint/lsp layers. This keeps the parser
arch-agnostic and its output deterministic.
### Optional preprocessing
With `Options{Expand: true}` the parser runs a pre-parse pass
(`preproc.go`) that splices `#include` files (the source directory, then the
`-I` directories), expands object and parameterised `#define` macros,
applies `#undef` and the `#ifdef`/`#ifndef`/`#else`/`#endif` family, and
folds constant expressions left in operands. The go command's platform
macros (`GOARCH_<arch>`, `GOOS_<goos>`) arrive through `Options.Predefines`.
The assembly path (`asm`, `diff`, `audit`) expands; `lint`, `fmt` and the
language server read the raw file. The command layer adds the go_asm.h
generator (`asmhdr.go`): a file that includes go_asm.h gets the package's
defines type-checked out of its Go files for the target architecture and
GOOS, with no compiler in the loop.
### `arch`
Register files are generated programmatically (the regular `R8`-`R15`,
+16 -4
View File
@@ -141,7 +141,7 @@ gasm lint kernel_amd64.s
## asm
```text
Usage: gasm asm [--format raw|elf|goobj] [-I dir] [-p pkg] [-GOARCH arch] [-o out] <file>
Usage: gasm asm [--format raw|elf|goobj] [-I dir] [-p pkg] [-GOARCH arch] [-GOOS os] [-o out] <file>
```
| Flag | Default | Effect |
@@ -150,6 +150,7 @@ Usage: gasm asm [--format raw|elf|goobj] [-I dir] [-p pkg] [-GOARCH arch] [-o ou
| `-I` | empty | directory to search for `#include` files; may be repeated, searched in order after the source directory |
| `-p` | empty | package path for `--format goobj`, qualifying the exported symbols |
| `-GOARCH` | empty | target architecture: `amd64`, `arm64`, `riscv64` or `loong64`; overrides the file-name suffix |
| `-GOOS` | empty | operating system for the generated `go_asm.h`: any GOOS `go/build` recognises in file names; default is the host's |
| `-o` | empty | write the output to this file instead of a hex dump on stdout |
Supported architectures: amd64 (VEX/AVX2 and EVEX/AVX-512 included), arm64,
@@ -163,6 +164,13 @@ system toolchain; `goobj` emits the Go toolchain's own object format, which
installed: the object preamble is captured from `go tool asm` and the format
version from `go version`. `raw` and `elf` need no toolchain at all.
A file that includes `go_asm.h` gets that header generated from the Go
files beside it, type-checked for the target. `-GOOS` selects the
type-checking GOOS for that header, because a GOOS-specific file needs its
platform's defines: `sys_darwin_arm64.s` fails against the ambient GOOS
(`machTimebaseInfo_numer` is missing from a linux type-check) and assembles
with `-GOOS darwin`.
Assembly preprocessing matches the toolchain's: `#define` macros (object and
parameterised) expand at the point of use, `#undef`, `#ifdef`, `#ifndef`,
`#else` and `#endif` behave as in `go tool asm`, `;` separates statements,
@@ -357,7 +365,7 @@ add: 16 bytes, args=24, frame=0 NOSPLIT
## audit-instructions
```text
Usage: gasm audit-instructions [--corpus [dir]] [-I dir] [amd64|arm64|riscv64|loong64]
Usage: gasm audit-instructions [--corpus [dir]] [--list] [-I dir] [amd64|arm64|riscv64|loong64]
```
Compare the gasm encoder for the given architecture (default amd64) against the
@@ -390,10 +398,14 @@ With `--corpus` the audit changes shape: it assembles every `.s` file under
DIR (default `GOROOT/src`) with the gasm encoder only, no toolchain probing.
A file whose name carries a recognisable `_arch` suffix is attempted for that
architecture; a file without one is attempted for all four, exactly as a
`GOARCH` build would compile it. The report gives the headline number (files
`GOARCH` build would compile it, and a name that names a GOOS
(`sys_darwin_arm64.s`) type-checks its generated `go_asm.h` for that GOOS.
The report gives the headline number (files
that assemble for every target architecture), the per-architecture pass rates
and the most common failure reasons with one representative file each, which
drive the encodability backlog by frequency rather than by table order. A run
drive the encodability backlog by frequency rather than by table order. With
`--list` the report additionally prints every failing file with its failure
reason, per architecture. A run
over GOROOT takes under a second.
```sh
+657
View File
@@ -0,0 +1,657 @@
# The GOOBJ object file format
This document is a complete specification of GOOBJ, the object file format
that the Go toolchain's assembler, compiler and linker exchange, written for
implementers of independent producers and consumers. It documents the format
as shipped by Go 1.27.1, identified by the magic string `"\x00go120ld"`.
No comparable document exists upstream. The format is defined only by the
source of the `cmd/internal/goobj` package inside the toolchain tree, it is an
internal interface with no stability promise, and it can change in any
release. This specification was therefore produced by reverse engineering
that source and by parsing real objects produced by `go tool asm` and
`go tool compile`, byte for byte, against the layout described here. Within
gasm-devkit it is kept honest by the differential tests in `asm/goobj_test.go`
and `asm/link_test.go`, which compare `gasm asm --format goobj` output against
the toolchain's own products and feed gasm objects to `go build`.
Every numeric value in this document, every block index, structure size, flag
bit, type code and relocation number, was read from the Go 1.27.1 source at
`/usr/local/go/src/cmd/internal/goobj`, `cmd/internal/obj` and
`cmd/internal/objabi`, and exercised against assembled objects.
## Containers
The unit this document specifies is the **object**: one package's worth of
symbols, relocations and data. An object is never consumed naked. Two
wrappers exist in practice, and the linker dispatches on the first bytes of
the file.
**The bare object**, written by `go tool asm`:
```text
"go object linux amd64 go1.27.1 GOAMD64=v1 X:regabiwrappers,...\n"
"!\n"
<GOOBJ blob>
```
The first line is the toolchain configuration string, produced by
`objabi.HeaderString`: `go object`, the GOOS, the GOARCH, the toolchain
version, an optional architecture qualifier such as `GOAMD64=v1`, and
`X:` followed by the enabled experiments, comma separated. The linker requires
this line to match its own configuration exactly and rejects the file
otherwise; the `-f` linker flag waives the check. Header lines may be
followed by export data delimited by `$$` markers; the header region always
ends at the first line consisting of exactly `!`, and the GOOBJ blob starts
immediately after that line.
**The package archive**, written by the compiler output pipeline and consumed
by `go build`: the classic `ar` format, magic `!<arch>\n`, with the export
data in a `__.PKGDEF` member and one or more objects as further members, each
carrying the bare-object structure above. `go tool pack` creates and
inspects such archives.
| Consumer | Role |
|---|---|
| `cmd/asm` | writes objects from `.s` files |
| `cmd/compile` | writes objects from Go source |
| `cmd/link` | reads objects and archives, produces executables |
| `cmd/nm`, `cmd/objdump` | read objects through `cmd/internal/objfile` |
## Conventions
- All integers are **little endian**.
- There is **no alignment or padding** anywhere in the file; structures follow
one another byte by byte.
- Every offset stored in the file is **relative to the first byte of the GOOBJ
blob**, not to the start of the container.
- The blob opens with a 96 byte header that carries the byte offset of every
block. A block's length is the difference between its own offset and the
next block's, so the offset array is the only index the format needs.
### Layout overview
```mermaid
flowchart TB
A[Container header line and ! terminator] --> B[File header, 96 bytes]
B --> C[String table, implicit region]
C --> D[Autolib]
D --> E[PkgIndex]
E --> F[Files]
F --> G[Symbol definition arrays: Symdef, Hashed64def, Hasheddef, Nonpkgdef, Nonpkgref]
G --> H[RefFlags]
H --> I[Hash64 and Hash]
I --> J[RelocIndex, AuxIndex, DataIndex]
J --> K[Relocs]
K --> L[Aux]
L --> M[Data]
M --> N[RefNames]
N --> O[BlkEnd marks the end of the blob]
```
## The file header
Exactly 96 bytes: 8 magic, 8 fingerprint, 4 flags, and 19 four byte block
offsets.
| Offset | Size | Field | Meaning |
|---|---|---|---|
| 0 | 8 | Magic | `"\x00go120ld"`. A reader rejects anything else. The digits are the format version and have moved before; a new toolchain release may move them again. |
| 8 | 8 | Fingerprint | Identifies the package build. The compiler writes a hash of the export data; the assembler leaves all zero. The linker compares this against the fingerprint recorded by importers. |
| 16 | 4 | Flags | Bit field, see below. |
| 20 | 76 | Offsets | 19 `uint32` entries, one per block index 0 to 18. |
Header flags:
| Bit | Value | Name | Meaning |
|---|---|---|---|
| 0 | 1 | ObjFlagShared | built with `-shared` |
| 1 | 2 | reserved | was `ObjFlagNeedNameExpansion`, now unused |
| 2 | 4 | ObjFlagFromAssembly | produced from assembly source; `go tool asm` and gasm set this |
| 3 | 8 | ObjFlagUnlinkable | package path is invalid, the linker refuses to link |
| 4 | 16 | ObjFlagStd | standard library package |
### Block indices
The offset array is indexed by these constants, in file order:
| Index | Constant | Contents |
|---|---|---|
| 0 | BlkAutolib | imported packages |
| 1 | BlkPkgIndex | referenced packages, indexed |
| 2 | BlkFile | source file names |
| 3 | BlkSymdef | symbol definitions, package scope |
| 4 | BlkHashed64def | short hashed definitions |
| 5 | BlkHasheddef | hashed definitions |
| 6 | BlkNonpkgdef | non-package definitions |
| 7 | BlkNonpkgref | non-package references |
| 8 | BlkRefFlags | flags of referenced symbols |
| 9 | BlkHash64 | 8 byte hashes for short hashed definitions |
| 10 | BlkHash | 16 byte hashes for hashed definitions |
| 11 | BlkRelocIndex | per symbol relocation start index |
| 12 | BlkAuxIndex | per symbol aux start index |
| 13 | BlkDataIndex | per symbol data offset |
| 14 | BlkReloc | relocations |
| 15 | BlkAux | aux symbol entries |
| 16 | BlkData | symbol payloads |
| 17 | BlkRefName | names of referenced symbols, for tools |
| 18 | BlkEnd | no contents; its offset is the end of the blob |
## The string table
There is no block index for strings. The table occupies the implicit region
between the end of the header (offset 96) and `Offsets[BlkAutolib]`, and every
string offset in the file points into that region. The writer de-duplicates:
each distinct string is stored once, in first-use order, and the empty string
is always the first entry, so its reference is length 0 and offset 96.
A **string reference** is 8 bytes: `uint32` length, then `uint32` absolute
offset of the bytes. The bytes are stored raw, with no terminator.
## Symbol references and the package index
A **symbol reference** (SymRef) is 8 bytes: two `uint32`, `PkgIdx` and
`SymIdx`. The pair `{0, 0}` means nil. `PkgIdx` says which array the symbol
lives in:
| Value | Constant | SymIdx indexes |
|---|---|---|
| 0 | PkgIdxInvalid | never valid in a written file |
| 1 and up, ascending | (imported packages) | the SymbolDefs array of the package named at PkgIndex entry `PkgIdx` |
| 0x7ffffffb | PkgIdxSelf | this object's Symdef array |
| 0x7ffffffc | PkgIdxBuiltin | the compiler's builtin table, see Builtins |
| 0x7ffffffd | PkgIdxHashed | this object's Hasheddef array |
| 0x7ffffffe | PkgIdxHashed64 | this object's Hashed64def array |
| 0x7fffffff | PkgIdxNone | NonPkgDefs, overflowing into NonPkgRefs |
Assignment rules, as the toolchain performs them:
- Every definition a package exports to the linker by index lands in Symdefs
with PkgIdxSelf. The compiler puts its functions and data here; the
assembler puts only its file-local static symbols here, everything else by
name, see below.
- External package references take indices 1, 2, 3, in order of first
reference during assembly; the package names go into PkgIndex at those
indices, entry 0 is the empty package and is never referenced.
- References to the compiler's builtin functions become PkgIdxBuiltin with
SymIdx set to the builtin's index.
- A symbol referenced **by name** rather than by index becomes PkgIdxNone and
its index counts through NonPkgDefs first, then continues into NonPkgRefs.
A producer must emit the definitions it made in NonPkgDefs and the pure
references in NonPkgRefs.
- The assembler's rule, from `cmd/internal/obj/sym.go`: every assembly symbol
is referenced by name, PkgIdxNone, **except** file-local static symbols,
whose names carry `<>` and which are referenced by index. The compiler also
forces references by name for symbols marked `//go:linkname` and for any
symbol with the DUPOK attribute, which the linker de-duplicates by name.
## Symbol definition entries
The five definition and reference arrays (block indices 3 to 7) share one
element layout, 21 bytes:
| Offset | Size | Field | Meaning |
|---|---|---|---|
| 0 | 8 | Name | string reference |
| 8 | 2 | ABI | see table below |
| 10 | 1 | Type | symbol kind, see the kind table |
| 11 | 1 | Flag | bit field, see below |
| 12 | 1 | Flag2 | second bit field, see below |
| 13 | 4 | Siz | payload size in bytes, `uint32` |
| 17 | 4 | Align | alignment the linker must honour, `uint32` |
The Name is a real string reference for hand-written symbols. The auxiliary
symbols the toolchain generates per function, the FuncInfo payload, the DWARF
entries, have empty names: length 0, and their identity is only via the Aux
entries that point at them by index.
### The ABI field
| Value | Meaning |
|---|---|
| 0 | ABI0, the stack based ABI, the ABI of every hand-written assembly function |
| 1 | ABIInternal, the register ABI of compiler-generated functions |
| 0xffff | static, a file-local symbol (`name<>(SB)`), `SymABIstatic` |
### The Flag byte
| Bit | Value | Name | Meaning |
|---|---|---|---|
| 0 | 1 | SymFlagDupok | duplicates allowed, the linker merges them |
| 1 | 2 | SymFlagLocal | file-local |
| 2 | 4 | SymFlagTypelink | belongs in the typelink table |
| 3 | 8 | SymFlagLeaf | leaf function |
| 4 | 16 | SymFlagNoSplit | no stack-split preamble |
| 5 | 32 | SymFlagReflectMethod | `//go:reflectmethod` reachability |
| 6 | 64 | SymFlagGoType | a Go type descriptor, `type:` name and SRODATA |
Note that NoSplit is not reserved for explicit `NOSPLIT` declarations. On
amd64 the assembler itself marks any function whose frame is below
`abi.StackSmall` and whose body calls nothing that needs stack as NoSplit and
omits the split check, so a `TEXT` without `NOSPLIT` can still carry the bit.
### The Flag2 byte
| Bit | Value | Name | Meaning |
|---|---|---|---|
| 0 | 1 | SymFlagUsedInIface | type or itab reachable through an interface |
| 1 | 2 | SymFlagItab | an itab, `go:itab.` name and SRODATA |
| 2 | 4 | SymFlagDict | a generic dictionary symbol |
| 3 | 8 | SymFlagPkgInit | package initialisation function |
| 4 | 16 | SymFlagLinkname | reachable through `//go:linkname`; the assembler also sets it on `main.main` |
| 5 | 32 | SymFlagLinknameStd | linkname into the standard library |
| 6 | 64 | SymFlagABIWrapper | ABI transition wrapper |
| 7 | 128 | SymFlagWasmExport | `//go:wasmexport` target |
### The Type byte: symbol kinds
Values of `objabi.SymKind`, in numeric order:
| Value | Name | Meaning |
|---|---|---|
| 0 | Sxxx | invalid zero value |
| 1 | STEXT | executable code |
| 2 | STEXTFIPS | executable code, FIPS section |
| 3 | SRODATA | read only data |
| 4 | SRODATAFIPS | read only data, FIPS section |
| 5 | SNOPTRDATA | data without pointers |
| 6 | SNOPTRDATAFIPS | data without pointers, FIPS section |
| 7 | SDATA | data, may contain pointers |
| 8 | SDATAFIPS | data, FIPS section |
| 9 | SBSS | zero initialised data |
| 10 | SNOPTRBSS | zero initialised data without pointers |
| 11 | STLSBSS | thread local zero initialised data |
| 12 | SDWARFCUINFO | DWARF compile unit information |
| 13 | SDWARFCONST | DWARF constants |
| 14 | SDWARFFCN | DWARF function entry |
| 15 | SDWARFABSFCN | DWARF absolute function entry |
| 16 | SDWARFTYPE | DWARF type information |
| 17 | SDWARFVAR | DWARF variable information |
| 18 | SDWARFRANGE | DWARF range lists |
| 19 | SDWARFLOC | DWARF location lists |
| 20 | SDWARFLINES | DWARF line programs |
| 21 | SDWARFADDR | DWARF address table |
| 22 | SLIBFUZZER_8BIT_COUNTER | libFuzzer coverage counter |
| 23 | SCOVERAGE_COUNTER | coverage counter |
| 24 | SCOVERAGE_AUXVAR | coverage auxiliary variable |
| 25 | SSEHUNWINDINFO | Windows SEH unwind information |
## Referenced symbol flags (RefFlags)
Element size 10 bytes, one per referenced external indexed symbol that
carries a non-zero Flag2:
| Offset | Size | Field |
|---|---|---|
| 0 | 8 | Sym, a SymRef into another package |
| 8 | 1 | Flag, always 0 in current writers |
| 9 | 1 | Flag2, only SymFlagUsedInIface is ever written |
The linker uses these to preserve reachability of interface conversions
across package boundaries. Entries with no flags are omitted entirely.
## Hashes
**Hash64**, block 9: one `uint64` per Hashed64def entry, in array order. Not
a hash at all: the writer copies the **first 8 bytes of the symbol's
payload**. Only symbols whose content-hash section byte is 0 may use the
short form.
**Hash**, block 10: 16 bytes per Hasheddef entry: the first 16 bytes of a
SHA-256 computation over a seed byte `0x01` followed by the hash input. The
input, from `cmd/internal/obj/objfile.go`:
1. the payload size, little endian `uint64`;
2. the section byte, one of `t` for STEXT, `f` for STEXTFIPS, `P` for pcdata,
`F` for the `go:func.*` and `go:funcrel.*` families, `T` for `type:`
symbols, otherwise 0;
3. for text symbols, the symbol name, which keeps distinct functions from
merging;
4. the payload with trailing zero bytes trimmed;
5. for each relocation: a 14 byte record, offset `uint32`, size `uint8`,
low type byte `uint8`, addend `int64`, followed by an encoding of the
target: tag byte 0 then the target's short hash, tag 1 then its full
hash, tag 2 then its expanded name, tag 3 then its builtin index, or,
for PkgIdxSelf and imported packages, no tag, then the package path
and the symbol index.
Two symbols with equal hashes are interchangeable at link time, which is what
makes content addressing work. A producer that computes these hashes wrongly
produces objects that link but de-duplicate wrongly; gasm verifies them by
byte comparison against `go tool asm`.
## The index arrays
Three arrays of `uint32`, one element per **defined** symbol plus one final
element, in the order Symdefs, Hashed64defs, Hasheddefs, NonPkgDefs. With N
defined symbols, each array holds N + 1 entries, and the entry at N is the
total.
- RelocIndex: entry i is where symbol i's relocations start in BlkReloc;
entry i + 1 minus entry i is its count.
- AuxIndex: the same construction over BlkAux.
- DataIndex: entry i is the byte offset of symbol i's payload within BlkData;
the count is the difference of neighbours.
The toolchain writes relocations grouped per symbol in definition order, and
sorts each symbol's relocations by their Off field first. A producer that
skips the sort produces objects the linker still accepts, but that no longer
compare byte-for-byte with the toolchain's output.
## Relocations
Element size 23 bytes:
| Offset | Size | Field | Meaning |
|---|---|---|---|
| 0 | 4 | Off | patch position, bytes from the start of the symbol's payload, `int32` |
| 4 | 1 | Siz | patch width in bytes |
| 5 | 2 | Type | relocation type, `uint16`, see the table |
| 7 | 8 | Add | addend, `int64` |
| 15 | 8 | Sym | target SymRef |
The computed value `payload[Off:Off+Siz] += address(Sym) + Add` in the
flavour the type prescribes is the linker's job; the object only records the
request. A size 0 relocation patches nothing and exists purely as a marker
for the linker's reachability analysis.
### Relocation types
Values of `objabi.RelocType`. The assembler and compiler emit the generic
ones plus their own architecture's family; the rest exist for other ports and
for the linker itself.
| Value | Name | Meaning |
|---|---|---|
| 1 | R_ADDR | absolute address |
| 2 | R_ADDRPOWER | ppc64: high adjusted plus low 16 bits across two D-form instructions |
| 3 | R_ADDRARM64 | arm64: adrp plus add pair |
| 4 | R_ADDRMIPS | mips: low 16 bits of an external address |
| 5 | R_ADDROFF | 32-bit offset from the section start to the symbol |
| 6 | R_SIZE | size of the referenced symbol |
| 7 | R_CALL | direct call, PC relative |
| 8 | R_CALLARM | arm: call with a shifted 24-bit field |
| 9 | R_CALLARM64 | arm64: BL |
| 10 | R_CALLIND | indirect call marker |
| 11 | R_CALLPOWER | ppc64: call |
| 12 | R_CALLMIPS | mips: non-PC-relative call target |
| 13 | R_CONST | constant value of the symbol |
| 14 | R_PCREL | PC relative displacement |
| 15 | R_TLS_LE | thread local, local exec offset |
| 16 | R_TLS_IE | thread local, initial exec GOT offset |
| 17 | R_GOTOFF | offset from the GOT base |
| 18 | R_PLT0 | PLT sequence, first instruction |
| 19 | R_PLT1 | PLT sequence, second instruction |
| 20 | R_PLT2 | PLT sequence, third instruction |
| 21 | R_USEFIELD | field reachability marker |
| 22 | R_USETYPE | type reachability marker, no bytes patched |
| 23 | R_USEIFACE | interface conversion marker, size 0 |
| 24 | R_USEIFACEMETHOD | interface method marker, size 0, addend is the method offset |
| 25 | R_USENAMEDMETHOD | keeps named methods alive |
| 26 | R_METHODOFF | like R_ADDROFF, the linker may zero it when the method is dead |
| 27 | R_KEEP | keeps the target alive if the source survives |
| 28 | R_POWER_TOC | ppc64: TOC relative |
| 29 | R_GOTPCREL | 32-bit PC relative GOT slot |
| 30 | R_JMPMIPS | mips: non-PC-relative jump target |
| 31 | R_DWARFSECREF | offset of the symbol from its section, DWARF use |
| 32 | R_ARM64_TLS_LE | arm64: MOV[NZ] immediate, TLS local exec |
| 33 | R_ARM64_TLS_IE | arm64: adrp plus ldr, TLS initial exec |
| 34 | R_ARM64_GOTPCREL | arm64: adrp plus ldr GOT slot |
| 35 | R_ARM64_GOT | arm64: GOT relative sequence |
| 36 | R_ARM64_PCREL | arm64: adrp plus add PC relative |
| 37 | R_ARM64_PCREL_LDST8 | arm64: adrp plus 8-bit load or store |
| 38 | R_ARM64_PCREL_LDST16 | arm64: adrp plus 16-bit load or store |
| 39 | R_ARM64_PCREL_LDST32 | arm64: adrp plus 32-bit load or store |
| 40 | R_ARM64_PCREL_LDST64 | arm64: adrp plus 64-bit load or store |
| 41 | R_ARM64_LDST8 | arm64: 12-bit load or store immediate, byte |
| 42 | R_ARM64_LDST16 | arm64: bits 11 to 1 of the address |
| 43 | R_ARM64_LDST32 | arm64: bits 11 to 2 |
| 44 | R_ARM64_LDST64 | arm64: bits 11 to 3 |
| 45 | R_ARM64_LDST128 | arm64: bits 11 to 4 |
| 46 | R_POWER_TLS_LE | ppc64: TLS local exec across two instructions |
| 47 | R_POWER_TLS_IE | ppc64: TLS initial exec via GOT |
| 48 | R_POWER_TLS | ppc64: marks the X-form instruction completing a TLS sequence |
| 49 | R_POWER_TLS_IE_PCREL34 | ppc64: prefixed TLS initial exec load |
| 50 | R_POWER_TLS_LE_TPREL34 | ppc64: prefixed TLS local exec |
| 51 | R_ADDRPOWER_DS | ppc64: DS-form second instruction, bits 15 to 2 |
| 52 | R_ADDRPOWER_GOT | ppc64: GOT entry relative to TOC |
| 53 | R_ADDRPOWER_GOT_PCREL34 | ppc64: PC relative GOT, prefixed |
| 54 | R_ADDRPOWER_PCREL | ppc64: PC relative across two D-form instructions |
| 55 | R_ADDRPOWER_TOCREL | ppc64: TOC relative across two D-form instructions |
| 56 | R_ADDRPOWER_TOCREL_DS | ppc64: TOC relative, DS form |
| 57 | R_ADDRPOWER_D34 | ppc64: prefixed absolute, 34 bits |
| 58 | R_ADDRPOWER_PCREL34 | ppc64: prefixed PC relative, 34 bits |
| 59 | R_RISCV_JAL | riscv64: 20-bit J-type offset |
| 60 | R_RISCV_JAL_TRAMP | riscv64: as R_RISCV_JAL, linker-generated trampolines only |
| 61 | R_RISCV_CALL | riscv64: AUIPC plus JALR pair |
| 62 | R_RISCV_PCREL_ITYPE | riscv64: AUIPC plus I-type pair |
| 63 | R_RISCV_PCREL_STYPE | riscv64: AUIPC plus S-type pair |
| 64 | R_RISCV_TLS_IE | riscv64: TLS initial exec, AUIPC plus I-type |
| 65 | R_RISCV_TLS_LE | riscv64: TLS local exec, LUI plus I-type |
| 66 | R_RISCV_GOT_HI20 | riscv64: high 20 bits of a GOT address |
| 67 | R_RISCV_GOT_PCREL_ITYPE | riscv64: GOT entry, AUIPC plus I-type |
| 68 | R_RISCV_PCREL_HI20 | riscv64: high 20 bits of a PC relative address |
| 69 | R_RISCV_PCREL_LO12_I | riscv64: low 12 bits, I-type |
| 70 | R_RISCV_PCREL_LO12_S | riscv64: low 12 bits, S-type |
| 71 | R_RISCV_BRANCH | riscv64: 12-bit branch offset |
| 72 | R_RISCV_ADD32 | riscv64: in-place addition, V + S + A |
| 73 | R_RISCV_SUB32 | riscv64: in-place subtraction, V - S - A |
| 74 | R_RISCV_RVC_BRANCH | riscv64: 8-bit compressed branch offset |
| 75 | R_RISCV_RVC_JUMP | riscv64: 11-bit compressed jump offset |
| 76 | R_PCRELDBL | s390x: PC relative, 2-byte aligned |
| 77 | R_LOONG64_ADDR_HI | loong64: bits 31 to 12 of an address |
| 78 | R_LOONG64_ADDR_LO | loong64: low 12 bits |
| 79 | R_LOONG64_ADDR64_HI | loong64: bits 63 to 52 |
| 80 | R_LOONG64_ADDR64_LO | loong64: bits 51 to 32 |
| 81 | R_LOONG64_ADDR_PCREL20_S2 | loong64: 22-bit aligned PC relative, PCADDI |
| 82 | R_LOONG64_TLS_LE_HI | loong64: TLS local exec, high bits |
| 83 | R_LOONG64_TLS_LE_LO | loong64: TLS local exec, low bits |
| 84 | R_CALLLOONG64 | loong64: 28-bit aligned BL |
| 85 | R_LOONG64_CALL36 | loong64: 38-bit aligned PCADDU18I plus JIRL |
| 86 | R_LOONG64_TLS_IE_HI | loong64: TLS initial exec via GOT, high |
| 87 | R_LOONG64_TLS_IE_LO | loong64: TLS initial exec via GOT, low |
| 88 | R_LOONG64_GOT_HI | loong64: GOT entry, high bits |
| 89 | R_LOONG64_GOT_LO | loong64: GOT entry, low bits |
| 90 | R_LOONG64_GOT64_HI | loong64: 64-bit GOT entry, high |
| 91 | R_LOONG64_GOT64_LO | loong64: 64-bit GOT entry, low |
| 92 | R_LOONG64_ADD64 | loong64: 64-bit in-place addition |
| 93 | R_LOONG64_SUB64 | loong64: 64-bit in-place subtraction |
| 94 | R_JMP16LOONG64 | loong64: 18-bit aligned conditional jump |
| 95 | R_JMP21LOONG64 | loong64: 23-bit aligned BEQZ or BNEZ |
| 96 | R_ADDRMIPSU | mips: sign-adjusted upper 16 bits |
| 97 | R_ADDRMIPSTLS | mips: TLS low 16 bits |
| 98 | R_ADDRCUOFF | pointer-sized offset from the DWARF compile unit start |
| 99 | R_WASMIMPORT | wasm: import module and name indices |
| 100 | R_XCOFFREF | aix: keeps the target alive, patches nothing |
| 101 | R_PEIMAGEOFF | windows: offset from the image base |
| 102 | R_INITORDER | orders inittask records, patches nothing |
| 103 | R_DWTXTADDR_U1 | writes a 1-byte ULEB .debug_addr index for the target function |
| 104 | R_DWTXTADDR_U2 | as above, 2 bytes |
| 105 | R_DWTXTADDR_U3 | as above, 3 bytes |
| 106 | R_DWTXTADDR_U4 | as above, 4 bytes; the assembler always picks this one |
| -32768 | R_WEAK | mask: the target need not be reachable, see below |
| -32767 | R_WEAKADDR | R_WEAK or R_ADDR |
| -32763 | R_WEAKADDROFF | R_WEAK or R_ADDROFF |
R_WEAK is bit 15 set on a negative `int16`: a weak relocation is the base
type's value with bit 15 set. The linker strips the bit before dispatch.
## Aux symbol entries
Element size 9 bytes: a `uint8` type then a SymRef. Aux entries attach
auxiliary symbols to a definition; the arrays run per symbol in the order
given by AuxIndex.
| Value | Name | Attaches |
|---|---|---|
| 0 | AuxGotype | the Go type of a data symbol |
| 1 | AuxFuncInfo | the FuncInfo payload of a text symbol |
| 2 | AuxFuncdata | one funcdata symbol; one entry per slot, nil slots carry the {0,0} reference |
| 3 | AuxDwarfInfo | DWARF debug info for the function |
| 4 | AuxDwarfLoc | DWARF location lists |
| 5 | AuxDwarfRanges | DWARF range lists |
| 6 | AuxDwarfLines | DWARF line program |
| 7 | AuxPcsp | pc-value table: SP adjustments |
| 8 | AuxPcfile | pc-value table: source file indices |
| 9 | AuxPcline | pc-value table: line numbers |
| 10 | AuxPcinline | pc-value table: inlining tree positions |
| 11 | AuxPcdata | one pc-value table per live variable slot |
| 12 | AuxWasmImport | wasm import description |
| 13 | AuxWasmType | wasm export type description |
| 14 | AuxSehUnwindInfo | Windows SEH unwind info |
The writer emits them in the order Gotype, FuncInfo, Funcdata entries,
DwarfInfo, DwarfLoc, DwarfRanges, DwarfLines, Pcsp, Pcfile, Pcline, Pcinline,
SehUnwindInfo, Pcdata entries, WasmImport, WasmType, and skips any whose
payload would be empty. A function assembled from `.s` source by Go 1.27.1
carries exactly: FuncInfo, the Funcdata slots including nils, DwarfInfo,
DwarfLines, Pcsp, Pcfile, Pcline and Pcinline; gasm's writer produces the
same set.
The aux targets are either PkgIdxSelf definitions, PkgIdxHashed pcdata
symbols, or, for the funcdata of assembly functions, PkgIdxNone references
carrying names such as `pkg.Fn.args_stackmap` and `pkg.Fn.arginfo0`, which
resolve to definitions in the package's compiled Go code when there is any.
## Symbol payloads (BlkData)
The payloads of all defined symbols, in definition order, concatenated with
no padding; DataIndex gives each symbol's slice. A text symbol's payload is
its machine code, with the stack-split preamble and any morestack block
already included. A data symbol's payload is the bytes laid down by its DATA
directives, zero filled to its declared size. If a symbol was created from an
embedded file, the file's bytes follow the payload and count towards its
DataIndex extent; assembly producers never write this extension.
### The FuncInfo payload
An SDATA symbol with no name, referenced by AuxFuncInfo. 28 bytes minimum,
little endian:
| Offset | Size | Field | Meaning |
|---|---|---|---|
| 0 | 4 | Args | argument area in bytes; 0x80000000 when the producer declared none |
| 4 | 4 | Locals | frame size in bytes |
| 8 | 1 | FuncID | runtime function classification, 0 means normal |
| 9 | 1 | FuncFlag | TopFrame = 1, SPWrite = 2, Asm = 4 |
| 10 | 2 | padding | zero, reserved to a 4 byte boundary |
| 12 | 4 | StartLine | source line of the TEXT declaration |
| 16 | 4 | NumFile | count of file indices that follow |
| 20 | 4 × NumFile | Files | indices into the Files block, ascending |
| then | 4 | NumInlTree | count of inlining tree nodes that follow |
| then | 24 × NumInlTree | InlTree | nodes, see below |
One InlTree node, 24 bytes: `int32` parent index, `uint32` file index,
`int32` line, `uint32` PkgIdx and `uint32` SymIdx of the inlined function, and
`int32` parent PC.
The assembler derives FuncID from the symbol name through
`objabi.GetFuncID`, so a runtime function with a name the runtime treats
specially gets that classification even when defined in assembly; an ordinary
name yields 0. FuncFlag carries the Asm bit, 4, for every assembly function.
### The pc-value tables
The AuxPcsp, AuxPcfile, AuxPcline, AuxPcinline and AuxPcdata payloads are
pc-value tables, each a sequence of value deltas and PC deltas:
- a signed value delta, zig-zag encoded, `binary.PutVarint` form;
- an unsigned PC delta in ULEB128 form, counted in instruction units, the
raw delta divided by the architecture's minimum instruction length;
- the table ends with a final PC delta to the end of the function followed by
a zero byte.
The first value applies from function entry. The encoding is the one
`cmd/internal/obj/pcln.go` calls funcpctab, and it is the same encoding the
final runtime pclntable carries.
### The DWARF payloads
AuxDwarfInfo, AuxDwarfLoc, AuxDwarfRanges and AuxDwarfLines reference SDWARF
symbols whose payloads are DWARF byte streams. The object format treats them
as opaque: the linker concatenates them into the final `.debug_*` sections
and resolves the relocations recorded inside them. The compiler produces
DWARF content per its own generation; gasm produces DWARF5 streams in
`asm/goobj_dwarf.go`.
## Builtins
Frequently referenced runtime functions are referenced by index rather than
by name: PkgIdxBuiltin with SymIdx set to the position in the generated table
`cmd/internal/goobj/builtinlist.go`, 299 entries in Go 1.27.1, names such as
`runtime.newobject` at index 0; 232 entries carry ABI 1 and the remaining 67
ABI 0. Builtin names never enter the string table. The mapping only applies
while the object is not linked against shared libraries, and a linkname'd
symbol never counts as a builtin even when its name matches.
## Fingerprints
The 8 byte fingerprint identifies one build of a package. The compiler fills
it with a hash of the package's export data; the assembler leaves it zero.
The linker checks a package's fingerprint against the fingerprints its
importers recorded in their Autolib entries and rejects a mismatched build,
which is how stale objects are caught.
## What a producer must do
The checklist a third-party writer must satisfy for `go build` to accept its
objects, in one place:
1. Write the container exactly: the `go object` line matching the target
toolchain's configuration string, the `!\n` terminator, then the blob.
2. Emit the 19 block offsets, in order, and make BlkEnd the blob length.
3. Deduplicate the string table, keep the empty string at offset 96, and
reference it everywhere a name appears.
4. Index relocations, aux entries and data per symbol with the N + 1 arrays,
definitions ordered Symdefs, Hashed64defs, Hasheddefs, NonPkgDefs.
5. Sort relocations by offset within each symbol.
6. Fill Siz with the true payload length, set Align for every
content-addressable symbol, and keep symbols under 2 GB.
7. Reference symbols by the package-index rules. An assembly producer
references everything outside the object by name, PkgIdxNone,
except its own file-local statics and the builtins; PkgIdxSelf is
reserved for definitions in this object. Assembly TEXT symbols
carry ABI 0.
8. Compute the content hashes exactly as the toolchain does, or emit no
hashed definitions at all.
## How gasm-devkit implements and verifies it
The writer lives in `asm/goobj.go`, which carries the shared container and the
amd64 relocation emission, with per-architecture relocation emitters in
`asm/goobjarm64.go`, `asm/goobjriscv.go` and `asm/goobjloong64.go`, symbol
resolution in `asm/goobj_resolve.go` and DWARF generation in
`asm/goobj_dwarf.go`. `gasm asm --format goobj -p pkg/path` writes objects
that `go build` consumes in place of the toolchain's own.
Verification is differential and continuous:
- `asm/goobj_test.go` compares gasm's GOOBJ output against `go tool asm`
output for the same source, byte for byte;
- `asm/link_test.go` builds real Go programs whose assembly comes from gasm
objects and runs them;
- `gasm verify` keeps the machine code itself identical to the toolchain's,
which is the precondition for the object comparison to be meaningful.
## Versioning and drift
The magic string carries the format generation, `go120ld` in Go 1.27.1. When
a toolchain release changes the format, it changes that string first, and the
linker refuses blobs whose magic it does not know. The watch points for a new
release are, in order: the magic, the block index list, the Aux type list,
the tail of the relocation table, the FuncInfo layout, and the builtin table
count. gasm's tests fail against any of these changes, which is the mechanism
that keeps this document and the writer current.
The authoritative sources, for the release this document covers:
- `cmd/internal/goobj/objfile.go`: the format, every structure in this
document;
- `cmd/internal/goobj/funcinfo.go`: FuncInfo and the inlining tree;
- `cmd/internal/goobj/builtinlist.go`: the builtin table;
- `cmd/internal/obj/objfile.go`: the writer, hash inputs and aux order;
- `cmd/internal/obj/sym.go`: package index assignment and the by-name rule;
- `cmd/internal/obj/pcln.go`: the pc-value encoding;
- `cmd/internal/objabi/reloctype.go`: relocation types;
- `cmd/internal/objabi/symkind.go`: symbol kinds;
- `cmd/link/internal/ld/lib.go`: container parsing and fingerprint checks.
+124
View File
@@ -0,0 +1,124 @@
# AMD64
Layer 1, target page. Verified against `go tool asm` of Go 1.27.1 and against
gasm's encoder, whose output is compared byte for byte with the toolchain's
and executed on real hardware (`gasm verify`). The complete mnemonic
inventory lives in the generated appendix
[INSTRUCTIONS-AMD64.md](INSTRUCTIONS-AMD64.md); this page is the grammar and
the conventions.
## Registers
| Group | Names | Notes |
|---|---|---|
| General purpose, 64-bit | `AX` `BX` `CX` `DX` `SI` `DI` `BP` `SP` `R8` to `R15` | bare names, no prefix |
| Sub-registers | `AL` `CL` `DL` `BL` `AH` family; `R8B` `R8W` `R8D` for the byte, word and double word of `R8` | width rides the mnemonic as well |
| Vector | `X0` to `X15` (128-bit), `Y0` to `Y15` (256-bit), `Z0` to `Z31` (512-bit) | SSE, AVX and AVX-512 |
| Mask | `K0` to `K7` | AVX-512 opmask |
| System | `TLS` | the thread pointer, see below |
Roles the calling convention fixes, which assembly must respect and can rely
on:
- `SP` is the hardware stack pointer; the virtual frame pointer of the
common language is the pseudo-register SP of OPERANDS.md, a different
spelling with a different meaning.
- `BP` is callee-save. The assembler inserts the save and restore whenever
the function has a non-zero frame, so using BP as a general register
interferes with sampling profilers that walk the frame chain.
- `R14` holds `g`, the goroutine pointer, in the register ABI; `RDX` holds
the closure context; `R12` and `R13` are the register ABI's scratch pair
and `R15` its GOT temporary; `X15` is the zeroing register the compiler
uses. An ABI0 assembly function called from Go sees none of these live
across the call, but runtime assembly reads them directly.
- The legacy spellings for the goroutine pointer are the macros of
`runtime/go_tls.h`: `get_tls(r)` expands to `MOVQ TLS, r` and `g(r)` to
`0(r)(TLS*1)`, the segment base riding the index field.
## Addressing
The common forms of OPERANDS.md, with the amd64 specifics:
```text
offset(base) MOVQ 16(BX), AX
offset(base)(index*scale) MOVL foo+32(SP)(R9*8), CX
scale is 1, 2, 4 or 8
name±offset(SB) MOVQ ·table(SB), CX
```
- Global references assemble as absolute addresses and produce R_ADDR
relocations; branch targets produce R_PCREL.
- Vector indexed memory, the VSIB form with an X, Y or Z register in the
index position, exists for the gather and scatter families.
- There are no segment overrides in source; the one segment-flavoured form
is the TLS base in the index field shown above.
## The frame and the split check
The assembler manages the frame, not the programmer:
- It inserts the `BP` save and restore for any non-zero frame.
- It inserts the stack-split check for any function that is not NoSplit:
the check compares SP against the guard, and on exhaustion calls
`runtime.morestack_noctxt`. Frames at or below 128 bytes, StackSmall, use
the small compare; frames at or below 4096 bytes, StackBig, use the
adjusted form; larger frames compare in two steps.
- On amd64 the assembler marks a function NoSplit itself when the frame is
under StackSmall and the body calls nothing that needs stack: such a
function carries the NoSplit flag in the object without the source ever
writing NOSPLIT.
Results and arguments are stack-only in ABI0: the caller's frame carries
them at FP offsets, per the Go prototype.
## Instructions
The inventory counts 1654 recognised mnemonics today, of which the encoder
emits 1113; both numbers are generated in the appendix, and the gap is the
encoder backlog that `gasm audit-instructions` measures. The families:
- **Integer base.** The ALU and move set with width suffixes, `MOVB`,
`MOVW`, `MOVL`, `MOVQ`; the extension moves `MOVBLZX`, `MOVWLSX`,
`MOVLQSX` and their siblings, which the compiler's output leans on;
`LEA`; `PUSH` and `POP`; the shifts and rotates; the bit operations `BT`
through `BTC`, `BSF`, `BSR`, `LZCNT`, `TZCNT`, `POPCNT`, `BSWAP`; the
string primitives `MOVS` and `STOS`.
- **Exchange and atomics.** `XCHG`, `CMPXCHG`, `XADD`; the extended-carry
pair `ADCX` and `ADOX`; `CRC32`.
- **Scalar floating point.** The SSE2 scalar moves and arithmetic
(`MOVSD`, `MOVSS`, `ADDSD`, and the `CVT` family). Floating-point
immediates are not encodable on this target, so the assembler
materialises them: the constant lands in a synthesised read-only pool,
and a positive zero collapses to `XORPS` of the register with itself,
exactly as the toolchain does.
- **Legacy SIMD, SSE.** The `MOVO`, `MOVOU`, `MOVAPS` family and the packed
integer and floating operations, shuffles, lane extracts and inserts and
the imm8-controlled forms.
- **VEX and EVEX.** The `V`-prefixed forms for 256 and 512-bit work,
opmask operations on `K0` to `K7`, gathers and scatters, and the
quad-register families 4FMAPS, 4FNMADD, 4VNNIW, VP4DPWSSD and VP4DPWSSDS,
whose register list rides the inverted V′VVV field. Mixing VEX and legacy
SSE in one loop pays the AVX-SSE transition penalty on every switch: keep
a loop in one dialect.
- **Cryptographic and counting extensions.** AES-NI, SHA-1 and SHA-256,
PCLMULQDQ, GFNI.
- **System.** `CPUID`, `RDTSC`, `SYSCALL`, the fences, `LDMXCSR` and
`STMXCSR`, the prefetch family.
- **Pseudo-operations.** `BYTE`, `WORD`, `LONG`, `QUAD` lay raw bytes or
words into the stream for encodings the assembler does not know; `ADJSP`
adjusts the stack pointer; `DUFFCOPY` and `DUFFZERO` and `GETCALLERPC`
are compiler-side names the table recognises but an encoder need not
emit.
A mnemonic the appendix lists with `gasm encodes: no` assembles nowhere:
gasm reports it as an explicit error, never as wrong bytes, and the
`unencodable-instruction` lint flags it at edit time.
## Relocations
The relocations an amd64 object carries, all specified in
[GOOBJ.md](../GOOBJ.md): `R_ADDR` for absolute globals, `R_PCREL` for
relative addresses, `R_CALL` for direct calls, `R_TLS_LE` and `R_TLS_IE` for
thread local access and `R_GOTPCREL` for GOT relative sequences, plus
`R_DWTXTADDR_U4` inside the DWARF records, which the assembler always
emits in the four-byte flavour.
+122
View File
@@ -0,0 +1,122 @@
# ARM64
Layer 1, target page. Verified against `go tool asm` of Go 1.27.1, against the
toolchain's own arm64 assembler manual (`cmd/internal/obj/arm64/doc.go`) and
against gasm's encoder, whose output is compared byte for byte with the
toolchain's. The complete mnemonic inventory lives in the generated appendix
[INSTRUCTIONS-ARM64.md](INSTRUCTIONS-ARM64.md).
## Registers
- General purpose: `R0` to `R30`, plus `ZR`, the zero register, and `RSP`,
the stack pointer. There is no R31: thirty-one names and ZR.
- Floating-point and SIMD share one file written `Vn`; where an instruction
is scalar floating point the operand may be written `Fn` (`F0` to `F31`).
- SVE register names (`Z0` to `Z31`, `P0` to `P15`) exist in the assembler's
tables.
- Roles the convention fixes: `RSP` is the stack pointer, `R29` the frame
pointer, `R30` the link register, `R26` the closure context and `R27` the
assembler's scratch register. The goroutine pointer lives in `R28` and is
written `g` in source, its fields as `g_m(g)`, `g_sched(g)`; `R18` is the
platform-reserved register and the Go toolchain never addresses it.
## Loads, stores and the width suffixes
The MOV series is the load and store interface, with the width in the
mnemonic rather than the register name:
| Mnemonic | Machine instruction |
|---|---|
| `MOVD` | ldr, str, stur, 64-bit |
| `MOVW` | ldrsw, str, stur, 32-bit sign extending |
| `MOVWU` | ldr, 32-bit zero extending |
| `MOVH` | ldrsh, strh, sturh |
| `MOVHU` | ldrh |
| `MOVB` | ldrsb, strb, sturb |
| `MOVBU` | ldrb |
Post-index and pre-index addressing take the `.P` and `.W` suffixes on the
mnemonic: `MOVD.P -8(R10), R8` is `ldr x8, [x10],#-8`, and `MOVB.W
16(R16), R10` is `ldrsb x10, [x16,#16]!`.
## Addressing
```text
imm(Rn|RSP) 28(R17)
(Rn|RSP) (R22)
(Rn)(Rm) (R27)(R23)
(Rn)(Rm<<scale) (R4)(R12<<2)
(Rn)(Rm.UXTW<<3) extended and shifted index
(Rt1, Rt2) register pair for LDP, STP and the exclusive pair forms
```
Branch targets are labels, `(R3)` for indirect, `name(SB)` for static.
## Operand order and the special forms
Most instructions appear in left-to-right assignment order: `ADD R11,
RSP, R25` computes into R25. The exceptions the toolchain's manual lists,
each with its own order:
- stores and `CBZ`, `CBNZ` keep the GNU order: `MOVD R29, 384(R19)`.
- The multiply-accumulate family `MADD`, `MSUB`, `SMADDL` and friends are
`<Rm>, <Ra>, <Rn>, <Rd>`.
- The scalar FMA family `FMADDD` and friends are `<Fm>, <Fa>, <Fn>, <Fd>`.
- The bitfield family `BFI`, `BFXIL`, `SBFIZ`, `SBFX`, `UBFIZ`, `UBFX` is
`$<lsb>, <Rn>, $<width>, <Rd>`.
- The conditional compare and select families carry the condition as the
**first** operand: `CSEL GT, R0, R19, R1`, `CCMP MI, R22, $12, $13`,
`FCCMPD AL, F8, F26, $0`.
- The exclusive stores are `<Rf>, (<Rn>), <Rs>` with the status register
last: `STLXR ZR, (R15), R16`.
- `TBZ` and `TBNZ` are `$<imm>, <Rt>, <label>`.
Shifted and extended register operands ride the register: `R19>>30`,
`R26->24` for arithmetic right shift, `@>` for rotate, and the extend forms
`R19.UXTB<<4`, `R14.SXTX` with extend operators UXTB, UXTH, UXTW, UXTX,
SXTB, SXTH, SXTW, SXTX.
## Conditions, branches and names
- Conditions ride the branch mnemonic: `B.EQ`, or the canonical
per-condition names such as `BEQ`. Both spellings exist; the canonical
names are what the generated inventory lists.
- `br` is `JMP` and `blr` is `CALL` in this dialect; indirect branches are
`JMP (R3)` and `CALL (R17)`.
- `NOP` is a zero-width pseudo-instruction; the hardware nop is `NOOP`,
an alias of `HINT $0`.
- `umov` is written as `VMOV`.
## Constants
- A 16-bit immediate optionally shifted: `MOVK $(10<<32), R20`, with
`MOVZ`, `MOVN` and their W variants; a zero shift is rejected by the
assembler.
- Large integer constants: `MOV` materialises any 64-bit constant, the
closest-instruction way.
- Vector constants: `VMOVS`, `VMOVD` and `VMOVQ`, the last taking two
64-bit halves for a 128-bit value:
`VMOVQ $0x1122334455667788, $0x99aabbccddeeff00, V2`.
## SIMD
Floating-point and SIMD instructions mostly carry a `V` prefix
(`VADD`, `VFMLA`), the cryptographic extensions (`AESD`, `SHA256H`) and the
scalar floating-point instructions being the exceptions. Operands carry an
arrangement suffix, `V5.H8`, and structure loads and stores use bracket
lists, `[V21.B16]`, with element selection as `V9.S[1]`.
## Alignment
`PCALIGN $n` pads to a power-of-two boundary between 8 and 2048 and also
raises the function's alignment to the coarsest boundary any of its PCALIGN
directives asks for. Functions default to 16-byte alignment on this target.
## Relocations
`R_ADDRARM64` for the adrp-plus-add pair, `R_ARM64_PCREL` and the
`R_ARM64_PCREL_LDST` family for PC relative addressing, `R_ARM64_LDST` for
the load and store immediates, `R_ARM64_GOTPCREL` and `R_ARM64_GOT` for the
GOT, `R_ARM64_TLS_LE` and `R_ARM64_TLS_IE` for thread local storage and
`R_CALLARM64` for direct calls, all specified in
[GOOBJ.md](../GOOBJ.md).
+148
View File
@@ -0,0 +1,148 @@
# Directives: TEXT, DATA, GLOBL and the annotations
Layer 1, the common language, with the flag vocabulary both layers share.
Verified against `go tool asm` of Go 1.27.1, against the shipped headers
`textflag.h` and `funcdata.h` in `$GOROOT/pkg/include`, and against gasm's
parser. Where gasm extends a directive, the extension says so and is marked.
Six directives exist. Three define things: TEXT, DATA, GLOBL. Three
annotate: FUNCDATA, PCDATA, PCALIGN.
## TEXT
```text
// func Add(a, b int64) int64
TEXT ·Add(SB), NOSPLIT, $0-24
...instructions...
RET
```
```text
TEXT symbol(SB), [flags,] $framesize[-argsize]
```
- The symbol is an `·Name(SB)` reference into the current package, or a
fully qualified name.
- The optional flag argument is a constant expression, normally an OR of the
names from `textflag.h`, the table below. Without `#include "textflag.h"`
the names are not macros and the assembler reports the misleading error
`illegal or missing addressing mode for symbol NOSPLIT`: include the
header first.
- `$framesize-argsize` is two constants, not a subtraction: the local frame
size in bytes, and the caller's argument area in bytes. The argument size
may be omitted entirely, `$16`, which marks the argument size unknown
(0x80000000 in the object, the value of `ArgsSizeUnknown` from
`funcdata.h`); a frame size may be negative only in the generated ABI
wrappers.
- A function whose last instruction is not a branch cannot fall through into
the next TEXT: the toolchain appends a jump to itself, so end functions
with `RET` deliberately.
- One TEXT per symbol; redeclaring is an error. The TEXT line also fixes the
function's source line for traceback: it is the line number that pcln
reports for the function's start.
The framesize and argsize fields do real work: the framesize drives the
stack-split preamble (RUNTIME.md carries the contract), and both travel into
the FuncInfo record of the object (GOOBJ.md carries its layout).
### The flag table
Values from `textflag.h`, in agreement with `cmd/internal/obj/textflag.go`:
| Name | Value | Applies to | Meaning |
|---|---|---|---|
| NOPROF | 1 | both | do not profile; deprecated |
| DUPOK | 2 | both | the linker may keep one of several duplicates |
| NOSPLIT | 4 | TEXT | no stack-split preamble |
| RODATA | 8 | data | put the data in a read-only section |
| NOPTR | 16 | data | the data contains no pointers |
| WRAPPER | 32 | TEXT | a wrapper; must not disable `recover` |
| NEEDCTXT | 64 | TEXT | a closure consuming the context register |
| TLSBSS | 256 | data | a thread local word in BSS |
| NOFRAME | 512 | TEXT | no frame setup; only valid with a frame size of 0 |
| REFLECTMETHOD | 1024 | TEXT | the function calls `reflect.Type.Method` or `MethodByName` |
| TOPFRAME | 2048 | TEXT | the outermost frame; unwinders stop here |
| ABIWRAPPER | 4096 | TEXT | an ABI transition wrapper |
Rules with teeth:
- `NOSPLIT` removes the split check, so the frame plus everything the
function calls must fit in the stack segment that remains. It exists to
protect the splitting code itself; reaching for it to save two instructions
is how stack overflows corrupt memory. On amd64 the assembler additionally
marks small leaf functions NoSplit itself and omits the check, so the
absence of the preamble is not proof the flag was written.
- A TEXT whose symbol is declared `ABIInternal` must carry NOSPLIT: the
assembler rejects it otherwise, because it cannot generate
the split path for a register-ABI function.
- `RODATA` implies NOPTR for the garbage collector.
## DATA
```text
DATA ·table+0(SB)/8, $0x0102030405060708
DATA ·msg+0(SB)/14, $"hello, world\n"
GLOBL ·msg(SB), RODATA, $14
```
```text
DATA symbol+offset(SB)/width, value
```
- `width` is exactly 1, 2, 4 or 8: the initialiser is written into the data
image at `symbol+offset` in that many bytes.
- The value is an integer or character constant of the width, or a string
literal whose byte length equals the width exactly; escapes count. Long
data is written as successive DATA lines at increasing offsets; bytes the
directives never name are zero.
- Every symbol initialised with DATA ends with a GLOBL line declaring its
total size, after all of its DATA lines.
A symbol containing pointers cannot be defined in assembly, because the
collector cannot see into it: define it in Go and refer to it by name. As a
rule, data that is not read-only belongs in Go.
Extension, gasm only: a DATA initialiser may name a symbol,
`DATA ·fn+0(SB)/8, $·handler(SB)`, which gasm lays down as an absolute
relocation on that field. The toolchain offers no ground truth for this
form; gasm's behaviour is verified by linking and execution.
## GLOBL
```text
GLOBL symbol(SB), [flags,] $size
```
Declares the symbol global with its total size in bytes. The useful flags
are RODATA, NOPTR, DUPOK and TLSBSS from the table above. Uninitialised
bytes are zero, which makes GLOBL with no DATA the language's BSS.
## FUNCDATA and PCDATA
```text
FUNCDATA $functypeid, symbol(SB)
PCDATA $pctypeid, $value
```
The compiler's annotations for the garbage collector and traceback, named by
the ids in `funcdata.h`: FUNCDATA 0 to 7 (args pointer maps, locals pointer
maps, stack objects, inline tree, open-coded defer info, argument info,
argument liveness, wrap info), PCDATA 0 to 4 (unsafe point, stack map index,
inline tree index, argument liveness index, panic bounds). Assembly code
normally reaches them only through the macro forms in `funcdata.h`, which
RUNTIME.md explains. Outside the macros, hand-written PCDATA is meaningless:
the values are pc-value tables the compiler builds from its own view of the
program.
## PCALIGN
```text
PCALIGN $32
```
Pads the code so that the next instruction lands on the given boundary,
which must be a power of two and at least the target's instruction
alignment. Supported on amd64, arm64, ppc64, loong64 and riscv64. The
padding instructions are the target's NOP encoding, so the bytes between
functions differ from what the instruction stream alone would produce, which
matters to anyone comparing encodings byte for byte.
File diff suppressed because it is too large Load Diff
+570
View File
@@ -0,0 +1,570 @@
# ARM64: instruction inventory
Generated by gasm-devkit's `_gen` from the Go toolchain's instruction table
(`cmd/internal/obj/arm64/anames.go`, go1.27.1); DO NOT EDIT. This page lists every mnemonic
`go tool asm` accepts on this target, which is the upper bound of the
language on it: a name absent here is not an instruction of the target,
and a name present here may still be one gasm's encoder cannot emit yet.
The inventory carries no per-mnemonic encoder column: on this target
encodability is decided per operand shape, and the live measured
coverage is reported by `gasm audit-instructions`.
| Mnemonic | Notes |
|---|---|
| `CALL` | |
| `DUFFCOPY` | |
| `DUFFZERO` | |
| `END` | |
| `FUNCDATA` | |
| `GETCALLERPC` | |
| `JMP` | |
| `NOP` | No operation |
| `PCALIGN` | |
| `PCALIGNMAX` | |
| `PCDATA` | |
| `RET` | Return |
| `TEXT` | |
| `UNDEF` | |
| `ADC` | ADC (64-bit) |
| `ADCS` | ADCS (64-bit) |
| `ADCSW` | ADCS (32-bit) |
| `ADCW` | ADC (32-bit) |
| `ADD` | ADD (64-bit) |
| `ADDS` | ADDS (64-bit) |
| `ADDSW` | ADDS (32-bit) |
| `ADDW` | ADD (32-bit) |
| `ADR` | Address of label/page |
| `ADRP` | Address of label/page |
| `AESD` | AES round |
| `AESE` | AES round |
| `AESIMC` | AES round |
| `AESMC` | AES round |
| `AND` | AND (64-bit) |
| `ANDS` | ANDS (64-bit) |
| `ANDSW` | ANDS (32-bit) |
| `ANDW` | AND (32-bit) |
| `ASR` | ASR shift |
| `ASRW` | ASR shift (32-bit) |
| `AT` | |
| `AUTIA1716` | |
| `AUTIASP` | |
| `AUTIB1716` | |
| `AUTIBSP` | |
| `BCC` | Conditional branch |
| `BCS` | Conditional branch |
| `BEQ` | Conditional branch |
| `BFI` | |
| `BFIW` | |
| `BFM` | |
| `BFMW` | |
| `BFXIL` | Bitfield extract |
| `BFXILW` | |
| `BGE` | Conditional branch |
| `BGT` | Conditional branch |
| `BHI` | Conditional branch |
| `BHS` | Conditional branch |
| `BIC` | BIC (64-bit) |
| `BICS` | BICS (64-bit) |
| `BICSW` | BICS (32-bit) |
| `BICW` | BIC (32-bit) |
| `BLE` | Conditional branch |
| `BLO` | Conditional branch |
| `BLS` | Conditional branch |
| `BLT` | Conditional branch |
| `BMI` | Conditional branch |
| `BNE` | Conditional branch |
| `BPL` | Conditional branch |
| `BRK` | Breakpoint |
| `BTI` | |
| `BVC` | Conditional branch |
| `BVS` | Conditional branch |
| `CASAD` | |
| `CASALB` | |
| `CASALD` | |
| `CASALH` | |
| `CASALW` | |
| `CASAW` | |
| `CASB` | |
| `CASD` | |
| `CASH` | |
| `CASLD` | |
| `CASLW` | |
| `CASPD` | |
| `CASPW` | |
| `CASW` | |
| `CBNZ` | Compare/test and branch |
| `CBNZW` | Compare/test and branch (32-bit) |
| `CBZ` | Compare/test and branch |
| `CBZW` | Compare/test and branch (32-bit) |
| `CCMN` | Conditional compare |
| `CCMNW` | Conditional compare |
| `CCMP` | Conditional compare |
| `CCMPW` | Conditional compare |
| `CINC` | Conditional select |
| `CINCW` | Conditional select (32-bit) |
| `CINV` | Conditional select |
| `CINVW` | Conditional select (32-bit) |
| `CLREX` | |
| `CLS` | Bit manipulation |
| `CLSW` | Bit manipulation |
| `CLZ` | Bit manipulation |
| `CLZW` | Bit manipulation |
| `CMN` | CMN (64-bit) |
| `CMNW` | CMN (32-bit) |
| `CMP` | CMP (64-bit) |
| `CMPW` | CMP (32-bit) |
| `CNEG` | Conditional select |
| `CNEGW` | Conditional select (32-bit) |
| `CRC32B` | |
| `CRC32CB` | |
| `CRC32CH` | |
| `CRC32CW` | |
| `CRC32CX` | |
| `CRC32H` | |
| `CRC32W` | |
| `CRC32X` | |
| `CSEL` | Conditional select |
| `CSELW` | Conditional select (32-bit) |
| `CSET` | Conditional select |
| `CSETM` | Conditional select |
| `CSETMW` | Conditional select (32-bit) |
| `CSETW` | Conditional select (32-bit) |
| `CSINC` | Conditional select |
| `CSINCW` | Conditional select (32-bit) |
| `CSINV` | Conditional select |
| `CSINVW` | Conditional select (32-bit) |
| `CSNEG` | Conditional select |
| `CSNEGW` | Conditional select (32-bit) |
| `DC` | Data cache maintenance |
| `DCPS1` | |
| `DCPS2` | |
| `DCPS3` | |
| `DMB` | Barrier |
| `DRPS` | |
| `DSB` | Barrier |
| `DWORD` | |
| `EON` | EON (64-bit) |
| `EONW` | EON (32-bit) |
| `EOR` | EOR (64-bit) |
| `EORW` | EOR (32-bit) |
| `ERET` | |
| `EXTR` | Bitfield extract |
| `EXTRW` | |
| `FABSD` | |
| `FABSS` | |
| `FADDD` | |
| `FADDS` | |
| `FCCMPD` | |
| `FCCMPED` | |
| `FCCMPES` | |
| `FCCMPS` | |
| `FCMPD` | |
| `FCMPED` | |
| `FCMPES` | |
| `FCMPS` | |
| `FCSELD` | |
| `FCSELS` | |
| `FCVTDH` | |
| `FCVTDS` | |
| `FCVTHD` | |
| `FCVTHS` | |
| `FCVTSD` | |
| `FCVTSH` | |
| `FCVTZSD` | |
| `FCVTZSDW` | |
| `FCVTZSS` | |
| `FCVTZSSW` | |
| `FCVTZUD` | |
| `FCVTZUDW` | |
| `FCVTZUS` | |
| `FCVTZUSW` | |
| `FDIVD` | |
| `FDIVS` | |
| `FLDPD` | Register-pair load or store |
| `FLDPQ` | |
| `FLDPS` | |
| `FMADDD` | |
| `FMADDS` | |
| `FMAXD` | |
| `FMAXNMD` | |
| `FMAXNMS` | |
| `FMAXS` | |
| `FMIND` | |
| `FMINNMD` | |
| `FMINNMS` | |
| `FMINS` | |
| `FMOVD` | Move / load / store |
| `FMOVQ` | |
| `FMOVS` | Move / load / store |
| `FMSUBD` | |
| `FMSUBS` | |
| `FMULD` | |
| `FMULS` | |
| `FNEGD` | |
| `FNEGS` | |
| `FNMADDD` | |
| `FNMADDS` | |
| `FNMSUBD` | |
| `FNMSUBS` | |
| `FNMULD` | |
| `FNMULS` | |
| `FRINTAD` | |
| `FRINTAS` | |
| `FRINTID` | |
| `FRINTIS` | |
| `FRINTMD` | |
| `FRINTMS` | |
| `FRINTND` | |
| `FRINTNS` | |
| `FRINTPD` | |
| `FRINTPS` | |
| `FRINTXD` | |
| `FRINTXS` | |
| `FRINTZD` | |
| `FRINTZS` | |
| `FSQRTD` | |
| `FSQRTS` | |
| `FSTPD` | Register-pair load or store |
| `FSTPQ` | |
| `FSTPS` | |
| `FSUBD` | |
| `FSUBS` | |
| `HINT` | |
| `HLT` | |
| `HVC` | Exception generation |
| `IC` | |
| `ISB` | Barrier |
| `LDADDAB` | |
| `LDADDAD` | |
| `LDADDAH` | |
| `LDADDALB` | |
| `LDADDALD` | |
| `LDADDALH` | |
| `LDADDALW` | |
| `LDADDAW` | |
| `LDADDB` | |
| `LDADDD` | |
| `LDADDH` | |
| `LDADDLB` | |
| `LDADDLD` | |
| `LDADDLH` | |
| `LDADDLW` | |
| `LDADDW` | |
| `LDAR` | Atomic memory operation |
| `LDARB` | Atomic memory operation |
| `LDARH` | Atomic memory operation |
| `LDARW` | Atomic memory operation |
| `LDAXP` | |
| `LDAXPW` | |
| `LDAXR` | Atomic memory operation |
| `LDAXRB` | Atomic memory operation |
| `LDAXRH` | Atomic memory operation |
| `LDAXRW` | Atomic memory operation |
| `LDCLRAB` | |
| `LDCLRAD` | |
| `LDCLRAH` | |
| `LDCLRALB` | |
| `LDCLRALD` | |
| `LDCLRALH` | |
| `LDCLRALW` | |
| `LDCLRAW` | |
| `LDCLRB` | |
| `LDCLRD` | |
| `LDCLRH` | |
| `LDCLRLB` | |
| `LDCLRLD` | |
| `LDCLRLH` | |
| `LDCLRLW` | |
| `LDCLRW` | |
| `LDEORAB` | |
| `LDEORAD` | |
| `LDEORAH` | |
| `LDEORALB` | |
| `LDEORALD` | |
| `LDEORALH` | |
| `LDEORALW` | |
| `LDEORAW` | |
| `LDEORB` | |
| `LDEORD` | |
| `LDEORH` | |
| `LDEORLB` | |
| `LDEORLD` | |
| `LDEORLH` | |
| `LDEORLW` | |
| `LDEORW` | |
| `LDORAB` | |
| `LDORAD` | |
| `LDORAH` | |
| `LDORALB` | |
| `LDORALD` | |
| `LDORALH` | |
| `LDORALW` | |
| `LDORAW` | |
| `LDORB` | |
| `LDORD` | |
| `LDORH` | |
| `LDORLB` | |
| `LDORLD` | |
| `LDORLH` | |
| `LDORLW` | |
| `LDORW` | |
| `LDP` | Register-pair load or store |
| `LDPSW` | |
| `LDPW` | Register-pair load or store |
| `LDXP` | |
| `LDXPW` | |
| `LDXR` | |
| `LDXRB` | |
| `LDXRH` | |
| `LDXRW` | |
| `LSL` | LSL shift |
| `LSLW` | LSL shift (32-bit) |
| `LSR` | LSR shift |
| `LSRW` | LSR shift (32-bit) |
| `MADD` | Multiply / multiply-accumulate |
| `MADDW` | |
| `MNEG` | Multiply / multiply-accumulate |
| `MNEGW` | |
| `MOVB` | Move / load / store |
| `MOVBU` | Move / load / store |
| `MOVD` | Move / load / store |
| `MOVH` | Move / load / store |
| `MOVHU` | Move / load / store |
| `MOVK` | Move wide constant |
| `MOVKW` | Move wide constant |
| `MOVN` | Move wide constant |
| `MOVNW` | Move wide constant |
| `MOVP` | |
| `MOVPD` | |
| `MOVPQ` | |
| `MOVPS` | |
| `MOVPSW` | |
| `MOVPW` | |
| `MOVW` | Move / load / store |
| `MOVWU` | Move / load / store |
| `MOVZ` | Move wide constant |
| `MOVZW` | Move wide constant |
| `MRS` | System register access |
| `MSR` | System register access |
| `MSUB` | Multiply / multiply-accumulate |
| `MSUBW` | |
| `MUL` | Multiply / multiply-accumulate |
| `MULW` | |
| `MVN` | MVN (64-bit) |
| `MVNW` | MVN (32-bit) |
| `NEG` | NEG (64-bit) |
| `NEGS` | |
| `NEGSW` | |
| `NEGW` | NEG (32-bit) |
| `NGC` | NGC (64-bit) |
| `NGCS` | |
| `NGCSW` | |
| `NGCW` | NGC (32-bit) |
| `NOOP` | |
| `ORN` | ORN (64-bit) |
| `ORNW` | ORN (32-bit) |
| `ORR` | ORR (64-bit) |
| `ORRW` | ORR (32-bit) |
| `PACIASP` | |
| `PACIBSP` | |
| `PRFM` | Memory prefetch |
| `PRFUM` | |
| `RBIT` | Bit manipulation |
| `RBITW` | Bit manipulation |
| `REM` | |
| `REMW` | |
| `REV` | Bit manipulation |
| `REV16` | Bit manipulation |
| `REV16W` | |
| `REV32` | Bit manipulation |
| `REVW` | Bit manipulation |
| `ROR` | ROR shift |
| `RORW` | ROR shift (32-bit) |
| `SBC` | SBC (64-bit) |
| `SBCS` | SBCS (64-bit) |
| `SBCSW` | SBCS (32-bit) |
| `SBCW` | SBC (32-bit) |
| `SBFIZ` | |
| `SBFIZW` | |
| `SBFM` | Bitfield extract |
| `SBFMW` | |
| `SBFX` | Bitfield extract |
| `SBFXW` | |
| `SCVTFD` | |
| `SCVTFS` | |
| `SCVTFWD` | |
| `SCVTFWS` | |
| `SDIV` | Divide |
| `SDIVW` | Divide |
| `SEV` | |
| `SEVL` | |
| `SHA1C` | SHA round |
| `SHA1H` | SHA round |
| `SHA1M` | SHA round |
| `SHA1P` | SHA round |
| `SHA1SU0` | SHA round |
| `SHA1SU1` | SHA round |
| `SHA256H` | SHA round |
| `SHA256H2` | SHA round |
| `SHA256SU0` | SHA round |
| `SHA256SU1` | SHA round |
| `SHA512H` | SHA round |
| `SHA512H2` | SHA round |
| `SHA512SU0` | SHA round |
| `SHA512SU1` | SHA round |
| `SMADDL` | Multiply / multiply-accumulate |
| `SMC` | Exception generation |
| `SMNEGL` | |
| `SMSUBL` | Multiply / multiply-accumulate |
| `SMULH` | Multiply / multiply-accumulate |
| `SMULL` | Multiply / multiply-accumulate |
| `STLR` | Atomic memory operation |
| `STLRB` | Atomic memory operation |
| `STLRH` | Atomic memory operation |
| `STLRW` | Atomic memory operation |
| `STLXP` | |
| `STLXPW` | |
| `STLXR` | |
| `STLXRB` | |
| `STLXRH` | |
| `STLXRW` | |
| `STP` | Register-pair load or store |
| `STPW` | Register-pair load or store |
| `STXP` | |
| `STXPW` | |
| `STXR` | Atomic memory operation |
| `STXRB` | Atomic memory operation |
| `STXRH` | Atomic memory operation |
| `STXRW` | Atomic memory operation |
| `SUB` | SUB (64-bit) |
| `SUBS` | SUBS (64-bit) |
| `SUBSW` | SUBS (32-bit) |
| `SUBW` | SUB (32-bit) |
| `SVC` | Exception generation |
| `SWPAB` | |
| `SWPAD` | |
| `SWPAH` | |
| `SWPALB` | |
| `SWPALD` | |
| `SWPALH` | |
| `SWPALW` | |
| `SWPAW` | |
| `SWPB` | |
| `SWPD` | |
| `SWPH` | |
| `SWPLB` | |
| `SWPLD` | |
| `SWPLH` | |
| `SWPLW` | |
| `SWPW` | |
| `SXTB` | |
| `SXTBW` | |
| `SXTH` | |
| `SXTHW` | |
| `SXTW` | |
| `SYS` | |
| `SYSL` | |
| `TBNZ` | Compare/test and branch |
| `TBZ` | Compare/test and branch |
| `TLBI` | |
| `TST` | TST (64-bit) |
| `TSTW` | TST (32-bit) |
| `UBFIZ` | |
| `UBFIZW` | |
| `UBFM` | Bitfield extract |
| `UBFMW` | |
| `UBFX` | Bitfield extract |
| `UBFXW` | |
| `UCVTFD` | |
| `UCVTFS` | |
| `UCVTFWD` | |
| `UCVTFWS` | |
| `UDIV` | Divide |
| `UDIVW` | Divide |
| `UMADDL` | Multiply / multiply-accumulate |
| `UMNEGL` | |
| `UMSUBL` | Multiply / multiply-accumulate |
| `UMULH` | Multiply / multiply-accumulate |
| `UMULL` | Multiply / multiply-accumulate |
| `UREM` | |
| `UREMW` | |
| `UXTB` | |
| `UXTBW` | |
| `UXTH` | |
| `UXTHW` | |
| `UXTW` | |
| `VADD` | NEON SIMD vector operation |
| `VADDP` | |
| `VADDV` | NEON SIMD vector operation |
| `VAND` | NEON SIMD vector operation |
| `VBCAX` | Three-way XOR / rotate crypto vector operation |
| `VBIF` | NEON SIMD vector operation |
| `VBIT` | |
| `VBSL` | NEON SIMD vector operation |
| `VCMEQ` | |
| `VCMTST` | |
| `VCNT` | NEON SIMD vector operation |
| `VDUP` | NEON SIMD vector operation |
| `VEOR` | NEON SIMD vector operation |
| `VEOR3` | Three-way XOR / rotate crypto vector operation |
| `VEXT` | NEON SIMD vector operation |
| `VFMLA` | NEON SIMD vector operation |
| `VFMLS` | NEON SIMD vector operation |
| `VLD1` | NEON SIMD vector operation |
| `VLD1R` | |
| `VLD2` | NEON SIMD vector operation |
| `VLD2R` | |
| `VLD3` | NEON SIMD vector operation |
| `VLD3R` | |
| `VLD4` | NEON SIMD vector operation |
| `VLD4R` | |
| `VMOV` | NEON SIMD vector operation |
| `VMOVD` | |
| `VMOVI` | NEON SIMD vector operation |
| `VMOVQ` | NEON SIMD vector operation |
| `VMOVS` | |
| `VORR` | NEON SIMD vector operation |
| `VPMULL` | |
| `VPMULL2` | |
| `VRAX1` | Three-way XOR / rotate crypto vector operation |
| `VRBIT` | |
| `VREV16` | NEON SIMD vector operation |
| `VREV32` | NEON SIMD vector operation |
| `VREV64` | NEON SIMD vector operation |
| `VSHL` | NEON SIMD vector operation |
| `VSLI` | |
| `VSRI` | |
| `VST1` | NEON SIMD vector operation |
| `VST2` | NEON SIMD vector operation |
| `VST3` | NEON SIMD vector operation |
| `VST4` | NEON SIMD vector operation |
| `VSUB` | NEON SIMD vector operation |
| `VTBL` | NEON SIMD vector operation |
| `VTBX` | NEON SIMD vector operation |
| `VTRN1` | NEON SIMD vector operation |
| `VTRN2` | NEON SIMD vector operation |
| `VUADDLV` | |
| `VUADDW` | |
| `VUADDW2` | |
| `VUMAX` | |
| `VUMIN` | |
| `VUSHLL` | |
| `VUSHLL2` | |
| `VUSHR` | NEON SIMD vector operation |
| `VUSRA` | |
| `VUXTL` | |
| `VUXTL2` | |
| `VUZP1` | NEON SIMD vector operation |
| `VUZP2` | NEON SIMD vector operation |
| `VXAR` | Three-way XOR / rotate crypto vector operation |
| `VZIP1` | NEON SIMD vector operation |
| `VZIP2` | NEON SIMD vector operation |
| `WFE` | |
| `WFI` | |
| `WORD` | |
| `YIELD` | |
| `B` | Unconditional branch |
| `BL` | Branch with link |
Recognised: 554 mnemonics.
+830
View File
@@ -0,0 +1,830 @@
# LoongArch 64: instruction inventory
Generated by gasm-devkit's `_gen` from the Go toolchain's instruction table
(`cmd/internal/obj/loong64/anames.go`, go1.27.1); DO NOT EDIT. This page lists every mnemonic
`go tool asm` accepts on this target, which is the upper bound of the
language on it: a name absent here is not an instruction of the target,
and a name present here may still be one gasm's encoder cannot emit yet.
The inventory carries no per-mnemonic encoder column: on this target
encodability is decided per operand shape, and the live measured
coverage is reported by `gasm audit-instructions`.
| Mnemonic | Notes |
|---|---|
| `CALL` | |
| `DUFFCOPY` | |
| `DUFFZERO` | |
| `END` | |
| `FUNCDATA` | |
| `GETCALLERPC` | |
| `JMP` | |
| `NOP` | No operation |
| `PCALIGN` | |
| `PCALIGNMAX` | |
| `PCDATA` | |
| `RET` | Return |
| `TEXT` | |
| `UNDEF` | |
| `ABSD` | |
| `ABSF` | |
| `ADD` | Integer add (word) |
| `ADDD` | Add doubleword |
| `ADDF` | |
| `ADDV` | |
| `ADDV16` | |
| `ADDVU` | |
| `ADDW` | Add word |
| `ALSLV` | |
| `ALSLW` | |
| `ALSLWU` | |
| `AMADDDBV` | |
| `AMADDDBW` | |
| `AMADDV` | |
| `AMADDW` | |
| `AMANDDBV` | |
| `AMANDDBW` | |
| `AMANDV` | |
| `AMANDW` | |
| `AMCASB` | |
| `AMCASDBB` | |
| `AMCASDBH` | |
| `AMCASDBV` | |
| `AMCASDBW` | |
| `AMCASH` | |
| `AMCASV` | |
| `AMCASW` | |
| `AMMAXDBV` | |
| `AMMAXDBVU` | |
| `AMMAXDBW` | |
| `AMMAXDBWU` | |
| `AMMAXV` | |
| `AMMAXVU` | |
| `AMMAXW` | |
| `AMMAXWU` | |
| `AMMINDBV` | |
| `AMMINDBVU` | |
| `AMMINDBW` | |
| `AMMINDBWU` | |
| `AMMINV` | |
| `AMMINVU` | |
| `AMMINW` | |
| `AMMINWU` | |
| `AMORDBV` | |
| `AMORDBW` | |
| `AMORV` | |
| `AMORW` | |
| `AMSWAPB` | |
| `AMSWAPDBB` | |
| `AMSWAPDBH` | |
| `AMSWAPDBV` | |
| `AMSWAPDBW` | |
| `AMSWAPH` | |
| `AMSWAPV` | |
| `AMSWAPW` | |
| `AMXORDBV` | |
| `AMXORDBW` | |
| `AMXORV` | |
| `AMXORW` | |
| `AND` | Bitwise AND |
| `ANDN` | |
| `BEQ` | Branch if equal |
| `BFPF` | |
| `BFPT` | |
| `BGE` | Branch if greater or equal |
| `BGEU` | Branch if greater or equal unsigned |
| `BGEZ` | |
| `BGTZ` | |
| `BITREV4B` | |
| `BITREV8B` | |
| `BITREVV` | |
| `BITREVW` | |
| `BLEZ` | |
| `BLT` | Branch if less than |
| `BLTU` | Branch if less than unsigned |
| `BLTZ` | |
| `BNE` | Branch if not equal |
| `BREAK` | Breakpoint |
| `BSTRINSV` | |
| `BSTRINSW` | |
| `BSTRPICKV` | |
| `BSTRPICKW` | |
| `CLOV` | |
| `CLOW` | |
| `CLZV` | |
| `CLZW` | |
| `CMPEQD` | |
| `CMPEQF` | |
| `CMPGED` | |
| `CMPGEF` | |
| `CMPGTD` | |
| `CMPGTF` | |
| `CPUCFG` | |
| `CRCCWBW` | |
| `CRCCWHW` | |
| `CRCCWVW` | |
| `CRCCWWW` | |
| `CRCWBW` | |
| `CRCWHW` | |
| `CRCWVW` | |
| `CRCWWW` | |
| `CTOV` | |
| `CTOW` | |
| `CTZV` | |
| `CTZW` | |
| `DBAR` | Barrier |
| `DIV` | Divide (word) |
| `DIVD` | Divide doubleword |
| `DIVF` | |
| `DIVU` | |
| `DIVV` | |
| `DIVVU` | |
| `DIVW` | Divide word |
| `DIVWU` | |
| `EXTWB` | |
| `EXTWH` | |
| `FCLASSD` | |
| `FCLASSF` | |
| `FCOPYSGD` | |
| `FCOPYSGF` | |
| `FFINTDV` | |
| `FFINTDW` | |
| `FFINTFV` | |
| `FFINTFW` | |
| `FLOGBD` | |
| `FLOGBF` | |
| `FMADDD` | |
| `FMADDF` | |
| `FMAXAD` | |
| `FMAXAF` | |
| `FMAXD` | |
| `FMAXF` | |
| `FMINAD` | |
| `FMINAF` | |
| `FMIND` | |
| `FMINF` | |
| `FMSUBD` | |
| `FMSUBF` | |
| `FNMADDD` | |
| `FNMADDF` | |
| `FNMSUBD` | |
| `FNMSUBF` | |
| `FSCALEBD` | |
| `FSCALEBF` | |
| `FSEL` | |
| `FTINTRMVD` | |
| `FTINTRMVF` | |
| `FTINTRMWD` | |
| `FTINTRMWF` | |
| `FTINTRNEVD` | |
| `FTINTRNEVF` | |
| `FTINTRNEWD` | |
| `FTINTRNEWF` | |
| `FTINTRPVD` | |
| `FTINTRPVF` | |
| `FTINTRPWD` | |
| `FTINTRPWF` | |
| `FTINTRZVD` | |
| `FTINTRZVF` | |
| `FTINTRZWD` | |
| `FTINTRZWF` | |
| `FTINTVD` | |
| `FTINTVF` | |
| `FTINTWD` | |
| `FTINTWF` | |
| `JIRL` | Jump indirect with link |
| `LL` | |
| `LLV` | |
| `LU12IW` | |
| `LU32ID` | |
| `LU52ID` | |
| `LUI` | |
| `MASKEQZ` | |
| `MASKNEZ` | |
| `MOVB` | |
| `MOVBU` | |
| `MOVD` | |
| `MOVDF` | |
| `MOVDV` | |
| `MOVDW` | |
| `MOVF` | |
| `MOVFD` | |
| `MOVFV` | |
| `MOVFW` | |
| `MOVH` | |
| `MOVHU` | |
| `MOVV` | |
| `MOVVD` | |
| `MOVVF` | |
| `MOVVP` | |
| `MOVW` | |
| `MOVWD` | |
| `MOVWF` | |
| `MOVWP` | |
| `MOVWU` | |
| `MUL` | Multiply (word) |
| `MULD` | Multiply doubleword |
| `MULF` | |
| `MULH` | |
| `MULHU` | |
| `MULHV` | |
| `MULHVU` | |
| `MULV` | |
| `MULVU` | |
| `MULW` | Multiply word |
| `MULWVW` | |
| `MULWVWU` | |
| `NEGD` | |
| `NEGF` | |
| `NEGV` | |
| `NEGW` | |
| `NOOP` | |
| `NOR` | Bitwise NOR |
| `OR` | Bitwise OR |
| `ORN` | |
| `PCADDU12I` | |
| `PCALAU12I` | |
| `PRELD` | |
| `PRELDX` | |
| `RDTIMED` | |
| `RDTIMEHW` | |
| `RDTIMELW` | |
| `REM` | |
| `REMU` | |
| `REMV` | |
| `REMVU` | |
| `REMW` | |
| `REMWU` | |
| `REVB2H` | |
| `REVB2W` | |
| `REVB4H` | |
| `REVBV` | |
| `REVH2W` | |
| `REVHV` | |
| `RFE` | |
| `ROTR` | Rotate right |
| `ROTRV` | |
| `SC` | |
| `SCV` | |
| `SGT` | |
| `SGTU` | |
| `SLL` | Shift left logical |
| `SLLV` | |
| `SQRTD` | |
| `SQRTF` | |
| `SRA` | Shift right arithmetic |
| `SRAV` | |
| `SRL` | Shift right logical |
| `SRLV` | |
| `SUB` | Subtract (word) |
| `SUBD` | Subtract doubleword |
| `SUBF` | |
| `SUBV` | |
| `SUBVU` | |
| `SUBW` | Subtract word |
| `SYSCALL` | System call |
| `TEQ` | |
| `TNE` | |
| `TRUNCDV` | |
| `TRUNCDW` | |
| `TRUNCFV` | |
| `TRUNCFW` | |
| `VADDB` | |
| `VADDBU` | |
| `VADDD` | |
| `VADDF` | |
| `VADDH` | |
| `VADDHU` | |
| `VADDQ` | |
| `VADDV` | |
| `VADDVU` | |
| `VADDW` | |
| `VADDWEVHB` | |
| `VADDWEVHBU` | |
| `VADDWEVQV` | |
| `VADDWEVQVU` | |
| `VADDWEVVW` | |
| `VADDWEVVWU` | |
| `VADDWEVWH` | |
| `VADDWEVWHU` | |
| `VADDWODHB` | |
| `VADDWODHBU` | |
| `VADDWODQV` | |
| `VADDWODQVU` | |
| `VADDWODVW` | |
| `VADDWODVWU` | |
| `VADDWODWH` | |
| `VADDWODWHU` | |
| `VADDWU` | |
| `VANDB` | |
| `VANDNV` | |
| `VANDV` | |
| `VBITCLRB` | |
| `VBITCLRH` | |
| `VBITCLRV` | |
| `VBITCLRW` | |
| `VBITREVB` | |
| `VBITREVH` | |
| `VBITREVV` | |
| `VBITREVW` | |
| `VBITSETB` | |
| `VBITSETH` | |
| `VBITSETV` | |
| `VBITSETW` | |
| `VDIVB` | |
| `VDIVBU` | |
| `VDIVD` | |
| `VDIVF` | |
| `VDIVH` | |
| `VDIVHU` | |
| `VDIVV` | |
| `VDIVVU` | |
| `VDIVW` | |
| `VDIVWU` | |
| `VEXTRINSB` | |
| `VEXTRINSH` | |
| `VEXTRINSV` | |
| `VEXTRINSW` | |
| `VFCLASSD` | |
| `VFCLASSF` | |
| `VFRECIPD` | |
| `VFRECIPF` | |
| `VFRINTD` | |
| `VFRINTF` | |
| `VFRINTRMD` | |
| `VFRINTRMF` | |
| `VFRINTRNED` | |
| `VFRINTRNEF` | |
| `VFRINTRPD` | |
| `VFRINTRPF` | |
| `VFRINTRZD` | |
| `VFRINTRZF` | |
| `VFRSQRTD` | |
| `VFRSQRTF` | |
| `VFSQRTD` | |
| `VFSQRTF` | |
| `VILVHB` | |
| `VILVHH` | |
| `VILVHV` | |
| `VILVHW` | |
| `VILVLB` | |
| `VILVLH` | |
| `VILVLV` | |
| `VILVLW` | |
| `VMADDB` | |
| `VMADDH` | |
| `VMADDV` | |
| `VMADDW` | |
| `VMADDWEVHB` | |
| `VMADDWEVHBU` | |
| `VMADDWEVHBUB` | |
| `VMADDWEVQV` | |
| `VMADDWEVQVU` | |
| `VMADDWEVQVUV` | |
| `VMADDWEVVW` | |
| `VMADDWEVVWU` | |
| `VMADDWEVVWUW` | |
| `VMADDWEVWH` | |
| `VMADDWEVWHU` | |
| `VMADDWEVWHUH` | |
| `VMADDWODHB` | |
| `VMADDWODHBU` | |
| `VMADDWODHBUB` | |
| `VMADDWODQV` | |
| `VMADDWODQVU` | |
| `VMADDWODQVUV` | |
| `VMADDWODVW` | |
| `VMADDWODVWU` | |
| `VMADDWODVWUW` | |
| `VMADDWODWH` | |
| `VMADDWODWHU` | |
| `VMADDWODWHUH` | |
| `VMODB` | |
| `VMODBU` | |
| `VMODH` | |
| `VMODHU` | |
| `VMODV` | |
| `VMODVU` | |
| `VMODW` | |
| `VMODWU` | |
| `VMOVQ` | |
| `VMSUBB` | |
| `VMSUBH` | |
| `VMSUBV` | |
| `VMSUBW` | |
| `VMUHB` | |
| `VMUHBU` | |
| `VMUHH` | |
| `VMUHHU` | |
| `VMUHV` | |
| `VMUHVU` | |
| `VMUHW` | |
| `VMUHWU` | |
| `VMULB` | |
| `VMULD` | |
| `VMULF` | |
| `VMULH` | |
| `VMULV` | |
| `VMULW` | |
| `VMULWEVHB` | |
| `VMULWEVHBU` | |
| `VMULWEVHBUB` | |
| `VMULWEVQV` | |
| `VMULWEVQVU` | |
| `VMULWEVQVUV` | |
| `VMULWEVVW` | |
| `VMULWEVVWU` | |
| `VMULWEVVWUW` | |
| `VMULWEVWH` | |
| `VMULWEVWHU` | |
| `VMULWEVWHUH` | |
| `VMULWODHB` | |
| `VMULWODHBU` | |
| `VMULWODHBUB` | |
| `VMULWODQV` | |
| `VMULWODQVU` | |
| `VMULWODQVUV` | |
| `VMULWODVW` | |
| `VMULWODVWU` | |
| `VMULWODVWUW` | |
| `VMULWODWH` | |
| `VMULWODWHU` | |
| `VMULWODWHUH` | |
| `VNEGB` | |
| `VNEGH` | |
| `VNEGV` | |
| `VNEGW` | |
| `VNORB` | |
| `VNORV` | |
| `VORB` | |
| `VORNV` | |
| `VORV` | |
| `VPCNTB` | |
| `VPCNTH` | |
| `VPCNTV` | |
| `VPCNTW` | |
| `VPERMIW` | |
| `VROTRB` | |
| `VROTRH` | |
| `VROTRV` | |
| `VROTRW` | |
| `VSADDB` | |
| `VSADDBU` | |
| `VSADDH` | |
| `VSADDHU` | |
| `VSADDV` | |
| `VSADDVU` | |
| `VSADDW` | |
| `VSADDWU` | |
| `VSEQB` | |
| `VSEQH` | |
| `VSEQV` | |
| `VSEQW` | |
| `VSETALLNEB` | |
| `VSETALLNEH` | |
| `VSETALLNEV` | |
| `VSETALLNEW` | |
| `VSETANYEQB` | |
| `VSETANYEQH` | |
| `VSETANYEQV` | |
| `VSETANYEQW` | |
| `VSETEQV` | |
| `VSETNEV` | |
| `VSHUF4IB` | |
| `VSHUF4IH` | |
| `VSHUF4IV` | |
| `VSHUF4IW` | |
| `VSHUFB` | |
| `VSHUFH` | |
| `VSHUFV` | |
| `VSHUFW` | |
| `VSLLB` | |
| `VSLLH` | |
| `VSLLV` | |
| `VSLLW` | |
| `VSLTB` | |
| `VSLTBU` | |
| `VSLTH` | |
| `VSLTHU` | |
| `VSLTV` | |
| `VSLTVU` | |
| `VSLTW` | |
| `VSLTWU` | |
| `VSRAB` | |
| `VSRAH` | |
| `VSRAV` | |
| `VSRAW` | |
| `VSRLB` | |
| `VSRLH` | |
| `VSRLV` | |
| `VSRLW` | |
| `VSSUBB` | |
| `VSSUBBU` | |
| `VSSUBH` | |
| `VSSUBHU` | |
| `VSSUBV` | |
| `VSSUBVU` | |
| `VSSUBW` | |
| `VSSUBWU` | |
| `VSUBB` | |
| `VSUBBU` | |
| `VSUBD` | |
| `VSUBF` | |
| `VSUBH` | |
| `VSUBHU` | |
| `VSUBQ` | |
| `VSUBV` | |
| `VSUBVU` | |
| `VSUBW` | |
| `VSUBWEVHB` | |
| `VSUBWEVHBU` | |
| `VSUBWEVQV` | |
| `VSUBWEVQVU` | |
| `VSUBWEVVW` | |
| `VSUBWEVVWU` | |
| `VSUBWEVWH` | |
| `VSUBWEVWHU` | |
| `VSUBWODHB` | |
| `VSUBWODHBU` | |
| `VSUBWODQV` | |
| `VSUBWODQVU` | |
| `VSUBWODVW` | |
| `VSUBWODVWU` | |
| `VSUBWODWH` | |
| `VSUBWODWHU` | |
| `VSUBWU` | |
| `VXORB` | |
| `VXORV` | |
| `WORD` | |
| `XOR` | Bitwise XOR |
| `XVADDB` | |
| `XVADDBU` | |
| `XVADDD` | |
| `XVADDF` | |
| `XVADDH` | |
| `XVADDHU` | |
| `XVADDQ` | |
| `XVADDV` | |
| `XVADDVU` | |
| `XVADDW` | |
| `XVADDWEVHB` | |
| `XVADDWEVHBU` | |
| `XVADDWEVQV` | |
| `XVADDWEVQVU` | |
| `XVADDWEVVW` | |
| `XVADDWEVVWU` | |
| `XVADDWEVWH` | |
| `XVADDWEVWHU` | |
| `XVADDWODHB` | |
| `XVADDWODHBU` | |
| `XVADDWODQV` | |
| `XVADDWODQVU` | |
| `XVADDWODVW` | |
| `XVADDWODVWU` | |
| `XVADDWODWH` | |
| `XVADDWODWHU` | |
| `XVADDWU` | |
| `XVANDB` | |
| `XVANDNV` | |
| `XVANDV` | |
| `XVBITCLRB` | |
| `XVBITCLRH` | |
| `XVBITCLRV` | |
| `XVBITCLRW` | |
| `XVBITREVB` | |
| `XVBITREVH` | |
| `XVBITREVV` | |
| `XVBITREVW` | |
| `XVBITSETB` | |
| `XVBITSETH` | |
| `XVBITSETV` | |
| `XVBITSETW` | |
| `XVDIVB` | |
| `XVDIVBU` | |
| `XVDIVD` | |
| `XVDIVF` | |
| `XVDIVH` | |
| `XVDIVHU` | |
| `XVDIVV` | |
| `XVDIVVU` | |
| `XVDIVW` | |
| `XVDIVWU` | |
| `XVEXTRINSB` | |
| `XVEXTRINSH` | |
| `XVEXTRINSV` | |
| `XVEXTRINSW` | |
| `XVFCLASSD` | |
| `XVFCLASSF` | |
| `XVFRECIPD` | |
| `XVFRECIPF` | |
| `XVFRINTD` | |
| `XVFRINTF` | |
| `XVFRINTRMD` | |
| `XVFRINTRMF` | |
| `XVFRINTRNED` | |
| `XVFRINTRNEF` | |
| `XVFRINTRPD` | |
| `XVFRINTRPF` | |
| `XVFRINTRZD` | |
| `XVFRINTRZF` | |
| `XVFRSQRTD` | |
| `XVFRSQRTF` | |
| `XVFSQRTD` | |
| `XVFSQRTF` | |
| `XVILVHB` | |
| `XVILVHH` | |
| `XVILVHV` | |
| `XVILVHW` | |
| `XVILVLB` | |
| `XVILVLH` | |
| `XVILVLV` | |
| `XVILVLW` | |
| `XVMADDB` | |
| `XVMADDH` | |
| `XVMADDV` | |
| `XVMADDW` | |
| `XVMADDWEVHB` | |
| `XVMADDWEVHBU` | |
| `XVMADDWEVHBUB` | |
| `XVMADDWEVQV` | |
| `XVMADDWEVQVU` | |
| `XVMADDWEVQVUV` | |
| `XVMADDWEVVW` | |
| `XVMADDWEVVWU` | |
| `XVMADDWEVVWUW` | |
| `XVMADDWEVWH` | |
| `XVMADDWEVWHU` | |
| `XVMADDWEVWHUH` | |
| `XVMADDWODHB` | |
| `XVMADDWODHBU` | |
| `XVMADDWODHBUB` | |
| `XVMADDWODQV` | |
| `XVMADDWODQVU` | |
| `XVMADDWODQVUV` | |
| `XVMADDWODVW` | |
| `XVMADDWODVWU` | |
| `XVMADDWODVWUW` | |
| `XVMADDWODWH` | |
| `XVMADDWODWHU` | |
| `XVMADDWODWHUH` | |
| `XVMODB` | |
| `XVMODBU` | |
| `XVMODH` | |
| `XVMODHU` | |
| `XVMODV` | |
| `XVMODVU` | |
| `XVMODW` | |
| `XVMODWU` | |
| `XVMOVQ` | |
| `XVMSUBB` | |
| `XVMSUBH` | |
| `XVMSUBV` | |
| `XVMSUBW` | |
| `XVMUHB` | |
| `XVMUHBU` | |
| `XVMUHH` | |
| `XVMUHHU` | |
| `XVMUHV` | |
| `XVMUHVU` | |
| `XVMUHW` | |
| `XVMUHWU` | |
| `XVMULB` | |
| `XVMULD` | |
| `XVMULF` | |
| `XVMULH` | |
| `XVMULV` | |
| `XVMULW` | |
| `XVMULWEVHB` | |
| `XVMULWEVHBU` | |
| `XVMULWEVHBUB` | |
| `XVMULWEVQV` | |
| `XVMULWEVQVU` | |
| `XVMULWEVQVUV` | |
| `XVMULWEVVW` | |
| `XVMULWEVVWU` | |
| `XVMULWEVVWUW` | |
| `XVMULWEVWH` | |
| `XVMULWEVWHU` | |
| `XVMULWEVWHUH` | |
| `XVMULWODHB` | |
| `XVMULWODHBU` | |
| `XVMULWODHBUB` | |
| `XVMULWODQV` | |
| `XVMULWODQVU` | |
| `XVMULWODQVUV` | |
| `XVMULWODVW` | |
| `XVMULWODVWU` | |
| `XVMULWODVWUW` | |
| `XVMULWODWH` | |
| `XVMULWODWHU` | |
| `XVMULWODWHUH` | |
| `XVNEGB` | |
| `XVNEGH` | |
| `XVNEGV` | |
| `XVNEGW` | |
| `XVNORB` | |
| `XVNORV` | |
| `XVORB` | |
| `XVORNV` | |
| `XVORV` | |
| `XVPCNTB` | |
| `XVPCNTH` | |
| `XVPCNTV` | |
| `XVPCNTW` | |
| `XVPERMIQ` | |
| `XVPERMIV` | |
| `XVPERMIW` | |
| `XVROTRB` | |
| `XVROTRH` | |
| `XVROTRV` | |
| `XVROTRW` | |
| `XVSADDB` | |
| `XVSADDBU` | |
| `XVSADDH` | |
| `XVSADDHU` | |
| `XVSADDV` | |
| `XVSADDVU` | |
| `XVSADDW` | |
| `XVSADDWU` | |
| `XVSEQB` | |
| `XVSEQH` | |
| `XVSEQV` | |
| `XVSEQW` | |
| `XVSETALLNEB` | |
| `XVSETALLNEH` | |
| `XVSETALLNEV` | |
| `XVSETALLNEW` | |
| `XVSETANYEQB` | |
| `XVSETANYEQH` | |
| `XVSETANYEQV` | |
| `XVSETANYEQW` | |
| `XVSETEQV` | |
| `XVSETNEV` | |
| `XVSHUF4IB` | |
| `XVSHUF4IH` | |
| `XVSHUF4IV` | |
| `XVSHUF4IW` | |
| `XVSHUFB` | |
| `XVSHUFH` | |
| `XVSHUFV` | |
| `XVSHUFW` | |
| `XVSLLB` | |
| `XVSLLH` | |
| `XVSLLV` | |
| `XVSLLW` | |
| `XVSLTB` | |
| `XVSLTBU` | |
| `XVSLTH` | |
| `XVSLTHU` | |
| `XVSLTV` | |
| `XVSLTVU` | |
| `XVSLTW` | |
| `XVSLTWU` | |
| `XVSRAB` | |
| `XVSRAH` | |
| `XVSRAV` | |
| `XVSRAW` | |
| `XVSRLB` | |
| `XVSRLH` | |
| `XVSRLV` | |
| `XVSRLW` | |
| `XVSSUBB` | |
| `XVSSUBBU` | |
| `XVSSUBH` | |
| `XVSSUBHU` | |
| `XVSSUBV` | |
| `XVSSUBVU` | |
| `XVSSUBW` | |
| `XVSSUBWU` | |
| `XVSUBB` | |
| `XVSUBBU` | |
| `XVSUBD` | |
| `XVSUBF` | |
| `XVSUBH` | |
| `XVSUBHU` | |
| `XVSUBQ` | |
| `XVSUBV` | |
| `XVSUBVU` | |
| `XVSUBW` | |
| `XVSUBWEVHB` | |
| `XVSUBWEVHBU` | |
| `XVSUBWEVQV` | |
| `XVSUBWEVQVU` | |
| `XVSUBWEVVW` | |
| `XVSUBWEVVWU` | |
| `XVSUBWEVWH` | |
| `XVSUBWEVWHU` | |
| `XVSUBWODHB` | |
| `XVSUBWODHBU` | |
| `XVSUBWODQV` | |
| `XVSUBWODQVU` | |
| `XVSUBWODVW` | |
| `XVSUBWODVWU` | |
| `XVSUBWODWH` | |
| `XVSUBWODWHU` | |
| `XVSUBWU` | |
| `XVXORB` | |
| `XVXORV` | |
| `JAL` | |
Recognised: 814 mnemonics.
+991
View File
@@ -0,0 +1,991 @@
# RISC-V 64: instruction inventory
Generated by gasm-devkit's `_gen` from the Go toolchain's instruction table
(`cmd/internal/obj/riscv/anames.go`, go1.27.1); DO NOT EDIT. This page lists every mnemonic
`go tool asm` accepts on this target, which is the upper bound of the
language on it: a name absent here is not an instruction of the target,
and a name present here may still be one gasm's encoder cannot emit yet.
The inventory carries no per-mnemonic encoder column: on this target
encodability is decided per operand shape, and the live measured
coverage is reported by `gasm audit-instructions`.
| Mnemonic | Notes |
|---|---|
| `CALL` | Call subroutine |
| `DUFFCOPY` | |
| `DUFFZERO` | |
| `END` | |
| `FUNCDATA` | |
| `GETCALLERPC` | |
| `JMP` | Unconditional jump |
| `NOP` | |
| `PCALIGN` | |
| `PCALIGNMAX` | |
| `PCDATA` | |
| `RET` | Return |
| `TEXT` | |
| `UNDEF` | |
| `ADD` | Integer add |
| `ADDI` | Add immediate |
| `ADDIW` | Add immediate (32-bit) |
| `ADDUW` | |
| `ADDW` | Add (32-bit) |
| `AMOADDD` | Atomic add doubleword |
| `AMOADDW` | Atomic add word |
| `AMOANDD` | |
| `AMOANDW` | |
| `AMOMAXD` | |
| `AMOMAXUD` | |
| `AMOMAXUW` | |
| `AMOMAXW` | |
| `AMOMIND` | |
| `AMOMINUD` | |
| `AMOMINUW` | |
| `AMOMINW` | |
| `AMOORD` | |
| `AMOORW` | |
| `AMOSWAPD` | Atomic swap doubleword |
| `AMOSWAPW` | Atomic swap word |
| `AMOXORD` | |
| `AMOXORW` | |
| `AND` | Bitwise AND |
| `ANDI` | AND immediate |
| `ANDN` | |
| `AUIPC` | Add upper immediate to PC |
| `BCLR` | |
| `BCLRI` | |
| `BEQ` | Branch if equal |
| `BEQZ` | |
| `BEXT` | |
| `BEXTI` | |
| `BGE` | Branch if greater or equal |
| `BGEU` | Branch if greater or equal unsigned |
| `BGEZ` | |
| `BGT` | |
| `BGTU` | |
| `BGTZ` | |
| `BINV` | |
| `BINVI` | |
| `BLE` | |
| `BLEU` | |
| `BLEZ` | |
| `BLT` | Branch if less than |
| `BLTU` | Branch if less than unsigned |
| `BLTZ` | |
| `BNE` | Branch if not equal |
| `BNEZ` | |
| `BSET` | |
| `BSETI` | |
| `CADD` | |
| `CADDI` | |
| `CADDI16SP` | |
| `CADDI4SPN` | |
| `CADDIW` | |
| `CADDW` | |
| `CAND` | |
| `CANDI` | |
| `CBEQZ` | |
| `CBNEZ` | |
| `CEBREAK` | |
| `CFLD` | |
| `CFLDSP` | |
| `CFSD` | |
| `CFSDSP` | |
| `CJ` | |
| `CJALR` | |
| `CJR` | |
| `CLD` | |
| `CLDSP` | |
| `CLI` | |
| `CLUI` | |
| `CLW` | |
| `CLWSP` | |
| `CLZ` | |
| `CLZW` | |
| `CMV` | |
| `CNOP` | |
| `COR` | |
| `CPOP` | |
| `CPOPW` | |
| `CSD` | |
| `CSDSP` | |
| `CSLLI` | |
| `CSRAI` | |
| `CSRLI` | |
| `CSRRC` | |
| `CSRRCI` | |
| `CSRRS` | |
| `CSRRSI` | |
| `CSRRW` | |
| `CSRRWI` | |
| `CSUB` | |
| `CSUBW` | |
| `CSW` | |
| `CSWSP` | |
| `CTZ` | |
| `CTZW` | |
| `CXOR` | |
| `CZEROEQZ` | |
| `CZERONEZ` | |
| `DIV` | Divide |
| `DIVU` | Divide unsigned |
| `DIVUW` | |
| `DIVW` | Divide (32-bit) |
| `DRET` | |
| `EBREAK` | Breakpoint |
| `ECALL` | Environment call |
| `FABSD` | |
| `FABSS` | |
| `FADDD` | FP add (double) |
| `FADDQ` | |
| `FADDS` | FP add (single) |
| `FCLASSD` | |
| `FCLASSQ` | |
| `FCLASSS` | |
| `FCVTDL` | |
| `FCVTDLU` | |
| `FCVTDQ` | |
| `FCVTDS` | |
| `FCVTDW` | |
| `FCVTDWU` | |
| `FCVTLD` | |
| `FCVTLQ` | |
| `FCVTLS` | |
| `FCVTLUD` | |
| `FCVTLUQ` | |
| `FCVTLUS` | |
| `FCVTQD` | |
| `FCVTQL` | |
| `FCVTQLU` | |
| `FCVTQS` | |
| `FCVTQW` | |
| `FCVTQWU` | |
| `FCVTSD` | |
| `FCVTSL` | |
| `FCVTSLU` | |
| `FCVTSQ` | |
| `FCVTSW` | |
| `FCVTSWU` | |
| `FCVTWD` | |
| `FCVTWQ` | |
| `FCVTWS` | |
| `FCVTWUD` | |
| `FCVTWUQ` | |
| `FCVTWUS` | |
| `FDIVD` | FP divide (double) |
| `FDIVQ` | |
| `FDIVS` | FP divide (single) |
| `FENCE` | Memory barrier |
| `FEQD` | |
| `FEQQ` | |
| `FEQS` | |
| `FLD` | FP load doubleword |
| `FLED` | |
| `FLEQ` | |
| `FLES` | |
| `FLQ` | |
| `FLTD` | |
| `FLTQ` | |
| `FLTS` | |
| `FLW` | FP load word |
| `FMADDD` | |
| `FMADDQ` | |
| `FMADDS` | |
| `FMAXD` | |
| `FMAXQ` | |
| `FMAXS` | |
| `FMIND` | |
| `FMINQ` | |
| `FMINS` | |
| `FMSUBD` | |
| `FMSUBQ` | |
| `FMSUBS` | |
| `FMULD` | FP multiply (double) |
| `FMULQ` | |
| `FMULS` | FP multiply (single) |
| `FMVDX` | |
| `FMVSX` | |
| `FMVWX` | |
| `FMVXD` | |
| `FMVXS` | |
| `FMVXW` | |
| `FNED` | |
| `FNEGD` | |
| `FNEGS` | |
| `FNES` | |
| `FNMADDD` | |
| `FNMADDQ` | |
| `FNMADDS` | |
| `FNMSUBD` | |
| `FNMSUBQ` | |
| `FNMSUBS` | |
| `FSD` | FP store doubleword |
| `FSGNJD` | |
| `FSGNJND` | |
| `FSGNJNQ` | |
| `FSGNJNS` | |
| `FSGNJQ` | |
| `FSGNJS` | |
| `FSGNJXD` | |
| `FSGNJXQ` | |
| `FSGNJXS` | |
| `FSQ` | |
| `FSQRTD` | |
| `FSQRTQ` | |
| `FSQRTS` | |
| `FSUBD` | FP subtract (double) |
| `FSUBQ` | |
| `FSUBS` | FP subtract (single) |
| `FSW` | FP store word |
| `JAL` | Jump and link |
| `JALR` | Jump and link register |
| `LB` | Load byte |
| `LBU` | Load byte unsigned |
| `LD` | Load doubleword |
| `LH` | Load halfword |
| `LHU` | Load halfword unsigned |
| `LRD` | Load-reserved doubleword |
| `LRW` | Load-reserved word |
| `LUI` | Load upper immediate |
| `LW` | Load word |
| `LWU` | Load word unsigned |
| `MAX` | |
| `MAXU` | |
| `MIN` | |
| `MINU` | |
| `MOV` | |
| `MOVB` | |
| `MOVBU` | |
| `MOVD` | |
| `MOVF` | |
| `MOVH` | |
| `MOVHU` | |
| `MOVW` | |
| `MOVWU` | |
| `MRET` | |
| `MUL` | Multiply |
| `MULH` | Multiply high |
| `MULHSU` | Multiply high signed/unsigned |
| `MULHU` | Multiply high unsigned |
| `MULW` | Multiply (32-bit) |
| `NEG` | |
| `NEGW` | |
| `NOT` | |
| `OR` | Bitwise OR |
| `ORCB` | |
| `ORI` | OR immediate |
| `ORN` | |
| `RDCYCLE` | |
| `RDINSTRET` | |
| `RDTIME` | |
| `REM` | Remainder |
| `REMU` | Remainder unsigned |
| `REMUW` | |
| `REMW` | |
| `REV8` | |
| `ROL` | |
| `ROLW` | |
| `ROR` | |
| `RORI` | |
| `RORIW` | |
| `RORW` | |
| `SB` | Store byte |
| `SBREAK` | |
| `SCALL` | |
| `SCD` | Store-conditional doubleword |
| `SCW` | Store-conditional word |
| `SD` | Store doubleword |
| `SEQZ` | |
| `SEXTB` | |
| `SEXTH` | |
| `SFENCEVMA` | |
| `SH` | Store halfword |
| `SH1ADD` | |
| `SH1ADDUW` | |
| `SH2ADD` | |
| `SH2ADDUW` | |
| `SH3ADD` | |
| `SH3ADDUW` | |
| `SLL` | Shift left logical |
| `SLLI` | Shift left logical immediate |
| `SLLIUW` | |
| `SLLIW` | |
| `SLLW` | |
| `SLT` | Set if less than |
| `SLTI` | Set if less than immediate |
| `SLTIU` | Set if less than unsigned immediate |
| `SLTU` | Set if less than unsigned |
| `SNEZ` | |
| `SRA` | Shift right arithmetic |
| `SRAI` | Shift right arithmetic immediate |
| `SRAIW` | |
| `SRAW` | |
| `SRET` | |
| `SRL` | Shift right logical |
| `SRLI` | Shift right logical immediate |
| `SRLIW` | |
| `SRLW` | |
| `SUB` | Integer subtract |
| `SUBW` | Subtract (32-bit) |
| `SW` | Store word |
| `VAADDUVV` | |
| `VAADDUVX` | |
| `VAADDVV` | |
| `VAADDVX` | |
| `VADCVIM` | |
| `VADCVVM` | |
| `VADCVXM` | |
| `VADDVI` | |
| `VADDVV` | |
| `VADDVX` | |
| `VANDVI` | |
| `VANDVV` | |
| `VANDVX` | |
| `VASUBUVV` | |
| `VASUBUVX` | |
| `VASUBVV` | |
| `VASUBVX` | |
| `VCOMPRESSVM` | |
| `VCPOPM` | |
| `VDIVUVV` | |
| `VDIVUVX` | |
| `VDIVVV` | |
| `VDIVVX` | |
| `VFABSV` | |
| `VFADDVF` | |
| `VFADDVV` | |
| `VFCLASSV` | |
| `VFCVTFXUV` | |
| `VFCVTFXV` | |
| `VFCVTRTZXFV` | |
| `VFCVTRTZXUFV` | |
| `VFCVTXFV` | |
| `VFCVTXUFV` | |
| `VFDIVVF` | |
| `VFDIVVV` | |
| `VFIRSTM` | |
| `VFMACCVF` | |
| `VFMACCVV` | |
| `VFMADDVF` | |
| `VFMADDVV` | |
| `VFMAXVF` | |
| `VFMAXVV` | |
| `VFMERGEVFM` | |
| `VFMINVF` | |
| `VFMINVV` | |
| `VFMSACVF` | |
| `VFMSACVV` | |
| `VFMSUBVF` | |
| `VFMSUBVV` | |
| `VFMULVF` | |
| `VFMULVV` | |
| `VFMVFS` | |
| `VFMVSF` | |
| `VFMVVF` | |
| `VFNCVTFFW` | |
| `VFNCVTFXUW` | |
| `VFNCVTFXW` | |
| `VFNCVTRODFFW` | |
| `VFNCVTRTZXFW` | |
| `VFNCVTRTZXUFW` | |
| `VFNCVTXFW` | |
| `VFNCVTXUFW` | |
| `VFNEGV` | |
| `VFNMACCVF` | |
| `VFNMACCVV` | |
| `VFNMADDVF` | |
| `VFNMADDVV` | |
| `VFNMSACVF` | |
| `VFNMSACVV` | |
| `VFNMSUBVF` | |
| `VFNMSUBVV` | |
| `VFRDIVVF` | |
| `VFREC7V` | |
| `VFREDMAXVS` | |
| `VFREDMINVS` | |
| `VFREDOSUMVS` | |
| `VFREDUSUMVS` | |
| `VFRSQRT7V` | |
| `VFRSUBVF` | |
| `VFSGNJNVF` | |
| `VFSGNJNVV` | |
| `VFSGNJVF` | |
| `VFSGNJVV` | |
| `VFSGNJXVF` | |
| `VFSGNJXVV` | |
| `VFSLIDE1DOWNVF` | |
| `VFSLIDE1UPVF` | |
| `VFSQRTV` | |
| `VFSUBVF` | |
| `VFSUBVV` | |
| `VFWADDVF` | |
| `VFWADDVV` | |
| `VFWADDWF` | |
| `VFWADDWV` | |
| `VFWCVTFFV` | |
| `VFWCVTFXUV` | |
| `VFWCVTFXV` | |
| `VFWCVTRTZXFV` | |
| `VFWCVTRTZXUFV` | |
| `VFWCVTXFV` | |
| `VFWCVTXUFV` | |
| `VFWMACCVF` | |
| `VFWMACCVV` | |
| `VFWMSACVF` | |
| `VFWMSACVV` | |
| `VFWMULVF` | |
| `VFWMULVV` | |
| `VFWNMACCVF` | |
| `VFWNMACCVV` | |
| `VFWNMSACVF` | |
| `VFWNMSACVV` | |
| `VFWREDOSUMVS` | |
| `VFWREDUSUMVS` | |
| `VFWSUBVF` | |
| `VFWSUBVV` | |
| `VFWSUBWF` | |
| `VFWSUBWV` | |
| `VIDV` | |
| `VIOTAM` | |
| `VL1RE16V` | |
| `VL1RE32V` | |
| `VL1RE64V` | |
| `VL1RE8V` | |
| `VL1RV` | |
| `VL2RE16V` | |
| `VL2RE32V` | |
| `VL2RE64V` | |
| `VL2RE8V` | |
| `VL2RV` | |
| `VL4RE16V` | |
| `VL4RE32V` | |
| `VL4RE64V` | |
| `VL4RE8V` | |
| `VL4RV` | |
| `VL8RE16V` | |
| `VL8RE32V` | |
| `VL8RE64V` | |
| `VL8RE8V` | |
| `VL8RV` | |
| `VLE16FFV` | |
| `VLE16V` | |
| `VLE32FFV` | |
| `VLE32V` | |
| `VLE64FFV` | |
| `VLE64V` | |
| `VLE8FFV` | |
| `VLE8V` | |
| `VLMV` | |
| `VLOXEI16V` | |
| `VLOXEI32V` | |
| `VLOXEI64V` | |
| `VLOXEI8V` | |
| `VLOXSEG2EI16V` | |
| `VLOXSEG2EI32V` | |
| `VLOXSEG2EI64V` | |
| `VLOXSEG2EI8V` | |
| `VLOXSEG3EI16V` | |
| `VLOXSEG3EI32V` | |
| `VLOXSEG3EI64V` | |
| `VLOXSEG3EI8V` | |
| `VLOXSEG4EI16V` | |
| `VLOXSEG4EI32V` | |
| `VLOXSEG4EI64V` | |
| `VLOXSEG4EI8V` | |
| `VLOXSEG5EI16V` | |
| `VLOXSEG5EI32V` | |
| `VLOXSEG5EI64V` | |
| `VLOXSEG5EI8V` | |
| `VLOXSEG6EI16V` | |
| `VLOXSEG6EI32V` | |
| `VLOXSEG6EI64V` | |
| `VLOXSEG6EI8V` | |
| `VLOXSEG7EI16V` | |
| `VLOXSEG7EI32V` | |
| `VLOXSEG7EI64V` | |
| `VLOXSEG7EI8V` | |
| `VLOXSEG8EI16V` | |
| `VLOXSEG8EI32V` | |
| `VLOXSEG8EI64V` | |
| `VLOXSEG8EI8V` | |
| `VLSE16V` | |
| `VLSE32V` | |
| `VLSE64V` | |
| `VLSE8V` | |
| `VLSEG2E16FFV` | |
| `VLSEG2E16V` | |
| `VLSEG2E32FFV` | |
| `VLSEG2E32V` | |
| `VLSEG2E64FFV` | |
| `VLSEG2E64V` | |
| `VLSEG2E8FFV` | |
| `VLSEG2E8V` | |
| `VLSEG3E16FFV` | |
| `VLSEG3E16V` | |
| `VLSEG3E32FFV` | |
| `VLSEG3E32V` | |
| `VLSEG3E64FFV` | |
| `VLSEG3E64V` | |
| `VLSEG3E8FFV` | |
| `VLSEG3E8V` | |
| `VLSEG4E16FFV` | |
| `VLSEG4E16V` | |
| `VLSEG4E32FFV` | |
| `VLSEG4E32V` | |
| `VLSEG4E64FFV` | |
| `VLSEG4E64V` | |
| `VLSEG4E8FFV` | |
| `VLSEG4E8V` | |
| `VLSEG5E16FFV` | |
| `VLSEG5E16V` | |
| `VLSEG5E32FFV` | |
| `VLSEG5E32V` | |
| `VLSEG5E64FFV` | |
| `VLSEG5E64V` | |
| `VLSEG5E8FFV` | |
| `VLSEG5E8V` | |
| `VLSEG6E16FFV` | |
| `VLSEG6E16V` | |
| `VLSEG6E32FFV` | |
| `VLSEG6E32V` | |
| `VLSEG6E64FFV` | |
| `VLSEG6E64V` | |
| `VLSEG6E8FFV` | |
| `VLSEG6E8V` | |
| `VLSEG7E16FFV` | |
| `VLSEG7E16V` | |
| `VLSEG7E32FFV` | |
| `VLSEG7E32V` | |
| `VLSEG7E64FFV` | |
| `VLSEG7E64V` | |
| `VLSEG7E8FFV` | |
| `VLSEG7E8V` | |
| `VLSEG8E16FFV` | |
| `VLSEG8E16V` | |
| `VLSEG8E32FFV` | |
| `VLSEG8E32V` | |
| `VLSEG8E64FFV` | |
| `VLSEG8E64V` | |
| `VLSEG8E8FFV` | |
| `VLSEG8E8V` | |
| `VLSSEG2E16V` | |
| `VLSSEG2E32V` | |
| `VLSSEG2E64V` | |
| `VLSSEG2E8V` | |
| `VLSSEG3E16V` | |
| `VLSSEG3E32V` | |
| `VLSSEG3E64V` | |
| `VLSSEG3E8V` | |
| `VLSSEG4E16V` | |
| `VLSSEG4E32V` | |
| `VLSSEG4E64V` | |
| `VLSSEG4E8V` | |
| `VLSSEG5E16V` | |
| `VLSSEG5E32V` | |
| `VLSSEG5E64V` | |
| `VLSSEG5E8V` | |
| `VLSSEG6E16V` | |
| `VLSSEG6E32V` | |
| `VLSSEG6E64V` | |
| `VLSSEG6E8V` | |
| `VLSSEG7E16V` | |
| `VLSSEG7E32V` | |
| `VLSSEG7E64V` | |
| `VLSSEG7E8V` | |
| `VLSSEG8E16V` | |
| `VLSSEG8E32V` | |
| `VLSSEG8E64V` | |
| `VLSSEG8E8V` | |
| `VLUXEI16V` | |
| `VLUXEI32V` | |
| `VLUXEI64V` | |
| `VLUXEI8V` | |
| `VLUXSEG2EI16V` | |
| `VLUXSEG2EI32V` | |
| `VLUXSEG2EI64V` | |
| `VLUXSEG2EI8V` | |
| `VLUXSEG3EI16V` | |
| `VLUXSEG3EI32V` | |
| `VLUXSEG3EI64V` | |
| `VLUXSEG3EI8V` | |
| `VLUXSEG4EI16V` | |
| `VLUXSEG4EI32V` | |
| `VLUXSEG4EI64V` | |
| `VLUXSEG4EI8V` | |
| `VLUXSEG5EI16V` | |
| `VLUXSEG5EI32V` | |
| `VLUXSEG5EI64V` | |
| `VLUXSEG5EI8V` | |
| `VLUXSEG6EI16V` | |
| `VLUXSEG6EI32V` | |
| `VLUXSEG6EI64V` | |
| `VLUXSEG6EI8V` | |
| `VLUXSEG7EI16V` | |
| `VLUXSEG7EI32V` | |
| `VLUXSEG7EI64V` | |
| `VLUXSEG7EI8V` | |
| `VLUXSEG8EI16V` | |
| `VLUXSEG8EI32V` | |
| `VLUXSEG8EI64V` | |
| `VLUXSEG8EI8V` | |
| `VMACCVV` | |
| `VMACCVX` | |
| `VMADCVI` | |
| `VMADCVIM` | |
| `VMADCVV` | |
| `VMADCVVM` | |
| `VMADCVX` | |
| `VMADCVXM` | |
| `VMADDVV` | |
| `VMADDVX` | |
| `VMANDMM` | |
| `VMANDNMM` | |
| `VMAXUVV` | |
| `VMAXUVX` | |
| `VMAXVV` | |
| `VMAXVX` | |
| `VMCLRM` | |
| `VMERGEVIM` | |
| `VMERGEVVM` | |
| `VMERGEVXM` | |
| `VMFEQVF` | |
| `VMFEQVV` | |
| `VMFGEVF` | |
| `VMFGEVV` | |
| `VMFGTVF` | |
| `VMFGTVV` | |
| `VMFLEVF` | |
| `VMFLEVV` | |
| `VMFLTVF` | |
| `VMFLTVV` | |
| `VMFNEVF` | |
| `VMFNEVV` | |
| `VMINUVV` | |
| `VMINUVX` | |
| `VMINVV` | |
| `VMINVX` | |
| `VMMVM` | |
| `VMNANDMM` | |
| `VMNORMM` | |
| `VMNOTM` | |
| `VMORMM` | |
| `VMORNMM` | |
| `VMSBCVV` | |
| `VMSBCVVM` | |
| `VMSBCVX` | |
| `VMSBCVXM` | |
| `VMSBFM` | |
| `VMSEQVI` | |
| `VMSEQVV` | |
| `VMSEQVX` | |
| `VMSETM` | |
| `VMSGEUVI` | |
| `VMSGEUVV` | |
| `VMSGEVI` | |
| `VMSGEVV` | |
| `VMSGTUVI` | |
| `VMSGTUVV` | |
| `VMSGTUVX` | |
| `VMSGTVI` | |
| `VMSGTVV` | |
| `VMSGTVX` | |
| `VMSIFM` | |
| `VMSLEUVI` | |
| `VMSLEUVV` | |
| `VMSLEUVX` | |
| `VMSLEVI` | |
| `VMSLEVV` | |
| `VMSLEVX` | |
| `VMSLTUVI` | |
| `VMSLTUVV` | |
| `VMSLTUVX` | |
| `VMSLTVI` | |
| `VMSLTVV` | |
| `VMSLTVX` | |
| `VMSNEVI` | |
| `VMSNEVV` | |
| `VMSNEVX` | |
| `VMSOFM` | |
| `VMULHSUVV` | |
| `VMULHSUVX` | |
| `VMULHUVV` | |
| `VMULHUVX` | |
| `VMULHVV` | |
| `VMULHVX` | |
| `VMULVV` | |
| `VMULVX` | |
| `VMV1RV` | |
| `VMV2RV` | |
| `VMV4RV` | |
| `VMV8RV` | |
| `VMVSX` | |
| `VMVVI` | |
| `VMVVV` | |
| `VMVVX` | |
| `VMVXS` | |
| `VMXNORMM` | |
| `VMXORMM` | |
| `VNCLIPUWI` | |
| `VNCLIPUWV` | |
| `VNCLIPUWX` | |
| `VNCLIPWI` | |
| `VNCLIPWV` | |
| `VNCLIPWX` | |
| `VNCVTXXW` | |
| `VNEGV` | |
| `VNMSACVV` | |
| `VNMSACVX` | |
| `VNMSUBVV` | |
| `VNMSUBVX` | |
| `VNOTV` | |
| `VNSRAWI` | |
| `VNSRAWV` | |
| `VNSRAWX` | |
| `VNSRLWI` | |
| `VNSRLWV` | |
| `VNSRLWX` | |
| `VORVI` | |
| `VORVV` | |
| `VORVX` | |
| `VREDANDVS` | |
| `VREDMAXUVS` | |
| `VREDMAXVS` | |
| `VREDMINUVS` | |
| `VREDMINVS` | |
| `VREDORVS` | |
| `VREDSUMVS` | |
| `VREDXORVS` | |
| `VREMUVV` | |
| `VREMUVX` | |
| `VREMVV` | |
| `VREMVX` | |
| `VRGATHEREI16VV` | |
| `VRGATHERVI` | |
| `VRGATHERVV` | |
| `VRGATHERVX` | |
| `VRSUBVI` | |
| `VRSUBVX` | |
| `VS1RV` | |
| `VS2RV` | |
| `VS4RV` | |
| `VS8RV` | |
| `VSADDUVI` | |
| `VSADDUVV` | |
| `VSADDUVX` | |
| `VSADDVI` | |
| `VSADDVV` | |
| `VSADDVX` | |
| `VSBCVVM` | |
| `VSBCVXM` | |
| `VSE16V` | |
| `VSE32V` | |
| `VSE64V` | |
| `VSE8V` | |
| `VSETIVLI` | |
| `VSETVL` | |
| `VSETVLI` | |
| `VSEXTVF2` | |
| `VSEXTVF4` | |
| `VSEXTVF8` | |
| `VSLIDE1DOWNVX` | |
| `VSLIDE1UPVX` | |
| `VSLIDEDOWNVI` | |
| `VSLIDEDOWNVX` | |
| `VSLIDEUPVI` | |
| `VSLIDEUPVX` | |
| `VSLLVI` | |
| `VSLLVV` | |
| `VSLLVX` | |
| `VSMULVV` | |
| `VSMULVX` | |
| `VSMV` | |
| `VSOXEI16V` | |
| `VSOXEI32V` | |
| `VSOXEI64V` | |
| `VSOXEI8V` | |
| `VSOXSEG2EI16V` | |
| `VSOXSEG2EI32V` | |
| `VSOXSEG2EI64V` | |
| `VSOXSEG2EI8V` | |
| `VSOXSEG3EI16V` | |
| `VSOXSEG3EI32V` | |
| `VSOXSEG3EI64V` | |
| `VSOXSEG3EI8V` | |
| `VSOXSEG4EI16V` | |
| `VSOXSEG4EI32V` | |
| `VSOXSEG4EI64V` | |
| `VSOXSEG4EI8V` | |
| `VSOXSEG5EI16V` | |
| `VSOXSEG5EI32V` | |
| `VSOXSEG5EI64V` | |
| `VSOXSEG5EI8V` | |
| `VSOXSEG6EI16V` | |
| `VSOXSEG6EI32V` | |
| `VSOXSEG6EI64V` | |
| `VSOXSEG6EI8V` | |
| `VSOXSEG7EI16V` | |
| `VSOXSEG7EI32V` | |
| `VSOXSEG7EI64V` | |
| `VSOXSEG7EI8V` | |
| `VSOXSEG8EI16V` | |
| `VSOXSEG8EI32V` | |
| `VSOXSEG8EI64V` | |
| `VSOXSEG8EI8V` | |
| `VSRAVI` | |
| `VSRAVV` | |
| `VSRAVX` | |
| `VSRLVI` | |
| `VSRLVV` | |
| `VSRLVX` | |
| `VSSE16V` | |
| `VSSE32V` | |
| `VSSE64V` | |
| `VSSE8V` | |
| `VSSEG2E16V` | |
| `VSSEG2E32V` | |
| `VSSEG2E64V` | |
| `VSSEG2E8V` | |
| `VSSEG3E16V` | |
| `VSSEG3E32V` | |
| `VSSEG3E64V` | |
| `VSSEG3E8V` | |
| `VSSEG4E16V` | |
| `VSSEG4E32V` | |
| `VSSEG4E64V` | |
| `VSSEG4E8V` | |
| `VSSEG5E16V` | |
| `VSSEG5E32V` | |
| `VSSEG5E64V` | |
| `VSSEG5E8V` | |
| `VSSEG6E16V` | |
| `VSSEG6E32V` | |
| `VSSEG6E64V` | |
| `VSSEG6E8V` | |
| `VSSEG7E16V` | |
| `VSSEG7E32V` | |
| `VSSEG7E64V` | |
| `VSSEG7E8V` | |
| `VSSEG8E16V` | |
| `VSSEG8E32V` | |
| `VSSEG8E64V` | |
| `VSSEG8E8V` | |
| `VSSRAVI` | |
| `VSSRAVV` | |
| `VSSRAVX` | |
| `VSSRLVI` | |
| `VSSRLVV` | |
| `VSSRLVX` | |
| `VSSSEG2E16V` | |
| `VSSSEG2E32V` | |
| `VSSSEG2E64V` | |
| `VSSSEG2E8V` | |
| `VSSSEG3E16V` | |
| `VSSSEG3E32V` | |
| `VSSSEG3E64V` | |
| `VSSSEG3E8V` | |
| `VSSSEG4E16V` | |
| `VSSSEG4E32V` | |
| `VSSSEG4E64V` | |
| `VSSSEG4E8V` | |
| `VSSSEG5E16V` | |
| `VSSSEG5E32V` | |
| `VSSSEG5E64V` | |
| `VSSSEG5E8V` | |
| `VSSSEG6E16V` | |
| `VSSSEG6E32V` | |
| `VSSSEG6E64V` | |
| `VSSSEG6E8V` | |
| `VSSSEG7E16V` | |
| `VSSSEG7E32V` | |
| `VSSSEG7E64V` | |
| `VSSSEG7E8V` | |
| `VSSSEG8E16V` | |
| `VSSSEG8E32V` | |
| `VSSSEG8E64V` | |
| `VSSSEG8E8V` | |
| `VSSUBUVV` | |
| `VSSUBUVX` | |
| `VSSUBVV` | |
| `VSSUBVX` | |
| `VSUBVV` | |
| `VSUBVX` | |
| `VSUXEI16V` | |
| `VSUXEI32V` | |
| `VSUXEI64V` | |
| `VSUXEI8V` | |
| `VSUXSEG2EI16V` | |
| `VSUXSEG2EI32V` | |
| `VSUXSEG2EI64V` | |
| `VSUXSEG2EI8V` | |
| `VSUXSEG3EI16V` | |
| `VSUXSEG3EI32V` | |
| `VSUXSEG3EI64V` | |
| `VSUXSEG3EI8V` | |
| `VSUXSEG4EI16V` | |
| `VSUXSEG4EI32V` | |
| `VSUXSEG4EI64V` | |
| `VSUXSEG4EI8V` | |
| `VSUXSEG5EI16V` | |
| `VSUXSEG5EI32V` | |
| `VSUXSEG5EI64V` | |
| `VSUXSEG5EI8V` | |
| `VSUXSEG6EI16V` | |
| `VSUXSEG6EI32V` | |
| `VSUXSEG6EI64V` | |
| `VSUXSEG6EI8V` | |
| `VSUXSEG7EI16V` | |
| `VSUXSEG7EI32V` | |
| `VSUXSEG7EI64V` | |
| `VSUXSEG7EI8V` | |
| `VSUXSEG8EI16V` | |
| `VSUXSEG8EI32V` | |
| `VSUXSEG8EI64V` | |
| `VSUXSEG8EI8V` | |
| `VWADDUVV` | |
| `VWADDUVX` | |
| `VWADDUWV` | |
| `VWADDUWX` | |
| `VWADDVV` | |
| `VWADDVX` | |
| `VWADDWV` | |
| `VWADDWX` | |
| `VWCVTUXXV` | |
| `VWCVTXXV` | |
| `VWMACCSUVV` | |
| `VWMACCSUVX` | |
| `VWMACCUSVX` | |
| `VWMACCUVV` | |
| `VWMACCUVX` | |
| `VWMACCVV` | |
| `VWMACCVX` | |
| `VWMULSUVV` | |
| `VWMULSUVX` | |
| `VWMULUVV` | |
| `VWMULUVX` | |
| `VWMULVV` | |
| `VWMULVX` | |
| `VWREDSUMUVS` | |
| `VWREDSUMVS` | |
| `VWSUBUVV` | |
| `VWSUBUVX` | |
| `VWSUBUWV` | |
| `VWSUBUWX` | |
| `VWSUBVV` | |
| `VWSUBVX` | |
| `VWSUBWV` | |
| `VWSUBWX` | |
| `VXORVI` | |
| `VXORVV` | |
| `VXORVX` | |
| `VZEXTVF2` | |
| `VZEXTVF4` | |
| `VZEXTVF8` | |
| `WFI` | |
| `WORD` | |
| `XNOR` | |
| `XOR` | Bitwise XOR |
| `XORI` | XOR immediate |
| `ZEXTH` | |
Recognised: 975 mnemonics.
+140
View File
@@ -0,0 +1,140 @@
# Language: lexicon, statements and expressions
Layer 1, the common language, the same on every target. Verified against
`go tool asm` of Go 1.27.1 and against gasm's parser, which is differentially
tested against the toolchain. The authoritative sources behind this page are
the assembler's lexer (`cmd/asm/internal/lex`), its parser
(`cmd/asm/internal/asm/parse.go`) and the toolchain's own test data.
## Source files and targets
An assembly source is a `.s` file. The Go build convention names a
target-specific file with the architecture suffix, `_amd64.s`, `_arm64.s`,
`_riscv64.s` or `_loong64.s`; files without a suffix are portable across
targets. The same assembler program assembles every target: `go tool asm`
picks the target from the `GOOS` and `GOARCH` environment variables, and gasm
from the file name suffix or the `--arch` flag.
## Character set and identifiers
Sources are ASCII text. An identifier is a sequence of ASCII letters, digits
and underscores, digits never first, with exactly two additions:
- U+00B7, the middle dot `·`, stands for the period in a symbol's
package-qualified name;
- U+2215, the division slash `∕`, stands for the slash in a package path.
The two substitutions exist because the parser treats a real period and a
real slash as punctuation. The syntax is otherwise uppercase throughout:
instructions, registers and directives are written in upper case. The one
inherited exception is the `g` register name on 32-bit ARM.
## Comments
Two comment forms, both Go's:
```text
// a line comment
/* a block comment */
```
A comment of the form `//go:build` or the legacy `+build` comment is not a
plain comment: the lexer reports it to the build system as a build
constraint.
## Statements
The grammar of one line, from the parser:
```text
{label:} WORD[.qualifier] [ arg {, arg} ] (';' | '\n')
```
- A **label** is an identifier followed by a colon. Labels are
function-local: two functions in one file may reuse the same name, and a
reference resolves within the function that contains it. A branch
instruction names its target with a bare label operand, and the assembler
resolves it PC-relative. The explicit forms `offset(PC)`, a constant
counting instructions from the branch, and `name(SB)`, a cross-function
static reference, appear as branch targets as well.
- **WORD** is the instruction or directive name, upper case. On the ARM
family the word may carry a dot qualifier selecting a condition or shift
mode, such as the condition suffixes on 32-bit ARM; the amd64, arm64,
riscv64 and loong64 assemblies carry no instruction qualifiers apart from
their own width suffixes, which are part of the mnemonic.
- **Arguments** are separated by commas, with no trailing comma.
- A statement ends at a newline or at a semicolon, so several statements fit
on one line separated by `;`. Blank lines are free.
The first word of a line is a directive if it is one of the directive names
(TEXT, DATA, GLOBL, FUNCDATA, PCDATA, PCALIGN) and an instruction otherwise.
Unknown instruction names are errors; the instruction set is the set the
toolchain itself defines per target, plus the common pseudo-instructions.
## Literals
| Form | Examples | Notes |
|---|---|---|
| Integer | `0`, `42`, `0x2a`, `0o52`, `0b101010`, `1_000` | decimal, hexadecimal, octal and binary forms with Go's digit separators |
| Character | `'a'`, `'\n'`, `'\x41'` | single quoted, Go escape rules |
| String | `"this program can only run\n"` | double quoted, Go escape rules; accepted where an operand takes raw bytes, in practice a DATA initialiser |
| Float | `1.5`, `1e9` | accepted by the lexer; only meaningful where the target's encoding takes a float operand |
## Expressions
Constant expressions may appear wherever a constant is expected: in
immediates after `$`, in memory offsets, in frame and data sizes. The
evaluator works on unsigned 64-bit values with Go's operator precedence, and
the parser states its grammar in exactly those terms:
```text
expr = term { '+' term | '-' term | '|' term | '^' term }
term = factor { '*' factor | '/' factor | '%' factor | '<<' factor | '>>' factor | '&' factor }
factor = const | '+' factor | '-' factor | '~' factor | '(' expr ')'
```
Two consequences are worth naming, because the arithmetic surprises people
who read it as C:
- Shifts bind at the multiplicative level, next to `*` and `&`, while `|`
and `^` bind at the additive level. `$x<<1|3` computes `(x<<1)|3`, which
differs from `x*2+3` whenever `x` is odd. Plan 9 arithmetic is Go
precedence applied to a byte-oriented language, not the C expression it
resembles.
- The evaluator is unsigned and guarded: division or modulo by zero is an
error, and so is dividing a value with the high bit set; shift counts must
be non-negative; and a right shift of a value with the high bit set is
rejected rather than sign-extended.
An address expression such as `(index*4)(base)` is evaluated at assembly
time only if every name in it is a constant; a name that resolves to a
symbol turns the expression into a relocation request, never into a folded
constant.
Named constants enter expressions through the preprocessor (`#define`,
`-D`) and, in Go-embedded packages, through the generated `go_asm.h`; see
PREPROCESSOR.md and RUNTIME.md.
## The common pseudo-instructions
A handful of instructions exist on every target, assembled by the assembler
itself rather than the encoder: `NOP`, which emits the target's no-operation
encoding, and the frame-management pseudo-instructions the compiler emits
(`FUNCDATA`, `PCDATA`) which DIRECTIVES.md specifies. Everything else is the
target's own instruction set, and the assembler knows only the instructions
the toolchain's compiler emits; a hand-written kernel wanting more lays the
encoding down with `BYTE` on amd64 or waits for the extended layer.
## Case study: three lines, decomposed
```text
B.EQ 1(PC) // arm64: condition qualifier on the mnemonic,
// target one instruction past the branch
JMP done // every target: bare label, function-local,
// resolved PC-relative
MOVQ $reader__size>>3, CX // amd64: expression over a go_asm.h constant
```
The first shows a qualifier and the explicit relative target form; the second
the ordinary label reference; the third an expression over a generated
constant. Labels are reusable between functions without conflict.
+94
View File
@@ -0,0 +1,94 @@
# LoongArch 64
Layer 1, target page. Verified against `go tool asm` of Go 1.27.1, against
the toolchain's own loong64 assembler manual (`cmd/internal/obj/loong64/doc.go`)
and against gasm's encoder, whose output is compared byte for byte with the
toolchain's. The complete mnemonic inventory lives in the generated appendix
[INSTRUCTIONS-LOONG64.md](INSTRUCTIONS-LOONG64.md).
## Registers
- General purpose `R0` to `R31`, floating point `F0` to `F31`, LSX vectors
`V0` to `V31` and LASX vectors `X0` to `X31`.
- Fixed roles from the toolchain's table: `R0` is the constant zero, `R1`
the return address, `R3` the stack pointer, `R22` the goroutine pointer,
`R29` the closure context and `R30` the assembler's temporary. `R12`,
`R13`, `R14`, `R15` and `R20` serve the PLT and trampoline sequences:
usable in assembly, but saved before any call.
## Widths ride the mnemonic
| Suffix | Width |
|---|---|
| `B`, `BU` | 8-bit, 8-bit unsigned |
| `H`, `HU` | 16-bit, 16-bit unsigned |
| `W`, `WU` | 32-bit, 32-bit unsigned |
| `V` | 64-bit |
| `F`, `D` | 32-bit and 64-bit float |
| `V` prefix (LSX) | 128-bit vector |
| `XV` prefix (LASX) | 256-bit vector |
The MOV series is the load and store interface: `MOVB (R2), R3` loads a
byte, `MOVV (R2), R3` a double word, `VMOVQ (R2), V1` a 128-bit vector and
`XVMOVQ (R2), X1` a 256-bit one.
## Operand order
Most instructions appear in left-to-right assignment order: `ADDV R11, R12,
R13` is `add.d R13, R12, R11`, and the two-operand form
`OR R5, R6` assigns into R6. Exceptions:
- Jump and branch instructions keep the GNU order: `BEQ R0, R4, label1`.
- The bitfield family is `BSTRINSW`, `BSTRINSV`, `BSTRPICKW`, `BSTRPICKV`
`$<msb>, <Rj>, $<lsb>, <Rd>`.
## Addressing
- Plain: `offset(Rbase)`.
- Base plus offset **register**, no scale: `(R4)(R5)`, as in
`MOVB (R4)(R5), R6`, the `ldx` family.
- The pointer loads and stores `MOVWP` and `MOVVP` take a source-level
16-bit offset that the encoder halves into the 14-bit field, writing
`MOVWP 8(R4), R5` as `ldptr.w r5, r4, $2`.
## Vector element syntax
The `VMOVQ` and `XVMOVQ` transfer family covers register-to-vector moves
with arrangement and index suffixes: `VMOVQ Rj, Vd.B[index]` inserts a
general register into one lane, `VMOVQ Vj.B[index], Rd` extracts one,
`VMOVQ Rj, Vd.B16` broadcasts across all sixteen, and `VMOVQ Vj.B[index],
Vd.B16` replicates one lane. The broadcast-from-memory form takes the true
byte offset at source level, which the encoder rescales per arrangement.
The permute and extract families take their 8-bit control word first:
`VPERMIW ui8, Vj, Vd`, `VEXTRINSB ui8, Vj, Vd`.
## Alignment
`PCALIGN $n` pads with NOOP to a power-of-two boundary between 8 and 2048,
and this target additionally auto-aligns loop heads to 16 bytes.
## Atomics, barriers and prefetch
- The `AM` atomic family comes in plain and `_DB` flavours; the `_DB`
forms, such as `AMSWAPDBW`, complete the atomic sequence and act as a
full data barrier. Within the AM family the destination and base
registers may not coincide and the destination may not equal the operand
register: one is an exception, the other silently unspecified.
- `DBAR` carries the graded hint encoding documented for LA664 and later,
with hint 0x700 as the read-after-read lightweight barrier; older cores
treat every hint as the full barrier.
- `PRELD offset(Rbase), $hint` prefetches with the documented hints (0
load to L1, 2 load to L3, 8 store to L1); `PRELDX` adds the encoded
block descriptor.
- `ALSL`-family shift-and-add writes the desired shift amount in source and
encodes one less: `ALSLV $4, R4, R5, R6` shifts by 4.
- `ADDV16 si16<<16, Rj, Rd` is the high-immediate add paired with the
pointer loads for GOT relative access.
## Relocations
`R_CALLLOONG64` for the 28-bit BL, `R_LOONG64_CALL36` for the
PCADDU18I-plus-JIRL pair, the `R_LOONG64_ADDR`, `ADDR64`, `TLS_LE`, `TLS_IE`,
`GOT` and `GOT64` high and low pairs, the aligned conditional jump forms
`R_JMP16LOONG64` and `R_JMP21LOONG64`, and `R_LOONG64_ADD64` and `SUB64`
for in-place arithmetic, all specified in [GOOBJ.md](../GOOBJ.md).
+114
View File
@@ -0,0 +1,114 @@
# Operands: grammar, pseudo-registers, addressing and symbols
Layer 1, the common language. Verified against `go tool asm` of Go 1.27.1 and
against gasm's parser. The operand grammar is the part of the language that
varies most between targets, so this page fixes the common grammar and the
pseudo-registers; the per architecture pages carry the register names and the
addressing quirks each target adds.
## The four operand kinds
Every operand is one of four kinds:
```text
R1 register
$4 immediate
label branch target or symbol
-8(BX)(DI*4) memory
```
**Operands go source first, destination last**: `MOVQ x+0(FP), AX` loads the
argument into AX. This is the opposite of Intel order and the same order as
AT&T, with the sigils removed: registers are bare names, immediates take
`$`, memory is `offset(base)`.
## Registers
A register operand is its bare name, with no prefix: `AX`, `X15`, `R14` on
amd64; `R0` to `R30`, `ZR`, `V0` to `V31` on arm64; `X0` to `X31`, `F0` to
`F31`, `V0` on riscv64; `R0` to `R31`, `F0` to `F31`, `V0` on loong64.
Sub-register and width selection rides the mnemonic, not the operand: the
amd64 family spells `MOVB`, `MOVW`, `MOVL`, `MOVQ`, and the arm64 family
suffices `B`, `H`, `S`, `D`, `Q` on the shared forms. Each architecture page
lists its registers and the reserved ones.
## Immediates
`$` introduces a constant: `$42`, `$-1`, `$0x2a`, `$'A'`, `$bufSize`. The
`$` applies to the whole constant expression that follows, so
`$(4*8+reader__size)` is one immediate. Without the `$`, a number in operand
position is an address, not a value; the classic error `ADDQ 1, AX` asks the
assembler for the byte at address 1.
The one place a `$` number is not an immediate is the frame and argument
size field of TEXT, `$16-24`, which is two separate constants and not a
subtraction; DIRECTIVES.md specifies it.
## Memory
```text
offset(base)
offset(base)(index*scale)
```
Both parts are optional where the target allows them: `(BX)` is the memory
at BX, `foo+16(SB)` is a global, and on amd64 `foo+32(SP)(R9*8)` adds a
scaled index. `offset` is a constant expression, optionally carrying a
symbol name. The extensions beyond `offset(base)` are where the targets
diverge, and each belongs to its architecture page: amd64 carries the
`index*scale` form with scale 1, 2, 4 or 8 and its own rules on which
registers may index; loong64 writes base plus index as `(R4)(R5)`; the ARM
family attaches shift amounts to the index register in its own spelling.
The address arithmetic is on **byte addresses**: the offset is added to the
base as it stands, whatever the operand width of the instruction. Loading
the third 8-byte word of an array at BX is `16(BX)`, not `2(BX)`.
## The four pseudo-registers
Four names denote locations no target register holds, and they mean the same
on every architecture:
- **FP**, the frame pointer: the arguments and results of the current
function, at positive offsets, in the order the Go prototype declares
them. Every FP reference must carry a name: `x+0(FP)`, and an unnamed
`0(FP)` is rejected. Results follow arguments; an unnamed result is called
`ret`.
- **SP**, the virtual stack pointer: the high end of the function's local
frame, so locals live at negative offsets, `x-8(SP)`. A reference without
a name and without a plus, `-8(SP)`, addresses the **hardware** stack
pointer instead: the two spellings are one character apart and mean
different registers. That is the sharpest edge in the language and the
source of the deepest bugs.
- **SB**, the static base: the origin of memory, used for globals and
cross-package symbols, always with a name: `foo(SB)`, `foo+4(SB)`.
- **PC**, the program counter: branch targets, and the explicit relative
form `1(PC)`.
## Symbol names
A symbol's full name is the package path, a period, and the base name. In
source, the period is written U+00B7 (`·`) and a slash in the path U+2215
(`∕`), because the parser treats the ASCII forms as punctuation. Inside the
package's own file, `·Name` is enough and is the preferred spelling, since
it survives a rename of the import path.
| Spelling | Meaning |
|---|---|
| `·Name(SB)` | this package's Name |
| `runtime·morestack(SB)` | another package's morestack |
| `sourcedock.dev∕petrbalvin∕pkg·Name(SB)` | fully qualified |
| `msg<>(SB)` | file-local, the static of this language; `<>` also makes the ABI field static in the object |
| `Name<ABIInternal>(SB)` | ABI-qualified reference, the ABI in angle brackets after the name |
The object file these symbols produce, with the index rules that decide what
is referenced by name and what by index, is specified in
[GOOBJ.md](../GOOBJ.md).
## What vet adds in Go
Inside a Go package, `go vet`'s asmdecl analyzer checks every FP offset and
name against the Go prototype, and checks the declared argument area against
the frame. That layer, the prototype requirement and `go_asm.h`, belongs to
RUNTIME.md; the grammar above is the whole of what the assembler itself
requires.
+79
View File
@@ -0,0 +1,79 @@
# Preprocessing: include, define and selection
Layer 1, the common language. Verified against the preprocessor inside
`go tool asm` of Go 1.27.1 (`cmd/asm/internal/lex`), whose directives are
`#define`, `#undef`, `#include`, `#ifdef`, `#ifndef`, `#else`, `#endif` and
`#line`, and against gasm's implementation, which is differentially tested
against the toolchain's.
Input runs through a simplified C preprocessor before the parser sees it.
The set is deliberately small: there is no `#if` with constant expressions
and no token pasting with `##`. `#line` is honoured, so it changes the
positions the assembler reports and records.
## #include
```text
#include "textflag.h"
#include "go_asm.h"
#include "defs_linux_amd64.h"
```
The search path, in order: the directory of the including file, then the
directories given by repeatable `-I` flags. The assembler seeds no default
of its own: a bare `go tool asm` invocation finds none of the standard
headers, and it is the `go` build system that passes `$GOROOT/pkg/include`
among the `-I` directories when it drives the build. That directory ships
`textflag.h`, `funcdata.h` and the per architecture register headers.
Includes nest; a file included twice through different paths is processed
twice, which is why headers guard their defines.
## #define and #undef
```text
#define bufSize 1024
#define MOVD(d, s) MOVQ s, d
#undef bufSize
```
- An object macro replaces its name with its token sequence at the point of
use.
- A parameterised macro takes its arguments in parentheses and substitutes
them into the body. Macro parameters compose with the rest of the
language: an argument used with an element suffix, as in `A.S4` on the
vector forms, substitutes correctly.
- Redefinition is an error; `#undef` first, or pick a new name.
- The `-D name[=value]` flag predefines an object macro from the command
line, repeatable, exactly as `#define` would; a `-D` without a value
defines the name as `1`.
- Expansion happens when the name is used, so a macro may expand to
instructions, operands or fragments of either, and a macro body may use
macros defined before it.
`textflag.h` and `funcdata.h` are themselves ordinary `#define` files: the
flag names and the runtime macros are preprocessor definitions, not language
keywords. That is why a missing include produces a parser error at the first
use of `NOSPLIT` rather than a complaint about the name.
## #ifdef, #ifndef, #else, #endif
```text
#ifdef GOOS_windows
#define SYSCALL_INT 0x2b
#endif
```
Selection is by defined-name only: `#ifdef`, `#ifndef`, `#else`, `#endif`,
nesting freely. There is no `#if defined(x) && y`, because the preprocessor
evaluates no expressions; reach that with a build-tag Go file generating a
header, which is exactly how the runtime's own `go_asm.h` and defs headers
are produced.
## What preprocessing does not cover
The preprocessor is textual and runs first, so it knows nothing of assembly
semantics: it does not check that a macro expansion is a legal instruction,
and it does not participate in the constant expression evaluator, which runs
later, in the parser. A constant folded with `#define` and a constant folded
in an operand expression end at the same value through different doors;
GOOBJ.md records both in the object identically.
+56
View File
@@ -0,0 +1,56 @@
# The Plan 9 assembly language
This directory is the reference for the Plan 9 assembly language as the Go
toolchain and gasm accept it, written to be complete enough to implement
against. It exists because no such reference exists upstream: Go documents
the language on a single page, and the rest of the knowledge lives in the
toolchain's source and in the practice of reading it.
Every page carries the same conformance statement: which layer of the system
it describes, which toolchain release it was verified against, and how the
claims were checked. Pages in this directory are verified against Go 1.27.1
and against gasm's own differential test suite, which compares gasm's
behaviour with `go tool asm` byte for byte and output for output.
## The three layers
The reference deliberately separates three layers, because their rules have
different owners and different lifetimes:
1. **The common language** (LANGUAGE, OPERANDS, DIRECTIVES,
PREPROCESSOR): the syntax, operands, directives and preprocessing, the
same on every target and meaningful without a Go runtime.
2. **The Go-embedded layer** (RUNTIME): everything that exists only because
the code runs inside a Go program: the ABI0 contract, generated wrappers,
`go_asm.h`, the garbage collector annotations and `go vet` checks.
3. **The standalone layer** (STANDALONE, planned with the standalone
compilation phase): using the language outside Go, through gasm's ELF
output and the extended instruction set, where the toolchain offers no
ground truth and execution testing is the only verification.
A rule stated in layer 1 holds on every target. A rule stated in layer 2
says which part of the Go machinery imposes it. Nothing in layer 3 changes
layers 1 or 2; it extends them.
## Pages
| Page | Layer | Contents |
|---|---|---|
| [LANGUAGE.md](LANGUAGE.md) | 1 | lexicon, statement structure, labels, literals, expressions |
| [OPERANDS.md](OPERANDS.md) | 1 | operand grammar, pseudo-registers, addressing modes, symbol naming |
| [DIRECTIVES.md](DIRECTIVES.md) | 1 | TEXT, DATA, GLOBL, FUNCDATA, PCDATA, PCALIGN and the function flags |
| [PREPROCESSOR.md](PREPROCESSOR.md) | 1 | `#include`, `#define`, `#ifdef` and friends, `-D`, `-I` |
| [RUNTIME.md](RUNTIME.md) | 2 | ABI0, prototypes, `go_asm.h`, `funcdata.h`, `go vet` |
| [AMD64.md](AMD64.md) | 1 | registers, addressing, the frame and split check, families, relocations |
| [ARM64.md](ARM64.md) | 1 | registers, the MOV load and store series, special operand orders, SIMD |
| [RISCV64.md](RISCV64.md) | 1 | registers and their constrained names, per class operand order, profiles, vector extension |
| [LOONG64.md](LOONG64.md) | 1 | registers, width suffixes, vector element syntax, atomics and barriers |
| INSTRUCTIONS-AMD64.md and the other three | 1 | generated per architecture inventory of every accepted mnemonic |
| STANDALONE.md | 3 | the language outside Go |
## Status
The common-language core, the Go-embedded layer, all four per-architecture
pages and the generated instruction appendices are written and verified.
STANDALONE.md lands with the standalone compilation phase. The object format
these pages feed is specified in [GOOBJ.md](../GOOBJ.md).
+104
View File
@@ -0,0 +1,104 @@
# RISC-V 64
Layer 1, target page. Verified against `go tool asm` of Go 1.27.1, against
the toolchain's own riscv64 assembler manual (`cmd/internal/obj/riscv/doc.go`)
and against gasm's encoder, whose output is compared byte for byte with the
toolchain's. The complete mnemonic inventory lives in the generated appendix
[INSTRUCTIONS-RISCV64.md](INSTRUCTIONS-RISCV64.md).
## Registers
- Integer: `X0` to `X31`. `X0` is hardwired zero. Three names the toolchain
constrains: `X4` must be written through its ABI name `TP`; `X27`, the
goroutine pointer, must be written `g` and may not be written `S11`; in
shared builds `X3` is off limits and must be written `GP`.
- The other integer registers may be written `Xn` or by their ABI names
(`A0`, `T0`, `S1`, and so on).
- Floating point: `F0` to `F31`. Vector: `V0` to `V31`.
- `X26` is the closure pointer and `X31` is the assembler's own scratch
register: its value may be clobbered by instruction sequences the
assembler inserts, so hand-written code must not rely on it.
- There is no reserved frame pointer register on this target.
## Operand order
The ordering differs from the ISA manual, and per instruction class:
- **R-type** is reversed: `ADD X10, X11, X12` is `add x12, x11, x10`.
- **I-type arithmetic** keeps that shape with the immediate first:
`ADDI $1, X11, X12`.
- **Loads and stores** are source first, like every Plan 9 dialect:
`MOV 16(X2), X10` loads and `MOV X10, (X2)` stores. The MOV series hides
the width; `MOVB` through `MOVD` spell it out.
- **Branches** keep the ISA order: `BLT X12, X23, loop1`, which jumps when
X12 < X23, the reverse of the SLT operand order.
- **FMA** is rotated one place left so the destination comes last:
`FMADDS F1, F2, F3, F4`.
- **AMO** is likewise rotated: `AMOSWAPW X5, (X6), X7`.
- **Ternary abbreviation** is supported and encouraged: `ADD X10, X12` means
`ADD X10, X12, X12`.
Where an R-type instruction has an I-type sibling, the assembler picks the
immediate form from the operand: `AND $3, X12, X13` assembles as `ANDI`.
## Names, suffixes and rounding
Dots are removed and suffixes are upper-cased: the ISA's `fmv.w.x` is
`FMVWX`. Floating-point rounding modes become suffixes, `FCVTLUS.RNE F0,
X5`, with RTZ assumed when the suffix is omitted; the toolchain never sets
the FCSR.
## Constants
- `MOV` materialises any 64-bit integer constant, synthesising it from a
few arithmetic instructions where possible and otherwise loading it from
a literal pool in the binary.
- A 32-bit constant is accepted by `ADDI`, `ANDI`, `ORI` and `XORI`, and
the assembler synthesises values that exceed the 12-bit encoding window.
- `MOVF` and `MOVD` materialise floating-point constants, encoding them as
`FLW` and `FLD` from a pool location unless the constant is exactly 0.0.
## Extensions and profiles
The default target profile is rva20u64, selected or raised with the
GORISCV64 environment variable. A short list of instructions outside the
default profile is synthesised by the assembler when the profile does not
provide them, so they are safe without guards: `ANDN`, `MAX`, `MAXU`, `MIN`,
`MINU`, `MOVB`, `MOVH`, `MOVHU`, `MOVWU`, `ORN`, `ROL`, `ROLW`, `ROR`,
`RORI`, `RORIW`, `RORW`, `XNOR`. The header `asm_riscv64.h` defines the
`hasZba`, `hasZbb`, `hasZbs` and `hasV` macros for guarding everything else.
## Fences and atomics
`FENCE` takes predecessor and successor sets in that order, uppercase
letters, `FENCE R, RW`; a bare `FENCE` is a full fence, as is
`FENCE IORW, IORW`. `FENCE.TSO` exists. The ordering bits of `LR`, `SC`
and the AMO instructions are not specifiable in source: the assembler sets
acquire and release on the AMO instructions, acquire on `LR` and release on
`SC`, always.
## Compressed instructions
The assembler converts 32-bit instructions to their compressed encodings
automatically; the conversion is a property of the emitted machine code, not
of the source, and register choice influences how much compresses.
Hand-writing compressed instructions in source is accepted but discouraged.
The debug flag `compressinstructions=0` turns the automatic conversion off.
## Vector extension
`VSETVLI` writes its vtype components in uppercase with the destination
last: `VSETVLI X10, E8, M1, TU, MU, X12`. Vector loads and stores are
source first like the scalar ones, with an optional stride or index register
second and the mask register, when present, always penultimate:
`VLE8V (X10), V3`, `VLE8V (X10), V0, V3` for the masked form. Vector
arithmetic reverses its operands, `VADDVV V1, V2, V3`, with the mask again
penultimate.
## Relocations
`R_RISCV_JAL`, `R_RISCV_CALL`, the `R_RISCV_PCREL_ITYPE` and `STYPE` pairs,
`R_RISCV_BRANCH`, the compressed branch and jump forms, the TLS and GOT
families and `R_RISCV_ADD32` and `SUB32`, all specified in
[GOOBJ.md](../GOOBJ.md). The assembler always emits the four-byte
`R_DWTXTADDR_U4` flavour inside its DWARF records.
+120
View File
@@ -0,0 +1,120 @@
# The Go-embedded layer: ABI0, prototypes and the runtime contract
Layer 2: everything that exists only because the assembly runs inside a Go
program. Without a Go runtime this page does not apply; the language of
OPERANDS.md and DIRECTIVES.md still does. Verified against Go 1.27.1, against
the shipped `funcdata.h` header, and against the object files the toolchain
produces, which were parsed and checked field by field while writing
[GOOBJ.md](../GOOBJ.md).
## Hand-written assembly is ABI0
Go functions compiled from source use ABIInternal, the register-based
calling convention, which the toolchain documents as unstable and free to
change between releases. A `.s` function is written against ABI0, the stack
based convention: arguments and results live in the caller's frame at
positive FP offsets, byte-addressed, in declaration order, with no registers
assigned at all. The toolchain generates the wrapper that translates between
the two; a caller in Go calling an assembly function goes through it, and it
is marked `ABIWRAPPER` in the object. Hand-writing a bridge is never needed
and never correct.
## Every assembly function carries a Go prototype
```go
package add
func Add(x, y int64) int64
```
The body-less declaration is not optional, and not only for the linker: it
is what tells the garbage collector which arguments and results hold
pointers, and what `go vet` checks the assembly against. Even a function
nothing in Go calls gets one. Consequences:
- The FP operand names and offsets are checked by vet's asmdecl analyzer
against the prototype: `x+0(FP)` must name an argument that exists, at the
offset the prototype says. A file that assembles and links can still fail
vet.
- The declared argument area in `$framesize-argsize` is checked against the
prototype's size. An omitted argsize marks the argument size unknown
(0x80000000 in the object, the value of `ArgsSizeUnknown` from
`funcdata.h`), which is the normal spelling for functions with no Go
callers.
- `//go:noescape` on the declaration tells the compiler that a pointer
argument does not escape, for assembly that keeps the pointer beyond the
call.
## The frame, the stack and the collector
The runtime owns the stack and the pointer map, and assembly must hold up
its end of four rules:
1. **Arguments are initialised on entry; results are not.** A function whose
results hold live pointers across a call must zero them and then execute
`GO_RESULTS_INITIALIZED`. Designing functions that return no pointers
avoids the problem.
2. **A frame with calls and no local pointers says so** with
`NO_LOCAL_POINTERS`. A frame with local pointers that the runtime cannot
see is not allowed at all: assembly cannot describe a pointer-containing
local, so it must not have one. Data symbols containing pointers are the
same: define them in Go.
3. **The stack may move.** Stack growth copies the frame, so no pointer into
the frame may be held across a call, and the raw hardware SP register may
not be cached across a call either.
4. **The split check is not optional by default.** Without NOSPLIT, the
assembler inserts the stack-growth preamble, including the morestack
block for framed functions; NOSPLIT is a contract that the frame and
everything below it fit in the remaining stack segment. On amd64 the
assembler also marks small leaf functions NoSplit itself and skips the
preamble, so silence is not a promise.
The simplest safe shape is a leaf function with no local frame and no calls:
it needs no annotation beyond the prototype.
## go_asm.h: Go constants and layout in assembly
A package with `.s` files gets a generated header. Include it and use the
generated names instead of hard-coding layouts, which lie silently when the
Go side changes:
| Go declaration | Assembly name |
|---|---|
| `const bufSize = 1024` | `const_bufSize` |
| field `r` of `type reader struct` | `reader_r` |
| size of `type reader struct` | `reader__size` |
The constants arrive as macros, usable as immediates and offsets, computed
from the Go declarations. An ambiguous name, such as a struct that really
has a `_size` field, fails the generation with a redefinition error.
## funcdata.h: the runtime macros
`$GOROOT/pkg/include/funcdata.h` defines the PCDATA and FUNCDATA ids and the
three macros assembly normally uses instead:
| Macro | Expands to | Meaning |
|---|---|---|
| `GO_ARGS` | `FUNCDATA $FUNCDATA_ArgsPointerMaps, go_args_stackmap(SB)` | the Go prototype defines the argument pointer map |
| `GO_RESULTS_INITIALIZED` | `PCDATA $PCDATA_StackMapIndex, $1` | results are initialised; treat them as live from here |
| `NO_LOCAL_POINTERS` | `FUNCDATA $FUNCDATA_LocalsPointerMaps, no_pointers_stackmap(SB)` | the frame holds no pointers |
`GO_ARGS` is inserted implicitly by the assembler for any function whose
package-qualified name belongs to the current package, which is why most
assembly never writes it. `NOSPLIT` leaf functions that call nothing need
none of the three.
The underlying ids, for reading toolchain output rather than for writing
source: FUNCDATA 0 to 7 are args pointer maps, locals pointer maps, stack
objects, inline tree, open-coded defer info, argument info, argument
liveness and wrap info; PCDATA 0 to 4 are unsafe point, stack map index,
inline tree index, argument liveness index and panic bounds.
## What the runtime does with all of this
The object file records the annotations as aux symbols and FuncInfo records;
GOOBJ.md specifies the encoding. The linker assembles them into the runtime's
pclntable, which traceback and the collector consume. An assembly function
that misdeclares its frame is not a compile error and usually not a link
error: it is a wrong collector decision or a wrong traceback at runtime,
which is why the annotations are a contract and not documentation.
+14 -1
View File
@@ -2,7 +2,7 @@
.SH NAME
gasm-asm \- assemble Plan 9 assembly without the Go toolchain
.SH SYNOPSIS
.B gasm asm [\-\-format raw|elf|goobj] [\-I dir] [\-p pkg] [\-GOARCH arch] [\-o out] <file>
.B gasm asm [\-\-format raw|elf|goobj] [\-I dir] [\-p pkg] [\-GOARCH arch] [\-GOOS os] [\-o out] <file>
.SH DESCRIPTION
Assemble FILE without the Go toolchain: every TEXT function is encoded
to machine code and printed as a hex dump. Supported architectures:
@@ -42,6 +42,15 @@ need no toolchain at all.
Framed functions receive the stack-split guard and the trailing
morestack block, byte-identical to the toolchain's output, so split
functions link too.
.PP
A file that includes go_asm.h gets that header generated from the Go
files beside it, type-checked for the target.
.B \-GOOS
selects the type-checking GOOS for that header, because a GOOS-specific
file needs its platform's defines: sys_darwin_arm64.s fails against the
ambient GOOS (machTimebaseInfo_numer is missing from a linux type-check)
and assembles with
.BR "\-GOOS darwin" .
.SH OPTIONS
.TP
.B \-\-format \fIraw|elf|goobj\fR
@@ -59,6 +68,10 @@ Target architecture: amd64, arm64, riscv64 or loong64; overrides the
file-name suffix, which is how the suffix-less majority of GOROOT's
files (cpu_x86.s, stub.s, ...) become assemblable.
.TP
.B \-GOOS \fIos\fR
Operating system for the generated go_asm.h: any GOOS go/build
recognises in file names; the default is the host's.
.TP
.B \-o \fIfile\fR
Write the output to this file instead of a hex dump on stdout.
.SH EXIT STATUS
+8 -2
View File
@@ -1,8 +1,8 @@
.TH GASM-AUDIT-INSTRUCTIONS 1 "2026-09-19" "gasm" "User Commands"
.TH GASM-AUDIT-INSTRUCTIONS 1 "2026-09-21" "gasm" "User Commands"
.SH NAME
gasm-audit-instructions \- diff the encoder against the Go toolchain, or measure a corpus
.SH SYNOPSIS
.B gasm audit\-instructions [\-\-corpus [\fIdir\fR]] [\-I dir] [amd64|arm64|riscv64|loong64]
.B gasm audit\-instructions [\-\-corpus [\fIdir\fR]] [\-\-list] [\-I dir] [amd64|arm64|riscv64|loong64]
.SH DESCRIPTION
Compare the gasm encoder for the given architecture (default amd64)
against
@@ -39,6 +39,12 @@ second.
Assemble a corpus of .s files and report pass rates and failure
reasons.
.TP
.B \-\-list
With
.BR \-\-corpus ,
print every failing file with its failure reason, per architecture,
instead of one representative file per reason.
.TP
.B \-I \fIdir\fR
Directory to search for #include files; may be repeated, searched in
order after the source directory. A corpus run whose files include
+23 -6
View File
@@ -241,6 +241,13 @@ func renderInstr(line []token.Token, width int) string {
if line[0].Kind != token.Ident {
return "\t" + mnem + " " + ops
}
// A statement separator belongs to the statement it ends: when the
// operands open with a ';', the alignment padding would land between
// the mnemonic and its own separator (REP ; MOVSQ), so such a line
// renders with a single space whatever the function's width.
if strings.HasPrefix(ops, ";") {
return "\t" + mnem + " " + ops
}
if width < len(mnem) {
width = len(mnem)
}
@@ -254,11 +261,12 @@ func renderPreproc(line []token.Token) string {
line[2].Kind == token.String {
return "#include " + line[2].Text
}
parts := make([]string, 0, len(line)-1)
for _, t := range line[1:] {
parts = append(parts, t.Text)
}
return "#" + strings.Join(parts, " ")
// The body of a directive, a macro definition included, is an ordinary
// token run: rendering it through renderOps applies the same punctuation
// rules as everywhere else, so a macro body keeps its canonical spelling
// ($v, (a, b), the ';' separators between statements) instead of being
// spread with a space between every token.
return "#" + renderOps(line[1:])
}
// renderOps re-spaces a run of operand tokens into canonical form. It never
@@ -348,6 +356,13 @@ func spaceBetween(prev, cur token.Token) bool {
return false
case token.Comma:
return false
case token.Semicolon:
// A ';' is a statement separator on the assembly path, not an
// operand: dropping it would fuse two statements into a line the
// assembler rejects, so it must survive as punctuation. It glues
// to the statement it ends and the next statement takes one space,
// matching the toolchain's listing style.
return false
case token.Star, token.Plus, token.Minus, token.Slash, token.Pipe:
return false
case token.LShift, token.RShift, token.Arrow, token.At:
@@ -376,7 +391,9 @@ func spaceBetween(prev, cur token.Token) bool {
return false
case token.LAngle, token.RAngle:
return false
case token.Comma:
case token.Comma, token.Semicolon:
// The statement after a ';' separator takes its own space, exactly
// like the operand after a comma.
return true
}
return true
+192
View File
@@ -5,9 +5,11 @@ package format
import (
"os"
"slices"
"strings"
"testing"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
"sourcedock.dev/petrbalvin/gasm-devkit/lexer"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
"sourcedock.dev/petrbalvin/gasm-devkit/token"
@@ -298,6 +300,196 @@ func TestCRLFInputIsNormalisedToLF(t *testing.T) {
}
}
// TestSemicolonSeparators pins the treatment of ';' statement separators.
// The separator is load-bearing on the assembly path, where the parser reads
// semicolon-separated statements: a formatter that drops it fuses two
// statements into a line the assembler rejects, which is data corruption.
// Each row pins the canonical spelling, one space after the ';', tight
// before it, the way the toolchain's own sources and listings write it.
func TestSemicolonSeparators(t *testing.T) {
cases := []struct {
name string
in string
want string
}{
{
name: "between instructions, tight",
in: "TEXT ·f(SB), $0\nBYTE $0x48;BYTE $0xc7\nRET\n",
want: "TEXT ·f(SB), $0\n\tBYTE $0x48; BYTE $0xc7\n\tRET\n",
},
{
name: "between instructions, spaced",
in: "TEXT ·f(SB), $0\nBYTE $0x48 ; BYTE $0xc7\nRET\n",
want: "TEXT ·f(SB), $0\n\tBYTE $0x48; BYTE $0xc7\n\tRET\n",
},
{
name: "after a label",
in: "TEXT ·f(SB), $0\nlabel: BYTE $1; BYTE $2\nRET\n",
want: "TEXT ·f(SB), $0\nlabel:\n\tBYTE $1; BYTE $2\n\tRET\n",
},
{
// The continuation-spliced macro shape of the runtime sources:
// the lexer makes one logical line of the backslash continuations.
name: "inside a macro body, continued",
in: "#define MOVLTOREG(v, off) \\\n\tMOVL $v, AX; \\\n\tMOVL AX, ret+off(FP)\n",
want: "#define MOVLTOREG(v, off) MOVL $v, AX; MOVL AX, ret+off(FP)\n",
},
{
name: "inside a macro body, one line",
in: "#define PEAS BYTE $0x0a; BYTE $0x0b\n",
want: "#define PEAS BYTE $0x0a; BYTE $0x0b\n",
},
{
name: "several separators in one line",
in: "TEXT ·f(SB), $0\nBYTE $1; BYTE $2; BYTE $3\nRET\n",
want: "TEXT ·f(SB), $0\n\tBYTE $1; BYTE $2; BYTE $3\n\tRET\n",
},
{
name: "two separators back to back",
in: "TEXT ·f(SB), $0\nBYTE $1;; BYTE $2\nRET\n",
want: "TEXT ·f(SB), $0\n\tBYTE $1;; BYTE $2\n\tRET\n",
},
{
name: "inside a line comment, untouched",
in: "TEXT ·f(SB), $0\n// keep; the; separators\nBYTE $1\nRET\n",
want: "TEXT ·f(SB), $0\n\t// keep; the; separators\n\tBYTE $1\n\tRET\n",
},
{
name: "after a statement, before a comment",
in: "TEXT ·f(SB), $0\nMOVQ AX, BX; // tail\nRET\n",
want: "TEXT ·f(SB), $0\n\tMOVQ AX, BX; // tail\n\tRET\n",
},
{
name: "last character on a line",
in: "TEXT ·f(SB), $0\nBYTE $1;\nRET\n",
want: "TEXT ·f(SB), $0\n\tBYTE $1;\n\tRET\n",
},
{
// The REP shape: a prefix-style zero-operand statement
// followed by the instruction it prefixes. The separator
// belongs to the statement it ends, so the function's
// alignment width (MOVSQ is the widest mnemonic here) must
// not open a gap before it: one space after the mnemonic
// whatever the neighbours' lengths.
name: "after a prefix-style statement",
in: "TEXT ·f(SB), $0\nMOVQ AX, BX\nREP; MOVSQ\nRET\n",
want: "TEXT ·f(SB), $0\n\tMOVQ AX, BX\n\tREP ; MOVSQ\n\tRET\n",
},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
got := Source(tc.in)
if got != tc.want {
t.Fatalf("formatting mismatch:\n--- got ---\n%q\n--- want ---\n%q", got, tc.want)
}
if again := Source(got); again != got {
t.Fatalf("not idempotent:\n%q", again)
}
if in, out := strings.Count(tc.in, ";"), strings.Count(got, ";"); in != out {
t.Fatalf("semicolon count changed: %d -> %d\n%s", in, out, got)
}
if _, errs := parser.Parse("in.s", got); len(errs) > 0 {
t.Fatalf("formatted output no longer parses: %v", errs)
}
})
}
}
// TestSemicolonStatementRoundTrip proves the formatter's contract on the
// path where ';' separates statements: parse the source the way the
// assembler does, format it, re-parse the formatted text and compare the
// statement sequence. Raw operand texts are token-joined, so they are
// insensitive to the whitespace a format pass chooses, and the comparison
// can only fail when a token is lost: dropping a ';' fuses two statements
// into one, exactly the corruption the released formatter committed.
func TestSemicolonStatementRoundTrip(t *testing.T) {
src := "#define MOVLTOREG(v, off) \\\n" +
"\tMOVL $v, AX; \\\n" +
"\tMOVL AX, ret+off(FP)\n" +
"\n" +
"TEXT ·f(SB), NOSPLIT, $0\n" +
"BYTE $0x48; BYTE $0xc7\n" +
"first: BYTE $1; BYTE $2\n" +
"MOVLTOREG($42, 0)\n" +
"RET\n"
before, errs := parser.ParseWithOptions("in.s", src, parser.Options{Expand: true})
if len(errs) > 0 {
t.Fatalf("source does not parse: %v", errs)
}
formatted := Source(src)
after, errs := parser.ParseWithOptions("in.s", formatted, parser.Options{Expand: true})
if len(errs) > 0 {
t.Fatalf("formatted source does not parse: %v", errs)
}
want, got := stmtSignature(before), stmtSignature(after)
if !slices.Equal(got, want) {
t.Fatalf("statement sequence changed:\n--- before ---\n%q\n--- after ---\n%q", want, got)
}
if again := Source(formatted); again != formatted {
t.Fatalf("not idempotent:\n%q", again)
}
// The two BYTE statements on the first line must stay two: one fused
// statement here is the exact defect this package once shipped.
var bytes []string
for _, stmt := range stmtSignature(after) {
if rest, ok := strings.CutPrefix(stmt, "instr BYTE "); ok {
bytes = append(bytes, rest)
}
}
if want := []string{"$ 0x48", "$ 0xc7", "$ 1", "$ 2"}; !slices.Equal(bytes, want) {
t.Fatalf("BYTE statements after expansion = %q, want %q", bytes, want)
}
}
// stmtSignature flattens a parsed file into one string per declaration and
// statement, in source order. Every component is token-derived, so the
// signature is stable across format passes and moves only when a token is
// lost or gained.
func stmtSignature(f *ast.File) []string {
var out []string
for _, d := range f.Decls {
switch d := d.(type) {
case *ast.Text:
out = append(out, "text "+d.Name.Raw)
for _, s := range d.Body {
out = append(out, stmtText(s))
}
case *ast.Globl:
out = append(out, "globl "+d.Name.Raw)
case *ast.Data:
out = append(out, "data "+d.Name.Raw)
case *ast.Include:
out = append(out, "include "+d.Header.Text)
case *ast.Preproc:
out = append(out, "preproc "+d.Raw)
}
}
for _, s := range f.Orphans {
out = append(out, stmtText(s))
}
return out
}
// stmtText renders one statement for stmtSignature.
func stmtText(s ast.Stmt) string {
switch s := s.(type) {
case *ast.Label:
return "label " + s.Name.Text
case *ast.Instr:
parts := make([]string, 0, len(s.Operands)+1)
parts = append(parts, s.Mnemonic.Text)
for _, op := range s.Operands {
parts = append(parts, op.Raw)
}
return "instr " + strings.Join(parts, " ")
default:
return "stmt"
}
}
// lexOperands lexes a single operand string and drops the EOF token.
func lexOperands(s string) []token.Token {
toks := lexer.Tokenize(s)
+11 -3
View File
@@ -16,8 +16,16 @@ import (
"sourcedock.dev/petrbalvin/gasm-devkit/token"
)
// middleDot is the Plan 9 symbol separator (U+00B7), used in ·funcName(SB).
const middleDot = '\u00B7'
const (
// middleDot is the Plan 9 symbol separator (U+00B7), used in
// ·funcName(SB): it stands for the period between package path and name.
middleDot = '\u00B7'
// divisionSlash is the Plan 9 path separator (U+2215), used inside the
// package path of a symbol: internal∕runtime∕atomic·Xchg. Like the
// middle dot it is an identifier character, so a package path containing
// it lexes as one name; the ordinary slash (U+002F) stays punctuation.
divisionSlash = '\u2215'
)
// Lexer scans a source string one token at a time.
type Lexer struct {
@@ -460,7 +468,7 @@ func isHexDigit(r rune) bool {
}
func isIdentStart(r rune) bool {
return r == '_' || r == middleDot || unicode.IsLetter(r)
return r == '_' || r == middleDot || r == divisionSlash || unicode.IsLetter(r)
}
func isIdentChar(r rune) bool {
+13
View File
@@ -182,6 +182,19 @@ func TestNulIsIllegal(t *testing.T) {
eq(t, texts("MOVQ \x00 AX"), []string{"MOVQ", "\x00", "AX"})
}
func TestDivisionSlashInIdentifiers(t *testing.T) {
// U+2215 DIVISION SLASH is an identifier character, the way the
// toolchain's tokenizer treats it: the package path of a symbol is
// written with it (internal∕runtime∕atomic·Xchg) and must lex as one
// name. The ordinary slash (U+002F) stays punctuation.
eq(t, texts("CALL internal∕runtime∕atomic·Xchg(SB)"),
[]string{"CALL", "internal∕runtime∕atomic·Xchg", "(", "SB", ")"})
eq(t, texts("MOVQ sync∕atomic·Align(SB), AX"),
[]string{"MOVQ", "sync∕atomic·Align", "(", "SB", ")", ",", "AX"})
// It may also begin a name, like any letter of the toolchain's rule.
eq(t, kinds("∕x"), []token.Kind{token.Ident})
}
// TestOffsetsAroundInvalidByte pins Position.Offset against the original
// bytes: an invalid UTF-8 byte decodes to RuneError but advances the offset
// table by exactly one byte, so every later position stays a true byte
+146 -8
View File
@@ -32,9 +32,8 @@ func (e Error) Error() string {
// returned file is usable even when errors is non-empty.
func Parse(path, src string) (*ast.File, []error) {
tokens := lexer.Tokenize(src)
lines := splitLines(tokens)
p := &state{path: path}
p.parse(lines)
p.parse(statementLines(tokens))
return p.file, p.errs
}
@@ -73,6 +72,46 @@ func splitLines(tokens []token.Token) [][]token.Token {
return lines
}
// statementLines turns the token stream into the logical lines the parser
// reads: physical lines split at the ';' statement separators, exactly the
// way the expansion path treats the expanded bodies. The runtime writes
// "ROLQ $3, DI; ROLQ $13, DI" and "REP; MOVSB" in plain files, and the
// separator carries no meaning beyond the break. Comments are statement
// text, not structure: the lexer delivers a whole comment as one token, so
// a ';' inside a comment is never a separator; a comment after a statement
// stays on that statement's line; and a comment that sits between
// statements (the runtime's "NO_LOCAL_POINTERS; /* … */" style) stands as
// its own logical line, like a whole-line comment.
func statementLines(tokens []token.Token) [][]token.Token {
var out [][]token.Token
var cur []token.Token
flush := func() {
if len(cur) > 0 {
out = append(out, cur)
cur = nil
}
}
for _, t := range tokens {
switch t.Kind {
case token.EOF:
// The stream's terminator is not statement content.
case token.Newline, token.Semicolon:
flush()
case token.Comment:
if len(cur) > 0 {
cur = append(cur, t)
} else {
out = append(out, []token.Token{t})
}
flush()
default:
cur = append(cur, t)
}
}
flush()
return out
}
func (p *state) parse(lines [][]token.Token) {
p.file = &ast.File{Path: p.path, Macros: map[string]bool{}}
for _, line := range lines {
@@ -265,7 +304,7 @@ func (p *state) parseGlobl(line []token.Token) *ast.Globl {
rest = rest[1:]
}
if len(rest) > 0 && rest[0].Kind == token.Dollar {
g.Size = parseOperand(rest)
g.Size = parseOperand(rest, false)
}
return g
}
@@ -282,7 +321,7 @@ func (p *state) parseData(line []token.Token) *ast.Data {
d.Name = sym
d.Width = width
if len(valuePart) > 0 {
d.Value = parseOperand(stripComment(valuePart))
d.Value = parseOperand(stripComment(valuePart), false)
}
return d
}
@@ -293,8 +332,13 @@ func (p *state) parseInstr(line []token.Token) {
return
}
instr := &ast.Instr{Mnemonic: body[0], Comment: comment}
for _, grp := range splitOperands(body[1:]) {
if op := parseOperand(grp); op != nil {
grps := splitOperands(body[1:])
for i, grp := range grps {
// Only the final operand slot may carry a bare constant: the
// toolchain reads the trailing 1 of CMPSD X1, X0, 1 as $1
// (math/floor_amd64.s), while an earlier bare number names an
// absolute address, a form this parser keeps out of the tree.
if op := parseOperand(grp, i == len(grps)-1); op != nil {
instr.Operands = append(instr.Operands, op)
}
}
@@ -378,8 +422,10 @@ func setName(raw string, sym *ast.Symbol) {
// --- operand parsing --------------------------------------------------------
// parseOperand parses one operand group into an Operand.
func parseOperand(g []token.Token) *ast.Operand {
// parseOperand parses one operand group into an Operand. allowBare marks
// the final operand slot of an instruction, where the toolchain reads a
// bare constant expression as an immediate.
func parseOperand(g []token.Token, allowBare bool) *ast.Operand {
g = stripComment(g)
if len(g) == 0 {
return nil
@@ -392,9 +438,25 @@ func parseOperand(g []token.Token) *ast.Operand {
}
op.Kind = ast.OpAddr
op.Addr = parseAddress(g)
// A trailing bare constant leaves every address field empty: the
// grammar sees no register, memory reference or symbol, and the closed
// constant expression is the whole group. Read it as the immediate it
// names, exactly what the $ spelling would produce.
if allowBare && isEmptyAddress(op.Addr) {
if v, rest, ok := foldExpr(g); ok && len(rest) == 0 {
op.Kind = ast.OpImmediate
op.Imm = ast.Immediate{Val: v, HasVal: true}
}
}
return op
}
// isEmptyAddress reports whether parseAddress populated nothing, its sign
// that the group is no register, memory reference, symbol or register range.
func isEmptyAddress(a ast.Address) bool {
return a.Sym == nil && a.Base == "" && a.Index == "" && a.Range == nil && a.Shift == ""
}
// parseImmediate parses the tokens following a '$'.
func parseImmediate(g []token.Token) ast.Immediate {
var imm ast.Immediate
@@ -427,6 +489,17 @@ func parseImmediate(g []token.Token) ast.Immediate {
} else if g[i].Kind == token.Plus {
i++
}
// A constant expression after the sign: $-(R - 8), $+(32-shift). The
// toolchain folds the negated value in place (the cgo ABI macros write
// ADJSP $-(REGS_HOST_TO_ABI0_STACK - 8)), so the sign applies to the
// folded value exactly as it does to a bare literal.
if i < len(g) && (g[i].Kind == token.LParen || g[i].Kind == token.Tilde) {
if v, rest, ok := foldExpr(g[i:]); ok && len(rest) == 0 {
imm.Val = v
imm.HasVal = true
return imm
}
}
if i < len(g) && g[i].Kind == token.Number {
text := g[i].Text
if v, ok := tryInt(text); ok {
@@ -458,6 +531,14 @@ func parseAddress(g []token.Token) ast.Address {
if len(g) == 0 {
return addr
}
// A bracketed register range, [Z0-Z3]: the amd64 4FMAPS/4VNNIW
// multi-source operand. The bracket runes arrive as Illegal tokens
// (the lexer has no bracket kind), so the shape matches on their text.
if isBracket(g[0], "[") && len(g) == 5 && g[1].Kind == token.Ident &&
g[2].Kind == token.Minus && g[3].Kind == token.Ident && isBracket(g[4], "]") {
addr.Range = &ast.RegRange{Lo: g[1].Text, Hi: g[3].Text, Pos: g[0].Pos}
return addr
}
// Symbol-with-pseudo form: name[<>][+off](PSEUDO).
// When the prefix is not a valid symbol name (e.g. a bare number like
// 0(SP) in RISC-V), sym is nil, and we fall through to regular memory
@@ -482,6 +563,20 @@ func parseAddress(g []token.Token) ast.Address {
i = len(g) - len(rest)
}
}
// The same expression under a leading sign: -(24+8)(X6) puts the sign
// outside the fold. The base group must follow for the value to
// commit, exactly as in the unsigned branch above.
if i < len(g) && (g[i].Kind == token.Minus || g[i].Kind == token.Plus) &&
i+1 < len(g) && g[i+1].Kind == token.LParen {
if v, rest, ok := foldExpr(g[i+1:]); ok && len(rest) > 0 && rest[0].Kind == token.LParen {
if g[i].Kind == token.Minus {
v = -v
}
addr.Offset = v
addr.HasOff = true
i = len(g) - len(rest)
}
}
// Optional leading displacement before a '(' base group. A sign pushes
// the parenthesis one token further out: -4(DX) has it at i+2.
if isSignedNumber(g, i) {
@@ -538,6 +633,16 @@ func parseAddress(g []token.Token) ast.Address {
}
}
}
// A lone (index*scale) group is the VSIB index-only form: the
// gather/scatter families address memory through a scaled vector index
// with no base register, 8(X4*1). The two-group grammar below reads
// (base)(index*scale), so a first group whose member carries a scale
// factor can only be an index.
if isIndexGroup(g[i:]) {
addr.Index = g[i+1].Text
addr.Scale = int(parseInt(g[i+3].Text))
i += 5
}
// First parenthesised group: the base register.
if i < len(g) && g[i].Kind == token.LParen {
i++
@@ -579,6 +684,26 @@ func parseAddress(g []token.Token) ast.Address {
if i > 0 && i < len(g) {
addr.Shift = joinRaw(g[i:])
}
// A lone (possibly signed) number is an absolute address: MOVL $0xf1,
// 0xf1 stores through the bare displacement with no base at all. In
// operand position a number without $ is an address, never a value.
if addr.Sym == nil && addr.Base == "" && addr.Index == "" && !addr.HasOff {
neg := false
j := 0
if j < len(g) && (g[j].Kind == token.Minus || g[j].Kind == token.Plus) {
neg = g[j].Kind == token.Minus
j++
}
if j == len(g)-1 && g[j].Kind == token.Number {
v := parseInt(g[j].Text)
if neg {
v = -v
}
addr.Offset = v
addr.HasOff = true
return addr
}
}
return addr
}
@@ -594,6 +719,19 @@ func findPseudoParen(g []token.Token) int {
return -1
}
// isBracket reports whether t is a square bracket. The lexer has no bracket
// kind, so '[' and ']' arrive as Illegal tokens.
func isBracket(t token.Token, text string) bool {
return t.Kind == token.Illegal && t.Text == text
}
// isIndexGroup reports whether g begins with a complete (index*scale) group:
// one identifier followed by a scale factor, all inside a single parenthesis.
func isIndexGroup(g []token.Token) bool {
return len(g) >= 5 && g[0].Kind == token.LParen && g[1].Kind == token.Ident &&
g[2].Kind == token.Star && g[3].Kind == token.Number && g[4].Kind == token.RParen
}
// --- token helpers ----------------------------------------------------------
// splitOperands splits a token slice on top-level commas (commas outside any
+257
View File
@@ -401,3 +401,260 @@ func TestInt64MinimumImmediate(t *testing.T) {
t.Errorf("imm.Float = %q, want empty", imm.Float)
}
}
// TestDivisionSlashPackagePath covers the runtime's package-path spelling:
// U+2215 DIVISION SLASH separates the elements of an import path inside a
// symbol (internal∕runtime∕atomic·Xchg), and the middle dot still separates
// the package from the name. The whole spelling must reach the symbol, not
// stop at the first slash.
func TestDivisionSlashPackagePath(t *testing.T) {
file, errs := Parse("t.s", "TEXT \u00b7f(SB), $0\n\tCALL internal∕runtime∕atomic·Xchg(SB)\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse errors: %v", errs)
}
txt := file.Decls[0].(*ast.Text)
instr := txt.Body[0].(*ast.Instr)
sym := instr.Operands[0].Addr.Sym
if sym == nil {
t.Fatal("operand carries no symbol")
}
if sym.Pkg != "internal∕runtime∕atomic" {
t.Errorf("pkg = %q, want internal∕runtime∕atomic", sym.Pkg)
}
if sym.Name != "Xchg" {
t.Errorf("name = %q, want Xchg", sym.Name)
}
if sym.Raw != "internal∕runtime∕atomic·Xchg(SB)" {
t.Errorf("raw = %q", sym.Raw)
}
}
// TestSemicolonStatements covers the plain parse path: ';' separates
// statements on one line exactly as it does inside macro expansion, and a
// ';' inside a comment is comment text.
func TestSemicolonStatements(t *testing.T) {
file, errs := Parse("t.s", "TEXT \u00b7f(SB), $0\n\tROLQ $3, DI; ROLQ $13, DI\n\tMOVQ AX, BX // note; still comment\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse errors: %v", errs)
}
txt := file.Decls[0].(*ast.Text)
if len(txt.Body) != 4 {
t.Fatalf("body = %d statements, want 4", len(txt.Body))
}
first := txt.Body[0].(*ast.Instr)
if first.Mnemonic.Text != "ROLQ" || len(first.Operands) != 2 {
t.Errorf("first statement = %+v, want ROLQ with two operands", first.Mnemonic)
}
second := txt.Body[1].(*ast.Instr)
if second.Mnemonic.Text != "ROLQ" || len(second.Operands) != 2 {
t.Errorf("second statement = %s, want ROLQ with two operands", second.Mnemonic.Text)
}
// The trailing comment belongs to the second MOVQ, semicolon included.
third := txt.Body[2].(*ast.Instr)
if third.Mnemonic.Text != "MOVQ" || third.Comment != "note; still comment" {
t.Errorf("third = %s, comment %q", third.Mnemonic.Text, third.Comment)
}
}
// TestSemicolonAfterLabel covers a label sharing its line with two
// statements.
func TestSemicolonAfterLabel(t *testing.T) {
file, errs := Parse("t.s", "TEXT \u00b7f(SB), $0\nloop: NOP; NOP\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse errors: %v", errs)
}
txt := file.Decls[0].(*ast.Text)
if len(txt.Body) != 4 {
t.Fatalf("body = %d statements, want 4 (label, two instructions, RET)", len(txt.Body))
}
if _, ok := txt.Body[0].(*ast.Label); !ok {
t.Errorf("first statement = %T, want *ast.Label", txt.Body[0])
}
for i, want := range []string{"NOP", "NOP", "RET"} {
in, ok := txt.Body[i+1].(*ast.Instr)
if !ok || in.Mnemonic.Text != want {
t.Errorf("statement %d = %v, want %s", i+1, txt.Body[i+1], want)
}
}
}
// TestParseEqualsZeroOptions pins the contract that ParseWithOptions with
// the zero Options reproduces Parse, here for the semicolon split.
func TestParseEqualsZeroOptions(t *testing.T) {
src := "TEXT \u00b7f(SB), $0\n\tNOP; NOP\n\tRET\n"
a, errsA := Parse("t.s", src)
b, errsB := ParseWithOptions("t.s", src, Options{})
if len(errsA) > 0 || len(errsB) > 0 {
t.Fatalf("errors: %v / %v", errsA, errsB)
}
ta, tb := texts(a), texts(b)
if len(ta) != len(tb) {
t.Fatalf("decl counts differ: %d vs %d", len(ta), len(tb))
}
for i := range ta {
if len(ta[i].Body) != len(tb[i].Body) {
t.Fatalf("TEXT %d: body lengths differ: %d vs %d", i, len(ta[i].Body), len(tb[i].Body))
}
}
}
// TestBracketRegisterRange pins the amd64 multi-source operand of the
// 4FMAPS/4VNNIW families: the bracket group [Z0-Z3] names four consecutive
// source registers and must reach the AST as a register range instead of an
// empty address.
func TestBracketRegisterRange(t *testing.T) {
file, errs := Parse("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tV4FMADDPS 17(SP), [Z0-Z3], K2, Z0\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse errors: %v", errs)
}
fn := file.Decls[0].(*ast.Text)
in := fn.Body[0].(*ast.Instr)
if len(in.Operands) != 4 {
t.Fatalf("operands = %d, want 4", len(in.Operands))
}
rng := in.Operands[1]
if rng.Kind != ast.OpAddr {
t.Errorf("range operand kind = %v, want OpAddr", rng.Kind)
}
if rng.Addr.Range == nil {
t.Fatalf("range operand = %+v, want a register range", rng.Addr)
}
if rng.Addr.Range.Lo != "Z0" || rng.Addr.Range.Hi != "Z3" {
t.Errorf("range = %s-%s, want Z0-Z3", rng.Addr.Range.Lo, rng.Addr.Range.Hi)
}
if rng.Addr.Sym != nil || rng.Addr.Base != "" || rng.Addr.Index != "" || rng.Addr.Shift != "" {
t.Errorf("range operand carries stray address fields: %+v", rng.Addr)
}
if rng.Raw != "[ Z0 - Z3 ]" {
t.Errorf("range raw = %q, want the verbatim spelling", rng.Raw)
}
}
// TestBracketRegisterRangeNotList pins that arm64-style register lists, whose
// members carry arrangements, stay out of the simple range shape: they remain
// plain bracketed groups the arm64 encoder reads from Raw. A comma inside
// brackets is a top-level comma, so a multi-member list spans several
// operands, exactly the shape the arm64 encoder's list scan stitches back.
func TestBracketRegisterRangeNotList(t *testing.T) {
file, errs := Parse("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tVLD1 (R2), [V21.B16]\n\tVLD1 (R1), [V2.B16, V3.B16]\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse errors: %v", errs)
}
fn := file.Decls[0].(*ast.Text)
for i, want := range []string{"[ V21.B16 ]", "V3.B16 ]"} {
in := fn.Body[i].(*ast.Instr)
op := in.Operands[len(in.Operands)-1]
if op.Addr.Range != nil {
t.Errorf("%s: range = %v, want nil", in.Mnemonic.Text, op.Addr.Range)
}
if op.Raw != want {
t.Errorf("operand %d raw = %q, want %q", i, op.Raw, want)
}
}
}
// TestVSIBIndexOnly pins the gather/scatter memory operand with a scaled
// vector index and no base register: 8(X4*1) must carry index and scale and
// leave the base empty, not strand the scale in the shift suffix.
func TestVSIBIndexOnly(t *testing.T) {
file, errs := Parse("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tVPGATHERDQ Y0, 8(X4*1), Y6\n\tVPGATHERDQ Y0, (X4*2), Y6\n\tVPGATHERDQ Y0, -8(X4*1), Y6\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse errors: %v", errs)
}
fn := file.Decls[0].(*ast.Text)
want := []ast.Address{
{Index: "X4", Scale: 1, Offset: 8, HasOff: true},
{Index: "X4", Scale: 2},
{Index: "X4", Scale: 1, Offset: -8, HasOff: true},
}
for i, w := range want {
in := fn.Body[i].(*ast.Instr)
a := in.Operands[1].Addr
if a.Base != "" || a.Index != w.Index || a.Scale != w.Scale || a.Offset != w.Offset || a.HasOff != w.HasOff || a.Shift != "" {
t.Errorf("operand %d = %+v, want %+v", i, a, w)
}
}
}
// TestVSIBTwoGroupKeepsBase pins that the ordinary (base)(index*scale)
// grammar is untouched by the index-only recognition.
func TestVSIBTwoGroupKeepsBase(t *testing.T) {
file, errs := Parse("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tVP4DPWSSD 7(SI)(DI*1), [Z2-Z5], K4, Z17\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse errors: %v", errs)
}
fn := file.Decls[0].(*ast.Text)
in := fn.Body[0].(*ast.Instr)
a := in.Operands[0].Addr
if a.Base != "SI" || a.Index != "DI" || a.Scale != 1 || a.Offset != 7 || !a.HasOff {
t.Errorf("address = %+v, want base SI index DI scale 1 offset 7", a)
}
if in.Operands[1].Addr.Range == nil || in.Operands[1].Addr.Range.Lo != "Z2" || in.Operands[1].Addr.Range.Hi != "Z5" {
t.Errorf("second operand = %+v, want range Z2-Z5", in.Operands[1].Addr)
}
}
// TestBareTrailingImmediate pins the toolchain's bare constant spelling in
// the final operand slot: CMPSD X1, X0, 1 reads as $1 (math/floor_amd64.s).
// Earlier slots keep the strict grammar, so a bare number there stays an
// address rather than becoming an immediate.
func TestBareTrailingImmediate(t *testing.T) {
file, errs := Parse("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tCMPSD X1, X0, 1\n\tCMPSD X1, X0, -1\n\tADDQ AX, 1+2\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse errors: %v", errs)
}
fn := file.Decls[0].(*ast.Text)
for i, want := range []int64{1, -1, 3} {
in := fn.Body[i].(*ast.Instr)
last := in.Operands[len(in.Operands)-1]
if last.Kind != ast.OpImmediate || !last.Imm.HasVal || last.Imm.Val != want {
t.Errorf("operand %d = %+v, want immediate %d", i, last, want)
}
}
// A bare number outside the final slot is not an immediate.
file2, errs2 := Parse("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tADDQ 1, AX\n\tRET\n")
if len(errs2) > 0 {
t.Fatalf("parse errors: %v", errs2)
}
fn2 := file2.Decls[0].(*ast.Text)
first := fn2.Body[0].(*ast.Instr).Operands[0]
if first.Kind != ast.OpAddr {
t.Errorf("non-final bare number kind = %v, want OpAddr", first.Kind)
}
// A bare name in the final slot stays a symbol: labels are names, not
// constants, and jump targets depend on the distinction.
file3, errs3 := Parse("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tJMP loop\nloop: NOP\n\tRET\n")
if len(errs3) > 0 {
t.Fatalf("parse errors: %v", errs3)
}
fn3 := file3.Decls[0].(*ast.Text)
jmp := fn3.Body[0].(*ast.Instr)
if jmp.Operands[0].Kind != ast.OpAddr || jmp.Operands[0].Addr.Sym == nil || jmp.Operands[0].Addr.Sym.Name != "loop" {
t.Errorf("jump target = %+v, want label loop", jmp.Operands[0])
}
}
// TestSignedParenDisplacement pins a sign before a parenthesised
// displacement expression: -(24+8)(X6) negates the folded value and keeps
// the base group, the shape GOROOT's riscv64 and loong64 files use.
func TestSignedParenDisplacement(t *testing.T) {
file, errs := Parse("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0-0\n\tMOV X7, -(24+8)(X6)\n\tMOV X7, +(16)(X6)\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
text := file.Decls[0].(*ast.Text)
ins := text.Body[0].(*ast.Instr)
op := ins.Operands[1] // Plan 9 order: the destination address is last
if !op.Addr.HasOff || op.Addr.Offset != -32 {
t.Errorf("-(24+8): offset = %v hasOff=%v, want -32 true", op.Addr.Offset, op.Addr.HasOff)
}
if op.Addr.Base != "X6" {
t.Errorf("-(24+8): base = %q, want X6", op.Addr.Base)
}
ins = text.Body[1].(*ast.Instr)
op = ins.Operands[1]
if !op.Addr.HasOff || op.Addr.Offset != 16 || op.Addr.Base != "X6" {
t.Errorf("+(16): offset = %v hasOff=%v base=%q, want 16 true X6", op.Addr.Offset, op.Addr.HasOff, op.Addr.Base)
}
}
+99 -3
View File
@@ -34,6 +34,12 @@ type Options struct {
// Expand enables macro expansion, include splicing and the
// statement-separator reading of ';' that the expanded bodies rely on.
Expand bool
// Predefines names the macros defined before the file is read. The
// go command drives go tool asm with -D GOOS_<goos> -D GOARCH_<arch>,
// and GOROOT's own headers (go_tls.h, asm_riscv64.h) select their
// platform blocks with #ifdef on exactly those names, so an assembler
// without them cannot see the platform definitions at all.
Predefines map[string]string
}
// ParseWithOptions parses src like Parse, optionally preprocessing it first.
@@ -44,10 +50,13 @@ func ParseWithOptions(path, src string, opts Options) (*ast.File, []error) {
var errs []error
if opts.Expand {
pp := &preproc{opts: opts, macros: map[string]*macroDef{}}
for name, value := range opts.Predefines {
pp.macros[name] = &macroDef{name: name, body: lexer.Tokenize(value)}
}
lines = pp.fileLines(path, tokens, token.Position{})
errs = pp.errs
} else {
lines = splitLines(tokens)
lines = statementLines(tokens)
}
p := &state{path: path}
p.parse(lines)
@@ -349,7 +358,7 @@ func (pp *preproc) expandTokens(in []token.Token) []token.Token {
consecutive = 0
continue
}
def := pp.macros[t.Text]
def, suffix := pp.macroFor(t.Text)
if def == nil {
i++
consecutive = 0
@@ -363,7 +372,14 @@ func (pp *preproc) expandTokens(in []token.Token) []token.Token {
return nil
}
if def.args == nil {
s = append(s[:i], append(restamp(def.body, t.Pos), s[i+1:]...)...)
body := restamp(def.body, t.Pos)
if suffix != "" {
// The macro was reached only through a compound spelling
// (ACC0.B16 over "#define ACC0 V8"), so the selector has
// to travel with the expansion.
body = appendSelector(body, suffix, t.Pos)
}
s = append(s[:i], append(body, s[i+1:]...)...)
continue
}
// A parameterised macro invoked without its parentheses stands
@@ -394,6 +410,17 @@ func (pp *preproc) expandTokens(in []token.Token) []token.Token {
sub = append(sub, restamp(args[k], t.Pos)...)
continue
}
// A parameter used with an element or lane selector: the
// lexer folds A.S4 into one identifier, so the whole-token
// match above cannot see the parameter. The toolchain
// lexes the period separately and substitutes the name
// alone; splitting at the FIRST period and pasting the
// argument back in front of the selector is the equivalent
// for this lexer.
if k, sel := parameterSelector(bt.Text, def.args); k >= 0 {
sub = append(sub, restamp(pasteSelector(args[k], sel), t.Pos)...)
continue
}
}
sub = append(sub, bt)
}
@@ -402,6 +429,75 @@ func (pp *preproc) expandTokens(in []token.Token) []token.Token {
return s
}
// macroFor finds the macro a use names. The lexer folds NAME.selector into
// one identifier token, so a macro written behind a selector suffix
// (ACC0.B16 over "#define ACC0 V8") never matches a whole-token table
// lookup; the toolchain splits on the period and reads the two halves, so
// the prefix before the FIRST period is tried here as well and the caller
// re-attaches the suffix to whatever the macro expands to. Only a whole
// name counts: AB.S4 does not reach a macro named A, and a parameterised
// macro is not hidden behind a selector, because its invocation would need
// the parentheses to follow the bare name.
func (pp *preproc) macroFor(text string) (*macroDef, string) {
if def := pp.macros[text]; def != nil {
return def, ""
}
if j := strings.IndexByte(text, '.'); j > 0 {
if def := pp.macros[text[:j]]; def != nil && def.args == nil {
return def, text[j:]
}
}
return nil, ""
}
// appendSelector glues a selector suffix onto an object macro's expansion:
// the selector binds to the identifier the expansion ends with, the way the
// toolchain's operand parser reads V0 and .B16 back as one register
// spelling. An expansion that does not end in an identifier carries the
// selector as its own token, which the parser then reports where it cannot
// parse it.
func appendSelector(body []token.Token, suffix string, pos token.Position) []token.Token {
if n := len(body); n > 0 && body[n-1].Kind == token.Ident {
body[n-1].Text += suffix
return body
}
return append(body, token.Token{Kind: token.Ident, Text: suffix, Pos: pos, End: pos})
}
// parameterSelector reports the argument a compound body token names: the
// parameter whose whole name occupies the text before the token's FIRST
// period, with the selector that follows. k is negative when no parameter
// matches, which leaves tokens like AB.S4 untouched even though a parameter
// A is bound.
func parameterSelector(text string, args []string) (int, string) {
j := strings.IndexByte(text, '.')
if j <= 0 {
return -1, ""
}
if k := slices.Index(args, text[:j]); k >= 0 {
return k, text[j:]
}
return -1, ""
}
// pasteSelector joins an argument with the selector a compound body token
// carries, textually: the selector binds to the identifier the argument
// ends with, so A.S4 over the argument V0.B16 spells V0.B16.S4, exactly the
// operand the toolchain's split-then-substitute leaves behind. An argument
// with no trailing identifier carries the selector as a separate token,
// which the parser then reports where it cannot parse it.
func pasteSelector(val []token.Token, suffix string) []token.Token {
if len(val) == 0 {
return []token.Token{{Kind: token.Ident, Text: suffix}}
}
out := slices.Clone(val)
if n := len(out); out[n-1].Kind == token.Ident {
out[n-1].Text += suffix
return out
}
return append(out, token.Token{Kind: token.Ident, Text: suffix})
}
// collectArgs reads the actual argument tokens of an invocation; the opening
// parenthesis is at start. Commas separate arguments except inside nested
// parentheses. A nil result means the list was unterminated, which is a
+125 -2
View File
@@ -462,7 +462,9 @@ TEXT ·f(SB), NOSPLIT, $0
func TestParseUnchangedWithoutExpand(t *testing.T) {
// Without Expand the preprocessor must not exist: a macro invocation
// stays an unexpanded instruction line and ';' keeps the old parse.
// stays an unexpanded instruction line. The ';' statement separator is
// not part of the preprocessor: the plain parse path splits on it the
// same way the expansion path does, so both spellings agree.
f, errs := Parse("t_amd64.s", `
#define TWICE ADDQ AX, AX
TEXT ·f(SB), NOSPLIT, $0
@@ -480,7 +482,7 @@ TEXT ·f(SB), NOSPLIT, $0
mnemonics = append(mnemonics, in.Mnemonic.Text)
}
}
if strings.Join(mnemonics, " ") != "TWICE BYTE RET" {
if strings.Join(mnemonics, " ") != "TWICE BYTE BYTE RET" {
t.Errorf("non-expanding parse changed: %v", mnemonics)
}
}
@@ -534,6 +536,127 @@ func TestConstantExpressionFoldsWithoutExpand(t *testing.T) {
}
}
func TestParameterWithSelectorSubstitutes(t *testing.T) {
// The lexer folds A.S4 into one identifier token, so a parameter used
// with an element or lane selector never matched the whole-token
// substitution; the toolchain's lexer splits on the period and its
// substitution sees the name alone. Several parameters carry selectors
// in one body here, which is the chacha8_arm64.s QR shape in miniature.
_, got := expand(t, `
#define QR(A, B, C, D) VADD A.S4, B.S4, C.S4; VEOR D.B16, A.B16, D.B16
TEXT ·f(SB), NOSPLIT, $0
QR(V0, V1, V2, V3)
RET
`)
wantLines(t, got,
"VADD V0.S4, V1.S4, V2.S4",
"VEOR V3.B16, V0.B16, V3.B16",
"RET",
)
}
func TestSelectorWithCompoundArgumentPastesTextually(t *testing.T) {
// An argument that is itself one compound identifier pastes verbatim:
// A.S4 over V0.B16 spells V0.B16.S4, the operand the toolchain's
// split-then-substitute leaves behind.
_, got := expand(t, `
#define M(A) VADD A.S4, A.S4, A.S4
TEXT ·f(SB), NOSPLIT, $0
M(V0.B16)
RET
`)
wantLines(t, got, "VADD V0.B16.S4, V0.B16.S4, V0.B16.S4", "RET")
}
func TestSelectorAlongsideBareParameter(t *testing.T) {
// A body may use the parameter bare and suffixed, and the argument may
// itself end in a selector; neither disturbs the other.
_, got := expand(t, `
#define M(A) VADD A, A.S4, A
TEXT ·f(SB), NOSPLIT, $0
M(V0)
M(V1.B16)
RET
`)
wantLines(t, got,
"VADD V0, V0.S4, V0",
"VADD V1.B16, V1.B16.S4, V1.B16",
"RET",
)
}
func TestSelectorKeepsNonParameterPrefixes(t *testing.T) {
// The prefix before the period must be the whole parameter name:
// AB.S4 never reaches a parameter A.
_, got := expand(t, `
#define M(A) VADD AB.S4, A.S4, AB.S4
TEXT ·f(SB), NOSPLIT, $0
M(V0)
RET
`)
wantLines(t, got, "VADD AB.S4, V0.S4, AB.S4", "RET")
}
func TestSelectorExpandsMacroValuedArgument(t *testing.T) {
// gcm_arm64.s invokes mulRound(B1) where B1 is itself an object macro:
// the paste stays rescannable, so B1.D1 still expands to V1.D1 the way
// the toolchain's rescan of substituted tokens does.
_, got := expand(t, `
#define B1 V1
#define mulRound(X) VPMULL X.D1, T1.D1, T3.Q1
TEXT ·f(SB), NOSPLIT, $0
mulRound(B1)
RET
`)
wantLines(t, got, "VPMULL V1.D1, T1.D1, T3.Q1", "RET")
}
func TestObjectMacroBehindSelectorExpands(t *testing.T) {
// Ordinary code writes ACC0.B16 where ACC0 is an object macro; the
// toolchain expands the alias because its lexer reads the selector as
// its own token, and the lookup here must reach the macro through the
// compound spelling the same way.
_, got := expand(t, `
#define ACC0 V8
TEXT ·f(SB), NOSPLIT, $0
VEOR ACC0.B16, ACC0.B16, ACC0.B16
RET
`)
wantLines(t, got, "VEOR V8.B16, V8.B16, V8.B16", "RET")
}
func TestChacha8QRMacroExpands(t *testing.T) {
// The real QR round of chacha8_arm64.s end to end: every parameter
// carries a selector somewhere, and the round is sixteen instructions.
_, got := expand(t, `
#define QR(A, B, C, D) \
VADD A.S4, B.S4, A.S4; VEOR D.B16, A.B16, D.B16; VREV32 D.H8, D.H8; \
VADD C.S4, D.S4, C.S4; VEOR B.B16, C.B16, V30.B16; VSHL $12, V30.S4, B.S4; VSRI $20, V30.S4, B.S4; \
VADD A.S4, B.S4, A.S4; VEOR D.B16, A.B16, D.B16; VTBL V31.B16, [D.B16], D.B16; \
VADD C.S4, D.S4, C.S4; VEOR B.B16, C.B16, V30.B16; VSHL $7, V30.S4, B.S4; VSRI $25, V30.S4, B.S4
TEXT ·f(SB), NOSPLIT, $0
QR(V0, V1, V2, V3)
RET
`)
wantLines(t, got,
"VADD V0.S4, V1.S4, V0.S4",
"VEOR V3.B16, V0.B16, V3.B16",
"VREV32 V3.H8, V3.H8",
"VADD V2.S4, V3.S4, V2.S4",
"VEOR V1.B16, V2.B16, V30.B16",
"VSHL $12, V30.S4, V1.S4",
"VSRI $20, V30.S4, V1.S4",
"VADD V0.S4, V1.S4, V0.S4",
"VEOR V3.B16, V0.B16, V3.B16",
"VTBL V31.B16, [V3.B16], V3.B16",
"VADD V2.S4, V3.S4, V2.S4",
"VEOR V1.B16, V2.B16, V30.B16",
"VSHL $7, V30.S4, V1.S4",
"VSRI $25, V30.S4, V1.S4",
"RET",
)
}
func TestNotAnExpressionFallsBack(t *testing.T) {
// Symbol immediates and floats must keep their ordinary parse.
f, errs := ParseWithOptions("t_amd64.s", "TEXT ·f(SB), NOSPLIT, $0\n\tMOVQ $1.5, AX\n\tMOVQ $·sym(SB), AX\n\tRET\n", Options{})
+44
View File
@@ -0,0 +1,44 @@
// Differential kernel: the _dbar (acquire/release) atomic exchange
// variants against the Go toolchain's loong64enc1.s rows.
#include "textflag.h"
TEXT ·AMXORDBW(SB), NOSPLIT, $0
AMXORDBW R14, (R13), R12
RET
TEXT ·AMXORDBV(SB), NOSPLIT, $0
AMXORDBV R14, (R13), R12
RET
TEXT ·AMMAXDBW(SB), NOSPLIT, $0
AMMAXDBW R14, (R13), R12
RET
TEXT ·AMMAXDBV(SB), NOSPLIT, $0
AMMAXDBV R14, (R13), R12
RET
TEXT ·AMMINDBW(SB), NOSPLIT, $0
AMMINDBW R14, (R13), R12
RET
TEXT ·AMMINDBV(SB), NOSPLIT, $0
AMMINDBV R14, (R13), R12
RET
TEXT ·AMMAXDBWU(SB), NOSPLIT, $0
AMMAXDBWU R14, (R13), R12
RET
TEXT ·AMMAXDBVU(SB), NOSPLIT, $0
AMMAXDBVU R14, (R13), R12
RET
TEXT ·AMMINDBWU(SB), NOSPLIT, $0
AMMINDBWU R14, (R13), R12
RET
TEXT ·AMMINDBVU(SB), NOSPLIT, $0
AMMINDBVU R14, (R13), R12
RET
+191
View File
@@ -0,0 +1,191 @@
// The AVX-512 families behind the avx512enc gap: AES round ops, integer
// VNNI and bit algorithms, word shifts and permutes with an immediate or a
// register count, lane broadcasts and extracts, gather and scatter prefetch
// hints, opmask broadcasts, the high/low half moves and the non-temporal
// stores. Every result is folded back so no instruction is dead.
#include "textflag.h"
// func avx512int(p *byte, n int) uint64
TEXT ·avx512int(SB), NOSPLIT, $0-24
MOVQ p+0(FP), SI
MOVQ n+16(FP), CX
// AES rounds through the EVEX spellings, masks included.
VAESENC Z20, Z21, Z22
VAESENCLAST Z23, Z24, Z25
VAESDEC (SI), Z26, Z27
VAESDECLAST Z28, Z29, Z30
// Integer VNNI and the bit algorithm group.
VPDPBUSD Z1, Z2, K2, Z3
VPDPBUSDS Z4, Z5, K2, Z6
VPDPWSSD Z7, Z8, Z9
VPDPWSSDS Z10, Z11, K2, Z12
VPOPCNTW Z12, K3, Z13
VPOPCNTB Z14, Z15
VGF2P8MULB Z16, Z17, K4, Z18
VGF2P8AFFINEQB $7, Z18, Z19, K5, Z20
// Byte/word arithmetic with saturation and masks.
VPADDSB Z1, Z2, K1, Z3
VPADDUSW Z3, Z4, K1, Z5
VPSUBSW Z5, Z6, K1, Z7
VPSUBUSB Z7, Z8, K1, Z9
VPSADBW Z9, Z10, Z11
VPMULHRSW Z11, Z12, Z13
VPMULHW Z13, Z14, Z15
VPUNPCKLBW Z15, Z16, K2, Z17
VPUNPCKHBW Z17, Z18, K2, Z19
VPUNPCKLWD Z19, Z20, K2, Z21
VPUNPCKHWD Z21, Z22, K2, Z23
VPCMPEQB Z23, Z24, K2, K3
VPCMPGTW Z25, Z26, K2, K3
VPCMPEQQ Z27, Z28, K2
VPMULTISHIFTQB Z29, Z30, K3, Z31
VDBPSADBW $3, Z1, Z2, K3, Z3
MOVQ CX, ret+16(FP)
RET
// func avx512perm(p *byte) uint64
TEXT ·avx512perm(SB), NOSPLIT, $0-16
MOVQ p+0(FP), SI
// Permutations: immediate and register counts, ternary logic.
VALIGNQ $3, Z1, Z2, K1, Z3
VPERMT2B Z3, Z4, K1, Z5
VPERMT2W Z5, Z6, K1, Z7
VPERMT2PS Z7, Z8, K1, Z9
VPERMI2W Z9, Z10, K1, Z11
VPERMI2PS Z11, Z12, K1, Z13
VPERMI2PD Z13, Z14, K1, Z15
VPERMB Z15, Z16, K1, Z17
VPERMW Z17, Z18, K1, Z19
VPERMPS Z19, Z20, Z21
VPERMD Z20, Z21, Z22
VPERMQ $1, Z1, K2, Z2
VPERMQ Z3, Z4, K2, Z5
VPERMPD $1, Z5, K2, Z6
VPERMPD Z7, Z8, K2, Z9
VPERMILPS $5, Z9, K2, Z10
VPERMILPS Z11, Z12, K2, Z13
VPERMILPD $1, Z13, K2, Z14
VPERMILPD Z15, Z16, K2, Z17
VPTERNLOGD $6, Z17, Z18, K2, Z19
VPTERNLOGQ $9, Z19, Z20, K2, Z21
// Lane shuffle and blend families.
VSHUFPD $1, Z1, Z2, K1, Z3
VSHUFPS $2, Z4, Z5, K1, Z6
VBLENDMPD Z7, Z8, K1, Z9
VBLENDMPS Z9, Z10, K1, Z11
VPBLENDMB Z11, Z12, K1, Z13
VPBLENDMW Z13, Z14, K1, Z15
VPBLENDMD Z15, Z16, K1, Z17
VPBLENDMQ Z17, Z18, K1, Z19
// Conflicts and leading zero counts.
VPCONFLICTD Z1, K1, Z2
VPCONFLICTQ Z3, K1, Z4
VPLZCNTD Z5, K1, Z6
VPLZCNTQ Z7, K1, Z8
// Compress and expand, byte and word widths.
VPCOMPRESSB Z1, K1, (SI)
VPCOMPRESSW Z2, K1, (SI)
VPEXPANDB (SI), K1, Z3
VPEXPANDW (SI), K1, Z4
MOVQ SI, ret+8(FP)
RET
// func avx512shift(p *byte) uint64
TEXT ·avx512shift(SB), NOSPLIT, $0-16
MOVQ p+0(FP), SI
// Variable shifts and shuffles with masks.
VPSLLVW Z1, Z2, K1, Z3
VPSRLVW Z3, Z4, K1, Z5
VPSRAVW Z5, Z6, K1, Z7
VPSHLDVW Z7, Z8, K1, Z9
VPSHRDVW Z9, Z10, K1, Z11
VPSHLDVD Z11, Z12, K1, Z13
VPSHLDVQ Z13, Z14, K1, Z15
VPSHRDVD Z15, Z16, K1, Z17
VPSHRDVQ Z17, Z18, K1, Z19
// Immediate shifts, the word/byte-quad widths and masks.
VPSLLW $3, Z1, K2, Z2
VPSRLW $5, Z3, K2, Z4
VPSRAW $7, Z5, K2, Z6
VPSLLDQ $9, Z7, Z8
VPSRLDQ $11, Z9, Z10
// Register-count shifts and their memory-count forms.
VPSLLD X1, Z2, K1, Z3
VPSRLD 16(SI), Z4, K1, Z5
VPSLLQ X6, Z7, K1, Z8
VPSRLQ X9, Z10, K1, Z11
VPSLLW X12, Z13, K1, Z14
VPSRAW X15, Z16, K1, Z17
VPSRAQ $13, Z12, K1, Z13
VPSRAD X14, Z15, K1, Z16
// Lane shuffles in and out.
VPSHLDW $2, Z1, Z2, K1, Z3
VPSHLDQ $4, Z3, Z4, K1, Z5
VPSHRDW $6, Z5, Z6, K1, Z7
VPSHRDQ $8, Z7, Z8, K1, Z9
VPSHUFBITQMB Z9, Z10, K3
VPTESTMB Z11, Z12, K4
VPTESTNMQ Z13, Z14, K5
MOVQ SI, ret+8(FP)
RET
// func avx512float(x float64) float64
TEXT ·avx512float(SB), NOSPLIT, $0-16
// Square roots, compares and the EXP2/RCP28 helpers.
MOVQ x+0(FP), AX
VSQRTPD Z1, K1, Z2
VSQRTPS Z3, K1, Z4
VSQRTSD X1, X2, K1, X3
VSQRTSS X3, X4, X5
VCOMISD X5, X6
VUCOMISS X7, X8
VEXP2PD Z5, K1, Z6
VRCP28PD Z7, K1, Z8
VRCP28SD X9, X8, K1, X10
VRSQRT28PS Z11, K1, Z12
VRSQRT28SS X11, X10, K1, X12
VCVTSD2SS X1, X2, X3
VCVTSS2SD X3, X2, K1, X4
VFMADD132PD Z1, Z2, K1, Z3
VFMADD231SD X1, X2, K1, X3
VFMSUBADD213PS Z3, Z4, K1, Z5
VFNMSUB231PD Z5, Z6, K1, Z7
// Broadcasts and masked moves.
VBROADCASTF32X2 X1, K1, Z2
VBROADCASTI64X2 (SI), K1, Z3
VMOVUPS Z1, K2, Z3
VMOVSD X14, X5, K3, X22
VMOVSS X18, X3, K2, X25
VMOVHPS (SI), X18, X19
VMOVHPS X20, 8(SI)
VMOVLHPS X16, X5, X17
VMOVNTDQ Z7, (SI)
VMOVNTDQA 64(SI), Z8
VMOVNTPD Z9, (SI)
MOVQ SI, ret+8(FP)
RET
// func avx512mask(p *byte) uint64
TEXT ·avx512mask(SB), NOSPLIT, $0-16
MOVQ p+0(FP), SI
// Omask broadcasts and the K register logic.
VPBROADCASTMB2Q K1, Z2
VPBROADCASTMW2D K3, Z4
KUNPCKWD K6, K4, K1
KADDB K2, K3, K5
KORW K1, K2, K7
// Gather and scatter prefetch hints.
VGATHERPF0DPD K5, (SI)(Y29*8)
VSCATTERPF1DPS K2, (SI)(Z28*4)
// Masked gathers ride the EVEX spelling; the data length wins L'L.
VGATHERDPD (SI)(X10*4), K7, Y22
VPSCATTERDQ Y6, K2, (SI)(X4*1)
// Lane extracts to general registers.
VPEXTRB $3, X1, AX
VPEXTRD $1, X2, DI
VPINSRQ $1, SI, X3, X4
VEXTRACTI32X4 $1, Z1, X5
VINSERTI64X2 $1, X6, Z7, K2, Z8
MOVQ SI, ret+8(FP)
RET
+32
View File
@@ -0,0 +1,32 @@
// The runtime bookkeeping statements: FUNCDATA and PCDATA contribute no
// text bytes on any architecture, and amd64 now matches. They sit between
// real instructions here, with plain, static and offset symbol references
// on the FUNCDATA lines, so the byte counts prove the zero contribution.
#include "textflag.h"
// func bookkeep(x int64) int64
TEXT ·bookkeep(SB), NOSPLIT, $0-16
PCDATA $0, $-1
MOVQ x+0(FP), AX
PCDATA $1, $-2
FUNCDATA $0, args_stackmap(SB)
ADDQ $1, AX
FUNCDATA $5, arginfo0(SB)
PCDATA $1, $3
MOVQ AX, ret+8(FP)
FUNCDATA $1, externalfuncdata(SB)
PCDATA $0, $0
RET
// func bookkeepstatic() int64
TEXT ·bookkeepstatic(SB), NOSPLIT, $0-8
// A static symbol and a defined data symbol as the funcdata target.
// (A symbol+offset target the toolchain itself refuses.)
FUNCDATA $2, fdtable<>(SB)
FUNCDATA $3, undefsym(SB)
MOVQ $7, AX
MOVQ AX, ret+0(FP)
RET
GLOBL fdtable<>(SB), NOPTR, $16
+31
View File
@@ -0,0 +1,31 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// Differential kernel for the arm64 bookkeeping statements: the funcdata.h
// pseudo-directives (GO_ARGS, NO_LOCAL_POINTERS, FUNCDATA, PCDATA) contribute
// no instruction bytes, and every function is byte-compared against
// go tool asm.
#include "textflag.h"
#include "funcdata.h"
// func bookkeep()
TEXT ·bookkeep(SB), NOSPLIT, $8-0
GO_ARGS
FUNCDATA $3, inline_tree(SB)
PCDATA $1, $2
MOVD R1, 0(RSP)
RET
// func bookkeepNoLocals()
TEXT ·bookkeepNoLocals(SB), NOSPLIT, $16-0
NO_LOCAL_POINTERS
PCDATA $0, $0
PCDATA $1, $1
MOVD R2, 8(RSP)
RET
// func bookkeepPlain()
TEXT ·bookkeepPlain(SB), NOSPLIT, $0-0
MOVD R3, R4
RET
+40
View File
@@ -0,0 +1,40 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// Differential kernel for the riscv64 bookkeeping statements and the
// slot-relative branches: FUNCDATA and PCDATA (the expanded forms of the
// funcdata.h macros, contributing no bytes), UNDEF (the toolchain's ebreak),
// and the JMP N(PC) slot jumps including the self-loop and the backward form.
#include "textflag.h"
TEXT ·bookkeep(SB), NOSPLIT, $8-8
FUNCDATA $1, marks<>(SB)
PCDATA $1, $-1
MOV ZERO, ret+0(FP)
PCDATA $1, $1
UNDEF
MOV $1, X10
RET
TEXT ·slots(SB), NOSPLIT, $0-0
MOV $1, X10
JMP 2(PC)
MOV $64, X11
MOV $128, X12
MOV $2, X11
MOV $3, X12
BEQ X10, X11, skip
JMP -2(PC)
skip:
JMP 0(PC)
TEXT ·marksreader(SB), NOSPLIT, $0-8
MOV $marks<>(SB), X10
MOV (X10), X11
MOV X11, ret+0(FP)
RET
GLOBL marks<>(SB), RODATA, $8
DATA marks<>+0(SB)/8, $1234605616436508552
+21
View File
@@ -0,0 +1,21 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// Differential kernel for the loong64 two-operand BEQ/BNE spellings the
// msan trampolines use: BEQ Rj, target compares against R0 (the beqz form).
#include "textflag.h"
TEXT ·branch2(SB), NOSPLIT, $0-8
MOVV arg+0(FP), R4
BEQ R4, zero
ADDV $1, R4, R4
zero:
MOVV $16, R5
BNE R4, done
ADDV $2, R4, R4
done:
MOVV R4, ret+0(FP)
RET
+27
View File
@@ -0,0 +1,27 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// Differential kernel for the loong64 DATA value forms the runtime's exp and
// asm files use: floating-point initialisers stored as IEEE-754 bits and
// string initialisers zero-padded within their declared width.
#include "textflag.h"
TEXT ·floatbits(SB), NOSPLIT, $0-8
MOVV $floats<>(SB), R12
MOVD 8(R12), F0
MOVD F0, ret+0(FP)
RET
TEXT ·stringhead(SB), NOSPLIT, $0-8
MOVV $msg<>(SB), R12
MOVV (R12), R13
MOVV R13, ret+0(FP)
RET
GLOBL floats<>(SB), RODATA, $16
DATA floats<>+0(SB)/8, $0.0
DATA floats<>+8(SB)/8, $0.5
GLOBL msg<>(SB), RODATA, $20
DATA msg<>+0(SB)/20, $"call frame too large"
+25
View File
@@ -0,0 +1,25 @@
#include "textflag.h"
// The kernel exercises the symbol-valued DATA spelling the runtime's rt0
// files use: a data word holding the address of a symbol, resolved by the
// linker through a relocation at the field.
// func lookup() ptr
TEXT ·lookup(SB), NOSPLIT, $0-8
MOVQ handlers+8(SB), AX
MOVQ AX, ret+0(FP)
RET
// func handler() int64
TEXT ·handler(SB), NOSPLIT, $0-8
MOVQ $42, AX
MOVQ AX, ret+0(FP)
RET
GLOBL handlers(SB), NOPTR, $24
DATA handlers+0(SB)/8, $·handler(SB)
DATA handlers+8(SB)/8, $table(SB)
DATA handlers+16(SB)/8, $·handler+5(SB)
GLOBL table(SB), RODATA, $8
DATA table+0(SB)/8, $0x123456789abcdef0
+25
View File
@@ -0,0 +1,25 @@
#include "textflag.h"
// The kernel exercises the symbol-valued DATA spelling the runtime's rt0
// files use: a data word holding the address of a symbol, resolved by the
// linker through a relocation at the field.
// func lookup() ptr
TEXT ·lookup(SB), NOSPLIT, $0-8
MOVD handlers+8(SB), R4
MOVD R4, ret+0(FP)
RET
// func handler() int64
TEXT ·handler(SB), NOSPLIT, $0-8
MOVZ $42, R4
MOVD R4, ret+0(FP)
RET
GLOBL handlers(SB), NOPTR, $24
DATA handlers+0(SB)/8, $·handler(SB)
DATA handlers+8(SB)/8, $table(SB)
DATA handlers+16(SB)/8, $extentry(SB)
GLOBL table(SB), RODATA, $8
DATA table+0(SB)/8, $0x123456789abcdef0
+19
View File
@@ -0,0 +1,19 @@
#include "textflag.h"
// The kernel exercises the U+2215 DIVISION SLASH inside a symbol's package
// path: internal∕runtime∕atomic·Xchg, the spelling sync/atomic/asm.s uses.
// The middle dot (U+00B7) still separates the package path from the name.
// func swap(a, b int64) int64
TEXT ·swap(SB), NOSPLIT, $0-24
MOVQ a+0(FP), DI
MOVQ b+8(FP), SI
CALL internal∕runtime∕atomic·Xchg(SB)
MOVQ AX, ret+16(FP)
RET
// func note() int64
TEXT ·note(SB), NOSPLIT, $0-8
CALL runtime∕debug·SetGCPercent(SB)
MOVQ AX, ret+0(FP)
RET
+19
View File
@@ -0,0 +1,19 @@
#include "textflag.h"
// The kernel exercises the U+2215 DIVISION SLASH inside a symbol's package
// path: internal∕runtime∕atomic·Xchg, the spelling sync/atomic/asm.s uses.
// The middle dot (U+00B7) still separates the package path from the name.
// func swap(a, b int64) int64
TEXT ·swap(SB), NOSPLIT, $0-24
MOVD a+0(FP), R4
MOVD b+8(FP), R5
CALL internal∕runtime∕atomic·Xchg(SB)
MOVD R4, ret+16(FP)
RET
// func note() int64
TEXT ·note(SB), NOSPLIT, $0-8
CALL runtime∕debug·SetGCPercent(SB)
MOVD R0, ret+0(FP)
RET
+55
View File
@@ -0,0 +1,55 @@
// Floating-point immediates on the SSE scalar paths: the constant is
// rewritten into a read from a read-only pool symbol ($f64.<hex> or
// $f32.<hex>, the IEEE-754 bits in the name), RIP-relative with the
// displacement left to the relocation. A positive zero on the moves
// collapses to XORPS dst, dst; a negative zero keeps its sign bit and
// takes the pool. The parenthesised $(-1.0) spelling is the one
// math/floor_amd64.s uses. Every result is folded back so no
// instruction is dead.
#include "textflag.h"
// func floatimm(x float64) float64
TEXT ·floatimm(SB), NOSPLIT, $0-16
MOVQ x+0(FP), AX
MOVQ AX, X0
// The floor kernel's sign fold: the parenthesised negative spelling.
MOVSD $ (-1.0), X2
ANDPD X2, X0
// Positive and fractional constants on the scalar moves.
MOVSD $1.5, X3
MOVSD $0.5, X4
MOVSS $2.5, X5
MOVSS $-0.5, X6
// A positive zero collapses to XORPS; a negative zero does not.
MOVSD $0.0, X7
MOVSS $0.0, X8
MOVSD $-0.0, X9
// The scalar arithmetic reads the pool through r/m (hypot's shape).
ADDSD $1.0, X3
SUBSD $0.5, X4
MULSD $-2.5, X4
DIVSD $2.0, X3
ADDSS $0.25, X5
// Fold everything into one double.
ADDSD X5, X3
ADDSD X6, X3
ADDSD X7, X3
ADDSD X8, X3
ADDSD X9, X3
ADDSD X4, X3
ADDSD X0, X3
MOVSD X3, ret+8(FP)
RET
// func floatimmfloat32() float32
TEXT ·floatimmfloat32(SB), NOSPLIT, $0-4
// The single-width pool constants ride the F3 prefix.
MOVSS $1.0, X0
MOVSS $-1.0, X1
MOVSS $0.0, X2
ADDSS $0.5, X0
ADDSS X1, X0
ADDSS X2, X0
MOVSS X0, ret+0(FP)
RET
+52
View File
@@ -0,0 +1,52 @@
// Kernel: the operand forms the GOROOT campaign surfaced — numeric
// PC-relative jumps, symbol-immediate materialisation (the toolchain rewrites
// MOVQ $sym(SB) into a RIP-relative LEA) and the negated constant-expression
// ADJSP the cgo ABI macros write. Bytes are pinned against go tool asm by
// TestDifferentialKernels.
#include "textflag.h"
DATA sd<>(SB)/4, $7
GLOBL sd<>(SB), RODATA, $4
// func Jumps(flag int64) int64
TEXT ·Jumps(SB), NOSPLIT, $0-16
MOVQ flag+0(FP), AX
TESTQ AX, AX
JEQ 2(PC)
MOVQ $1, AX
JMP 3(PC)
MOVQ $2, AX
MOVQ AX, ret+0(FP)
RET
// func SymImm() int64
TEXT ·SymImm(SB), NOSPLIT, $0-16
MOVQ $sd<>(SB), AX
MOVQ $·SymImm(SB), CX
MOVQ AX, ret+0(FP)
RET
// func Frame()
TEXT ·Frame(SB), NOSPLIT, $0
PUSHFQ
CLD
ADJSP $(64 - 8)
ADJSP $-(64 - 8)
POPFQ
RET
// func Tls() int64
TEXT ·Tls(SB), NOSPLIT, $0-8
MOVQ TLS, BX
MOVQ 0(BX)(TLS*1), AX
MOVQ AX, ret+0(FP)
RET
// func Aligned() int64
TEXT ·Aligned(SB), NOSPLIT, $0-8
MOVQ $1, AX
PCALIGN $16
MOVQ $2, AX
PCALIGN $32
MOVQ AX, ret+0(FP)
RET
+31
View File
@@ -0,0 +1,31 @@
// Differential kernel: the bookkeeping statements the assembler accepts and
// encodes to nothing (gasm v. go tool asm, byte for byte).
#include "textflag.h"
TEXT ·end(SB), NOSPLIT, $0
END
RET
TEXT ·funcdata(SB), NOSPLIT, $0
FUNCDATA $0, ref(SB)
RET
TEXT ·pcdata(SB), NOSPLIT, $0
PCDATA $0, $1
PCDATA $1, $-2
RET
TEXT ·getcallerpc(SB), NOSPLIT, $0
GETCALLERPC R4
RET
TEXT ·mixed(SB), NOSPLIT, $0
PCDATA $0, $1
ADDV R4, R5, R6
FUNCDATA $1, ref(SB)
GETCALLERPC R7
RET
ref:
RET
+44
View File
@@ -0,0 +1,44 @@
// The quad-register instructions: the 4FMAPS family (V4FMADDPS,
// V4FMADDSS, V4FNMADDPS, V4FNMADDSS) and the 4VNNIW pair (VP4DPWSSD,
// VP4DPWSSDS). The bracketed list's low register travels the inverted
// V'VVVV field, the memory source keeps r/m, the opmask rides aaa and the
// vector length follows the destination (512-bit for the ZMM forms,
// 128-bit for the scalar ones) while the disp8xN multiplier stays 16 for
// every member. Every result is folded back so no instruction is dead.
#include "textflag.h"
// func quadf4(src *[16]uint32, n int) float32
TEXT ·quadf4(SB), NOSPLIT, $0-20
MOVQ src+0(FP), SI
MOVQ n+8(FP), CX
// The packed 4-FMA form over four consecutive ZMM accumulators,
// masked with K2, K3 and unmasked alike; the displacements exercise
// the disp32 form and the disp8x16 compressed form.
V4FMADDPS 17(SI), [Z0-Z3], K2, Z0
V4FMADDPS 64(SI), [Z10-Z13], K2, Z1
V4FMADDPS (SI), [Z20-Z23], Z2
V4FNMADDPS 96(SI), [Z1-Z4], K3, Z5
// The scalar form reads XMM lists and takes the 128-bit length; the
// displacement compresses by 16.
V4FMADDSS 7(AX), [X0-X3], K5, X22
V4FMADDSS (DI), [X10-X13], K5, X23
V4FNMADDSS 16(SI), [X20-X23], K1, X24
// The 4-VNNI dot products, indexed source included.
VP4DPWSSD 15(DX)(BX*8), [Z2-Z5], K4, Z17
VP4DPWSSDS -7(DI)(R8*1), [Z4-Z7], K1, Z31
VP4DPWSSD (SI), [Z12-Z15], Z6
// Zeroing keeps the usual rule: a mask register must ride along.
V4FMADDPS.Z 128(SI), [Z24-Z27], K4, Z3
// Fold every accumulator into one scalar.
VPADDD Z0, Z1, Z9
VPADDD Z2, Z5, Z10
VPADDD Z9, Z17, Z11
VPADDD Z10, Z31, Z12
VPADDD Z11, Z12, Z13
VPADDD Z13, Z14, Z15
VADDSS X22, X23, X0
VADDSS X24, X0, X1
VADDSS X1, X2, X3
VMOVSS X3, ret+16(FP)
RET
+21
View File
@@ -0,0 +1,21 @@
#include "textflag.h"
// The kernel exercises the ';' statement separator in a plain file, the way
// the runtime writes it ("ROLQ $3, DI; ROLQ $13, DI", "REP; MOVSQ"). Each
// statement assembles exactly as it would on a line of its own.
// func rol(x int64) int64
TEXT ·rol(SB), NOSPLIT, $0-16
ROLQ $3, DI; ROLQ $13, DI
MOVQ DI, ret+0(FP)
RET
// func move(dst, src unsafe.Pointer)
TEXT ·move(SB), NOSPLIT, $0-16
REP ; MOVSQ
RET
TEXT ·paired(SB), NOSPLIT, $0-8
XORQ AX, AX; XORQ CX, CX
MOVQ AX, ret+0(FP)
RET
+57
View File
@@ -0,0 +1,57 @@
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
// SPDX-License-Identifier: BSD-3-Clause
// Differential kernel for the arm64 whole-vector moves between a general
// register and an arranged vector (VMOV/VDUP Rs, Vd.<T>), the two-operand
// accumulate spellings VADD/VSUB Vm, Vn, and the toolchain-reserved
// R18_PLATFORM register name. Every function is byte-compared against
// go tool asm.
#include "textflag.h"
// func gpIntoVector()
TEXT ·gpIntoVector(SB), NOSPLIT, $0-0
VMOV R1, V2.B8
VMOV R3, V4.B16
VMOV R5, V6.H4
VMOV R7, V8.H8
VMOV R9, V10.S2
VMOV R11, V12.S4
VMOV R13, V14.D2
VDUP R15, V16.B8
VDUP R17, V18.B16
VDUP R19, V20.H8
VDUP R21, V22.S4
VDUP R23, V24.D2
RET
// func simdAccumulate()
TEXT ·simdAccumulate(SB), NOSPLIT, $0-0
VADD V7, V8
VSUB V7, V8
VADD V1, V2
VSUB V30, V31
VADD V0.B16, V1.B16, V2.B16
VSUB V0.S4, V1.S4, V2.S4
RET
// func truncMove()
TEXT ·truncMove(SB), NOSPLIT, $0-0
MOVB R3, R4
MOVH R5, R6
MOVW R9, R10
MOVBU R3, R4
MOVHU R3, R4
MOVWU R3, R4
MOVD R3, R4
RET
// func platformRegister()
TEXT ·platformRegister(SB), NOSPLIT, $0-0
MOVD R18_PLATFORM, R3
MOVW R18_PLATFORM, R4
MOVD R3, R18_PLATFORM
MOVD 0x68(R18_PLATFORM), R5
MOVD R5, 0x68(R18_PLATFORM)
MOVW 8(R18_PLATFORM), R6
RET
+209
View File
@@ -0,0 +1,209 @@
// Differential kernel: the arith_add vector slice against the Go
// toolchain's loong64enc1.s rows (gasm v. go tool asm, byte for byte).
#include "textflag.h"
TEXT ·VADDB(SB), NOSPLIT, $0
VADDB V1, V2, V3
RET
TEXT ·VADDH(SB), NOSPLIT, $0
VADDH V1, V2, V3
RET
TEXT ·VADDQ(SB), NOSPLIT, $0
VADDQ V1, V2, V3
RET
TEXT ·XVADDB(SB), NOSPLIT, $0
XVADDB X3, X2, X1
RET
TEXT ·XVADDH(SB), NOSPLIT, $0
XVADDH X3, X2, X1
RET
TEXT ·XVADDW(SB), NOSPLIT, $0
XVADDW X3, X2, X1
RET
TEXT ·XVADDQ(SB), NOSPLIT, $0
XVADDQ X3, X2, X1
RET
TEXT ·VADDBU(SB), NOSPLIT, $0
VADDBU $1, V2
VADDBU $1, V2, V1
RET
TEXT ·VADDHU(SB), NOSPLIT, $0
VADDHU $2, V2, V1
RET
TEXT ·VADDWU(SB), NOSPLIT, $0
VADDWU $3, V2, V1
RET
TEXT ·VADDVU(SB), NOSPLIT, $0
VADDVU $4, V2, V1
RET
TEXT ·XVADDBU(SB), NOSPLIT, $0
XVADDBU $9, X1, X2
RET
TEXT ·XVADDHU(SB), NOSPLIT, $0
XVADDHU $10, X1, X2
RET
TEXT ·XVADDWU(SB), NOSPLIT, $0
XVADDWU $11, X1, X2
RET
TEXT ·XVADDVU(SB), NOSPLIT, $0
XVADDVU $12, X1, X2
RET
TEXT ·VADDWEVHB(SB), NOSPLIT, $0
VADDWEVHB V1, V2, V3
RET
TEXT ·VADDWEVWH(SB), NOSPLIT, $0
VADDWEVWH V1, V2, V3
RET
TEXT ·VADDWEVVW(SB), NOSPLIT, $0
VADDWEVVW V1, V2, V3
RET
TEXT ·VADDWEVQV(SB), NOSPLIT, $0
VADDWEVQV V1, V2, V3
RET
TEXT ·VADDWODHB(SB), NOSPLIT, $0
VADDWODHB V1, V2, V3
RET
TEXT ·VADDWODWH(SB), NOSPLIT, $0
VADDWODWH V1, V2, V3
RET
TEXT ·VADDWODVW(SB), NOSPLIT, $0
VADDWODVW V1, V2, V3
RET
TEXT ·VADDWODQV(SB), NOSPLIT, $0
VADDWODQV V1, V2, V3
RET
TEXT ·XVADDWEVHB(SB), NOSPLIT, $0
XVADDWEVHB X1, X2, X3
RET
TEXT ·XVADDWEVWH(SB), NOSPLIT, $0
XVADDWEVWH X1, X2, X3
RET
TEXT ·XVADDWEVVW(SB), NOSPLIT, $0
XVADDWEVVW X1, X2, X3
RET
TEXT ·XVADDWEVQV(SB), NOSPLIT, $0
XVADDWEVQV X1, X2, X3
RET
TEXT ·XVADDWODHB(SB), NOSPLIT, $0
XVADDWODHB X1, X2, X3
RET
TEXT ·XVADDWODWH(SB), NOSPLIT, $0
XVADDWODWH X1, X2, X3
RET
TEXT ·XVADDWODVW(SB), NOSPLIT, $0
XVADDWODVW X1, X2, X3
RET
TEXT ·XVADDWODQV(SB), NOSPLIT, $0
XVADDWODQV X1, X2, X3
RET
TEXT ·VADDWEVHBU(SB), NOSPLIT, $0
VADDWEVHBU V1, V2, V3
RET
TEXT ·VADDWEVWHU(SB), NOSPLIT, $0
VADDWEVWHU V1, V2, V3
RET
TEXT ·VADDWEVVWU(SB), NOSPLIT, $0
VADDWEVVWU V1, V2, V3
RET
TEXT ·VADDWEVQVU(SB), NOSPLIT, $0
VADDWEVQVU V1, V2, V3
RET
TEXT ·VADDWODHBU(SB), NOSPLIT, $0
VADDWODHBU V1, V2, V3
RET
TEXT ·VADDWODWHU(SB), NOSPLIT, $0
VADDWODWHU V1, V2, V3
RET
TEXT ·VADDWODVWU(SB), NOSPLIT, $0
VADDWODVWU V1, V2, V3
RET
TEXT ·VADDWODQVU(SB), NOSPLIT, $0
VADDWODQVU V1, V2, V3
RET
TEXT ·XVADDWEVHBU(SB), NOSPLIT, $0
XVADDWEVHBU X1, X2, X3
RET
TEXT ·XVADDWEVWHU(SB), NOSPLIT, $0
XVADDWEVWHU X1, X2, X3
RET
TEXT ·XVADDWEVVWU(SB), NOSPLIT, $0
XVADDWEVVWU X1, X2, X3
RET
TEXT ·XVADDWEVQVU(SB), NOSPLIT, $0
XVADDWEVQVU X1, X2, X3
RET
TEXT ·XVADDWODHBU(SB), NOSPLIT, $0
XVADDWODHBU X1, X2, X3
RET
TEXT ·XVADDWODWHU(SB), NOSPLIT, $0
XVADDWODWHU X1, X2, X3
RET
TEXT ·XVADDWODVWU(SB), NOSPLIT, $0
XVADDWODVWU X1, X2, X3
RET
TEXT ·XVADDWODQVU(SB), NOSPLIT, $0
XVADDWODQVU X1, X2, X3
RET
TEXT ·VADDF(SB), NOSPLIT, $0
VADDF V1, V2, V3
RET
TEXT ·VADDD(SB), NOSPLIT, $0
VADDD V1, V2, V3
RET
TEXT ·XVADDF(SB), NOSPLIT, $0
XVADDF X1, X2, X3
RET
TEXT ·XVADDD(SB), NOSPLIT, $0
XVADDD X1, X2, X3
RET
+132
View File
@@ -0,0 +1,132 @@
// Differential kernel: the arith_sat vector slice against the Go
// toolchain's loong64enc1.s rows (gasm v. go tool asm, byte for byte).
#include "textflag.h"
TEXT ·VSADDB(SB), NOSPLIT, $0
VSADDB V1, V2, V3
RET
TEXT ·VSADDH(SB), NOSPLIT, $0
VSADDH V1, V2, V3
RET
TEXT ·VSADDW(SB), NOSPLIT, $0
VSADDW V1, V2, V3
RET
TEXT ·VSADDV(SB), NOSPLIT, $0
VSADDV V1, V2, V3
RET
TEXT ·VSSUBB(SB), NOSPLIT, $0
VSSUBB V1, V2, V3
RET
TEXT ·VSSUBH(SB), NOSPLIT, $0
VSSUBH V1, V2, V3
RET
TEXT ·VSSUBW(SB), NOSPLIT, $0
VSSUBW V1, V2, V3
RET
TEXT ·VSSUBV(SB), NOSPLIT, $0
VSSUBV V1, V2, V3
RET
TEXT ·XVSADDB(SB), NOSPLIT, $0
XVSADDB X3, X2, X1
RET
TEXT ·XVSADDH(SB), NOSPLIT, $0
XVSADDH X3, X2, X1
RET
TEXT ·XVSADDW(SB), NOSPLIT, $0
XVSADDW X3, X2, X1
RET
TEXT ·XVSADDV(SB), NOSPLIT, $0
XVSADDV X3, X2, X1
RET
TEXT ·XVSSUBB(SB), NOSPLIT, $0
XVSSUBB X3, X2, X1
RET
TEXT ·XVSSUBH(SB), NOSPLIT, $0
XVSSUBH X3, X2, X1
RET
TEXT ·XVSSUBW(SB), NOSPLIT, $0
XVSSUBW X3, X2, X1
RET
TEXT ·XVSSUBV(SB), NOSPLIT, $0
XVSSUBV X3, X2, X1
RET
TEXT ·VSADDBU(SB), NOSPLIT, $0
VSADDBU V1, V2, V3
RET
TEXT ·VSADDHU(SB), NOSPLIT, $0
VSADDHU V1, V2, V3
RET
TEXT ·VSADDWU(SB), NOSPLIT, $0
VSADDWU V1, V2, V3
RET
TEXT ·VSADDVU(SB), NOSPLIT, $0
VSADDVU V1, V2, V3
RET
TEXT ·VSSUBBU(SB), NOSPLIT, $0
VSSUBBU V1, V2, V3
RET
TEXT ·VSSUBHU(SB), NOSPLIT, $0
VSSUBHU V1, V2, V3
RET
TEXT ·VSSUBWU(SB), NOSPLIT, $0
VSSUBWU V1, V2, V3
RET
TEXT ·VSSUBVU(SB), NOSPLIT, $0
VSSUBVU V1, V2, V3
RET
TEXT ·XVSADDBU(SB), NOSPLIT, $0
XVSADDBU X1, X2, X3
RET
TEXT ·XVSADDHU(SB), NOSPLIT, $0
XVSADDHU X1, X2, X3
RET
TEXT ·XVSADDWU(SB), NOSPLIT, $0
XVSADDWU X1, X2, X3
RET
TEXT ·XVSADDVU(SB), NOSPLIT, $0
XVSADDVU X1, X2, X3
RET
TEXT ·XVSSUBBU(SB), NOSPLIT, $0
XVSSUBBU X1, X2, X3
RET
TEXT ·XVSSUBHU(SB), NOSPLIT, $0
XVSSUBHU X1, X2, X3
RET
TEXT ·XVSSUBWU(SB), NOSPLIT, $0
XVSSUBWU X1, X2, X3
RET
TEXT ·XVSSUBVU(SB), NOSPLIT, $0
XVSSUBVU X1, X2, X3
RET
+220
View File
@@ -0,0 +1,220 @@
// Differential kernel: the arith_sub vector slice against the Go
// toolchain's loong64enc1.s rows (gasm v. go tool asm, byte for byte).
#include "textflag.h"
TEXT ·VSUBB(SB), NOSPLIT, $0
VSUBB V1, V2, V3
RET
TEXT ·VSUBH(SB), NOSPLIT, $0
VSUBH V1, V2, V3
RET
TEXT ·VSUBW(SB), NOSPLIT, $0
VSUBW V1, V2, V3
RET
TEXT ·VSUBV(SB), NOSPLIT, $0
VSUBV V1, V2, V3
RET
TEXT ·VSUBQ(SB), NOSPLIT, $0
VSUBQ V1, V2, V3
RET
TEXT ·XVSUBB(SB), NOSPLIT, $0
XVSUBB X3, X2, X1
RET
TEXT ·XVSUBH(SB), NOSPLIT, $0
XVSUBH X3, X2, X1
RET
TEXT ·XVSUBW(SB), NOSPLIT, $0
XVSUBW X3, X2, X1
RET
TEXT ·XVSUBV(SB), NOSPLIT, $0
XVSUBV X3, X2, X1
RET
TEXT ·XVSUBQ(SB), NOSPLIT, $0
XVSUBQ X3, X2, X1
RET
TEXT ·VSUBBU(SB), NOSPLIT, $0
VSUBBU $5, V2, V1
RET
TEXT ·VSUBHU(SB), NOSPLIT, $0
VSUBHU $6, V2, V1
RET
TEXT ·VSUBWU(SB), NOSPLIT, $0
VSUBWU $7, V2, V1
RET
TEXT ·VSUBVU(SB), NOSPLIT, $0
VSUBVU $8, V2, V1
RET
TEXT ·XVSUBBU(SB), NOSPLIT, $0
XVSUBBU $13, X1, X2
RET
TEXT ·XVSUBHU(SB), NOSPLIT, $0
XVSUBHU $14, X1, X2
RET
TEXT ·XVSUBWU(SB), NOSPLIT, $0
XVSUBWU $15, X1, X2
RET
TEXT ·XVSUBVU(SB), NOSPLIT, $0
XVSUBVU $16, X1, X2
RET
TEXT ·VSUBWEVHB(SB), NOSPLIT, $0
VSUBWEVHB V1, V2, V3
RET
TEXT ·VSUBWEVWH(SB), NOSPLIT, $0
VSUBWEVWH V1, V2, V3
RET
TEXT ·VSUBWEVVW(SB), NOSPLIT, $0
VSUBWEVVW V1, V2, V3
RET
TEXT ·VSUBWEVQV(SB), NOSPLIT, $0
VSUBWEVQV V1, V2, V3
RET
TEXT ·VSUBWODHB(SB), NOSPLIT, $0
VSUBWODHB V1, V2, V3
RET
TEXT ·VSUBWODWH(SB), NOSPLIT, $0
VSUBWODWH V1, V2, V3
RET
TEXT ·VSUBWODVW(SB), NOSPLIT, $0
VSUBWODVW V1, V2, V3
RET
TEXT ·VSUBWODQV(SB), NOSPLIT, $0
VSUBWODQV V1, V2, V3
RET
TEXT ·XVSUBWEVHB(SB), NOSPLIT, $0
XVSUBWEVHB X1, X2, X3
RET
TEXT ·XVSUBWEVWH(SB), NOSPLIT, $0
XVSUBWEVWH X1, X2, X3
RET
TEXT ·XVSUBWEVVW(SB), NOSPLIT, $0
XVSUBWEVVW X1, X2, X3
RET
TEXT ·XVSUBWEVQV(SB), NOSPLIT, $0
XVSUBWEVQV X1, X2, X3
RET
TEXT ·XVSUBWODHB(SB), NOSPLIT, $0
XVSUBWODHB X1, X2, X3
RET
TEXT ·XVSUBWODWH(SB), NOSPLIT, $0
XVSUBWODWH X1, X2, X3
RET
TEXT ·XVSUBWODVW(SB), NOSPLIT, $0
XVSUBWODVW X1, X2, X3
RET
TEXT ·XVSUBWODQV(SB), NOSPLIT, $0
XVSUBWODQV X1, X2, X3
RET
TEXT ·VSUBWEVHBU(SB), NOSPLIT, $0
VSUBWEVHBU V1, V2, V3
RET
TEXT ·VSUBWEVWHU(SB), NOSPLIT, $0
VSUBWEVWHU V1, V2, V3
RET
TEXT ·VSUBWEVVWU(SB), NOSPLIT, $0
VSUBWEVVWU V1, V2, V3
RET
TEXT ·VSUBWEVQVU(SB), NOSPLIT, $0
VSUBWEVQVU V1, V2, V3
RET
TEXT ·VSUBWODHBU(SB), NOSPLIT, $0
VSUBWODHBU V1, V2, V3
RET
TEXT ·VSUBWODWHU(SB), NOSPLIT, $0
VSUBWODWHU V1, V2, V3
RET
TEXT ·VSUBWODVWU(SB), NOSPLIT, $0
VSUBWODVWU V1, V2, V3
RET
TEXT ·VSUBWODQVU(SB), NOSPLIT, $0
VSUBWODQVU V1, V2, V3
RET
TEXT ·XVSUBWEVHBU(SB), NOSPLIT, $0
XVSUBWEVHBU X1, X2, X3
RET
TEXT ·XVSUBWEVWHU(SB), NOSPLIT, $0
XVSUBWEVWHU X1, X2, X3
RET
TEXT ·XVSUBWEVVWU(SB), NOSPLIT, $0
XVSUBWEVVWU X1, X2, X3
RET
TEXT ·XVSUBWEVQVU(SB), NOSPLIT, $0
XVSUBWEVQVU X1, X2, X3
RET
TEXT ·XVSUBWODHBU(SB), NOSPLIT, $0
XVSUBWODHBU X1, X2, X3
RET
TEXT ·XVSUBWODWHU(SB), NOSPLIT, $0
XVSUBWODWHU X1, X2, X3
RET
TEXT ·XVSUBWODVWU(SB), NOSPLIT, $0
XVSUBWODVWU X1, X2, X3
RET
TEXT ·XVSUBWODQVU(SB), NOSPLIT, $0
XVSUBWODQVU X1, X2, X3
RET
TEXT ·VSUBF(SB), NOSPLIT, $0
VSUBF V1, V2, V3
RET
TEXT ·VSUBD(SB), NOSPLIT, $0
VSUBD V1, V2, V3
RET
TEXT ·XVSUBF(SB), NOSPLIT, $0
XVSUBF X1, X2, X3
RET
TEXT ·XVSUBD(SB), NOSPLIT, $0
XVSUBD X1, X2, X3
RET
+124
View File
@@ -0,0 +1,124 @@
// Differential kernel: the bitops vector slice against the Go
// toolchain's loong64enc1.s rows (gasm v. go tool asm, byte for byte).
#include "textflag.h"
TEXT ·VBITCLRB(SB), NOSPLIT, $0
VBITCLRB V1, V2, V3
VBITCLRB $7, V2, V3
RET
TEXT ·VBITCLRH(SB), NOSPLIT, $0
VBITCLRH V1, V2, V3
VBITCLRH $15, V2, V3
RET
TEXT ·VBITCLRW(SB), NOSPLIT, $0
VBITCLRW V1, V2, V3
VBITCLRW $31, V2, V3
RET
TEXT ·VBITCLRV(SB), NOSPLIT, $0
VBITCLRV V1, V2, V3
VBITCLRV $63, V2, V3
RET
TEXT ·VBITSETB(SB), NOSPLIT, $0
VBITSETB V1, V2, V3
VBITSETB $7, V2, V3
RET
TEXT ·VBITSETH(SB), NOSPLIT, $0
VBITSETH V1, V2, V3
VBITSETH $15, V2, V3
RET
TEXT ·VBITSETW(SB), NOSPLIT, $0
VBITSETW V1, V2, V3
VBITSETW $31, V2, V3
RET
TEXT ·VBITSETV(SB), NOSPLIT, $0
VBITSETV V1, V2, V3
VBITSETV $63, V2, V3
RET
TEXT ·VBITREVB(SB), NOSPLIT, $0
VBITREVB V1, V2, V3
VBITREVB $7, V2, V3
RET
TEXT ·VBITREVH(SB), NOSPLIT, $0
VBITREVH V1, V2, V3
VBITREVH $15, V2, V3
RET
TEXT ·VBITREVW(SB), NOSPLIT, $0
VBITREVW V1, V2, V3
VBITREVW $31, V2, V3
RET
TEXT ·VBITREVV(SB), NOSPLIT, $0
VBITREVV V1, V2, V3
VBITREVV $63, V2, V3
RET
TEXT ·XVBITCLRB(SB), NOSPLIT, $0
XVBITCLRB X3, X2, X1
XVBITCLRB $7, X2, X1
RET
TEXT ·XVBITCLRH(SB), NOSPLIT, $0
XVBITCLRH X3, X2, X1
XVBITCLRH $15, X2, X1
RET
TEXT ·XVBITCLRW(SB), NOSPLIT, $0
XVBITCLRW X3, X2, X1
XVBITCLRW $31, X2, X1
RET
TEXT ·XVBITCLRV(SB), NOSPLIT, $0
XVBITCLRV X3, X2, X1
XVBITCLRV $63, X2, X1
RET
TEXT ·XVBITSETB(SB), NOSPLIT, $0
XVBITSETB X3, X2, X1
XVBITSETB $7, X2, X1
RET
TEXT ·XVBITSETH(SB), NOSPLIT, $0
XVBITSETH X3, X2, X1
XVBITSETH $15, X2, X1
RET
TEXT ·XVBITSETW(SB), NOSPLIT, $0
XVBITSETW X3, X2, X1
XVBITSETW $31, X2, X1
RET
TEXT ·XVBITSETV(SB), NOSPLIT, $0
XVBITSETV X3, X2, X1
XVBITSETV $63, X2, X1
RET
TEXT ·XVBITREVB(SB), NOSPLIT, $0
XVBITREVB X3, X2, X1
XVBITREVB $7, X2, X1
RET
TEXT ·XVBITREVH(SB), NOSPLIT, $0
XVBITREVH X3, X2, X1
XVBITREVH $15, X2, X1
RET
TEXT ·XVBITREVW(SB), NOSPLIT, $0
XVBITREVW X3, X2, X1
XVBITREVW $31, X2, X1
RET
TEXT ·XVBITREVV(SB), NOSPLIT, $0
XVBITREVV X3, X2, X1
XVBITREVV $63, X2, X1
RET
+419
View File
@@ -0,0 +1,419 @@
// Differential kernel: immediate range boundaries for the vector
// immediate forms (min, mid and max of each accepted window).
#include "textflag.h"
TEXT ·VSEQBbounds(SB), NOSPLIT, $0
VSEQB $-16, V2, V3
VSEQB $15, V2, V3
VSEQB $0, V2, V3
RET
TEXT ·VSEQHbounds(SB), NOSPLIT, $0
VSEQH $-16, V2, V3
VSEQH $15, V2, V3
VSEQH $0, V2, V3
RET
TEXT ·VSEQWbounds(SB), NOSPLIT, $0
VSEQW $-16, V2, V3
VSEQW $15, V2, V3
VSEQW $0, V2, V3
RET
TEXT ·VSEQVbounds(SB), NOSPLIT, $0
VSEQV $-16, V2, V3
VSEQV $15, V2, V3
VSEQV $0, V2, V3
RET
TEXT ·VSLTBbounds(SB), NOSPLIT, $0
VSLTB $-16, V2, V3
VSLTB $15, V2, V3
VSLTB $0, V2, V3
RET
TEXT ·VSLTHbounds(SB), NOSPLIT, $0
VSLTH $-16, V2, V3
VSLTH $15, V2, V3
VSLTH $0, V2, V3
RET
TEXT ·VSLTWbounds(SB), NOSPLIT, $0
VSLTW $-16, V2, V3
VSLTW $15, V2, V3
VSLTW $0, V2, V3
RET
TEXT ·VSLTVbounds(SB), NOSPLIT, $0
VSLTV $-16, V2, V3
VSLTV $15, V2, V3
VSLTV $0, V2, V3
RET
TEXT ·VSLTBUbounds(SB), NOSPLIT, $0
VSLTBU $0, V2, V3
VSLTBU $31, V2, V3
VSLTBU $15, V2, V3
RET
TEXT ·VSLTHUbounds(SB), NOSPLIT, $0
VSLTHU $0, V2, V3
VSLTHU $31, V2, V3
VSLTHU $15, V2, V3
RET
TEXT ·VSLTWUbounds(SB), NOSPLIT, $0
VSLTWU $0, V2, V3
VSLTWU $31, V2, V3
VSLTWU $15, V2, V3
RET
TEXT ·VSLTVUbounds(SB), NOSPLIT, $0
VSLTVU $0, V2, V3
VSLTVU $31, V2, V3
VSLTVU $15, V2, V3
RET
TEXT ·VADDBUbounds(SB), NOSPLIT, $0
VADDBU $0, V2, V3
VADDBU $31, V2, V3
VADDBU $15, V2, V3
VADDBU $15, V2
RET
TEXT ·VADDHUbounds(SB), NOSPLIT, $0
VADDHU $0, V2, V3
VADDHU $31, V2, V3
VADDHU $15, V2, V3
VADDHU $15, V2
RET
TEXT ·VADDWUbounds(SB), NOSPLIT, $0
VADDWU $0, V2, V3
VADDWU $31, V2, V3
VADDWU $15, V2, V3
VADDWU $15, V2
RET
TEXT ·VADDVUbounds(SB), NOSPLIT, $0
VADDVU $0, V2, V3
VADDVU $31, V2, V3
VADDVU $15, V2, V3
VADDVU $15, V2
RET
TEXT ·VSUBBUbounds(SB), NOSPLIT, $0
VSUBBU $0, V2, V3
VSUBBU $31, V2, V3
VSUBBU $15, V2, V3
VSUBBU $15, V2
RET
TEXT ·VSUBHUbounds(SB), NOSPLIT, $0
VSUBHU $0, V2, V3
VSUBHU $31, V2, V3
VSUBHU $15, V2, V3
VSUBHU $15, V2
RET
TEXT ·VSUBWUbounds(SB), NOSPLIT, $0
VSUBWU $0, V2, V3
VSUBWU $31, V2, V3
VSUBWU $15, V2, V3
VSUBWU $15, V2
RET
TEXT ·VSUBVUbounds(SB), NOSPLIT, $0
VSUBVU $0, V2, V3
VSUBVU $31, V2, V3
VSUBVU $15, V2, V3
VSUBVU $15, V2
RET
TEXT ·VANDBbounds(SB), NOSPLIT, $0
VANDB $0, V2, V3
VANDB $255, V2, V3
VANDB $127, V2, V3
VANDB $127, V2
RET
TEXT ·VBITCLRBbounds(SB), NOSPLIT, $0
VBITCLRB $0, V2, V3
VBITCLRB $7, V2, V3
VBITCLRB $3, V2, V3
VBITCLRB $3, V2
RET
TEXT ·VBITCLRHbounds(SB), NOSPLIT, $0
VBITCLRH $0, V2, V3
VBITCLRH $15, V2, V3
VBITCLRH $7, V2, V3
VBITCLRH $7, V2
RET
TEXT ·VBITCLRVbounds(SB), NOSPLIT, $0
VBITCLRV $0, V2, V3
VBITCLRV $63, V2, V3
VBITCLRV $31, V2, V3
VBITCLRV $31, V2
RET
TEXT ·VBITCLRWbounds(SB), NOSPLIT, $0
VBITCLRW $0, V2, V3
VBITCLRW $31, V2, V3
VBITCLRW $15, V2, V3
VBITCLRW $15, V2
RET
TEXT ·VBITREVBbounds(SB), NOSPLIT, $0
VBITREVB $0, V2, V3
VBITREVB $7, V2, V3
VBITREVB $3, V2, V3
VBITREVB $3, V2
RET
TEXT ·VBITREVHbounds(SB), NOSPLIT, $0
VBITREVH $0, V2, V3
VBITREVH $15, V2, V3
VBITREVH $7, V2, V3
VBITREVH $7, V2
RET
TEXT ·VBITREVVbounds(SB), NOSPLIT, $0
VBITREVV $0, V2, V3
VBITREVV $63, V2, V3
VBITREVV $31, V2, V3
VBITREVV $31, V2
RET
TEXT ·VBITREVWbounds(SB), NOSPLIT, $0
VBITREVW $0, V2, V3
VBITREVW $31, V2, V3
VBITREVW $15, V2, V3
VBITREVW $15, V2
RET
TEXT ·VBITSETBbounds(SB), NOSPLIT, $0
VBITSETB $0, V2, V3
VBITSETB $7, V2, V3
VBITSETB $3, V2, V3
VBITSETB $3, V2
RET
TEXT ·VBITSETHbounds(SB), NOSPLIT, $0
VBITSETH $0, V2, V3
VBITSETH $15, V2, V3
VBITSETH $7, V2, V3
VBITSETH $7, V2
RET
TEXT ·VBITSETVbounds(SB), NOSPLIT, $0
VBITSETV $0, V2, V3
VBITSETV $63, V2, V3
VBITSETV $31, V2, V3
VBITSETV $31, V2
RET
TEXT ·VBITSETWbounds(SB), NOSPLIT, $0
VBITSETW $0, V2, V3
VBITSETW $31, V2, V3
VBITSETW $15, V2, V3
VBITSETW $15, V2
RET
TEXT ·VEXTRINSBbounds(SB), NOSPLIT, $0
VEXTRINSB $0, V2, V3
VEXTRINSB $255, V2, V3
VEXTRINSB $127, V2, V3
VEXTRINSB $127, V2
RET
TEXT ·VEXTRINSHbounds(SB), NOSPLIT, $0
VEXTRINSH $0, V2, V3
VEXTRINSH $255, V2, V3
VEXTRINSH $127, V2, V3
VEXTRINSH $127, V2
RET
TEXT ·VEXTRINSVbounds(SB), NOSPLIT, $0
VEXTRINSV $0, V2, V3
VEXTRINSV $255, V2, V3
VEXTRINSV $127, V2, V3
VEXTRINSV $127, V2
RET
TEXT ·VEXTRINSWbounds(SB), NOSPLIT, $0
VEXTRINSW $0, V2, V3
VEXTRINSW $255, V2, V3
VEXTRINSW $127, V2, V3
VEXTRINSW $127, V2
RET
TEXT ·VNORBbounds(SB), NOSPLIT, $0
VNORB $0, V2, V3
VNORB $255, V2, V3
VNORB $127, V2, V3
VNORB $127, V2
RET
TEXT ·VORBbounds(SB), NOSPLIT, $0
VORB $0, V2, V3
VORB $255, V2, V3
VORB $127, V2, V3
VORB $127, V2
RET
TEXT ·VPERMIWbounds(SB), NOSPLIT, $0
VPERMIW $0, V2, V3
VPERMIW $255, V2, V3
VPERMIW $127, V2, V3
VPERMIW $127, V2
RET
TEXT ·VROTRBbounds(SB), NOSPLIT, $0
VROTRB $0, V2, V3
VROTRB $7, V2, V3
VROTRB $3, V2, V3
VROTRB $3, V2
RET
TEXT ·VROTRHbounds(SB), NOSPLIT, $0
VROTRH $0, V2, V3
VROTRH $15, V2, V3
VROTRH $7, V2, V3
VROTRH $7, V2
RET
TEXT ·VROTRVbounds(SB), NOSPLIT, $0
VROTRV $0, V2, V3
VROTRV $63, V2, V3
VROTRV $31, V2, V3
VROTRV $31, V2
RET
TEXT ·VROTRWbounds(SB), NOSPLIT, $0
VROTRW $0, V2, V3
VROTRW $31, V2, V3
VROTRW $15, V2, V3
VROTRW $15, V2
RET
TEXT ·VSHUF4IBbounds(SB), NOSPLIT, $0
VSHUF4IB $0, V2, V3
VSHUF4IB $255, V2, V3
VSHUF4IB $127, V2, V3
VSHUF4IB $127, V2
RET
TEXT ·VSHUF4IHbounds(SB), NOSPLIT, $0
VSHUF4IH $0, V2, V3
VSHUF4IH $255, V2, V3
VSHUF4IH $127, V2, V3
VSHUF4IH $127, V2
RET
TEXT ·VSHUF4IVbounds(SB), NOSPLIT, $0
VSHUF4IV $0, V2, V3
VSHUF4IV $15, V2, V3
VSHUF4IV $7, V2, V3
VSHUF4IV $7, V2
RET
TEXT ·VSHUF4IWbounds(SB), NOSPLIT, $0
VSHUF4IW $0, V2, V3
VSHUF4IW $255, V2, V3
VSHUF4IW $127, V2, V3
VSHUF4IW $127, V2
RET
TEXT ·VSLLBbounds(SB), NOSPLIT, $0
VSLLB $0, V2, V3
VSLLB $7, V2, V3
VSLLB $3, V2, V3
VSLLB $3, V2
RET
TEXT ·VSLLHbounds(SB), NOSPLIT, $0
VSLLH $0, V2, V3
VSLLH $15, V2, V3
VSLLH $7, V2, V3
VSLLH $7, V2
RET
TEXT ·VSLLVbounds(SB), NOSPLIT, $0
VSLLV $0, V2, V3
VSLLV $63, V2, V3
VSLLV $31, V2, V3
VSLLV $31, V2
RET
TEXT ·VSLLWbounds(SB), NOSPLIT, $0
VSLLW $0, V2, V3
VSLLW $31, V2, V3
VSLLW $15, V2, V3
VSLLW $15, V2
RET
TEXT ·VSRABbounds(SB), NOSPLIT, $0
VSRAB $0, V2, V3
VSRAB $7, V2, V3
VSRAB $3, V2, V3
VSRAB $3, V2
RET
TEXT ·VSRAHbounds(SB), NOSPLIT, $0
VSRAH $0, V2, V3
VSRAH $15, V2, V3
VSRAH $7, V2, V3
VSRAH $7, V2
RET
TEXT ·VSRAVbounds(SB), NOSPLIT, $0
VSRAV $0, V2, V3
VSRAV $63, V2, V3
VSRAV $31, V2, V3
VSRAV $31, V2
RET
TEXT ·VSRAWbounds(SB), NOSPLIT, $0
VSRAW $0, V2, V3
VSRAW $31, V2, V3
VSRAW $15, V2, V3
VSRAW $15, V2
RET
TEXT ·VSRLBbounds(SB), NOSPLIT, $0
VSRLB $0, V2, V3
VSRLB $7, V2, V3
VSRLB $3, V2, V3
VSRLB $3, V2
RET
TEXT ·VSRLHbounds(SB), NOSPLIT, $0
VSRLH $0, V2, V3
VSRLH $15, V2, V3
VSRLH $7, V2, V3
VSRLH $7, V2
RET
TEXT ·VSRLVbounds(SB), NOSPLIT, $0
VSRLV $0, V2, V3
VSRLV $63, V2, V3
VSRLV $31, V2, V3
VSRLV $31, V2
RET
TEXT ·VSRLWbounds(SB), NOSPLIT, $0
VSRLW $0, V2, V3
VSRLW $31, V2, V3
VSRLW $15, V2, V3
VSRLW $15, V2
RET
TEXT ·VXORBbounds(SB), NOSPLIT, $0
VXORB $0, V2, V3
VXORB $255, V2, V3
VXORB $127, V2, V3
VXORB $127, V2
RET
+148
View File
@@ -0,0 +1,148 @@
// Differential kernel: the divmod vector slice against the Go
// toolchain's loong64enc1.s rows (gasm v. go tool asm, byte for byte).
#include "textflag.h"
TEXT ·VDIVB(SB), NOSPLIT, $0
VDIVB V1, V2, V3
RET
TEXT ·VDIVH(SB), NOSPLIT, $0
VDIVH V1, V2, V3
RET
TEXT ·VDIVW(SB), NOSPLIT, $0
VDIVW V1, V2, V3
RET
TEXT ·VDIVV(SB), NOSPLIT, $0
VDIVV V1, V2, V3
RET
TEXT ·VDIVBU(SB), NOSPLIT, $0
VDIVBU V1, V2, V3
RET
TEXT ·VDIVHU(SB), NOSPLIT, $0
VDIVHU V1, V2, V3
RET
TEXT ·VDIVWU(SB), NOSPLIT, $0
VDIVWU V1, V2, V3
RET
TEXT ·VDIVVU(SB), NOSPLIT, $0
VDIVVU V1, V2, V3
RET
TEXT ·VMODB(SB), NOSPLIT, $0
VMODB V1, V2, V3
RET
TEXT ·VMODH(SB), NOSPLIT, $0
VMODH V1, V2, V3
RET
TEXT ·VMODW(SB), NOSPLIT, $0
VMODW V1, V2, V3
RET
TEXT ·VMODV(SB), NOSPLIT, $0
VMODV V1, V2, V3
RET
TEXT ·VMODBU(SB), NOSPLIT, $0
VMODBU V1, V2, V3
RET
TEXT ·VMODHU(SB), NOSPLIT, $0
VMODHU V1, V2, V3
RET
TEXT ·VMODWU(SB), NOSPLIT, $0
VMODWU V1, V2, V3
RET
TEXT ·VMODVU(SB), NOSPLIT, $0
VMODVU V1, V2, V3
RET
TEXT ·XVDIVB(SB), NOSPLIT, $0
XVDIVB X3, X2, X1
RET
TEXT ·XVDIVH(SB), NOSPLIT, $0
XVDIVH X3, X2, X1
RET
TEXT ·XVDIVW(SB), NOSPLIT, $0
XVDIVW X3, X2, X1
RET
TEXT ·XVDIVV(SB), NOSPLIT, $0
XVDIVV X3, X2, X1
RET
TEXT ·XVDIVBU(SB), NOSPLIT, $0
XVDIVBU X3, X2, X1
RET
TEXT ·XVDIVHU(SB), NOSPLIT, $0
XVDIVHU X3, X2, X1
RET
TEXT ·XVDIVWU(SB), NOSPLIT, $0
XVDIVWU X3, X2, X1
RET
TEXT ·XVDIVVU(SB), NOSPLIT, $0
XVDIVVU X3, X2, X1
RET
TEXT ·XVMODB(SB), NOSPLIT, $0
XVMODB X3, X2, X1
RET
TEXT ·XVMODH(SB), NOSPLIT, $0
XVMODH X3, X2, X1
RET
TEXT ·XVMODW(SB), NOSPLIT, $0
XVMODW X3, X2, X1
RET
TEXT ·XVMODV(SB), NOSPLIT, $0
XVMODV X3, X2, X1
RET
TEXT ·XVMODBU(SB), NOSPLIT, $0
XVMODBU X3, X2, X1
RET
TEXT ·XVMODHU(SB), NOSPLIT, $0
XVMODHU X3, X2, X1
RET
TEXT ·XVMODWU(SB), NOSPLIT, $0
XVMODWU X3, X2, X1
RET
TEXT ·XVMODVU(SB), NOSPLIT, $0
XVMODVU X3, X2, X1
RET
TEXT ·VDIVF(SB), NOSPLIT, $0
VDIVF V1, V2, V3
RET
TEXT ·VDIVD(SB), NOSPLIT, $0
VDIVD V1, V2, V3
RET
TEXT ·XVDIVF(SB), NOSPLIT, $0
XVDIVF X1, X2, X3
RET
TEXT ·XVDIVD(SB), NOSPLIT, $0
XVDIVD X1, X2, X3
RET
+176
View File
@@ -0,0 +1,176 @@
// Differential kernel: the fp vector slice against the Go
// toolchain's loong64enc1.s rows (gasm v. go tool asm, byte for byte).
#include "textflag.h"
TEXT ·FFINTFW(SB), NOSPLIT, $0
FFINTFW F0, F1
RET
TEXT ·FFINTFV(SB), NOSPLIT, $0
FFINTFV F0, F1
RET
TEXT ·FFINTDW(SB), NOSPLIT, $0
FFINTDW F0, F1
RET
TEXT ·FTINTWF(SB), NOSPLIT, $0
FTINTWF F0, F1
RET
TEXT ·FTINTWD(SB), NOSPLIT, $0
FTINTWD F0, F1
RET
TEXT ·FTINTVF(SB), NOSPLIT, $0
FTINTVF F0, F1
RET
TEXT ·FTINTVD(SB), NOSPLIT, $0
FTINTVD F0, F1
RET
TEXT ·VFSQRTF(SB), NOSPLIT, $0
VFSQRTF V1, V2
RET
TEXT ·VFSQRTD(SB), NOSPLIT, $0
VFSQRTD V1, V2
RET
TEXT ·VFRECIPF(SB), NOSPLIT, $0
VFRECIPF V1, V2
RET
TEXT ·VFRECIPD(SB), NOSPLIT, $0
VFRECIPD V1, V2
RET
TEXT ·VFRSQRTF(SB), NOSPLIT, $0
VFRSQRTF V1, V2
RET
TEXT ·VFRSQRTD(SB), NOSPLIT, $0
VFRSQRTD V1, V2
RET
TEXT ·XVFSQRTF(SB), NOSPLIT, $0
XVFSQRTF X2, X1
RET
TEXT ·XVFSQRTD(SB), NOSPLIT, $0
XVFSQRTD X2, X1
RET
TEXT ·XVFRECIPF(SB), NOSPLIT, $0
XVFRECIPF X2, X1
RET
TEXT ·XVFRECIPD(SB), NOSPLIT, $0
XVFRECIPD X2, X1
RET
TEXT ·XVFRSQRTF(SB), NOSPLIT, $0
XVFRSQRTF X2, X1
RET
TEXT ·XVFRSQRTD(SB), NOSPLIT, $0
XVFRSQRTD X2, X1
RET
TEXT ·VFRINTRNEF(SB), NOSPLIT, $0
VFRINTRNEF V1, V2
RET
TEXT ·VFRINTRNED(SB), NOSPLIT, $0
VFRINTRNED V1, V2
RET
TEXT ·VFRINTRZF(SB), NOSPLIT, $0
VFRINTRZF V1, V2
RET
TEXT ·VFRINTRZD(SB), NOSPLIT, $0
VFRINTRZD V1, V2
RET
TEXT ·VFRINTRPF(SB), NOSPLIT, $0
VFRINTRPF V1, V2
RET
TEXT ·VFRINTRPD(SB), NOSPLIT, $0
VFRINTRPD V1, V2
RET
TEXT ·VFRINTRMF(SB), NOSPLIT, $0
VFRINTRMF V1, V2
RET
TEXT ·VFRINTRMD(SB), NOSPLIT, $0
VFRINTRMD V1, V2
RET
TEXT ·VFRINTF(SB), NOSPLIT, $0
VFRINTF V1, V2
RET
TEXT ·VFRINTD(SB), NOSPLIT, $0
VFRINTD V1, V2
RET
TEXT ·XVFRINTRNEF(SB), NOSPLIT, $0
XVFRINTRNEF X1, X2
RET
TEXT ·XVFRINTRNED(SB), NOSPLIT, $0
XVFRINTRNED X1, X2
RET
TEXT ·XVFRINTRZF(SB), NOSPLIT, $0
XVFRINTRZF X1, X2
RET
TEXT ·XVFRINTRZD(SB), NOSPLIT, $0
XVFRINTRZD X1, X2
RET
TEXT ·XVFRINTRPF(SB), NOSPLIT, $0
XVFRINTRPF X1, X2
RET
TEXT ·XVFRINTRPD(SB), NOSPLIT, $0
XVFRINTRPD X1, X2
RET
TEXT ·XVFRINTRMF(SB), NOSPLIT, $0
XVFRINTRMF X1, X2
RET
TEXT ·XVFRINTRMD(SB), NOSPLIT, $0
XVFRINTRMD X1, X2
RET
TEXT ·XVFRINTF(SB), NOSPLIT, $0
XVFRINTF X1, X2
RET
TEXT ·XVFRINTD(SB), NOSPLIT, $0
XVFRINTD X1, X2
RET
TEXT ·VFCLASSF(SB), NOSPLIT, $0
VFCLASSF V1, V2
RET
TEXT ·VFCLASSD(SB), NOSPLIT, $0
VFCLASSD V1, V2
RET
TEXT ·XVFCLASSF(SB), NOSPLIT, $0
XVFCLASSF X1, X2
RET
TEXT ·XVFCLASSD(SB), NOSPLIT, $0
XVFCLASSD X1, X2
RET

Some files were not shown because too many files have changed in this diff Show More