Compare commits
47
Commits
c66a47973a
..
main
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
cf6bc6987e | ||
|
|
ff7b1452b1 | ||
|
|
517c1cea25 | ||
|
|
a3e3010e0f | ||
|
|
057c4eb545 | ||
|
|
f720381d43 | ||
|
|
2c9042d62c | ||
|
|
82ef289d3a | ||
|
|
7246b0e002 | ||
|
|
8cfd40aac8 | ||
|
|
5382c9a8e4 | ||
|
|
53de91b2df | ||
|
|
8a36af7c7d | ||
|
|
e9789ce3f4 | ||
|
|
837231c068 | ||
|
|
95025be1bc | ||
|
|
03a964bb2d | ||
|
|
123a16e346 | ||
|
|
9701812bee | ||
|
|
29ac03468e | ||
|
|
bfb7701db1 | ||
|
|
e8b6ff5d7c | ||
|
|
1456907000 | ||
|
|
ec1c521187 | ||
|
|
a7744c24bd | ||
|
|
522e6f2ae8 | ||
|
|
81d4bd81e4 | ||
|
|
687678a2ea | ||
|
|
b0f9071bf5 | ||
|
|
81e2673923 | ||
|
|
75e9fd771b | ||
|
|
863926abd6 | ||
|
|
241e7256f6 | ||
|
|
6556b85abf | ||
|
|
289cabe993 | ||
|
|
d6cf7cfa44 | ||
|
|
4cc2f0eba5 | ||
|
|
97dfaa7526 | ||
|
|
66aa4dbc8b | ||
|
|
dce5d31462 | ||
|
|
9dc3987e02 | ||
|
|
9b238a525a | ||
|
|
ad82aac663 | ||
|
|
0629f5e2df | ||
|
|
ecb203dcf5 | ||
|
|
6c672567f3 | ||
|
|
cc6e416c59 |
@@ -342,7 +342,10 @@ jobs:
|
||||
my @cmd = (q{curl}, q{-sS}, q{-o}, q{/dev/null}, q{-w}, q{%{http_code}},
|
||||
q{-H}, qq{Authorization: token $ENV{GITEA_TOKEN}},
|
||||
q{-H}, q{Content-Type: application/octet-stream},
|
||||
q{-X}, q{POST}, q{--data-binary}, qq{@$path},
|
||||
# The @ must not sit inside a qq{} string: there it starts an
|
||||
# array interpolation and the upload body collapses to empty,
|
||||
# which Gitea stores as a 201-created zero-byte attachment.
|
||||
q{-X}, q{POST}, q{--data-binary}, q{@} . $path,
|
||||
qq{$ENV{GITEA_SERVER_URL}/api/v1/repos/$ENV{GITEA_REPOSITORY}/releases/$id/assets?name=$name});
|
||||
open(my $curl, q{-|}, @cmd) or die qq{curl: $!};
|
||||
my $code = <$curl>;
|
||||
|
||||
+126
-9
@@ -9,7 +9,43 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
||||
|
||||
### Added
|
||||
|
||||
- **The GOROOT instruction wave, part 1.** The encoder now covers the
|
||||
-
|
||||
|
||||
## [0.35.0] - 2026-09-22
|
||||
|
||||
### Added
|
||||
|
||||
- **The go_asm.h generator.** `gasm asm` generates the package's go_asm.h
|
||||
itself when an assembly file includes it: the Go files beside the source
|
||||
are type-checked for the target architecture and the constants and field
|
||||
offsets become assembler defines, so package-context files assemble with
|
||||
no compiler and no `go build` in the loop. `-GOOS` selects the
|
||||
type-checking GOOS for GOOS-specific files, and the corpus audit derives
|
||||
the GOOS from the file name.
|
||||
- **ELF data relocations on arm64, riscv64 and loong64.** `gasm asm
|
||||
--format elf` emits `.rela.data` for symbol-valued DATA initialisers on
|
||||
every architecture (amd64 carried them already), so standalone ELF
|
||||
objects link on all four targets.
|
||||
- **Corpus failure listing.** `gasm audit-instructions --corpus --list`
|
||||
prints every failing file with its failure reason, per architecture,
|
||||
instead of one representative file per reason.
|
||||
- **DATA with symbol values and relaxed symbol spellings.** DATA
|
||||
initialisers accept `$symbol(SB)` values, laid down as an absolute
|
||||
relocation at the data field (GOOBJ on all four architectures and ELF
|
||||
on all four as of this release), and U+2215 is accepted inside symbol
|
||||
package paths.
|
||||
- **Macro expansion and include splicing.** `gasm asm`, `gasm diff` and
|
||||
`gasm audit-instructions` now preprocess assembly the way the
|
||||
toolchain does: object and parameterised `#define` macros expand at
|
||||
the point of use, `#undef` and the `#ifdef`/`#ifndef`/`#else`/
|
||||
`#endif` family select branches, `#include` splices headers resolved
|
||||
through the source directory and the new repeatable `-I` flag, `;`
|
||||
separates statements, and constant expressions left in operands
|
||||
(`$(32-7)`, `$~63`, `(index*4)(base)`) fold at parse. Expansion
|
||||
happens only on the assembly path: `gasm lint`, `gasm fmt` and the
|
||||
language server keep reading the raw file.
|
||||
- **Encoder coverage: the instruction families GOROOT's real code
|
||||
uses.** The encoder now covers the
|
||||
instruction families GOROOT's real code uses that gasm lacked,
|
||||
byte-verified against `go tool asm`: on amd64 the carry ALU, the
|
||||
atomics (CMPXCHG, XADD, XCHG), AES-NI, SHA-1/256, PCLMULQDQ, CRC32,
|
||||
@@ -25,14 +61,90 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
||||
Also fixed on the way: arm64 `CASD`/`CASW` lacked an opcode bit, and
|
||||
riscv64 `VSETVLI` with an immediate length now canonicalises to
|
||||
`vsetivli` as the toolchain does.
|
||||
- **The corpus audit measures honestly.** Files named for Go ports gasm
|
||||
does not target (arm, 386, s390x, ...) are no longer attempted for the
|
||||
four supported architectures (no supported build compiles them), and
|
||||
the headline rate is reported over attemptable files: 136 of 433 on
|
||||
the full corpus (31.4 %), 135 of 383 on real code (35.2 %), from the
|
||||
127 that the previous release measured. The probe battery that
|
||||
decides encodability gained the operand shapes the new families use.
|
||||
-
|
||||
- **Encoder coverage: quad-register AVX-512 and floating-point
|
||||
immediates.** The encoder gains the
|
||||
quad-register AVX-512 families (4FMAPS, 4FNMADD, 4VNNIW, VP4DPWSSD,
|
||||
VP4DPWSSDS) with the register list riding the inverted V'VVVV field,
|
||||
floating-point immediates on the SSE scalar moves and arithmetic
|
||||
(the constant lands in a synthesised read-only pool, a positive zero
|
||||
collapses to XORPS exactly as the toolchain does), accept-and-ignore
|
||||
FUNCDATA and PCDATA, three-operand double shifts, static-symbol
|
||||
operands for the legacy SSE moves, and the pooled 64-bit immediate
|
||||
materialisation on riscv64. The parser carries bracketed register
|
||||
ranges, index-only VSIB memory operands and bare trailing immediates;
|
||||
macro substitution reaches parameters used with element suffixes
|
||||
(`A.S4`), and `;` separates statements in plain files.
|
||||
- **Per-architecture reference pages.** [docs/asm/](docs/asm/README.md)
|
||||
gains AMD64, ARM64, RISCV64 and LOONG64: the register files and the
|
||||
roles the ABI fixes, addressing, operand order with every special form,
|
||||
constants and materialisation, alignment, fences and the relocations
|
||||
each target emits. An instruction inventory appendix per architecture
|
||||
is generated from the toolchain's own tables by `just gen`, and the
|
||||
regenerated tables recognise 147 more mnemonics than the previous
|
||||
release carried (arm64 107, riscv64 31, loong64 9).
|
||||
- **The Plan 9 assembly language reference.** [docs/asm/](docs/asm/README.md)
|
||||
opens the complete language reference with its common core: the lexicon,
|
||||
statement structure and constant expressions, the operand grammar with
|
||||
the pseudo-registers and symbol naming, the directives and the function
|
||||
flag vocabulary, preprocessing with `#define` and `#include`, and the
|
||||
Go-embedded layer (ABI0, prototypes, `go_asm.h`, `funcdata.h` and the
|
||||
runtime contract). Every claim is verified against `go tool asm` of
|
||||
Go 1.27.1 and gasm's differential tests; the per-architecture pages and
|
||||
generated instruction appendices follow.
|
||||
- **GOOBJ format specification.** [docs/GOOBJ.md](docs/GOOBJ.md)
|
||||
documents the Go object file format in full: both containers, the 96
|
||||
byte header and all 19 blocks, every structure with its byte
|
||||
offsets, symbol kinds and flag bits, all 106 relocation types with
|
||||
the weak variants, aux symbols, the FuncInfo payload, the pc-value
|
||||
table encoding, the content hashes and the builtin table, all
|
||||
verified byte for byte against objects produced by Go 1.27.1's own
|
||||
tools.
|
||||
|
||||
### Changed
|
||||
|
||||
- **The corpus audit measures like a build.** Files named for a Go port
|
||||
gasm does not target (arm, 386, s390x, ...) are never attempted, because
|
||||
no supported build compiles them; the GOOS comes from the file name; and
|
||||
each target's go_asm.h is generated on the fly. The headline is reported
|
||||
over attemptable files: 291 of 353 on the full corpus (82.4 %) assemble
|
||||
for every target architecture and 295 of 303 on real code (97.4 %),
|
||||
against 108 of 627 over all files (17.2 %) that the previous release
|
||||
measured.
|
||||
|
||||
### Fixed
|
||||
|
||||
- **The operand forms GOROOT writes.** Numeric PC-relative jumps
|
||||
(`JEQ 2(PC)`, the park loop `JMP 0(PC)`) resolve with the toolchain's
|
||||
own instruction counting and fold jump-to-jump chains exactly as its
|
||||
branch optimiser does; symbol immediates (`MOVQ $sym(SB), AX`)
|
||||
assemble to the toolchain's RIP-relative LEA with an R_PCREL
|
||||
relocation; negated constant expressions in operands (`ADJSP
|
||||
$-(REGS - 8)`, the shape the cgo ABI macros write) fold; the immediate
|
||||
multiply (`IMULQ $1000000000, AX`) encodes with the toolchain's
|
||||
0x69/0x6B selection; the TLS access pair assembles as the toolchain's
|
||||
one-instruction form (the bare `MOVQ TLS, r` load nops out and
|
||||
`off(r)(TLS*1)` folds to the segment-prefixed absolute whose disp32
|
||||
carries the R_TLSLE relocation, per-GOOS); arm64 accepts the
|
||||
bare-register indirect branch (`BL R9` beside `BL (R9)`, both BLR) and
|
||||
the zero-immediate store (`MOVD $0, mem` through the zero register,
|
||||
rejecting non-zero immediates as the toolchain does); `PCALIGN` now
|
||||
aligns on amd64, padding with the toolchain's greedy
|
||||
single-instruction NOPs; the segment-absolute forms (`MOVQ 0x30(GS),
|
||||
AX` and the store direction) and the absolute crash-store
|
||||
(`MOVL $0xf1, 0xf1`) encode; and `gasm asm` predefines the
|
||||
`GOARCH_<arch>` and `GOOS_<goos>` macros the go command passes to
|
||||
`go tool asm`, so GOROOT headers' `#ifdef GOARCH_amd64` platform
|
||||
blocks (`go_tls.h`'s `get_tls` and friends) select as intended. The
|
||||
GOROOT corpus measure moves to 291 of 353 files assembling for every
|
||||
target architecture (82.4 %), 97.4 % of the real-code corpus, from
|
||||
70.8 % and 82.2 %.
|
||||
- **Tool corrections across the pipeline.** The formatter keeps square
|
||||
brackets in SIMD operands, statement separators and canonical macro
|
||||
bodies; the linter drops false positives on shift counts, SETcc
|
||||
spellings and ABIInternal references; the lexer treats a trailing
|
||||
carriage return as a line end so comment text stays idempotent; and
|
||||
arm64 rejects bare BTI with a diagnostic while accepting the full
|
||||
family.
|
||||
|
||||
## [0.34.0] - 2026-09-20
|
||||
|
||||
@@ -129,6 +241,11 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
||||
|
||||
### Fixed
|
||||
|
||||
- **The corpus audit attempts fewer files that no build would compile.**
|
||||
Files named for Go ports gasm does not target (arm, 386, s390x, ...)
|
||||
are reported as other-port and never attempted, the headline rate is
|
||||
computed over attemptable files, and the audit searches the
|
||||
toolchain's shipped headers (funcdata.h and friends) automatically.
|
||||
- **riscv64 JALR silently jumped to the wrong register.** The trampoline
|
||||
form `JALR X0, 0(X5)` read the memory operand's base as the destination,
|
||||
encoding a jump to X0 with no diagnostic; the destination is the first
|
||||
|
||||
@@ -81,7 +81,10 @@ to give that syntax the tooling it deserves.
|
||||
GOOBJ format, which needs the installed toolchain and which `go build`
|
||||
consumes in place of the toolchain's output. Framed functions get the
|
||||
stack-split guard and the morestack block, byte-identical to the
|
||||
toolchain's, so split functions link too.
|
||||
toolchain's, so split functions link too. The assembler preprocesses
|
||||
like the toolchain (`#define`, `#include` with `-I`, `#ifdef`), generates
|
||||
`go_asm.h` from the package's Go files, and carries `PCALIGN`, the
|
||||
`LOCK`/`REP` prefixes and the literal-data pseudo-ops.
|
||||
- **Disassembler.** `gasm dis` lists a `.s` file's functions at their real
|
||||
offsets after assembling, or disassembles raw bytes from a file or stdin.
|
||||
- **Dynamic verification.** `gasm verify` JIT-loads assembled functions into
|
||||
@@ -110,9 +113,9 @@ Four architectures, the four that matter in practice:
|
||||
| Architecture | GOARCH | File suffix | Instructions recognised |
|
||||
|--------------|-------------|--------------|---------------------------------------------|
|
||||
| AMD64 | `amd64` | `_amd64.s` | 1600 + common opcodes + traditional aliases |
|
||||
| ARM64 | `arm64` | `_arm64.s` | 538 + common opcodes |
|
||||
| RISC-V | `riscv64` | `_riscv64.s` | 961 + common opcodes |
|
||||
| LoongArch | `loong64` | `_loong64.s` | 799 + common opcodes |
|
||||
| ARM64 | `arm64` | `_arm64.s` | 645 + common opcodes |
|
||||
| RISC-V | `riscv64` | `_riscv64.s` | 992 + common opcodes |
|
||||
| LoongArch | `loong64` | `_loong64.s` | 808 + common opcodes |
|
||||
|
||||
"Common opcodes" are the instructions shared by every architecture (`RET`,
|
||||
`JMP`, `NOP`, `CALL`, `TEXT`, `FUNCDATA`, `PCDATA`, ...). AMD64 additionally
|
||||
@@ -124,8 +127,8 @@ can emit today is narrower, and a recognised but unencodable instruction is
|
||||
reported as an explicit error, never as a wrong byte.
|
||||
|
||||
The same measurement runs over GOROOT's whole assembly corpus:
|
||||
`gasm audit-instructions --corpus` reports 136 of 433 attemptable files
|
||||
(31.4 %) assembling for every target architecture today (files named for
|
||||
`gasm audit-instructions --corpus` reports 291 of 353 attemptable files
|
||||
(82.4 %) assembling for every target architecture today (files named for
|
||||
other Go ports are counted but never attempted), with the top failure
|
||||
reasons per architecture; the number moves with every release.
|
||||
|
||||
@@ -157,6 +160,39 @@ been compiled and read, never executed. Its architecture-neutral units
|
||||
run under `go test ./...`, which the race workflow and a manual run
|
||||
perform; the default `just test` gate does not sweep `./debug/...`.
|
||||
|
||||
## The documentation goal
|
||||
|
||||
The toolkit is the primary goal. The secondary one is documentation: a
|
||||
specification of the Plan 9 assembly language and of the GOOBJ object
|
||||
format that is 100 % complete, detailed enough to implement against,
|
||||
and written to a professional standard. These are the two subjects this
|
||||
project works with every day, and they are the two for which no usable
|
||||
documentation exists.
|
||||
|
||||
Go documents the language on a single page, "A Quick Guide to Go's
|
||||
Assembler", which carries no section for loong64, one of the four
|
||||
architectures gasm supports, and covers a fraction of what each
|
||||
assembler accepts. What exists beyond it lives as comments inside the
|
||||
toolchain's internal source: per-architecture reference manuals for
|
||||
arm64, ppc64, riscv64 and loong64, written for the toolchain's own
|
||||
maintainers rather than for an outside reader, and none at all for
|
||||
amd64. GOOBJ fares worst of all. The format that `go build` consumes
|
||||
has no specification anywhere: it is described by a comment in an
|
||||
internal package, it is not a stable interface, and it can change with
|
||||
any toolchain release.
|
||||
|
||||
The gap is therefore filled the only way it can be filled: by reverse
|
||||
engineering the toolchain itself, the same work the encoders already
|
||||
perform. Most of the documentation can come from nowhere else, and it
|
||||
is written as that knowledge is produced during development. It is
|
||||
verified the way the code is verified: an encoding documented here is
|
||||
one that differential tests against `go tool asm` confirm
|
||||
byte-for-byte, and a format field documented here is one the linker
|
||||
demonstrably reads. The work has begun: [docs/GOOBJ.md](docs/GOOBJ.md)
|
||||
specifies the object file format completely, and
|
||||
[docs/asm/README.md](docs/asm/README.md) opens the language reference
|
||||
with its common core. The per-architecture pages follow.
|
||||
|
||||
## Direction
|
||||
|
||||
The plan, in the order it is being worked:
|
||||
@@ -282,6 +318,8 @@ recipe.
|
||||
~/.local/share/man (MANDIR overrides); `just uninstall-man` removes
|
||||
them
|
||||
- [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md): components and data flow
|
||||
- [docs/GOOBJ.md](docs/GOOBJ.md): the GOOBJ object file format specification
|
||||
- [docs/asm/](docs/asm/README.md): the Plan 9 assembly language reference
|
||||
- [docs/DEVELOPMENT.md](docs/DEVELOPMENT.md): development setup and recipes
|
||||
- [CHANGELOG.md](CHANGELOG.md): release history
|
||||
|
||||
|
||||
+1
-1
@@ -7,7 +7,7 @@ releases do not receive them.
|
||||
|
||||
| Version | Supported |
|
||||
|---|---|
|
||||
| 0.34.0 | yes |
|
||||
| 0.35.0 | yes |
|
||||
| older releases | no |
|
||||
|
||||
## Reporting a vulnerability
|
||||
|
||||
+101
-3
@@ -8,6 +8,10 @@
|
||||
// names so gasm-devkit supports every instruction the real assembler does,
|
||||
// with no hand-maintained (and therefore inevitably incomplete) lists.
|
||||
//
|
||||
// The same data feeds the generated instruction appendices of the assembly
|
||||
// language reference, docs/asm/INSTRUCTIONS-<ARCH>.md, so that the reference
|
||||
// cannot drift from the tables it documents.
|
||||
//
|
||||
// Usage (via the justfile):
|
||||
//
|
||||
// just gen
|
||||
@@ -26,6 +30,9 @@ import (
|
||||
"path/filepath"
|
||||
"sort"
|
||||
"strings"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/asm"
|
||||
)
|
||||
|
||||
// archDirs maps a gasm-devkit architecture name to its obj sub-directory.
|
||||
@@ -39,11 +46,30 @@ var archDirs = []struct {
|
||||
{"loong64", "loong64"},
|
||||
}
|
||||
|
||||
// docPages maps an architecture to its generated appendix in the language
|
||||
// reference. The amd64 page carries a per-mnemonic encodability column,
|
||||
// decided by asm.Encodable, which mirrors the encoder's own dispatch; the
|
||||
// other targets have no single cheap predicate, so their pages carry the
|
||||
// inventory and point at the live measurement instead.
|
||||
var docPages = []struct {
|
||||
arch arch.Arch
|
||||
title string
|
||||
file string
|
||||
anames string
|
||||
encodable bool
|
||||
}{
|
||||
{arch.AMD64, "AMD64", "INSTRUCTIONS-AMD64.md", "cmd/internal/obj/x86/anames.go", true},
|
||||
{arch.ARM64, "ARM64", "INSTRUCTIONS-ARM64.md", "cmd/internal/obj/arm64/anames.go", false},
|
||||
{arch.RISCV, "RISC-V 64", "INSTRUCTIONS-RISCV64.md", "cmd/internal/obj/riscv/anames.go", false},
|
||||
{arch.LOONG64, "LoongArch 64", "INSTRUCTIONS-LOONG64.md", "cmd/internal/obj/loong64/anames.go", false},
|
||||
}
|
||||
|
||||
func main() {
|
||||
goroot := strings.TrimSpace(runGoEnvGOROOT())
|
||||
if goroot == "" {
|
||||
fatal("could not determine GOROOT")
|
||||
}
|
||||
version := strings.TrimSpace(runGoEnv("GOVERSION"))
|
||||
// The common opcodes shared by every architecture (RET, JMP, NOP, CALL,
|
||||
// TEXT, FUNCDATA, …) live in cmd/internal/obj/util.go.
|
||||
commonPath := filepath.Join(goroot, "src", "cmd", "internal", "obj", "util.go")
|
||||
@@ -57,16 +83,24 @@ func main() {
|
||||
}
|
||||
fmt.Printf("%-8s %4d instructions -> arch/common_gen.go\n", "common", len(common))
|
||||
|
||||
names := map[string][]string{}
|
||||
for _, a := range archDirs {
|
||||
path := filepath.Join(goroot, "src", "cmd", "internal", "obj", a.sub, "anames.go")
|
||||
names, err := extractInstrs(path)
|
||||
names[a.arch], err = extractInstrs(path)
|
||||
if err != nil {
|
||||
fatal("extract %s: %v", a.arch, err)
|
||||
}
|
||||
if err := writeGen(a.arch, a.sub, names); err != nil {
|
||||
if err := writeGen(a.arch, a.sub, names[a.arch]); err != nil {
|
||||
fatal("write %s: %v", a.arch, err)
|
||||
}
|
||||
fmt.Printf("%-8s %4d instructions -> arch/%s_gen.go\n", a.arch, len(names), a.arch)
|
||||
fmt.Printf("%-8s %4d instructions -> arch/%s_gen.go\n", a.arch, len(names[a.arch]), a.arch)
|
||||
}
|
||||
|
||||
for _, p := range docPages {
|
||||
if err := writeDocPage(p.arch, p.title, p.file, p.anames, version, p.encodable); err != nil {
|
||||
fatal("write %s: %v", p.file, err)
|
||||
}
|
||||
fmt.Printf("%-8s -> docs/asm/%s\n", p.arch, p.file)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -172,6 +206,61 @@ func writeGen(arch, sub string, names []string) error {
|
||||
return os.WriteFile(filepath.Join("arch", arch+"_gen.go"), []byte(b.String()), 0o644)
|
||||
}
|
||||
|
||||
// writeDocPage emits docs/asm/<file>, the generated instruction appendix of
|
||||
// the language reference for one architecture: every mnemonic the toolchain
|
||||
// accepts, with the curated summary where the architecture table carries one
|
||||
// and, on amd64, a per-mnemonic encodability column.
|
||||
func writeDocPage(a arch.Arch, title, file, anames, version string, encodable bool) error {
|
||||
table := arch.ForArch(a)
|
||||
instrs := table.Instructions()
|
||||
|
||||
var b strings.Builder
|
||||
b.WriteString("# " + title + ": instruction inventory\n\n")
|
||||
b.WriteString("Generated by gasm-devkit's `_gen` from the Go toolchain's instruction table\n")
|
||||
b.WriteString("(`" + anames + "`, " + version + "); DO NOT EDIT. This page lists every mnemonic\n")
|
||||
b.WriteString("`go tool asm` accepts on this target, which is the upper bound of the\n")
|
||||
b.WriteString("language on it: a name absent here is not an instruction of the target,\n")
|
||||
b.WriteString("and a name present here may still be one gasm's encoder cannot emit yet.\n\n")
|
||||
|
||||
encodableCount := 0
|
||||
if encodable {
|
||||
b.WriteString("The `gasm encodes` column reports whether gasm's encoder can emit the\n")
|
||||
b.WriteString("mnemonic today; the gap is the encoder backlog, measured live by\n")
|
||||
b.WriteString("`gasm audit-instructions`.\n\n")
|
||||
b.WriteString("| Mnemonic | gasm encodes | Notes |\n")
|
||||
b.WriteString("|---|---|---|\n")
|
||||
for _, in := range instrs {
|
||||
ok := asm.Encodable(in.Name)
|
||||
if ok {
|
||||
encodableCount++
|
||||
}
|
||||
b.WriteString("| `" + in.Name + "` | " + yesNo(ok) + " | " + in.Summary + " |\n")
|
||||
}
|
||||
b.WriteString("\n")
|
||||
fmt.Fprintf(&b, "Recognised: %d mnemonics. gasm encodes: %d.\n", len(instrs), encodableCount)
|
||||
} else {
|
||||
b.WriteString("The inventory carries no per-mnemonic encoder column: on this target\n")
|
||||
b.WriteString("encodability is decided per operand shape, and the live measured\n")
|
||||
b.WriteString("coverage is reported by `gasm audit-instructions`.\n\n")
|
||||
b.WriteString("| Mnemonic | Notes |\n")
|
||||
b.WriteString("|---|---|\n")
|
||||
for _, in := range instrs {
|
||||
b.WriteString("| `" + in.Name + "` | " + in.Summary + " |\n")
|
||||
}
|
||||
b.WriteString("\n")
|
||||
fmt.Fprintf(&b, "Recognised: %d mnemonics.\n", len(instrs))
|
||||
}
|
||||
return os.WriteFile(filepath.Join("docs", "asm", file), []byte(b.String()), 0o644)
|
||||
}
|
||||
|
||||
// yesNo renders a boolean as the word the appendix tables use.
|
||||
func yesNo(v bool) string {
|
||||
if v {
|
||||
return "yes"
|
||||
}
|
||||
return "no"
|
||||
}
|
||||
|
||||
func runGoEnvGOROOT() string {
|
||||
out, err := exec.Command("go", "env", "GOROOT").Output()
|
||||
if err != nil {
|
||||
@@ -180,6 +269,15 @@ func runGoEnvGOROOT() string {
|
||||
return string(out)
|
||||
}
|
||||
|
||||
// runGoEnv runs `go env` for a single variable.
|
||||
func runGoEnv(name string) string {
|
||||
out, err := exec.Command("go", "env", name).Output()
|
||||
if err != nil {
|
||||
return ""
|
||||
}
|
||||
return string(out)
|
||||
}
|
||||
|
||||
func fatal(format string, args ...any) {
|
||||
fmt.Fprintf(os.Stderr, "gen: "+format+"\n", args...)
|
||||
os.Exit(1)
|
||||
|
||||
@@ -33,6 +33,7 @@ func arm64Registers() []Register {
|
||||
for i := 0; i <= 30; i++ {
|
||||
add(fmt.Sprintf("R%d", i), GPR, "64-bit general-purpose register")
|
||||
}
|
||||
add("R18_PLATFORM", GPR, "R18 under its toolchain-reserved Windows name (an alias of R18)")
|
||||
add("ZR", Special, "zero register (reads as 0)")
|
||||
add("SP", Special, "stack pointer")
|
||||
add("LR", Special, "link register (alias of R30)")
|
||||
|
||||
@@ -364,6 +364,8 @@ var arm64GeneratedInstrs = []string{
|
||||
"REVW",
|
||||
"ROR",
|
||||
"RORW",
|
||||
"RPRFM",
|
||||
"SB",
|
||||
"SBC",
|
||||
"SBCS",
|
||||
"SBCSW",
|
||||
@@ -477,23 +479,68 @@ var arm64GeneratedInstrs = []string{
|
||||
"UXTH",
|
||||
"UXTHW",
|
||||
"UXTW",
|
||||
"VABS",
|
||||
"VADD",
|
||||
"VADDP",
|
||||
"VADDV",
|
||||
"VAND",
|
||||
"VBCAX",
|
||||
"VBIC",
|
||||
"VBIF",
|
||||
"VBIT",
|
||||
"VBSL",
|
||||
"VCLS",
|
||||
"VCLZ",
|
||||
"VCMEQ",
|
||||
"VCMGE",
|
||||
"VCMGT",
|
||||
"VCMHI",
|
||||
"VCMHS",
|
||||
"VCMLE",
|
||||
"VCMLT",
|
||||
"VCMTST",
|
||||
"VCNT",
|
||||
"VDUP",
|
||||
"VEOR",
|
||||
"VEOR3",
|
||||
"VEXT",
|
||||
"VFABS",
|
||||
"VFADD",
|
||||
"VFADDP",
|
||||
"VFCMEQ",
|
||||
"VFCMGE",
|
||||
"VFCMGT",
|
||||
"VFCMLE",
|
||||
"VFCMLT",
|
||||
"VFCVTL",
|
||||
"VFCVTL2",
|
||||
"VFCVTN",
|
||||
"VFCVTN2",
|
||||
"VFCVTZS",
|
||||
"VFCVTZU",
|
||||
"VFDIV",
|
||||
"VFMAX",
|
||||
"VFMAXNM",
|
||||
"VFMAXNMP",
|
||||
"VFMAXNMV",
|
||||
"VFMAXP",
|
||||
"VFMAXV",
|
||||
"VFMIN",
|
||||
"VFMINNM",
|
||||
"VFMINNMP",
|
||||
"VFMINNMV",
|
||||
"VFMINP",
|
||||
"VFMINV",
|
||||
"VFMLA",
|
||||
"VFMLS",
|
||||
"VFMUL",
|
||||
"VFNEG",
|
||||
"VFRINTM",
|
||||
"VFRINTN",
|
||||
"VFRINTP",
|
||||
"VFRINTZ",
|
||||
"VFSQRT",
|
||||
"VFSUB",
|
||||
"VLD1",
|
||||
"VLD1R",
|
||||
"VLD2",
|
||||
@@ -502,11 +549,17 @@ var arm64GeneratedInstrs = []string{
|
||||
"VLD3R",
|
||||
"VLD4",
|
||||
"VLD4R",
|
||||
"VMLA",
|
||||
"VMLS",
|
||||
"VMOV",
|
||||
"VMOVD",
|
||||
"VMOVI",
|
||||
"VMOVQ",
|
||||
"VMOVS",
|
||||
"VMUL",
|
||||
"VNEG",
|
||||
"VNOT",
|
||||
"VORN",
|
||||
"VORR",
|
||||
"VPMULL",
|
||||
"VPMULL2",
|
||||
@@ -515,14 +568,47 @@ var arm64GeneratedInstrs = []string{
|
||||
"VREV16",
|
||||
"VREV32",
|
||||
"VREV64",
|
||||
"VSCVTF",
|
||||
"VSHADD",
|
||||
"VSHL",
|
||||
"VSHRN",
|
||||
"VSHRN2",
|
||||
"VSLI",
|
||||
"VSMAX",
|
||||
"VSMAXP",
|
||||
"VSMAXV",
|
||||
"VSMIN",
|
||||
"VSMINP",
|
||||
"VSMINV",
|
||||
"VSMLAL",
|
||||
"VSMLAL2",
|
||||
"VSMLSL",
|
||||
"VSMLSL2",
|
||||
"VSMULL",
|
||||
"VSMULL2",
|
||||
"VSQABS",
|
||||
"VSQADD",
|
||||
"VSQNEG",
|
||||
"VSQSHL",
|
||||
"VSQSUB",
|
||||
"VSQXTN",
|
||||
"VSQXTN2",
|
||||
"VSQXTUN",
|
||||
"VSQXTUN2",
|
||||
"VSRHADD",
|
||||
"VSRI",
|
||||
"VSRSHR",
|
||||
"VSSHL",
|
||||
"VSSHLL",
|
||||
"VSSHLL2",
|
||||
"VSSHR",
|
||||
"VST1",
|
||||
"VST2",
|
||||
"VST3",
|
||||
"VST4",
|
||||
"VSUB",
|
||||
"VSXTL",
|
||||
"VSXTL2",
|
||||
"VTBL",
|
||||
"VTBX",
|
||||
"VTRN1",
|
||||
@@ -530,8 +616,27 @@ var arm64GeneratedInstrs = []string{
|
||||
"VUADDLV",
|
||||
"VUADDW",
|
||||
"VUADDW2",
|
||||
"VUCVTF",
|
||||
"VUHADD",
|
||||
"VUMAX",
|
||||
"VUMAXP",
|
||||
"VUMAXV",
|
||||
"VUMIN",
|
||||
"VUMINP",
|
||||
"VUMINV",
|
||||
"VUMLAL",
|
||||
"VUMLAL2",
|
||||
"VUMLSL",
|
||||
"VUMLSL2",
|
||||
"VUMULL",
|
||||
"VUMULL2",
|
||||
"VUQADD",
|
||||
"VUQSHL",
|
||||
"VUQSUB",
|
||||
"VUQXTN",
|
||||
"VUQXTN2",
|
||||
"VURHADD",
|
||||
"VUSHL",
|
||||
"VUSHLL",
|
||||
"VUSHLL2",
|
||||
"VUSHR",
|
||||
@@ -541,6 +646,8 @@ var arm64GeneratedInstrs = []string{
|
||||
"VUZP1",
|
||||
"VUZP2",
|
||||
"VXAR",
|
||||
"VXTN",
|
||||
"VXTN2",
|
||||
"VZIP1",
|
||||
"VZIP2",
|
||||
"WFE",
|
||||
|
||||
@@ -152,6 +152,8 @@ var loong64GeneratedInstrs = []string{
|
||||
"FNMADDF",
|
||||
"FNMSUBD",
|
||||
"FNMSUBF",
|
||||
"FRINTD",
|
||||
"FRINTF",
|
||||
"FSCALEBD",
|
||||
"FSCALEBF",
|
||||
"FSEL",
|
||||
@@ -177,7 +179,10 @@ var loong64GeneratedInstrs = []string{
|
||||
"FTINTWF",
|
||||
"JIRL",
|
||||
"LL",
|
||||
"LLACQV",
|
||||
"LLACQW",
|
||||
"LLV",
|
||||
"LLW",
|
||||
"LU12IW",
|
||||
"LU32ID",
|
||||
"LU52ID",
|
||||
@@ -248,7 +253,11 @@ var loong64GeneratedInstrs = []string{
|
||||
"ROTR",
|
||||
"ROTRV",
|
||||
"SC",
|
||||
"SCQ",
|
||||
"SCRELV",
|
||||
"SCRELW",
|
||||
"SCV",
|
||||
"SCW",
|
||||
"SGT",
|
||||
"SGTU",
|
||||
"SLL",
|
||||
|
||||
@@ -81,6 +81,9 @@ var riscvGeneratedInstrs = []string{
|
||||
"CLD",
|
||||
"CLDSP",
|
||||
"CLI",
|
||||
"CLMUL",
|
||||
"CLMULH",
|
||||
"CLMULR",
|
||||
"CLUI",
|
||||
"CLW",
|
||||
"CLWSP",
|
||||
@@ -95,13 +98,20 @@ var riscvGeneratedInstrs = []string{
|
||||
"CSDSP",
|
||||
"CSLLI",
|
||||
"CSRAI",
|
||||
"CSRC",
|
||||
"CSRCI",
|
||||
"CSRLI",
|
||||
"CSRR",
|
||||
"CSRRC",
|
||||
"CSRRCI",
|
||||
"CSRRS",
|
||||
"CSRRSI",
|
||||
"CSRRW",
|
||||
"CSRRWI",
|
||||
"CSRS",
|
||||
"CSRSI",
|
||||
"CSRW",
|
||||
"CSRWI",
|
||||
"CSUB",
|
||||
"CSUBW",
|
||||
"CSW",
|
||||
@@ -259,6 +269,7 @@ var riscvGeneratedInstrs = []string{
|
||||
"ORCB",
|
||||
"ORI",
|
||||
"ORN",
|
||||
"PAUSE",
|
||||
"RDCYCLE",
|
||||
"RDINSTRET",
|
||||
"RDTIME",
|
||||
@@ -322,6 +333,8 @@ var riscvGeneratedInstrs = []string{
|
||||
"VADDVI",
|
||||
"VADDVV",
|
||||
"VADDVX",
|
||||
"VANDNVV",
|
||||
"VANDNVX",
|
||||
"VANDVI",
|
||||
"VANDVV",
|
||||
"VANDVX",
|
||||
@@ -329,8 +342,17 @@ var riscvGeneratedInstrs = []string{
|
||||
"VASUBUVX",
|
||||
"VASUBVV",
|
||||
"VASUBVX",
|
||||
"VBREV8V",
|
||||
"VBREVV",
|
||||
"VCLMULHVV",
|
||||
"VCLMULHVX",
|
||||
"VCLMULVV",
|
||||
"VCLMULVX",
|
||||
"VCLZV",
|
||||
"VCOMPRESSVM",
|
||||
"VCPOPM",
|
||||
"VCPOPV",
|
||||
"VCTZV",
|
||||
"VDIVUVV",
|
||||
"VDIVUVX",
|
||||
"VDIVVV",
|
||||
@@ -743,10 +765,16 @@ var riscvGeneratedInstrs = []string{
|
||||
"VREMUVX",
|
||||
"VREMVV",
|
||||
"VREMVX",
|
||||
"VREV8V",
|
||||
"VRGATHEREI16VV",
|
||||
"VRGATHERVI",
|
||||
"VRGATHERVV",
|
||||
"VRGATHERVX",
|
||||
"VROLVV",
|
||||
"VROLVX",
|
||||
"VRORVI",
|
||||
"VRORVV",
|
||||
"VRORVX",
|
||||
"VRSUBVI",
|
||||
"VRSUBVX",
|
||||
"VS1RV",
|
||||
@@ -950,6 +978,9 @@ var riscvGeneratedInstrs = []string{
|
||||
"VWMULVX",
|
||||
"VWREDSUMUVS",
|
||||
"VWREDSUMVS",
|
||||
"VWSLLVI",
|
||||
"VWSLLVV",
|
||||
"VWSLLVX",
|
||||
"VWSUBUVV",
|
||||
"VWSUBUVX",
|
||||
"VWSUBUWV",
|
||||
|
||||
@@ -245,3 +245,104 @@ func main() {
|
||||
t.Error("binary does not contain expected symbol")
|
||||
}
|
||||
}
|
||||
|
||||
// TestGOObjectAARCH64DataSymbolLink does for symbol-valued DATA fields what
|
||||
// the rt0 files do ("DATA _rt0…lib+0(SB)/8, $_rt0…lib(SB)"): the gasm object
|
||||
// carries an R_ADDR against the file's own TEXT symbol, the toolchain links
|
||||
// it, and the binary is checked for the symbol (no arm64 host to run it).
|
||||
func TestGOObjectAARCH64DataSymbolLink(t *testing.T) {
|
||||
goBin, err := exec.LookPath("go")
|
||||
if err != nil {
|
||||
t.Skip("no Go toolchain available")
|
||||
}
|
||||
dir := t.TempDir()
|
||||
asmSrc := `#include "textflag.h"
|
||||
GLOBL entry(SB), NOPTR, $8
|
||||
DATA entry+0(SB)/8, $·keepme(SB)
|
||||
|
||||
TEXT ·keepme(SB), NOSPLIT, $0-0
|
||||
RET
|
||||
|
||||
TEXT ·entryptr(SB), NOSPLIT, $0-8
|
||||
MOVD entry+0(SB), R4
|
||||
MOVD R4, ret+0(FP)
|
||||
RET
|
||||
`
|
||||
if err := os.WriteFile(filepath.Join(dir, "main_arm64.s"), []byte(asmSrc), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
mainSrc := `package main
|
||||
|
||||
func keepme()
|
||||
func entryptr() uintptr
|
||||
|
||||
func main() {
|
||||
if entryptr() == 0 {
|
||||
panic("the entry word is empty")
|
||||
}
|
||||
}
|
||||
`
|
||||
if err := os.WriteFile(filepath.Join(dir, "main.go"), []byte(mainSrc), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := os.WriteFile(filepath.Join(dir, "go.mod"), []byte("module a64dlink\n\ngo 1.21\n"), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
build := exec.Command(goBin, "build", "-x", "-work", "-o", filepath.Join(dir, "prog"), ".")
|
||||
build.Dir = dir
|
||||
build.Env = append(os.Environ(), "GOARCH=arm64")
|
||||
buildLog, err := build.CombinedOutput()
|
||||
if err != nil {
|
||||
t.Fatalf("baseline build: %v\n%s", err, buildLog)
|
||||
}
|
||||
var work, linkLine, asmObj string
|
||||
for line := range strings.SplitSeq(string(buildLog), "\n") {
|
||||
switch {
|
||||
case strings.HasPrefix(line, "WORK="):
|
||||
work = strings.TrimPrefix(line, "WORK=")
|
||||
case strings.Contains(line, "/asm ") && strings.Contains(line, "main_arm64.s") && !strings.Contains(line, "-gensymabis"):
|
||||
asmObj = fieldAfter(line, "-o")
|
||||
case strings.Contains(line, "/link ") && strings.Contains(line, "-importcfg"):
|
||||
linkLine = line
|
||||
}
|
||||
}
|
||||
if work == "" || asmObj == "" || linkLine == "" {
|
||||
t.Skipf("could not parse build log (work=%q asmObj=%q link=%q)", work, asmObj, linkLine)
|
||||
}
|
||||
defer os.RemoveAll(work)
|
||||
asmObj = strings.ReplaceAll(asmObj, "$WORK", work)
|
||||
linkLine = strings.ReplaceAll(linkLine, "$WORK", work)
|
||||
|
||||
src, err := os.ReadFile(filepath.Join(dir, "main_arm64.s"))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
f, errs := parser.Parse("main_arm64.s", string(src))
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
img, err := AssembleFileARM64(f)
|
||||
if err != nil {
|
||||
t.Fatalf("AssembleFileARM64: %v", err)
|
||||
}
|
||||
gasmObj, err := img.GOObjectAARCH64("a64dlink", "main_arm64.s")
|
||||
if err != nil {
|
||||
t.Fatalf("GOObjectAARCH64: %v", err)
|
||||
}
|
||||
if err := os.WriteFile(asmObj, gasmObj, 0o644); err != nil {
|
||||
t.Fatalf("write gasm object: %v", err)
|
||||
}
|
||||
linkCmd := exec.Command("bash", "-c", "cd "+dir+" && "+linkLine)
|
||||
linkCmd.Env = append(os.Environ(), "GOARCH=arm64")
|
||||
if out, err := linkCmd.CombinedOutput(); err != nil {
|
||||
t.Fatalf("re-link with gasm object: %v\n%s", err, out)
|
||||
}
|
||||
binData, err := os.ReadFile(filepath.Join(dir, "prog"))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if !strings.Contains(string(binData), "keepme") {
|
||||
t.Error("binary does not contain the keepme symbol")
|
||||
}
|
||||
}
|
||||
|
||||
+1698
-129
File diff suppressed because it is too large
Load Diff
+369
-55
@@ -29,6 +29,7 @@ package asm
|
||||
|
||||
import (
|
||||
"maps"
|
||||
"math/bits"
|
||||
"strconv"
|
||||
"strings"
|
||||
|
||||
@@ -78,6 +79,11 @@ func arm64RegNum(name string) int {
|
||||
return 17
|
||||
case "R18":
|
||||
return 18
|
||||
case "R18_PLATFORM":
|
||||
// The toolchain's Windows spelling: R18 is renamed R18_PLATFORM in
|
||||
// cmd/asm/internal/arch so assembly cannot use it by accident, and
|
||||
// sys_windows_arm64.s references it only through this name.
|
||||
return 18
|
||||
case "R19":
|
||||
return 19
|
||||
case "R20":
|
||||
@@ -159,6 +165,67 @@ func a64MoveWide(sf, opc, hw, imm16, rd uint32) uint32 {
|
||||
return sf<<31 | opc<<29 | 0x25<<23 | hw<<21 | imm16<<5 | rd
|
||||
}
|
||||
|
||||
// ---- logical immediate ----
|
||||
|
||||
// a64LogicalImm encodes v as the AArch64 logical (bitmask) immediate for the
|
||||
// given lane width (32 or 64): it returns the N, immr and imms fields of the
|
||||
// imm13 encoding. The algorithm mirrors cmd/internal/obj/arm64's
|
||||
// encodeLogicalImmArrEncoding: replicate the value, shrink it to the smallest
|
||||
// repeating element, find the run of ones and its rotation. ok is false when
|
||||
// v is not expressible (all zeros, all ones, or not a single cyclic run).
|
||||
func a64LogicalImm(v int64, width int) (n, immr, imms uint32, ok bool) {
|
||||
u := uint64(v)
|
||||
if width == 32 {
|
||||
u &= 0xFFFFFFFF
|
||||
}
|
||||
size := uint64(width)
|
||||
mask := ^uint64(0)
|
||||
if size < 64 {
|
||||
mask = uint64(1)<<size - 1
|
||||
}
|
||||
u &= mask
|
||||
// All zeros and all ones are MOV territory, not bitmask immediates.
|
||||
if u == 0 || u == mask {
|
||||
return 0, 0, 0, false
|
||||
}
|
||||
// Shrink to the smallest repeating element.
|
||||
for size > 2 {
|
||||
half := size / 2
|
||||
hm := uint64(1)<<half - 1
|
||||
if u&hm == u>>half&hm {
|
||||
size = half
|
||||
u &= hm
|
||||
} else {
|
||||
break
|
||||
}
|
||||
}
|
||||
ones := bits.OnesCount64(u)
|
||||
// Find the right-rotation that lays the ones out contiguously at the
|
||||
// bottom of the element; the hardware applies the inverse rotation.
|
||||
em := uint64(1)<<size - 1
|
||||
expected := uint64(1)<<ones - 1
|
||||
rot := -1
|
||||
for r := 0; r < int(size); r++ {
|
||||
rotated := u>>r | u<<(int(size)-r)
|
||||
if size < 64 {
|
||||
rotated &= em
|
||||
}
|
||||
if rotated == expected {
|
||||
rot = r
|
||||
break
|
||||
}
|
||||
}
|
||||
if rot < 0 {
|
||||
return 0, 0, 0, false
|
||||
}
|
||||
if size == 64 {
|
||||
n = 1
|
||||
}
|
||||
immr = uint32((int(size) - rot) % int(size))
|
||||
imms = ^uint32(uint32(size*2-1))&0x3F | uint32(ones-1)
|
||||
return n, immr, imms, true
|
||||
}
|
||||
|
||||
// ---- load/store (unsigned immediate, scaled) ----
|
||||
|
||||
// a64LSU encodes a load/store register (unsigned immediate, scaled):
|
||||
@@ -246,6 +313,8 @@ const (
|
||||
a64CondLT = 0xb
|
||||
a64CondGT = 0xc
|
||||
a64CondLE = 0xd
|
||||
a64CondAL = 0xe
|
||||
a64CondNV = 0xf
|
||||
)
|
||||
|
||||
// arm64CondMap maps Go assembler condition mnemonics to AArch64 condition codes.
|
||||
@@ -266,6 +335,8 @@ var arm64CondMap = map[string]uint32{
|
||||
"LT": a64CondLT,
|
||||
"GT": a64CondGT,
|
||||
"LE": a64CondLE,
|
||||
"AL": a64CondAL,
|
||||
"NV": a64CondNV,
|
||||
}
|
||||
|
||||
// ---- instruction format tags ----
|
||||
@@ -273,46 +344,47 @@ var arm64CondMap = map[string]uint32{
|
||||
type a64Format uint8
|
||||
|
||||
const (
|
||||
a64FDPSR a64Format = iota // data-processing (shifted register): ADD, SUB, AND, ORR, EOR, etc.
|
||||
a64FMovWide // move wide: MOVZ, MOVN, MOVK
|
||||
a64FBranch // unconditional branch (B/BL)
|
||||
a64FBranchCond // conditional branch (B.cond)
|
||||
a64FUncondBranch // unconditional branch register (BR/BLR/RET)
|
||||
a64FADR // ADR/ADRP
|
||||
a64FEXTR // EXTR
|
||||
a64FBitfield // bitfield: BFI/BFXIL/SBFM/UBFM/BFM
|
||||
a64FShift // shifts: LSL/LSR/ASR alias SBFM/UBFM, ROR aliases EXTR; register forms are two-source
|
||||
a64FDPR4 // data-processing 4-register: MADD/MSUB, Ra in bits 14:10
|
||||
a64FFP3 // FP 3-operand (Rm, Rn, Rd): FADD, FSUB, FMUL, FDIV, etc.
|
||||
a64FFPUnary // FP unary (Rn, Rd): FMOV, FABS, FNEG, FSQRT, FCVT, FRINT*
|
||||
a64FFP4 // FP 4-operand FMA (Ra, Rm, Rn, Rd): FMADD, FMSUB, etc.
|
||||
a64FFPCmp // FP compare (Rm, Rn): FCMP, FCMPE
|
||||
a64FFPCCmp // FP conditional compare (Rm, Rn, nzcv, cond): FCCMP, FCCMPE
|
||||
a64FFPCvt // FP↔integer conversion: FCVTZS, SCVTF, etc.
|
||||
a64FFPSel // FP conditional select (Rm, Rn, Rd, cond): FCSEL
|
||||
a64FCRC32 // CRC32
|
||||
a64FCSEL // conditional select: CSEL, CSINC, CSINV, CSNEG
|
||||
a64FExcl // exclusive load/store: LDXR, STXR, LDAXR, STLXR and pair forms LDXP, STXP
|
||||
a64FLSE // LSE atomics: LDADD, CAS, SWP
|
||||
a64FDP1 // data-processing (1 source): RBIT, REV, CLZ, CLS
|
||||
a64FBitfield2 // bitfield extract: UBFX, SBFX and the W forms
|
||||
a64FCondCmp // conditional compare: CCMP, CCMN
|
||||
a64FBranch19 // compare-and-branch: CBZ, CBNZ and the W forms
|
||||
a64FTestBranch // test-and-branch: TBZ, TBNZ and the W forms
|
||||
a64FPair // load/store pair: LDP, STP, LDPW, STPW, FLDPD, FSTPD
|
||||
a64FAcqRel // acquire/release: LDAR family, STLR family
|
||||
a64FSys // system: BRK, SVC, DMB, DSB, ISB, DC, MRS, MSR, PRFM
|
||||
a64FCrypto2 // crypto 2-register: AESD, AESE, AESIMC, AESMC, SHA1H, ...
|
||||
a64FCrypto3 // crypto 3-register: SHA1C, SHA256H, SHA512SU1, ...
|
||||
a64FSIMDV // SIMD 3-register with arrangement: VADD, VAND, VCMEQ, VZIP1, ...
|
||||
a64FSIMDVZero // SIMD compare against zero: VCMEQ $0, Vn, Vd
|
||||
a64FSIMDV2 // SIMD 2-register with arrangement: VREV32, VREV64, VUADDLV, VMOV
|
||||
a64FSIMDV4 // SIMD 4-register / imm 3-register: VEOR3, VBCAX, VXAR, VEXT
|
||||
a64FVTBL // SIMD table lookup: VTBL
|
||||
a64FDUP // SIMD element moves: VDUP, VMOV with element indices
|
||||
a64FVLDST // SIMD structure loads/stores: VLD1, VST1, VLD1R, VLD4R
|
||||
a64FShiftImm // SIMD shift by immediate: VSHL, VUSHR, VSRI
|
||||
a64FMoviLit // VMOVS/VMOVD/VMOVQ with a large constant (literal pool)
|
||||
a64FDPSR a64Format = iota // data-processing (shifted register): ADD, SUB, AND, ORR, EOR, etc.
|
||||
a64FMovWide // move wide: MOVZ, MOVN, MOVK
|
||||
a64FBranch // unconditional branch (B/BL)
|
||||
a64FBranchCond // conditional branch (B.cond)
|
||||
a64FUncondBranch // unconditional branch register (BR/BLR/RET)
|
||||
a64FADR // ADR/ADRP
|
||||
a64FEXTR // EXTR
|
||||
a64FBitfield // bitfield: BFI/BFXIL/SBFM/UBFM/BFM
|
||||
a64FBitfieldAlias // bitfield alias: BFI/BFXIL/SBFIZ/UBFIZ, ($lsb, Rn, $width, Rd)
|
||||
a64FShift // shifts: LSL/LSR/ASR alias SBFM/UBFM, ROR aliases EXTR; register forms are two-source
|
||||
a64FDPR4 // data-processing 4-register: MADD/MSUB, Ra in bits 14:10
|
||||
a64FFP3 // FP 3-operand (Rm, Rn, Rd): FADD, FSUB, FMUL, FDIV, etc.
|
||||
a64FFPUnary // FP unary (Rn, Rd): FMOV, FABS, FNEG, FSQRT, FCVT, FRINT*
|
||||
a64FFP4 // FP 4-operand FMA (Ra, Rm, Rn, Rd): FMADD, FMSUB, etc.
|
||||
a64FFPCmp // FP compare (Rm, Rn): FCMP, FCMPE
|
||||
a64FFPCCmp // FP conditional compare (Rm, Rn, nzcv, cond): FCCMP, FCCMPE
|
||||
a64FFPCvt // FP↔integer conversion: FCVTZS, SCVTF, etc.
|
||||
a64FFPSel // FP conditional select (Rm, Rn, Rd, cond): FCSEL
|
||||
a64FCRC32 // CRC32
|
||||
a64FCSEL // conditional select: CSEL, CSINC, CSINV, CSNEG
|
||||
a64FExcl // exclusive load/store: LDXR, STXR, LDAXR, STLXR and pair forms LDXP, STXP
|
||||
a64FLSE // LSE atomics: LDADD, CAS, SWP
|
||||
a64FDP1 // data-processing (1 source): RBIT, REV, CLZ, CLS
|
||||
a64FBitfield2 // bitfield extract: UBFX, SBFX and the W forms
|
||||
a64FCondCmp // conditional compare: CCMP, CCMN
|
||||
a64FBranch19 // compare-and-branch: CBZ, CBNZ and the W forms
|
||||
a64FTestBranch // test-and-branch: TBZ, TBNZ and the W forms
|
||||
a64FPair // load/store pair: LDP, STP, LDPW, STPW, FLDPD, FSTPD
|
||||
a64FAcqRel // acquire/release: LDAR family, STLR family
|
||||
a64FSys // system: BRK, SVC, DMB, DSB, ISB, DC, MRS, MSR, PRFM
|
||||
a64FCrypto2 // crypto 2-register: AESD, AESE, AESIMC, AESMC, SHA1H, ...
|
||||
a64FCrypto3 // crypto 3-register: SHA1C, SHA256H, SHA512SU1, ...
|
||||
a64FSIMDV // SIMD 3-register with arrangement: VADD, VAND, VCMEQ, VZIP1, ...
|
||||
a64FSIMDVZero // SIMD compare against zero: VCMEQ $0, Vn, Vd
|
||||
a64FSIMDV2 // SIMD 2-register with arrangement: VREV32, VREV64, VUADDLV, VMOV
|
||||
a64FSIMDV4 // SIMD 4-register / imm 3-register: VEOR3, VBCAX, VXAR, VEXT
|
||||
a64FVTBL // SIMD table lookup: VTBL
|
||||
a64FDUP // SIMD element moves: VDUP, VMOV with element indices
|
||||
a64FVLDST // SIMD structure loads/stores: VLD1, VST1, VLD1R, VLD4R
|
||||
a64FShiftImm // SIMD shift by immediate: VSHL, VUSHR, VSRI
|
||||
a64FMoviLit // VMOVS/VMOVD/VMOVQ with a large constant (literal pool)
|
||||
)
|
||||
|
||||
// a64Enc is one instruction's encoding: its bit layout (format) and the
|
||||
@@ -425,6 +497,16 @@ func init() {
|
||||
a64InstrTable["MADDW"] = a64Enc{format: a64FDPR4, op: 0<<31 | 0x1b<<24}
|
||||
a64InstrTable["MSUB"] = a64Enc{format: a64FDPR4, op: 1<<31 | 0x1b<<24 | 1<<15}
|
||||
a64InstrTable["MSUBW"] = a64Enc{format: a64FDPR4, op: 0<<31 | 0x1b<<24 | 1<<15}
|
||||
// The widening multiplies: a 64-bit result riding the same layout, the
|
||||
// three-operand forms reading the accumulate register as ZR.
|
||||
a64InstrTable["SMADDL"] = a64Enc{format: a64FDPR4, op: 1<<31 | 0x1b<<24 | 1<<21}
|
||||
a64InstrTable["UMADDL"] = a64Enc{format: a64FDPR4, op: 1<<31 | 0x1b<<24 | 1<<21 | 1<<23}
|
||||
a64InstrTable["SMSUBL"] = a64Enc{format: a64FDPR4, op: 1<<31 | 0x1b<<24 | 1<<21 | 1<<15}
|
||||
a64InstrTable["UMSUBL"] = a64Enc{format: a64FDPR4, op: 1<<31 | 0x1b<<24 | 1<<21 | 1<<23 | 1<<15}
|
||||
a64InstrTable["SMULL"] = a64Enc{format: a64FDPR4, op: 1<<31 | 0x1b<<24 | 1<<21 | 31<<10}
|
||||
a64InstrTable["UMULL"] = a64Enc{format: a64FDPR4, op: 1<<31 | 0x1b<<24 | 1<<21 | 1<<23 | 31<<10}
|
||||
a64InstrTable["SMNEGL"] = a64Enc{format: a64FDPR4, op: 1<<31 | 0x1b<<24 | 1<<21 | 1<<15 | 31<<10}
|
||||
a64InstrTable["UMNEGL"] = a64Enc{format: a64FDPR4, op: 1<<31 | 0x1b<<24 | 1<<21 | 1<<23 | 1<<15 | 31<<10}
|
||||
|
||||
// ---- move wide ----
|
||||
// MOVZ/MOVN/MOVK
|
||||
@@ -474,6 +556,15 @@ func init() {
|
||||
// ---- bitfield ----
|
||||
a64InstrTable["BFM"] = a64Enc{format: a64FBitfield, op: 1<<31 | 1<<29 | 0x26<<23 | 1<<22}
|
||||
a64InstrTable["BFMW"] = a64Enc{format: a64FBitfield, op: 0<<31 | 1<<29 | 0x26<<23 | 0<<22}
|
||||
// The four-operand bitfield aliases: ($lsb, Rn, $width, Rd).
|
||||
a64InstrTable["BFI"] = a64Enc{format: a64FBitfieldAlias, op: 1<<31 | 1<<29 | 0x26<<23 | 1<<22}
|
||||
a64InstrTable["BFIW"] = a64Enc{format: a64FBitfieldAlias, op: 0<<31 | 1<<29 | 0x26<<23}
|
||||
a64InstrTable["BFXIL"] = a64Enc{format: a64FBitfieldAlias, op: 1<<31 | 1<<29 | 0x26<<23 | 1<<22}
|
||||
a64InstrTable["BFXILW"] = a64Enc{format: a64FBitfieldAlias, op: 0<<31 | 1<<29 | 0x26<<23}
|
||||
a64InstrTable["SBFIZ"] = a64Enc{format: a64FBitfieldAlias, op: 0x93400000}
|
||||
a64InstrTable["SBFIZW"] = a64Enc{format: a64FBitfieldAlias, op: 0x13000000}
|
||||
a64InstrTable["UBFIZ"] = a64Enc{format: a64FBitfieldAlias, op: 0x53000000}
|
||||
a64InstrTable["UBFIZW"] = a64Enc{format: a64FBitfieldAlias, op: 0x33000000}
|
||||
a64InstrTable["SBFM"] = a64Enc{format: a64FBitfield, op: 1<<31 | 0<<29 | 0x26<<23 | 1<<22}
|
||||
a64InstrTable["SBFMW"] = a64Enc{format: a64FBitfield, op: 0<<31 | 0<<29 | 0x26<<23 | 0<<22}
|
||||
a64InstrTable["UBFM"] = a64Enc{format: a64FBitfield, op: 1<<31 | 2<<29 | 0x26<<23 | 1<<22}
|
||||
@@ -654,6 +745,13 @@ func init() {
|
||||
"RBIT": 0xdac00000, "REV16": 0xdac00400, "REV32": 0xdac00800,
|
||||
"REV": 0xdac00c00, "CLZ": 0xdac01000, "CLS": 0xdac01400,
|
||||
"RBITW": 0x5ac00000, "REVW": 0x5ac00800, "CLZW": 0x5ac01000, "CLSW": 0x5ac01400,
|
||||
// Extend and byte-reverse: the UBFM/SBFM aliases with imms fixing
|
||||
// the source width.
|
||||
"SXTB": 0x93401c00, "SXTBW": 0x13001c00, "SXTH": 0x93403c00,
|
||||
"SXTHW": 0x13003c00, "SXTW": 0x93407c00,
|
||||
"UXTB": 0x53001c00, "UXTBW": 0x53001c00, "UXTH": 0x53403c00,
|
||||
"UXTHW": 0x53003c00, "UXTW": 0x53407c00,
|
||||
"REV16W": 0x5ac00400,
|
||||
}
|
||||
for m, op := range dp1 {
|
||||
a64InstrTable[m] = a64Enc{format: a64FDP1, op: op}
|
||||
@@ -672,7 +770,7 @@ func init() {
|
||||
a64InstrTable["CCMNW"] = a64Enc{format: a64FCondCmp, op: 0x3a400000}
|
||||
|
||||
// ---- system operations ----
|
||||
for _, m := range []string{"BRK", "SVC", "DMB", "DSB", "ISB", "DC", "MRS", "MSR", "PRFM"} {
|
||||
for _, m := range []string{"BRK", "SVC", "DMB", "DSB", "ISB", "CLREX", "HINT", "BTI", "HLT", "SMC", "HVC", "DCPS1", "DCPS2", "DCPS3", "DRPS", "ERET", "AUTIASP", "AUTIBSP", "AUTIA1716", "AUTIB1716", "SEVL", "SEV", "WFE", "WFI", "YIELD", "DC", "MRS", "MSR", "PRFM"} {
|
||||
a64InstrTable[m] = a64Enc{format: a64FSys}
|
||||
}
|
||||
|
||||
@@ -723,6 +821,73 @@ func init() {
|
||||
for m, op := range lse {
|
||||
a64InstrTable[m] = a64Enc{format: a64FLSE, op: op}
|
||||
}
|
||||
// The remaining width and ordering spellings of the same shapes, and the
|
||||
// CAS compare-and-swap family, word-verified against go tool asm.
|
||||
lseMore := map[string]uint32{
|
||||
"LDADDAB": 0x38a00000,
|
||||
"LDADDAH": 0x78a00000,
|
||||
"LDADDALB": 0x38e00000,
|
||||
"LDADDALH": 0x78e00000,
|
||||
"LDADDLB": 0x38600000,
|
||||
"LDADDLD": 0xf8600000,
|
||||
"LDADDLH": 0x78600000,
|
||||
"LDADDLW": 0xb8600000,
|
||||
"LDCLRAB": 0x38a01000,
|
||||
"LDCLRAH": 0x78a01000,
|
||||
"LDCLRALH": 0x78e01000,
|
||||
"LDCLRB": 0x38201000,
|
||||
"LDCLRD": 0xf8201000,
|
||||
"LDCLRH": 0x78201000,
|
||||
"LDCLRLB": 0x38601000,
|
||||
"LDCLRLD": 0xf8601000,
|
||||
"LDCLRLH": 0x78601000,
|
||||
"LDCLRLW": 0xb8601000,
|
||||
"LDCLRW": 0xb8201000,
|
||||
"LDEORAB": 0x38a02000,
|
||||
"LDEORAD": 0xf8a02000,
|
||||
"LDEORAH": 0x78a02000,
|
||||
"LDEORALB": 0x38e02000,
|
||||
"LDEORALH": 0x78e02000,
|
||||
"LDEORAW": 0xb8a02000,
|
||||
"LDEORB": 0x38202000,
|
||||
"LDEORD": 0xf8202000,
|
||||
"LDEORH": 0x78202000,
|
||||
"LDEORLB": 0x38602000,
|
||||
"LDEORLD": 0xf8602000,
|
||||
"LDEORLH": 0x78602000,
|
||||
"LDEORLW": 0xb8602000,
|
||||
"LDEORW": 0xb8202000,
|
||||
"LDORAB": 0x38a03000,
|
||||
"LDORAD": 0xf8a03000,
|
||||
"LDORAH": 0x78a03000,
|
||||
"LDORALH": 0x78e03000,
|
||||
"LDORAW": 0xb8a03000,
|
||||
"LDORB": 0x38203000,
|
||||
"LDORD": 0xf8203000,
|
||||
"LDORH": 0x78203000,
|
||||
"LDORLB": 0x38603000,
|
||||
"LDORLD": 0xf8603000,
|
||||
"LDORLH": 0x78603000,
|
||||
"LDORLW": 0xb8603000,
|
||||
"LDORW": 0xb8203000,
|
||||
"SWPAB": 0x38a08000,
|
||||
"SWPAD": 0xf8a08000,
|
||||
"SWPAH": 0x78a08000,
|
||||
"SWPALH": 0x78e08000,
|
||||
"SWPAW": 0xb8a08000,
|
||||
"SWPB": 0x38208000,
|
||||
"SWPH": 0x78208000,
|
||||
"SWPLB": 0x38608000,
|
||||
"SWPLD": 0xf8608000,
|
||||
"SWPLH": 0x78608000,
|
||||
"SWPLW": 0xb8608000,
|
||||
"CASAD": 0xc8e07c00,
|
||||
"CASALB": 0x08e0fc00,
|
||||
"CASLW": 0x88a0fc00,
|
||||
}
|
||||
for m, op := range lseMore {
|
||||
a64InstrTable[m] = a64Enc{format: a64FLSE, op: op}
|
||||
}
|
||||
|
||||
// ---- carry-setting/carry-using arithmetic and widening multiply ----
|
||||
// MUL and SMULH/UMULH are the MADD/MSUB layout with the accumulate
|
||||
@@ -732,7 +897,12 @@ func init() {
|
||||
"ADCS": 0xba000000, "ADCSW": 0x3a000000,
|
||||
"SBC": 0xda000000, "SBCW": 0x5a000000,
|
||||
"SBCS": 0xfa000000, "SBCSW": 0x7a000000,
|
||||
"MUL": 0x9b007c00, "MULW": 0x1b007c00,
|
||||
// MNEG/MSUB and NGC/SBC with the complementing register preset to ZR.
|
||||
"MNEG": 0x9b00fc00, "MNEGW": 0x1b00fc00,
|
||||
"NGC": 0xda000000, "NGCW": 0x5a000000,
|
||||
"NGCS": 0xfa000000, "NGCSW": 0x7a000000,
|
||||
"NEGSW": 0x6b000000,
|
||||
"MUL": 0x9b007c00, "MULW": 0x1b007c00,
|
||||
"SMULH": 0x9b407c00, "UMULH": 0x9bc07c00,
|
||||
}
|
||||
for m, op := range dpsrExtra {
|
||||
@@ -773,12 +943,20 @@ func init() {
|
||||
a64InstrTable["VSHL"] = a64Enc{format: a64FShiftImm, op: 0x0f000000 | 21<<10}
|
||||
a64InstrTable["VUSHR"] = a64Enc{format: a64FShiftImm, op: 0x2f000000 | 1<<10}
|
||||
a64InstrTable["VSRI"] = a64Enc{format: a64FShiftImm, op: 0x2f000000 | 17<<10}
|
||||
a64InstrTable["VSSHR"] = a64Enc{format: a64FShiftImm, op: 0x0f000000 | 1<<10}
|
||||
a64InstrTable["VSRA"] = a64Enc{format: a64FShiftImm, op: 0x0f000000 | 17<<10}
|
||||
a64InstrTable["VSRSHR"] = a64Enc{format: a64FShiftImm, op: 0x0f000000 | 9<<10}
|
||||
a64InstrTable["VSLI"] = a64Enc{format: a64FShiftImm, op: 0x2f000000 | 21<<10}
|
||||
a64InstrTable["VSQSHL"] = a64Enc{format: a64FShiftImm, op: 0x0f000000 | 29<<10}
|
||||
a64InstrTable["VUQSHL"] = a64Enc{format: a64FShiftImm, op: 0x2f000000 | 29<<10}
|
||||
a64InstrTable["VLD1"] = a64Enc{format: a64FVLDST}
|
||||
a64InstrTable["VLD1.P"] = a64Enc{format: a64FVLDST, op: 1}
|
||||
a64InstrTable["VST1"] = a64Enc{format: a64FVLDST}
|
||||
a64InstrTable["VST1.P"] = a64Enc{format: a64FVLDST, op: 1}
|
||||
a64InstrTable["VLD1R"] = a64Enc{format: a64FVLDST}
|
||||
a64InstrTable["VLD1R.P"] = a64Enc{format: a64FVLDST, op: 1}
|
||||
a64InstrTable["VLD4R"] = a64Enc{format: a64FVLDST}
|
||||
a64InstrTable["VLD4R.P"] = a64Enc{format: a64FVLDST, op: 1}
|
||||
}
|
||||
|
||||
// a64SimdVSpec is one arrangement-aware SIMD instruction: the 8B base word,
|
||||
@@ -835,6 +1013,23 @@ func a64ElemLetter(s string) bool {
|
||||
return false
|
||||
}
|
||||
|
||||
// fpSimdArrs and fpAcrossArrs bound the arrangements the FP SIMD forms
|
||||
// accept: H, S and D widths for the pairwise data-processing, H and S for
|
||||
// the across-vector reductions.
|
||||
var fpSimdArrs = uint16(1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D)
|
||||
var fpAcrossArrs = uint16(1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S)
|
||||
|
||||
// a64SimdQOnly names the forms whose arrangement contributes the 128-bit
|
||||
// flag alone, without the size bits: the FP converts, the FP round-to-integral
|
||||
// and pairwise compares among them. Word-verified against go tool asm.
|
||||
var a64SimdQOnly = map[string]bool{
|
||||
"VSCVTF": true, "VUCVTF": true, "VFCVTZS": true, "VFCVTZU": true,
|
||||
"VFABS": true, "VFNEG": true, "VFSQRT": true,
|
||||
"VFRINTN": true, "VFRINTP": true, "VFRINTM": true, "VFRINTZ": true,
|
||||
"VFADDP": true, "VFMAXP": true, "VFMAXNMP": true,
|
||||
"VFMAXV": true, "VFMAXNMV": true,
|
||||
}
|
||||
|
||||
// a64ArrBits carries the fixed bits an arrangement contributes to the
|
||||
// three-same word shape: the element size at bits 23:22 and the 128-bit
|
||||
// flag at bit 30. Bit 29 belongs to the instruction's own base.
|
||||
@@ -854,19 +1049,96 @@ var a64ArrBits = [a64ArrCount]uint32{
|
||||
// instructions (word = base | arrBits | Rm<<16 | Rn<<5 | Rd). Every base
|
||||
// word and arrangement bit was read off go tool asm.
|
||||
var a64SimdVTable = map[string]a64SimdVSpec{
|
||||
"VADD": {0x0e208400, 0x7f, false},
|
||||
"VSUB": {0x2e208400, 0x7f, false},
|
||||
"VMUL": {0x0e209c00, 0x3f, false}, // no 2D: integer multiply stops at 4S
|
||||
"VAND": {0x0e201c00, 0x03, false}, // logical ops accept 8B and 16B only
|
||||
"VEOR": {0x2e201c00, 0x03, false},
|
||||
"VORR": {0x0ea01c00, 0x03, false},
|
||||
"VADDP": {0x0e20bc00, 0x7f, false},
|
||||
"VZIP1": {0x0e003800, 0x7f, false},
|
||||
"VZIP2": {0x0e007800, 0x7f, false},
|
||||
"VCMEQ": {0x2e208c00, 0x7f, false},
|
||||
"VRAX1": {0xce608c00, 1 << a64Arr2D, true}, // SHA3 group, D2 only
|
||||
"VPMULL": {0x0e20e000, 1<<a64Arr8B | 1<<a64ArrD1, false},
|
||||
"VPMULL2": {0x0e20e000, 1<<a64Arr16B | 1<<a64Arr2D, false},
|
||||
"VADD": {0x0e208400, 0x7f, false},
|
||||
"VSUB": {0x2e208400, 0x7f, false},
|
||||
"VMUL": {0x0e209c00, 0x3f, false}, // no 2D: integer multiply stops at 4S
|
||||
"VAND": {0x0e201c00, 0x03, false}, // logical ops accept 8B and 16B only
|
||||
"VEOR": {0x2e201c00, 0x03, false},
|
||||
"VORR": {0x0ea01c00, 0x03, false},
|
||||
"VADDP": {0x0e20bc00, 0x7f, false},
|
||||
"VZIP1": {0x0e003800, 0x7f, false},
|
||||
"VZIP2": {0x0e007800, 0x7f, false},
|
||||
"VCMEQ": {0x2e208c00, 0x7f, false},
|
||||
"VCMGE": {0x0e203c00, 0x7f, false},
|
||||
"VCMGT": {0x0e203400, 0x7f, false},
|
||||
"VCMHI": {0x2e203400, 0x7f, false},
|
||||
"VCMHS": {0x2e203c00, 0x7f, false},
|
||||
// FP compares take H, S and D arrangements only (the toolchain rejects
|
||||
// the byte forms), and VFCMLE/VFCMLT have no register form at all.
|
||||
"VFCMEQ": {0x0e20e400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
|
||||
"VFCMGE": {0x2e20e400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
|
||||
"VFCMGT": {0x2ea0e400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
|
||||
// FP arithmetic shares the same arrangement restriction.
|
||||
"VFADD": {0x0e20d400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
|
||||
"VFSUB": {0x0ea0d400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
|
||||
"VFMUL": {0x2e20dc00, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
|
||||
"VFDIV": {0x2e20fc00, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
|
||||
"VFMAX": {0x0e20f400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
|
||||
"VFMIN": {0x0ea0f400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
|
||||
"VFMAXNM": {0x0e20c400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
|
||||
"VFMINNM": {0x0ea0c400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
|
||||
"VFMLA": {0x0e20cc00, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
|
||||
"VFMLS": {0x0ea0cc00, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
|
||||
// Saturating, halving, polynomial and pairwise arithmetic, the logical
|
||||
// VBIT/VBSL family and the FP pairwise forms: word-verified against go
|
||||
// tool asm.
|
||||
"VBIC": {0x0e601c00, 0x7f, false},
|
||||
"VBIF": {0x2ee01c00, 0x7f, false},
|
||||
"VBIT": {0x6ea01c00, 0x7f, false},
|
||||
"VBSL": {0x6e601c00, 0x7f, false},
|
||||
"VCMTST": {0x0e208c00, 0x7f, false},
|
||||
"VFADDP": {0x2e20d400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
|
||||
"VFMAXP": {0x2e20f400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
|
||||
"VFMINP": {0x6ea0f400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
|
||||
"VFMAXNMP": {0x2e20c400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
|
||||
"VFMINNMP": {0x6ea0c400, 1<<a64Arr4H | 1<<a64Arr8H | 1<<a64Arr2S | 1<<a64Arr4S | 1<<a64Arr2D, false},
|
||||
"VMLA": {0x4ea09400, 0x7f, false},
|
||||
"VMLS": {0x6ea09400, 0x7f, false},
|
||||
"VORN": {0x4ee01c00, 0x7f, false},
|
||||
"VSHADD": {0x4ea00400, 0x7f, false},
|
||||
"VSRHADD": {0x4ea01400, 0x7f, false},
|
||||
"VUHADD": {0x6ea00400, 0x7f, false},
|
||||
"VURHADD": {0x6ea01400, 0x7f, false},
|
||||
"VSMAX": {0x4ea06400, 0x7f, false},
|
||||
"VSMIN": {0x4ea06c00, 0x7f, false},
|
||||
"VSMAXP": {0x4ea0a400, 0x7f, false},
|
||||
"VSMINP": {0x4ea0ac00, 0x7f, false},
|
||||
"VUMAX": {0x2e206400, 0x7f, false},
|
||||
"VUMIN": {0x2e206c00, 0x7f, false},
|
||||
"VUMAXP": {0x6ea0a400, 0x7f, false},
|
||||
"VUMINP": {0x6ea0ac00, 0x7f, false},
|
||||
"VSQADD": {0x4ea00c00, 0x7f, false},
|
||||
"VUQADD": {0x6ea00c00, 0x7f, false},
|
||||
"VSQSUB": {0x4ea02c00, 0x7f, false},
|
||||
"VUQSUB": {0x6ea02c00, 0x7f, false},
|
||||
"VSSHL": {0x4ee04400, 0x7f, false},
|
||||
"VUSHL": {0x6ee04400, 0x7f, false},
|
||||
"VUZP1": {0x0e001800, 0x7f, false},
|
||||
"VUZP2": {0x4ec05800, 0x7f, false},
|
||||
"VTRN1": {0x4ec02800, 0x7f, false},
|
||||
"VTRN2": {0x4ec06800, 0x7f, false},
|
||||
"VRAX1": {0xce608c00, 1 << a64Arr2D, true}, // SHA3 group, D2 only
|
||||
"VPMULL": {0x0e20e000, 1<<a64Arr8B | 1<<a64ArrD1, false},
|
||||
"VPMULL2": {0x0e20e000, 1<<a64Arr16B | 1<<a64Arr2D, false},
|
||||
}
|
||||
|
||||
// a64SimdVZero holds the compare-against-zero words of the SIMD compares
|
||||
// spelled with a $0 first operand (word = base | arrBits | Rn<<5 | Rd).
|
||||
// VCMHI and VCMHS have no zero form: the toolchain reports an illegal
|
||||
// combination for them, so they stay out and the encoder rejects the shape.
|
||||
var a64SimdVZero = map[string]uint32{
|
||||
"VCMEQ": 0x0e209800,
|
||||
"VCMGT": 0x0e208800,
|
||||
"VCMGE": 0x2e208800,
|
||||
"VCMLT": 0x0e20a800,
|
||||
"VCMLE": 0x2e209800,
|
||||
// FP compares against (0.0): the register forms above carry the U and op
|
||||
// bits; the zero forms reshape them.
|
||||
"VFCMEQ": 0x0ea0d800,
|
||||
"VFCMGE": 0x2ea0c800,
|
||||
"VFCMGT": 0x0ea0c800,
|
||||
"VFCMLE": 0x2ea0d800,
|
||||
"VFCMLT": 0x0ea0e800,
|
||||
}
|
||||
|
||||
// a64SimdV2Table holds the arrangement-aware two-register SIMD instructions
|
||||
@@ -875,8 +1147,41 @@ var a64SimdVTable = map[string]a64SimdVSpec{
|
||||
var a64SimdV2Table = map[string]a64SimdVSpec{
|
||||
"VREV32": {0x2e200800, 1<<a64Arr8B | 1<<a64Arr16B | 1<<a64Arr4H | 1<<a64Arr8H, false},
|
||||
"VREV64": {0x0e200800, 0x3f, false},
|
||||
"VREV16": {0x0e201800, 1<<a64Arr8B | 1<<a64Arr16B, false},
|
||||
"VUADDLV": {0x2e303800, 0x3f, false},
|
||||
"VMOV": {0x0ea01c00, 1<<a64Arr8B | 1<<a64Arr16B, false},
|
||||
// Two-register data-processing across one arrangement.
|
||||
"VABS": {0x0e20b800, 0x7f, false},
|
||||
"VNEG": {0x2e20b800, 0x7f, false},
|
||||
"VCLS": {0x0e204800, 0x7f, false},
|
||||
"VCLZ": {0x2e204800, 0x7f, false},
|
||||
"VCNT": {0x0e205800, 0x7f, false},
|
||||
"VNOT": {0x2e205800, 0x7f, false},
|
||||
"VSQABS": {0x0e207800, 0x7f, false},
|
||||
"VSQNEG": {0x2e207800, 0x7f, false},
|
||||
"VRBIT": {0x6e605800, 0x7f, false},
|
||||
"VSCVTF": {0x4e21d800, fpSimdArrs, false},
|
||||
"VUCVTF": {0x6e21d800, fpSimdArrs, false},
|
||||
"VFCVTZS": {0x4ea1b800, fpSimdArrs, false},
|
||||
"VFCVTZU": {0x6ea1b800, fpSimdArrs, false},
|
||||
"VFABS": {0x0ea0f800, fpSimdArrs, false},
|
||||
"VFNEG": {0x2ea0f800, fpSimdArrs, false},
|
||||
"VFSQRT": {0x2ea1f800, fpSimdArrs, false},
|
||||
"VFRINTN": {0x0e218800, fpSimdArrs, false},
|
||||
"VFRINTP": {0x0ea18800, fpSimdArrs, false},
|
||||
"VFRINTM": {0x0e219800, fpSimdArrs, false},
|
||||
"VFRINTZ": {0x0ea19800, fpSimdArrs, false},
|
||||
// Across-vector reductions: the operand arrangement rides as usual and
|
||||
// the destination stays a bare V register.
|
||||
"VADDV": {0x0e31b800, 0x3f, false},
|
||||
"VSMAXV": {0x0e30a800, 0x3f, false},
|
||||
"VSMINV": {0x0e31a800, 0x3f, false},
|
||||
"VUMAXV": {0x2e30a800, 0x3f, false},
|
||||
"VUMINV": {0x2e31a800, 0x3f, false},
|
||||
"VFMAXV": {0x2e30f800, fpAcrossArrs, false},
|
||||
"VFMINV": {0x2eb0f800, fpAcrossArrs, false},
|
||||
"VFMAXNMV": {0x2e30c800, fpAcrossArrs, false},
|
||||
"VFMINNMV": {0x2eb0c800, fpAcrossArrs, false},
|
||||
}
|
||||
|
||||
// a64CryptoArr is the arrangement each crypto instruction's operands must
|
||||
@@ -904,6 +1209,15 @@ var a64MRSOps = map[string]uint32{
|
||||
"ID_AA64ISAR1_EL1": 0xd5380620, "CNTFRQ_EL0": 0xd53be000,
|
||||
"CNTPCT_EL0": 0xd53be020, "CNTVCT_EL0": 0xd53be040,
|
||||
"DCZID_EL0": 0xd53b00e0, "DIT": 0xd53b42a0, "ID_AA64ZFR0_EL1": 0xd5380480,
|
||||
"NZCV": 0xd53b4200, "FPCR": 0xd53b4400, "FPSR": 0xd53b4420,
|
||||
}
|
||||
|
||||
// a64MSRRegOps maps the system register names GOROOT writes through the
|
||||
// MSR (register) form, spelled in Go assembly as MOVD Rn, <sysreg> or
|
||||
// MSR Rn, <sysreg>; the source register rides bits 4:0.
|
||||
var a64MSRRegOps = map[string]uint32{
|
||||
"NZCV": 0xd51b4200, "FPCR": 0xd51b4400, "FPSR": 0xd51b4420,
|
||||
"ELR_EL1": 0xd5184020,
|
||||
}
|
||||
|
||||
// a64MSROps maps the system register names GOROOT writes to their fixed
|
||||
|
||||
+288
-13
@@ -4,6 +4,7 @@
|
||||
package asm
|
||||
|
||||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
|
||||
@@ -104,6 +105,7 @@ func TestArm64RegNum(t *testing.T) {
|
||||
}{
|
||||
{"R0", 0}, {"R4", 4}, {"R29", 29}, {"R30", 30}, {"R31", 31},
|
||||
{"FP", 29}, {"LR", 30}, {"LINK", 30}, {"SP", 31}, {"ZR", 31},
|
||||
{"R18_PLATFORM", 18},
|
||||
{"F0", 0}, {"F4", 4}, {"F31", 31},
|
||||
{"INVALID", -1}, {"X0", -1}, {"", -1},
|
||||
}
|
||||
@@ -674,6 +676,36 @@ func TestArm64AcquireRelease(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// TestArm64BTI pins the landing-pad family against the toolchain words:
|
||||
// only the uppercase C/J/JC spellings assemble, and bare BTI is a
|
||||
// diagnostic, never a panic.
|
||||
func TestArm64BTI(t *testing.T) {
|
||||
got := arm64Words(t, "\tBTI C\n\tBTI J\n\tBTI JC\n")
|
||||
want := []uint32{
|
||||
0xd503245f, // BTI C
|
||||
0xd503249f, // BTI J
|
||||
0xd50324df, // BTI JC
|
||||
0xd65f03c0, // RET
|
||||
}
|
||||
if len(got) != len(want) {
|
||||
t.Fatalf("got %d words, want %d", len(got), len(want))
|
||||
}
|
||||
for i := range want {
|
||||
if got[i] != want[i] {
|
||||
t.Errorf("word %d = %#x, want %#x", i, got[i], want[i])
|
||||
}
|
||||
}
|
||||
for _, src := range []string{"\tBTI\n", "\tBTI c\n", "\tBTI B\n"} {
|
||||
f, errs := parser.Parse("test_arm64.s", "#include \"textflag.h\"\n\nTEXT ·f(SB), NOSPLIT, $0-0\n"+src+"\tRET\n")
|
||||
if len(errs) > 0 {
|
||||
continue
|
||||
}
|
||||
if _, err := AssembleFileARM64(f); err == nil {
|
||||
t.Errorf("BTI spelling %q should be rejected, as go tool asm rejects it", src)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestArm64System pins BRK, SVC, the barriers, cache maintenance and the
|
||||
// system register accesses.
|
||||
func TestArm64System(t *testing.T) {
|
||||
@@ -830,6 +862,99 @@ func TestArm64SIMDElement(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// TestArm64GPIntoVector pins the whole-vector moves VMOV/VDUP Rs, Vd.<T>
|
||||
// against `go tool asm -S` output (Go 1.27, arm64): word = Q | 7<<25 |
|
||||
// imm5<<16 | 3<<10 | rs<<5 | rd, shared by both mnemonics, the form
|
||||
// sys_windows_arm64.s and the bytealg loops use. The D1 destination is
|
||||
// rejected, as the toolchain rejects it.
|
||||
func TestArm64GPIntoVector(t *testing.T) {
|
||||
got := arm64Words(t, "\tVMOV R5, V5.B16\n\tVMOV R1, V2.B8\n\tVMOV R3, V4.H4\n"+
|
||||
"\tVMOV R9, V10.S4\n\tVMOV R7, V31.H8\n\tVMOV R11, V12.D2\n"+
|
||||
"\tVDUP R5, V5.B16\n\tVDUP R9, V10.H8\n\tVMOV V4.B16, V20.B16\n")
|
||||
want := []uint32{
|
||||
0x4e010ca5, // VMOV R5, V5.B16
|
||||
0x0e010c22, // VMOV R1, V2.B8
|
||||
0x0e020c64, // VMOV R3, V4.H4
|
||||
0x4e040d2a, // VMOV R9, V10.S4
|
||||
0x4e020cff, // VMOV R7, V31.H8
|
||||
0x4e080d6c, // VMOV R11, V12.D2
|
||||
0x4e010ca5, // VDUP R5, V5.B16 (same word as VMOV)
|
||||
0x4e020d2a, // VDUP R9, V10.H8
|
||||
0x4ea41c94, // VMOV V4.B16, V20.B16 (vector to vector stays ORR)
|
||||
0xd65f03c0,
|
||||
}
|
||||
if len(got) != len(want) {
|
||||
t.Fatalf("word count = %d, want %d", len(got), len(want))
|
||||
}
|
||||
for i := range want {
|
||||
if got[i] != want[i] {
|
||||
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
|
||||
}
|
||||
}
|
||||
f, errs := parser.Parse("test_arm64.s", "#include \"textflag.h\"\n\nTEXT ·f(SB), NOSPLIT, $0-0\n\tVMOV R7, V8.D1\n\tRET\n")
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
if _, err := AssembleFileARM64(f); err == nil {
|
||||
t.Errorf("VMOV R7, V8.D1 assembled, want an arrangement error")
|
||||
}
|
||||
}
|
||||
|
||||
// TestArm64SimdTwoOperand pins the two-operand accumulate spellings
|
||||
// VADD/VSUB Vm, Vn against `go tool asm -S` output (Go 1.27, arm64):
|
||||
// word = 5<<28|7<<25|7<<21|1<<15|1<<10 for VADD (7<<28 for VSUB) with
|
||||
// rf<<16 | rn<<5 | rn, bare V registers only (asm7.go case 89).
|
||||
func TestArm64SimdTwoOperand(t *testing.T) {
|
||||
got := arm64Words(t, "\tVADD V7, V8\n\tVSUB V7, V8\n\tVADD V1, V2\n\tVADD V0.B16, V1.B16, V2.B16\n")
|
||||
want := []uint32{
|
||||
0x5ee78508, // VADD V7, V8
|
||||
0x7ee78508, // VSUB V7, V8
|
||||
0x5ee18442, // VADD V1, V2
|
||||
0x4e208422, // VADD arranged: the ordinary three-register path
|
||||
0xd65f03c0,
|
||||
}
|
||||
if len(got) != len(want) {
|
||||
t.Fatalf("word count = %d, want %d", len(got), len(want))
|
||||
}
|
||||
for i := range want {
|
||||
if got[i] != want[i] {
|
||||
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestArm64TruncMove pins the truncating register moves against
|
||||
// `go tool asm -S` output (Go 1.27, arm64): the signed forms lower to SXTB,
|
||||
// SXTH and SXTW (SBFM), the unsigned byte and halfword forms to UXTB and
|
||||
// UXTH (UBFM), MOVWU to a W ORR, and a narrow move out of the zero register
|
||||
// drops to the W ORR too (asm7.go case 45).
|
||||
func TestArm64TruncMove(t *testing.T) {
|
||||
got := arm64Words(t, "\tMOVB R3, R4\n\tMOVH R5, R6\n\tMOVW R9, R10\n"+
|
||||
"\tMOVBU R3, R4\n\tMOVHU R3, R4\n\tMOVWU R3, R4\n\tMOVD R3, R4\n"+
|
||||
"\tMOVD ZR, R4\n\tMOVB ZR, R4\n\tMOVWU ZR, R5\n")
|
||||
want := []uint32{
|
||||
0x93401c64, // MOVB = SXTB
|
||||
0x93403ca6, // MOVH = SXTH
|
||||
0x93407d2a, // MOVW = SXTW
|
||||
0xd3401c64, // MOVBU = UXTB
|
||||
0xd3403c64, // MOVHU = UXTH
|
||||
0x2a0303e4, // MOVWU = ORR W
|
||||
0xaa0303e4, // MOVD = ORR X
|
||||
0xaa1f03e4, // MOVD ZR, R4 keeps the X form
|
||||
0x2a1f03e4, // MOVB ZR, R4 drops to the W form
|
||||
0x2a1f03e5, // MOVWU ZR, R5
|
||||
0xd65f03c0,
|
||||
}
|
||||
if len(got) != len(want) {
|
||||
t.Fatalf("word count = %d, want %d", len(got), len(want))
|
||||
}
|
||||
for i := range want {
|
||||
if got[i] != want[i] {
|
||||
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestArm64SIMDLoadStore pins the structure loads and stores.
|
||||
func TestArm64SIMDLoadStore(t *testing.T) {
|
||||
got := arm64Words(t, "\tVLD1 (R2), [V21.B16]\n\tVLD1 (R1), [V2.B16, V3.B16]\n\tVLD1 (R29), [V14.D1, V15.D1, V16.D1, V17.D1]\n"+
|
||||
@@ -1295,20 +1420,170 @@ func TestArm64ExclNoOffset(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// TestArm64AddSubImmRange: immediates that cannot ride the imm12 field are
|
||||
// rejected instead of wrapping through int32.
|
||||
func TestArm64AddSubImmRange(t *testing.T) {
|
||||
for _, body := range []string{
|
||||
"\tADD $0x100000000, R0, R1\n",
|
||||
"\tSUB $-0x100000000, R0, R1\n",
|
||||
"\tCMP $0x100000000, R0\n",
|
||||
} {
|
||||
f, errs := parser.Parse("test_arm64.s", "#include \"textflag.h\"\n\nTEXT ·f(SB), NOSPLIT, $0-0\n"+body+"\tRET\n")
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
// TestArm64AddSubImmWide pins the wide-immediate classification the toolchain
|
||||
// applies to the ADD/SUB family (asm7.go cases 48, 62, 13): the ADDCON2 split
|
||||
// into two imm12 instructions for plain ADD/SUB, the bitmask ORR into REGTMP,
|
||||
// and the MOVZ/MOVN/MOVK materialisations followed by the register form.
|
||||
// Comparisons never split, and the W forms classify the 32-bit value. Every
|
||||
// word is go tool asm's own for the same source.
|
||||
func TestArm64AddSubImmWide(t *testing.T) {
|
||||
got := arm64Words(t, strings.Join([]string{
|
||||
"\tADD $0xaaaaaa, R2, R3",
|
||||
"\tSUB $0xaaaaaa, R2",
|
||||
"\tADD $0x186a0, R2, R5",
|
||||
"\tADD $0x1ffe00, R2, R3",
|
||||
"\tADD $0x3fffffffc000, R5",
|
||||
"\tADD $-100000, R2, R3",
|
||||
"\tADD $-2048, R2, R3",
|
||||
"\tCMP $0xaaaaaa, R2",
|
||||
"\tCMP $0xffffffffffa0, R3",
|
||||
"\tCMPW $27745, R2",
|
||||
"\tCMPW $0x60060, R2",
|
||||
"\tADDS $0xaaaaaa, R2, R3",
|
||||
"\tADD $0x12345678, R2, R3",
|
||||
"\tADDW $0x60060, R2",
|
||||
"\tSUB $0xe7791f700, R3, R1",
|
||||
"\tADDW $0x12345678, R2, R3",
|
||||
"\tCMN $0x1000000, R2",
|
||||
}, "\n")+"\n")
|
||||
want := []uint32{
|
||||
0x912aa843, 0x916aa863, // ADD $0xaaaaaa, R2, R3: ADDCON2 split
|
||||
0xd12aa842, 0xd16aa842, // SUB $0xaaaaaa, R2: split with Rd = Rn
|
||||
0x911a8045, 0x914060a5, // ADD $0x186a0, R2, R5: split
|
||||
0xb2772ffb, 0x8b1b0043, // ADD $0x1ffe00: bitmask beats the split
|
||||
0xb2727ffb, 0x8b1b00a5, // ADD $0x3fffffffc000: bitmask into REGTMP
|
||||
0x9290d3fb, 0xf2bfffdb, 0x8b1b0043, // ADD $-100000: MOVN + MOVK
|
||||
0x9280fffb, 0x8b1b0043, // ADD $-2048: single MOVN + ADD
|
||||
0xd295555b, 0xf2a0155b, 0xeb1b005f, // CMP: never split, MOVZ + MOVK
|
||||
0x92800bfb, 0xf2e0001b, 0xeb1b007f, // CMP $0xffffffffffa0: MOVN + fixup
|
||||
0x528d8c3b, 0x6b1b005f, // CMPW $27745: W movcon, single MOVZW
|
||||
0x52800c1b, 0x72a000db, 0x6b1b005f, // CMPW $0x60060: S form skips the split
|
||||
0xd295555b, 0xf2a0155b, 0xab1b0043, // ADDS $0xaaaaaa: MOVZ + MOVK + ADDS
|
||||
0xd28acf1b, 0xf2a2469b, 0x8b1b0043, // ADD $0x12345678: MOVZ + MOVK
|
||||
0x11018042, 0x11418042, // ADDW $0x60060: W split
|
||||
0xd29ee01b, 0xf2aef23b, 0xf2c001db, 0xcb1b0061, // SUB $0xe7791f700
|
||||
0x528acf1b, 0x72a2469b, 0x0b1b0043, // ADDW $0x12345678: MOVZW + MOVKW
|
||||
0xd2a0201b, 0xab1b005f, // CMN $0x1000000: single MOVZ + CMN
|
||||
0xd65f03c0, // RET
|
||||
}
|
||||
if len(got) != len(want) {
|
||||
t.Fatalf("word count = %d, want %d", len(got), len(want))
|
||||
}
|
||||
for i := range want {
|
||||
if got[i] != want[i] {
|
||||
t.Errorf("wide word %d = %08x, want %08x", i, got[i], want[i])
|
||||
}
|
||||
if _, err := AssembleFileARM64(f); err == nil {
|
||||
t.Errorf("%s: expected an error, got none", body)
|
||||
}
|
||||
}
|
||||
|
||||
// TestArm64CarryImmWide pins the carry family's $0 spellings in two and
|
||||
// three operands, the ROR shift on the logical group (and its rejection for
|
||||
// the arithmetic forms), the NGC/MNEG zero-register aliases and the vector
|
||||
// alias with an element selector. Words are go tool asm's own.
|
||||
func TestArm64CarryShiftAlias(t *testing.T) {
|
||||
got := arm64Words(t, "\tADC $0, R20\n\tADC $0, R20, R4\n\tSBCS $0, R4, R12\n"+
|
||||
"\tSBCS R15, R4, R12\n\tANDW R9@>7, R19, R26\n\tAND R1@>33, R2, R3\n"+
|
||||
"\tNEGSW R23<<1, R30\n\tNGC R2, R7\n\tMNEG R14, R27, R23\n")
|
||||
want := []uint32{
|
||||
0x9a1f0294, // ADC ZR, R20, R20
|
||||
0x9a1f0284, // ADC ZR, R20, R4
|
||||
0xfa1f008c, // SBCS ZR, R4, R12
|
||||
0xfa0f008c, // SBCS R15, R4, R12
|
||||
0x0ac91e7a, // ANDW R9 ROR 7, R19, R26
|
||||
0x8ac18443, // AND R1 ROR 33, R2, R3
|
||||
0x6b1707fe, // SUBSW ZR, R30, R23 LSL 1
|
||||
0xda0203e7, // SBC ZR, R7, R2
|
||||
0x9b0eff77, // MSUB ZR, R27, R14, R23
|
||||
0xd65f03c0, // RET
|
||||
}
|
||||
if len(got) != len(want) {
|
||||
t.Fatalf("word count = %d, want %d", len(got), len(want))
|
||||
}
|
||||
for i := range want {
|
||||
if got[i] != want[i] {
|
||||
t.Errorf("carry word %d = %08x, want %08x", i, got[i], want[i])
|
||||
}
|
||||
}
|
||||
|
||||
// ROR on an arithmetic form is unallocated: the toolchain reports an
|
||||
// unsupported shift operator.
|
||||
f, errs := parser.Parse("test_arm64.s", "#include \"textflag.h\"\n\nTEXT ·f(SB), NOSPLIT, $0-0\n\tADD R1@>33, R2, R3\n\tRET\n")
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
if _, err := AssembleFileARM64(f); err == nil {
|
||||
t.Error("ADD R1@>33: expected an error, got none")
|
||||
}
|
||||
}
|
||||
|
||||
// TestArm64VecAliasElement pins the register-alias rewrite inside a vector
|
||||
// operand with an element selector and inside a split register list: the
|
||||
// aliases resolve textually where the parser carries the selector apart from
|
||||
// the name. Words are go tool asm's own.
|
||||
func TestArm64VecAliasElement(t *testing.T) {
|
||||
src := `#include "textflag.h"
|
||||
|
||||
#define POLY V15
|
||||
#define ACC0 V8
|
||||
#define ACC1 V9
|
||||
|
||||
TEXT ·f(SB), NOSPLIT, $0-0
|
||||
VMOV R1, POLY.D[0]
|
||||
VEOR POLY.B16, POLY.B16, POLY.B16
|
||||
VLD1 (R0), [ACC0.B16]
|
||||
VLD1.P (R0), [ACC0.B16, ACC1.B16]
|
||||
VST1.P [ACC0.B16, ACC1.B16], 32(R1)
|
||||
RET
|
||||
`
|
||||
f, errs := parser.Parse("test_arm64.s", src)
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
img, err := AssembleFileARM64(f)
|
||||
if err != nil {
|
||||
t.Fatalf("AssembleFileARM64: %v", err)
|
||||
}
|
||||
got := leWords(img.Code)
|
||||
want := []uint32{
|
||||
0x4e081c2f, // INS V15.D[0], R1
|
||||
0x6e2f1def, // VEOR V15.B16, V15.B16, V15.B16
|
||||
0x4c407008, // VLD1 (R0), [V8.B16]
|
||||
0x4cdfa008, // VLD1.P (R0), [V8.B16, V9.B16]
|
||||
0x4c9fa028, // VST1.P [V8.B16, V9.B16], 32(R1)
|
||||
0xd65f03c0, // RET
|
||||
}
|
||||
if len(got) != len(want) {
|
||||
t.Fatalf("word count = %d, want %d", len(got), len(want))
|
||||
}
|
||||
for i := range want {
|
||||
if got[i] != want[i] {
|
||||
t.Errorf("vecalias word %d = %08x, want %08x", i, got[i], want[i])
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestArm64AddSubImmBeyond32 pins the materialisation the toolchain applies
|
||||
// once the value leaves every imm12 form: a constant sequence into REGTMP
|
||||
// (R27) followed by the register form. SUB $-0x100000000 is a bitmask
|
||||
// immediate, so it rides the ORR form; the others take MOVZ. Words are go
|
||||
// tool asm's own.
|
||||
func TestArm64AddSubImmBeyond32(t *testing.T) {
|
||||
got := arm64Words(t, "\tADD $0x100000000, R0, R1\n\tSUB $-0x100000000, R0, R1\n\tCMP $0x100000000, R0\n")
|
||||
want := []uint32{
|
||||
0xd2c0003b, // MOVZ $(1<<32>>16), R27 (hw=2)
|
||||
0x8b1b0001, // ADD R27, R0, R1
|
||||
0xb2607ffb, // ORR $-4294967296, ZR, R27 (bitmask)
|
||||
0xcb1b0001, // SUB R27, R0, R1
|
||||
0xd2c0003b, // MOVZ $(1<<32>>16), R27 (hw=2)
|
||||
0xeb1b001f, // CMP R27, R0
|
||||
0xd65f03c0, // RET
|
||||
}
|
||||
if len(got) != len(want) {
|
||||
t.Fatalf("word count = %d, want %d", len(got), len(want))
|
||||
}
|
||||
for i := range want {
|
||||
if got[i] != want[i] {
|
||||
t.Errorf("word %d = %08x, want %08x", i, got[i], want[i])
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
+551
-38
@@ -5,6 +5,7 @@ package asm
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"strconv"
|
||||
"strings"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
|
||||
@@ -28,7 +29,7 @@ import (
|
||||
// emitted: the bytes match go tool asm only for NOSPLIT functions or
|
||||
// zero-frame leaves, where the toolchain emits no guard either.
|
||||
func Assemble(t *ast.Text) ([]byte, map[string]int, error) {
|
||||
code, _, labels, _, _, err := assemble(t, nil)
|
||||
code, _, labels, _, _, _, err := assemble(t, nil)
|
||||
return code, labels, err
|
||||
}
|
||||
|
||||
@@ -37,10 +38,25 @@ func Assemble(t *ast.Text) ([]byte, map[string]int, error) {
|
||||
// rejects SB operands outright (single-function assembly cannot resolve
|
||||
// them). When allowExternal is set, a reference to a symbol no GLOBL in the
|
||||
// file defines is recorded as an external relocation instead of failing
|
||||
// the object-file emitters resolve it at link time.
|
||||
// the object-file emitters resolve it at link time. goos selects the TLS
|
||||
// access form: the empty default behaves as linux.
|
||||
type linkInfo struct {
|
||||
symbols map[string]bool
|
||||
allowExternal bool
|
||||
goos string
|
||||
}
|
||||
|
||||
// tlsOneInsn reports the one-instruction TLS form, obj6.go's
|
||||
// CanUse1InsnTLS for the GOOS gasm supports: the bare TLS load nops out and
|
||||
// the (TLS*1) index folds to a segment-absolute access. Windows and plan9
|
||||
// keep the two-instruction form; shared linux does too, which gasm's raw
|
||||
// path does not model and therefore does not select.
|
||||
func (l *linkInfo) tlsOneInsn() bool {
|
||||
switch l.goos {
|
||||
case "", "linux", "freebsd":
|
||||
return true
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// sbPatch is a function-relative static-symbol relocation: the disp32 field
|
||||
@@ -66,7 +82,10 @@ type spadjStep struct {
|
||||
// assemble encodes a TEXT body, returning the machine code, the static-symbol
|
||||
// patch sites (for the file-level layout to resolve), the label table and the
|
||||
// stack-adjustment boundaries.
|
||||
func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, []spadjStep, []LineEntry, error) {
|
||||
func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, []spadjStep, []LineEntry, []floatPoolEntry, error) {
|
||||
if err := checkAdjspBalance(t); err != nil {
|
||||
return nil, nil, nil, nil, nil, nil, err
|
||||
}
|
||||
fi := computeFrame(t)
|
||||
chain := jumpChain(t)
|
||||
resolve := func(name string) string {
|
||||
@@ -82,29 +101,118 @@ func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, [
|
||||
// outgrows the short form.
|
||||
long := make([]bool, len(t.Body))
|
||||
sizes := make([]int, len(t.Body))
|
||||
numTargets := make([]int, len(t.Body))
|
||||
for i := range numTargets {
|
||||
numTargets[i] = -1
|
||||
}
|
||||
offsets := map[string]int{}
|
||||
pcs := make([]int, len(t.Body))
|
||||
var guardJBlong, guardJBElong, moreJMPlong bool
|
||||
poolSeen := map[string]bool{}
|
||||
var poolList []floatPoolEntry
|
||||
for {
|
||||
guard := fi.guardLen(guardJBlong, guardJBElong)
|
||||
pos := guard + len(fi.prologue)
|
||||
for i := range numTargets {
|
||||
numTargets[i] = -1
|
||||
}
|
||||
idxAtPc := map[int]int{}
|
||||
for i, stmt := range t.Body {
|
||||
switch s := stmt.(type) {
|
||||
case *ast.Label:
|
||||
offsets[s.Name.Text] = pos
|
||||
case *ast.Instr:
|
||||
if strings.ToUpper(s.Mnemonic.Text) == "PCALIGN" {
|
||||
// The alignment pseudo-statement: its size is the
|
||||
// padding to the next boundary at this very position,
|
||||
// filled with NOPs at emission.
|
||||
pad, err := pcAlignPad(pcAlignValue(s), pos)
|
||||
if err != nil {
|
||||
return nil, nil, nil, nil, nil, nil, fmt.Errorf("PCALIGN: %w", err)
|
||||
}
|
||||
sizes[i] = pad
|
||||
pcs[i] = pos
|
||||
pos += pad
|
||||
continue
|
||||
}
|
||||
sz, err := instrSize(s, fi, long[i], link)
|
||||
if err != nil {
|
||||
return nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", s.Mnemonic.Text, err)
|
||||
return nil, nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", s.Mnemonic.Text, err)
|
||||
}
|
||||
sizes[i] = sz
|
||||
pcs[i] = pos
|
||||
idxAtPc[pos] = i
|
||||
pos += sz
|
||||
}
|
||||
}
|
||||
bodyLen := pos - (guard + len(fi.prologue))
|
||||
// Expand any short jump whose displacement no longer fits rel8.
|
||||
changed := false
|
||||
// Numeric ±N(PC) jumps resolve against this iteration's layout; the
|
||||
// emission pass reads the same table after the loop converges. A
|
||||
// target that is itself an unconditional local JMP is chased to the
|
||||
// ultimate target: the toolchain's brloop pass collapses branch-to-
|
||||
// branch chains before it encodes, so matching its bytes requires
|
||||
// the same redirection.
|
||||
for i := range numTargets {
|
||||
numTargets[i] = -1
|
||||
}
|
||||
for i, stmt := range t.Body {
|
||||
s, ok := stmt.(*ast.Instr)
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
if len(s.Operands) == 1 {
|
||||
if n, isNum := pcJumpOffset(s.Operands[0]); isNum {
|
||||
if target, okT := pcJumpTarget(t, i, n, pcs); okT {
|
||||
numTargets[i] = target
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
for i := range numTargets {
|
||||
if numTargets[i] < 0 {
|
||||
continue
|
||||
}
|
||||
tgt := numTargets[i]
|
||||
for hop := 0; hop < len(t.Body); hop++ {
|
||||
idx, ok := idxAtPc[tgt]
|
||||
if !ok {
|
||||
break
|
||||
}
|
||||
in, ok := t.Body[idx].(*ast.Instr)
|
||||
if !ok || strings.ToUpper(in.Mnemonic.Text) != "JMP" || len(in.Operands) != 1 {
|
||||
break
|
||||
}
|
||||
if name, isLabel := labelName(in.Operands[0]); isLabel {
|
||||
tgt = offsets[resolve(name)]
|
||||
continue
|
||||
}
|
||||
if n, isNum := pcJumpOffset(in.Operands[0]); isNum {
|
||||
next, okT := pcJumpTarget(t, idx, n, pcs)
|
||||
if !okT {
|
||||
break
|
||||
}
|
||||
tgt = next
|
||||
continue
|
||||
}
|
||||
break // JMP through a register or memory: the chain ends
|
||||
}
|
||||
numTargets[i] = tgt
|
||||
}
|
||||
for i, stmt := range t.Body {
|
||||
s, ok := stmt.(*ast.Instr)
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
if numTargets[i] >= 0 && !long[i] {
|
||||
rel := int64(numTargets[i] - (pcs[i] + jumpSize(strings.ToUpper(s.Mnemonic.Text), false)))
|
||||
if !fits8(rel) {
|
||||
long[i] = true
|
||||
changed = true
|
||||
}
|
||||
}
|
||||
}
|
||||
for i, stmt := range t.Body {
|
||||
s, ok := stmt.(*ast.Instr)
|
||||
if !ok {
|
||||
@@ -203,6 +311,14 @@ func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, [
|
||||
spadjStep{guardLen + len(fi.prologue), 8 + fi.size},
|
||||
)
|
||||
}
|
||||
// frameBase is the SP delta the prologue leaves: 8 for the saved base
|
||||
// pointer plus the frame, 0 frameless. bodyDelta tracks the ADJSP
|
||||
// statements' straight-line sum, so a mid-body step's value is the
|
||||
// frame base plus what the body has opened so far.
|
||||
frameBase, bodyDelta := 0, 0
|
||||
if fi.useFP {
|
||||
frameBase = 8 + fi.size
|
||||
}
|
||||
pos := guardLen + len(fi.prologue)
|
||||
for i, stmt := range t.Body {
|
||||
s, ok := stmt.(*ast.Instr)
|
||||
@@ -218,18 +334,34 @@ func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, [
|
||||
spadjStep{pos + epi, 0},
|
||||
)
|
||||
}
|
||||
code, ps, err := encodeInstr(s, pos, offsets, fi, long[i], resolve, link)
|
||||
code, ps, pool, err := encodeInstr(s, pos, offsets, fi, long[i], resolve, link, numTargets[i])
|
||||
if err != nil {
|
||||
return nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", s.Mnemonic.Text, err)
|
||||
return nil, nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", s.Mnemonic.Text, err)
|
||||
}
|
||||
for _, entry := range pool {
|
||||
if !poolSeen[entry.name] {
|
||||
poolSeen[entry.name] = true
|
||||
poolList = append(poolList, entry)
|
||||
}
|
||||
}
|
||||
if len(code) != sizes[i] {
|
||||
return nil, nil, nil, nil, nil, fmt.Errorf("%s: size mismatch (%d vs %d)", s.Mnemonic.Text, len(code), sizes[i])
|
||||
return nil, nil, nil, nil, nil, nil, fmt.Errorf("%s: size mismatch (%d vs %d)", s.Mnemonic.Text, len(code), sizes[i])
|
||||
}
|
||||
if strings.ToUpper(s.Mnemonic.Text) == "CALL" {
|
||||
for k := range ps {
|
||||
ps[k].kind = RelCall
|
||||
}
|
||||
}
|
||||
if strings.ToUpper(s.Mnemonic.Text) == "ADJSP" && len(s.Operands) == 1 && s.Operands[0].Imm.HasVal {
|
||||
// The statement shifted SP mid-body: record the new running
|
||||
// delta as the value in effect from just past the instruction.
|
||||
v := s.Operands[0].Imm.Val
|
||||
if s.Operands[0].Imm.Neg {
|
||||
v = -v
|
||||
}
|
||||
bodyDelta += int(v)
|
||||
steps = append(steps, spadjStep{pos + len(code), frameBase + bodyDelta})
|
||||
}
|
||||
patches = append(patches, ps...)
|
||||
lines = append(lines, LineEntry{Offset: pos, Line: s.Pos().Line})
|
||||
out = append(out, code...)
|
||||
@@ -251,7 +383,7 @@ func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, [
|
||||
pos += len(suffix)
|
||||
}
|
||||
_ = pos
|
||||
return out, patches, offsets, steps, lines, nil
|
||||
return out, patches, offsets, steps, lines, poolList, nil
|
||||
}
|
||||
|
||||
// jumpChain precomputes jump-to-jump folding: a label whose first instruction
|
||||
@@ -398,6 +530,83 @@ func computeFrame(t *ast.Text) frameInfo {
|
||||
return fi
|
||||
}
|
||||
|
||||
// pcJumpOffset recognises the numeric relative jump operand ±N(PC) and
|
||||
// returns N: the toolchain counts instructions, not bytes, so +2(PC) targets
|
||||
// the second instruction boundary after the branch.
|
||||
func pcJumpOffset(op *ast.Operand) (int, bool) {
|
||||
if op.Kind != ast.OpAddr || op.Addr.Base != "PC" {
|
||||
return 0, false
|
||||
}
|
||||
return int(op.Addr.Offset), true
|
||||
}
|
||||
|
||||
// pcJumpTarget resolves a numeric jump at statement index j: N counts the
|
||||
// instruction statements after the jump itself (N = 0 is the jump's own
|
||||
// address, the classic park loop), and the target is the start of the Nth
|
||||
// one. It reports false when the count runs past the end of the function.
|
||||
func pcJumpTarget(t *ast.Text, j, n int, pcs []int) (int, bool) {
|
||||
if n == 0 {
|
||||
return pcs[j], true
|
||||
}
|
||||
seen := 0
|
||||
for k := j + 1; k < len(t.Body); k++ {
|
||||
if _, ok := t.Body[k].(*ast.Instr); !ok {
|
||||
continue
|
||||
}
|
||||
seen++
|
||||
if seen == n {
|
||||
return pcs[k], true
|
||||
}
|
||||
}
|
||||
return 0, false
|
||||
}
|
||||
|
||||
// x86 NOP encodings, single-instruction no-ops of lengths 1 to 9 (the
|
||||
// toolchain's asm6.go nop table); longer padding repeats the largest that
|
||||
// fits, greedy from the end.
|
||||
var x86Nops = [][]byte{
|
||||
{0x90},
|
||||
{0x66, 0x90},
|
||||
{0x0F, 0x1F, 0x00},
|
||||
{0x0F, 0x1F, 0x40, 0x00},
|
||||
{0x0F, 0x1F, 0x44, 0x00, 0x00},
|
||||
{0x66, 0x0F, 0x1F, 0x44, 0x00, 0x00},
|
||||
{0x0F, 0x1F, 0x80, 0x00, 0x00, 0x00, 0x00},
|
||||
{0x0F, 0x1F, 0x84, 0x00, 0x00, 0x00, 0x00, 0x00},
|
||||
{0x66, 0x0F, 0x1F, 0x84, 0x00, 0x00, 0x00, 0x00, 0x00},
|
||||
}
|
||||
|
||||
// fillNOPs fills p with the greedy largest single-instruction NOPs, exactly
|
||||
// the toolchain's fillnop.
|
||||
func fillNOPs(p []byte) {
|
||||
for len(p) > 0 {
|
||||
m := min(len(p), len(x86Nops))
|
||||
copy(p[:m], x86Nops[m-1])
|
||||
p = p[m:]
|
||||
}
|
||||
}
|
||||
|
||||
// pcAlignPad computes the padding PCALIGN $align inserts at pos: the
|
||||
// alignment must be a power of two in [8, 2048] and the padding runs to the
|
||||
// next boundary (zero when the position is already aligned).
|
||||
func pcAlignPad(align, pos int) (int, error) {
|
||||
if align <= 0 || align&(align-1) != 0 || align < 8 || align > 2048 {
|
||||
return 0, fmt.Errorf("alignment value of an instruction must be a power of two and in the range [8, 2048], got %d", align)
|
||||
}
|
||||
if lob := pos & (align - 1); lob != 0 {
|
||||
return align - lob, nil
|
||||
}
|
||||
return 0, nil
|
||||
}
|
||||
|
||||
// pcAlignValue reads a PCALIGN statement's alignment operand.
|
||||
func pcAlignValue(s *ast.Instr) int {
|
||||
if len(s.Operands) == 1 && s.Operands[0].Kind == ast.OpImmediate && s.Operands[0].Imm.HasVal {
|
||||
return int(s.Operands[0].Imm.Val)
|
||||
}
|
||||
return 0 // rejected by pcAlignPad's range check
|
||||
}
|
||||
|
||||
// hasCall reports whether the function body contains a CALL instruction.
|
||||
func hasCall(t *ast.Text) bool {
|
||||
for _, stmt := range t.Body {
|
||||
@@ -412,6 +621,40 @@ func hasCall(t *ast.Text) bool {
|
||||
return false
|
||||
}
|
||||
|
||||
// checkAdjspBalance mirrors the toolchain's push/pop walk: every ADJSP
|
||||
// shifts SP away from the entry state and every RET must see the shifts
|
||||
// closed. The assembler's own prologue and epilogue contribute matching
|
||||
// deltas on both sides, so the statements' straight-line sum must be zero
|
||||
// at each RET; branches do not reset the walk, which runs over the program
|
||||
// list in source order. go tool asm reports an offender as "unbalanced
|
||||
// PUSH/POP" (verified against ADJSP $16 before a RET, accepted as a
|
||||
// $16/$-16 pair, per-RET rather than per-function).
|
||||
func checkAdjspBalance(t *ast.Text) error {
|
||||
delta := 0
|
||||
for _, stmt := range t.Body {
|
||||
in, ok := stmt.(*ast.Instr)
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
switch strings.ToUpper(in.Mnemonic.Text) {
|
||||
case "ADJSP":
|
||||
if len(in.Operands) != 1 || !in.Operands[0].Imm.HasVal {
|
||||
continue // reported during emission
|
||||
}
|
||||
v := in.Operands[0].Imm.Val
|
||||
if in.Operands[0].Imm.Neg {
|
||||
v = -v
|
||||
}
|
||||
delta += int(v)
|
||||
case "RET":
|
||||
if delta != 0 {
|
||||
return fmt.Errorf("unbalanced PUSH/POP")
|
||||
}
|
||||
}
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// guardLen returns the byte length of the stack-split guard prefix. The
|
||||
// final conditional branch (JBE, and JB in the big class) is 2 bytes in the
|
||||
// short form and 6 in the long form.
|
||||
@@ -538,7 +781,7 @@ func instrSize(s *ast.Instr, fi frameInfo, long bool, link *linkInfo) (int, erro
|
||||
}
|
||||
return jumpSize(mnem, long), nil
|
||||
}
|
||||
code, _, err := encodeInstr(s, 0, nil, fi, false, nil, link)
|
||||
code, _, _, err := encodeInstr(s, 0, nil, fi, false, nil, link, -1)
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
@@ -573,9 +816,21 @@ func jumpSize(mnem string, long bool) int {
|
||||
// (relative to pc, the instruction's own offset). A RET in a frame-pointer
|
||||
// function is prefixed with the epilogue. resolve, when non-nil, redirects a
|
||||
// jump label through the jump-to-jump chain before the offset lookup.
|
||||
func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, long bool, resolve func(string) string, link *linkInfo) ([]byte, []sbPatch, error) {
|
||||
func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, long bool, resolve func(string) string, link *linkInfo, numTarget int) ([]byte, []sbPatch, []floatPoolEntry, error) {
|
||||
mnem := strings.ToUpper(s.Mnemonic.Text)
|
||||
|
||||
if mnem == "PCALIGN" {
|
||||
// The layout pass already accounted the padding; emit the same
|
||||
// amount of NOP bytes for the statement's own position.
|
||||
pad, err := pcAlignPad(pcAlignValue(s), pc)
|
||||
if err != nil {
|
||||
return nil, nil, nil, err
|
||||
}
|
||||
out := make([]byte, pad)
|
||||
fillNOPs(out)
|
||||
return out, nil, nil, nil
|
||||
}
|
||||
|
||||
var prefix []byte
|
||||
if mnem == "RET" && fi.useFP {
|
||||
prefix = fi.epilogue
|
||||
@@ -583,6 +838,7 @@ func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, lon
|
||||
|
||||
var code []byte
|
||||
var ps []sbPatch
|
||||
var pool []floatPoolEntry
|
||||
var err error
|
||||
if isJumpMnemonic(mnem) {
|
||||
if (mnem == "CALL" || mnem == "JMP") && isSBCall(s) {
|
||||
@@ -591,7 +847,7 @@ func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, lon
|
||||
// or the linker.
|
||||
code, ps, err = encodeSBCall(s, link)
|
||||
if err != nil {
|
||||
return nil, nil, err
|
||||
return nil, nil, nil, err
|
||||
}
|
||||
for i := range ps {
|
||||
ps[i].kind = RelCall
|
||||
@@ -601,23 +857,23 @@ func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, lon
|
||||
ps[i].off += body
|
||||
ps[i].after = body + len(code)
|
||||
}
|
||||
return append(prefix, code...), ps, nil
|
||||
return append(prefix, code...), ps, nil, nil
|
||||
}
|
||||
if (mnem == "CALL" || mnem == "JMP") && indirectJumpTarget(s) {
|
||||
// JMP/CALL through a register or memory: no relocation and no
|
||||
// label to resolve, the operand fully determines the bytes.
|
||||
code, err = encodeIndirectJump(s, mnem)
|
||||
if err != nil {
|
||||
return nil, nil, err
|
||||
return nil, nil, nil, err
|
||||
}
|
||||
return append(prefix, code...), nil, nil
|
||||
return append(prefix, code...), nil, nil, nil
|
||||
}
|
||||
code, err = encodeJump(s, mnem, pc+len(prefix), offsets, long, resolve)
|
||||
code, err = encodeJump(s, mnem, pc+len(prefix), offsets, long, resolve, numTarget)
|
||||
} else {
|
||||
code, ps, err = encodeNormal(s, fi, link)
|
||||
code, ps, pool, err = encodeNormal(s, fi, link)
|
||||
}
|
||||
if err != nil {
|
||||
return nil, nil, err
|
||||
return nil, nil, nil, err
|
||||
}
|
||||
// Anchor the patch fields at function-relative positions: off indexes the
|
||||
// disp32 field, after is the address just past the instruction.
|
||||
@@ -626,49 +882,183 @@ func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, lon
|
||||
ps[i].off += body
|
||||
ps[i].after = body + len(code)
|
||||
}
|
||||
return append(prefix, code...), ps, nil
|
||||
return append(prefix, code...), ps, pool, nil
|
||||
}
|
||||
|
||||
func encodeNormal(s *ast.Instr, fi frameInfo, link *linkInfo) ([]byte, []sbPatch, error) {
|
||||
_, size := splitSize(strings.ToUpper(s.Mnemonic.Text))
|
||||
func encodeNormal(s *ast.Instr, fi frameInfo, link *linkInfo) ([]byte, []sbPatch, []floatPoolEntry, error) {
|
||||
mnemUpper := strings.ToUpper(s.Mnemonic.Text)
|
||||
if mnemUpper == "FUNCDATA" || mnemUpper == "PCDATA" {
|
||||
code, err := encodeBookkeeping(mnemUpper, s)
|
||||
if err != nil {
|
||||
return nil, nil, nil, err
|
||||
}
|
||||
return code, nil, nil, nil
|
||||
}
|
||||
// MOVQ $sym±off(SB), r64: the toolchain assembles a symbol immediate as
|
||||
// LEAQ disp32(RIP), r64 with an R_PCREL relocation at the disp32 field,
|
||||
// never as a 64-bit absolute immediate (verified against go tool asm).
|
||||
// MOVD is the MOVQ alias; the narrower widths reject the form outright.
|
||||
if (mnemUpper == "MOVQ" || mnemUpper == "MOVD") && len(s.Operands) == 2 &&
|
||||
s.Operands[0].Kind == ast.OpImmediate && s.Operands[0].Imm.Sym != nil &&
|
||||
s.Operands[0].Imm.Sym.Pseudo == "SB" {
|
||||
mem := &ast.Operand{Kind: ast.OpAddr, Addr: ast.Address{Sym: s.Operands[0].Imm.Sym}}
|
||||
src, err := operandFromAST(mnemUpper, mem, 8, fi, link)
|
||||
if err != nil {
|
||||
return nil, nil, nil, err
|
||||
}
|
||||
dst, err := operandFromAST(mnemUpper, s.Operands[1], 8, fi, link)
|
||||
if err != nil {
|
||||
return nil, nil, nil, err
|
||||
}
|
||||
e := &enc{}
|
||||
if err := e.encodeLea([]Operand{src, dst}, 8); err != nil {
|
||||
return nil, nil, nil, err
|
||||
}
|
||||
ps := make([]sbPatch, len(e.patches))
|
||||
for i, p := range e.patches {
|
||||
ps[i] = sbPatch{off: p.off, name: p.name, addend: p.addend}
|
||||
}
|
||||
return e.out, ps, nil, nil
|
||||
}
|
||||
// MOVQ/MOVL TLS, r: the bare TLS load. The toolchain's progedit nops
|
||||
// it out on the one-instruction TLS systems (linux and freebsd, not
|
||||
// shared) and encodes the segment-prefixed load elsewhere; get_tls(r),
|
||||
// the macro GOROOT's go_tls.h defines, expands to exactly this
|
||||
// statement, and the toolchain's pairing pass removes it whenever the
|
||||
// following instruction's (TLS*1) index folds.
|
||||
if (mnemUpper == "MOVQ" || mnemUpper == "MOVL") && len(s.Operands) == 2 && isBareTLS(s.Operands[0]) {
|
||||
return encodeTLSBaseLoad(s, fi, link)
|
||||
}
|
||||
_, size := splitSize(mnemUpper)
|
||||
if size == 0 {
|
||||
size = 8
|
||||
}
|
||||
ops := make([]Operand, len(s.Operands))
|
||||
for i, op := range s.Operands {
|
||||
o, err := operandFromAST(op, size, fi, link)
|
||||
o, err := operandFromAST(mnemUpper, op, size, fi, link)
|
||||
if err != nil {
|
||||
return nil, nil, err
|
||||
return nil, nil, nil, err
|
||||
}
|
||||
ops[i] = o
|
||||
}
|
||||
e := &enc{}
|
||||
if err := e.encode(s.Mnemonic.Text, ops); err != nil {
|
||||
return nil, nil, err
|
||||
return nil, nil, nil, err
|
||||
}
|
||||
ps := make([]sbPatch, len(e.patches))
|
||||
for i, p := range e.patches {
|
||||
ps[i] = sbPatch{off: p.off, name: p.name, addend: p.addend}
|
||||
if p.tls {
|
||||
ps[i].kind = RelTLSLE
|
||||
}
|
||||
}
|
||||
return e.out, ps, nil
|
||||
return e.out, ps, e.floatPoolList(), nil
|
||||
}
|
||||
|
||||
// isBareTLS reports whether the operand is the bare TLS pseudo-register
|
||||
// load source, the expansion of go_tls.h's get_tls(r) macro.
|
||||
func isBareTLS(op *ast.Operand) bool {
|
||||
return op.Kind == ast.OpAddr && op.Addr.Sym != nil &&
|
||||
op.Addr.Sym.Pseudo == "" && op.Addr.Sym.Name == "TLS" &&
|
||||
op.Addr.Base == "" && op.Addr.Index == ""
|
||||
}
|
||||
|
||||
// encodeTLSBaseLoad assembles MOVQ/MOVL TLS, r. On the one-instruction TLS
|
||||
// systems (linux and freebsd outside -shared, obj6.go's CanUse1InsnTLS) the
|
||||
// statement nops out: the following (TLS*1) access folds to a direct
|
||||
// segment-absolute load. The two-instruction systems keep the segment load,
|
||||
// nine bytes with the R_TLSLE patch site at the disp32.
|
||||
func encodeTLSBaseLoad(s *ast.Instr, fi frameInfo, link *linkInfo) ([]byte, []sbPatch, []floatPoolEntry, error) {
|
||||
_, size := splitSize(strings.ToUpper(s.Mnemonic.Text))
|
||||
if size == 0 {
|
||||
size = 8
|
||||
}
|
||||
dst, err := operandFromAST("MOVQ", s.Operands[1], 8, fi, link)
|
||||
if err != nil {
|
||||
return nil, nil, nil, err
|
||||
}
|
||||
reg, ok := dst.(Reg)
|
||||
if !ok || reg.isVec() {
|
||||
return nil, nil, nil, fmt.Errorf("TLS: destination must be a general register")
|
||||
}
|
||||
if link == nil || link.tlsOneInsn() {
|
||||
return nil, nil, nil, nil // noped out
|
||||
}
|
||||
seg := byte(0x64) // FS
|
||||
if link.goos == "windows" {
|
||||
seg = 0x65 // GS
|
||||
}
|
||||
e := &enc{}
|
||||
i := &instr{
|
||||
prefix: seg,
|
||||
rexW: size == 8,
|
||||
rexR: reg.idx >= 8,
|
||||
opcode: []byte{0x8B},
|
||||
modrm: 0x04 | (reg.idx&7)<<3,
|
||||
sib: 0x25,
|
||||
disp: le32(0),
|
||||
tls: true,
|
||||
}
|
||||
if err := e.emit(i); err != nil {
|
||||
return nil, nil, nil, err
|
||||
}
|
||||
ps := make([]sbPatch, len(e.patches))
|
||||
for i, p := range e.patches {
|
||||
ps[i] = sbPatch{off: p.off, name: p.name, addend: p.addend, kind: RelTLSLE}
|
||||
}
|
||||
return e.out, ps, nil, nil
|
||||
}
|
||||
|
||||
// encodeBookkeeping accepts-and-ignores FUNCDATA and PCDATA at the statement
|
||||
// level, before operand conversion: the toolchain's shapes are FUNCDATA
|
||||
// $n, sym(SB) and PCDATA $n, $m, and neither contributes a byte to the
|
||||
// function body. The symbol reference must not run through the SB-operand
|
||||
// path, which demands file-level resolution the statement never needs.
|
||||
func encodeBookkeeping(upper string, s *ast.Instr) ([]byte, error) {
|
||||
if len(s.Operands) != 2 {
|
||||
return nil, fmt.Errorf("%s expects 2 operands, got %d", upper, len(s.Operands))
|
||||
}
|
||||
a, b := s.Operands[0], s.Operands[1]
|
||||
if a.Kind != ast.OpImmediate || !a.Imm.HasVal {
|
||||
return nil, fmt.Errorf("%s: first operand must be an integer immediate", upper)
|
||||
}
|
||||
switch upper {
|
||||
case "FUNCDATA":
|
||||
if b.Kind != ast.OpAddr || b.Addr.Sym == nil || b.Addr.Sym.Pseudo != "SB" {
|
||||
return nil, fmt.Errorf("FUNCDATA: second operand must be a symbol reference")
|
||||
}
|
||||
case "PCDATA":
|
||||
if b.Kind != ast.OpImmediate || !b.Imm.HasVal {
|
||||
return nil, fmt.Errorf("PCDATA: second operand must be an integer immediate")
|
||||
}
|
||||
}
|
||||
return nil, nil
|
||||
}
|
||||
|
||||
// encodeJump encodes a JMP/CALL/Jcc with a relative offset resolved from the
|
||||
// target label, in the short (rel8) or long (rel32) form.
|
||||
func encodeJump(s *ast.Instr, mnem string, pc int, offsets map[string]int, long bool, resolve func(string) string) ([]byte, error) {
|
||||
// target label or from a numeric ±N(PC) instruction count, in the short
|
||||
// (rel8) or long (rel32) form. numTarget is the resolved byte offset of a
|
||||
// numeric operand, negative when the operand is not one.
|
||||
func encodeJump(s *ast.Instr, mnem string, pc int, offsets map[string]int, long bool, resolve func(string) string, numTarget int) ([]byte, error) {
|
||||
if len(s.Operands) != 1 {
|
||||
return nil, fmt.Errorf("jump expects 1 operand, got %d", len(s.Operands))
|
||||
}
|
||||
name, ok := labelName(s.Operands[0])
|
||||
if !ok {
|
||||
name, isLabel := labelName(s.Operands[0])
|
||||
if !isLabel && numTarget < 0 {
|
||||
return nil, fmt.Errorf("jump target must be a local label")
|
||||
}
|
||||
if resolve != nil && mnem != "CALL" {
|
||||
name = resolve(name)
|
||||
}
|
||||
target, ok := offsets[name]
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("undefined label %q", name)
|
||||
var target int
|
||||
if isLabel {
|
||||
if resolve != nil && mnem != "CALL" {
|
||||
name = resolve(name)
|
||||
}
|
||||
t, ok := offsets[name]
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("undefined label %q", name)
|
||||
}
|
||||
target = t
|
||||
} else {
|
||||
target = numTarget
|
||||
}
|
||||
rel := int64(target - (pc + jumpSize(mnem, long)))
|
||||
|
||||
@@ -701,7 +1091,7 @@ func isSBCall(s *ast.Instr) bool {
|
||||
|
||||
// encodeSBCall encodes CALL sym(SB) as E8 rel32 with a patch site.
|
||||
func encodeSBCall(s *ast.Instr, link *linkInfo) ([]byte, []sbPatch, error) {
|
||||
o, err := operandFromAST(s.Operands[0], 8, frameInfo{}, link)
|
||||
o, err := operandFromAST(strings.ToUpper(s.Mnemonic.Text), s.Operands[0], 8, frameInfo{}, link)
|
||||
if err != nil {
|
||||
return nil, nil, err
|
||||
}
|
||||
@@ -742,6 +1132,11 @@ func indirectJumpTarget(s *ast.Instr) bool {
|
||||
return false
|
||||
}
|
||||
a := s.Operands[0].Addr
|
||||
// ±N(PC) is the numeric relative form, the PC counts instructions from
|
||||
// the branch: relative, not indirect.
|
||||
if a.Base == "PC" || a.Index == "PC" {
|
||||
return false
|
||||
}
|
||||
if a.Base != "" || a.Index != "" {
|
||||
return true
|
||||
}
|
||||
@@ -758,7 +1153,7 @@ func indirectJumpTarget(s *ast.Instr) bool {
|
||||
func encodeIndirectJump(s *ast.Instr, mnem string) ([]byte, error) {
|
||||
ops := make([]Operand, len(s.Operands))
|
||||
for i, op := range s.Operands {
|
||||
o, err := operandFromAST(op, 8, frameInfo{}, nil)
|
||||
o, err := operandFromAST(mnem, op, 8, frameInfo{}, nil)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
@@ -775,8 +1170,11 @@ func encodeIndirectJump(s *ast.Instr, mnem string) ([]byte, error) {
|
||||
var spReg = Reg{idx: 4, size: 8}
|
||||
|
||||
// operandFromAST converts a parsed operand into an encoder Operand, applying
|
||||
// the frame translation to FP/SP pseudo-register operands.
|
||||
func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Operand, error) {
|
||||
// the frame translation to FP/SP pseudo-register operands. mnemUpper is the
|
||||
// instruction's upper-case mnemonic, which the floating-point immediate gate
|
||||
// needs: only the SSE mnemonics whose encoding takes an XMM/memory source
|
||||
// accept one.
|
||||
func operandFromAST(mnemUpper string, op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Operand, error) {
|
||||
switch op.Kind {
|
||||
case ast.OpImmediate:
|
||||
if op.Imm.HasVal {
|
||||
@@ -786,11 +1184,45 @@ func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Op
|
||||
}
|
||||
return Imm(v), nil
|
||||
}
|
||||
// A floating-point immediate: $1.5, $-1.0 or the parenthesised
|
||||
// $(-1.0) spelling (the constant-expression folder only folds
|
||||
// integers, so that shape arrives with an empty Immediate and only
|
||||
// the raw spelling carries the value). The toolchain rewrites it
|
||||
// into a pooled-constant read on the SSE scalar paths and rejects
|
||||
// it everywhere else.
|
||||
if text, neg, ok := floatImmText(op); ok {
|
||||
if !sseFloatImm[mnemUpper] {
|
||||
return nil, fmt.Errorf("%s does not take a floating-point immediate", mnemUpper)
|
||||
}
|
||||
return FloatImm{Text: text, Neg: neg}, nil
|
||||
}
|
||||
return nil, fmt.Errorf("non-integer immediate not supported")
|
||||
|
||||
case ast.OpAddr:
|
||||
a := op.Addr
|
||||
|
||||
// A bracketed register range, [Z0-Z3]: the four-register source of
|
||||
// the 4FMAPS/4VNNIW families. The range must span four consecutive
|
||||
// same-width vector registers, exactly what the toolchain's parser
|
||||
// takes; the EVEX quad-register emit path reads the low end.
|
||||
if a.Range != nil {
|
||||
lo, ok := ParseReg(a.Range.Lo)
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("unknown register %q in range", a.Range.Lo)
|
||||
}
|
||||
hi, ok := ParseReg(a.Range.Hi)
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("unknown register %q in range", a.Range.Hi)
|
||||
}
|
||||
if !lo.isVec() || lo.size != hi.size {
|
||||
return nil, fmt.Errorf("register range %q must span four same-width vector registers", op.Raw)
|
||||
}
|
||||
if hi.idx != lo.idx+3 {
|
||||
return nil, fmt.Errorf("register range %q must span four consecutive registers", op.Raw)
|
||||
}
|
||||
return RegList{Lo: lo, Hi: hi}, nil
|
||||
}
|
||||
|
||||
// FP-relative: x+N(FP) → (N + fpAdjust)(SP). The offset N lives in the
|
||||
// symbol, not the address displacement.
|
||||
if a.Sym != nil && a.Sym.Pseudo == "FP" {
|
||||
@@ -822,12 +1254,44 @@ func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Op
|
||||
|
||||
// Memory with a real base register: (base), off(base), (base)(index*scale).
|
||||
if a.Base != "" {
|
||||
// Segment-absolute: 0x30(GS) and 0x28(FS), the windows TLS
|
||||
// spellings. The segment override prefixes a disp32 absolute
|
||||
// reference with no relocation.
|
||||
if a.Base == "GS" || a.Base == "FS" {
|
||||
seg := byte(0x64)
|
||||
if a.Base == "GS" {
|
||||
seg = 0x65
|
||||
}
|
||||
return SegAbs{Disp: a.Offset, Size: size, Seg: seg}, nil
|
||||
}
|
||||
base, ok := ParseReg(a.Base)
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("unknown base register %q", a.Base)
|
||||
}
|
||||
m := Mem{Base: base, Disp: a.Offset, HasBase: true, Size: size}
|
||||
if a.Index != "" {
|
||||
if a.Index == "TLS" {
|
||||
// off(base)(TLS*1): the thread-local annotation. The
|
||||
// one-instruction TLS form folds it to off(TLS), the
|
||||
// segment-prefixed absolute whose disp32 carries an
|
||||
// R_TLS_LE patch site; the base register disappears
|
||||
// from the encoding, exactly as the toolchain's
|
||||
// progedit rewrites the address.
|
||||
seg := byte(0x64) // FS on linux, freebsd, plan9
|
||||
if link != nil && link.goos == "windows" {
|
||||
seg = 0x65 // GS
|
||||
}
|
||||
return TLSMem{Disp: a.Offset, Size: size, Seg: seg}, nil
|
||||
}
|
||||
if a.Index == "GS" || a.Index == "FS" {
|
||||
// 0(CX)(GS): the segment annotation rides the base
|
||||
// access as the override prefix.
|
||||
m.Seg = 0x64
|
||||
if a.Index == "GS" {
|
||||
m.Seg = 0x65
|
||||
}
|
||||
return m, nil
|
||||
}
|
||||
idx, ok := ParseReg(a.Index)
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("unknown index register %q", a.Index)
|
||||
@@ -838,6 +1302,21 @@ func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Op
|
||||
}
|
||||
return m, nil
|
||||
}
|
||||
// Index-only memory: the VSIB form the gather/scatter families
|
||||
// read, 8(X4*1). A scaled vector index addresses memory with no
|
||||
// base register; the mod=00 SIB with base field 101 carries it.
|
||||
if a.Index != "" {
|
||||
idx, ok := ParseReg(a.Index)
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("unknown index register %q", a.Index)
|
||||
}
|
||||
return Mem{Index: idx, Scale: a.Scale, Disp: a.Offset, HasIndex: true, Size: size}, nil
|
||||
}
|
||||
// A bare displacement with no base: the absolute address form,
|
||||
// MOVL $0xf1, 0xf1. No segment and no relocation.
|
||||
if a.Sym == nil && a.Base == "" && a.Index == "" && a.HasOff {
|
||||
return SegAbs{Disp: a.Offset, Size: size}, nil
|
||||
}
|
||||
// Bare register.
|
||||
if a.Sym != nil && a.Sym.Pseudo == "" && a.Sym.Name != "" {
|
||||
if r, ok := ParseReg(a.Sym.Name); ok {
|
||||
@@ -848,3 +1327,37 @@ func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Op
|
||||
}
|
||||
return nil, fmt.Errorf("unsupported operand")
|
||||
}
|
||||
|
||||
// floatImmText recovers a floating-point immediate's magnitude and sign from
|
||||
// the parsed operand. The ordinary spellings arrive in Imm.Float; the
|
||||
// parenthesised $(-1.0) leaves the Immediate empty, because the integer
|
||||
// folder cannot read it, and only the verbatim operand text still carries
|
||||
// the value. Anything that is not a number a float parser accepts reports
|
||||
// not-ok, so every other shape keeps its existing diagnostic.
|
||||
func floatImmText(op *ast.Operand) (text string, neg bool, ok bool) {
|
||||
if op.Imm.Float != "" {
|
||||
return op.Imm.Float, op.Imm.Neg, true
|
||||
}
|
||||
if op.Imm.HasVal || op.Imm.Str != "" || op.Imm.Sym != nil {
|
||||
return "", false, false
|
||||
}
|
||||
// joinRaw spaced the token texts; the compact spelling is what matters.
|
||||
compact := strings.ReplaceAll(op.Raw, " ", "")
|
||||
inner, ok := strings.CutPrefix(compact, "$(")
|
||||
if !ok || !strings.HasSuffix(inner, ")") {
|
||||
return "", false, false
|
||||
}
|
||||
inner = strings.TrimSuffix(inner, ")")
|
||||
inner = strings.TrimPrefix(inner, "+")
|
||||
if s, ok := strings.CutPrefix(inner, "-"); ok {
|
||||
neg = true
|
||||
inner = s
|
||||
}
|
||||
if inner == "" || !strings.ContainsAny(inner, "0123456789") {
|
||||
return "", false, false
|
||||
}
|
||||
if _, err := strconv.ParseFloat(inner, 64); err != nil {
|
||||
return "", false, false
|
||||
}
|
||||
return inner, neg, true
|
||||
}
|
||||
|
||||
@@ -439,3 +439,157 @@ func TestSubSPEncodings(t *testing.T) {
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestAssemblePseudoStatements runs LOCK/REP, BYTE/WORD and END through the
|
||||
// full statement pipeline, pinned against go tool asm (Go 1.27, amd64). It
|
||||
// asserts the three behaviours the toolchain shows: each prefix statement is
|
||||
// a standalone byte with a PC of its own (so a label placed on the LOCK
|
||||
// points at the F0), the data pseudo-ops write their literal bytes inline,
|
||||
// and END terminates nothing (the statements after it still belong to the
|
||||
// function and carry no trace of it).
|
||||
func TestAssemblePseudoStatements(t *testing.T) {
|
||||
fn := firstText(t, `
|
||||
#include "textflag.h"
|
||||
TEXT ·pseudo(SB), NOSPLIT, $0-0
|
||||
pfx:
|
||||
LOCK
|
||||
CMPXCHGQ AX, (BX)
|
||||
REP
|
||||
MOVSQ
|
||||
BYTE $0x0f
|
||||
BYTE $0x1f
|
||||
WORD $0x1234
|
||||
END
|
||||
BYTE $0x02
|
||||
RET
|
||||
`)
|
||||
code, labels, err := Assemble(fn)
|
||||
if err != nil {
|
||||
t.Fatalf("Assemble: %v", err)
|
||||
}
|
||||
// go tool asm: f0 480fb103 f3 48a5 0f 1f 3412 02 c3
|
||||
want := []byte{
|
||||
0xf0,
|
||||
0x48, 0x0f, 0xb1, 0x03,
|
||||
0xf3, 0x48, 0xa5,
|
||||
0x0f, 0x1f, 0x34, 0x12,
|
||||
0x02, 0xc3,
|
||||
}
|
||||
if hexBytes(code) != hexBytes(want) {
|
||||
t.Errorf("pseudo statements:\n got: %s\n want: %s", hexBytes(code), hexBytes(want))
|
||||
}
|
||||
// The label sits on the LOCK byte, exactly where the toolchain's PC
|
||||
// listing puts it.
|
||||
if off := labels["pfx"]; off != 0 {
|
||||
t.Errorf("label pfx = %d, want 0 (the LOCK's own byte)", off)
|
||||
}
|
||||
// The trailing BYTE lands where the layout says: after the 8 bytes of
|
||||
// LOCK, CMPXCHGQ, REP and MOVSQ plus the 4 data bytes, END contributing
|
||||
// none.
|
||||
if code[12] != 0x02 {
|
||||
t.Errorf("byte at 12 = %02x, want 02 (the BYTE after END)", code[12])
|
||||
}
|
||||
}
|
||||
|
||||
// TestAssembleAdjspBalance pins the toolchain's push/pop balance rule over
|
||||
// ADJSP: the straight-line sum of the adjustments must be zero at each
|
||||
// RET, branches in between counting for nothing (verified against go tool
|
||||
// asm: ADJSP $16 before a RET is reported as "unbalanced PUSH/POP", a
|
||||
// $16/$-16 pair with a JMP in between assembles).
|
||||
func TestAssembleAdjspBalance(t *testing.T) {
|
||||
// Balanced pair with a branch in between, bytes pinned from go tool asm.
|
||||
fn := firstText(t, `
|
||||
#include "textflag.h"
|
||||
TEXT ·adjsp(SB), NOSPLIT, $0-0
|
||||
ADJSP $16
|
||||
JMP body
|
||||
body:
|
||||
ADJSP $-16
|
||||
RET
|
||||
`)
|
||||
code, _, err := Assemble(fn)
|
||||
if err != nil {
|
||||
t.Fatalf("Assemble: %v", err)
|
||||
}
|
||||
want := []byte{0x48, 0x83, 0xEC, 0x10, 0xEB, 0x00, 0x48, 0x83, 0xC4, 0x10, 0xC3}
|
||||
if hexBytes(code) != hexBytes(want) {
|
||||
t.Errorf("adjsp pair:\n got: %s\n want: %s", hexBytes(code), hexBytes(want))
|
||||
}
|
||||
|
||||
// Unbalanced at the RET: the toolchain diagnoses, so must we.
|
||||
_, _, err = Assemble(firstText(t, `
|
||||
#include "textflag.h"
|
||||
TEXT ·unbalanced(SB), NOSPLIT, $0-0
|
||||
ADJSP $16
|
||||
RET
|
||||
`))
|
||||
if err == nil || !strings.Contains(err.Error(), "unbalanced PUSH/POP") {
|
||||
t.Errorf("unbalanced ADJSP: err = %v, want unbalanced PUSH/POP", err)
|
||||
}
|
||||
|
||||
// The check runs per RET: a closed pair before the first RET does not
|
||||
// excuse an open adjustment before the second.
|
||||
_, _, err = Assemble(firstText(t, `
|
||||
#include "textflag.h"
|
||||
TEXT ·tworet(SB), NOSPLIT, $0-0
|
||||
ADJSP $8
|
||||
ADJSP $-8
|
||||
RET
|
||||
mid:
|
||||
ADJSP $8
|
||||
RET
|
||||
`))
|
||||
if err == nil || !strings.Contains(err.Error(), "unbalanced PUSH/POP") {
|
||||
t.Errorf("second RET with open ADJSP: err = %v, want unbalanced PUSH/POP", err)
|
||||
}
|
||||
|
||||
// A framed function: the assembler's own prologue and epilogue
|
||||
// contribute matching deltas, so the pair in the body still balances,
|
||||
// and the bytes match go tool asm end to end.
|
||||
fn = firstText(t, `
|
||||
#include "textflag.h"
|
||||
TEXT ·framed(SB), $16-8
|
||||
ADJSP $8
|
||||
ADJSP $-8
|
||||
RET
|
||||
`)
|
||||
code, _, err = Assemble(fn)
|
||||
if err != nil {
|
||||
t.Fatalf("Assemble framed: %v", err)
|
||||
}
|
||||
want = []byte{
|
||||
0x55, 0x48, 0x89, 0xE5, 0x48, 0x83, 0xEC, 0x10, // prologue
|
||||
0x48, 0x83, 0xEC, 0x08, // ADJSP $8
|
||||
0x48, 0x83, 0xC4, 0x08, // ADJSP $-8
|
||||
0x48, 0x83, 0xC4, 0x10, 0x5D, // epilogue
|
||||
0xC3,
|
||||
}
|
||||
if hexBytes(code) != hexBytes(want) {
|
||||
t.Errorf("framed adjsp:\n got: %s\n want: %s", hexBytes(code), hexBytes(want))
|
||||
}
|
||||
}
|
||||
|
||||
// TestAssembleRegRange pins the bracketed register range at the statement
|
||||
// level: exactly four consecutive same-width vector registers assemble, the
|
||||
// toolchain's rejected shapes all report an error.
|
||||
func TestAssembleRegRange(t *testing.T) {
|
||||
asm := func(t *testing.T, op string) ([]byte, error) {
|
||||
t.Helper()
|
||||
f, errs := parser.Parse("f_amd64.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tV4FMADDPS 17(SP), "+op+", K2, Z0\n\tRET\n")
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse %s: %v", op, errs)
|
||||
}
|
||||
code, _, err := Assemble(f.Decls[0].(*ast.Text))
|
||||
return code, err
|
||||
}
|
||||
for _, op := range []string{"[Z0-Z3]", "[Z4-Z7]", "[Z28-Z31]"} {
|
||||
if _, err := asm(t, op); err != nil {
|
||||
t.Errorf("%s: %v", op, err)
|
||||
}
|
||||
}
|
||||
for _, op := range []string{"[Z0-Z4]", "[Z0-Z2]", "[Z0-Z0]", "[Z4-Z0]", "[Z1-Z0]", "[AX-Z3]", "[Z0-AX]"} {
|
||||
if _, err := asm(t, op); err == nil {
|
||||
t.Errorf("%s: assembled, want an error", op)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
+65
-10
@@ -43,6 +43,9 @@ const (
|
||||
stInfoShift = 4
|
||||
|
||||
rX8664PC32 = 2
|
||||
// R_X86_64_32 (debug/elf): the absolute 32-bit address of a symbol, the
|
||||
// R_ADDR shape a 4-byte DATA field carries.
|
||||
rX8664Abs32 = 10
|
||||
// R_X86_64_TPOFF32 (debug/elf): the local-exec TLS offset the stack
|
||||
// guard loads from FS. 20 is R_X86_64_TLSLD, a different relocation.
|
||||
rX8664TPOFF32 = 23
|
||||
@@ -158,6 +161,50 @@ func (img *Image) ELFObject() ([]byte, error) {
|
||||
}
|
||||
}
|
||||
|
||||
// The data symbols' symbol-valued DATA fields ("DATA s+0(SB)/8,
|
||||
// $other(SB)") become .rela.data entries: an absolute relocation of the
|
||||
// DATA line's width at the field's data-section offset, S + A with no
|
||||
// PC term. Widths 4 and 8 have ELF relocation shapes; narrower fields
|
||||
// cannot hold an address, so they are refused rather than truncated.
|
||||
var dataRelas []elfRela
|
||||
for _, d := range img.DataSyms {
|
||||
for _, r := range d.Relocs {
|
||||
idx, ok := symIdx[r.Name]
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("data relocation references unknown symbol %q", r.Name)
|
||||
}
|
||||
var typ uint32
|
||||
switch r.Siz {
|
||||
case 8:
|
||||
typ = rX8664Abs64
|
||||
case 4:
|
||||
typ = rX8664Abs32
|
||||
default:
|
||||
return nil, fmt.Errorf("DATA %q: a symbol value of width %d has no ELF relocation", d.Name, r.Siz)
|
||||
}
|
||||
dataRelas = append(dataRelas, elfRela{
|
||||
off: uint64(d.Offset + r.Off),
|
||||
sym: idx,
|
||||
typ: typ,
|
||||
addend: r.Addend,
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// Section presence: .rela.text only when there are code relocations,
|
||||
// .rela.data only when a DATA line holds a symbol value.
|
||||
hasRela := len(relas) > 0
|
||||
hasDataRela := len(dataRelas) > 0
|
||||
nSections := 6 // NULL, .text, .data, .symtab, .strtab, .shstrtab
|
||||
if hasRela {
|
||||
nSections++
|
||||
}
|
||||
if hasDataRela {
|
||||
nSections++
|
||||
}
|
||||
secSymtab, secStrtab := 3, 4
|
||||
secShstr := nSections - 1
|
||||
|
||||
// Serialise the string tables.
|
||||
stNames := newElfStrtab()
|
||||
for _, s := range syms {
|
||||
@@ -167,19 +214,13 @@ func (img *Image) ELFObject() ([]byte, error) {
|
||||
for _, n := range []string{".text", ".data", ".symtab", ".strtab", ".rela.text", ".shstrtab"} {
|
||||
stSections.add(n)
|
||||
}
|
||||
if hasDataRela {
|
||||
stSections.add(".rela.data")
|
||||
}
|
||||
for _, n := range dwarfSectionNames {
|
||||
stSections.add(n)
|
||||
}
|
||||
|
||||
// Section presence: .rela.text only when there are relocations.
|
||||
hasRela := len(relas) > 0
|
||||
nSections := 6 // NULL, .text, .data, .symtab, .strtab, .shstrtab
|
||||
if hasRela {
|
||||
nSections = 7
|
||||
}
|
||||
secSymtab, secStrtab := 3, 4
|
||||
secShstr := nSections - 1
|
||||
|
||||
// Lay the file out: header, section data, section headers.
|
||||
var out []byte
|
||||
out = append(out, make([]byte, 64)...) // ELF header, filled last
|
||||
@@ -214,7 +255,7 @@ func (img *Image) ELFObject() ([]byte, error) {
|
||||
strtabOff := len(out)
|
||||
out = append(out, stNames.bytes()...)
|
||||
|
||||
var relaOff int
|
||||
var relaOff, relaDataOff int
|
||||
if hasRela {
|
||||
align(8)
|
||||
relaOff = len(out)
|
||||
@@ -226,6 +267,17 @@ func (img *Image) ELFObject() ([]byte, error) {
|
||||
out = append(out, b[:]...)
|
||||
}
|
||||
}
|
||||
if hasDataRela {
|
||||
align(8)
|
||||
relaDataOff = len(out)
|
||||
for _, r := range dataRelas {
|
||||
var b [24]byte
|
||||
le.PutUint64(b[0:], r.off)
|
||||
le.PutUint64(b[8:], uint64(r.sym)<<32|uint64(r.typ))
|
||||
le.PutUint64(b[16:], uint64(r.addend))
|
||||
out = append(out, b[:]...)
|
||||
}
|
||||
}
|
||||
|
||||
shstrOff := len(out)
|
||||
out = append(out, stSections.bytes()...)
|
||||
@@ -284,6 +336,9 @@ func (img *Image) ELFObject() ([]byte, error) {
|
||||
if hasRela {
|
||||
putSh(".rela.text", shtRela, 0, relaOff, 24*len(relas), secSymtab, secText, 8, 24)
|
||||
}
|
||||
if hasDataRela {
|
||||
putSh(".rela.data", shtRela, 0, relaDataOff, 24*len(dataRelas), secSymtab, secData, 8, 24)
|
||||
}
|
||||
putSh(".shstrtab", shtStrtab, 0, shstrOff, len(stSections.bytes()), 0, 0, 1, 0)
|
||||
|
||||
// DWARF section headers; their indices follow the write order.
|
||||
|
||||
+108
@@ -750,3 +750,111 @@ func readFormSkip(t *testing.T, r *ulebIter, form uint64) {
|
||||
t.Fatalf("unsupported form %#x", form)
|
||||
}
|
||||
}
|
||||
|
||||
// TestELFObjectDataRelocation checks that a symbol-valued DATA field ("DATA
|
||||
// s+0(SB)/8, $other(SB)") reaches the ELF object as a .rela.data entry: an
|
||||
// absolute 64-bit relocation at the field's offset within .data, against
|
||||
// the named symbol, external targets included.
|
||||
func TestELFObjectDataRelocation(t *testing.T) {
|
||||
f, errs := parser.Parse("t_amd64.s", `#include "textflag.h"
|
||||
TEXT ·Keep(SB), NOSPLIT, $0-8
|
||||
RET
|
||||
GLOBL holder(SB), NOPTR, $24
|
||||
DATA holder+0(SB)/8, $·Keep+5(SB)
|
||||
DATA holder+8(SB)/8, $holder(SB)
|
||||
DATA holder+16(SB)/8, $extvar(SB)
|
||||
`)
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
img, err := AssembleFile(f)
|
||||
if err != nil {
|
||||
t.Fatalf("assemble: %v", err)
|
||||
}
|
||||
obj, err := img.ELFObject()
|
||||
if err != nil {
|
||||
t.Fatalf("ELFObject: %v", err)
|
||||
}
|
||||
ef, err := elf.NewFile(bytes.NewReader(obj))
|
||||
if err != nil {
|
||||
t.Fatalf("parse emitted object: %v", err)
|
||||
}
|
||||
defer ef.Close()
|
||||
relaData := ef.Section(".rela.data")
|
||||
if relaData == nil {
|
||||
t.Fatal("missing .rela.data section")
|
||||
}
|
||||
if relaData.Link == 0 || ef.Sections[relaData.Link].Name != ".symtab" {
|
||||
t.Errorf(".rela.data sh_link = %d, want the .symtab index", relaData.Link)
|
||||
}
|
||||
if ef.Sections[relaData.Info].Name != ".data" {
|
||||
t.Errorf(".rela.data sh_info = %d, want the .data index", relaData.Info)
|
||||
}
|
||||
relas, err := relaData.Data()
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
var got []struct {
|
||||
off uint64
|
||||
sym uint32
|
||||
typ uint32
|
||||
addend int64
|
||||
}
|
||||
for i := 0; i+24 <= len(relas); i += 24 {
|
||||
got = append(got, struct {
|
||||
off uint64
|
||||
sym uint32
|
||||
typ uint32
|
||||
addend int64
|
||||
}{
|
||||
off: binary.LittleEndian.Uint64(relas[i:]),
|
||||
// r_info packs the type in the low dword and the symbol index
|
||||
// in the high dword.
|
||||
typ: binary.LittleEndian.Uint32(relas[i+8:]),
|
||||
sym: binary.LittleEndian.Uint32(relas[i+12:]),
|
||||
addend: int64(binary.LittleEndian.Uint64(relas[i+16:])),
|
||||
})
|
||||
}
|
||||
// debug/elf hides the table's null entry, so raw index s names syms[s-1].
|
||||
syms, err := ef.Symbols()
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
name := func(idx uint32) string {
|
||||
if idx >= 1 && int(idx) <= len(syms) {
|
||||
return syms[idx-1].Name
|
||||
}
|
||||
return ""
|
||||
}
|
||||
// The offsets are data-section-relative: the field's DATA offset plus
|
||||
// the symbol's position in .data (the layout aligns each symbol to 16).
|
||||
base := uint64(0)
|
||||
for _, d := range img.DataSyms {
|
||||
if d.Name == "holder" {
|
||||
base = uint64(d.Offset)
|
||||
}
|
||||
}
|
||||
want := []struct {
|
||||
off uint64
|
||||
typ uint32
|
||||
addend int64
|
||||
target string
|
||||
}{
|
||||
{off: base + 0, typ: uint32(elf.R_X86_64_64), addend: 5, target: "Keep"},
|
||||
{off: base + 8, typ: uint32(elf.R_X86_64_64), addend: 0, target: "holder"},
|
||||
{off: base + 16, typ: uint32(elf.R_X86_64_64), addend: 0, target: "extvar"},
|
||||
}
|
||||
if len(got) != len(want) {
|
||||
t.Fatalf(".rela.data entries = %d, want %d", len(got), len(want))
|
||||
}
|
||||
for i, w := range want {
|
||||
g := got[i]
|
||||
if g.off != w.off || g.typ != w.typ || g.addend != w.addend {
|
||||
t.Errorf("entry %d = {off %d typ %d addend %d}, want {off %d typ %d addend %d}",
|
||||
i, g.off, g.typ, g.addend, w.off, w.typ, w.addend)
|
||||
}
|
||||
if n := name(g.sym); n != w.target {
|
||||
t.Errorf("entry %d names %q, want %q", i, n, w.target)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
+68
-11
@@ -18,12 +18,16 @@ const (
|
||||
rArm64AddAbsLo12NC = 277 // R_AARCH64_ADD_ABS_LO12_NC (ADD page offset)
|
||||
rArm64Call26 = 283 // R_AARCH64_CALL26 (BL instruction)
|
||||
rArm64Ldst64Lo12NC = 286 // R_AARCH64_LDST64_ABS_LO12_NC (64-bit LDR/STR page offset)
|
||||
// R_AARCH64_ABS32 (debug/elf 258): the absolute 32-bit address of a
|
||||
// symbol, the R_ADDR shape a 4-byte DATA field carries. ABS64 (257)
|
||||
// lives with the DWARF fixup constants as rAARCH64Abs64.
|
||||
rArm64Abs32 = 258
|
||||
)
|
||||
|
||||
// ELFAARCH64Object returns the image as an ELF64 relocatable object file for
|
||||
// AArch64 (EM_AARCH64, 64-bit, little-endian). The structure mirrors the
|
||||
// amd64 and RISC-V ELF emitters: .text, .data, .symtab, .strtab and an
|
||||
// optional .rela.text.
|
||||
// amd64 and RISC-V ELF emitters: .text, .data, .symtab, .strtab, an
|
||||
// optional .rela.text and an optional .rela.data.
|
||||
func (img *Image) ELFAARCH64Object() ([]byte, error) {
|
||||
le := binary.LittleEndian
|
||||
|
||||
@@ -133,6 +137,50 @@ func (img *Image) ELFAARCH64Object() ([]byte, error) {
|
||||
}
|
||||
}
|
||||
|
||||
// The data symbols' symbol-valued DATA fields ("DATA s+0(SB)/8,
|
||||
// $other(SB)") become .rela.data entries: an absolute relocation of the
|
||||
// DATA line's width at the field's data-section offset, S + A with no
|
||||
// PC term. Widths 4 and 8 have ELF relocation shapes; narrower fields
|
||||
// cannot hold an address, so they are refused rather than truncated.
|
||||
var dataRelas []elfRela
|
||||
for _, d := range img.DataSyms {
|
||||
for _, r := range d.Relocs {
|
||||
idx, ok := symIdx[r.Name]
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("data relocation references unknown symbol %q", r.Name)
|
||||
}
|
||||
var typ uint32
|
||||
switch r.Siz {
|
||||
case 8:
|
||||
typ = rAARCH64Abs64
|
||||
case 4:
|
||||
typ = rArm64Abs32
|
||||
default:
|
||||
return nil, fmt.Errorf("DATA %q: a symbol value of width %d has no ELF relocation", d.Name, r.Siz)
|
||||
}
|
||||
dataRelas = append(dataRelas, elfRela{
|
||||
off: uint64(d.Offset + r.Off),
|
||||
sym: idx,
|
||||
typ: typ,
|
||||
addend: r.Addend,
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// Section presence: .rela.text only when there are code relocations,
|
||||
// .rela.data only when a DATA line holds a symbol value.
|
||||
hasRela := len(relas) > 0
|
||||
hasDataRela := len(dataRelas) > 0
|
||||
nSections := 6
|
||||
if hasRela {
|
||||
nSections++
|
||||
}
|
||||
if hasDataRela {
|
||||
nSections++
|
||||
}
|
||||
secSymtab, secStrtab := 3, 4
|
||||
secShstr := nSections - 1
|
||||
|
||||
// String tables.
|
||||
stNames := newElfStrtab()
|
||||
for _, s := range syms {
|
||||
@@ -142,18 +190,13 @@ func (img *Image) ELFAARCH64Object() ([]byte, error) {
|
||||
for _, n := range []string{".text", ".data", ".symtab", ".strtab", ".rela.text", ".shstrtab"} {
|
||||
stSections.add(n)
|
||||
}
|
||||
if hasDataRela {
|
||||
stSections.add(".rela.data")
|
||||
}
|
||||
for _, n := range dwarfSectionNames {
|
||||
stSections.add(n)
|
||||
}
|
||||
|
||||
hasRela := len(relas) > 0
|
||||
nSections := 6
|
||||
if hasRela {
|
||||
nSections = 7
|
||||
}
|
||||
secSymtab, secStrtab := 3, 4
|
||||
secShstr := nSections - 1
|
||||
|
||||
// Layout.
|
||||
var out []byte
|
||||
out = append(out, make([]byte, 64)...)
|
||||
@@ -188,7 +231,7 @@ func (img *Image) ELFAARCH64Object() ([]byte, error) {
|
||||
strtabOff := len(out)
|
||||
out = append(out, stNames.bytes()...)
|
||||
|
||||
var relaOff int
|
||||
var relaOff, relaDataOff int
|
||||
if hasRela {
|
||||
align(8)
|
||||
relaOff = len(out)
|
||||
@@ -200,6 +243,17 @@ func (img *Image) ELFAARCH64Object() ([]byte, error) {
|
||||
out = append(out, b[:]...)
|
||||
}
|
||||
}
|
||||
if hasDataRela {
|
||||
align(8)
|
||||
relaDataOff = len(out)
|
||||
for _, r := range dataRelas {
|
||||
var b [24]byte
|
||||
le.PutUint64(b[0:], r.off)
|
||||
le.PutUint64(b[8:], uint64(r.sym)<<32|uint64(r.typ))
|
||||
le.PutUint64(b[16:], uint64(r.addend))
|
||||
out = append(out, b[:]...)
|
||||
}
|
||||
}
|
||||
|
||||
shstrOff := len(out)
|
||||
out = append(out, stSections.bytes()...)
|
||||
@@ -257,6 +311,9 @@ func (img *Image) ELFAARCH64Object() ([]byte, error) {
|
||||
if hasRela {
|
||||
putSh(".rela.text", shtRela, 0, relaOff, 24*len(relas), secSymtab, secText, 8, 24)
|
||||
}
|
||||
if hasDataRela {
|
||||
putSh(".rela.data", shtRela, 0, relaDataOff, 24*len(dataRelas), secSymtab, secData, 8, 24)
|
||||
}
|
||||
putSh(".shstrtab", shtStrtab, 0, shstrOff, len(stSections.bytes()), 0, 0, 1, 0)
|
||||
// DWARF section headers; their indices follow the write order.
|
||||
if dw != nil {
|
||||
|
||||
@@ -197,3 +197,118 @@ TEXT ·add(SB), NOSPLIT, $0-24
|
||||
t.Error("unexpected .rela.text section when there are no relocations")
|
||||
}
|
||||
}
|
||||
|
||||
// TestELFAARCH64ObjectDataRelocation checks that a symbol-valued DATA field
|
||||
// ("DATA s+0(SB)/8, $other(SB)") reaches the AArch64 ELF object as a
|
||||
// .rela.data entry: an R_AARCH64_ABS64 (ABS32 for a width-4 field) at the
|
||||
// field's offset within .data, against the named symbol, external targets
|
||||
// included.
|
||||
func TestELFAARCH64ObjectDataRelocation(t *testing.T) {
|
||||
f, errs := parser.Parse("t_arm64.s", `#include "textflag.h"
|
||||
TEXT ·Keep(SB), NOSPLIT, $0-0
|
||||
RET
|
||||
GLOBL holder(SB), NOPTR, $32
|
||||
DATA holder+0(SB)/8, $·Keep+5(SB)
|
||||
DATA holder+8(SB)/8, $holder(SB)
|
||||
DATA holder+16(SB)/8, $extvar(SB)
|
||||
DATA holder+24(SB)/4, $Keep(SB)
|
||||
`)
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
img, err := AssembleFileARM64(f)
|
||||
if err != nil {
|
||||
t.Fatalf("AssembleFileARM64: %v", err)
|
||||
}
|
||||
obj, err := img.ELFAARCH64Object()
|
||||
if err != nil {
|
||||
t.Fatalf("ELFAARCH64Object: %v", err)
|
||||
}
|
||||
checkELFSectionAccounting(t, obj)
|
||||
ef, err := elf.NewFile(bytes.NewReader(obj))
|
||||
if err != nil {
|
||||
t.Fatalf("parse emitted object: %v", err)
|
||||
}
|
||||
defer ef.Close()
|
||||
relaData := ef.Section(".rela.data")
|
||||
if relaData == nil {
|
||||
t.Fatal("missing .rela.data section")
|
||||
}
|
||||
if relaData.Type != elf.SHT_RELA {
|
||||
t.Errorf(".rela.data type = %v, want SHT_RELA", relaData.Type)
|
||||
}
|
||||
if relaData.Link == 0 || ef.Sections[relaData.Link].Name != ".symtab" {
|
||||
t.Errorf(".rela.data sh_link = %d, want the .symtab index", relaData.Link)
|
||||
}
|
||||
if ef.Sections[relaData.Info].Name != ".data" {
|
||||
t.Errorf(".rela.data sh_info = %d, want the .data index", relaData.Info)
|
||||
}
|
||||
relas, err := relaData.Data()
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
var got []struct {
|
||||
off uint64
|
||||
sym uint32
|
||||
typ uint32
|
||||
addend int64
|
||||
}
|
||||
for i := 0; i+24 <= len(relas); i += 24 {
|
||||
got = append(got, struct {
|
||||
off uint64
|
||||
sym uint32
|
||||
typ uint32
|
||||
addend int64
|
||||
}{
|
||||
off: binary.LittleEndian.Uint64(relas[i:]),
|
||||
// r_info packs the type in the low dword and the symbol index
|
||||
// in the high dword.
|
||||
typ: binary.LittleEndian.Uint32(relas[i+8:]),
|
||||
sym: binary.LittleEndian.Uint32(relas[i+12:]),
|
||||
addend: int64(binary.LittleEndian.Uint64(relas[i+16:])),
|
||||
})
|
||||
}
|
||||
// debug/elf hides the table's null entry, so raw index s names syms[s-1].
|
||||
syms, err := ef.Symbols()
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
name := func(idx uint32) string {
|
||||
if idx >= 1 && int(idx) <= len(syms) {
|
||||
return syms[idx-1].Name
|
||||
}
|
||||
return ""
|
||||
}
|
||||
// The offsets are data-section-relative: the field's DATA offset plus
|
||||
// the symbol's position in .data (the layout aligns each symbol to 16).
|
||||
base := uint64(0)
|
||||
for _, d := range img.DataSyms {
|
||||
if d.Name == "holder" {
|
||||
base = uint64(d.Offset)
|
||||
}
|
||||
}
|
||||
want := []struct {
|
||||
off uint64
|
||||
typ uint32
|
||||
addend int64
|
||||
target string
|
||||
}{
|
||||
{off: base + 0, typ: uint32(elf.R_AARCH64_ABS64), addend: 5, target: "Keep"},
|
||||
{off: base + 8, typ: uint32(elf.R_AARCH64_ABS64), addend: 0, target: "holder"},
|
||||
{off: base + 16, typ: uint32(elf.R_AARCH64_ABS64), addend: 0, target: "extvar"},
|
||||
{off: base + 24, typ: uint32(elf.R_AARCH64_ABS32), addend: 0, target: "Keep"},
|
||||
}
|
||||
if len(got) != len(want) {
|
||||
t.Fatalf(".rela.data entries = %d, want %d", len(got), len(want))
|
||||
}
|
||||
for i, w := range want {
|
||||
g := got[i]
|
||||
if g.off != w.off || g.typ != w.typ || g.addend != w.addend {
|
||||
t.Errorf("entry %d = {off %d typ %d addend %d}, want {off %d typ %d addend %d}",
|
||||
i, g.off, g.typ, g.addend, w.off, w.typ, w.addend)
|
||||
}
|
||||
if n := name(g.sym); n != w.target {
|
||||
t.Errorf("entry %d names %q, want %q", i, n, w.target)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
+68
-11
@@ -23,12 +23,16 @@ const (
|
||||
rLarchPCALAHI20 = 71 // R_LARCH_PCALA_HI20 (pcalau12i)
|
||||
rLarchPCALALO12 = 72 // R_LARCH_PCALA_LO12 (addi.d/ld/st)
|
||||
rLarchB26 = 66 // R_LARCH_B26 (b/bl, matches the Go linker's mapping)
|
||||
// R_LARCH_32 (debug/elf 1): the absolute 32-bit address of a symbol,
|
||||
// the R_ADDR shape a 4-byte DATA field carries. R_LARCH_64 (2) lives
|
||||
// with the DWARF fixup constants as rLarchAbs64.
|
||||
rLarchAbs32 = 1
|
||||
)
|
||||
|
||||
// ELFLOONG64Object returns the image as an ELF64 relocatable object file for
|
||||
// LoongArch (EM_LOONGARCH, 64-bit, little-endian). The structure mirrors the
|
||||
// amd64 and RISC-V ELF emitters: .text, .data, .symtab, .strtab and an
|
||||
// optional .rela.text.
|
||||
// amd64 and RISC-V ELF emitters: .text, .data, .symtab, .strtab, an
|
||||
// optional .rela.text and an optional .rela.data.
|
||||
func (img *Image) ELFLOONG64Object() ([]byte, error) {
|
||||
le := binary.LittleEndian
|
||||
|
||||
@@ -117,6 +121,50 @@ func (img *Image) ELFLOONG64Object() ([]byte, error) {
|
||||
}
|
||||
}
|
||||
|
||||
// The data symbols' symbol-valued DATA fields ("DATA s+0(SB)/8,
|
||||
// $other(SB)") become .rela.data entries: an absolute relocation of the
|
||||
// DATA line's width at the field's data-section offset, S + A with no
|
||||
// PC term. Widths 4 and 8 have ELF relocation shapes; narrower fields
|
||||
// cannot hold an address, so they are refused rather than truncated.
|
||||
var dataRelas []elfRela
|
||||
for _, d := range img.DataSyms {
|
||||
for _, r := range d.Relocs {
|
||||
idx, ok := symIdx[r.Name]
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("data relocation references unknown symbol %q", r.Name)
|
||||
}
|
||||
var typ uint32
|
||||
switch r.Siz {
|
||||
case 8:
|
||||
typ = rLarchAbs64
|
||||
case 4:
|
||||
typ = rLarchAbs32
|
||||
default:
|
||||
return nil, fmt.Errorf("DATA %q: a symbol value of width %d has no ELF relocation", d.Name, r.Siz)
|
||||
}
|
||||
dataRelas = append(dataRelas, elfRela{
|
||||
off: uint64(d.Offset + r.Off),
|
||||
sym: idx,
|
||||
typ: typ,
|
||||
addend: r.Addend,
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// Section presence: .rela.text only when there are code relocations,
|
||||
// .rela.data only when a DATA line holds a symbol value.
|
||||
hasRela := len(relas) > 0
|
||||
hasDataRela := len(dataRelas) > 0
|
||||
nSections := 6
|
||||
if hasRela {
|
||||
nSections++
|
||||
}
|
||||
if hasDataRela {
|
||||
nSections++
|
||||
}
|
||||
secSymtab, secStrtab := 3, 4
|
||||
secShstr := nSections - 1
|
||||
|
||||
// String tables.
|
||||
stNames := newElfStrtab()
|
||||
for _, s := range syms {
|
||||
@@ -126,18 +174,13 @@ func (img *Image) ELFLOONG64Object() ([]byte, error) {
|
||||
for _, n := range []string{".text", ".data", ".symtab", ".strtab", ".rela.text", ".shstrtab"} {
|
||||
stSections.add(n)
|
||||
}
|
||||
if hasDataRela {
|
||||
stSections.add(".rela.data")
|
||||
}
|
||||
for _, n := range dwarfSectionNames {
|
||||
stSections.add(n)
|
||||
}
|
||||
|
||||
hasRela := len(relas) > 0
|
||||
nSections := 6
|
||||
if hasRela {
|
||||
nSections = 7
|
||||
}
|
||||
secSymtab, secStrtab := 3, 4
|
||||
secShstr := nSections - 1
|
||||
|
||||
// Layout.
|
||||
var out []byte
|
||||
out = append(out, make([]byte, 64)...)
|
||||
@@ -172,7 +215,7 @@ func (img *Image) ELFLOONG64Object() ([]byte, error) {
|
||||
strtabOff := len(out)
|
||||
out = append(out, stNames.bytes()...)
|
||||
|
||||
var relaOff int
|
||||
var relaOff, relaDataOff int
|
||||
if hasRela {
|
||||
align(8)
|
||||
relaOff = len(out)
|
||||
@@ -184,6 +227,17 @@ func (img *Image) ELFLOONG64Object() ([]byte, error) {
|
||||
out = append(out, b[:]...)
|
||||
}
|
||||
}
|
||||
if hasDataRela {
|
||||
align(8)
|
||||
relaDataOff = len(out)
|
||||
for _, r := range dataRelas {
|
||||
var b [24]byte
|
||||
le.PutUint64(b[0:], r.off)
|
||||
le.PutUint64(b[8:], uint64(r.sym)<<32|uint64(r.typ))
|
||||
le.PutUint64(b[16:], uint64(r.addend))
|
||||
out = append(out, b[:]...)
|
||||
}
|
||||
}
|
||||
|
||||
shstrOff := len(out)
|
||||
out = append(out, stSections.bytes()...)
|
||||
@@ -239,6 +293,9 @@ func (img *Image) ELFLOONG64Object() ([]byte, error) {
|
||||
if hasRela {
|
||||
putSh(".rela.text", shtRela, 0, relaOff, 24*len(relas), secSymtab, secText, 8, 24)
|
||||
}
|
||||
if hasDataRela {
|
||||
putSh(".rela.data", shtRela, 0, relaDataOff, 24*len(dataRelas), secSymtab, secData, 8, 24)
|
||||
}
|
||||
putSh(".shstrtab", shtStrtab, 0, shstrOff, len(stSections.bytes()), 0, 0, 1, 0)
|
||||
// DWARF section headers; their indices follow the write order.
|
||||
if dw != nil {
|
||||
|
||||
@@ -245,3 +245,118 @@ func TestELFLOONG64BranchRelocation(t *testing.T) {
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestELFLOONG64ObjectDataRelocation checks that a symbol-valued DATA field
|
||||
// ("DATA s+0(SB)/8, $other(SB)") reaches the LoongArch ELF object as a
|
||||
// .rela.data entry: an R_LARCH_64 (R_LARCH_32 for a width-4 field) at the
|
||||
// field's offset within .data, against the named symbol, external targets
|
||||
// included.
|
||||
func TestELFLOONG64ObjectDataRelocation(t *testing.T) {
|
||||
f, errs := parser.Parse("t_loong64.s", `#include "textflag.h"
|
||||
TEXT ·Keep(SB), NOSPLIT, $0-0
|
||||
RET
|
||||
GLOBL holder(SB), NOPTR, $32
|
||||
DATA holder+0(SB)/8, $·Keep+5(SB)
|
||||
DATA holder+8(SB)/8, $holder(SB)
|
||||
DATA holder+16(SB)/8, $extvar(SB)
|
||||
DATA holder+24(SB)/4, $Keep(SB)
|
||||
`)
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
img, err := AssembleFileLOONG64(f)
|
||||
if err != nil {
|
||||
t.Fatalf("AssembleFileLOONG64: %v", err)
|
||||
}
|
||||
obj, err := img.ELFLOONG64Object()
|
||||
if err != nil {
|
||||
t.Fatalf("ELFLOONG64Object: %v", err)
|
||||
}
|
||||
checkELFSectionAccounting(t, obj)
|
||||
ef, err := elf.NewFile(bytes.NewReader(obj))
|
||||
if err != nil {
|
||||
t.Fatalf("parse emitted object: %v", err)
|
||||
}
|
||||
defer ef.Close()
|
||||
relaData := ef.Section(".rela.data")
|
||||
if relaData == nil {
|
||||
t.Fatal("missing .rela.data section")
|
||||
}
|
||||
if relaData.Type != elf.SHT_RELA {
|
||||
t.Errorf(".rela.data type = %v, want SHT_RELA", relaData.Type)
|
||||
}
|
||||
if relaData.Link == 0 || ef.Sections[relaData.Link].Name != ".symtab" {
|
||||
t.Errorf(".rela.data sh_link = %d, want the .symtab index", relaData.Link)
|
||||
}
|
||||
if ef.Sections[relaData.Info].Name != ".data" {
|
||||
t.Errorf(".rela.data sh_info = %d, want the .data index", relaData.Info)
|
||||
}
|
||||
relas, err := relaData.Data()
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
var got []struct {
|
||||
off uint64
|
||||
sym uint32
|
||||
typ uint32
|
||||
addend int64
|
||||
}
|
||||
for i := 0; i+24 <= len(relas); i += 24 {
|
||||
got = append(got, struct {
|
||||
off uint64
|
||||
sym uint32
|
||||
typ uint32
|
||||
addend int64
|
||||
}{
|
||||
off: binary.LittleEndian.Uint64(relas[i:]),
|
||||
// r_info packs the type in the low dword and the symbol index
|
||||
// in the high dword.
|
||||
typ: binary.LittleEndian.Uint32(relas[i+8:]),
|
||||
sym: binary.LittleEndian.Uint32(relas[i+12:]),
|
||||
addend: int64(binary.LittleEndian.Uint64(relas[i+16:])),
|
||||
})
|
||||
}
|
||||
// debug/elf hides the table's null entry, so raw index s names syms[s-1].
|
||||
syms, err := ef.Symbols()
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
name := func(idx uint32) string {
|
||||
if idx >= 1 && int(idx) <= len(syms) {
|
||||
return syms[idx-1].Name
|
||||
}
|
||||
return ""
|
||||
}
|
||||
// The offsets are data-section-relative: the field's DATA offset plus
|
||||
// the symbol's position in .data (the layout aligns each symbol to 16).
|
||||
base := uint64(0)
|
||||
for _, d := range img.DataSyms {
|
||||
if d.Name == "holder" {
|
||||
base = uint64(d.Offset)
|
||||
}
|
||||
}
|
||||
want := []struct {
|
||||
off uint64
|
||||
typ uint32
|
||||
addend int64
|
||||
target string
|
||||
}{
|
||||
{off: base + 0, typ: uint32(elf.R_LARCH_64), addend: 5, target: "Keep"},
|
||||
{off: base + 8, typ: uint32(elf.R_LARCH_64), addend: 0, target: "holder"},
|
||||
{off: base + 16, typ: uint32(elf.R_LARCH_64), addend: 0, target: "extvar"},
|
||||
{off: base + 24, typ: uint32(elf.R_LARCH_32), addend: 0, target: "Keep"},
|
||||
}
|
||||
if len(got) != len(want) {
|
||||
t.Fatalf(".rela.data entries = %d, want %d", len(got), len(want))
|
||||
}
|
||||
for i, w := range want {
|
||||
g := got[i]
|
||||
if g.off != w.off || g.typ != w.typ || g.addend != w.addend {
|
||||
t.Errorf("entry %d = {off %d typ %d addend %d}, want {off %d typ %d addend %d}",
|
||||
i, g.off, g.typ, g.addend, w.off, w.typ, w.addend)
|
||||
}
|
||||
if n := name(g.sym); n != w.target {
|
||||
t.Errorf("entry %d names %q, want %q", i, n, w.target)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
+68
-10
@@ -24,11 +24,16 @@ const (
|
||||
rRISCVPCRELHI20 = 23 // R_RISCV_PCREL_HI20
|
||||
rRISCVPCRELLO12I = 24 // R_RISCV_PCREL_LO12_I
|
||||
rRISCVPCRELLO12S = 25 // R_RISCV_PCREL_LO12_S
|
||||
// R_RISCV_32 (debug/elf 1): the absolute 32-bit address of a symbol,
|
||||
// the R_ADDR shape a 4-byte DATA field carries. R_RISCV_64 (2) lives
|
||||
// with the DWARF fixup constants as rRISCVAbs64.
|
||||
rRISVCAbs32 = 1
|
||||
)
|
||||
|
||||
// ELFRISCVObject returns the image as an ELF64 relocatable object file for
|
||||
// RISC-V (EM_RISCV, 64-bit, little-endian). The structure mirrors the amd64
|
||||
// ELF emission: .text, .data, .symtab, .strtab and optional .rela.text.
|
||||
// ELF emission: .text, .data, .symtab, .strtab, an optional .rela.text and
|
||||
// an optional .rela.data.
|
||||
func (img *Image) ELFRISCVObject() ([]byte, error) {
|
||||
le := binary.LittleEndian
|
||||
|
||||
@@ -129,6 +134,50 @@ func (img *Image) ELFRISCVObject() ([]byte, error) {
|
||||
}
|
||||
}
|
||||
|
||||
// The data symbols' symbol-valued DATA fields ("DATA s+0(SB)/8,
|
||||
// $other(SB)") become .rela.data entries: an absolute relocation of the
|
||||
// DATA line's width at the field's data-section offset, S + A with no
|
||||
// PC term. Widths 4 and 8 have ELF relocation shapes; narrower fields
|
||||
// cannot hold an address, so they are refused rather than truncated.
|
||||
var dataRelas []elfRela
|
||||
for _, d := range img.DataSyms {
|
||||
for _, r := range d.Relocs {
|
||||
idx, ok := symIdx[r.Name]
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("data relocation references unknown symbol %q", r.Name)
|
||||
}
|
||||
var typ uint32
|
||||
switch r.Siz {
|
||||
case 8:
|
||||
typ = rRISCVAbs64
|
||||
case 4:
|
||||
typ = rRISVCAbs32
|
||||
default:
|
||||
return nil, fmt.Errorf("DATA %q: a symbol value of width %d has no ELF relocation", d.Name, r.Siz)
|
||||
}
|
||||
dataRelas = append(dataRelas, elfRela{
|
||||
off: uint64(d.Offset + r.Off),
|
||||
sym: idx,
|
||||
typ: typ,
|
||||
addend: r.Addend,
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// Section presence: .rela.text only when there are code relocations,
|
||||
// .rela.data only when a DATA line holds a symbol value.
|
||||
hasRela := len(relas) > 0
|
||||
hasDataRela := len(dataRelas) > 0
|
||||
nSections := 6
|
||||
if hasRela {
|
||||
nSections++
|
||||
}
|
||||
if hasDataRela {
|
||||
nSections++
|
||||
}
|
||||
secSymtab, secStrtab := 3, 4
|
||||
secShstr := nSections - 1
|
||||
|
||||
// String tables.
|
||||
stNames := newElfStrtab()
|
||||
for _, s := range syms {
|
||||
@@ -138,18 +187,13 @@ func (img *Image) ELFRISCVObject() ([]byte, error) {
|
||||
for _, n := range []string{".text", ".data", ".symtab", ".strtab", ".rela.text", ".shstrtab"} {
|
||||
stSections.add(n)
|
||||
}
|
||||
if hasDataRela {
|
||||
stSections.add(".rela.data")
|
||||
}
|
||||
for _, n := range dwarfSectionNames {
|
||||
stSections.add(n)
|
||||
}
|
||||
|
||||
hasRela := len(relas) > 0
|
||||
nSections := 6
|
||||
if hasRela {
|
||||
nSections = 7
|
||||
}
|
||||
secSymtab, secStrtab := 3, 4
|
||||
secShstr := nSections - 1
|
||||
|
||||
// Layout.
|
||||
var out []byte
|
||||
out = append(out, make([]byte, 64)...)
|
||||
@@ -184,7 +228,7 @@ func (img *Image) ELFRISCVObject() ([]byte, error) {
|
||||
strtabOff := len(out)
|
||||
out = append(out, stNames.bytes()...)
|
||||
|
||||
var relaOff int
|
||||
var relaOff, relaDataOff int
|
||||
if hasRela {
|
||||
align(8)
|
||||
relaOff = len(out)
|
||||
@@ -196,6 +240,17 @@ func (img *Image) ELFRISCVObject() ([]byte, error) {
|
||||
out = append(out, b[:]...)
|
||||
}
|
||||
}
|
||||
if hasDataRela {
|
||||
align(8)
|
||||
relaDataOff = len(out)
|
||||
for _, r := range dataRelas {
|
||||
var b [24]byte
|
||||
le.PutUint64(b[0:], r.off)
|
||||
le.PutUint64(b[8:], uint64(r.sym)<<32|uint64(r.typ))
|
||||
le.PutUint64(b[16:], uint64(r.addend))
|
||||
out = append(out, b[:]...)
|
||||
}
|
||||
}
|
||||
|
||||
shstrOff := len(out)
|
||||
out = append(out, stSections.bytes()...)
|
||||
@@ -251,6 +306,9 @@ func (img *Image) ELFRISCVObject() ([]byte, error) {
|
||||
if hasRela {
|
||||
putSh(".rela.text", shtRela, 0, relaOff, 24*len(relas), secSymtab, secText, 8, 24)
|
||||
}
|
||||
if hasDataRela {
|
||||
putSh(".rela.data", shtRela, 0, relaDataOff, 24*len(dataRelas), secSymtab, secData, 8, 24)
|
||||
}
|
||||
putSh(".shstrtab", shtStrtab, 0, shstrOff, len(stSections.bytes()), 0, 0, 1, 0)
|
||||
// DWARF section headers; their indices follow the write order.
|
||||
if dw != nil {
|
||||
|
||||
@@ -0,0 +1,128 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
package asm
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"debug/elf"
|
||||
"encoding/binary"
|
||||
"testing"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
)
|
||||
|
||||
// TestELFRISCVObjectDataRelocation checks that a symbol-valued DATA field
|
||||
// ("DATA s+0(SB)/8, $other(SB)") reaches the RISC-V ELF object as a
|
||||
// .rela.data entry: an R_RISCV_64 (R_RISCV_32 for a width-4 field) at the
|
||||
// field's offset within .data, against the named symbol, external targets
|
||||
// included.
|
||||
func TestELFRISCVObjectDataRelocation(t *testing.T) {
|
||||
f, errs := parser.Parse("t_riscv64.s", `#include "textflag.h"
|
||||
TEXT ·Keep(SB), NOSPLIT, $0-0
|
||||
RET
|
||||
GLOBL holder(SB), NOPTR, $32
|
||||
DATA holder+0(SB)/8, $·Keep+5(SB)
|
||||
DATA holder+8(SB)/8, $holder(SB)
|
||||
DATA holder+16(SB)/8, $extvar(SB)
|
||||
DATA holder+24(SB)/4, $Keep(SB)
|
||||
`)
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
img, err := AssembleFileRISCV(f)
|
||||
if err != nil {
|
||||
t.Fatalf("AssembleFileRISCV: %v", err)
|
||||
}
|
||||
obj, err := img.ELFRISCVObject()
|
||||
if err != nil {
|
||||
t.Fatalf("ELFRISCVObject: %v", err)
|
||||
}
|
||||
checkELFSectionAccounting(t, obj)
|
||||
ef, err := elf.NewFile(bytes.NewReader(obj))
|
||||
if err != nil {
|
||||
t.Fatalf("parse emitted object: %v", err)
|
||||
}
|
||||
defer ef.Close()
|
||||
relaData := ef.Section(".rela.data")
|
||||
if relaData == nil {
|
||||
t.Fatal("missing .rela.data section")
|
||||
}
|
||||
if relaData.Type != elf.SHT_RELA {
|
||||
t.Errorf(".rela.data type = %v, want SHT_RELA", relaData.Type)
|
||||
}
|
||||
if relaData.Link == 0 || ef.Sections[relaData.Link].Name != ".symtab" {
|
||||
t.Errorf(".rela.data sh_link = %d, want the .symtab index", relaData.Link)
|
||||
}
|
||||
if ef.Sections[relaData.Info].Name != ".data" {
|
||||
t.Errorf(".rela.data sh_info = %d, want the .data index", relaData.Info)
|
||||
}
|
||||
relas, err := relaData.Data()
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
var got []struct {
|
||||
off uint64
|
||||
sym uint32
|
||||
typ uint32
|
||||
addend int64
|
||||
}
|
||||
for i := 0; i+24 <= len(relas); i += 24 {
|
||||
got = append(got, struct {
|
||||
off uint64
|
||||
sym uint32
|
||||
typ uint32
|
||||
addend int64
|
||||
}{
|
||||
off: binary.LittleEndian.Uint64(relas[i:]),
|
||||
// r_info packs the type in the low dword and the symbol index
|
||||
// in the high dword.
|
||||
typ: binary.LittleEndian.Uint32(relas[i+8:]),
|
||||
sym: binary.LittleEndian.Uint32(relas[i+12:]),
|
||||
addend: int64(binary.LittleEndian.Uint64(relas[i+16:])),
|
||||
})
|
||||
}
|
||||
// debug/elf hides the table's null entry, so raw index s names syms[s-1].
|
||||
syms, err := ef.Symbols()
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
name := func(idx uint32) string {
|
||||
if idx >= 1 && int(idx) <= len(syms) {
|
||||
return syms[idx-1].Name
|
||||
}
|
||||
return ""
|
||||
}
|
||||
// The offsets are data-section-relative: the field's DATA offset plus
|
||||
// the symbol's position in .data (the layout aligns each symbol to 16).
|
||||
base := uint64(0)
|
||||
for _, d := range img.DataSyms {
|
||||
if d.Name == "holder" {
|
||||
base = uint64(d.Offset)
|
||||
}
|
||||
}
|
||||
want := []struct {
|
||||
off uint64
|
||||
typ uint32
|
||||
addend int64
|
||||
target string
|
||||
}{
|
||||
{off: base + 0, typ: uint32(elf.R_RISCV_64), addend: 5, target: "Keep"},
|
||||
{off: base + 8, typ: uint32(elf.R_RISCV_64), addend: 0, target: "holder"},
|
||||
{off: base + 16, typ: uint32(elf.R_RISCV_64), addend: 0, target: "extvar"},
|
||||
{off: base + 24, typ: uint32(elf.R_RISCV_32), addend: 0, target: "Keep"},
|
||||
}
|
||||
if len(got) != len(want) {
|
||||
t.Fatalf(".rela.data entries = %d, want %d", len(got), len(want))
|
||||
}
|
||||
for i, w := range want {
|
||||
g := got[i]
|
||||
if g.off != w.off || g.typ != w.typ || g.addend != w.addend {
|
||||
t.Errorf("entry %d = {off %d typ %d addend %d}, want {off %d typ %d addend %d}",
|
||||
i, g.off, g.typ, g.addend, w.off, w.typ, w.addend)
|
||||
}
|
||||
if n := name(g.sym); n != w.target {
|
||||
t.Errorf("entry %d names %q, want %q", i, n, w.target)
|
||||
}
|
||||
}
|
||||
}
|
||||
+4
-1
@@ -19,7 +19,10 @@ func Encodable(mnemonic string) bool {
|
||||
// Fixed-name instructions (no size suffix).
|
||||
switch upper {
|
||||
case "RET", "NOP", "CALL", "JMP",
|
||||
"POPFQ", "PUSHFQ", "INT", "LDMXCSR", "STMXCSR", "CMPSD", "SHA256RNDS2":
|
||||
"POPFQ", "PUSHFQ", "INT", "LDMXCSR", "STMXCSR", "CMPSD", "SHA256RNDS2",
|
||||
// The literal-data pseudo-ops, the accepted-and-ignored END and
|
||||
// bookkeeping statements, and the SP adjust.
|
||||
"BYTE", "WORD", "LONG", "QUAD", "END", "ADJSP", "FUNCDATA", "PCDATA":
|
||||
return true
|
||||
}
|
||||
if _, ok := noOperandTable[upper]; ok {
|
||||
|
||||
+307
-3
@@ -5,6 +5,8 @@ package asm
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"math"
|
||||
"strconv"
|
||||
"strings"
|
||||
)
|
||||
|
||||
@@ -21,6 +23,39 @@ func Encode(mnemonic string, ops ...Operand) ([]byte, error) {
|
||||
type enc struct {
|
||||
out []byte
|
||||
patches []encPatch // disp32 fields awaiting static-symbol resolution
|
||||
|
||||
// FloatPool collects the pooled constants the floating-point
|
||||
// immediates reference, in first-use order.
|
||||
floatPool []floatPoolEntry
|
||||
floatPoolSeen map[string]bool
|
||||
}
|
||||
|
||||
// floatPoolEntry is one pooled floating-point constant: the symbol name
|
||||
// the emitted RIP-relative load refers to and its IEEE-754 bytes.
|
||||
type floatPoolEntry struct {
|
||||
name string
|
||||
data []byte
|
||||
}
|
||||
|
||||
// addFloatPool records a pooled constant, deduplicated by symbol name.
|
||||
func (e *enc) addFloatPool(name string, bits uint64, width int) {
|
||||
if e.floatPoolSeen == nil {
|
||||
e.floatPoolSeen = map[string]bool{}
|
||||
}
|
||||
if e.floatPoolSeen[name] {
|
||||
return
|
||||
}
|
||||
e.floatPoolSeen[name] = true
|
||||
data := make([]byte, width)
|
||||
for i := range width {
|
||||
data[i] = byte(bits >> (8 * i))
|
||||
}
|
||||
e.floatPool = append(e.floatPool, floatPoolEntry{name: name, data: data})
|
||||
}
|
||||
|
||||
// floatPoolList returns the pooled constants in first-use order.
|
||||
func (e *enc) floatPoolList() []floatPoolEntry {
|
||||
return e.floatPool
|
||||
}
|
||||
|
||||
// encPatch marks a 4-byte displacement field in enc.out that must receive the
|
||||
@@ -29,6 +64,7 @@ type encPatch struct {
|
||||
off int
|
||||
name string
|
||||
addend int64
|
||||
tls bool // a TLS slot offset: the patch is R_TLSLE with no symbol
|
||||
}
|
||||
|
||||
func (e *enc) encode(mnem string, ops []Operand) error {
|
||||
@@ -92,6 +128,21 @@ func (e *enc) encode(mnem string, ops []Operand) error {
|
||||
// SHA256RNDS2 carries the round constant in a literal X0 first operand.
|
||||
case "SHA256RNDS2":
|
||||
return e.encodeSha256rnds2(ops)
|
||||
// BYTE, WORD, LONG and QUAD write the immediate into the text stream
|
||||
// itself: 1, 2, 4 or 8 literal bytes, little-endian. END is accepted
|
||||
// and ignored. ADJSP adjusts SP by the immediate, sign-chosen between
|
||||
// the SUBQ and ADDQ forms.
|
||||
case "BYTE", "WORD", "LONG", "QUAD":
|
||||
return e.encodeData(upper, ops)
|
||||
case "END":
|
||||
return e.encodeEnd(ops)
|
||||
case "ADJSP":
|
||||
return e.encodeAdjsp(ops)
|
||||
// The runtime's bookkeeping statements carry no text bytes: go tool asm
|
||||
// records FUNCDATA and PCDATA in the program list only, so the encoded
|
||||
// body shows nothing, on every architecture.
|
||||
case "FUNCDATA", "PCDATA":
|
||||
return e.encodeFuncdata(upper, ops)
|
||||
}
|
||||
|
||||
// VEX (AVX/AVX2) and EVEX (AVX-512) instructions: the trailing
|
||||
@@ -103,6 +154,7 @@ func (e *enc) encode(mnem string, ops []Operand) error {
|
||||
return err
|
||||
}
|
||||
if isVex(base) || isEvex(base) || isKOp(base) || isGather(base) || isScatter(base) ||
|
||||
isEvexPrefGather(base) ||
|
||||
base == "KMOVW" || base == "KMOVQ" || base == "KMOVB" || base == "KMOVD" {
|
||||
return e.encodeVec(base, ops, sfx)
|
||||
}
|
||||
@@ -130,11 +182,18 @@ func (e *enc) encode(mnem string, ops []Operand) error {
|
||||
}
|
||||
// Legacy SSE packed binaries dispatch on the full name: the packed
|
||||
// integer mnemonics carry real width suffixes (PADDB/PCMPGTW/...),
|
||||
// which the size split must not eat.
|
||||
// which the size split must not eat. A floating-point immediate
|
||||
// rewrites into a pooled-constant read on the scalar members.
|
||||
if m, ok := sseBinTable[upper]; ok {
|
||||
if f, isFloat := floatImmOperand(ops); isFloat {
|
||||
return e.encodeSSEFloatBin(upper, m, f, ops)
|
||||
}
|
||||
return e.encodeSSEBin(m, ops)
|
||||
}
|
||||
if m, ok := sseBinTable[base]; ok {
|
||||
if f, isFloat := floatImmOperand(ops); isFloat {
|
||||
return e.encodeSSEFloatBin(upper, m, f, ops)
|
||||
}
|
||||
return e.encodeSSEBin(m, ops)
|
||||
}
|
||||
// The imm8-controlled legacy instructions, the lane extracts and inserts
|
||||
@@ -173,7 +232,7 @@ func (e *enc) encode(mnem string, ops []Operand) error {
|
||||
case "INC", "DEC", "NEG", "NOT", "MUL", "DIV", "IDIV":
|
||||
return e.encodeUnary(unaryOp[base], ops, size)
|
||||
case "SHL", "SHR", "SAR", "SAL", "ROL", "ROR", "RCL", "RCR":
|
||||
return e.encodeShift(shiftOp[base], ops, size)
|
||||
return e.encodeShift(base, ops, size)
|
||||
case "BT", "BTS", "BTR", "BTC":
|
||||
return e.encodeBitTest(base, ops, size)
|
||||
case "XCHG":
|
||||
@@ -211,7 +270,12 @@ func (e *enc) encode(mnem string, ops []Operand) error {
|
||||
return e.encodeCvtInt(base, ops, size)
|
||||
case "FMOVD":
|
||||
return e.encodeFmov(ops)
|
||||
case "MOVOU", "MOVO", "MOVOA", "MOVUPS", "MOVAPS", "MOVUPD", "MOVAPD", "MOVSD", "MOVSS":
|
||||
case "MOVSD", "MOVSS":
|
||||
if f, isFloat := floatImmOperand(ops); isFloat {
|
||||
return e.encodeSSEFloatMove(upper, f, ops)
|
||||
}
|
||||
return e.encodeSSEMove(sseMoveTable[base], ops)
|
||||
case "MOVOU", "MOVO", "MOVOA", "MOVUPS", "MOVAPS", "MOVUPD", "MOVAPD":
|
||||
return e.encodeSSEMove(sseMoveTable[base], ops)
|
||||
}
|
||||
return fmt.Errorf("unsupported instruction %q", mnem)
|
||||
@@ -241,6 +305,214 @@ var prefetchVariant = map[string]int{
|
||||
"PREFETCHT2": 3,
|
||||
}
|
||||
|
||||
// dataWidth is the literal byte count of each data-emission pseudo-op.
|
||||
var dataWidth = map[string]int{
|
||||
"BYTE": 1,
|
||||
"WORD": 2,
|
||||
"LONG": 4,
|
||||
"QUAD": 8,
|
||||
}
|
||||
|
||||
// encodeData emits the literal-data pseudo-ops: BYTE, WORD, LONG and QUAD
|
||||
// write the immediate into the text stream as 1, 2, 4 or 8 bytes,
|
||||
// little-endian, with no opcode lookup. The value is truncated to the
|
||||
// width rather than range-checked, exactly as go tool asm behaves (BYTE
|
||||
// $0x1FF emits FF, WORD $0x12345 emits 45 23, both without an error), and
|
||||
// exactly one immediate is accepted: the toolchain rejects a list such as
|
||||
// BYTE $1, $2, $3.
|
||||
func (e *enc) encodeData(mnem string, ops []Operand) error {
|
||||
if len(ops) != 1 {
|
||||
return fmt.Errorf("%s expects 1 immediate operand, got %d", mnem, len(ops))
|
||||
}
|
||||
imm, ok := ops[0].(Imm)
|
||||
if !ok {
|
||||
return fmt.Errorf("%s requires an integer immediate", mnem)
|
||||
}
|
||||
width := dataWidth[mnem]
|
||||
out := make([]byte, width)
|
||||
u := uint64(imm)
|
||||
for i := range width {
|
||||
out[i] = byte(u >> (8 * i))
|
||||
}
|
||||
e.out = append(e.out, out...)
|
||||
return nil
|
||||
}
|
||||
|
||||
// encodeFuncdata accepts-and-ignores the runtime bookkeeping statements:
|
||||
// FUNCDATA $n, sym(SB) and PCDATA $n, $m. go tool asm emits no text bytes
|
||||
// for either (the entries live in the object's ancillary tables, not the
|
||||
// function body), and the operand shapes it takes are exactly these: an
|
||||
// integer count first, then a symbol reference for FUNCDATA and an integer
|
||||
// value for PCDATA. The other architectures accept-and-ignore the same
|
||||
// statements; amd64 now matches.
|
||||
func (e *enc) encodeFuncdata(upper string, ops []Operand) error {
|
||||
if len(ops) != 2 {
|
||||
return fmt.Errorf("%s expects 2 operands, got %d", upper, len(ops))
|
||||
}
|
||||
if _, ok := ops[0].(Imm); !ok {
|
||||
return fmt.Errorf("%s: first operand must be an integer immediate", upper)
|
||||
}
|
||||
switch upper {
|
||||
case "FUNCDATA":
|
||||
if _, ok := ops[1].(sbMem); !ok {
|
||||
return fmt.Errorf("FUNCDATA: second operand must be a symbol reference")
|
||||
}
|
||||
case "PCDATA":
|
||||
if _, ok := ops[1].(Imm); !ok {
|
||||
return fmt.Errorf("PCDATA: second operand must be an integer immediate")
|
||||
}
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// encodeEnd accepts-and-ignores END. go tool asm drops the statement
|
||||
// entirely: the AEND Prog is skipped when the program list is flushed, so
|
||||
// the statements after an END still belong to the same function and the
|
||||
// encoded body carries no trace of it, whatever operands follow the name
|
||||
// (the toolchain takes END $0 and END AX alike). Zero bytes, no effect.
|
||||
func (e *enc) encodeEnd(ops []Operand) error {
|
||||
return nil
|
||||
}
|
||||
|
||||
// encodeAdjsp emits ADJSP $imm: a positive value is SUBQ $imm, SP, a
|
||||
// negative one ADDQ $-imm, SP, in the imm8 or imm32 form the magnitude
|
||||
// picks (the same selection subSP and addSP make for the frame). go tool
|
||||
// asm refuses ADJSP $0 outright, so a zero value is an error here too; the
|
||||
// statement's effect on the SP balance is checked by the function-level
|
||||
// assembly (checkAdjspBalance), as the toolchain's push/pop walk does.
|
||||
func (e *enc) encodeAdjsp(ops []Operand) error {
|
||||
if len(ops) != 1 {
|
||||
return fmt.Errorf("ADJSP expects 1 immediate operand, got %d", len(ops))
|
||||
}
|
||||
imm, ok := ops[0].(Imm)
|
||||
if !ok {
|
||||
return fmt.Errorf("ADJSP requires an integer immediate")
|
||||
}
|
||||
switch v := int(imm); {
|
||||
case v > 0:
|
||||
e.out = append(e.out, subSP(v)...)
|
||||
case v < 0:
|
||||
e.out = append(e.out, addSP(-v)...)
|
||||
default:
|
||||
return fmt.Errorf("ADJSP $0 has no encoding")
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// --- floating-point immediates ----------------------------------------------
|
||||
|
||||
// sseFloatImm lists the mnemonics whose first operand may be a floating-point
|
||||
// immediate, the set go tool asm rewrites into a pooled-constant read: the
|
||||
// scalar moves, the four scalar arithmetic pairs and the scalar compares.
|
||||
// The packed members and the uniform forms (MAXSD, MINSD, SQRTSD, CMPSD)
|
||||
// reject the immediate in the toolchain and are absent here on purpose.
|
||||
var sseFloatImm = map[string]bool{
|
||||
"MOVSD": true, "MOVSS": true,
|
||||
"ADDSD": true, "ADDSS": true,
|
||||
"SUBSD": true, "SUBSS": true,
|
||||
"MULSD": true, "MULSS": true,
|
||||
"DIVSD": true, "DIVSS": true,
|
||||
"COMISD": true, "COMISS": true,
|
||||
"UCOMISD": true, "UCOMISS": true,
|
||||
}
|
||||
|
||||
// floatImmOperand reports whether the operand list opens with a
|
||||
// floating-point immediate in the two-operand spelling (imm, dst).
|
||||
func floatImmOperand(ops []Operand) (FloatImm, bool) {
|
||||
if len(ops) != 2 {
|
||||
return FloatImm{}, false
|
||||
}
|
||||
f, ok := ops[0].(FloatImm)
|
||||
return f, ok
|
||||
}
|
||||
|
||||
// floatPoolValue evaluates a floating-point immediate at the width its
|
||||
// mnemonic encodes and names the pool constant the toolchain synthesises:
|
||||
// $f64.<16 hex> for the doubles, $f32.<8 hex> for the singles (the float32
|
||||
// rounding of the parsed value). The name carries the IEEE-754 bits; the
|
||||
// section holds them little-endian.
|
||||
func floatPoolValue(mnem string, f FloatImm) (bits uint64, name string, err error) {
|
||||
v, err := strconv.ParseFloat(f.Text, 64)
|
||||
if err != nil {
|
||||
return 0, "", fmt.Errorf("invalid floating-point immediate %q", f.Text)
|
||||
}
|
||||
if f.Neg {
|
||||
v = -v
|
||||
}
|
||||
if strings.HasSuffix(mnem, "D") {
|
||||
bits = math.Float64bits(v)
|
||||
return bits, fmt.Sprintf("$f64.%016x", bits), nil
|
||||
}
|
||||
bits = uint64(math.Float32bits(float32(v)))
|
||||
return bits, fmt.Sprintf("$f32.%08x", bits), nil
|
||||
}
|
||||
|
||||
// encodeSSEFloatMove encodes MOVSD/MOVSS with a floating-point immediate
|
||||
// source. A positive zero needs no memory read: the toolchain emits
|
||||
// XORPS dst, dst. Anything else loads the pooled constant RIP-relative
|
||||
// ($f64.<hex>(SB) / $f32.<hex>(SB)), the displacement a patch site the
|
||||
// file-level layout or the linker resolves.
|
||||
func (e *enc) encodeSSEFloatMove(mnem string, f FloatImm, ops []Operand) error {
|
||||
if !sseFloatImm[mnem] {
|
||||
return fmt.Errorf("%s does not take a floating-point immediate", mnem)
|
||||
}
|
||||
dst, ok := ops[1].(Reg)
|
||||
if !ok || !dst.isVec() {
|
||||
return fmt.Errorf("%s: destination must be a vector register", mnem)
|
||||
}
|
||||
bits, name, err := floatPoolValue(mnem, f)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
e.addFloatPool(name, bits, mwidth(mnem))
|
||||
if bits == 0 {
|
||||
i := &instr{opcode: []byte{0x0F, 0x57}, modrm: -1, sib: -1} // XORPS
|
||||
if err := setRM(i, dst, dst, 8); err != nil {
|
||||
return err
|
||||
}
|
||||
return e.emit(i)
|
||||
}
|
||||
m := sseMoveTable[mnem]
|
||||
i := &instr{prefix: m.prefix, opcode: []byte{0x0F, m.load}, modrm: -1, sib: -1}
|
||||
if err := setRM(i, dst, sbMem{size: mwidth(mnem), name: name}, 8); err != nil {
|
||||
return err
|
||||
}
|
||||
return e.emit(i)
|
||||
}
|
||||
|
||||
// encodeSSEFloatBin encodes the scalar arithmetic and compare mnemonics with
|
||||
// a floating-point immediate source: the constant is read from the pool into
|
||||
// the instruction's r/m side (reg = destination), the rewrite go tool asm
|
||||
// performs at the source level.
|
||||
func (e *enc) encodeSSEFloatBin(mnem string, m sseBin, f FloatImm, ops []Operand) error {
|
||||
if !sseFloatImm[mnem] {
|
||||
return fmt.Errorf("%s does not take a floating-point immediate", mnem)
|
||||
}
|
||||
dst, ok := ops[1].(Reg)
|
||||
if !ok || !dst.isVec() {
|
||||
return fmt.Errorf("%s: destination must be a vector register", mnem)
|
||||
}
|
||||
bits, name, err := floatPoolValue(mnem, f)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
e.addFloatPool(name, bits, mwidth(mnem))
|
||||
i := &instr{prefix: m.prefix, opcode: []byte{0x0F, m.op}, modrm: -1, sib: -1}
|
||||
if err := setRM(i, dst, sbMem{size: mwidth(mnem), name: name}, 8); err != nil {
|
||||
return err
|
||||
}
|
||||
return e.emit(i)
|
||||
}
|
||||
|
||||
// mwidth returns the operand width a scalar SSE mnemonic encodes: the double
|
||||
// spellings end in D, the single spellings in S.
|
||||
func mwidth(mnem string) int {
|
||||
if strings.HasSuffix(mnem, "D") {
|
||||
return 8
|
||||
}
|
||||
return 4
|
||||
}
|
||||
|
||||
// splitSize separates a trailing B/W/L/Q size suffix from the mnemonic.
|
||||
func splitSize(upper string) (base string, size int) {
|
||||
if upper == "" {
|
||||
@@ -307,6 +579,7 @@ type instr struct {
|
||||
disp []byte
|
||||
imm []byte
|
||||
sb *sbRef // static-symbol displacement in disp, awaiting resolution
|
||||
tls bool // the displacement is a TLS slot offset, patched R_TLSLE
|
||||
}
|
||||
|
||||
// sbRef records that an instruction's displacement refers to a static symbol
|
||||
@@ -349,6 +622,9 @@ func (e *enc) emit(i *instr) error {
|
||||
if i.sb != nil {
|
||||
e.patches = append(e.patches, encPatch{off: len(e.out), name: i.sb.name, addend: i.sb.addend})
|
||||
}
|
||||
if i.tls {
|
||||
e.patches = append(e.patches, encPatch{off: len(e.out), tls: true})
|
||||
}
|
||||
e.out = append(e.out, i.disp...)
|
||||
e.out = append(e.out, i.imm...)
|
||||
return nil
|
||||
@@ -403,12 +679,30 @@ func setRMReg(i *instr, regField int, rexR, regForced bool, rm Operand, opSize i
|
||||
i.disp = le32(0)
|
||||
i.sb = &sbRef{name: r.name, addend: r.addend}
|
||||
return nil
|
||||
case TLSMem:
|
||||
// off(TLS): the segment-prefixed absolute access, mod=00 with the
|
||||
// SIB escape's disp32 absolute form. The displacement is the TLS
|
||||
// slot offset, patched by the linker's TLS relocation.
|
||||
i.prefix = r.Seg
|
||||
i.modrm = 0x04 | regField<<3
|
||||
i.sib = 0x25
|
||||
i.disp = le32(r.Disp)
|
||||
i.tls = true
|
||||
return nil
|
||||
case SegAbs:
|
||||
// 0x30(GS): the segment override with the SIB escape's disp32
|
||||
// absolute form, no relocation.
|
||||
setSegAbs(i, regField, r)
|
||||
return nil
|
||||
default:
|
||||
return fmt.Errorf("invalid r/m operand %T", rm)
|
||||
}
|
||||
}
|
||||
|
||||
func setMem(i *instr, regField int, m Mem) error {
|
||||
if m.Seg != 0 {
|
||||
i.prefix = m.Seg
|
||||
}
|
||||
modrm, sib, disp, xBit, bBit, err := memComponents(regField, m)
|
||||
if err != nil {
|
||||
return err
|
||||
@@ -421,6 +715,16 @@ func setMem(i *instr, regField int, m Mem) error {
|
||||
return nil
|
||||
}
|
||||
|
||||
// setSegAbs assembles a segment-absolute operand, 0x30(GS): the segment
|
||||
// override with the mod=00 SIB escape's disp32 absolute form and no
|
||||
// relocation.
|
||||
func setSegAbs(i *instr, regField int, m SegAbs) {
|
||||
i.prefix = m.Seg
|
||||
i.modrm = 0x04 | regField<<3
|
||||
i.sib = 0x25
|
||||
i.disp = le32(m.Disp)
|
||||
}
|
||||
|
||||
// memComponents computes the ModR/M byte (with the given reg field), the SIB
|
||||
// byte (-1 if none), the displacement bytes, and the high index/base bits, for
|
||||
// a memory operand. It is shared by the REX (scalar) and VEX (vector) paths.
|
||||
|
||||
@@ -9,6 +9,9 @@ import (
|
||||
"testing"
|
||||
|
||||
"golang.org/x/arch/x86/x86asm"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
)
|
||||
|
||||
// decode encodes an instruction and decodes it back, returning the decoded
|
||||
@@ -200,10 +203,60 @@ func TestUnary(t *testing.T) {
|
||||
func TestShift(t *testing.T) {
|
||||
checkSyntax(t, "shl rdx, 0x2", "SHLQ", Imm(2), DX)
|
||||
checkSyntax(t, "shl rdx, cl", "SHLQ", CL, DX)
|
||||
checkSyntax(t, "shl rdx, cl", "SHLQ", CX, DX)
|
||||
checkSyntax(t, "shl rdx, 0x1", "SHLQ", Imm(1), DX)
|
||||
checkSyntax(t, "sar rcx, 0x1f", "SARQ", Imm(31), CX)
|
||||
}
|
||||
|
||||
// TestDoubleShift pins the three-operand SHL/SHR form, which encodes as
|
||||
// SHLD/SHRD: go tool asm accepts it for SHL/SHR at W/L/Q widths and rejects
|
||||
// it for SAR, SAL, the rotates and the B width. The byte pins mirror the
|
||||
// oracle's objdump output (48 0f a4 fe 0d for the first case, and so on).
|
||||
func TestDoubleShift(t *testing.T) {
|
||||
cases := []struct {
|
||||
name string
|
||||
mnem string
|
||||
ops []Operand
|
||||
want string // hex encoding
|
||||
}{
|
||||
{"SHLQ imm", "SHLQ", []Operand{Imm(0x0d), DI, SI}, "480fa4fe0d"},
|
||||
{"SHLQ CX high regs", "SHLQ", []Operand{CX, Reg{idx: 8, size: 8}, Reg{idx: 9, size: 8}}, "4d0fa5c1"},
|
||||
{"SHRQ imm", "SHRQ", []Operand{Imm(1), AX, CX}, "480facc101"},
|
||||
{"SHLW imm", "SHLW", []Operand{Imm(1), AX, CX}, "660fa4c101"},
|
||||
{"SHRD CL", "SHRQ", []Operand{CL, AX, CX}, "480fadc1"},
|
||||
{"SHLD imm high regs", "SHLQ", []Operand{Imm(2), Reg{idx: 10, size: 8}, Reg{idx: 11, size: 8}}, "4d0fa4d302"},
|
||||
{"SHRD imm max", "SHRQ", []Operand{Imm(63), Reg{idx: 9, size: 8}, Reg{idx: 15, size: 8}}, "4d0faccf3f"},
|
||||
}
|
||||
for _, c := range cases {
|
||||
code, err := Encode(c.mnem, c.ops...)
|
||||
if err != nil {
|
||||
t.Errorf("%s: Encode: %v", c.name, err)
|
||||
continue
|
||||
}
|
||||
if got := hexCompact(code); got != c.want {
|
||||
t.Errorf("%s: bytes %s, want %s", c.name, got, c.want)
|
||||
}
|
||||
}
|
||||
// Rejected forms: the oracle rejects every one of these.
|
||||
rejected := []struct {
|
||||
name string
|
||||
mnem string
|
||||
ops []Operand
|
||||
}{
|
||||
{"SARQ three operands", "SARQ", []Operand{Imm(1), AX, CX}},
|
||||
{"SALQ three operands", "SALQ", []Operand{Imm(1), AX, CX}},
|
||||
{"ROLQ three operands", "ROLQ", []Operand{Imm(1), AX, CX}},
|
||||
{"SHLB three operands", "SHLB", []Operand{Imm(1), AL, CL}},
|
||||
{"SHRQ memory source", "SHRQ", []Operand{Imm(1), Ptr(AX, 0, 8), CX}},
|
||||
{"SHRQ ECX count", "SHRQ", []Operand{Reg{idx: 1, size: 4}, AX, CX}},
|
||||
}
|
||||
for _, c := range rejected {
|
||||
if _, err := Encode(c.mnem, c.ops...); err == nil {
|
||||
t.Errorf("%s: Encode succeeded, want rejection", c.name)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestImul(t *testing.T) {
|
||||
checkSyntax(t, "imul rdx, rcx", "IMULQ", CX, DX)
|
||||
checkSyntax(t, "imul edx, edx, 0x3", "IMULL", Imm(3), DX, DX)
|
||||
@@ -277,6 +330,12 @@ func TestSSEMoveGroundTruth(t *testing.T) {
|
||||
{"MOVSD (SI),X1", "MOVSD", []Operand{Ptr(SI, 0, 8), vreg(t, "X1")}, "f20f100e", "MOVSD_XMM"},
|
||||
{"MOVSD X1,X2", "MOVSD", []Operand{vreg(t, "X1"), vreg(t, "X2")}, "f20f10d1", "MOVSD_XMM"},
|
||||
{"MOVSS X3,(DI)", "MOVSS", []Operand{vreg(t, "X3"), Ptr(DI, 0, 4)}, "f30f111f", "MOVSS"},
|
||||
// Static-symbol (SB) references: the GOROOT crypto kernels load and
|
||||
// store octa constants by name (MOVOU bswapMask<>+0(SB), X0).
|
||||
{"MOVOU sym,X0", "MOVOU", []Operand{sbMem{size: 16, name: "bswapMask"}, vreg(t, "X0")}, "f30f6f0500000000", "MOVDQU"},
|
||||
{"MOVOU X0,sym+8", "MOVOU", []Operand{vreg(t, "X0"), sbMem{size: 16, name: "bswapMask", addend: 8}}, "f30f7f0500000000", "MOVDQU"},
|
||||
{"MOVO sym,X1", "MOVO", []Operand{sbMem{size: 16, name: "gcmPoly"}, vreg(t, "X1")}, "660f6f0d00000000", "MOVDQA"},
|
||||
{"MOVO X2,sym", "MOVO", []Operand{vreg(t, "X2"), sbMem{size: 16, name: "gcmPoly"}}, "660f7f1500000000", "MOVDQA"},
|
||||
}
|
||||
for _, c := range cases {
|
||||
code, err := Encode(c.mnem, c.ops...)
|
||||
@@ -869,3 +928,296 @@ func TestMOVQXMMGroundTruth(t *testing.T) {
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestPrefixStatements pins LOCK, REP and REPN. go tool asm encodes each as
|
||||
// a standalone one-byte instruction with a PC of its own (F0, F3, F2), not a
|
||||
// prefix field merged into the following instruction, and it validates
|
||||
// nothing about the pairing (LOCK before NOP assembles). The prefixed
|
||||
// atomic and string shapes are the bytes the runtime's own kernels need.
|
||||
func TestPrefixStatements(t *testing.T) {
|
||||
cases := []struct {
|
||||
name string
|
||||
mnem string
|
||||
ops []Operand
|
||||
want string
|
||||
}{
|
||||
{"LOCK", "LOCK", nil, "f0"},
|
||||
{"REP", "REP", nil, "f3"},
|
||||
{"REPN", "REPN", nil, "f2"},
|
||||
// LOCK; CMPXCHGQ AX, (BX)
|
||||
{"LOCK CMPXCHGQ", "CMPXCHGQ", []Operand{AX, Ptr(BX, 0, 8)}, "480fb103"},
|
||||
// REP; MOVSQ
|
||||
{"REP MOVSQ", "MOVSQ", nil, "48a5"},
|
||||
// REPN; MOVSB
|
||||
{"REPN MOVSB", "MOVSB", nil, "a4"},
|
||||
}
|
||||
for _, c := range cases {
|
||||
code, err := Encode(c.mnem, c.ops...)
|
||||
if err != nil {
|
||||
t.Errorf("%s: %v", c.name, err)
|
||||
continue
|
||||
}
|
||||
if got := fmt.Sprintf("%x", code); got != c.want {
|
||||
t.Errorf("%s = %s, want %s", c.name, got, c.want)
|
||||
}
|
||||
}
|
||||
// The prefix statements take no operands, as the toolchain reports for
|
||||
// LOCK AX.
|
||||
if _, err := Encode("LOCK", AX); err == nil {
|
||||
t.Error("LOCK AX assembled, want an error")
|
||||
}
|
||||
if _, err := Encode("REP", Imm(1)); err == nil {
|
||||
t.Error("REP $1 assembled, want an error")
|
||||
}
|
||||
}
|
||||
|
||||
// TestDataEmission pins BYTE, WORD, LONG and QUAD: the immediate lands in
|
||||
// the text stream as 1, 2, 4 or 8 little-endian bytes with no opcode
|
||||
// lookup, truncated to the width rather than range-checked (go tool asm
|
||||
// emits FF for BYTE $0x1FF and 45 23 for WORD $0x12345, both silently).
|
||||
func TestDataEmission(t *testing.T) {
|
||||
cases := []struct {
|
||||
name string
|
||||
mnem string
|
||||
imm Imm
|
||||
want string
|
||||
}{
|
||||
{"BYTE", "BYTE", 0x0f, "0f"},
|
||||
{"BYTE negative", "BYTE", -1, "ff"},
|
||||
{"BYTE truncated", "BYTE", 0x1ff, "ff"},
|
||||
{"WORD", "WORD", 0x1234, "3412"},
|
||||
{"WORD negative", "WORD", -1, "ffff"},
|
||||
{"WORD truncated", "WORD", 0x12345, "4523"},
|
||||
{"LONG", "LONG", 0x11223344, "44332211"},
|
||||
{"LONG negative", "LONG", -1, "ffffffff"},
|
||||
{"QUAD", "QUAD", 0x1122334455667788, "8877665544332211"},
|
||||
{"QUAD negative", "QUAD", -2, "feffffffffffffff"},
|
||||
}
|
||||
for _, c := range cases {
|
||||
code, err := Encode(c.mnem, c.imm)
|
||||
if err != nil {
|
||||
t.Errorf("%s: %v", c.name, err)
|
||||
continue
|
||||
}
|
||||
if got := fmt.Sprintf("%x", code); got != c.want {
|
||||
t.Errorf("%s = %s, want %s", c.name, got, c.want)
|
||||
}
|
||||
}
|
||||
// Exactly one immediate: the toolchain rejects BYTE $1, $2, $3, and a
|
||||
// register or a missing operand is no immediate at all.
|
||||
if _, err := Encode("BYTE"); err == nil {
|
||||
t.Error("BYTE with no operand assembled, want an error")
|
||||
}
|
||||
if _, err := Encode("BYTE", Imm(1), Imm(2)); err == nil {
|
||||
t.Error("BYTE $1, $2 assembled, want an error")
|
||||
}
|
||||
if _, err := Encode("WORD", AX); err == nil {
|
||||
t.Error("WORD AX assembled, want an error")
|
||||
}
|
||||
}
|
||||
|
||||
// TestEndIgnored pins END: go tool asm drops the statement entirely, so it
|
||||
// encodes to zero bytes and takes any operands without complaint (the
|
||||
// toolchain accepts END $0 and END AX alike).
|
||||
func TestEndIgnored(t *testing.T) {
|
||||
for _, ops := range [][]Operand{nil, {Imm(0)}, {AX}} {
|
||||
code, err := Encode("END", ops...)
|
||||
if err != nil {
|
||||
t.Errorf("END: %v", err)
|
||||
continue
|
||||
}
|
||||
if len(code) != 0 {
|
||||
t.Errorf("END = %x, want no bytes", code)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestAdjsp pins ADJSP: a positive immediate is SUBQ $imm, SP, a negative
|
||||
// one ADDQ $-imm, SP, in the imm8 or imm32 form the magnitude picks; $0
|
||||
// has no encoding (go tool asm refuses ADJSP $0 outright).
|
||||
func TestAdjsp(t *testing.T) {
|
||||
cases := []struct {
|
||||
name string
|
||||
imm Imm
|
||||
want string
|
||||
}{
|
||||
{"imm8", 112, "4883ec70"},
|
||||
{"imm8 negative", -112, "4883c470"},
|
||||
{"imm32", 200, "4881ecc8000000"},
|
||||
{"imm32 negative", -200, "4881c4c8000000"},
|
||||
{"small", 8, "4883ec08"},
|
||||
}
|
||||
for _, c := range cases {
|
||||
code, err := Encode("ADJSP", c.imm)
|
||||
if err != nil {
|
||||
t.Errorf("%s: %v", c.name, err)
|
||||
continue
|
||||
}
|
||||
if got := fmt.Sprintf("%x", code); got != c.want {
|
||||
t.Errorf("ADJSP %d = %s, want %s", int64(c.imm), got, c.want)
|
||||
}
|
||||
}
|
||||
if _, err := Encode("ADJSP", Imm(0)); err == nil {
|
||||
t.Error("ADJSP $0 assembled, want an error")
|
||||
}
|
||||
if _, err := Encode("ADJSP"); err == nil {
|
||||
t.Error("ADJSP with no operand assembled, want an error")
|
||||
}
|
||||
if _, err := Encode("ADJSP", AX); err == nil {
|
||||
t.Error("ADJSP AX assembled, want an error")
|
||||
}
|
||||
}
|
||||
|
||||
// TestFloatImmediateGroundTruth pins the floating-point immediate rewrite
|
||||
// byte for byte against go tool asm: the scalar moves and the scalar
|
||||
// arithmetic read the constant from a synthesised read-only pool symbol
|
||||
// ($f64.<hex>, $f32.<hex>) RIP-relative with the displacement left to the
|
||||
// relocation, and a positive zero on the moves collapses to XORPS dst, dst.
|
||||
func TestFloatImmediateGroundTruth(t *testing.T) {
|
||||
cases := []struct {
|
||||
name string
|
||||
mnem string
|
||||
ops []Operand
|
||||
want string
|
||||
}{
|
||||
{"MOVSD -1.0", "MOVSD", []Operand{FloatImm{Text: "1.0", Neg: true}, vreg(t, "X2")}, "f20f101500000000"},
|
||||
{"MOVSD 1.5", "MOVSD", []Operand{FloatImm{Text: "1.5"}, vreg(t, "X3")}, "f20f101d00000000"},
|
||||
{"MOVSS 2.5", "MOVSS", []Operand{FloatImm{Text: "2.5"}, vreg(t, "X4")}, "f30f102500000000"},
|
||||
{"MOVSS -0.5", "MOVSS", []Operand{FloatImm{Text: "0.5", Neg: true}, vreg(t, "X5")}, "f30f102d00000000"},
|
||||
{"MOVSS +0.0 is XORPS", "MOVSS", []Operand{FloatImm{Text: "0.0"}, vreg(t, "X10")}, "450f57d2"},
|
||||
{"MOVSD +0.0 is XORPS", "MOVSD", []Operand{FloatImm{Text: "0.0"}, vreg(t, "X6")}, "0f57f6"},
|
||||
{"ADDSD 1.0", "ADDSD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}, "f20f580500000000"},
|
||||
{"ADDSS 0.5", "ADDSS", []Operand{FloatImm{Text: "0.5"}, vreg(t, "X1")}, "f30f580d00000000"},
|
||||
{"SUBSD 2.0", "SUBSD", []Operand{FloatImm{Text: "2.0"}, vreg(t, "X3")}, "f20f5c1d00000000"},
|
||||
{"MULSD -2.5", "MULSD", []Operand{FloatImm{Text: "2.5", Neg: true}, vreg(t, "X3")}, "f20f591d00000000"},
|
||||
{"DIVSD 1.0", "DIVSD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}, "f20f5e0500000000"},
|
||||
{"COMISD 1.0", "COMISD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}, "660f2f0500000000"},
|
||||
}
|
||||
for _, c := range cases {
|
||||
code, err := Encode(c.mnem, c.ops...)
|
||||
if err != nil {
|
||||
t.Errorf("%s: Encode: %v", c.name, err)
|
||||
continue
|
||||
}
|
||||
if got := hexCompact(code); got != c.want {
|
||||
t.Errorf("%s: got %s, want %s", c.name, got, c.want)
|
||||
}
|
||||
}
|
||||
|
||||
// The pool names carry the IEEE-754 bits, the float32 narrowing for the
|
||||
// single spellings; negative zero keeps its sign bit and never takes the
|
||||
// XORPS shortcut.
|
||||
for _, c := range []struct {
|
||||
mnem string
|
||||
imm FloatImm
|
||||
want string
|
||||
}{
|
||||
{"MOVSD", FloatImm{Text: "1.0", Neg: true}, "$f64.bff0000000000000"},
|
||||
{"MOVSD", FloatImm{Text: "0.5"}, "$f64.3fe0000000000000"},
|
||||
{"MOVSS", FloatImm{Text: "2.5"}, "$f32.40200000"},
|
||||
{"MOVSS", FloatImm{Text: "0.5", Neg: true}, "$f32.bf000000"},
|
||||
{"MOVSD", FloatImm{Text: "0.0", Neg: true}, "$f64.8000000000000000"},
|
||||
} {
|
||||
_, name, err := floatPoolValue(c.mnem, c.imm)
|
||||
if err != nil {
|
||||
t.Errorf("%s %s: %v", c.mnem, c.imm.Text, err)
|
||||
continue
|
||||
}
|
||||
if name != c.want {
|
||||
t.Errorf("%s $%s: pool name %s, want %s", c.mnem, c.imm.Text, name, c.want)
|
||||
}
|
||||
}
|
||||
|
||||
// The shapes the toolchain's parser rejects: the packed and uniform
|
||||
// forms, a non-vector destination, and the integer spellings.
|
||||
for _, c := range []struct {
|
||||
name string
|
||||
mnem string
|
||||
ops []Operand
|
||||
}{
|
||||
{"MAXSD rejects the immediate", "MAXSD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}},
|
||||
{"MINSD rejects the immediate", "MINSD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}},
|
||||
{"SQRTSD rejects the immediate", "SQRTSD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}},
|
||||
{"integer destination", "MOVSD", []Operand{FloatImm{Text: "1.0"}, AX}},
|
||||
} {
|
||||
if _, err := Encode(c.mnem, c.ops...); err == nil {
|
||||
t.Errorf("%s: expected an error, got none", c.name)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestBookkeepingGroundTruth pins FUNCDATA and PCDATA as accept-and-ignore:
|
||||
// go tool asm emits no text bytes for either, on every architecture.
|
||||
func TestBookkeepingGroundTruth(t *testing.T) {
|
||||
for _, c := range []struct {
|
||||
name string
|
||||
mnem string
|
||||
ops []Operand
|
||||
}{
|
||||
{"FUNCDATA", "FUNCDATA", []Operand{Imm(3), sbMem{name: "\u00b7f.arginfo0"}}},
|
||||
{"PCDATA", "PCDATA", []Operand{Imm(1), Imm(-1)}},
|
||||
} {
|
||||
code, err := Encode(c.mnem, c.ops...)
|
||||
if err != nil {
|
||||
t.Errorf("%s: Encode: %v", c.name, err)
|
||||
continue
|
||||
}
|
||||
if len(code) != 0 {
|
||||
t.Errorf("%s: emitted %x, want no bytes", c.name, code)
|
||||
}
|
||||
}
|
||||
for _, c := range []struct {
|
||||
name string
|
||||
mnem string
|
||||
ops []Operand
|
||||
}{
|
||||
{"FUNCDATA arity", "FUNCDATA", []Operand{Imm(3)}},
|
||||
{"FUNCDATA missing the count", "FUNCDATA", []Operand{sbMem{name: "x"}}},
|
||||
{"FUNCDATA integer value", "FUNCDATA", []Operand{Imm(3), Imm(4)}},
|
||||
{"PCDATA arity", "PCDATA", []Operand{Imm(1)}},
|
||||
{"PCDATA register value", "PCDATA", []Operand{Imm(1), AX}},
|
||||
} {
|
||||
if _, err := Encode(c.mnem, c.ops...); err == nil {
|
||||
t.Errorf("%s: expected an error, got none", c.name)
|
||||
}
|
||||
}
|
||||
|
||||
// At the statement level the bookkeeping lines sit between real
|
||||
// instructions and contribute nothing to the body, symbol reference
|
||||
// included: the FUNCDATA operand never needs file-level resolution.
|
||||
f, errs := parser.Parse("t_amd64.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tNOP\n\tFUNCDATA $3, \u00b7f.arginfo0(SB)\n\tPCDATA $1, $-1\n\tFUNCDATA $0, x<>(SB)\n\tRET\n")
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
img, err := AssembleFile(f)
|
||||
if err != nil {
|
||||
t.Fatalf("assemble: %v", err)
|
||||
}
|
||||
want := "90c3"
|
||||
if got := hexCompact(img.Code); got != want {
|
||||
t.Errorf("body %s, want %s (the bookkeeping lines contribute nothing)", got, want)
|
||||
}
|
||||
if _, err := AssembleFile(mustParse(t, "TEXT \u00b7f(SB), NOSPLIT, $0\n\tFUNCDATA $1, X0\n\tRET\n")); err == nil {
|
||||
t.Error("FUNCDATA $1, X0 assembled, want an error")
|
||||
}
|
||||
if _, err := AssembleFile(mustParse(t, "TEXT \u00b7f(SB), NOSPLIT, $0\n\tPCDATA $1, X0\n\tRET\n")); err == nil {
|
||||
t.Error("PCDATA $1, X0 assembled, want an error")
|
||||
}
|
||||
|
||||
// Encodable mirrors Encode for the names this work touched.
|
||||
for _, mnem := range []string{"FUNCDATA", "PCDATA", "V4FMADDPS", "V4FMADDSS", "V4FNMADDPS", "V4FNMADDSS", "VP4DPWSSD", "VP4DPWSSDS"} {
|
||||
if !Encodable(mnem) {
|
||||
t.Errorf("Encodable(%s) = false, want true", mnem)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// mustParse parses src or fails the test.
|
||||
func mustParse(t *testing.T, src string) *ast.File {
|
||||
t.Helper()
|
||||
f, errs := parser.Parse("t_amd64.s", src)
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
return f
|
||||
}
|
||||
|
||||
+593
-24
@@ -5,6 +5,7 @@ package asm
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"slices"
|
||||
"strings"
|
||||
)
|
||||
|
||||
@@ -91,7 +92,7 @@ var evexTable = map[string]evexSpec{
|
||||
"VPSRAD": {1, 0x72, 0, 1, 4, vexShiftImm, [3]int{16, 32, 64}},
|
||||
// EVEX.128/256/512.66.0F.W1, variable shift with an XMM count (VPSRAQ;
|
||||
// the W bit distinguishes it from VPSRAD's E2 form).
|
||||
"VPSRAQ": {1, 0xE2, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPSRAQ": {1, 0x72, 1, 1, 4, vexShiftImm, [3]int{16, 32, 64}},
|
||||
|
||||
// EVEX.128/256/512.F3.0F.W1, signed qword to packed double (reg=dst,
|
||||
// rm=src, no vvvv).
|
||||
@@ -138,7 +139,7 @@ var evexTable = map[string]evexSpec{
|
||||
|
||||
// EVEX.66.0F, the EVEX forms of the VEX two-source shuffle.
|
||||
"VSHUFPD": {1, 0xC6, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
|
||||
"VSHUFPS": {1, 0xC6, 0, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
|
||||
"VSHUFPS": {1, 0xC6, 0, 0, -1, vexNDS3Imm, [3]int{16, 32, 64}},
|
||||
|
||||
// EVEX.66.0F3A, lane insert ($imm, xsrc, zsrc1, zdst).
|
||||
"VINSERTF32X4": {3, 0x18, 0, 1, -1, vexNDS3Imm, [3]int{0, 16, 32}},
|
||||
@@ -202,7 +203,7 @@ var evexTable = map[string]evexSpec{
|
||||
"VPMULHUW": {1, 0xE4, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPMADDUBSW": {2, 0x04, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPSLLVW": {2, 0x12, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPSRLVW": {2, 0x11, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPSRLVW": {2, 0x10, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPACKSSWB": {1, 0x63, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPACKUSWB": {1, 0x67, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPACKSSDW": {1, 0x6B, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
@@ -324,7 +325,7 @@ var evexTable = map[string]evexSpec{
|
||||
"VCVTPD2UQQ": {1, 0x79, 1, 1, -1, vexRM, [3]int{16, 32, 64}},
|
||||
"VCVTPS2QQ": {1, 0x7B, 0, 1, -1, vexRM, [3]int{8, 16, 32}},
|
||||
"VCVTUDQ2PD": {1, 0x7A, 0, 2, -1, vexRM, [3]int{8, 16, 32}},
|
||||
"VCVTUDQ2PS": {1, 0x7A, 0, 0, -1, vexRM, [3]int{8, 16, 32}},
|
||||
"VCVTUDQ2PS": {1, 0x7A, 0, 3, -1, vexRM, [3]int{8, 16, 32}},
|
||||
// EVEX.66.0F38, half-precision convert (half-width source).
|
||||
"VCVTPH2PS": {2, 0x13, 0, 1, -1, vexRM, [3]int{8, 16, 32}},
|
||||
// EVEX.66.0F3A, half-precision convert back ($imm, src, dst: reg=src,
|
||||
@@ -494,6 +495,331 @@ var evexTable = map[string]evexSpec{
|
||||
// destination (VPMOVDW dword→word, VPMOVQD qword→dword).
|
||||
"VPMOVDW": {2, 0x33, 0, 2, -1, vexRMRev, [3]int{8, 16, 32}},
|
||||
"VPMOVQD": {2, 0x35, 0, 2, -1, vexRMRev, [3]int{8, 16, 32}},
|
||||
|
||||
// --- the AVX-512 families the avx512enc corpus exercises, read off
|
||||
// the toolchain opcodetables ---
|
||||
"VAESDEC": {2, 0xDE, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VAESDECLAST": {2, 0xDF, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VAESENC": {2, 0xDC, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VAESENCLAST": {2, 0xDD, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VALIGNQ": {3, 0x03, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
|
||||
"VANDNPD": {1, 0x55, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VANDPD": {1, 0x54, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VBLENDMPD": {2, 0x65, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VBLENDMPS": {2, 0x65, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VBROADCASTF32X2": {2, 0x19, 0, 1, -1, vexRM, [3]int{0, 8, 8}},
|
||||
"VBROADCASTF32X4": {2, 0x1A, 0, 1, -1, vexRM, [3]int{0, 16, 16}},
|
||||
"VBROADCASTF32X8": {2, 0x1B, 0, 1, -1, vexRM, [3]int{0, 0, 32}},
|
||||
"VBROADCASTF64X2": {2, 0x1A, 1, 1, -1, vexRM, [3]int{0, 16, 16}},
|
||||
"VBROADCASTF64X4": {2, 0x1B, 1, 1, -1, vexRM, [3]int{0, 0, 32}},
|
||||
"VBROADCASTI32X2": {2, 0x59, 0, 1, -1, vexRM, [3]int{8, 8, 8}},
|
||||
"VBROADCASTI32X4": {2, 0x5A, 0, 1, -1, vexRM, [3]int{0, 16, 16}},
|
||||
"VBROADCASTI32X8": {2, 0x5B, 0, 1, -1, vexRM, [3]int{0, 0, 32}},
|
||||
"VBROADCASTI64X2": {2, 0x5A, 1, 1, -1, vexRM, [3]int{0, 16, 16}},
|
||||
"VBROADCASTI64X4": {2, 0x5B, 1, 1, -1, vexRM, [3]int{0, 0, 32}},
|
||||
"VCOMISD": {1, 0x2F, 1, 1, -1, vexRM, [3]int{8, 0, 0}},
|
||||
"VCVTSD2SS": {1, 0x5A, 1, 3, -1, vexNDS3, [3]int{8, 0, 0}},
|
||||
"VCVTSS2SD": {1, 0x5A, 0, 2, -1, vexNDS3, [3]int{4, 0, 0}},
|
||||
"VDBPSADBW": {3, 0x42, 0, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
|
||||
"VEXP2PD": {2, 0xC8, 1, 1, -1, vexRM, [3]int{0, 0, 64}},
|
||||
"VEXP2PS": {2, 0xC8, 0, 1, -1, vexRM, [3]int{0, 0, 64}},
|
||||
"VFMADD132PD": {2, 0x98, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFMADD132PS": {2, 0x98, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFMADD132SD": {2, 0x99, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
|
||||
"VFMADD132SS": {2, 0x99, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
|
||||
"VFMADD213PD": {2, 0xA8, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFMADD213PS": {2, 0xA8, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFMADD213SD": {2, 0xA9, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
|
||||
"VFMADD213SS": {2, 0xA9, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
|
||||
"VFMADD231PS": {2, 0xB8, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFMADD231SD": {2, 0xB9, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
|
||||
"VFMADD231SS": {2, 0xB9, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
|
||||
"VFMADDSUB132PD": {2, 0x96, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFMADDSUB132PS": {2, 0x96, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFMADDSUB213PD": {2, 0xA6, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFMADDSUB213PS": {2, 0xA6, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFMADDSUB231PD": {2, 0xB6, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFMADDSUB231PS": {2, 0xB6, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFMSUB132PD": {2, 0x9A, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFMSUB132PS": {2, 0x9A, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFMSUB132SD": {2, 0x9B, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
|
||||
"VFMSUB132SS": {2, 0x9B, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
|
||||
"VFMSUB213PD": {2, 0xAA, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFMSUB213PS": {2, 0xAA, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFMSUB213SD": {2, 0xAB, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
|
||||
"VFMSUB213SS": {2, 0xAB, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
|
||||
"VFMSUB231PD": {2, 0xBA, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFMSUB231PS": {2, 0xBA, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFMSUB231SD": {2, 0xBB, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
|
||||
"VFMSUB231SS": {2, 0xBB, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
|
||||
"VFMSUBADD132PD": {2, 0x97, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFMSUBADD132PS": {2, 0x97, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFMSUBADD213PD": {2, 0xA7, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFMSUBADD213PS": {2, 0xA7, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFMSUBADD231PD": {2, 0xB7, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFMSUBADD231PS": {2, 0xB7, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFNMADD132PD": {2, 0x9C, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFNMADD132PS": {2, 0x9C, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFNMADD132SD": {2, 0x9D, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
|
||||
"VFNMADD132SS": {2, 0x9D, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
|
||||
"VFNMADD213PD": {2, 0xAC, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFNMADD213PS": {2, 0xAC, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFNMADD213SD": {2, 0xAD, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
|
||||
"VFNMADD213SS": {2, 0xAD, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
|
||||
"VFNMADD231PD": {2, 0xBC, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFNMADD231PS": {2, 0xBC, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFNMADD231SD": {2, 0xBD, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
|
||||
"VFNMADD231SS": {2, 0xBD, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
|
||||
"VFNMSUB132PD": {2, 0x9E, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFNMSUB132PS": {2, 0x9E, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFNMSUB132SD": {2, 0x9F, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
|
||||
"VFNMSUB132SS": {2, 0x9F, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
|
||||
"VFNMSUB213PD": {2, 0xAE, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFNMSUB213PS": {2, 0xAE, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFNMSUB213SD": {2, 0xAF, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
|
||||
"VFNMSUB213SS": {2, 0xAF, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
|
||||
"VFNMSUB231PD": {2, 0xBE, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFNMSUB231PS": {2, 0xBE, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VFNMSUB231SD": {2, 0xBF, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
|
||||
"VFNMSUB231SS": {2, 0xBF, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
|
||||
"VGF2P8AFFINEINVQB": {3, 0xCF, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
|
||||
"VGF2P8AFFINEQB": {3, 0xCE, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
|
||||
"VGF2P8MULB": {2, 0xCF, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VMOVNTDQ": {1, 0xE7, 0, 1, -1, vexRMRev, [3]int{16, 32, 64}},
|
||||
"VMOVNTDQA": {2, 0x2A, 0, 1, -1, vexRM, [3]int{16, 32, 64}},
|
||||
"VMOVNTPD": {1, 0x2B, 1, 1, -1, vexRMRev, [3]int{16, 32, 64}},
|
||||
"VORPD": {1, 0x56, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPADDSB": {1, 0xEC, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPADDSW": {1, 0xED, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPADDUSB": {1, 0xDC, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPADDUSW": {1, 0xDD, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPBLENDMB": {2, 0x66, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPBLENDMD": {2, 0x64, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPBLENDMQ": {2, 0x64, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPBLENDMW": {2, 0x66, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPBROADCASTMB2Q": {2, 0x2A, 1, 2, -1, vexRM, [3]int{0, 0, 0}},
|
||||
"VPBROADCASTMW2D": {2, 0x3A, 0, 2, -1, vexRM, [3]int{0, 0, 0}},
|
||||
"VPCLMULQDQ": {3, 0x44, 0, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
|
||||
"VPCMPEQB": {1, 0x74, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPCMPEQQ": {2, 0x29, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPCMPEQW": {1, 0x75, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPCMPGTB": {1, 0x64, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPCMPGTD": {1, 0x66, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPCMPGTQ": {2, 0x37, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPCMPGTW": {1, 0x65, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPCOMPRESSB": {2, 0x63, 0, 1, -1, vexRMRev, [3]int{1, 1, 1}},
|
||||
"VPCOMPRESSW": {2, 0x63, 1, 1, -1, vexRMRev, [3]int{2, 2, 2}},
|
||||
"VPCONFLICTD": {2, 0xC4, 0, 1, -1, vexRM, [3]int{16, 32, 64}},
|
||||
"VPCONFLICTQ": {2, 0xC4, 1, 1, -1, vexRM, [3]int{16, 32, 64}},
|
||||
"VPDPBUSD": {2, 0x50, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPDPBUSDS": {2, 0x51, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPDPWSSD": {2, 0x52, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPDPWSSDS": {2, 0x53, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPERMI2PD": {2, 0x77, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPERMI2PS": {2, 0x77, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPERMI2W": {2, 0x75, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPERMPS": {2, 0x16, 0, 1, -1, vexNDS3, [3]int{0, 32, 64}},
|
||||
"VPERMT2B": {2, 0x7D, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPERMT2PS": {2, 0x7F, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPERMT2W": {2, 0x7D, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPEXPANDB": {2, 0x62, 0, 1, -1, vexRM, [3]int{1, 1, 1}},
|
||||
"VPEXPANDW": {2, 0x62, 1, 1, -1, vexRM, [3]int{2, 2, 2}},
|
||||
"VPINSRD": {3, 0x22, 0, 1, -1, vexNDS3Imm, [3]int{4, 0, 0}},
|
||||
"VPINSRQ": {3, 0x22, 1, 1, -1, vexNDS3Imm, [3]int{8, 0, 0}},
|
||||
"VPLZCNTD": {2, 0x44, 0, 1, -1, vexRM, [3]int{16, 32, 64}},
|
||||
"VPLZCNTQ": {2, 0x44, 1, 1, -1, vexRM, [3]int{16, 32, 64}},
|
||||
"VPMADD52HUQ": {2, 0xB5, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPMADD52LUQ": {2, 0xB4, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPMULDQ": {2, 0x28, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPMULHRSW": {2, 0x0B, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPMULHW": {1, 0xE5, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPMULTISHIFTQB": {2, 0x83, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPMULUDQ": {1, 0xF4, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPOPCNTW": {2, 0x54, 1, 1, -1, vexRM, [3]int{16, 32, 64}},
|
||||
"VPORD": {1, 0xEB, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPROLVD": {2, 0x15, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPROLVQ": {2, 0x15, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPRORVD": {2, 0x14, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPRORVQ": {2, 0x14, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPSADBW": {1, 0xF6, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPSHLDD": {3, 0x71, 0, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
|
||||
"VPSHLDQ": {3, 0x71, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
|
||||
"VPSHLDVD": {2, 0x71, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPSHLDVQ": {2, 0x71, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPSHLDVW": {2, 0x70, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPSHLDW": {3, 0x70, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
|
||||
"VPSHRDD": {3, 0x73, 0, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
|
||||
"VPSHRDQ": {3, 0x73, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
|
||||
"VPSHRDVD": {2, 0x73, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPSHRDVQ": {2, 0x73, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPSHRDVW": {2, 0x72, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPSHRDW": {3, 0x72, 1, 1, -1, vexNDS3Imm, [3]int{16, 32, 64}},
|
||||
"VPSHUFBITQMB": {2, 0x8F, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPSRAVW": {2, 0x11, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPSRLD": {1, 0x72, 0, 1, 2, vexShiftImm, [3]int{16, 32, 64}},
|
||||
"VPSRLDQ": {1, 0x73, 0, 1, 3, vexShiftImm, [3]int{16, 32, 64}},
|
||||
// EVEX.66.0F73 /7, the byte-quad shift left (the count is always an
|
||||
// immediate; there is no register-count twin).
|
||||
"VPSLLDQ": {1, 0x73, 0, 1, 7, vexShiftImm, [3]int{16, 32, 64}},
|
||||
|
||||
// EVEX.128/256/512.0F.W0, the plain-prefix (no 66) packed spellings
|
||||
// whose EVEX form drops the legacy prefix entirely.
|
||||
"VANDNPS": {1, 0x55, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VANDPS": {1, 0x54, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VORPS": {1, 0x56, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VXORPS": {1, 0x57, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VUNPCKLPS": {1, 0x14, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VUNPCKHPS": {1, 0x15, 0, 0, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VSQRTPS": {1, 0x51, 0, 0, -1, vexRM, [3]int{16, 32, 64}},
|
||||
"VCOMISS": {1, 0x2F, 0, 0, -1, vexRM, [3]int{4, 0, 0}},
|
||||
"VUCOMISS": {1, 0x2E, 0, 0, -1, vexRM, [3]int{4, 0, 0}},
|
||||
"VMOVNTPS": {1, 0x2B, 0, 0, -1, vexRMRev, [3]int{16, 32, 64}},
|
||||
"VPSUBSB": {1, 0xE8, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPSUBSW": {1, 0xE9, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPSUBUSB": {1, 0xD8, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPSUBUSW": {1, 0xD9, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPTESTMB": {2, 0x26, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPTESTMD": {2, 0x27, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPTESTMQ": {2, 0x27, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPTESTMW": {2, 0x26, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPTESTNMB": {2, 0x26, 0, 2, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPTESTNMD": {2, 0x27, 0, 2, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPTESTNMQ": {2, 0x27, 1, 2, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPTESTNMW": {2, 0x26, 1, 2, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPUNPCKHBW": {1, 0x68, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPUNPCKHQDQ": {1, 0x6D, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPUNPCKHWD": {1, 0x69, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPUNPCKLBW": {1, 0x60, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPUNPCKLQDQ": {1, 0x6C, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPUNPCKLWD": {1, 0x61, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VRCP28PD": {2, 0xCA, 1, 1, -1, vexRM, [3]int{0, 0, 64}},
|
||||
"VRCP28PS": {2, 0xCA, 0, 1, -1, vexRM, [3]int{0, 0, 64}},
|
||||
"VRCP28SD": {2, 0xCB, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
|
||||
"VRCP28SS": {2, 0xCB, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
|
||||
"VRSQRT28PD": {2, 0xCC, 1, 1, -1, vexRM, [3]int{0, 0, 64}},
|
||||
"VRSQRT28PS": {2, 0xCC, 0, 1, -1, vexRM, [3]int{0, 0, 64}},
|
||||
"VRSQRT28SD": {2, 0xCD, 1, 1, -1, vexNDS3, [3]int{8, 0, 0}},
|
||||
"VRSQRT28SS": {2, 0xCD, 0, 1, -1, vexNDS3, [3]int{4, 0, 0}},
|
||||
"VSQRTPD": {1, 0x51, 1, 1, -1, vexRM, [3]int{16, 32, 64}},
|
||||
"VSQRTSD": {1, 0x51, 1, 3, -1, vexNDS3, [3]int{8, 0, 0}},
|
||||
"VSQRTSS": {1, 0x51, 0, 2, -1, vexNDS3, [3]int{4, 0, 0}},
|
||||
"VUCOMISD": {1, 0x2E, 1, 1, -1, vexRM, [3]int{8, 0, 0}},
|
||||
"VXORPD": {1, 0x57, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
|
||||
// EVEX.128/256/512.0F.F3/F2.W0, word shuffles with an immediate
|
||||
// ($imm, src, dst: reg = dst, rm = src, imm8). The F3/F2 prefixes
|
||||
// split the high/low lane spellings.
|
||||
"VPSHUFHW": {1, 0x70, 0, 2, -1, vexImmRM, [3]int{16, 32, 64}},
|
||||
"VPSHUFLW": {1, 0x70, 0, 3, -1, vexImmRM, [3]int{16, 32, 64}},
|
||||
|
||||
// EVEX.128.66.0F3A, lane extract to a general-purpose register or
|
||||
// memory ($imm, xsrc, GPR/mem dst: reg = source, rm = destination).
|
||||
"VPEXTRB": {3, 0x14, 0, 1, -1, vexExtractGPR, [3]int{1, 1, 1}},
|
||||
"VPEXTRW": {3, 0x15, 0, 1, -1, vexExtractGPR, [3]int{2, 2, 2}},
|
||||
"VPEXTRD": {3, 0x16, 0, 1, -1, vexExtractGPR, [3]int{4, 4, 4}},
|
||||
"VPEXTRQ": {3, 0x16, 1, 1, -1, vexExtractGPR, [3]int{8, 8, 8}},
|
||||
|
||||
// EVEX.66.0F3A.W1, the qword permutes with an immediate control
|
||||
// ($imm, src, dst: reg = dst, rm = src, imm8); the register-count
|
||||
// forms live in evexRegFormTable.
|
||||
"VPERMQ": {3, 0x00, 1, 1, -1, vexImmRM, [3]int{16, 32, 64}},
|
||||
"VPERMPD": {3, 0x01, 1, 1, -1, vexImmRM, [3]int{16, 32, 64}},
|
||||
// EVEX.66.0F3A, the packed permute shuffles with an immediate control.
|
||||
"VPERMILPS": {3, 0x04, 0, 1, -1, vexImmRM, [3]int{16, 32, 64}},
|
||||
"VPERMILPD": {3, 0x05, 1, 1, -1, vexImmRM, [3]int{16, 32, 64}},
|
||||
|
||||
// EVEX.128.0F.W0, high/low half moves. VMOVHPS carries the
|
||||
// three-operand insert form (rm = m64 source, vvvv = preserved,
|
||||
// reg = dst) and the two-operand store (reg = source, rm = m64);
|
||||
// the encoder splits on the operand count. VMOVLHPS is the
|
||||
// three-operand form alone.
|
||||
"VMOVHPS": {1, 0x16, 0, 0, -1, vexNDS3, [3]int{8, 0, 0}},
|
||||
"VMOVLHPS": {1, 0x16, 0, 0, -1, vexNDS3, [3]int{8, 0, 0}},
|
||||
}
|
||||
|
||||
// evexQuad describes one quad-register instruction: the opcode under
|
||||
// EVEX.0F38.W0 with the F2 mandatory prefix, and the width of the vector
|
||||
// registers the bracketed list and the destination take (512-bit ZMM for
|
||||
// the packed forms, 128-bit XMM for the scalar ones).
|
||||
type evexQuad struct {
|
||||
opcode byte
|
||||
width int // register width in bytes: 64 (ZMM) or 16 (XMM)
|
||||
}
|
||||
|
||||
// evexQuadTable maps the quad-register instructions (the 4FMAPS and 4VNNIW
|
||||
// families) to their encoding. The operand shape is fixed: a single memory
|
||||
// source in r/m, the bracketed register list whose LOW register travels the
|
||||
// inverted 5-bit V'VVVV field, an optional opmask in aaa and the vector
|
||||
// destination in reg. The vector length follows the destination (512-bit
|
||||
// for the ZMM list forms, 128-bit for the scalar ones) while the disp8×N
|
||||
// multiplier stays 16 for every member, the toolchain's own tuple choice.
|
||||
var evexQuadTable = map[string]evexQuad{
|
||||
"V4FMADDPS": {0x9A, 64},
|
||||
"V4FMADDSS": {0x9B, 16},
|
||||
"V4FNMADDPS": {0xAA, 64},
|
||||
"V4FNMADDSS": {0xAB, 16},
|
||||
"VP4DPWSSD": {0x52, 64},
|
||||
"VP4DPWSSDS": {0x53, 64},
|
||||
}
|
||||
|
||||
// isEvexQuad reports whether the mnemonic is a quad-register instruction.
|
||||
func isEvexQuad(upper string) bool {
|
||||
_, ok := evexQuadTable[upper]
|
||||
return ok
|
||||
}
|
||||
|
||||
// encodeEvexQuad encodes the quad-register form: OP mem, [Zn-Zn+3], (K), dst.
|
||||
// The register list is the VVVV-side source: its low register fills the
|
||||
// inverted V'VVVV bits, which is why an indexed memory source above Z15 (no
|
||||
// spare EVEX.X bit once V' is taken) is refused. Masking rides the standard
|
||||
// aaa field, zeroing keeps the usual requires-a-mask rule, and no other
|
||||
// suffix applies.
|
||||
func (e *enc) encodeEvexQuad(mnem string, q evexQuad, ops []Operand, sfx evexSuffix) error {
|
||||
if len(ops) != 3 && len(ops) != 4 {
|
||||
return fmt.Errorf("%s expects 3 or 4 operands (mem, [Zn-Zn+3], (K), dst), got %d", mnem, len(ops))
|
||||
}
|
||||
mem, lst := ops[0], ops[1]
|
||||
dst := ops[len(ops)-1]
|
||||
mask := 0
|
||||
if len(ops) == 4 {
|
||||
k, ok := ops[2].(Reg)
|
||||
if !ok || !k.mask {
|
||||
return fmt.Errorf("%s: third operand must be an opmask register", mnem)
|
||||
}
|
||||
if k.idx == 0 {
|
||||
return fmt.Errorf("k0 is not a usable mask register")
|
||||
}
|
||||
mask = k.idx
|
||||
}
|
||||
list, ok := lst.(RegList)
|
||||
if !ok {
|
||||
return fmt.Errorf("%s: second operand must be a four-register list", mnem)
|
||||
}
|
||||
if list.Lo.size != q.width {
|
||||
return fmt.Errorf("%s: the register list must hold %d-bit vector registers", mnem, q.width*8)
|
||||
}
|
||||
dstReg, ok := dst.(Reg)
|
||||
if !ok || !dstReg.isVec() {
|
||||
return fmt.Errorf("%s: destination must be a vector register", mnem)
|
||||
}
|
||||
if dstReg.size != q.width {
|
||||
return fmt.Errorf("%s: the destination must be a %d-bit vector register", mnem, q.width*8)
|
||||
}
|
||||
if !memOperand(mem) {
|
||||
return fmt.Errorf("%s: the source must be a memory operand", mnem)
|
||||
}
|
||||
// The list owns V'VVVV; a scaled index in the EVEX-only half would fold
|
||||
// its fifth bit into the same field the list's low register occupies.
|
||||
if m, ok := mem.(Mem); ok && m.HasIndex && m.Index.idx >= 16 {
|
||||
return fmt.Errorf("%s: an index register above Z15 has no EVEX bit free", mnem)
|
||||
}
|
||||
if sfx.zeroing && mask == 0 {
|
||||
return fmt.Errorf("%s: zeroing (.Z) requires a mask register", mnem)
|
||||
}
|
||||
spec := evexSpec{mapSel: 2, opcode: q.opcode, w: 0, pp: 3, opdigit: -1, n: [3]int{16, 16, 16}}
|
||||
// The vector length follows the destination (512-bit for the ZMM forms,
|
||||
// 128-bit for the scalar ones), exactly as the oracle encodes it.
|
||||
return e.emitEvexFields(spec, dstReg.vecLenBit(), dstReg.idx, list.Lo.idx, mem, mask, sfx)
|
||||
}
|
||||
|
||||
// evexBcastSpec describes an EVEX broadcast (VPBROADCASTD/Q): the opcode
|
||||
@@ -529,32 +855,38 @@ type evexMoveSpec struct {
|
||||
n [3]int
|
||||
vecOK bool // the non-memory operand may be a vector register
|
||||
xmmOnly bool // wider than XMM registers are rejected
|
||||
nds3 bool // a three-operand register form exists (VMOVSD/VMOVSS)
|
||||
}
|
||||
|
||||
// evexMoveTable maps an upper-case EVEX move mnemonic to its encoding.
|
||||
var evexMoveTable = map[string]evexMoveSpec{
|
||||
// EVEX.128/256/512.F3.0F.W0, unaligned integer move.
|
||||
"VMOVDQU32": {1, 2, 0x6F, 0x7F, 0, [3]int{16, 32, 64}, true, false},
|
||||
"VMOVDQU32": {1, 2, 0x6F, 0x7F, 0, [3]int{16, 32, 64}, true, false, false},
|
||||
// EVEX.128/256/512.F3.0F.W1, unaligned qword move.
|
||||
"VMOVDQU64": {1, 2, 0x6F, 0x7F, 1, [3]int{16, 32, 64}, true, false},
|
||||
"VMOVDQU64": {1, 2, 0x6F, 0x7F, 1, [3]int{16, 32, 64}, true, false, false},
|
||||
// EVEX.128/256/512.F2.0F.W0, unaligned byte move (byte/word moves use the
|
||||
// F2 prefix, dword/qword moves F3; the element size only changes the tuple
|
||||
// semantics).
|
||||
"VMOVDQU8": {1, 3, 0x6F, 0x7F, 0, [3]int{16, 32, 64}, true, false},
|
||||
"VMOVDQU8": {1, 3, 0x6F, 0x7F, 0, [3]int{16, 32, 64}, true, false, false},
|
||||
// EVEX.128/256/512.F2.0F.W1, unaligned word move (shares the qword
|
||||
// encoding).
|
||||
"VMOVDQU16": {1, 3, 0x6F, 0x7F, 1, [3]int{16, 32, 64}, true, false},
|
||||
"VMOVDQU16": {1, 3, 0x6F, 0x7F, 1, [3]int{16, 32, 64}, true, false, false},
|
||||
// EVEX.128/256/512.66.0F.W1, unaligned packed double move.
|
||||
"VMOVUPD": {1, 1, 0x10, 0x11, 1, [3]int{16, 32, 64}, true, false},
|
||||
"VMOVUPD": {1, 1, 0x10, 0x11, 1, [3]int{16, 32, 64}, true, false, false},
|
||||
// EVEX.128/256/512, aligned packed moves.
|
||||
"VMOVAPS": {1, 0, 0x28, 0x29, 0, [3]int{16, 32, 64}, true, false},
|
||||
"VMOVAPD": {1, 1, 0x28, 0x29, 1, [3]int{16, 32, 64}, true, false},
|
||||
"VMOVAPS": {1, 0, 0x28, 0x29, 0, [3]int{16, 32, 64}, true, false, false},
|
||||
"VMOVAPD": {1, 1, 0x28, 0x29, 1, [3]int{16, 32, 64}, true, false, false},
|
||||
// EVEX.128/256/512.66.0F, aligned integer moves.
|
||||
"VMOVDQA32": {1, 1, 0x6F, 0x7F, 0, [3]int{16, 32, 64}, true, false},
|
||||
"VMOVDQA64": {1, 1, 0x6F, 0x7F, 1, [3]int{16, 32, 64}, true, false},
|
||||
"VMOVDQA32": {1, 1, 0x6F, 0x7F, 0, [3]int{16, 32, 64}, true, false, false},
|
||||
"VMOVDQA64": {1, 1, 0x6F, 0x7F, 1, [3]int{16, 32, 64}, true, false, false},
|
||||
// EVEX.128.F3.0F.W0, scalar single move, memory operands (the
|
||||
// three-operand register form is not supported).
|
||||
"VMOVSS": {1, 2, 0x10, 0x11, 0, [3]int{4, 4, 4}, false, true},
|
||||
"VMOVSS": {1, 2, 0x10, 0x11, 0, [3]int{4, 4, 4}, false, true, true},
|
||||
// EVEX.128.F2.0F.W1, scalar double move: memory operands and the
|
||||
// three-operand register form (VMOVSD dst, src1, src2).
|
||||
"VMOVSD": {1, 3, 0x10, 0x11, 1, [3]int{8, 8, 8}, false, true, true},
|
||||
// EVEX.128/256/512.0F.W0, unaligned packed single move.
|
||||
"VMOVUPS": {1, 0, 0x10, 0x11, 0, [3]int{16, 32, 64}, true, false, false},
|
||||
}
|
||||
|
||||
// isEvex reports whether the mnemonic has an EVEX encoding we handle.
|
||||
@@ -565,8 +897,10 @@ func isEvex(mnemUpper string) bool {
|
||||
if _, ok := evexBcastTable[mnemUpper]; ok {
|
||||
return true
|
||||
}
|
||||
_, ok := evexMoveTable[mnemUpper]
|
||||
return ok
|
||||
if _, ok := evexMoveTable[mnemUpper]; ok {
|
||||
return true
|
||||
}
|
||||
return isEvexQuad(mnemUpper)
|
||||
}
|
||||
|
||||
// evexRequired reports whether the operands force the EVEX encoding of a
|
||||
@@ -579,6 +913,13 @@ func evexRequired(upper string, ops []Operand) bool {
|
||||
if !inVex && !inVexMove {
|
||||
return true // EVEX-only mnemonic
|
||||
}
|
||||
// The byte-quad shifts have VEX register forms but EVEX-only memory
|
||||
// forms: a memory count source forces the EVEX encoding.
|
||||
if upper == "VPSLLDQ" || upper == "VPSRLDQ" {
|
||||
if slices.ContainsFunc(ops, memOperand) {
|
||||
return true
|
||||
}
|
||||
}
|
||||
for _, op := range ops {
|
||||
if r, ok := op.(Reg); ok && (r.size == 64 || r.mask || (r.isVec() && r.idx >= 16)) {
|
||||
return true
|
||||
@@ -741,6 +1082,14 @@ func (e *enc) encodeEvex(mnemUpper string, ops []Operand, sfx evexSuffix) error
|
||||
return e.encodeEvexRM(spec, ops, 0, sfx)
|
||||
}
|
||||
spec, inTable := evexTable[mnemUpper]
|
||||
if q, ok := evexQuadTable[mnemUpper]; ok {
|
||||
// The quad-register family carries no rounding, SAE or broadcast;
|
||||
// only masking and zeroing apply.
|
||||
if sfx.sae || sfx.bcst || sfx.rounding >= 0 {
|
||||
return fmt.Errorf("%s takes no rounding/SAE/broadcast suffix", mnemUpper)
|
||||
}
|
||||
return e.encodeEvexQuad(mnemUpper, q, ops, sfx)
|
||||
}
|
||||
if inTable {
|
||||
if (sfx.rounding >= 0 || sfx.sae) && !evexRound[mnemUpper] {
|
||||
return fmt.Errorf("%s: rounding/SAE is not supported for this instruction", mnemUpper)
|
||||
@@ -752,6 +1101,27 @@ func (e *enc) encodeEvex(mnemUpper string, ops []Operand, sfx evexSuffix) error
|
||||
}
|
||||
spec.n = [3]int{n, n, n}
|
||||
}
|
||||
// A mnemonic with an immediate and a register spelling (the
|
||||
// variable-count shifts, the permutes) encodes the register one
|
||||
// when the first operand is not an immediate.
|
||||
if len(ops) > 0 {
|
||||
if _, isImm := ops[0].(Imm); !isImm {
|
||||
if alt, ok := evexRegFormTable[mnemUpper]; ok {
|
||||
spec, inTable = alt, true
|
||||
}
|
||||
}
|
||||
}
|
||||
// The high/low half moves split by operand count: three operands
|
||||
// insert, two store (VMOVHPS m64, X1).
|
||||
if hs, ok := evexHptrTable[mnemUpper]; ok {
|
||||
if len(ops) == 2 {
|
||||
if hs.store.opcode == 0 {
|
||||
return fmt.Errorf("%s has no two-operand form", mnemUpper)
|
||||
}
|
||||
return e.encodeEvexRMRev(hs.store, ops, 0, sfx)
|
||||
}
|
||||
spec = hs.insert
|
||||
}
|
||||
} else if sfx.evexOnly() {
|
||||
return fmt.Errorf("%s: the instruction does not take rounding/SAE/broadcast suffixes", mnemUpper)
|
||||
}
|
||||
@@ -811,6 +1181,12 @@ func (e *enc) encodeEvex(mnemUpper string, ops []Operand, sfx evexSuffix) error
|
||||
}
|
||||
return e.encodeEvexMove(mnemUpper, ms, ops, mask, sfx)
|
||||
}
|
||||
if ps, ok := evexPrefGatherTable[mnemUpper]; ok {
|
||||
if sfx.any() {
|
||||
return fmt.Errorf("%s takes no EVEX suffixes", mnemUpper)
|
||||
}
|
||||
return e.encodeEvexPrefGather(mnemUpper, ps, ops, mask, sfx)
|
||||
}
|
||||
if !inTable {
|
||||
return fmt.Errorf("unsupported instruction %q for ZMM/K operands", mnemUpper)
|
||||
}
|
||||
@@ -829,6 +1205,8 @@ func (e *enc) encodeEvex(mnemUpper string, ops []Operand, sfx evexSuffix) error
|
||||
return e.encodeEvexNDS3Imm(spec, ops, mask, sfx)
|
||||
case vexExtract:
|
||||
return e.encodeEvexExtract(spec, ops, mask, sfx)
|
||||
case vexExtractGPR:
|
||||
return e.encodeEvexExtractGPR(spec, ops, mask, sfx)
|
||||
case vexRMSrcLen:
|
||||
return e.encodeEvexRMSrcLen(spec, ops, mask, sfx)
|
||||
}
|
||||
@@ -908,6 +1286,11 @@ func (e *enc) encodeEvexImmRM(spec evexSpec, ops []Operand, mask int, sfx evexSu
|
||||
if dstReg.mask {
|
||||
if r, ok := src.(Reg); ok && r.isVec() {
|
||||
ll = r.vecLenBit()
|
||||
} else if l, err := soleLen(spec.n); err == nil {
|
||||
// A memory source with a length-fixed mnemonic
|
||||
// (VFPCLASSPDX/Y/Z): the length comes from the table's
|
||||
// single valid slot, not from the operand.
|
||||
ll = l
|
||||
}
|
||||
} else if r, ok := src.(Reg); ok && r.isVec() {
|
||||
ll = r.vecLenBit()
|
||||
@@ -934,9 +1317,11 @@ func (e *enc) encodeEvexShiftImm(spec evexSpec, ops []Operand, mask int, sfx eve
|
||||
if !ok {
|
||||
return fmt.Errorf("shift count must be an immediate")
|
||||
}
|
||||
srcReg, ok := src.(Reg)
|
||||
if !ok || !srcReg.isVec() {
|
||||
return fmt.Errorf("shift source must be a vector register")
|
||||
// The count source is a vector register or memory; the length the L'L
|
||||
// field and the disp8×N multiplier follow is the destination's either
|
||||
// way.
|
||||
if !vecOrMem(src) {
|
||||
return fmt.Errorf("shift source must be a vector register or memory")
|
||||
}
|
||||
dstReg, ok := dst.(Reg)
|
||||
if !ok || !dstReg.isVec() {
|
||||
@@ -946,7 +1331,7 @@ func (e *enc) encodeEvexShiftImm(spec evexSpec, ops []Operand, mask int, sfx eve
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
if err := e.emitEvexFields(spec, dstReg.vecLenBit(), spec.opdigit, dstReg.idx, srcReg, mask, sfx); err != nil {
|
||||
if err := e.emitEvexFields(spec, dstReg.vecLenBit(), spec.opdigit, dstReg.idx, src, mask, sfx); err != nil {
|
||||
return err
|
||||
}
|
||||
e.out = append(e.out, immByte)
|
||||
@@ -1018,10 +1403,77 @@ func (e *enc) encodeEvexExtract(spec evexSpec, ops []Operand, mask int, sfx evex
|
||||
return nil
|
||||
}
|
||||
|
||||
// encodeEvexExtractGPR encodes the lane extract to a general-purpose
|
||||
// register or memory: OP $imm, xsrc, dst (reg = the XMM source, rm = the
|
||||
// destination, imm8). The encoding is 128-bit regardless of register
|
||||
// numbers, so L'L is fixed at 0 and the disp8×N multiplier is the extracted
|
||||
// element size the table carries.
|
||||
func (e *enc) encodeEvexExtractGPR(spec evexSpec, ops []Operand, mask int, sfx evexSuffix) error {
|
||||
if len(ops) != 3 {
|
||||
return fmt.Errorf("extract expects 3 operands ($imm, xsrc, dst), got %d", len(ops))
|
||||
}
|
||||
imm, src, dst := ops[0], ops[1], ops[2]
|
||||
immVal, ok := imm.(Imm)
|
||||
if !ok {
|
||||
return fmt.Errorf("extract lane must be an immediate")
|
||||
}
|
||||
srcReg, ok := src.(Reg)
|
||||
if !ok || !srcReg.isVec() {
|
||||
return fmt.Errorf("extract source must be a vector register")
|
||||
}
|
||||
switch dst.(type) {
|
||||
case Reg:
|
||||
if dst.(Reg).isVec() {
|
||||
return fmt.Errorf("extract destination must be a general-purpose register or memory")
|
||||
}
|
||||
case Mem, sbMem:
|
||||
default:
|
||||
return fmt.Errorf("extract destination must be a general-purpose register or memory")
|
||||
}
|
||||
immByte, err := imm8(int64(immVal))
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
if err := e.emitEvexFields(spec, 0, srcReg.idx, -1, dst, mask, sfx); err != nil {
|
||||
return err
|
||||
}
|
||||
e.out = append(e.out, immByte)
|
||||
return nil
|
||||
}
|
||||
|
||||
// encodeEvexMove encodes a two-operand EVEX move; a vector→vector move uses
|
||||
// the store-form opcode (reg = source, rm = destination), matching the Go
|
||||
// assembler.
|
||||
// assembler. The scalar moves also carry a three-operand register form
|
||||
// (VMOVSD dst, src1, src2: the load opcode with vvvv = src1), which ms.nds3
|
||||
// opens.
|
||||
func (e *enc) encodeEvexMove(mnem string, ms evexMoveSpec, ops []Operand, mask int, sfx evexSuffix) error {
|
||||
if len(ops) == 3 {
|
||||
if !ms.nds3 {
|
||||
return fmt.Errorf("EVEX move expects 2 operands, got %d", len(ops))
|
||||
}
|
||||
// The masked scalar register form keeps the Go assembler's own
|
||||
// layout: the store opcode with reg = op0, vvvv = op1 and the
|
||||
// destination in r/m (op2) — the bytes go tool asm emits, not
|
||||
// the manual's NDS reading.
|
||||
src, src1, dst := ops[0], ops[1], ops[2]
|
||||
reg, ok := src.(Reg)
|
||||
if !ok || !reg.isVec() {
|
||||
return fmt.Errorf("%s: first operand must be a vector register", mnem)
|
||||
}
|
||||
vvvvReg, ok := src1.(Reg)
|
||||
if !ok || !vvvvReg.isVec() {
|
||||
return fmt.Errorf("%s: second operand must be a vector register", mnem)
|
||||
}
|
||||
dstReg, ok := dst.(Reg)
|
||||
if !ok || !dstReg.isVec() {
|
||||
return fmt.Errorf("%s: destination must be a vector register", mnem)
|
||||
}
|
||||
if ms.xmmOnly && (reg.size != 16 || vvvvReg.size != 16 || dstReg.size != 16) {
|
||||
return fmt.Errorf("%s operates on XMM registers only", mnem)
|
||||
}
|
||||
spec := evexSpec{mapSel: ms.mapSel, opcode: ms.store, w: ms.w, pp: ms.pp, opdigit: -1, n: ms.n}
|
||||
return e.emitEvexFields(spec, dstReg.vecLenBit(), reg.idx, vvvvReg.idx, dst, mask, sfx)
|
||||
}
|
||||
if len(ops) != 2 {
|
||||
return fmt.Errorf("EVEX move expects 2 operands, got %d", len(ops))
|
||||
}
|
||||
@@ -1141,12 +1593,20 @@ func (e *enc) encodeEvexBcast(bs evexBcastSpec, ops []Operand, mask int, sfx eve
|
||||
return fmt.Errorf("broadcast destination must be a vector register")
|
||||
}
|
||||
spec := evexSpec{mapSel: bs.mapSel, w: bs.w, pp: 1, opdigit: -1}
|
||||
switch src.(type) {
|
||||
switch r := src.(type) {
|
||||
case Mem, sbMem:
|
||||
spec.opcode = bs.opMem
|
||||
spec.n = [3]int{bs.n, bs.n, bs.n}
|
||||
case Reg:
|
||||
spec.opcode = bs.opReg
|
||||
// A GPR source uses the register broadcast opcode; a vector
|
||||
// source shares the xmm/mem one (the low byte is copied from
|
||||
// the lane or from the memory operand).
|
||||
if r.isVec() {
|
||||
spec.opcode = bs.opMem
|
||||
spec.n = [3]int{bs.n, bs.n, bs.n}
|
||||
} else {
|
||||
spec.opcode = bs.opReg
|
||||
}
|
||||
default:
|
||||
return fmt.Errorf("broadcast source must be a register or memory")
|
||||
}
|
||||
@@ -1349,6 +1809,102 @@ func isScatter(upper string) bool {
|
||||
return ok
|
||||
}
|
||||
|
||||
// isEvexPrefGather reports whether the mnemonic is a gather/scatter
|
||||
// prefetch hint.
|
||||
func isEvexPrefGather(upper string) bool {
|
||||
_, ok := evexPrefGatherTable[upper]
|
||||
return ok
|
||||
}
|
||||
|
||||
// evexRegFormTable holds the register-count twin of the immediate-form
|
||||
// entries in evexTable. Several mnemonics name two encodings: an immediate
|
||||
// count or control ($imm, src, dst …) and a register-count one whose second
|
||||
// operand is a vector register or memory (count, src2, src1, dst). The
|
||||
// immediate spelling lives in evexTable, this table carries the register
|
||||
// spelling, and encodeEvex picks by whether the first operand is an
|
||||
// immediate, the way vexVarShift does on the VEX side.
|
||||
var evexRegFormTable = map[string]evexSpec{
|
||||
"VPSLLD": {1, 0xF2, 0, 1, -1, vexNDS3, [3]int{16, 16, 16}},
|
||||
"VPSLLQ": {1, 0xF3, 1, 1, -1, vexNDS3, [3]int{16, 16, 16}},
|
||||
"VPSLLW": {1, 0xF1, 0, 1, -1, vexNDS3, [3]int{16, 16, 16}},
|
||||
"VPSRAD": {1, 0xE2, 0, 1, -1, vexNDS3, [3]int{16, 16, 16}},
|
||||
"VPSRAQ": {1, 0xE2, 1, 1, -1, vexNDS3, [3]int{16, 16, 16}},
|
||||
"VPSRAW": {1, 0xE1, 0, 1, -1, vexNDS3, [3]int{16, 16, 16}},
|
||||
"VPSRLD": {1, 0xD2, 0, 1, -1, vexNDS3, [3]int{16, 16, 16}},
|
||||
"VPSRLQ": {1, 0xD3, 1, 1, -1, vexNDS3, [3]int{16, 16, 16}},
|
||||
"VPSRLW": {1, 0xD1, 0, 1, -1, vexNDS3, [3]int{16, 16, 16}},
|
||||
// EVEX.NDS.0F38.W1, the register-count permutes (the immediate
|
||||
// controls live in evexTable under 0F3A).
|
||||
"VPERMQ": {2, 0x36, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPERMPD": {2, 0x16, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
// EVEX.NDS.0F38, the register-count permil shuffles.
|
||||
"VPERMILPS": {2, 0x0C, 0, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
"VPERMILPD": {2, 0x0D, 1, 1, -1, vexNDS3, [3]int{16, 32, 64}},
|
||||
}
|
||||
|
||||
// evexPrefGatherSpec describes a gather/scatter prefetch hint: one memory
|
||||
// operand with a VSIB index and an opmask register, no destination. The
|
||||
// ModRM.reg field carries a fixed /digit, the L'L field is fixed at 512, and
|
||||
// the mask register is the instruction's only register operand.
|
||||
type evexPrefGatherSpec struct {
|
||||
mapSel int
|
||||
opcode byte
|
||||
w int
|
||||
pp int
|
||||
opdigit int
|
||||
n int
|
||||
}
|
||||
|
||||
var evexPrefGatherTable = map[string]evexPrefGatherSpec{
|
||||
"VGATHERPF0DPD": {2, 0xC6, 1, 1, 1, 8},
|
||||
"VGATHERPF0DPS": {2, 0xC6, 0, 1, 1, 4},
|
||||
"VGATHERPF0QPD": {2, 0xC7, 1, 1, 1, 8},
|
||||
"VGATHERPF0QPS": {2, 0xC7, 0, 1, 1, 4},
|
||||
"VGATHERPF1DPD": {2, 0xC6, 1, 1, 2, 8},
|
||||
"VGATHERPF1DPS": {2, 0xC6, 0, 1, 2, 4},
|
||||
"VGATHERPF1QPD": {2, 0xC7, 1, 1, 2, 8},
|
||||
"VGATHERPF1QPS": {2, 0xC7, 0, 1, 2, 4},
|
||||
"VSCATTERPF0DPD": {2, 0xC6, 1, 1, 5, 8},
|
||||
"VSCATTERPF0DPS": {2, 0xC6, 0, 1, 5, 4},
|
||||
"VSCATTERPF0QPD": {2, 0xC7, 1, 1, 5, 8},
|
||||
"VSCATTERPF0QPS": {2, 0xC7, 0, 1, 5, 4},
|
||||
"VSCATTERPF1DPD": {2, 0xC6, 1, 1, 6, 8},
|
||||
"VSCATTERPF1DPS": {2, 0xC6, 0, 1, 6, 4},
|
||||
"VSCATTERPF1QPD": {2, 0xC7, 1, 1, 6, 8},
|
||||
"VSCATTERPF1QPS": {2, 0xC7, 0, 1, 6, 4},
|
||||
}
|
||||
|
||||
// evexHptrSpec describes the high/low half moves (VMOVHPS family): the
|
||||
// three-operand insert shares an opcode with a two-operand store whose
|
||||
// source is the vector register and whose destination is m64.
|
||||
type evexHptrSpec struct {
|
||||
insert evexSpec
|
||||
store evexSpec // store.opcode == 0 when the mnemonic has no store form
|
||||
}
|
||||
|
||||
var evexHptrTable = map[string]evexHptrSpec{
|
||||
"VMOVHPS": {
|
||||
insert: evexSpec{mapSel: 1, opcode: 0x16, w: 0, pp: 0, opdigit: -1, form: vexNDS3, n: [3]int{8, 0, 0}},
|
||||
store: evexSpec{mapSel: 1, opcode: 0x17, w: 0, pp: 0, opdigit: -1, form: vexRMRev, n: [3]int{8, 0, 0}},
|
||||
},
|
||||
"VMOVLHPS": {
|
||||
insert: evexSpec{mapSel: 1, opcode: 0x16, w: 0, pp: 0, opdigit: -1, form: vexNDS3, n: [3]int{8, 0, 0}},
|
||||
},
|
||||
}
|
||||
|
||||
// encodeEvexPrefGather encodes a gather/scatter prefetch hint: OP K, vsib.
|
||||
func (e *enc) encodeEvexPrefGather(upper string, ps evexPrefGatherSpec, ops []Operand, mask int, sfx evexSuffix) error {
|
||||
if len(ops) != 1 {
|
||||
return fmt.Errorf("%s expects 2 operands (K, vsib memory), got %d", upper, len(ops)+1)
|
||||
}
|
||||
m, ok := ops[0].(Mem)
|
||||
if !ok || !m.HasIndex || !m.Index.isVec() {
|
||||
return fmt.Errorf("%s: operand must be a VSIB memory reference with a vector index", upper)
|
||||
}
|
||||
spec := evexSpec{mapSel: ps.mapSel, opcode: ps.opcode, w: ps.w, pp: ps.pp, opdigit: ps.opdigit, n: [3]int{ps.n, ps.n, ps.n}}
|
||||
return e.emitEvexFields(spec, 2, ps.opdigit, -1, m, mask, sfx)
|
||||
}
|
||||
|
||||
// vsibLen validates a VSIB memory operand (the index must be a vector
|
||||
// register) and returns it with the vector length the index selects, the
|
||||
// EVEX L'L field follows the index register, not the data register.
|
||||
@@ -1370,7 +1926,9 @@ func (e *enc) encodeGather(upper string, gs gatherSpec, ops []Operand, sfx evexS
|
||||
return err
|
||||
}
|
||||
if mask != 0 || sfx.any() {
|
||||
// EVEX form: OP vsib, K, dst.
|
||||
// EVEX form: OP vsib, K, dst. The L'L field is the wider of the
|
||||
// index and the data register lengths (the Go assembler's
|
||||
// layout); the disp8×N multiplier stays the index element size.
|
||||
if len(rest) != 2 {
|
||||
return fmt.Errorf("%s expects 3 operands (vsib, K, dst), got %d", upper, len(ops))
|
||||
}
|
||||
@@ -1382,6 +1940,9 @@ func (e *enc) encodeGather(upper string, gs gatherSpec, ops []Operand, sfx evexS
|
||||
if !ok || !dst.isVec() {
|
||||
return fmt.Errorf("%s: destination must be a vector register", upper)
|
||||
}
|
||||
if d := dst.vecLenBit(); d > ll {
|
||||
ll = d
|
||||
}
|
||||
evex := evexSpec{mapSel: 2, opcode: gs.opcode, w: gs.w, pp: 1, opdigit: -1, n: [3]int{gs.n, gs.n, gs.n}}
|
||||
return e.emitEvexFields(evex, ll, dst.idx, -1, vsib, mask, sfx)
|
||||
}
|
||||
@@ -1431,6 +1992,11 @@ func (e *enc) encodeScatter(upper string, ss gatherSpec, ops []Operand, sfx evex
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
// The L'L field is the wider of the data register and the VSIB index
|
||||
// lengths, the bytes go tool asm emits.
|
||||
if d := src.vecLenBit(); d > ll {
|
||||
ll = d
|
||||
}
|
||||
evex := evexSpec{mapSel: 2, opcode: ss.opcode, w: ss.w, pp: 1, opdigit: -1, n: [3]int{ss.n, ss.n, ss.n}}
|
||||
return e.emitEvexFields(evex, ll, src.idx, -1, vsib, mask, sfx)
|
||||
}
|
||||
@@ -1441,6 +2007,8 @@ func (e *enc) encodeScatter(upper string, ss gatherSpec, ops []Operand, sfx evex
|
||||
var evexKOperand = map[string]bool{
|
||||
"VPMOVM2B": true, "VPMOVM2W": true, "VPMOVM2D": true, "VPMOVM2Q": true,
|
||||
"VPMOVB2M": true, "VPMOVW2M": true, "VPMOVD2M": true, "VPMOVQ2M": true,
|
||||
// The K-to-vector broadcast reads its opmask source from r/m.
|
||||
"VPBROADCASTMB2Q": true, "VPBROADCASTMW2D": true,
|
||||
}
|
||||
|
||||
// kmovSpec describes a KMOV width: the opcode depends on the operand
|
||||
@@ -1541,6 +2109,7 @@ var kOpsTable = map[string]kOpSpec{
|
||||
"KXORD": {1, 0x47, 1, 1, 1, vexNDS3},
|
||||
"KXORQ": {1, 0x47, 1, 0, 1, vexNDS3},
|
||||
"KUNPCKBW": {1, 0x4B, 0, 1, 1, vexNDS3},
|
||||
"KUNPCKWD": {1, 0x4B, 0, 0, 1, vexNDS3},
|
||||
"KUNPCKDQ": {1, 0x4B, 1, 0, 1, vexNDS3},
|
||||
"KADDB": {1, 0x4A, 0, 1, 1, vexNDS3},
|
||||
"KADDW": {1, 0x4A, 0, 0, 1, vexNDS3},
|
||||
|
||||
@@ -721,3 +721,216 @@ func hexCompact(b []byte) string {
|
||||
}
|
||||
return string(out)
|
||||
}
|
||||
|
||||
// TestAvx512CorpusFamilies pins representative encodings of the AVX-512
|
||||
// families the toolchain's avx512enc corpus exercises: the bytes are the
|
||||
// go tool asm output for exactly these operands, and the same families are
|
||||
// covered end to end by the avx512_amd64.s differential kernel.
|
||||
func TestAvx512CorpusFamilies(t *testing.T) {
|
||||
vsib := func(base, idx string, scale int) Operand {
|
||||
return Idx(vreg(t, base), vreg(t, idx), scale, 0, 0)
|
||||
}
|
||||
cases := []struct {
|
||||
name string
|
||||
mnem string
|
||||
ops []Operand
|
||||
want string
|
||||
}{
|
||||
// AES rounds (EVEX NDS, VEX twin routed by operand width).
|
||||
{"VAESDEC Z", "VAESDEC", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "Z3")}, "62f26d48ded9"},
|
||||
// Integer VNNI and the bit algorithm group.
|
||||
{"VPDPBUSD", "VPDPBUSD", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "K2"), vreg(t, "Z3")}, "62f26d4a50d9"},
|
||||
{"VPOPCNTW", "VPOPCNTW", []Operand{vreg(t, "Z1"), vreg(t, "K3"), vreg(t, "Z2")}, "62f2fd4b54d1"},
|
||||
{"VPCONFLICTD", "VPCONFLICTD", []Operand{vreg(t, "Z1"), vreg(t, "K1"), vreg(t, "Z2")}, "62f27d49c4d1"},
|
||||
{"VPLZCNTQ masked", "VPLZCNTQ", []Operand{vreg(t, "Z7"), vreg(t, "K1"), vreg(t, "Z8")}, "6272fd4944c7"},
|
||||
{"VPERMT2B", "VPERMT2B", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "K1"), vreg(t, "Z3")}, "62f26d497dd9"},
|
||||
{"VPMULTISHIFTQB", "VPMULTISHIFTQB", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "K3"), vreg(t, "Z4")}, "62f2ed4b83e1"},
|
||||
{"VDBPSADBW", "VDBPSADBW", []Operand{Imm(3), vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "K3"), vreg(t, "Z3")}, "62f36d4b42d903"},
|
||||
{"VPSHUFBITQMB", "VPSHUFBITQMB", []Operand{vreg(t, "Z9"), vreg(t, "Z10"), vreg(t, "K3")}, "62d22d488fd9"},
|
||||
{"VPTESTNMQ", "VPTESTNMQ", []Operand{vreg(t, "Z13"), vreg(t, "Z14"), vreg(t, "K5")}, "62d28e4827ed"},
|
||||
// Permutations: immediate and register counts.
|
||||
{"VALIGNQ", "VALIGNQ", []Operand{Imm(3), vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "K1"), vreg(t, "Z3")}, "62f3ed4903d903"},
|
||||
{"VPERMQ imm", "VPERMQ", []Operand{Imm(1), vreg(t, "Z1"), vreg(t, "K2"), vreg(t, "Z2")}, "62f3fd4a00d101"},
|
||||
{"VPERMQ reg", "VPERMQ", []Operand{vreg(t, "Z3"), vreg(t, "Z4"), vreg(t, "K2"), vreg(t, "Z5")}, "62f2dd4a36eb"},
|
||||
{"VPERMPD reg", "VPERMPD", []Operand{vreg(t, "Z1"), vreg(t, "Z2"), vreg(t, "Z3")}, "62f2ed4816d9"},
|
||||
{"VPERMILPS imm", "VPERMILPS", []Operand{Imm(5), vreg(t, "Z9"), vreg(t, "K2"), vreg(t, "Z10")}, "62537d4a04d105"},
|
||||
{"VPERMILPS reg", "VPERMILPS", []Operand{vreg(t, "Z11"), vreg(t, "Z12"), vreg(t, "K2"), vreg(t, "Z13")}, "62521d4a0ceb"},
|
||||
// Shifts: immediate, register-count and memory-count forms; the
|
||||
// count source carries its own XMM tuple width.
|
||||
{"VPSLLW imm mask", "VPSLLW", []Operand{Imm(3), vreg(t, "Z1"), vreg(t, "K2"), vreg(t, "Z2")}, "62f16d4a71f103"},
|
||||
{"VPSLLD reg count", "VPSLLD", []Operand{vreg(t, "X1"), vreg(t, "Z2"), vreg(t, "K1"), vreg(t, "Z3")}, "62f16d49f2d9"},
|
||||
{"VPSLLDQ", "VPSLLDQ", []Operand{Imm(9), vreg(t, "Z7"), vreg(t, "Z8")}, "62f13d4873ff09"},
|
||||
{"VPSRLDQ mem", "VPSRLDQ", []Operand{Imm(11), Ptr(SI, 16, 16), vreg(t, "Z4")}, "62f15d48739e100000000b"},
|
||||
{"VPSRLVW", "VPSRLVW", []Operand{vreg(t, "Z3"), vreg(t, "Z4"), vreg(t, "K1"), vreg(t, "Z5")}, "62f2dd4910eb"},
|
||||
// Conversions and shuffles with the F2 prefix and no prefix.
|
||||
{"VCVTUDQ2PS", "VCVTUDQ2PS", []Operand{vreg(t, "Z1"), vreg(t, "K1"), vreg(t, "Z2")}, "62f17f497ad1"},
|
||||
{"VSHUFPS", "VSHUFPS", []Operand{Imm(2), vreg(t, "Z4"), vreg(t, "Z5"), vreg(t, "K1"), vreg(t, "Z6")}, "62f15449c6f402"},
|
||||
// Gather and scatter prefetch hints (memory-only, /digit in reg).
|
||||
{"VGATHERPF0DPD", "VGATHERPF0DPD", []Operand{vreg(t, "K5"), vsib("R10", "Y29", 8)}, "6292fd45c60cea"},
|
||||
{"VSCATTERPF1DPS", "VSCATTERPF1DPS", []Operand{vreg(t, "K2"), vsib("R10", "Z28", 4)}, "62927d42c634a2"},
|
||||
// Opmask broadcasts and the K logic.
|
||||
{"VPBROADCASTMB2Q", "VPBROADCASTMB2Q", []Operand{vreg(t, "K1"), vreg(t, "Z2")}, "62f2fe482ad1"},
|
||||
{"VPBROADCASTMW2D", "VPBROADCASTMW2D", []Operand{vreg(t, "K3"), vreg(t, "Z4")}, "62f27e483ae3"},
|
||||
{"KUNPCKWD", "KUNPCKWD", []Operand{vreg(t, "K6"), vreg(t, "K4"), vreg(t, "K1")}, "c5dc4bce"},
|
||||
{"KADDB", "KADDB", []Operand{vreg(t, "K2"), vreg(t, "K3"), vreg(t, "K5")}, "c5e54aea"},
|
||||
// Lane extracts to general registers (EVEX and VEX routes).
|
||||
{"VPEXTRB", "VPEXTRB", []Operand{Imm(3), vreg(t, "X26"), AX}, "62637d0814d003"},
|
||||
{"VPEXTRD", "VPEXTRD", []Operand{Imm(1), vreg(t, "X26"), vreg(t, "R9")}, "62437d0816d101"},
|
||||
{"VPEXTRD vex", "VPEXTRD", []Operand{Imm(1), vreg(t, "X2"), DI}, "c4e37916d701"},
|
||||
{"VPINSRQ", "VPINSRQ", []Operand{Imm(1), DI, vreg(t, "X3"), vreg(t, "X4")}, "c4e3e122e701"},
|
||||
// Moves: masked unaligned, masked scalar register form, half moves
|
||||
// and non-temporal stores.
|
||||
{"VMOVUPS mask", "VMOVUPS", []Operand{vreg(t, "Z1"), vreg(t, "K2"), vreg(t, "Z3")}, "62f17c4a11cb"},
|
||||
{"VMOVSD 3op", "VMOVSD", []Operand{vreg(t, "X14"), vreg(t, "X5"), vreg(t, "K3"), vreg(t, "X22")}, "6231d70b11f6"},
|
||||
{"VMOVSS 3op", "VMOVSS", []Operand{vreg(t, "X18"), vreg(t, "X3"), vreg(t, "K2"), vreg(t, "X25")}, "6281660a11d1"},
|
||||
{"VMOVHPS insert", "VMOVHPS", []Operand{Ptr(SI, 0, 8), vreg(t, "X18"), vreg(t, "X19")}, "62e16c00161e"},
|
||||
{"VMOVHPS store", "VMOVHPS", []Operand{vreg(t, "X20"), Ptr(SI, 8, 8)}, "62e17c08176601"},
|
||||
{"VMOVLHPS", "VMOVLHPS", []Operand{vreg(t, "X16"), vreg(t, "X5"), vreg(t, "X17")}, "62a1540816c8"},
|
||||
{"VMOVNTDQ", "VMOVNTDQ", []Operand{vreg(t, "Z7"), Ptr(SI, 0, 64)}, "62f17d48e73e"},
|
||||
{"VMOVNTDQA", "VMOVNTDQA", []Operand{Ptr(SI, 64, 64), vreg(t, "Z8")}, "62727d482a4601"},
|
||||
{"VMOVNTPS", "VMOVNTPS", []Operand{vreg(t, "Z9"), Ptr(SI, 0, 64)}, "62717c482b0e"},
|
||||
// Scalar compares with and without the 66 prefix.
|
||||
{"VCOMISD", "VCOMISD", []Operand{vreg(t, "X5"), vreg(t, "X6")}, "c5f92ff5"},
|
||||
{"VUCOMISS", "VUCOMISS", []Operand{vreg(t, "X7"), vreg(t, "X8")}, "c5782ec7"},
|
||||
// Floating point helpers.
|
||||
{"VSQRTSD", "VSQRTSD", []Operand{vreg(t, "X1"), vreg(t, "X2"), vreg(t, "K1"), vreg(t, "X3")}, "62f1ef0951d9"},
|
||||
{"VEXP2PD", "VEXP2PD", []Operand{vreg(t, "Z5"), vreg(t, "K1"), vreg(t, "Z6")}, "62f2fd49c8f5"},
|
||||
{"VRCP28SD", "VRCP28SD", []Operand{vreg(t, "X9"), vreg(t, "X8"), vreg(t, "K1"), vreg(t, "X10")}, "6252bd09cbd1"},
|
||||
{"VBROADCASTF32X2", "VBROADCASTF32X2", []Operand{vreg(t, "X1"), vreg(t, "K1"), vreg(t, "Z2")}, "62f27d4919d1"},
|
||||
{"VPCOMPRESSB", "VPCOMPRESSB", []Operand{vreg(t, "Z1"), vreg(t, "K1"), Ptr(SI, 0, 64)}, "62f27d49630e"},
|
||||
}
|
||||
for _, c := range cases {
|
||||
code, err := Encode(c.mnem, c.ops...)
|
||||
if err != nil {
|
||||
t.Errorf("%s: Encode: %v", c.name, err)
|
||||
continue
|
||||
}
|
||||
if got := hexCompact(code); got != c.want {
|
||||
t.Errorf("%s: got %s, want %s", c.name, got, c.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestEvexQuadRegisterGroundTruth pins the quad-register instructions (the
|
||||
// 4FMAPS and 4VNNIW families) byte for byte against go tool asm: the memory
|
||||
// source keeps r/m, the bracketed list's LOW register travels the inverted
|
||||
// 5-bit V'VVVV field, the destination sits in reg, the opmask rides aaa and
|
||||
// the vector length follows the destination (L'L=512 for the ZMM forms,
|
||||
// 128 for the scalar ones) while the disp8×N multiplier stays 16 for every
|
||||
// member. The x86 decoder has no view of these forms, so no decode check
|
||||
// runs.
|
||||
func TestEvexQuadRegisterGroundTruth(t *testing.T) {
|
||||
sp := vreg(t, "RSP")
|
||||
cases := []struct {
|
||||
name string
|
||||
mnem string
|
||||
ops []Operand
|
||||
want string
|
||||
}{
|
||||
{"V4FMADDPS 17(SP) [Z0-Z3] K2 Z0", "V4FMADDPS",
|
||||
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "K2"), vreg(t, "Z0")},
|
||||
"62f27f4a9a842411000000"},
|
||||
{"V4FMADDPS [Z10-Z13]", "V4FMADDPS",
|
||||
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z10"), vreg(t, "Z13")}, vreg(t, "K2"), vreg(t, "Z0")},
|
||||
"62f22f4a9a842411000000"},
|
||||
{"V4FMADDPS [Z20-Z23]", "V4FMADDPS",
|
||||
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z20"), vreg(t, "Z23")}, vreg(t, "K2"), vreg(t, "Z0")},
|
||||
"62f25f429a842411000000"},
|
||||
{"V4FMADDPS Z8 dst", "V4FMADDPS",
|
||||
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "K2"), vreg(t, "Z8")},
|
||||
"62727f4a9a842411000000"},
|
||||
{"V4FMADDPS disp8x16", "V4FMADDPS",
|
||||
[]Operand{Ptr(sp, 64, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "K2"), vreg(t, "Z0")},
|
||||
"62f27f4a9a442404"},
|
||||
{"V4FMADDPS unmasked", "V4FMADDPS",
|
||||
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "Z0")},
|
||||
"62f27f489a842411000000"},
|
||||
{"V4FMADDSS 7(AX) [X0-X3] K5 X22", "V4FMADDSS",
|
||||
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X0"), vreg(t, "X3")}, vreg(t, "K5"), vreg(t, "X22")},
|
||||
"62e27f0d9bb007000000"},
|
||||
{"V4FMADDSS (DI)", "V4FMADDSS",
|
||||
[]Operand{Ptr(DI, 0, 8), RegList{vreg(t, "X0"), vreg(t, "X3")}, vreg(t, "K5"), vreg(t, "X22")},
|
||||
"62e27f0d9b37"},
|
||||
{"V4FMADDSS [X10-X13]", "V4FMADDSS",
|
||||
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X10"), vreg(t, "X13")}, vreg(t, "K5"), vreg(t, "X22")},
|
||||
"62e22f0d9bb007000000"},
|
||||
{"V4FMADDSS [X20-X23]", "V4FMADDSS",
|
||||
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X20"), vreg(t, "X23")}, vreg(t, "K5"), vreg(t, "X22")},
|
||||
"62e25f059bb007000000"},
|
||||
{"V4FMADDSS X30 dst", "V4FMADDSS",
|
||||
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X0"), vreg(t, "X3")}, vreg(t, "K5"), vreg(t, "X30")},
|
||||
"62627f0d9bb007000000"},
|
||||
{"V4FMADDSS X3 dst", "V4FMADDSS",
|
||||
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X0"), vreg(t, "X3")}, vreg(t, "K5"), vreg(t, "X3")},
|
||||
"62f27f0d9b9807000000"},
|
||||
{"V4FMADDSS disp8x16", "V4FMADDSS",
|
||||
[]Operand{Ptr(AX, 16, 8), RegList{vreg(t, "X20"), vreg(t, "X23")}, vreg(t, "K5"), vreg(t, "X30")},
|
||||
"62625f059b7001"},
|
||||
{"V4FNMADDPS", "V4FNMADDPS",
|
||||
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "K2"), vreg(t, "Z0")},
|
||||
"62f27f4aaa842411000000"},
|
||||
{"V4FNMADDSS", "V4FNMADDSS",
|
||||
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X0"), vreg(t, "X3")}, vreg(t, "K5"), vreg(t, "X22")},
|
||||
"62e27f0dabb007000000"},
|
||||
{"VP4DPWSSD", "VP4DPWSSD",
|
||||
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "K2"), vreg(t, "Z0")},
|
||||
"62f27f4a52842411000000"},
|
||||
{"VP4DPWSSDS unmasked", "VP4DPWSSDS",
|
||||
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "Z0")},
|
||||
"62f27f4853842411000000"},
|
||||
}
|
||||
for _, c := range cases {
|
||||
code, err := Encode(c.mnem, c.ops...)
|
||||
if err != nil {
|
||||
t.Errorf("%s: Encode: %v", c.name, err)
|
||||
continue
|
||||
}
|
||||
if got := hexCompact(code); got != c.want {
|
||||
t.Errorf("%s: got %s, want %s", c.name, got, c.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestEvexQuadRegisterErrors pins the operand shapes the toolchain rejects:
|
||||
// the register class the list and the destination take is fixed per
|
||||
// instruction, the source is memory only, the opmask slot is positional and
|
||||
// the list's low register owns V'VVVV.
|
||||
func TestEvexQuadRegisterErrors(t *testing.T) {
|
||||
sp := vreg(t, "RSP")
|
||||
list := func(lo, hi string) RegList {
|
||||
return RegList{vreg(t, lo), vreg(t, hi)}
|
||||
}
|
||||
cases := []struct {
|
||||
name string
|
||||
mnem string
|
||||
ops []Operand
|
||||
}{
|
||||
{"X list on the PS form", "V4FMADDPS",
|
||||
[]Operand{Ptr(sp, 0, 8), list("X0", "X3"), vreg(t, "K2"), vreg(t, "Z0")}},
|
||||
{"Z list on the SS form", "V4FMADDSS",
|
||||
[]Operand{Ptr(AX, 0, 8), list("Z0", "Z3"), vreg(t, "K5"), vreg(t, "X22")}},
|
||||
{"Y destination", "V4FMADDPS",
|
||||
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "K2"), vreg(t, "Y0")}},
|
||||
{"register source", "V4FMADDPS",
|
||||
[]Operand{vreg(t, "Z1"), list("Z0", "Z3"), vreg(t, "K2"), vreg(t, "Z0")}},
|
||||
{"non-mask third operand", "V4FMADDPS",
|
||||
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "Z4"), vreg(t, "Z0")}},
|
||||
{"k0 mask", "V4FMADDPS",
|
||||
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "K0"), vreg(t, "Z0")}},
|
||||
{"K after the destination", "V4FMADDPS",
|
||||
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "Z0"), vreg(t, "K2")}},
|
||||
{"zeroing without a mask", "V4FMADDPS.Z",
|
||||
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "Z0")}},
|
||||
{"SAE suffix", "V4FMADDPS.SAE",
|
||||
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "K2"), vreg(t, "Z0")}},
|
||||
{"high index source", "VP4DPWSSD",
|
||||
[]Operand{Idx(DI, vreg(t, "X16"), 1, 0, 8), list("Z0", "Z3"), vreg(t, "K2"), vreg(t, "Z0")}},
|
||||
{"short operand list", "V4FMADDPS",
|
||||
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3")}},
|
||||
}
|
||||
for _, c := range cases {
|
||||
if _, err := Encode(c.mnem, c.ops...); err == nil {
|
||||
t.Errorf("%s: expected an error, got none", c.name)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -463,6 +463,61 @@ func (img *Image) emitGOObject(pkgPath, srcPath string, pre []byte, minLC int, r
|
||||
symRelocs[si] = append(symRelocs[si], rec[:]...)
|
||||
}
|
||||
}
|
||||
// The data symbols' own relocations: the symbol-valued DATA fields
|
||||
// ("DATA s+0(SB)/8, $other(SB)"). The toolchain patches each field
|
||||
// with the target's absolute address through an R_ADDR of the DATA
|
||||
// line's width, on every architecture (the code relocations are
|
||||
// per-architecture PC-relative shapes; a data pointer word is not), so
|
||||
// this mapping bypasses relocField. The definitions were appended in
|
||||
// DataSyms order, so data symbol i is definition index i.
|
||||
for i, d := range img.DataSyms {
|
||||
for _, r := range d.Relocs {
|
||||
if r.Kind != RelAddr {
|
||||
return nil, fmt.Errorf("GOOBJ emission: data symbol %q carries a non-data relocation", d.Name)
|
||||
}
|
||||
var rec [23]byte
|
||||
binary.LittleEndian.PutUint32(rec[0:], uint32(int32(r.Off)))
|
||||
rec[4] = r.Siz
|
||||
binary.LittleEndian.PutUint16(rec[5:], relocAddr)
|
||||
binary.LittleEndian.PutUint64(rec[7:], uint64(r.Addend))
|
||||
switch {
|
||||
case r.External && r.Name == goobjBuiltinMorestack:
|
||||
binary.LittleEndian.PutUint32(rec[15:], pkgIdxBuiltin)
|
||||
binary.LittleEndian.PutUint32(rec[19:], goobjBuiltinMorestackNoctxt)
|
||||
case r.External:
|
||||
pkg, name := splitQualified(r.Name)
|
||||
if pkg == "" {
|
||||
return nil, fmt.Errorf("GOOBJ emission: external symbol %q has no package prefix", r.Name)
|
||||
}
|
||||
pIdx, ok := extPkgIdx[pkg]
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("GOOBJ emission: package %q not resolved", pkg)
|
||||
}
|
||||
sIdx, ok := extSymIdx[pkg+"·"+name]
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("GOOBJ emission: symbol %s·%s not resolved", pkg, name)
|
||||
}
|
||||
binary.LittleEndian.PutUint32(rec[15:], uint32(pIdx))
|
||||
binary.LittleEndian.PutUint32(rec[19:], uint32(sIdx))
|
||||
default:
|
||||
if di, ok := defIdx[r.Name]; ok {
|
||||
binary.LittleEndian.PutUint32(rec[15:], pkgIdxSelf)
|
||||
binary.LittleEndian.PutUint32(rec[19:], uint32(di))
|
||||
break
|
||||
}
|
||||
// A DATA field may hold the address of a TEXT function of
|
||||
// the same file (the rt0 lib entry spelling), which is a
|
||||
// non-package definition.
|
||||
ni, isText := textNpIdx[r.Name]
|
||||
if !isText {
|
||||
return nil, fmt.Errorf("GOOBJ emission: reference to unknown symbol %q", r.Name)
|
||||
}
|
||||
binary.LittleEndian.PutUint32(rec[15:], pkgIdxNone)
|
||||
binary.LittleEndian.PutUint32(rec[19:], uint32(ni))
|
||||
}
|
||||
symRelocs[i] = append(symRelocs[i], rec[:]...)
|
||||
}
|
||||
}
|
||||
// The DWARF symbols' own relocations (the function address references).
|
||||
for _, ds := range dwarfRelocs {
|
||||
for _, r := range ds.relocs {
|
||||
|
||||
+4
-4
@@ -307,12 +307,12 @@ func TestStackGuardBytesLOONG64(t *testing.T) {
|
||||
func TestStackGuardGOObjInternalCall(t *testing.T) {
|
||||
for _, tt := range []struct {
|
||||
src string
|
||||
assemble func(*ast.File) (*Image, error)
|
||||
assemble func(*ast.File, ...AssembleOption) (*Image, error)
|
||||
}{
|
||||
{"g_amd64.s", AssembleFile},
|
||||
{"g_arm64.s", AssembleFileARM64},
|
||||
{"g_riscv64.s", AssembleFileRISCV},
|
||||
{"g_loong64.s", AssembleFileLOONG64},
|
||||
{"g_arm64.s", func(f *ast.File, _ ...AssembleOption) (*Image, error) { return AssembleFileARM64(f) }},
|
||||
{"g_riscv64.s", func(f *ast.File, _ ...AssembleOption) (*Image, error) { return AssembleFileRISCV(f) }},
|
||||
{"g_loong64.s", func(f *ast.File, _ ...AssembleOption) (*Image, error) { return AssembleFileLOONG64(f) }},
|
||||
} {
|
||||
f, errs := parser.Parse(tt.src, "TEXT \u00b7callsmall(SB), $16-0\n\tCALL \u00b7other(SB)\n\tRET\nTEXT \u00b7other(SB), NOSPLIT, $0\n\tRET\n")
|
||||
if len(errs) > 0 {
|
||||
|
||||
+150
-26
@@ -65,6 +65,15 @@ var bitTestOp = map[string]int{
|
||||
// noOperandTable maps a fixed no-operand mnemonic to its opcode bytes. The
|
||||
// fence names carry their opcode inside the 0F AE /digit group spelled out in
|
||||
// full (E8/F0/F8), and PAUSE is F3 90.
|
||||
//
|
||||
// LOCK, REP and REPN are the prefix statements. go tool asm encodes each as
|
||||
// a standalone one-byte instruction with a PC of its own (F0, F3 and F2
|
||||
// respectively), not as a prefix field merged into the next instruction: the
|
||||
// statement that follows is encoded unaware of it, and nothing validates
|
||||
// that the pairing is a legal one (LOCK before NOP assembles without
|
||||
// complaint, each byte pinned against the toolchain). Because the bytes
|
||||
// land in the stream before the following statement anyway, a LOCKed
|
||||
// CMPXCHGQ encodes identically to a prefixed form.
|
||||
var noOperandTable = map[string][]byte{
|
||||
"CPUID": {0x0F, 0xA2},
|
||||
"RDTSC": {0x0F, 0x31},
|
||||
@@ -78,6 +87,9 @@ var noOperandTable = map[string][]byte{
|
||||
"MFENCE": {0x0F, 0xAE, 0xF0},
|
||||
"SFENCE": {0x0F, 0xAE, 0xF8},
|
||||
"UNDEF": {0x0F, 0x0B},
|
||||
"LOCK": {0xF0},
|
||||
"REP": {0xF3},
|
||||
"REPN": {0xF2},
|
||||
}
|
||||
|
||||
// --- MOV --------------------------------------------------------------------
|
||||
@@ -180,6 +192,30 @@ func (e *enc) encodeMov(ops []Operand, size int) error {
|
||||
}
|
||||
return e.emit(i)
|
||||
|
||||
case TLSMem:
|
||||
if !dstIsReg {
|
||||
return fmt.Errorf("MOV: two memory operands")
|
||||
}
|
||||
// MOV r, off(TLS): the segment-prefixed absolute load, reg=dst,
|
||||
// rm=src(tlsMem) through the SIB escape; the disp32 is the TLS slot
|
||||
// offset with its R_TLSLE patch site.
|
||||
i := newInstr(size, []byte{movRR(size)})
|
||||
if err := setRM(i, dstReg, src, size); err != nil {
|
||||
return err
|
||||
}
|
||||
return e.emit(i)
|
||||
|
||||
case SegAbs:
|
||||
if !dstIsReg {
|
||||
return fmt.Errorf("MOV: two memory operands")
|
||||
}
|
||||
// MOV r, 0x30(GS): the segment-absolute load.
|
||||
i := newInstr(size, []byte{movRR(size)})
|
||||
if err := setRM(i, dstReg, src, size); err != nil {
|
||||
return err
|
||||
}
|
||||
return e.emit(i)
|
||||
|
||||
case Imm:
|
||||
if dstIsReg {
|
||||
v := int64(src)
|
||||
@@ -220,11 +256,24 @@ func (e *enc) encodeMov(ops []Operand, size int) error {
|
||||
i.imm = imm
|
||||
return e.emit(i)
|
||||
}
|
||||
// MOV r/m, imm: 0xC6 (8-bit) / 0xC7 /0.
|
||||
// MOV r/m, imm: 0xC6 (8-bit) / 0xC7 /0. An immediate in the
|
||||
// destination slot is the absolute-address crash-store spelling,
|
||||
// MOVL $0xf1, 0xf1: the parser reads the trailing bare constant
|
||||
// as an immediate, and the store's disp32 carries the address.
|
||||
op := byte(0xC7)
|
||||
if size == 1 {
|
||||
op = 0xC6
|
||||
}
|
||||
if d, ok := dst.(Imm); ok {
|
||||
i := newInstr(size, []byte{op})
|
||||
setSegAbs(i, 0, SegAbs{Disp: int64(d)})
|
||||
immBytes, err := immediate(int64(src), size, false)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
i.imm = immBytes
|
||||
return e.emit(i)
|
||||
}
|
||||
i := newInstr(size, []byte{op})
|
||||
if err := setRMDigit(i, 0, dst, size); err != nil {
|
||||
return err
|
||||
@@ -498,13 +547,34 @@ func (e *enc) encodeUnary(op struct {
|
||||
|
||||
// --- SHL/SHR/SAR ------------------------------------------------------------
|
||||
|
||||
func (e *enc) encodeShift(digit int, ops []Operand, size int) error {
|
||||
// doubleShiftOp maps the two mnemonics whose three-operand form go tool asm
|
||||
// accepts to the SHLD/SHRD opcode pair (imm8 form, CL form). SAR, SAL and
|
||||
// the rotates have no such form: the oracle rejects SARQ/ROLQ with three
|
||||
// operands, and so do we.
|
||||
var doubleShiftOp = map[string][2]byte{
|
||||
"SHL": {0xA4, 0xA5}, // SHLD
|
||||
"SHR": {0xAC, 0xAD}, // SHRD
|
||||
}
|
||||
|
||||
// isShiftCountCL reports whether a count operand is the CL register or its
|
||||
// CX spelling: go tool asm accepts both (CX names the same low byte) and
|
||||
// rejects ECX/RCX.
|
||||
func isShiftCountCL(o Operand) bool {
|
||||
reg, ok := o.(Reg)
|
||||
return ok && reg.idx == 1 && (reg.size == 1 || reg.size == 2)
|
||||
}
|
||||
|
||||
func (e *enc) encodeShift(base string, ops []Operand, size int) error {
|
||||
digit := shiftOp[base]
|
||||
if len(ops) == 3 {
|
||||
return e.encodeDoubleShift(base, ops, size)
|
||||
}
|
||||
if len(ops) != 2 {
|
||||
return fmt.Errorf("shift expects 2 operands, got %d", len(ops))
|
||||
}
|
||||
count, dst := ops[0], ops[1]
|
||||
// Count is $1, %CL, or an imm8.
|
||||
if reg, ok := count.(Reg); ok && reg.idx == 1 && reg.size <= 1 {
|
||||
// Count is $1, CL (or its CX spelling), or an imm8.
|
||||
if isShiftCountCL(count) {
|
||||
// CL: 0xD2 (8-bit) / 0xD3.
|
||||
op := byte(0xD3)
|
||||
if size == 1 {
|
||||
@@ -551,12 +621,59 @@ func (e *enc) encodeShift(digit int, ops []Operand, size int) error {
|
||||
return e.emit(i)
|
||||
}
|
||||
|
||||
// encodeDoubleShift emits the three-operand SHL/SHR form, which the Go
|
||||
// assembler spells as a shift but encodes as SHLD/SHRD (0F A4/A5, 0F AC/AD):
|
||||
// the first operand is the count ($imm or CL), the second feeds the vacated
|
||||
// bits (the reg field) and the third is the shifted value (the r/m field),
|
||||
// matching go tool asm byte for byte. The W/L/Q widths exist; the oracle
|
||||
// rejects the three-operand B form and every SAR/rotate one.
|
||||
func (e *enc) encodeDoubleShift(base string, ops []Operand, size int) error {
|
||||
opc, ok := doubleShiftOp[base]
|
||||
if !ok || size == 1 {
|
||||
return fmt.Errorf("%s: shift expects 2 operands, got %d", base, len(ops))
|
||||
}
|
||||
count, src, dst := ops[0], ops[1], ops[2]
|
||||
srcReg, ok := src.(Reg)
|
||||
if !ok {
|
||||
return fmt.Errorf("%s: middle operand must be a register, like go tool asm", base)
|
||||
}
|
||||
i := newInstr(size, []byte{0x0F, opc[0]})
|
||||
if isShiftCountCL(count) {
|
||||
// CL (or CX) form: 0F A5/AD.
|
||||
i.opcode[1] = opc[1]
|
||||
} else {
|
||||
imm, ok := count.(Imm)
|
||||
if !ok {
|
||||
return fmt.Errorf("shift count must be $1, CL or an immediate")
|
||||
}
|
||||
// The count is an unsigned imm8: the same range convention as the
|
||||
// two-operand shift above.
|
||||
if imm < 0 || imm > 255 {
|
||||
return fmt.Errorf("shift count $%d is out of the 0..255 range", int64(imm))
|
||||
}
|
||||
i.imm = []byte{byte(imm)}
|
||||
}
|
||||
if err := setRMReg(i, srcReg.idx, srcReg.idx >= 8, false, dst, size); err != nil {
|
||||
return err
|
||||
}
|
||||
return e.emit(i)
|
||||
}
|
||||
|
||||
// --- IMUL -------------------------------------------------------------------
|
||||
|
||||
func (e *enc) encodeImul(ops []Operand, size int) error {
|
||||
switch len(ops) {
|
||||
case 2:
|
||||
// IMUL r, r/m: 0x0F 0xAF.
|
||||
// Two shapes. The leading-immediate spelling IMUL $imm, r multiplies
|
||||
// r in place (dst = rm = r): the shape GOROOT's clock code writes.
|
||||
// Otherwise IMUL r, r/m: 0x0F 0xAF.
|
||||
if imm, ok := ops[0].(Imm); ok {
|
||||
dstReg, isReg := ops[1].(Reg)
|
||||
if !isReg {
|
||||
return fmt.Errorf("IMUL: destination must be a register")
|
||||
}
|
||||
return e.encodeImulImm(imm, dstReg, dstReg, size)
|
||||
}
|
||||
dstReg, ok := ops[1].(Reg)
|
||||
if !ok {
|
||||
return fmt.Errorf("IMUL: destination must be a register")
|
||||
@@ -576,29 +693,36 @@ func (e *enc) encodeImul(ops []Operand, size int) error {
|
||||
if !ok {
|
||||
return fmt.Errorf("IMUL: immediate operand expected first")
|
||||
}
|
||||
// Plan 9 order: IMUL $imm, src, dst.
|
||||
if fits8(int64(imm)) {
|
||||
i := newInstr(size, []byte{0x6B})
|
||||
if err := setRM(i, dstReg, ops[1], size); err != nil {
|
||||
return err
|
||||
}
|
||||
i.imm = []byte{byte(int8(imm))}
|
||||
return e.emit(i)
|
||||
}
|
||||
i := newInstr(size, []byte{0x69})
|
||||
if err := setRM(i, dstReg, ops[1], size); err != nil {
|
||||
return err
|
||||
}
|
||||
immBytes, err := immediate(int64(imm), size, false)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
i.imm = immBytes
|
||||
return e.emit(i)
|
||||
// Plan 9 order: IMUL $imm, src, dst; the source stays a general
|
||||
// r/m operand (setRM takes registers and memory alike).
|
||||
return e.encodeImulImm(imm, ops[1], dstReg, size)
|
||||
}
|
||||
return fmt.Errorf("IMUL expects 2 or 3 operands, got %d", len(ops))
|
||||
}
|
||||
|
||||
// encodeImulImm emits the immediate multiply: 0x6B with a sign-extended imm8
|
||||
// when the value fits, 0x69 with a 32-bit immediate otherwise.
|
||||
func (e *enc) encodeImulImm(imm Imm, rm Operand, dst Reg, size int) error {
|
||||
if fits8(int64(imm)) {
|
||||
i := newInstr(size, []byte{0x6B})
|
||||
if err := setRM(i, dst, rm, size); err != nil {
|
||||
return err
|
||||
}
|
||||
i.imm = []byte{byte(int8(imm))}
|
||||
return e.emit(i)
|
||||
}
|
||||
i := newInstr(size, []byte{0x69})
|
||||
if err := setRM(i, dst, rm, size); err != nil {
|
||||
return err
|
||||
}
|
||||
immBytes, err := immediate(int64(imm), size, false)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
i.imm = immBytes
|
||||
return e.emit(i)
|
||||
}
|
||||
|
||||
// --- PUSH / POP -------------------------------------------------------------
|
||||
|
||||
func (e *enc) encodePushPop(ops []Operand, size int, push bool) error {
|
||||
@@ -986,12 +1110,12 @@ func (e *enc) encodeSSEMove(m sseMove, ops []Operand) error {
|
||||
op = m.load
|
||||
reg, rm = dstReg, src
|
||||
case srcVec:
|
||||
if _, ok := dst.(Mem); !ok {
|
||||
if !isX86Mem(dst) {
|
||||
return fmt.Errorf("SSE move: invalid destination operand")
|
||||
}
|
||||
reg, rm = srcReg, dst
|
||||
case dstVec:
|
||||
if _, ok := src.(Mem); !ok {
|
||||
if !isX86Mem(src) {
|
||||
return fmt.Errorf("SSE move: invalid source operand")
|
||||
}
|
||||
op = m.load
|
||||
|
||||
@@ -0,0 +1,219 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
package asm
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"encoding/binary"
|
||||
"os"
|
||||
"os/exec"
|
||||
"path/filepath"
|
||||
"runtime"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
)
|
||||
|
||||
// The differential kernels for the DATA-path and front-end gaps are kept in
|
||||
// testdata/verify beside the campaign's other kernels; the verify package's
|
||||
// suites are not open to the asm package, so this test is their runner: each
|
||||
// kernel assembles through gasm and through go tool asm, and the functions'
|
||||
// bytes must agree with the relocation sites masked on both sides.
|
||||
|
||||
// toolAsmObject assembles path with the installed toolchain's assembler for
|
||||
// goarch ("" = the host) and returns the object bytes.
|
||||
func toolAsmObject(t *testing.T, path, goarch string) []byte {
|
||||
t.Helper()
|
||||
goBin, err := exec.LookPath("go")
|
||||
if err != nil {
|
||||
t.Skip("no Go toolchain available")
|
||||
}
|
||||
out, err := exec.Command(goBin, "env", "GOROOT").Output()
|
||||
if err != nil {
|
||||
t.Fatalf("go env GOROOT: %v", err)
|
||||
}
|
||||
includeDir := filepath.Join(strings.TrimSpace(string(out)), "pkg", "include")
|
||||
|
||||
pkg := strings.TrimSuffix(filepath.Base(path), ".s")
|
||||
pkg = strings.TrimSuffix(pkg, "_amd64")
|
||||
pkg = strings.TrimSuffix(pkg, "_arm64")
|
||||
|
||||
objPath := filepath.Join(t.TempDir(), "oracle.o")
|
||||
cmd := exec.Command(goBin, "tool", "asm", "-I", includeDir, "-p", pkg, "-o", objPath, path)
|
||||
if goarch != "" {
|
||||
environ := os.Environ()
|
||||
env := make([]string, 0, len(environ)+1)
|
||||
for _, e := range environ {
|
||||
if !strings.HasPrefix(e, "GOARCH=") {
|
||||
env = append(env, e)
|
||||
}
|
||||
}
|
||||
cmd.Env = append(env, "GOARCH="+goarch)
|
||||
}
|
||||
if out, err := cmd.CombinedOutput(); err != nil {
|
||||
t.Fatalf("go tool asm %s: %v\n%s", filepath.Base(path), err, out)
|
||||
}
|
||||
obj, err := os.ReadFile(objPath)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return obj
|
||||
}
|
||||
|
||||
// oracleFuncCode extracts the non-package TEXT functions' code bytes from a
|
||||
// toolchain object, keyed by the name the object records (pkg.name). Each
|
||||
// function's span is its own symbol size: a toolchain object that follows
|
||||
// the text with data symbols (the synthesised float-constant pool) would
|
||||
// otherwise fold them into the last function's bytes.
|
||||
func oracleFuncCode(t *testing.T, obj []byte) map[string][]byte {
|
||||
t.Helper()
|
||||
v := openGoobj(t, obj)
|
||||
le := binary.LittleEndian
|
||||
const symSize = 21
|
||||
nps := v.syms(blkNonpkgdef)
|
||||
data := v.blk(blkData)
|
||||
didx := v.blk(blkDataIdx)
|
||||
preceding := 0
|
||||
for _, bi := range []int{blkSymdef, blkHashed64def, blkHasheddef} {
|
||||
preceding += len(v.blk(bi)) / symSize
|
||||
}
|
||||
out := make(map[string][]byte, len(nps))
|
||||
for i, s := range nps {
|
||||
if s.typ != kindSTEXT {
|
||||
continue
|
||||
}
|
||||
start := le.Uint32(didx[4*(preceding+i):])
|
||||
out[s.name] = data[start : start+s.size]
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// maskCode zeroes every relocation field, the way the toolchain's object
|
||||
// leaves them for the linker.
|
||||
func maskCode(code []byte, relocs []Reloc) []byte {
|
||||
for _, r := range relocs {
|
||||
for j := r.Off; j < r.Off+4 && j < len(code); j++ {
|
||||
code[j] = 0
|
||||
}
|
||||
}
|
||||
return code
|
||||
}
|
||||
|
||||
// code assembles src for amd64 and returns the image's code bytes.
|
||||
func code(path, src string) []byte {
|
||||
f, errs := parser.Parse(path, src)
|
||||
if len(errs) > 0 {
|
||||
return nil
|
||||
}
|
||||
img, err := AssembleFile(f)
|
||||
if err != nil {
|
||||
return nil
|
||||
}
|
||||
return img.Code
|
||||
}
|
||||
|
||||
// TestDifferentialKernels pins the new kernels against the oracle.
|
||||
func TestDifferentialKernels(t *testing.T) {
|
||||
if runtime.GOARCH != "amd64" {
|
||||
t.Skip("the amd64 kernels assume an amd64 host assembler default")
|
||||
}
|
||||
for _, k := range []struct {
|
||||
path string
|
||||
goarch string
|
||||
arm64 bool
|
||||
}{
|
||||
{filepath.Join("..", "testdata", "verify", "datarel_amd64.s"), "", false},
|
||||
{filepath.Join("..", "testdata", "verify", "divslash_amd64.s"), "", false},
|
||||
{filepath.Join("..", "testdata", "verify", "semicolons_amd64.s"), "", false},
|
||||
{filepath.Join("..", "testdata", "verify", "quadreg_amd64.s"), "", false},
|
||||
{filepath.Join("..", "testdata", "verify", "floatimm_amd64.s"), "", false},
|
||||
{filepath.Join("..", "testdata", "verify", "bookkeep_amd64.s"), "", false},
|
||||
{filepath.Join("..", "testdata", "verify", "forms_amd64.s"), "", false},
|
||||
{filepath.Join("..", "testdata", "verify", "datarel_arm64.s"), "arm64", true},
|
||||
{filepath.Join("..", "testdata", "verify", "divslash_arm64.s"), "arm64", true},
|
||||
} {
|
||||
t.Run(filepath.Base(k.path), func(t *testing.T) {
|
||||
src, err := os.ReadFile(k.path)
|
||||
if err != nil {
|
||||
t.Fatalf("read: %v", err)
|
||||
}
|
||||
f, errs := parser.Parse(k.path, string(src))
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
var img *Image
|
||||
if k.arm64 {
|
||||
img, err = AssembleFileARM64(f)
|
||||
} else {
|
||||
img, err = AssembleFile(f)
|
||||
}
|
||||
if err != nil {
|
||||
t.Fatalf("assemble: %v", err)
|
||||
}
|
||||
gt := oracleFuncCode(t, toolAsmObject(t, k.path, k.goarch))
|
||||
// The oracle keys its functions by the qualified object name
|
||||
// (pkg.name); match on the local part.
|
||||
byLocal := make(map[string][]byte, len(gt))
|
||||
for name, code := range gt {
|
||||
if _, after, ok := strings.Cut(name, "."); ok {
|
||||
name = after
|
||||
}
|
||||
byLocal[name] = code
|
||||
}
|
||||
|
||||
matched := 0
|
||||
for _, fn := range img.Funcs {
|
||||
gasmCode := maskCode(append([]byte(nil), img.Code[fn.Offset:fn.Offset+fn.Size]...), fn.Relocs)
|
||||
goCode, ok := byLocal[fn.Name]
|
||||
if !ok {
|
||||
t.Errorf("%s: not in ground truth (%d functions: %v)", fn.Name, len(gt), keysOf(byLocal))
|
||||
continue
|
||||
}
|
||||
goCode = maskCode(append([]byte(nil), goCode...), fn.Relocs)
|
||||
cmpLen := min(len(goCode), len(gasmCode))
|
||||
if !bytes.Equal(gasmCode[:cmpLen], goCode[:cmpLen]) {
|
||||
t.Errorf("%s: MISMATCH gasm=%d go=%d bytes\ngasm %x\ngo %x", fn.Name, len(gasmCode), len(goCode), gasmCode, goCode)
|
||||
continue
|
||||
}
|
||||
for _, b := range goCode[len(gasmCode):] {
|
||||
if b != 0 {
|
||||
t.Errorf("%s: non-zero trailing bytes in go tool asm output", fn.Name)
|
||||
break
|
||||
}
|
||||
}
|
||||
matched++
|
||||
t.Logf("%s: MATCH (%d bytes)", fn.Name, len(gasmCode))
|
||||
}
|
||||
if matched == 0 {
|
||||
t.Fatal("no functions matched")
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func keysOf(m map[string][]byte) []string {
|
||||
out := make([]string, 0, len(m))
|
||||
for k := range m {
|
||||
out = append(out, k)
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// TestSemicolonSpellingParity pins that the ';' statement separator changes
|
||||
// nothing about the encoding: the one-line spelling assembles to exactly the
|
||||
// bytes of the same statements written one per line.
|
||||
func TestSemicolonSpellingParity(t *testing.T) {
|
||||
for _, tt := range []struct{ one, two string }{
|
||||
{"\tROLQ $3, DI; ROLQ $13, DI\n", "\tROLQ $3, DI\n\tROLQ $13, DI\n"},
|
||||
{"\tREP; MOVSQ\n", "\tREP\n\tMOVSQ\n"},
|
||||
{"\tXORQ AX, AX; XORQ CX, CX\n", "\tXORQ AX, AX\n\tXORQ CX, CX\n"},
|
||||
} {
|
||||
one := code("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n"+tt.one+"\tRET\n")
|
||||
two := code("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n"+tt.two+"\tRET\n")
|
||||
if !bytes.Equal(one, two) {
|
||||
t.Errorf("semicolon spelling %q: %x, want the two-line bytes %x", tt.one, one, two)
|
||||
}
|
||||
}
|
||||
}
|
||||
+176
-11
@@ -5,6 +5,7 @@ package asm
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"math"
|
||||
"sort"
|
||||
"strconv"
|
||||
|
||||
@@ -103,6 +104,7 @@ const (
|
||||
RelArm64Branch // R_CALLARM64 (BL instruction)
|
||||
RelArm64LDST64 // R_ARM64_PCREL_LDST64 (ADRP + 64-bit LDR/STR pair)
|
||||
RelLoong64Branch // R_CALLLOONG64 (BL instruction)
|
||||
RelAddr // R_ADDR: the absolute address of a symbol held in a DATA field
|
||||
)
|
||||
|
||||
type Reloc struct {
|
||||
@@ -112,13 +114,17 @@ type Reloc struct {
|
||||
// Addend select the target: the symbol plus the byte offset. An
|
||||
// External relocation names a symbol no GLOBL in the file defines;
|
||||
// the object-file emitters carry it into the output's relocation
|
||||
// table.
|
||||
// table. Siz is the width of the patched field and is set only for
|
||||
// data-field relocations (RelAddr, Off relative to the data symbol),
|
||||
// whose width is the DATA line's; code relocations take their width
|
||||
// from the architecture's instruction encoding.
|
||||
Off int
|
||||
After int
|
||||
Name string
|
||||
Addend int64
|
||||
External bool
|
||||
Kind RelocKind
|
||||
Siz uint8
|
||||
}
|
||||
|
||||
// DataSymbol describes one GLOBL symbol laid out in the data section.
|
||||
@@ -130,6 +136,11 @@ type DataSymbol struct {
|
||||
Static bool // the <> marker: file-local, not exported
|
||||
Rodata bool // the RODATA flag: read-only data
|
||||
Dupok bool // the DUPOK flag: duplicate-OK
|
||||
// Relocs carries the symbol-valued DATA initialisers ("DATA s+0(SB)/8,
|
||||
// $other(SB)"): fields of this symbol's data that hold another symbol's
|
||||
// address, resolved by the linker. Off is relative to the symbol's
|
||||
// data start.
|
||||
Relocs []Reloc
|
||||
}
|
||||
|
||||
// Bytes returns the whole image: code, then data.
|
||||
@@ -139,6 +150,18 @@ func (img *Image) Bytes() []byte {
|
||||
return append(out, img.Data...)
|
||||
}
|
||||
|
||||
// AssembleOption adjusts the file-level assembly context.
|
||||
type AssembleOption func(*linkInfo)
|
||||
|
||||
// WithGOOS selects the target operating system for the forms that depend on
|
||||
// it, the TLS access shape above all: linux and freebsd take the
|
||||
// one-instruction form, windows and plan9 keep the two-instruction load.
|
||||
func WithGOOS(goos string) AssembleOption {
|
||||
return func(l *linkInfo) {
|
||||
l.goos = goos
|
||||
}
|
||||
}
|
||||
|
||||
// AssembleFile assembles every TEXT function of a parsed file and lays out
|
||||
// its static symbols (GLOBL/DATA) in a data section behind the code. Each
|
||||
// reference to a file-local static symbol becomes a RIP-relative load whose
|
||||
@@ -146,7 +169,7 @@ func (img *Image) Bytes() []byte {
|
||||
// GLOBL defines is recorded as an external relocation (Externals) with its
|
||||
// displacement left zero, the object-file emitters resolve it at link
|
||||
// time, while the raw image (Bytes) cannot represent it.
|
||||
func AssembleFile(f *ast.File) (*Image, error) {
|
||||
func AssembleFile(f *ast.File, opts ...AssembleOption) (*Image, error) {
|
||||
dataSyms, err := collectData(f)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
@@ -155,7 +178,18 @@ func AssembleFile(f *ast.File) (*Image, error) {
|
||||
for _, d := range dataSyms {
|
||||
known[d.name] = true
|
||||
}
|
||||
// TEXT symbols are file-level definitions too: a symbol immediate
|
||||
// ($fn(SB)) may name one, exactly as a data reference names a GLOBL.
|
||||
for _, d := range f.Decls {
|
||||
if t, ok := d.(*ast.Text); ok {
|
||||
known[t.Name.Name] = true
|
||||
}
|
||||
}
|
||||
link := &linkInfo{symbols: known, allowExternal: true}
|
||||
for _, o := range opts {
|
||||
o(link)
|
||||
}
|
||||
poolSeen := map[string]bool{}
|
||||
|
||||
img := &Image{Symbols: map[string]int{}, SourcePath: f.Path}
|
||||
textOff := map[string]int{}
|
||||
@@ -169,7 +203,26 @@ func AssembleFile(f *ast.File) (*Image, error) {
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
code, patches, labels, steps, lines, err := assemble(t, link)
|
||||
code, patches, labels, steps, lines, pool, err := assemble(t, link)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("%s: %w", t.Name.Name, err)
|
||||
}
|
||||
// The pooled floating-point constants join the declared data as
|
||||
// read-only symbols, deduplicated across the file (the toolchain
|
||||
// synthesises the same symbols into its rodata).
|
||||
for _, entry := range pool {
|
||||
if poolSeen[entry.name] {
|
||||
continue
|
||||
}
|
||||
poolSeen[entry.name] = true
|
||||
dataSyms = append(dataSyms, dataSym{
|
||||
name: entry.name,
|
||||
buf: entry.data,
|
||||
size: len(entry.data),
|
||||
rodata: true,
|
||||
dupok: true,
|
||||
})
|
||||
}
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("%s: %w", t.Name.Name, err)
|
||||
}
|
||||
@@ -257,6 +310,22 @@ func AssembleFile(f *ast.File) (*Image, error) {
|
||||
img.Funcs[i].Relocs = append(img.Funcs[i].Relocs, reloc)
|
||||
}
|
||||
}
|
||||
// The data symbols' symbol-valued DATA fields resolve the same way the
|
||||
// code references do: a name the file defines (GLOBL or TEXT) stays an
|
||||
// internal reference the emitters resolve, anything else is external.
|
||||
// img.DataSyms was laid out in dataSyms order, so the indexes line up.
|
||||
for i := range img.DataSyms {
|
||||
for _, r := range dataSyms[i].relocs {
|
||||
reloc := r
|
||||
if _, ok := img.Symbols[reloc.Name]; !ok {
|
||||
if _, ok := textOff[reloc.Name]; !ok {
|
||||
reloc.External = true
|
||||
externals[reloc.Name] = true
|
||||
}
|
||||
}
|
||||
img.DataSyms[i].Relocs = append(img.DataSyms[i].Relocs, reloc)
|
||||
}
|
||||
}
|
||||
for name := range externals {
|
||||
img.Externals = append(img.Externals, name)
|
||||
}
|
||||
@@ -273,6 +342,10 @@ func AssembleFileRISCV(f *ast.File) (*Image, error) {
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
// The pooled $i64 constants the wide MOV immediate loads refer to join
|
||||
// the declared data as read-only symbols, deduplicated across the file
|
||||
// (the toolchain synthesises the same symbols into its rodata).
|
||||
litSeen := map[string]bool{}
|
||||
|
||||
img := &Image{Symbols: map[string]int{}, SourcePath: f.Path}
|
||||
for _, d := range f.Decls {
|
||||
@@ -280,10 +353,23 @@ func AssembleFileRISCV(f *ast.File) (*Image, error) {
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
code, labels, relocs, lines, spadj, err := assembleRISCV(t)
|
||||
code, labels, relocs, lines, spadj, lits, err := assembleRISCV(t)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("%s: %w", t.Name.Name, err)
|
||||
}
|
||||
for _, lit := range lits {
|
||||
if litSeen[lit.Name] {
|
||||
continue
|
||||
}
|
||||
litSeen[lit.Name] = true
|
||||
dataSyms = append(dataSyms, dataSym{
|
||||
name: lit.Name,
|
||||
buf: lit.Data,
|
||||
size: len(lit.Data),
|
||||
rodata: true,
|
||||
dupok: true,
|
||||
})
|
||||
}
|
||||
fl := FuncLayout{
|
||||
Name: t.Name.Name,
|
||||
Pkg: t.Name.Pkg,
|
||||
@@ -429,6 +515,23 @@ func markExternals(img *Image, dataSyms []dataSym) {
|
||||
}
|
||||
}
|
||||
}
|
||||
// The declared data symbols carry the file's own relocations (the
|
||||
// symbol-valued DATA fields); the layouts appended img.DataSyms in
|
||||
// dataSyms order, so the indexes line up. The trailing entries (the
|
||||
// pooled arm64 literals) have no source relocations.
|
||||
for i := range img.DataSyms {
|
||||
if i >= len(dataSyms) {
|
||||
break
|
||||
}
|
||||
for _, r := range dataSyms[i].relocs {
|
||||
reloc := r
|
||||
if !known[reloc.Name] {
|
||||
reloc.External = true
|
||||
externals[reloc.Name] = true
|
||||
}
|
||||
img.DataSyms[i].Relocs = append(img.DataSyms[i].Relocs, reloc)
|
||||
}
|
||||
}
|
||||
for name := range externals {
|
||||
img.Externals = append(img.Externals, name)
|
||||
}
|
||||
@@ -444,6 +547,9 @@ type dataSym struct {
|
||||
static bool
|
||||
rodata bool
|
||||
dupok bool
|
||||
// relocs are the symbol-valued DATA fields, in declaration order; Off
|
||||
// is relative to the symbol's data start.
|
||||
relocs []Reloc
|
||||
}
|
||||
|
||||
// collectData gathers the file's static symbols (GLOBL) and their initial
|
||||
@@ -511,20 +617,79 @@ func collectData(f *ast.File) ([]dataSym, error) {
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("DATA %q: no matching GLOBL", dd.Name.Name)
|
||||
}
|
||||
if dd.Value == nil || !dd.Value.Imm.HasVal {
|
||||
return nil, fmt.Errorf("DATA %q: value must be an integer immediate", dd.Name.Name)
|
||||
if dd.Value == nil {
|
||||
return nil, fmt.Errorf("DATA %q: missing value", dd.Name.Name)
|
||||
}
|
||||
w := dd.Width
|
||||
switch w {
|
||||
case 1, 2, 4, 8:
|
||||
default:
|
||||
return nil, fmt.Errorf("DATA %q: invalid width %d (want 1, 2, 4 or 8)", dd.Name.Name, w)
|
||||
}
|
||||
off := dd.Name.Offset
|
||||
buf := syms[i].buf
|
||||
if off < 0 || off+int64(w) > int64(len(buf)) {
|
||||
return nil, fmt.Errorf("DATA %q+%d/%d exceeds GLOBL size %d", dd.Name.Name, off, w, len(buf))
|
||||
}
|
||||
// A symbol value ("DATA s+0(SB)/8, $other(SB)", the rt0 spelling)
|
||||
// leaves the field zero and records a relocation against the named
|
||||
// symbol: the linker patches the absolute address at this data
|
||||
// offset. The toolchain emits the same shape, an R_ADDR of the
|
||||
// DATA width with the value's offset as the addend, on every
|
||||
// architecture.
|
||||
if sym := dd.Value.Imm.Sym; !dd.Value.Imm.HasVal && sym != nil {
|
||||
syms[i].relocs = append(syms[i].relocs, Reloc{
|
||||
Off: int(off),
|
||||
Name: sym.Name,
|
||||
Addend: sym.Offset,
|
||||
Kind: RelAddr,
|
||||
Siz: uint8(w),
|
||||
})
|
||||
continue
|
||||
}
|
||||
// A string or rune value ("DATA s+0(SB)/20, $"text"") writes its
|
||||
// bytes into the field and leaves the rest zero, the toolchain's
|
||||
// WriteString: the declared width must hold every byte, and any
|
||||
// width is legal.
|
||||
if s := dd.Value.Imm.Str; s != "" && !dd.Value.Imm.HasVal {
|
||||
text, err := strconv.Unquote(s)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("DATA %q: invalid string value %s", dd.Name.Name, s)
|
||||
}
|
||||
if len(text) > w {
|
||||
return nil, fmt.Errorf("DATA %q: string of %d bytes does not fit width %d", dd.Name.Name, len(text), w)
|
||||
}
|
||||
copy(buf[off:], text)
|
||||
continue
|
||||
}
|
||||
// A floating-point value stores its IEEE-754 bits: /4 the float32
|
||||
// rounding of the parsed double, /8 the full 64 bits, the
|
||||
// toolchain's WriteFloat32 and WriteFloat64.
|
||||
if f := dd.Value.Imm.Float; f != "" && !dd.Value.Imm.HasVal {
|
||||
num, err := strconv.ParseFloat(f, 64)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("DATA %q: invalid floating-point value %q", dd.Name.Name, f)
|
||||
}
|
||||
if dd.Value.Imm.Neg {
|
||||
num = -num
|
||||
}
|
||||
var v uint64
|
||||
switch w {
|
||||
case 4:
|
||||
v = uint64(math.Float32bits(float32(num)))
|
||||
case 8:
|
||||
v = math.Float64bits(num)
|
||||
default:
|
||||
return nil, fmt.Errorf("DATA %q: invalid width %d for a float (want 4 or 8)", dd.Name.Name, w)
|
||||
}
|
||||
for j := range w {
|
||||
buf[off+int64(j)] = byte(v >> (8 * j))
|
||||
}
|
||||
continue
|
||||
}
|
||||
if !dd.Value.Imm.HasVal {
|
||||
return nil, fmt.Errorf("DATA %q: value must be an integer immediate or a symbol address", dd.Name.Name)
|
||||
}
|
||||
switch w {
|
||||
case 1, 2, 4, 8:
|
||||
default:
|
||||
return nil, fmt.Errorf("DATA %q: invalid width %d (want 1, 2, 4 or 8)", dd.Name.Name, w)
|
||||
}
|
||||
v := dd.Value.Imm.Val
|
||||
if dd.Value.Imm.Neg {
|
||||
v = -v
|
||||
|
||||
@@ -4,6 +4,11 @@
|
||||
package asm
|
||||
|
||||
import (
|
||||
"encoding/binary"
|
||||
"fmt"
|
||||
"os"
|
||||
"os/exec"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
@@ -166,3 +171,323 @@ func TestCollectDataNumericFlags(t *testing.T) {
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestCollectDataSymbolValue covers the symbol-valued DATA field ("DATA
|
||||
// s+0(SB)/8, $other(SB)", the rt0 spelling): the field stays zero in the
|
||||
// image and the relocation is recorded against the named symbol, whatever
|
||||
// the file defines (a TEXT function, a GLOBL) or leaves external.
|
||||
func TestCollectDataSymbolValue(t *testing.T) {
|
||||
src := `#include "textflag.h"
|
||||
TEXT ·Keep(SB), NOSPLIT, $0-8
|
||||
MOVQ target+0(FP), AX
|
||||
RET
|
||||
GLOBL holder(SB), NOPTR, $32
|
||||
DATA holder+0(SB)/8, $·Keep(SB)
|
||||
DATA holder+8(SB)/8, $·Keep+5(SB)
|
||||
DATA holder+16(SB)/8, $holder(SB)
|
||||
GLOBL spare(SB), NOPTR, $8
|
||||
DATA spare+0(SB)/8, $extvar(SB)
|
||||
`
|
||||
f, errs := parser.Parse("f_amd64.s", src)
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
img, err := AssembleFile(f)
|
||||
if err != nil {
|
||||
t.Fatalf("assemble: %v", err)
|
||||
}
|
||||
byName := map[string]DataSymbol{}
|
||||
for _, d := range img.DataSyms {
|
||||
byName[d.Name] = d
|
||||
}
|
||||
want := []struct {
|
||||
sym string
|
||||
off int
|
||||
name string
|
||||
addend int64
|
||||
ext bool
|
||||
}{
|
||||
{"holder", 0, "Keep", 0, false},
|
||||
{"holder", 8, "Keep", 5, false},
|
||||
{"holder", 16, "holder", 0, false},
|
||||
{"spare", 0, "extvar", 0, true},
|
||||
}
|
||||
var flat []struct {
|
||||
sym string
|
||||
r Reloc
|
||||
}
|
||||
for _, d := range img.DataSyms {
|
||||
for _, r := range d.Relocs {
|
||||
flat = append(flat, struct {
|
||||
sym string
|
||||
r Reloc
|
||||
}{d.Name, r})
|
||||
}
|
||||
}
|
||||
if len(flat) != len(want) {
|
||||
t.Fatalf("data relocations = %d, want %d", len(flat), len(want))
|
||||
}
|
||||
for i, w := range want {
|
||||
g := flat[i]
|
||||
r := g.r
|
||||
if g.sym != w.sym {
|
||||
t.Errorf("relocation %d sits on %q, want %q", i, g.sym, w.sym)
|
||||
continue
|
||||
}
|
||||
if r.Off != w.off || r.Name != w.name || r.Addend != w.addend || r.External != w.ext {
|
||||
t.Errorf("relocation %d = {+%d %q addend %d ext %v}, want {+%d %q addend %d ext %v}",
|
||||
i, r.Off, r.Name, r.Addend, r.External, w.off, w.name, w.addend, w.ext)
|
||||
}
|
||||
if r.Kind != RelAddr {
|
||||
t.Errorf("relocation %d kind = %v, want RelAddr", i, r.Kind)
|
||||
}
|
||||
if r.Siz != 8 {
|
||||
t.Errorf("relocation %d siz = %d, want 8", i, r.Siz)
|
||||
}
|
||||
}
|
||||
// The fields themselves stay zero: only the linker fills them.
|
||||
for _, b := range img.Data {
|
||||
if b != 0 {
|
||||
t.Fatal("data section is not all zero before relocation")
|
||||
}
|
||||
}
|
||||
if len(img.Externals) != 1 || img.Externals[0] != "extvar" {
|
||||
t.Errorf("Externals = %v, want [extvar]", img.Externals)
|
||||
}
|
||||
}
|
||||
|
||||
// TestGOObjectDataSymbolReloc pins the GOOBJ record a symbol-valued DATA
|
||||
// field produces, against the shape the toolchain emits for the same
|
||||
// source: an R_ADDR of the DATA width at the field offset, pkgIdxNone plus
|
||||
// the non-package definition index when the target is the file's own TEXT
|
||||
// function (the rt0 lib entry spelling).
|
||||
func TestGOObjectDataSymbolReloc(t *testing.T) {
|
||||
f, errs := parser.Parse("f_amd64.s", `#include "textflag.h"
|
||||
TEXT ·Keep(SB), NOSPLIT, $0-8
|
||||
RET
|
||||
GLOBL holder(SB), NOPTR, $16
|
||||
DATA holder+0(SB)/8, $·Keep+5(SB)
|
||||
`)
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
img, err := AssembleFile(f)
|
||||
if err != nil {
|
||||
t.Fatalf("assemble: %v", err)
|
||||
}
|
||||
obj, err := img.GOObject("main", "f_amd64.s")
|
||||
if err != nil {
|
||||
t.Fatalf("GOObject: %v", err)
|
||||
}
|
||||
v := openGoobj(t, obj)
|
||||
// Walk every relocation record; the data record is the one of Siz 8
|
||||
// and type R_ADDR.
|
||||
var off, add int64
|
||||
var pkg, sym uint32
|
||||
found := false
|
||||
for data := v.blk(blkReloc); len(data) >= 23; data = data[23:] {
|
||||
if data[4] != 8 || binary.LittleEndian.Uint16(data[5:]) != relocAddr {
|
||||
continue
|
||||
}
|
||||
found = true
|
||||
off = int64(int32(binary.LittleEndian.Uint32(data[0:])))
|
||||
add = int64(binary.LittleEndian.Uint64(data[7:]))
|
||||
pkg = binary.LittleEndian.Uint32(data[15:])
|
||||
sym = binary.LittleEndian.Uint32(data[19:])
|
||||
break
|
||||
}
|
||||
if !found {
|
||||
t.Fatal("no data relocation record in the object")
|
||||
}
|
||||
if off != 0 || add != 5 {
|
||||
t.Errorf("data reloc = {off %d addend %d}, want {off 0 addend 5}", off, add)
|
||||
}
|
||||
if pkg != pkgIdxNone {
|
||||
t.Errorf("data reloc pkg = %#x, want pkgIdxNone (the TEXT function)", pkg)
|
||||
}
|
||||
// The function's non-package definition index: the four pc tables
|
||||
// precede it, so index 4.
|
||||
if sym != 4 {
|
||||
t.Errorf("data reloc sym = %d, want 4", sym)
|
||||
}
|
||||
}
|
||||
|
||||
// TestGOObjectDataSymbolLink is the end-to-end proof for symbol-valued DATA
|
||||
// fields: the gasm object is substituted for the toolchain's and re-linked,
|
||||
// then executed, and the linked data word must hold the real address of the
|
||||
// function the DATA line named (runtime.FuncForPC identifies it).
|
||||
func TestGOObjectDataSymbolLink(t *testing.T) {
|
||||
goBin, err := exec.LookPath("go")
|
||||
if err != nil {
|
||||
t.Skip("no Go toolchain available")
|
||||
}
|
||||
dir := t.TempDir()
|
||||
asmSrc := `#include "textflag.h"
|
||||
GLOBL entry(SB), NOPTR, $8
|
||||
DATA entry+0(SB)/8, $·keepme(SB)
|
||||
|
||||
TEXT ·keepme(SB), NOSPLIT, $0-0
|
||||
RET
|
||||
|
||||
TEXT ·entryptr(SB), NOSPLIT, $0-8
|
||||
MOVQ entry+0(SB), AX
|
||||
MOVQ AX, ret+0(FP)
|
||||
RET
|
||||
`
|
||||
if err := os.WriteFile(filepath.Join(dir, "main_amd64.s"), []byte(asmSrc), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
mainSrc := `package main
|
||||
|
||||
import "runtime"
|
||||
|
||||
func keepme()
|
||||
func entryptr() uintptr
|
||||
|
||||
func main() {
|
||||
pc := entryptr()
|
||||
fn := runtime.FuncForPC(pc)
|
||||
if fn == nil {
|
||||
panic("the entry word does not point at a function")
|
||||
}
|
||||
if fn.Name() != "main.keepme" {
|
||||
panic("the entry word points at " + fn.Name())
|
||||
}
|
||||
}
|
||||
`
|
||||
if err := os.WriteFile(filepath.Join(dir, "main.go"), []byte(mainSrc), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := os.WriteFile(filepath.Join(dir, "go.mod"), []byte("module dlink\n\ngo 1.21\n"), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
// Capture the build: the package archive's asm object and the link line.
|
||||
build := exec.Command(goBin, "build", "-x", "-work", "-o", filepath.Join(dir, "prog"), ".")
|
||||
build.Dir = dir
|
||||
buildLog, err := build.CombinedOutput()
|
||||
if err != nil {
|
||||
t.Fatalf("baseline build: %v\n%s", err, buildLog)
|
||||
}
|
||||
var work, linkLine, asmObj string
|
||||
for line := range strings.SplitSeq(string(buildLog), "\n") {
|
||||
switch {
|
||||
case strings.HasPrefix(line, "WORK="):
|
||||
work = strings.TrimPrefix(line, "WORK=")
|
||||
case strings.Contains(line, "/asm ") && strings.Contains(line, "main_amd64.s") && !strings.Contains(line, "-gensymabis"):
|
||||
asmObj = fieldAfter(line, "-o")
|
||||
case strings.Contains(line, "/link ") && strings.Contains(line, "-importcfg"):
|
||||
linkLine = line
|
||||
}
|
||||
}
|
||||
if work == "" || asmObj == "" || linkLine == "" {
|
||||
t.Skipf("could not parse build log (work=%q asmObj=%q link=%q)", work, asmObj, linkLine)
|
||||
}
|
||||
defer os.RemoveAll(work)
|
||||
asmObj = strings.ReplaceAll(asmObj, "$WORK", work)
|
||||
linkLine = strings.ReplaceAll(linkLine, "$WORK", work)
|
||||
|
||||
// Assemble the same source with gasm and substitute the object.
|
||||
src, err := os.ReadFile(filepath.Join(dir, "main_amd64.s"))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
f, errs := parser.Parse("main_amd64.s", string(src))
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
img, err := AssembleFile(f)
|
||||
if err != nil {
|
||||
t.Fatalf("AssembleFile: %v", err)
|
||||
}
|
||||
gasmObj, err := img.GOObject("dlink", "main_amd64.s")
|
||||
if err != nil {
|
||||
t.Fatalf("GOObject: %v", err)
|
||||
}
|
||||
if err := os.WriteFile(asmObj, gasmObj, 0o644); err != nil {
|
||||
t.Fatalf("write gasm object: %v", err)
|
||||
}
|
||||
linkCmd := exec.Command("bash", "-c", "cd "+dir+" && "+linkLine)
|
||||
if out, err := linkCmd.CombinedOutput(); err != nil {
|
||||
t.Fatalf("re-link with gasm object: %v\n%s", err, out)
|
||||
}
|
||||
|
||||
// The linked program must run and find the right function behind the
|
||||
// data word.
|
||||
out, err := exec.Command(filepath.Join(dir, "prog")).CombinedOutput()
|
||||
if err != nil {
|
||||
t.Fatalf("linked program failed: %v\n%s", err, out)
|
||||
}
|
||||
}
|
||||
|
||||
// TestCollectDataFloatAndStringValues covers the non-integer DATA values the
|
||||
// runtime's math and asm files use: floating-point initialisers store their
|
||||
// IEEE-754 bits (/4 the float32 rounding, /8 the full double) and string
|
||||
// initialisers write their bytes zero-padded within the declared width.
|
||||
func TestCollectDataFloatAndStringValues(t *testing.T) {
|
||||
src := `#include "textflag.h"
|
||||
TEXT ·Keep(SB), NOSPLIT, $0-8
|
||||
RET
|
||||
GLOBL vals<>(SB), RODATA, $44
|
||||
DATA vals<>+0(SB)/8, $0.5
|
||||
DATA vals<>+8(SB)/8, $-1.0
|
||||
DATA vals<>+16(SB)/4, $1.5
|
||||
DATA vals<>+20(SB)/16, $"call frame too "
|
||||
DATA vals<>+36(SB)/4, $"hi"
|
||||
`
|
||||
f, errs := parser.Parse("fvals_amd64.s", src)
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
img, err := AssembleFile(f)
|
||||
if err != nil {
|
||||
t.Fatalf("AssembleFile: %v", err)
|
||||
}
|
||||
byName := map[string]DataSymbol{}
|
||||
for _, d := range img.DataSyms {
|
||||
byName[d.Name] = d
|
||||
}
|
||||
d := byName["vals"]
|
||||
if d.Size != 44 {
|
||||
t.Fatalf("vals size = %d, want 44", d.Size)
|
||||
}
|
||||
buf := img.Data[d.Offset : d.Offset+44]
|
||||
// 0.5 = 0x3FE0000000000000, -1.0 = 0xBFF0000000000000 (float64);
|
||||
// 1.5 = 0x3FC00000 (float32).
|
||||
for _, c := range []struct {
|
||||
off int
|
||||
want []byte
|
||||
}{
|
||||
{0, []byte{0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0xE0, 0x3F}},
|
||||
{8, []byte{0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0xF0, 0xBF}},
|
||||
{16, []byte{0x00, 0x00, 0xC0, 0x3F}},
|
||||
{20, []byte("call frame too ")},
|
||||
{36, []byte{'h', 'i', 0x00, 0x00}},
|
||||
} {
|
||||
if string(buf[c.off:c.off+len(c.want)]) != string(c.want) {
|
||||
t.Errorf("vals+%d: got % x, want % x", c.off, buf[c.off:c.off+len(c.want)], c.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestCollectDataValueErrors pins the value-kind width rules: a float needs
|
||||
// width 4 or 8, a string must fit its declared width, and a bad float
|
||||
// literal is diagnosed rather than stored.
|
||||
func TestCollectDataValueErrors(t *testing.T) {
|
||||
cases := []string{
|
||||
`GLOBL v<>(SB), RODATA, $4
|
||||
DATA v<>+0(SB)/1, $0.5`,
|
||||
`GLOBL v<>(SB), RODATA, $2
|
||||
DATA v<>+0(SB)/2, $"toolarge"`,
|
||||
}
|
||||
for i, src := range cases {
|
||||
full := "#include \"textflag.h\"\nTEXT ·Keep(SB), NOSPLIT, $0-8\n\tRET\n" + src
|
||||
f, errs := parser.Parse(fmt.Sprintf("verr%d_amd64.s", i), full)
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("case %d parse: %v", i, errs)
|
||||
}
|
||||
if _, err := AssembleFile(f); err == nil {
|
||||
t.Errorf("case %d: expected an error, got none", i)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
+450
-45
@@ -46,30 +46,70 @@ func assembleLOONG64(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry,
|
||||
spadj = append(spadj, SpadjStep{PC: guardLen + (loong64StoreWords(fi.autosize)+loong64AdjustWords(-int64(fi.autosize)))*4, Value: fi.autosize})
|
||||
}
|
||||
|
||||
// Pass 1: label offsets from the instruction sizes.
|
||||
offsets := map[string]int{}
|
||||
pos := guardLen + len(prologue)
|
||||
// The toolchain's parser counts N(PC) displacements over the source
|
||||
// instructions at a uniform 4 bytes each, so a PC-relative branch
|
||||
// resolves to the instruction N slots away in body order; the resolved
|
||||
// target then participates in layout and loop-head padding like any
|
||||
// branch target.
|
||||
instrs := make([]*ast.Instr, 0, len(t.Body))
|
||||
for _, stmt := range t.Body {
|
||||
switch s := stmt.(type) {
|
||||
case *ast.Label:
|
||||
offsets[s.Name.Text] = pos
|
||||
case *ast.Instr:
|
||||
pos += loong64InstrSize(s, fi)
|
||||
if in, ok := stmt.(*ast.Instr); ok && strings.ToUpper(in.Mnemonic.Text) != "PCALIGN" {
|
||||
instrs = append(instrs, in)
|
||||
}
|
||||
}
|
||||
|
||||
// Pass 2: encode. The guard prefix precedes the prologue; its branches
|
||||
// target the morestack block at the end of the function, which the first
|
||||
// pass has sized.
|
||||
bodyLen := 0
|
||||
{
|
||||
p := guardLen + len(prologue)
|
||||
for _, stmt := range t.Body {
|
||||
if in, ok := stmt.(*ast.Instr); ok {
|
||||
p += loong64InstrSize(in, fi)
|
||||
}
|
||||
parseIndex := make(map[*ast.Instr]int, len(instrs))
|
||||
for i, in := range instrs {
|
||||
parseIndex[in] = i
|
||||
}
|
||||
pcRelTarget := make(map[*ast.Instr]*ast.Instr)
|
||||
for _, in := range instrs {
|
||||
off, ok := loong64PCRelOffset(in)
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
bodyLen = p - (guardLen + len(prologue))
|
||||
tgt := parseIndex[in] + off
|
||||
if tgt < 0 || tgt >= len(instrs) {
|
||||
continue
|
||||
}
|
||||
pcRelTarget[in] = instrs[tgt]
|
||||
}
|
||||
|
||||
// Pass 1: label offsets from the instruction sizes. PCALIGN contributes
|
||||
// only its padding. On top of the explicit PCALIGNs, the toolchain pads
|
||||
// every backward-branch target (loop head) to a 16-byte boundary, so the
|
||||
// layout runs to a fixpoint over the alignment set.
|
||||
loopAligns := map[string]bool{}
|
||||
alignInstrs := map[*ast.Instr]bool{}
|
||||
for {
|
||||
offsets, _, pcs, _ := loong64Layout(t, guardLen+len(prologue), fi, loopAligns, alignInstrs)
|
||||
changed := false
|
||||
for _, in := range instrs {
|
||||
// A backward PC-relative target is the resolved instruction.
|
||||
if tgt, ok := pcRelTarget[in]; ok && pcs[tgt] < pcs[in] && !alignInstrs[tgt] {
|
||||
alignInstrs[tgt] = true
|
||||
changed = true
|
||||
}
|
||||
target, ok := loong64BranchTarget(in)
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
tOff, ok := offsets[target]
|
||||
if !ok || tOff >= pcs[in] || loopAligns[target] {
|
||||
continue
|
||||
}
|
||||
loopAligns[target] = true
|
||||
changed = true
|
||||
}
|
||||
if !changed {
|
||||
break
|
||||
}
|
||||
}
|
||||
// Final layout with the complete alignment set.
|
||||
offsets, alignPad, pcs, bodyEnd := loong64Layout(t, guardLen+len(prologue), fi, loopAligns, alignInstrs)
|
||||
bodyLen := bodyEnd - (guardLen + len(prologue))
|
||||
pcRelPcs := make(map[*ast.Instr]int, len(pcRelTarget))
|
||||
for in, tgt := range pcRelTarget {
|
||||
pcRelPcs[in] = pcs[tgt]
|
||||
}
|
||||
var out []byte
|
||||
if fi.needSplit {
|
||||
@@ -84,7 +124,20 @@ func assembleLOONG64(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry,
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
code, err := encodeLOONG64Instr(in, pc, offsets, fi, &relocs, resolve)
|
||||
// PCALIGN pads to the requested boundary with andi $0, $0, 0, the
|
||||
// architecture's NOP, and encodes to nothing itself.
|
||||
if strings.ToUpper(in.Mnemonic.Text) == "PCALIGN" {
|
||||
pad := loong64PCAlignPad(pc, in)
|
||||
out = append(out, loong64PadBytes(pad)...)
|
||||
pc += pad
|
||||
continue
|
||||
}
|
||||
// Loop-head alignment padding precedes the instruction.
|
||||
if pad := alignPad[in]; pad > 0 {
|
||||
out = append(out, loong64PadBytes(pad)...)
|
||||
pc += pad
|
||||
}
|
||||
code, err := encodeLOONG64Instr(in, pc, offsets, fi, &relocs, resolve, pcRelPcs)
|
||||
if err != nil {
|
||||
return nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", in.Mnemonic.Text, err)
|
||||
}
|
||||
@@ -115,6 +168,115 @@ func assembleLOONG64(t *ast.Text) ([]byte, map[string]int, []Reloc, []LineEntry,
|
||||
return out, offsets, relocs, lines, spadj, nil
|
||||
}
|
||||
|
||||
// loong64PCRelOffset reports the N of a branch operand spelled N(PC): the
|
||||
// displacement counted in source instructions from the branch itself.
|
||||
func loong64PCRelOffset(instr *ast.Instr) (int, bool) {
|
||||
mnem := strings.ToUpper(instr.Mnemonic.Text)
|
||||
branch := false
|
||||
switch mnem {
|
||||
case "JMP":
|
||||
branch = len(instr.Operands) == 1
|
||||
case "JAL", "CALL", "BL":
|
||||
branch = len(instr.Operands) == 1 || len(instr.Operands) == 2
|
||||
case "BFPT", "BFPF":
|
||||
branch = len(instr.Operands) == 1
|
||||
case "BEQ", "BNE", "BLT", "BGE", "BLTU", "BGEU",
|
||||
"BEQZ", "BNEZ", "BLTZ", "BGEZ", "BLEZ", "BGTZ":
|
||||
branch = len(instr.Operands) >= 2
|
||||
}
|
||||
if !branch {
|
||||
return 0, false
|
||||
}
|
||||
op := instr.Operands[len(instr.Operands)-1]
|
||||
if op.Kind == ast.OpAddr && op.Addr.Sym == nil && op.Addr.Base == "PC" {
|
||||
return int(op.Addr.Offset), true
|
||||
}
|
||||
return 0, false
|
||||
}
|
||||
|
||||
// loong64Layout walks the function body once and returns the label offsets,
|
||||
// the loop-alignment padding due before each instruction (a pad of 0 needs
|
||||
// nothing), the pc each instruction starts at (its padding included) and the
|
||||
// first pc past the body. Explicit PCALIGN pads, the alignment pads for the
|
||||
// labels in aligns and those for the instructions in alignInstrs (backward
|
||||
// PC-relative targets) all contribute, mirroring the toolchain's layout
|
||||
// pass.
|
||||
func loong64Layout(t *ast.Text, start int, fi loong64FrameInfo, aligns map[string]bool, alignInstrs map[*ast.Instr]bool) (map[string]int, map[*ast.Instr]int, map[*ast.Instr]int, int) {
|
||||
offsets := map[string]int{}
|
||||
alignPad := map[*ast.Instr]int{}
|
||||
pcs := map[*ast.Instr]int{}
|
||||
pos := start
|
||||
pendingAlign := false
|
||||
var pendingNames []string
|
||||
explicit := false
|
||||
for _, stmt := range t.Body {
|
||||
switch s := stmt.(type) {
|
||||
case *ast.Label:
|
||||
if aligns[s.Name.Text] {
|
||||
pendingAlign = true
|
||||
}
|
||||
pendingNames = append(pendingNames, s.Name.Text)
|
||||
// Provisional: a branch to the label lands here unless a loop
|
||||
// alignment pad follows, in which case the label resolves to the
|
||||
// padded instruction (the toolchain's labels bind to the branch
|
||||
// target instruction, which the padding pass precedes).
|
||||
offsets[s.Name.Text] = pos
|
||||
case *ast.Instr:
|
||||
if strings.ToUpper(s.Mnemonic.Text) == "PCALIGN" {
|
||||
pos += loong64PCAlignPad(pos, s)
|
||||
explicit = true
|
||||
continue
|
||||
}
|
||||
if pendingAlign {
|
||||
pendingAlign = false
|
||||
if pos&15 != 0 {
|
||||
alignPad[s] = 16 - pos&15
|
||||
}
|
||||
}
|
||||
if alignInstrs[s] && pos&15 != 0 {
|
||||
alignPad[s] = 16 - pos&15
|
||||
}
|
||||
if !explicit {
|
||||
for _, n := range pendingNames {
|
||||
offsets[n] = pos + alignPad[s]
|
||||
}
|
||||
}
|
||||
pendingNames = nil
|
||||
explicit = false
|
||||
pcs[s] = pos + alignPad[s]
|
||||
pos += alignPad[s] + loong64InstrSize(s, fi)
|
||||
}
|
||||
}
|
||||
return offsets, alignPad, pcs, pos
|
||||
}
|
||||
|
||||
// loong64BranchTarget reports the local label a branch-like instruction
|
||||
// transfers to, the loop-head signal the toolchain derives from backward
|
||||
// branch targets.
|
||||
func loong64BranchTarget(instr *ast.Instr) (string, bool) {
|
||||
mnem := strings.ToUpper(instr.Mnemonic.Text)
|
||||
ops := instr.Operands
|
||||
var op *ast.Operand
|
||||
switch {
|
||||
case mnem == "JMP" || mnem == "JAL" || mnem == "BFPT" || mnem == "BFPF":
|
||||
if len(ops) != 1 {
|
||||
return "", false
|
||||
}
|
||||
op = ops[0]
|
||||
case mnem == "TEQ" || mnem == "TNE":
|
||||
return "", false
|
||||
case len(ops) >= 2:
|
||||
op = ops[len(ops)-1]
|
||||
default:
|
||||
return "", false
|
||||
}
|
||||
if op.Kind == ast.OpAddr && op.Addr.Sym != nil && op.Addr.Sym.Pseudo == "" &&
|
||||
op.Addr.Base == "" && op.Addr.Sym.Name != "" {
|
||||
return op.Addr.Sym.Name, true
|
||||
}
|
||||
return "", false
|
||||
}
|
||||
|
||||
// loong64JumpChain precomputes jump-to-jump folding, mirroring the linker's
|
||||
// branch-chasing pass: a label whose first instruction is an unconditional
|
||||
// local jump redirects its own jumpers to the ultimate target. The Go
|
||||
@@ -177,21 +339,55 @@ func l64LabelOK(op *ast.Operand) (string, bool) {
|
||||
return "", false
|
||||
}
|
||||
|
||||
// l64SubToAdd rewrites the SUB family with an immediate first operand onto
|
||||
// its ADD counterpart with the negated immediate: LoongArch has no
|
||||
// subtract-immediate instructions, and the toolchain folds SUB $v into the
|
||||
// ADD immediate form through the same optab matching (the $0 fold into 3R
|
||||
// and the large-constant materialisations included). The negation is the
|
||||
// second result; the operand is left untouched because the size pass
|
||||
// normalises the same instruction.
|
||||
func l64SubToAdd(mnem string, ops []*ast.Operand) (string, bool) {
|
||||
if len(ops) >= 2 && isImmOperand(ops[0]) {
|
||||
switch mnem {
|
||||
case "SUB":
|
||||
return "ADD", true
|
||||
case "SUBW":
|
||||
return "ADDW", true
|
||||
case "SUBV", "SUBVU":
|
||||
return "ADDV", true
|
||||
}
|
||||
}
|
||||
return mnem, false
|
||||
}
|
||||
|
||||
// loong64InstrSize returns the encoded size of an instruction: 4 bytes for
|
||||
// most, more for the multi-instruction expansions.
|
||||
func loong64InstrSize(instr *ast.Instr, fi loong64FrameInfo) int {
|
||||
mnem := strings.ToUpper(instr.Mnemonic.Text)
|
||||
ops := instr.Operands
|
||||
var neg bool
|
||||
mnem, neg = l64SubToAdd(mnem, ops)
|
||||
|
||||
if mnem == "RET" {
|
||||
return len(loong64Return(fi))
|
||||
}
|
||||
switch mnem {
|
||||
case "END", "FUNCDATA", "PCDATA":
|
||||
return 0 // bookkeeping statements contribute no bytes
|
||||
case "GETCALLERPC":
|
||||
return 4 // or rd, r1, r0
|
||||
case "TEQ", "TNE":
|
||||
return 8 // bne/beq over the BREAK, then BREAK
|
||||
case "PRELDX":
|
||||
return 20 // the four-instruction constant materialisation + preldx
|
||||
case "MOV", "MOVB", "MOVH", "MOVW", "MOVV", "MOVBU", "MOVHU", "MOVWU", "MOVF", "MOVD":
|
||||
return loong64MovSize(mnem, ops, fi)
|
||||
case "ADD", "ADDW", "ADDV", "ADDVU", "AND", "OR", "XOR", "SGT", "SGTU":
|
||||
if len(ops) >= 2 && isImmOperand(ops[0]) {
|
||||
v := l64Imm64(ops[0])
|
||||
if neg {
|
||||
v = -v
|
||||
}
|
||||
if v == 0 {
|
||||
return 4 // folds into the 3R form (rk = R0)
|
||||
}
|
||||
@@ -229,11 +425,50 @@ func loong64InstrSize(instr *ast.Instr, fi loong64FrameInfo) int {
|
||||
return 4
|
||||
}
|
||||
|
||||
// loong64PCAlignPad returns the padding PCALIGN inserts before the next
|
||||
// instruction so that it starts at the requested boundary relative to the
|
||||
// function start. The boundary must be a power of two between 8 and 2048, as
|
||||
// the toolchain requires; anything else pads nothing.
|
||||
func loong64PCAlignPad(pos int, instr *ast.Instr) int {
|
||||
if len(instr.Operands) != 1 || !isImmOperand(instr.Operands[0]) {
|
||||
return 0
|
||||
}
|
||||
align := int(immFromOperand(instr.Operands[0]))
|
||||
if align < 8 || align > 2048 || align&(align-1) != 0 {
|
||||
return 0
|
||||
}
|
||||
return (align - pos%align) % align
|
||||
}
|
||||
|
||||
// loong64PadBytes renders PCALIGN padding: the toolchain emits andi $0, $0, 0
|
||||
// (the architecture's NOP) for every full 4 bytes of pad.
|
||||
func loong64PadBytes(pad int) []byte {
|
||||
nop := l64wordLE(l64irr(l64DualTable["AND"].imm, 0, 0, 0))
|
||||
out := make([]byte, 0, pad/4*len(nop))
|
||||
for i := 0; i < pad/4; i++ {
|
||||
out = append(out, nop...)
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// encodeLOONG64Instr encodes a single LoongArch instruction.
|
||||
func encodeLOONG64Instr(instr *ast.Instr, pc int, offsets map[string]int, fi loong64FrameInfo, relocs *[]Reloc, resolve func(string) string) ([]byte, error) {
|
||||
func encodeLOONG64Instr(instr *ast.Instr, pc int, offsets map[string]int, fi loong64FrameInfo, relocs *[]Reloc, resolve func(string) string, pcRelPcs map[*ast.Instr]int) ([]byte, error) {
|
||||
mnem := strings.ToUpper(instr.Mnemonic.Text)
|
||||
ops := instr.Operands
|
||||
|
||||
// The SUB family with an immediate first operand folds onto the ADD
|
||||
// immediate form with the negated immediate; the negation happens on a
|
||||
// copy of the operand, never on the shared syntax tree.
|
||||
mnem, neg := l64SubToAdd(mnem, ops)
|
||||
if neg {
|
||||
c := *ops[0]
|
||||
c.Imm.Val = -c.Imm.Val
|
||||
ops2 := make([]*ast.Operand, len(ops))
|
||||
ops2[0] = &c
|
||||
copy(ops2[1:], ops[1:])
|
||||
ops = ops2
|
||||
}
|
||||
|
||||
// Pseudo-instructions and the branches first.
|
||||
switch mnem {
|
||||
case "RET":
|
||||
@@ -249,10 +484,110 @@ func encodeLOONG64Instr(instr *ast.Instr, pc int, offsets map[string]int, fi loo
|
||||
return nil, fmt.Errorf("WORD expects 1 operand, got %d", len(ops))
|
||||
}
|
||||
return l64wordLE(uint32(immFromOperand(ops[0]))), nil
|
||||
case "END", "FUNCDATA", "PCDATA", "GETCALLERPC":
|
||||
// The assembler's bookkeeping statements. END, FUNCDATA and PCDATA
|
||||
// contribute no bytes, the same shapes GOARCH=loong64 go tool asm
|
||||
// accepts and emits nothing for; GETCALLERPC reads the caller's
|
||||
// address out of R1 (RA) as or rd, r1, r0.
|
||||
switch mnem {
|
||||
case "END":
|
||||
if len(ops) != 0 {
|
||||
return nil, fmt.Errorf("END expects no operands, got %d", len(ops))
|
||||
}
|
||||
return nil, nil
|
||||
case "FUNCDATA":
|
||||
if len(ops) != 2 || !isImmOperand(ops[0]) {
|
||||
return nil, fmt.Errorf("FUNCDATA expects $n, sym(SB)")
|
||||
}
|
||||
return nil, nil
|
||||
case "PCDATA":
|
||||
if len(ops) != 2 || !isImmOperand(ops[0]) || !isImmOperand(ops[1]) {
|
||||
return nil, fmt.Errorf("PCDATA expects $n, $n")
|
||||
}
|
||||
return nil, nil
|
||||
}
|
||||
if len(ops) != 1 || isMemOperand(ops[0]) || loong64RegClass(operandRegName(ops[0])) != l64ClsGR {
|
||||
return nil, fmt.Errorf("GETCALLERPC expects a general register")
|
||||
}
|
||||
rd := loong64RegNum(operandRegName(ops[0]))
|
||||
if rd < 0 {
|
||||
return nil, fmt.Errorf("GETCALLERPC: invalid register operand")
|
||||
}
|
||||
return l64wordLE(l64rrr(l64movRegTable["MOVV"].op, 0, 1, rd)), nil
|
||||
case "NEGW", "NEGV":
|
||||
// The integer negation pseudo is a subtract from zero:
|
||||
// NEGW src, dst → sub.w r0, src, dst.
|
||||
if len(ops) != 2 {
|
||||
return nil, fmt.Errorf("%s expects 2 operands, got %d", mnem, len(ops))
|
||||
}
|
||||
src, dst := l64Reg(ops[0]), l64Reg(ops[1])
|
||||
if src < 0 || dst < 0 {
|
||||
return nil, fmt.Errorf("invalid register operand")
|
||||
}
|
||||
sub := l64InstrTable["SUBW"].op
|
||||
if mnem == "NEGV" {
|
||||
sub = l64InstrTable["SUBV"].op
|
||||
}
|
||||
return l64wordLE(l64rrr(sub, src, 0, dst)), nil
|
||||
case "TEQ", "TNE":
|
||||
// The trap pseudo expands to two instructions: bne/beq rj, rd over
|
||||
// the BREAK (offset 2 instruction units), then BREAK $code.
|
||||
if len(ops) != 2 && len(ops) != 3 {
|
||||
return nil, fmt.Errorf("%s expects 2 or 3 operands, got %d", mnem, len(ops))
|
||||
}
|
||||
code := int(immFromOperand(ops[0]))
|
||||
rj, rd := 0, l64Reg(ops[len(ops)-1])
|
||||
if len(ops) == 3 {
|
||||
rj = l64Reg(ops[1])
|
||||
}
|
||||
if rj < 0 || rd < 0 {
|
||||
return nil, fmt.Errorf("invalid register operand")
|
||||
}
|
||||
bop := l64branchTable["BNE"]
|
||||
if mnem == "TNE" {
|
||||
bop = l64branchTable["BEQ"]
|
||||
}
|
||||
return l64WordsLE(
|
||||
l64irr16(bop, 2, rj, rd),
|
||||
l64i15(l64InstrTable["BREAK"].op, code),
|
||||
), nil
|
||||
case "PRELDX":
|
||||
// preldx offset(Rbase), $n, $hint: the 64-bit descriptor n packs
|
||||
// (addrSeq, blockSize, blockNums, stride); the constant v built from
|
||||
// it materialises in R30 across four instructions, then the preldx.
|
||||
if len(ops) != 3 || !isMemOperand(ops[0]) || !isImmOperand(ops[1]) || !isImmOperand(ops[2]) {
|
||||
return nil, fmt.Errorf("PRELDX expects offset(reg), $n, $hint")
|
||||
}
|
||||
rj := loong64RegNum(ops[0].Addr.Base)
|
||||
if rj < 0 {
|
||||
return nil, fmt.Errorf("invalid register operand")
|
||||
}
|
||||
n := uint64(l64Imm64(ops[1]))
|
||||
hint := int(l64Imm64(ops[2]))
|
||||
addrSeq := (n >> 0) & 0x1
|
||||
blkSize := (n >> 1) & 0x7ff
|
||||
blkNums := (n >> 12) & 0x1ff
|
||||
stride := (n >> 21) & 0xffff
|
||||
v := uint64(ops[0].Addr.Offset)&0xffff + addrSeq<<16 +
|
||||
((blkSize/16)-1)<<20 + (blkNums-1)<<32 + stride<<44
|
||||
const (
|
||||
lu12iw = 0x0a << 25
|
||||
lu32id = 0x0b << 25
|
||||
lu52id = 0x00c << 22
|
||||
ori = 0x00e << 22
|
||||
preldx = 0x7058 << 15
|
||||
)
|
||||
return l64WordsLE(
|
||||
l64ir(lu12iw, int(uint32(v>>12)), 30),
|
||||
l64irr(ori, int(uint32(v)), 30, 30),
|
||||
l64ir(lu32id, int(uint32(v>>32)), 30),
|
||||
l64irr(lu52id, int(uint32(v>>52)), 30, 30),
|
||||
l64rrr(preldx, 30, rj, hint),
|
||||
), nil
|
||||
case "JMP", "B":
|
||||
return encodeLOONG64Branch(instr, mnem, pc, offsets, false, resolve, relocs)
|
||||
return encodeLOONG64Branch(instr, mnem, pc, offsets, false, resolve, relocs, pcRelPcs)
|
||||
case "JAL", "CALL", "BL":
|
||||
return encodeLOONG64Branch(instr, mnem, pc, offsets, true, resolve, relocs)
|
||||
return encodeLOONG64Branch(instr, mnem, pc, offsets, true, resolve, relocs, pcRelPcs)
|
||||
case "MOV", "MOVB", "MOVH", "MOVW", "MOVV", "MOVBU", "MOVHU", "MOVWU", "MOVF", "MOVD":
|
||||
return encodeLOONG64Mov(instr, mnem, fi, relocs)
|
||||
}
|
||||
@@ -262,12 +597,12 @@ func encodeLOONG64Instr(instr *ast.Instr, pc int, offsets map[string]int, fi loo
|
||||
if mnem == "JIRL" {
|
||||
return encodeLOONG64Jirl(op, ops)
|
||||
}
|
||||
return encodeLOONG64Branch16(mnem, op, ops, pc, offsets, resolve)
|
||||
return encodeLOONG64Branch16(instr, mnem, op, ops, pc, offsets, resolve, pcRelPcs)
|
||||
}
|
||||
// Single-register branches with 21-bit offsets (BLTZ/BGEZ/BLEZ/BGTZ,
|
||||
// BFPT/BFPF; BEQZ/BNEZ are reached through BEQ/BNE with R0).
|
||||
if op, ok := l64branch21Table[mnem]; ok {
|
||||
return encodeLOONG64Branch21(mnem, op, ops, pc, offsets, resolve)
|
||||
return encodeLOONG64Branch21(instr, mnem, op, ops, pc, offsets, resolve, pcRelPcs)
|
||||
}
|
||||
// B/BL aliases reached only via JMP/JAL above.
|
||||
|
||||
@@ -533,11 +868,23 @@ func encodeLOONG64Instr(instr *ast.Instr, pc int, offsets map[string]int, fi loo
|
||||
//
|
||||
// JMP/B label → b label JMP/B (rj) → jirl r0, rj, 0
|
||||
// JAL/CALL/BL label → bl label JAL/CALL/BL (rj) → jirl r1, rj, 0
|
||||
func encodeLOONG64Branch(instr *ast.Instr, mnem string, pc int, offsets map[string]int, link bool, resolve func(string) string, relocs *[]Reloc) ([]byte, error) {
|
||||
func encodeLOONG64Branch(instr *ast.Instr, mnem string, pc int, offsets map[string]int, link bool, resolve func(string) string, relocs *[]Reloc, pcRelPcs map[*ast.Instr]int) ([]byte, error) {
|
||||
if len(instr.Operands) != 1 {
|
||||
return nil, fmt.Errorf("%s expects 1 operand, got %d", mnem, len(instr.Operands))
|
||||
}
|
||||
op := instr.Operands[0]
|
||||
// PC-relative displacement: N(PC) resolves to the instruction N slots
|
||||
// away in source order (the toolchain's parse-time count), and the field
|
||||
// carries the final pc distance in instruction units.
|
||||
if op.Addr.Sym == nil && op.Addr.Base == "PC" {
|
||||
targetPc, ok := pcRelPcs[instr]
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("%s: PC-relative target %d out of range", mnem, op.Addr.Offset)
|
||||
}
|
||||
v := (targetPc - pc) >> 2
|
||||
opc := l64jumpTable[mnem]
|
||||
return l64wordLE(l64bbl(opc, v)), nil
|
||||
}
|
||||
if isMemOperand(op) && op.Addr.Base != "" && op.Addr.Index == "" && op.Addr.Sym == nil {
|
||||
// Indirect: (rj) → jirl.
|
||||
rj := loong64RegNum(op.Addr.Base)
|
||||
@@ -620,16 +967,28 @@ func l64offsetOperand(op *ast.Operand) (int32, bool) {
|
||||
// encodeLOONG64Branch16 encodes a 16-bit branch (BEQ/BNE/BLT/BGE/BLTU/BGEU):
|
||||
// INSTR rj, rd, label, or INSTR rj, label with rd = R0, which the toolchain
|
||||
// turns into the 21-bit BEQZ/BNEZ form when the register is the only operand.
|
||||
func encodeLOONG64Branch16(mnem string, op uint32, ops []*ast.Operand, pc int, offsets map[string]int, resolve func(string) string) ([]byte, error) {
|
||||
func encodeLOONG64Branch16(instr *ast.Instr, mnem string, op uint32, ops []*ast.Operand, pc int, offsets map[string]int, resolve func(string) string, pcRelPcs map[*ast.Instr]int) ([]byte, error) {
|
||||
if len(ops) != 2 && len(ops) != 3 {
|
||||
return nil, fmt.Errorf("%s expects 2 or 3 operands, got %d", mnem, len(ops))
|
||||
}
|
||||
target := resolve(l64Label(ops[len(ops)-1]))
|
||||
targetOff, ok := offsets[target]
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("undefined label %q%s", target, suggestLabel(target, offsets))
|
||||
var target string
|
||||
var v int
|
||||
lastOp := ops[len(ops)-1]
|
||||
if lastOp.Kind == ast.OpAddr && lastOp.Addr.Sym == nil && lastOp.Addr.Base == "PC" {
|
||||
// N(PC) resolves to the instruction N slots away in source order.
|
||||
targetPc, ok := pcRelPcs[instr]
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("%s: PC-relative target %d out of range", mnem, lastOp.Addr.Offset)
|
||||
}
|
||||
v = (targetPc - pc) >> 2
|
||||
} else {
|
||||
target = resolve(l64Label(lastOp))
|
||||
targetOff, ok := offsets[target]
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("undefined label %q%s", target, suggestLabel(target, offsets))
|
||||
}
|
||||
v = (targetOff - pc) >> 2
|
||||
}
|
||||
v := (targetOff - pc) >> 2
|
||||
if len(ops) == 2 {
|
||||
// Single register: BEQ rj, label → beqz (21-bit), and the BLTZ/
|
||||
// BGEZ-family aliases encoded with rj in the rj field.
|
||||
@@ -690,33 +1049,55 @@ func encodeLOONG64Branch16(mnem string, op uint32, ops []*ast.Operand, pc int, o
|
||||
// BFPT/BFPF use the 21-bit offset form (register in the rj field), while
|
||||
// BGTZ/BLEZ, which the toolchain encodes with the register in the rd field
|
||||
// and a 16-bit offset, are handled separately.
|
||||
func encodeLOONG64Branch21(mnem string, op uint32, ops []*ast.Operand, pc int, offsets map[string]int, resolve func(string) string) ([]byte, error) {
|
||||
if len(ops) != 2 {
|
||||
func encodeLOONG64Branch21(instr *ast.Instr, mnem string, op uint32, ops []*ast.Operand, pc int, offsets map[string]int, resolve func(string) string, pcRelPcs map[*ast.Instr]int) ([]byte, error) {
|
||||
isBF := mnem == "BFPT" || mnem == "BFPF"
|
||||
if len(ops) != 2 && !(isBF && (len(ops) == 1 || len(ops) == 2)) {
|
||||
return nil, fmt.Errorf("%s expects 2 operands, got %d", mnem, len(ops))
|
||||
}
|
||||
target := resolve(l64Label(ops[1]))
|
||||
targetOff, ok := offsets[target]
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("undefined label %q%s", target, suggestLabel(target, offsets))
|
||||
}
|
||||
v := (targetOff - pc) >> 2
|
||||
rj := 0 // BFPT/BFPF default to FCC0
|
||||
if mnem != "BFPT" && mnem != "BFPF" {
|
||||
var rj int
|
||||
tgtOp := ops[len(ops)-1]
|
||||
if isBF {
|
||||
// BFPT/BFPF test an FCC condition register, defaulting to FCC0 when
|
||||
// spelled without one.
|
||||
rj = 0
|
||||
if len(ops) == 2 {
|
||||
rj = l64Reg(ops[0])
|
||||
if rj < 0 {
|
||||
return nil, fmt.Errorf("invalid register operand")
|
||||
}
|
||||
}
|
||||
} else {
|
||||
rj = l64Reg(ops[0])
|
||||
if rj < 0 {
|
||||
return nil, fmt.Errorf("invalid register operand")
|
||||
}
|
||||
}
|
||||
var v int
|
||||
if tgtOp.Kind == ast.OpAddr && tgtOp.Addr.Sym == nil && tgtOp.Addr.Base == "PC" {
|
||||
// N(PC) resolves to the instruction N slots away in source order.
|
||||
targetPc, ok := pcRelPcs[instr]
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("%s: PC-relative target %d out of range", mnem, tgtOp.Addr.Offset)
|
||||
}
|
||||
v = (targetPc - pc) >> 2
|
||||
} else {
|
||||
target := resolve(l64Label(tgtOp))
|
||||
targetOff, ok := offsets[target]
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("undefined label %q%s", target, suggestLabel(target, offsets))
|
||||
}
|
||||
v = (targetOff - pc) >> 2
|
||||
}
|
||||
if mnem == "BGTZ" || mnem == "BLEZ" {
|
||||
// The toolchain swaps the register into the rd field and keeps the
|
||||
// 16-bit offset form.
|
||||
if (v<<16)>>16 != v {
|
||||
return nil, fmt.Errorf("branch to %q too far (16-bit range)", target)
|
||||
return nil, fmt.Errorf("branch %d too far (16-bit range)", v)
|
||||
}
|
||||
return l64wordLE(l64irr16(op, v, 0, rj)), nil
|
||||
}
|
||||
if (v<<11)>>11 != v {
|
||||
return nil, fmt.Errorf("branch to %q too far (21-bit range)", target)
|
||||
return nil, fmt.Errorf("branch %d too far (21-bit range)", v)
|
||||
}
|
||||
return l64wordLE(l64ir21(op, v, rj)), nil
|
||||
}
|
||||
@@ -1737,6 +2118,30 @@ func encodeLOONG64Vector(instr *ast.Instr, mnem string, fi loong64FrameInfo) ([]
|
||||
return l64wordLE(l64rr(l64InstrTable[mnem].op, vj, fcc)), true, nil
|
||||
}
|
||||
|
||||
// Four-register forms (vshuf.b): INSTR va, vk, vj, vd.
|
||||
if l64Vec4R[mnem] {
|
||||
if len(ops) != 4 {
|
||||
return nil, true, fmt.Errorf("%s expects 4 operands, got %d", mnem, len(ops))
|
||||
}
|
||||
va, err := vec(ops[0])
|
||||
if err != nil {
|
||||
return nil, true, err
|
||||
}
|
||||
vk, err := vec(ops[1])
|
||||
if err != nil {
|
||||
return nil, true, err
|
||||
}
|
||||
vj, err := vec(ops[2])
|
||||
if err != nil {
|
||||
return nil, true, err
|
||||
}
|
||||
vd, err := vec(ops[3])
|
||||
if err != nil {
|
||||
return nil, true, err
|
||||
}
|
||||
return l64wordLE(l64InstrTable[mnem].op | uint32(va&0x1f)<<15 | uint32(vk&0x1f)<<10 | uint32(vj&0x1f)<<5 | uint32(vd&0x1f)), true, nil
|
||||
}
|
||||
|
||||
// Three-register forms: INSTR vk, vj, vd or INSTR vk, vd (vj = vd).
|
||||
if len(ops) != 2 && len(ops) != 3 {
|
||||
return nil, true, fmt.Errorf("%s expects 2 or 3 operands, got %d", mnem, len(ops))
|
||||
|
||||
+449
-12
@@ -278,6 +278,7 @@ const (
|
||||
l64Fpreld // preld (2RI12 + 5-bit hint)
|
||||
l64Fvvv // 3R vector (LSX/LASX): op | vk<<10 | vj<<5 | vd
|
||||
l64Fvcf // vector-to-condition: op | subop<<10 | vj<<5 | fcc
|
||||
l64Fvvvv // 4R vector shuffle: op | va<<15 | vk<<10 | vj<<5 | vd
|
||||
)
|
||||
|
||||
// l64Enc is one instruction's encoding: its bit layout (format) and the
|
||||
@@ -336,6 +337,10 @@ var l64VecImmInfo = map[string]l64VecImmEnc{}
|
||||
// vpcnt.v).
|
||||
var l64Vec2R = map[string]bool{}
|
||||
|
||||
// l64Vec4R marks the four-operand vector mnemonics (INSTR va, vk, vj, vd,
|
||||
// such as vshuf.b).
|
||||
var l64Vec4R = map[string]bool{}
|
||||
|
||||
// l64VmovqOps holds the VMOVQ/XVMOVQ opcode constants (pre-shifted to bit
|
||||
// 15), read off `go tool objdump` of GOARCH=loong64 `go tool asm` kernels.
|
||||
type l64VmovqEnc struct {
|
||||
@@ -445,6 +450,14 @@ func init() {
|
||||
// bank is the FP registers (the toolchain spells it `FFINTDV F0, F1`),
|
||||
// so the entry stays on the 2R integer/FP format.
|
||||
"FFINTDV": 0x474a << 10,
|
||||
// The rest of the scalar conversions (all F-bank, 2R).
|
||||
"FFINTFW": 0x4744 << 10, // ffint.s.w
|
||||
"FFINTFV": 0x4746 << 10, // ffint.s.l
|
||||
"FFINTDW": 0x4748 << 10, // ffint.d.w
|
||||
"FTINTWF": 0x46c1 << 10, // ftint.w.s
|
||||
"FTINTWD": 0x46c2 << 10, // ftint.w.d
|
||||
"FTINTVF": 0x46c9 << 10, // ftint.l.s
|
||||
"FTINTVD": 0x46ca << 10, // ftint.l.d
|
||||
}
|
||||
for m, op := range rr {
|
||||
l64InstrTable[m] = l64Enc{format: l64Frr, op: op}
|
||||
@@ -568,6 +581,12 @@ func init() {
|
||||
"AMADDDBW": 0x070D4 << 15, "AMADDDBV": 0x070D5 << 15,
|
||||
"AMANDDBW": 0x070D6 << 15, "AMANDDBV": 0x070D7 << 15,
|
||||
"AMORDBW": 0x070D8 << 15, "AMORDBV": 0x070D9 << 15,
|
||||
// The remaining _dbar exchange variants (loong64enc1.s).
|
||||
"AMXORDBW": 0x070DA << 15, "AMXORDBV": 0x070DB << 15,
|
||||
"AMMAXDBW": 0x070DC << 15, "AMMAXDBV": 0x070DD << 15,
|
||||
"AMMINDBW": 0x070DE << 15, "AMMINDBV": 0x070DF << 15,
|
||||
"AMMAXDBWU": 0x070E0 << 15, "AMMAXDBVU": 0x070E1 << 15,
|
||||
"AMMINDBWU": 0x070E2 << 15, "AMMINDBVU": 0x070E3 << 15,
|
||||
}
|
||||
for m, op := range am {
|
||||
l64InstrTable[m] = l64Enc{format: l64Fam, op: op}
|
||||
@@ -590,6 +609,240 @@ func init() {
|
||||
"XVANDV": {0xEA4C << 15, true}, "XVXORV": {0xEA4E << 15, true},
|
||||
"XVSEQB": {0xE800 << 15, true}, "XVSEQV": {0xE803 << 15, true},
|
||||
}
|
||||
|
||||
// The integer and FP add/subtract families: [X]VADD and [X]VSUB by lane
|
||||
// width, plus the [X]VSADD/[X]VSSUB saturating pairs.
|
||||
// Opcodes transcribed from the toolchain's loong64enc1.s.
|
||||
addsub := map[string]l64Vec3Enc{
|
||||
"VADDB": {0xE014 << 15, false}, "VADDH": {0xE015 << 15, false},
|
||||
"VADDD": {0xE262 << 15, false}, "VADDF": {0xE261 << 15, false},
|
||||
"VADDQ": {0xE25A << 15, false},
|
||||
"VSUBB": {0xE018 << 15, false}, "VSUBH": {0xE019 << 15, false},
|
||||
"VSUBW": {0xE01A << 15, false}, "VSUBV": {0xE01B << 15, false},
|
||||
"VSUBQ": {0xE25B << 15, false},
|
||||
"VSUBF": {0xE265 << 15, false}, "VSUBD": {0xE266 << 15, false},
|
||||
"VSADDB": {0xE08C << 15, false}, "VSADDH": {0xE08D << 15, false},
|
||||
"VSADDW": {0xE08E << 15, false}, "VSADDV": {0xE08F << 15, false},
|
||||
"VSADDBU": {0xE094 << 15, false}, "VSADDHU": {0xE095 << 15, false},
|
||||
"VSADDWU": {0xE096 << 15, false}, "VSADDVU": {0xE097 << 15, false},
|
||||
"VSSUBB": {0xE090 << 15, false}, "VSSUBH": {0xE091 << 15, false},
|
||||
"VSSUBW": {0xE092 << 15, false}, "VSSUBV": {0xE093 << 15, false},
|
||||
"VSSUBBU": {0xE098 << 15, false}, "VSSUBHU": {0xE099 << 15, false},
|
||||
"VSSUBWU": {0xE09A << 15, false}, "VSSUBVU": {0xE09B << 15, false},
|
||||
"XVADDB": {0xE814 << 15, true}, "XVADDH": {0xE815 << 15, true},
|
||||
"XVADDW": {0xE816 << 15, true},
|
||||
"XVADDD": {0xEA62 << 15, true}, "XVADDF": {0xEA61 << 15, true},
|
||||
"XVADDQ": {0xEA5A << 15, true},
|
||||
"XVSUBB": {0xE818 << 15, true}, "XVSUBH": {0xE819 << 15, true},
|
||||
"XVSUBW": {0xE81A << 15, true}, "XVSUBV": {0xE81B << 15, true},
|
||||
"XVSUBQ": {0xEA5B << 15, true},
|
||||
"XVSUBF": {0xEA65 << 15, true}, "XVSUBD": {0xEA66 << 15, true},
|
||||
"XVSADDB": {0xE88C << 15, true}, "XVSADDH": {0xE88D << 15, true},
|
||||
"XVSADDW": {0xE88E << 15, true}, "XVSADDV": {0xE88F << 15, true},
|
||||
"XVSADDBU": {0xE894 << 15, true}, "XVSADDHU": {0xE895 << 15, true},
|
||||
"XVSADDWU": {0xE896 << 15, true}, "XVSADDVU": {0xE897 << 15, true},
|
||||
"XVSSUBB": {0xE890 << 15, true}, "XVSSUBH": {0xE891 << 15, true},
|
||||
"XVSSUBW": {0xE892 << 15, true}, "XVSSUBV": {0xE893 << 15, true},
|
||||
"XVSSUBBU": {0xE898 << 15, true}, "XVSSUBHU": {0xE899 << 15, true},
|
||||
"XVSSUBWU": {0xE89A << 15, true}, "XVSSUBVU": {0xE89B << 15, true},
|
||||
}
|
||||
|
||||
// The multiply families: plain and high-half [X]VMUL/[X]VMUH, the
|
||||
// widening [X]VMULW{EV,OD} ladder and its accumulating [X]VMADDW twins,
|
||||
// plus the [X]VMADD/[X]VMSUB fused multiply-add and the [X]VDIV/[X]VMOD
|
||||
// divide and modulo pairs.
|
||||
muldiv := map[string]l64Vec3Enc{
|
||||
"VMULB": {0xE108 << 15, false}, "VMULH": {0xE109 << 15, false},
|
||||
"VMULW": {0xE10A << 15, false}, "VMULV": {0xE10B << 15, false},
|
||||
"VMUHB": {0xE10C << 15, false}, "VMUHH": {0xE10D << 15, false},
|
||||
"VMUHW": {0xE10E << 15, false}, "VMUHV": {0xE10F << 15, false},
|
||||
"VMUHBU": {0xE110 << 15, false}, "VMUHHU": {0xE111 << 15, false},
|
||||
"VMUHWU": {0xE112 << 15, false}, "VMUHVU": {0xE113 << 15, false},
|
||||
"VMULWEVHB": {0xE120 << 15, false}, "VMULWEVWH": {0xE121 << 15, false},
|
||||
"VMULWEVVW": {0xE122 << 15, false}, "VMULWEVQV": {0xE123 << 15, false},
|
||||
"VMULWODHB": {0xE124 << 15, false}, "VMULWODWH": {0xE125 << 15, false},
|
||||
"VMULWODVW": {0xE126 << 15, false}, "VMULWODQV": {0xE127 << 15, false},
|
||||
"VMULWEVHBU": {0xE130 << 15, false}, "VMULWEVWHU": {0xE131 << 15, false},
|
||||
"VMULWEVVWU": {0xE132 << 15, false}, "VMULWEVQVU": {0xE133 << 15, false},
|
||||
"VMULWODHBU": {0xE134 << 15, false}, "VMULWODWHU": {0xE135 << 15, false},
|
||||
"VMULWODVWU": {0xE136 << 15, false}, "VMULWODQVU": {0xE137 << 15, false},
|
||||
"VMULWEVHBUB": {0xE140 << 15, false}, "VMULWEVWHUH": {0xE141 << 15, false},
|
||||
"VMULWEVVWUW": {0xE142 << 15, false}, "VMULWEVQVUV": {0xE143 << 15, false},
|
||||
"VMULWODHBUB": {0xE144 << 15, false}, "VMULWODWHUH": {0xE145 << 15, false},
|
||||
"VMULWODVWUW": {0xE146 << 15, false}, "VMULWODQVUV": {0xE147 << 15, false},
|
||||
"VMADDB": {0xE150 << 15, false}, "VMADDH": {0xE151 << 15, false},
|
||||
"VMADDW": {0xE152 << 15, false}, "VMADDV": {0xE153 << 15, false},
|
||||
"VMSUBB": {0xE154 << 15, false}, "VMSUBH": {0xE155 << 15, false},
|
||||
"VMSUBW": {0xE156 << 15, false}, "VMSUBV": {0xE157 << 15, false},
|
||||
"VMADDWEVHB": {0xE158 << 15, false}, "VMADDWEVWH": {0xE159 << 15, false},
|
||||
"VMADDWEVVW": {0xE15A << 15, false}, "VMADDWEVQV": {0xE15B << 15, false},
|
||||
"VMADDWODHB": {0xE15C << 15, false}, "VMADDWODWH": {0xE15D << 15, false},
|
||||
"VMADDWODVW": {0xE15E << 15, false}, "VMADDWODQV": {0xE15F << 15, false},
|
||||
"VMADDWEVHBU": {0xE168 << 15, false}, "VMADDWEVWHU": {0xE169 << 15, false},
|
||||
"VMADDWEVVWU": {0xE16A << 15, false}, "VMADDWEVQVU": {0xE16B << 15, false},
|
||||
"VMADDWODHBU": {0xE16C << 15, false}, "VMADDWODWHU": {0xE16D << 15, false},
|
||||
"VMADDWODVWU": {0xE16E << 15, false}, "VMADDWODQVU": {0xE16F << 15, false},
|
||||
"VMADDWEVHBUB": {0xE178 << 15, false}, "VMADDWEVWHUH": {0xE179 << 15, false},
|
||||
"VMADDWEVVWUW": {0xE17A << 15, false}, "VMADDWEVQVUV": {0xE17B << 15, false},
|
||||
"VMADDWODHBUB": {0xE17C << 15, false}, "VMADDWODWHUH": {0xE17D << 15, false},
|
||||
"VMADDWODVWUW": {0xE17E << 15, false}, "VMADDWODQVUV": {0xE17F << 15, false},
|
||||
"VDIVB": {0xE1C0 << 15, false}, "VDIVH": {0xE1C1 << 15, false},
|
||||
"VDIVW": {0xE1C2 << 15, false}, "VDIVV": {0xE1C3 << 15, false},
|
||||
"VMODB": {0xE1C4 << 15, false}, "VMODH": {0xE1C5 << 15, false},
|
||||
"VMODW": {0xE1C6 << 15, false}, "VMODV": {0xE1C7 << 15, false},
|
||||
"VDIVBU": {0xE1C8 << 15, false}, "VDIVHU": {0xE1C9 << 15, false},
|
||||
"VDIVWU": {0xE1CA << 15, false}, "VDIVVU": {0xE1CB << 15, false},
|
||||
"VMODBU": {0xE1CC << 15, false}, "VMODHU": {0xE1CD << 15, false},
|
||||
"VMODWU": {0xE1CE << 15, false}, "VMODVU": {0xE1CF << 15, false},
|
||||
"VMULF": {0xE271 << 15, false}, "VMULD": {0xE272 << 15, false},
|
||||
"VDIVF": {0xE275 << 15, false}, "VDIVD": {0xE276 << 15, false},
|
||||
"XVMULB": {0xE908 << 15, true}, "XVMULH": {0xE909 << 15, true},
|
||||
"XVMULW": {0xE90A << 15, true}, "XVMULV": {0xE90B << 15, true},
|
||||
"XVMUHB": {0xE90C << 15, true}, "XVMUHH": {0xE90D << 15, true},
|
||||
"XVMUHW": {0xE90E << 15, true}, "XVMUHV": {0xE90F << 15, true},
|
||||
"XVMUHBU": {0xE910 << 15, true}, "XVMUHHU": {0xE911 << 15, true},
|
||||
"XVMUHWU": {0xE912 << 15, true}, "XVMUHVU": {0xE913 << 15, true},
|
||||
"XVMULWEVHB": {0xE920 << 15, true}, "XVMULWEVWH": {0xE921 << 15, true},
|
||||
"XVMULWEVVW": {0xE922 << 15, true}, "XVMULWEVQV": {0xE923 << 15, true},
|
||||
"XVMULWODHB": {0xE924 << 15, true}, "XVMULWODWH": {0xE925 << 15, true},
|
||||
"XVMULWODVW": {0xE926 << 15, true}, "XVMULWODQV": {0xE927 << 15, true},
|
||||
"XVMULWEVHBU": {0xE930 << 15, true}, "XVMULWEVWHU": {0xE931 << 15, true},
|
||||
"XVMULWEVVWU": {0xE932 << 15, true}, "XVMULWEVQVU": {0xE933 << 15, true},
|
||||
"XVMULWODHBU": {0xE934 << 15, true}, "XVMULWODWHU": {0xE935 << 15, true},
|
||||
"XVMULWODVWU": {0xE936 << 15, true}, "XVMULWODQVU": {0xE937 << 15, true},
|
||||
"XVMULWEVHBUB": {0xE940 << 15, true}, "XVMULWEVWHUH": {0xE941 << 15, true},
|
||||
"XVMULWEVVWUW": {0xE942 << 15, true}, "XVMULWEVQVUV": {0xE943 << 15, true},
|
||||
"XVMULWODHBUB": {0xE944 << 15, true}, "XVMULWODWHUH": {0xE945 << 15, true},
|
||||
"XVMULWODVWUW": {0xE946 << 15, true}, "XVMULWODQVUV": {0xE947 << 15, true},
|
||||
"XVMADDB": {0xE950 << 15, true}, "XVMADDH": {0xE951 << 15, true},
|
||||
"XVMADDW": {0xE952 << 15, true}, "XVMADDV": {0xE953 << 15, true},
|
||||
"XVMSUBB": {0xE954 << 15, true}, "XVMSUBH": {0xE955 << 15, true},
|
||||
"XVMSUBW": {0xE956 << 15, true}, "XVMSUBV": {0xE957 << 15, true},
|
||||
"XVMADDWEVHB": {0xE958 << 15, true}, "XVMADDWEVWH": {0xE959 << 15, true},
|
||||
"XVMADDWEVVW": {0xE95A << 15, true}, "XVMADDWEVQV": {0xE95B << 15, true},
|
||||
"XVMADDWODHB": {0xE95C << 15, true}, "XVMADDWODWH": {0xE95D << 15, true},
|
||||
"XVMADDWODVW": {0xE95E << 15, true}, "XVMADDWODQV": {0xE95F << 15, true},
|
||||
"XVMADDWEVHBU": {0xE968 << 15, true}, "XVMADDWEVWHU": {0xE969 << 15, true},
|
||||
"XVMADDWEVVWU": {0xE96A << 15, true}, "XVMADDWEVQVU": {0xE96B << 15, true},
|
||||
"XVMADDWODHBU": {0xE96C << 15, true}, "XVMADDWODWHU": {0xE96D << 15, true},
|
||||
"XVMADDWODVWU": {0xE96E << 15, true}, "XVMADDWODQVU": {0xE96F << 15, true},
|
||||
"XVMADDWEVHBUB": {0xE978 << 15, true}, "XVMADDWEVWHUH": {0xE979 << 15, true},
|
||||
"XVMADDWEVVWUW": {0xE97A << 15, true}, "XVMADDWEVQVUV": {0xE97B << 15, true},
|
||||
"XVMADDWODHBUB": {0xE97C << 15, true}, "XVMADDWODWHUH": {0xE97D << 15, true},
|
||||
"XVMADDWODVWUW": {0xE97E << 15, true}, "XVMADDWODQVUV": {0xE97F << 15, true},
|
||||
"XVDIVB": {0xE9C0 << 15, true}, "XVDIVH": {0xE9C1 << 15, true},
|
||||
"XVDIVW": {0xE9C2 << 15, true}, "XVDIVV": {0xE9C3 << 15, true},
|
||||
"XVMODB": {0xE9C4 << 15, true}, "XVMODH": {0xE9C5 << 15, true},
|
||||
"XVMODW": {0xE9C6 << 15, true}, "XVMODV": {0xE9C7 << 15, true},
|
||||
"XVDIVBU": {0xE9C8 << 15, true}, "XVDIVHU": {0xE9C9 << 15, true},
|
||||
"XVDIVWU": {0xE9CA << 15, true}, "XVDIVVU": {0xE9CB << 15, true},
|
||||
"XVMODBU": {0xE9CC << 15, true}, "XVMODHU": {0xE9CD << 15, true},
|
||||
"XVMODWU": {0xE9CE << 15, true}, "XVMODVU": {0xE9CF << 15, true},
|
||||
"XVMULF": {0xEA71 << 15, true}, "XVMULD": {0xEA72 << 15, true},
|
||||
"XVDIVF": {0xEA75 << 15, true}, "XVDIVD": {0xEA76 << 15, true},
|
||||
}
|
||||
|
||||
// The lane-wise shifts and rotates (three-register forms; the immediate
|
||||
// forms live in l64VecImmInfo), the interleave families, the bit
|
||||
// clear/set/rev register forms, the remaining logic and compare
|
||||
// spellings, the widening add/subtract ladder and the vector FP
|
||||
// arithmetic.
|
||||
vecmisc := map[string]l64Vec3Enc{
|
||||
"VSLLB": {0xE1D0 << 15, false}, "VSLLH": {0xE1D1 << 15, false},
|
||||
"VSLLW": {0xE1D2 << 15, false}, "VSLLV": {0xE1D3 << 15, false},
|
||||
"VSRLB": {0xE1D4 << 15, false}, "VSRLH": {0xE1D5 << 15, false},
|
||||
"VSRLW": {0xE1D6 << 15, false}, "VSRLV": {0xE1D7 << 15, false},
|
||||
"VSRAH": {0xE1D9 << 15, false}, "VSRAW": {0xE1DA << 15, false},
|
||||
"VSRAV": {0xE1DB << 15, false},
|
||||
"VROTRB": {0xE1DC << 15, false}, "VROTRH": {0xE1DD << 15, false},
|
||||
"VROTRV": {0xE1DF << 15, false},
|
||||
"VILVLB": {0xE234 << 15, false}, "VILVLH": {0xE235 << 15, false},
|
||||
"VILVLW": {0xE236 << 15, false}, "VILVLV": {0xE237 << 15, false},
|
||||
"VILVHB": {0xE238 << 15, false}, "VILVHH": {0xE239 << 15, false},
|
||||
"VILVHW": {0xE23A << 15, false}, "VILVHV": {0xE23B << 15, false},
|
||||
"VBITCLRB": {0xE218 << 15, false}, "VBITCLRH": {0xE219 << 15, false},
|
||||
"VBITCLRW": {0xE21A << 15, false}, "VBITCLRV": {0xE21B << 15, false},
|
||||
"VBITSETB": {0xE21C << 15, false}, "VBITSETH": {0xE21D << 15, false},
|
||||
"VBITSETW": {0xE21E << 15, false}, "VBITSETV": {0xE21F << 15, false},
|
||||
"VBITREVB": {0xE220 << 15, false}, "VBITREVH": {0xE221 << 15, false},
|
||||
"VBITREVW": {0xE222 << 15, false}, "VBITREVV": {0xE223 << 15, false},
|
||||
"VORV": {0xE24D << 15, false}, "VNORV": {0xE24F << 15, false},
|
||||
"VANDNV": {0xE250 << 15, false}, "VORNV": {0xE251 << 15, false},
|
||||
"VSEQH": {0xE001 << 15, false}, "VSEQW": {0xE002 << 15, false},
|
||||
"VSLTB": {0xE00C << 15, false}, "VSLTH": {0xE00D << 15, false},
|
||||
"VSLTW": {0xE00E << 15, false}, "VSLTV": {0xE00F << 15, false},
|
||||
"VSLTBU": {0xE010 << 15, false}, "VSLTHU": {0xE011 << 15, false},
|
||||
"VSLTWU": {0xE012 << 15, false}, "VSLTVU": {0xE013 << 15, false},
|
||||
"VADDWEVHB": {0xE03C << 15, false}, "VADDWEVWH": {0xE03D << 15, false},
|
||||
"VADDWEVVW": {0xE03E << 15, false}, "VADDWEVQV": {0xE03F << 15, false},
|
||||
"VSUBWEVHB": {0xE040 << 15, false}, "VSUBWEVWH": {0xE041 << 15, false},
|
||||
"VSUBWEVVW": {0xE042 << 15, false}, "VSUBWEVQV": {0xE043 << 15, false},
|
||||
"VADDWODHB": {0xE044 << 15, false}, "VADDWODWH": {0xE045 << 15, false},
|
||||
"VADDWODVW": {0xE046 << 15, false}, "VADDWODQV": {0xE047 << 15, false},
|
||||
"VSUBWODHB": {0xE048 << 15, false}, "VSUBWODWH": {0xE049 << 15, false},
|
||||
"VSUBWODVW": {0xE04A << 15, false}, "VSUBWODQV": {0xE04B << 15, false},
|
||||
"VSUBWEVHBU": {0xE060 << 15, false}, "VSUBWEVWHU": {0xE061 << 15, false},
|
||||
"VSUBWEVVWU": {0xE062 << 15, false}, "VSUBWEVQVU": {0xE063 << 15, false},
|
||||
"VADDWEVHBU": {0xE05C << 15, false}, "VADDWEVWHU": {0xE05D << 15, false},
|
||||
"VADDWEVVWU": {0xE05E << 15, false}, "VADDWEVQVU": {0xE05F << 15, false},
|
||||
"VADDWODHBU": {0xE064 << 15, false}, "VADDWODWHU": {0xE065 << 15, false},
|
||||
"VADDWODVWU": {0xE066 << 15, false}, "VADDWODQVU": {0xE067 << 15, false},
|
||||
"VSUBWODHBU": {0xE068 << 15, false}, "VSUBWODWHU": {0xE069 << 15, false},
|
||||
"VSUBWODVWU": {0xE06A << 15, false}, "VSUBWODQVU": {0xE06B << 15, false},
|
||||
"VSHUFH": {0xE2F5 << 15, false}, "VSHUFW": {0xE2F6 << 15, false},
|
||||
"VSHUFV": {0xE2F7 << 15, false},
|
||||
"XVSLLB": {0xE9D0 << 15, true}, "XVSLLH": {0xE9D1 << 15, true},
|
||||
"XVSLLW": {0xE9D2 << 15, true}, "XVSLLV": {0xE9D3 << 15, true},
|
||||
"XVSRLB": {0xE9D4 << 15, true}, "XVSRLH": {0xE9D5 << 15, true},
|
||||
"XVSRLW": {0xE9D6 << 15, true}, "XVSRLV": {0xE9D7 << 15, true},
|
||||
"XVSRAB": {0xE9D8 << 15, true}, "XVSRAH": {0xE9D9 << 15, true},
|
||||
"XVSRAW": {0xE9DA << 15, true}, "XVSRAV": {0xE9DB << 15, true},
|
||||
"XVROTRB": {0xE9DC << 15, true}, "XVROTRH": {0xE9DD << 15, true},
|
||||
"XVROTRW": {0xE9DE << 15, true}, "XVROTRV": {0xE9DF << 15, true},
|
||||
"XVILVLB": {0xEA34 << 15, true}, "XVILVLH": {0xEA35 << 15, true},
|
||||
"XVILVLW": {0xEA36 << 15, true}, "XVILVLV": {0xEA37 << 15, true},
|
||||
"XVILVHB": {0xEA38 << 15, true}, "XVILVHH": {0xEA39 << 15, true},
|
||||
"XVILVHW": {0xEA3A << 15, true}, "XVILVHV": {0xEA3B << 15, true},
|
||||
"XVBITCLRB": {0xEA18 << 15, true}, "XVBITCLRH": {0xEA19 << 15, true},
|
||||
"XVBITCLRW": {0xEA1A << 15, true}, "XVBITCLRV": {0xEA1B << 15, true},
|
||||
"XVBITSETB": {0xEA1C << 15, true}, "XVBITSETH": {0xEA1D << 15, true},
|
||||
"XVBITSETW": {0xEA1E << 15, true}, "XVBITSETV": {0xEA1F << 15, true},
|
||||
"XVBITREVB": {0xEA20 << 15, true}, "XVBITREVH": {0xEA21 << 15, true},
|
||||
"XVBITREVW": {0xEA22 << 15, true}, "XVBITREVV": {0xEA23 << 15, true},
|
||||
"XVORV": {0xEA4D << 15, true}, "XVNORV": {0xEA4F << 15, true},
|
||||
"XVANDNV": {0xEA50 << 15, true}, "XVORNV": {0xEA51 << 15, true},
|
||||
"XVSEQH": {0xE801 << 15, true}, "XVSEQW": {0xE802 << 15, true},
|
||||
"XVSLTB": {0xE80C << 15, true}, "XVSLTH": {0xE80D << 15, true},
|
||||
"XVSLTW": {0xE80E << 15, true}, "XVSLTV": {0xE80F << 15, true},
|
||||
"XVSLTBU": {0xE810 << 15, true}, "XVSLTHU": {0xE811 << 15, true},
|
||||
"XVSLTWU": {0xE812 << 15, true}, "XVSLTVU": {0xE813 << 15, true},
|
||||
"XVADDWEVHB": {0xE83C << 15, true}, "XVADDWEVWH": {0xE83D << 15, true},
|
||||
"XVADDWEVVW": {0xE83E << 15, true}, "XVADDWEVQV": {0xE83F << 15, true},
|
||||
"XVSUBWEVHB": {0xE840 << 15, true}, "XVSUBWEVWH": {0xE841 << 15, true},
|
||||
"XVSUBWEVVW": {0xE842 << 15, true}, "XVSUBWEVQV": {0xE843 << 15, true},
|
||||
"XVADDWODHB": {0xE844 << 15, true}, "XVADDWODWH": {0xE845 << 15, true},
|
||||
"XVADDWODVW": {0xE846 << 15, true}, "XVADDWODQV": {0xE847 << 15, true},
|
||||
"XVSUBWODHB": {0xE848 << 15, true}, "XVSUBWODWH": {0xE849 << 15, true},
|
||||
"XVSUBWODVW": {0xE84A << 15, true}, "XVSUBWODQV": {0xE84B << 15, true},
|
||||
"XVADDWEVHBU": {0xE85C << 15, true}, "XVADDWEVWHU": {0xE85D << 15, true},
|
||||
"XVADDWEVVWU": {0xE85E << 15, true}, "XVADDWEVQVU": {0xE85F << 15, true},
|
||||
"XVSUBWEVHBU": {0xE860 << 15, true}, "XVSUBWEVWHU": {0xE861 << 15, true},
|
||||
"XVSUBWEVVWU": {0xE862 << 15, true}, "XVSUBWEVQVU": {0xE863 << 15, true},
|
||||
"XVADDWODHBU": {0xE864 << 15, true}, "XVADDWODWHU": {0xE865 << 15, true},
|
||||
"XVADDWODVWU": {0xE866 << 15, true}, "XVADDWODQVU": {0xE867 << 15, true},
|
||||
"XVSUBWODHBU": {0xE868 << 15, true}, "XVSUBWODWHU": {0xE869 << 15, true},
|
||||
"XVSUBWODVWU": {0xE86A << 15, true}, "XVSUBWODQVU": {0xE86B << 15, true},
|
||||
"XVSHUFH": {0xEAF5 << 15, true}, "XVSHUFW": {0xEAF6 << 15, true},
|
||||
"XVSHUFV": {0xEAF7 << 15, true},
|
||||
}
|
||||
for _, tab := range []map[string]l64Vec3Enc{addsub, muldiv, vecmisc} {
|
||||
for m, e := range tab {
|
||||
if _, dup := vec3[m]; dup {
|
||||
panic("loong64: duplicate vector mnemonic " + m)
|
||||
}
|
||||
vec3[m] = e
|
||||
}
|
||||
}
|
||||
for m, e := range vec3 {
|
||||
l64InstrTable[m] = l64Enc{format: l64Fvvv, op: e.op}
|
||||
l64VecBank[m] = e.lasx
|
||||
@@ -597,21 +850,156 @@ func init() {
|
||||
|
||||
// Immediate forms: INSTR $imm, vj, vd (or INSTR $imm, vd). The immediate
|
||||
// range, bias and field mask are the ones the toolchain encodes: vandi.b
|
||||
// stores the raw 8-bit constant, vsrai.b stores imm+8 (byte-lane bias),
|
||||
// vseqi.b and vseqi.d store 5-bit and 7-bit two's-complement values.
|
||||
// The mnemonics that also have a register form (VSEQB, VSEQV, VSRAB,
|
||||
// VROTRW) keep their three-register entry in l64InstrTable; the
|
||||
// dispatcher picks the immediate opcode from l64VecImmInfo by operand
|
||||
// kind, so the immediate entries must not overwrite the table.
|
||||
// stores the raw 8-bit constant, vsrari.b stores imm+8 (lane-width
|
||||
// bias), the si5 compares store 5-bit two's-complement values and vseqi.d
|
||||
// a 7-bit field the toolchain range-checks down to si5.
|
||||
// The mnemonics that also have a register form (the shifts, the bit
|
||||
// clear/set/rev families, VSEQ and the logic immediates) keep their
|
||||
// three-register entry in l64InstrTable; the dispatcher picks the
|
||||
// immediate opcode from l64VecImmInfo by operand kind, so the immediate
|
||||
// entries must not overwrite the table.
|
||||
vecImm := map[string]l64VecImmEnc{
|
||||
"VANDB": {0xE7A0 << 15, false, 0, 255, 0, 0xFF},
|
||||
"XVANDB": {0xEFA0 << 15, true, 0, 255, 0, 0xFF},
|
||||
"VORB": {0xE7A8 << 15, false, 0, 255, 0, 0xFF},
|
||||
"XVORB": {0xEFA8 << 15, true, 0, 255, 0, 0xFF},
|
||||
"VXORB": {0xE7B0 << 15, false, 0, 255, 0, 0xFF},
|
||||
"XVXORB": {0xEFB0 << 15, true, 0, 255, 0, 0xFF},
|
||||
"VNORB": {0xE7B8 << 15, false, 0, 255, 0, 0xFF},
|
||||
"XVNORB": {0xEFB8 << 15, true, 0, 255, 0, 0xFF},
|
||||
"VSEQB": {0xE500 << 15, false, -16, 15, 0, 0x1F},
|
||||
"XVSEQB": {0xE900 << 15, true, -16, 15, 0, 0x1F},
|
||||
"VSEQV": {0xE503 << 15, false, -64, 63, 0, 0x7F},
|
||||
"XVSEQV": {0xE903 << 15, true, -64, 63, 0, 0x7F},
|
||||
"VSRAB": {0xE668 << 15, false, 0, 7, 8, 0x1F},
|
||||
"VROTRW": {0xE541 << 15, false, 0, 31, 0, 0x1F},
|
||||
// vseqi.h/w accept the same si5 window as vseqi.b; vseqi.d carries a
|
||||
// 7-bit field, but the toolchain range-checks it down to si5 as well
|
||||
// (GOARCH=loong64 go tool asm rejects VSEQV $32 and VSEQV $-64).
|
||||
"VSEQH": {0xE501 << 15, false, -16, 15, 0, 0x1F},
|
||||
"XVSEQH": {0xED01 << 15, true, -16, 15, 0, 0x1F},
|
||||
"VSEQW": {0xE502 << 15, false, -16, 15, 0, 0x1F},
|
||||
"XVSEQW": {0xED02 << 15, true, -16, 15, 0, 0x1F},
|
||||
"VSEQV": {0xE503 << 15, false, -16, 15, 0, 0x7F},
|
||||
"XVSEQV": {0xE903 << 15, true, -16, 15, 0, 0x7F},
|
||||
// vslti compares against a signed (or, in the U spellings, unsigned)
|
||||
// si5/ui5 constant.
|
||||
"VSLTB": {0xE50C << 15, false, -16, 15, 0, 0x1F},
|
||||
"XVSLTB": {0xED0C << 15, true, -16, 15, 0, 0x1F},
|
||||
"VSLTH": {0xE50D << 15, false, -16, 15, 0, 0x1F},
|
||||
"XVSLTH": {0xED0D << 15, true, -16, 15, 0, 0x1F},
|
||||
"VSLTW": {0xE50E << 15, false, -16, 15, 0, 0x1F},
|
||||
"XVSLTW": {0xED0E << 15, true, -16, 15, 0, 0x1F},
|
||||
"VSLTV": {0xE50F << 15, false, -16, 15, 0, 0x1F},
|
||||
"XVSLTV": {0xED0F << 15, true, -16, 15, 0, 0x1F},
|
||||
"VSLTBU": {0xE510 << 15, false, 0, 31, 0, 0x1F},
|
||||
"XVSLTBU": {0xED10 << 15, true, 0, 31, 0, 0x1F},
|
||||
"VSLTHU": {0xE511 << 15, false, 0, 31, 0, 0x1F},
|
||||
"XVSLTHU": {0xED11 << 15, true, 0, 31, 0, 0x1F},
|
||||
"VSLTWU": {0xE512 << 15, false, 0, 31, 0, 0x1F},
|
||||
"XVSLTWU": {0xED12 << 15, true, 0, 31, 0, 0x1F},
|
||||
"VSLTVU": {0xE513 << 15, false, 0, 31, 0, 0x1F},
|
||||
"XVSLTVU": {0xED13 << 15, true, 0, 31, 0, 0x1F},
|
||||
// vaddi/vsubi take ui5 constants for every width on this toolchain
|
||||
// (VADDVU $32 is rejected by the oracle although the field is ui8).
|
||||
"VADDBU": {0xE514 << 15, false, 0, 31, 0, 0x1F},
|
||||
"XVADDBU": {0xED14 << 15, true, 0, 31, 0, 0x1F},
|
||||
"VADDHU": {0xE515 << 15, false, 0, 31, 0, 0x1F},
|
||||
"XVADDHU": {0xED15 << 15, true, 0, 31, 0, 0x1F},
|
||||
"VADDWU": {0xE516 << 15, false, 0, 31, 0, 0x1F},
|
||||
"XVADDWU": {0xED16 << 15, true, 0, 31, 0, 0x1F},
|
||||
"VADDVU": {0xE517 << 15, false, 0, 31, 0, 0x1F},
|
||||
"XVADDVU": {0xED17 << 15, true, 0, 31, 0, 0x1F},
|
||||
"VSUBBU": {0xE518 << 15, false, 0, 31, 0, 0x1F},
|
||||
"XVSUBBU": {0xED18 << 15, true, 0, 31, 0, 0x1F},
|
||||
"VSUBHU": {0xE519 << 15, false, 0, 31, 0, 0x1F},
|
||||
"XVSUBHU": {0xED19 << 15, true, 0, 31, 0, 0x1F},
|
||||
"VSUBWU": {0xE51A << 15, false, 0, 31, 0, 0x1F},
|
||||
"XVSUBWU": {0xED1A << 15, true, 0, 31, 0, 0x1F},
|
||||
"VSUBVU": {0xE51B << 15, false, 0, 31, 0, 0x1F},
|
||||
"XVSUBVU": {0xED1B << 15, true, 0, 31, 0, 0x1F},
|
||||
// The shift/rotate immediates ride in a width-sized field whose upper
|
||||
// bits carry the lane-width code: vslli.b stores ui3 at [12:0] with
|
||||
// bits [14:13] inside the opcode, vslli.h ui4 under a 4 bit mask, and
|
||||
// the .w/.d spellings a raw ui5/ui6.
|
||||
"VSLLB": {0x732C2000, false, 0, 7, 0, 0x7},
|
||||
"XVSLLB": {0x772C2000, true, 0, 7, 0, 0x7},
|
||||
"VSLLH": {0x732C4000, false, 0, 15, 0, 0xF},
|
||||
"XVSLLH": {0x772C4000, true, 0, 15, 0, 0xF},
|
||||
"VSLLW": {0xE659 << 15, false, 0, 31, 0, 0x1F},
|
||||
"XVSLLW": {0xEE59 << 15, true, 0, 31, 0, 0x1F},
|
||||
"VSLLV": {0xE65A << 15, false, 0, 63, 0, 0x3F},
|
||||
"XVSLLV": {0xEE5A << 15, true, 0, 63, 0, 0x3F},
|
||||
"VSRLB": {0x73302000, false, 0, 7, 0, 0x7},
|
||||
"XVSRLB": {0x77302000, true, 0, 7, 0, 0x7},
|
||||
"VSRLH": {0x73304000, false, 0, 15, 0, 0xF},
|
||||
"XVSRLH": {0x77304000, true, 0, 15, 0, 0xF},
|
||||
"VSRLW": {0xE661 << 15, false, 0, 31, 0, 0x1F},
|
||||
"XVSRLW": {0xEE61 << 15, true, 0, 31, 0, 0x1F},
|
||||
"VSRLV": {0xE662 << 15, false, 0, 63, 0, 0x3F},
|
||||
"XVSRLV": {0xEE62 << 15, true, 0, 63, 0, 0x3F},
|
||||
// vsrari/vrotri bias the field so the lane-width code rides above the
|
||||
// shift amount (.b adds 8, .h 16, .w 32; .d is a raw ui6).
|
||||
"VSRAB": {0xE668 << 15, false, 0, 7, 8, 0x1F},
|
||||
"XVSRAB": {0xEE68 << 15, true, 0, 7, 8, 0x1F},
|
||||
"VSRAH": {0x73344000, false, 0, 15, 0, 0xF},
|
||||
"XVSRAH": {0x77344000, true, 0, 15, 0, 0xF},
|
||||
"VSRAW": {0xE669 << 15, false, 0, 31, 0, 0x1F},
|
||||
"XVSRAW": {0xEE69 << 15, true, 0, 31, 0, 0x1F},
|
||||
"VSRAV": {0xE66A << 15, false, 0, 63, 0, 0x3F},
|
||||
"XVSRAV": {0xEE6A << 15, true, 0, 63, 0, 0x3F},
|
||||
"VROTRB": {0x72A02000, false, 0, 7, 0, 0x7},
|
||||
"XVROTRB": {0x76A02000, true, 0, 7, 0, 0x7},
|
||||
"VROTRH": {0x72A04000, false, 0, 15, 0, 0xF},
|
||||
"XVROTRH": {0x76A04000, true, 0, 15, 0, 0xF},
|
||||
"VROTRW": {0xE541 << 15, false, 0, 31, 0, 0x1F},
|
||||
"XVROTRW": {0xED41 << 15, true, 0, 31, 0, 0x1F},
|
||||
"VROTRV": {0xE542 << 15, false, 0, 63, 0, 0x3F},
|
||||
"XVROTRV": {0xED42 << 15, true, 0, 63, 0, 0x3F},
|
||||
// vbitclri/vbitseti/vbitrevi follow the same width-coded layout.
|
||||
"VBITCLRB": {0x73102000, false, 0, 7, 0, 0x7},
|
||||
"XVBITCLRB": {0x77102000, true, 0, 7, 0, 0x7},
|
||||
"VBITCLRH": {0x73104000, false, 0, 15, 0, 0xF},
|
||||
"XVBITCLRH": {0x77104000, true, 0, 15, 0, 0xF},
|
||||
"VBITCLRW": {0xE621 << 15, false, 0, 31, 0, 0x1F},
|
||||
"XVBITCLRW": {0xEE21 << 15, true, 0, 31, 0, 0x1F},
|
||||
"VBITCLRV": {0xE622 << 15, false, 0, 63, 0, 0x3F},
|
||||
"XVBITCLRV": {0xEE22 << 15, true, 0, 63, 0, 0x3F},
|
||||
"VBITSETB": {0x73142000, false, 0, 7, 0, 0x7},
|
||||
"XVBITSETB": {0x77142000, true, 0, 7, 0, 0x7},
|
||||
"VBITSETH": {0x73144000, false, 0, 15, 0, 0xF},
|
||||
"XVBITSETH": {0x77144000, true, 0, 15, 0, 0xF},
|
||||
"VBITSETW": {0xE629 << 15, false, 0, 31, 0, 0x1F},
|
||||
"XVBITSETW": {0xEE29 << 15, true, 0, 31, 0, 0x1F},
|
||||
"VBITSETV": {0xE62A << 15, false, 0, 63, 0, 0x3F},
|
||||
"XVBITSETV": {0xEE2A << 15, true, 0, 63, 0, 0x3F},
|
||||
"VBITREVB": {0x73182000, false, 0, 7, 0, 0x7},
|
||||
"XVBITREVB": {0x77182000, true, 0, 7, 0, 0x7},
|
||||
"VBITREVH": {0x73184000, false, 0, 15, 0, 0xF},
|
||||
"XVBITREVH": {0x77184000, true, 0, 15, 0, 0xF},
|
||||
"VBITREVW": {0xE631 << 15, false, 0, 31, 0, 0x1F},
|
||||
"XVBITREVW": {0xEE31 << 15, true, 0, 31, 0, 0x1F},
|
||||
"VBITREVV": {0xE632 << 15, false, 0, 63, 0, 0x3F},
|
||||
"XVBITREVV": {0xEE32 << 15, true, 0, 63, 0, 0x3F},
|
||||
// The 4-bit-select shuffles and the byte-extract/insert permutations
|
||||
// take ui8 (the .d shuffle ui4 range-checked to 0..15 by the
|
||||
// toolchain) packing both position nibbles.
|
||||
"VSHUF4IB": {0xE720 << 15, false, 0, 255, 0, 0xFF},
|
||||
"XVSHUF4IB": {0xEF20 << 15, true, 0, 255, 0, 0xFF},
|
||||
"VSHUF4IH": {0xE728 << 15, false, 0, 255, 0, 0xFF},
|
||||
"XVSHUF4IH": {0xEF28 << 15, true, 0, 255, 0, 0xFF},
|
||||
"VSHUF4IW": {0xE730 << 15, false, 0, 255, 0, 0xFF},
|
||||
"XVSHUF4IW": {0xEF30 << 15, true, 0, 255, 0, 0xFF},
|
||||
"VSHUF4IV": {0xE738 << 15, false, 0, 15, 0, 0xFF},
|
||||
"XVSHUF4IV": {0xEF38 << 15, true, 0, 15, 0, 0xFF},
|
||||
"VPERMIW": {0xE7C8 << 15, false, 0, 255, 0, 0xFF},
|
||||
"XVPERMIW": {0xEFC8 << 15, true, 0, 255, 0, 0xFF},
|
||||
"XVPERMIV": {0xEFD0 << 15, true, 0, 255, 0, 0xFF},
|
||||
"XVPERMIQ": {0xEFD8 << 15, true, 0, 255, 0, 0xFF},
|
||||
"VEXTRINSB": {0xE718 << 15, false, 0, 255, 0, 0xFF},
|
||||
"XVEXTRINSB": {0xEF18 << 15, true, 0, 255, 0, 0xFF},
|
||||
"VEXTRINSH": {0xE710 << 15, false, 0, 255, 0, 0xFF},
|
||||
"XVEXTRINSH": {0xEF10 << 15, true, 0, 255, 0, 0xFF},
|
||||
"VEXTRINSW": {0xE708 << 15, false, 0, 255, 0, 0xFF},
|
||||
"XVEXTRINSW": {0xEF08 << 15, true, 0, 255, 0, 0xFF},
|
||||
"VEXTRINSV": {0xE700 << 15, false, 0, 255, 0, 0xFF},
|
||||
"XVEXTRINSV": {0xEF00 << 15, true, 0, 255, 0, 0xFF},
|
||||
}
|
||||
for m, e := range vecImm {
|
||||
l64VecImmInfo[m] = e
|
||||
@@ -625,22 +1013,71 @@ func init() {
|
||||
"VSETANYEQB": 0xE539<<15 | 8<<10, "XVSETANYEQB": 0xED39<<15 | 8<<10,
|
||||
"VSETANYEQV": 0xE539<<15 | 11<<10, "XVSETANYEQV": 0xED39<<15 | 11<<10,
|
||||
"VSETALLNEV": 0xE539<<15 | 15<<10, "XVSETALLNEV": 0xED39<<15 | 15<<10,
|
||||
"VSETEQV": 0xE539<<15 | 6<<10, "XVSETEQV": 0xED39<<15 | 6<<10,
|
||||
"VSETANYEQH": 0xE539<<15 | 9<<10, "XVSETANYEQH": 0xED39<<15 | 9<<10,
|
||||
"VSETANYEQW": 0xE539<<15 | 10<<10, "XVSETANYEQW": 0xED39<<15 | 10<<10,
|
||||
"VSETALLNEB": 0xE539<<15 | 12<<10, "XVSETALLNEB": 0xED39<<15 | 12<<10,
|
||||
"VSETALLNEH": 0xE539<<15 | 13<<10, "XVSETALLNEH": 0xED39<<15 | 13<<10,
|
||||
"VSETALLNEW": 0xE539<<15 | 14<<10, "XVSETALLNEW": 0xED39<<15 | 14<<10,
|
||||
}
|
||||
for m, op := range vecCf {
|
||||
l64InstrTable[m] = l64Enc{format: l64Fvcf, op: op}
|
||||
l64VecBank[m] = strings.HasPrefix(m, "XV")
|
||||
}
|
||||
|
||||
// Lane popcount: INSTR vj, vd (the 2R layout with the opcode extending
|
||||
// over the unused vk field).
|
||||
// Lane popcount and the two-operand vector FP/unary spellings: INSTR vj,
|
||||
// vd (the 2R layout with the opcode extending over the unused vk field;
|
||||
// the low byte of each constant is the instruction's own sub-op).
|
||||
vec2r := map[string]l64Vec3Enc{
|
||||
"VPCNTV": {0x1CA70B << 10, false}, "XVPCNTV": {0x1DA70B << 10, true},
|
||||
}
|
||||
// The rest of the lane popcounts, the vector negations and the vector FP
|
||||
// unary conversions (loong64enc1.s).
|
||||
vec2rMore := map[string]l64Vec3Enc{
|
||||
"VPCNTB": {0x1CA708 << 10, false}, "VPCNTH": {0x1CA709 << 10, false},
|
||||
"VPCNTW": {0x1CA70A << 10, false},
|
||||
"VNEGB": {0x1CA70C << 10, false}, "VNEGH": {0x1CA70D << 10, false},
|
||||
"VNEGW": {0x1CA70E << 10, false}, "VNEGV": {0x1CA70F << 10, false},
|
||||
"VFCLASSF": {0x1CA735 << 10, false}, "VFCLASSD": {0x1CA736 << 10, false},
|
||||
"VFSQRTF": {0x1CA739 << 10, false}, "VFSQRTD": {0x1CA73A << 10, false},
|
||||
"VFRECIPF": {0x1CA73D << 10, false}, "VFRECIPD": {0x1CA73E << 10, false},
|
||||
"VFRSQRTF": {0x1CA741 << 10, false}, "VFRSQRTD": {0x1CA742 << 10, false},
|
||||
"VFRINTF": {0x1CA74D << 10, false}, "VFRINTD": {0x1CA74E << 10, false},
|
||||
"VFRINTRMF": {0x1CA751 << 10, false}, "VFRINTRMD": {0x1CA752 << 10, false},
|
||||
"VFRINTRPF": {0x1CA755 << 10, false}, "VFRINTRPD": {0x1CA756 << 10, false},
|
||||
"VFRINTRZF": {0x1CA759 << 10, false}, "VFRINTRZD": {0x1CA75A << 10, false},
|
||||
"VFRINTRNEF": {0x1CA75D << 10, false}, "VFRINTRNED": {0x1CA75E << 10, false},
|
||||
"XVPCNTB": {0x1DA708 << 10, true}, "XVPCNTH": {0x1DA709 << 10, true},
|
||||
"XVPCNTW": {0x1DA70A << 10, true},
|
||||
"XVNEGB": {0x1DA70C << 10, true}, "XVNEGH": {0x1DA70D << 10, true},
|
||||
"XVNEGW": {0x1DA70E << 10, true}, "XVNEGV": {0x1DA70F << 10, true},
|
||||
"XVFCLASSF": {0x1DA735 << 10, true}, "XVFCLASSD": {0x1DA736 << 10, true},
|
||||
"XVFSQRTF": {0x1DA739 << 10, true}, "XVFSQRTD": {0x1DA73A << 10, true},
|
||||
"XVFRECIPF": {0x1DA73D << 10, true}, "XVFRECIPD": {0x1DA73E << 10, true},
|
||||
"XVFRSQRTF": {0x1DA741 << 10, true}, "XVFRSQRTD": {0x1DA742 << 10, true},
|
||||
"XVFRINTF": {0x1DA74D << 10, true}, "XVFRINTD": {0x1DA74E << 10, true},
|
||||
"XVFRINTRMF": {0x1DA751 << 10, true}, "XVFRINTRMD": {0x1DA752 << 10, true},
|
||||
"XVFRINTRPF": {0x1DA755 << 10, true}, "XVFRINTRPD": {0x1DA756 << 10, true},
|
||||
"XVFRINTRZF": {0x1DA759 << 10, true}, "XVFRINTRZD": {0x1DA75A << 10, true},
|
||||
"XVFRINTRNEF": {0x1DA75D << 10, true}, "XVFRINTRNED": {0x1DA75E << 10, true},
|
||||
}
|
||||
maps.Copy(vec2r, vec2rMore)
|
||||
for m, e := range vec2r {
|
||||
l64InstrTable[m] = l64Enc{format: l64Frr, op: e.op}
|
||||
l64VecBank[m] = e.lasx
|
||||
l64Vec2R[m] = true
|
||||
}
|
||||
|
||||
// The four-register byte shuffle: INSTR va, vk, vj, vd (the operand the
|
||||
// table reads in each field position, va at bits [19:15]).
|
||||
vec4r := map[string]l64Vec3Enc{
|
||||
"VSHUFB": {0x0D50 << 16, false}, "XVSHUFB": {0x0D60 << 16, true},
|
||||
}
|
||||
for m, e := range vec4r {
|
||||
l64InstrTable[m] = l64Enc{format: l64Fvvvv, op: e.op}
|
||||
l64VecBank[m] = e.lasx
|
||||
l64Vec4R[m] = true
|
||||
}
|
||||
}
|
||||
|
||||
// l64FpMovTable maps (mnemonic, from-class, to-class) to the 2R opcode of the
|
||||
|
||||
@@ -507,6 +507,218 @@ TEXT ·v(SB), NOSPLIT, $0
|
||||
0x4C000020,
|
||||
)
|
||||
})
|
||||
|
||||
// The integer and FP add/subtract families with their saturating pairs
|
||||
// and immediate spellings (loong64enc1.s words).
|
||||
t.Run("add and subtract families", func(t *testing.T) {
|
||||
fn := firstTextLOONG64(t, `#include "textflag.h"
|
||||
TEXT ·v(SB), NOSPLIT, $0
|
||||
VADDB V1, V2, V3
|
||||
VADDF V1, V2, V3
|
||||
VADDD V1, V2, V3
|
||||
VSUBD V1, V2, V3
|
||||
VSADDV V1, V2, V3
|
||||
VSSUBVU V1, V2, V3
|
||||
VADDBU $1, V2, V1
|
||||
VADDBU $1, V2
|
||||
VSUBVU $31, V2
|
||||
XVSADDV X3, X2, X1
|
||||
XVSUBD X1, X2, X3
|
||||
RET
|
||||
`)
|
||||
code := assembleLOONG64Helper(t, fn)
|
||||
wantWords(t, code,
|
||||
0x700A0443, // vadd.b
|
||||
0x71308443, // vadd.f
|
||||
0x71310443, // vadd.d
|
||||
0x71330443, // vsub.d
|
||||
0x70478443, // vsadd.v
|
||||
0x704D8443, // vssub.u.d
|
||||
0x728A0441, // vaddi.bu v1, v2, 1
|
||||
0x728A0442, // vaddi.bu v2, v2, 1 (two-operand form)
|
||||
0x728DFC42, // vsubi.du v2, v2, 31 (two-operand form)
|
||||
0x74478C41, // xvsadd.d x1, x2, x3
|
||||
0x75330443, // xvsub.d x3, x2, x1
|
||||
0x4C000020,
|
||||
)
|
||||
})
|
||||
|
||||
// The multiply, divide and accumulate families.
|
||||
t.Run("multiply and divide families", func(t *testing.T) {
|
||||
fn := firstTextLOONG64(t, `#include "textflag.h"
|
||||
TEXT ·v(SB), NOSPLIT, $0
|
||||
VMULV V1, V2, V3
|
||||
VMUHHU V1, V2, V3
|
||||
VDIVBU V1, V2, V3
|
||||
VMODV V1, V2, V3
|
||||
VMADDB V1, V2, V3
|
||||
VMSUBV V1, V2, V3
|
||||
VMULWEVHB V1, V2, V3
|
||||
VMULWODQV V1, V2, V3
|
||||
VMADDWEVHBUB V1, V2, V3
|
||||
XVDIVD X1, X2, X3
|
||||
RET
|
||||
`)
|
||||
code := assembleLOONG64Helper(t, fn)
|
||||
wantWords(t, code,
|
||||
0x70858443, // vmul.v
|
||||
0x70888443, // vmuh.u.d
|
||||
0x70E40443, // vdiv.u.b
|
||||
0x70E38443, // vmod.d
|
||||
0x70A80443, // vmadd.b
|
||||
0x70AB8443, // vmsub.d
|
||||
0x70900443, // vmulwev.h.b
|
||||
0x70938443, // vmulwod.q.d
|
||||
0x70BC0443, // vmaddwev.h.bu.b
|
||||
0x753B0443, // xvdiv.d
|
||||
0x4C000020,
|
||||
)
|
||||
})
|
||||
|
||||
// The shift, bit and interleave families in register and immediate
|
||||
// spellings, with the width-coded shift immediates.
|
||||
t.Run("shift, bit and interleave families", func(t *testing.T) {
|
||||
fn := firstTextLOONG64(t, `#include "textflag.h"
|
||||
TEXT ·v(SB), NOSPLIT, $0
|
||||
VSLLV V1, V2, V3
|
||||
VROTRB V1, V2, V3
|
||||
VBITCLRV V1, V2, V3
|
||||
VBITSETW V1, V2, V3
|
||||
VBITREVV V1, V2, V3
|
||||
VILVLB V1, V2, V3
|
||||
VILVHV V1, V2, V3
|
||||
VSLLB $7, V1, V2
|
||||
VSLLB $5, V1
|
||||
VSRLH $15, V1, V2
|
||||
VSRAW $31, V1, V2
|
||||
VSRAV $63, V1, V2
|
||||
VROTRV $63, V1, V2
|
||||
VBITCLRB $7, V2, V3
|
||||
VBITREVV $63, V2, V3
|
||||
VSEQH $-16, V2, V3
|
||||
VSLTB $1, V2, V3
|
||||
VSLTHU $31, V2, V3
|
||||
XVILVLV X3, X2, X1
|
||||
XVSLLB $7, X2, X1
|
||||
XVSRAV $63, X2, X1
|
||||
XVBITREVV $63, X2, X1
|
||||
RET
|
||||
`)
|
||||
code := assembleLOONG64Helper(t, fn)
|
||||
wantWords(t, code,
|
||||
0x70E98443, // vsll.d
|
||||
0x70EE0443, // vrotr.b
|
||||
0x710D8443, // vbitclr.d
|
||||
0x710F0443, // vbitset.w
|
||||
0x71118443, // vbitrev.d
|
||||
0x711A0443, // vilvl.b
|
||||
0x711D8443, // vilvh.d
|
||||
0x732C3C22, // vslli.b v2, v1, 7
|
||||
0x732C3421, // vslli.b v1, v1, 5 (two-operand form)
|
||||
0x73307C22, // vsrli.h v2, v1, 15
|
||||
0x7334FC22, // vsrai.w v2, v1, 31
|
||||
0x7335FC22, // vsrai.d v2, v1, 63
|
||||
0x72A1FC22, // vrotri.d v2, v1, 63
|
||||
0x73103C43, // vbitclri.b v3, v2, 7
|
||||
0x7319FC43, // vbitrevi.d v3, v2, 63
|
||||
0x7280C043, // vseqi.h v3, v2, -16
|
||||
0x72860443, // vslti.b v3, v2, 1
|
||||
0x7288FC43, // vslti.hu v3, v2, 31
|
||||
0x751B8C41, // xvilvl.d x1, x2, x3
|
||||
0x772C3C41, // xvslli.b x1, x2, 7
|
||||
0x7735FC41, // xvsrai.d x1, x2, 63
|
||||
0x7719FC41, // xvbitrevi.d x1, x2, 63
|
||||
0x4C000020,
|
||||
)
|
||||
})
|
||||
|
||||
// The shuffle, select and permutation families, including the
|
||||
// four-register byte shuffle.
|
||||
t.Run("shuffle and permutation families", func(t *testing.T) {
|
||||
fn := firstTextLOONG64(t, `#include "textflag.h"
|
||||
TEXT ·v(SB), NOSPLIT, $0
|
||||
VSHUFH V1, V2, V3
|
||||
VSHUFW V1, V2, V3
|
||||
VSHUFV V1, V2, V3
|
||||
VSHUFB V1, V2, V3, V4
|
||||
XVSHUFB X1, X2, X3, X4
|
||||
VSHUF4IB $255, V2, V1
|
||||
VSHUF4IV $15, V2, V1
|
||||
XVSHUF4IV $15, X1, X2
|
||||
VEXTRINSB $0x18, V1, V2
|
||||
XVEXTRINSV $0x81, X1, X2
|
||||
VPERMIW $0x1B, V1, V2
|
||||
XVPERMIQ $0x4B, X1, X2
|
||||
RET
|
||||
`)
|
||||
code := assembleLOONG64Helper(t, fn)
|
||||
wantWords(t, code,
|
||||
0x717A8443, // vshuf.h
|
||||
0x717B0443, // vshuf.w
|
||||
0x717B8443, // vshuf.d
|
||||
0x0D508864, // vshuf.b v4, v3, v2, v1
|
||||
0x0D608864, // xvshuf.b
|
||||
0x7393FC41, // vshuf4i.b v1, v2, 255
|
||||
0x739C3C41, // vshuf4i.d v1, v2, 15
|
||||
0x779C3C22, // xvshuf4i.d x2, x1, 15
|
||||
0x738C6022, // vextrins.b v2, v1, 0x18
|
||||
0x77820422, // xvextrins.d x2, x1, 0x81
|
||||
0x73E46C22, // vpermi.w v2, v1, 0x1b
|
||||
0x77ED2C22, // xvpermi.q x2, x1, 0x4b
|
||||
0x4C000020,
|
||||
)
|
||||
})
|
||||
|
||||
// The vector FP families, the unary spellings, the compare-to-flag
|
||||
// additions and the scalar int/float conversions.
|
||||
t.Run("FP and conversion families", func(t *testing.T) {
|
||||
fn := firstTextLOONG64(t, `#include "textflag.h"
|
||||
TEXT ·v(SB), NOSPLIT, $0
|
||||
VADDF V1, V2, V3
|
||||
VMULF V1, V2, V3
|
||||
VFCLASSD V1, V2
|
||||
VFSQRTF V1, V2
|
||||
VFRECIPD V1, V2
|
||||
VFRSQRTF V1, V2
|
||||
VFRINTF V1, V2
|
||||
VFRINTRNED V1, V2
|
||||
VNEGB V1, V2
|
||||
VPCNTB V1, V2
|
||||
XVNEGV X2, X1
|
||||
XVPCNTW X3, X2
|
||||
XVFRINTRNEF X1, X2
|
||||
VSETEQV V1, FCC0
|
||||
VSETANYEQH V1, FCC0
|
||||
VSETALLNEB V1, FCC0
|
||||
XVSETALLNEW X1, FCC0
|
||||
FFINTFW F0, F1
|
||||
FTINTVD F0, F1
|
||||
RET
|
||||
`)
|
||||
code := assembleLOONG64Helper(t, fn)
|
||||
wantWords(t, code,
|
||||
0x71308443, // vfadd.s
|
||||
0x71388443, // vfmul.s
|
||||
0x729CD822, // vfclass.d
|
||||
0x729CE422, // vfsqrt.s
|
||||
0x729CF822, // vfrecip.d
|
||||
0x729D0422, // vfrsqrt.s
|
||||
0x729D3422, // vfrint.s
|
||||
0x729D7822, // vfrintne.s
|
||||
0x729C3022, // vneg.b
|
||||
0x729C2022, // vpcnt.b
|
||||
0x769C3C41, // xvneg.d x1, x2
|
||||
0x769C2862, // xvpcnt.w x2, x3
|
||||
0x769D7422, // xvfrintne.s x2, x1
|
||||
0x729C9820, // vseteqz.d fcc0, v1
|
||||
0x729CA420, // vsetanyeqz.h
|
||||
0x729CB020, // vsetallnez.b
|
||||
0x769CB820, // xvsetallnez.w
|
||||
0x011D1001, // ffint.s.w f1, f0
|
||||
0x011B2801, // ftint.l.d f1, f0
|
||||
0x4C000020,
|
||||
)
|
||||
})
|
||||
}
|
||||
|
||||
// TestLOONG64_vectorErrors pins the register-class and range diagnostics of
|
||||
@@ -545,6 +757,36 @@ func TestLOONG64_vectorErrors(t *testing.T) {
|
||||
`TEXT ·e(SB), NOSPLIT, $0
|
||||
VROTRW $32, V1, V2
|
||||
RET
|
||||
`,
|
||||
`TEXT ·e(SB), NOSPLIT, $0
|
||||
VADDVU $32, V2
|
||||
RET
|
||||
`,
|
||||
`TEXT ·e(SB), NOSPLIT, $0
|
||||
VSEQV $32, V2, V3
|
||||
RET
|
||||
`,
|
||||
`TEXT ·e(SB), NOSPLIT, $0
|
||||
VSHUF4IV $16, V2, V1
|
||||
RET
|
||||
`,
|
||||
`TEXT ·e(SB), NOSPLIT, $0
|
||||
VEXTRINSB $256, V1, V2
|
||||
RET
|
||||
`,
|
||||
`TEXT ·e(SB), NOSPLIT, $0
|
||||
VSLTV $-17, V2, V3
|
||||
RET
|
||||
`,
|
||||
// VSHUFB wants four vector registers.
|
||||
`TEXT ·e(SB), NOSPLIT, $0
|
||||
VSHUFB V1, V2, V3
|
||||
RET
|
||||
`,
|
||||
// The FCC forms still refuse vector registers.
|
||||
`TEXT ·e(SB), NOSPLIT, $0
|
||||
VSETEQV V1, V2
|
||||
RET
|
||||
`,
|
||||
// VSET* wants an FCC flag, not a vector register.
|
||||
`TEXT ·e(SB), NOSPLIT, $0
|
||||
|
||||
@@ -14,6 +14,51 @@ type Imm int64
|
||||
|
||||
func (Imm) isOperand() {}
|
||||
|
||||
// RegList is a bracketed register range, [Z0-Z3]: the four-register source
|
||||
// of the 4FMAPS and 4VNNIW families. The EVEX emit path carries the list's
|
||||
// low register through the inverted 5-bit V'VVVV field; the three higher
|
||||
// registers are implied by the instruction, so only the pair travels here.
|
||||
type RegList struct {
|
||||
Lo Reg
|
||||
Hi Reg // implied by the encoding; Lo.idx+3 by construction
|
||||
}
|
||||
|
||||
func (RegList) isOperand() {}
|
||||
|
||||
// FloatImm is a floating-point immediate ($-1.0). The SSE mnemonics whose
|
||||
// encoding takes an XMM/memory source at that position rewrite it as a read
|
||||
// from a read-only pool constant ($f64.<hex> or $f32.<hex>), the toolchain's
|
||||
// own behaviour; every other instruction rejects it.
|
||||
type FloatImm struct {
|
||||
Text string // the numeric text as written, sign excluded
|
||||
Neg bool // a leading minus
|
||||
}
|
||||
|
||||
func (FloatImm) isOperand() {}
|
||||
|
||||
// TLSMem is a thread-local access, the source form off(base)(TLS*1) with the
|
||||
// base dropped: the toolchain's one-instruction TLS rewrite assembles it as
|
||||
// the segment-prefixed absolute whose disp32 carries an R_TLS_LE patch site
|
||||
// (the linker fills the TLS slot offset).
|
||||
type TLSMem struct {
|
||||
Disp int64
|
||||
Size int
|
||||
Seg byte // the segment override: FS (0x64) or GS (0x65) on windows
|
||||
}
|
||||
|
||||
func (TLSMem) isOperand() {}
|
||||
|
||||
// SegAbs is a segment-absolute access, 0x30(GS): the segment override
|
||||
// prefixes a disp32 absolute reference with no relocation. The base
|
||||
// register spellings GS and FS produce it.
|
||||
type SegAbs struct {
|
||||
Disp int64
|
||||
Size int
|
||||
Seg byte // 0x64 FS, 0x65 GS
|
||||
}
|
||||
|
||||
func (SegAbs) isOperand() {}
|
||||
|
||||
// Mem is a memory operand of the form disp(base)(index*scale).
|
||||
type Mem struct {
|
||||
Base Reg
|
||||
@@ -23,6 +68,7 @@ type Mem struct {
|
||||
Size int // operand width in bytes
|
||||
HasBase bool
|
||||
HasIndex bool
|
||||
Seg byte // segment override prefix (0x64 FS, 0x65 GS); 0 = none
|
||||
}
|
||||
|
||||
func (Mem) isOperand() {}
|
||||
@@ -48,3 +94,16 @@ type sbMem struct {
|
||||
}
|
||||
|
||||
func (sbMem) isOperand() {}
|
||||
|
||||
// isX86Mem reports whether the operand is an amd64 memory reference: a base
|
||||
// or indexed Mem, or an SB-relative sbMem. Encoders that gate on "memory in
|
||||
// this position" must accept both; the r/m emitters distinguish the two
|
||||
// themselves.
|
||||
func isX86Mem(o Operand) bool {
|
||||
switch o.(type) {
|
||||
case Mem, sbMem:
|
||||
return true
|
||||
default:
|
||||
return false
|
||||
}
|
||||
}
|
||||
|
||||
+828
-58
File diff suppressed because it is too large
Load Diff
+27
-5
@@ -63,7 +63,7 @@ func riscvRegNum(name string) int {
|
||||
return 24
|
||||
case "X25", "S9":
|
||||
return 25
|
||||
case "X26", "S10":
|
||||
case "X26", "S10", "CTXT":
|
||||
return 26
|
||||
case "X27", "S11", "g":
|
||||
return 27
|
||||
@@ -219,6 +219,9 @@ var riscvInstrTable = map[string]riscvEnc{
|
||||
"DIVUW": {0x3B, 0x5, 0x01},
|
||||
"REMW": {0x3B, 0x6, 0x01},
|
||||
"REMUW": {0x3B, 0x7, 0x01},
|
||||
// Zicond conditional zeroing.
|
||||
"CZEROEQZ": {0x33, 0x5, 0x07},
|
||||
"CZERONEZ": {0x33, 0x7, 0x07},
|
||||
// RV64I, I-type arithmetic.
|
||||
"ADDI": {0x13, 0x0, 0x00},
|
||||
"ADDIW": {0x1B, 0x0, 0x00},
|
||||
@@ -247,13 +250,21 @@ var riscvInstrTable = map[string]riscvEnc{
|
||||
"BGE": {0x63, 0x5, 0x00},
|
||||
"BLTU": {0x63, 0x6, 0x00},
|
||||
"BGEU": {0x63, 0x7, 0x00},
|
||||
// The swapped-spelling comparison forms: encoded as BLT/BGE/BLTU/BGEU
|
||||
// with the register operands swapped.
|
||||
"BGT": {0x63, 0x4, 0x00},
|
||||
"BLE": {0x63, 0x5, 0x00},
|
||||
"BGTU": {0x63, 0x6, 0x00},
|
||||
"BLEU": {0x63, 0x7, 0x00},
|
||||
// U-type.
|
||||
"LUI": {0x37, 0x0, 0x00},
|
||||
"AUIPC": {0x17, 0x0, 0x00},
|
||||
// System.
|
||||
"ECALL": {0x73, 0x0, 0x00},
|
||||
"EBREAK": {0x73, 0x0, 0x00},
|
||||
"FENCE": {0x0F, 0x0, 0x00},
|
||||
"ECALL": {0x73, 0x0, 0x00},
|
||||
"EBREAK": {0x73, 0x0, 0x00},
|
||||
"FENCE": {0x0F, 0x0, 0x00},
|
||||
"FENCE.TSO": {0x0F, 0x0, 0x00},
|
||||
"PAUSE": {0x0F, 0x0, 0x00},
|
||||
// JALR, indirect jump/call (I-type).
|
||||
"JALR": {0x67, 0x0, 0x00},
|
||||
|
||||
@@ -303,7 +314,14 @@ var riscvInstrTable = map[string]riscvEnc{
|
||||
"FMIND": {0x53, 0x0, 0x15},
|
||||
"FMAXD": {0x53, 0x1, 0x15},
|
||||
// FP sign injection (double): rs2 carries the sign source.
|
||||
"FSGNJD": {0x53, 0x0, 0x11},
|
||||
"FSGNJD": {0x53, 0x0, 0x11},
|
||||
"FSGNJS": {0x53, 0x0, 0x10},
|
||||
"FSGNJX": {0x53, 0x0, 0x14},
|
||||
"FSGNJXD": {0x53, 0x0, 0x15},
|
||||
"FSGNJXS": {0x53, 0x0, 0x14},
|
||||
"FSGNJND": {0x53, 0x1, 0x11},
|
||||
"FSGNJNS": {0x53, 0x1, 0x10},
|
||||
"FSGNJNX": {0x53, 0x1, 0x14},
|
||||
|
||||
// RV64A, load-reserved / store-conditional (funct5 0x02 / 0x03).
|
||||
// The toolchain gives LR acquire ordering (aq = 1) and SC release
|
||||
@@ -375,6 +393,10 @@ var riscvCvtTable = map[string]riscvCvtEnc{
|
||||
"FMVDX": {0x79, 0x0, 0x53}, // int64 → float64 (bit move)
|
||||
"FMVXW": {0x70, 0x0, 0x53}, // float32 → int32 (bit move)
|
||||
"FMVWX": {0x78, 0x0, 0x53}, // int32 → float32 (bit move)
|
||||
// The toolchain's W/D suffix spellings of the same moves.
|
||||
"FMVXS": {0x70, 0x0, 0x53},
|
||||
"FMVFS": {0x78, 0x0, 0x53},
|
||||
"FMVSX": {0x79, 0x0, 0x53},
|
||||
}
|
||||
|
||||
// riscvCvtType encodes an FP conversion instruction.
|
||||
|
||||
+169
-23
@@ -33,7 +33,7 @@ func firstTextRISCV(t *testing.T, src string) *ast.Text {
|
||||
// assembleRISCVHelper assembles one TEXT function and returns its code bytes.
|
||||
func assembleRISCVHelper(t *testing.T, fn *ast.Text) []byte {
|
||||
t.Helper()
|
||||
code, _, _, _, _, err := assembleRISCV(fn)
|
||||
code, _, _, _, _, _, err := assembleRISCV(fn)
|
||||
if err != nil {
|
||||
t.Fatalf("assemble: %v", err)
|
||||
}
|
||||
@@ -785,16 +785,140 @@ TEXT ·sys(SB), NOSPLIT, $0
|
||||
}
|
||||
}
|
||||
|
||||
func TestRISCV_MOV_sym_FP_error(t *testing.T) {
|
||||
// MOV $sym(FP), rd should return an error (unsupported).
|
||||
func TestRISCV_MOV_sym_FP(t *testing.T) {
|
||||
// MOV $sym(FP), rd lowers to the frame-adjusted ADDI against SP: the
|
||||
// toolchain's argframe spelling. A zero frame leaves the offset at the
|
||||
// 8-byte link slot, compressed to C.ADDI4SPN.
|
||||
fn := firstTextRISCV(t, `#include "textflag.h"
|
||||
TEXT ·badfp(SB), NOSPLIT, $0
|
||||
TEXT ·argfp(SB), NOSPLIT, $0
|
||||
MOV $arg(FP), X10
|
||||
RET
|
||||
`)
|
||||
_, _, _, _, _, err := assembleRISCV(fn)
|
||||
if err == nil {
|
||||
t.Error("expected error for MOV $arg(FP), got nil")
|
||||
code, _, _, _, _, _, err := assembleRISCV(fn)
|
||||
if err != nil {
|
||||
t.Fatalf("assemble: %v", err)
|
||||
}
|
||||
// prologue (0: leaf, zero frame) + C.ADDI4SPN (2) + RET (4) = 6
|
||||
want := []byte{0x28, 0x00, 0x67, 0x80, 0x00, 0x00}
|
||||
if string(code) != string(want) {
|
||||
t.Errorf("got % x, want % x", code, want)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRISCV_Bookkeeping(t *testing.T) {
|
||||
// FUNCDATA and PCDATA contribute no bytes; UNDEF is the toolchain's
|
||||
// ebreak, compressed to C.EBREAK under RVC.
|
||||
fn := firstTextRISCV(t, `#include "textflag.h"
|
||||
TEXT ·book(SB), NOSPLIT, $0-8
|
||||
FUNCDATA $0, marks<>(SB)
|
||||
PCDATA $1, $1
|
||||
UNDEF
|
||||
MOV $1, X10
|
||||
MOV X10, ret+0(FP)
|
||||
RET
|
||||
`)
|
||||
code, _, _, _, _, _, err := assembleRISCV(fn)
|
||||
if err != nil {
|
||||
t.Fatalf("assemble: %v", err)
|
||||
}
|
||||
// C.EBREAK (2) + C.LI X10, 1 (2) + C.SWSP (2) + RET (4) = 10: the
|
||||
// FUNCDATA and PCDATA statements contribute nothing.
|
||||
want := []byte{0x02, 0x90, 0x05, 0x45, 0x2a, 0xe4, 0x67, 0x80, 0x00, 0x00}
|
||||
if string(code) != string(want) {
|
||||
t.Errorf("got % x, want % x", code, want)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRISCV_JMPPCRel(t *testing.T) {
|
||||
// JMP N(PC): the displacement tracks the instruction N source slots
|
||||
// away in the final layout (0 the jump itself, negative backwards).
|
||||
fn := firstTextRISCV(t, `#include "textflag.h"
|
||||
TEXT ·slots(SB), NOSPLIT, $0-0
|
||||
JMP 2(PC)
|
||||
MOV $1, X11
|
||||
MOV $2, X12
|
||||
MOV X12, X11
|
||||
JMP -3(PC)
|
||||
RET
|
||||
`)
|
||||
code, _, _, _, _, _, err := assembleRISCV(fn)
|
||||
if err != nil {
|
||||
t.Fatalf("assemble: %v", err)
|
||||
}
|
||||
// JMP 2(PC) lands on the C.MV six bytes ahead; JMP -3(PC) lands back on
|
||||
// the first C.LI, six bytes behind.
|
||||
want := []byte{
|
||||
0x6f, 0x00, 0x60, 0x00, // JAL X0, 6
|
||||
0x85, 0x45, // C.LI X11, 1
|
||||
0x09, 0x46, // C.LI X12, 2
|
||||
0xb2, 0x85, // C.MV X11, X12
|
||||
0x6f, 0xf0, 0xbf, 0xff, // JAL X0, -6
|
||||
0x67, 0x80, 0x00, 0x00, // RET
|
||||
}
|
||||
if string(code) != string(want) {
|
||||
t.Errorf("got % x, want % x", code, want)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRISCV_MOVWideImm(t *testing.T) {
|
||||
// Shift-sequence constants compress like the toolchain's expansion.
|
||||
fn := firstTextRISCV(t, `#include "textflag.h"
|
||||
TEXT ·wide(SB), NOSPLIT, $0-0
|
||||
MOV $0x8000000000000000, X5
|
||||
MOV $0x100000000, X5
|
||||
MOV $0x000fffffffffffda, X5
|
||||
RET
|
||||
`)
|
||||
code, _, _, _, _, _, err := assembleRISCV(fn)
|
||||
if err != nil {
|
||||
t.Fatalf("assemble: %v", err)
|
||||
}
|
||||
// C.LI -1, C.SLLI 63; C.LI 1, C.SLLI 32; C.LI -19, C.SLLI 13, SRLI 12.
|
||||
want := []byte{
|
||||
0xfd, 0x52, 0xfe, 0x12,
|
||||
0x85, 0x42, 0x82, 0x12,
|
||||
0xb5, 0x52, 0xb6, 0x02, 0x93, 0xd2, 0xc2, 0x00,
|
||||
0x67, 0x80, 0x00, 0x00,
|
||||
}
|
||||
if string(code) != string(want) {
|
||||
t.Errorf("got % x, want % x", code, want)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRISCV_MOVImmPool(t *testing.T) {
|
||||
// A constant outside the shift shapes loads from the pooled $i64 data
|
||||
// symbol via AUIPC+LD, named like the toolchain's pool.
|
||||
src := `#include "textflag.h"
|
||||
TEXT ·pool(SB), NOSPLIT, $0-8
|
||||
MOV $0x0101010101010101, X16
|
||||
MOV X16, ret+0(FP)
|
||||
RET
|
||||
`
|
||||
f, errs := parser.Parse("pool_riscv64.s", src)
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
img, err := AssembleFileRISCV(f)
|
||||
if err != nil {
|
||||
t.Fatalf("AssembleFileRISCV: %v", err)
|
||||
}
|
||||
// AUIPC X16, 0 + LD X16, 0(X16): the relocation pair carries the symbol.
|
||||
wantCode := []byte{0x17, 0x08, 0x00, 0x00, 0x03, 0x38, 0x08, 0x00}
|
||||
if string(img.Code[0:8]) != string(wantCode) {
|
||||
t.Errorf("pool load: got % x", img.Code[0:8])
|
||||
}
|
||||
var lit *DataSymbol
|
||||
for i := range img.DataSyms {
|
||||
if img.DataSyms[i].Name == "$i64.0101010101010101" {
|
||||
lit = &img.DataSyms[i]
|
||||
}
|
||||
}
|
||||
if lit == nil {
|
||||
t.Fatalf("pool symbol missing: %v", img.DataSyms)
|
||||
}
|
||||
wantData := []byte{0x01, 0x01, 0x01, 0x01, 0x01, 0x01, 0x01, 0x01}
|
||||
if string(img.Data[lit.Offset:lit.Offset+8]) != string(wantData) {
|
||||
t.Errorf("pool bytes: got % x", img.Data[lit.Offset:lit.Offset+8])
|
||||
}
|
||||
}
|
||||
|
||||
@@ -805,7 +929,7 @@ TEXT ·calltest(SB), NOSPLIT, $0
|
||||
CALL ext(SB)
|
||||
RET
|
||||
`)
|
||||
code, _, relocs, _, _, err := assembleRISCV(fn)
|
||||
code, _, relocs, _, _, _, err := assembleRISCV(fn)
|
||||
if err != nil {
|
||||
t.Fatalf("assemble: %v", err)
|
||||
}
|
||||
@@ -834,7 +958,7 @@ TEXT ·calllocal(SB), NOSPLIT, $0
|
||||
sub:
|
||||
RET
|
||||
`)
|
||||
_, _, _, _, _, err := assembleRISCV(fn)
|
||||
_, _, _, _, _, _, err := assembleRISCV(fn)
|
||||
if err == nil {
|
||||
t.Error("expected error for CALL to local label, got nil")
|
||||
}
|
||||
@@ -868,7 +992,7 @@ func encodeOneInstrRISCV(t *testing.T, src string, pc int, offsets map[string]in
|
||||
t.Helper()
|
||||
fn := firstTextRISCV(t, "#include \"textflag.h\"\n"+src)
|
||||
instr := fn.Body[0].(*ast.Instr)
|
||||
return encodeRISCVInstr(instr, pc, offsets, riscvFrameInfo{}, nil)
|
||||
return encodeRISCVInstr(instr, pc, offsets, riscvFrameInfo{}, nil, nil, nil)
|
||||
}
|
||||
|
||||
// TestRISCVBranchJumpRange checks that displacements beyond the B-type span
|
||||
@@ -905,9 +1029,10 @@ func TestRISCVBranchJumpRange(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// TestRISCVBranchFarBody drives the range check through the full two-pass
|
||||
// assembler: a forward branch over a body larger than the B-type span must
|
||||
// error rather than wrap.
|
||||
// TestRISCVBranchFarBody drives the relaxation pass through the full
|
||||
// assembler: a forward branch over a body larger than the B-type span is
|
||||
// rewritten as an inverted branch over an inserted JMP, the same layout the
|
||||
// toolchain produces, instead of wrapping to a wrong target.
|
||||
func TestRISCVBranchFarBody(t *testing.T) {
|
||||
var sb strings.Builder
|
||||
sb.WriteString("#include \"textflag.h\"\nTEXT ·far(SB), NOSPLIT, $0\n\tBEQ X10, X11, done\n")
|
||||
@@ -916,8 +1041,20 @@ func TestRISCVBranchFarBody(t *testing.T) {
|
||||
}
|
||||
sb.WriteString("done:\n\tRET\n")
|
||||
fn := firstTextRISCV(t, sb.String())
|
||||
if _, _, _, _, _, err := assembleRISCV(fn); err == nil {
|
||||
t.Error("expected a branch-out-of-range error, got none")
|
||||
out, _, _, _, _, _, err := assembleRISCV(fn)
|
||||
if err != nil {
|
||||
t.Fatalf("unexpected error: %v", err)
|
||||
}
|
||||
// The relaxed branch at offset 0 targets the inserted JMP at 4 (bne
|
||||
// x10, x11, +4); the JMP at 4 carries the far forward displacement.
|
||||
wantBranch := wordLE(riscvBType(riscvEnc{0x63, 0x1, 0x00}, 10, 11, 4))
|
||||
if !bytes.Equal(out[0:4], wantBranch) {
|
||||
t.Errorf("relaxed branch = %x, want %x", out[0:4], wantBranch)
|
||||
}
|
||||
// done sits after 1100 ADDs: 4 + 4400, i.e. offset 4404 from the JMP at 4.
|
||||
wantJmp := wordLE(riscvJType(0, 4404))
|
||||
if !bytes.Equal(out[4:8], wantJmp) {
|
||||
t.Errorf("inserted JMP = %x, want %x", out[4:8], wantJmp)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -930,7 +1067,7 @@ TEXT ·csrhi(SB), NOSPLIT, $0
|
||||
CSRRW $4096, X10, X11
|
||||
RET
|
||||
`)
|
||||
if _, _, _, _, _, err := assembleRISCV(fn); err == nil {
|
||||
if _, _, _, _, _, _, err := assembleRISCV(fn); err == nil {
|
||||
t.Error("expected an out-of-range error for CSR $4096, got none")
|
||||
}
|
||||
fn = firstTextRISCV(t, `#include "textflag.h"
|
||||
@@ -938,25 +1075,24 @@ TEXT ·csrmax(SB), NOSPLIT, $0
|
||||
CSRRW $4095, X10, X11
|
||||
RET
|
||||
`)
|
||||
if _, _, _, _, _, err := assembleRISCV(fn); err != nil {
|
||||
if _, _, _, _, _, _, err := assembleRISCV(fn); err != nil {
|
||||
t.Errorf("CSR $4095 must assemble: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
// TestRISCV_Imm64Rejected checks that immediates outside the signed 32-bit
|
||||
// span are diagnosed instead of silently truncated to their low 32 bits (the
|
||||
// toolchain materialises such constants via SLLI expansion, which this
|
||||
// assembler does not implement).
|
||||
// span are diagnosed instead of silently truncated to their low 32 bits for
|
||||
// the I-type arithmetic; the MOV forms materialise the wide constant instead
|
||||
// (shift sequence or pooled load), like the toolchain.
|
||||
func TestRISCV_Imm64Rejected(t *testing.T) {
|
||||
cases := []string{
|
||||
"MOV $0x123456789, X10",
|
||||
"ADDI $0x100000000, X10, X11",
|
||||
"ANDI $-0x800000001, X10, X11",
|
||||
"SUB $0x100000000, X10, X11",
|
||||
}
|
||||
for _, src := range cases {
|
||||
fn := firstTextRISCV(t, "#include \"textflag.h\"\nTEXT ·wide(SB), NOSPLIT, $0\n\t"+src+"\n\tRET\n")
|
||||
if _, _, _, _, _, err := assembleRISCV(fn); err == nil {
|
||||
if _, _, _, _, _, _, err := assembleRISCV(fn); err == nil {
|
||||
t.Errorf("%s: expected an out-of-range error, got none", src)
|
||||
}
|
||||
}
|
||||
@@ -969,9 +1105,19 @@ TEXT ·edge(SB), NOSPLIT, $0
|
||||
SUB $0x80000000, X12, X13
|
||||
RET
|
||||
`)
|
||||
if _, _, _, _, _, err := assembleRISCV(fn); err != nil {
|
||||
if _, _, _, _, _, _, err := assembleRISCV(fn); err != nil {
|
||||
t.Errorf("int32-span immediates must assemble: %v", err)
|
||||
}
|
||||
// Beyond the span the MOV forms materialise the constant like the
|
||||
// toolchain instead of diagnosing it.
|
||||
fn = firstTextRISCV(t, `#include "textflag.h"
|
||||
TEXT ·pool(SB), NOSPLIT, $0
|
||||
MOV $0x123456789, X10
|
||||
RET
|
||||
`)
|
||||
if _, _, _, _, _, _, err := assembleRISCV(fn); err != nil {
|
||||
t.Errorf("MOV with a 64-bit immediate must assemble: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
// riscvWants decodes code as little-endian words and pins each one; the
|
||||
|
||||
+221
-7
@@ -61,6 +61,21 @@ const (
|
||||
// vexImmRMGPR is the immediate form over general-purpose registers
|
||||
// (RORX): reg = dst, rm = src, imm8 = op0, L = 0.
|
||||
vexImmRMGPR
|
||||
// vexRMOpGPR is the two-operand /digit form over general-purpose
|
||||
// registers (BLSI, BLSMSK, BLSR): ModRM.reg = /digit, ModRM.rm = src
|
||||
// (op0), VEX.vvvv = dst (op1), L = 0.
|
||||
vexRMOpGPR
|
||||
// vexCountGPR is the three-operand count form over general-purpose
|
||||
// registers (SHLX, SHRX, SARX, BEXTR, BZHI): the first operand rides
|
||||
// VEX.vvvv and the second is r/m, the opposite pairing of the ANDN
|
||||
// family, with reg = dst (op2), L = 0.
|
||||
vexCountGPR
|
||||
// vexExtractGPR is the lane-extract-to-GPR form `OP $imm, xsrc, GPR/mem
|
||||
// dst`: ModRM.reg = xsrc (op1), ModRM.rm = destination (op2), imm8 =
|
||||
// op0, the VPEXTRB/W/D/Q layout. EVEX only; the destination never
|
||||
// carries a vector length, so the register the L'L field follows is the
|
||||
// XMM source.
|
||||
vexExtractGPR
|
||||
)
|
||||
|
||||
// vexSpec describes one VEX instruction's encoding parameters.
|
||||
@@ -230,8 +245,34 @@ var vexTable = map[string]vexSpec{
|
||||
"ANDNQ": {2, 0xF2, 1, 0, -1, vexNDS3GPR},
|
||||
"MULXL": {2, 0xF6, 0, 3, -1, vexNDS3GPR},
|
||||
"MULXQ": {2, 0xF6, 1, 3, -1, vexNDS3GPR},
|
||||
"RORXL": {3, 0xF0, 0, 3, -1, vexImmRMGPR},
|
||||
"RORXQ": {3, 0xF0, 1, 3, -1, vexImmRMGPR},
|
||||
// VEX.NDS.LZ.0F38, the BMI2 three-operand bit ops: BEXTR and BZHI
|
||||
// share the F7/F5 opcodes across W, the variable shifts carry their
|
||||
// direction in the prefix (SHLX 66, SHRX F2, SARX F3) and PDEP/PEXT
|
||||
// in F2/F3.
|
||||
"BEXTRL": {2, 0xF7, 0, 0, -1, vexCountGPR},
|
||||
"BEXTRQ": {2, 0xF7, 1, 0, -1, vexCountGPR},
|
||||
"BZHIL": {2, 0xF5, 0, 0, -1, vexCountGPR},
|
||||
"BZHIQ": {2, 0xF5, 1, 0, -1, vexCountGPR},
|
||||
"SARXL": {2, 0xF7, 0, 2, -1, vexCountGPR},
|
||||
"SARXQ": {2, 0xF7, 1, 2, -1, vexCountGPR},
|
||||
"SHLXL": {2, 0xF7, 0, 1, -1, vexCountGPR},
|
||||
"SHLXQ": {2, 0xF7, 1, 1, -1, vexCountGPR},
|
||||
"SHRXL": {2, 0xF7, 0, 3, -1, vexCountGPR},
|
||||
"SHRXQ": {2, 0xF7, 1, 3, -1, vexCountGPR},
|
||||
"PDEPL": {2, 0xF5, 0, 3, -1, vexNDS3GPR},
|
||||
"PDEPQ": {2, 0xF5, 1, 3, -1, vexNDS3GPR},
|
||||
"PEXTL": {2, 0xF5, 0, 2, -1, vexNDS3GPR},
|
||||
"PEXTQ": {2, 0xF5, 1, 2, -1, vexNDS3GPR},
|
||||
// VEX.LZ.0F38.W, the BMI1 unary bit ops (src, dst: ModRM.reg = /digit,
|
||||
// rm = src, vvvv = dst).
|
||||
"BLSIL": {2, 0xF3, 0, 0, 3, vexRMOpGPR},
|
||||
"BLSIQ": {2, 0xF3, 1, 0, 3, vexRMOpGPR},
|
||||
"BLSMSKL": {2, 0xF3, 0, 0, 2, vexRMOpGPR},
|
||||
"BLSMSKQ": {2, 0xF3, 1, 0, 2, vexRMOpGPR},
|
||||
"BLSRL": {2, 0xF3, 0, 0, 1, vexRMOpGPR},
|
||||
"BLSRQ": {2, 0xF3, 1, 0, 1, vexRMOpGPR},
|
||||
"RORXL": {3, 0xF0, 0, 3, -1, vexImmRMGPR},
|
||||
"RORXQ": {3, 0xF0, 1, 3, -1, vexImmRMGPR},
|
||||
|
||||
// VEX.128.0F.W0, mask-register test (KTESTW k1, k2: reg = dst, rm = src).
|
||||
"KTESTW": {1, 0x99, 0, 0, -1, vexRM},
|
||||
@@ -290,6 +331,128 @@ var vexTable = map[string]vexSpec{
|
||||
"VCVTPD2DQY": {1, 0xE6, 0, 3, -1, vexRMSrcLen},
|
||||
"VCVTTPD2DQX": {1, 0xE6, 0, 1, -1, vexRMSrcLen},
|
||||
"VCVTTPD2DQY": {1, 0xE6, 0, 1, -1, vexRMSrcLen},
|
||||
|
||||
// --- the VEX forms the avx512enc corpus exercises alongside the EVEX
|
||||
// spellings, read off the toolchain opcode tables ---
|
||||
"VAESDEC": {2, 0xDE, 0, 1, -1, vexNDS3},
|
||||
"VAESDECLAST": {2, 0xDF, 0, 1, -1, vexNDS3},
|
||||
"VAESENC": {2, 0xDC, 0, 1, -1, vexNDS3},
|
||||
"VAESENCLAST": {2, 0xDD, 0, 1, -1, vexNDS3},
|
||||
"VANDNPD": {1, 0x55, 0, 1, -1, vexNDS3},
|
||||
"VANDPD": {1, 0x54, 0, 1, -1, vexNDS3},
|
||||
"VCOMISD": {1, 0x2F, 0, 1, -1, vexRM},
|
||||
"VCVTSD2SS": {1, 0x5A, 0, 3, -1, vexNDS3},
|
||||
"VCVTSS2SD": {1, 0x5A, 0, 2, -1, vexNDS3},
|
||||
"VFMADD132PD": {2, 0x98, 1, 1, -1, vexNDS3},
|
||||
"VFMADD132PS": {2, 0x98, 0, 1, -1, vexNDS3},
|
||||
"VFMADD132SD": {2, 0x99, 1, 1, -1, vexNDS3},
|
||||
"VFMADD132SS": {2, 0x99, 0, 1, -1, vexNDS3},
|
||||
"VFMADD213PD": {2, 0xA8, 1, 1, -1, vexNDS3},
|
||||
"VFMADD213PS": {2, 0xA8, 0, 1, -1, vexNDS3},
|
||||
"VFMADD213SS": {2, 0xA9, 0, 1, -1, vexNDS3},
|
||||
"VFMADD231PS": {2, 0xB8, 0, 1, -1, vexNDS3},
|
||||
"VFMADD231SD": {2, 0xB9, 1, 1, -1, vexNDS3},
|
||||
"VFMADD231SS": {2, 0xB9, 0, 1, -1, vexNDS3},
|
||||
"VFMADDSUB132PD": {2, 0x96, 1, 1, -1, vexNDS3},
|
||||
"VFMADDSUB132PS": {2, 0x96, 0, 1, -1, vexNDS3},
|
||||
"VFMADDSUB213PD": {2, 0xA6, 1, 1, -1, vexNDS3},
|
||||
"VFMADDSUB213PS": {2, 0xA6, 0, 1, -1, vexNDS3},
|
||||
"VFMADDSUB231PD": {2, 0xB6, 1, 1, -1, vexNDS3},
|
||||
"VFMADDSUB231PS": {2, 0xB6, 0, 1, -1, vexNDS3},
|
||||
"VFMSUB132PD": {2, 0x9A, 1, 1, -1, vexNDS3},
|
||||
"VFMSUB132PS": {2, 0x9A, 0, 1, -1, vexNDS3},
|
||||
"VFMSUB132SD": {2, 0x9B, 1, 1, -1, vexNDS3},
|
||||
"VFMSUB132SS": {2, 0x9B, 0, 1, -1, vexNDS3},
|
||||
"VFMSUB213PD": {2, 0xAA, 1, 1, -1, vexNDS3},
|
||||
"VFMSUB213PS": {2, 0xAA, 0, 1, -1, vexNDS3},
|
||||
"VFMSUB213SD": {2, 0xAB, 1, 1, -1, vexNDS3},
|
||||
"VFMSUB213SS": {2, 0xAB, 0, 1, -1, vexNDS3},
|
||||
"VFMSUB231PD": {2, 0xBA, 1, 1, -1, vexNDS3},
|
||||
"VFMSUB231PS": {2, 0xBA, 0, 1, -1, vexNDS3},
|
||||
"VFMSUB231SD": {2, 0xBB, 1, 1, -1, vexNDS3},
|
||||
"VFMSUB231SS": {2, 0xBB, 0, 1, -1, vexNDS3},
|
||||
"VFMSUBADD132PD": {2, 0x97, 1, 1, -1, vexNDS3},
|
||||
"VFMSUBADD132PS": {2, 0x97, 0, 1, -1, vexNDS3},
|
||||
"VFMSUBADD213PD": {2, 0xA7, 1, 1, -1, vexNDS3},
|
||||
"VFMSUBADD213PS": {2, 0xA7, 0, 1, -1, vexNDS3},
|
||||
"VFMSUBADD231PD": {2, 0xB7, 1, 1, -1, vexNDS3},
|
||||
"VFMSUBADD231PS": {2, 0xB7, 0, 1, -1, vexNDS3},
|
||||
"VFNMADD132PD": {2, 0x9C, 1, 1, -1, vexNDS3},
|
||||
"VFNMADD132PS": {2, 0x9C, 0, 1, -1, vexNDS3},
|
||||
"VFNMADD132SD": {2, 0x9D, 1, 1, -1, vexNDS3},
|
||||
"VFNMADD132SS": {2, 0x9D, 0, 1, -1, vexNDS3},
|
||||
"VFNMADD213PD": {2, 0xAC, 1, 1, -1, vexNDS3},
|
||||
"VFNMADD213PS": {2, 0xAC, 0, 1, -1, vexNDS3},
|
||||
"VFNMADD213SD": {2, 0xAD, 1, 1, -1, vexNDS3},
|
||||
"VFNMADD213SS": {2, 0xAD, 0, 1, -1, vexNDS3},
|
||||
"VFNMADD231PD": {2, 0xBC, 1, 1, -1, vexNDS3},
|
||||
"VFNMADD231PS": {2, 0xBC, 0, 1, -1, vexNDS3},
|
||||
"VFNMADD231SS": {2, 0xBD, 0, 1, -1, vexNDS3},
|
||||
"VFNMSUB132PD": {2, 0x9E, 1, 1, -1, vexNDS3},
|
||||
"VFNMSUB132PS": {2, 0x9E, 0, 1, -1, vexNDS3},
|
||||
"VFNMSUB132SD": {2, 0x9F, 1, 1, -1, vexNDS3},
|
||||
"VFNMSUB132SS": {2, 0x9F, 0, 1, -1, vexNDS3},
|
||||
"VFNMSUB213PD": {2, 0xAE, 1, 1, -1, vexNDS3},
|
||||
"VFNMSUB213PS": {2, 0xAE, 0, 1, -1, vexNDS3},
|
||||
"VFNMSUB213SD": {2, 0xAF, 1, 1, -1, vexNDS3},
|
||||
"VFNMSUB213SS": {2, 0xAF, 0, 1, -1, vexNDS3},
|
||||
"VFNMSUB231PD": {2, 0xBE, 1, 1, -1, vexNDS3},
|
||||
"VFNMSUB231PS": {2, 0xBE, 0, 1, -1, vexNDS3},
|
||||
"VFNMSUB231SD": {2, 0xBF, 1, 1, -1, vexNDS3},
|
||||
"VFNMSUB231SS": {2, 0xBF, 0, 1, -1, vexNDS3},
|
||||
"VGF2P8AFFINEINVQB": {3, 0xCF, 1, 1, -1, vexNDS3Imm},
|
||||
"VGF2P8MULB": {2, 0xCF, 0, 1, -1, vexNDS3},
|
||||
"VMOVNTDQA": {2, 0x2A, 0, 1, -1, vexRM},
|
||||
"VMOVNTPD": {1, 0x2B, 0, 1, -1, vexRMRev},
|
||||
"VORPD": {1, 0x56, 0, 1, -1, vexNDS3},
|
||||
"VPADDSB": {1, 0xEC, 0, 1, -1, vexNDS3},
|
||||
"VPADDSW": {1, 0xED, 0, 1, -1, vexNDS3},
|
||||
"VPADDUSB": {1, 0xDC, 0, 1, -1, vexNDS3},
|
||||
"VPADDUSW": {1, 0xDD, 0, 1, -1, vexNDS3},
|
||||
"VPCMPEQQ": {2, 0x29, 0, 1, -1, vexNDS3},
|
||||
"VPCMPEQW": {1, 0x75, 0, 1, -1, vexNDS3},
|
||||
"VPCMPGTB": {1, 0x64, 0, 1, -1, vexNDS3},
|
||||
"VPCMPGTD": {1, 0x66, 0, 1, -1, vexNDS3},
|
||||
"VPCMPGTW": {1, 0x65, 0, 1, -1, vexNDS3},
|
||||
"VPERMPS": {2, 0x16, 0, 1, -1, vexNDS3},
|
||||
"VPEXTRB": {3, 0x14, 0, 1, -1, vexExtract},
|
||||
"VPEXTRD": {3, 0x16, 0, 1, -1, vexExtract},
|
||||
"VPEXTRQ": {3, 0x16, 1, 1, -1, vexExtract},
|
||||
"VPINSRD": {3, 0x22, 0, 1, -1, vexNDS3Imm},
|
||||
"VPINSRQ": {3, 0x22, 1, 1, -1, vexNDS3Imm},
|
||||
"VPMULHRSW": {2, 0x0B, 0, 1, -1, vexNDS3},
|
||||
"VPMULHW": {1, 0xE5, 0, 1, -1, vexNDS3},
|
||||
"VPMULUDQ": {1, 0xF4, 0, 1, -1, vexNDS3},
|
||||
"VPSADBW": {1, 0xF6, 0, 1, -1, vexNDS3},
|
||||
"VPSUBSB": {1, 0xE8, 0, 1, -1, vexNDS3},
|
||||
"VPSUBSW": {1, 0xE9, 0, 1, -1, vexNDS3},
|
||||
"VPSUBUSB": {1, 0xD8, 0, 1, -1, vexNDS3},
|
||||
"VPSUBUSW": {1, 0xD9, 0, 1, -1, vexNDS3},
|
||||
"VPUNPCKHBW": {1, 0x68, 0, 1, -1, vexNDS3},
|
||||
"VPUNPCKHQDQ": {1, 0x6D, 0, 1, -1, vexNDS3},
|
||||
"VPUNPCKHWD": {1, 0x69, 0, 1, -1, vexNDS3},
|
||||
"VPUNPCKLBW": {1, 0x60, 0, 1, -1, vexNDS3},
|
||||
"VPUNPCKLWD": {1, 0x61, 0, 1, -1, vexNDS3},
|
||||
"VSQRTPD": {1, 0x51, 0, 1, -1, vexRM},
|
||||
"VSQRTSD": {1, 0x51, 0, 3, -1, vexNDS3},
|
||||
"VSQRTSS": {1, 0x51, 0, 2, -1, vexNDS3},
|
||||
"VUCOMISD": {1, 0x2E, 0, 1, -1, vexRM},
|
||||
|
||||
// VEX.0F.WIG, the plain-prefix single/double arithmetic and unpack
|
||||
// spellings (no 66 prefix; WIG, so W = 0).
|
||||
"VANDNPS": {1, 0x55, 0, 0, -1, vexNDS3},
|
||||
"VANDPS": {1, 0x54, 0, 0, -1, vexNDS3},
|
||||
"VORPS": {1, 0x56, 0, 0, -1, vexNDS3},
|
||||
"VUNPCKLPS": {1, 0x14, 0, 0, -1, vexNDS3},
|
||||
"VUNPCKHPS": {1, 0x15, 0, 0, -1, vexNDS3},
|
||||
"VSQRTPS": {1, 0x51, 0, 0, -1, vexRM},
|
||||
"VMOVNTPS": {1, 0x2B, 0, 0, -1, vexRMRev},
|
||||
// VEX.128.66.0F, the scalar and packed compare forms.
|
||||
"VCOMISS": {1, 0x2F, 0, 1, -1, vexRM},
|
||||
"VUCOMISS": {1, 0x2E, 0, 0, -1, vexRM},
|
||||
// VEX.128.0F.F3/F2.W0, the high/low word shuffles ($imm, src, dst).
|
||||
"VPSHUFHW": {1, 0x70, 0, 2, -1, vexImmRM},
|
||||
"VPSHUFLW": {1, 0x70, 0, 3, -1, vexImmRM},
|
||||
}
|
||||
|
||||
// vexSrcLen maps a source-length conversion mnemonic (the X/Y spellings of
|
||||
@@ -420,6 +583,10 @@ func (e *enc) encodeVex(mnemUpper string, ops []Operand) error {
|
||||
return e.encodeVexNDS3GPR(spec, ops)
|
||||
case vexImmRMGPR:
|
||||
return e.encodeVexImmRMGPR(spec, ops)
|
||||
case vexRMOpGPR:
|
||||
return e.encodeVexRMOpGPR(spec, ops)
|
||||
case vexCountGPR:
|
||||
return e.encodeVexCountGPR(spec, ops)
|
||||
case vexRMRev:
|
||||
return e.encodeVexRMRev(spec, ops)
|
||||
}
|
||||
@@ -523,9 +690,10 @@ func (e *enc) encodeVexShiftImm(spec vexSpec, ops []Operand) error {
|
||||
if !ok {
|
||||
return fmt.Errorf("shift count must be an immediate")
|
||||
}
|
||||
srcReg, ok := src.(Reg)
|
||||
if !ok || !srcReg.isVec() {
|
||||
return fmt.Errorf("shift source must be a vector register")
|
||||
// The count source is a vector register or memory; the VEX length
|
||||
// follows the destination register either way.
|
||||
if !vecOrMem(src) {
|
||||
return fmt.Errorf("shift source must be a vector register or memory")
|
||||
}
|
||||
dstReg, ok := dst.(Reg)
|
||||
if !ok || !dstReg.isVec() {
|
||||
@@ -533,7 +701,7 @@ func (e *enc) encodeVexShiftImm(spec vexSpec, ops []Operand) error {
|
||||
}
|
||||
|
||||
vvvvBar := 15 - (dstReg.idx & 15)
|
||||
if err := e.emitVexFields(spec, dstReg.vecLenBit(), spec.opdigit, 0, vvvvBar, srcReg); err != nil {
|
||||
if err := e.emitVexFields(spec, dstReg.vecLenBit(), spec.opdigit, 0, vvvvBar, src); err != nil {
|
||||
return err
|
||||
}
|
||||
immByte, err := imm8(int64(immVal))
|
||||
@@ -700,7 +868,11 @@ func (e *enc) encodeVexNDS3GPR(spec vexSpec, ops []Operand) error {
|
||||
if !ok || vvvvReg.isVec() {
|
||||
return fmt.Errorf("VEX vvvv operand must be a general-purpose register")
|
||||
}
|
||||
return e.emitVexFields(spec, 0, dstReg.idx&7, 0, 15-(vvvvReg.idx&15), src2)
|
||||
rBit := 0
|
||||
if dstReg.idx >= 8 {
|
||||
rBit = 1
|
||||
}
|
||||
return e.emitVexFields(spec, 0, dstReg.idx&7, rBit, 15-(vvvvReg.idx&15), src2)
|
||||
}
|
||||
|
||||
// encodeVexImmRMGPR encodes the immediate form over general-purpose
|
||||
@@ -729,6 +901,48 @@ func (e *enc) encodeVexImmRMGPR(spec vexSpec, ops []Operand) error {
|
||||
return nil
|
||||
}
|
||||
|
||||
// encodeVexRMOpGPR encodes the two-operand /digit form over general-purpose
|
||||
// registers (BLSI, BLSMSK, BLSR): OP src, dst with ModRM.reg = /digit,
|
||||
// ModRM.rm = src and VEX.vvvv = dst.
|
||||
func (e *enc) encodeVexRMOpGPR(spec vexSpec, ops []Operand) error {
|
||||
if len(ops) != 2 {
|
||||
return fmt.Errorf("instruction expects 2 operands (src, dst), got %d", len(ops))
|
||||
}
|
||||
src, dst := ops[0], ops[1]
|
||||
dstReg, ok := dst.(Reg)
|
||||
if !ok || dstReg.isVec() {
|
||||
return fmt.Errorf("VEX destination must be a general-purpose register")
|
||||
}
|
||||
return e.emitVexFields(spec, 0, spec.opdigit, 0, 15-(dstReg.idx&15), src)
|
||||
}
|
||||
|
||||
// encodeVexCountGPR encodes the three-operand count form over general-purpose
|
||||
// registers (SHLX, SHRX, SARX, BEXTR, BZHI): OP src, count, dst with
|
||||
// VEX.vvvv = src (op0), ModRM.rm = count (op1), ModRM.reg = dst (op2).
|
||||
func (e *enc) encodeVexCountGPR(spec vexSpec, ops []Operand) error {
|
||||
if len(ops) != 3 {
|
||||
return fmt.Errorf("VEX count instruction expects 3 operands, got %d", len(ops))
|
||||
}
|
||||
src, count, dst := ops[0], ops[1], ops[2]
|
||||
dstReg, ok := dst.(Reg)
|
||||
if !ok || dstReg.isVec() {
|
||||
return fmt.Errorf("VEX destination must be a general-purpose register")
|
||||
}
|
||||
countReg, ok := count.(Reg)
|
||||
if !ok || countReg.isVec() {
|
||||
return fmt.Errorf("VEX count operand must be a general-purpose register")
|
||||
}
|
||||
srcReg, ok := src.(Reg)
|
||||
if !ok || srcReg.isVec() {
|
||||
return fmt.Errorf("VEX count source must be a general-purpose register")
|
||||
}
|
||||
rBit := 0
|
||||
if dstReg.idx >= 8 {
|
||||
rBit = 1
|
||||
}
|
||||
return e.emitVexFields(spec, 0, dstReg.idx&7, rBit, 15-(srcReg.idx&15), count)
|
||||
}
|
||||
|
||||
// encodeVexRMRev encodes the reversed two-operand form: OP src, dst with the
|
||||
// vector source in ModRM.reg and the memory destination in r/m (VMOVNTDQ,
|
||||
// a store with no register-destination form).
|
||||
|
||||
@@ -31,6 +31,51 @@ var x86asmUnrecognised = map[string]bool{
|
||||
"RORXQ": true,
|
||||
"VFMADD213SD": true,
|
||||
"VFNMADD231SD": true,
|
||||
// The scalar FMA spellings the decoder's tables lack entirely.
|
||||
"VFMADD132SD": true,
|
||||
"VFMADD132SS": true,
|
||||
"VFMADD213SS": true,
|
||||
"VFMADD231SD": true,
|
||||
"VFMADD231SS": true,
|
||||
"VFMSUB132SD": true,
|
||||
"VFMSUB132SS": true,
|
||||
"VFMSUB213SD": true,
|
||||
"VFMSUB213SS": true,
|
||||
"VFMSUB231SD": true,
|
||||
"VFMSUB231SS": true,
|
||||
"VFNMADD132SD": true,
|
||||
"VFNMADD132SS": true,
|
||||
"VFNMADD213SD": true,
|
||||
"VFNMADD213SS": true,
|
||||
"VFNMADD231SS": true,
|
||||
"VFNMSUB132SD": true,
|
||||
"VFNMSUB132SS": true,
|
||||
"VFNMSUB213SD": true,
|
||||
"VFNMSUB213SS": true,
|
||||
"VFNMSUB231SD": true,
|
||||
"VFNMSUB231SS": true,
|
||||
// The BMI1 unary bit ops the decoder's AVX tables lack.
|
||||
"BLSIL": true,
|
||||
"BLSIQ": true,
|
||||
"BLSMSKL": true,
|
||||
"BLSMSKQ": true,
|
||||
"BLSRL": true,
|
||||
"BLSRQ": true,
|
||||
// The BMI2 bit ops whose W1/LZ rows the decoder misses.
|
||||
"BEXTRL": true,
|
||||
"BEXTRQ": true,
|
||||
"BZHIL": true,
|
||||
"BZHIQ": true,
|
||||
"PDEPL": true,
|
||||
"PDEPQ": true,
|
||||
"PEXTL": true,
|
||||
"PEXTQ": true,
|
||||
"SARXL": true,
|
||||
"SARXQ": true,
|
||||
"SHLXL": true,
|
||||
"SHLXQ": true,
|
||||
"SHRXL": true,
|
||||
"SHRXQ": true,
|
||||
}
|
||||
|
||||
// TestVexNDS3 encodes `mnem Y0, Y1, Y2` for every three-operand NDS
|
||||
@@ -237,6 +282,18 @@ func TestVexGroundTruth(t *testing.T) {
|
||||
{"MULXQ AX,BX,CX", "MULXQ", []Operand{AX, BX, CX}, "c4e2e3f6c8", ""},
|
||||
{"RORXL $3,AX,CX", "RORXL", []Operand{Imm(3), AX, CX}, "c4e37bf0c803", ""},
|
||||
{"RORXQ $3,AX,CX", "RORXQ", []Operand{Imm(3), AX, CX}, "c4e3fbf0c803", ""},
|
||||
// BMI2 variable shifts and bit ops (three general registers).
|
||||
{"SHLXL AX,CX,R15", "SHLXL", []Operand{AX, CX, vreg(t, "R15")}, "c46279f7f9", ""},
|
||||
{"SHRXQ R8,DX,AX", "SHRXQ", []Operand{vreg(t, "R8"), DX, AX}, "c4e2bbf7c2", ""},
|
||||
{"SARXQ AX,DX,R9", "SARXQ", []Operand{AX, DX, vreg(t, "R9")}, "c462faf7ca", ""},
|
||||
{"BEXTRL AX,CX,R15", "BEXTRL", []Operand{AX, CX, vreg(t, "R15")}, "c46278f7f9", ""},
|
||||
{"BZHIQ AX,CX,R15", "BZHIQ", []Operand{AX, CX, vreg(t, "R15")}, "c462f8f5f9", ""},
|
||||
{"PDEPQ AX,CX,R15", "PDEPQ", []Operand{AX, CX, vreg(t, "R15")}, "c462f3f5f8", ""},
|
||||
{"PEXTQ AX,CX,R15", "PEXTQ", []Operand{AX, CX, vreg(t, "R15")}, "c462f2f5f8", ""},
|
||||
// BMI1 unary bit ops (src, dst: /digit in ModRM.reg, dst in vvvv).
|
||||
{"BLSIL AX,CX", "BLSIL", []Operand{AX, CX}, "c4e270f3d8", ""},
|
||||
{"BLSRQ AX,CX", "BLSRQ", []Operand{AX, CX}, "c4e2f0f3c8", ""},
|
||||
{"BLSMSKQ AX,CX", "BLSMSKQ", []Operand{AX, CX}, "c4e2f0f3d0", ""},
|
||||
// Two-operand reg/rm form (v̄vvv must be 1111).
|
||||
{"VPMOVSXDQ X0,Y4", "VPMOVSXDQ", []Operand{vreg(t, "X0"), vreg(t, "Y4")}, "c4e27d25e0", ""},
|
||||
{"VPMOVSXWD (SI),Y0", "VPMOVSXWD", []Operand{Ptr(SI, 0, 8), vreg(t, "Y0")}, "c4e27d2306", ""},
|
||||
|
||||
+17
-7
@@ -151,11 +151,21 @@ type Immediate struct {
|
||||
// Address is a non-immediate operand: a register, a memory reference, a symbol
|
||||
// reference or a label. Fields are populated best-effort from the syntax.
|
||||
type Address struct {
|
||||
Sym *Symbol // name reference (bare ident, or name+off(pseudo))
|
||||
Base string // base register, from (base)
|
||||
Index string // index register, from (index*scale)
|
||||
Scale int // index scale; 0 when absent
|
||||
Offset int64 // leading displacement, from off(base)
|
||||
HasOff bool // a leading displacement is present
|
||||
Shift string // verbatim arm64 shift suffix, e.g. "<< 2"
|
||||
Sym *Symbol // name reference (bare ident, or name+off(pseudo))
|
||||
Base string // base register, from (base)
|
||||
Index string // index register, from (index*scale)
|
||||
Scale int // index scale; 0 when absent
|
||||
Offset int64 // leading displacement, from off(base)
|
||||
HasOff bool // a leading displacement is present
|
||||
Shift string // verbatim arm64 shift suffix, e.g. "<< 2"
|
||||
Range *RegRange // bracketed register range; nil for every other form
|
||||
}
|
||||
|
||||
// RegRange is a bracketed register range, [Z0-Z3]: the amd64 spelling of
|
||||
// the four-register source of the 4FMAPS/4VNNIW families. Lo and Hi carry
|
||||
// the verbatim register spellings; the range is inclusive at both ends.
|
||||
type RegRange struct {
|
||||
Lo string
|
||||
Hi string
|
||||
Pos token.Position
|
||||
}
|
||||
|
||||
@@ -0,0 +1,367 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
package main
|
||||
|
||||
import (
|
||||
"errors"
|
||||
"fmt"
|
||||
"go/ast"
|
||||
"go/build"
|
||||
"go/constant"
|
||||
"go/parser"
|
||||
"go/token"
|
||||
"go/types"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"regexp"
|
||||
"strings"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
|
||||
)
|
||||
|
||||
// go_asm.h is the header the Go compiler writes for every package that
|
||||
// carries assembly (the compiler's -asmhdr output): "#define const_NAME
|
||||
// value" for each package constant, and for each named struct type
|
||||
// "#define TYPE__size size" plus one "#define TYPE_field offset" per field.
|
||||
// GOROOT assembly includes it, and a standalone assembler has no compiler
|
||||
// to have produced it, so gasm generates the equivalent itself: the package
|
||||
// the .s file lives in is parsed and type-checked here, with the target
|
||||
// architecture's own sizes, and the same defines are written out. The
|
||||
// type-checking GOOS is selected by the caller: a GOOS-specific file
|
||||
// (sys_darwin_arm64.s) needs its platform's defines, which a header from
|
||||
// the ambient GOOS silently omits.
|
||||
//
|
||||
// The emitter mirrors cmd/compile's dumpasmhdr exactly: constants come out
|
||||
// as "const_NAME", struct entries as "NAME__size" followed by the fields in
|
||||
// declaration order, blank names are skipped, and float and complex
|
||||
// constants are omitted (the assembler carries integers, bools and strings
|
||||
// only). Aliases to structs are emitted, generic types are not: they have
|
||||
// no fixed size. A define the assembly references but this header does not
|
||||
// carry surfaces later as the assembler's own "undefined" diagnostic naming
|
||||
// the define, which is the honest failure.
|
||||
|
||||
// goAsmInclude matches the #include "go_asm.h" directive, tolerant of
|
||||
// whitespace, so the wiring knows which files need a generated header
|
||||
// before the preprocessor runs and would report the header as missing.
|
||||
var goAsmInclude = regexp.MustCompile(`(?m)^\s*#\s*include\s+"go_asm\.h"`)
|
||||
|
||||
// needsGoAsmHeader reports whether src includes go_asm.h.
|
||||
func needsGoAsmHeader(src string) bool {
|
||||
return goAsmInclude.MatchString(src)
|
||||
}
|
||||
|
||||
// goAsmHeaderResolved reports whether the include of go_asm.h from a file in
|
||||
// asmDir already resolves: to a header in the package directory itself, or
|
||||
// in one of the -I directories, the way the preprocessor searches. Only an
|
||||
// unresolved include is generated for; a header someone placed by hand is
|
||||
// the tool the author chose, and it also wins the preprocessor's own search
|
||||
// order, so generating a second copy would be dead weight at best.
|
||||
func goAsmHeaderResolved(asmDir string, dirs []string) bool {
|
||||
candidates := []string{filepath.Join(asmDir, "go_asm.h")}
|
||||
for _, d := range dirs {
|
||||
candidates = append(candidates, filepath.Join(d, "go_asm.h"))
|
||||
}
|
||||
for _, candidate := range candidates {
|
||||
if st, err := os.Stat(candidate); err == nil && !st.IsDir() {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// generateGoAsmHeader type-checks the Go package in pkgDir for goos and
|
||||
// goarch, writes its go_asm.h equivalent into dir, and returns dir. An
|
||||
// empty goos means the ambient one. The caller owns the directory and its
|
||||
// removal.
|
||||
func generateGoAsmHeader(pkgDir, goos, goarch, dir string) (string, error) {
|
||||
if goos == "" {
|
||||
goos = build.Default.GOOS
|
||||
}
|
||||
imp := newSourceImporter(goos, goarch)
|
||||
if imp.sizes == nil {
|
||||
return "", fmt.Errorf("go_asm.h: unknown GOARCH %q", goarch)
|
||||
}
|
||||
bp, err := imp.ctxt.ImportDir(pkgDir, 0)
|
||||
if err != nil {
|
||||
return "", fmt.Errorf("go_asm.h for GOARCH %s in %s: %w", goarch, pkgDir, err)
|
||||
}
|
||||
files, errs := imp.parse(bp)
|
||||
if len(errs) > 0 {
|
||||
return "", fmt.Errorf("go_asm.h for GOARCH %s in %s: %s", goarch, pkgDir, errorList(errs))
|
||||
}
|
||||
_, info, errs := imp.checkPackage(bp, files)
|
||||
if len(errs) > 0 {
|
||||
return "", fmt.Errorf("go_asm.h for GOARCH %s in %s: package does not type-check: %s", goarch, pkgDir, errorList(errs))
|
||||
}
|
||||
|
||||
var b strings.Builder
|
||||
fmt.Fprintf(&b, "// generated by gasm from package %s (GOOS %s, GOARCH %s)\n\n", bp.Name, goos, goarch)
|
||||
// Files in the build's own order and declarations in source order: the
|
||||
// same walk the compiler's reader makes, so the header reads the same
|
||||
// way the toolchain's does. Order carries no meaning to the assembler
|
||||
// (defines form a table), only to a human diffing against one.
|
||||
for _, f := range files {
|
||||
for _, decl := range f.Decls {
|
||||
gd, ok := decl.(*ast.GenDecl)
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
for _, spec := range gd.Specs {
|
||||
switch gd.Tok {
|
||||
case token.CONST:
|
||||
vs, ok := spec.(*ast.ValueSpec)
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
for _, name := range vs.Names {
|
||||
emitConst(&b, info.Defs[name], name.Name)
|
||||
}
|
||||
case token.TYPE:
|
||||
ts, ok := spec.(*ast.TypeSpec)
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
emitStruct(&b, imp.sizes, info.Defs[ts.Name], ts.Name.Name)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
if err := os.MkdirAll(dir, 0o755); err != nil {
|
||||
return "", fmt.Errorf("go_asm.h for GOARCH %s in %s: %w", goarch, pkgDir, err)
|
||||
}
|
||||
out := filepath.Join(dir, "go_asm.h")
|
||||
if err := os.WriteFile(out, []byte(b.String()), 0o644); err != nil {
|
||||
return "", fmt.Errorf("go_asm.h for GOARCH %s in %s: %w", goarch, pkgDir, err)
|
||||
}
|
||||
return dir, nil
|
||||
}
|
||||
|
||||
// emitConst writes one const define, skipping what the toolchain skips:
|
||||
// blank names, and float and complex values the assembler has no syntax for.
|
||||
func emitConst(b *strings.Builder, obj types.Object, name string) {
|
||||
c, ok := obj.(*types.Const)
|
||||
if !ok || name == "_" {
|
||||
return
|
||||
}
|
||||
switch c.Val().Kind() {
|
||||
case constant.Float, constant.Complex, constant.Unknown:
|
||||
return
|
||||
}
|
||||
fmt.Fprintf(b, "#define const_%s %s\n", name, c.Val().ExactString())
|
||||
}
|
||||
|
||||
// emitStruct writes one named struct type's size and field offsets,
|
||||
// skipping what the toolchain skips: blank names, non-struct types, and
|
||||
// generic types, whose size depends on their instantiation.
|
||||
func emitStruct(b *strings.Builder, sizes types.Sizes, obj types.Object, name string) {
|
||||
tn, ok := obj.(*types.TypeName)
|
||||
if !ok || name == "_" {
|
||||
return
|
||||
}
|
||||
t := types.Unalias(tn.Type())
|
||||
// Generic types are spelled *types.Named with a type-parameter list;
|
||||
// a plain struct type or an instantiated one carries none.
|
||||
if named, ok := t.(*types.Named); ok && named.TypeParams().Len() > 0 {
|
||||
return
|
||||
}
|
||||
st, ok := t.Underlying().(*types.Struct)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
fmt.Fprintf(b, "#define %s__size %d\n", name, sizes.Sizeof(t))
|
||||
fields := make([]*types.Var, st.NumFields())
|
||||
for i := range st.NumFields() {
|
||||
fields[i] = st.Field(i)
|
||||
}
|
||||
for i, off := range sizes.Offsetsof(fields) {
|
||||
fld := fields[i]
|
||||
if fld.Name() == "_" {
|
||||
continue
|
||||
}
|
||||
fmt.Fprintf(b, "#define %s_%s %d\n", name, fld.Name(), off)
|
||||
}
|
||||
}
|
||||
|
||||
// errorList renders at most three errors, enough to say what is wrong
|
||||
// without burying the diagnostic the caller actually reads.
|
||||
func errorList(errs []error) string {
|
||||
if len(errs) > 3 {
|
||||
errs = errs[:3]
|
||||
}
|
||||
msgs := make([]string, len(errs))
|
||||
for i, err := range errs {
|
||||
msgs[i] = err.Error()
|
||||
}
|
||||
return strings.Join(msgs, "; ")
|
||||
}
|
||||
|
||||
// sourceImporter type-checks imported packages from source with the target
|
||||
// architecture's sizes. go/importer's "source" importer pins the host
|
||||
// GOARCH, which would lay out imported types (internal/cpu, internal/abi)
|
||||
// for the wrong target on a cross-architecture header, so the recursion is
|
||||
// carried here with one build context and one sizes instance per
|
||||
// architecture.
|
||||
type sourceImporter struct {
|
||||
fset *token.FileSet
|
||||
ctxt *build.Context
|
||||
sizes types.Sizes
|
||||
pkgs map[string]*types.Package
|
||||
}
|
||||
|
||||
// newSourceImporter returns the importer for one target GOOS and GOARCH.
|
||||
// Cgo is disabled so the file set is deterministic and independent of the
|
||||
// host's C toolchain: cgo-tagged files drop out of the build exactly as
|
||||
// they do from a CGO_ENABLED=0 build, whose assembly is what gasm targets.
|
||||
func newSourceImporter(goos, goarch string) *sourceImporter {
|
||||
ctxt := new(build.Context)
|
||||
*ctxt = build.Default
|
||||
ctxt.GOOS = goos
|
||||
ctxt.GOARCH = goarch
|
||||
ctxt.CgoEnabled = false
|
||||
return &sourceImporter{
|
||||
fset: token.NewFileSet(),
|
||||
ctxt: ctxt,
|
||||
sizes: types.SizesFor("gc", goarch),
|
||||
pkgs: map[string]*types.Package{},
|
||||
}
|
||||
}
|
||||
|
||||
// Import type-checks one imported package and memoises it. "unsafe" must
|
||||
// resolve to go/types' own package, never to the source in GOROOT/src/unsafe:
|
||||
// the source declares Sizeof and Offsetof as ordinary functions over
|
||||
// ArbitraryType, and checking against that signature rejects half the
|
||||
// unsafe arithmetic the gc compiler accepts, which is exactly the divergence
|
||||
// srcimporter guards against the same way.
|
||||
func (im *sourceImporter) Import(path string) (*types.Package, error) {
|
||||
if path == "unsafe" {
|
||||
return types.Unsafe, nil
|
||||
}
|
||||
if p, ok := im.pkgs[path]; ok {
|
||||
return p, nil
|
||||
}
|
||||
bp, err := im.ctxt.Import(path, "", 0)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
files, errs := im.parse(bp)
|
||||
if len(errs) > 0 {
|
||||
return nil, errors.New(errorList(errs))
|
||||
}
|
||||
pkg, _, _ := im.checkPackage(bp, files)
|
||||
im.pkgs[path] = pkg
|
||||
return pkg, nil
|
||||
}
|
||||
|
||||
// parse reads the build package's Go files. Import-level failures (no Go
|
||||
// files for the target, unreadable files) come back as errors, and the
|
||||
// type-check decides the rest.
|
||||
func (im *sourceImporter) parse(bp *build.Package) ([]*ast.File, []error) {
|
||||
if len(bp.GoFiles) == 0 {
|
||||
return nil, []error{fmt.Errorf("no Go source files for GOOS=%s GOARCH=%s", im.ctxt.GOOS, im.ctxt.GOARCH)}
|
||||
}
|
||||
var (
|
||||
files []*ast.File
|
||||
errs []error
|
||||
)
|
||||
for _, name := range bp.GoFiles {
|
||||
f, err := parser.ParseFile(im.fset, filepath.Join(bp.Dir, name), nil, parser.SkipObjectResolution)
|
||||
if err != nil {
|
||||
errs = append(errs, err)
|
||||
continue
|
||||
}
|
||||
files = append(files, f)
|
||||
}
|
||||
return files, errs
|
||||
}
|
||||
|
||||
// checkPackage type-checks one package's files with the importer's sizes,
|
||||
// recording every error: a header from a package that does not type-check
|
||||
// could silently mis-state an offset, so the caller refuses the header
|
||||
// rather than trusting it. The returned Defs map backs the root package's
|
||||
// emission walk; imports only need the checked package itself.
|
||||
func (im *sourceImporter) checkPackage(bp *build.Package, files []*ast.File) (*types.Package, *types.Info, []error) {
|
||||
var errs []error
|
||||
conf := &types.Config{
|
||||
Importer: im,
|
||||
Sizes: im.sizes,
|
||||
Error: func(err error) { errs = append(errs, err) },
|
||||
}
|
||||
info := &types.Info{Defs: map[*ast.Ident]types.Object{}}
|
||||
pkg, _ := conf.Check(bp.ImportPath, im.fset, files, info)
|
||||
return pkg, info, errs
|
||||
}
|
||||
|
||||
// asmhdrCache generates one go_asm.h per package directory and target
|
||||
// architecture under one temp root, for callers that assemble many files
|
||||
// (the corpus audit). Failures are cached too: a package that does not
|
||||
// type-check must not be re-checked once per file.
|
||||
type asmhdrCache struct {
|
||||
root string
|
||||
dirs map[string]string // "pkgDir\x00goos\x00goarch" -> directory holding go_asm.h
|
||||
errs map[string]error
|
||||
}
|
||||
|
||||
func newAsmhdrCache() (*asmhdrCache, error) {
|
||||
root, err := os.MkdirTemp("", "gasm-asmhdr")
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
return &asmhdrCache{root: root, dirs: map[string]string{}, errs: map[string]error{}}, nil
|
||||
}
|
||||
|
||||
// dirFor returns the directory holding the generated go_asm.h for pkgDir
|
||||
// under goos and goarch, generating it on first use. An empty goos means
|
||||
// the ambient one, resolved here so that one package cannot generate twice
|
||||
// under an explicit and an implicit spelling of the same GOOS.
|
||||
func (c *asmhdrCache) dirFor(pkgDir, goos, goarch string) (string, error) {
|
||||
if goos == "" {
|
||||
goos = build.Default.GOOS
|
||||
}
|
||||
key := pkgDir + "\x00" + goos + "\x00" + goarch
|
||||
if dir, ok := c.dirs[key]; ok {
|
||||
return dir, nil
|
||||
}
|
||||
if err, ok := c.errs[key]; ok {
|
||||
return "", err
|
||||
}
|
||||
dir := filepath.Join(c.root, fmt.Sprintf("h%d_%s_%s", len(c.dirs), goos, goarch))
|
||||
if _, err := generateGoAsmHeader(pkgDir, goos, goarch, dir); err != nil {
|
||||
c.errs[key] = err
|
||||
return "", err
|
||||
}
|
||||
c.dirs[key] = dir
|
||||
return dir, nil
|
||||
}
|
||||
|
||||
// close removes the temp root.
|
||||
func (c *asmhdrCache) close() { os.RemoveAll(c.root) }
|
||||
|
||||
// ensureGoAsmHeader prepares the include directory a file that includes
|
||||
// go_asm.h needs: the generated header for the package in path's directory,
|
||||
// for the file's target GOOS and architecture. It reports a usage error
|
||||
// when the architecture cannot be determined, and passes through the
|
||||
// generator's diagnostics, which name the package.
|
||||
func ensureGoAsmHeader(path string, target arch.Arch, goos string, cache *asmhdrCache) (string, func(), error) {
|
||||
if path == "-" {
|
||||
return "", nil, errors.New("cannot generate go_asm.h for standard input (no package directory)")
|
||||
}
|
||||
if target == arch.Unknown {
|
||||
return "", nil, errors.New("a file that includes go_asm.h needs a target architecture: name the file _<arch>.s or pass -GOARCH")
|
||||
}
|
||||
if cache != nil {
|
||||
dir, err := cache.dirFor(filepath.Dir(path), goos, goarchName(target))
|
||||
return dir, func() {}, err
|
||||
}
|
||||
root, err := os.MkdirTemp("", "gasm-asmhdr")
|
||||
if err != nil {
|
||||
return "", nil, err
|
||||
}
|
||||
dir, err := generateGoAsmHeader(filepath.Dir(path), goos, goarchName(target), root)
|
||||
if err != nil {
|
||||
os.RemoveAll(root)
|
||||
return "", nil, err
|
||||
}
|
||||
return dir, func() { os.RemoveAll(root) }, nil
|
||||
}
|
||||
@@ -0,0 +1,430 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
package main
|
||||
|
||||
import (
|
||||
"os"
|
||||
"os/exec"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// writePkg lays out a minimal Go package in a temp directory.
|
||||
func writePkg(t *testing.T, files map[string]string) string {
|
||||
t.Helper()
|
||||
dir := t.TempDir()
|
||||
for name, src := range files {
|
||||
if err := os.WriteFile(filepath.Join(dir, name), []byte(src), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
return dir
|
||||
}
|
||||
|
||||
// generateFor generates the header for dir and returns its text. An empty
|
||||
// goos means the ambient one.
|
||||
func generateFor(t *testing.T, dir, goos, goarch string) string {
|
||||
t.Helper()
|
||||
hdrDir, err := generateGoAsmHeader(dir, goos, goarch, t.TempDir())
|
||||
if err != nil {
|
||||
t.Fatalf("generateGoAsmHeader(%q, %s, %s): %v", dir, goos, goarch, err)
|
||||
}
|
||||
b, err := os.ReadFile(filepath.Join(hdrDir, "go_asm.h"))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return string(b)
|
||||
}
|
||||
|
||||
func TestGenerateGoAsmHeaderShape(t *testing.T) {
|
||||
dir := writePkg(t, map[string]string{"sample.go": `package sample
|
||||
|
||||
const bufSize = 1024
|
||||
|
||||
const (
|
||||
a = iota * 8
|
||||
b
|
||||
c
|
||||
)
|
||||
|
||||
const (
|
||||
strConst = "hello"
|
||||
boolConst = true
|
||||
floatConst = 1.5
|
||||
_ = "the blank identifier is skipped"
|
||||
)
|
||||
|
||||
const shift = 1 << 20
|
||||
|
||||
type reader struct {
|
||||
r int64
|
||||
w int64
|
||||
_ [4]byte
|
||||
name string
|
||||
}
|
||||
|
||||
type scalar int
|
||||
|
||||
type aliased struct {
|
||||
k uint32
|
||||
v uint32
|
||||
}
|
||||
|
||||
type alias = aliased
|
||||
`})
|
||||
hdr := generateFor(t, dir, "", "amd64")
|
||||
want := []string{
|
||||
"#define const_bufSize 1024",
|
||||
// iota resolves through go/types, one define per name.
|
||||
"#define const_a 0",
|
||||
"#define const_b 8",
|
||||
"#define const_c 16",
|
||||
`#define const_strConst "hello"`,
|
||||
"#define const_boolConst true",
|
||||
// Floats are the toolchain's own skip, as are blank names.
|
||||
"#define const_shift 1048576",
|
||||
// The blank field still occupies its bytes: the pad after w runs to
|
||||
// the string's 8-byte alignment.
|
||||
"#define reader__size 40",
|
||||
"#define reader_r 0",
|
||||
"#define reader_w 8",
|
||||
"#define reader_name 24",
|
||||
// Non-struct named types carry no defines; aliases to structs do.
|
||||
"#define aliased__size 8",
|
||||
"#define aliased_k 0",
|
||||
"#define aliased_v 4",
|
||||
"#define alias__size 8",
|
||||
"#define alias_k 0",
|
||||
"#define alias_v 4",
|
||||
}
|
||||
for _, w := range want {
|
||||
if !strings.Contains(hdr, w+"\n") {
|
||||
t.Errorf("header misses %q\ngot:\n%s", w, hdr)
|
||||
}
|
||||
}
|
||||
for _, banned := range []string{"#define const_floatConst", "#define _ ", "#define scalar"} {
|
||||
if strings.Contains(hdr, banned) {
|
||||
t.Errorf("header must not carry %s\ngot:\n%s", banned, hdr)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestGenerateGoAsmHeaderPerArch(t *testing.T) {
|
||||
dir := writePkg(t, map[string]string{
|
||||
"common.go": `package perarch
|
||||
|
||||
type layout struct {
|
||||
a int32
|
||||
p uintptr
|
||||
}
|
||||
`,
|
||||
// The build-tagged file set is part of the contract: a per-arch
|
||||
// package is exactly how internal/cpu declares its layouts.
|
||||
"const_amd64.go": `//go:build amd64
|
||||
|
||||
package perarch
|
||||
|
||||
const flavour = 1
|
||||
`,
|
||||
"const_arm64.go": `//go:build arm64
|
||||
|
||||
package perarch
|
||||
|
||||
const flavour = 2
|
||||
`,
|
||||
})
|
||||
amd64 := generateFor(t, dir, "", "amd64")
|
||||
arm64 := generateFor(t, dir, "", "arm64")
|
||||
if !strings.Contains(amd64, "#define const_flavour 1\n") {
|
||||
t.Errorf("amd64 header misses const_flavour 1:\n%s", amd64)
|
||||
}
|
||||
if !strings.Contains(arm64, "#define const_flavour 2\n") {
|
||||
t.Errorf("arm64 header misses const_flavour 2:\n%s", arm64)
|
||||
}
|
||||
if strings.Contains(arm64, "#define const_flavour 1\n") {
|
||||
t.Errorf("arm64 header must not carry the amd64 file's value")
|
||||
}
|
||||
// SizesFor makes the layout the target's: uintptr is 4 bytes wide on
|
||||
// 386 and 8 on amd64, which must move p and grow the struct.
|
||||
if !strings.Contains(amd64, "#define layout__size 16\n") || !strings.Contains(amd64, "#define layout_p 8\n") {
|
||||
t.Errorf("amd64 layout wrong:\n%s", amd64)
|
||||
}
|
||||
w386 := generateFor(t, dir, "", "386")
|
||||
if !strings.Contains(w386, "#define layout__size 8\n") || !strings.Contains(w386, "#define layout_p 4\n") {
|
||||
t.Errorf("386 layout wrong:\n%s", w386)
|
||||
}
|
||||
}
|
||||
|
||||
// TestGenerateGoAsmHeaderGOOS pins the GOOS half of the target: only the
|
||||
// platform's own files type-check into the header, which is why
|
||||
// sys_darwin_arm64.s cannot assemble against a linux-generated one.
|
||||
func TestGenerateGoAsmHeaderGOOS(t *testing.T) {
|
||||
dir := writePkg(t, map[string]string{
|
||||
"common.go": `package goosaware
|
||||
|
||||
type shared struct {
|
||||
a int32
|
||||
}
|
||||
`,
|
||||
"plat_darwin.go": `//go:build darwin
|
||||
|
||||
package goosaware
|
||||
|
||||
type platform struct {
|
||||
trampoline_numer int64
|
||||
}
|
||||
`,
|
||||
"plat_windows.go": `//go:build windows
|
||||
|
||||
package goosaware
|
||||
|
||||
type platform struct {
|
||||
callbackArgs__size int32
|
||||
}
|
||||
`,
|
||||
})
|
||||
darwin := generateFor(t, dir, "darwin", "arm64")
|
||||
if !strings.Contains(darwin, "#define platform__size 8\n") || !strings.Contains(darwin, "#define platform_trampoline_numer 0\n") {
|
||||
t.Errorf("darwin header misses the darwin layout:\n%s", darwin)
|
||||
}
|
||||
if strings.Contains(darwin, "callbackArgs") {
|
||||
t.Errorf("darwin header must not carry the windows layout:\n%s", darwin)
|
||||
}
|
||||
windows := generateFor(t, dir, "windows", "arm64")
|
||||
if !strings.Contains(windows, "#define platform_callbackArgs__size 0\n") {
|
||||
t.Errorf("windows header misses the windows layout:\n%s", windows)
|
||||
}
|
||||
if strings.Contains(windows, "trampoline_numer") {
|
||||
t.Errorf("windows header must not carry the darwin layout:\n%s", windows)
|
||||
}
|
||||
// The ambient GOOS is neither of the two, so only shared's defines are
|
||||
// emitted; the shared type keeps its layout there.
|
||||
ambient := generateFor(t, dir, "", "arm64")
|
||||
if !strings.Contains(ambient, "#define shared__size 4\n") {
|
||||
t.Errorf("ambient header misses the shared layout:\n%s", ambient)
|
||||
}
|
||||
if strings.Contains(ambient, "#define platform_") {
|
||||
t.Errorf("ambient header must not carry either platform layout:\n%s", ambient)
|
||||
}
|
||||
}
|
||||
|
||||
func TestGoosFromFilename(t *testing.T) {
|
||||
for path, want := range map[string]string{
|
||||
"/x/sys_darwin_arm64.s": "darwin",
|
||||
"/x/sys_windows_arm64.s": "windows",
|
||||
"/x/asm_linux_amd64.s": "linux",
|
||||
"/x/rt0_darwin_arm64.s": "darwin",
|
||||
"/x/vgetrandom_zos_s390x.s": "zos",
|
||||
"/x/rt0_js_wasm.s": "js",
|
||||
"/x/memmove_amd64.s": "",
|
||||
"/x/vlop_arm.s": "",
|
||||
"/x/stubs.s": "",
|
||||
} {
|
||||
if got := goosFromFilename(path); got != want {
|
||||
t.Errorf("goosFromFilename(%q) = %q, want %q", path, got, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestGenerateGoAsmHeaderErrors(t *testing.T) {
|
||||
t.Run("type error", func(t *testing.T) {
|
||||
dir := writePkg(t, map[string]string{"bad.go": `package bad
|
||||
|
||||
const x = undefinedIdent
|
||||
`})
|
||||
_, err := generateGoAsmHeader(dir, "", "amd64", t.TempDir())
|
||||
if err == nil {
|
||||
t.Fatal("generation must fail for a package that does not type-check")
|
||||
}
|
||||
if !strings.Contains(err.Error(), dir) {
|
||||
t.Errorf("error must name the package directory: %v", err)
|
||||
}
|
||||
if !strings.Contains(err.Error(), "type-check") {
|
||||
t.Errorf("error must say the package does not type-check: %v", err)
|
||||
}
|
||||
})
|
||||
t.Run("no go files", func(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
_, err := generateGoAsmHeader(dir, "", "amd64", t.TempDir())
|
||||
if err == nil {
|
||||
t.Fatal("generation must fail without Go files")
|
||||
}
|
||||
if !strings.Contains(err.Error(), dir) {
|
||||
t.Errorf("error must name the package directory: %v", err)
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
func TestNeedsGoAsmHeader(t *testing.T) {
|
||||
yes := "#include \"go_asm.h\"\n#include \"textflag.h\"\n"
|
||||
no := "#include \"textflag.h\"\n#include \"funcdata.h\"\n"
|
||||
if !needsGoAsmHeader(yes) {
|
||||
t.Error("needsGoAsmHeader(missing on a go_asm.h include)")
|
||||
}
|
||||
if needsGoAsmHeader(no) {
|
||||
t.Error("needsGoAsmHeader claims other headers need generation")
|
||||
}
|
||||
}
|
||||
|
||||
func TestGoAsmHeaderResolved(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
if goAsmHeaderResolved(dir, nil) {
|
||||
t.Error("resolved with no header anywhere")
|
||||
}
|
||||
other := t.TempDir()
|
||||
if goAsmHeaderResolved(dir, []string{other}) {
|
||||
t.Error("resolved with an empty -I directory")
|
||||
}
|
||||
if err := os.WriteFile(filepath.Join(dir, "go_asm.h"), nil, 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if !goAsmHeaderResolved(dir, nil) {
|
||||
t.Error("not resolved with the header in the package directory")
|
||||
}
|
||||
}
|
||||
|
||||
func TestOtherGOOSFile(t *testing.T) {
|
||||
for path, want := range map[string]bool{
|
||||
"/x/sys_windows_amd64.s": true,
|
||||
"/x/rt0_js_wasm.s": true,
|
||||
"/x/sys_darwin_arm64.s": true,
|
||||
"/x/sys_linux_amd64.s": false,
|
||||
"/x/time_linux_amd64.s": false,
|
||||
"/x/memmove_amd64.s": false,
|
||||
"/x/generic.s": false,
|
||||
} {
|
||||
if got := otherGOOSFile(path); got != want {
|
||||
t.Errorf("otherGOOSFile(%q) = %v, want %v", path, got, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestRunCorpusAuditGoAsm covers the audit wiring end to end: a package
|
||||
// beside its kernel, the kernel living off the generated defines, and the
|
||||
// histogram recording a generation failure as its own reason.
|
||||
func TestRunCorpusAuditGoAsm(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
write := func(name, src string) {
|
||||
t.Helper()
|
||||
if err := os.WriteFile(filepath.Join(dir, name), []byte(src), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
write("pkg.go", `package corpus
|
||||
|
||||
const pageSize = 4096
|
||||
|
||||
type header struct {
|
||||
magic uint64
|
||||
flags uint64
|
||||
}
|
||||
`)
|
||||
write("kern_amd64.s", "#include \"go_asm.h\"\nTEXT \xc2\xb7f(SB), NOSPLIT, $0-16\n\tMOVQ\t$const_pageSize, AX\n\tMOVQ\t$header__size, BX\n\tRET\n")
|
||||
// The defines live in the file's own package; a kernel in a directory
|
||||
// without Go files has no package to generate from.
|
||||
if err := os.MkdirAll(filepath.Join(dir, "sub"), 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
write(filepath.Join("sub", "lonely_arm64.s"), "#include \"go_asm.h\"\nTEXT \xc2\xb7g(SB), NOSPLIT, $0-0\n\tRET\n")
|
||||
|
||||
stats, err := runCorpusAudit(dir, nil)
|
||||
if err != nil {
|
||||
t.Fatalf("runCorpusAudit: %v", err)
|
||||
}
|
||||
get := func(name string) *corpusTally {
|
||||
for i, tg := range stats.targets {
|
||||
if tg.name == name {
|
||||
return stats.tallies[i]
|
||||
}
|
||||
}
|
||||
t.Fatalf("no tally for %s", name)
|
||||
return nil
|
||||
}
|
||||
if a := get("amd64"); a.attempted != 1 || a.assembled != 1 {
|
||||
t.Errorf("amd64 = %d/%d, want 1/1", a.assembled, a.attempted)
|
||||
}
|
||||
// lonely_arm64.s is an arm64 file whose package cannot be generated.
|
||||
if a := get("arm64"); a.attempted != 1 || a.assembled != 0 {
|
||||
t.Errorf("arm64 = %d/%d, want 0/1", a.assembled, a.attempted)
|
||||
}
|
||||
if r := get("arm64").reasons["go_asm.h generation failed"]; r != 1 {
|
||||
t.Errorf("arm64 go_asm.h failure count = %d, want 1", r)
|
||||
}
|
||||
}
|
||||
|
||||
// TestRunCorpusAuditGOOS covers the filename-derived GOOS end to end: a
|
||||
// kernel whose name names darwin must have its header type-checked with
|
||||
// GOOS=darwin, so the darwin-only constant it offsets with is defined. The
|
||||
// operand mirrors sys_darwin_arm64.s's trampoline, where a missing define
|
||||
// leaves an unexpanded symbol in the offset and fails.
|
||||
func TestRunCorpusAuditGOOS(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
write := func(name, src string) {
|
||||
t.Helper()
|
||||
if err := os.WriteFile(filepath.Join(dir, name), []byte(src), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
write("pkg.go", "package corpus\n")
|
||||
write("plat_darwin.go", "//go:build darwin\n\npackage corpus\n\nconst trampolineNumer = 8\n")
|
||||
write("kern_darwin_arm64.s", "#include \"go_asm.h\"\n"+
|
||||
"GLOBL timebase<>(SB), NOPTR, $16\n"+
|
||||
"TEXT \xc2\xb7g(SB), NOSPLIT, $0-0\n"+
|
||||
"\tMOVD\ttimebase<>+const_trampolineNumer(SB), R0\n"+
|
||||
"\tRET\n")
|
||||
|
||||
stats, err := runCorpusAudit(dir, nil)
|
||||
if err != nil {
|
||||
t.Fatalf("runCorpusAudit: %v", err)
|
||||
}
|
||||
var arm *corpusTally
|
||||
for i, tg := range stats.targets {
|
||||
if tg.name == "arm64" {
|
||||
arm = stats.tallies[i]
|
||||
}
|
||||
}
|
||||
if arm == nil {
|
||||
t.Fatal("no arm64 tally")
|
||||
}
|
||||
if arm.attempted != 1 || arm.assembled != 1 {
|
||||
t.Errorf("arm64 = %d/%d, want 1/1; reasons: %v", arm.assembled, arm.attempted, arm.reasons)
|
||||
}
|
||||
}
|
||||
|
||||
// TestGenerateGoAsmHeaderRuntime pins the generator against the real thing:
|
||||
// the runtime package of the ambient toolchain, whose header the toolchain's
|
||||
// own -asmhdr output was sampled from. Skipped in short mode: it type-checks
|
||||
// the whole package. The GOROOT comes from the go command itself, so the
|
||||
// test follows whatever toolchain the host provides.
|
||||
func TestGenerateGoAsmHeaderRuntime(t *testing.T) {
|
||||
if testing.Short() {
|
||||
t.Skip("type-checks the whole runtime package")
|
||||
}
|
||||
out, err := exec.Command("go", "env", "GOROOT").Output()
|
||||
if err != nil {
|
||||
t.Skipf("no Go toolchain: %v", err)
|
||||
}
|
||||
runtimeDir := filepath.Join(strings.TrimSpace(string(out)), "src", "runtime")
|
||||
dir, err := generateGoAsmHeader(runtimeDir, "", "amd64", t.TempDir())
|
||||
if err != nil {
|
||||
t.Fatalf("generateGoAsmHeader(runtime): %v", err)
|
||||
}
|
||||
b, err := os.ReadFile(dir + "/go_asm.h")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
hdr := string(b)
|
||||
for _, want := range []string{
|
||||
"#define const_hashSize 8\n",
|
||||
"#define const_avxSupported 1\n",
|
||||
"#define const_pageSize 8192\n",
|
||||
"#define g_stackguard0 16\n",
|
||||
"#define m__size ",
|
||||
} {
|
||||
if !strings.Contains(hdr, want) {
|
||||
t.Errorf("runtime header misses %q", want)
|
||||
}
|
||||
}
|
||||
}
|
||||
+221
-21
@@ -5,6 +5,7 @@ package main
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"maps"
|
||||
"os"
|
||||
"os/exec"
|
||||
"path/filepath"
|
||||
@@ -37,7 +38,7 @@ import (
|
||||
// construction and are excluded from the diff; the other architectures list
|
||||
// their conditional branches outright.
|
||||
func cmdAuditInstructions(args []string) error {
|
||||
fs := newCommand("audit-instructions", "gasm audit-instructions [--corpus [dir]] [amd64|arm64|riscv64|loong64]", `
|
||||
fs := newCommand("audit-instructions", "gasm audit-instructions [--corpus [dir]] [--list] [-I dir] [amd64|arm64|riscv64|loong64]", `
|
||||
Compare the gasm encoder for the given architecture (default amd64) against
|
||||
go tool asm and print the diff: superset encodings (gasm-only, shippable via
|
||||
gasm asm --format goobj) and known-but-unencodable names (the backlog). The
|
||||
@@ -54,14 +55,19 @@ toolchain probing. A file whose name carries a recognisable _arch suffix is
|
||||
attempted for that architecture; a file without one is attempted for all
|
||||
four, exactly as a GOARCH build would compile it. The report gives the
|
||||
per-architecture pass rates and the most common failure reasons, which drive
|
||||
the encodability backlog by frequency rather than by table order.
|
||||
the encodability backlog by frequency rather than by table order. With
|
||||
-list the report also prints every failing file with its reason, per
|
||||
architecture.
|
||||
`)
|
||||
corpus := fs.Bool("corpus", false, "assemble a corpus of .s files and report pass rates and failure reasons")
|
||||
list := fs.Bool("list", false, "with --corpus, list every failing file with its reason, per architecture")
|
||||
var dirs includeDirs
|
||||
fs.Var(&dirs, "I", "directory to search for #include files (may be repeated)")
|
||||
if err := fs.Parse(args); err != nil {
|
||||
return err
|
||||
}
|
||||
if *corpus {
|
||||
return cmdAuditCorpus(fs.Args())
|
||||
return cmdAuditCorpus(fs.Args(), dirs, *list)
|
||||
}
|
||||
archName := "amd64"
|
||||
switch n := len(fs.Args()); {
|
||||
@@ -286,6 +292,11 @@ func probeShapes(a arch.Arch) []string {
|
||||
"V1.B16, [V2.B16], V3.B16", "V1.B8, [V2.B16, V3.B16], V4.B8",
|
||||
"$4, V1.B16, V2.B16, V3.B16", "$15, V1", "V1, V2, p2",
|
||||
"R0, R1, $1, $4, p2",
|
||||
// The landing-pad kind, the compiler's PCDATA
|
||||
// bookkeeping and the four-operand bitfield
|
||||
// insert/extract family, as the toolchain's own
|
||||
// testdata spells them.
|
||||
"C", "$1, $0", "$0, R1, $1, R2",
|
||||
}
|
||||
case arch.RISCV:
|
||||
return []string{
|
||||
@@ -305,6 +316,9 @@ func probeShapes(a arch.Arch) []string {
|
||||
"X5, X6, p2", "R5, R6, p2",
|
||||
"X5, E8, M8, TA, MA, X6", "$4, E32, M1, TA, MA, X1",
|
||||
"(X5), X6, V1, V2",
|
||||
// The CSR immediate forms the toolchain's testdata spells:
|
||||
// immediate, CSR name, destination.
|
||||
"$2, TIME, X5",
|
||||
"",
|
||||
}
|
||||
case arch.LOONG64:
|
||||
@@ -321,6 +335,12 @@ func probeShapes(a arch.Arch) []string {
|
||||
"V1, V2, V3", "X1, X2, X3", "V1, V2", "X1, X2", "V1", "X1",
|
||||
// The vector compare-to-flag forms land in an FCC register.
|
||||
"V1, FCC0", "X1, FCC0",
|
||||
// The compiler's bookkeeping pair and the raw spellings the
|
||||
// toolchain's own testdata carries: JIRL rd, rj, offset (the
|
||||
// form RET lowers to), the prefetch with a 32-bit address and
|
||||
// hint, and the byte-shuffle quads.
|
||||
"$1, $0", "R1, R5, 0", "0(R7), $5, $0", "(R7), $5, $0",
|
||||
"V1, V2, V3, V4", "X1, X2, X3, X4",
|
||||
"",
|
||||
}
|
||||
}
|
||||
@@ -386,17 +406,30 @@ type corpusTally struct {
|
||||
assembled int
|
||||
reasons map[string]int // failure reason → count
|
||||
example map[string]string // failure reason → one representative file
|
||||
fails []corpusFailure // every failure, in file order, for --list
|
||||
}
|
||||
|
||||
func (t *corpusTally) fail(path, reason string) {
|
||||
// corpusFailure is one failed attempt, recorded for the --list report.
|
||||
type corpusFailure struct {
|
||||
path string
|
||||
reason string
|
||||
detail string
|
||||
}
|
||||
|
||||
func (t *corpusTally) fail(path string, err error) {
|
||||
reason := corpusReason(err)
|
||||
t.reasons[reason]++
|
||||
if t.example[reason] == "" {
|
||||
t.example[reason] = path
|
||||
}
|
||||
t.fails = append(t.fails, corpusFailure{path: path, reason: reason, detail: firstLine(err.Error())})
|
||||
}
|
||||
|
||||
// cmdAuditCorpus implements audit-instructions --corpus.
|
||||
func cmdAuditCorpus(args []string) error {
|
||||
// cmdAuditCorpus implements audit-instructions --corpus. The include
|
||||
// directories carry #include resolution over a corpus whose files refer to
|
||||
// headers such as GOROOT/pkg/include, the same -I a toolchain comparison
|
||||
// needs.
|
||||
func cmdAuditCorpus(args []string, dirs includeDirs, list bool) error {
|
||||
if len(args) > 1 {
|
||||
return &usageError{fmt.Errorf("audit-instructions --corpus takes at most one directory argument")}
|
||||
}
|
||||
@@ -410,11 +443,31 @@ func cmdAuditCorpus(args []string) error {
|
||||
}
|
||||
root = filepath.Join(strings.TrimSpace(string(out)), "src")
|
||||
}
|
||||
stats, err := runCorpusAudit(root)
|
||||
// The toolchain's shipped headers (funcdata.h and friends) define the
|
||||
// macros GOROOT files include; a corpus audit measures those files, so
|
||||
// the header directory joins the search path automatically. go_asm.h
|
||||
// is compiler-generated per package, so it is not resolved from here:
|
||||
// files that include it get one generated per target architecture,
|
||||
// which runCorpusAudit arranges.
|
||||
if out, err := exec.Command("go", "env", "GOROOT").Output(); err == nil {
|
||||
pkgInclude := filepath.Join(strings.TrimSpace(string(out)), "pkg", "include")
|
||||
if fi, err := os.Stat(pkgInclude); err == nil && fi.IsDir() {
|
||||
seen := false
|
||||
for _, d := range dirs {
|
||||
if d == pkgInclude {
|
||||
seen = true
|
||||
}
|
||||
}
|
||||
if !seen {
|
||||
dirs = append(dirs, pkgInclude)
|
||||
}
|
||||
}
|
||||
}
|
||||
stats, err := runCorpusAudit(root, dirs)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
printCorpusStats(stats)
|
||||
printCorpusStats(stats, list)
|
||||
return nil
|
||||
}
|
||||
|
||||
@@ -435,13 +488,19 @@ type corpusStats struct {
|
||||
// set, even when gasm does not support the architecture.
|
||||
var goPortSuffixes = []string{
|
||||
"386", "amd64", "arm", "arm64", "loong64", "mips", "mips64",
|
||||
"mips64le", "mipsle", "ppc64", "ppc64le", "riscv", "riscv64",
|
||||
"s390x", "wasm",
|
||||
"mips64le", "mipsle", "mips64x", "mipsx", "ppc64", "ppc64le",
|
||||
"ppc64x", "riscv", "riscv64", "s390x", "wasm",
|
||||
}
|
||||
|
||||
// otherPortFile reports whether the file's name carries a Go-architecture
|
||||
// suffix gasm does not support.
|
||||
// otherPortFile reports whether the file belongs to a build no supported
|
||||
// target ever compiles: either its name carries a Go-architecture suffix
|
||||
// gasm does not support, or, for a file with no architecture suffix at all,
|
||||
// it names another GOOS, which go/build drops from the file set
|
||||
// (rt0_js_wasm.s is a javascript build, not a generic one).
|
||||
func otherPortFile(path string) bool {
|
||||
if otherGOOSFile(path) {
|
||||
return true
|
||||
}
|
||||
base := path
|
||||
if i := strings.LastIndexByte(base, '/'); i >= 0 {
|
||||
base = base[i+1:]
|
||||
@@ -454,7 +513,67 @@ func otherPortFile(path string) bool {
|
||||
return false
|
||||
}
|
||||
|
||||
func runCorpusAudit(root string) (*corpusStats, error) {
|
||||
// goOSNames are the GOOS values go/build recognises in file names.
|
||||
var goOSNames = map[string]bool{
|
||||
"aix": true, "android": true, "darwin": true, "dragonfly": true,
|
||||
"freebsd": true, "hurd": true, "illumos": true, "ios": true,
|
||||
"js": true, "linux": true, "nacl": true, "netbsd": true,
|
||||
"openbsd": true, "plan9": true, "solaris": true, "wasip1": true,
|
||||
"windows": true, "zos": true,
|
||||
}
|
||||
|
||||
// resolveGOOS validates a -GOOS flag value, mirroring the architecture
|
||||
// check's surface: a usage error naming what the tool accepts.
|
||||
func resolveGOOS(name string) (string, error) {
|
||||
lower := strings.ToLower(name)
|
||||
if goOSNames[lower] {
|
||||
return lower, nil
|
||||
}
|
||||
return "", &usageError{fmt.Errorf("unknown GOOS %q: want one of %s", name, strings.Join(slices.Sorted(maps.Keys(goOSNames)), ", "))}
|
||||
}
|
||||
|
||||
// goosFromFilename returns the GOOS the file's name carries, by go/build's
|
||||
// goodOSArchFile rule: the GOOS segment sits last, or last before the
|
||||
// architecture segment (sys_darwin_arm64.s, vlop_arm.s carries none). An
|
||||
// empty result means the name names no GOOS and the ambient one applies.
|
||||
func goosFromFilename(path string) string {
|
||||
base := path
|
||||
if i := strings.LastIndexByte(base, '/'); i >= 0 {
|
||||
base = base[i+1:]
|
||||
}
|
||||
base = strings.TrimSuffix(base, ".s")
|
||||
// go/build ignores everything before the first underscore, so a GOOS
|
||||
// segment is only ever looked for from there on.
|
||||
i := strings.IndexByte(base, '_')
|
||||
if i < 0 {
|
||||
return ""
|
||||
}
|
||||
segs := strings.Split(base[i:], "_")
|
||||
if n := len(segs); n >= 2 && goOSNames[segs[n-2]] && slices.Contains(goPortSuffixes, segs[n-1]) {
|
||||
return segs[n-2]
|
||||
}
|
||||
if goOSNames[segs[len(segs)-1]] {
|
||||
return segs[len(segs)-1]
|
||||
}
|
||||
return ""
|
||||
}
|
||||
|
||||
// otherGOOSFile reports whether the file's name names a GOOS other than the
|
||||
// host's, by go/build's file-name rules.
|
||||
func otherGOOSFile(path string) bool {
|
||||
base := path
|
||||
if i := strings.LastIndexByte(base, '/'); i >= 0 {
|
||||
base = base[i+1:]
|
||||
}
|
||||
for seg := range strings.SplitSeq(strings.TrimSuffix(base, ".s"), "_") {
|
||||
if goOSNames[seg] && seg != runtime.GOOS {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
func runCorpusAudit(root string, dirs includeDirs) (*corpusStats, error) {
|
||||
files, err := asmFiles(root)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
@@ -474,12 +593,26 @@ func runCorpusAudit(root string) (*corpusStats, error) {
|
||||
// its name allows assembles it.
|
||||
full, generic, otherPort := 0, 0, 0
|
||||
|
||||
// Header generation is created on first use, so a corpus with no
|
||||
// go_asm.h includes never pays for a temp directory.
|
||||
var hdr *asmhdrCache
|
||||
defer func() {
|
||||
if hdr != nil {
|
||||
hdr.close()
|
||||
}
|
||||
}()
|
||||
|
||||
for _, path := range files {
|
||||
src, err := readSource(path)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
f, errs := parser.Parse(path, src)
|
||||
|
||||
// The GOOS the header generation type-checks under follows the
|
||||
// file's name when the name carries one; the ambient GOOS is the
|
||||
// honest guess otherwise (a build tag naming another GOOS is
|
||||
// invisible to a file-name rule).
|
||||
goos := goosFromFilename(path)
|
||||
|
||||
var wanted []int // indexes into targets
|
||||
if a := arch.FromFilename(path); a != arch.Unknown {
|
||||
@@ -490,10 +623,11 @@ func runCorpusAudit(root string) (*corpusStats, error) {
|
||||
}
|
||||
} else if otherPortFile(path) {
|
||||
// A file named for a Go port gasm does not support (arm,
|
||||
// 386, s390x, ...) is compiled by no supported-arch build,
|
||||
// so it is neither generic nor a per-arch attempt: counting
|
||||
// it as generic would make the headline unreachably low
|
||||
// for reasons no supported target can fix.
|
||||
// 386, s390x, ...) or for another GOOS is compiled by no
|
||||
// supported-arch build, so it is neither generic nor a
|
||||
// per-arch attempt: counting it as generic would make the
|
||||
// headline unreachably low for reasons no supported target
|
||||
// can fix.
|
||||
otherPort++
|
||||
} else {
|
||||
generic++
|
||||
@@ -502,19 +636,76 @@ func runCorpusAudit(root string) (*corpusStats, error) {
|
||||
}
|
||||
}
|
||||
|
||||
// A file that includes go_asm.h parses against a per-target header:
|
||||
// the defines differ per architecture (internal/cpu's layout, for
|
||||
// one) and per GOOS (sys_darwin_arm64.s's trampoline constants,
|
||||
// for another), so the parse cannot be shared the way a
|
||||
// header-free file's can. A generation failure is a failure for
|
||||
// every target, named for the package rather than a bare "include
|
||||
// not found". A header already resolvable in the package
|
||||
// directory or the -I list is left alone.
|
||||
if len(wanted) > 0 && needsGoAsmHeader(src) && !goAsmHeaderResolved(filepath.Dir(path), dirs) {
|
||||
if hdr == nil {
|
||||
if hdr, err = newAsmhdrCache(); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
}
|
||||
pkgDir := filepath.Dir(path)
|
||||
ok := true
|
||||
for _, i := range wanted {
|
||||
tg, t := targets[i], tallies[i]
|
||||
t.attempted++
|
||||
hdrDir, err := hdr.dirFor(pkgDir, goos, goarchName(tg.a))
|
||||
if err != nil {
|
||||
ok = false
|
||||
t.fail(path, err)
|
||||
continue
|
||||
}
|
||||
f, errs := parser.ParseWithOptions(path, src, parser.Options{
|
||||
Expand: true,
|
||||
IncludeDirs: append(slices.Clone(dirs), hdrDir),
|
||||
Predefines: platformPredefinesFor(goarchName(tg.a), goos),
|
||||
})
|
||||
if len(errs) > 0 {
|
||||
ok = false
|
||||
t.fail(path, errs[0])
|
||||
continue
|
||||
}
|
||||
if _, err := assembleFile(tg.a, f, goos); err != nil {
|
||||
ok = false
|
||||
t.fail(path, err)
|
||||
continue
|
||||
}
|
||||
t.assembled++
|
||||
}
|
||||
if ok && len(wanted) > 0 {
|
||||
full++
|
||||
}
|
||||
continue
|
||||
}
|
||||
|
||||
ok := true
|
||||
for _, i := range wanted {
|
||||
tg, t := targets[i], tallies[i]
|
||||
t.attempted++
|
||||
// The parse carries the target's platform predefines, so it
|
||||
// cannot be shared across targets the way a header-free file's
|
||||
// could: a #ifdef GOARCH_arm block must be live on arm64 and
|
||||
// dead everywhere else.
|
||||
f, errs := parser.ParseWithOptions(path, src, parser.Options{
|
||||
Expand: true,
|
||||
IncludeDirs: dirs,
|
||||
Predefines: platformPredefinesFor(goarchName(tg.a), goos),
|
||||
})
|
||||
var err error
|
||||
if len(errs) > 0 {
|
||||
err = errs[0] // a parse failure is a failure for every target
|
||||
} else {
|
||||
_, err = assembleFile(tg.a, f)
|
||||
_, err = assembleFile(tg.a, f, goos)
|
||||
}
|
||||
if err != nil {
|
||||
ok = false
|
||||
t.fail(path, corpusReason(err))
|
||||
t.fail(path, err)
|
||||
continue
|
||||
}
|
||||
t.assembled++
|
||||
@@ -536,7 +727,7 @@ func runCorpusAudit(root string) (*corpusStats, error) {
|
||||
}
|
||||
|
||||
// printCorpusStats renders the corpus audit report.
|
||||
func printCorpusStats(s *corpusStats) {
|
||||
func printCorpusStats(s *corpusStats, list bool) {
|
||||
fmt.Printf("corpus %s: %d files (%d generic, attempted for all architectures; %d named for other Go ports, never attempted)\n", s.root, s.files, s.generic, s.otherPort)
|
||||
// The rate is over the files a supported build would attempt: the
|
||||
// other ports' files sit in the count for completeness but can never
|
||||
@@ -551,6 +742,13 @@ func printCorpusStats(s *corpusStats) {
|
||||
fmt.Printf(" %4d %s\n", t.reasons[r], r)
|
||||
fmt.Printf(" e.g. %s\n", t.example[r])
|
||||
}
|
||||
if !list {
|
||||
continue
|
||||
}
|
||||
for _, f := range t.fails {
|
||||
fmt.Printf(" FAIL %s\n", f.path)
|
||||
fmt.Printf(" %s: %s\n", f.reason, f.detail)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -558,6 +756,8 @@ func printCorpusStats(s *corpusStats) {
|
||||
func corpusReason(err error) string {
|
||||
msg := err.Error()
|
||||
switch {
|
||||
case strings.Contains(msg, "go_asm.h for GOARCH"):
|
||||
return "go_asm.h generation failed"
|
||||
case strings.Contains(msg, "unsupported"), strings.Contains(msg, "cannot encode"):
|
||||
return "instruction not encodable"
|
||||
case strings.Contains(msg, "undefined label"):
|
||||
|
||||
+1
-1
@@ -84,7 +84,7 @@ func disSource(path string, target arch.Arch) int {
|
||||
if len(errs) > 0 {
|
||||
return 1
|
||||
}
|
||||
img, err := assembleFile(target, f)
|
||||
img, err := assembleFile(target, f, "")
|
||||
if err != nil {
|
||||
fmt.Fprintf(os.Stderr, "gasm dis: %v\n", err)
|
||||
return 1
|
||||
|
||||
+94
-20
@@ -240,6 +240,16 @@ func readSource(path string) (string, error) {
|
||||
return string(b), err
|
||||
}
|
||||
|
||||
// includeDirs collects repeatable -I flags: the directories searched for
|
||||
// #include files during macro expansion and include splicing.
|
||||
type includeDirs []string
|
||||
|
||||
func (d *includeDirs) String() string { return strings.Join(*d, ",") }
|
||||
func (d *includeDirs) Set(v string) error {
|
||||
*d = append(*d, v)
|
||||
return nil
|
||||
}
|
||||
|
||||
func cmdTokens(args []string) int {
|
||||
fs := newCommand("tokens", "gasm tokens <file>", `
|
||||
Print the lexical token stream of FILE: position, token kind and text, one
|
||||
@@ -476,7 +486,7 @@ hover, document symbols, diagnostics and semantic-token highlighting.
|
||||
}
|
||||
|
||||
func cmdAsm(args []string) int {
|
||||
fs := newCommand("asm", "gasm asm [--format raw|elf|goobj] [-p pkg] [-GOARCH arch] [-o out] <file>", `
|
||||
fs := newCommand("asm", "gasm asm [--format raw|elf|goobj] [-I dir] [-p pkg] [-GOARCH arch] [-GOOS os] [-o out] <file>", `
|
||||
Assemble FILE without the Go toolchain: every TEXT function is encoded to
|
||||
machine code and printed as a hex dump. Supported architectures: amd64
|
||||
(including VEX/AVX2 and EVEX/AVX-512), arm64 (AArch64 integer, FP,
|
||||
@@ -493,14 +503,26 @@ system toolchain; goobj emits the Go toolchain's own object format, which
|
||||
cmd/link consumes directly (it requires -p, the package path, and the
|
||||
installed Go toolchain: the object preamble is captured from go tool asm
|
||||
and the format version from go version).
|
||||
|
||||
A file that includes go_asm.h gets that header generated automatically from
|
||||
the package it lives in (the .go files beside it, type-checked for the
|
||||
target architecture, the toolchain's own defines), so GOROOT assembly
|
||||
assembles without a compiler. -GOOS selects the type-checking GOOS for
|
||||
that header: a GOOS-specific file (sys_darwin_arm64.s) needs its platform's
|
||||
defines, which a header from the ambient GOOS silently omits. A package
|
||||
that has no Go files for the target or does not type-check is a hard error
|
||||
naming the package.
|
||||
`)
|
||||
out := fs.String("o", "", "write the output to this file")
|
||||
format := fs.String("format", "raw", "output format: raw (concatenated image), elf or goobj (Go object)")
|
||||
pkg := fs.String("p", "", "package path for --format goobj (qualifies the exported symbols)")
|
||||
archName := fs.String("GOARCH", "", "target architecture: amd64, arm64, riscv64 or loong64 (overrides the file-name suffix)")
|
||||
goosName := fs.String("GOOS", "", "operating system for go_asm.h generation: a GOOS go/build recognises (default: the host's)")
|
||||
var dirs includeDirs
|
||||
fs.Var(&dirs, "I", "directory to search for #include files (may be repeated)")
|
||||
fs.Parse(args)
|
||||
if fs.NArg() != 1 {
|
||||
fmt.Fprintln(os.Stderr, "usage: gasm asm [--format raw|elf|goobj] [-p pkg] [-GOARCH arch] [-o out] <file>")
|
||||
fmt.Fprintln(os.Stderr, "usage: gasm asm [--format raw|elf|goobj] [-I dir] [-p pkg] [-GOARCH arch] [-GOOS os] [-o out] <file>")
|
||||
return 2
|
||||
}
|
||||
// The format is validated before anything else, so a bogus value exits 2
|
||||
@@ -521,12 +543,38 @@ and the format version from go version).
|
||||
}
|
||||
targetArch = a
|
||||
}
|
||||
goos := ""
|
||||
if *goosName != "" {
|
||||
g, err := resolveGOOS(*goosName)
|
||||
if err != nil {
|
||||
fmt.Fprintf(os.Stderr, "gasm asm: %v\n", err)
|
||||
return 2
|
||||
}
|
||||
goos = g
|
||||
}
|
||||
src, err := readSource(path)
|
||||
if err != nil {
|
||||
fmt.Fprintln(os.Stderr, "gasm:", err)
|
||||
return 1
|
||||
}
|
||||
f, errs := parser.Parse(path, src)
|
||||
// A file that includes go_asm.h cannot assemble without the package's
|
||||
// defines, and without a compiler nothing else has generated them, so
|
||||
// gasm produces the equivalent itself: automatic, because the compiler
|
||||
// behaves the same way and a flag would only ever be forgotten. A
|
||||
// generation failure is fatal and names the package: assembling against
|
||||
// a missing header would fail later with a bare "undefined" instead.
|
||||
// A go_asm.h that already resolves (placed by hand, or passed with -I)
|
||||
// is left alone.
|
||||
if needsGoAsmHeader(src) && !goAsmHeaderResolved(filepath.Dir(path), dirs) {
|
||||
hdrDir, cleanup, err := ensureGoAsmHeader(path, targetArch, goos, nil)
|
||||
if err != nil {
|
||||
fmt.Fprintln(os.Stderr, "gasm asm:", err)
|
||||
return 1
|
||||
}
|
||||
defer cleanup()
|
||||
dirs = append(dirs, hdrDir)
|
||||
}
|
||||
f, errs := parser.ParseWithOptions(path, src, parser.Options{Expand: true, IncludeDirs: dirs, Predefines: platformPredefinesFor(string(targetArch), goos)})
|
||||
for _, e := range errs {
|
||||
fmt.Fprintf(os.Stderr, "%s: %v\n", path, e)
|
||||
}
|
||||
@@ -534,7 +582,7 @@ and the format version from go version).
|
||||
return 1
|
||||
}
|
||||
|
||||
img, err := assembleFile(targetArch, f)
|
||||
img, err := assembleFile(targetArch, f, goos)
|
||||
if err != nil {
|
||||
fmt.Fprintf(os.Stderr, "%s: %v\n", path, err)
|
||||
return 1
|
||||
@@ -634,7 +682,7 @@ and the format version from go version).
|
||||
|
||||
// cmdDiff compares the machine code of two assembly files.
|
||||
func cmdDiff(args []string) int {
|
||||
set := newCommand("diff", "gasm diff [-GOARCH arch] <file1.s> <file2.s>", `
|
||||
set := newCommand("diff", "gasm diff [-GOARCH arch] [-I dir] <file1.s> <file2.s>", `
|
||||
Compare the machine code produced by assembling two files.
|
||||
Shows which functions differ and the byte-level differences.
|
||||
Useful for verifying that two implementations produce identical code,
|
||||
@@ -645,9 +693,11 @@ e.g. --map wideCopyAVX2=wideCopyAVX512 pairs the two regardless of suffix.
|
||||
`)
|
||||
mapSpec := set.String("map", "", "comma-separated old=new pairs to match functions with different names")
|
||||
archName := set.String("GOARCH", "", "target architecture for both files: amd64, arm64, riscv64 or loong64")
|
||||
var dirs includeDirs
|
||||
set.Var(&dirs, "I", "directory to search for #include files (may be repeated)")
|
||||
set.Parse(args)
|
||||
if set.NArg() != 2 {
|
||||
fmt.Fprintln(os.Stderr, "usage: gasm diff [-GOARCH arch] <file1.s> <file2.s>")
|
||||
fmt.Fprintln(os.Stderr, "usage: gasm diff [-GOARCH arch] [-I dir] <file1.s> <file2.s>")
|
||||
return 2
|
||||
}
|
||||
path1, path2 := set.Arg(0), set.Arg(1)
|
||||
@@ -675,12 +725,12 @@ e.g. --map wideCopyAVX2=wideCopyAVX512 pairs the two regardless of suffix.
|
||||
}
|
||||
|
||||
// Assemble both files.
|
||||
img1, err := assemblePath(path1, forced)
|
||||
img1, err := assemblePath(path1, forced, dirs)
|
||||
if err != nil {
|
||||
fmt.Fprintf(os.Stderr, "gasm diff: %s: %v\n", path1, err)
|
||||
return 1
|
||||
}
|
||||
img2, err := assemblePath(path2, forced)
|
||||
img2, err := assemblePath(path2, forced, dirs)
|
||||
if err != nil {
|
||||
fmt.Fprintf(os.Stderr, "gasm diff: %s: %v\n", path2, err)
|
||||
return 1
|
||||
@@ -739,11 +789,34 @@ e.g. --map wideCopyAVX2=wideCopyAVX512 pairs the two regardless of suffix.
|
||||
return 1
|
||||
}
|
||||
|
||||
// platformPredefines mirrors the go command's assembler invocation, which
|
||||
// defines GOOS_<goos> and GOARCH_<arch> as -D macros: GOROOT headers
|
||||
// (go_tls.h, asm_riscv64.h) select their platform blocks with #ifdef on
|
||||
// exactly those names, so an assembler without them cannot see the platform
|
||||
// definitions at all.
|
||||
func platformPredefines(goarch, goos string) map[string]string {
|
||||
return map[string]string{
|
||||
"GOARCH_" + goarch: "1",
|
||||
"GOOS_" + goos: "1",
|
||||
}
|
||||
}
|
||||
|
||||
// platformPredefinesFor resolves the ambient GOOS the way a build would: a
|
||||
// file whose name carries one (sys_darwin_arm64.s) is compiled for that GOOS
|
||||
// and nothing else.
|
||||
func platformPredefinesFor(goarch string, fileGoos string) map[string]string {
|
||||
goos := fileGoos
|
||||
if goos == "" {
|
||||
goos = runtime.GOOS
|
||||
}
|
||||
return platformPredefines(goarch, goos)
|
||||
}
|
||||
|
||||
// assembleFile assembles a parsed file for the given architecture and returns the image.
|
||||
func assembleFile(targetArch arch.Arch, f *ast.File) (*asm.Image, error) {
|
||||
func assembleFile(targetArch arch.Arch, f *ast.File, goos string) (*asm.Image, error) {
|
||||
switch targetArch {
|
||||
case arch.AMD64:
|
||||
return asm.AssembleFile(f)
|
||||
return asm.AssembleFile(f, asm.WithGOOS(goos))
|
||||
case arch.RISCV:
|
||||
return asm.AssembleFileRISCV(f)
|
||||
case arch.ARM64:
|
||||
@@ -755,25 +828,26 @@ func assembleFile(targetArch arch.Arch, f *ast.File) (*asm.Image, error) {
|
||||
}
|
||||
}
|
||||
|
||||
// assemblePath reads, parses and assembles a file (used by cmdDiff). A
|
||||
// non-Unknown forced architecture overrides the file-name suffix.
|
||||
func assemblePath(path string, forced arch.Arch) (*asm.Image, error) {
|
||||
// assemblePath reads, preprocesses, parses and assembles a file (used by
|
||||
// cmdDiff). A non-Unknown forced architecture overrides the file-name
|
||||
// suffix.
|
||||
func assemblePath(path string, forced arch.Arch, dirs includeDirs) (*asm.Image, error) {
|
||||
src, err := readSource(path)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
f, errs := parser.Parse(path, src)
|
||||
target := forced
|
||||
if target == arch.Unknown {
|
||||
target = arch.FromFilename(path)
|
||||
}
|
||||
f, errs := parser.ParseWithOptions(path, src, parser.Options{Expand: true, IncludeDirs: dirs, Predefines: platformPredefinesFor(string(target), "")})
|
||||
for _, e := range errs {
|
||||
fmt.Fprintf(os.Stderr, "%s: %v\n", path, e)
|
||||
}
|
||||
if len(errs) > 0 {
|
||||
return nil, fmt.Errorf("parse errors")
|
||||
}
|
||||
target := forced
|
||||
if target == arch.Unknown {
|
||||
target = arch.FromFilename(path)
|
||||
}
|
||||
return assembleFile(target, f)
|
||||
return assembleFile(target, f, "")
|
||||
}
|
||||
|
||||
// printByteDiff shows the first few byte differences between two code blocks.
|
||||
@@ -869,7 +943,7 @@ func cmdVerifyNonJIT(path string, targetArch arch.Arch, groundTruth, profile boo
|
||||
if len(errs) > 0 {
|
||||
return 1
|
||||
}
|
||||
img, err := assembleFile(targetArch, f)
|
||||
img, err := assembleFile(targetArch, f, "")
|
||||
if err != nil {
|
||||
fmt.Fprintf(os.Stderr, "gasm verify: %v\n", err)
|
||||
return 1
|
||||
|
||||
@@ -403,7 +403,7 @@ func TestRunCorpusAudit(t *testing.T) {
|
||||
write("generic.s", "#include \"textflag.h\"\nTEXT ·g(SB), NOSPLIT, $0-0\n\tRET\n")
|
||||
write("broken.s", "#include \"textflag.h\"\nTEXT ·b(SB), NOSPLIT, $0-0\n\tJMP nowhere\n\tRET\n")
|
||||
|
||||
stats, err := runCorpusAudit(dir)
|
||||
stats, err := runCorpusAudit(dir, nil)
|
||||
if err != nil {
|
||||
t.Fatalf("runCorpusAudit: %v", err)
|
||||
}
|
||||
|
||||
@@ -0,0 +1,109 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
package main
|
||||
|
||||
import (
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// writeTree writes a directory of files and returns its root.
|
||||
func writeTree(t *testing.T, files map[string]string) string {
|
||||
t.Helper()
|
||||
dir := t.TempDir()
|
||||
for name, content := range files {
|
||||
path := filepath.Join(dir, name)
|
||||
if err := os.MkdirAll(filepath.Dir(path), 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := os.WriteFile(path, []byte(content), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
return dir
|
||||
}
|
||||
|
||||
// TestAsmMacroAndIncludeEndToEnd drives `gasm asm` over a source with an
|
||||
// in-file parameterised macro and an include resolved through -I, and checks
|
||||
// the assembled bytes came from the expansion (the loop body counts six
|
||||
// increments, two per expanded iteration).
|
||||
func TestAsmMacroAndIncludeEndToEnd(t *testing.T) {
|
||||
if testing.Short() {
|
||||
t.Skip("runs the assembler end to end")
|
||||
}
|
||||
dir := writeTree(t, map[string]string{
|
||||
"inc/consts.h": "#define NITER 3\n",
|
||||
"main_amd64.s": "#include \"textflag.h\"\n" +
|
||||
"#include \"consts.h\"\n" +
|
||||
"#define STEP(r) ADDQ $1, r; ADDQ $1, r\n" +
|
||||
"TEXT ·f(SB), NOSPLIT, $0-8\n" +
|
||||
"\tXORQ AX, AX\n" +
|
||||
"\tMOVQ $NITER, CX\n" +
|
||||
"loop:\n" +
|
||||
"\tSTEP(AX)\n" +
|
||||
"\tDECQ CX\n" +
|
||||
"\tJNZ loop\n" +
|
||||
"\tMOVQ AX, ret+0(FP)\n" +
|
||||
"\tRET\n",
|
||||
})
|
||||
stdout, stderr, code := capture(func() int {
|
||||
return cmdAsm([]string{"-I", filepath.Join(dir, "inc"), "-GOARCH", "amd64", filepath.Join(dir, "main_amd64.s")})
|
||||
})
|
||||
if code != 0 {
|
||||
t.Fatalf("gasm asm exited %d: %s%s", code, stdout, stderr)
|
||||
}
|
||||
// The macro expanded to two ADDQ $1 encodings in the static body; the
|
||||
// iteration count lives in the runtime loop.
|
||||
if n := strings.Count(stdout, "83 c0 01"); n != 2 {
|
||||
t.Errorf("found %d ADDQ $1 encodings in the image, want 2:\n%s", n, stdout)
|
||||
}
|
||||
}
|
||||
|
||||
// TestAsmIncludeResolutionOrder pins the -I search order end to end: the
|
||||
// including file's directory wins over the -I directories.
|
||||
func TestAsmIncludeResolutionOrder(t *testing.T) {
|
||||
if testing.Short() {
|
||||
t.Skip("runs the assembler end to end")
|
||||
}
|
||||
dir := writeTree(t, map[string]string{
|
||||
"src/main_amd64.s": "#include \"textflag.h\"\n" +
|
||||
"#include \"vals.h\"\n" +
|
||||
"TEXT ·f(SB), NOSPLIT, $0\n" +
|
||||
"\tMOVQ $VAL, AX\n" +
|
||||
"\tRET\n",
|
||||
"src/vals.h": "#define VAL 1\n",
|
||||
"late/vals.h": "#define VAL 2\n",
|
||||
"early/vals.h": "#define VAL 3\n",
|
||||
})
|
||||
stdout, stderr, code := capture(func() int {
|
||||
return cmdAsm([]string{"-I", filepath.Join(dir, "early"), "-I", filepath.Join(dir, "late"),
|
||||
"-GOARCH", "amd64", filepath.Join(dir, "src", "main_amd64.s")})
|
||||
})
|
||||
if code != 0 {
|
||||
t.Fatalf("gasm asm exited %d: %s%s", code, stdout, stderr)
|
||||
}
|
||||
// VAL came from src/vals.h, not from either -I directory: the image
|
||||
// loads the immediate 1.
|
||||
if !strings.Contains(stdout, "b8 01 00 00 00") {
|
||||
t.Errorf("expected the source-directory VAL (immediate 1) in:\n%s", stdout)
|
||||
}
|
||||
}
|
||||
|
||||
// TestAsmMissingIncludeIsAnError pins the diagnostic for an include that
|
||||
// resolves nowhere on the assembly path.
|
||||
func TestAsmMissingIncludeIsAnError(t *testing.T) {
|
||||
if testing.Short() {
|
||||
t.Skip("runs the assembler end to end")
|
||||
}
|
||||
path := writeTemp(t, "main_amd64.s", "#include \"textflag.h\"\n#include \"nothere.h\"\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n")
|
||||
_, stderr, code := capture(func() int { return cmdAsm([]string{"-GOARCH", "amd64", path}) })
|
||||
if code == 0 {
|
||||
t.Fatal("gasm asm accepted a file whose include resolves nowhere")
|
||||
}
|
||||
if !strings.Contains(stderr, `#include "nothere.h"`) {
|
||||
t.Errorf("stderr does not name the failing include: %s", stderr)
|
||||
}
|
||||
}
|
||||
@@ -119,6 +119,20 @@ identifier is a register or a label is an *architecture* question, so it is
|
||||
left to `arch` and resolved in the lint/lsp layers. This keeps the parser
|
||||
arch-agnostic and its output deterministic.
|
||||
|
||||
### Optional preprocessing
|
||||
|
||||
With `Options{Expand: true}` the parser runs a pre-parse pass
|
||||
(`preproc.go`) that splices `#include` files (the source directory, then the
|
||||
`-I` directories), expands object and parameterised `#define` macros,
|
||||
applies `#undef` and the `#ifdef`/`#ifndef`/`#else`/`#endif` family, and
|
||||
folds constant expressions left in operands. The go command's platform
|
||||
macros (`GOARCH_<arch>`, `GOOS_<goos>`) arrive through `Options.Predefines`.
|
||||
The assembly path (`asm`, `diff`, `audit`) expands; `lint`, `fmt` and the
|
||||
language server read the raw file. The command layer adds the go_asm.h
|
||||
generator (`asmhdr.go`): a file that includes go_asm.h gets the package's
|
||||
defines type-checked out of its Go files for the target architecture and
|
||||
GOOS, with no compiler in the loop.
|
||||
|
||||
### `arch`
|
||||
|
||||
Register files are generated programmatically (the regular `R8`-`R15`,
|
||||
|
||||
+26
-5
@@ -141,14 +141,16 @@ gasm lint kernel_amd64.s
|
||||
## asm
|
||||
|
||||
```text
|
||||
Usage: gasm asm [--format raw|elf|goobj] [-p pkg] [-GOARCH arch] [-o out] <file>
|
||||
Usage: gasm asm [--format raw|elf|goobj] [-I dir] [-p pkg] [-GOARCH arch] [-GOOS os] [-o out] <file>
|
||||
```
|
||||
|
||||
| Flag | Default | Effect |
|
||||
|---|---|---|
|
||||
| `-format` | `raw` | output format: `raw` (concatenated image), `elf` or `goobj` (Go object) |
|
||||
| `-I` | empty | directory to search for `#include` files; may be repeated, searched in order after the source directory |
|
||||
| `-p` | empty | package path for `--format goobj`, qualifying the exported symbols |
|
||||
| `-GOARCH` | empty | target architecture: `amd64`, `arm64`, `riscv64` or `loong64`; overrides the file-name suffix |
|
||||
| `-GOOS` | empty | operating system for the generated `go_asm.h`: any GOOS `go/build` recognises in file names; default is the host's |
|
||||
| `-o` | empty | write the output to this file instead of a hex dump on stdout |
|
||||
|
||||
Supported architectures: amd64 (VEX/AVX2 and EVEX/AVX-512 included), arm64,
|
||||
@@ -162,6 +164,20 @@ system toolchain; `goobj` emits the Go toolchain's own object format, which
|
||||
installed: the object preamble is captured from `go tool asm` and the format
|
||||
version from `go version`. `raw` and `elf` need no toolchain at all.
|
||||
|
||||
A file that includes `go_asm.h` gets that header generated from the Go
|
||||
files beside it, type-checked for the target. `-GOOS` selects the
|
||||
type-checking GOOS for that header, because a GOOS-specific file needs its
|
||||
platform's defines: `sys_darwin_arm64.s` fails against the ambient GOOS
|
||||
(`machTimebaseInfo_numer` is missing from a linux type-check) and assembles
|
||||
with `-GOOS darwin`.
|
||||
|
||||
Assembly preprocessing matches the toolchain's: `#define` macros (object and
|
||||
parameterised) expand at the point of use, `#undef`, `#ifdef`, `#ifndef`,
|
||||
`#else` and `#endif` behave as in `go tool asm`, `;` separates statements,
|
||||
and `#include "file"` splices the named file in, resolved against the source
|
||||
directory and then each `-I` directory in order. `textflag.h` is the one
|
||||
header that is not spliced: gasm consumes its flag names natively.
|
||||
|
||||
```sh
|
||||
gasm asm hello_amd64.s
|
||||
```
|
||||
@@ -305,12 +321,13 @@ gasm debug --func add --cover hello_amd64.s
|
||||
## diff
|
||||
|
||||
```text
|
||||
Usage: gasm diff [-GOARCH arch] <file1.s> <file2.s>
|
||||
Usage: gasm diff [-GOARCH arch] [-I dir] <file1.s> <file2.s>
|
||||
```
|
||||
|
||||
| Flag | Default | Effect |
|
||||
|---|---|---|
|
||||
| `-GOARCH` | empty | target architecture for both files, overriding the file-name suffixes |
|
||||
| `-I` | empty | directory to search for `#include` files; may be repeated, searched in order after the source directory |
|
||||
| `-map` | empty | comma-separated `old=new` pairs to match functions with different names |
|
||||
|
||||
Functions are paired by exact name unless `--map` says otherwise, so
|
||||
@@ -348,7 +365,7 @@ add: 16 bytes, args=24, frame=0 NOSPLIT
|
||||
## audit-instructions
|
||||
|
||||
```text
|
||||
Usage: gasm audit-instructions [--corpus [dir]] [amd64|arm64|riscv64|loong64]
|
||||
Usage: gasm audit-instructions [--corpus [dir]] [--list] [-I dir] [amd64|arm64|riscv64|loong64]
|
||||
```
|
||||
|
||||
Compare the gasm encoder for the given architecture (default amd64) against the
|
||||
@@ -381,10 +398,14 @@ With `--corpus` the audit changes shape: it assembles every `.s` file under
|
||||
DIR (default `GOROOT/src`) with the gasm encoder only, no toolchain probing.
|
||||
A file whose name carries a recognisable `_arch` suffix is attempted for that
|
||||
architecture; a file without one is attempted for all four, exactly as a
|
||||
`GOARCH` build would compile it. The report gives the headline number (files
|
||||
`GOARCH` build would compile it, and a name that names a GOOS
|
||||
(`sys_darwin_arm64.s`) type-checks its generated `go_asm.h` for that GOOS.
|
||||
The report gives the headline number (files
|
||||
that assemble for every target architecture), the per-architecture pass rates
|
||||
and the most common failure reasons with one representative file each, which
|
||||
drive the encodability backlog by frequency rather than by table order. A run
|
||||
drive the encodability backlog by frequency rather than by table order. With
|
||||
`--list` the report additionally prints every failing file with its failure
|
||||
reason, per architecture. A run
|
||||
over GOROOT takes under a second.
|
||||
|
||||
```sh
|
||||
|
||||
+657
@@ -0,0 +1,657 @@
|
||||
# The GOOBJ object file format
|
||||
|
||||
This document is a complete specification of GOOBJ, the object file format
|
||||
that the Go toolchain's assembler, compiler and linker exchange, written for
|
||||
implementers of independent producers and consumers. It documents the format
|
||||
as shipped by Go 1.27.1, identified by the magic string `"\x00go120ld"`.
|
||||
|
||||
No comparable document exists upstream. The format is defined only by the
|
||||
source of the `cmd/internal/goobj` package inside the toolchain tree, it is an
|
||||
internal interface with no stability promise, and it can change in any
|
||||
release. This specification was therefore produced by reverse engineering
|
||||
that source and by parsing real objects produced by `go tool asm` and
|
||||
`go tool compile`, byte for byte, against the layout described here. Within
|
||||
gasm-devkit it is kept honest by the differential tests in `asm/goobj_test.go`
|
||||
and `asm/link_test.go`, which compare `gasm asm --format goobj` output against
|
||||
the toolchain's own products and feed gasm objects to `go build`.
|
||||
|
||||
Every numeric value in this document, every block index, structure size, flag
|
||||
bit, type code and relocation number, was read from the Go 1.27.1 source at
|
||||
`/usr/local/go/src/cmd/internal/goobj`, `cmd/internal/obj` and
|
||||
`cmd/internal/objabi`, and exercised against assembled objects.
|
||||
|
||||
## Containers
|
||||
|
||||
The unit this document specifies is the **object**: one package's worth of
|
||||
symbols, relocations and data. An object is never consumed naked. Two
|
||||
wrappers exist in practice, and the linker dispatches on the first bytes of
|
||||
the file.
|
||||
|
||||
**The bare object**, written by `go tool asm`:
|
||||
|
||||
```text
|
||||
"go object linux amd64 go1.27.1 GOAMD64=v1 X:regabiwrappers,...\n"
|
||||
"!\n"
|
||||
<GOOBJ blob>
|
||||
```
|
||||
|
||||
The first line is the toolchain configuration string, produced by
|
||||
`objabi.HeaderString`: `go object`, the GOOS, the GOARCH, the toolchain
|
||||
version, an optional architecture qualifier such as `GOAMD64=v1`, and
|
||||
`X:` followed by the enabled experiments, comma separated. The linker requires
|
||||
this line to match its own configuration exactly and rejects the file
|
||||
otherwise; the `-f` linker flag waives the check. Header lines may be
|
||||
followed by export data delimited by `$$` markers; the header region always
|
||||
ends at the first line consisting of exactly `!`, and the GOOBJ blob starts
|
||||
immediately after that line.
|
||||
|
||||
**The package archive**, written by the compiler output pipeline and consumed
|
||||
by `go build`: the classic `ar` format, magic `!<arch>\n`, with the export
|
||||
data in a `__.PKGDEF` member and one or more objects as further members, each
|
||||
carrying the bare-object structure above. `go tool pack` creates and
|
||||
inspects such archives.
|
||||
|
||||
| Consumer | Role |
|
||||
|---|---|
|
||||
| `cmd/asm` | writes objects from `.s` files |
|
||||
| `cmd/compile` | writes objects from Go source |
|
||||
| `cmd/link` | reads objects and archives, produces executables |
|
||||
| `cmd/nm`, `cmd/objdump` | read objects through `cmd/internal/objfile` |
|
||||
|
||||
## Conventions
|
||||
|
||||
- All integers are **little endian**.
|
||||
- There is **no alignment or padding** anywhere in the file; structures follow
|
||||
one another byte by byte.
|
||||
- Every offset stored in the file is **relative to the first byte of the GOOBJ
|
||||
blob**, not to the start of the container.
|
||||
- The blob opens with a 96 byte header that carries the byte offset of every
|
||||
block. A block's length is the difference between its own offset and the
|
||||
next block's, so the offset array is the only index the format needs.
|
||||
|
||||
### Layout overview
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
A[Container header line and ! terminator] --> B[File header, 96 bytes]
|
||||
B --> C[String table, implicit region]
|
||||
C --> D[Autolib]
|
||||
D --> E[PkgIndex]
|
||||
E --> F[Files]
|
||||
F --> G[Symbol definition arrays: Symdef, Hashed64def, Hasheddef, Nonpkgdef, Nonpkgref]
|
||||
G --> H[RefFlags]
|
||||
H --> I[Hash64 and Hash]
|
||||
I --> J[RelocIndex, AuxIndex, DataIndex]
|
||||
J --> K[Relocs]
|
||||
K --> L[Aux]
|
||||
L --> M[Data]
|
||||
M --> N[RefNames]
|
||||
N --> O[BlkEnd marks the end of the blob]
|
||||
```
|
||||
|
||||
## The file header
|
||||
|
||||
Exactly 96 bytes: 8 magic, 8 fingerprint, 4 flags, and 19 four byte block
|
||||
offsets.
|
||||
|
||||
| Offset | Size | Field | Meaning |
|
||||
|---|---|---|---|
|
||||
| 0 | 8 | Magic | `"\x00go120ld"`. A reader rejects anything else. The digits are the format version and have moved before; a new toolchain release may move them again. |
|
||||
| 8 | 8 | Fingerprint | Identifies the package build. The compiler writes a hash of the export data; the assembler leaves all zero. The linker compares this against the fingerprint recorded by importers. |
|
||||
| 16 | 4 | Flags | Bit field, see below. |
|
||||
| 20 | 76 | Offsets | 19 `uint32` entries, one per block index 0 to 18. |
|
||||
|
||||
Header flags:
|
||||
|
||||
| Bit | Value | Name | Meaning |
|
||||
|---|---|---|---|
|
||||
| 0 | 1 | ObjFlagShared | built with `-shared` |
|
||||
| 1 | 2 | reserved | was `ObjFlagNeedNameExpansion`, now unused |
|
||||
| 2 | 4 | ObjFlagFromAssembly | produced from assembly source; `go tool asm` and gasm set this |
|
||||
| 3 | 8 | ObjFlagUnlinkable | package path is invalid, the linker refuses to link |
|
||||
| 4 | 16 | ObjFlagStd | standard library package |
|
||||
|
||||
### Block indices
|
||||
|
||||
The offset array is indexed by these constants, in file order:
|
||||
|
||||
| Index | Constant | Contents |
|
||||
|---|---|---|
|
||||
| 0 | BlkAutolib | imported packages |
|
||||
| 1 | BlkPkgIndex | referenced packages, indexed |
|
||||
| 2 | BlkFile | source file names |
|
||||
| 3 | BlkSymdef | symbol definitions, package scope |
|
||||
| 4 | BlkHashed64def | short hashed definitions |
|
||||
| 5 | BlkHasheddef | hashed definitions |
|
||||
| 6 | BlkNonpkgdef | non-package definitions |
|
||||
| 7 | BlkNonpkgref | non-package references |
|
||||
| 8 | BlkRefFlags | flags of referenced symbols |
|
||||
| 9 | BlkHash64 | 8 byte hashes for short hashed definitions |
|
||||
| 10 | BlkHash | 16 byte hashes for hashed definitions |
|
||||
| 11 | BlkRelocIndex | per symbol relocation start index |
|
||||
| 12 | BlkAuxIndex | per symbol aux start index |
|
||||
| 13 | BlkDataIndex | per symbol data offset |
|
||||
| 14 | BlkReloc | relocations |
|
||||
| 15 | BlkAux | aux symbol entries |
|
||||
| 16 | BlkData | symbol payloads |
|
||||
| 17 | BlkRefName | names of referenced symbols, for tools |
|
||||
| 18 | BlkEnd | no contents; its offset is the end of the blob |
|
||||
|
||||
## The string table
|
||||
|
||||
There is no block index for strings. The table occupies the implicit region
|
||||
between the end of the header (offset 96) and `Offsets[BlkAutolib]`, and every
|
||||
string offset in the file points into that region. The writer de-duplicates:
|
||||
each distinct string is stored once, in first-use order, and the empty string
|
||||
is always the first entry, so its reference is length 0 and offset 96.
|
||||
|
||||
A **string reference** is 8 bytes: `uint32` length, then `uint32` absolute
|
||||
offset of the bytes. The bytes are stored raw, with no terminator.
|
||||
|
||||
## Symbol references and the package index
|
||||
|
||||
A **symbol reference** (SymRef) is 8 bytes: two `uint32`, `PkgIdx` and
|
||||
`SymIdx`. The pair `{0, 0}` means nil. `PkgIdx` says which array the symbol
|
||||
lives in:
|
||||
|
||||
| Value | Constant | SymIdx indexes |
|
||||
|---|---|---|
|
||||
| 0 | PkgIdxInvalid | never valid in a written file |
|
||||
| 1 and up, ascending | (imported packages) | the SymbolDefs array of the package named at PkgIndex entry `PkgIdx` |
|
||||
| 0x7ffffffb | PkgIdxSelf | this object's Symdef array |
|
||||
| 0x7ffffffc | PkgIdxBuiltin | the compiler's builtin table, see Builtins |
|
||||
| 0x7ffffffd | PkgIdxHashed | this object's Hasheddef array |
|
||||
| 0x7ffffffe | PkgIdxHashed64 | this object's Hashed64def array |
|
||||
| 0x7fffffff | PkgIdxNone | NonPkgDefs, overflowing into NonPkgRefs |
|
||||
|
||||
Assignment rules, as the toolchain performs them:
|
||||
|
||||
- Every definition a package exports to the linker by index lands in Symdefs
|
||||
with PkgIdxSelf. The compiler puts its functions and data here; the
|
||||
assembler puts only its file-local static symbols here, everything else by
|
||||
name, see below.
|
||||
- External package references take indices 1, 2, 3, in order of first
|
||||
reference during assembly; the package names go into PkgIndex at those
|
||||
indices, entry 0 is the empty package and is never referenced.
|
||||
- References to the compiler's builtin functions become PkgIdxBuiltin with
|
||||
SymIdx set to the builtin's index.
|
||||
- A symbol referenced **by name** rather than by index becomes PkgIdxNone and
|
||||
its index counts through NonPkgDefs first, then continues into NonPkgRefs.
|
||||
A producer must emit the definitions it made in NonPkgDefs and the pure
|
||||
references in NonPkgRefs.
|
||||
- The assembler's rule, from `cmd/internal/obj/sym.go`: every assembly symbol
|
||||
is referenced by name, PkgIdxNone, **except** file-local static symbols,
|
||||
whose names carry `<>` and which are referenced by index. The compiler also
|
||||
forces references by name for symbols marked `//go:linkname` and for any
|
||||
symbol with the DUPOK attribute, which the linker de-duplicates by name.
|
||||
|
||||
## Symbol definition entries
|
||||
|
||||
The five definition and reference arrays (block indices 3 to 7) share one
|
||||
element layout, 21 bytes:
|
||||
|
||||
| Offset | Size | Field | Meaning |
|
||||
|---|---|---|---|
|
||||
| 0 | 8 | Name | string reference |
|
||||
| 8 | 2 | ABI | see table below |
|
||||
| 10 | 1 | Type | symbol kind, see the kind table |
|
||||
| 11 | 1 | Flag | bit field, see below |
|
||||
| 12 | 1 | Flag2 | second bit field, see below |
|
||||
| 13 | 4 | Siz | payload size in bytes, `uint32` |
|
||||
| 17 | 4 | Align | alignment the linker must honour, `uint32` |
|
||||
|
||||
The Name is a real string reference for hand-written symbols. The auxiliary
|
||||
symbols the toolchain generates per function, the FuncInfo payload, the DWARF
|
||||
entries, have empty names: length 0, and their identity is only via the Aux
|
||||
entries that point at them by index.
|
||||
|
||||
### The ABI field
|
||||
|
||||
| Value | Meaning |
|
||||
|---|---|
|
||||
| 0 | ABI0, the stack based ABI, the ABI of every hand-written assembly function |
|
||||
| 1 | ABIInternal, the register ABI of compiler-generated functions |
|
||||
| 0xffff | static, a file-local symbol (`name<>(SB)`), `SymABIstatic` |
|
||||
|
||||
### The Flag byte
|
||||
|
||||
| Bit | Value | Name | Meaning |
|
||||
|---|---|---|---|
|
||||
| 0 | 1 | SymFlagDupok | duplicates allowed, the linker merges them |
|
||||
| 1 | 2 | SymFlagLocal | file-local |
|
||||
| 2 | 4 | SymFlagTypelink | belongs in the typelink table |
|
||||
| 3 | 8 | SymFlagLeaf | leaf function |
|
||||
| 4 | 16 | SymFlagNoSplit | no stack-split preamble |
|
||||
| 5 | 32 | SymFlagReflectMethod | `//go:reflectmethod` reachability |
|
||||
| 6 | 64 | SymFlagGoType | a Go type descriptor, `type:` name and SRODATA |
|
||||
|
||||
Note that NoSplit is not reserved for explicit `NOSPLIT` declarations. On
|
||||
amd64 the assembler itself marks any function whose frame is below
|
||||
`abi.StackSmall` and whose body calls nothing that needs stack as NoSplit and
|
||||
omits the split check, so a `TEXT` without `NOSPLIT` can still carry the bit.
|
||||
|
||||
### The Flag2 byte
|
||||
|
||||
| Bit | Value | Name | Meaning |
|
||||
|---|---|---|---|
|
||||
| 0 | 1 | SymFlagUsedInIface | type or itab reachable through an interface |
|
||||
| 1 | 2 | SymFlagItab | an itab, `go:itab.` name and SRODATA |
|
||||
| 2 | 4 | SymFlagDict | a generic dictionary symbol |
|
||||
| 3 | 8 | SymFlagPkgInit | package initialisation function |
|
||||
| 4 | 16 | SymFlagLinkname | reachable through `//go:linkname`; the assembler also sets it on `main.main` |
|
||||
| 5 | 32 | SymFlagLinknameStd | linkname into the standard library |
|
||||
| 6 | 64 | SymFlagABIWrapper | ABI transition wrapper |
|
||||
| 7 | 128 | SymFlagWasmExport | `//go:wasmexport` target |
|
||||
|
||||
### The Type byte: symbol kinds
|
||||
|
||||
Values of `objabi.SymKind`, in numeric order:
|
||||
|
||||
| Value | Name | Meaning |
|
||||
|---|---|---|
|
||||
| 0 | Sxxx | invalid zero value |
|
||||
| 1 | STEXT | executable code |
|
||||
| 2 | STEXTFIPS | executable code, FIPS section |
|
||||
| 3 | SRODATA | read only data |
|
||||
| 4 | SRODATAFIPS | read only data, FIPS section |
|
||||
| 5 | SNOPTRDATA | data without pointers |
|
||||
| 6 | SNOPTRDATAFIPS | data without pointers, FIPS section |
|
||||
| 7 | SDATA | data, may contain pointers |
|
||||
| 8 | SDATAFIPS | data, FIPS section |
|
||||
| 9 | SBSS | zero initialised data |
|
||||
| 10 | SNOPTRBSS | zero initialised data without pointers |
|
||||
| 11 | STLSBSS | thread local zero initialised data |
|
||||
| 12 | SDWARFCUINFO | DWARF compile unit information |
|
||||
| 13 | SDWARFCONST | DWARF constants |
|
||||
| 14 | SDWARFFCN | DWARF function entry |
|
||||
| 15 | SDWARFABSFCN | DWARF absolute function entry |
|
||||
| 16 | SDWARFTYPE | DWARF type information |
|
||||
| 17 | SDWARFVAR | DWARF variable information |
|
||||
| 18 | SDWARFRANGE | DWARF range lists |
|
||||
| 19 | SDWARFLOC | DWARF location lists |
|
||||
| 20 | SDWARFLINES | DWARF line programs |
|
||||
| 21 | SDWARFADDR | DWARF address table |
|
||||
| 22 | SLIBFUZZER_8BIT_COUNTER | libFuzzer coverage counter |
|
||||
| 23 | SCOVERAGE_COUNTER | coverage counter |
|
||||
| 24 | SCOVERAGE_AUXVAR | coverage auxiliary variable |
|
||||
| 25 | SSEHUNWINDINFO | Windows SEH unwind information |
|
||||
|
||||
## Referenced symbol flags (RefFlags)
|
||||
|
||||
Element size 10 bytes, one per referenced external indexed symbol that
|
||||
carries a non-zero Flag2:
|
||||
|
||||
| Offset | Size | Field |
|
||||
|---|---|---|
|
||||
| 0 | 8 | Sym, a SymRef into another package |
|
||||
| 8 | 1 | Flag, always 0 in current writers |
|
||||
| 9 | 1 | Flag2, only SymFlagUsedInIface is ever written |
|
||||
|
||||
The linker uses these to preserve reachability of interface conversions
|
||||
across package boundaries. Entries with no flags are omitted entirely.
|
||||
|
||||
## Hashes
|
||||
|
||||
**Hash64**, block 9: one `uint64` per Hashed64def entry, in array order. Not
|
||||
a hash at all: the writer copies the **first 8 bytes of the symbol's
|
||||
payload**. Only symbols whose content-hash section byte is 0 may use the
|
||||
short form.
|
||||
|
||||
**Hash**, block 10: 16 bytes per Hasheddef entry: the first 16 bytes of a
|
||||
SHA-256 computation over a seed byte `0x01` followed by the hash input. The
|
||||
input, from `cmd/internal/obj/objfile.go`:
|
||||
|
||||
1. the payload size, little endian `uint64`;
|
||||
2. the section byte, one of `t` for STEXT, `f` for STEXTFIPS, `P` for pcdata,
|
||||
`F` for the `go:func.*` and `go:funcrel.*` families, `T` for `type:`
|
||||
symbols, otherwise 0;
|
||||
3. for text symbols, the symbol name, which keeps distinct functions from
|
||||
merging;
|
||||
4. the payload with trailing zero bytes trimmed;
|
||||
5. for each relocation: a 14 byte record, offset `uint32`, size `uint8`,
|
||||
low type byte `uint8`, addend `int64`, followed by an encoding of the
|
||||
target: tag byte 0 then the target's short hash, tag 1 then its full
|
||||
hash, tag 2 then its expanded name, tag 3 then its builtin index, or,
|
||||
for PkgIdxSelf and imported packages, no tag, then the package path
|
||||
and the symbol index.
|
||||
|
||||
Two symbols with equal hashes are interchangeable at link time, which is what
|
||||
makes content addressing work. A producer that computes these hashes wrongly
|
||||
produces objects that link but de-duplicate wrongly; gasm verifies them by
|
||||
byte comparison against `go tool asm`.
|
||||
|
||||
## The index arrays
|
||||
|
||||
Three arrays of `uint32`, one element per **defined** symbol plus one final
|
||||
element, in the order Symdefs, Hashed64defs, Hasheddefs, NonPkgDefs. With N
|
||||
defined symbols, each array holds N + 1 entries, and the entry at N is the
|
||||
total.
|
||||
|
||||
- RelocIndex: entry i is where symbol i's relocations start in BlkReloc;
|
||||
entry i + 1 minus entry i is its count.
|
||||
- AuxIndex: the same construction over BlkAux.
|
||||
- DataIndex: entry i is the byte offset of symbol i's payload within BlkData;
|
||||
the count is the difference of neighbours.
|
||||
|
||||
The toolchain writes relocations grouped per symbol in definition order, and
|
||||
sorts each symbol's relocations by their Off field first. A producer that
|
||||
skips the sort produces objects the linker still accepts, but that no longer
|
||||
compare byte-for-byte with the toolchain's output.
|
||||
|
||||
## Relocations
|
||||
|
||||
Element size 23 bytes:
|
||||
|
||||
| Offset | Size | Field | Meaning |
|
||||
|---|---|---|---|
|
||||
| 0 | 4 | Off | patch position, bytes from the start of the symbol's payload, `int32` |
|
||||
| 4 | 1 | Siz | patch width in bytes |
|
||||
| 5 | 2 | Type | relocation type, `uint16`, see the table |
|
||||
| 7 | 8 | Add | addend, `int64` |
|
||||
| 15 | 8 | Sym | target SymRef |
|
||||
|
||||
The computed value `payload[Off:Off+Siz] += address(Sym) + Add` in the
|
||||
flavour the type prescribes is the linker's job; the object only records the
|
||||
request. A size 0 relocation patches nothing and exists purely as a marker
|
||||
for the linker's reachability analysis.
|
||||
|
||||
### Relocation types
|
||||
|
||||
Values of `objabi.RelocType`. The assembler and compiler emit the generic
|
||||
ones plus their own architecture's family; the rest exist for other ports and
|
||||
for the linker itself.
|
||||
|
||||
| Value | Name | Meaning |
|
||||
|---|---|---|
|
||||
| 1 | R_ADDR | absolute address |
|
||||
| 2 | R_ADDRPOWER | ppc64: high adjusted plus low 16 bits across two D-form instructions |
|
||||
| 3 | R_ADDRARM64 | arm64: adrp plus add pair |
|
||||
| 4 | R_ADDRMIPS | mips: low 16 bits of an external address |
|
||||
| 5 | R_ADDROFF | 32-bit offset from the section start to the symbol |
|
||||
| 6 | R_SIZE | size of the referenced symbol |
|
||||
| 7 | R_CALL | direct call, PC relative |
|
||||
| 8 | R_CALLARM | arm: call with a shifted 24-bit field |
|
||||
| 9 | R_CALLARM64 | arm64: BL |
|
||||
| 10 | R_CALLIND | indirect call marker |
|
||||
| 11 | R_CALLPOWER | ppc64: call |
|
||||
| 12 | R_CALLMIPS | mips: non-PC-relative call target |
|
||||
| 13 | R_CONST | constant value of the symbol |
|
||||
| 14 | R_PCREL | PC relative displacement |
|
||||
| 15 | R_TLS_LE | thread local, local exec offset |
|
||||
| 16 | R_TLS_IE | thread local, initial exec GOT offset |
|
||||
| 17 | R_GOTOFF | offset from the GOT base |
|
||||
| 18 | R_PLT0 | PLT sequence, first instruction |
|
||||
| 19 | R_PLT1 | PLT sequence, second instruction |
|
||||
| 20 | R_PLT2 | PLT sequence, third instruction |
|
||||
| 21 | R_USEFIELD | field reachability marker |
|
||||
| 22 | R_USETYPE | type reachability marker, no bytes patched |
|
||||
| 23 | R_USEIFACE | interface conversion marker, size 0 |
|
||||
| 24 | R_USEIFACEMETHOD | interface method marker, size 0, addend is the method offset |
|
||||
| 25 | R_USENAMEDMETHOD | keeps named methods alive |
|
||||
| 26 | R_METHODOFF | like R_ADDROFF, the linker may zero it when the method is dead |
|
||||
| 27 | R_KEEP | keeps the target alive if the source survives |
|
||||
| 28 | R_POWER_TOC | ppc64: TOC relative |
|
||||
| 29 | R_GOTPCREL | 32-bit PC relative GOT slot |
|
||||
| 30 | R_JMPMIPS | mips: non-PC-relative jump target |
|
||||
| 31 | R_DWARFSECREF | offset of the symbol from its section, DWARF use |
|
||||
| 32 | R_ARM64_TLS_LE | arm64: MOV[NZ] immediate, TLS local exec |
|
||||
| 33 | R_ARM64_TLS_IE | arm64: adrp plus ldr, TLS initial exec |
|
||||
| 34 | R_ARM64_GOTPCREL | arm64: adrp plus ldr GOT slot |
|
||||
| 35 | R_ARM64_GOT | arm64: GOT relative sequence |
|
||||
| 36 | R_ARM64_PCREL | arm64: adrp plus add PC relative |
|
||||
| 37 | R_ARM64_PCREL_LDST8 | arm64: adrp plus 8-bit load or store |
|
||||
| 38 | R_ARM64_PCREL_LDST16 | arm64: adrp plus 16-bit load or store |
|
||||
| 39 | R_ARM64_PCREL_LDST32 | arm64: adrp plus 32-bit load or store |
|
||||
| 40 | R_ARM64_PCREL_LDST64 | arm64: adrp plus 64-bit load or store |
|
||||
| 41 | R_ARM64_LDST8 | arm64: 12-bit load or store immediate, byte |
|
||||
| 42 | R_ARM64_LDST16 | arm64: bits 11 to 1 of the address |
|
||||
| 43 | R_ARM64_LDST32 | arm64: bits 11 to 2 |
|
||||
| 44 | R_ARM64_LDST64 | arm64: bits 11 to 3 |
|
||||
| 45 | R_ARM64_LDST128 | arm64: bits 11 to 4 |
|
||||
| 46 | R_POWER_TLS_LE | ppc64: TLS local exec across two instructions |
|
||||
| 47 | R_POWER_TLS_IE | ppc64: TLS initial exec via GOT |
|
||||
| 48 | R_POWER_TLS | ppc64: marks the X-form instruction completing a TLS sequence |
|
||||
| 49 | R_POWER_TLS_IE_PCREL34 | ppc64: prefixed TLS initial exec load |
|
||||
| 50 | R_POWER_TLS_LE_TPREL34 | ppc64: prefixed TLS local exec |
|
||||
| 51 | R_ADDRPOWER_DS | ppc64: DS-form second instruction, bits 15 to 2 |
|
||||
| 52 | R_ADDRPOWER_GOT | ppc64: GOT entry relative to TOC |
|
||||
| 53 | R_ADDRPOWER_GOT_PCREL34 | ppc64: PC relative GOT, prefixed |
|
||||
| 54 | R_ADDRPOWER_PCREL | ppc64: PC relative across two D-form instructions |
|
||||
| 55 | R_ADDRPOWER_TOCREL | ppc64: TOC relative across two D-form instructions |
|
||||
| 56 | R_ADDRPOWER_TOCREL_DS | ppc64: TOC relative, DS form |
|
||||
| 57 | R_ADDRPOWER_D34 | ppc64: prefixed absolute, 34 bits |
|
||||
| 58 | R_ADDRPOWER_PCREL34 | ppc64: prefixed PC relative, 34 bits |
|
||||
| 59 | R_RISCV_JAL | riscv64: 20-bit J-type offset |
|
||||
| 60 | R_RISCV_JAL_TRAMP | riscv64: as R_RISCV_JAL, linker-generated trampolines only |
|
||||
| 61 | R_RISCV_CALL | riscv64: AUIPC plus JALR pair |
|
||||
| 62 | R_RISCV_PCREL_ITYPE | riscv64: AUIPC plus I-type pair |
|
||||
| 63 | R_RISCV_PCREL_STYPE | riscv64: AUIPC plus S-type pair |
|
||||
| 64 | R_RISCV_TLS_IE | riscv64: TLS initial exec, AUIPC plus I-type |
|
||||
| 65 | R_RISCV_TLS_LE | riscv64: TLS local exec, LUI plus I-type |
|
||||
| 66 | R_RISCV_GOT_HI20 | riscv64: high 20 bits of a GOT address |
|
||||
| 67 | R_RISCV_GOT_PCREL_ITYPE | riscv64: GOT entry, AUIPC plus I-type |
|
||||
| 68 | R_RISCV_PCREL_HI20 | riscv64: high 20 bits of a PC relative address |
|
||||
| 69 | R_RISCV_PCREL_LO12_I | riscv64: low 12 bits, I-type |
|
||||
| 70 | R_RISCV_PCREL_LO12_S | riscv64: low 12 bits, S-type |
|
||||
| 71 | R_RISCV_BRANCH | riscv64: 12-bit branch offset |
|
||||
| 72 | R_RISCV_ADD32 | riscv64: in-place addition, V + S + A |
|
||||
| 73 | R_RISCV_SUB32 | riscv64: in-place subtraction, V - S - A |
|
||||
| 74 | R_RISCV_RVC_BRANCH | riscv64: 8-bit compressed branch offset |
|
||||
| 75 | R_RISCV_RVC_JUMP | riscv64: 11-bit compressed jump offset |
|
||||
| 76 | R_PCRELDBL | s390x: PC relative, 2-byte aligned |
|
||||
| 77 | R_LOONG64_ADDR_HI | loong64: bits 31 to 12 of an address |
|
||||
| 78 | R_LOONG64_ADDR_LO | loong64: low 12 bits |
|
||||
| 79 | R_LOONG64_ADDR64_HI | loong64: bits 63 to 52 |
|
||||
| 80 | R_LOONG64_ADDR64_LO | loong64: bits 51 to 32 |
|
||||
| 81 | R_LOONG64_ADDR_PCREL20_S2 | loong64: 22-bit aligned PC relative, PCADDI |
|
||||
| 82 | R_LOONG64_TLS_LE_HI | loong64: TLS local exec, high bits |
|
||||
| 83 | R_LOONG64_TLS_LE_LO | loong64: TLS local exec, low bits |
|
||||
| 84 | R_CALLLOONG64 | loong64: 28-bit aligned BL |
|
||||
| 85 | R_LOONG64_CALL36 | loong64: 38-bit aligned PCADDU18I plus JIRL |
|
||||
| 86 | R_LOONG64_TLS_IE_HI | loong64: TLS initial exec via GOT, high |
|
||||
| 87 | R_LOONG64_TLS_IE_LO | loong64: TLS initial exec via GOT, low |
|
||||
| 88 | R_LOONG64_GOT_HI | loong64: GOT entry, high bits |
|
||||
| 89 | R_LOONG64_GOT_LO | loong64: GOT entry, low bits |
|
||||
| 90 | R_LOONG64_GOT64_HI | loong64: 64-bit GOT entry, high |
|
||||
| 91 | R_LOONG64_GOT64_LO | loong64: 64-bit GOT entry, low |
|
||||
| 92 | R_LOONG64_ADD64 | loong64: 64-bit in-place addition |
|
||||
| 93 | R_LOONG64_SUB64 | loong64: 64-bit in-place subtraction |
|
||||
| 94 | R_JMP16LOONG64 | loong64: 18-bit aligned conditional jump |
|
||||
| 95 | R_JMP21LOONG64 | loong64: 23-bit aligned BEQZ or BNEZ |
|
||||
| 96 | R_ADDRMIPSU | mips: sign-adjusted upper 16 bits |
|
||||
| 97 | R_ADDRMIPSTLS | mips: TLS low 16 bits |
|
||||
| 98 | R_ADDRCUOFF | pointer-sized offset from the DWARF compile unit start |
|
||||
| 99 | R_WASMIMPORT | wasm: import module and name indices |
|
||||
| 100 | R_XCOFFREF | aix: keeps the target alive, patches nothing |
|
||||
| 101 | R_PEIMAGEOFF | windows: offset from the image base |
|
||||
| 102 | R_INITORDER | orders inittask records, patches nothing |
|
||||
| 103 | R_DWTXTADDR_U1 | writes a 1-byte ULEB .debug_addr index for the target function |
|
||||
| 104 | R_DWTXTADDR_U2 | as above, 2 bytes |
|
||||
| 105 | R_DWTXTADDR_U3 | as above, 3 bytes |
|
||||
| 106 | R_DWTXTADDR_U4 | as above, 4 bytes; the assembler always picks this one |
|
||||
| -32768 | R_WEAK | mask: the target need not be reachable, see below |
|
||||
| -32767 | R_WEAKADDR | R_WEAK or R_ADDR |
|
||||
| -32763 | R_WEAKADDROFF | R_WEAK or R_ADDROFF |
|
||||
|
||||
R_WEAK is bit 15 set on a negative `int16`: a weak relocation is the base
|
||||
type's value with bit 15 set. The linker strips the bit before dispatch.
|
||||
|
||||
## Aux symbol entries
|
||||
|
||||
Element size 9 bytes: a `uint8` type then a SymRef. Aux entries attach
|
||||
auxiliary symbols to a definition; the arrays run per symbol in the order
|
||||
given by AuxIndex.
|
||||
|
||||
| Value | Name | Attaches |
|
||||
|---|---|---|
|
||||
| 0 | AuxGotype | the Go type of a data symbol |
|
||||
| 1 | AuxFuncInfo | the FuncInfo payload of a text symbol |
|
||||
| 2 | AuxFuncdata | one funcdata symbol; one entry per slot, nil slots carry the {0,0} reference |
|
||||
| 3 | AuxDwarfInfo | DWARF debug info for the function |
|
||||
| 4 | AuxDwarfLoc | DWARF location lists |
|
||||
| 5 | AuxDwarfRanges | DWARF range lists |
|
||||
| 6 | AuxDwarfLines | DWARF line program |
|
||||
| 7 | AuxPcsp | pc-value table: SP adjustments |
|
||||
| 8 | AuxPcfile | pc-value table: source file indices |
|
||||
| 9 | AuxPcline | pc-value table: line numbers |
|
||||
| 10 | AuxPcinline | pc-value table: inlining tree positions |
|
||||
| 11 | AuxPcdata | one pc-value table per live variable slot |
|
||||
| 12 | AuxWasmImport | wasm import description |
|
||||
| 13 | AuxWasmType | wasm export type description |
|
||||
| 14 | AuxSehUnwindInfo | Windows SEH unwind info |
|
||||
|
||||
The writer emits them in the order Gotype, FuncInfo, Funcdata entries,
|
||||
DwarfInfo, DwarfLoc, DwarfRanges, DwarfLines, Pcsp, Pcfile, Pcline, Pcinline,
|
||||
SehUnwindInfo, Pcdata entries, WasmImport, WasmType, and skips any whose
|
||||
payload would be empty. A function assembled from `.s` source by Go 1.27.1
|
||||
carries exactly: FuncInfo, the Funcdata slots including nils, DwarfInfo,
|
||||
DwarfLines, Pcsp, Pcfile, Pcline and Pcinline; gasm's writer produces the
|
||||
same set.
|
||||
|
||||
The aux targets are either PkgIdxSelf definitions, PkgIdxHashed pcdata
|
||||
symbols, or, for the funcdata of assembly functions, PkgIdxNone references
|
||||
carrying names such as `pkg.Fn.args_stackmap` and `pkg.Fn.arginfo0`, which
|
||||
resolve to definitions in the package's compiled Go code when there is any.
|
||||
|
||||
## Symbol payloads (BlkData)
|
||||
|
||||
The payloads of all defined symbols, in definition order, concatenated with
|
||||
no padding; DataIndex gives each symbol's slice. A text symbol's payload is
|
||||
its machine code, with the stack-split preamble and any morestack block
|
||||
already included. A data symbol's payload is the bytes laid down by its DATA
|
||||
directives, zero filled to its declared size. If a symbol was created from an
|
||||
embedded file, the file's bytes follow the payload and count towards its
|
||||
DataIndex extent; assembly producers never write this extension.
|
||||
|
||||
### The FuncInfo payload
|
||||
|
||||
An SDATA symbol with no name, referenced by AuxFuncInfo. 28 bytes minimum,
|
||||
little endian:
|
||||
|
||||
| Offset | Size | Field | Meaning |
|
||||
|---|---|---|---|
|
||||
| 0 | 4 | Args | argument area in bytes; 0x80000000 when the producer declared none |
|
||||
| 4 | 4 | Locals | frame size in bytes |
|
||||
| 8 | 1 | FuncID | runtime function classification, 0 means normal |
|
||||
| 9 | 1 | FuncFlag | TopFrame = 1, SPWrite = 2, Asm = 4 |
|
||||
| 10 | 2 | padding | zero, reserved to a 4 byte boundary |
|
||||
| 12 | 4 | StartLine | source line of the TEXT declaration |
|
||||
| 16 | 4 | NumFile | count of file indices that follow |
|
||||
| 20 | 4 × NumFile | Files | indices into the Files block, ascending |
|
||||
| then | 4 | NumInlTree | count of inlining tree nodes that follow |
|
||||
| then | 24 × NumInlTree | InlTree | nodes, see below |
|
||||
|
||||
One InlTree node, 24 bytes: `int32` parent index, `uint32` file index,
|
||||
`int32` line, `uint32` PkgIdx and `uint32` SymIdx of the inlined function, and
|
||||
`int32` parent PC.
|
||||
|
||||
The assembler derives FuncID from the symbol name through
|
||||
`objabi.GetFuncID`, so a runtime function with a name the runtime treats
|
||||
specially gets that classification even when defined in assembly; an ordinary
|
||||
name yields 0. FuncFlag carries the Asm bit, 4, for every assembly function.
|
||||
|
||||
### The pc-value tables
|
||||
|
||||
The AuxPcsp, AuxPcfile, AuxPcline, AuxPcinline and AuxPcdata payloads are
|
||||
pc-value tables, each a sequence of value deltas and PC deltas:
|
||||
|
||||
- a signed value delta, zig-zag encoded, `binary.PutVarint` form;
|
||||
- an unsigned PC delta in ULEB128 form, counted in instruction units, the
|
||||
raw delta divided by the architecture's minimum instruction length;
|
||||
- the table ends with a final PC delta to the end of the function followed by
|
||||
a zero byte.
|
||||
|
||||
The first value applies from function entry. The encoding is the one
|
||||
`cmd/internal/obj/pcln.go` calls funcpctab, and it is the same encoding the
|
||||
final runtime pclntable carries.
|
||||
|
||||
### The DWARF payloads
|
||||
|
||||
AuxDwarfInfo, AuxDwarfLoc, AuxDwarfRanges and AuxDwarfLines reference SDWARF
|
||||
symbols whose payloads are DWARF byte streams. The object format treats them
|
||||
as opaque: the linker concatenates them into the final `.debug_*` sections
|
||||
and resolves the relocations recorded inside them. The compiler produces
|
||||
DWARF content per its own generation; gasm produces DWARF5 streams in
|
||||
`asm/goobj_dwarf.go`.
|
||||
|
||||
## Builtins
|
||||
|
||||
Frequently referenced runtime functions are referenced by index rather than
|
||||
by name: PkgIdxBuiltin with SymIdx set to the position in the generated table
|
||||
`cmd/internal/goobj/builtinlist.go`, 299 entries in Go 1.27.1, names such as
|
||||
`runtime.newobject` at index 0; 232 entries carry ABI 1 and the remaining 67
|
||||
ABI 0. Builtin names never enter the string table. The mapping only applies
|
||||
while the object is not linked against shared libraries, and a linkname'd
|
||||
symbol never counts as a builtin even when its name matches.
|
||||
|
||||
## Fingerprints
|
||||
|
||||
The 8 byte fingerprint identifies one build of a package. The compiler fills
|
||||
it with a hash of the package's export data; the assembler leaves it zero.
|
||||
The linker checks a package's fingerprint against the fingerprints its
|
||||
importers recorded in their Autolib entries and rejects a mismatched build,
|
||||
which is how stale objects are caught.
|
||||
|
||||
## What a producer must do
|
||||
|
||||
The checklist a third-party writer must satisfy for `go build` to accept its
|
||||
objects, in one place:
|
||||
|
||||
1. Write the container exactly: the `go object` line matching the target
|
||||
toolchain's configuration string, the `!\n` terminator, then the blob.
|
||||
2. Emit the 19 block offsets, in order, and make BlkEnd the blob length.
|
||||
3. Deduplicate the string table, keep the empty string at offset 96, and
|
||||
reference it everywhere a name appears.
|
||||
4. Index relocations, aux entries and data per symbol with the N + 1 arrays,
|
||||
definitions ordered Symdefs, Hashed64defs, Hasheddefs, NonPkgDefs.
|
||||
5. Sort relocations by offset within each symbol.
|
||||
6. Fill Siz with the true payload length, set Align for every
|
||||
content-addressable symbol, and keep symbols under 2 GB.
|
||||
7. Reference symbols by the package-index rules. An assembly producer
|
||||
references everything outside the object by name, PkgIdxNone,
|
||||
except its own file-local statics and the builtins; PkgIdxSelf is
|
||||
reserved for definitions in this object. Assembly TEXT symbols
|
||||
carry ABI 0.
|
||||
8. Compute the content hashes exactly as the toolchain does, or emit no
|
||||
hashed definitions at all.
|
||||
|
||||
## How gasm-devkit implements and verifies it
|
||||
|
||||
The writer lives in `asm/goobj.go`, which carries the shared container and the
|
||||
amd64 relocation emission, with per-architecture relocation emitters in
|
||||
`asm/goobjarm64.go`, `asm/goobjriscv.go` and `asm/goobjloong64.go`, symbol
|
||||
resolution in `asm/goobj_resolve.go` and DWARF generation in
|
||||
`asm/goobj_dwarf.go`. `gasm asm --format goobj -p pkg/path` writes objects
|
||||
that `go build` consumes in place of the toolchain's own.
|
||||
|
||||
Verification is differential and continuous:
|
||||
|
||||
- `asm/goobj_test.go` compares gasm's GOOBJ output against `go tool asm`
|
||||
output for the same source, byte for byte;
|
||||
- `asm/link_test.go` builds real Go programs whose assembly comes from gasm
|
||||
objects and runs them;
|
||||
- `gasm verify` keeps the machine code itself identical to the toolchain's,
|
||||
which is the precondition for the object comparison to be meaningful.
|
||||
|
||||
## Versioning and drift
|
||||
|
||||
The magic string carries the format generation, `go120ld` in Go 1.27.1. When
|
||||
a toolchain release changes the format, it changes that string first, and the
|
||||
linker refuses blobs whose magic it does not know. The watch points for a new
|
||||
release are, in order: the magic, the block index list, the Aux type list,
|
||||
the tail of the relocation table, the FuncInfo layout, and the builtin table
|
||||
count. gasm's tests fail against any of these changes, which is the mechanism
|
||||
that keeps this document and the writer current.
|
||||
|
||||
The authoritative sources, for the release this document covers:
|
||||
|
||||
- `cmd/internal/goobj/objfile.go`: the format, every structure in this
|
||||
document;
|
||||
- `cmd/internal/goobj/funcinfo.go`: FuncInfo and the inlining tree;
|
||||
- `cmd/internal/goobj/builtinlist.go`: the builtin table;
|
||||
- `cmd/internal/obj/objfile.go`: the writer, hash inputs and aux order;
|
||||
- `cmd/internal/obj/sym.go`: package index assignment and the by-name rule;
|
||||
- `cmd/internal/obj/pcln.go`: the pc-value encoding;
|
||||
- `cmd/internal/objabi/reloctype.go`: relocation types;
|
||||
- `cmd/internal/objabi/symkind.go`: symbol kinds;
|
||||
- `cmd/link/internal/ld/lib.go`: container parsing and fingerprint checks.
|
||||
@@ -0,0 +1,124 @@
|
||||
# AMD64
|
||||
|
||||
Layer 1, target page. Verified against `go tool asm` of Go 1.27.1 and against
|
||||
gasm's encoder, whose output is compared byte for byte with the toolchain's
|
||||
and executed on real hardware (`gasm verify`). The complete mnemonic
|
||||
inventory lives in the generated appendix
|
||||
[INSTRUCTIONS-AMD64.md](INSTRUCTIONS-AMD64.md); this page is the grammar and
|
||||
the conventions.
|
||||
|
||||
## Registers
|
||||
|
||||
| Group | Names | Notes |
|
||||
|---|---|---|
|
||||
| General purpose, 64-bit | `AX` `BX` `CX` `DX` `SI` `DI` `BP` `SP` `R8` to `R15` | bare names, no prefix |
|
||||
| Sub-registers | `AL` `CL` `DL` `BL` `AH` family; `R8B` `R8W` `R8D` for the byte, word and double word of `R8` | width rides the mnemonic as well |
|
||||
| Vector | `X0` to `X15` (128-bit), `Y0` to `Y15` (256-bit), `Z0` to `Z31` (512-bit) | SSE, AVX and AVX-512 |
|
||||
| Mask | `K0` to `K7` | AVX-512 opmask |
|
||||
| System | `TLS` | the thread pointer, see below |
|
||||
|
||||
Roles the calling convention fixes, which assembly must respect and can rely
|
||||
on:
|
||||
|
||||
- `SP` is the hardware stack pointer; the virtual frame pointer of the
|
||||
common language is the pseudo-register SP of OPERANDS.md, a different
|
||||
spelling with a different meaning.
|
||||
- `BP` is callee-save. The assembler inserts the save and restore whenever
|
||||
the function has a non-zero frame, so using BP as a general register
|
||||
interferes with sampling profilers that walk the frame chain.
|
||||
- `R14` holds `g`, the goroutine pointer, in the register ABI; `RDX` holds
|
||||
the closure context; `R12` and `R13` are the register ABI's scratch pair
|
||||
and `R15` its GOT temporary; `X15` is the zeroing register the compiler
|
||||
uses. An ABI0 assembly function called from Go sees none of these live
|
||||
across the call, but runtime assembly reads them directly.
|
||||
- The legacy spellings for the goroutine pointer are the macros of
|
||||
`runtime/go_tls.h`: `get_tls(r)` expands to `MOVQ TLS, r` and `g(r)` to
|
||||
`0(r)(TLS*1)`, the segment base riding the index field.
|
||||
|
||||
## Addressing
|
||||
|
||||
The common forms of OPERANDS.md, with the amd64 specifics:
|
||||
|
||||
```text
|
||||
offset(base) MOVQ 16(BX), AX
|
||||
offset(base)(index*scale) MOVL foo+32(SP)(R9*8), CX
|
||||
scale is 1, 2, 4 or 8
|
||||
name±offset(SB) MOVQ ·table(SB), CX
|
||||
```
|
||||
|
||||
- Global references assemble as absolute addresses and produce R_ADDR
|
||||
relocations; branch targets produce R_PCREL.
|
||||
- Vector indexed memory, the VSIB form with an X, Y or Z register in the
|
||||
index position, exists for the gather and scatter families.
|
||||
- There are no segment overrides in source; the one segment-flavoured form
|
||||
is the TLS base in the index field shown above.
|
||||
|
||||
## The frame and the split check
|
||||
|
||||
The assembler manages the frame, not the programmer:
|
||||
|
||||
- It inserts the `BP` save and restore for any non-zero frame.
|
||||
- It inserts the stack-split check for any function that is not NoSplit:
|
||||
the check compares SP against the guard, and on exhaustion calls
|
||||
`runtime.morestack_noctxt`. Frames at or below 128 bytes, StackSmall, use
|
||||
the small compare; frames at or below 4096 bytes, StackBig, use the
|
||||
adjusted form; larger frames compare in two steps.
|
||||
- On amd64 the assembler marks a function NoSplit itself when the frame is
|
||||
under StackSmall and the body calls nothing that needs stack: such a
|
||||
function carries the NoSplit flag in the object without the source ever
|
||||
writing NOSPLIT.
|
||||
|
||||
Results and arguments are stack-only in ABI0: the caller's frame carries
|
||||
them at FP offsets, per the Go prototype.
|
||||
|
||||
## Instructions
|
||||
|
||||
The inventory counts 1654 recognised mnemonics today, of which the encoder
|
||||
emits 1113; both numbers are generated in the appendix, and the gap is the
|
||||
encoder backlog that `gasm audit-instructions` measures. The families:
|
||||
|
||||
- **Integer base.** The ALU and move set with width suffixes, `MOVB`,
|
||||
`MOVW`, `MOVL`, `MOVQ`; the extension moves `MOVBLZX`, `MOVWLSX`,
|
||||
`MOVLQSX` and their siblings, which the compiler's output leans on;
|
||||
`LEA`; `PUSH` and `POP`; the shifts and rotates; the bit operations `BT`
|
||||
through `BTC`, `BSF`, `BSR`, `LZCNT`, `TZCNT`, `POPCNT`, `BSWAP`; the
|
||||
string primitives `MOVS` and `STOS`.
|
||||
- **Exchange and atomics.** `XCHG`, `CMPXCHG`, `XADD`; the extended-carry
|
||||
pair `ADCX` and `ADOX`; `CRC32`.
|
||||
- **Scalar floating point.** The SSE2 scalar moves and arithmetic
|
||||
(`MOVSD`, `MOVSS`, `ADDSD`, and the `CVT` family). Floating-point
|
||||
immediates are not encodable on this target, so the assembler
|
||||
materialises them: the constant lands in a synthesised read-only pool,
|
||||
and a positive zero collapses to `XORPS` of the register with itself,
|
||||
exactly as the toolchain does.
|
||||
- **Legacy SIMD, SSE.** The `MOVO`, `MOVOU`, `MOVAPS` family and the packed
|
||||
integer and floating operations, shuffles, lane extracts and inserts and
|
||||
the imm8-controlled forms.
|
||||
- **VEX and EVEX.** The `V`-prefixed forms for 256 and 512-bit work,
|
||||
opmask operations on `K0` to `K7`, gathers and scatters, and the
|
||||
quad-register families 4FMAPS, 4FNMADD, 4VNNIW, VP4DPWSSD and VP4DPWSSDS,
|
||||
whose register list rides the inverted V′VVV field. Mixing VEX and legacy
|
||||
SSE in one loop pays the AVX-SSE transition penalty on every switch: keep
|
||||
a loop in one dialect.
|
||||
- **Cryptographic and counting extensions.** AES-NI, SHA-1 and SHA-256,
|
||||
PCLMULQDQ, GFNI.
|
||||
- **System.** `CPUID`, `RDTSC`, `SYSCALL`, the fences, `LDMXCSR` and
|
||||
`STMXCSR`, the prefetch family.
|
||||
- **Pseudo-operations.** `BYTE`, `WORD`, `LONG`, `QUAD` lay raw bytes or
|
||||
words into the stream for encodings the assembler does not know; `ADJSP`
|
||||
adjusts the stack pointer; `DUFFCOPY` and `DUFFZERO` and `GETCALLERPC`
|
||||
are compiler-side names the table recognises but an encoder need not
|
||||
emit.
|
||||
|
||||
A mnemonic the appendix lists with `gasm encodes: no` assembles nowhere:
|
||||
gasm reports it as an explicit error, never as wrong bytes, and the
|
||||
`unencodable-instruction` lint flags it at edit time.
|
||||
|
||||
## Relocations
|
||||
|
||||
The relocations an amd64 object carries, all specified in
|
||||
[GOOBJ.md](../GOOBJ.md): `R_ADDR` for absolute globals, `R_PCREL` for
|
||||
relative addresses, `R_CALL` for direct calls, `R_TLS_LE` and `R_TLS_IE` for
|
||||
thread local access and `R_GOTPCREL` for GOT relative sequences, plus
|
||||
`R_DWTXTADDR_U4` inside the DWARF records, which the assembler always
|
||||
emits in the four-byte flavour.
|
||||
@@ -0,0 +1,122 @@
|
||||
# ARM64
|
||||
|
||||
Layer 1, target page. Verified against `go tool asm` of Go 1.27.1, against the
|
||||
toolchain's own arm64 assembler manual (`cmd/internal/obj/arm64/doc.go`) and
|
||||
against gasm's encoder, whose output is compared byte for byte with the
|
||||
toolchain's. The complete mnemonic inventory lives in the generated appendix
|
||||
[INSTRUCTIONS-ARM64.md](INSTRUCTIONS-ARM64.md).
|
||||
|
||||
## Registers
|
||||
|
||||
- General purpose: `R0` to `R30`, plus `ZR`, the zero register, and `RSP`,
|
||||
the stack pointer. There is no R31: thirty-one names and ZR.
|
||||
- Floating-point and SIMD share one file written `Vn`; where an instruction
|
||||
is scalar floating point the operand may be written `Fn` (`F0` to `F31`).
|
||||
- SVE register names (`Z0` to `Z31`, `P0` to `P15`) exist in the assembler's
|
||||
tables.
|
||||
- Roles the convention fixes: `RSP` is the stack pointer, `R29` the frame
|
||||
pointer, `R30` the link register, `R26` the closure context and `R27` the
|
||||
assembler's scratch register. The goroutine pointer lives in `R28` and is
|
||||
written `g` in source, its fields as `g_m(g)`, `g_sched(g)`; `R18` is the
|
||||
platform-reserved register and the Go toolchain never addresses it.
|
||||
|
||||
## Loads, stores and the width suffixes
|
||||
|
||||
The MOV series is the load and store interface, with the width in the
|
||||
mnemonic rather than the register name:
|
||||
|
||||
| Mnemonic | Machine instruction |
|
||||
|---|---|
|
||||
| `MOVD` | ldr, str, stur, 64-bit |
|
||||
| `MOVW` | ldrsw, str, stur, 32-bit sign extending |
|
||||
| `MOVWU` | ldr, 32-bit zero extending |
|
||||
| `MOVH` | ldrsh, strh, sturh |
|
||||
| `MOVHU` | ldrh |
|
||||
| `MOVB` | ldrsb, strb, sturb |
|
||||
| `MOVBU` | ldrb |
|
||||
|
||||
Post-index and pre-index addressing take the `.P` and `.W` suffixes on the
|
||||
mnemonic: `MOVD.P -8(R10), R8` is `ldr x8, [x10],#-8`, and `MOVB.W
|
||||
16(R16), R10` is `ldrsb x10, [x16,#16]!`.
|
||||
|
||||
## Addressing
|
||||
|
||||
```text
|
||||
imm(Rn|RSP) 28(R17)
|
||||
(Rn|RSP) (R22)
|
||||
(Rn)(Rm) (R27)(R23)
|
||||
(Rn)(Rm<<scale) (R4)(R12<<2)
|
||||
(Rn)(Rm.UXTW<<3) extended and shifted index
|
||||
(Rt1, Rt2) register pair for LDP, STP and the exclusive pair forms
|
||||
```
|
||||
|
||||
Branch targets are labels, `(R3)` for indirect, `name(SB)` for static.
|
||||
|
||||
## Operand order and the special forms
|
||||
|
||||
Most instructions appear in left-to-right assignment order: `ADD R11,
|
||||
RSP, R25` computes into R25. The exceptions the toolchain's manual lists,
|
||||
each with its own order:
|
||||
|
||||
- stores and `CBZ`, `CBNZ` keep the GNU order: `MOVD R29, 384(R19)`.
|
||||
- The multiply-accumulate family `MADD`, `MSUB`, `SMADDL` and friends are
|
||||
`<Rm>, <Ra>, <Rn>, <Rd>`.
|
||||
- The scalar FMA family `FMADDD` and friends are `<Fm>, <Fa>, <Fn>, <Fd>`.
|
||||
- The bitfield family `BFI`, `BFXIL`, `SBFIZ`, `SBFX`, `UBFIZ`, `UBFX` is
|
||||
`$<lsb>, <Rn>, $<width>, <Rd>`.
|
||||
- The conditional compare and select families carry the condition as the
|
||||
**first** operand: `CSEL GT, R0, R19, R1`, `CCMP MI, R22, $12, $13`,
|
||||
`FCCMPD AL, F8, F26, $0`.
|
||||
- The exclusive stores are `<Rf>, (<Rn>), <Rs>` with the status register
|
||||
last: `STLXR ZR, (R15), R16`.
|
||||
- `TBZ` and `TBNZ` are `$<imm>, <Rt>, <label>`.
|
||||
|
||||
Shifted and extended register operands ride the register: `R19>>30`,
|
||||
`R26->24` for arithmetic right shift, `@>` for rotate, and the extend forms
|
||||
`R19.UXTB<<4`, `R14.SXTX` with extend operators UXTB, UXTH, UXTW, UXTX,
|
||||
SXTB, SXTH, SXTW, SXTX.
|
||||
|
||||
## Conditions, branches and names
|
||||
|
||||
- Conditions ride the branch mnemonic: `B.EQ`, or the canonical
|
||||
per-condition names such as `BEQ`. Both spellings exist; the canonical
|
||||
names are what the generated inventory lists.
|
||||
- `br` is `JMP` and `blr` is `CALL` in this dialect; indirect branches are
|
||||
`JMP (R3)` and `CALL (R17)`.
|
||||
- `NOP` is a zero-width pseudo-instruction; the hardware nop is `NOOP`,
|
||||
an alias of `HINT $0`.
|
||||
- `umov` is written as `VMOV`.
|
||||
|
||||
## Constants
|
||||
|
||||
- A 16-bit immediate optionally shifted: `MOVK $(10<<32), R20`, with
|
||||
`MOVZ`, `MOVN` and their W variants; a zero shift is rejected by the
|
||||
assembler.
|
||||
- Large integer constants: `MOV` materialises any 64-bit constant, the
|
||||
closest-instruction way.
|
||||
- Vector constants: `VMOVS`, `VMOVD` and `VMOVQ`, the last taking two
|
||||
64-bit halves for a 128-bit value:
|
||||
`VMOVQ $0x1122334455667788, $0x99aabbccddeeff00, V2`.
|
||||
|
||||
## SIMD
|
||||
|
||||
Floating-point and SIMD instructions mostly carry a `V` prefix
|
||||
(`VADD`, `VFMLA`), the cryptographic extensions (`AESD`, `SHA256H`) and the
|
||||
scalar floating-point instructions being the exceptions. Operands carry an
|
||||
arrangement suffix, `V5.H8`, and structure loads and stores use bracket
|
||||
lists, `[V21.B16]`, with element selection as `V9.S[1]`.
|
||||
|
||||
## Alignment
|
||||
|
||||
`PCALIGN $n` pads to a power-of-two boundary between 8 and 2048 and also
|
||||
raises the function's alignment to the coarsest boundary any of its PCALIGN
|
||||
directives asks for. Functions default to 16-byte alignment on this target.
|
||||
|
||||
## Relocations
|
||||
|
||||
`R_ADDRARM64` for the adrp-plus-add pair, `R_ARM64_PCREL` and the
|
||||
`R_ARM64_PCREL_LDST` family for PC relative addressing, `R_ARM64_LDST` for
|
||||
the load and store immediates, `R_ARM64_GOTPCREL` and `R_ARM64_GOT` for the
|
||||
GOT, `R_ARM64_TLS_LE` and `R_ARM64_TLS_IE` for thread local storage and
|
||||
`R_CALLARM64` for direct calls, all specified in
|
||||
[GOOBJ.md](../GOOBJ.md).
|
||||
@@ -0,0 +1,148 @@
|
||||
# Directives: TEXT, DATA, GLOBL and the annotations
|
||||
|
||||
Layer 1, the common language, with the flag vocabulary both layers share.
|
||||
Verified against `go tool asm` of Go 1.27.1, against the shipped headers
|
||||
`textflag.h` and `funcdata.h` in `$GOROOT/pkg/include`, and against gasm's
|
||||
parser. Where gasm extends a directive, the extension says so and is marked.
|
||||
|
||||
Six directives exist. Three define things: TEXT, DATA, GLOBL. Three
|
||||
annotate: FUNCDATA, PCDATA, PCALIGN.
|
||||
|
||||
## TEXT
|
||||
|
||||
```text
|
||||
// func Add(a, b int64) int64
|
||||
TEXT ·Add(SB), NOSPLIT, $0-24
|
||||
...instructions...
|
||||
RET
|
||||
```
|
||||
|
||||
```text
|
||||
TEXT symbol(SB), [flags,] $framesize[-argsize]
|
||||
```
|
||||
|
||||
- The symbol is an `·Name(SB)` reference into the current package, or a
|
||||
fully qualified name.
|
||||
- The optional flag argument is a constant expression, normally an OR of the
|
||||
names from `textflag.h`, the table below. Without `#include "textflag.h"`
|
||||
the names are not macros and the assembler reports the misleading error
|
||||
`illegal or missing addressing mode for symbol NOSPLIT`: include the
|
||||
header first.
|
||||
- `$framesize-argsize` is two constants, not a subtraction: the local frame
|
||||
size in bytes, and the caller's argument area in bytes. The argument size
|
||||
may be omitted entirely, `$16`, which marks the argument size unknown
|
||||
(0x80000000 in the object, the value of `ArgsSizeUnknown` from
|
||||
`funcdata.h`); a frame size may be negative only in the generated ABI
|
||||
wrappers.
|
||||
- A function whose last instruction is not a branch cannot fall through into
|
||||
the next TEXT: the toolchain appends a jump to itself, so end functions
|
||||
with `RET` deliberately.
|
||||
- One TEXT per symbol; redeclaring is an error. The TEXT line also fixes the
|
||||
function's source line for traceback: it is the line number that pcln
|
||||
reports for the function's start.
|
||||
|
||||
The framesize and argsize fields do real work: the framesize drives the
|
||||
stack-split preamble (RUNTIME.md carries the contract), and both travel into
|
||||
the FuncInfo record of the object (GOOBJ.md carries its layout).
|
||||
|
||||
### The flag table
|
||||
|
||||
Values from `textflag.h`, in agreement with `cmd/internal/obj/textflag.go`:
|
||||
|
||||
| Name | Value | Applies to | Meaning |
|
||||
|---|---|---|---|
|
||||
| NOPROF | 1 | both | do not profile; deprecated |
|
||||
| DUPOK | 2 | both | the linker may keep one of several duplicates |
|
||||
| NOSPLIT | 4 | TEXT | no stack-split preamble |
|
||||
| RODATA | 8 | data | put the data in a read-only section |
|
||||
| NOPTR | 16 | data | the data contains no pointers |
|
||||
| WRAPPER | 32 | TEXT | a wrapper; must not disable `recover` |
|
||||
| NEEDCTXT | 64 | TEXT | a closure consuming the context register |
|
||||
| TLSBSS | 256 | data | a thread local word in BSS |
|
||||
| NOFRAME | 512 | TEXT | no frame setup; only valid with a frame size of 0 |
|
||||
| REFLECTMETHOD | 1024 | TEXT | the function calls `reflect.Type.Method` or `MethodByName` |
|
||||
| TOPFRAME | 2048 | TEXT | the outermost frame; unwinders stop here |
|
||||
| ABIWRAPPER | 4096 | TEXT | an ABI transition wrapper |
|
||||
|
||||
Rules with teeth:
|
||||
|
||||
- `NOSPLIT` removes the split check, so the frame plus everything the
|
||||
function calls must fit in the stack segment that remains. It exists to
|
||||
protect the splitting code itself; reaching for it to save two instructions
|
||||
is how stack overflows corrupt memory. On amd64 the assembler additionally
|
||||
marks small leaf functions NoSplit itself and omits the check, so the
|
||||
absence of the preamble is not proof the flag was written.
|
||||
- A TEXT whose symbol is declared `ABIInternal` must carry NOSPLIT: the
|
||||
assembler rejects it otherwise, because it cannot generate
|
||||
the split path for a register-ABI function.
|
||||
- `RODATA` implies NOPTR for the garbage collector.
|
||||
|
||||
## DATA
|
||||
|
||||
```text
|
||||
DATA ·table+0(SB)/8, $0x0102030405060708
|
||||
DATA ·msg+0(SB)/14, $"hello, world\n"
|
||||
GLOBL ·msg(SB), RODATA, $14
|
||||
```
|
||||
|
||||
```text
|
||||
DATA symbol+offset(SB)/width, value
|
||||
```
|
||||
|
||||
- `width` is exactly 1, 2, 4 or 8: the initialiser is written into the data
|
||||
image at `symbol+offset` in that many bytes.
|
||||
- The value is an integer or character constant of the width, or a string
|
||||
literal whose byte length equals the width exactly; escapes count. Long
|
||||
data is written as successive DATA lines at increasing offsets; bytes the
|
||||
directives never name are zero.
|
||||
- Every symbol initialised with DATA ends with a GLOBL line declaring its
|
||||
total size, after all of its DATA lines.
|
||||
|
||||
A symbol containing pointers cannot be defined in assembly, because the
|
||||
collector cannot see into it: define it in Go and refer to it by name. As a
|
||||
rule, data that is not read-only belongs in Go.
|
||||
|
||||
Extension, gasm only: a DATA initialiser may name a symbol,
|
||||
`DATA ·fn+0(SB)/8, $·handler(SB)`, which gasm lays down as an absolute
|
||||
relocation on that field. The toolchain offers no ground truth for this
|
||||
form; gasm's behaviour is verified by linking and execution.
|
||||
|
||||
## GLOBL
|
||||
|
||||
```text
|
||||
GLOBL symbol(SB), [flags,] $size
|
||||
```
|
||||
|
||||
Declares the symbol global with its total size in bytes. The useful flags
|
||||
are RODATA, NOPTR, DUPOK and TLSBSS from the table above. Uninitialised
|
||||
bytes are zero, which makes GLOBL with no DATA the language's BSS.
|
||||
|
||||
## FUNCDATA and PCDATA
|
||||
|
||||
```text
|
||||
FUNCDATA $functypeid, symbol(SB)
|
||||
PCDATA $pctypeid, $value
|
||||
```
|
||||
|
||||
The compiler's annotations for the garbage collector and traceback, named by
|
||||
the ids in `funcdata.h`: FUNCDATA 0 to 7 (args pointer maps, locals pointer
|
||||
maps, stack objects, inline tree, open-coded defer info, argument info,
|
||||
argument liveness, wrap info), PCDATA 0 to 4 (unsafe point, stack map index,
|
||||
inline tree index, argument liveness index, panic bounds). Assembly code
|
||||
normally reaches them only through the macro forms in `funcdata.h`, which
|
||||
RUNTIME.md explains. Outside the macros, hand-written PCDATA is meaningless:
|
||||
the values are pc-value tables the compiler builds from its own view of the
|
||||
program.
|
||||
|
||||
## PCALIGN
|
||||
|
||||
```text
|
||||
PCALIGN $32
|
||||
```
|
||||
|
||||
Pads the code so that the next instruction lands on the given boundary,
|
||||
which must be a power of two and at least the target's instruction
|
||||
alignment. Supported on amd64, arm64, ppc64, loong64 and riscv64. The
|
||||
padding instructions are the target's NOP encoding, so the bytes between
|
||||
functions differ from what the instruction stream alone would produce, which
|
||||
matters to anyone comparing encodings byte for byte.
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,570 @@
|
||||
# ARM64: instruction inventory
|
||||
|
||||
Generated by gasm-devkit's `_gen` from the Go toolchain's instruction table
|
||||
(`cmd/internal/obj/arm64/anames.go`, go1.27.1); DO NOT EDIT. This page lists every mnemonic
|
||||
`go tool asm` accepts on this target, which is the upper bound of the
|
||||
language on it: a name absent here is not an instruction of the target,
|
||||
and a name present here may still be one gasm's encoder cannot emit yet.
|
||||
|
||||
The inventory carries no per-mnemonic encoder column: on this target
|
||||
encodability is decided per operand shape, and the live measured
|
||||
coverage is reported by `gasm audit-instructions`.
|
||||
|
||||
| Mnemonic | Notes |
|
||||
|---|---|
|
||||
| `CALL` | |
|
||||
| `DUFFCOPY` | |
|
||||
| `DUFFZERO` | |
|
||||
| `END` | |
|
||||
| `FUNCDATA` | |
|
||||
| `GETCALLERPC` | |
|
||||
| `JMP` | |
|
||||
| `NOP` | No operation |
|
||||
| `PCALIGN` | |
|
||||
| `PCALIGNMAX` | |
|
||||
| `PCDATA` | |
|
||||
| `RET` | Return |
|
||||
| `TEXT` | |
|
||||
| `UNDEF` | |
|
||||
| `ADC` | ADC (64-bit) |
|
||||
| `ADCS` | ADCS (64-bit) |
|
||||
| `ADCSW` | ADCS (32-bit) |
|
||||
| `ADCW` | ADC (32-bit) |
|
||||
| `ADD` | ADD (64-bit) |
|
||||
| `ADDS` | ADDS (64-bit) |
|
||||
| `ADDSW` | ADDS (32-bit) |
|
||||
| `ADDW` | ADD (32-bit) |
|
||||
| `ADR` | Address of label/page |
|
||||
| `ADRP` | Address of label/page |
|
||||
| `AESD` | AES round |
|
||||
| `AESE` | AES round |
|
||||
| `AESIMC` | AES round |
|
||||
| `AESMC` | AES round |
|
||||
| `AND` | AND (64-bit) |
|
||||
| `ANDS` | ANDS (64-bit) |
|
||||
| `ANDSW` | ANDS (32-bit) |
|
||||
| `ANDW` | AND (32-bit) |
|
||||
| `ASR` | ASR shift |
|
||||
| `ASRW` | ASR shift (32-bit) |
|
||||
| `AT` | |
|
||||
| `AUTIA1716` | |
|
||||
| `AUTIASP` | |
|
||||
| `AUTIB1716` | |
|
||||
| `AUTIBSP` | |
|
||||
| `BCC` | Conditional branch |
|
||||
| `BCS` | Conditional branch |
|
||||
| `BEQ` | Conditional branch |
|
||||
| `BFI` | |
|
||||
| `BFIW` | |
|
||||
| `BFM` | |
|
||||
| `BFMW` | |
|
||||
| `BFXIL` | Bitfield extract |
|
||||
| `BFXILW` | |
|
||||
| `BGE` | Conditional branch |
|
||||
| `BGT` | Conditional branch |
|
||||
| `BHI` | Conditional branch |
|
||||
| `BHS` | Conditional branch |
|
||||
| `BIC` | BIC (64-bit) |
|
||||
| `BICS` | BICS (64-bit) |
|
||||
| `BICSW` | BICS (32-bit) |
|
||||
| `BICW` | BIC (32-bit) |
|
||||
| `BLE` | Conditional branch |
|
||||
| `BLO` | Conditional branch |
|
||||
| `BLS` | Conditional branch |
|
||||
| `BLT` | Conditional branch |
|
||||
| `BMI` | Conditional branch |
|
||||
| `BNE` | Conditional branch |
|
||||
| `BPL` | Conditional branch |
|
||||
| `BRK` | Breakpoint |
|
||||
| `BTI` | |
|
||||
| `BVC` | Conditional branch |
|
||||
| `BVS` | Conditional branch |
|
||||
| `CASAD` | |
|
||||
| `CASALB` | |
|
||||
| `CASALD` | |
|
||||
| `CASALH` | |
|
||||
| `CASALW` | |
|
||||
| `CASAW` | |
|
||||
| `CASB` | |
|
||||
| `CASD` | |
|
||||
| `CASH` | |
|
||||
| `CASLD` | |
|
||||
| `CASLW` | |
|
||||
| `CASPD` | |
|
||||
| `CASPW` | |
|
||||
| `CASW` | |
|
||||
| `CBNZ` | Compare/test and branch |
|
||||
| `CBNZW` | Compare/test and branch (32-bit) |
|
||||
| `CBZ` | Compare/test and branch |
|
||||
| `CBZW` | Compare/test and branch (32-bit) |
|
||||
| `CCMN` | Conditional compare |
|
||||
| `CCMNW` | Conditional compare |
|
||||
| `CCMP` | Conditional compare |
|
||||
| `CCMPW` | Conditional compare |
|
||||
| `CINC` | Conditional select |
|
||||
| `CINCW` | Conditional select (32-bit) |
|
||||
| `CINV` | Conditional select |
|
||||
| `CINVW` | Conditional select (32-bit) |
|
||||
| `CLREX` | |
|
||||
| `CLS` | Bit manipulation |
|
||||
| `CLSW` | Bit manipulation |
|
||||
| `CLZ` | Bit manipulation |
|
||||
| `CLZW` | Bit manipulation |
|
||||
| `CMN` | CMN (64-bit) |
|
||||
| `CMNW` | CMN (32-bit) |
|
||||
| `CMP` | CMP (64-bit) |
|
||||
| `CMPW` | CMP (32-bit) |
|
||||
| `CNEG` | Conditional select |
|
||||
| `CNEGW` | Conditional select (32-bit) |
|
||||
| `CRC32B` | |
|
||||
| `CRC32CB` | |
|
||||
| `CRC32CH` | |
|
||||
| `CRC32CW` | |
|
||||
| `CRC32CX` | |
|
||||
| `CRC32H` | |
|
||||
| `CRC32W` | |
|
||||
| `CRC32X` | |
|
||||
| `CSEL` | Conditional select |
|
||||
| `CSELW` | Conditional select (32-bit) |
|
||||
| `CSET` | Conditional select |
|
||||
| `CSETM` | Conditional select |
|
||||
| `CSETMW` | Conditional select (32-bit) |
|
||||
| `CSETW` | Conditional select (32-bit) |
|
||||
| `CSINC` | Conditional select |
|
||||
| `CSINCW` | Conditional select (32-bit) |
|
||||
| `CSINV` | Conditional select |
|
||||
| `CSINVW` | Conditional select (32-bit) |
|
||||
| `CSNEG` | Conditional select |
|
||||
| `CSNEGW` | Conditional select (32-bit) |
|
||||
| `DC` | Data cache maintenance |
|
||||
| `DCPS1` | |
|
||||
| `DCPS2` | |
|
||||
| `DCPS3` | |
|
||||
| `DMB` | Barrier |
|
||||
| `DRPS` | |
|
||||
| `DSB` | Barrier |
|
||||
| `DWORD` | |
|
||||
| `EON` | EON (64-bit) |
|
||||
| `EONW` | EON (32-bit) |
|
||||
| `EOR` | EOR (64-bit) |
|
||||
| `EORW` | EOR (32-bit) |
|
||||
| `ERET` | |
|
||||
| `EXTR` | Bitfield extract |
|
||||
| `EXTRW` | |
|
||||
| `FABSD` | |
|
||||
| `FABSS` | |
|
||||
| `FADDD` | |
|
||||
| `FADDS` | |
|
||||
| `FCCMPD` | |
|
||||
| `FCCMPED` | |
|
||||
| `FCCMPES` | |
|
||||
| `FCCMPS` | |
|
||||
| `FCMPD` | |
|
||||
| `FCMPED` | |
|
||||
| `FCMPES` | |
|
||||
| `FCMPS` | |
|
||||
| `FCSELD` | |
|
||||
| `FCSELS` | |
|
||||
| `FCVTDH` | |
|
||||
| `FCVTDS` | |
|
||||
| `FCVTHD` | |
|
||||
| `FCVTHS` | |
|
||||
| `FCVTSD` | |
|
||||
| `FCVTSH` | |
|
||||
| `FCVTZSD` | |
|
||||
| `FCVTZSDW` | |
|
||||
| `FCVTZSS` | |
|
||||
| `FCVTZSSW` | |
|
||||
| `FCVTZUD` | |
|
||||
| `FCVTZUDW` | |
|
||||
| `FCVTZUS` | |
|
||||
| `FCVTZUSW` | |
|
||||
| `FDIVD` | |
|
||||
| `FDIVS` | |
|
||||
| `FLDPD` | Register-pair load or store |
|
||||
| `FLDPQ` | |
|
||||
| `FLDPS` | |
|
||||
| `FMADDD` | |
|
||||
| `FMADDS` | |
|
||||
| `FMAXD` | |
|
||||
| `FMAXNMD` | |
|
||||
| `FMAXNMS` | |
|
||||
| `FMAXS` | |
|
||||
| `FMIND` | |
|
||||
| `FMINNMD` | |
|
||||
| `FMINNMS` | |
|
||||
| `FMINS` | |
|
||||
| `FMOVD` | Move / load / store |
|
||||
| `FMOVQ` | |
|
||||
| `FMOVS` | Move / load / store |
|
||||
| `FMSUBD` | |
|
||||
| `FMSUBS` | |
|
||||
| `FMULD` | |
|
||||
| `FMULS` | |
|
||||
| `FNEGD` | |
|
||||
| `FNEGS` | |
|
||||
| `FNMADDD` | |
|
||||
| `FNMADDS` | |
|
||||
| `FNMSUBD` | |
|
||||
| `FNMSUBS` | |
|
||||
| `FNMULD` | |
|
||||
| `FNMULS` | |
|
||||
| `FRINTAD` | |
|
||||
| `FRINTAS` | |
|
||||
| `FRINTID` | |
|
||||
| `FRINTIS` | |
|
||||
| `FRINTMD` | |
|
||||
| `FRINTMS` | |
|
||||
| `FRINTND` | |
|
||||
| `FRINTNS` | |
|
||||
| `FRINTPD` | |
|
||||
| `FRINTPS` | |
|
||||
| `FRINTXD` | |
|
||||
| `FRINTXS` | |
|
||||
| `FRINTZD` | |
|
||||
| `FRINTZS` | |
|
||||
| `FSQRTD` | |
|
||||
| `FSQRTS` | |
|
||||
| `FSTPD` | Register-pair load or store |
|
||||
| `FSTPQ` | |
|
||||
| `FSTPS` | |
|
||||
| `FSUBD` | |
|
||||
| `FSUBS` | |
|
||||
| `HINT` | |
|
||||
| `HLT` | |
|
||||
| `HVC` | Exception generation |
|
||||
| `IC` | |
|
||||
| `ISB` | Barrier |
|
||||
| `LDADDAB` | |
|
||||
| `LDADDAD` | |
|
||||
| `LDADDAH` | |
|
||||
| `LDADDALB` | |
|
||||
| `LDADDALD` | |
|
||||
| `LDADDALH` | |
|
||||
| `LDADDALW` | |
|
||||
| `LDADDAW` | |
|
||||
| `LDADDB` | |
|
||||
| `LDADDD` | |
|
||||
| `LDADDH` | |
|
||||
| `LDADDLB` | |
|
||||
| `LDADDLD` | |
|
||||
| `LDADDLH` | |
|
||||
| `LDADDLW` | |
|
||||
| `LDADDW` | |
|
||||
| `LDAR` | Atomic memory operation |
|
||||
| `LDARB` | Atomic memory operation |
|
||||
| `LDARH` | Atomic memory operation |
|
||||
| `LDARW` | Atomic memory operation |
|
||||
| `LDAXP` | |
|
||||
| `LDAXPW` | |
|
||||
| `LDAXR` | Atomic memory operation |
|
||||
| `LDAXRB` | Atomic memory operation |
|
||||
| `LDAXRH` | Atomic memory operation |
|
||||
| `LDAXRW` | Atomic memory operation |
|
||||
| `LDCLRAB` | |
|
||||
| `LDCLRAD` | |
|
||||
| `LDCLRAH` | |
|
||||
| `LDCLRALB` | |
|
||||
| `LDCLRALD` | |
|
||||
| `LDCLRALH` | |
|
||||
| `LDCLRALW` | |
|
||||
| `LDCLRAW` | |
|
||||
| `LDCLRB` | |
|
||||
| `LDCLRD` | |
|
||||
| `LDCLRH` | |
|
||||
| `LDCLRLB` | |
|
||||
| `LDCLRLD` | |
|
||||
| `LDCLRLH` | |
|
||||
| `LDCLRLW` | |
|
||||
| `LDCLRW` | |
|
||||
| `LDEORAB` | |
|
||||
| `LDEORAD` | |
|
||||
| `LDEORAH` | |
|
||||
| `LDEORALB` | |
|
||||
| `LDEORALD` | |
|
||||
| `LDEORALH` | |
|
||||
| `LDEORALW` | |
|
||||
| `LDEORAW` | |
|
||||
| `LDEORB` | |
|
||||
| `LDEORD` | |
|
||||
| `LDEORH` | |
|
||||
| `LDEORLB` | |
|
||||
| `LDEORLD` | |
|
||||
| `LDEORLH` | |
|
||||
| `LDEORLW` | |
|
||||
| `LDEORW` | |
|
||||
| `LDORAB` | |
|
||||
| `LDORAD` | |
|
||||
| `LDORAH` | |
|
||||
| `LDORALB` | |
|
||||
| `LDORALD` | |
|
||||
| `LDORALH` | |
|
||||
| `LDORALW` | |
|
||||
| `LDORAW` | |
|
||||
| `LDORB` | |
|
||||
| `LDORD` | |
|
||||
| `LDORH` | |
|
||||
| `LDORLB` | |
|
||||
| `LDORLD` | |
|
||||
| `LDORLH` | |
|
||||
| `LDORLW` | |
|
||||
| `LDORW` | |
|
||||
| `LDP` | Register-pair load or store |
|
||||
| `LDPSW` | |
|
||||
| `LDPW` | Register-pair load or store |
|
||||
| `LDXP` | |
|
||||
| `LDXPW` | |
|
||||
| `LDXR` | |
|
||||
| `LDXRB` | |
|
||||
| `LDXRH` | |
|
||||
| `LDXRW` | |
|
||||
| `LSL` | LSL shift |
|
||||
| `LSLW` | LSL shift (32-bit) |
|
||||
| `LSR` | LSR shift |
|
||||
| `LSRW` | LSR shift (32-bit) |
|
||||
| `MADD` | Multiply / multiply-accumulate |
|
||||
| `MADDW` | |
|
||||
| `MNEG` | Multiply / multiply-accumulate |
|
||||
| `MNEGW` | |
|
||||
| `MOVB` | Move / load / store |
|
||||
| `MOVBU` | Move / load / store |
|
||||
| `MOVD` | Move / load / store |
|
||||
| `MOVH` | Move / load / store |
|
||||
| `MOVHU` | Move / load / store |
|
||||
| `MOVK` | Move wide constant |
|
||||
| `MOVKW` | Move wide constant |
|
||||
| `MOVN` | Move wide constant |
|
||||
| `MOVNW` | Move wide constant |
|
||||
| `MOVP` | |
|
||||
| `MOVPD` | |
|
||||
| `MOVPQ` | |
|
||||
| `MOVPS` | |
|
||||
| `MOVPSW` | |
|
||||
| `MOVPW` | |
|
||||
| `MOVW` | Move / load / store |
|
||||
| `MOVWU` | Move / load / store |
|
||||
| `MOVZ` | Move wide constant |
|
||||
| `MOVZW` | Move wide constant |
|
||||
| `MRS` | System register access |
|
||||
| `MSR` | System register access |
|
||||
| `MSUB` | Multiply / multiply-accumulate |
|
||||
| `MSUBW` | |
|
||||
| `MUL` | Multiply / multiply-accumulate |
|
||||
| `MULW` | |
|
||||
| `MVN` | MVN (64-bit) |
|
||||
| `MVNW` | MVN (32-bit) |
|
||||
| `NEG` | NEG (64-bit) |
|
||||
| `NEGS` | |
|
||||
| `NEGSW` | |
|
||||
| `NEGW` | NEG (32-bit) |
|
||||
| `NGC` | NGC (64-bit) |
|
||||
| `NGCS` | |
|
||||
| `NGCSW` | |
|
||||
| `NGCW` | NGC (32-bit) |
|
||||
| `NOOP` | |
|
||||
| `ORN` | ORN (64-bit) |
|
||||
| `ORNW` | ORN (32-bit) |
|
||||
| `ORR` | ORR (64-bit) |
|
||||
| `ORRW` | ORR (32-bit) |
|
||||
| `PACIASP` | |
|
||||
| `PACIBSP` | |
|
||||
| `PRFM` | Memory prefetch |
|
||||
| `PRFUM` | |
|
||||
| `RBIT` | Bit manipulation |
|
||||
| `RBITW` | Bit manipulation |
|
||||
| `REM` | |
|
||||
| `REMW` | |
|
||||
| `REV` | Bit manipulation |
|
||||
| `REV16` | Bit manipulation |
|
||||
| `REV16W` | |
|
||||
| `REV32` | Bit manipulation |
|
||||
| `REVW` | Bit manipulation |
|
||||
| `ROR` | ROR shift |
|
||||
| `RORW` | ROR shift (32-bit) |
|
||||
| `SBC` | SBC (64-bit) |
|
||||
| `SBCS` | SBCS (64-bit) |
|
||||
| `SBCSW` | SBCS (32-bit) |
|
||||
| `SBCW` | SBC (32-bit) |
|
||||
| `SBFIZ` | |
|
||||
| `SBFIZW` | |
|
||||
| `SBFM` | Bitfield extract |
|
||||
| `SBFMW` | |
|
||||
| `SBFX` | Bitfield extract |
|
||||
| `SBFXW` | |
|
||||
| `SCVTFD` | |
|
||||
| `SCVTFS` | |
|
||||
| `SCVTFWD` | |
|
||||
| `SCVTFWS` | |
|
||||
| `SDIV` | Divide |
|
||||
| `SDIVW` | Divide |
|
||||
| `SEV` | |
|
||||
| `SEVL` | |
|
||||
| `SHA1C` | SHA round |
|
||||
| `SHA1H` | SHA round |
|
||||
| `SHA1M` | SHA round |
|
||||
| `SHA1P` | SHA round |
|
||||
| `SHA1SU0` | SHA round |
|
||||
| `SHA1SU1` | SHA round |
|
||||
| `SHA256H` | SHA round |
|
||||
| `SHA256H2` | SHA round |
|
||||
| `SHA256SU0` | SHA round |
|
||||
| `SHA256SU1` | SHA round |
|
||||
| `SHA512H` | SHA round |
|
||||
| `SHA512H2` | SHA round |
|
||||
| `SHA512SU0` | SHA round |
|
||||
| `SHA512SU1` | SHA round |
|
||||
| `SMADDL` | Multiply / multiply-accumulate |
|
||||
| `SMC` | Exception generation |
|
||||
| `SMNEGL` | |
|
||||
| `SMSUBL` | Multiply / multiply-accumulate |
|
||||
| `SMULH` | Multiply / multiply-accumulate |
|
||||
| `SMULL` | Multiply / multiply-accumulate |
|
||||
| `STLR` | Atomic memory operation |
|
||||
| `STLRB` | Atomic memory operation |
|
||||
| `STLRH` | Atomic memory operation |
|
||||
| `STLRW` | Atomic memory operation |
|
||||
| `STLXP` | |
|
||||
| `STLXPW` | |
|
||||
| `STLXR` | |
|
||||
| `STLXRB` | |
|
||||
| `STLXRH` | |
|
||||
| `STLXRW` | |
|
||||
| `STP` | Register-pair load or store |
|
||||
| `STPW` | Register-pair load or store |
|
||||
| `STXP` | |
|
||||
| `STXPW` | |
|
||||
| `STXR` | Atomic memory operation |
|
||||
| `STXRB` | Atomic memory operation |
|
||||
| `STXRH` | Atomic memory operation |
|
||||
| `STXRW` | Atomic memory operation |
|
||||
| `SUB` | SUB (64-bit) |
|
||||
| `SUBS` | SUBS (64-bit) |
|
||||
| `SUBSW` | SUBS (32-bit) |
|
||||
| `SUBW` | SUB (32-bit) |
|
||||
| `SVC` | Exception generation |
|
||||
| `SWPAB` | |
|
||||
| `SWPAD` | |
|
||||
| `SWPAH` | |
|
||||
| `SWPALB` | |
|
||||
| `SWPALD` | |
|
||||
| `SWPALH` | |
|
||||
| `SWPALW` | |
|
||||
| `SWPAW` | |
|
||||
| `SWPB` | |
|
||||
| `SWPD` | |
|
||||
| `SWPH` | |
|
||||
| `SWPLB` | |
|
||||
| `SWPLD` | |
|
||||
| `SWPLH` | |
|
||||
| `SWPLW` | |
|
||||
| `SWPW` | |
|
||||
| `SXTB` | |
|
||||
| `SXTBW` | |
|
||||
| `SXTH` | |
|
||||
| `SXTHW` | |
|
||||
| `SXTW` | |
|
||||
| `SYS` | |
|
||||
| `SYSL` | |
|
||||
| `TBNZ` | Compare/test and branch |
|
||||
| `TBZ` | Compare/test and branch |
|
||||
| `TLBI` | |
|
||||
| `TST` | TST (64-bit) |
|
||||
| `TSTW` | TST (32-bit) |
|
||||
| `UBFIZ` | |
|
||||
| `UBFIZW` | |
|
||||
| `UBFM` | Bitfield extract |
|
||||
| `UBFMW` | |
|
||||
| `UBFX` | Bitfield extract |
|
||||
| `UBFXW` | |
|
||||
| `UCVTFD` | |
|
||||
| `UCVTFS` | |
|
||||
| `UCVTFWD` | |
|
||||
| `UCVTFWS` | |
|
||||
| `UDIV` | Divide |
|
||||
| `UDIVW` | Divide |
|
||||
| `UMADDL` | Multiply / multiply-accumulate |
|
||||
| `UMNEGL` | |
|
||||
| `UMSUBL` | Multiply / multiply-accumulate |
|
||||
| `UMULH` | Multiply / multiply-accumulate |
|
||||
| `UMULL` | Multiply / multiply-accumulate |
|
||||
| `UREM` | |
|
||||
| `UREMW` | |
|
||||
| `UXTB` | |
|
||||
| `UXTBW` | |
|
||||
| `UXTH` | |
|
||||
| `UXTHW` | |
|
||||
| `UXTW` | |
|
||||
| `VADD` | NEON SIMD vector operation |
|
||||
| `VADDP` | |
|
||||
| `VADDV` | NEON SIMD vector operation |
|
||||
| `VAND` | NEON SIMD vector operation |
|
||||
| `VBCAX` | Three-way XOR / rotate crypto vector operation |
|
||||
| `VBIF` | NEON SIMD vector operation |
|
||||
| `VBIT` | |
|
||||
| `VBSL` | NEON SIMD vector operation |
|
||||
| `VCMEQ` | |
|
||||
| `VCMTST` | |
|
||||
| `VCNT` | NEON SIMD vector operation |
|
||||
| `VDUP` | NEON SIMD vector operation |
|
||||
| `VEOR` | NEON SIMD vector operation |
|
||||
| `VEOR3` | Three-way XOR / rotate crypto vector operation |
|
||||
| `VEXT` | NEON SIMD vector operation |
|
||||
| `VFMLA` | NEON SIMD vector operation |
|
||||
| `VFMLS` | NEON SIMD vector operation |
|
||||
| `VLD1` | NEON SIMD vector operation |
|
||||
| `VLD1R` | |
|
||||
| `VLD2` | NEON SIMD vector operation |
|
||||
| `VLD2R` | |
|
||||
| `VLD3` | NEON SIMD vector operation |
|
||||
| `VLD3R` | |
|
||||
| `VLD4` | NEON SIMD vector operation |
|
||||
| `VLD4R` | |
|
||||
| `VMOV` | NEON SIMD vector operation |
|
||||
| `VMOVD` | |
|
||||
| `VMOVI` | NEON SIMD vector operation |
|
||||
| `VMOVQ` | NEON SIMD vector operation |
|
||||
| `VMOVS` | |
|
||||
| `VORR` | NEON SIMD vector operation |
|
||||
| `VPMULL` | |
|
||||
| `VPMULL2` | |
|
||||
| `VRAX1` | Three-way XOR / rotate crypto vector operation |
|
||||
| `VRBIT` | |
|
||||
| `VREV16` | NEON SIMD vector operation |
|
||||
| `VREV32` | NEON SIMD vector operation |
|
||||
| `VREV64` | NEON SIMD vector operation |
|
||||
| `VSHL` | NEON SIMD vector operation |
|
||||
| `VSLI` | |
|
||||
| `VSRI` | |
|
||||
| `VST1` | NEON SIMD vector operation |
|
||||
| `VST2` | NEON SIMD vector operation |
|
||||
| `VST3` | NEON SIMD vector operation |
|
||||
| `VST4` | NEON SIMD vector operation |
|
||||
| `VSUB` | NEON SIMD vector operation |
|
||||
| `VTBL` | NEON SIMD vector operation |
|
||||
| `VTBX` | NEON SIMD vector operation |
|
||||
| `VTRN1` | NEON SIMD vector operation |
|
||||
| `VTRN2` | NEON SIMD vector operation |
|
||||
| `VUADDLV` | |
|
||||
| `VUADDW` | |
|
||||
| `VUADDW2` | |
|
||||
| `VUMAX` | |
|
||||
| `VUMIN` | |
|
||||
| `VUSHLL` | |
|
||||
| `VUSHLL2` | |
|
||||
| `VUSHR` | NEON SIMD vector operation |
|
||||
| `VUSRA` | |
|
||||
| `VUXTL` | |
|
||||
| `VUXTL2` | |
|
||||
| `VUZP1` | NEON SIMD vector operation |
|
||||
| `VUZP2` | NEON SIMD vector operation |
|
||||
| `VXAR` | Three-way XOR / rotate crypto vector operation |
|
||||
| `VZIP1` | NEON SIMD vector operation |
|
||||
| `VZIP2` | NEON SIMD vector operation |
|
||||
| `WFE` | |
|
||||
| `WFI` | |
|
||||
| `WORD` | |
|
||||
| `YIELD` | |
|
||||
| `B` | Unconditional branch |
|
||||
| `BL` | Branch with link |
|
||||
|
||||
Recognised: 554 mnemonics.
|
||||
@@ -0,0 +1,830 @@
|
||||
# LoongArch 64: instruction inventory
|
||||
|
||||
Generated by gasm-devkit's `_gen` from the Go toolchain's instruction table
|
||||
(`cmd/internal/obj/loong64/anames.go`, go1.27.1); DO NOT EDIT. This page lists every mnemonic
|
||||
`go tool asm` accepts on this target, which is the upper bound of the
|
||||
language on it: a name absent here is not an instruction of the target,
|
||||
and a name present here may still be one gasm's encoder cannot emit yet.
|
||||
|
||||
The inventory carries no per-mnemonic encoder column: on this target
|
||||
encodability is decided per operand shape, and the live measured
|
||||
coverage is reported by `gasm audit-instructions`.
|
||||
|
||||
| Mnemonic | Notes |
|
||||
|---|---|
|
||||
| `CALL` | |
|
||||
| `DUFFCOPY` | |
|
||||
| `DUFFZERO` | |
|
||||
| `END` | |
|
||||
| `FUNCDATA` | |
|
||||
| `GETCALLERPC` | |
|
||||
| `JMP` | |
|
||||
| `NOP` | No operation |
|
||||
| `PCALIGN` | |
|
||||
| `PCALIGNMAX` | |
|
||||
| `PCDATA` | |
|
||||
| `RET` | Return |
|
||||
| `TEXT` | |
|
||||
| `UNDEF` | |
|
||||
| `ABSD` | |
|
||||
| `ABSF` | |
|
||||
| `ADD` | Integer add (word) |
|
||||
| `ADDD` | Add doubleword |
|
||||
| `ADDF` | |
|
||||
| `ADDV` | |
|
||||
| `ADDV16` | |
|
||||
| `ADDVU` | |
|
||||
| `ADDW` | Add word |
|
||||
| `ALSLV` | |
|
||||
| `ALSLW` | |
|
||||
| `ALSLWU` | |
|
||||
| `AMADDDBV` | |
|
||||
| `AMADDDBW` | |
|
||||
| `AMADDV` | |
|
||||
| `AMADDW` | |
|
||||
| `AMANDDBV` | |
|
||||
| `AMANDDBW` | |
|
||||
| `AMANDV` | |
|
||||
| `AMANDW` | |
|
||||
| `AMCASB` | |
|
||||
| `AMCASDBB` | |
|
||||
| `AMCASDBH` | |
|
||||
| `AMCASDBV` | |
|
||||
| `AMCASDBW` | |
|
||||
| `AMCASH` | |
|
||||
| `AMCASV` | |
|
||||
| `AMCASW` | |
|
||||
| `AMMAXDBV` | |
|
||||
| `AMMAXDBVU` | |
|
||||
| `AMMAXDBW` | |
|
||||
| `AMMAXDBWU` | |
|
||||
| `AMMAXV` | |
|
||||
| `AMMAXVU` | |
|
||||
| `AMMAXW` | |
|
||||
| `AMMAXWU` | |
|
||||
| `AMMINDBV` | |
|
||||
| `AMMINDBVU` | |
|
||||
| `AMMINDBW` | |
|
||||
| `AMMINDBWU` | |
|
||||
| `AMMINV` | |
|
||||
| `AMMINVU` | |
|
||||
| `AMMINW` | |
|
||||
| `AMMINWU` | |
|
||||
| `AMORDBV` | |
|
||||
| `AMORDBW` | |
|
||||
| `AMORV` | |
|
||||
| `AMORW` | |
|
||||
| `AMSWAPB` | |
|
||||
| `AMSWAPDBB` | |
|
||||
| `AMSWAPDBH` | |
|
||||
| `AMSWAPDBV` | |
|
||||
| `AMSWAPDBW` | |
|
||||
| `AMSWAPH` | |
|
||||
| `AMSWAPV` | |
|
||||
| `AMSWAPW` | |
|
||||
| `AMXORDBV` | |
|
||||
| `AMXORDBW` | |
|
||||
| `AMXORV` | |
|
||||
| `AMXORW` | |
|
||||
| `AND` | Bitwise AND |
|
||||
| `ANDN` | |
|
||||
| `BEQ` | Branch if equal |
|
||||
| `BFPF` | |
|
||||
| `BFPT` | |
|
||||
| `BGE` | Branch if greater or equal |
|
||||
| `BGEU` | Branch if greater or equal unsigned |
|
||||
| `BGEZ` | |
|
||||
| `BGTZ` | |
|
||||
| `BITREV4B` | |
|
||||
| `BITREV8B` | |
|
||||
| `BITREVV` | |
|
||||
| `BITREVW` | |
|
||||
| `BLEZ` | |
|
||||
| `BLT` | Branch if less than |
|
||||
| `BLTU` | Branch if less than unsigned |
|
||||
| `BLTZ` | |
|
||||
| `BNE` | Branch if not equal |
|
||||
| `BREAK` | Breakpoint |
|
||||
| `BSTRINSV` | |
|
||||
| `BSTRINSW` | |
|
||||
| `BSTRPICKV` | |
|
||||
| `BSTRPICKW` | |
|
||||
| `CLOV` | |
|
||||
| `CLOW` | |
|
||||
| `CLZV` | |
|
||||
| `CLZW` | |
|
||||
| `CMPEQD` | |
|
||||
| `CMPEQF` | |
|
||||
| `CMPGED` | |
|
||||
| `CMPGEF` | |
|
||||
| `CMPGTD` | |
|
||||
| `CMPGTF` | |
|
||||
| `CPUCFG` | |
|
||||
| `CRCCWBW` | |
|
||||
| `CRCCWHW` | |
|
||||
| `CRCCWVW` | |
|
||||
| `CRCCWWW` | |
|
||||
| `CRCWBW` | |
|
||||
| `CRCWHW` | |
|
||||
| `CRCWVW` | |
|
||||
| `CRCWWW` | |
|
||||
| `CTOV` | |
|
||||
| `CTOW` | |
|
||||
| `CTZV` | |
|
||||
| `CTZW` | |
|
||||
| `DBAR` | Barrier |
|
||||
| `DIV` | Divide (word) |
|
||||
| `DIVD` | Divide doubleword |
|
||||
| `DIVF` | |
|
||||
| `DIVU` | |
|
||||
| `DIVV` | |
|
||||
| `DIVVU` | |
|
||||
| `DIVW` | Divide word |
|
||||
| `DIVWU` | |
|
||||
| `EXTWB` | |
|
||||
| `EXTWH` | |
|
||||
| `FCLASSD` | |
|
||||
| `FCLASSF` | |
|
||||
| `FCOPYSGD` | |
|
||||
| `FCOPYSGF` | |
|
||||
| `FFINTDV` | |
|
||||
| `FFINTDW` | |
|
||||
| `FFINTFV` | |
|
||||
| `FFINTFW` | |
|
||||
| `FLOGBD` | |
|
||||
| `FLOGBF` | |
|
||||
| `FMADDD` | |
|
||||
| `FMADDF` | |
|
||||
| `FMAXAD` | |
|
||||
| `FMAXAF` | |
|
||||
| `FMAXD` | |
|
||||
| `FMAXF` | |
|
||||
| `FMINAD` | |
|
||||
| `FMINAF` | |
|
||||
| `FMIND` | |
|
||||
| `FMINF` | |
|
||||
| `FMSUBD` | |
|
||||
| `FMSUBF` | |
|
||||
| `FNMADDD` | |
|
||||
| `FNMADDF` | |
|
||||
| `FNMSUBD` | |
|
||||
| `FNMSUBF` | |
|
||||
| `FSCALEBD` | |
|
||||
| `FSCALEBF` | |
|
||||
| `FSEL` | |
|
||||
| `FTINTRMVD` | |
|
||||
| `FTINTRMVF` | |
|
||||
| `FTINTRMWD` | |
|
||||
| `FTINTRMWF` | |
|
||||
| `FTINTRNEVD` | |
|
||||
| `FTINTRNEVF` | |
|
||||
| `FTINTRNEWD` | |
|
||||
| `FTINTRNEWF` | |
|
||||
| `FTINTRPVD` | |
|
||||
| `FTINTRPVF` | |
|
||||
| `FTINTRPWD` | |
|
||||
| `FTINTRPWF` | |
|
||||
| `FTINTRZVD` | |
|
||||
| `FTINTRZVF` | |
|
||||
| `FTINTRZWD` | |
|
||||
| `FTINTRZWF` | |
|
||||
| `FTINTVD` | |
|
||||
| `FTINTVF` | |
|
||||
| `FTINTWD` | |
|
||||
| `FTINTWF` | |
|
||||
| `JIRL` | Jump indirect with link |
|
||||
| `LL` | |
|
||||
| `LLV` | |
|
||||
| `LU12IW` | |
|
||||
| `LU32ID` | |
|
||||
| `LU52ID` | |
|
||||
| `LUI` | |
|
||||
| `MASKEQZ` | |
|
||||
| `MASKNEZ` | |
|
||||
| `MOVB` | |
|
||||
| `MOVBU` | |
|
||||
| `MOVD` | |
|
||||
| `MOVDF` | |
|
||||
| `MOVDV` | |
|
||||
| `MOVDW` | |
|
||||
| `MOVF` | |
|
||||
| `MOVFD` | |
|
||||
| `MOVFV` | |
|
||||
| `MOVFW` | |
|
||||
| `MOVH` | |
|
||||
| `MOVHU` | |
|
||||
| `MOVV` | |
|
||||
| `MOVVD` | |
|
||||
| `MOVVF` | |
|
||||
| `MOVVP` | |
|
||||
| `MOVW` | |
|
||||
| `MOVWD` | |
|
||||
| `MOVWF` | |
|
||||
| `MOVWP` | |
|
||||
| `MOVWU` | |
|
||||
| `MUL` | Multiply (word) |
|
||||
| `MULD` | Multiply doubleword |
|
||||
| `MULF` | |
|
||||
| `MULH` | |
|
||||
| `MULHU` | |
|
||||
| `MULHV` | |
|
||||
| `MULHVU` | |
|
||||
| `MULV` | |
|
||||
| `MULVU` | |
|
||||
| `MULW` | Multiply word |
|
||||
| `MULWVW` | |
|
||||
| `MULWVWU` | |
|
||||
| `NEGD` | |
|
||||
| `NEGF` | |
|
||||
| `NEGV` | |
|
||||
| `NEGW` | |
|
||||
| `NOOP` | |
|
||||
| `NOR` | Bitwise NOR |
|
||||
| `OR` | Bitwise OR |
|
||||
| `ORN` | |
|
||||
| `PCADDU12I` | |
|
||||
| `PCALAU12I` | |
|
||||
| `PRELD` | |
|
||||
| `PRELDX` | |
|
||||
| `RDTIMED` | |
|
||||
| `RDTIMEHW` | |
|
||||
| `RDTIMELW` | |
|
||||
| `REM` | |
|
||||
| `REMU` | |
|
||||
| `REMV` | |
|
||||
| `REMVU` | |
|
||||
| `REMW` | |
|
||||
| `REMWU` | |
|
||||
| `REVB2H` | |
|
||||
| `REVB2W` | |
|
||||
| `REVB4H` | |
|
||||
| `REVBV` | |
|
||||
| `REVH2W` | |
|
||||
| `REVHV` | |
|
||||
| `RFE` | |
|
||||
| `ROTR` | Rotate right |
|
||||
| `ROTRV` | |
|
||||
| `SC` | |
|
||||
| `SCV` | |
|
||||
| `SGT` | |
|
||||
| `SGTU` | |
|
||||
| `SLL` | Shift left logical |
|
||||
| `SLLV` | |
|
||||
| `SQRTD` | |
|
||||
| `SQRTF` | |
|
||||
| `SRA` | Shift right arithmetic |
|
||||
| `SRAV` | |
|
||||
| `SRL` | Shift right logical |
|
||||
| `SRLV` | |
|
||||
| `SUB` | Subtract (word) |
|
||||
| `SUBD` | Subtract doubleword |
|
||||
| `SUBF` | |
|
||||
| `SUBV` | |
|
||||
| `SUBVU` | |
|
||||
| `SUBW` | Subtract word |
|
||||
| `SYSCALL` | System call |
|
||||
| `TEQ` | |
|
||||
| `TNE` | |
|
||||
| `TRUNCDV` | |
|
||||
| `TRUNCDW` | |
|
||||
| `TRUNCFV` | |
|
||||
| `TRUNCFW` | |
|
||||
| `VADDB` | |
|
||||
| `VADDBU` | |
|
||||
| `VADDD` | |
|
||||
| `VADDF` | |
|
||||
| `VADDH` | |
|
||||
| `VADDHU` | |
|
||||
| `VADDQ` | |
|
||||
| `VADDV` | |
|
||||
| `VADDVU` | |
|
||||
| `VADDW` | |
|
||||
| `VADDWEVHB` | |
|
||||
| `VADDWEVHBU` | |
|
||||
| `VADDWEVQV` | |
|
||||
| `VADDWEVQVU` | |
|
||||
| `VADDWEVVW` | |
|
||||
| `VADDWEVVWU` | |
|
||||
| `VADDWEVWH` | |
|
||||
| `VADDWEVWHU` | |
|
||||
| `VADDWODHB` | |
|
||||
| `VADDWODHBU` | |
|
||||
| `VADDWODQV` | |
|
||||
| `VADDWODQVU` | |
|
||||
| `VADDWODVW` | |
|
||||
| `VADDWODVWU` | |
|
||||
| `VADDWODWH` | |
|
||||
| `VADDWODWHU` | |
|
||||
| `VADDWU` | |
|
||||
| `VANDB` | |
|
||||
| `VANDNV` | |
|
||||
| `VANDV` | |
|
||||
| `VBITCLRB` | |
|
||||
| `VBITCLRH` | |
|
||||
| `VBITCLRV` | |
|
||||
| `VBITCLRW` | |
|
||||
| `VBITREVB` | |
|
||||
| `VBITREVH` | |
|
||||
| `VBITREVV` | |
|
||||
| `VBITREVW` | |
|
||||
| `VBITSETB` | |
|
||||
| `VBITSETH` | |
|
||||
| `VBITSETV` | |
|
||||
| `VBITSETW` | |
|
||||
| `VDIVB` | |
|
||||
| `VDIVBU` | |
|
||||
| `VDIVD` | |
|
||||
| `VDIVF` | |
|
||||
| `VDIVH` | |
|
||||
| `VDIVHU` | |
|
||||
| `VDIVV` | |
|
||||
| `VDIVVU` | |
|
||||
| `VDIVW` | |
|
||||
| `VDIVWU` | |
|
||||
| `VEXTRINSB` | |
|
||||
| `VEXTRINSH` | |
|
||||
| `VEXTRINSV` | |
|
||||
| `VEXTRINSW` | |
|
||||
| `VFCLASSD` | |
|
||||
| `VFCLASSF` | |
|
||||
| `VFRECIPD` | |
|
||||
| `VFRECIPF` | |
|
||||
| `VFRINTD` | |
|
||||
| `VFRINTF` | |
|
||||
| `VFRINTRMD` | |
|
||||
| `VFRINTRMF` | |
|
||||
| `VFRINTRNED` | |
|
||||
| `VFRINTRNEF` | |
|
||||
| `VFRINTRPD` | |
|
||||
| `VFRINTRPF` | |
|
||||
| `VFRINTRZD` | |
|
||||
| `VFRINTRZF` | |
|
||||
| `VFRSQRTD` | |
|
||||
| `VFRSQRTF` | |
|
||||
| `VFSQRTD` | |
|
||||
| `VFSQRTF` | |
|
||||
| `VILVHB` | |
|
||||
| `VILVHH` | |
|
||||
| `VILVHV` | |
|
||||
| `VILVHW` | |
|
||||
| `VILVLB` | |
|
||||
| `VILVLH` | |
|
||||
| `VILVLV` | |
|
||||
| `VILVLW` | |
|
||||
| `VMADDB` | |
|
||||
| `VMADDH` | |
|
||||
| `VMADDV` | |
|
||||
| `VMADDW` | |
|
||||
| `VMADDWEVHB` | |
|
||||
| `VMADDWEVHBU` | |
|
||||
| `VMADDWEVHBUB` | |
|
||||
| `VMADDWEVQV` | |
|
||||
| `VMADDWEVQVU` | |
|
||||
| `VMADDWEVQVUV` | |
|
||||
| `VMADDWEVVW` | |
|
||||
| `VMADDWEVVWU` | |
|
||||
| `VMADDWEVVWUW` | |
|
||||
| `VMADDWEVWH` | |
|
||||
| `VMADDWEVWHU` | |
|
||||
| `VMADDWEVWHUH` | |
|
||||
| `VMADDWODHB` | |
|
||||
| `VMADDWODHBU` | |
|
||||
| `VMADDWODHBUB` | |
|
||||
| `VMADDWODQV` | |
|
||||
| `VMADDWODQVU` | |
|
||||
| `VMADDWODQVUV` | |
|
||||
| `VMADDWODVW` | |
|
||||
| `VMADDWODVWU` | |
|
||||
| `VMADDWODVWUW` | |
|
||||
| `VMADDWODWH` | |
|
||||
| `VMADDWODWHU` | |
|
||||
| `VMADDWODWHUH` | |
|
||||
| `VMODB` | |
|
||||
| `VMODBU` | |
|
||||
| `VMODH` | |
|
||||
| `VMODHU` | |
|
||||
| `VMODV` | |
|
||||
| `VMODVU` | |
|
||||
| `VMODW` | |
|
||||
| `VMODWU` | |
|
||||
| `VMOVQ` | |
|
||||
| `VMSUBB` | |
|
||||
| `VMSUBH` | |
|
||||
| `VMSUBV` | |
|
||||
| `VMSUBW` | |
|
||||
| `VMUHB` | |
|
||||
| `VMUHBU` | |
|
||||
| `VMUHH` | |
|
||||
| `VMUHHU` | |
|
||||
| `VMUHV` | |
|
||||
| `VMUHVU` | |
|
||||
| `VMUHW` | |
|
||||
| `VMUHWU` | |
|
||||
| `VMULB` | |
|
||||
| `VMULD` | |
|
||||
| `VMULF` | |
|
||||
| `VMULH` | |
|
||||
| `VMULV` | |
|
||||
| `VMULW` | |
|
||||
| `VMULWEVHB` | |
|
||||
| `VMULWEVHBU` | |
|
||||
| `VMULWEVHBUB` | |
|
||||
| `VMULWEVQV` | |
|
||||
| `VMULWEVQVU` | |
|
||||
| `VMULWEVQVUV` | |
|
||||
| `VMULWEVVW` | |
|
||||
| `VMULWEVVWU` | |
|
||||
| `VMULWEVVWUW` | |
|
||||
| `VMULWEVWH` | |
|
||||
| `VMULWEVWHU` | |
|
||||
| `VMULWEVWHUH` | |
|
||||
| `VMULWODHB` | |
|
||||
| `VMULWODHBU` | |
|
||||
| `VMULWODHBUB` | |
|
||||
| `VMULWODQV` | |
|
||||
| `VMULWODQVU` | |
|
||||
| `VMULWODQVUV` | |
|
||||
| `VMULWODVW` | |
|
||||
| `VMULWODVWU` | |
|
||||
| `VMULWODVWUW` | |
|
||||
| `VMULWODWH` | |
|
||||
| `VMULWODWHU` | |
|
||||
| `VMULWODWHUH` | |
|
||||
| `VNEGB` | |
|
||||
| `VNEGH` | |
|
||||
| `VNEGV` | |
|
||||
| `VNEGW` | |
|
||||
| `VNORB` | |
|
||||
| `VNORV` | |
|
||||
| `VORB` | |
|
||||
| `VORNV` | |
|
||||
| `VORV` | |
|
||||
| `VPCNTB` | |
|
||||
| `VPCNTH` | |
|
||||
| `VPCNTV` | |
|
||||
| `VPCNTW` | |
|
||||
| `VPERMIW` | |
|
||||
| `VROTRB` | |
|
||||
| `VROTRH` | |
|
||||
| `VROTRV` | |
|
||||
| `VROTRW` | |
|
||||
| `VSADDB` | |
|
||||
| `VSADDBU` | |
|
||||
| `VSADDH` | |
|
||||
| `VSADDHU` | |
|
||||
| `VSADDV` | |
|
||||
| `VSADDVU` | |
|
||||
| `VSADDW` | |
|
||||
| `VSADDWU` | |
|
||||
| `VSEQB` | |
|
||||
| `VSEQH` | |
|
||||
| `VSEQV` | |
|
||||
| `VSEQW` | |
|
||||
| `VSETALLNEB` | |
|
||||
| `VSETALLNEH` | |
|
||||
| `VSETALLNEV` | |
|
||||
| `VSETALLNEW` | |
|
||||
| `VSETANYEQB` | |
|
||||
| `VSETANYEQH` | |
|
||||
| `VSETANYEQV` | |
|
||||
| `VSETANYEQW` | |
|
||||
| `VSETEQV` | |
|
||||
| `VSETNEV` | |
|
||||
| `VSHUF4IB` | |
|
||||
| `VSHUF4IH` | |
|
||||
| `VSHUF4IV` | |
|
||||
| `VSHUF4IW` | |
|
||||
| `VSHUFB` | |
|
||||
| `VSHUFH` | |
|
||||
| `VSHUFV` | |
|
||||
| `VSHUFW` | |
|
||||
| `VSLLB` | |
|
||||
| `VSLLH` | |
|
||||
| `VSLLV` | |
|
||||
| `VSLLW` | |
|
||||
| `VSLTB` | |
|
||||
| `VSLTBU` | |
|
||||
| `VSLTH` | |
|
||||
| `VSLTHU` | |
|
||||
| `VSLTV` | |
|
||||
| `VSLTVU` | |
|
||||
| `VSLTW` | |
|
||||
| `VSLTWU` | |
|
||||
| `VSRAB` | |
|
||||
| `VSRAH` | |
|
||||
| `VSRAV` | |
|
||||
| `VSRAW` | |
|
||||
| `VSRLB` | |
|
||||
| `VSRLH` | |
|
||||
| `VSRLV` | |
|
||||
| `VSRLW` | |
|
||||
| `VSSUBB` | |
|
||||
| `VSSUBBU` | |
|
||||
| `VSSUBH` | |
|
||||
| `VSSUBHU` | |
|
||||
| `VSSUBV` | |
|
||||
| `VSSUBVU` | |
|
||||
| `VSSUBW` | |
|
||||
| `VSSUBWU` | |
|
||||
| `VSUBB` | |
|
||||
| `VSUBBU` | |
|
||||
| `VSUBD` | |
|
||||
| `VSUBF` | |
|
||||
| `VSUBH` | |
|
||||
| `VSUBHU` | |
|
||||
| `VSUBQ` | |
|
||||
| `VSUBV` | |
|
||||
| `VSUBVU` | |
|
||||
| `VSUBW` | |
|
||||
| `VSUBWEVHB` | |
|
||||
| `VSUBWEVHBU` | |
|
||||
| `VSUBWEVQV` | |
|
||||
| `VSUBWEVQVU` | |
|
||||
| `VSUBWEVVW` | |
|
||||
| `VSUBWEVVWU` | |
|
||||
| `VSUBWEVWH` | |
|
||||
| `VSUBWEVWHU` | |
|
||||
| `VSUBWODHB` | |
|
||||
| `VSUBWODHBU` | |
|
||||
| `VSUBWODQV` | |
|
||||
| `VSUBWODQVU` | |
|
||||
| `VSUBWODVW` | |
|
||||
| `VSUBWODVWU` | |
|
||||
| `VSUBWODWH` | |
|
||||
| `VSUBWODWHU` | |
|
||||
| `VSUBWU` | |
|
||||
| `VXORB` | |
|
||||
| `VXORV` | |
|
||||
| `WORD` | |
|
||||
| `XOR` | Bitwise XOR |
|
||||
| `XVADDB` | |
|
||||
| `XVADDBU` | |
|
||||
| `XVADDD` | |
|
||||
| `XVADDF` | |
|
||||
| `XVADDH` | |
|
||||
| `XVADDHU` | |
|
||||
| `XVADDQ` | |
|
||||
| `XVADDV` | |
|
||||
| `XVADDVU` | |
|
||||
| `XVADDW` | |
|
||||
| `XVADDWEVHB` | |
|
||||
| `XVADDWEVHBU` | |
|
||||
| `XVADDWEVQV` | |
|
||||
| `XVADDWEVQVU` | |
|
||||
| `XVADDWEVVW` | |
|
||||
| `XVADDWEVVWU` | |
|
||||
| `XVADDWEVWH` | |
|
||||
| `XVADDWEVWHU` | |
|
||||
| `XVADDWODHB` | |
|
||||
| `XVADDWODHBU` | |
|
||||
| `XVADDWODQV` | |
|
||||
| `XVADDWODQVU` | |
|
||||
| `XVADDWODVW` | |
|
||||
| `XVADDWODVWU` | |
|
||||
| `XVADDWODWH` | |
|
||||
| `XVADDWODWHU` | |
|
||||
| `XVADDWU` | |
|
||||
| `XVANDB` | |
|
||||
| `XVANDNV` | |
|
||||
| `XVANDV` | |
|
||||
| `XVBITCLRB` | |
|
||||
| `XVBITCLRH` | |
|
||||
| `XVBITCLRV` | |
|
||||
| `XVBITCLRW` | |
|
||||
| `XVBITREVB` | |
|
||||
| `XVBITREVH` | |
|
||||
| `XVBITREVV` | |
|
||||
| `XVBITREVW` | |
|
||||
| `XVBITSETB` | |
|
||||
| `XVBITSETH` | |
|
||||
| `XVBITSETV` | |
|
||||
| `XVBITSETW` | |
|
||||
| `XVDIVB` | |
|
||||
| `XVDIVBU` | |
|
||||
| `XVDIVD` | |
|
||||
| `XVDIVF` | |
|
||||
| `XVDIVH` | |
|
||||
| `XVDIVHU` | |
|
||||
| `XVDIVV` | |
|
||||
| `XVDIVVU` | |
|
||||
| `XVDIVW` | |
|
||||
| `XVDIVWU` | |
|
||||
| `XVEXTRINSB` | |
|
||||
| `XVEXTRINSH` | |
|
||||
| `XVEXTRINSV` | |
|
||||
| `XVEXTRINSW` | |
|
||||
| `XVFCLASSD` | |
|
||||
| `XVFCLASSF` | |
|
||||
| `XVFRECIPD` | |
|
||||
| `XVFRECIPF` | |
|
||||
| `XVFRINTD` | |
|
||||
| `XVFRINTF` | |
|
||||
| `XVFRINTRMD` | |
|
||||
| `XVFRINTRMF` | |
|
||||
| `XVFRINTRNED` | |
|
||||
| `XVFRINTRNEF` | |
|
||||
| `XVFRINTRPD` | |
|
||||
| `XVFRINTRPF` | |
|
||||
| `XVFRINTRZD` | |
|
||||
| `XVFRINTRZF` | |
|
||||
| `XVFRSQRTD` | |
|
||||
| `XVFRSQRTF` | |
|
||||
| `XVFSQRTD` | |
|
||||
| `XVFSQRTF` | |
|
||||
| `XVILVHB` | |
|
||||
| `XVILVHH` | |
|
||||
| `XVILVHV` | |
|
||||
| `XVILVHW` | |
|
||||
| `XVILVLB` | |
|
||||
| `XVILVLH` | |
|
||||
| `XVILVLV` | |
|
||||
| `XVILVLW` | |
|
||||
| `XVMADDB` | |
|
||||
| `XVMADDH` | |
|
||||
| `XVMADDV` | |
|
||||
| `XVMADDW` | |
|
||||
| `XVMADDWEVHB` | |
|
||||
| `XVMADDWEVHBU` | |
|
||||
| `XVMADDWEVHBUB` | |
|
||||
| `XVMADDWEVQV` | |
|
||||
| `XVMADDWEVQVU` | |
|
||||
| `XVMADDWEVQVUV` | |
|
||||
| `XVMADDWEVVW` | |
|
||||
| `XVMADDWEVVWU` | |
|
||||
| `XVMADDWEVVWUW` | |
|
||||
| `XVMADDWEVWH` | |
|
||||
| `XVMADDWEVWHU` | |
|
||||
| `XVMADDWEVWHUH` | |
|
||||
| `XVMADDWODHB` | |
|
||||
| `XVMADDWODHBU` | |
|
||||
| `XVMADDWODHBUB` | |
|
||||
| `XVMADDWODQV` | |
|
||||
| `XVMADDWODQVU` | |
|
||||
| `XVMADDWODQVUV` | |
|
||||
| `XVMADDWODVW` | |
|
||||
| `XVMADDWODVWU` | |
|
||||
| `XVMADDWODVWUW` | |
|
||||
| `XVMADDWODWH` | |
|
||||
| `XVMADDWODWHU` | |
|
||||
| `XVMADDWODWHUH` | |
|
||||
| `XVMODB` | |
|
||||
| `XVMODBU` | |
|
||||
| `XVMODH` | |
|
||||
| `XVMODHU` | |
|
||||
| `XVMODV` | |
|
||||
| `XVMODVU` | |
|
||||
| `XVMODW` | |
|
||||
| `XVMODWU` | |
|
||||
| `XVMOVQ` | |
|
||||
| `XVMSUBB` | |
|
||||
| `XVMSUBH` | |
|
||||
| `XVMSUBV` | |
|
||||
| `XVMSUBW` | |
|
||||
| `XVMUHB` | |
|
||||
| `XVMUHBU` | |
|
||||
| `XVMUHH` | |
|
||||
| `XVMUHHU` | |
|
||||
| `XVMUHV` | |
|
||||
| `XVMUHVU` | |
|
||||
| `XVMUHW` | |
|
||||
| `XVMUHWU` | |
|
||||
| `XVMULB` | |
|
||||
| `XVMULD` | |
|
||||
| `XVMULF` | |
|
||||
| `XVMULH` | |
|
||||
| `XVMULV` | |
|
||||
| `XVMULW` | |
|
||||
| `XVMULWEVHB` | |
|
||||
| `XVMULWEVHBU` | |
|
||||
| `XVMULWEVHBUB` | |
|
||||
| `XVMULWEVQV` | |
|
||||
| `XVMULWEVQVU` | |
|
||||
| `XVMULWEVQVUV` | |
|
||||
| `XVMULWEVVW` | |
|
||||
| `XVMULWEVVWU` | |
|
||||
| `XVMULWEVVWUW` | |
|
||||
| `XVMULWEVWH` | |
|
||||
| `XVMULWEVWHU` | |
|
||||
| `XVMULWEVWHUH` | |
|
||||
| `XVMULWODHB` | |
|
||||
| `XVMULWODHBU` | |
|
||||
| `XVMULWODHBUB` | |
|
||||
| `XVMULWODQV` | |
|
||||
| `XVMULWODQVU` | |
|
||||
| `XVMULWODQVUV` | |
|
||||
| `XVMULWODVW` | |
|
||||
| `XVMULWODVWU` | |
|
||||
| `XVMULWODVWUW` | |
|
||||
| `XVMULWODWH` | |
|
||||
| `XVMULWODWHU` | |
|
||||
| `XVMULWODWHUH` | |
|
||||
| `XVNEGB` | |
|
||||
| `XVNEGH` | |
|
||||
| `XVNEGV` | |
|
||||
| `XVNEGW` | |
|
||||
| `XVNORB` | |
|
||||
| `XVNORV` | |
|
||||
| `XVORB` | |
|
||||
| `XVORNV` | |
|
||||
| `XVORV` | |
|
||||
| `XVPCNTB` | |
|
||||
| `XVPCNTH` | |
|
||||
| `XVPCNTV` | |
|
||||
| `XVPCNTW` | |
|
||||
| `XVPERMIQ` | |
|
||||
| `XVPERMIV` | |
|
||||
| `XVPERMIW` | |
|
||||
| `XVROTRB` | |
|
||||
| `XVROTRH` | |
|
||||
| `XVROTRV` | |
|
||||
| `XVROTRW` | |
|
||||
| `XVSADDB` | |
|
||||
| `XVSADDBU` | |
|
||||
| `XVSADDH` | |
|
||||
| `XVSADDHU` | |
|
||||
| `XVSADDV` | |
|
||||
| `XVSADDVU` | |
|
||||
| `XVSADDW` | |
|
||||
| `XVSADDWU` | |
|
||||
| `XVSEQB` | |
|
||||
| `XVSEQH` | |
|
||||
| `XVSEQV` | |
|
||||
| `XVSEQW` | |
|
||||
| `XVSETALLNEB` | |
|
||||
| `XVSETALLNEH` | |
|
||||
| `XVSETALLNEV` | |
|
||||
| `XVSETALLNEW` | |
|
||||
| `XVSETANYEQB` | |
|
||||
| `XVSETANYEQH` | |
|
||||
| `XVSETANYEQV` | |
|
||||
| `XVSETANYEQW` | |
|
||||
| `XVSETEQV` | |
|
||||
| `XVSETNEV` | |
|
||||
| `XVSHUF4IB` | |
|
||||
| `XVSHUF4IH` | |
|
||||
| `XVSHUF4IV` | |
|
||||
| `XVSHUF4IW` | |
|
||||
| `XVSHUFB` | |
|
||||
| `XVSHUFH` | |
|
||||
| `XVSHUFV` | |
|
||||
| `XVSHUFW` | |
|
||||
| `XVSLLB` | |
|
||||
| `XVSLLH` | |
|
||||
| `XVSLLV` | |
|
||||
| `XVSLLW` | |
|
||||
| `XVSLTB` | |
|
||||
| `XVSLTBU` | |
|
||||
| `XVSLTH` | |
|
||||
| `XVSLTHU` | |
|
||||
| `XVSLTV` | |
|
||||
| `XVSLTVU` | |
|
||||
| `XVSLTW` | |
|
||||
| `XVSLTWU` | |
|
||||
| `XVSRAB` | |
|
||||
| `XVSRAH` | |
|
||||
| `XVSRAV` | |
|
||||
| `XVSRAW` | |
|
||||
| `XVSRLB` | |
|
||||
| `XVSRLH` | |
|
||||
| `XVSRLV` | |
|
||||
| `XVSRLW` | |
|
||||
| `XVSSUBB` | |
|
||||
| `XVSSUBBU` | |
|
||||
| `XVSSUBH` | |
|
||||
| `XVSSUBHU` | |
|
||||
| `XVSSUBV` | |
|
||||
| `XVSSUBVU` | |
|
||||
| `XVSSUBW` | |
|
||||
| `XVSSUBWU` | |
|
||||
| `XVSUBB` | |
|
||||
| `XVSUBBU` | |
|
||||
| `XVSUBD` | |
|
||||
| `XVSUBF` | |
|
||||
| `XVSUBH` | |
|
||||
| `XVSUBHU` | |
|
||||
| `XVSUBQ` | |
|
||||
| `XVSUBV` | |
|
||||
| `XVSUBVU` | |
|
||||
| `XVSUBW` | |
|
||||
| `XVSUBWEVHB` | |
|
||||
| `XVSUBWEVHBU` | |
|
||||
| `XVSUBWEVQV` | |
|
||||
| `XVSUBWEVQVU` | |
|
||||
| `XVSUBWEVVW` | |
|
||||
| `XVSUBWEVVWU` | |
|
||||
| `XVSUBWEVWH` | |
|
||||
| `XVSUBWEVWHU` | |
|
||||
| `XVSUBWODHB` | |
|
||||
| `XVSUBWODHBU` | |
|
||||
| `XVSUBWODQV` | |
|
||||
| `XVSUBWODQVU` | |
|
||||
| `XVSUBWODVW` | |
|
||||
| `XVSUBWODVWU` | |
|
||||
| `XVSUBWODWH` | |
|
||||
| `XVSUBWODWHU` | |
|
||||
| `XVSUBWU` | |
|
||||
| `XVXORB` | |
|
||||
| `XVXORV` | |
|
||||
| `JAL` | |
|
||||
|
||||
Recognised: 814 mnemonics.
|
||||
@@ -0,0 +1,991 @@
|
||||
# RISC-V 64: instruction inventory
|
||||
|
||||
Generated by gasm-devkit's `_gen` from the Go toolchain's instruction table
|
||||
(`cmd/internal/obj/riscv/anames.go`, go1.27.1); DO NOT EDIT. This page lists every mnemonic
|
||||
`go tool asm` accepts on this target, which is the upper bound of the
|
||||
language on it: a name absent here is not an instruction of the target,
|
||||
and a name present here may still be one gasm's encoder cannot emit yet.
|
||||
|
||||
The inventory carries no per-mnemonic encoder column: on this target
|
||||
encodability is decided per operand shape, and the live measured
|
||||
coverage is reported by `gasm audit-instructions`.
|
||||
|
||||
| Mnemonic | Notes |
|
||||
|---|---|
|
||||
| `CALL` | Call subroutine |
|
||||
| `DUFFCOPY` | |
|
||||
| `DUFFZERO` | |
|
||||
| `END` | |
|
||||
| `FUNCDATA` | |
|
||||
| `GETCALLERPC` | |
|
||||
| `JMP` | Unconditional jump |
|
||||
| `NOP` | |
|
||||
| `PCALIGN` | |
|
||||
| `PCALIGNMAX` | |
|
||||
| `PCDATA` | |
|
||||
| `RET` | Return |
|
||||
| `TEXT` | |
|
||||
| `UNDEF` | |
|
||||
| `ADD` | Integer add |
|
||||
| `ADDI` | Add immediate |
|
||||
| `ADDIW` | Add immediate (32-bit) |
|
||||
| `ADDUW` | |
|
||||
| `ADDW` | Add (32-bit) |
|
||||
| `AMOADDD` | Atomic add doubleword |
|
||||
| `AMOADDW` | Atomic add word |
|
||||
| `AMOANDD` | |
|
||||
| `AMOANDW` | |
|
||||
| `AMOMAXD` | |
|
||||
| `AMOMAXUD` | |
|
||||
| `AMOMAXUW` | |
|
||||
| `AMOMAXW` | |
|
||||
| `AMOMIND` | |
|
||||
| `AMOMINUD` | |
|
||||
| `AMOMINUW` | |
|
||||
| `AMOMINW` | |
|
||||
| `AMOORD` | |
|
||||
| `AMOORW` | |
|
||||
| `AMOSWAPD` | Atomic swap doubleword |
|
||||
| `AMOSWAPW` | Atomic swap word |
|
||||
| `AMOXORD` | |
|
||||
| `AMOXORW` | |
|
||||
| `AND` | Bitwise AND |
|
||||
| `ANDI` | AND immediate |
|
||||
| `ANDN` | |
|
||||
| `AUIPC` | Add upper immediate to PC |
|
||||
| `BCLR` | |
|
||||
| `BCLRI` | |
|
||||
| `BEQ` | Branch if equal |
|
||||
| `BEQZ` | |
|
||||
| `BEXT` | |
|
||||
| `BEXTI` | |
|
||||
| `BGE` | Branch if greater or equal |
|
||||
| `BGEU` | Branch if greater or equal unsigned |
|
||||
| `BGEZ` | |
|
||||
| `BGT` | |
|
||||
| `BGTU` | |
|
||||
| `BGTZ` | |
|
||||
| `BINV` | |
|
||||
| `BINVI` | |
|
||||
| `BLE` | |
|
||||
| `BLEU` | |
|
||||
| `BLEZ` | |
|
||||
| `BLT` | Branch if less than |
|
||||
| `BLTU` | Branch if less than unsigned |
|
||||
| `BLTZ` | |
|
||||
| `BNE` | Branch if not equal |
|
||||
| `BNEZ` | |
|
||||
| `BSET` | |
|
||||
| `BSETI` | |
|
||||
| `CADD` | |
|
||||
| `CADDI` | |
|
||||
| `CADDI16SP` | |
|
||||
| `CADDI4SPN` | |
|
||||
| `CADDIW` | |
|
||||
| `CADDW` | |
|
||||
| `CAND` | |
|
||||
| `CANDI` | |
|
||||
| `CBEQZ` | |
|
||||
| `CBNEZ` | |
|
||||
| `CEBREAK` | |
|
||||
| `CFLD` | |
|
||||
| `CFLDSP` | |
|
||||
| `CFSD` | |
|
||||
| `CFSDSP` | |
|
||||
| `CJ` | |
|
||||
| `CJALR` | |
|
||||
| `CJR` | |
|
||||
| `CLD` | |
|
||||
| `CLDSP` | |
|
||||
| `CLI` | |
|
||||
| `CLUI` | |
|
||||
| `CLW` | |
|
||||
| `CLWSP` | |
|
||||
| `CLZ` | |
|
||||
| `CLZW` | |
|
||||
| `CMV` | |
|
||||
| `CNOP` | |
|
||||
| `COR` | |
|
||||
| `CPOP` | |
|
||||
| `CPOPW` | |
|
||||
| `CSD` | |
|
||||
| `CSDSP` | |
|
||||
| `CSLLI` | |
|
||||
| `CSRAI` | |
|
||||
| `CSRLI` | |
|
||||
| `CSRRC` | |
|
||||
| `CSRRCI` | |
|
||||
| `CSRRS` | |
|
||||
| `CSRRSI` | |
|
||||
| `CSRRW` | |
|
||||
| `CSRRWI` | |
|
||||
| `CSUB` | |
|
||||
| `CSUBW` | |
|
||||
| `CSW` | |
|
||||
| `CSWSP` | |
|
||||
| `CTZ` | |
|
||||
| `CTZW` | |
|
||||
| `CXOR` | |
|
||||
| `CZEROEQZ` | |
|
||||
| `CZERONEZ` | |
|
||||
| `DIV` | Divide |
|
||||
| `DIVU` | Divide unsigned |
|
||||
| `DIVUW` | |
|
||||
| `DIVW` | Divide (32-bit) |
|
||||
| `DRET` | |
|
||||
| `EBREAK` | Breakpoint |
|
||||
| `ECALL` | Environment call |
|
||||
| `FABSD` | |
|
||||
| `FABSS` | |
|
||||
| `FADDD` | FP add (double) |
|
||||
| `FADDQ` | |
|
||||
| `FADDS` | FP add (single) |
|
||||
| `FCLASSD` | |
|
||||
| `FCLASSQ` | |
|
||||
| `FCLASSS` | |
|
||||
| `FCVTDL` | |
|
||||
| `FCVTDLU` | |
|
||||
| `FCVTDQ` | |
|
||||
| `FCVTDS` | |
|
||||
| `FCVTDW` | |
|
||||
| `FCVTDWU` | |
|
||||
| `FCVTLD` | |
|
||||
| `FCVTLQ` | |
|
||||
| `FCVTLS` | |
|
||||
| `FCVTLUD` | |
|
||||
| `FCVTLUQ` | |
|
||||
| `FCVTLUS` | |
|
||||
| `FCVTQD` | |
|
||||
| `FCVTQL` | |
|
||||
| `FCVTQLU` | |
|
||||
| `FCVTQS` | |
|
||||
| `FCVTQW` | |
|
||||
| `FCVTQWU` | |
|
||||
| `FCVTSD` | |
|
||||
| `FCVTSL` | |
|
||||
| `FCVTSLU` | |
|
||||
| `FCVTSQ` | |
|
||||
| `FCVTSW` | |
|
||||
| `FCVTSWU` | |
|
||||
| `FCVTWD` | |
|
||||
| `FCVTWQ` | |
|
||||
| `FCVTWS` | |
|
||||
| `FCVTWUD` | |
|
||||
| `FCVTWUQ` | |
|
||||
| `FCVTWUS` | |
|
||||
| `FDIVD` | FP divide (double) |
|
||||
| `FDIVQ` | |
|
||||
| `FDIVS` | FP divide (single) |
|
||||
| `FENCE` | Memory barrier |
|
||||
| `FEQD` | |
|
||||
| `FEQQ` | |
|
||||
| `FEQS` | |
|
||||
| `FLD` | FP load doubleword |
|
||||
| `FLED` | |
|
||||
| `FLEQ` | |
|
||||
| `FLES` | |
|
||||
| `FLQ` | |
|
||||
| `FLTD` | |
|
||||
| `FLTQ` | |
|
||||
| `FLTS` | |
|
||||
| `FLW` | FP load word |
|
||||
| `FMADDD` | |
|
||||
| `FMADDQ` | |
|
||||
| `FMADDS` | |
|
||||
| `FMAXD` | |
|
||||
| `FMAXQ` | |
|
||||
| `FMAXS` | |
|
||||
| `FMIND` | |
|
||||
| `FMINQ` | |
|
||||
| `FMINS` | |
|
||||
| `FMSUBD` | |
|
||||
| `FMSUBQ` | |
|
||||
| `FMSUBS` | |
|
||||
| `FMULD` | FP multiply (double) |
|
||||
| `FMULQ` | |
|
||||
| `FMULS` | FP multiply (single) |
|
||||
| `FMVDX` | |
|
||||
| `FMVSX` | |
|
||||
| `FMVWX` | |
|
||||
| `FMVXD` | |
|
||||
| `FMVXS` | |
|
||||
| `FMVXW` | |
|
||||
| `FNED` | |
|
||||
| `FNEGD` | |
|
||||
| `FNEGS` | |
|
||||
| `FNES` | |
|
||||
| `FNMADDD` | |
|
||||
| `FNMADDQ` | |
|
||||
| `FNMADDS` | |
|
||||
| `FNMSUBD` | |
|
||||
| `FNMSUBQ` | |
|
||||
| `FNMSUBS` | |
|
||||
| `FSD` | FP store doubleword |
|
||||
| `FSGNJD` | |
|
||||
| `FSGNJND` | |
|
||||
| `FSGNJNQ` | |
|
||||
| `FSGNJNS` | |
|
||||
| `FSGNJQ` | |
|
||||
| `FSGNJS` | |
|
||||
| `FSGNJXD` | |
|
||||
| `FSGNJXQ` | |
|
||||
| `FSGNJXS` | |
|
||||
| `FSQ` | |
|
||||
| `FSQRTD` | |
|
||||
| `FSQRTQ` | |
|
||||
| `FSQRTS` | |
|
||||
| `FSUBD` | FP subtract (double) |
|
||||
| `FSUBQ` | |
|
||||
| `FSUBS` | FP subtract (single) |
|
||||
| `FSW` | FP store word |
|
||||
| `JAL` | Jump and link |
|
||||
| `JALR` | Jump and link register |
|
||||
| `LB` | Load byte |
|
||||
| `LBU` | Load byte unsigned |
|
||||
| `LD` | Load doubleword |
|
||||
| `LH` | Load halfword |
|
||||
| `LHU` | Load halfword unsigned |
|
||||
| `LRD` | Load-reserved doubleword |
|
||||
| `LRW` | Load-reserved word |
|
||||
| `LUI` | Load upper immediate |
|
||||
| `LW` | Load word |
|
||||
| `LWU` | Load word unsigned |
|
||||
| `MAX` | |
|
||||
| `MAXU` | |
|
||||
| `MIN` | |
|
||||
| `MINU` | |
|
||||
| `MOV` | |
|
||||
| `MOVB` | |
|
||||
| `MOVBU` | |
|
||||
| `MOVD` | |
|
||||
| `MOVF` | |
|
||||
| `MOVH` | |
|
||||
| `MOVHU` | |
|
||||
| `MOVW` | |
|
||||
| `MOVWU` | |
|
||||
| `MRET` | |
|
||||
| `MUL` | Multiply |
|
||||
| `MULH` | Multiply high |
|
||||
| `MULHSU` | Multiply high signed/unsigned |
|
||||
| `MULHU` | Multiply high unsigned |
|
||||
| `MULW` | Multiply (32-bit) |
|
||||
| `NEG` | |
|
||||
| `NEGW` | |
|
||||
| `NOT` | |
|
||||
| `OR` | Bitwise OR |
|
||||
| `ORCB` | |
|
||||
| `ORI` | OR immediate |
|
||||
| `ORN` | |
|
||||
| `RDCYCLE` | |
|
||||
| `RDINSTRET` | |
|
||||
| `RDTIME` | |
|
||||
| `REM` | Remainder |
|
||||
| `REMU` | Remainder unsigned |
|
||||
| `REMUW` | |
|
||||
| `REMW` | |
|
||||
| `REV8` | |
|
||||
| `ROL` | |
|
||||
| `ROLW` | |
|
||||
| `ROR` | |
|
||||
| `RORI` | |
|
||||
| `RORIW` | |
|
||||
| `RORW` | |
|
||||
| `SB` | Store byte |
|
||||
| `SBREAK` | |
|
||||
| `SCALL` | |
|
||||
| `SCD` | Store-conditional doubleword |
|
||||
| `SCW` | Store-conditional word |
|
||||
| `SD` | Store doubleword |
|
||||
| `SEQZ` | |
|
||||
| `SEXTB` | |
|
||||
| `SEXTH` | |
|
||||
| `SFENCEVMA` | |
|
||||
| `SH` | Store halfword |
|
||||
| `SH1ADD` | |
|
||||
| `SH1ADDUW` | |
|
||||
| `SH2ADD` | |
|
||||
| `SH2ADDUW` | |
|
||||
| `SH3ADD` | |
|
||||
| `SH3ADDUW` | |
|
||||
| `SLL` | Shift left logical |
|
||||
| `SLLI` | Shift left logical immediate |
|
||||
| `SLLIUW` | |
|
||||
| `SLLIW` | |
|
||||
| `SLLW` | |
|
||||
| `SLT` | Set if less than |
|
||||
| `SLTI` | Set if less than immediate |
|
||||
| `SLTIU` | Set if less than unsigned immediate |
|
||||
| `SLTU` | Set if less than unsigned |
|
||||
| `SNEZ` | |
|
||||
| `SRA` | Shift right arithmetic |
|
||||
| `SRAI` | Shift right arithmetic immediate |
|
||||
| `SRAIW` | |
|
||||
| `SRAW` | |
|
||||
| `SRET` | |
|
||||
| `SRL` | Shift right logical |
|
||||
| `SRLI` | Shift right logical immediate |
|
||||
| `SRLIW` | |
|
||||
| `SRLW` | |
|
||||
| `SUB` | Integer subtract |
|
||||
| `SUBW` | Subtract (32-bit) |
|
||||
| `SW` | Store word |
|
||||
| `VAADDUVV` | |
|
||||
| `VAADDUVX` | |
|
||||
| `VAADDVV` | |
|
||||
| `VAADDVX` | |
|
||||
| `VADCVIM` | |
|
||||
| `VADCVVM` | |
|
||||
| `VADCVXM` | |
|
||||
| `VADDVI` | |
|
||||
| `VADDVV` | |
|
||||
| `VADDVX` | |
|
||||
| `VANDVI` | |
|
||||
| `VANDVV` | |
|
||||
| `VANDVX` | |
|
||||
| `VASUBUVV` | |
|
||||
| `VASUBUVX` | |
|
||||
| `VASUBVV` | |
|
||||
| `VASUBVX` | |
|
||||
| `VCOMPRESSVM` | |
|
||||
| `VCPOPM` | |
|
||||
| `VDIVUVV` | |
|
||||
| `VDIVUVX` | |
|
||||
| `VDIVVV` | |
|
||||
| `VDIVVX` | |
|
||||
| `VFABSV` | |
|
||||
| `VFADDVF` | |
|
||||
| `VFADDVV` | |
|
||||
| `VFCLASSV` | |
|
||||
| `VFCVTFXUV` | |
|
||||
| `VFCVTFXV` | |
|
||||
| `VFCVTRTZXFV` | |
|
||||
| `VFCVTRTZXUFV` | |
|
||||
| `VFCVTXFV` | |
|
||||
| `VFCVTXUFV` | |
|
||||
| `VFDIVVF` | |
|
||||
| `VFDIVVV` | |
|
||||
| `VFIRSTM` | |
|
||||
| `VFMACCVF` | |
|
||||
| `VFMACCVV` | |
|
||||
| `VFMADDVF` | |
|
||||
| `VFMADDVV` | |
|
||||
| `VFMAXVF` | |
|
||||
| `VFMAXVV` | |
|
||||
| `VFMERGEVFM` | |
|
||||
| `VFMINVF` | |
|
||||
| `VFMINVV` | |
|
||||
| `VFMSACVF` | |
|
||||
| `VFMSACVV` | |
|
||||
| `VFMSUBVF` | |
|
||||
| `VFMSUBVV` | |
|
||||
| `VFMULVF` | |
|
||||
| `VFMULVV` | |
|
||||
| `VFMVFS` | |
|
||||
| `VFMVSF` | |
|
||||
| `VFMVVF` | |
|
||||
| `VFNCVTFFW` | |
|
||||
| `VFNCVTFXUW` | |
|
||||
| `VFNCVTFXW` | |
|
||||
| `VFNCVTRODFFW` | |
|
||||
| `VFNCVTRTZXFW` | |
|
||||
| `VFNCVTRTZXUFW` | |
|
||||
| `VFNCVTXFW` | |
|
||||
| `VFNCVTXUFW` | |
|
||||
| `VFNEGV` | |
|
||||
| `VFNMACCVF` | |
|
||||
| `VFNMACCVV` | |
|
||||
| `VFNMADDVF` | |
|
||||
| `VFNMADDVV` | |
|
||||
| `VFNMSACVF` | |
|
||||
| `VFNMSACVV` | |
|
||||
| `VFNMSUBVF` | |
|
||||
| `VFNMSUBVV` | |
|
||||
| `VFRDIVVF` | |
|
||||
| `VFREC7V` | |
|
||||
| `VFREDMAXVS` | |
|
||||
| `VFREDMINVS` | |
|
||||
| `VFREDOSUMVS` | |
|
||||
| `VFREDUSUMVS` | |
|
||||
| `VFRSQRT7V` | |
|
||||
| `VFRSUBVF` | |
|
||||
| `VFSGNJNVF` | |
|
||||
| `VFSGNJNVV` | |
|
||||
| `VFSGNJVF` | |
|
||||
| `VFSGNJVV` | |
|
||||
| `VFSGNJXVF` | |
|
||||
| `VFSGNJXVV` | |
|
||||
| `VFSLIDE1DOWNVF` | |
|
||||
| `VFSLIDE1UPVF` | |
|
||||
| `VFSQRTV` | |
|
||||
| `VFSUBVF` | |
|
||||
| `VFSUBVV` | |
|
||||
| `VFWADDVF` | |
|
||||
| `VFWADDVV` | |
|
||||
| `VFWADDWF` | |
|
||||
| `VFWADDWV` | |
|
||||
| `VFWCVTFFV` | |
|
||||
| `VFWCVTFXUV` | |
|
||||
| `VFWCVTFXV` | |
|
||||
| `VFWCVTRTZXFV` | |
|
||||
| `VFWCVTRTZXUFV` | |
|
||||
| `VFWCVTXFV` | |
|
||||
| `VFWCVTXUFV` | |
|
||||
| `VFWMACCVF` | |
|
||||
| `VFWMACCVV` | |
|
||||
| `VFWMSACVF` | |
|
||||
| `VFWMSACVV` | |
|
||||
| `VFWMULVF` | |
|
||||
| `VFWMULVV` | |
|
||||
| `VFWNMACCVF` | |
|
||||
| `VFWNMACCVV` | |
|
||||
| `VFWNMSACVF` | |
|
||||
| `VFWNMSACVV` | |
|
||||
| `VFWREDOSUMVS` | |
|
||||
| `VFWREDUSUMVS` | |
|
||||
| `VFWSUBVF` | |
|
||||
| `VFWSUBVV` | |
|
||||
| `VFWSUBWF` | |
|
||||
| `VFWSUBWV` | |
|
||||
| `VIDV` | |
|
||||
| `VIOTAM` | |
|
||||
| `VL1RE16V` | |
|
||||
| `VL1RE32V` | |
|
||||
| `VL1RE64V` | |
|
||||
| `VL1RE8V` | |
|
||||
| `VL1RV` | |
|
||||
| `VL2RE16V` | |
|
||||
| `VL2RE32V` | |
|
||||
| `VL2RE64V` | |
|
||||
| `VL2RE8V` | |
|
||||
| `VL2RV` | |
|
||||
| `VL4RE16V` | |
|
||||
| `VL4RE32V` | |
|
||||
| `VL4RE64V` | |
|
||||
| `VL4RE8V` | |
|
||||
| `VL4RV` | |
|
||||
| `VL8RE16V` | |
|
||||
| `VL8RE32V` | |
|
||||
| `VL8RE64V` | |
|
||||
| `VL8RE8V` | |
|
||||
| `VL8RV` | |
|
||||
| `VLE16FFV` | |
|
||||
| `VLE16V` | |
|
||||
| `VLE32FFV` | |
|
||||
| `VLE32V` | |
|
||||
| `VLE64FFV` | |
|
||||
| `VLE64V` | |
|
||||
| `VLE8FFV` | |
|
||||
| `VLE8V` | |
|
||||
| `VLMV` | |
|
||||
| `VLOXEI16V` | |
|
||||
| `VLOXEI32V` | |
|
||||
| `VLOXEI64V` | |
|
||||
| `VLOXEI8V` | |
|
||||
| `VLOXSEG2EI16V` | |
|
||||
| `VLOXSEG2EI32V` | |
|
||||
| `VLOXSEG2EI64V` | |
|
||||
| `VLOXSEG2EI8V` | |
|
||||
| `VLOXSEG3EI16V` | |
|
||||
| `VLOXSEG3EI32V` | |
|
||||
| `VLOXSEG3EI64V` | |
|
||||
| `VLOXSEG3EI8V` | |
|
||||
| `VLOXSEG4EI16V` | |
|
||||
| `VLOXSEG4EI32V` | |
|
||||
| `VLOXSEG4EI64V` | |
|
||||
| `VLOXSEG4EI8V` | |
|
||||
| `VLOXSEG5EI16V` | |
|
||||
| `VLOXSEG5EI32V` | |
|
||||
| `VLOXSEG5EI64V` | |
|
||||
| `VLOXSEG5EI8V` | |
|
||||
| `VLOXSEG6EI16V` | |
|
||||
| `VLOXSEG6EI32V` | |
|
||||
| `VLOXSEG6EI64V` | |
|
||||
| `VLOXSEG6EI8V` | |
|
||||
| `VLOXSEG7EI16V` | |
|
||||
| `VLOXSEG7EI32V` | |
|
||||
| `VLOXSEG7EI64V` | |
|
||||
| `VLOXSEG7EI8V` | |
|
||||
| `VLOXSEG8EI16V` | |
|
||||
| `VLOXSEG8EI32V` | |
|
||||
| `VLOXSEG8EI64V` | |
|
||||
| `VLOXSEG8EI8V` | |
|
||||
| `VLSE16V` | |
|
||||
| `VLSE32V` | |
|
||||
| `VLSE64V` | |
|
||||
| `VLSE8V` | |
|
||||
| `VLSEG2E16FFV` | |
|
||||
| `VLSEG2E16V` | |
|
||||
| `VLSEG2E32FFV` | |
|
||||
| `VLSEG2E32V` | |
|
||||
| `VLSEG2E64FFV` | |
|
||||
| `VLSEG2E64V` | |
|
||||
| `VLSEG2E8FFV` | |
|
||||
| `VLSEG2E8V` | |
|
||||
| `VLSEG3E16FFV` | |
|
||||
| `VLSEG3E16V` | |
|
||||
| `VLSEG3E32FFV` | |
|
||||
| `VLSEG3E32V` | |
|
||||
| `VLSEG3E64FFV` | |
|
||||
| `VLSEG3E64V` | |
|
||||
| `VLSEG3E8FFV` | |
|
||||
| `VLSEG3E8V` | |
|
||||
| `VLSEG4E16FFV` | |
|
||||
| `VLSEG4E16V` | |
|
||||
| `VLSEG4E32FFV` | |
|
||||
| `VLSEG4E32V` | |
|
||||
| `VLSEG4E64FFV` | |
|
||||
| `VLSEG4E64V` | |
|
||||
| `VLSEG4E8FFV` | |
|
||||
| `VLSEG4E8V` | |
|
||||
| `VLSEG5E16FFV` | |
|
||||
| `VLSEG5E16V` | |
|
||||
| `VLSEG5E32FFV` | |
|
||||
| `VLSEG5E32V` | |
|
||||
| `VLSEG5E64FFV` | |
|
||||
| `VLSEG5E64V` | |
|
||||
| `VLSEG5E8FFV` | |
|
||||
| `VLSEG5E8V` | |
|
||||
| `VLSEG6E16FFV` | |
|
||||
| `VLSEG6E16V` | |
|
||||
| `VLSEG6E32FFV` | |
|
||||
| `VLSEG6E32V` | |
|
||||
| `VLSEG6E64FFV` | |
|
||||
| `VLSEG6E64V` | |
|
||||
| `VLSEG6E8FFV` | |
|
||||
| `VLSEG6E8V` | |
|
||||
| `VLSEG7E16FFV` | |
|
||||
| `VLSEG7E16V` | |
|
||||
| `VLSEG7E32FFV` | |
|
||||
| `VLSEG7E32V` | |
|
||||
| `VLSEG7E64FFV` | |
|
||||
| `VLSEG7E64V` | |
|
||||
| `VLSEG7E8FFV` | |
|
||||
| `VLSEG7E8V` | |
|
||||
| `VLSEG8E16FFV` | |
|
||||
| `VLSEG8E16V` | |
|
||||
| `VLSEG8E32FFV` | |
|
||||
| `VLSEG8E32V` | |
|
||||
| `VLSEG8E64FFV` | |
|
||||
| `VLSEG8E64V` | |
|
||||
| `VLSEG8E8FFV` | |
|
||||
| `VLSEG8E8V` | |
|
||||
| `VLSSEG2E16V` | |
|
||||
| `VLSSEG2E32V` | |
|
||||
| `VLSSEG2E64V` | |
|
||||
| `VLSSEG2E8V` | |
|
||||
| `VLSSEG3E16V` | |
|
||||
| `VLSSEG3E32V` | |
|
||||
| `VLSSEG3E64V` | |
|
||||
| `VLSSEG3E8V` | |
|
||||
| `VLSSEG4E16V` | |
|
||||
| `VLSSEG4E32V` | |
|
||||
| `VLSSEG4E64V` | |
|
||||
| `VLSSEG4E8V` | |
|
||||
| `VLSSEG5E16V` | |
|
||||
| `VLSSEG5E32V` | |
|
||||
| `VLSSEG5E64V` | |
|
||||
| `VLSSEG5E8V` | |
|
||||
| `VLSSEG6E16V` | |
|
||||
| `VLSSEG6E32V` | |
|
||||
| `VLSSEG6E64V` | |
|
||||
| `VLSSEG6E8V` | |
|
||||
| `VLSSEG7E16V` | |
|
||||
| `VLSSEG7E32V` | |
|
||||
| `VLSSEG7E64V` | |
|
||||
| `VLSSEG7E8V` | |
|
||||
| `VLSSEG8E16V` | |
|
||||
| `VLSSEG8E32V` | |
|
||||
| `VLSSEG8E64V` | |
|
||||
| `VLSSEG8E8V` | |
|
||||
| `VLUXEI16V` | |
|
||||
| `VLUXEI32V` | |
|
||||
| `VLUXEI64V` | |
|
||||
| `VLUXEI8V` | |
|
||||
| `VLUXSEG2EI16V` | |
|
||||
| `VLUXSEG2EI32V` | |
|
||||
| `VLUXSEG2EI64V` | |
|
||||
| `VLUXSEG2EI8V` | |
|
||||
| `VLUXSEG3EI16V` | |
|
||||
| `VLUXSEG3EI32V` | |
|
||||
| `VLUXSEG3EI64V` | |
|
||||
| `VLUXSEG3EI8V` | |
|
||||
| `VLUXSEG4EI16V` | |
|
||||
| `VLUXSEG4EI32V` | |
|
||||
| `VLUXSEG4EI64V` | |
|
||||
| `VLUXSEG4EI8V` | |
|
||||
| `VLUXSEG5EI16V` | |
|
||||
| `VLUXSEG5EI32V` | |
|
||||
| `VLUXSEG5EI64V` | |
|
||||
| `VLUXSEG5EI8V` | |
|
||||
| `VLUXSEG6EI16V` | |
|
||||
| `VLUXSEG6EI32V` | |
|
||||
| `VLUXSEG6EI64V` | |
|
||||
| `VLUXSEG6EI8V` | |
|
||||
| `VLUXSEG7EI16V` | |
|
||||
| `VLUXSEG7EI32V` | |
|
||||
| `VLUXSEG7EI64V` | |
|
||||
| `VLUXSEG7EI8V` | |
|
||||
| `VLUXSEG8EI16V` | |
|
||||
| `VLUXSEG8EI32V` | |
|
||||
| `VLUXSEG8EI64V` | |
|
||||
| `VLUXSEG8EI8V` | |
|
||||
| `VMACCVV` | |
|
||||
| `VMACCVX` | |
|
||||
| `VMADCVI` | |
|
||||
| `VMADCVIM` | |
|
||||
| `VMADCVV` | |
|
||||
| `VMADCVVM` | |
|
||||
| `VMADCVX` | |
|
||||
| `VMADCVXM` | |
|
||||
| `VMADDVV` | |
|
||||
| `VMADDVX` | |
|
||||
| `VMANDMM` | |
|
||||
| `VMANDNMM` | |
|
||||
| `VMAXUVV` | |
|
||||
| `VMAXUVX` | |
|
||||
| `VMAXVV` | |
|
||||
| `VMAXVX` | |
|
||||
| `VMCLRM` | |
|
||||
| `VMERGEVIM` | |
|
||||
| `VMERGEVVM` | |
|
||||
| `VMERGEVXM` | |
|
||||
| `VMFEQVF` | |
|
||||
| `VMFEQVV` | |
|
||||
| `VMFGEVF` | |
|
||||
| `VMFGEVV` | |
|
||||
| `VMFGTVF` | |
|
||||
| `VMFGTVV` | |
|
||||
| `VMFLEVF` | |
|
||||
| `VMFLEVV` | |
|
||||
| `VMFLTVF` | |
|
||||
| `VMFLTVV` | |
|
||||
| `VMFNEVF` | |
|
||||
| `VMFNEVV` | |
|
||||
| `VMINUVV` | |
|
||||
| `VMINUVX` | |
|
||||
| `VMINVV` | |
|
||||
| `VMINVX` | |
|
||||
| `VMMVM` | |
|
||||
| `VMNANDMM` | |
|
||||
| `VMNORMM` | |
|
||||
| `VMNOTM` | |
|
||||
| `VMORMM` | |
|
||||
| `VMORNMM` | |
|
||||
| `VMSBCVV` | |
|
||||
| `VMSBCVVM` | |
|
||||
| `VMSBCVX` | |
|
||||
| `VMSBCVXM` | |
|
||||
| `VMSBFM` | |
|
||||
| `VMSEQVI` | |
|
||||
| `VMSEQVV` | |
|
||||
| `VMSEQVX` | |
|
||||
| `VMSETM` | |
|
||||
| `VMSGEUVI` | |
|
||||
| `VMSGEUVV` | |
|
||||
| `VMSGEVI` | |
|
||||
| `VMSGEVV` | |
|
||||
| `VMSGTUVI` | |
|
||||
| `VMSGTUVV` | |
|
||||
| `VMSGTUVX` | |
|
||||
| `VMSGTVI` | |
|
||||
| `VMSGTVV` | |
|
||||
| `VMSGTVX` | |
|
||||
| `VMSIFM` | |
|
||||
| `VMSLEUVI` | |
|
||||
| `VMSLEUVV` | |
|
||||
| `VMSLEUVX` | |
|
||||
| `VMSLEVI` | |
|
||||
| `VMSLEVV` | |
|
||||
| `VMSLEVX` | |
|
||||
| `VMSLTUVI` | |
|
||||
| `VMSLTUVV` | |
|
||||
| `VMSLTUVX` | |
|
||||
| `VMSLTVI` | |
|
||||
| `VMSLTVV` | |
|
||||
| `VMSLTVX` | |
|
||||
| `VMSNEVI` | |
|
||||
| `VMSNEVV` | |
|
||||
| `VMSNEVX` | |
|
||||
| `VMSOFM` | |
|
||||
| `VMULHSUVV` | |
|
||||
| `VMULHSUVX` | |
|
||||
| `VMULHUVV` | |
|
||||
| `VMULHUVX` | |
|
||||
| `VMULHVV` | |
|
||||
| `VMULHVX` | |
|
||||
| `VMULVV` | |
|
||||
| `VMULVX` | |
|
||||
| `VMV1RV` | |
|
||||
| `VMV2RV` | |
|
||||
| `VMV4RV` | |
|
||||
| `VMV8RV` | |
|
||||
| `VMVSX` | |
|
||||
| `VMVVI` | |
|
||||
| `VMVVV` | |
|
||||
| `VMVVX` | |
|
||||
| `VMVXS` | |
|
||||
| `VMXNORMM` | |
|
||||
| `VMXORMM` | |
|
||||
| `VNCLIPUWI` | |
|
||||
| `VNCLIPUWV` | |
|
||||
| `VNCLIPUWX` | |
|
||||
| `VNCLIPWI` | |
|
||||
| `VNCLIPWV` | |
|
||||
| `VNCLIPWX` | |
|
||||
| `VNCVTXXW` | |
|
||||
| `VNEGV` | |
|
||||
| `VNMSACVV` | |
|
||||
| `VNMSACVX` | |
|
||||
| `VNMSUBVV` | |
|
||||
| `VNMSUBVX` | |
|
||||
| `VNOTV` | |
|
||||
| `VNSRAWI` | |
|
||||
| `VNSRAWV` | |
|
||||
| `VNSRAWX` | |
|
||||
| `VNSRLWI` | |
|
||||
| `VNSRLWV` | |
|
||||
| `VNSRLWX` | |
|
||||
| `VORVI` | |
|
||||
| `VORVV` | |
|
||||
| `VORVX` | |
|
||||
| `VREDANDVS` | |
|
||||
| `VREDMAXUVS` | |
|
||||
| `VREDMAXVS` | |
|
||||
| `VREDMINUVS` | |
|
||||
| `VREDMINVS` | |
|
||||
| `VREDORVS` | |
|
||||
| `VREDSUMVS` | |
|
||||
| `VREDXORVS` | |
|
||||
| `VREMUVV` | |
|
||||
| `VREMUVX` | |
|
||||
| `VREMVV` | |
|
||||
| `VREMVX` | |
|
||||
| `VRGATHEREI16VV` | |
|
||||
| `VRGATHERVI` | |
|
||||
| `VRGATHERVV` | |
|
||||
| `VRGATHERVX` | |
|
||||
| `VRSUBVI` | |
|
||||
| `VRSUBVX` | |
|
||||
| `VS1RV` | |
|
||||
| `VS2RV` | |
|
||||
| `VS4RV` | |
|
||||
| `VS8RV` | |
|
||||
| `VSADDUVI` | |
|
||||
| `VSADDUVV` | |
|
||||
| `VSADDUVX` | |
|
||||
| `VSADDVI` | |
|
||||
| `VSADDVV` | |
|
||||
| `VSADDVX` | |
|
||||
| `VSBCVVM` | |
|
||||
| `VSBCVXM` | |
|
||||
| `VSE16V` | |
|
||||
| `VSE32V` | |
|
||||
| `VSE64V` | |
|
||||
| `VSE8V` | |
|
||||
| `VSETIVLI` | |
|
||||
| `VSETVL` | |
|
||||
| `VSETVLI` | |
|
||||
| `VSEXTVF2` | |
|
||||
| `VSEXTVF4` | |
|
||||
| `VSEXTVF8` | |
|
||||
| `VSLIDE1DOWNVX` | |
|
||||
| `VSLIDE1UPVX` | |
|
||||
| `VSLIDEDOWNVI` | |
|
||||
| `VSLIDEDOWNVX` | |
|
||||
| `VSLIDEUPVI` | |
|
||||
| `VSLIDEUPVX` | |
|
||||
| `VSLLVI` | |
|
||||
| `VSLLVV` | |
|
||||
| `VSLLVX` | |
|
||||
| `VSMULVV` | |
|
||||
| `VSMULVX` | |
|
||||
| `VSMV` | |
|
||||
| `VSOXEI16V` | |
|
||||
| `VSOXEI32V` | |
|
||||
| `VSOXEI64V` | |
|
||||
| `VSOXEI8V` | |
|
||||
| `VSOXSEG2EI16V` | |
|
||||
| `VSOXSEG2EI32V` | |
|
||||
| `VSOXSEG2EI64V` | |
|
||||
| `VSOXSEG2EI8V` | |
|
||||
| `VSOXSEG3EI16V` | |
|
||||
| `VSOXSEG3EI32V` | |
|
||||
| `VSOXSEG3EI64V` | |
|
||||
| `VSOXSEG3EI8V` | |
|
||||
| `VSOXSEG4EI16V` | |
|
||||
| `VSOXSEG4EI32V` | |
|
||||
| `VSOXSEG4EI64V` | |
|
||||
| `VSOXSEG4EI8V` | |
|
||||
| `VSOXSEG5EI16V` | |
|
||||
| `VSOXSEG5EI32V` | |
|
||||
| `VSOXSEG5EI64V` | |
|
||||
| `VSOXSEG5EI8V` | |
|
||||
| `VSOXSEG6EI16V` | |
|
||||
| `VSOXSEG6EI32V` | |
|
||||
| `VSOXSEG6EI64V` | |
|
||||
| `VSOXSEG6EI8V` | |
|
||||
| `VSOXSEG7EI16V` | |
|
||||
| `VSOXSEG7EI32V` | |
|
||||
| `VSOXSEG7EI64V` | |
|
||||
| `VSOXSEG7EI8V` | |
|
||||
| `VSOXSEG8EI16V` | |
|
||||
| `VSOXSEG8EI32V` | |
|
||||
| `VSOXSEG8EI64V` | |
|
||||
| `VSOXSEG8EI8V` | |
|
||||
| `VSRAVI` | |
|
||||
| `VSRAVV` | |
|
||||
| `VSRAVX` | |
|
||||
| `VSRLVI` | |
|
||||
| `VSRLVV` | |
|
||||
| `VSRLVX` | |
|
||||
| `VSSE16V` | |
|
||||
| `VSSE32V` | |
|
||||
| `VSSE64V` | |
|
||||
| `VSSE8V` | |
|
||||
| `VSSEG2E16V` | |
|
||||
| `VSSEG2E32V` | |
|
||||
| `VSSEG2E64V` | |
|
||||
| `VSSEG2E8V` | |
|
||||
| `VSSEG3E16V` | |
|
||||
| `VSSEG3E32V` | |
|
||||
| `VSSEG3E64V` | |
|
||||
| `VSSEG3E8V` | |
|
||||
| `VSSEG4E16V` | |
|
||||
| `VSSEG4E32V` | |
|
||||
| `VSSEG4E64V` | |
|
||||
| `VSSEG4E8V` | |
|
||||
| `VSSEG5E16V` | |
|
||||
| `VSSEG5E32V` | |
|
||||
| `VSSEG5E64V` | |
|
||||
| `VSSEG5E8V` | |
|
||||
| `VSSEG6E16V` | |
|
||||
| `VSSEG6E32V` | |
|
||||
| `VSSEG6E64V` | |
|
||||
| `VSSEG6E8V` | |
|
||||
| `VSSEG7E16V` | |
|
||||
| `VSSEG7E32V` | |
|
||||
| `VSSEG7E64V` | |
|
||||
| `VSSEG7E8V` | |
|
||||
| `VSSEG8E16V` | |
|
||||
| `VSSEG8E32V` | |
|
||||
| `VSSEG8E64V` | |
|
||||
| `VSSEG8E8V` | |
|
||||
| `VSSRAVI` | |
|
||||
| `VSSRAVV` | |
|
||||
| `VSSRAVX` | |
|
||||
| `VSSRLVI` | |
|
||||
| `VSSRLVV` | |
|
||||
| `VSSRLVX` | |
|
||||
| `VSSSEG2E16V` | |
|
||||
| `VSSSEG2E32V` | |
|
||||
| `VSSSEG2E64V` | |
|
||||
| `VSSSEG2E8V` | |
|
||||
| `VSSSEG3E16V` | |
|
||||
| `VSSSEG3E32V` | |
|
||||
| `VSSSEG3E64V` | |
|
||||
| `VSSSEG3E8V` | |
|
||||
| `VSSSEG4E16V` | |
|
||||
| `VSSSEG4E32V` | |
|
||||
| `VSSSEG4E64V` | |
|
||||
| `VSSSEG4E8V` | |
|
||||
| `VSSSEG5E16V` | |
|
||||
| `VSSSEG5E32V` | |
|
||||
| `VSSSEG5E64V` | |
|
||||
| `VSSSEG5E8V` | |
|
||||
| `VSSSEG6E16V` | |
|
||||
| `VSSSEG6E32V` | |
|
||||
| `VSSSEG6E64V` | |
|
||||
| `VSSSEG6E8V` | |
|
||||
| `VSSSEG7E16V` | |
|
||||
| `VSSSEG7E32V` | |
|
||||
| `VSSSEG7E64V` | |
|
||||
| `VSSSEG7E8V` | |
|
||||
| `VSSSEG8E16V` | |
|
||||
| `VSSSEG8E32V` | |
|
||||
| `VSSSEG8E64V` | |
|
||||
| `VSSSEG8E8V` | |
|
||||
| `VSSUBUVV` | |
|
||||
| `VSSUBUVX` | |
|
||||
| `VSSUBVV` | |
|
||||
| `VSSUBVX` | |
|
||||
| `VSUBVV` | |
|
||||
| `VSUBVX` | |
|
||||
| `VSUXEI16V` | |
|
||||
| `VSUXEI32V` | |
|
||||
| `VSUXEI64V` | |
|
||||
| `VSUXEI8V` | |
|
||||
| `VSUXSEG2EI16V` | |
|
||||
| `VSUXSEG2EI32V` | |
|
||||
| `VSUXSEG2EI64V` | |
|
||||
| `VSUXSEG2EI8V` | |
|
||||
| `VSUXSEG3EI16V` | |
|
||||
| `VSUXSEG3EI32V` | |
|
||||
| `VSUXSEG3EI64V` | |
|
||||
| `VSUXSEG3EI8V` | |
|
||||
| `VSUXSEG4EI16V` | |
|
||||
| `VSUXSEG4EI32V` | |
|
||||
| `VSUXSEG4EI64V` | |
|
||||
| `VSUXSEG4EI8V` | |
|
||||
| `VSUXSEG5EI16V` | |
|
||||
| `VSUXSEG5EI32V` | |
|
||||
| `VSUXSEG5EI64V` | |
|
||||
| `VSUXSEG5EI8V` | |
|
||||
| `VSUXSEG6EI16V` | |
|
||||
| `VSUXSEG6EI32V` | |
|
||||
| `VSUXSEG6EI64V` | |
|
||||
| `VSUXSEG6EI8V` | |
|
||||
| `VSUXSEG7EI16V` | |
|
||||
| `VSUXSEG7EI32V` | |
|
||||
| `VSUXSEG7EI64V` | |
|
||||
| `VSUXSEG7EI8V` | |
|
||||
| `VSUXSEG8EI16V` | |
|
||||
| `VSUXSEG8EI32V` | |
|
||||
| `VSUXSEG8EI64V` | |
|
||||
| `VSUXSEG8EI8V` | |
|
||||
| `VWADDUVV` | |
|
||||
| `VWADDUVX` | |
|
||||
| `VWADDUWV` | |
|
||||
| `VWADDUWX` | |
|
||||
| `VWADDVV` | |
|
||||
| `VWADDVX` | |
|
||||
| `VWADDWV` | |
|
||||
| `VWADDWX` | |
|
||||
| `VWCVTUXXV` | |
|
||||
| `VWCVTXXV` | |
|
||||
| `VWMACCSUVV` | |
|
||||
| `VWMACCSUVX` | |
|
||||
| `VWMACCUSVX` | |
|
||||
| `VWMACCUVV` | |
|
||||
| `VWMACCUVX` | |
|
||||
| `VWMACCVV` | |
|
||||
| `VWMACCVX` | |
|
||||
| `VWMULSUVV` | |
|
||||
| `VWMULSUVX` | |
|
||||
| `VWMULUVV` | |
|
||||
| `VWMULUVX` | |
|
||||
| `VWMULVV` | |
|
||||
| `VWMULVX` | |
|
||||
| `VWREDSUMUVS` | |
|
||||
| `VWREDSUMVS` | |
|
||||
| `VWSUBUVV` | |
|
||||
| `VWSUBUVX` | |
|
||||
| `VWSUBUWV` | |
|
||||
| `VWSUBUWX` | |
|
||||
| `VWSUBVV` | |
|
||||
| `VWSUBVX` | |
|
||||
| `VWSUBWV` | |
|
||||
| `VWSUBWX` | |
|
||||
| `VXORVI` | |
|
||||
| `VXORVV` | |
|
||||
| `VXORVX` | |
|
||||
| `VZEXTVF2` | |
|
||||
| `VZEXTVF4` | |
|
||||
| `VZEXTVF8` | |
|
||||
| `WFI` | |
|
||||
| `WORD` | |
|
||||
| `XNOR` | |
|
||||
| `XOR` | Bitwise XOR |
|
||||
| `XORI` | XOR immediate |
|
||||
| `ZEXTH` | |
|
||||
|
||||
Recognised: 975 mnemonics.
|
||||
@@ -0,0 +1,140 @@
|
||||
# Language: lexicon, statements and expressions
|
||||
|
||||
Layer 1, the common language, the same on every target. Verified against
|
||||
`go tool asm` of Go 1.27.1 and against gasm's parser, which is differentially
|
||||
tested against the toolchain. The authoritative sources behind this page are
|
||||
the assembler's lexer (`cmd/asm/internal/lex`), its parser
|
||||
(`cmd/asm/internal/asm/parse.go`) and the toolchain's own test data.
|
||||
|
||||
## Source files and targets
|
||||
|
||||
An assembly source is a `.s` file. The Go build convention names a
|
||||
target-specific file with the architecture suffix, `_amd64.s`, `_arm64.s`,
|
||||
`_riscv64.s` or `_loong64.s`; files without a suffix are portable across
|
||||
targets. The same assembler program assembles every target: `go tool asm`
|
||||
picks the target from the `GOOS` and `GOARCH` environment variables, and gasm
|
||||
from the file name suffix or the `--arch` flag.
|
||||
|
||||
## Character set and identifiers
|
||||
|
||||
Sources are ASCII text. An identifier is a sequence of ASCII letters, digits
|
||||
and underscores, digits never first, with exactly two additions:
|
||||
|
||||
- U+00B7, the middle dot `·`, stands for the period in a symbol's
|
||||
package-qualified name;
|
||||
- U+2215, the division slash `∕`, stands for the slash in a package path.
|
||||
|
||||
The two substitutions exist because the parser treats a real period and a
|
||||
real slash as punctuation. The syntax is otherwise uppercase throughout:
|
||||
instructions, registers and directives are written in upper case. The one
|
||||
inherited exception is the `g` register name on 32-bit ARM.
|
||||
|
||||
## Comments
|
||||
|
||||
Two comment forms, both Go's:
|
||||
|
||||
```text
|
||||
// a line comment
|
||||
/* a block comment */
|
||||
```
|
||||
|
||||
A comment of the form `//go:build` or the legacy `+build` comment is not a
|
||||
plain comment: the lexer reports it to the build system as a build
|
||||
constraint.
|
||||
|
||||
## Statements
|
||||
|
||||
The grammar of one line, from the parser:
|
||||
|
||||
```text
|
||||
{label:} WORD[.qualifier] [ arg {, arg} ] (';' | '\n')
|
||||
```
|
||||
|
||||
- A **label** is an identifier followed by a colon. Labels are
|
||||
function-local: two functions in one file may reuse the same name, and a
|
||||
reference resolves within the function that contains it. A branch
|
||||
instruction names its target with a bare label operand, and the assembler
|
||||
resolves it PC-relative. The explicit forms `offset(PC)`, a constant
|
||||
counting instructions from the branch, and `name(SB)`, a cross-function
|
||||
static reference, appear as branch targets as well.
|
||||
- **WORD** is the instruction or directive name, upper case. On the ARM
|
||||
family the word may carry a dot qualifier selecting a condition or shift
|
||||
mode, such as the condition suffixes on 32-bit ARM; the amd64, arm64,
|
||||
riscv64 and loong64 assemblies carry no instruction qualifiers apart from
|
||||
their own width suffixes, which are part of the mnemonic.
|
||||
- **Arguments** are separated by commas, with no trailing comma.
|
||||
- A statement ends at a newline or at a semicolon, so several statements fit
|
||||
on one line separated by `;`. Blank lines are free.
|
||||
|
||||
The first word of a line is a directive if it is one of the directive names
|
||||
(TEXT, DATA, GLOBL, FUNCDATA, PCDATA, PCALIGN) and an instruction otherwise.
|
||||
Unknown instruction names are errors; the instruction set is the set the
|
||||
toolchain itself defines per target, plus the common pseudo-instructions.
|
||||
|
||||
## Literals
|
||||
|
||||
| Form | Examples | Notes |
|
||||
|---|---|---|
|
||||
| Integer | `0`, `42`, `0x2a`, `0o52`, `0b101010`, `1_000` | decimal, hexadecimal, octal and binary forms with Go's digit separators |
|
||||
| Character | `'a'`, `'\n'`, `'\x41'` | single quoted, Go escape rules |
|
||||
| String | `"this program can only run\n"` | double quoted, Go escape rules; accepted where an operand takes raw bytes, in practice a DATA initialiser |
|
||||
| Float | `1.5`, `1e9` | accepted by the lexer; only meaningful where the target's encoding takes a float operand |
|
||||
|
||||
## Expressions
|
||||
|
||||
Constant expressions may appear wherever a constant is expected: in
|
||||
immediates after `$`, in memory offsets, in frame and data sizes. The
|
||||
evaluator works on unsigned 64-bit values with Go's operator precedence, and
|
||||
the parser states its grammar in exactly those terms:
|
||||
|
||||
```text
|
||||
expr = term { '+' term | '-' term | '|' term | '^' term }
|
||||
term = factor { '*' factor | '/' factor | '%' factor | '<<' factor | '>>' factor | '&' factor }
|
||||
factor = const | '+' factor | '-' factor | '~' factor | '(' expr ')'
|
||||
```
|
||||
|
||||
Two consequences are worth naming, because the arithmetic surprises people
|
||||
who read it as C:
|
||||
|
||||
- Shifts bind at the multiplicative level, next to `*` and `&`, while `|`
|
||||
and `^` bind at the additive level. `$x<<1|3` computes `(x<<1)|3`, which
|
||||
differs from `x*2+3` whenever `x` is odd. Plan 9 arithmetic is Go
|
||||
precedence applied to a byte-oriented language, not the C expression it
|
||||
resembles.
|
||||
- The evaluator is unsigned and guarded: division or modulo by zero is an
|
||||
error, and so is dividing a value with the high bit set; shift counts must
|
||||
be non-negative; and a right shift of a value with the high bit set is
|
||||
rejected rather than sign-extended.
|
||||
|
||||
An address expression such as `(index*4)(base)` is evaluated at assembly
|
||||
time only if every name in it is a constant; a name that resolves to a
|
||||
symbol turns the expression into a relocation request, never into a folded
|
||||
constant.
|
||||
|
||||
Named constants enter expressions through the preprocessor (`#define`,
|
||||
`-D`) and, in Go-embedded packages, through the generated `go_asm.h`; see
|
||||
PREPROCESSOR.md and RUNTIME.md.
|
||||
|
||||
## The common pseudo-instructions
|
||||
|
||||
A handful of instructions exist on every target, assembled by the assembler
|
||||
itself rather than the encoder: `NOP`, which emits the target's no-operation
|
||||
encoding, and the frame-management pseudo-instructions the compiler emits
|
||||
(`FUNCDATA`, `PCDATA`) which DIRECTIVES.md specifies. Everything else is the
|
||||
target's own instruction set, and the assembler knows only the instructions
|
||||
the toolchain's compiler emits; a hand-written kernel wanting more lays the
|
||||
encoding down with `BYTE` on amd64 or waits for the extended layer.
|
||||
|
||||
## Case study: three lines, decomposed
|
||||
|
||||
```text
|
||||
B.EQ 1(PC) // arm64: condition qualifier on the mnemonic,
|
||||
// target one instruction past the branch
|
||||
JMP done // every target: bare label, function-local,
|
||||
// resolved PC-relative
|
||||
MOVQ $reader__size>>3, CX // amd64: expression over a go_asm.h constant
|
||||
```
|
||||
|
||||
The first shows a qualifier and the explicit relative target form; the second
|
||||
the ordinary label reference; the third an expression over a generated
|
||||
constant. Labels are reusable between functions without conflict.
|
||||
@@ -0,0 +1,94 @@
|
||||
# LoongArch 64
|
||||
|
||||
Layer 1, target page. Verified against `go tool asm` of Go 1.27.1, against
|
||||
the toolchain's own loong64 assembler manual (`cmd/internal/obj/loong64/doc.go`)
|
||||
and against gasm's encoder, whose output is compared byte for byte with the
|
||||
toolchain's. The complete mnemonic inventory lives in the generated appendix
|
||||
[INSTRUCTIONS-LOONG64.md](INSTRUCTIONS-LOONG64.md).
|
||||
|
||||
## Registers
|
||||
|
||||
- General purpose `R0` to `R31`, floating point `F0` to `F31`, LSX vectors
|
||||
`V0` to `V31` and LASX vectors `X0` to `X31`.
|
||||
- Fixed roles from the toolchain's table: `R0` is the constant zero, `R1`
|
||||
the return address, `R3` the stack pointer, `R22` the goroutine pointer,
|
||||
`R29` the closure context and `R30` the assembler's temporary. `R12`,
|
||||
`R13`, `R14`, `R15` and `R20` serve the PLT and trampoline sequences:
|
||||
usable in assembly, but saved before any call.
|
||||
|
||||
## Widths ride the mnemonic
|
||||
|
||||
| Suffix | Width |
|
||||
|---|---|
|
||||
| `B`, `BU` | 8-bit, 8-bit unsigned |
|
||||
| `H`, `HU` | 16-bit, 16-bit unsigned |
|
||||
| `W`, `WU` | 32-bit, 32-bit unsigned |
|
||||
| `V` | 64-bit |
|
||||
| `F`, `D` | 32-bit and 64-bit float |
|
||||
| `V` prefix (LSX) | 128-bit vector |
|
||||
| `XV` prefix (LASX) | 256-bit vector |
|
||||
|
||||
The MOV series is the load and store interface: `MOVB (R2), R3` loads a
|
||||
byte, `MOVV (R2), R3` a double word, `VMOVQ (R2), V1` a 128-bit vector and
|
||||
`XVMOVQ (R2), X1` a 256-bit one.
|
||||
|
||||
## Operand order
|
||||
|
||||
Most instructions appear in left-to-right assignment order: `ADDV R11, R12,
|
||||
R13` is `add.d R13, R12, R11`, and the two-operand form
|
||||
`OR R5, R6` assigns into R6. Exceptions:
|
||||
|
||||
- Jump and branch instructions keep the GNU order: `BEQ R0, R4, label1`.
|
||||
- The bitfield family is `BSTRINSW`, `BSTRINSV`, `BSTRPICKW`, `BSTRPICKV`
|
||||
`$<msb>, <Rj>, $<lsb>, <Rd>`.
|
||||
|
||||
## Addressing
|
||||
|
||||
- Plain: `offset(Rbase)`.
|
||||
- Base plus offset **register**, no scale: `(R4)(R5)`, as in
|
||||
`MOVB (R4)(R5), R6`, the `ldx` family.
|
||||
- The pointer loads and stores `MOVWP` and `MOVVP` take a source-level
|
||||
16-bit offset that the encoder halves into the 14-bit field, writing
|
||||
`MOVWP 8(R4), R5` as `ldptr.w r5, r4, $2`.
|
||||
|
||||
## Vector element syntax
|
||||
|
||||
The `VMOVQ` and `XVMOVQ` transfer family covers register-to-vector moves
|
||||
with arrangement and index suffixes: `VMOVQ Rj, Vd.B[index]` inserts a
|
||||
general register into one lane, `VMOVQ Vj.B[index], Rd` extracts one,
|
||||
`VMOVQ Rj, Vd.B16` broadcasts across all sixteen, and `VMOVQ Vj.B[index],
|
||||
Vd.B16` replicates one lane. The broadcast-from-memory form takes the true
|
||||
byte offset at source level, which the encoder rescales per arrangement.
|
||||
The permute and extract families take their 8-bit control word first:
|
||||
`VPERMIW ui8, Vj, Vd`, `VEXTRINSB ui8, Vj, Vd`.
|
||||
|
||||
## Alignment
|
||||
|
||||
`PCALIGN $n` pads with NOOP to a power-of-two boundary between 8 and 2048,
|
||||
and this target additionally auto-aligns loop heads to 16 bytes.
|
||||
|
||||
## Atomics, barriers and prefetch
|
||||
|
||||
- The `AM` atomic family comes in plain and `_DB` flavours; the `_DB`
|
||||
forms, such as `AMSWAPDBW`, complete the atomic sequence and act as a
|
||||
full data barrier. Within the AM family the destination and base
|
||||
registers may not coincide and the destination may not equal the operand
|
||||
register: one is an exception, the other silently unspecified.
|
||||
- `DBAR` carries the graded hint encoding documented for LA664 and later,
|
||||
with hint 0x700 as the read-after-read lightweight barrier; older cores
|
||||
treat every hint as the full barrier.
|
||||
- `PRELD offset(Rbase), $hint` prefetches with the documented hints (0
|
||||
load to L1, 2 load to L3, 8 store to L1); `PRELDX` adds the encoded
|
||||
block descriptor.
|
||||
- `ALSL`-family shift-and-add writes the desired shift amount in source and
|
||||
encodes one less: `ALSLV $4, R4, R5, R6` shifts by 4.
|
||||
- `ADDV16 si16<<16, Rj, Rd` is the high-immediate add paired with the
|
||||
pointer loads for GOT relative access.
|
||||
|
||||
## Relocations
|
||||
|
||||
`R_CALLLOONG64` for the 28-bit BL, `R_LOONG64_CALL36` for the
|
||||
PCADDU18I-plus-JIRL pair, the `R_LOONG64_ADDR`, `ADDR64`, `TLS_LE`, `TLS_IE`,
|
||||
`GOT` and `GOT64` high and low pairs, the aligned conditional jump forms
|
||||
`R_JMP16LOONG64` and `R_JMP21LOONG64`, and `R_LOONG64_ADD64` and `SUB64`
|
||||
for in-place arithmetic, all specified in [GOOBJ.md](../GOOBJ.md).
|
||||
@@ -0,0 +1,114 @@
|
||||
# Operands: grammar, pseudo-registers, addressing and symbols
|
||||
|
||||
Layer 1, the common language. Verified against `go tool asm` of Go 1.27.1 and
|
||||
against gasm's parser. The operand grammar is the part of the language that
|
||||
varies most between targets, so this page fixes the common grammar and the
|
||||
pseudo-registers; the per architecture pages carry the register names and the
|
||||
addressing quirks each target adds.
|
||||
|
||||
## The four operand kinds
|
||||
|
||||
Every operand is one of four kinds:
|
||||
|
||||
```text
|
||||
R1 register
|
||||
$4 immediate
|
||||
label branch target or symbol
|
||||
-8(BX)(DI*4) memory
|
||||
```
|
||||
|
||||
**Operands go source first, destination last**: `MOVQ x+0(FP), AX` loads the
|
||||
argument into AX. This is the opposite of Intel order and the same order as
|
||||
AT&T, with the sigils removed: registers are bare names, immediates take
|
||||
`$`, memory is `offset(base)`.
|
||||
|
||||
## Registers
|
||||
|
||||
A register operand is its bare name, with no prefix: `AX`, `X15`, `R14` on
|
||||
amd64; `R0` to `R30`, `ZR`, `V0` to `V31` on arm64; `X0` to `X31`, `F0` to
|
||||
`F31`, `V0` on riscv64; `R0` to `R31`, `F0` to `F31`, `V0` on loong64.
|
||||
Sub-register and width selection rides the mnemonic, not the operand: the
|
||||
amd64 family spells `MOVB`, `MOVW`, `MOVL`, `MOVQ`, and the arm64 family
|
||||
suffices `B`, `H`, `S`, `D`, `Q` on the shared forms. Each architecture page
|
||||
lists its registers and the reserved ones.
|
||||
|
||||
## Immediates
|
||||
|
||||
`$` introduces a constant: `$42`, `$-1`, `$0x2a`, `$'A'`, `$bufSize`. The
|
||||
`$` applies to the whole constant expression that follows, so
|
||||
`$(4*8+reader__size)` is one immediate. Without the `$`, a number in operand
|
||||
position is an address, not a value; the classic error `ADDQ 1, AX` asks the
|
||||
assembler for the byte at address 1.
|
||||
|
||||
The one place a `$` number is not an immediate is the frame and argument
|
||||
size field of TEXT, `$16-24`, which is two separate constants and not a
|
||||
subtraction; DIRECTIVES.md specifies it.
|
||||
|
||||
## Memory
|
||||
|
||||
```text
|
||||
offset(base)
|
||||
offset(base)(index*scale)
|
||||
```
|
||||
|
||||
Both parts are optional where the target allows them: `(BX)` is the memory
|
||||
at BX, `foo+16(SB)` is a global, and on amd64 `foo+32(SP)(R9*8)` adds a
|
||||
scaled index. `offset` is a constant expression, optionally carrying a
|
||||
symbol name. The extensions beyond `offset(base)` are where the targets
|
||||
diverge, and each belongs to its architecture page: amd64 carries the
|
||||
`index*scale` form with scale 1, 2, 4 or 8 and its own rules on which
|
||||
registers may index; loong64 writes base plus index as `(R4)(R5)`; the ARM
|
||||
family attaches shift amounts to the index register in its own spelling.
|
||||
|
||||
The address arithmetic is on **byte addresses**: the offset is added to the
|
||||
base as it stands, whatever the operand width of the instruction. Loading
|
||||
the third 8-byte word of an array at BX is `16(BX)`, not `2(BX)`.
|
||||
|
||||
## The four pseudo-registers
|
||||
|
||||
Four names denote locations no target register holds, and they mean the same
|
||||
on every architecture:
|
||||
|
||||
- **FP**, the frame pointer: the arguments and results of the current
|
||||
function, at positive offsets, in the order the Go prototype declares
|
||||
them. Every FP reference must carry a name: `x+0(FP)`, and an unnamed
|
||||
`0(FP)` is rejected. Results follow arguments; an unnamed result is called
|
||||
`ret`.
|
||||
- **SP**, the virtual stack pointer: the high end of the function's local
|
||||
frame, so locals live at negative offsets, `x-8(SP)`. A reference without
|
||||
a name and without a plus, `-8(SP)`, addresses the **hardware** stack
|
||||
pointer instead: the two spellings are one character apart and mean
|
||||
different registers. That is the sharpest edge in the language and the
|
||||
source of the deepest bugs.
|
||||
- **SB**, the static base: the origin of memory, used for globals and
|
||||
cross-package symbols, always with a name: `foo(SB)`, `foo+4(SB)`.
|
||||
- **PC**, the program counter: branch targets, and the explicit relative
|
||||
form `1(PC)`.
|
||||
|
||||
## Symbol names
|
||||
|
||||
A symbol's full name is the package path, a period, and the base name. In
|
||||
source, the period is written U+00B7 (`·`) and a slash in the path U+2215
|
||||
(`∕`), because the parser treats the ASCII forms as punctuation. Inside the
|
||||
package's own file, `·Name` is enough and is the preferred spelling, since
|
||||
it survives a rename of the import path.
|
||||
|
||||
| Spelling | Meaning |
|
||||
|---|---|
|
||||
| `·Name(SB)` | this package's Name |
|
||||
| `runtime·morestack(SB)` | another package's morestack |
|
||||
| `sourcedock.dev∕petrbalvin∕pkg·Name(SB)` | fully qualified |
|
||||
| `msg<>(SB)` | file-local, the static of this language; `<>` also makes the ABI field static in the object |
|
||||
| `Name<ABIInternal>(SB)` | ABI-qualified reference, the ABI in angle brackets after the name |
|
||||
|
||||
The object file these symbols produce, with the index rules that decide what
|
||||
is referenced by name and what by index, is specified in
|
||||
[GOOBJ.md](../GOOBJ.md).
|
||||
|
||||
## What vet adds in Go
|
||||
|
||||
Inside a Go package, `go vet`'s asmdecl analyzer checks every FP offset and
|
||||
name against the Go prototype, and checks the declared argument area against
|
||||
the frame. That layer, the prototype requirement and `go_asm.h`, belongs to
|
||||
RUNTIME.md; the grammar above is the whole of what the assembler itself
|
||||
requires.
|
||||
@@ -0,0 +1,79 @@
|
||||
# Preprocessing: include, define and selection
|
||||
|
||||
Layer 1, the common language. Verified against the preprocessor inside
|
||||
`go tool asm` of Go 1.27.1 (`cmd/asm/internal/lex`), whose directives are
|
||||
`#define`, `#undef`, `#include`, `#ifdef`, `#ifndef`, `#else`, `#endif` and
|
||||
`#line`, and against gasm's implementation, which is differentially tested
|
||||
against the toolchain's.
|
||||
|
||||
Input runs through a simplified C preprocessor before the parser sees it.
|
||||
The set is deliberately small: there is no `#if` with constant expressions
|
||||
and no token pasting with `##`. `#line` is honoured, so it changes the
|
||||
positions the assembler reports and records.
|
||||
|
||||
## #include
|
||||
|
||||
```text
|
||||
#include "textflag.h"
|
||||
#include "go_asm.h"
|
||||
#include "defs_linux_amd64.h"
|
||||
```
|
||||
|
||||
The search path, in order: the directory of the including file, then the
|
||||
directories given by repeatable `-I` flags. The assembler seeds no default
|
||||
of its own: a bare `go tool asm` invocation finds none of the standard
|
||||
headers, and it is the `go` build system that passes `$GOROOT/pkg/include`
|
||||
among the `-I` directories when it drives the build. That directory ships
|
||||
`textflag.h`, `funcdata.h` and the per architecture register headers.
|
||||
Includes nest; a file included twice through different paths is processed
|
||||
twice, which is why headers guard their defines.
|
||||
|
||||
## #define and #undef
|
||||
|
||||
```text
|
||||
#define bufSize 1024
|
||||
#define MOVD(d, s) MOVQ s, d
|
||||
#undef bufSize
|
||||
```
|
||||
|
||||
- An object macro replaces its name with its token sequence at the point of
|
||||
use.
|
||||
- A parameterised macro takes its arguments in parentheses and substitutes
|
||||
them into the body. Macro parameters compose with the rest of the
|
||||
language: an argument used with an element suffix, as in `A.S4` on the
|
||||
vector forms, substitutes correctly.
|
||||
- Redefinition is an error; `#undef` first, or pick a new name.
|
||||
- The `-D name[=value]` flag predefines an object macro from the command
|
||||
line, repeatable, exactly as `#define` would; a `-D` without a value
|
||||
defines the name as `1`.
|
||||
- Expansion happens when the name is used, so a macro may expand to
|
||||
instructions, operands or fragments of either, and a macro body may use
|
||||
macros defined before it.
|
||||
|
||||
`textflag.h` and `funcdata.h` are themselves ordinary `#define` files: the
|
||||
flag names and the runtime macros are preprocessor definitions, not language
|
||||
keywords. That is why a missing include produces a parser error at the first
|
||||
use of `NOSPLIT` rather than a complaint about the name.
|
||||
|
||||
## #ifdef, #ifndef, #else, #endif
|
||||
|
||||
```text
|
||||
#ifdef GOOS_windows
|
||||
#define SYSCALL_INT 0x2b
|
||||
#endif
|
||||
```
|
||||
|
||||
Selection is by defined-name only: `#ifdef`, `#ifndef`, `#else`, `#endif`,
|
||||
nesting freely. There is no `#if defined(x) && y`, because the preprocessor
|
||||
evaluates no expressions; reach that with a build-tag Go file generating a
|
||||
header, which is exactly how the runtime's own `go_asm.h` and defs headers
|
||||
are produced.
|
||||
|
||||
## What preprocessing does not cover
|
||||
|
||||
The preprocessor is textual and runs first, so it knows nothing of assembly
|
||||
semantics: it does not check that a macro expansion is a legal instruction,
|
||||
and it does not participate in the constant expression evaluator, which runs
|
||||
later, in the parser. A constant folded with `#define` and a constant folded
|
||||
in an operand expression end at the same value through different doors;
|
||||
GOOBJ.md records both in the object identically.
|
||||
@@ -0,0 +1,56 @@
|
||||
# The Plan 9 assembly language
|
||||
|
||||
This directory is the reference for the Plan 9 assembly language as the Go
|
||||
toolchain and gasm accept it, written to be complete enough to implement
|
||||
against. It exists because no such reference exists upstream: Go documents
|
||||
the language on a single page, and the rest of the knowledge lives in the
|
||||
toolchain's source and in the practice of reading it.
|
||||
|
||||
Every page carries the same conformance statement: which layer of the system
|
||||
it describes, which toolchain release it was verified against, and how the
|
||||
claims were checked. Pages in this directory are verified against Go 1.27.1
|
||||
and against gasm's own differential test suite, which compares gasm's
|
||||
behaviour with `go tool asm` byte for byte and output for output.
|
||||
|
||||
## The three layers
|
||||
|
||||
The reference deliberately separates three layers, because their rules have
|
||||
different owners and different lifetimes:
|
||||
|
||||
1. **The common language** (LANGUAGE, OPERANDS, DIRECTIVES,
|
||||
PREPROCESSOR): the syntax, operands, directives and preprocessing, the
|
||||
same on every target and meaningful without a Go runtime.
|
||||
2. **The Go-embedded layer** (RUNTIME): everything that exists only because
|
||||
the code runs inside a Go program: the ABI0 contract, generated wrappers,
|
||||
`go_asm.h`, the garbage collector annotations and `go vet` checks.
|
||||
3. **The standalone layer** (STANDALONE, planned with the standalone
|
||||
compilation phase): using the language outside Go, through gasm's ELF
|
||||
output and the extended instruction set, where the toolchain offers no
|
||||
ground truth and execution testing is the only verification.
|
||||
|
||||
A rule stated in layer 1 holds on every target. A rule stated in layer 2
|
||||
says which part of the Go machinery imposes it. Nothing in layer 3 changes
|
||||
layers 1 or 2; it extends them.
|
||||
|
||||
## Pages
|
||||
|
||||
| Page | Layer | Contents |
|
||||
|---|---|---|
|
||||
| [LANGUAGE.md](LANGUAGE.md) | 1 | lexicon, statement structure, labels, literals, expressions |
|
||||
| [OPERANDS.md](OPERANDS.md) | 1 | operand grammar, pseudo-registers, addressing modes, symbol naming |
|
||||
| [DIRECTIVES.md](DIRECTIVES.md) | 1 | TEXT, DATA, GLOBL, FUNCDATA, PCDATA, PCALIGN and the function flags |
|
||||
| [PREPROCESSOR.md](PREPROCESSOR.md) | 1 | `#include`, `#define`, `#ifdef` and friends, `-D`, `-I` |
|
||||
| [RUNTIME.md](RUNTIME.md) | 2 | ABI0, prototypes, `go_asm.h`, `funcdata.h`, `go vet` |
|
||||
| [AMD64.md](AMD64.md) | 1 | registers, addressing, the frame and split check, families, relocations |
|
||||
| [ARM64.md](ARM64.md) | 1 | registers, the MOV load and store series, special operand orders, SIMD |
|
||||
| [RISCV64.md](RISCV64.md) | 1 | registers and their constrained names, per class operand order, profiles, vector extension |
|
||||
| [LOONG64.md](LOONG64.md) | 1 | registers, width suffixes, vector element syntax, atomics and barriers |
|
||||
| INSTRUCTIONS-AMD64.md and the other three | 1 | generated per architecture inventory of every accepted mnemonic |
|
||||
| STANDALONE.md | 3 | the language outside Go |
|
||||
|
||||
## Status
|
||||
|
||||
The common-language core, the Go-embedded layer, all four per-architecture
|
||||
pages and the generated instruction appendices are written and verified.
|
||||
STANDALONE.md lands with the standalone compilation phase. The object format
|
||||
these pages feed is specified in [GOOBJ.md](../GOOBJ.md).
|
||||
@@ -0,0 +1,104 @@
|
||||
# RISC-V 64
|
||||
|
||||
Layer 1, target page. Verified against `go tool asm` of Go 1.27.1, against
|
||||
the toolchain's own riscv64 assembler manual (`cmd/internal/obj/riscv/doc.go`)
|
||||
and against gasm's encoder, whose output is compared byte for byte with the
|
||||
toolchain's. The complete mnemonic inventory lives in the generated appendix
|
||||
[INSTRUCTIONS-RISCV64.md](INSTRUCTIONS-RISCV64.md).
|
||||
|
||||
## Registers
|
||||
|
||||
- Integer: `X0` to `X31`. `X0` is hardwired zero. Three names the toolchain
|
||||
constrains: `X4` must be written through its ABI name `TP`; `X27`, the
|
||||
goroutine pointer, must be written `g` and may not be written `S11`; in
|
||||
shared builds `X3` is off limits and must be written `GP`.
|
||||
- The other integer registers may be written `Xn` or by their ABI names
|
||||
(`A0`, `T0`, `S1`, and so on).
|
||||
- Floating point: `F0` to `F31`. Vector: `V0` to `V31`.
|
||||
- `X26` is the closure pointer and `X31` is the assembler's own scratch
|
||||
register: its value may be clobbered by instruction sequences the
|
||||
assembler inserts, so hand-written code must not rely on it.
|
||||
- There is no reserved frame pointer register on this target.
|
||||
|
||||
## Operand order
|
||||
|
||||
The ordering differs from the ISA manual, and per instruction class:
|
||||
|
||||
- **R-type** is reversed: `ADD X10, X11, X12` is `add x12, x11, x10`.
|
||||
- **I-type arithmetic** keeps that shape with the immediate first:
|
||||
`ADDI $1, X11, X12`.
|
||||
- **Loads and stores** are source first, like every Plan 9 dialect:
|
||||
`MOV 16(X2), X10` loads and `MOV X10, (X2)` stores. The MOV series hides
|
||||
the width; `MOVB` through `MOVD` spell it out.
|
||||
- **Branches** keep the ISA order: `BLT X12, X23, loop1`, which jumps when
|
||||
X12 < X23, the reverse of the SLT operand order.
|
||||
- **FMA** is rotated one place left so the destination comes last:
|
||||
`FMADDS F1, F2, F3, F4`.
|
||||
- **AMO** is likewise rotated: `AMOSWAPW X5, (X6), X7`.
|
||||
- **Ternary abbreviation** is supported and encouraged: `ADD X10, X12` means
|
||||
`ADD X10, X12, X12`.
|
||||
|
||||
Where an R-type instruction has an I-type sibling, the assembler picks the
|
||||
immediate form from the operand: `AND $3, X12, X13` assembles as `ANDI`.
|
||||
|
||||
## Names, suffixes and rounding
|
||||
|
||||
Dots are removed and suffixes are upper-cased: the ISA's `fmv.w.x` is
|
||||
`FMVWX`. Floating-point rounding modes become suffixes, `FCVTLUS.RNE F0,
|
||||
X5`, with RTZ assumed when the suffix is omitted; the toolchain never sets
|
||||
the FCSR.
|
||||
|
||||
## Constants
|
||||
|
||||
- `MOV` materialises any 64-bit integer constant, synthesising it from a
|
||||
few arithmetic instructions where possible and otherwise loading it from
|
||||
a literal pool in the binary.
|
||||
- A 32-bit constant is accepted by `ADDI`, `ANDI`, `ORI` and `XORI`, and
|
||||
the assembler synthesises values that exceed the 12-bit encoding window.
|
||||
- `MOVF` and `MOVD` materialise floating-point constants, encoding them as
|
||||
`FLW` and `FLD` from a pool location unless the constant is exactly 0.0.
|
||||
|
||||
## Extensions and profiles
|
||||
|
||||
The default target profile is rva20u64, selected or raised with the
|
||||
GORISCV64 environment variable. A short list of instructions outside the
|
||||
default profile is synthesised by the assembler when the profile does not
|
||||
provide them, so they are safe without guards: `ANDN`, `MAX`, `MAXU`, `MIN`,
|
||||
`MINU`, `MOVB`, `MOVH`, `MOVHU`, `MOVWU`, `ORN`, `ROL`, `ROLW`, `ROR`,
|
||||
`RORI`, `RORIW`, `RORW`, `XNOR`. The header `asm_riscv64.h` defines the
|
||||
`hasZba`, `hasZbb`, `hasZbs` and `hasV` macros for guarding everything else.
|
||||
|
||||
## Fences and atomics
|
||||
|
||||
`FENCE` takes predecessor and successor sets in that order, uppercase
|
||||
letters, `FENCE R, RW`; a bare `FENCE` is a full fence, as is
|
||||
`FENCE IORW, IORW`. `FENCE.TSO` exists. The ordering bits of `LR`, `SC`
|
||||
and the AMO instructions are not specifiable in source: the assembler sets
|
||||
acquire and release on the AMO instructions, acquire on `LR` and release on
|
||||
`SC`, always.
|
||||
|
||||
## Compressed instructions
|
||||
|
||||
The assembler converts 32-bit instructions to their compressed encodings
|
||||
automatically; the conversion is a property of the emitted machine code, not
|
||||
of the source, and register choice influences how much compresses.
|
||||
Hand-writing compressed instructions in source is accepted but discouraged.
|
||||
The debug flag `compressinstructions=0` turns the automatic conversion off.
|
||||
|
||||
## Vector extension
|
||||
|
||||
`VSETVLI` writes its vtype components in uppercase with the destination
|
||||
last: `VSETVLI X10, E8, M1, TU, MU, X12`. Vector loads and stores are
|
||||
source first like the scalar ones, with an optional stride or index register
|
||||
second and the mask register, when present, always penultimate:
|
||||
`VLE8V (X10), V3`, `VLE8V (X10), V0, V3` for the masked form. Vector
|
||||
arithmetic reverses its operands, `VADDVV V1, V2, V3`, with the mask again
|
||||
penultimate.
|
||||
|
||||
## Relocations
|
||||
|
||||
`R_RISCV_JAL`, `R_RISCV_CALL`, the `R_RISCV_PCREL_ITYPE` and `STYPE` pairs,
|
||||
`R_RISCV_BRANCH`, the compressed branch and jump forms, the TLS and GOT
|
||||
families and `R_RISCV_ADD32` and `SUB32`, all specified in
|
||||
[GOOBJ.md](../GOOBJ.md). The assembler always emits the four-byte
|
||||
`R_DWTXTADDR_U4` flavour inside its DWARF records.
|
||||
@@ -0,0 +1,120 @@
|
||||
# The Go-embedded layer: ABI0, prototypes and the runtime contract
|
||||
|
||||
Layer 2: everything that exists only because the assembly runs inside a Go
|
||||
program. Without a Go runtime this page does not apply; the language of
|
||||
OPERANDS.md and DIRECTIVES.md still does. Verified against Go 1.27.1, against
|
||||
the shipped `funcdata.h` header, and against the object files the toolchain
|
||||
produces, which were parsed and checked field by field while writing
|
||||
[GOOBJ.md](../GOOBJ.md).
|
||||
|
||||
## Hand-written assembly is ABI0
|
||||
|
||||
Go functions compiled from source use ABIInternal, the register-based
|
||||
calling convention, which the toolchain documents as unstable and free to
|
||||
change between releases. A `.s` function is written against ABI0, the stack
|
||||
based convention: arguments and results live in the caller's frame at
|
||||
positive FP offsets, byte-addressed, in declaration order, with no registers
|
||||
assigned at all. The toolchain generates the wrapper that translates between
|
||||
the two; a caller in Go calling an assembly function goes through it, and it
|
||||
is marked `ABIWRAPPER` in the object. Hand-writing a bridge is never needed
|
||||
and never correct.
|
||||
|
||||
## Every assembly function carries a Go prototype
|
||||
|
||||
```go
|
||||
package add
|
||||
|
||||
func Add(x, y int64) int64
|
||||
```
|
||||
|
||||
The body-less declaration is not optional, and not only for the linker: it
|
||||
is what tells the garbage collector which arguments and results hold
|
||||
pointers, and what `go vet` checks the assembly against. Even a function
|
||||
nothing in Go calls gets one. Consequences:
|
||||
|
||||
- The FP operand names and offsets are checked by vet's asmdecl analyzer
|
||||
against the prototype: `x+0(FP)` must name an argument that exists, at the
|
||||
offset the prototype says. A file that assembles and links can still fail
|
||||
vet.
|
||||
- The declared argument area in `$framesize-argsize` is checked against the
|
||||
prototype's size. An omitted argsize marks the argument size unknown
|
||||
(0x80000000 in the object, the value of `ArgsSizeUnknown` from
|
||||
`funcdata.h`), which is the normal spelling for functions with no Go
|
||||
callers.
|
||||
- `//go:noescape` on the declaration tells the compiler that a pointer
|
||||
argument does not escape, for assembly that keeps the pointer beyond the
|
||||
call.
|
||||
|
||||
## The frame, the stack and the collector
|
||||
|
||||
The runtime owns the stack and the pointer map, and assembly must hold up
|
||||
its end of four rules:
|
||||
|
||||
1. **Arguments are initialised on entry; results are not.** A function whose
|
||||
results hold live pointers across a call must zero them and then execute
|
||||
`GO_RESULTS_INITIALIZED`. Designing functions that return no pointers
|
||||
avoids the problem.
|
||||
2. **A frame with calls and no local pointers says so** with
|
||||
`NO_LOCAL_POINTERS`. A frame with local pointers that the runtime cannot
|
||||
see is not allowed at all: assembly cannot describe a pointer-containing
|
||||
local, so it must not have one. Data symbols containing pointers are the
|
||||
same: define them in Go.
|
||||
3. **The stack may move.** Stack growth copies the frame, so no pointer into
|
||||
the frame may be held across a call, and the raw hardware SP register may
|
||||
not be cached across a call either.
|
||||
4. **The split check is not optional by default.** Without NOSPLIT, the
|
||||
assembler inserts the stack-growth preamble, including the morestack
|
||||
block for framed functions; NOSPLIT is a contract that the frame and
|
||||
everything below it fit in the remaining stack segment. On amd64 the
|
||||
assembler also marks small leaf functions NoSplit itself and skips the
|
||||
preamble, so silence is not a promise.
|
||||
|
||||
The simplest safe shape is a leaf function with no local frame and no calls:
|
||||
it needs no annotation beyond the prototype.
|
||||
|
||||
## go_asm.h: Go constants and layout in assembly
|
||||
|
||||
A package with `.s` files gets a generated header. Include it and use the
|
||||
generated names instead of hard-coding layouts, which lie silently when the
|
||||
Go side changes:
|
||||
|
||||
| Go declaration | Assembly name |
|
||||
|---|---|
|
||||
| `const bufSize = 1024` | `const_bufSize` |
|
||||
| field `r` of `type reader struct` | `reader_r` |
|
||||
| size of `type reader struct` | `reader__size` |
|
||||
|
||||
The constants arrive as macros, usable as immediates and offsets, computed
|
||||
from the Go declarations. An ambiguous name, such as a struct that really
|
||||
has a `_size` field, fails the generation with a redefinition error.
|
||||
|
||||
## funcdata.h: the runtime macros
|
||||
|
||||
`$GOROOT/pkg/include/funcdata.h` defines the PCDATA and FUNCDATA ids and the
|
||||
three macros assembly normally uses instead:
|
||||
|
||||
| Macro | Expands to | Meaning |
|
||||
|---|---|---|
|
||||
| `GO_ARGS` | `FUNCDATA $FUNCDATA_ArgsPointerMaps, go_args_stackmap(SB)` | the Go prototype defines the argument pointer map |
|
||||
| `GO_RESULTS_INITIALIZED` | `PCDATA $PCDATA_StackMapIndex, $1` | results are initialised; treat them as live from here |
|
||||
| `NO_LOCAL_POINTERS` | `FUNCDATA $FUNCDATA_LocalsPointerMaps, no_pointers_stackmap(SB)` | the frame holds no pointers |
|
||||
|
||||
`GO_ARGS` is inserted implicitly by the assembler for any function whose
|
||||
package-qualified name belongs to the current package, which is why most
|
||||
assembly never writes it. `NOSPLIT` leaf functions that call nothing need
|
||||
none of the three.
|
||||
|
||||
The underlying ids, for reading toolchain output rather than for writing
|
||||
source: FUNCDATA 0 to 7 are args pointer maps, locals pointer maps, stack
|
||||
objects, inline tree, open-coded defer info, argument info, argument
|
||||
liveness and wrap info; PCDATA 0 to 4 are unsafe point, stack map index,
|
||||
inline tree index, argument liveness index and panic bounds.
|
||||
|
||||
## What the runtime does with all of this
|
||||
|
||||
The object file records the annotations as aux symbols and FuncInfo records;
|
||||
GOOBJ.md specifies the encoding. The linker assembles them into the runtime's
|
||||
pclntable, which traceback and the collector consume. An assembly function
|
||||
that misdeclares its frame is not a compile error and usually not a link
|
||||
error: it is a wrong collector decision or a wrong traceback at runtime,
|
||||
which is why the annotations are a contract and not documentation.
|
||||
+18
-1
@@ -2,7 +2,7 @@
|
||||
.SH NAME
|
||||
gasm-asm \- assemble Plan 9 assembly without the Go toolchain
|
||||
.SH SYNOPSIS
|
||||
.B gasm asm [\-\-format raw|elf|goobj] [\-p pkg] [\-GOARCH arch] [\-o out] <file>
|
||||
.B gasm asm [\-\-format raw|elf|goobj] [\-I dir] [\-p pkg] [\-GOARCH arch] [\-GOOS os] [\-o out] <file>
|
||||
.SH DESCRIPTION
|
||||
Assemble FILE without the Go toolchain: every TEXT function is encoded
|
||||
to machine code and printed as a hex dump. Supported architectures:
|
||||
@@ -42,11 +42,24 @@ need no toolchain at all.
|
||||
Framed functions receive the stack-split guard and the trailing
|
||||
morestack block, byte-identical to the toolchain's output, so split
|
||||
functions link too.
|
||||
.PP
|
||||
A file that includes go_asm.h gets that header generated from the Go
|
||||
files beside it, type-checked for the target.
|
||||
.B \-GOOS
|
||||
selects the type-checking GOOS for that header, because a GOOS-specific
|
||||
file needs its platform's defines: sys_darwin_arm64.s fails against the
|
||||
ambient GOOS (machTimebaseInfo_numer is missing from a linux type-check)
|
||||
and assembles with
|
||||
.BR "\-GOOS darwin" .
|
||||
.SH OPTIONS
|
||||
.TP
|
||||
.B \-\-format \fIraw|elf|goobj\fR
|
||||
Output format; the default is raw.
|
||||
.TP
|
||||
.B \-I \fIdir\fR
|
||||
Directory to search for #include files; may be repeated, searched in
|
||||
order after the source directory.
|
||||
.TP
|
||||
.B \-p \fIpkg\fR
|
||||
Package path for --format goobj, qualifying the exported symbols.
|
||||
.TP
|
||||
@@ -55,6 +68,10 @@ Target architecture: amd64, arm64, riscv64 or loong64; overrides the
|
||||
file-name suffix, which is how the suffix-less majority of GOROOT's
|
||||
files (cpu_x86.s, stub.s, ...) become assemblable.
|
||||
.TP
|
||||
.B \-GOOS \fIos\fR
|
||||
Operating system for the generated go_asm.h: any GOOS go/build
|
||||
recognises in file names; the default is the host's.
|
||||
.TP
|
||||
.B \-o \fIfile\fR
|
||||
Write the output to this file instead of a hex dump on stdout.
|
||||
.SH EXIT STATUS
|
||||
|
||||
@@ -1,8 +1,8 @@
|
||||
.TH GASM-AUDIT-INSTRUCTIONS 1 "2026-09-19" "gasm" "User Commands"
|
||||
.TH GASM-AUDIT-INSTRUCTIONS 1 "2026-09-21" "gasm" "User Commands"
|
||||
.SH NAME
|
||||
gasm-audit-instructions \- diff the encoder against the Go toolchain, or measure a corpus
|
||||
.SH SYNOPSIS
|
||||
.B gasm audit\-instructions [\-\-corpus [\fIdir\fR]] [amd64|arm64|riscv64|loong64]
|
||||
.B gasm audit\-instructions [\-\-corpus [\fIdir\fR]] [\-\-list] [\-I dir] [amd64|arm64|riscv64|loong64]
|
||||
.SH DESCRIPTION
|
||||
Compare the gasm encoder for the given architecture (default amd64)
|
||||
against
|
||||
@@ -38,6 +38,18 @@ second.
|
||||
.B \-\-corpus [\fIdir\fR]
|
||||
Assemble a corpus of .s files and report pass rates and failure
|
||||
reasons.
|
||||
.TP
|
||||
.B \-\-list
|
||||
With
|
||||
.BR \-\-corpus ,
|
||||
print every failing file with its failure reason, per architecture,
|
||||
instead of one representative file per reason.
|
||||
.TP
|
||||
.B \-I \fIdir\fR
|
||||
Directory to search for #include files; may be repeated, searched in
|
||||
order after the source directory. A corpus run whose files include
|
||||
toolchain headers (such as GOROOT/pkg/include) needs it, the same -I a
|
||||
toolchain comparison takes.
|
||||
.SH EXIT STATUS
|
||||
The mnemonic-diff mode reports through its output and exits 0; a failed
|
||||
probe or an unknown architecture exits non-zero.
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
.SH NAME
|
||||
gasm-diff \- compare the machine code of two assembly files
|
||||
.SH SYNOPSIS
|
||||
.B gasm diff [\-GOARCH arch] <file1.s> <file2.s>
|
||||
.B gasm diff [\-GOARCH arch] [\-I dir] <file1.s> <file2.s>
|
||||
.SH DESCRIPTION
|
||||
Compare the machine code produced by assembling two files. Shows which
|
||||
functions differ and the byte-level differences. Useful for verifying
|
||||
@@ -20,6 +20,10 @@ pairs two variants regardless of suffix.
|
||||
Target architecture for both files: amd64, arm64, riscv64 or loong64;
|
||||
overrides the file-name suffixes.
|
||||
.TP
|
||||
.B \-I \fIdir\fR
|
||||
Directory to search for #include files; may be repeated, searched in
|
||||
order after the source directory.
|
||||
.TP
|
||||
.B \-\-map \fIspec\fR
|
||||
Comma-separated old=new pairs to match functions with different names.
|
||||
.SH EXIT STATUS
|
||||
|
||||
+23
-6
@@ -241,6 +241,13 @@ func renderInstr(line []token.Token, width int) string {
|
||||
if line[0].Kind != token.Ident {
|
||||
return "\t" + mnem + " " + ops
|
||||
}
|
||||
// A statement separator belongs to the statement it ends: when the
|
||||
// operands open with a ';', the alignment padding would land between
|
||||
// the mnemonic and its own separator (REP ; MOVSQ), so such a line
|
||||
// renders with a single space whatever the function's width.
|
||||
if strings.HasPrefix(ops, ";") {
|
||||
return "\t" + mnem + " " + ops
|
||||
}
|
||||
if width < len(mnem) {
|
||||
width = len(mnem)
|
||||
}
|
||||
@@ -254,11 +261,12 @@ func renderPreproc(line []token.Token) string {
|
||||
line[2].Kind == token.String {
|
||||
return "#include " + line[2].Text
|
||||
}
|
||||
parts := make([]string, 0, len(line)-1)
|
||||
for _, t := range line[1:] {
|
||||
parts = append(parts, t.Text)
|
||||
}
|
||||
return "#" + strings.Join(parts, " ")
|
||||
// The body of a directive, a macro definition included, is an ordinary
|
||||
// token run: rendering it through renderOps applies the same punctuation
|
||||
// rules as everywhere else, so a macro body keeps its canonical spelling
|
||||
// ($v, (a, b), the ';' separators between statements) instead of being
|
||||
// spread with a space between every token.
|
||||
return "#" + renderOps(line[1:])
|
||||
}
|
||||
|
||||
// renderOps re-spaces a run of operand tokens into canonical form. It never
|
||||
@@ -348,6 +356,13 @@ func spaceBetween(prev, cur token.Token) bool {
|
||||
return false
|
||||
case token.Comma:
|
||||
return false
|
||||
case token.Semicolon:
|
||||
// A ';' is a statement separator on the assembly path, not an
|
||||
// operand: dropping it would fuse two statements into a line the
|
||||
// assembler rejects, so it must survive as punctuation. It glues
|
||||
// to the statement it ends and the next statement takes one space,
|
||||
// matching the toolchain's listing style.
|
||||
return false
|
||||
case token.Star, token.Plus, token.Minus, token.Slash, token.Pipe:
|
||||
return false
|
||||
case token.LShift, token.RShift, token.Arrow, token.At:
|
||||
@@ -376,7 +391,9 @@ func spaceBetween(prev, cur token.Token) bool {
|
||||
return false
|
||||
case token.LAngle, token.RAngle:
|
||||
return false
|
||||
case token.Comma:
|
||||
case token.Comma, token.Semicolon:
|
||||
// The statement after a ';' separator takes its own space, exactly
|
||||
// like the operand after a comma.
|
||||
return true
|
||||
}
|
||||
return true
|
||||
|
||||
@@ -5,9 +5,11 @@ package format
|
||||
|
||||
import (
|
||||
"os"
|
||||
"slices"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/lexer"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/token"
|
||||
@@ -298,6 +300,196 @@ func TestCRLFInputIsNormalisedToLF(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// TestSemicolonSeparators pins the treatment of ';' statement separators.
|
||||
// The separator is load-bearing on the assembly path, where the parser reads
|
||||
// semicolon-separated statements: a formatter that drops it fuses two
|
||||
// statements into a line the assembler rejects, which is data corruption.
|
||||
// Each row pins the canonical spelling, one space after the ';', tight
|
||||
// before it, the way the toolchain's own sources and listings write it.
|
||||
func TestSemicolonSeparators(t *testing.T) {
|
||||
cases := []struct {
|
||||
name string
|
||||
in string
|
||||
want string
|
||||
}{
|
||||
{
|
||||
name: "between instructions, tight",
|
||||
in: "TEXT ·f(SB), $0\nBYTE $0x48;BYTE $0xc7\nRET\n",
|
||||
want: "TEXT ·f(SB), $0\n\tBYTE $0x48; BYTE $0xc7\n\tRET\n",
|
||||
},
|
||||
{
|
||||
name: "between instructions, spaced",
|
||||
in: "TEXT ·f(SB), $0\nBYTE $0x48 ; BYTE $0xc7\nRET\n",
|
||||
want: "TEXT ·f(SB), $0\n\tBYTE $0x48; BYTE $0xc7\n\tRET\n",
|
||||
},
|
||||
{
|
||||
name: "after a label",
|
||||
in: "TEXT ·f(SB), $0\nlabel: BYTE $1; BYTE $2\nRET\n",
|
||||
want: "TEXT ·f(SB), $0\nlabel:\n\tBYTE $1; BYTE $2\n\tRET\n",
|
||||
},
|
||||
{
|
||||
// The continuation-spliced macro shape of the runtime sources:
|
||||
// the lexer makes one logical line of the backslash continuations.
|
||||
name: "inside a macro body, continued",
|
||||
in: "#define MOVLTOREG(v, off) \\\n\tMOVL $v, AX; \\\n\tMOVL AX, ret+off(FP)\n",
|
||||
want: "#define MOVLTOREG(v, off) MOVL $v, AX; MOVL AX, ret+off(FP)\n",
|
||||
},
|
||||
{
|
||||
name: "inside a macro body, one line",
|
||||
in: "#define PEAS BYTE $0x0a; BYTE $0x0b\n",
|
||||
want: "#define PEAS BYTE $0x0a; BYTE $0x0b\n",
|
||||
},
|
||||
{
|
||||
name: "several separators in one line",
|
||||
in: "TEXT ·f(SB), $0\nBYTE $1; BYTE $2; BYTE $3\nRET\n",
|
||||
want: "TEXT ·f(SB), $0\n\tBYTE $1; BYTE $2; BYTE $3\n\tRET\n",
|
||||
},
|
||||
{
|
||||
name: "two separators back to back",
|
||||
in: "TEXT ·f(SB), $0\nBYTE $1;; BYTE $2\nRET\n",
|
||||
want: "TEXT ·f(SB), $0\n\tBYTE $1;; BYTE $2\n\tRET\n",
|
||||
},
|
||||
{
|
||||
name: "inside a line comment, untouched",
|
||||
in: "TEXT ·f(SB), $0\n// keep; the; separators\nBYTE $1\nRET\n",
|
||||
want: "TEXT ·f(SB), $0\n\t// keep; the; separators\n\tBYTE $1\n\tRET\n",
|
||||
},
|
||||
{
|
||||
name: "after a statement, before a comment",
|
||||
in: "TEXT ·f(SB), $0\nMOVQ AX, BX; // tail\nRET\n",
|
||||
want: "TEXT ·f(SB), $0\n\tMOVQ AX, BX; // tail\n\tRET\n",
|
||||
},
|
||||
{
|
||||
name: "last character on a line",
|
||||
in: "TEXT ·f(SB), $0\nBYTE $1;\nRET\n",
|
||||
want: "TEXT ·f(SB), $0\n\tBYTE $1;\n\tRET\n",
|
||||
},
|
||||
{
|
||||
// The REP shape: a prefix-style zero-operand statement
|
||||
// followed by the instruction it prefixes. The separator
|
||||
// belongs to the statement it ends, so the function's
|
||||
// alignment width (MOVSQ is the widest mnemonic here) must
|
||||
// not open a gap before it: one space after the mnemonic
|
||||
// whatever the neighbours' lengths.
|
||||
name: "after a prefix-style statement",
|
||||
in: "TEXT ·f(SB), $0\nMOVQ AX, BX\nREP; MOVSQ\nRET\n",
|
||||
want: "TEXT ·f(SB), $0\n\tMOVQ AX, BX\n\tREP ; MOVSQ\n\tRET\n",
|
||||
},
|
||||
}
|
||||
for _, tc := range cases {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
got := Source(tc.in)
|
||||
if got != tc.want {
|
||||
t.Fatalf("formatting mismatch:\n--- got ---\n%q\n--- want ---\n%q", got, tc.want)
|
||||
}
|
||||
if again := Source(got); again != got {
|
||||
t.Fatalf("not idempotent:\n%q", again)
|
||||
}
|
||||
if in, out := strings.Count(tc.in, ";"), strings.Count(got, ";"); in != out {
|
||||
t.Fatalf("semicolon count changed: %d -> %d\n%s", in, out, got)
|
||||
}
|
||||
if _, errs := parser.Parse("in.s", got); len(errs) > 0 {
|
||||
t.Fatalf("formatted output no longer parses: %v", errs)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestSemicolonStatementRoundTrip proves the formatter's contract on the
|
||||
// path where ';' separates statements: parse the source the way the
|
||||
// assembler does, format it, re-parse the formatted text and compare the
|
||||
// statement sequence. Raw operand texts are token-joined, so they are
|
||||
// insensitive to the whitespace a format pass chooses, and the comparison
|
||||
// can only fail when a token is lost: dropping a ';' fuses two statements
|
||||
// into one, exactly the corruption the released formatter committed.
|
||||
func TestSemicolonStatementRoundTrip(t *testing.T) {
|
||||
src := "#define MOVLTOREG(v, off) \\\n" +
|
||||
"\tMOVL $v, AX; \\\n" +
|
||||
"\tMOVL AX, ret+off(FP)\n" +
|
||||
"\n" +
|
||||
"TEXT ·f(SB), NOSPLIT, $0\n" +
|
||||
"BYTE $0x48; BYTE $0xc7\n" +
|
||||
"first: BYTE $1; BYTE $2\n" +
|
||||
"MOVLTOREG($42, 0)\n" +
|
||||
"RET\n"
|
||||
|
||||
before, errs := parser.ParseWithOptions("in.s", src, parser.Options{Expand: true})
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("source does not parse: %v", errs)
|
||||
}
|
||||
formatted := Source(src)
|
||||
after, errs := parser.ParseWithOptions("in.s", formatted, parser.Options{Expand: true})
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("formatted source does not parse: %v", errs)
|
||||
}
|
||||
|
||||
want, got := stmtSignature(before), stmtSignature(after)
|
||||
if !slices.Equal(got, want) {
|
||||
t.Fatalf("statement sequence changed:\n--- before ---\n%q\n--- after ---\n%q", want, got)
|
||||
}
|
||||
if again := Source(formatted); again != formatted {
|
||||
t.Fatalf("not idempotent:\n%q", again)
|
||||
}
|
||||
|
||||
// The two BYTE statements on the first line must stay two: one fused
|
||||
// statement here is the exact defect this package once shipped.
|
||||
var bytes []string
|
||||
for _, stmt := range stmtSignature(after) {
|
||||
if rest, ok := strings.CutPrefix(stmt, "instr BYTE "); ok {
|
||||
bytes = append(bytes, rest)
|
||||
}
|
||||
}
|
||||
if want := []string{"$ 0x48", "$ 0xc7", "$ 1", "$ 2"}; !slices.Equal(bytes, want) {
|
||||
t.Fatalf("BYTE statements after expansion = %q, want %q", bytes, want)
|
||||
}
|
||||
}
|
||||
|
||||
// stmtSignature flattens a parsed file into one string per declaration and
|
||||
// statement, in source order. Every component is token-derived, so the
|
||||
// signature is stable across format passes and moves only when a token is
|
||||
// lost or gained.
|
||||
func stmtSignature(f *ast.File) []string {
|
||||
var out []string
|
||||
for _, d := range f.Decls {
|
||||
switch d := d.(type) {
|
||||
case *ast.Text:
|
||||
out = append(out, "text "+d.Name.Raw)
|
||||
for _, s := range d.Body {
|
||||
out = append(out, stmtText(s))
|
||||
}
|
||||
case *ast.Globl:
|
||||
out = append(out, "globl "+d.Name.Raw)
|
||||
case *ast.Data:
|
||||
out = append(out, "data "+d.Name.Raw)
|
||||
case *ast.Include:
|
||||
out = append(out, "include "+d.Header.Text)
|
||||
case *ast.Preproc:
|
||||
out = append(out, "preproc "+d.Raw)
|
||||
}
|
||||
}
|
||||
for _, s := range f.Orphans {
|
||||
out = append(out, stmtText(s))
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// stmtText renders one statement for stmtSignature.
|
||||
func stmtText(s ast.Stmt) string {
|
||||
switch s := s.(type) {
|
||||
case *ast.Label:
|
||||
return "label " + s.Name.Text
|
||||
case *ast.Instr:
|
||||
parts := make([]string, 0, len(s.Operands)+1)
|
||||
parts = append(parts, s.Mnemonic.Text)
|
||||
for _, op := range s.Operands {
|
||||
parts = append(parts, op.Raw)
|
||||
}
|
||||
return "instr " + strings.Join(parts, " ")
|
||||
default:
|
||||
return "stmt"
|
||||
}
|
||||
}
|
||||
|
||||
// lexOperands lexes a single operand string and drops the EOF token.
|
||||
func lexOperands(s string) []token.Token {
|
||||
toks := lexer.Tokenize(s)
|
||||
|
||||
@@ -31,6 +31,12 @@ func FuzzFormatIdempotency(f *testing.F) {
|
||||
f.Add("TEXT ·f(SB), NOSPLIT, $0\n\tMOVQ AX, BX\n\tRET\n")
|
||||
f.Add("TEXT ·f(SB),NOSPLIT,$0\n\tMOVQ AX,BX\n\n\n\tRET\n")
|
||||
f.Add("garbage ### ???\n")
|
||||
// Line-ending whitespace at the edge of a comment: a CR followed by more
|
||||
// trailing whitespace once survived the first pass and disappeared on
|
||||
// re-lexing, so formatting was not idempotent.
|
||||
f.Add("//\r ")
|
||||
f.Add("// loop \r\t\nMOVQ AX, BX\n")
|
||||
f.Add("TEXT ·f(SB), NOSPLIT, $0 // tail\r\n\tMOVQ AX, BX\r\n\tRET\r\n")
|
||||
|
||||
f.Fuzz(func(t *testing.T, src string) {
|
||||
once := Source(src)
|
||||
|
||||
@@ -0,0 +1,2 @@
|
||||
go test fuzz v1
|
||||
string("//\r ")
|
||||
+62
-7
@@ -16,8 +16,16 @@ import (
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/token"
|
||||
)
|
||||
|
||||
// middleDot is the Plan 9 symbol separator (U+00B7), used in ·funcName(SB).
|
||||
const middleDot = '\u00B7'
|
||||
const (
|
||||
// middleDot is the Plan 9 symbol separator (U+00B7), used in
|
||||
// ·funcName(SB): it stands for the period between package path and name.
|
||||
middleDot = '\u00B7'
|
||||
// divisionSlash is the Plan 9 path separator (U+2215), used inside the
|
||||
// package path of a symbol: internal∕runtime∕atomic·Xchg. Like the
|
||||
// middle dot it is an identifier character, so a package path containing
|
||||
// it lexes as one name; the ordinary slash (U+002F) stays punctuation.
|
||||
divisionSlash = '\u2215'
|
||||
)
|
||||
|
||||
// Lexer scans a source string one token at a time.
|
||||
type Lexer struct {
|
||||
@@ -119,7 +127,9 @@ func (l *Lexer) Next() token.Token {
|
||||
// is a C-preprocessor line continuation (used by #define macros in the
|
||||
// runtime .s files): splice the lines together by consuming both, so
|
||||
// the whole macro becomes one logical line that the parser treats as an
|
||||
// opaque preprocessor directive.
|
||||
// opaque preprocessor directive. The backslash may also reach its
|
||||
// newline across whitespace and a trailing comment ("…; \ // note\n"),
|
||||
// which the toolchain's scanner skips the same way.
|
||||
for {
|
||||
c := l.cur()
|
||||
if c == ' ' || c == '\t' || c == '\r' {
|
||||
@@ -136,6 +146,16 @@ func (l *Lexer) Next() token.Token {
|
||||
}
|
||||
continue
|
||||
}
|
||||
if c == '\\' && l.continuationAhead() {
|
||||
l.advance() // backslash, then the runes the scan saw
|
||||
for !l.atEnd() && l.cur() != '\n' {
|
||||
l.advance()
|
||||
}
|
||||
if !l.atEnd() {
|
||||
l.advance() // the newline that closes the continuation
|
||||
}
|
||||
continue
|
||||
}
|
||||
break
|
||||
}
|
||||
|
||||
@@ -185,16 +205,42 @@ func (l *Lexer) Next() token.Token {
|
||||
}
|
||||
}
|
||||
|
||||
// continuationAhead reports, without consuming anything, whether the
|
||||
// backslash at the current position closes onto a newline through nothing
|
||||
// but horizontal whitespace and one line comment. Positions after the
|
||||
// backslash are inspected directly on the rune slice so a non-match leaves
|
||||
// the scanner state untouched.
|
||||
func (l *Lexer) continuationAhead() bool {
|
||||
i := l.i + 1
|
||||
for i < len(l.src) {
|
||||
switch r := l.src[i]; {
|
||||
case r == ' ' || r == '\t' || r == '\r':
|
||||
i++
|
||||
case r == '/' && i+1 < len(l.src) && l.src[i+1] == '/':
|
||||
for i < len(l.src) && l.src[i] != '\n' {
|
||||
i++
|
||||
}
|
||||
default:
|
||||
return r == '\n'
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// lineComment consumes a // comment up to, but not including, the newline. A
|
||||
// trailing \r is part of a CRLF line ending rather than comment content:
|
||||
// dropping it keeps the formatter's output uniformly LF-terminated.
|
||||
// trailing run of \r, spaces and tabs is line-ending whitespace rather than
|
||||
// comment content, so it never enters the token text. Trimming only a \r
|
||||
// directly before the token's end would make the text depend on what follows
|
||||
// the comment (a newline or the end of the input): "//x\r " would carry the
|
||||
// "\r " while "//x\r\n" would not, and a formatter that terminates the line
|
||||
// with \n would then re-lex its own output to a shorter comment.
|
||||
func (l *Lexer) lineComment(start token.Position) token.Token {
|
||||
var b strings.Builder
|
||||
for !l.atEnd() && l.cur() != '\n' {
|
||||
b.WriteRune(l.cur())
|
||||
l.advance()
|
||||
}
|
||||
return l.make(token.Comment, start, strings.TrimSuffix(b.String(), "\r"))
|
||||
return l.make(token.Comment, start, strings.TrimRight(b.String(), " \t\r"))
|
||||
}
|
||||
|
||||
// blockComment consumes a /* ... */ comment, tolerating an unterminated one.
|
||||
@@ -399,6 +445,15 @@ func (l *Lexer) punct(start token.Position) token.Token {
|
||||
case '|':
|
||||
l.advance()
|
||||
return l.make(token.Pipe, start, "|")
|
||||
case ';':
|
||||
l.advance()
|
||||
return l.make(token.Semicolon, start, ";")
|
||||
case '&':
|
||||
l.advance()
|
||||
return l.make(token.Ampersand, start, "&")
|
||||
case '~':
|
||||
l.advance()
|
||||
return l.make(token.Tilde, start, "~")
|
||||
default:
|
||||
// Unknown rune: emit it as Illegal and move on.
|
||||
l.advance()
|
||||
@@ -413,7 +468,7 @@ func isHexDigit(r rune) bool {
|
||||
}
|
||||
|
||||
func isIdentStart(r rune) bool {
|
||||
return r == '_' || r == middleDot || unicode.IsLetter(r)
|
||||
return r == '_' || r == middleDot || r == divisionSlash || unicode.IsLetter(r)
|
||||
}
|
||||
|
||||
func isIdentChar(r rune) bool {
|
||||
|
||||
@@ -85,6 +85,19 @@ func TestLabelAndComment(t *testing.T) {
|
||||
[]token.Kind{token.Ident, token.Colon, token.Ident, token.Ident, token.Comment})
|
||||
}
|
||||
|
||||
func TestLineCommentTrailingWhitespace(t *testing.T) {
|
||||
// A trailing run of CR, spaces and tabs is line-ending whitespace, not
|
||||
// comment content. The token text must not depend on what follows the
|
||||
// comment: before the trim covered only a CR directly before the token's
|
||||
// end, "// loop\r " kept the CR while "// loop\r\n" dropped it, and the
|
||||
// formatter re-lexed its own output to a shorter comment.
|
||||
eq(t, texts("// loop\r"), []string{"// loop"})
|
||||
eq(t, texts("// loop\r "), []string{"// loop"})
|
||||
eq(t, texts("// loop \r\t\nMOVQ AX, BX"), []string{"// loop", "MOVQ", "AX", ",", "BX"})
|
||||
// A CR inside the comment is content and stays.
|
||||
eq(t, texts("// loops\rall"), []string{"// loops\rall"})
|
||||
}
|
||||
|
||||
func TestAVX512Mnemonics(t *testing.T) {
|
||||
eq(t, texts("VFMADD231PD Z14, Z12, Z10"),
|
||||
[]string{"VFMADD231PD", "Z14", ",", "Z12", ",", "Z10"})
|
||||
@@ -169,6 +182,19 @@ func TestNulIsIllegal(t *testing.T) {
|
||||
eq(t, texts("MOVQ \x00 AX"), []string{"MOVQ", "\x00", "AX"})
|
||||
}
|
||||
|
||||
func TestDivisionSlashInIdentifiers(t *testing.T) {
|
||||
// U+2215 DIVISION SLASH is an identifier character, the way the
|
||||
// toolchain's tokenizer treats it: the package path of a symbol is
|
||||
// written with it (internal∕runtime∕atomic·Xchg) and must lex as one
|
||||
// name. The ordinary slash (U+002F) stays punctuation.
|
||||
eq(t, texts("CALL internal∕runtime∕atomic·Xchg(SB)"),
|
||||
[]string{"CALL", "internal∕runtime∕atomic·Xchg", "(", "SB", ")"})
|
||||
eq(t, texts("MOVQ sync∕atomic·Align(SB), AX"),
|
||||
[]string{"MOVQ", "sync∕atomic·Align", "(", "SB", ")", ",", "AX"})
|
||||
// It may also begin a name, like any letter of the toolchain's rule.
|
||||
eq(t, kinds("∕x"), []token.Kind{token.Ident})
|
||||
}
|
||||
|
||||
// TestOffsetsAroundInvalidByte pins Position.Offset against the original
|
||||
// bytes: an invalid UTF-8 byte decodes to RuneError but advances the offset
|
||||
// table by exactly one byte, so every later position stays a true byte
|
||||
|
||||
@@ -21,6 +21,12 @@ import (
|
||||
// The check requires a parseable signature; functions without one, and
|
||||
// functions whose parameters are all covered by frame reads, stay silent.
|
||||
func checkABI0Args(t *ast.Text) []Diagnostic {
|
||||
// An explicit <ABIInternal> TEXT reads its arguments from the register
|
||||
// file by declaration (runtime·memmove<ABIInternal> is the canonical
|
||||
// example), so the ABI0 frame contract does not apply to it.
|
||||
if t.Name != nil && t.Name.ABI != "" {
|
||||
return nil
|
||||
}
|
||||
params, ok := abiParamNames(t.Doc)
|
||||
if !ok || len(params) == 0 {
|
||||
return nil
|
||||
|
||||
@@ -82,6 +82,23 @@ func TestABIArgSizeSkipsRegisterABI(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// TestABI0ArgsSkipsABIInternal verifies the frame-read check does not fire for
|
||||
// a TEXT declared <ABIInternal>: runtime·memmove<ABIInternal> and friends read
|
||||
// their arguments from the register file by declaration, which is the correct
|
||||
// spelling there, not the register-args port bug the rule hunts.
|
||||
func TestABI0ArgsSkipsABIInternal(t *testing.T) {
|
||||
diags := lintSrc(t, "#include \"textflag.h\"\n"+
|
||||
"// func memmove(to, from unsafe.Pointer, n uintptr)\n"+
|
||||
"TEXT ·memmove<ABIInternal>(SB), NOSPLIT, $0-24\n"+
|
||||
"\tMOVQ AX, DI\n"+
|
||||
"\tMOVQ BX, SI\n"+
|
||||
"\tMOVQ CX, BX\n"+
|
||||
"\tRET\n")
|
||||
if codes(diags)[CodeABI0RegisterArgs] != 0 {
|
||||
t.Fatalf("ABIInternal TEXT must not be checked against the FP frame: %+v", diags)
|
||||
}
|
||||
}
|
||||
|
||||
// TestUnreachableCode exercises the dead-code detection and its guard rails.
|
||||
func TestUnreachableCode(t *testing.T) {
|
||||
// Code after a RET is unreachable.
|
||||
|
||||
+111
-8
@@ -338,10 +338,8 @@ func lintText(t *ast.Text, tab *arch.Table, archKnown bool, cfg Config, macros m
|
||||
}
|
||||
|
||||
if isJump(cfg.Arch, upper) {
|
||||
for _, op := range st.Operands {
|
||||
if name, pos, ok := localLabelRef(op); ok && !tab.IsRegister(name) && !arch.IsPseudoReg(name) {
|
||||
referenced[name] = pos
|
||||
}
|
||||
if name, pos, ok := branchTargetRef(cfg.Arch, upper, st.Operands, tab); ok {
|
||||
referenced[name] = pos
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -649,17 +647,65 @@ func localLabelRef(op *ast.Operand) (string, token.Position, bool) {
|
||||
return sym.Name, op.Pos, true
|
||||
}
|
||||
|
||||
// branchTargetRef returns the local label a branch transfers control to: the
|
||||
// bare symbol in the destination position, the last operand, since that is
|
||||
// where the Plan 9 branch target sits. A register-named target is a
|
||||
// register-indirect branch (JMP AX, arm64 BR R5, riscv64 JALR X6, loong64
|
||||
// JIRL R1) and yields no reference, unless the encoder reads the target
|
||||
// positionally (positionalBranchTarget): there a label may legitimately
|
||||
// collide with a register alias, riscv64 ZERO being the ABI name of X0, and
|
||||
// a label named zero is ordinary code.
|
||||
func branchTargetRef(a arch.Arch, upper string, ops []*ast.Operand, tab *arch.Table) (string, token.Position, bool) {
|
||||
if len(ops) == 0 {
|
||||
return "", token.Position{}, false
|
||||
}
|
||||
name, pos, ok := localLabelRef(ops[len(ops)-1])
|
||||
if !ok {
|
||||
return "", token.Position{}, false
|
||||
}
|
||||
if !positionalBranchTarget(a, upper) && (tab.IsRegister(name) || arch.IsPseudoReg(name)) {
|
||||
return "", token.Position{}, false
|
||||
}
|
||||
return name, pos, true
|
||||
}
|
||||
|
||||
// positionalBranchTarget reports whether the encoder reads a bare-symbol
|
||||
// operand of the branch as its label target from a fixed position, without
|
||||
// consulting the register file. The riscv64 branch, JMP and JAL encoders do
|
||||
// (labelFromOperand in asm/riscv_assemble.go), as do the loong64 branch,
|
||||
// BFPT/BFPF and jump encoders (l64Label in asm/loong64_assemble.go). amd64
|
||||
// never does, because a bare register operand to JMP/CALL/Jcc is a
|
||||
// register-indirect branch; nor do the register-indirect forms of the RISC
|
||||
// families (arm64 BR/BLR, riscv64 JALR/JR, loong64 JIRL).
|
||||
func positionalBranchTarget(a arch.Arch, upper string) bool {
|
||||
switch a {
|
||||
case arch.RISCV:
|
||||
return riscvBranches[upper] || upper == "JMP" || upper == "JAL"
|
||||
case arch.LOONG64:
|
||||
return loong64Branches[upper] || upper == "JMP" || upper == "B" ||
|
||||
upper == "JAL" || upper == "BL"
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// riscvBranches and loong64Branches are the conditional-branch mnemonics; they
|
||||
// are listed explicitly rather than matched by a "B" prefix so that bit-manip
|
||||
// instructions (BCLR, BSET, …) are never mistaken for branches.
|
||||
// instructions (BCLR, BSET, …) are never mistaken for branches. The sets
|
||||
// mirror the encoder's own branch cases: the B-type table entries
|
||||
// (riscv_encode.go), the branch-zero pseudos and the reversed branches
|
||||
// BGT/BGTU/BLE/BLEU (riscv_assemble.go), and for loong64 the 16-bit branch
|
||||
// table plus the single-register forms of l64branch21Table (BEQZ/BNEZ and the
|
||||
// floating-point branches BFPT/BFPF).
|
||||
var riscvBranches = map[string]bool{
|
||||
"BEQ": true, "BNE": true, "BLT": true, "BGE": true, "BLTU": true, "BGEU": true,
|
||||
"BEQZ": true, "BNEZ": true, "BLEZ": true, "BGEZ": true, "BLTZ": true, "BGTZ": true,
|
||||
"BGT": true, "BGTU": true, "BLE": true, "BLEU": true,
|
||||
}
|
||||
|
||||
var loong64Branches = map[string]bool{
|
||||
"BEQ": true, "BNE": true, "BLT": true, "BGE": true, "BLTU": true, "BGEU": true,
|
||||
"BLEZ": true, "BLTZ": true, "BGEZ": true, "BGTZ": true,
|
||||
"BEQZ": true, "BNEZ": true, "BFPT": true, "BFPF": true,
|
||||
}
|
||||
|
||||
// isJump reports whether the mnemonic is any branch.
|
||||
@@ -676,7 +722,8 @@ func isJump(a arch.Arch, upper string) bool {
|
||||
upper == "JR" || upper == "BR"
|
||||
case arch.LOONG64:
|
||||
return upper == "CALL" || loong64Branches[upper] ||
|
||||
upper == "JIRL" || upper == "JMP" || upper == "BR"
|
||||
upper == "JIRL" || upper == "JMP" || upper == "BR" ||
|
||||
upper == "B" || upper == "JAL" || upper == "BL"
|
||||
default: // amd64
|
||||
return upper == "CALL" || strings.HasPrefix(upper, "J")
|
||||
}
|
||||
@@ -692,7 +739,8 @@ func isUnconditionalJump(a arch.Arch, upper string) bool {
|
||||
return upper == "JMP" || upper == "J" || upper == "JAL" ||
|
||||
upper == "JALR" || upper == "JR" || upper == "BR"
|
||||
case arch.LOONG64:
|
||||
return upper == "JMP" || upper == "JIRL" || upper == "BR"
|
||||
return upper == "JMP" || upper == "JIRL" || upper == "BR" || upper == "B" ||
|
||||
upper == "JAL" || upper == "BL"
|
||||
default:
|
||||
return upper == "JMP"
|
||||
}
|
||||
@@ -792,6 +840,43 @@ func isSPReg(op *ast.Operand, a arch.Arch) bool {
|
||||
return false
|
||||
}
|
||||
|
||||
// shiftRotateBases are the shift and rotate mnemonics without their width
|
||||
// suffix. These are the instructions whose encoder path (encodeShift) reads
|
||||
// the count from the first operand.
|
||||
var shiftRotateBases = map[string]bool{
|
||||
"SHL": true, "SHR": true, "SAR": true, "SAL": true,
|
||||
"ROL": true, "ROR": true, "RCL": true, "RCR": true,
|
||||
}
|
||||
|
||||
// isShiftCountOperand reports whether operand i of mnem is the shift count.
|
||||
// The ISA fixes the shift/rotate count register at CL: the D2/D3 group (and
|
||||
// C0/C1 for immediates) encode the count outside the ModRM register field,
|
||||
// so the count operand is 8-bit by definition no matter how wide the data is.
|
||||
// The count arrives as the first of the two operands; the one-operand form
|
||||
// does not exist.
|
||||
func isShiftCountOperand(mnem string, i, nops int) bool {
|
||||
if nops != 2 || i != 0 {
|
||||
return false
|
||||
}
|
||||
if shiftRotateBases[mnem] {
|
||||
return true
|
||||
}
|
||||
if len(mnem) > 1 {
|
||||
switch mnem[len(mnem)-1] {
|
||||
case 'Q', 'L', 'W', 'B':
|
||||
return shiftRotateBases[mnem[:len(mnem)-1]]
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// isSetcc reports whether the mnemonic is a SETcc: SET plus a condition code.
|
||||
// The membership test is the encoder's own SET dispatch, which asm.Encodable
|
||||
// mirrors.
|
||||
func isSetcc(mnem string) bool {
|
||||
return strings.HasPrefix(mnem, "SET") && asm.Encodable(mnem)
|
||||
}
|
||||
|
||||
// checkRegisterWidth detects amd64 register-width mismatches. The naming
|
||||
// truth of the Go assembler governs: AX, BX, CX, DX, SI, DI, BP, SP and
|
||||
// R8-R15 ARE the 64-bit register names (there are no separate EAX/RAX
|
||||
@@ -802,6 +887,14 @@ func isSPReg(op *ast.Operand, a arch.Arch) bool {
|
||||
// register (EAX under the gasm alias extension, or a byte form), and byte
|
||||
// registers in L/W operations.
|
||||
func checkRegisterWidth(mnem string, ops []*ast.Operand) string {
|
||||
// A SETcc stores one byte: the destination is an 8-bit register or an
|
||||
// 8-bit memory location by definition (0F 90+cc), whichever condition it
|
||||
// tests. The trailing letter of spellings like SETPL or SETEQ is part of
|
||||
// the condition code, not an operand width, so the whole family is
|
||||
// exempt from the suffix logic.
|
||||
if isSetcc(mnem) {
|
||||
return ""
|
||||
}
|
||||
// Determine expected width from mnemonic suffix.
|
||||
var expected int // 0=unknown, 8/4/2/1=bytes
|
||||
switch {
|
||||
@@ -816,10 +909,20 @@ func checkRegisterWidth(mnem string, ops []*ast.Operand) string {
|
||||
default:
|
||||
return "" // no suffix, can't determine width
|
||||
}
|
||||
for _, op := range ops {
|
||||
for i, op := range ops {
|
||||
if op.Kind != ast.OpAddr || op.Addr.Sym == nil {
|
||||
continue
|
||||
}
|
||||
// Only a bare register carries a width to compare: frame and static
|
||||
// symbol references (ch+8(FP), foo(SB)) and memory operands are not
|
||||
// registers even when their name collides with one.
|
||||
if op.Addr.Sym.Pseudo != "" || op.Addr.Base != "" || op.Addr.Index != "" {
|
||||
continue
|
||||
}
|
||||
// The shift/rotate count is exempt: fixed at 8 bits by the ISA.
|
||||
if isShiftCountOperand(mnem, i, len(ops)) {
|
||||
continue
|
||||
}
|
||||
name := strings.ToLower(op.Addr.Sym.Name)
|
||||
regWidth := amd64RegWidth(name)
|
||||
if regWidth == 0 {
|
||||
|
||||
@@ -8,6 +8,7 @@ import (
|
||||
"testing"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/asm"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
|
||||
)
|
||||
@@ -21,6 +22,21 @@ func lintSrc(t *testing.T, src string) []Diagnostic {
|
||||
return File(f, Config{Arch: arch.AMD64})
|
||||
}
|
||||
|
||||
// lintArchFile parses and lints src under a, then hands the same file to
|
||||
// assemble so the assertion is pinned against the encoder: a kernel the
|
||||
// linter reasons about must also be one the encoder accepts.
|
||||
func lintArchFile(t *testing.T, filename, src string, a arch.Arch, assemble func(*ast.File) (*asm.Image, error)) []Diagnostic {
|
||||
t.Helper()
|
||||
f, errs := parser.Parse(filename, src)
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
if _, err := assemble(f); err != nil {
|
||||
t.Fatalf("encoder rejects the kernel: %v", err)
|
||||
}
|
||||
return File(f, Config{Arch: a})
|
||||
}
|
||||
|
||||
// lintSrcArch lints src under the architecture inferred from filename.
|
||||
func lintSrcArch(t *testing.T, filename, src string) []Diagnostic {
|
||||
t.Helper()
|
||||
@@ -262,6 +278,138 @@ loop:
|
||||
}
|
||||
}
|
||||
|
||||
func TestRiscvBranchFamilyRegistersLabels(t *testing.T) {
|
||||
// Every riscv64 pseudo-branch that references a label must register that
|
||||
// reference: the reversed branches BGT/BGTU/BLE/BLEU (GOROOT's
|
||||
// memmove_riscv64 branches with BGTU) and a label named like the ZERO
|
||||
// register alias (GOROOT's memclr_riscv64 carries a label named zero;
|
||||
// ZERO is the ABI name of X0) must not be reported unused.
|
||||
diags := lintSrcArch(t, "f_riscv64.s", `
|
||||
#include "textflag.h"
|
||||
TEXT ·f(SB), NOSPLIT, $0
|
||||
BGTU X10, X11, backward
|
||||
BGT X10, X11, zero
|
||||
BLE X10, X11, one
|
||||
BLEU X10, X11, two
|
||||
BEQZ X10, zero
|
||||
BNEZ X10, one
|
||||
JMP two
|
||||
backward:
|
||||
RET
|
||||
zero:
|
||||
RET
|
||||
one:
|
||||
RET
|
||||
two:
|
||||
RET
|
||||
`)
|
||||
if codes(diags)[CodeUnusedLabel] != 0 {
|
||||
t.Fatalf("branch-referenced labels must not be flagged unused: %+v", diags)
|
||||
}
|
||||
if codes(diags)[CodeUndefinedLabel] != 0 {
|
||||
t.Fatalf("defined labels must resolve: %+v", diags)
|
||||
}
|
||||
|
||||
// A branch to a truly undefined label still reports.
|
||||
diags = lintSrcArch(t, "f_riscv64.s", `
|
||||
#include "textflag.h"
|
||||
TEXT ·f(SB), NOSPLIT, $0
|
||||
BGT X10, X11, nowhere
|
||||
RET
|
||||
`)
|
||||
if codes(diags)[CodeUndefinedLabel] != 1 {
|
||||
t.Fatalf("undefined branch target must be flagged: %+v", diags)
|
||||
}
|
||||
|
||||
// A register-indirect JALR is not a label reference.
|
||||
diags = lintSrcArch(t, "f_riscv64.s", `
|
||||
#include "textflag.h"
|
||||
TEXT ·f(SB), NOSPLIT, $0
|
||||
JALR X1
|
||||
RET
|
||||
`)
|
||||
if codes(diags)[CodeUndefinedLabel] != 0 {
|
||||
t.Fatalf("register operand of JALR is not a label: %+v", diags)
|
||||
}
|
||||
}
|
||||
|
||||
func TestLoong64BranchFamilyRegistersLabels(t *testing.T) {
|
||||
// The loong64 jumps and single-register branches (JAL, B, BL, BEQZ/BNEZ,
|
||||
// BFPT/BFPF) all reference their label from the last operand; GOROOT's
|
||||
// own basic kernels tail-call with JAL, so the reference must register.
|
||||
diags := lintSrcArch(t, "f_loong64.s", `
|
||||
#include "textflag.h"
|
||||
TEXT ·f(SB), NOSPLIT, $0
|
||||
BEQZ R4, fin
|
||||
BNEZ R4, fin
|
||||
BLTZ R4, fin
|
||||
JAL fin
|
||||
BL fin
|
||||
B fin
|
||||
RET
|
||||
fin:
|
||||
RET
|
||||
`)
|
||||
if codes(diags)[CodeUnusedLabel] != 0 {
|
||||
t.Fatalf("branch-referenced labels must not be flagged unused: %+v", diags)
|
||||
}
|
||||
if codes(diags)[CodeUndefinedLabel] != 0 {
|
||||
t.Fatalf("defined labels must resolve: %+v", diags)
|
||||
}
|
||||
}
|
||||
|
||||
// TestBranchFamiliesAssemble pins the lint branch sets to the encoder: every
|
||||
// mnemonic the linter classifies as a riscv64 or loong64 label branch must be
|
||||
// a branch the encoder actually assembles, with the label in the last
|
||||
// operand. If the encoder gains or renames a branch, this test fails and the
|
||||
// set follows it.
|
||||
func TestBranchFamiliesAssemble(t *testing.T) {
|
||||
riscvForms := map[string]string{}
|
||||
for m := range riscvBranches {
|
||||
riscvForms[m] = m + " X10, X11, tgt"
|
||||
}
|
||||
for _, m := range []string{"BEQZ", "BNEZ", "BLTZ", "BGEZ", "BLEZ", "BGTZ"} {
|
||||
riscvForms[m] = m + " X10, tgt"
|
||||
}
|
||||
riscvForms["JMP"] = "JMP tgt"
|
||||
riscvForms["JAL"] = "JAL tgt"
|
||||
|
||||
loongForms := map[string]string{}
|
||||
for _, m := range []string{"BEQ", "BNE", "BLT", "BGE", "BLTU", "BGEU"} {
|
||||
loongForms[m] = m + " R4, R5, tgt"
|
||||
}
|
||||
for _, m := range []string{"BEQZ", "BNEZ", "BLTZ", "BGEZ", "BLEZ", "BGTZ", "BFPT", "BFPF"} {
|
||||
loongForms[m] = m + " R4, tgt"
|
||||
}
|
||||
loongForms["JMP"] = "JMP tgt"
|
||||
loongForms["B"] = "B tgt"
|
||||
loongForms["JAL"] = "JAL tgt"
|
||||
loongForms["BL"] = "BL tgt"
|
||||
|
||||
for m, form := range riscvForms {
|
||||
src := "#include \"textflag.h\"\n" +
|
||||
"TEXT ·f(SB), NOSPLIT, $0\n" +
|
||||
"\t" + form + "\n" +
|
||||
"tgt:\n" +
|
||||
"\tRET\n"
|
||||
diags := lintArchFile(t, "f_riscv64.s", src, arch.RISCV, asm.AssembleFileRISCV)
|
||||
if codes(diags)[CodeUnusedLabel] != 0 || codes(diags)[CodeUndefinedLabel] != 0 {
|
||||
t.Errorf("riscv64 %s: label reference not registered: %+v", m, diags)
|
||||
}
|
||||
}
|
||||
for m, form := range loongForms {
|
||||
src := "#include \"textflag.h\"\n" +
|
||||
"TEXT ·f(SB), NOSPLIT, $0\n" +
|
||||
"\t" + form + "\n" +
|
||||
"tgt:\n" +
|
||||
"\tRET\n"
|
||||
diags := lintArchFile(t, "f_loong64.s", src, arch.LOONG64, asm.AssembleFileLOONG64)
|
||||
if codes(diags)[CodeUnusedLabel] != 0 || codes(diags)[CodeUndefinedLabel] != 0 {
|
||||
t.Errorf("loong64 %s: label reference not registered: %+v", m, diags)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestInvalidTextflag(t *testing.T) {
|
||||
diags := lintSrc(t, `
|
||||
#include "textflag.h"
|
||||
@@ -413,6 +561,72 @@ TEXT ·f(SB), NOSPLIT, $0
|
||||
}
|
||||
}
|
||||
|
||||
func TestRegisterWidthShiftCount(t *testing.T) {
|
||||
// The shift and rotate count lives in CL by ISA definition (the D2/D3
|
||||
// group encodes the count outside the ModRM register field), so the count
|
||||
// operand is 8-bit no matter how wide the data is: SHLQ CL, AX is the
|
||||
// normal spelling of a 64-bit shift. The data operand keeps its check.
|
||||
diags := lintSrc(t, `
|
||||
#include "textflag.h"
|
||||
TEXT ·f(SB), NOSPLIT, $0
|
||||
SHLQ CL, AX
|
||||
SHRL CL, BX
|
||||
SARQ CL, CX
|
||||
ROLL CL, DX
|
||||
RORQ CL, R8
|
||||
RCLL CL, R9
|
||||
RCRQ CL, R10
|
||||
MOVQ CL, R10
|
||||
RET
|
||||
`)
|
||||
if codes(diags)[CodeRegisterWidthMismatch] != 1 {
|
||||
t.Fatalf("only the MOVQ CL data move must be flagged, got %+v", diags)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRegisterWidthSetcc(t *testing.T) {
|
||||
// A SETcc stores one byte whichever condition it tests (0F 90+cc), so
|
||||
// SETNE AL is always right and the trailing letters of SETEQ, SETPL and
|
||||
// SETLS are condition codes, not width suffixes.
|
||||
diags := lintSrc(t, `
|
||||
#include "textflag.h"
|
||||
TEXT ·f(SB), NOSPLIT, $0
|
||||
CMPQ AX, BX
|
||||
SETNE AL
|
||||
SETEQ AL
|
||||
SETPL AL
|
||||
SETLS AL
|
||||
SETCC (BX)
|
||||
SETGE (R8)
|
||||
RET
|
||||
`)
|
||||
if codes(diags)[CodeRegisterWidthMismatch] != 0 {
|
||||
t.Fatalf("SETcc destinations are 8-bit by definition: %+v", diags)
|
||||
}
|
||||
if codes(diags)[CodeUnknownInstr] != 0 {
|
||||
t.Fatalf("every SETcc spelling must be known: %+v", diags)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRegisterWidthFrameNames(t *testing.T) {
|
||||
// GOROOT's BSD syscall stubs carry frame parameters whose names collide
|
||||
// with byte register names (kevent's ch and nch): MOVQ ch+8(FP), SI is a
|
||||
// frame reference, not the CH register.
|
||||
diags := lintSrc(t, `
|
||||
#include "textflag.h"
|
||||
TEXT ·kevent(SB), NOSPLIT, $0-36
|
||||
MOVL kq+0(FP), DI
|
||||
MOVQ ch+8(FP), SI
|
||||
MOVL nch+16(FP), DX
|
||||
MOVQ ev+24(FP), R10
|
||||
MOVQ AX, ret+32(FP)
|
||||
RET
|
||||
`)
|
||||
if codes(diags)[CodeRegisterWidthMismatch] != 0 {
|
||||
t.Fatalf("frame and static symbol names are not registers: %+v", diags)
|
||||
}
|
||||
}
|
||||
|
||||
func TestNonportableRegisterName(t *testing.T) {
|
||||
diags := lintSrc(t, `
|
||||
#include "textflag.h"
|
||||
|
||||
+146
@@ -0,0 +1,146 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
// Constant-expression folding for operands. The toolchain's assembler
|
||||
// evaluates arithmetic in every operand position, and macro-heavy GOROOT
|
||||
// sources lean on it: parameterised bodies carry offsets like
|
||||
// ((index*4)+0)(base), immediates like $(32-shift) and masks like
|
||||
// $~63 or $(1<<0|1<<9). Substituting the parameters textually therefore
|
||||
// leaves constant arithmetic behind, and the parser folds it here, keeping
|
||||
// the operand AST identical to what the same literals written out would
|
||||
// produce. Anything that is not a closed integer expression fails to fold
|
||||
// and falls through to the ordinary operand paths.
|
||||
package parser
|
||||
|
||||
import (
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/token"
|
||||
)
|
||||
|
||||
// foldExpr evaluates the constant integer expression at the head of ts and
|
||||
// returns its value together with the unconsumed tokens. ok is false when
|
||||
// the tokens do not form an expression, which is the callers' signal to use
|
||||
// the ordinary parsing paths.
|
||||
func foldExpr(ts []token.Token) (val int64, rest []token.Token, ok bool) {
|
||||
v, rest, ok := foldAdd(ts)
|
||||
if !ok {
|
||||
return 0, ts, false
|
||||
}
|
||||
return v, rest, true
|
||||
}
|
||||
|
||||
// foldAdd parses addition-level expressions: +, - and | bind loosest, the
|
||||
// Plan 9 convention that makes x<<1|3 read as (x<<1)|3.
|
||||
func foldAdd(ts []token.Token) (int64, []token.Token, bool) {
|
||||
v, rest, ok := foldMul(ts)
|
||||
if !ok {
|
||||
return 0, ts, false
|
||||
}
|
||||
for len(rest) > 0 {
|
||||
kind := rest[0].Kind
|
||||
if kind != token.Plus && kind != token.Minus && kind != token.Pipe {
|
||||
return v, rest, true
|
||||
}
|
||||
w, r2, ok := foldMul(rest[1:])
|
||||
if !ok {
|
||||
return v, rest, true
|
||||
}
|
||||
switch kind {
|
||||
case token.Plus:
|
||||
v += w
|
||||
case token.Minus:
|
||||
v -= w
|
||||
case token.Pipe:
|
||||
v |= w
|
||||
}
|
||||
rest = r2
|
||||
}
|
||||
return v, rest, true
|
||||
}
|
||||
|
||||
// foldMul parses multiplication-level expressions: *, / and the bit
|
||||
// operators &, << and >>.
|
||||
func foldMul(ts []token.Token) (int64, []token.Token, bool) {
|
||||
v, rest, ok := foldFactor(ts)
|
||||
if !ok {
|
||||
return 0, ts, false
|
||||
}
|
||||
for len(rest) > 0 {
|
||||
switch rest[0].Kind {
|
||||
case token.Star:
|
||||
w, r2, ok := foldFactor(rest[1:])
|
||||
if !ok {
|
||||
return v, rest, true
|
||||
}
|
||||
v *= w
|
||||
rest = r2
|
||||
case token.Slash:
|
||||
w, r2, ok := foldFactor(rest[1:])
|
||||
if !ok || w == 0 {
|
||||
return v, rest, true
|
||||
}
|
||||
v /= w
|
||||
rest = r2
|
||||
case token.Ampersand:
|
||||
w, r2, ok := foldFactor(rest[1:])
|
||||
if !ok {
|
||||
return v, rest, true
|
||||
}
|
||||
v &= w
|
||||
rest = r2
|
||||
case token.LShift:
|
||||
w, r2, ok := foldFactor(rest[1:])
|
||||
if !ok || w < 0 || w >= 64 {
|
||||
return v, rest, true
|
||||
}
|
||||
v <<= uint(w)
|
||||
rest = r2
|
||||
case token.RShift:
|
||||
w, r2, ok := foldFactor(rest[1:])
|
||||
if !ok || w < 0 || w >= 64 {
|
||||
return v, rest, true
|
||||
}
|
||||
v >>= uint(w)
|
||||
rest = r2
|
||||
default:
|
||||
return v, rest, true
|
||||
}
|
||||
}
|
||||
return v, rest, true
|
||||
}
|
||||
|
||||
// foldFactor parses a number, a parenthesised expression, or a unary sign
|
||||
// or complement.
|
||||
func foldFactor(ts []token.Token) (int64, []token.Token, bool) {
|
||||
if len(ts) == 0 {
|
||||
return 0, ts, false
|
||||
}
|
||||
switch ts[0].Kind {
|
||||
case token.Number:
|
||||
v, ok := tryInt(ts[0].Text)
|
||||
if !ok {
|
||||
return 0, ts, false
|
||||
}
|
||||
return v, ts[1:], true
|
||||
case token.LParen:
|
||||
v, rest, ok := foldAdd(ts[1:])
|
||||
if !ok || len(rest) == 0 || rest[0].Kind != token.RParen {
|
||||
return 0, ts, false
|
||||
}
|
||||
return v, rest[1:], true
|
||||
case token.Minus:
|
||||
v, rest, ok := foldFactor(ts[1:])
|
||||
if !ok {
|
||||
return 0, ts, false
|
||||
}
|
||||
return -v, rest, true
|
||||
case token.Plus:
|
||||
return foldFactor(ts[1:])
|
||||
case token.Tilde:
|
||||
v, rest, ok := foldFactor(ts[1:])
|
||||
if !ok {
|
||||
return 0, ts, false
|
||||
}
|
||||
return ^v, rest, true
|
||||
}
|
||||
return 0, ts, false
|
||||
}
|
||||
+169
-8
@@ -32,9 +32,8 @@ func (e Error) Error() string {
|
||||
// returned file is usable even when errors is non-empty.
|
||||
func Parse(path, src string) (*ast.File, []error) {
|
||||
tokens := lexer.Tokenize(src)
|
||||
lines := splitLines(tokens)
|
||||
p := &state{path: path}
|
||||
p.parse(lines)
|
||||
p.parse(statementLines(tokens))
|
||||
return p.file, p.errs
|
||||
}
|
||||
|
||||
@@ -73,6 +72,46 @@ func splitLines(tokens []token.Token) [][]token.Token {
|
||||
return lines
|
||||
}
|
||||
|
||||
// statementLines turns the token stream into the logical lines the parser
|
||||
// reads: physical lines split at the ';' statement separators, exactly the
|
||||
// way the expansion path treats the expanded bodies. The runtime writes
|
||||
// "ROLQ $3, DI; ROLQ $13, DI" and "REP; MOVSB" in plain files, and the
|
||||
// separator carries no meaning beyond the break. Comments are statement
|
||||
// text, not structure: the lexer delivers a whole comment as one token, so
|
||||
// a ';' inside a comment is never a separator; a comment after a statement
|
||||
// stays on that statement's line; and a comment that sits between
|
||||
// statements (the runtime's "NO_LOCAL_POINTERS; /* … */" style) stands as
|
||||
// its own logical line, like a whole-line comment.
|
||||
func statementLines(tokens []token.Token) [][]token.Token {
|
||||
var out [][]token.Token
|
||||
var cur []token.Token
|
||||
flush := func() {
|
||||
if len(cur) > 0 {
|
||||
out = append(out, cur)
|
||||
cur = nil
|
||||
}
|
||||
}
|
||||
for _, t := range tokens {
|
||||
switch t.Kind {
|
||||
case token.EOF:
|
||||
// The stream's terminator is not statement content.
|
||||
case token.Newline, token.Semicolon:
|
||||
flush()
|
||||
case token.Comment:
|
||||
if len(cur) > 0 {
|
||||
cur = append(cur, t)
|
||||
} else {
|
||||
out = append(out, []token.Token{t})
|
||||
}
|
||||
flush()
|
||||
default:
|
||||
cur = append(cur, t)
|
||||
}
|
||||
}
|
||||
flush()
|
||||
return out
|
||||
}
|
||||
|
||||
func (p *state) parse(lines [][]token.Token) {
|
||||
p.file = &ast.File{Path: p.path, Macros: map[string]bool{}}
|
||||
for _, line := range lines {
|
||||
@@ -265,7 +304,7 @@ func (p *state) parseGlobl(line []token.Token) *ast.Globl {
|
||||
rest = rest[1:]
|
||||
}
|
||||
if len(rest) > 0 && rest[0].Kind == token.Dollar {
|
||||
g.Size = parseOperand(rest)
|
||||
g.Size = parseOperand(rest, false)
|
||||
}
|
||||
return g
|
||||
}
|
||||
@@ -282,7 +321,7 @@ func (p *state) parseData(line []token.Token) *ast.Data {
|
||||
d.Name = sym
|
||||
d.Width = width
|
||||
if len(valuePart) > 0 {
|
||||
d.Value = parseOperand(stripComment(valuePart))
|
||||
d.Value = parseOperand(stripComment(valuePart), false)
|
||||
}
|
||||
return d
|
||||
}
|
||||
@@ -293,8 +332,13 @@ func (p *state) parseInstr(line []token.Token) {
|
||||
return
|
||||
}
|
||||
instr := &ast.Instr{Mnemonic: body[0], Comment: comment}
|
||||
for _, grp := range splitOperands(body[1:]) {
|
||||
if op := parseOperand(grp); op != nil {
|
||||
grps := splitOperands(body[1:])
|
||||
for i, grp := range grps {
|
||||
// Only the final operand slot may carry a bare constant: the
|
||||
// toolchain reads the trailing 1 of CMPSD X1, X0, 1 as $1
|
||||
// (math/floor_amd64.s), while an earlier bare number names an
|
||||
// absolute address, a form this parser keeps out of the tree.
|
||||
if op := parseOperand(grp, i == len(grps)-1); op != nil {
|
||||
instr.Operands = append(instr.Operands, op)
|
||||
}
|
||||
}
|
||||
@@ -378,8 +422,10 @@ func setName(raw string, sym *ast.Symbol) {
|
||||
|
||||
// --- operand parsing --------------------------------------------------------
|
||||
|
||||
// parseOperand parses one operand group into an Operand.
|
||||
func parseOperand(g []token.Token) *ast.Operand {
|
||||
// parseOperand parses one operand group into an Operand. allowBare marks
|
||||
// the final operand slot of an instruction, where the toolchain reads a
|
||||
// bare constant expression as an immediate.
|
||||
func parseOperand(g []token.Token, allowBare bool) *ast.Operand {
|
||||
g = stripComment(g)
|
||||
if len(g) == 0 {
|
||||
return nil
|
||||
@@ -392,9 +438,25 @@ func parseOperand(g []token.Token) *ast.Operand {
|
||||
}
|
||||
op.Kind = ast.OpAddr
|
||||
op.Addr = parseAddress(g)
|
||||
// A trailing bare constant leaves every address field empty: the
|
||||
// grammar sees no register, memory reference or symbol, and the closed
|
||||
// constant expression is the whole group. Read it as the immediate it
|
||||
// names, exactly what the $ spelling would produce.
|
||||
if allowBare && isEmptyAddress(op.Addr) {
|
||||
if v, rest, ok := foldExpr(g); ok && len(rest) == 0 {
|
||||
op.Kind = ast.OpImmediate
|
||||
op.Imm = ast.Immediate{Val: v, HasVal: true}
|
||||
}
|
||||
}
|
||||
return op
|
||||
}
|
||||
|
||||
// isEmptyAddress reports whether parseAddress populated nothing, its sign
|
||||
// that the group is no register, memory reference, symbol or register range.
|
||||
func isEmptyAddress(a ast.Address) bool {
|
||||
return a.Sym == nil && a.Base == "" && a.Index == "" && a.Range == nil && a.Shift == ""
|
||||
}
|
||||
|
||||
// parseImmediate parses the tokens following a '$'.
|
||||
func parseImmediate(g []token.Token) ast.Immediate {
|
||||
var imm ast.Immediate
|
||||
@@ -408,6 +470,18 @@ func parseImmediate(g []token.Token) ast.Immediate {
|
||||
return imm
|
||||
}
|
||||
}
|
||||
// A constant expression introduced by '(' or '~'. Textual macro
|
||||
// substitution leaves arithmetic such as $(32-shift) and $~63 behind,
|
||||
// and the toolchain evaluates it in place; only shapes the ordinary
|
||||
// paths below cannot read reach the folder, so every existing form
|
||||
// keeps its exact parse.
|
||||
if g[0].Kind == token.LParen || g[0].Kind == token.Tilde {
|
||||
if v, rest, ok := foldExpr(g); ok && len(rest) == 0 {
|
||||
imm.Val = v
|
||||
imm.HasVal = true
|
||||
return imm
|
||||
}
|
||||
}
|
||||
i := 0
|
||||
if g[i].Kind == token.Minus {
|
||||
imm.Neg = true
|
||||
@@ -415,6 +489,17 @@ func parseImmediate(g []token.Token) ast.Immediate {
|
||||
} else if g[i].Kind == token.Plus {
|
||||
i++
|
||||
}
|
||||
// A constant expression after the sign: $-(R - 8), $+(32-shift). The
|
||||
// toolchain folds the negated value in place (the cgo ABI macros write
|
||||
// ADJSP $-(REGS_HOST_TO_ABI0_STACK - 8)), so the sign applies to the
|
||||
// folded value exactly as it does to a bare literal.
|
||||
if i < len(g) && (g[i].Kind == token.LParen || g[i].Kind == token.Tilde) {
|
||||
if v, rest, ok := foldExpr(g[i:]); ok && len(rest) == 0 {
|
||||
imm.Val = v
|
||||
imm.HasVal = true
|
||||
return imm
|
||||
}
|
||||
}
|
||||
if i < len(g) && g[i].Kind == token.Number {
|
||||
text := g[i].Text
|
||||
if v, ok := tryInt(text); ok {
|
||||
@@ -446,6 +531,14 @@ func parseAddress(g []token.Token) ast.Address {
|
||||
if len(g) == 0 {
|
||||
return addr
|
||||
}
|
||||
// A bracketed register range, [Z0-Z3]: the amd64 4FMAPS/4VNNIW
|
||||
// multi-source operand. The bracket runes arrive as Illegal tokens
|
||||
// (the lexer has no bracket kind), so the shape matches on their text.
|
||||
if isBracket(g[0], "[") && len(g) == 5 && g[1].Kind == token.Ident &&
|
||||
g[2].Kind == token.Minus && g[3].Kind == token.Ident && isBracket(g[4], "]") {
|
||||
addr.Range = &ast.RegRange{Lo: g[1].Text, Hi: g[3].Text, Pos: g[0].Pos}
|
||||
return addr
|
||||
}
|
||||
// Symbol-with-pseudo form: name[<>][+off](PSEUDO).
|
||||
// When the prefix is not a valid symbol name (e.g. a bare number like
|
||||
// 0(SP) in RISC-V), sym is nil, and we fall through to regular memory
|
||||
@@ -459,6 +552,31 @@ func parseAddress(g []token.Token) ast.Address {
|
||||
}
|
||||
|
||||
i := 0
|
||||
// A parenthesised constant expression as the displacement: substituted
|
||||
// macro bodies carry ((index*4)+0)(base) shapes. As with the signed
|
||||
// number path below, the value is committed only when a base group
|
||||
// follows.
|
||||
if i < len(g) && g[i].Kind == token.LParen {
|
||||
if v, rest, ok := foldExpr(g[i:]); ok && len(rest) > 0 && rest[0].Kind == token.LParen {
|
||||
addr.Offset = v
|
||||
addr.HasOff = true
|
||||
i = len(g) - len(rest)
|
||||
}
|
||||
}
|
||||
// The same expression under a leading sign: -(24+8)(X6) puts the sign
|
||||
// outside the fold. The base group must follow for the value to
|
||||
// commit, exactly as in the unsigned branch above.
|
||||
if i < len(g) && (g[i].Kind == token.Minus || g[i].Kind == token.Plus) &&
|
||||
i+1 < len(g) && g[i+1].Kind == token.LParen {
|
||||
if v, rest, ok := foldExpr(g[i+1:]); ok && len(rest) > 0 && rest[0].Kind == token.LParen {
|
||||
if g[i].Kind == token.Minus {
|
||||
v = -v
|
||||
}
|
||||
addr.Offset = v
|
||||
addr.HasOff = true
|
||||
i = len(g) - len(rest)
|
||||
}
|
||||
}
|
||||
// Optional leading displacement before a '(' base group. A sign pushes
|
||||
// the parenthesis one token further out: -4(DX) has it at i+2.
|
||||
if isSignedNumber(g, i) {
|
||||
@@ -515,6 +633,16 @@ func parseAddress(g []token.Token) ast.Address {
|
||||
}
|
||||
}
|
||||
}
|
||||
// A lone (index*scale) group is the VSIB index-only form: the
|
||||
// gather/scatter families address memory through a scaled vector index
|
||||
// with no base register, 8(X4*1). The two-group grammar below reads
|
||||
// (base)(index*scale), so a first group whose member carries a scale
|
||||
// factor can only be an index.
|
||||
if isIndexGroup(g[i:]) {
|
||||
addr.Index = g[i+1].Text
|
||||
addr.Scale = int(parseInt(g[i+3].Text))
|
||||
i += 5
|
||||
}
|
||||
// First parenthesised group: the base register.
|
||||
if i < len(g) && g[i].Kind == token.LParen {
|
||||
i++
|
||||
@@ -556,6 +684,26 @@ func parseAddress(g []token.Token) ast.Address {
|
||||
if i > 0 && i < len(g) {
|
||||
addr.Shift = joinRaw(g[i:])
|
||||
}
|
||||
// A lone (possibly signed) number is an absolute address: MOVL $0xf1,
|
||||
// 0xf1 stores through the bare displacement with no base at all. In
|
||||
// operand position a number without $ is an address, never a value.
|
||||
if addr.Sym == nil && addr.Base == "" && addr.Index == "" && !addr.HasOff {
|
||||
neg := false
|
||||
j := 0
|
||||
if j < len(g) && (g[j].Kind == token.Minus || g[j].Kind == token.Plus) {
|
||||
neg = g[j].Kind == token.Minus
|
||||
j++
|
||||
}
|
||||
if j == len(g)-1 && g[j].Kind == token.Number {
|
||||
v := parseInt(g[j].Text)
|
||||
if neg {
|
||||
v = -v
|
||||
}
|
||||
addr.Offset = v
|
||||
addr.HasOff = true
|
||||
return addr
|
||||
}
|
||||
}
|
||||
return addr
|
||||
}
|
||||
|
||||
@@ -571,6 +719,19 @@ func findPseudoParen(g []token.Token) int {
|
||||
return -1
|
||||
}
|
||||
|
||||
// isBracket reports whether t is a square bracket. The lexer has no bracket
|
||||
// kind, so '[' and ']' arrive as Illegal tokens.
|
||||
func isBracket(t token.Token, text string) bool {
|
||||
return t.Kind == token.Illegal && t.Text == text
|
||||
}
|
||||
|
||||
// isIndexGroup reports whether g begins with a complete (index*scale) group:
|
||||
// one identifier followed by a scale factor, all inside a single parenthesis.
|
||||
func isIndexGroup(g []token.Token) bool {
|
||||
return len(g) >= 5 && g[0].Kind == token.LParen && g[1].Kind == token.Ident &&
|
||||
g[2].Kind == token.Star && g[3].Kind == token.Number && g[4].Kind == token.RParen
|
||||
}
|
||||
|
||||
// --- token helpers ----------------------------------------------------------
|
||||
|
||||
// splitOperands splits a token slice on top-level commas (commas outside any
|
||||
|
||||
@@ -401,3 +401,260 @@ func TestInt64MinimumImmediate(t *testing.T) {
|
||||
t.Errorf("imm.Float = %q, want empty", imm.Float)
|
||||
}
|
||||
}
|
||||
|
||||
// TestDivisionSlashPackagePath covers the runtime's package-path spelling:
|
||||
// U+2215 DIVISION SLASH separates the elements of an import path inside a
|
||||
// symbol (internal∕runtime∕atomic·Xchg), and the middle dot still separates
|
||||
// the package from the name. The whole spelling must reach the symbol, not
|
||||
// stop at the first slash.
|
||||
func TestDivisionSlashPackagePath(t *testing.T) {
|
||||
file, errs := Parse("t.s", "TEXT \u00b7f(SB), $0\n\tCALL internal∕runtime∕atomic·Xchg(SB)\n\tRET\n")
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse errors: %v", errs)
|
||||
}
|
||||
txt := file.Decls[0].(*ast.Text)
|
||||
instr := txt.Body[0].(*ast.Instr)
|
||||
sym := instr.Operands[0].Addr.Sym
|
||||
if sym == nil {
|
||||
t.Fatal("operand carries no symbol")
|
||||
}
|
||||
if sym.Pkg != "internal∕runtime∕atomic" {
|
||||
t.Errorf("pkg = %q, want internal∕runtime∕atomic", sym.Pkg)
|
||||
}
|
||||
if sym.Name != "Xchg" {
|
||||
t.Errorf("name = %q, want Xchg", sym.Name)
|
||||
}
|
||||
if sym.Raw != "internal∕runtime∕atomic·Xchg(SB)" {
|
||||
t.Errorf("raw = %q", sym.Raw)
|
||||
}
|
||||
}
|
||||
|
||||
// TestSemicolonStatements covers the plain parse path: ';' separates
|
||||
// statements on one line exactly as it does inside macro expansion, and a
|
||||
// ';' inside a comment is comment text.
|
||||
func TestSemicolonStatements(t *testing.T) {
|
||||
file, errs := Parse("t.s", "TEXT \u00b7f(SB), $0\n\tROLQ $3, DI; ROLQ $13, DI\n\tMOVQ AX, BX // note; still comment\n\tRET\n")
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse errors: %v", errs)
|
||||
}
|
||||
txt := file.Decls[0].(*ast.Text)
|
||||
if len(txt.Body) != 4 {
|
||||
t.Fatalf("body = %d statements, want 4", len(txt.Body))
|
||||
}
|
||||
first := txt.Body[0].(*ast.Instr)
|
||||
if first.Mnemonic.Text != "ROLQ" || len(first.Operands) != 2 {
|
||||
t.Errorf("first statement = %+v, want ROLQ with two operands", first.Mnemonic)
|
||||
}
|
||||
second := txt.Body[1].(*ast.Instr)
|
||||
if second.Mnemonic.Text != "ROLQ" || len(second.Operands) != 2 {
|
||||
t.Errorf("second statement = %s, want ROLQ with two operands", second.Mnemonic.Text)
|
||||
}
|
||||
// The trailing comment belongs to the second MOVQ, semicolon included.
|
||||
third := txt.Body[2].(*ast.Instr)
|
||||
if third.Mnemonic.Text != "MOVQ" || third.Comment != "note; still comment" {
|
||||
t.Errorf("third = %s, comment %q", third.Mnemonic.Text, third.Comment)
|
||||
}
|
||||
}
|
||||
|
||||
// TestSemicolonAfterLabel covers a label sharing its line with two
|
||||
// statements.
|
||||
func TestSemicolonAfterLabel(t *testing.T) {
|
||||
file, errs := Parse("t.s", "TEXT \u00b7f(SB), $0\nloop: NOP; NOP\n\tRET\n")
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse errors: %v", errs)
|
||||
}
|
||||
txt := file.Decls[0].(*ast.Text)
|
||||
if len(txt.Body) != 4 {
|
||||
t.Fatalf("body = %d statements, want 4 (label, two instructions, RET)", len(txt.Body))
|
||||
}
|
||||
if _, ok := txt.Body[0].(*ast.Label); !ok {
|
||||
t.Errorf("first statement = %T, want *ast.Label", txt.Body[0])
|
||||
}
|
||||
for i, want := range []string{"NOP", "NOP", "RET"} {
|
||||
in, ok := txt.Body[i+1].(*ast.Instr)
|
||||
if !ok || in.Mnemonic.Text != want {
|
||||
t.Errorf("statement %d = %v, want %s", i+1, txt.Body[i+1], want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestParseEqualsZeroOptions pins the contract that ParseWithOptions with
|
||||
// the zero Options reproduces Parse, here for the semicolon split.
|
||||
func TestParseEqualsZeroOptions(t *testing.T) {
|
||||
src := "TEXT \u00b7f(SB), $0\n\tNOP; NOP\n\tRET\n"
|
||||
a, errsA := Parse("t.s", src)
|
||||
b, errsB := ParseWithOptions("t.s", src, Options{})
|
||||
if len(errsA) > 0 || len(errsB) > 0 {
|
||||
t.Fatalf("errors: %v / %v", errsA, errsB)
|
||||
}
|
||||
ta, tb := texts(a), texts(b)
|
||||
if len(ta) != len(tb) {
|
||||
t.Fatalf("decl counts differ: %d vs %d", len(ta), len(tb))
|
||||
}
|
||||
for i := range ta {
|
||||
if len(ta[i].Body) != len(tb[i].Body) {
|
||||
t.Fatalf("TEXT %d: body lengths differ: %d vs %d", i, len(ta[i].Body), len(tb[i].Body))
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestBracketRegisterRange pins the amd64 multi-source operand of the
|
||||
// 4FMAPS/4VNNIW families: the bracket group [Z0-Z3] names four consecutive
|
||||
// source registers and must reach the AST as a register range instead of an
|
||||
// empty address.
|
||||
func TestBracketRegisterRange(t *testing.T) {
|
||||
file, errs := Parse("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tV4FMADDPS 17(SP), [Z0-Z3], K2, Z0\n\tRET\n")
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse errors: %v", errs)
|
||||
}
|
||||
fn := file.Decls[0].(*ast.Text)
|
||||
in := fn.Body[0].(*ast.Instr)
|
||||
if len(in.Operands) != 4 {
|
||||
t.Fatalf("operands = %d, want 4", len(in.Operands))
|
||||
}
|
||||
rng := in.Operands[1]
|
||||
if rng.Kind != ast.OpAddr {
|
||||
t.Errorf("range operand kind = %v, want OpAddr", rng.Kind)
|
||||
}
|
||||
if rng.Addr.Range == nil {
|
||||
t.Fatalf("range operand = %+v, want a register range", rng.Addr)
|
||||
}
|
||||
if rng.Addr.Range.Lo != "Z0" || rng.Addr.Range.Hi != "Z3" {
|
||||
t.Errorf("range = %s-%s, want Z0-Z3", rng.Addr.Range.Lo, rng.Addr.Range.Hi)
|
||||
}
|
||||
if rng.Addr.Sym != nil || rng.Addr.Base != "" || rng.Addr.Index != "" || rng.Addr.Shift != "" {
|
||||
t.Errorf("range operand carries stray address fields: %+v", rng.Addr)
|
||||
}
|
||||
if rng.Raw != "[ Z0 - Z3 ]" {
|
||||
t.Errorf("range raw = %q, want the verbatim spelling", rng.Raw)
|
||||
}
|
||||
}
|
||||
|
||||
// TestBracketRegisterRangeNotList pins that arm64-style register lists, whose
|
||||
// members carry arrangements, stay out of the simple range shape: they remain
|
||||
// plain bracketed groups the arm64 encoder reads from Raw. A comma inside
|
||||
// brackets is a top-level comma, so a multi-member list spans several
|
||||
// operands, exactly the shape the arm64 encoder's list scan stitches back.
|
||||
func TestBracketRegisterRangeNotList(t *testing.T) {
|
||||
file, errs := Parse("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tVLD1 (R2), [V21.B16]\n\tVLD1 (R1), [V2.B16, V3.B16]\n\tRET\n")
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse errors: %v", errs)
|
||||
}
|
||||
fn := file.Decls[0].(*ast.Text)
|
||||
for i, want := range []string{"[ V21.B16 ]", "V3.B16 ]"} {
|
||||
in := fn.Body[i].(*ast.Instr)
|
||||
op := in.Operands[len(in.Operands)-1]
|
||||
if op.Addr.Range != nil {
|
||||
t.Errorf("%s: range = %v, want nil", in.Mnemonic.Text, op.Addr.Range)
|
||||
}
|
||||
if op.Raw != want {
|
||||
t.Errorf("operand %d raw = %q, want %q", i, op.Raw, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestVSIBIndexOnly pins the gather/scatter memory operand with a scaled
|
||||
// vector index and no base register: 8(X4*1) must carry index and scale and
|
||||
// leave the base empty, not strand the scale in the shift suffix.
|
||||
func TestVSIBIndexOnly(t *testing.T) {
|
||||
file, errs := Parse("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tVPGATHERDQ Y0, 8(X4*1), Y6\n\tVPGATHERDQ Y0, (X4*2), Y6\n\tVPGATHERDQ Y0, -8(X4*1), Y6\n\tRET\n")
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse errors: %v", errs)
|
||||
}
|
||||
fn := file.Decls[0].(*ast.Text)
|
||||
want := []ast.Address{
|
||||
{Index: "X4", Scale: 1, Offset: 8, HasOff: true},
|
||||
{Index: "X4", Scale: 2},
|
||||
{Index: "X4", Scale: 1, Offset: -8, HasOff: true},
|
||||
}
|
||||
for i, w := range want {
|
||||
in := fn.Body[i].(*ast.Instr)
|
||||
a := in.Operands[1].Addr
|
||||
if a.Base != "" || a.Index != w.Index || a.Scale != w.Scale || a.Offset != w.Offset || a.HasOff != w.HasOff || a.Shift != "" {
|
||||
t.Errorf("operand %d = %+v, want %+v", i, a, w)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestVSIBTwoGroupKeepsBase pins that the ordinary (base)(index*scale)
|
||||
// grammar is untouched by the index-only recognition.
|
||||
func TestVSIBTwoGroupKeepsBase(t *testing.T) {
|
||||
file, errs := Parse("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tVP4DPWSSD 7(SI)(DI*1), [Z2-Z5], K4, Z17\n\tRET\n")
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse errors: %v", errs)
|
||||
}
|
||||
fn := file.Decls[0].(*ast.Text)
|
||||
in := fn.Body[0].(*ast.Instr)
|
||||
a := in.Operands[0].Addr
|
||||
if a.Base != "SI" || a.Index != "DI" || a.Scale != 1 || a.Offset != 7 || !a.HasOff {
|
||||
t.Errorf("address = %+v, want base SI index DI scale 1 offset 7", a)
|
||||
}
|
||||
if in.Operands[1].Addr.Range == nil || in.Operands[1].Addr.Range.Lo != "Z2" || in.Operands[1].Addr.Range.Hi != "Z5" {
|
||||
t.Errorf("second operand = %+v, want range Z2-Z5", in.Operands[1].Addr)
|
||||
}
|
||||
}
|
||||
|
||||
// TestBareTrailingImmediate pins the toolchain's bare constant spelling in
|
||||
// the final operand slot: CMPSD X1, X0, 1 reads as $1 (math/floor_amd64.s).
|
||||
// Earlier slots keep the strict grammar, so a bare number there stays an
|
||||
// address rather than becoming an immediate.
|
||||
func TestBareTrailingImmediate(t *testing.T) {
|
||||
file, errs := Parse("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tCMPSD X1, X0, 1\n\tCMPSD X1, X0, -1\n\tADDQ AX, 1+2\n\tRET\n")
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse errors: %v", errs)
|
||||
}
|
||||
fn := file.Decls[0].(*ast.Text)
|
||||
for i, want := range []int64{1, -1, 3} {
|
||||
in := fn.Body[i].(*ast.Instr)
|
||||
last := in.Operands[len(in.Operands)-1]
|
||||
if last.Kind != ast.OpImmediate || !last.Imm.HasVal || last.Imm.Val != want {
|
||||
t.Errorf("operand %d = %+v, want immediate %d", i, last, want)
|
||||
}
|
||||
}
|
||||
|
||||
// A bare number outside the final slot is not an immediate.
|
||||
file2, errs2 := Parse("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tADDQ 1, AX\n\tRET\n")
|
||||
if len(errs2) > 0 {
|
||||
t.Fatalf("parse errors: %v", errs2)
|
||||
}
|
||||
fn2 := file2.Decls[0].(*ast.Text)
|
||||
first := fn2.Body[0].(*ast.Instr).Operands[0]
|
||||
if first.Kind != ast.OpAddr {
|
||||
t.Errorf("non-final bare number kind = %v, want OpAddr", first.Kind)
|
||||
}
|
||||
// A bare name in the final slot stays a symbol: labels are names, not
|
||||
// constants, and jump targets depend on the distinction.
|
||||
file3, errs3 := Parse("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tJMP loop\nloop: NOP\n\tRET\n")
|
||||
if len(errs3) > 0 {
|
||||
t.Fatalf("parse errors: %v", errs3)
|
||||
}
|
||||
fn3 := file3.Decls[0].(*ast.Text)
|
||||
jmp := fn3.Body[0].(*ast.Instr)
|
||||
if jmp.Operands[0].Kind != ast.OpAddr || jmp.Operands[0].Addr.Sym == nil || jmp.Operands[0].Addr.Sym.Name != "loop" {
|
||||
t.Errorf("jump target = %+v, want label loop", jmp.Operands[0])
|
||||
}
|
||||
}
|
||||
|
||||
// TestSignedParenDisplacement pins a sign before a parenthesised
|
||||
// displacement expression: -(24+8)(X6) negates the folded value and keeps
|
||||
// the base group, the shape GOROOT's riscv64 and loong64 files use.
|
||||
func TestSignedParenDisplacement(t *testing.T) {
|
||||
file, errs := Parse("t.s", "TEXT \u00b7f(SB), NOSPLIT, $0-0\n\tMOV X7, -(24+8)(X6)\n\tMOV X7, +(16)(X6)\n\tRET\n")
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
text := file.Decls[0].(*ast.Text)
|
||||
ins := text.Body[0].(*ast.Instr)
|
||||
op := ins.Operands[1] // Plan 9 order: the destination address is last
|
||||
if !op.Addr.HasOff || op.Addr.Offset != -32 {
|
||||
t.Errorf("-(24+8): offset = %v hasOff=%v, want -32 true", op.Addr.Offset, op.Addr.HasOff)
|
||||
}
|
||||
if op.Addr.Base != "X6" {
|
||||
t.Errorf("-(24+8): base = %q, want X6", op.Addr.Base)
|
||||
}
|
||||
ins = text.Body[1].(*ast.Instr)
|
||||
op = ins.Operands[1]
|
||||
if !op.Addr.HasOff || op.Addr.Offset != 16 || op.Addr.Base != "X6" {
|
||||
t.Errorf("+(16): offset = %v hasOff=%v base=%q, want 16 true X6", op.Addr.Offset, op.Addr.HasOff, op.Addr.Base)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,571 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
// The preprocessor turns #define and #include directives into the token
|
||||
// stream the parser really sees, the way the Go toolchain's assembler does:
|
||||
// object and parameterised macros expand at the point of use, and an
|
||||
// #include splices the named file's lines in place of the directive. The
|
||||
// pass runs only on the assembly path (gasm asm, diff, the corpus audit),
|
||||
// where the result is machine code; parsing for the linter, formatter and
|
||||
// language server keeps the raw file so their view of #define lines, and
|
||||
// therefore their macro-aware behaviour, is unchanged.
|
||||
package parser
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"slices"
|
||||
"strconv"
|
||||
"strings"
|
||||
"unicode/utf8"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/lexer"
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/token"
|
||||
)
|
||||
|
||||
// Options controls the optional preprocessing applied before a file is
|
||||
// parsed. The zero value reproduces Parse exactly.
|
||||
type Options struct {
|
||||
// IncludeDirs lists the -I directories searched for #include files,
|
||||
// in order, after the including file's own directory.
|
||||
IncludeDirs []string
|
||||
// Expand enables macro expansion, include splicing and the
|
||||
// statement-separator reading of ';' that the expanded bodies rely on.
|
||||
Expand bool
|
||||
// Predefines names the macros defined before the file is read. The
|
||||
// go command drives go tool asm with -D GOOS_<goos> -D GOARCH_<arch>,
|
||||
// and GOROOT's own headers (go_tls.h, asm_riscv64.h) select their
|
||||
// platform blocks with #ifdef on exactly those names, so an assembler
|
||||
// without them cannot see the platform definitions at all.
|
||||
Predefines map[string]string
|
||||
}
|
||||
|
||||
// ParseWithOptions parses src like Parse, optionally preprocessing it first.
|
||||
// The returned file is usable even when errors is non-empty.
|
||||
func ParseWithOptions(path, src string, opts Options) (*ast.File, []error) {
|
||||
tokens := lexer.Tokenize(src)
|
||||
var lines [][]token.Token
|
||||
var errs []error
|
||||
if opts.Expand {
|
||||
pp := &preproc{opts: opts, macros: map[string]*macroDef{}}
|
||||
for name, value := range opts.Predefines {
|
||||
pp.macros[name] = ¯oDef{name: name, body: lexer.Tokenize(value)}
|
||||
}
|
||||
lines = pp.fileLines(path, tokens, token.Position{})
|
||||
errs = pp.errs
|
||||
} else {
|
||||
lines = statementLines(tokens)
|
||||
}
|
||||
p := &state{path: path}
|
||||
p.parse(lines)
|
||||
return p.file, append(errs, p.errs...)
|
||||
}
|
||||
|
||||
// maxExpansionDepth bounds recursive macro expansion; the toolchain's
|
||||
// assembler gives up after 100 nested invocations without producing a token.
|
||||
const maxExpansionDepth = 100
|
||||
|
||||
// textflagHeader names the one header gasm does not splice: its flag macros
|
||||
// (NOSPLIT, RODATA, …) are consumed by name throughout gasm's parser,
|
||||
// encoders and linter, and expanding them to their numeric constants would
|
||||
// leave every consumer blind to them.
|
||||
const textflagHeader = "textflag.h"
|
||||
|
||||
// macroDef is one #define. A nil args slice is an object macro; a non-nil
|
||||
// (possibly empty) one is parameterised, the C distinction between
|
||||
// "#define A(x)" and "#define A (x)".
|
||||
type macroDef struct {
|
||||
name string
|
||||
args []string
|
||||
body []token.Token
|
||||
}
|
||||
|
||||
// preproc carries the state of one expansion pass: the live macro table, the
|
||||
// chain of files currently being read, for cycle detection, and the
|
||||
// conditional-inclusion stack of #ifdef regions.
|
||||
type preproc struct {
|
||||
opts Options
|
||||
macros map[string]*macroDef
|
||||
errs []error
|
||||
stack []string // absolute paths of files being read, innermost last
|
||||
ifdefStack []bool // one entry per open #ifdef/#ifndef, its truth
|
||||
}
|
||||
|
||||
// enabled reports whether the position being read is inside a live
|
||||
// conditional branch. Directives inside a disabled branch contribute
|
||||
// nothing, and its content lines are dropped, exactly as the toolchain's
|
||||
// input stack does.
|
||||
func (pp *preproc) enabled() bool {
|
||||
return len(pp.ifdefStack) == 0 || pp.ifdefStack[len(pp.ifdefStack)-1]
|
||||
}
|
||||
|
||||
func (pp *preproc) errorf(pos token.Position, format string, args ...any) {
|
||||
pp.errs = append(pp.errs, Error{Pos: pos, Msg: fmt.Sprintf(format, args...)})
|
||||
}
|
||||
|
||||
// fileLines tokenizes and preprocesses one file into logical lines.
|
||||
// Directive lines are kept (the parser records them for the tooling);
|
||||
// #include lines are replaced by the included file's lines. includePos is
|
||||
// the position of the #include that pulled this file in, zero for the
|
||||
// top-level file, and only serves cycle diagnostics.
|
||||
func (pp *preproc) fileLines(path string, tokens []token.Token, includePos token.Position) [][]token.Token {
|
||||
abs, err := filepath.Abs(path)
|
||||
if err != nil {
|
||||
abs = filepath.Clean(path)
|
||||
}
|
||||
if slices.Contains(pp.stack, abs) {
|
||||
if includePos.IsValid() {
|
||||
pp.errorf(includePos, "#include %q: include cycle (%s is already being read)", path, filepath.Base(path))
|
||||
}
|
||||
return nil
|
||||
}
|
||||
pp.stack = append(pp.stack, abs)
|
||||
|
||||
var out [][]token.Token
|
||||
for _, line := range splitLines(tokens) {
|
||||
if len(line) == 0 {
|
||||
out = append(out, line)
|
||||
continue
|
||||
}
|
||||
if line[0].Kind == token.Hash {
|
||||
out = append(out, pp.directive(line, filepath.Dir(path))...)
|
||||
continue
|
||||
}
|
||||
if !pp.enabled() {
|
||||
continue
|
||||
}
|
||||
out = append(out, splitOnSemicolons(pp.expandTokens(line))...)
|
||||
}
|
||||
pp.stack = pp.stack[:len(pp.stack)-1]
|
||||
if len(pp.stack) == 0 && len(pp.ifdefStack) > 0 {
|
||||
// The stack is per-input, shared across includes, so only the
|
||||
// top-level file's end can decide the input was left unclosed.
|
||||
pp.errorf(token.Position{Line: 1, Column: 1}, "unclosed #ifdef or #ifndef")
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// directive processes one '#' line and returns the lines to keep in the
|
||||
// stream: every directive line is kept as-is for the parser (which records
|
||||
// it), except #include, which is replaced by the spliced content.
|
||||
// Conditionals are tracked on every line; every other directive is inert
|
||||
// inside a disabled branch.
|
||||
func (pp *preproc) directive(line []token.Token, dir string) [][]token.Token {
|
||||
if len(line) < 2 || line[1].Kind != token.Ident {
|
||||
return [][]token.Token{line}
|
||||
}
|
||||
switch line[1].Text {
|
||||
case "ifdef", "ifndef":
|
||||
pp.ifdef(line, line[1].Text == "ifndef")
|
||||
case "else":
|
||||
pp.elseBranch(line)
|
||||
case "endif":
|
||||
pp.endif(line)
|
||||
case "define":
|
||||
if pp.enabled() {
|
||||
pp.define(line)
|
||||
}
|
||||
case "undef":
|
||||
if pp.enabled() {
|
||||
pp.undef(line)
|
||||
}
|
||||
case "include":
|
||||
if pp.enabled() {
|
||||
return pp.include(line, dir)
|
||||
}
|
||||
default:
|
||||
// #line and unknown directives are recorded but not interpreted:
|
||||
// conservative support keeps the parser's view intact and files
|
||||
// using them fail on their content, not silently.
|
||||
}
|
||||
return [][]token.Token{line}
|
||||
}
|
||||
|
||||
// ifdef handles "#ifdef NAME" and "#ifndef NAME", pushing the branch's truth
|
||||
// onto the conditional stack. A branch opened inside a disabled region is
|
||||
// itself disabled, however the name resolves.
|
||||
func (pp *preproc) ifdef(line []token.Token, inverted bool) {
|
||||
truth := false
|
||||
if len(line) >= 3 && line[2].Kind == token.Ident {
|
||||
_, defined := pp.macros[line[2].Text]
|
||||
truth = defined != inverted
|
||||
} else {
|
||||
pp.errorf(line[0].Pos, "expected identifier after #%s", line[1].Text)
|
||||
}
|
||||
if !pp.enabled() {
|
||||
truth = false
|
||||
}
|
||||
pp.ifdefStack = append(pp.ifdefStack, truth)
|
||||
}
|
||||
|
||||
// elseBranch flips the innermost conditional's truth, but only when the
|
||||
// region enclosing it is itself live: the toolchain keeps outer overrides.
|
||||
func (pp *preproc) elseBranch(line []token.Token) {
|
||||
if len(pp.ifdefStack) == 0 {
|
||||
pp.errorf(line[0].Pos, "unmatched #else")
|
||||
return
|
||||
}
|
||||
if len(pp.ifdefStack) == 1 || pp.ifdefStack[len(pp.ifdefStack)-2] {
|
||||
pp.ifdefStack[len(pp.ifdefStack)-1] = !pp.ifdefStack[len(pp.ifdefStack)-1]
|
||||
}
|
||||
}
|
||||
|
||||
// endif closes the innermost conditional.
|
||||
func (pp *preproc) endif(line []token.Token) {
|
||||
if len(pp.ifdefStack) == 0 {
|
||||
pp.errorf(line[0].Pos, "unmatched #endif")
|
||||
return
|
||||
}
|
||||
pp.ifdefStack = pp.ifdefStack[:len(pp.ifdefStack)-1]
|
||||
}
|
||||
|
||||
// define parses "#define NAME[(formals)] body" into the macro table. The
|
||||
// body runs to the end of the logical line (the lexer has already spliced
|
||||
// backslash continuations) and stops at a comment, which never expands.
|
||||
func (pp *preproc) define(line []token.Token) {
|
||||
if len(line) < 3 || line[2].Kind != token.Ident {
|
||||
return
|
||||
}
|
||||
name := line[2]
|
||||
args := []string(nil)
|
||||
body := line[3:]
|
||||
// The definition is parameterised only when '(' follows the name
|
||||
// directly; the toolchain separates "#define A(x)" from
|
||||
// "#define A (x)" by adjacency, and so does the column check here.
|
||||
if len(body) > 0 && body[0].Kind == token.LParen &&
|
||||
body[0].Pos.Column == name.Pos.Column+utf8.RuneCountInString(name.Text) {
|
||||
args = []string{}
|
||||
i := 1
|
||||
for i < len(body) && body[i].Kind != token.RParen {
|
||||
if body[i].Kind == token.Ident {
|
||||
args = append(args, body[i].Text)
|
||||
}
|
||||
i++
|
||||
}
|
||||
if i < len(body) {
|
||||
body = body[i+1:]
|
||||
} else {
|
||||
body = nil
|
||||
}
|
||||
}
|
||||
if i := slices.IndexFunc(body, func(t token.Token) bool { return t.Kind == token.Comment }); i >= 0 {
|
||||
body = body[:i]
|
||||
}
|
||||
if _, exists := pp.macros[name.Text]; exists {
|
||||
// The toolchain refuses redefinition, so a file the oracle accepts
|
||||
// never redefines; failing here keeps that contract visible.
|
||||
pp.errorf(name.Pos, "redefinition of macro %s", name.Text)
|
||||
}
|
||||
pp.macros[name.Text] = ¯oDef{name: name.Text, args: args, body: pp.bodyWithBreaks(body)}
|
||||
|
||||
}
|
||||
|
||||
// bodyWithBreaks records the statement boundaries the continuations carry.
|
||||
// The lexer splices backslash-continued lines into one logical line, but the
|
||||
// toolchain keeps the newline as a token in the stored body, which is how a
|
||||
// multi-instruction body without semicolons (the arm64 style) still splits
|
||||
// into statements on expansion. A line change inside the logical line is
|
||||
// exactly a continuation, so the boundary is restored from the positions.
|
||||
func (pp *preproc) bodyWithBreaks(body []token.Token) []token.Token {
|
||||
out := make([]token.Token, 0, len(body))
|
||||
for i, t := range body {
|
||||
if i > 0 && t.Pos.Line != body[i-1].Pos.Line {
|
||||
out = append(out, token.Token{Kind: token.Newline, Text: "\n", Pos: t.Pos, End: t.Pos})
|
||||
}
|
||||
out = append(out, t)
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// undef handles "#undef NAME", which the toolchain honours and requires to
|
||||
// name a defined macro.
|
||||
func (pp *preproc) undef(line []token.Token) {
|
||||
if len(line) < 3 || line[2].Kind != token.Ident {
|
||||
return
|
||||
}
|
||||
if _, ok := pp.macros[line[2].Text]; !ok {
|
||||
pp.errorf(line[2].Pos, "#undef for undefined macro %s", line[2].Text)
|
||||
return
|
||||
}
|
||||
delete(pp.macros, line[2].Text)
|
||||
}
|
||||
|
||||
// include resolves and splices "#include \"file\"". A header that cannot be
|
||||
// read keeps the directive line in the stream, with a diagnostic.
|
||||
func (pp *preproc) include(line []token.Token, dir string) [][]token.Token {
|
||||
if len(line) < 3 || line[2].Kind != token.String {
|
||||
return [][]token.Token{line}
|
||||
}
|
||||
header := line[2]
|
||||
name, err := strconv.Unquote(header.Text)
|
||||
if err != nil {
|
||||
pp.errorf(header.Pos, "unquoting include file name: %v", err)
|
||||
return [][]token.Token{line}
|
||||
}
|
||||
if filepath.Base(name) == textflagHeader {
|
||||
// Flag macros are handled natively (see textflagHeader); the
|
||||
// directive stays so tools still see the include.
|
||||
return [][]token.Token{line}
|
||||
}
|
||||
resolved, ok := pp.resolve(name, dir)
|
||||
if !ok {
|
||||
searched := append([]string{dir}, pp.opts.IncludeDirs...)
|
||||
pp.errorf(header.Pos, "#include %q: file not found (searched %s)", name, strings.Join(searched, ", "))
|
||||
return [][]token.Token{line}
|
||||
}
|
||||
src, err := os.ReadFile(resolved)
|
||||
if err != nil {
|
||||
pp.errorf(header.Pos, "#include %q: %v", name, err)
|
||||
return [][]token.Token{line}
|
||||
}
|
||||
return pp.fileLines(resolved, lexer.Tokenize(string(src)), header.Pos)
|
||||
}
|
||||
|
||||
// resolve looks an include name up the way the toolchain does: as written
|
||||
// (relative to the working directory), then relative to the including
|
||||
// file's directory, then in each -I directory in order.
|
||||
func (pp *preproc) resolve(name, dir string) (string, bool) {
|
||||
candidates := []string{name}
|
||||
if !filepath.IsAbs(name) {
|
||||
candidates = append(candidates, filepath.Join(dir, name))
|
||||
for _, d := range pp.opts.IncludeDirs {
|
||||
candidates = append(candidates, filepath.Join(d, name))
|
||||
}
|
||||
}
|
||||
for _, c := range candidates {
|
||||
if st, err := os.Stat(c); err == nil && !st.IsDir() {
|
||||
return c, true
|
||||
}
|
||||
}
|
||||
return "", false
|
||||
}
|
||||
|
||||
// expandTokens expands every macro invocation in a token sequence,
|
||||
// recursively, with a depth guard. A body is spliced into the sequence in
|
||||
// place and rescanned, the way the toolchain's input stack re-reads pushed
|
||||
// tokens: an object macro may name a parameterised one, and the argument
|
||||
// list of the expansion may then come from the tokens that follow.
|
||||
func (pp *preproc) expandTokens(in []token.Token) []token.Token {
|
||||
s := in
|
||||
i := 0
|
||||
consecutive := 0
|
||||
for i < len(s) {
|
||||
t := s[i]
|
||||
if t.Kind != token.Ident {
|
||||
i++
|
||||
consecutive = 0
|
||||
continue
|
||||
}
|
||||
def, suffix := pp.macroFor(t.Text)
|
||||
if def == nil {
|
||||
i++
|
||||
consecutive = 0
|
||||
continue
|
||||
}
|
||||
// The guard mirrors the toolchain's: 100 nested invocations in a
|
||||
// row without a plain token between them means recursion.
|
||||
consecutive++
|
||||
if consecutive > maxExpansionDepth {
|
||||
pp.errorf(t.Pos, "recursive macro invocation (deeper than %d levels)", maxExpansionDepth)
|
||||
return nil
|
||||
}
|
||||
if def.args == nil {
|
||||
body := restamp(def.body, t.Pos)
|
||||
if suffix != "" {
|
||||
// The macro was reached only through a compound spelling
|
||||
// (ACC0.B16 over "#define ACC0 V8"), so the selector has
|
||||
// to travel with the expansion.
|
||||
body = appendSelector(body, suffix, t.Pos)
|
||||
}
|
||||
s = append(s[:i], append(body, s[i+1:]...)...)
|
||||
continue
|
||||
}
|
||||
// A parameterised macro invoked without its parentheses stands
|
||||
// unexpanded, naming itself, as in the toolchain.
|
||||
if i+1 >= len(s) || s[i+1].Kind != token.LParen {
|
||||
i++
|
||||
consecutive = 0
|
||||
continue
|
||||
}
|
||||
args, next := pp.collectArgs(s, i+1, t)
|
||||
if args == nil {
|
||||
return nil
|
||||
}
|
||||
// A zero-argument macro may be invoked as NAME().
|
||||
if len(def.args) == 0 && len(args) == 1 && len(args[0]) == 0 {
|
||||
args = nil
|
||||
}
|
||||
if len(args) != len(def.args) {
|
||||
pp.errorf(t.Pos, "wrong arg count for macro %s: got %d, want %d", t.Text, len(args), len(def.args))
|
||||
i = next
|
||||
consecutive = 0
|
||||
continue
|
||||
}
|
||||
sub := make([]token.Token, 0, len(def.body))
|
||||
for _, bt := range def.body {
|
||||
if bt.Kind == token.Ident {
|
||||
if k := slices.Index(def.args, bt.Text); k >= 0 {
|
||||
sub = append(sub, restamp(args[k], t.Pos)...)
|
||||
continue
|
||||
}
|
||||
// A parameter used with an element or lane selector: the
|
||||
// lexer folds A.S4 into one identifier, so the whole-token
|
||||
// match above cannot see the parameter. The toolchain
|
||||
// lexes the period separately and substitutes the name
|
||||
// alone; splitting at the FIRST period and pasting the
|
||||
// argument back in front of the selector is the equivalent
|
||||
// for this lexer.
|
||||
if k, sel := parameterSelector(bt.Text, def.args); k >= 0 {
|
||||
sub = append(sub, restamp(pasteSelector(args[k], sel), t.Pos)...)
|
||||
continue
|
||||
}
|
||||
}
|
||||
sub = append(sub, bt)
|
||||
}
|
||||
s = append(s[:i], append(sub, s[next:]...)...)
|
||||
}
|
||||
return s
|
||||
}
|
||||
|
||||
// macroFor finds the macro a use names. The lexer folds NAME.selector into
|
||||
// one identifier token, so a macro written behind a selector suffix
|
||||
// (ACC0.B16 over "#define ACC0 V8") never matches a whole-token table
|
||||
// lookup; the toolchain splits on the period and reads the two halves, so
|
||||
// the prefix before the FIRST period is tried here as well and the caller
|
||||
// re-attaches the suffix to whatever the macro expands to. Only a whole
|
||||
// name counts: AB.S4 does not reach a macro named A, and a parameterised
|
||||
// macro is not hidden behind a selector, because its invocation would need
|
||||
// the parentheses to follow the bare name.
|
||||
func (pp *preproc) macroFor(text string) (*macroDef, string) {
|
||||
if def := pp.macros[text]; def != nil {
|
||||
return def, ""
|
||||
}
|
||||
if j := strings.IndexByte(text, '.'); j > 0 {
|
||||
if def := pp.macros[text[:j]]; def != nil && def.args == nil {
|
||||
return def, text[j:]
|
||||
}
|
||||
}
|
||||
return nil, ""
|
||||
}
|
||||
|
||||
// appendSelector glues a selector suffix onto an object macro's expansion:
|
||||
// the selector binds to the identifier the expansion ends with, the way the
|
||||
// toolchain's operand parser reads V0 and .B16 back as one register
|
||||
// spelling. An expansion that does not end in an identifier carries the
|
||||
// selector as its own token, which the parser then reports where it cannot
|
||||
// parse it.
|
||||
func appendSelector(body []token.Token, suffix string, pos token.Position) []token.Token {
|
||||
if n := len(body); n > 0 && body[n-1].Kind == token.Ident {
|
||||
body[n-1].Text += suffix
|
||||
return body
|
||||
}
|
||||
return append(body, token.Token{Kind: token.Ident, Text: suffix, Pos: pos, End: pos})
|
||||
}
|
||||
|
||||
// parameterSelector reports the argument a compound body token names: the
|
||||
// parameter whose whole name occupies the text before the token's FIRST
|
||||
// period, with the selector that follows. k is negative when no parameter
|
||||
// matches, which leaves tokens like AB.S4 untouched even though a parameter
|
||||
// A is bound.
|
||||
func parameterSelector(text string, args []string) (int, string) {
|
||||
j := strings.IndexByte(text, '.')
|
||||
if j <= 0 {
|
||||
return -1, ""
|
||||
}
|
||||
if k := slices.Index(args, text[:j]); k >= 0 {
|
||||
return k, text[j:]
|
||||
}
|
||||
return -1, ""
|
||||
}
|
||||
|
||||
// pasteSelector joins an argument with the selector a compound body token
|
||||
// carries, textually: the selector binds to the identifier the argument
|
||||
// ends with, so A.S4 over the argument V0.B16 spells V0.B16.S4, exactly the
|
||||
// operand the toolchain's split-then-substitute leaves behind. An argument
|
||||
// with no trailing identifier carries the selector as a separate token,
|
||||
// which the parser then reports where it cannot parse it.
|
||||
func pasteSelector(val []token.Token, suffix string) []token.Token {
|
||||
if len(val) == 0 {
|
||||
return []token.Token{{Kind: token.Ident, Text: suffix}}
|
||||
}
|
||||
out := slices.Clone(val)
|
||||
if n := len(out); out[n-1].Kind == token.Ident {
|
||||
out[n-1].Text += suffix
|
||||
return out
|
||||
}
|
||||
return append(out, token.Token{Kind: token.Ident, Text: suffix})
|
||||
}
|
||||
|
||||
// collectArgs reads the actual argument tokens of an invocation; the opening
|
||||
// parenthesis is at start. Commas separate arguments except inside nested
|
||||
// parentheses. A nil result means the list was unterminated, which is a
|
||||
// diagnostic.
|
||||
func (pp *preproc) collectArgs(in []token.Token, start int, name token.Token) ([][]token.Token, int) {
|
||||
var args [][]token.Token
|
||||
var cur []token.Token
|
||||
nesting := 0
|
||||
for i := start + 1; i < len(in); i++ {
|
||||
t := in[i]
|
||||
switch t.Kind {
|
||||
case token.LParen:
|
||||
nesting++
|
||||
cur = append(cur, t)
|
||||
case token.RParen:
|
||||
if nesting == 0 {
|
||||
return append(args, cur), i + 1
|
||||
}
|
||||
nesting--
|
||||
cur = append(cur, t)
|
||||
case token.Comma:
|
||||
if nesting == 0 {
|
||||
args = append(args, cur)
|
||||
cur = nil
|
||||
continue
|
||||
}
|
||||
cur = append(cur, t)
|
||||
case token.Comment:
|
||||
pp.errorf(name.Pos, "unterminated arg list invoking macro %s", name.Text)
|
||||
return nil, i
|
||||
default:
|
||||
cur = append(cur, t)
|
||||
}
|
||||
}
|
||||
pp.errorf(name.Pos, "unterminated arg list invoking macro %s", name.Text)
|
||||
return nil, len(in)
|
||||
}
|
||||
|
||||
// restamp copies body tokens to the invocation's position, so diagnostics
|
||||
// and the line table point where the macro was used, as the toolchain's
|
||||
// input stack does.
|
||||
func restamp(body []token.Token, pos token.Position) []token.Token {
|
||||
out := make([]token.Token, len(body))
|
||||
for i, t := range body {
|
||||
t.Pos, t.End = pos, pos
|
||||
out[i] = t
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// splitOnSemicolons breaks a token sequence at ';' statement separators and
|
||||
// at the Newline markers that record continuation boundaries inside macro
|
||||
// bodies, producing the logical lines the parser expects. The separators
|
||||
// carry no meaning beyond the break, so the pieces are exactly what the same
|
||||
// statements on separate lines would produce.
|
||||
func splitOnSemicolons(ts []token.Token) [][]token.Token {
|
||||
var out [][]token.Token
|
||||
start := 0
|
||||
for i, t := range ts {
|
||||
if t.Kind == token.Semicolon || t.Kind == token.Newline {
|
||||
if i > start {
|
||||
out = append(out, ts[start:i])
|
||||
}
|
||||
start = i + 1
|
||||
}
|
||||
}
|
||||
if start < len(ts) {
|
||||
out = append(out, ts[start:])
|
||||
}
|
||||
return out
|
||||
}
|
||||
@@ -0,0 +1,675 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
package parser
|
||||
|
||||
import (
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
|
||||
)
|
||||
|
||||
// expand parses src with preprocessing enabled and returns the first TEXT's
|
||||
// body instructions as "MNEMONIC operand|operand" strings, the shape the
|
||||
// expansion assertions below compare against. Runs of spaces are
|
||||
// collapsed: Raw renders a token group as its tokens joined with single
|
||||
// spaces, so "$(32-7)" arrives as "$ ( 32 - 7 )" and the comparison must
|
||||
// not depend on that spelling.
|
||||
func expand(t *testing.T, src string) (*ast.File, []string) {
|
||||
t.Helper()
|
||||
f, errs := ParseWithOptions("t_amd64.s", src, Options{Expand: true})
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
ts := texts(f)
|
||||
if len(ts) == 0 {
|
||||
t.Fatalf("no TEXT in:\n%s", src)
|
||||
}
|
||||
var got []string
|
||||
for _, s := range ts[0].Body {
|
||||
in, ok := s.(*ast.Instr)
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
var ops []string
|
||||
for _, op := range in.Operands {
|
||||
ops = append(ops, op.Raw)
|
||||
}
|
||||
line := in.Mnemonic.Text + " " + strings.Join(ops, ", ")
|
||||
got = append(got, strings.ReplaceAll(line, " ", ""))
|
||||
}
|
||||
return f, got
|
||||
}
|
||||
|
||||
func wantLines(t *testing.T, got []string, want ...string) {
|
||||
t.Helper()
|
||||
strip := func(lines []string) string {
|
||||
var out []string
|
||||
for _, l := range lines {
|
||||
out = append(out, strings.ReplaceAll(l, " ", ""))
|
||||
}
|
||||
return strings.Join(out, "\n")
|
||||
}
|
||||
if strip(got) != strip(want) {
|
||||
t.Errorf("expanded body:\n %s\nwant:\n %s", strings.Join(got, "\n "), strings.Join(want, "\n "))
|
||||
}
|
||||
}
|
||||
|
||||
func TestObjectMacroExpandsAtUse(t *testing.T) {
|
||||
_, got := expand(t, `
|
||||
#define REGTMP CX
|
||||
#define TWICE ADDQ CX, AX; ADDQ CX, AX
|
||||
TEXT ·f(SB), NOSPLIT, $0
|
||||
MOVQ 8(SP), REGTMP
|
||||
TWICE
|
||||
RET
|
||||
`)
|
||||
wantLines(t, got,
|
||||
"MOVQ 8(SP), CX",
|
||||
"ADDQ CX, AX",
|
||||
"ADDQ CX, AX",
|
||||
"RET",
|
||||
)
|
||||
}
|
||||
|
||||
func TestParameterisedMacroSubstitutesArguments(t *testing.T) {
|
||||
f, errs := ParseWithOptions("t_amd64.s", `
|
||||
#define ROUND1(a, index, const, shift) \
|
||||
ADDQ $const, a; \
|
||||
MOVW (index*4)(SP), a; \
|
||||
RORQ $(32-shift), a
|
||||
TEXT ·f(SB), NOSPLIT, $0
|
||||
ROUND1(AX, 3, 0xd76aa478, 7)
|
||||
RET
|
||||
`, Options{Expand: true})
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
body := texts(f)[0].Body
|
||||
add := body[0].(*ast.Instr)
|
||||
if add.Mnemonic.Text != "ADDQ" || !add.Operands[0].Imm.HasVal ||
|
||||
add.Operands[0].Imm.Val != 0xd76aa478 || add.Operands[1].Addr.Sym == nil ||
|
||||
add.Operands[1].Addr.Sym.Name != "AX" {
|
||||
t.Errorf("ADDQ operands substituted wrong: %+v %+v", add.Operands[0].Imm, add.Operands[1].Addr)
|
||||
}
|
||||
mov := body[1].(*ast.Instr)
|
||||
if addr := mov.Operands[0].Addr; !addr.HasOff || addr.Offset != 12 {
|
||||
t.Errorf("MOVW offset = %+v, want 12 from 3*4", addr)
|
||||
}
|
||||
ror := body[2].(*ast.Instr)
|
||||
if !ror.Operands[0].Imm.HasVal || ror.Operands[0].Imm.Val != 25 {
|
||||
t.Errorf("RORQ immediate = %+v, want 25 from (32-7)", ror.Operands[0].Imm)
|
||||
}
|
||||
}
|
||||
|
||||
func TestMacroArgumentsKeepCommasInParens(t *testing.T) {
|
||||
// An argument may itself be an unparenthesised expression: the tokens
|
||||
// substitute verbatim and the parser folds the result, as the
|
||||
// toolchain's parser does.
|
||||
f, errs := ParseWithOptions("t_amd64.s", `
|
||||
#define LOAD(dst, off) MOVQ off(SP), dst
|
||||
TEXT ·f(SB), NOSPLIT, $0
|
||||
LOAD(AX, 1*8)
|
||||
RET
|
||||
`, Options{Expand: true})
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
in := texts(f)[0].Body[0].(*ast.Instr)
|
||||
addr := in.Operands[0].Addr
|
||||
if !addr.HasOff || addr.Offset != 8 {
|
||||
t.Errorf("offset = %+v, want 8", addr)
|
||||
}
|
||||
if sym := in.Operands[1].Addr.Sym; sym == nil || sym.Name != "AX" {
|
||||
t.Errorf("destination = %+v, want AX", in.Operands[1].Addr)
|
||||
}
|
||||
}
|
||||
|
||||
func TestNestedMacroInvocations(t *testing.T) {
|
||||
// An object macro naming a parameterised one, and a parameterised body
|
||||
// invoking another parameterised macro: the toolchain's input stack
|
||||
// rescans substituted tokens, and so does expansion here.
|
||||
_, got := expand(t, `
|
||||
#define DOUBLE(x) ADDQ x, x
|
||||
#define TWICE2 DOUBLE
|
||||
#define FOUR(a, b) DOUBLE(a); DOUBLE(b)
|
||||
TEXT ·f(SB), NOSPLIT, $0
|
||||
TWICE2(AX)
|
||||
FOUR(AX, CX)
|
||||
RET
|
||||
`)
|
||||
wantLines(t, got,
|
||||
"ADDQ AX, AX",
|
||||
"ADDQ AX, AX",
|
||||
"ADDQ CX, CX",
|
||||
"RET",
|
||||
)
|
||||
}
|
||||
|
||||
func TestMultiLineBodySplitsWithoutSemicolons(t *testing.T) {
|
||||
// The arm64 style: backslash-continued lines with no semicolons. The
|
||||
// continuation newline is a statement boundary, as in the toolchain.
|
||||
_, got := expand(t, `
|
||||
#define PAIR \
|
||||
ADDQ AX, AX \
|
||||
MOVQ AX, CX
|
||||
TEXT ·f(SB), NOSPLIT, $0
|
||||
PAIR
|
||||
RET
|
||||
`)
|
||||
wantLines(t, got,
|
||||
"ADDQ AX, AX",
|
||||
"MOVQ AX, CX",
|
||||
"RET",
|
||||
)
|
||||
}
|
||||
|
||||
func TestZeroArgumentMacro(t *testing.T) {
|
||||
_, got := expand(t, `
|
||||
#define BARRIER()
|
||||
TEXT ·f(SB), NOSPLIT, $0
|
||||
BARRIER()
|
||||
RET
|
||||
`)
|
||||
wantLines(t, got, "RET")
|
||||
}
|
||||
|
||||
func TestParameterisedWithoutParensStandsAsName(t *testing.T) {
|
||||
// A parameterised macro invoked without its parentheses names itself,
|
||||
// which the parser then reports as an unknown instruction rather than
|
||||
// silently expanding nothing.
|
||||
f, errs := ParseWithOptions("t_amd64.s", `
|
||||
#define M(x) ADDQ x, x
|
||||
TEXT ·f(SB), NOSPLIT, $0
|
||||
M
|
||||
RET
|
||||
`, Options{Expand: true})
|
||||
if len(errs) != 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
fn := texts(f)[0]
|
||||
if len(fn.Body) == 0 {
|
||||
t.Fatal("body empty")
|
||||
}
|
||||
in, ok := fn.Body[0].(*ast.Instr)
|
||||
if !ok || in.Mnemonic.Text != "M" {
|
||||
t.Fatalf("bare parameterised macro did not stand as its name: %+v", fn.Body[0])
|
||||
}
|
||||
}
|
||||
|
||||
func TestDefinitionScoping(t *testing.T) {
|
||||
// A definition applies from its point onward: the use before the
|
||||
// #define stays untouched.
|
||||
_, got := expand(t, `
|
||||
TEXT ·f(SB), NOSPLIT, $0
|
||||
SPECIAL
|
||||
#define SPECIAL ADDQ AX, AX
|
||||
SPECIAL
|
||||
RET
|
||||
`)
|
||||
wantLines(t, got,
|
||||
"SPECIAL",
|
||||
"ADDQ AX, AX",
|
||||
"RET",
|
||||
)
|
||||
}
|
||||
|
||||
func TestUndefRemovesMacro(t *testing.T) {
|
||||
_, got := expand(t, `
|
||||
#define TEMP AX
|
||||
TEXT ·f(SB), NOSPLIT, $0
|
||||
TEMP
|
||||
#undef TEMP
|
||||
TEMP
|
||||
RET
|
||||
`)
|
||||
wantLines(t, got,
|
||||
"AX",
|
||||
"TEMP",
|
||||
"RET",
|
||||
)
|
||||
}
|
||||
|
||||
func TestUndefUndefinedMacroIsAnError(t *testing.T) {
|
||||
_, errs := ParseWithOptions("t_amd64.s", "#undef NOSUCH\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n", Options{Expand: true})
|
||||
if len(errs) == 0 || !strings.Contains(errs[0].Error(), "undefined macro NOSUCH") {
|
||||
t.Fatalf("#undef of an undefined macro: got %v, want an error naming it", errs)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRedefinitionIsAnError(t *testing.T) {
|
||||
_, errs := ParseWithOptions("t_amd64.s", "#define A X\n#define A Y\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n", Options{Expand: true})
|
||||
if len(errs) == 0 || !strings.Contains(errs[0].Error(), "redefinition of macro A") {
|
||||
t.Fatalf("redefinition: got %v, want an error", errs)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRecursiveMacroIsAnError(t *testing.T) {
|
||||
_, errs := ParseWithOptions("t_amd64.s", "#define A B\n#define B A\nTEXT ·f(SB), NOSPLIT, $0\n\tA\n\tRET\n", Options{Expand: true})
|
||||
if len(errs) == 0 || !strings.Contains(errs[0].Error(), "recursive macro invocation") {
|
||||
t.Fatalf("recursion: got %v, want a recursive-macro error, not a hang", errs)
|
||||
}
|
||||
}
|
||||
|
||||
func TestWrongArgumentCountIsAnError(t *testing.T) {
|
||||
_, errs := ParseWithOptions("t_amd64.s", "#define M(a, b) ADDQ a, b\nTEXT ·f(SB), NOSPLIT, $0\n\tM(AX)\n\tRET\n", Options{Expand: true})
|
||||
if len(errs) == 0 || !strings.Contains(errs[0].Error(), "wrong arg count for macro M") {
|
||||
t.Fatalf("arg count: got %v, want an error", errs)
|
||||
}
|
||||
}
|
||||
|
||||
func TestConditionalsSelectOneBranch(t *testing.T) {
|
||||
_, got := expand(t, `
|
||||
#define MODE2
|
||||
TEXT ·f(SB), NOSPLIT, $0
|
||||
#ifdef MODE2
|
||||
ADDQ AX, AX
|
||||
#else
|
||||
SUBQ AX, AX
|
||||
#endif
|
||||
#ifndef MODE2
|
||||
SUBQ CX, CX
|
||||
#else
|
||||
ADDQ CX, CX
|
||||
#endif
|
||||
RET
|
||||
`)
|
||||
wantLines(t, got,
|
||||
"ADDQ AX, AX",
|
||||
"ADDQ CX, CX",
|
||||
"RET",
|
||||
)
|
||||
}
|
||||
|
||||
func TestConditionalsHideDefinitionsAndIncludes(t *testing.T) {
|
||||
// A definition inside a disabled branch must not exist, and an
|
||||
// unresolvable include there must not be followed.
|
||||
_, got := expand(t, `
|
||||
TEXT ·f(SB), NOSPLIT, $0
|
||||
#ifdef NOTDEFINED
|
||||
#define HIDEN ADDQ AX, AX
|
||||
#include "nowhere.h"
|
||||
#endif
|
||||
HIDEN
|
||||
RET
|
||||
`)
|
||||
wantLines(t, got, "HIDEN", "RET")
|
||||
}
|
||||
|
||||
func TestUnclosedConditionalIsAnError(t *testing.T) {
|
||||
_, errs := ParseWithOptions("t_amd64.s", "#ifdef X\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n", Options{Expand: true})
|
||||
if len(errs) == 0 || !strings.Contains(errs[0].Error(), "unclosed #ifdef") {
|
||||
t.Fatalf("unclosed conditional: got %v, want an error", errs)
|
||||
}
|
||||
}
|
||||
|
||||
func TestUnmatchedConditionalDelimitersAreErrors(t *testing.T) {
|
||||
_, errs := ParseWithOptions("t_amd64.s", "#endif\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n", Options{Expand: true})
|
||||
if len(errs) == 0 || !strings.Contains(errs[0].Error(), "unmatched #endif") {
|
||||
t.Fatalf("unmatched #endif: got %v, want an error", errs)
|
||||
}
|
||||
_, errs = ParseWithOptions("t_amd64.s", "#else\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n", Options{Expand: true})
|
||||
if len(errs) == 0 || !strings.Contains(errs[0].Error(), "unmatched #else") {
|
||||
t.Fatalf("unmatched #else: got %v, want an error", errs)
|
||||
}
|
||||
}
|
||||
|
||||
// includeTree writes a directory of include files and returns its path.
|
||||
func includeTree(t *testing.T, files map[string]string) string {
|
||||
t.Helper()
|
||||
dir := t.TempDir()
|
||||
for name, content := range files {
|
||||
path := filepath.Join(dir, name)
|
||||
if err := os.MkdirAll(filepath.Dir(path), 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := os.WriteFile(path, []byte(content), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
return dir
|
||||
}
|
||||
|
||||
func TestIncludeSplicesAndDefinesAreShared(t *testing.T) {
|
||||
dir := includeTree(t, map[string]string{
|
||||
"consts.h": "#define KONST $42\n",
|
||||
})
|
||||
f, errs := ParseWithOptions("t_amd64.s", `
|
||||
#include "consts.h"
|
||||
TEXT ·f(SB), NOSPLIT, $0
|
||||
MOVQ KONST, AX
|
||||
RET
|
||||
`, Options{Expand: true, IncludeDirs: []string{dir}})
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
in := texts(f)[0].Body[0].(*ast.Instr)
|
||||
if in.Mnemonic.Text != "MOVQ" || strings.ReplaceAll(in.Operands[0].Raw, " ", "") != "$42" {
|
||||
t.Fatalf("include splicing failed: %+v", in)
|
||||
}
|
||||
}
|
||||
|
||||
func TestIncludeResolutionOrder(t *testing.T) {
|
||||
// The including file's directory wins over the -I list, and the -I list
|
||||
// is searched in order.
|
||||
src := includeTree(t, map[string]string{
|
||||
"inc/main.s": "#include \"which.h\"\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n",
|
||||
"inc/which.h": "#define WHO ONE\n",
|
||||
"first/which.h": "#define WHO TWO\n",
|
||||
"second/which.h": "#define WHO THREE\n",
|
||||
})
|
||||
main := filepath.Join(src, "inc", "main.s")
|
||||
body, err := os.ReadFile(main)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
// The header exists in the including file's directory and in two -I
|
||||
// directories; the source-directory copy must win.
|
||||
f, errs := ParseWithOptions(main, string(body), Options{Expand: true, IncludeDirs: []string{
|
||||
filepath.Join(src, "first"), filepath.Join(src, "second"),
|
||||
}})
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
found := false
|
||||
for _, d := range f.Decls {
|
||||
if pp, ok := d.(*ast.Preproc); ok && strings.Contains(pp.Raw, "define WHO ONE") {
|
||||
found = true
|
||||
}
|
||||
}
|
||||
if !found {
|
||||
t.Error("the including file's directory did not win include resolution")
|
||||
}
|
||||
}
|
||||
|
||||
func TestIncludeSearchesIncludeDirsInOrder(t *testing.T) {
|
||||
src := includeTree(t, map[string]string{
|
||||
"inc/main.s": "#include \"which.h\"\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n",
|
||||
"first/which.h": "#define WHO TWO\n",
|
||||
"second/which.h": "#define WHO THREE\n",
|
||||
})
|
||||
main := filepath.Join(src, "inc", "main.s")
|
||||
body, err := os.ReadFile(main)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
f, errs := ParseWithOptions(main, string(body), Options{Expand: true, IncludeDirs: []string{
|
||||
filepath.Join(src, "first"), filepath.Join(src, "second"),
|
||||
}})
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
for _, d := range f.Decls {
|
||||
if pp, ok := d.(*ast.Preproc); ok && strings.Contains(pp.Raw, "define WHO THREE") {
|
||||
t.Error("the second -I directory was searched before the first")
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestIncludeCycleIsDetected(t *testing.T) {
|
||||
src := includeTree(t, map[string]string{
|
||||
"a.s": "#include \"b.s\"\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n",
|
||||
"b.s": "#include \"a.s\"\n",
|
||||
})
|
||||
_, errs := ParseWithOptions(filepath.Join(src, "a.s"), "#include \"b.s\"\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n",
|
||||
Options{Expand: true})
|
||||
if len(errs) == 0 || !strings.Contains(errs[0].Error(), "include cycle") {
|
||||
t.Fatalf("include cycle: got %v, want a cycle diagnostic, not a hang", errs)
|
||||
}
|
||||
}
|
||||
|
||||
func TestUnresolvableIncludeIsAnError(t *testing.T) {
|
||||
_, errs := ParseWithOptions("t_amd64.s", "#include \"nothere.h\"\nTEXT ·f(SB), NOSPLIT, $0\n\tRET\n",
|
||||
Options{Expand: true, IncludeDirs: []string{t.TempDir()}})
|
||||
if len(errs) == 0 || !strings.Contains(errs[0].Error(), `#include "nothere.h"`) {
|
||||
t.Fatalf("missing include: got %v, want a clear diagnostic", errs)
|
||||
}
|
||||
}
|
||||
|
||||
func TestTextflagHeaderIsNeverSpliced(t *testing.T) {
|
||||
// textflag.h resolves nowhere here, yet the file must parse: the flag
|
||||
// names are consumed natively and the include stays in the tree.
|
||||
f, errs := ParseWithOptions("t_amd64.s", `
|
||||
#include "textflag.h"
|
||||
TEXT ·f(SB), NOSPLIT, $0
|
||||
RET
|
||||
`, Options{Expand: true})
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
hasInclude := false
|
||||
for _, d := range f.Decls {
|
||||
if _, ok := d.(*ast.Include); ok {
|
||||
hasInclude = true
|
||||
}
|
||||
}
|
||||
if !hasInclude {
|
||||
t.Error("textflag.h include was dropped from the tree")
|
||||
}
|
||||
}
|
||||
|
||||
func TestSemicolonSplitsRawLinesToo(t *testing.T) {
|
||||
_, got := expand(t, `
|
||||
TEXT ·f(SB), NOSPLIT, $0
|
||||
BYTE $0x0f; BYTE $0x1f
|
||||
RET
|
||||
`)
|
||||
wantLines(t, got, "BYTE $0x0f", "BYTE $0x1f", "RET")
|
||||
}
|
||||
|
||||
func TestParseUnchangedWithoutExpand(t *testing.T) {
|
||||
// Without Expand the preprocessor must not exist: a macro invocation
|
||||
// stays an unexpanded instruction line. The ';' statement separator is
|
||||
// not part of the preprocessor: the plain parse path splits on it the
|
||||
// same way the expansion path does, so both spellings agree.
|
||||
f, errs := Parse("t_amd64.s", `
|
||||
#define TWICE ADDQ AX, AX
|
||||
TEXT ·f(SB), NOSPLIT, $0
|
||||
TWICE
|
||||
BYTE $0x0f; BYTE $0x1f
|
||||
RET
|
||||
`)
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
fn := texts(f)[0]
|
||||
var mnemonics []string
|
||||
for _, s := range fn.Body {
|
||||
if in, ok := s.(*ast.Instr); ok {
|
||||
mnemonics = append(mnemonics, in.Mnemonic.Text)
|
||||
}
|
||||
}
|
||||
if strings.Join(mnemonics, " ") != "TWICE BYTE BYTE RET" {
|
||||
t.Errorf("non-expanding parse changed: %v", mnemonics)
|
||||
}
|
||||
}
|
||||
|
||||
func TestConstantExpressionFolding(t *testing.T) {
|
||||
// The shapes substituted macro bodies leave behind: parenthesised
|
||||
// arithmetic in immediates and displacements, tilde complements. The
|
||||
// assertions read the semantic fields; Raw keeps the operand's tokens
|
||||
// in the canonicalised rendering, not the folded values.
|
||||
f, errs := ParseWithOptions("t_amd64.s", `
|
||||
TEXT ·f(SB), NOSPLIT, $0
|
||||
RORQ $(32-7), AX
|
||||
ANDQ $~63, AX
|
||||
MOVQ ((2*4)+0)(SP), AX
|
||||
MOVQ $((1<<3)|(1<<1)), AX
|
||||
RET
|
||||
`, Options{Expand: true})
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
body := texts(f)[0].Body
|
||||
ror := body[0].(*ast.Instr)
|
||||
if !ror.Operands[0].Imm.HasVal || ror.Operands[0].Imm.Val != 25 {
|
||||
t.Errorf("RORQ immediate = %+v, want 25", ror.Operands[0].Imm)
|
||||
}
|
||||
and := body[1].(*ast.Instr)
|
||||
if !and.Operands[0].Imm.HasVal || and.Operands[0].Imm.Val != -64 {
|
||||
t.Errorf("ANDQ immediate = %+v, want -64", and.Operands[0].Imm)
|
||||
}
|
||||
mov := body[2].(*ast.Instr)
|
||||
addr := mov.Operands[0].Addr
|
||||
if !addr.HasOff || addr.Offset != 8 || addr.Base != "SP" {
|
||||
t.Errorf("MOVQ address = %+v, want 8(SP)", addr)
|
||||
}
|
||||
mov2 := body[3].(*ast.Instr)
|
||||
if !mov2.Operands[0].Imm.HasVal || mov2.Operands[0].Imm.Val != 10 {
|
||||
t.Errorf("MOVQ immediate = %+v, want 10", mov2.Operands[0].Imm)
|
||||
}
|
||||
}
|
||||
|
||||
func TestConstantExpressionFoldsWithoutExpand(t *testing.T) {
|
||||
// Folding is a parser capability, not a preprocessing one: a
|
||||
// hand-written $(32-7) folds the same way with expansion off.
|
||||
f, errs := ParseWithOptions("t_amd64.s", "TEXT ·f(SB), NOSPLIT, $0\n\tRORQ $(32-7), AX\n\tRET\n", Options{})
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
in := texts(f)[0].Body[0].(*ast.Instr)
|
||||
if !in.Operands[0].Imm.HasVal || in.Operands[0].Imm.Val != 25 {
|
||||
t.Errorf("Imm = %+v, want 25", in.Operands[0].Imm)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParameterWithSelectorSubstitutes(t *testing.T) {
|
||||
// The lexer folds A.S4 into one identifier token, so a parameter used
|
||||
// with an element or lane selector never matched the whole-token
|
||||
// substitution; the toolchain's lexer splits on the period and its
|
||||
// substitution sees the name alone. Several parameters carry selectors
|
||||
// in one body here, which is the chacha8_arm64.s QR shape in miniature.
|
||||
_, got := expand(t, `
|
||||
#define QR(A, B, C, D) VADD A.S4, B.S4, C.S4; VEOR D.B16, A.B16, D.B16
|
||||
TEXT ·f(SB), NOSPLIT, $0
|
||||
QR(V0, V1, V2, V3)
|
||||
RET
|
||||
`)
|
||||
wantLines(t, got,
|
||||
"VADD V0.S4, V1.S4, V2.S4",
|
||||
"VEOR V3.B16, V0.B16, V3.B16",
|
||||
"RET",
|
||||
)
|
||||
}
|
||||
|
||||
func TestSelectorWithCompoundArgumentPastesTextually(t *testing.T) {
|
||||
// An argument that is itself one compound identifier pastes verbatim:
|
||||
// A.S4 over V0.B16 spells V0.B16.S4, the operand the toolchain's
|
||||
// split-then-substitute leaves behind.
|
||||
_, got := expand(t, `
|
||||
#define M(A) VADD A.S4, A.S4, A.S4
|
||||
TEXT ·f(SB), NOSPLIT, $0
|
||||
M(V0.B16)
|
||||
RET
|
||||
`)
|
||||
wantLines(t, got, "VADD V0.B16.S4, V0.B16.S4, V0.B16.S4", "RET")
|
||||
}
|
||||
|
||||
func TestSelectorAlongsideBareParameter(t *testing.T) {
|
||||
// A body may use the parameter bare and suffixed, and the argument may
|
||||
// itself end in a selector; neither disturbs the other.
|
||||
_, got := expand(t, `
|
||||
#define M(A) VADD A, A.S4, A
|
||||
TEXT ·f(SB), NOSPLIT, $0
|
||||
M(V0)
|
||||
M(V1.B16)
|
||||
RET
|
||||
`)
|
||||
wantLines(t, got,
|
||||
"VADD V0, V0.S4, V0",
|
||||
"VADD V1.B16, V1.B16.S4, V1.B16",
|
||||
"RET",
|
||||
)
|
||||
}
|
||||
|
||||
func TestSelectorKeepsNonParameterPrefixes(t *testing.T) {
|
||||
// The prefix before the period must be the whole parameter name:
|
||||
// AB.S4 never reaches a parameter A.
|
||||
_, got := expand(t, `
|
||||
#define M(A) VADD AB.S4, A.S4, AB.S4
|
||||
TEXT ·f(SB), NOSPLIT, $0
|
||||
M(V0)
|
||||
RET
|
||||
`)
|
||||
wantLines(t, got, "VADD AB.S4, V0.S4, AB.S4", "RET")
|
||||
}
|
||||
|
||||
func TestSelectorExpandsMacroValuedArgument(t *testing.T) {
|
||||
// gcm_arm64.s invokes mulRound(B1) where B1 is itself an object macro:
|
||||
// the paste stays rescannable, so B1.D1 still expands to V1.D1 the way
|
||||
// the toolchain's rescan of substituted tokens does.
|
||||
_, got := expand(t, `
|
||||
#define B1 V1
|
||||
#define mulRound(X) VPMULL X.D1, T1.D1, T3.Q1
|
||||
TEXT ·f(SB), NOSPLIT, $0
|
||||
mulRound(B1)
|
||||
RET
|
||||
`)
|
||||
wantLines(t, got, "VPMULL V1.D1, T1.D1, T3.Q1", "RET")
|
||||
}
|
||||
|
||||
func TestObjectMacroBehindSelectorExpands(t *testing.T) {
|
||||
// Ordinary code writes ACC0.B16 where ACC0 is an object macro; the
|
||||
// toolchain expands the alias because its lexer reads the selector as
|
||||
// its own token, and the lookup here must reach the macro through the
|
||||
// compound spelling the same way.
|
||||
_, got := expand(t, `
|
||||
#define ACC0 V8
|
||||
TEXT ·f(SB), NOSPLIT, $0
|
||||
VEOR ACC0.B16, ACC0.B16, ACC0.B16
|
||||
RET
|
||||
`)
|
||||
wantLines(t, got, "VEOR V8.B16, V8.B16, V8.B16", "RET")
|
||||
}
|
||||
|
||||
func TestChacha8QRMacroExpands(t *testing.T) {
|
||||
// The real QR round of chacha8_arm64.s end to end: every parameter
|
||||
// carries a selector somewhere, and the round is sixteen instructions.
|
||||
_, got := expand(t, `
|
||||
#define QR(A, B, C, D) \
|
||||
VADD A.S4, B.S4, A.S4; VEOR D.B16, A.B16, D.B16; VREV32 D.H8, D.H8; \
|
||||
VADD C.S4, D.S4, C.S4; VEOR B.B16, C.B16, V30.B16; VSHL $12, V30.S4, B.S4; VSRI $20, V30.S4, B.S4; \
|
||||
VADD A.S4, B.S4, A.S4; VEOR D.B16, A.B16, D.B16; VTBL V31.B16, [D.B16], D.B16; \
|
||||
VADD C.S4, D.S4, C.S4; VEOR B.B16, C.B16, V30.B16; VSHL $7, V30.S4, B.S4; VSRI $25, V30.S4, B.S4
|
||||
TEXT ·f(SB), NOSPLIT, $0
|
||||
QR(V0, V1, V2, V3)
|
||||
RET
|
||||
`)
|
||||
wantLines(t, got,
|
||||
"VADD V0.S4, V1.S4, V0.S4",
|
||||
"VEOR V3.B16, V0.B16, V3.B16",
|
||||
"VREV32 V3.H8, V3.H8",
|
||||
"VADD V2.S4, V3.S4, V2.S4",
|
||||
"VEOR V1.B16, V2.B16, V30.B16",
|
||||
"VSHL $12, V30.S4, V1.S4",
|
||||
"VSRI $20, V30.S4, V1.S4",
|
||||
"VADD V0.S4, V1.S4, V0.S4",
|
||||
"VEOR V3.B16, V0.B16, V3.B16",
|
||||
"VTBL V31.B16, [V3.B16], V3.B16",
|
||||
"VADD V2.S4, V3.S4, V2.S4",
|
||||
"VEOR V1.B16, V2.B16, V30.B16",
|
||||
"VSHL $7, V30.S4, V1.S4",
|
||||
"VSRI $25, V30.S4, V1.S4",
|
||||
"RET",
|
||||
)
|
||||
}
|
||||
|
||||
func TestNotAnExpressionFallsBack(t *testing.T) {
|
||||
// Symbol immediates and floats must keep their ordinary parse.
|
||||
f, errs := ParseWithOptions("t_amd64.s", "TEXT ·f(SB), NOSPLIT, $0\n\tMOVQ $1.5, AX\n\tMOVQ $·sym(SB), AX\n\tRET\n", Options{})
|
||||
if len(errs) > 0 {
|
||||
t.Fatalf("parse: %v", errs)
|
||||
}
|
||||
fn := texts(f)[0]
|
||||
mov1 := fn.Body[0].(*ast.Instr)
|
||||
if mov1.Operands[0].Imm.HasVal || mov1.Operands[0].Imm.Float != "1.5" {
|
||||
t.Errorf("float immediate parsed as %+v", mov1.Operands[0].Imm)
|
||||
}
|
||||
mov2 := fn.Body[1].(*ast.Instr)
|
||||
if mov2.Operands[0].Imm.Sym == nil {
|
||||
t.Errorf("symbol immediate parsed as %+v", mov2.Operands[0].Imm)
|
||||
}
|
||||
}
|
||||
Vendored
+44
@@ -0,0 +1,44 @@
|
||||
// Differential kernel: the _dbar (acquire/release) atomic exchange
|
||||
// variants against the Go toolchain's loong64enc1.s rows.
|
||||
|
||||
#include "textflag.h"
|
||||
|
||||
TEXT ·AMXORDBW(SB), NOSPLIT, $0
|
||||
AMXORDBW R14, (R13), R12
|
||||
RET
|
||||
|
||||
TEXT ·AMXORDBV(SB), NOSPLIT, $0
|
||||
AMXORDBV R14, (R13), R12
|
||||
RET
|
||||
|
||||
TEXT ·AMMAXDBW(SB), NOSPLIT, $0
|
||||
AMMAXDBW R14, (R13), R12
|
||||
RET
|
||||
|
||||
TEXT ·AMMAXDBV(SB), NOSPLIT, $0
|
||||
AMMAXDBV R14, (R13), R12
|
||||
RET
|
||||
|
||||
TEXT ·AMMINDBW(SB), NOSPLIT, $0
|
||||
AMMINDBW R14, (R13), R12
|
||||
RET
|
||||
|
||||
TEXT ·AMMINDBV(SB), NOSPLIT, $0
|
||||
AMMINDBV R14, (R13), R12
|
||||
RET
|
||||
|
||||
TEXT ·AMMAXDBWU(SB), NOSPLIT, $0
|
||||
AMMAXDBWU R14, (R13), R12
|
||||
RET
|
||||
|
||||
TEXT ·AMMAXDBVU(SB), NOSPLIT, $0
|
||||
AMMAXDBVU R14, (R13), R12
|
||||
RET
|
||||
|
||||
TEXT ·AMMINDBWU(SB), NOSPLIT, $0
|
||||
AMMINDBWU R14, (R13), R12
|
||||
RET
|
||||
|
||||
TEXT ·AMMINDBVU(SB), NOSPLIT, $0
|
||||
AMMINDBVU R14, (R13), R12
|
||||
RET
|
||||
Vendored
+191
@@ -0,0 +1,191 @@
|
||||
// The AVX-512 families behind the avx512enc gap: AES round ops, integer
|
||||
// VNNI and bit algorithms, word shifts and permutes with an immediate or a
|
||||
// register count, lane broadcasts and extracts, gather and scatter prefetch
|
||||
// hints, opmask broadcasts, the high/low half moves and the non-temporal
|
||||
// stores. Every result is folded back so no instruction is dead.
|
||||
|
||||
#include "textflag.h"
|
||||
|
||||
// func avx512int(p *byte, n int) uint64
|
||||
TEXT ·avx512int(SB), NOSPLIT, $0-24
|
||||
MOVQ p+0(FP), SI
|
||||
MOVQ n+16(FP), CX
|
||||
// AES rounds through the EVEX spellings, masks included.
|
||||
VAESENC Z20, Z21, Z22
|
||||
VAESENCLAST Z23, Z24, Z25
|
||||
VAESDEC (SI), Z26, Z27
|
||||
VAESDECLAST Z28, Z29, Z30
|
||||
// Integer VNNI and the bit algorithm group.
|
||||
VPDPBUSD Z1, Z2, K2, Z3
|
||||
VPDPBUSDS Z4, Z5, K2, Z6
|
||||
VPDPWSSD Z7, Z8, Z9
|
||||
VPDPWSSDS Z10, Z11, K2, Z12
|
||||
VPOPCNTW Z12, K3, Z13
|
||||
VPOPCNTB Z14, Z15
|
||||
VGF2P8MULB Z16, Z17, K4, Z18
|
||||
VGF2P8AFFINEQB $7, Z18, Z19, K5, Z20
|
||||
// Byte/word arithmetic with saturation and masks.
|
||||
VPADDSB Z1, Z2, K1, Z3
|
||||
VPADDUSW Z3, Z4, K1, Z5
|
||||
VPSUBSW Z5, Z6, K1, Z7
|
||||
VPSUBUSB Z7, Z8, K1, Z9
|
||||
VPSADBW Z9, Z10, Z11
|
||||
VPMULHRSW Z11, Z12, Z13
|
||||
VPMULHW Z13, Z14, Z15
|
||||
VPUNPCKLBW Z15, Z16, K2, Z17
|
||||
VPUNPCKHBW Z17, Z18, K2, Z19
|
||||
VPUNPCKLWD Z19, Z20, K2, Z21
|
||||
VPUNPCKHWD Z21, Z22, K2, Z23
|
||||
VPCMPEQB Z23, Z24, K2, K3
|
||||
VPCMPGTW Z25, Z26, K2, K3
|
||||
VPCMPEQQ Z27, Z28, K2
|
||||
VPMULTISHIFTQB Z29, Z30, K3, Z31
|
||||
VDBPSADBW $3, Z1, Z2, K3, Z3
|
||||
MOVQ CX, ret+16(FP)
|
||||
RET
|
||||
|
||||
// func avx512perm(p *byte) uint64
|
||||
TEXT ·avx512perm(SB), NOSPLIT, $0-16
|
||||
MOVQ p+0(FP), SI
|
||||
// Permutations: immediate and register counts, ternary logic.
|
||||
VALIGNQ $3, Z1, Z2, K1, Z3
|
||||
VPERMT2B Z3, Z4, K1, Z5
|
||||
VPERMT2W Z5, Z6, K1, Z7
|
||||
VPERMT2PS Z7, Z8, K1, Z9
|
||||
VPERMI2W Z9, Z10, K1, Z11
|
||||
VPERMI2PS Z11, Z12, K1, Z13
|
||||
VPERMI2PD Z13, Z14, K1, Z15
|
||||
VPERMB Z15, Z16, K1, Z17
|
||||
VPERMW Z17, Z18, K1, Z19
|
||||
VPERMPS Z19, Z20, Z21
|
||||
VPERMD Z20, Z21, Z22
|
||||
VPERMQ $1, Z1, K2, Z2
|
||||
VPERMQ Z3, Z4, K2, Z5
|
||||
VPERMPD $1, Z5, K2, Z6
|
||||
VPERMPD Z7, Z8, K2, Z9
|
||||
VPERMILPS $5, Z9, K2, Z10
|
||||
VPERMILPS Z11, Z12, K2, Z13
|
||||
VPERMILPD $1, Z13, K2, Z14
|
||||
VPERMILPD Z15, Z16, K2, Z17
|
||||
VPTERNLOGD $6, Z17, Z18, K2, Z19
|
||||
VPTERNLOGQ $9, Z19, Z20, K2, Z21
|
||||
// Lane shuffle and blend families.
|
||||
VSHUFPD $1, Z1, Z2, K1, Z3
|
||||
VSHUFPS $2, Z4, Z5, K1, Z6
|
||||
VBLENDMPD Z7, Z8, K1, Z9
|
||||
VBLENDMPS Z9, Z10, K1, Z11
|
||||
VPBLENDMB Z11, Z12, K1, Z13
|
||||
VPBLENDMW Z13, Z14, K1, Z15
|
||||
VPBLENDMD Z15, Z16, K1, Z17
|
||||
VPBLENDMQ Z17, Z18, K1, Z19
|
||||
// Conflicts and leading zero counts.
|
||||
VPCONFLICTD Z1, K1, Z2
|
||||
VPCONFLICTQ Z3, K1, Z4
|
||||
VPLZCNTD Z5, K1, Z6
|
||||
VPLZCNTQ Z7, K1, Z8
|
||||
// Compress and expand, byte and word widths.
|
||||
VPCOMPRESSB Z1, K1, (SI)
|
||||
VPCOMPRESSW Z2, K1, (SI)
|
||||
VPEXPANDB (SI), K1, Z3
|
||||
VPEXPANDW (SI), K1, Z4
|
||||
MOVQ SI, ret+8(FP)
|
||||
RET
|
||||
|
||||
// func avx512shift(p *byte) uint64
|
||||
TEXT ·avx512shift(SB), NOSPLIT, $0-16
|
||||
MOVQ p+0(FP), SI
|
||||
// Variable shifts and shuffles with masks.
|
||||
VPSLLVW Z1, Z2, K1, Z3
|
||||
VPSRLVW Z3, Z4, K1, Z5
|
||||
VPSRAVW Z5, Z6, K1, Z7
|
||||
VPSHLDVW Z7, Z8, K1, Z9
|
||||
VPSHRDVW Z9, Z10, K1, Z11
|
||||
VPSHLDVD Z11, Z12, K1, Z13
|
||||
VPSHLDVQ Z13, Z14, K1, Z15
|
||||
VPSHRDVD Z15, Z16, K1, Z17
|
||||
VPSHRDVQ Z17, Z18, K1, Z19
|
||||
// Immediate shifts, the word/byte-quad widths and masks.
|
||||
VPSLLW $3, Z1, K2, Z2
|
||||
VPSRLW $5, Z3, K2, Z4
|
||||
VPSRAW $7, Z5, K2, Z6
|
||||
VPSLLDQ $9, Z7, Z8
|
||||
VPSRLDQ $11, Z9, Z10
|
||||
// Register-count shifts and their memory-count forms.
|
||||
VPSLLD X1, Z2, K1, Z3
|
||||
VPSRLD 16(SI), Z4, K1, Z5
|
||||
VPSLLQ X6, Z7, K1, Z8
|
||||
VPSRLQ X9, Z10, K1, Z11
|
||||
VPSLLW X12, Z13, K1, Z14
|
||||
VPSRAW X15, Z16, K1, Z17
|
||||
VPSRAQ $13, Z12, K1, Z13
|
||||
VPSRAD X14, Z15, K1, Z16
|
||||
// Lane shuffles in and out.
|
||||
VPSHLDW $2, Z1, Z2, K1, Z3
|
||||
VPSHLDQ $4, Z3, Z4, K1, Z5
|
||||
VPSHRDW $6, Z5, Z6, K1, Z7
|
||||
VPSHRDQ $8, Z7, Z8, K1, Z9
|
||||
VPSHUFBITQMB Z9, Z10, K3
|
||||
VPTESTMB Z11, Z12, K4
|
||||
VPTESTNMQ Z13, Z14, K5
|
||||
MOVQ SI, ret+8(FP)
|
||||
RET
|
||||
|
||||
// func avx512float(x float64) float64
|
||||
TEXT ·avx512float(SB), NOSPLIT, $0-16
|
||||
// Square roots, compares and the EXP2/RCP28 helpers.
|
||||
MOVQ x+0(FP), AX
|
||||
VSQRTPD Z1, K1, Z2
|
||||
VSQRTPS Z3, K1, Z4
|
||||
VSQRTSD X1, X2, K1, X3
|
||||
VSQRTSS X3, X4, X5
|
||||
VCOMISD X5, X6
|
||||
VUCOMISS X7, X8
|
||||
VEXP2PD Z5, K1, Z6
|
||||
VRCP28PD Z7, K1, Z8
|
||||
VRCP28SD X9, X8, K1, X10
|
||||
VRSQRT28PS Z11, K1, Z12
|
||||
VRSQRT28SS X11, X10, K1, X12
|
||||
VCVTSD2SS X1, X2, X3
|
||||
VCVTSS2SD X3, X2, K1, X4
|
||||
VFMADD132PD Z1, Z2, K1, Z3
|
||||
VFMADD231SD X1, X2, K1, X3
|
||||
VFMSUBADD213PS Z3, Z4, K1, Z5
|
||||
VFNMSUB231PD Z5, Z6, K1, Z7
|
||||
// Broadcasts and masked moves.
|
||||
VBROADCASTF32X2 X1, K1, Z2
|
||||
VBROADCASTI64X2 (SI), K1, Z3
|
||||
VMOVUPS Z1, K2, Z3
|
||||
VMOVSD X14, X5, K3, X22
|
||||
VMOVSS X18, X3, K2, X25
|
||||
VMOVHPS (SI), X18, X19
|
||||
VMOVHPS X20, 8(SI)
|
||||
VMOVLHPS X16, X5, X17
|
||||
VMOVNTDQ Z7, (SI)
|
||||
VMOVNTDQA 64(SI), Z8
|
||||
VMOVNTPD Z9, (SI)
|
||||
MOVQ SI, ret+8(FP)
|
||||
RET
|
||||
|
||||
// func avx512mask(p *byte) uint64
|
||||
TEXT ·avx512mask(SB), NOSPLIT, $0-16
|
||||
MOVQ p+0(FP), SI
|
||||
// Omask broadcasts and the K register logic.
|
||||
VPBROADCASTMB2Q K1, Z2
|
||||
VPBROADCASTMW2D K3, Z4
|
||||
KUNPCKWD K6, K4, K1
|
||||
KADDB K2, K3, K5
|
||||
KORW K1, K2, K7
|
||||
// Gather and scatter prefetch hints.
|
||||
VGATHERPF0DPD K5, (SI)(Y29*8)
|
||||
VSCATTERPF1DPS K2, (SI)(Z28*4)
|
||||
// Masked gathers ride the EVEX spelling; the data length wins L'L.
|
||||
VGATHERDPD (SI)(X10*4), K7, Y22
|
||||
VPSCATTERDQ Y6, K2, (SI)(X4*1)
|
||||
// Lane extracts to general registers.
|
||||
VPEXTRB $3, X1, AX
|
||||
VPEXTRD $1, X2, DI
|
||||
VPINSRQ $1, SI, X3, X4
|
||||
VEXTRACTI32X4 $1, Z1, X5
|
||||
VINSERTI64X2 $1, X6, Z7, K2, Z8
|
||||
MOVQ SI, ret+8(FP)
|
||||
RET
|
||||
Vendored
+32
@@ -0,0 +1,32 @@
|
||||
// The runtime bookkeeping statements: FUNCDATA and PCDATA contribute no
|
||||
// text bytes on any architecture, and amd64 now matches. They sit between
|
||||
// real instructions here, with plain, static and offset symbol references
|
||||
// on the FUNCDATA lines, so the byte counts prove the zero contribution.
|
||||
|
||||
#include "textflag.h"
|
||||
|
||||
// func bookkeep(x int64) int64
|
||||
TEXT ·bookkeep(SB), NOSPLIT, $0-16
|
||||
PCDATA $0, $-1
|
||||
MOVQ x+0(FP), AX
|
||||
PCDATA $1, $-2
|
||||
FUNCDATA $0, args_stackmap(SB)
|
||||
ADDQ $1, AX
|
||||
FUNCDATA $5, arginfo0(SB)
|
||||
PCDATA $1, $3
|
||||
MOVQ AX, ret+8(FP)
|
||||
FUNCDATA $1, externalfuncdata(SB)
|
||||
PCDATA $0, $0
|
||||
RET
|
||||
|
||||
// func bookkeepstatic() int64
|
||||
TEXT ·bookkeepstatic(SB), NOSPLIT, $0-8
|
||||
// A static symbol and a defined data symbol as the funcdata target.
|
||||
// (A symbol+offset target the toolchain itself refuses.)
|
||||
FUNCDATA $2, fdtable<>(SB)
|
||||
FUNCDATA $3, undefsym(SB)
|
||||
MOVQ $7, AX
|
||||
MOVQ AX, ret+0(FP)
|
||||
RET
|
||||
|
||||
GLOBL fdtable<>(SB), NOPTR, $16
|
||||
Vendored
+31
@@ -0,0 +1,31 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
// Differential kernel for the arm64 bookkeeping statements: the funcdata.h
|
||||
// pseudo-directives (GO_ARGS, NO_LOCAL_POINTERS, FUNCDATA, PCDATA) contribute
|
||||
// no instruction bytes, and every function is byte-compared against
|
||||
// go tool asm.
|
||||
|
||||
#include "textflag.h"
|
||||
#include "funcdata.h"
|
||||
|
||||
// func bookkeep()
|
||||
TEXT ·bookkeep(SB), NOSPLIT, $8-0
|
||||
GO_ARGS
|
||||
FUNCDATA $3, inline_tree(SB)
|
||||
PCDATA $1, $2
|
||||
MOVD R1, 0(RSP)
|
||||
RET
|
||||
|
||||
// func bookkeepNoLocals()
|
||||
TEXT ·bookkeepNoLocals(SB), NOSPLIT, $16-0
|
||||
NO_LOCAL_POINTERS
|
||||
PCDATA $0, $0
|
||||
PCDATA $1, $1
|
||||
MOVD R2, 8(RSP)
|
||||
RET
|
||||
|
||||
// func bookkeepPlain()
|
||||
TEXT ·bookkeepPlain(SB), NOSPLIT, $0-0
|
||||
MOVD R3, R4
|
||||
RET
|
||||
Vendored
+40
@@ -0,0 +1,40 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
// Differential kernel for the riscv64 bookkeeping statements and the
|
||||
// slot-relative branches: FUNCDATA and PCDATA (the expanded forms of the
|
||||
// funcdata.h macros, contributing no bytes), UNDEF (the toolchain's ebreak),
|
||||
// and the JMP N(PC) slot jumps including the self-loop and the backward form.
|
||||
|
||||
#include "textflag.h"
|
||||
|
||||
TEXT ·bookkeep(SB), NOSPLIT, $8-8
|
||||
FUNCDATA $1, marks<>(SB)
|
||||
PCDATA $1, $-1
|
||||
MOV ZERO, ret+0(FP)
|
||||
PCDATA $1, $1
|
||||
UNDEF
|
||||
MOV $1, X10
|
||||
RET
|
||||
|
||||
TEXT ·slots(SB), NOSPLIT, $0-0
|
||||
MOV $1, X10
|
||||
JMP 2(PC)
|
||||
MOV $64, X11
|
||||
MOV $128, X12
|
||||
MOV $2, X11
|
||||
MOV $3, X12
|
||||
BEQ X10, X11, skip
|
||||
JMP -2(PC)
|
||||
|
||||
skip:
|
||||
JMP 0(PC)
|
||||
|
||||
TEXT ·marksreader(SB), NOSPLIT, $0-8
|
||||
MOV $marks<>(SB), X10
|
||||
MOV (X10), X11
|
||||
MOV X11, ret+0(FP)
|
||||
RET
|
||||
|
||||
GLOBL marks<>(SB), RODATA, $8
|
||||
DATA marks<>+0(SB)/8, $1234605616436508552
|
||||
Vendored
+21
@@ -0,0 +1,21 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
// Differential kernel for the loong64 two-operand BEQ/BNE spellings the
|
||||
// msan trampolines use: BEQ Rj, target compares against R0 (the beqz form).
|
||||
|
||||
#include "textflag.h"
|
||||
|
||||
TEXT ·branch2(SB), NOSPLIT, $0-8
|
||||
MOVV arg+0(FP), R4
|
||||
BEQ R4, zero
|
||||
ADDV $1, R4, R4
|
||||
|
||||
zero:
|
||||
MOVV $16, R5
|
||||
BNE R4, done
|
||||
ADDV $2, R4, R4
|
||||
|
||||
done:
|
||||
MOVV R4, ret+0(FP)
|
||||
RET
|
||||
Vendored
+2110
File diff suppressed because it is too large
Load Diff
Vendored
+48
@@ -0,0 +1,48 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
// Carry arithmetic, logical shifts, register aliases with element selectors
|
||||
// and the ADC/SBC immediate spellings: the shapes nat_arm64.s, p256 and
|
||||
// gcm_arm64.s exercise. Byte-for-byte against go tool asm.
|
||||
|
||||
#include "textflag.h"
|
||||
|
||||
#define acc0 V8
|
||||
#define acc1 V9
|
||||
#define const0 R15
|
||||
#define POLY V15
|
||||
|
||||
// carry pins the ADC/SBC family: the $0 spellings in two and three
|
||||
// operands, and the register-carry forms.
|
||||
TEXT ·carry(SB), NOSPLIT, $0-0
|
||||
ADC $0, R20
|
||||
ADC $0, R20, R4
|
||||
SBCS $0, R4
|
||||
SBCS $0, R4, R12
|
||||
SBCS R15, R4, R12
|
||||
SBC $0, R1
|
||||
ADCSW $0, R2, R3
|
||||
RET
|
||||
|
||||
// shift pins the shifted-register forms including ROR, which only the
|
||||
// logical family accepts.
|
||||
TEXT ·shift(SB), NOSPLIT, $0-0
|
||||
ANDW R9@>7, R19, R26
|
||||
AND R1@>33, R2, R3
|
||||
ADD R1<<11, R2, R3
|
||||
SUB R1->33, R2
|
||||
ORR R5<<2, R6, R7
|
||||
RET
|
||||
|
||||
// vecalias pins the vector aliases with element selectors and the
|
||||
// structure loads with aliased members.
|
||||
TEXT ·vecalias(SB), NOSPLIT, $0-0
|
||||
MOVD $0xC2, R1
|
||||
VMOV R1, POLY.D[0]
|
||||
VMOV R0, POLY.D[1]
|
||||
VEOR POLY.B16, POLY.B16, POLY.B16
|
||||
VLD1 (R0), [acc0.B16]
|
||||
VLD1.P (R0), [acc0.B16, acc1.B16]
|
||||
VST1 [acc0.B16, acc1.B16], (R1)
|
||||
VST1.P [acc0.B16, acc1.B16], 32(R1)
|
||||
RET
|
||||
Vendored
+27
@@ -0,0 +1,27 @@
|
||||
// Copyright (c) 2026 Petr Balvín <opensource@petrbalvin.org> (https://petrbalvin.org)
|
||||
// SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
// Differential kernel for the loong64 DATA value forms the runtime's exp and
|
||||
// asm files use: floating-point initialisers stored as IEEE-754 bits and
|
||||
// string initialisers zero-padded within their declared width.
|
||||
|
||||
#include "textflag.h"
|
||||
|
||||
TEXT ·floatbits(SB), NOSPLIT, $0-8
|
||||
MOVV $floats<>(SB), R12
|
||||
MOVD 8(R12), F0
|
||||
MOVD F0, ret+0(FP)
|
||||
RET
|
||||
|
||||
TEXT ·stringhead(SB), NOSPLIT, $0-8
|
||||
MOVV $msg<>(SB), R12
|
||||
MOVV (R12), R13
|
||||
MOVV R13, ret+0(FP)
|
||||
RET
|
||||
|
||||
GLOBL floats<>(SB), RODATA, $16
|
||||
DATA floats<>+0(SB)/8, $0.0
|
||||
DATA floats<>+8(SB)/8, $0.5
|
||||
|
||||
GLOBL msg<>(SB), RODATA, $20
|
||||
DATA msg<>+0(SB)/20, $"call frame too large"
|
||||
Vendored
+25
@@ -0,0 +1,25 @@
|
||||
#include "textflag.h"
|
||||
|
||||
// The kernel exercises the symbol-valued DATA spelling the runtime's rt0
|
||||
// files use: a data word holding the address of a symbol, resolved by the
|
||||
// linker through a relocation at the field.
|
||||
|
||||
// func lookup() ptr
|
||||
TEXT ·lookup(SB), NOSPLIT, $0-8
|
||||
MOVQ handlers+8(SB), AX
|
||||
MOVQ AX, ret+0(FP)
|
||||
RET
|
||||
|
||||
// func handler() int64
|
||||
TEXT ·handler(SB), NOSPLIT, $0-8
|
||||
MOVQ $42, AX
|
||||
MOVQ AX, ret+0(FP)
|
||||
RET
|
||||
|
||||
GLOBL handlers(SB), NOPTR, $24
|
||||
DATA handlers+0(SB)/8, $·handler(SB)
|
||||
DATA handlers+8(SB)/8, $table(SB)
|
||||
DATA handlers+16(SB)/8, $·handler+5(SB)
|
||||
|
||||
GLOBL table(SB), RODATA, $8
|
||||
DATA table+0(SB)/8, $0x123456789abcdef0
|
||||
Vendored
+25
@@ -0,0 +1,25 @@
|
||||
#include "textflag.h"
|
||||
|
||||
// The kernel exercises the symbol-valued DATA spelling the runtime's rt0
|
||||
// files use: a data word holding the address of a symbol, resolved by the
|
||||
// linker through a relocation at the field.
|
||||
|
||||
// func lookup() ptr
|
||||
TEXT ·lookup(SB), NOSPLIT, $0-8
|
||||
MOVD handlers+8(SB), R4
|
||||
MOVD R4, ret+0(FP)
|
||||
RET
|
||||
|
||||
// func handler() int64
|
||||
TEXT ·handler(SB), NOSPLIT, $0-8
|
||||
MOVZ $42, R4
|
||||
MOVD R4, ret+0(FP)
|
||||
RET
|
||||
|
||||
GLOBL handlers(SB), NOPTR, $24
|
||||
DATA handlers+0(SB)/8, $·handler(SB)
|
||||
DATA handlers+8(SB)/8, $table(SB)
|
||||
DATA handlers+16(SB)/8, $extentry(SB)
|
||||
|
||||
GLOBL table(SB), RODATA, $8
|
||||
DATA table+0(SB)/8, $0x123456789abcdef0
|
||||
Vendored
+19
@@ -0,0 +1,19 @@
|
||||
#include "textflag.h"
|
||||
|
||||
// The kernel exercises the U+2215 DIVISION SLASH inside a symbol's package
|
||||
// path: internal∕runtime∕atomic·Xchg, the spelling sync/atomic/asm.s uses.
|
||||
// The middle dot (U+00B7) still separates the package path from the name.
|
||||
|
||||
// func swap(a, b int64) int64
|
||||
TEXT ·swap(SB), NOSPLIT, $0-24
|
||||
MOVQ a+0(FP), DI
|
||||
MOVQ b+8(FP), SI
|
||||
CALL internal∕runtime∕atomic·Xchg(SB)
|
||||
MOVQ AX, ret+16(FP)
|
||||
RET
|
||||
|
||||
// func note() int64
|
||||
TEXT ·note(SB), NOSPLIT, $0-8
|
||||
CALL runtime∕debug·SetGCPercent(SB)
|
||||
MOVQ AX, ret+0(FP)
|
||||
RET
|
||||
Vendored
+19
@@ -0,0 +1,19 @@
|
||||
#include "textflag.h"
|
||||
|
||||
// The kernel exercises the U+2215 DIVISION SLASH inside a symbol's package
|
||||
// path: internal∕runtime∕atomic·Xchg, the spelling sync/atomic/asm.s uses.
|
||||
// The middle dot (U+00B7) still separates the package path from the name.
|
||||
|
||||
// func swap(a, b int64) int64
|
||||
TEXT ·swap(SB), NOSPLIT, $0-24
|
||||
MOVD a+0(FP), R4
|
||||
MOVD b+8(FP), R5
|
||||
CALL internal∕runtime∕atomic·Xchg(SB)
|
||||
MOVD R4, ret+16(FP)
|
||||
RET
|
||||
|
||||
// func note() int64
|
||||
TEXT ·note(SB), NOSPLIT, $0-8
|
||||
CALL runtime∕debug·SetGCPercent(SB)
|
||||
MOVD R0, ret+0(FP)
|
||||
RET
|
||||
Vendored
+33
@@ -0,0 +1,33 @@
|
||||
// The three-operand SHL/SHR forms, which go tool asm encodes as SHLD/SHRD:
|
||||
// immediate and CL (or its CX spelling) counts at the Q and W widths, next
|
||||
// to the two-operand CX-count spelling GOROOT's bignum kernels use. Every
|
||||
// result is folded back so no instruction is dead.
|
||||
|
||||
#include "textflag.h"
|
||||
|
||||
// func dblshift(x, y uint64) uint64
|
||||
TEXT ·dblshift(SB), NOSPLIT, $0-24
|
||||
MOVQ x+0(FP), SI
|
||||
MOVQ y+8(FP), DI
|
||||
MOVQ $12, CX
|
||||
SHLQ $13, SI, DI
|
||||
SHRQ $7, DI, SI
|
||||
SHLQ CX, SI, DI
|
||||
SHRQ CX, DI, SI
|
||||
SHLQ CX, SI
|
||||
SHLQ $9, DI
|
||||
SHLW $1, SI, DI
|
||||
SHRW $3, DI, SI
|
||||
XORQ DI, SI
|
||||
MOVQ SI, ret+16(FP)
|
||||
RET
|
||||
|
||||
// func dblshift32(a, b uint32) uint32
|
||||
TEXT ·dblshift32(SB), NOSPLIT, $0-12
|
||||
MOVL a+0(FP), SI
|
||||
MOVL b+4(FP), DI
|
||||
SHLL $5, SI, DI
|
||||
SHRL $2, DI, SI
|
||||
XORL SI, DI
|
||||
MOVL DI, ret+8(FP)
|
||||
RET
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user