Compare commits

...
21 Commits
Author SHA1 Message Date
petrbalvin cf6bc6987e fix(ci): pass the upload file to curl, not its interpolation
Test / test (push) Successful in 2m26s
Release / gates (push) Successful in 2m25s
Release / build (amd64, linux) (push) Successful in 1m15s
Release / build (arm64, linux) (push) Successful in 1m16s
Release / build (loong64, linux) (push) Successful in 1m16s
Release / build (riscv64, linux) (push) Successful in 1m16s
Release / release (push) Successful in 34s
2026-09-22 01:31:49 +02:00
petrbalvin ff7b1452b1 docs: name 0.35.0 as the supported release
Test / test (push) Successful in 2m33s
Release / gates (push) Successful in 2m29s
Release / build (amd64, linux) (push) Successful in 1m18s
Release / build (arm64, linux) (push) Successful in 1m20s
Release / build (loong64, linux) (push) Successful in 1m17s
Release / build (riscv64, linux) (push) Successful in 1m26s
Release / release (push) Failing after 35s
2026-09-22 00:52:56 +02:00
petrbalvin 517c1cea25 chore: prepare release v0.35.0
Test / test (push) Successful in 2m33s
Release / gates (push) Failing after 46s
Release / build (amd64, linux) (push) Skipped
Release / build (arm64, linux) (push) Skipped
Release / build (loong64, linux) (push) Skipped
Release / build (riscv64, linux) (push) Skipped
Release / release (push) Skipped
2026-09-22 00:44:10 +02:00
petrbalvin a3e3010e0f fix(cmd): resolve the runtime header test GOROOT from the go command
Test / test (push) Successful in 2m39s
2026-09-21 22:46:07 +02:00
petrbalvin 057c4eb545 docs: complete the release delta in the changelog and readme 2026-09-21 22:45:56 +02:00
petrbalvin f720381d43 feat(asm): the segment-absolute and crash-store forms GOROOT writes
Test / test (push) Failing after 2m28s
Assisted-by: GLM 5.3 Flash
2026-09-21 22:19:53 +02:00
petrbalvin 2c9042d62c feat(asm): PCALIGN alignment on amd64
Assisted-by: GLM 5.3 Flash
2026-09-21 22:00:30 +02:00
petrbalvin 82ef289d3a feat(asm): the immediate multiply and arm64 indirect branches GOROOT writes
Assisted-by: GLM 5.3 Flash
2026-09-21 21:50:11 +02:00
petrbalvin 7246b0e002 feat(asm): the TLS access pair in the toolchain's one-instruction form
Assisted-by: GLM 5.3 Flash
2026-09-21 21:35:15 +02:00
petrbalvin 8cfd40aac8 feat(asm): the operand forms and defines GOROOT writes
Assisted-by: GLM 5.3 Flash
2026-09-21 21:17:34 +02:00
petrbalvin 5382c9a8e4 feat(audit): list every corpus failure per architecture 2026-09-21 21:17:34 +02:00
petrbalvin 53de91b2df docs(asm): describe the four target architectures
Test / test (push) Failing after 2m23s
Assisted-by: GLM 5.3 Flash
2026-09-21 20:15:55 +02:00
petrbalvin 8a36af7c7d docs(asm): generate the instruction appendices
Assisted-by: GLM 5.3 Flash
2026-09-21 20:15:55 +02:00
petrbalvin e9789ce3f4 chore(arch): regenerate the instruction tables 2026-09-21 20:15:55 +02:00
petrbalvin 837231c068 docs(asm): open the assembly language reference
Assisted-by: GLM 5.3 Flash
2026-09-21 19:49:04 +02:00
petrbalvin 95025be1bc docs(changelog): describe the encoder entries by content
Test / test (push) Failing after 2m33s
2026-09-21 19:20:01 +02:00
petrbalvin 03a964bb2d docs(goobj): document the GOOBJ object file format 2026-09-21 19:19:53 +02:00
petrbalvin 123a16e346 docs(readme): state the documentation goal 2026-09-21 18:35:27 +02:00
petrbalvin 9701812bee docs: changelog for the completeness waves
Test / test (push) Failing after 3m6s
Assisted-by: GLM 5.3 Flash
2026-09-21 02:04:44 +02:00
petrbalvin 29ac03468e feat(amd64): floating-point immediates through a synthesised pool
Assisted-by: GLM 5.3 Flash
2026-09-21 02:04:44 +02:00
petrbalvin bfb7701db1 feat(amd64): emit the quad-register EVEX families
Assisted-by: GLM 5.3 Flash
2026-09-21 02:02:19 +02:00
50 changed files with 7888 additions and 138 deletions
+4 -1
View File
@@ -342,7 +342,10 @@ jobs:
my @cmd = (q{curl}, q{-sS}, q{-o}, q{/dev/null}, q{-w}, q{%{http_code}},
q{-H}, qq{Authorization: token $ENV{GITEA_TOKEN}},
q{-H}, q{Content-Type: application/octet-stream},
q{-X}, q{POST}, q{--data-binary}, qq{@$path},
# The @ must not sit inside a qq{} string: there it starts an
# array interpolation and the upload body collapses to empty,
# which Gitea stores as a 201-created zero-byte attachment.
q{-X}, q{POST}, q{--data-binary}, q{@} . $path,
qq{$ENV{GITEA_SERVER_URL}/api/v1/repos/$ENV{GITEA_REPOSITORY}/releases/$id/assets?name=$name});
open(my $curl, q{-|}, @cmd) or die qq{curl: $!};
my $code = <$curl>;
+111 -9
View File
@@ -9,6 +9,31 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
### Added
-
## [0.35.0] - 2026-09-22
### Added
- **The go_asm.h generator.** `gasm asm` generates the package's go_asm.h
itself when an assembly file includes it: the Go files beside the source
are type-checked for the target architecture and the constants and field
offsets become assembler defines, so package-context files assemble with
no compiler and no `go build` in the loop. `-GOOS` selects the
type-checking GOOS for GOOS-specific files, and the corpus audit derives
the GOOS from the file name.
- **ELF data relocations on arm64, riscv64 and loong64.** `gasm asm
--format elf` emits `.rela.data` for symbol-valued DATA initialisers on
every architecture (amd64 carried them already), so standalone ELF
objects link on all four targets.
- **Corpus failure listing.** `gasm audit-instructions --corpus --list`
prints every failing file with its failure reason, per architecture,
instead of one representative file per reason.
- **DATA with symbol values and relaxed symbol spellings.** DATA
initialisers accept `$symbol(SB)` values, laid down as an absolute
relocation at the data field (GOOBJ on all four architectures and ELF
on all four as of this release), and U+2215 is accepted inside symbol
package paths.
- **Macro expansion and include splicing.** `gasm asm`, `gasm diff` and
`gasm audit-instructions` now preprocess assembly the way the
toolchain does: object and parameterised `#define` macros expand at
@@ -19,7 +44,8 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
(`$(32-7)`, `$~63`, `(index*4)(base)`) fold at parse. Expansion
happens only on the assembly path: `gasm lint`, `gasm fmt` and the
language server keep reading the raw file.
- **The GOROOT instruction wave, part 1.** The encoder now covers the
- **Encoder coverage: the instruction families GOROOT's real code
uses.** The encoder now covers the
instruction families GOROOT's real code uses that gasm lacked,
byte-verified against `go tool asm`: on amd64 the carry ALU, the
atomics (CMPXCHG, XADD, XCHG), AES-NI, SHA-1/256, PCLMULQDQ, CRC32,
@@ -35,14 +61,90 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
Also fixed on the way: arm64 `CASD`/`CASW` lacked an opcode bit, and
riscv64 `VSETVLI` with an immediate length now canonicalises to
`vsetivli` as the toolchain does.
- **The corpus audit measures honestly.** Files named for Go ports gasm
does not target (arm, 386, s390x, ...) are no longer attempted for the
four supported architectures (no supported build compiles them), and
the headline rate is reported over attemptable files: 136 of 433 on
the full corpus (31.4 %), 135 of 383 on real code (35.2 %), from the
127 that the previous release measured. The probe battery that
decides encodability gained the operand shapes the new families use.
-
- **Encoder coverage: quad-register AVX-512 and floating-point
immediates.** The encoder gains the
quad-register AVX-512 families (4FMAPS, 4FNMADD, 4VNNIW, VP4DPWSSD,
VP4DPWSSDS) with the register list riding the inverted V'VVVV field,
floating-point immediates on the SSE scalar moves and arithmetic
(the constant lands in a synthesised read-only pool, a positive zero
collapses to XORPS exactly as the toolchain does), accept-and-ignore
FUNCDATA and PCDATA, three-operand double shifts, static-symbol
operands for the legacy SSE moves, and the pooled 64-bit immediate
materialisation on riscv64. The parser carries bracketed register
ranges, index-only VSIB memory operands and bare trailing immediates;
macro substitution reaches parameters used with element suffixes
(`A.S4`), and `;` separates statements in plain files.
- **Per-architecture reference pages.** [docs/asm/](docs/asm/README.md)
gains AMD64, ARM64, RISCV64 and LOONG64: the register files and the
roles the ABI fixes, addressing, operand order with every special form,
constants and materialisation, alignment, fences and the relocations
each target emits. An instruction inventory appendix per architecture
is generated from the toolchain's own tables by `just gen`, and the
regenerated tables recognise 147 more mnemonics than the previous
release carried (arm64 107, riscv64 31, loong64 9).
- **The Plan 9 assembly language reference.** [docs/asm/](docs/asm/README.md)
opens the complete language reference with its common core: the lexicon,
statement structure and constant expressions, the operand grammar with
the pseudo-registers and symbol naming, the directives and the function
flag vocabulary, preprocessing with `#define` and `#include`, and the
Go-embedded layer (ABI0, prototypes, `go_asm.h`, `funcdata.h` and the
runtime contract). Every claim is verified against `go tool asm` of
Go 1.27.1 and gasm's differential tests; the per-architecture pages and
generated instruction appendices follow.
- **GOOBJ format specification.** [docs/GOOBJ.md](docs/GOOBJ.md)
documents the Go object file format in full: both containers, the 96
byte header and all 19 blocks, every structure with its byte
offsets, symbol kinds and flag bits, all 106 relocation types with
the weak variants, aux symbols, the FuncInfo payload, the pc-value
table encoding, the content hashes and the builtin table, all
verified byte for byte against objects produced by Go 1.27.1's own
tools.
### Changed
- **The corpus audit measures like a build.** Files named for a Go port
gasm does not target (arm, 386, s390x, ...) are never attempted, because
no supported build compiles them; the GOOS comes from the file name; and
each target's go_asm.h is generated on the fly. The headline is reported
over attemptable files: 291 of 353 on the full corpus (82.4 %) assemble
for every target architecture and 295 of 303 on real code (97.4 %),
against 108 of 627 over all files (17.2 %) that the previous release
measured.
### Fixed
- **The operand forms GOROOT writes.** Numeric PC-relative jumps
(`JEQ 2(PC)`, the park loop `JMP 0(PC)`) resolve with the toolchain's
own instruction counting and fold jump-to-jump chains exactly as its
branch optimiser does; symbol immediates (`MOVQ $sym(SB), AX`)
assemble to the toolchain's RIP-relative LEA with an R_PCREL
relocation; negated constant expressions in operands (`ADJSP
$-(REGS - 8)`, the shape the cgo ABI macros write) fold; the immediate
multiply (`IMULQ $1000000000, AX`) encodes with the toolchain's
0x69/0x6B selection; the TLS access pair assembles as the toolchain's
one-instruction form (the bare `MOVQ TLS, r` load nops out and
`off(r)(TLS*1)` folds to the segment-prefixed absolute whose disp32
carries the R_TLSLE relocation, per-GOOS); arm64 accepts the
bare-register indirect branch (`BL R9` beside `BL (R9)`, both BLR) and
the zero-immediate store (`MOVD $0, mem` through the zero register,
rejecting non-zero immediates as the toolchain does); `PCALIGN` now
aligns on amd64, padding with the toolchain's greedy
single-instruction NOPs; the segment-absolute forms (`MOVQ 0x30(GS),
AX` and the store direction) and the absolute crash-store
(`MOVL $0xf1, 0xf1`) encode; and `gasm asm` predefines the
`GOARCH_<arch>` and `GOOS_<goos>` macros the go command passes to
`go tool asm`, so GOROOT headers' `#ifdef GOARCH_amd64` platform
blocks (`go_tls.h`'s `get_tls` and friends) select as intended. The
GOROOT corpus measure moves to 291 of 353 files assembling for every
target architecture (82.4 %), 97.4 % of the real-code corpus, from
70.8 % and 82.2 %.
- **Tool corrections across the pipeline.** The formatter keeps square
brackets in SIMD operands, statement separators and canonical macro
bodies; the linter drops false positives on shift counts, SETcc
spellings and ABIInternal references; the lexer treats a trailing
carriage return as a line end so comment text stays idempotent; and
arm64 rejects bare BTI with a diagnostic while accepting the full
family.
## [0.34.0] - 2026-09-20
+44 -6
View File
@@ -81,7 +81,10 @@ to give that syntax the tooling it deserves.
GOOBJ format, which needs the installed toolchain and which `go build`
consumes in place of the toolchain's output. Framed functions get the
stack-split guard and the morestack block, byte-identical to the
toolchain's, so split functions link too.
toolchain's, so split functions link too. The assembler preprocesses
like the toolchain (`#define`, `#include` with `-I`, `#ifdef`), generates
`go_asm.h` from the package's Go files, and carries `PCALIGN`, the
`LOCK`/`REP` prefixes and the literal-data pseudo-ops.
- **Disassembler.** `gasm dis` lists a `.s` file's functions at their real
offsets after assembling, or disassembles raw bytes from a file or stdin.
- **Dynamic verification.** `gasm verify` JIT-loads assembled functions into
@@ -110,9 +113,9 @@ Four architectures, the four that matter in practice:
| Architecture | GOARCH | File suffix | Instructions recognised |
|--------------|-------------|--------------|---------------------------------------------|
| AMD64 | `amd64` | `_amd64.s` | 1600 + common opcodes + traditional aliases |
| ARM64 | `arm64` | `_arm64.s` | 538 + common opcodes |
| RISC-V | `riscv64` | `_riscv64.s` | 961 + common opcodes |
| LoongArch | `loong64` | `_loong64.s` | 799 + common opcodes |
| ARM64 | `arm64` | `_arm64.s` | 645 + common opcodes |
| RISC-V | `riscv64` | `_riscv64.s` | 992 + common opcodes |
| LoongArch | `loong64` | `_loong64.s` | 808 + common opcodes |
"Common opcodes" are the instructions shared by every architecture (`RET`,
`JMP`, `NOP`, `CALL`, `TEXT`, `FUNCDATA`, `PCDATA`, ...). AMD64 additionally
@@ -124,8 +127,8 @@ can emit today is narrower, and a recognised but unencodable instruction is
reported as an explicit error, never as a wrong byte.
The same measurement runs over GOROOT's whole assembly corpus:
`gasm audit-instructions --corpus` reports 136 of 433 attemptable files
(31.4 %) assembling for every target architecture today (files named for
`gasm audit-instructions --corpus` reports 291 of 353 attemptable files
(82.4 %) assembling for every target architecture today (files named for
other Go ports are counted but never attempted), with the top failure
reasons per architecture; the number moves with every release.
@@ -157,6 +160,39 @@ been compiled and read, never executed. Its architecture-neutral units
run under `go test ./...`, which the race workflow and a manual run
perform; the default `just test` gate does not sweep `./debug/...`.
## The documentation goal
The toolkit is the primary goal. The secondary one is documentation: a
specification of the Plan 9 assembly language and of the GOOBJ object
format that is 100 % complete, detailed enough to implement against,
and written to a professional standard. These are the two subjects this
project works with every day, and they are the two for which no usable
documentation exists.
Go documents the language on a single page, "A Quick Guide to Go's
Assembler", which carries no section for loong64, one of the four
architectures gasm supports, and covers a fraction of what each
assembler accepts. What exists beyond it lives as comments inside the
toolchain's internal source: per-architecture reference manuals for
arm64, ppc64, riscv64 and loong64, written for the toolchain's own
maintainers rather than for an outside reader, and none at all for
amd64. GOOBJ fares worst of all. The format that `go build` consumes
has no specification anywhere: it is described by a comment in an
internal package, it is not a stable interface, and it can change with
any toolchain release.
The gap is therefore filled the only way it can be filled: by reverse
engineering the toolchain itself, the same work the encoders already
perform. Most of the documentation can come from nowhere else, and it
is written as that knowledge is produced during development. It is
verified the way the code is verified: an encoding documented here is
one that differential tests against `go tool asm` confirm
byte-for-byte, and a format field documented here is one the linker
demonstrably reads. The work has begun: [docs/GOOBJ.md](docs/GOOBJ.md)
specifies the object file format completely, and
[docs/asm/README.md](docs/asm/README.md) opens the language reference
with its common core. The per-architecture pages follow.
## Direction
The plan, in the order it is being worked:
@@ -282,6 +318,8 @@ recipe.
~/.local/share/man (MANDIR overrides); `just uninstall-man` removes
them
- [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md): components and data flow
- [docs/GOOBJ.md](docs/GOOBJ.md): the GOOBJ object file format specification
- [docs/asm/](docs/asm/README.md): the Plan 9 assembly language reference
- [docs/DEVELOPMENT.md](docs/DEVELOPMENT.md): development setup and recipes
- [CHANGELOG.md](CHANGELOG.md): release history
+1 -1
View File
@@ -7,7 +7,7 @@ releases do not receive them.
| Version | Supported |
|---|---|
| 0.34.0 | yes |
| 0.35.0 | yes |
| older releases | no |
## Reporting a vulnerability
+101 -3
View File
@@ -8,6 +8,10 @@
// names so gasm-devkit supports every instruction the real assembler does,
// with no hand-maintained (and therefore inevitably incomplete) lists.
//
// The same data feeds the generated instruction appendices of the assembly
// language reference, docs/asm/INSTRUCTIONS-<ARCH>.md, so that the reference
// cannot drift from the tables it documents.
//
// Usage (via the justfile):
//
// just gen
@@ -26,6 +30,9 @@ import (
"path/filepath"
"sort"
"strings"
"sourcedock.dev/petrbalvin/gasm-devkit/arch"
"sourcedock.dev/petrbalvin/gasm-devkit/asm"
)
// archDirs maps a gasm-devkit architecture name to its obj sub-directory.
@@ -39,11 +46,30 @@ var archDirs = []struct {
{"loong64", "loong64"},
}
// docPages maps an architecture to its generated appendix in the language
// reference. The amd64 page carries a per-mnemonic encodability column,
// decided by asm.Encodable, which mirrors the encoder's own dispatch; the
// other targets have no single cheap predicate, so their pages carry the
// inventory and point at the live measurement instead.
var docPages = []struct {
arch arch.Arch
title string
file string
anames string
encodable bool
}{
{arch.AMD64, "AMD64", "INSTRUCTIONS-AMD64.md", "cmd/internal/obj/x86/anames.go", true},
{arch.ARM64, "ARM64", "INSTRUCTIONS-ARM64.md", "cmd/internal/obj/arm64/anames.go", false},
{arch.RISCV, "RISC-V 64", "INSTRUCTIONS-RISCV64.md", "cmd/internal/obj/riscv/anames.go", false},
{arch.LOONG64, "LoongArch 64", "INSTRUCTIONS-LOONG64.md", "cmd/internal/obj/loong64/anames.go", false},
}
func main() {
goroot := strings.TrimSpace(runGoEnvGOROOT())
if goroot == "" {
fatal("could not determine GOROOT")
}
version := strings.TrimSpace(runGoEnv("GOVERSION"))
// The common opcodes shared by every architecture (RET, JMP, NOP, CALL,
// TEXT, FUNCDATA, …) live in cmd/internal/obj/util.go.
commonPath := filepath.Join(goroot, "src", "cmd", "internal", "obj", "util.go")
@@ -57,16 +83,24 @@ func main() {
}
fmt.Printf("%-8s %4d instructions -> arch/common_gen.go\n", "common", len(common))
names := map[string][]string{}
for _, a := range archDirs {
path := filepath.Join(goroot, "src", "cmd", "internal", "obj", a.sub, "anames.go")
names, err := extractInstrs(path)
names[a.arch], err = extractInstrs(path)
if err != nil {
fatal("extract %s: %v", a.arch, err)
}
if err := writeGen(a.arch, a.sub, names); err != nil {
if err := writeGen(a.arch, a.sub, names[a.arch]); err != nil {
fatal("write %s: %v", a.arch, err)
}
fmt.Printf("%-8s %4d instructions -> arch/%s_gen.go\n", a.arch, len(names), a.arch)
fmt.Printf("%-8s %4d instructions -> arch/%s_gen.go\n", a.arch, len(names[a.arch]), a.arch)
}
for _, p := range docPages {
if err := writeDocPage(p.arch, p.title, p.file, p.anames, version, p.encodable); err != nil {
fatal("write %s: %v", p.file, err)
}
fmt.Printf("%-8s -> docs/asm/%s\n", p.arch, p.file)
}
}
@@ -172,6 +206,61 @@ func writeGen(arch, sub string, names []string) error {
return os.WriteFile(filepath.Join("arch", arch+"_gen.go"), []byte(b.String()), 0o644)
}
// writeDocPage emits docs/asm/<file>, the generated instruction appendix of
// the language reference for one architecture: every mnemonic the toolchain
// accepts, with the curated summary where the architecture table carries one
// and, on amd64, a per-mnemonic encodability column.
func writeDocPage(a arch.Arch, title, file, anames, version string, encodable bool) error {
table := arch.ForArch(a)
instrs := table.Instructions()
var b strings.Builder
b.WriteString("# " + title + ": instruction inventory\n\n")
b.WriteString("Generated by gasm-devkit's `_gen` from the Go toolchain's instruction table\n")
b.WriteString("(`" + anames + "`, " + version + "); DO NOT EDIT. This page lists every mnemonic\n")
b.WriteString("`go tool asm` accepts on this target, which is the upper bound of the\n")
b.WriteString("language on it: a name absent here is not an instruction of the target,\n")
b.WriteString("and a name present here may still be one gasm's encoder cannot emit yet.\n\n")
encodableCount := 0
if encodable {
b.WriteString("The `gasm encodes` column reports whether gasm's encoder can emit the\n")
b.WriteString("mnemonic today; the gap is the encoder backlog, measured live by\n")
b.WriteString("`gasm audit-instructions`.\n\n")
b.WriteString("| Mnemonic | gasm encodes | Notes |\n")
b.WriteString("|---|---|---|\n")
for _, in := range instrs {
ok := asm.Encodable(in.Name)
if ok {
encodableCount++
}
b.WriteString("| `" + in.Name + "` | " + yesNo(ok) + " | " + in.Summary + " |\n")
}
b.WriteString("\n")
fmt.Fprintf(&b, "Recognised: %d mnemonics. gasm encodes: %d.\n", len(instrs), encodableCount)
} else {
b.WriteString("The inventory carries no per-mnemonic encoder column: on this target\n")
b.WriteString("encodability is decided per operand shape, and the live measured\n")
b.WriteString("coverage is reported by `gasm audit-instructions`.\n\n")
b.WriteString("| Mnemonic | Notes |\n")
b.WriteString("|---|---|\n")
for _, in := range instrs {
b.WriteString("| `" + in.Name + "` | " + in.Summary + " |\n")
}
b.WriteString("\n")
fmt.Fprintf(&b, "Recognised: %d mnemonics.\n", len(instrs))
}
return os.WriteFile(filepath.Join("docs", "asm", file), []byte(b.String()), 0o644)
}
// yesNo renders a boolean as the word the appendix tables use.
func yesNo(v bool) string {
if v {
return "yes"
}
return "no"
}
func runGoEnvGOROOT() string {
out, err := exec.Command("go", "env", "GOROOT").Output()
if err != nil {
@@ -180,6 +269,15 @@ func runGoEnvGOROOT() string {
return string(out)
}
// runGoEnv runs `go env` for a single variable.
func runGoEnv(name string) string {
out, err := exec.Command("go", "env", name).Output()
if err != nil {
return ""
}
return string(out)
}
func fatal(format string, args ...any) {
fmt.Fprintf(os.Stderr, "gen: "+format+"\n", args...)
os.Exit(1)
+107
View File
@@ -364,6 +364,8 @@ var arm64GeneratedInstrs = []string{
"REVW",
"ROR",
"RORW",
"RPRFM",
"SB",
"SBC",
"SBCS",
"SBCSW",
@@ -477,23 +479,68 @@ var arm64GeneratedInstrs = []string{
"UXTH",
"UXTHW",
"UXTW",
"VABS",
"VADD",
"VADDP",
"VADDV",
"VAND",
"VBCAX",
"VBIC",
"VBIF",
"VBIT",
"VBSL",
"VCLS",
"VCLZ",
"VCMEQ",
"VCMGE",
"VCMGT",
"VCMHI",
"VCMHS",
"VCMLE",
"VCMLT",
"VCMTST",
"VCNT",
"VDUP",
"VEOR",
"VEOR3",
"VEXT",
"VFABS",
"VFADD",
"VFADDP",
"VFCMEQ",
"VFCMGE",
"VFCMGT",
"VFCMLE",
"VFCMLT",
"VFCVTL",
"VFCVTL2",
"VFCVTN",
"VFCVTN2",
"VFCVTZS",
"VFCVTZU",
"VFDIV",
"VFMAX",
"VFMAXNM",
"VFMAXNMP",
"VFMAXNMV",
"VFMAXP",
"VFMAXV",
"VFMIN",
"VFMINNM",
"VFMINNMP",
"VFMINNMV",
"VFMINP",
"VFMINV",
"VFMLA",
"VFMLS",
"VFMUL",
"VFNEG",
"VFRINTM",
"VFRINTN",
"VFRINTP",
"VFRINTZ",
"VFSQRT",
"VFSUB",
"VLD1",
"VLD1R",
"VLD2",
@@ -502,11 +549,17 @@ var arm64GeneratedInstrs = []string{
"VLD3R",
"VLD4",
"VLD4R",
"VMLA",
"VMLS",
"VMOV",
"VMOVD",
"VMOVI",
"VMOVQ",
"VMOVS",
"VMUL",
"VNEG",
"VNOT",
"VORN",
"VORR",
"VPMULL",
"VPMULL2",
@@ -515,14 +568,47 @@ var arm64GeneratedInstrs = []string{
"VREV16",
"VREV32",
"VREV64",
"VSCVTF",
"VSHADD",
"VSHL",
"VSHRN",
"VSHRN2",
"VSLI",
"VSMAX",
"VSMAXP",
"VSMAXV",
"VSMIN",
"VSMINP",
"VSMINV",
"VSMLAL",
"VSMLAL2",
"VSMLSL",
"VSMLSL2",
"VSMULL",
"VSMULL2",
"VSQABS",
"VSQADD",
"VSQNEG",
"VSQSHL",
"VSQSUB",
"VSQXTN",
"VSQXTN2",
"VSQXTUN",
"VSQXTUN2",
"VSRHADD",
"VSRI",
"VSRSHR",
"VSSHL",
"VSSHLL",
"VSSHLL2",
"VSSHR",
"VST1",
"VST2",
"VST3",
"VST4",
"VSUB",
"VSXTL",
"VSXTL2",
"VTBL",
"VTBX",
"VTRN1",
@@ -530,8 +616,27 @@ var arm64GeneratedInstrs = []string{
"VUADDLV",
"VUADDW",
"VUADDW2",
"VUCVTF",
"VUHADD",
"VUMAX",
"VUMAXP",
"VUMAXV",
"VUMIN",
"VUMINP",
"VUMINV",
"VUMLAL",
"VUMLAL2",
"VUMLSL",
"VUMLSL2",
"VUMULL",
"VUMULL2",
"VUQADD",
"VUQSHL",
"VUQSUB",
"VUQXTN",
"VUQXTN2",
"VURHADD",
"VUSHL",
"VUSHLL",
"VUSHLL2",
"VUSHR",
@@ -541,6 +646,8 @@ var arm64GeneratedInstrs = []string{
"VUZP1",
"VUZP2",
"VXAR",
"VXTN",
"VXTN2",
"VZIP1",
"VZIP2",
"WFE",
+9
View File
@@ -152,6 +152,8 @@ var loong64GeneratedInstrs = []string{
"FNMADDF",
"FNMSUBD",
"FNMSUBF",
"FRINTD",
"FRINTF",
"FSCALEBD",
"FSCALEBF",
"FSEL",
@@ -177,7 +179,10 @@ var loong64GeneratedInstrs = []string{
"FTINTWF",
"JIRL",
"LL",
"LLACQV",
"LLACQW",
"LLV",
"LLW",
"LU12IW",
"LU32ID",
"LU52ID",
@@ -248,7 +253,11 @@ var loong64GeneratedInstrs = []string{
"ROTR",
"ROTRV",
"SC",
"SCQ",
"SCRELV",
"SCRELW",
"SCV",
"SCW",
"SGT",
"SGTU",
"SLL",
+31
View File
@@ -81,6 +81,9 @@ var riscvGeneratedInstrs = []string{
"CLD",
"CLDSP",
"CLI",
"CLMUL",
"CLMULH",
"CLMULR",
"CLUI",
"CLW",
"CLWSP",
@@ -95,13 +98,20 @@ var riscvGeneratedInstrs = []string{
"CSDSP",
"CSLLI",
"CSRAI",
"CSRC",
"CSRCI",
"CSRLI",
"CSRR",
"CSRRC",
"CSRRCI",
"CSRRS",
"CSRRSI",
"CSRRW",
"CSRRWI",
"CSRS",
"CSRSI",
"CSRW",
"CSRWI",
"CSUB",
"CSUBW",
"CSW",
@@ -259,6 +269,7 @@ var riscvGeneratedInstrs = []string{
"ORCB",
"ORI",
"ORN",
"PAUSE",
"RDCYCLE",
"RDINSTRET",
"RDTIME",
@@ -322,6 +333,8 @@ var riscvGeneratedInstrs = []string{
"VADDVI",
"VADDVV",
"VADDVX",
"VANDNVV",
"VANDNVX",
"VANDVI",
"VANDVV",
"VANDVX",
@@ -329,8 +342,17 @@ var riscvGeneratedInstrs = []string{
"VASUBUVX",
"VASUBVV",
"VASUBVX",
"VBREV8V",
"VBREVV",
"VCLMULHVV",
"VCLMULHVX",
"VCLMULVV",
"VCLMULVX",
"VCLZV",
"VCOMPRESSVM",
"VCPOPM",
"VCPOPV",
"VCTZV",
"VDIVUVV",
"VDIVUVX",
"VDIVVV",
@@ -743,10 +765,16 @@ var riscvGeneratedInstrs = []string{
"VREMUVX",
"VREMVV",
"VREMVX",
"VREV8V",
"VRGATHEREI16VV",
"VRGATHERVI",
"VRGATHERVV",
"VRGATHERVX",
"VROLVV",
"VROLVX",
"VRORVI",
"VRORVV",
"VRORVX",
"VRSUBVI",
"VRSUBVX",
"VS1RV",
@@ -950,6 +978,9 @@ var riscvGeneratedInstrs = []string{
"VWMULVX",
"VWREDSUMUVS",
"VWREDSUMVS",
"VWSLLVI",
"VWSLLVV",
"VWSLLVX",
"VWSUBUVV",
"VWSUBUVX",
"VWSUBUWV",
+23
View File
@@ -637,6 +637,20 @@ func encodeARM64Branch(mnem string, ops []*ast.Operand, pc int, offsets map[stri
return a64wordLE(a64UncondBranch(opc, uint32(rn), 0)), nil
}
// The bare spelling BL R9 is the same indirect branch: the parser reads
// a bare identifier as a symbol, and one named for a register is an
// indirect branch through it, which the toolchain accepts alongside the
// parenthesised form (BL (R3) and BL R3 both encode BLR R3).
if op.Addr.Sym != nil && op.Addr.Sym.Pseudo == "" && op.Addr.Base == "" && op.Addr.Index == "" {
if rn := arm64RegNum(op.Addr.Sym.Name); rn >= 0 {
opc := uint32(0) // BR
if link {
opc = 1 // BLR
}
return a64wordLE(a64UncondBranch(opc, uint32(rn), 0)), nil
}
}
// Symbol reference: BL sym(SB), or B sym(SB) for a tail call, against a
// relocation (R_CALLARM64 either way).
if op.Addr.Sym != nil && op.Addr.Sym.Pseudo == "SB" {
@@ -1454,6 +1468,15 @@ func encodeARM64Mov(instr *ast.Instr, mnem string, wb string, fi arm64FrameInfo,
}
return encodeARM64SBAddr(src.Imm.Sym, rd, relocs), nil
}
// Immediate → memory: only storing zero is encodable (the ZR
// register); the toolchain rejects any other immediate-to-memory
// combination ("illegal combination").
if isMemOperand(dst) {
if arm64Imm64(src) != 0 {
return nil, fmt.Errorf("%s: illegal combination: an immediate store must be zero", mnem)
}
return encodeARM64MemOp(mnem, dst, 31, false, fi, "")
}
rd := arm64RegNum(operandRegName(dst))
if rd < 0 {
return nil, fmt.Errorf("%s $imm: invalid destination register", mnem)
+483 -43
View File
@@ -5,6 +5,7 @@ package asm
import (
"fmt"
"strconv"
"strings"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
@@ -28,7 +29,7 @@ import (
// emitted: the bytes match go tool asm only for NOSPLIT functions or
// zero-frame leaves, where the toolchain emits no guard either.
func Assemble(t *ast.Text) ([]byte, map[string]int, error) {
code, _, labels, _, _, err := assemble(t, nil)
code, _, labels, _, _, _, err := assemble(t, nil)
return code, labels, err
}
@@ -37,10 +38,25 @@ func Assemble(t *ast.Text) ([]byte, map[string]int, error) {
// rejects SB operands outright (single-function assembly cannot resolve
// them). When allowExternal is set, a reference to a symbol no GLOBL in the
// file defines is recorded as an external relocation instead of failing
// the object-file emitters resolve it at link time.
// the object-file emitters resolve it at link time. goos selects the TLS
// access form: the empty default behaves as linux.
type linkInfo struct {
symbols map[string]bool
allowExternal bool
goos string
}
// tlsOneInsn reports the one-instruction TLS form, obj6.go's
// CanUse1InsnTLS for the GOOS gasm supports: the bare TLS load nops out and
// the (TLS*1) index folds to a segment-absolute access. Windows and plan9
// keep the two-instruction form; shared linux does too, which gasm's raw
// path does not model and therefore does not select.
func (l *linkInfo) tlsOneInsn() bool {
switch l.goos {
case "", "linux", "freebsd":
return true
}
return false
}
// sbPatch is a function-relative static-symbol relocation: the disp32 field
@@ -66,9 +82,9 @@ type spadjStep struct {
// assemble encodes a TEXT body, returning the machine code, the static-symbol
// patch sites (for the file-level layout to resolve), the label table and the
// stack-adjustment boundaries.
func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, []spadjStep, []LineEntry, error) {
func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, []spadjStep, []LineEntry, []floatPoolEntry, error) {
if err := checkAdjspBalance(t); err != nil {
return nil, nil, nil, nil, nil, err
return nil, nil, nil, nil, nil, nil, err
}
fi := computeFrame(t)
chain := jumpChain(t)
@@ -85,29 +101,118 @@ func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, [
// outgrows the short form.
long := make([]bool, len(t.Body))
sizes := make([]int, len(t.Body))
numTargets := make([]int, len(t.Body))
for i := range numTargets {
numTargets[i] = -1
}
offsets := map[string]int{}
pcs := make([]int, len(t.Body))
var guardJBlong, guardJBElong, moreJMPlong bool
poolSeen := map[string]bool{}
var poolList []floatPoolEntry
for {
guard := fi.guardLen(guardJBlong, guardJBElong)
pos := guard + len(fi.prologue)
for i := range numTargets {
numTargets[i] = -1
}
idxAtPc := map[int]int{}
for i, stmt := range t.Body {
switch s := stmt.(type) {
case *ast.Label:
offsets[s.Name.Text] = pos
case *ast.Instr:
if strings.ToUpper(s.Mnemonic.Text) == "PCALIGN" {
// The alignment pseudo-statement: its size is the
// padding to the next boundary at this very position,
// filled with NOPs at emission.
pad, err := pcAlignPad(pcAlignValue(s), pos)
if err != nil {
return nil, nil, nil, nil, nil, nil, fmt.Errorf("PCALIGN: %w", err)
}
sizes[i] = pad
pcs[i] = pos
pos += pad
continue
}
sz, err := instrSize(s, fi, long[i], link)
if err != nil {
return nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", s.Mnemonic.Text, err)
return nil, nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", s.Mnemonic.Text, err)
}
sizes[i] = sz
pcs[i] = pos
idxAtPc[pos] = i
pos += sz
}
}
bodyLen := pos - (guard + len(fi.prologue))
// Expand any short jump whose displacement no longer fits rel8.
changed := false
// Numeric ±N(PC) jumps resolve against this iteration's layout; the
// emission pass reads the same table after the loop converges. A
// target that is itself an unconditional local JMP is chased to the
// ultimate target: the toolchain's brloop pass collapses branch-to-
// branch chains before it encodes, so matching its bytes requires
// the same redirection.
for i := range numTargets {
numTargets[i] = -1
}
for i, stmt := range t.Body {
s, ok := stmt.(*ast.Instr)
if !ok {
continue
}
if len(s.Operands) == 1 {
if n, isNum := pcJumpOffset(s.Operands[0]); isNum {
if target, okT := pcJumpTarget(t, i, n, pcs); okT {
numTargets[i] = target
}
}
}
}
for i := range numTargets {
if numTargets[i] < 0 {
continue
}
tgt := numTargets[i]
for hop := 0; hop < len(t.Body); hop++ {
idx, ok := idxAtPc[tgt]
if !ok {
break
}
in, ok := t.Body[idx].(*ast.Instr)
if !ok || strings.ToUpper(in.Mnemonic.Text) != "JMP" || len(in.Operands) != 1 {
break
}
if name, isLabel := labelName(in.Operands[0]); isLabel {
tgt = offsets[resolve(name)]
continue
}
if n, isNum := pcJumpOffset(in.Operands[0]); isNum {
next, okT := pcJumpTarget(t, idx, n, pcs)
if !okT {
break
}
tgt = next
continue
}
break // JMP through a register or memory: the chain ends
}
numTargets[i] = tgt
}
for i, stmt := range t.Body {
s, ok := stmt.(*ast.Instr)
if !ok {
continue
}
if numTargets[i] >= 0 && !long[i] {
rel := int64(numTargets[i] - (pcs[i] + jumpSize(strings.ToUpper(s.Mnemonic.Text), false)))
if !fits8(rel) {
long[i] = true
changed = true
}
}
}
for i, stmt := range t.Body {
s, ok := stmt.(*ast.Instr)
if !ok {
@@ -229,12 +334,18 @@ func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, [
spadjStep{pos + epi, 0},
)
}
code, ps, err := encodeInstr(s, pos, offsets, fi, long[i], resolve, link)
code, ps, pool, err := encodeInstr(s, pos, offsets, fi, long[i], resolve, link, numTargets[i])
if err != nil {
return nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", s.Mnemonic.Text, err)
return nil, nil, nil, nil, nil, nil, fmt.Errorf("%s: %w", s.Mnemonic.Text, err)
}
for _, entry := range pool {
if !poolSeen[entry.name] {
poolSeen[entry.name] = true
poolList = append(poolList, entry)
}
}
if len(code) != sizes[i] {
return nil, nil, nil, nil, nil, fmt.Errorf("%s: size mismatch (%d vs %d)", s.Mnemonic.Text, len(code), sizes[i])
return nil, nil, nil, nil, nil, nil, fmt.Errorf("%s: size mismatch (%d vs %d)", s.Mnemonic.Text, len(code), sizes[i])
}
if strings.ToUpper(s.Mnemonic.Text) == "CALL" {
for k := range ps {
@@ -272,7 +383,7 @@ func assemble(t *ast.Text, link *linkInfo) ([]byte, []sbPatch, map[string]int, [
pos += len(suffix)
}
_ = pos
return out, patches, offsets, steps, lines, nil
return out, patches, offsets, steps, lines, poolList, nil
}
// jumpChain precomputes jump-to-jump folding: a label whose first instruction
@@ -419,6 +530,83 @@ func computeFrame(t *ast.Text) frameInfo {
return fi
}
// pcJumpOffset recognises the numeric relative jump operand ±N(PC) and
// returns N: the toolchain counts instructions, not bytes, so +2(PC) targets
// the second instruction boundary after the branch.
func pcJumpOffset(op *ast.Operand) (int, bool) {
if op.Kind != ast.OpAddr || op.Addr.Base != "PC" {
return 0, false
}
return int(op.Addr.Offset), true
}
// pcJumpTarget resolves a numeric jump at statement index j: N counts the
// instruction statements after the jump itself (N = 0 is the jump's own
// address, the classic park loop), and the target is the start of the Nth
// one. It reports false when the count runs past the end of the function.
func pcJumpTarget(t *ast.Text, j, n int, pcs []int) (int, bool) {
if n == 0 {
return pcs[j], true
}
seen := 0
for k := j + 1; k < len(t.Body); k++ {
if _, ok := t.Body[k].(*ast.Instr); !ok {
continue
}
seen++
if seen == n {
return pcs[k], true
}
}
return 0, false
}
// x86 NOP encodings, single-instruction no-ops of lengths 1 to 9 (the
// toolchain's asm6.go nop table); longer padding repeats the largest that
// fits, greedy from the end.
var x86Nops = [][]byte{
{0x90},
{0x66, 0x90},
{0x0F, 0x1F, 0x00},
{0x0F, 0x1F, 0x40, 0x00},
{0x0F, 0x1F, 0x44, 0x00, 0x00},
{0x66, 0x0F, 0x1F, 0x44, 0x00, 0x00},
{0x0F, 0x1F, 0x80, 0x00, 0x00, 0x00, 0x00},
{0x0F, 0x1F, 0x84, 0x00, 0x00, 0x00, 0x00, 0x00},
{0x66, 0x0F, 0x1F, 0x84, 0x00, 0x00, 0x00, 0x00, 0x00},
}
// fillNOPs fills p with the greedy largest single-instruction NOPs, exactly
// the toolchain's fillnop.
func fillNOPs(p []byte) {
for len(p) > 0 {
m := min(len(p), len(x86Nops))
copy(p[:m], x86Nops[m-1])
p = p[m:]
}
}
// pcAlignPad computes the padding PCALIGN $align inserts at pos: the
// alignment must be a power of two in [8, 2048] and the padding runs to the
// next boundary (zero when the position is already aligned).
func pcAlignPad(align, pos int) (int, error) {
if align <= 0 || align&(align-1) != 0 || align < 8 || align > 2048 {
return 0, fmt.Errorf("alignment value of an instruction must be a power of two and in the range [8, 2048], got %d", align)
}
if lob := pos & (align - 1); lob != 0 {
return align - lob, nil
}
return 0, nil
}
// pcAlignValue reads a PCALIGN statement's alignment operand.
func pcAlignValue(s *ast.Instr) int {
if len(s.Operands) == 1 && s.Operands[0].Kind == ast.OpImmediate && s.Operands[0].Imm.HasVal {
return int(s.Operands[0].Imm.Val)
}
return 0 // rejected by pcAlignPad's range check
}
// hasCall reports whether the function body contains a CALL instruction.
func hasCall(t *ast.Text) bool {
for _, stmt := range t.Body {
@@ -593,7 +781,7 @@ func instrSize(s *ast.Instr, fi frameInfo, long bool, link *linkInfo) (int, erro
}
return jumpSize(mnem, long), nil
}
code, _, err := encodeInstr(s, 0, nil, fi, false, nil, link)
code, _, _, err := encodeInstr(s, 0, nil, fi, false, nil, link, -1)
if err != nil {
return 0, err
}
@@ -628,9 +816,21 @@ func jumpSize(mnem string, long bool) int {
// (relative to pc, the instruction's own offset). A RET in a frame-pointer
// function is prefixed with the epilogue. resolve, when non-nil, redirects a
// jump label through the jump-to-jump chain before the offset lookup.
func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, long bool, resolve func(string) string, link *linkInfo) ([]byte, []sbPatch, error) {
func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, long bool, resolve func(string) string, link *linkInfo, numTarget int) ([]byte, []sbPatch, []floatPoolEntry, error) {
mnem := strings.ToUpper(s.Mnemonic.Text)
if mnem == "PCALIGN" {
// The layout pass already accounted the padding; emit the same
// amount of NOP bytes for the statement's own position.
pad, err := pcAlignPad(pcAlignValue(s), pc)
if err != nil {
return nil, nil, nil, err
}
out := make([]byte, pad)
fillNOPs(out)
return out, nil, nil, nil
}
var prefix []byte
if mnem == "RET" && fi.useFP {
prefix = fi.epilogue
@@ -638,6 +838,7 @@ func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, lon
var code []byte
var ps []sbPatch
var pool []floatPoolEntry
var err error
if isJumpMnemonic(mnem) {
if (mnem == "CALL" || mnem == "JMP") && isSBCall(s) {
@@ -646,7 +847,7 @@ func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, lon
// or the linker.
code, ps, err = encodeSBCall(s, link)
if err != nil {
return nil, nil, err
return nil, nil, nil, err
}
for i := range ps {
ps[i].kind = RelCall
@@ -656,23 +857,23 @@ func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, lon
ps[i].off += body
ps[i].after = body + len(code)
}
return append(prefix, code...), ps, nil
return append(prefix, code...), ps, nil, nil
}
if (mnem == "CALL" || mnem == "JMP") && indirectJumpTarget(s) {
// JMP/CALL through a register or memory: no relocation and no
// label to resolve, the operand fully determines the bytes.
code, err = encodeIndirectJump(s, mnem)
if err != nil {
return nil, nil, err
return nil, nil, nil, err
}
return append(prefix, code...), nil, nil
return append(prefix, code...), nil, nil, nil
}
code, err = encodeJump(s, mnem, pc+len(prefix), offsets, long, resolve)
code, err = encodeJump(s, mnem, pc+len(prefix), offsets, long, resolve, numTarget)
} else {
code, ps, err = encodeNormal(s, fi, link)
code, ps, pool, err = encodeNormal(s, fi, link)
}
if err != nil {
return nil, nil, err
return nil, nil, nil, err
}
// Anchor the patch fields at function-relative positions: off indexes the
// disp32 field, after is the address just past the instruction.
@@ -681,49 +882,183 @@ func encodeInstr(s *ast.Instr, pc int, offsets map[string]int, fi frameInfo, lon
ps[i].off += body
ps[i].after = body + len(code)
}
return append(prefix, code...), ps, nil
return append(prefix, code...), ps, pool, nil
}
func encodeNormal(s *ast.Instr, fi frameInfo, link *linkInfo) ([]byte, []sbPatch, error) {
_, size := splitSize(strings.ToUpper(s.Mnemonic.Text))
func encodeNormal(s *ast.Instr, fi frameInfo, link *linkInfo) ([]byte, []sbPatch, []floatPoolEntry, error) {
mnemUpper := strings.ToUpper(s.Mnemonic.Text)
if mnemUpper == "FUNCDATA" || mnemUpper == "PCDATA" {
code, err := encodeBookkeeping(mnemUpper, s)
if err != nil {
return nil, nil, nil, err
}
return code, nil, nil, nil
}
// MOVQ $sym±off(SB), r64: the toolchain assembles a symbol immediate as
// LEAQ disp32(RIP), r64 with an R_PCREL relocation at the disp32 field,
// never as a 64-bit absolute immediate (verified against go tool asm).
// MOVD is the MOVQ alias; the narrower widths reject the form outright.
if (mnemUpper == "MOVQ" || mnemUpper == "MOVD") && len(s.Operands) == 2 &&
s.Operands[0].Kind == ast.OpImmediate && s.Operands[0].Imm.Sym != nil &&
s.Operands[0].Imm.Sym.Pseudo == "SB" {
mem := &ast.Operand{Kind: ast.OpAddr, Addr: ast.Address{Sym: s.Operands[0].Imm.Sym}}
src, err := operandFromAST(mnemUpper, mem, 8, fi, link)
if err != nil {
return nil, nil, nil, err
}
dst, err := operandFromAST(mnemUpper, s.Operands[1], 8, fi, link)
if err != nil {
return nil, nil, nil, err
}
e := &enc{}
if err := e.encodeLea([]Operand{src, dst}, 8); err != nil {
return nil, nil, nil, err
}
ps := make([]sbPatch, len(e.patches))
for i, p := range e.patches {
ps[i] = sbPatch{off: p.off, name: p.name, addend: p.addend}
}
return e.out, ps, nil, nil
}
// MOVQ/MOVL TLS, r: the bare TLS load. The toolchain's progedit nops
// it out on the one-instruction TLS systems (linux and freebsd, not
// shared) and encodes the segment-prefixed load elsewhere; get_tls(r),
// the macro GOROOT's go_tls.h defines, expands to exactly this
// statement, and the toolchain's pairing pass removes it whenever the
// following instruction's (TLS*1) index folds.
if (mnemUpper == "MOVQ" || mnemUpper == "MOVL") && len(s.Operands) == 2 && isBareTLS(s.Operands[0]) {
return encodeTLSBaseLoad(s, fi, link)
}
_, size := splitSize(mnemUpper)
if size == 0 {
size = 8
}
ops := make([]Operand, len(s.Operands))
for i, op := range s.Operands {
o, err := operandFromAST(op, size, fi, link)
o, err := operandFromAST(mnemUpper, op, size, fi, link)
if err != nil {
return nil, nil, err
return nil, nil, nil, err
}
ops[i] = o
}
e := &enc{}
if err := e.encode(s.Mnemonic.Text, ops); err != nil {
return nil, nil, err
return nil, nil, nil, err
}
ps := make([]sbPatch, len(e.patches))
for i, p := range e.patches {
ps[i] = sbPatch{off: p.off, name: p.name, addend: p.addend}
if p.tls {
ps[i].kind = RelTLSLE
}
}
return e.out, ps, nil
return e.out, ps, e.floatPoolList(), nil
}
// isBareTLS reports whether the operand is the bare TLS pseudo-register
// load source, the expansion of go_tls.h's get_tls(r) macro.
func isBareTLS(op *ast.Operand) bool {
return op.Kind == ast.OpAddr && op.Addr.Sym != nil &&
op.Addr.Sym.Pseudo == "" && op.Addr.Sym.Name == "TLS" &&
op.Addr.Base == "" && op.Addr.Index == ""
}
// encodeTLSBaseLoad assembles MOVQ/MOVL TLS, r. On the one-instruction TLS
// systems (linux and freebsd outside -shared, obj6.go's CanUse1InsnTLS) the
// statement nops out: the following (TLS*1) access folds to a direct
// segment-absolute load. The two-instruction systems keep the segment load,
// nine bytes with the R_TLSLE patch site at the disp32.
func encodeTLSBaseLoad(s *ast.Instr, fi frameInfo, link *linkInfo) ([]byte, []sbPatch, []floatPoolEntry, error) {
_, size := splitSize(strings.ToUpper(s.Mnemonic.Text))
if size == 0 {
size = 8
}
dst, err := operandFromAST("MOVQ", s.Operands[1], 8, fi, link)
if err != nil {
return nil, nil, nil, err
}
reg, ok := dst.(Reg)
if !ok || reg.isVec() {
return nil, nil, nil, fmt.Errorf("TLS: destination must be a general register")
}
if link == nil || link.tlsOneInsn() {
return nil, nil, nil, nil // noped out
}
seg := byte(0x64) // FS
if link.goos == "windows" {
seg = 0x65 // GS
}
e := &enc{}
i := &instr{
prefix: seg,
rexW: size == 8,
rexR: reg.idx >= 8,
opcode: []byte{0x8B},
modrm: 0x04 | (reg.idx&7)<<3,
sib: 0x25,
disp: le32(0),
tls: true,
}
if err := e.emit(i); err != nil {
return nil, nil, nil, err
}
ps := make([]sbPatch, len(e.patches))
for i, p := range e.patches {
ps[i] = sbPatch{off: p.off, name: p.name, addend: p.addend, kind: RelTLSLE}
}
return e.out, ps, nil, nil
}
// encodeBookkeeping accepts-and-ignores FUNCDATA and PCDATA at the statement
// level, before operand conversion: the toolchain's shapes are FUNCDATA
// $n, sym(SB) and PCDATA $n, $m, and neither contributes a byte to the
// function body. The symbol reference must not run through the SB-operand
// path, which demands file-level resolution the statement never needs.
func encodeBookkeeping(upper string, s *ast.Instr) ([]byte, error) {
if len(s.Operands) != 2 {
return nil, fmt.Errorf("%s expects 2 operands, got %d", upper, len(s.Operands))
}
a, b := s.Operands[0], s.Operands[1]
if a.Kind != ast.OpImmediate || !a.Imm.HasVal {
return nil, fmt.Errorf("%s: first operand must be an integer immediate", upper)
}
switch upper {
case "FUNCDATA":
if b.Kind != ast.OpAddr || b.Addr.Sym == nil || b.Addr.Sym.Pseudo != "SB" {
return nil, fmt.Errorf("FUNCDATA: second operand must be a symbol reference")
}
case "PCDATA":
if b.Kind != ast.OpImmediate || !b.Imm.HasVal {
return nil, fmt.Errorf("PCDATA: second operand must be an integer immediate")
}
}
return nil, nil
}
// encodeJump encodes a JMP/CALL/Jcc with a relative offset resolved from the
// target label, in the short (rel8) or long (rel32) form.
func encodeJump(s *ast.Instr, mnem string, pc int, offsets map[string]int, long bool, resolve func(string) string) ([]byte, error) {
// target label or from a numeric ±N(PC) instruction count, in the short
// (rel8) or long (rel32) form. numTarget is the resolved byte offset of a
// numeric operand, negative when the operand is not one.
func encodeJump(s *ast.Instr, mnem string, pc int, offsets map[string]int, long bool, resolve func(string) string, numTarget int) ([]byte, error) {
if len(s.Operands) != 1 {
return nil, fmt.Errorf("jump expects 1 operand, got %d", len(s.Operands))
}
name, ok := labelName(s.Operands[0])
if !ok {
name, isLabel := labelName(s.Operands[0])
if !isLabel && numTarget < 0 {
return nil, fmt.Errorf("jump target must be a local label")
}
if resolve != nil && mnem != "CALL" {
name = resolve(name)
}
target, ok := offsets[name]
if !ok {
return nil, fmt.Errorf("undefined label %q", name)
var target int
if isLabel {
if resolve != nil && mnem != "CALL" {
name = resolve(name)
}
t, ok := offsets[name]
if !ok {
return nil, fmt.Errorf("undefined label %q", name)
}
target = t
} else {
target = numTarget
}
rel := int64(target - (pc + jumpSize(mnem, long)))
@@ -756,7 +1091,7 @@ func isSBCall(s *ast.Instr) bool {
// encodeSBCall encodes CALL sym(SB) as E8 rel32 with a patch site.
func encodeSBCall(s *ast.Instr, link *linkInfo) ([]byte, []sbPatch, error) {
o, err := operandFromAST(s.Operands[0], 8, frameInfo{}, link)
o, err := operandFromAST(strings.ToUpper(s.Mnemonic.Text), s.Operands[0], 8, frameInfo{}, link)
if err != nil {
return nil, nil, err
}
@@ -797,6 +1132,11 @@ func indirectJumpTarget(s *ast.Instr) bool {
return false
}
a := s.Operands[0].Addr
// ±N(PC) is the numeric relative form, the PC counts instructions from
// the branch: relative, not indirect.
if a.Base == "PC" || a.Index == "PC" {
return false
}
if a.Base != "" || a.Index != "" {
return true
}
@@ -813,7 +1153,7 @@ func indirectJumpTarget(s *ast.Instr) bool {
func encodeIndirectJump(s *ast.Instr, mnem string) ([]byte, error) {
ops := make([]Operand, len(s.Operands))
for i, op := range s.Operands {
o, err := operandFromAST(op, 8, frameInfo{}, nil)
o, err := operandFromAST(mnem, op, 8, frameInfo{}, nil)
if err != nil {
return nil, err
}
@@ -830,8 +1170,11 @@ func encodeIndirectJump(s *ast.Instr, mnem string) ([]byte, error) {
var spReg = Reg{idx: 4, size: 8}
// operandFromAST converts a parsed operand into an encoder Operand, applying
// the frame translation to FP/SP pseudo-register operands.
func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Operand, error) {
// the frame translation to FP/SP pseudo-register operands. mnemUpper is the
// instruction's upper-case mnemonic, which the floating-point immediate gate
// needs: only the SSE mnemonics whose encoding takes an XMM/memory source
// accept one.
func operandFromAST(mnemUpper string, op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Operand, error) {
switch op.Kind {
case ast.OpImmediate:
if op.Imm.HasVal {
@@ -841,17 +1184,43 @@ func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Op
}
return Imm(v), nil
}
// A floating-point immediate: $1.5, $-1.0 or the parenthesised
// $(-1.0) spelling (the constant-expression folder only folds
// integers, so that shape arrives with an empty Immediate and only
// the raw spelling carries the value). The toolchain rewrites it
// into a pooled-constant read on the SSE scalar paths and rejects
// it everywhere else.
if text, neg, ok := floatImmText(op); ok {
if !sseFloatImm[mnemUpper] {
return nil, fmt.Errorf("%s does not take a floating-point immediate", mnemUpper)
}
return FloatImm{Text: text, Neg: neg}, nil
}
return nil, fmt.Errorf("non-integer immediate not supported")
case ast.OpAddr:
a := op.Addr
// A bracketed register range, [Z0-Z3]: the four-register source of
// the 4FMAPS/4VNNIW families. The EVEX quad-register emit path
// needs an encoder operand of its own, so the shape stays a named
// gap rather than an encoding.
// the 4FMAPS/4VNNIW families. The range must span four consecutive
// same-width vector registers, exactly what the toolchain's parser
// takes; the EVEX quad-register emit path reads the low end.
if a.Range != nil {
return nil, fmt.Errorf("register range %q needs quad-register encoder support", op.Raw)
lo, ok := ParseReg(a.Range.Lo)
if !ok {
return nil, fmt.Errorf("unknown register %q in range", a.Range.Lo)
}
hi, ok := ParseReg(a.Range.Hi)
if !ok {
return nil, fmt.Errorf("unknown register %q in range", a.Range.Hi)
}
if !lo.isVec() || lo.size != hi.size {
return nil, fmt.Errorf("register range %q must span four same-width vector registers", op.Raw)
}
if hi.idx != lo.idx+3 {
return nil, fmt.Errorf("register range %q must span four consecutive registers", op.Raw)
}
return RegList{Lo: lo, Hi: hi}, nil
}
// FP-relative: x+N(FP) → (N + fpAdjust)(SP). The offset N lives in the
@@ -885,12 +1254,44 @@ func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Op
// Memory with a real base register: (base), off(base), (base)(index*scale).
if a.Base != "" {
// Segment-absolute: 0x30(GS) and 0x28(FS), the windows TLS
// spellings. The segment override prefixes a disp32 absolute
// reference with no relocation.
if a.Base == "GS" || a.Base == "FS" {
seg := byte(0x64)
if a.Base == "GS" {
seg = 0x65
}
return SegAbs{Disp: a.Offset, Size: size, Seg: seg}, nil
}
base, ok := ParseReg(a.Base)
if !ok {
return nil, fmt.Errorf("unknown base register %q", a.Base)
}
m := Mem{Base: base, Disp: a.Offset, HasBase: true, Size: size}
if a.Index != "" {
if a.Index == "TLS" {
// off(base)(TLS*1): the thread-local annotation. The
// one-instruction TLS form folds it to off(TLS), the
// segment-prefixed absolute whose disp32 carries an
// R_TLS_LE patch site; the base register disappears
// from the encoding, exactly as the toolchain's
// progedit rewrites the address.
seg := byte(0x64) // FS on linux, freebsd, plan9
if link != nil && link.goos == "windows" {
seg = 0x65 // GS
}
return TLSMem{Disp: a.Offset, Size: size, Seg: seg}, nil
}
if a.Index == "GS" || a.Index == "FS" {
// 0(CX)(GS): the segment annotation rides the base
// access as the override prefix.
m.Seg = 0x64
if a.Index == "GS" {
m.Seg = 0x65
}
return m, nil
}
idx, ok := ParseReg(a.Index)
if !ok {
return nil, fmt.Errorf("unknown index register %q", a.Index)
@@ -911,6 +1312,11 @@ func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Op
}
return Mem{Index: idx, Scale: a.Scale, Disp: a.Offset, HasIndex: true, Size: size}, nil
}
// A bare displacement with no base: the absolute address form,
// MOVL $0xf1, 0xf1. No segment and no relocation.
if a.Sym == nil && a.Base == "" && a.Index == "" && a.HasOff {
return SegAbs{Disp: a.Offset, Size: size}, nil
}
// Bare register.
if a.Sym != nil && a.Sym.Pseudo == "" && a.Sym.Name != "" {
if r, ok := ParseReg(a.Sym.Name); ok {
@@ -921,3 +1327,37 @@ func operandFromAST(op *ast.Operand, size int, fi frameInfo, link *linkInfo) (Op
}
return nil, fmt.Errorf("unsupported operand")
}
// floatImmText recovers a floating-point immediate's magnitude and sign from
// the parsed operand. The ordinary spellings arrive in Imm.Float; the
// parenthesised $(-1.0) leaves the Immediate empty, because the integer
// folder cannot read it, and only the verbatim operand text still carries
// the value. Anything that is not a number a float parser accepts reports
// not-ok, so every other shape keeps its existing diagnostic.
func floatImmText(op *ast.Operand) (text string, neg bool, ok bool) {
if op.Imm.Float != "" {
return op.Imm.Float, op.Imm.Neg, true
}
if op.Imm.HasVal || op.Imm.Str != "" || op.Imm.Sym != nil {
return "", false, false
}
// joinRaw spaced the token texts; the compact spelling is what matters.
compact := strings.ReplaceAll(op.Raw, " ", "")
inner, ok := strings.CutPrefix(compact, "$(")
if !ok || !strings.HasSuffix(inner, ")") {
return "", false, false
}
inner = strings.TrimSuffix(inner, ")")
inner = strings.TrimPrefix(inner, "+")
if s, ok := strings.CutPrefix(inner, "-"); ok {
neg = true
inner = s
}
if inner == "" || !strings.ContainsAny(inner, "0123456789") {
return "", false, false
}
if _, err := strconv.ParseFloat(inner, 64); err != nil {
return "", false, false
}
return inner, neg, true
}
+25
View File
@@ -568,3 +568,28 @@ TEXT ·framed(SB), $16-8
t.Errorf("framed adjsp:\n got: %s\n want: %s", hexBytes(code), hexBytes(want))
}
}
// TestAssembleRegRange pins the bracketed register range at the statement
// level: exactly four consecutive same-width vector registers assemble, the
// toolchain's rejected shapes all report an error.
func TestAssembleRegRange(t *testing.T) {
asm := func(t *testing.T, op string) ([]byte, error) {
t.Helper()
f, errs := parser.Parse("f_amd64.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tV4FMADDPS 17(SP), "+op+", K2, Z0\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse %s: %v", op, errs)
}
code, _, err := Assemble(f.Decls[0].(*ast.Text))
return code, err
}
for _, op := range []string{"[Z0-Z3]", "[Z4-Z7]", "[Z28-Z31]"} {
if _, err := asm(t, op); err != nil {
t.Errorf("%s: %v", op, err)
}
}
for _, op := range []string{"[Z0-Z4]", "[Z0-Z2]", "[Z0-Z0]", "[Z4-Z0]", "[Z1-Z0]", "[AX-Z3]", "[Z0-AX]"} {
if _, err := asm(t, op); err == nil {
t.Errorf("%s: assembled, want an error", op)
}
}
}
+3 -3
View File
@@ -20,9 +20,9 @@ func Encodable(mnemonic string) bool {
switch upper {
case "RET", "NOP", "CALL", "JMP",
"POPFQ", "PUSHFQ", "INT", "LDMXCSR", "STMXCSR", "CMPSD", "SHA256RNDS2",
// The literal-data pseudo-ops, the accepted-and-ignored END and the
// SP adjust.
"BYTE", "WORD", "LONG", "QUAD", "END", "ADJSP":
// The literal-data pseudo-ops, the accepted-and-ignored END and
// bookkeeping statements, and the SP adjust.
"BYTE", "WORD", "LONG", "QUAD", "END", "ADJSP", "FUNCDATA", "PCDATA":
return true
}
if _, ok := noOperandTable[upper]; ok {
+228 -2
View File
@@ -5,6 +5,8 @@ package asm
import (
"fmt"
"math"
"strconv"
"strings"
)
@@ -21,6 +23,39 @@ func Encode(mnemonic string, ops ...Operand) ([]byte, error) {
type enc struct {
out []byte
patches []encPatch // disp32 fields awaiting static-symbol resolution
// FloatPool collects the pooled constants the floating-point
// immediates reference, in first-use order.
floatPool []floatPoolEntry
floatPoolSeen map[string]bool
}
// floatPoolEntry is one pooled floating-point constant: the symbol name
// the emitted RIP-relative load refers to and its IEEE-754 bytes.
type floatPoolEntry struct {
name string
data []byte
}
// addFloatPool records a pooled constant, deduplicated by symbol name.
func (e *enc) addFloatPool(name string, bits uint64, width int) {
if e.floatPoolSeen == nil {
e.floatPoolSeen = map[string]bool{}
}
if e.floatPoolSeen[name] {
return
}
e.floatPoolSeen[name] = true
data := make([]byte, width)
for i := range width {
data[i] = byte(bits >> (8 * i))
}
e.floatPool = append(e.floatPool, floatPoolEntry{name: name, data: data})
}
// floatPoolList returns the pooled constants in first-use order.
func (e *enc) floatPoolList() []floatPoolEntry {
return e.floatPool
}
// encPatch marks a 4-byte displacement field in enc.out that must receive the
@@ -29,6 +64,7 @@ type encPatch struct {
off int
name string
addend int64
tls bool // a TLS slot offset: the patch is R_TLSLE with no symbol
}
func (e *enc) encode(mnem string, ops []Operand) error {
@@ -102,6 +138,11 @@ func (e *enc) encode(mnem string, ops []Operand) error {
return e.encodeEnd(ops)
case "ADJSP":
return e.encodeAdjsp(ops)
// The runtime's bookkeeping statements carry no text bytes: go tool asm
// records FUNCDATA and PCDATA in the program list only, so the encoded
// body shows nothing, on every architecture.
case "FUNCDATA", "PCDATA":
return e.encodeFuncdata(upper, ops)
}
// VEX (AVX/AVX2) and EVEX (AVX-512) instructions: the trailing
@@ -141,11 +182,18 @@ func (e *enc) encode(mnem string, ops []Operand) error {
}
// Legacy SSE packed binaries dispatch on the full name: the packed
// integer mnemonics carry real width suffixes (PADDB/PCMPGTW/...),
// which the size split must not eat.
// which the size split must not eat. A floating-point immediate
// rewrites into a pooled-constant read on the scalar members.
if m, ok := sseBinTable[upper]; ok {
if f, isFloat := floatImmOperand(ops); isFloat {
return e.encodeSSEFloatBin(upper, m, f, ops)
}
return e.encodeSSEBin(m, ops)
}
if m, ok := sseBinTable[base]; ok {
if f, isFloat := floatImmOperand(ops); isFloat {
return e.encodeSSEFloatBin(upper, m, f, ops)
}
return e.encodeSSEBin(m, ops)
}
// The imm8-controlled legacy instructions, the lane extracts and inserts
@@ -222,7 +270,12 @@ func (e *enc) encode(mnem string, ops []Operand) error {
return e.encodeCvtInt(base, ops, size)
case "FMOVD":
return e.encodeFmov(ops)
case "MOVOU", "MOVO", "MOVOA", "MOVUPS", "MOVAPS", "MOVUPD", "MOVAPD", "MOVSD", "MOVSS":
case "MOVSD", "MOVSS":
if f, isFloat := floatImmOperand(ops); isFloat {
return e.encodeSSEFloatMove(upper, f, ops)
}
return e.encodeSSEMove(sseMoveTable[base], ops)
case "MOVOU", "MOVO", "MOVOA", "MOVUPS", "MOVAPS", "MOVUPD", "MOVAPD":
return e.encodeSSEMove(sseMoveTable[base], ops)
}
return fmt.Errorf("unsupported instruction %q", mnem)
@@ -285,6 +338,33 @@ func (e *enc) encodeData(mnem string, ops []Operand) error {
return nil
}
// encodeFuncdata accepts-and-ignores the runtime bookkeeping statements:
// FUNCDATA $n, sym(SB) and PCDATA $n, $m. go tool asm emits no text bytes
// for either (the entries live in the object's ancillary tables, not the
// function body), and the operand shapes it takes are exactly these: an
// integer count first, then a symbol reference for FUNCDATA and an integer
// value for PCDATA. The other architectures accept-and-ignore the same
// statements; amd64 now matches.
func (e *enc) encodeFuncdata(upper string, ops []Operand) error {
if len(ops) != 2 {
return fmt.Errorf("%s expects 2 operands, got %d", upper, len(ops))
}
if _, ok := ops[0].(Imm); !ok {
return fmt.Errorf("%s: first operand must be an integer immediate", upper)
}
switch upper {
case "FUNCDATA":
if _, ok := ops[1].(sbMem); !ok {
return fmt.Errorf("FUNCDATA: second operand must be a symbol reference")
}
case "PCDATA":
if _, ok := ops[1].(Imm); !ok {
return fmt.Errorf("PCDATA: second operand must be an integer immediate")
}
}
return nil
}
// encodeEnd accepts-and-ignores END. go tool asm drops the statement
// entirely: the AEND Prog is skipped when the program list is flushed, so
// the statements after an END still belong to the same function and the
@@ -319,6 +399,120 @@ func (e *enc) encodeAdjsp(ops []Operand) error {
return nil
}
// --- floating-point immediates ----------------------------------------------
// sseFloatImm lists the mnemonics whose first operand may be a floating-point
// immediate, the set go tool asm rewrites into a pooled-constant read: the
// scalar moves, the four scalar arithmetic pairs and the scalar compares.
// The packed members and the uniform forms (MAXSD, MINSD, SQRTSD, CMPSD)
// reject the immediate in the toolchain and are absent here on purpose.
var sseFloatImm = map[string]bool{
"MOVSD": true, "MOVSS": true,
"ADDSD": true, "ADDSS": true,
"SUBSD": true, "SUBSS": true,
"MULSD": true, "MULSS": true,
"DIVSD": true, "DIVSS": true,
"COMISD": true, "COMISS": true,
"UCOMISD": true, "UCOMISS": true,
}
// floatImmOperand reports whether the operand list opens with a
// floating-point immediate in the two-operand spelling (imm, dst).
func floatImmOperand(ops []Operand) (FloatImm, bool) {
if len(ops) != 2 {
return FloatImm{}, false
}
f, ok := ops[0].(FloatImm)
return f, ok
}
// floatPoolValue evaluates a floating-point immediate at the width its
// mnemonic encodes and names the pool constant the toolchain synthesises:
// $f64.<16 hex> for the doubles, $f32.<8 hex> for the singles (the float32
// rounding of the parsed value). The name carries the IEEE-754 bits; the
// section holds them little-endian.
func floatPoolValue(mnem string, f FloatImm) (bits uint64, name string, err error) {
v, err := strconv.ParseFloat(f.Text, 64)
if err != nil {
return 0, "", fmt.Errorf("invalid floating-point immediate %q", f.Text)
}
if f.Neg {
v = -v
}
if strings.HasSuffix(mnem, "D") {
bits = math.Float64bits(v)
return bits, fmt.Sprintf("$f64.%016x", bits), nil
}
bits = uint64(math.Float32bits(float32(v)))
return bits, fmt.Sprintf("$f32.%08x", bits), nil
}
// encodeSSEFloatMove encodes MOVSD/MOVSS with a floating-point immediate
// source. A positive zero needs no memory read: the toolchain emits
// XORPS dst, dst. Anything else loads the pooled constant RIP-relative
// ($f64.<hex>(SB) / $f32.<hex>(SB)), the displacement a patch site the
// file-level layout or the linker resolves.
func (e *enc) encodeSSEFloatMove(mnem string, f FloatImm, ops []Operand) error {
if !sseFloatImm[mnem] {
return fmt.Errorf("%s does not take a floating-point immediate", mnem)
}
dst, ok := ops[1].(Reg)
if !ok || !dst.isVec() {
return fmt.Errorf("%s: destination must be a vector register", mnem)
}
bits, name, err := floatPoolValue(mnem, f)
if err != nil {
return err
}
e.addFloatPool(name, bits, mwidth(mnem))
if bits == 0 {
i := &instr{opcode: []byte{0x0F, 0x57}, modrm: -1, sib: -1} // XORPS
if err := setRM(i, dst, dst, 8); err != nil {
return err
}
return e.emit(i)
}
m := sseMoveTable[mnem]
i := &instr{prefix: m.prefix, opcode: []byte{0x0F, m.load}, modrm: -1, sib: -1}
if err := setRM(i, dst, sbMem{size: mwidth(mnem), name: name}, 8); err != nil {
return err
}
return e.emit(i)
}
// encodeSSEFloatBin encodes the scalar arithmetic and compare mnemonics with
// a floating-point immediate source: the constant is read from the pool into
// the instruction's r/m side (reg = destination), the rewrite go tool asm
// performs at the source level.
func (e *enc) encodeSSEFloatBin(mnem string, m sseBin, f FloatImm, ops []Operand) error {
if !sseFloatImm[mnem] {
return fmt.Errorf("%s does not take a floating-point immediate", mnem)
}
dst, ok := ops[1].(Reg)
if !ok || !dst.isVec() {
return fmt.Errorf("%s: destination must be a vector register", mnem)
}
bits, name, err := floatPoolValue(mnem, f)
if err != nil {
return err
}
e.addFloatPool(name, bits, mwidth(mnem))
i := &instr{prefix: m.prefix, opcode: []byte{0x0F, m.op}, modrm: -1, sib: -1}
if err := setRM(i, dst, sbMem{size: mwidth(mnem), name: name}, 8); err != nil {
return err
}
return e.emit(i)
}
// mwidth returns the operand width a scalar SSE mnemonic encodes: the double
// spellings end in D, the single spellings in S.
func mwidth(mnem string) int {
if strings.HasSuffix(mnem, "D") {
return 8
}
return 4
}
// splitSize separates a trailing B/W/L/Q size suffix from the mnemonic.
func splitSize(upper string) (base string, size int) {
if upper == "" {
@@ -385,6 +579,7 @@ type instr struct {
disp []byte
imm []byte
sb *sbRef // static-symbol displacement in disp, awaiting resolution
tls bool // the displacement is a TLS slot offset, patched R_TLSLE
}
// sbRef records that an instruction's displacement refers to a static symbol
@@ -427,6 +622,9 @@ func (e *enc) emit(i *instr) error {
if i.sb != nil {
e.patches = append(e.patches, encPatch{off: len(e.out), name: i.sb.name, addend: i.sb.addend})
}
if i.tls {
e.patches = append(e.patches, encPatch{off: len(e.out), tls: true})
}
e.out = append(e.out, i.disp...)
e.out = append(e.out, i.imm...)
return nil
@@ -481,12 +679,30 @@ func setRMReg(i *instr, regField int, rexR, regForced bool, rm Operand, opSize i
i.disp = le32(0)
i.sb = &sbRef{name: r.name, addend: r.addend}
return nil
case TLSMem:
// off(TLS): the segment-prefixed absolute access, mod=00 with the
// SIB escape's disp32 absolute form. The displacement is the TLS
// slot offset, patched by the linker's TLS relocation.
i.prefix = r.Seg
i.modrm = 0x04 | regField<<3
i.sib = 0x25
i.disp = le32(r.Disp)
i.tls = true
return nil
case SegAbs:
// 0x30(GS): the segment override with the SIB escape's disp32
// absolute form, no relocation.
setSegAbs(i, regField, r)
return nil
default:
return fmt.Errorf("invalid r/m operand %T", rm)
}
}
func setMem(i *instr, regField int, m Mem) error {
if m.Seg != 0 {
i.prefix = m.Seg
}
modrm, sib, disp, xBit, bBit, err := memComponents(regField, m)
if err != nil {
return err
@@ -499,6 +715,16 @@ func setMem(i *instr, regField int, m Mem) error {
return nil
}
// setSegAbs assembles a segment-absolute operand, 0x30(GS): the segment
// override with the mod=00 SIB escape's disp32 absolute form and no
// relocation.
func setSegAbs(i *instr, regField int, m SegAbs) {
i.prefix = m.Seg
i.modrm = 0x04 | regField<<3
i.sib = 0x25
i.disp = le32(m.Disp)
}
// memComponents computes the ModR/M byte (with the given reg field), the SIB
// byte (-1 if none), the displacement bytes, and the high index/base bits, for
// a memory operand. It is shared by the REX (scalar) and VEX (vector) paths.
+157
View File
@@ -9,6 +9,9 @@ import (
"testing"
"golang.org/x/arch/x86/x86asm"
"sourcedock.dev/petrbalvin/gasm-devkit/ast"
"sourcedock.dev/petrbalvin/gasm-devkit/parser"
)
// decode encodes an instruction and decodes it back, returning the decoded
@@ -1064,3 +1067,157 @@ func TestAdjsp(t *testing.T) {
t.Error("ADJSP AX assembled, want an error")
}
}
// TestFloatImmediateGroundTruth pins the floating-point immediate rewrite
// byte for byte against go tool asm: the scalar moves and the scalar
// arithmetic read the constant from a synthesised read-only pool symbol
// ($f64.<hex>, $f32.<hex>) RIP-relative with the displacement left to the
// relocation, and a positive zero on the moves collapses to XORPS dst, dst.
func TestFloatImmediateGroundTruth(t *testing.T) {
cases := []struct {
name string
mnem string
ops []Operand
want string
}{
{"MOVSD -1.0", "MOVSD", []Operand{FloatImm{Text: "1.0", Neg: true}, vreg(t, "X2")}, "f20f101500000000"},
{"MOVSD 1.5", "MOVSD", []Operand{FloatImm{Text: "1.5"}, vreg(t, "X3")}, "f20f101d00000000"},
{"MOVSS 2.5", "MOVSS", []Operand{FloatImm{Text: "2.5"}, vreg(t, "X4")}, "f30f102500000000"},
{"MOVSS -0.5", "MOVSS", []Operand{FloatImm{Text: "0.5", Neg: true}, vreg(t, "X5")}, "f30f102d00000000"},
{"MOVSS +0.0 is XORPS", "MOVSS", []Operand{FloatImm{Text: "0.0"}, vreg(t, "X10")}, "450f57d2"},
{"MOVSD +0.0 is XORPS", "MOVSD", []Operand{FloatImm{Text: "0.0"}, vreg(t, "X6")}, "0f57f6"},
{"ADDSD 1.0", "ADDSD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}, "f20f580500000000"},
{"ADDSS 0.5", "ADDSS", []Operand{FloatImm{Text: "0.5"}, vreg(t, "X1")}, "f30f580d00000000"},
{"SUBSD 2.0", "SUBSD", []Operand{FloatImm{Text: "2.0"}, vreg(t, "X3")}, "f20f5c1d00000000"},
{"MULSD -2.5", "MULSD", []Operand{FloatImm{Text: "2.5", Neg: true}, vreg(t, "X3")}, "f20f591d00000000"},
{"DIVSD 1.0", "DIVSD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}, "f20f5e0500000000"},
{"COMISD 1.0", "COMISD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}, "660f2f0500000000"},
}
for _, c := range cases {
code, err := Encode(c.mnem, c.ops...)
if err != nil {
t.Errorf("%s: Encode: %v", c.name, err)
continue
}
if got := hexCompact(code); got != c.want {
t.Errorf("%s: got %s, want %s", c.name, got, c.want)
}
}
// The pool names carry the IEEE-754 bits, the float32 narrowing for the
// single spellings; negative zero keeps its sign bit and never takes the
// XORPS shortcut.
for _, c := range []struct {
mnem string
imm FloatImm
want string
}{
{"MOVSD", FloatImm{Text: "1.0", Neg: true}, "$f64.bff0000000000000"},
{"MOVSD", FloatImm{Text: "0.5"}, "$f64.3fe0000000000000"},
{"MOVSS", FloatImm{Text: "2.5"}, "$f32.40200000"},
{"MOVSS", FloatImm{Text: "0.5", Neg: true}, "$f32.bf000000"},
{"MOVSD", FloatImm{Text: "0.0", Neg: true}, "$f64.8000000000000000"},
} {
_, name, err := floatPoolValue(c.mnem, c.imm)
if err != nil {
t.Errorf("%s %s: %v", c.mnem, c.imm.Text, err)
continue
}
if name != c.want {
t.Errorf("%s $%s: pool name %s, want %s", c.mnem, c.imm.Text, name, c.want)
}
}
// The shapes the toolchain's parser rejects: the packed and uniform
// forms, a non-vector destination, and the integer spellings.
for _, c := range []struct {
name string
mnem string
ops []Operand
}{
{"MAXSD rejects the immediate", "MAXSD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}},
{"MINSD rejects the immediate", "MINSD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}},
{"SQRTSD rejects the immediate", "SQRTSD", []Operand{FloatImm{Text: "1.0"}, vreg(t, "X0")}},
{"integer destination", "MOVSD", []Operand{FloatImm{Text: "1.0"}, AX}},
} {
if _, err := Encode(c.mnem, c.ops...); err == nil {
t.Errorf("%s: expected an error, got none", c.name)
}
}
}
// TestBookkeepingGroundTruth pins FUNCDATA and PCDATA as accept-and-ignore:
// go tool asm emits no text bytes for either, on every architecture.
func TestBookkeepingGroundTruth(t *testing.T) {
for _, c := range []struct {
name string
mnem string
ops []Operand
}{
{"FUNCDATA", "FUNCDATA", []Operand{Imm(3), sbMem{name: "\u00b7f.arginfo0"}}},
{"PCDATA", "PCDATA", []Operand{Imm(1), Imm(-1)}},
} {
code, err := Encode(c.mnem, c.ops...)
if err != nil {
t.Errorf("%s: Encode: %v", c.name, err)
continue
}
if len(code) != 0 {
t.Errorf("%s: emitted %x, want no bytes", c.name, code)
}
}
for _, c := range []struct {
name string
mnem string
ops []Operand
}{
{"FUNCDATA arity", "FUNCDATA", []Operand{Imm(3)}},
{"FUNCDATA missing the count", "FUNCDATA", []Operand{sbMem{name: "x"}}},
{"FUNCDATA integer value", "FUNCDATA", []Operand{Imm(3), Imm(4)}},
{"PCDATA arity", "PCDATA", []Operand{Imm(1)}},
{"PCDATA register value", "PCDATA", []Operand{Imm(1), AX}},
} {
if _, err := Encode(c.mnem, c.ops...); err == nil {
t.Errorf("%s: expected an error, got none", c.name)
}
}
// At the statement level the bookkeeping lines sit between real
// instructions and contribute nothing to the body, symbol reference
// included: the FUNCDATA operand never needs file-level resolution.
f, errs := parser.Parse("t_amd64.s", "TEXT \u00b7f(SB), NOSPLIT, $0\n\tNOP\n\tFUNCDATA $3, \u00b7f.arginfo0(SB)\n\tPCDATA $1, $-1\n\tFUNCDATA $0, x<>(SB)\n\tRET\n")
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
img, err := AssembleFile(f)
if err != nil {
t.Fatalf("assemble: %v", err)
}
want := "90c3"
if got := hexCompact(img.Code); got != want {
t.Errorf("body %s, want %s (the bookkeeping lines contribute nothing)", got, want)
}
if _, err := AssembleFile(mustParse(t, "TEXT \u00b7f(SB), NOSPLIT, $0\n\tFUNCDATA $1, X0\n\tRET\n")); err == nil {
t.Error("FUNCDATA $1, X0 assembled, want an error")
}
if _, err := AssembleFile(mustParse(t, "TEXT \u00b7f(SB), NOSPLIT, $0\n\tPCDATA $1, X0\n\tRET\n")); err == nil {
t.Error("PCDATA $1, X0 assembled, want an error")
}
// Encodable mirrors Encode for the names this work touched.
for _, mnem := range []string{"FUNCDATA", "PCDATA", "V4FMADDPS", "V4FMADDSS", "V4FNMADDPS", "V4FNMADDSS", "VP4DPWSSD", "VP4DPWSSDS"} {
if !Encodable(mnem) {
t.Errorf("Encodable(%s) = false, want true", mnem)
}
}
}
// mustParse parses src or fails the test.
func mustParse(t *testing.T, src string) *ast.File {
t.Helper()
f, errs := parser.Parse("t_amd64.s", src)
if len(errs) > 0 {
t.Fatalf("parse: %v", errs)
}
return f
}
+97 -2
View File
@@ -737,6 +737,91 @@ var evexTable = map[string]evexSpec{
"VMOVLHPS": {1, 0x16, 0, 0, -1, vexNDS3, [3]int{8, 0, 0}},
}
// evexQuad describes one quad-register instruction: the opcode under
// EVEX.0F38.W0 with the F2 mandatory prefix, and the width of the vector
// registers the bracketed list and the destination take (512-bit ZMM for
// the packed forms, 128-bit XMM for the scalar ones).
type evexQuad struct {
opcode byte
width int // register width in bytes: 64 (ZMM) or 16 (XMM)
}
// evexQuadTable maps the quad-register instructions (the 4FMAPS and 4VNNIW
// families) to their encoding. The operand shape is fixed: a single memory
// source in r/m, the bracketed register list whose LOW register travels the
// inverted 5-bit V'VVVV field, an optional opmask in aaa and the vector
// destination in reg. The vector length follows the destination (512-bit
// for the ZMM list forms, 128-bit for the scalar ones) while the disp8×N
// multiplier stays 16 for every member, the toolchain's own tuple choice.
var evexQuadTable = map[string]evexQuad{
"V4FMADDPS": {0x9A, 64},
"V4FMADDSS": {0x9B, 16},
"V4FNMADDPS": {0xAA, 64},
"V4FNMADDSS": {0xAB, 16},
"VP4DPWSSD": {0x52, 64},
"VP4DPWSSDS": {0x53, 64},
}
// isEvexQuad reports whether the mnemonic is a quad-register instruction.
func isEvexQuad(upper string) bool {
_, ok := evexQuadTable[upper]
return ok
}
// encodeEvexQuad encodes the quad-register form: OP mem, [Zn-Zn+3], (K), dst.
// The register list is the VVVV-side source: its low register fills the
// inverted V'VVVV bits, which is why an indexed memory source above Z15 (no
// spare EVEX.X bit once V' is taken) is refused. Masking rides the standard
// aaa field, zeroing keeps the usual requires-a-mask rule, and no other
// suffix applies.
func (e *enc) encodeEvexQuad(mnem string, q evexQuad, ops []Operand, sfx evexSuffix) error {
if len(ops) != 3 && len(ops) != 4 {
return fmt.Errorf("%s expects 3 or 4 operands (mem, [Zn-Zn+3], (K), dst), got %d", mnem, len(ops))
}
mem, lst := ops[0], ops[1]
dst := ops[len(ops)-1]
mask := 0
if len(ops) == 4 {
k, ok := ops[2].(Reg)
if !ok || !k.mask {
return fmt.Errorf("%s: third operand must be an opmask register", mnem)
}
if k.idx == 0 {
return fmt.Errorf("k0 is not a usable mask register")
}
mask = k.idx
}
list, ok := lst.(RegList)
if !ok {
return fmt.Errorf("%s: second operand must be a four-register list", mnem)
}
if list.Lo.size != q.width {
return fmt.Errorf("%s: the register list must hold %d-bit vector registers", mnem, q.width*8)
}
dstReg, ok := dst.(Reg)
if !ok || !dstReg.isVec() {
return fmt.Errorf("%s: destination must be a vector register", mnem)
}
if dstReg.size != q.width {
return fmt.Errorf("%s: the destination must be a %d-bit vector register", mnem, q.width*8)
}
if !memOperand(mem) {
return fmt.Errorf("%s: the source must be a memory operand", mnem)
}
// The list owns V'VVVV; a scaled index in the EVEX-only half would fold
// its fifth bit into the same field the list's low register occupies.
if m, ok := mem.(Mem); ok && m.HasIndex && m.Index.idx >= 16 {
return fmt.Errorf("%s: an index register above Z15 has no EVEX bit free", mnem)
}
if sfx.zeroing && mask == 0 {
return fmt.Errorf("%s: zeroing (.Z) requires a mask register", mnem)
}
spec := evexSpec{mapSel: 2, opcode: q.opcode, w: 0, pp: 3, opdigit: -1, n: [3]int{16, 16, 16}}
// The vector length follows the destination (512-bit for the ZMM forms,
// 128-bit for the scalar ones), exactly as the oracle encodes it.
return e.emitEvexFields(spec, dstReg.vecLenBit(), dstReg.idx, list.Lo.idx, mem, mask, sfx)
}
// evexBcastSpec describes an EVEX broadcast (VPBROADCASTD/Q): the opcode
// depends on the source kind, a GPR source uses opReg, a memory source uses
// opMem with a disp8×N of n.
@@ -812,8 +897,10 @@ func isEvex(mnemUpper string) bool {
if _, ok := evexBcastTable[mnemUpper]; ok {
return true
}
_, ok := evexMoveTable[mnemUpper]
return ok
if _, ok := evexMoveTable[mnemUpper]; ok {
return true
}
return isEvexQuad(mnemUpper)
}
// evexRequired reports whether the operands force the EVEX encoding of a
@@ -995,6 +1082,14 @@ func (e *enc) encodeEvex(mnemUpper string, ops []Operand, sfx evexSuffix) error
return e.encodeEvexRM(spec, ops, 0, sfx)
}
spec, inTable := evexTable[mnemUpper]
if q, ok := evexQuadTable[mnemUpper]; ok {
// The quad-register family carries no rounding, SAE or broadcast;
// only masking and zeroing apply.
if sfx.sae || sfx.bcst || sfx.rounding >= 0 {
return fmt.Errorf("%s takes no rounding/SAE/broadcast suffix", mnemUpper)
}
return e.encodeEvexQuad(mnemUpper, q, ops, sfx)
}
if inTable {
if (sfx.rounding >= 0 || sfx.sae) && !evexRound[mnemUpper] {
return fmt.Errorf("%s: rounding/SAE is not supported for this instruction", mnemUpper)
+124
View File
@@ -810,3 +810,127 @@ func TestAvx512CorpusFamilies(t *testing.T) {
}
}
}
// TestEvexQuadRegisterGroundTruth pins the quad-register instructions (the
// 4FMAPS and 4VNNIW families) byte for byte against go tool asm: the memory
// source keeps r/m, the bracketed list's LOW register travels the inverted
// 5-bit V'VVVV field, the destination sits in reg, the opmask rides aaa and
// the vector length follows the destination (L'L=512 for the ZMM forms,
// 128 for the scalar ones) while the disp8×N multiplier stays 16 for every
// member. The x86 decoder has no view of these forms, so no decode check
// runs.
func TestEvexQuadRegisterGroundTruth(t *testing.T) {
sp := vreg(t, "RSP")
cases := []struct {
name string
mnem string
ops []Operand
want string
}{
{"V4FMADDPS 17(SP) [Z0-Z3] K2 Z0", "V4FMADDPS",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "K2"), vreg(t, "Z0")},
"62f27f4a9a842411000000"},
{"V4FMADDPS [Z10-Z13]", "V4FMADDPS",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z10"), vreg(t, "Z13")}, vreg(t, "K2"), vreg(t, "Z0")},
"62f22f4a9a842411000000"},
{"V4FMADDPS [Z20-Z23]", "V4FMADDPS",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z20"), vreg(t, "Z23")}, vreg(t, "K2"), vreg(t, "Z0")},
"62f25f429a842411000000"},
{"V4FMADDPS Z8 dst", "V4FMADDPS",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "K2"), vreg(t, "Z8")},
"62727f4a9a842411000000"},
{"V4FMADDPS disp8x16", "V4FMADDPS",
[]Operand{Ptr(sp, 64, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "K2"), vreg(t, "Z0")},
"62f27f4a9a442404"},
{"V4FMADDPS unmasked", "V4FMADDPS",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "Z0")},
"62f27f489a842411000000"},
{"V4FMADDSS 7(AX) [X0-X3] K5 X22", "V4FMADDSS",
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X0"), vreg(t, "X3")}, vreg(t, "K5"), vreg(t, "X22")},
"62e27f0d9bb007000000"},
{"V4FMADDSS (DI)", "V4FMADDSS",
[]Operand{Ptr(DI, 0, 8), RegList{vreg(t, "X0"), vreg(t, "X3")}, vreg(t, "K5"), vreg(t, "X22")},
"62e27f0d9b37"},
{"V4FMADDSS [X10-X13]", "V4FMADDSS",
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X10"), vreg(t, "X13")}, vreg(t, "K5"), vreg(t, "X22")},
"62e22f0d9bb007000000"},
{"V4FMADDSS [X20-X23]", "V4FMADDSS",
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X20"), vreg(t, "X23")}, vreg(t, "K5"), vreg(t, "X22")},
"62e25f059bb007000000"},
{"V4FMADDSS X30 dst", "V4FMADDSS",
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X0"), vreg(t, "X3")}, vreg(t, "K5"), vreg(t, "X30")},
"62627f0d9bb007000000"},
{"V4FMADDSS X3 dst", "V4FMADDSS",
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X0"), vreg(t, "X3")}, vreg(t, "K5"), vreg(t, "X3")},
"62f27f0d9b9807000000"},
{"V4FMADDSS disp8x16", "V4FMADDSS",
[]Operand{Ptr(AX, 16, 8), RegList{vreg(t, "X20"), vreg(t, "X23")}, vreg(t, "K5"), vreg(t, "X30")},
"62625f059b7001"},
{"V4FNMADDPS", "V4FNMADDPS",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "K2"), vreg(t, "Z0")},
"62f27f4aaa842411000000"},
{"V4FNMADDSS", "V4FNMADDSS",
[]Operand{Ptr(AX, 7, 8), RegList{vreg(t, "X0"), vreg(t, "X3")}, vreg(t, "K5"), vreg(t, "X22")},
"62e27f0dabb007000000"},
{"VP4DPWSSD", "VP4DPWSSD",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "K2"), vreg(t, "Z0")},
"62f27f4a52842411000000"},
{"VP4DPWSSDS unmasked", "VP4DPWSSDS",
[]Operand{Ptr(sp, 17, 8), RegList{vreg(t, "Z0"), vreg(t, "Z3")}, vreg(t, "Z0")},
"62f27f4853842411000000"},
}
for _, c := range cases {
code, err := Encode(c.mnem, c.ops...)
if err != nil {
t.Errorf("%s: Encode: %v", c.name, err)
continue
}
if got := hexCompact(code); got != c.want {
t.Errorf("%s: got %s, want %s", c.name, got, c.want)
}
}
}
// TestEvexQuadRegisterErrors pins the operand shapes the toolchain rejects:
// the register class the list and the destination take is fixed per
// instruction, the source is memory only, the opmask slot is positional and
// the list's low register owns V'VVVV.
func TestEvexQuadRegisterErrors(t *testing.T) {
sp := vreg(t, "RSP")
list := func(lo, hi string) RegList {
return RegList{vreg(t, lo), vreg(t, hi)}
}
cases := []struct {
name string
mnem string
ops []Operand
}{
{"X list on the PS form", "V4FMADDPS",
[]Operand{Ptr(sp, 0, 8), list("X0", "X3"), vreg(t, "K2"), vreg(t, "Z0")}},
{"Z list on the SS form", "V4FMADDSS",
[]Operand{Ptr(AX, 0, 8), list("Z0", "Z3"), vreg(t, "K5"), vreg(t, "X22")}},
{"Y destination", "V4FMADDPS",
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "K2"), vreg(t, "Y0")}},
{"register source", "V4FMADDPS",
[]Operand{vreg(t, "Z1"), list("Z0", "Z3"), vreg(t, "K2"), vreg(t, "Z0")}},
{"non-mask third operand", "V4FMADDPS",
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "Z4"), vreg(t, "Z0")}},
{"k0 mask", "V4FMADDPS",
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "K0"), vreg(t, "Z0")}},
{"K after the destination", "V4FMADDPS",
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "Z0"), vreg(t, "K2")}},
{"zeroing without a mask", "V4FMADDPS.Z",
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "Z0")}},
{"SAE suffix", "V4FMADDPS.SAE",
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3"), vreg(t, "K2"), vreg(t, "Z0")}},
{"high index source", "VP4DPWSSD",
[]Operand{Idx(DI, vreg(t, "X16"), 1, 0, 8), list("Z0", "Z3"), vreg(t, "K2"), vreg(t, "Z0")}},
{"short operand list", "V4FMADDPS",
[]Operand{Ptr(sp, 0, 8), list("Z0", "Z3")}},
}
for _, c := range cases {
if _, err := Encode(c.mnem, c.ops...); err == nil {
t.Errorf("%s: expected an error, got none", c.name)
}
}
}
+4 -4
View File
@@ -307,12 +307,12 @@ func TestStackGuardBytesLOONG64(t *testing.T) {
func TestStackGuardGOObjInternalCall(t *testing.T) {
for _, tt := range []struct {
src string
assemble func(*ast.File) (*Image, error)
assemble func(*ast.File, ...AssembleOption) (*Image, error)
}{
{"g_amd64.s", AssembleFile},
{"g_arm64.s", AssembleFileARM64},
{"g_riscv64.s", AssembleFileRISCV},
{"g_loong64.s", AssembleFileLOONG64},
{"g_arm64.s", func(f *ast.File, _ ...AssembleOption) (*Image, error) { return AssembleFileARM64(f) }},
{"g_riscv64.s", func(f *ast.File, _ ...AssembleOption) (*Image, error) { return AssembleFileRISCV(f) }},
{"g_loong64.s", func(f *ast.File, _ ...AssembleOption) (*Image, error) { return AssembleFileLOONG64(f) }},
} {
f, errs := parser.Parse(tt.src, "TEXT \u00b7callsmall(SB), $16-0\n\tCALL \u00b7other(SB)\n\tRET\nTEXT \u00b7other(SB), NOSPLIT, $0\n\tRET\n")
if len(errs) > 0 {
+74 -21
View File
@@ -192,6 +192,30 @@ func (e *enc) encodeMov(ops []Operand, size int) error {
}
return e.emit(i)
case TLSMem:
if !dstIsReg {
return fmt.Errorf("MOV: two memory operands")
}
// MOV r, off(TLS): the segment-prefixed absolute load, reg=dst,
// rm=src(tlsMem) through the SIB escape; the disp32 is the TLS slot
// offset with its R_TLSLE patch site.
i := newInstr(size, []byte{movRR(size)})
if err := setRM(i, dstReg, src, size); err != nil {
return err
}
return e.emit(i)
case SegAbs:
if !dstIsReg {
return fmt.Errorf("MOV: two memory operands")
}
// MOV r, 0x30(GS): the segment-absolute load.
i := newInstr(size, []byte{movRR(size)})
if err := setRM(i, dstReg, src, size); err != nil {
return err
}
return e.emit(i)
case Imm:
if dstIsReg {
v := int64(src)
@@ -232,11 +256,24 @@ func (e *enc) encodeMov(ops []Operand, size int) error {
i.imm = imm
return e.emit(i)
}
// MOV r/m, imm: 0xC6 (8-bit) / 0xC7 /0.
// MOV r/m, imm: 0xC6 (8-bit) / 0xC7 /0. An immediate in the
// destination slot is the absolute-address crash-store spelling,
// MOVL $0xf1, 0xf1: the parser reads the trailing bare constant
// as an immediate, and the store's disp32 carries the address.
op := byte(0xC7)
if size == 1 {
op = 0xC6
}
if d, ok := dst.(Imm); ok {
i := newInstr(size, []byte{op})
setSegAbs(i, 0, SegAbs{Disp: int64(d)})
immBytes, err := immediate(int64(src), size, false)
if err != nil {
return err
}
i.imm = immBytes
return e.emit(i)
}
i := newInstr(size, []byte{op})
if err := setRMDigit(i, 0, dst, size); err != nil {
return err
@@ -627,7 +664,16 @@ func (e *enc) encodeDoubleShift(base string, ops []Operand, size int) error {
func (e *enc) encodeImul(ops []Operand, size int) error {
switch len(ops) {
case 2:
// IMUL r, r/m: 0x0F 0xAF.
// Two shapes. The leading-immediate spelling IMUL $imm, r multiplies
// r in place (dst = rm = r): the shape GOROOT's clock code writes.
// Otherwise IMUL r, r/m: 0x0F 0xAF.
if imm, ok := ops[0].(Imm); ok {
dstReg, isReg := ops[1].(Reg)
if !isReg {
return fmt.Errorf("IMUL: destination must be a register")
}
return e.encodeImulImm(imm, dstReg, dstReg, size)
}
dstReg, ok := ops[1].(Reg)
if !ok {
return fmt.Errorf("IMUL: destination must be a register")
@@ -647,29 +693,36 @@ func (e *enc) encodeImul(ops []Operand, size int) error {
if !ok {
return fmt.Errorf("IMUL: immediate operand expected first")
}
// Plan 9 order: IMUL $imm, src, dst.
if fits8(int64(imm)) {
i := newInstr(size, []byte{0x6B})
if err := setRM(i, dstReg, ops[1], size); err != nil {
return err
}
i.imm = []byte{byte(int8(imm))}
return e.emit(i)
}
i := newInstr(size, []byte{0x69})
if err := setRM(i, dstReg, ops[1], size); err != nil {
return err
}
immBytes, err := immediate(int64(imm), size, false)
if err != nil {
return err
}
i.imm = immBytes
return e.emit(i)
// Plan 9 order: IMUL $imm, src, dst; the source stays a general
// r/m operand (setRM takes registers and memory alike).
return e.encodeImulImm(imm, ops[1], dstReg, size)
}
return fmt.Errorf("IMUL expects 2 or 3 operands, got %d", len(ops))
}
// encodeImulImm emits the immediate multiply: 0x6B with a sign-extended imm8
// when the value fits, 0x69 with a 32-bit immediate otherwise.
func (e *enc) encodeImulImm(imm Imm, rm Operand, dst Reg, size int) error {
if fits8(int64(imm)) {
i := newInstr(size, []byte{0x6B})
if err := setRM(i, dst, rm, size); err != nil {
return err
}
i.imm = []byte{byte(int8(imm))}
return e.emit(i)
}
i := newInstr(size, []byte{0x69})
if err := setRM(i, dst, rm, size); err != nil {
return err
}
immBytes, err := immediate(int64(imm), size, false)
if err != nil {
return err
}
i.imm = immBytes
return e.emit(i)
}
// --- PUSH / POP -------------------------------------------------------------
func (e *enc) encodePushPop(ops []Operand, size int, push bool) error {
+9 -7
View File
@@ -63,7 +63,10 @@ func toolAsmObject(t *testing.T, path, goarch string) []byte {
}
// oracleFuncCode extracts the non-package TEXT functions' code bytes from a
// toolchain object, keyed by the name the object records (pkg.name).
// toolchain object, keyed by the name the object records (pkg.name). Each
// function's span is its own symbol size: a toolchain object that follows
// the text with data symbols (the synthesised float-constant pool) would
// otherwise fold them into the last function's bytes.
func oracleFuncCode(t *testing.T, obj []byte) map[string][]byte {
t.Helper()
v := openGoobj(t, obj)
@@ -76,18 +79,13 @@ func oracleFuncCode(t *testing.T, obj []byte) map[string][]byte {
for _, bi := range []int{blkSymdef, blkHashed64def, blkHasheddef} {
preceding += len(v.blk(bi)) / symSize
}
total := preceding + len(nps)
out := make(map[string][]byte, len(nps))
for i, s := range nps {
if s.typ != kindSTEXT {
continue
}
start := le.Uint32(didx[4*(preceding+i):])
end := uint32(len(data))
if preceding+i+1 < total {
end = le.Uint32(didx[4*(preceding+i+1):])
}
out[s.name] = data[start:end]
out[s.name] = data[start : start+s.size]
}
return out
}
@@ -129,6 +127,10 @@ func TestDifferentialKernels(t *testing.T) {
{filepath.Join("..", "testdata", "verify", "datarel_amd64.s"), "", false},
{filepath.Join("..", "testdata", "verify", "divslash_amd64.s"), "", false},
{filepath.Join("..", "testdata", "verify", "semicolons_amd64.s"), "", false},
{filepath.Join("..", "testdata", "verify", "quadreg_amd64.s"), "", false},
{filepath.Join("..", "testdata", "verify", "floatimm_amd64.s"), "", false},
{filepath.Join("..", "testdata", "verify", "bookkeep_amd64.s"), "", false},
{filepath.Join("..", "testdata", "verify", "forms_amd64.s"), "", false},
{filepath.Join("..", "testdata", "verify", "datarel_arm64.s"), "arm64", true},
{filepath.Join("..", "testdata", "verify", "divslash_arm64.s"), "arm64", true},
} {
+44 -2
View File
@@ -150,6 +150,18 @@ func (img *Image) Bytes() []byte {
return append(out, img.Data...)
}
// AssembleOption adjusts the file-level assembly context.
type AssembleOption func(*linkInfo)
// WithGOOS selects the target operating system for the forms that depend on
// it, the TLS access shape above all: linux and freebsd take the
// one-instruction form, windows and plan9 keep the two-instruction load.
func WithGOOS(goos string) AssembleOption {
return func(l *linkInfo) {
l.goos = goos
}
}
// AssembleFile assembles every TEXT function of a parsed file and lays out
// its static symbols (GLOBL/DATA) in a data section behind the code. Each
// reference to a file-local static symbol becomes a RIP-relative load whose
@@ -157,7 +169,7 @@ func (img *Image) Bytes() []byte {
// GLOBL defines is recorded as an external relocation (Externals) with its
// displacement left zero, the object-file emitters resolve it at link
// time, while the raw image (Bytes) cannot represent it.
func AssembleFile(f *ast.File) (*Image, error) {
func AssembleFile(f *ast.File, opts ...AssembleOption) (*Image, error) {
dataSyms, err := collectData(f)
if err != nil {
return nil, err
@@ -166,7 +178,18 @@ func AssembleFile(f *ast.File) (*Image, error) {
for _, d := range dataSyms {
known[d.name] = true
}
// TEXT symbols are file-level definitions too: a symbol immediate
// ($fn(SB)) may name one, exactly as a data reference names a GLOBL.
for _, d := range f.Decls {
if t, ok := d.(*ast.Text); ok {
known[t.Name.Name] = true
}
}
link := &linkInfo{symbols: known, allowExternal: true}
for _, o := range opts {
o(link)
}
poolSeen := map[string]bool{}
img := &Image{Symbols: map[string]int{}, SourcePath: f.Path}
textOff := map[string]int{}
@@ -180,7 +203,26 @@ func AssembleFile(f *ast.File) (*Image, error) {
if !ok {
continue
}
code, patches, labels, steps, lines, err := assemble(t, link)
code, patches, labels, steps, lines, pool, err := assemble(t, link)
if err != nil {
return nil, fmt.Errorf("%s: %w", t.Name.Name, err)
}
// The pooled floating-point constants join the declared data as
// read-only symbols, deduplicated across the file (the toolchain
// synthesises the same symbols into its rodata).
for _, entry := range pool {
if poolSeen[entry.name] {
continue
}
poolSeen[entry.name] = true
dataSyms = append(dataSyms, dataSym{
name: entry.name,
buf: entry.data,
size: len(entry.data),
rodata: true,
dupok: true,
})
}
if err != nil {
return nil, fmt.Errorf("%s: %w", t.Name.Name, err)
}
+46
View File
@@ -14,6 +14,51 @@ type Imm int64
func (Imm) isOperand() {}
// RegList is a bracketed register range, [Z0-Z3]: the four-register source
// of the 4FMAPS and 4VNNIW families. The EVEX emit path carries the list's
// low register through the inverted 5-bit V'VVVV field; the three higher
// registers are implied by the instruction, so only the pair travels here.
type RegList struct {
Lo Reg
Hi Reg // implied by the encoding; Lo.idx+3 by construction
}
func (RegList) isOperand() {}
// FloatImm is a floating-point immediate ($-1.0). The SSE mnemonics whose
// encoding takes an XMM/memory source at that position rewrite it as a read
// from a read-only pool constant ($f64.<hex> or $f32.<hex>), the toolchain's
// own behaviour; every other instruction rejects it.
type FloatImm struct {
Text string // the numeric text as written, sign excluded
Neg bool // a leading minus
}
func (FloatImm) isOperand() {}
// TLSMem is a thread-local access, the source form off(base)(TLS*1) with the
// base dropped: the toolchain's one-instruction TLS rewrite assembles it as
// the segment-prefixed absolute whose disp32 carries an R_TLS_LE patch site
// (the linker fills the TLS slot offset).
type TLSMem struct {
Disp int64
Size int
Seg byte // the segment override: FS (0x64) or GS (0x65) on windows
}
func (TLSMem) isOperand() {}
// SegAbs is a segment-absolute access, 0x30(GS): the segment override
// prefixes a disp32 absolute reference with no relocation. The base
// register spellings GS and FS produce it.
type SegAbs struct {
Disp int64
Size int
Seg byte // 0x64 FS, 0x65 GS
}
func (SegAbs) isOperand() {}
// Mem is a memory operand of the form disp(base)(index*scale).
type Mem struct {
Base Reg
@@ -23,6 +68,7 @@ type Mem struct {
Size int // operand width in bytes
HasBase bool
HasIndex bool
Seg byte // segment override prefix (0x64 FS, 0x65 GS); 0 = none
}
func (Mem) isOperand() {}
+11 -3
View File
@@ -5,6 +5,7 @@ package main
import (
"os"
"os/exec"
"path/filepath"
"strings"
"testing"
@@ -393,13 +394,20 @@ func TestRunCorpusAuditGOOS(t *testing.T) {
}
// TestGenerateGoAsmHeaderRuntime pins the generator against the real thing:
// the runtime package, whose header the toolchain's own -asmhdr output was
// sampled from. Skipped in short mode: it type-checks the whole package.
// the runtime package of the ambient toolchain, whose header the toolchain's
// own -asmhdr output was sampled from. Skipped in short mode: it type-checks
// the whole package. The GOROOT comes from the go command itself, so the
// test follows whatever toolchain the host provides.
func TestGenerateGoAsmHeaderRuntime(t *testing.T) {
if testing.Short() {
t.Skip("type-checks the whole runtime package")
}
dir, err := generateGoAsmHeader("/usr/local/go/src/runtime", "", "amd64", t.TempDir())
out, err := exec.Command("go", "env", "GOROOT").Output()
if err != nil {
t.Skipf("no Go toolchain: %v", err)
}
runtimeDir := filepath.Join(strings.TrimSpace(string(out)), "src", "runtime")
dir, err := generateGoAsmHeader(runtimeDir, "", "amd64", t.TempDir())
if err != nil {
t.Fatalf("generateGoAsmHeader(runtime): %v", err)
}
+43 -15
View File
@@ -38,7 +38,7 @@ import (
// construction and are excluded from the diff; the other architectures list
// their conditional branches outright.
func cmdAuditInstructions(args []string) error {
fs := newCommand("audit-instructions", "gasm audit-instructions [--corpus [dir]] [-I dir] [amd64|arm64|riscv64|loong64]", `
fs := newCommand("audit-instructions", "gasm audit-instructions [--corpus [dir]] [--list] [-I dir] [amd64|arm64|riscv64|loong64]", `
Compare the gasm encoder for the given architecture (default amd64) against
go tool asm and print the diff: superset encodings (gasm-only, shippable via
gasm asm --format goobj) and known-but-unencodable names (the backlog). The
@@ -55,16 +55,19 @@ toolchain probing. A file whose name carries a recognisable _arch suffix is
attempted for that architecture; a file without one is attempted for all
four, exactly as a GOARCH build would compile it. The report gives the
per-architecture pass rates and the most common failure reasons, which drive
the encodability backlog by frequency rather than by table order.
the encodability backlog by frequency rather than by table order. With
-list the report also prints every failing file with its reason, per
architecture.
`)
corpus := fs.Bool("corpus", false, "assemble a corpus of .s files and report pass rates and failure reasons")
list := fs.Bool("list", false, "with --corpus, list every failing file with its reason, per architecture")
var dirs includeDirs
fs.Var(&dirs, "I", "directory to search for #include files (may be repeated)")
if err := fs.Parse(args); err != nil {
return err
}
if *corpus {
return cmdAuditCorpus(fs.Args(), dirs)
return cmdAuditCorpus(fs.Args(), dirs, *list)
}
archName := "amd64"
switch n := len(fs.Args()); {
@@ -403,20 +406,30 @@ type corpusTally struct {
assembled int
reasons map[string]int // failure reason → count
example map[string]string // failure reason → one representative file
fails []corpusFailure // every failure, in file order, for --list
}
func (t *corpusTally) fail(path, reason string) {
// corpusFailure is one failed attempt, recorded for the --list report.
type corpusFailure struct {
path string
reason string
detail string
}
func (t *corpusTally) fail(path string, err error) {
reason := corpusReason(err)
t.reasons[reason]++
if t.example[reason] == "" {
t.example[reason] = path
}
t.fails = append(t.fails, corpusFailure{path: path, reason: reason, detail: firstLine(err.Error())})
}
// cmdAuditCorpus implements audit-instructions --corpus. The include
// directories carry #include resolution over a corpus whose files refer to
// headers such as GOROOT/pkg/include, the same -I a toolchain comparison
// needs.
func cmdAuditCorpus(args []string, dirs includeDirs) error {
func cmdAuditCorpus(args []string, dirs includeDirs, list bool) error {
if len(args) > 1 {
return &usageError{fmt.Errorf("audit-instructions --corpus takes at most one directory argument")}
}
@@ -454,7 +467,7 @@ func cmdAuditCorpus(args []string, dirs includeDirs) error {
if err != nil {
return err
}
printCorpusStats(stats)
printCorpusStats(stats, list)
return nil
}
@@ -645,21 +658,22 @@ func runCorpusAudit(root string, dirs includeDirs) (*corpusStats, error) {
hdrDir, err := hdr.dirFor(pkgDir, goos, goarchName(tg.a))
if err != nil {
ok = false
t.fail(path, corpusReason(err))
t.fail(path, err)
continue
}
f, errs := parser.ParseWithOptions(path, src, parser.Options{
Expand: true,
IncludeDirs: append(slices.Clone(dirs), hdrDir),
Predefines: platformPredefinesFor(goarchName(tg.a), goos),
})
if len(errs) > 0 {
ok = false
t.fail(path, corpusReason(errs[0]))
t.fail(path, errs[0])
continue
}
if _, err := assembleFile(tg.a, f); err != nil {
if _, err := assembleFile(tg.a, f, goos); err != nil {
ok = false
t.fail(path, corpusReason(err))
t.fail(path, err)
continue
}
t.assembled++
@@ -670,21 +684,28 @@ func runCorpusAudit(root string, dirs includeDirs) (*corpusStats, error) {
continue
}
f, errs := parser.ParseWithOptions(path, src, parser.Options{Expand: true, IncludeDirs: dirs})
ok := true
for _, i := range wanted {
tg, t := targets[i], tallies[i]
t.attempted++
// The parse carries the target's platform predefines, so it
// cannot be shared across targets the way a header-free file's
// could: a #ifdef GOARCH_arm block must be live on arm64 and
// dead everywhere else.
f, errs := parser.ParseWithOptions(path, src, parser.Options{
Expand: true,
IncludeDirs: dirs,
Predefines: platformPredefinesFor(goarchName(tg.a), goos),
})
var err error
if len(errs) > 0 {
err = errs[0] // a parse failure is a failure for every target
} else {
_, err = assembleFile(tg.a, f)
_, err = assembleFile(tg.a, f, goos)
}
if err != nil {
ok = false
t.fail(path, corpusReason(err))
t.fail(path, err)
continue
}
t.assembled++
@@ -706,7 +727,7 @@ func runCorpusAudit(root string, dirs includeDirs) (*corpusStats, error) {
}
// printCorpusStats renders the corpus audit report.
func printCorpusStats(s *corpusStats) {
func printCorpusStats(s *corpusStats, list bool) {
fmt.Printf("corpus %s: %d files (%d generic, attempted for all architectures; %d named for other Go ports, never attempted)\n", s.root, s.files, s.generic, s.otherPort)
// The rate is over the files a supported build would attempt: the
// other ports' files sit in the count for completeness but can never
@@ -721,6 +742,13 @@ func printCorpusStats(s *corpusStats) {
fmt.Printf(" %4d %s\n", t.reasons[r], r)
fmt.Printf(" e.g. %s\n", t.example[r])
}
if !list {
continue
}
for _, f := range t.fails {
fmt.Printf(" FAIL %s\n", f.path)
fmt.Printf(" %s: %s\n", f.reason, f.detail)
}
}
}
+1 -1
View File
@@ -84,7 +84,7 @@ func disSource(path string, target arch.Arch) int {
if len(errs) > 0 {
return 1
}
img, err := assembleFile(target, f)
img, err := assembleFile(target, f, "")
if err != nil {
fmt.Fprintf(os.Stderr, "gasm dis: %v\n", err)
return 1
+34 -11
View File
@@ -574,7 +574,7 @@ naming the package.
defer cleanup()
dirs = append(dirs, hdrDir)
}
f, errs := parser.ParseWithOptions(path, src, parser.Options{Expand: true, IncludeDirs: dirs})
f, errs := parser.ParseWithOptions(path, src, parser.Options{Expand: true, IncludeDirs: dirs, Predefines: platformPredefinesFor(string(targetArch), goos)})
for _, e := range errs {
fmt.Fprintf(os.Stderr, "%s: %v\n", path, e)
}
@@ -582,7 +582,7 @@ naming the package.
return 1
}
img, err := assembleFile(targetArch, f)
img, err := assembleFile(targetArch, f, goos)
if err != nil {
fmt.Fprintf(os.Stderr, "%s: %v\n", path, err)
return 1
@@ -789,11 +789,34 @@ e.g. --map wideCopyAVX2=wideCopyAVX512 pairs the two regardless of suffix.
return 1
}
// platformPredefines mirrors the go command's assembler invocation, which
// defines GOOS_<goos> and GOARCH_<arch> as -D macros: GOROOT headers
// (go_tls.h, asm_riscv64.h) select their platform blocks with #ifdef on
// exactly those names, so an assembler without them cannot see the platform
// definitions at all.
func platformPredefines(goarch, goos string) map[string]string {
return map[string]string{
"GOARCH_" + goarch: "1",
"GOOS_" + goos: "1",
}
}
// platformPredefinesFor resolves the ambient GOOS the way a build would: a
// file whose name carries one (sys_darwin_arm64.s) is compiled for that GOOS
// and nothing else.
func platformPredefinesFor(goarch string, fileGoos string) map[string]string {
goos := fileGoos
if goos == "" {
goos = runtime.GOOS
}
return platformPredefines(goarch, goos)
}
// assembleFile assembles a parsed file for the given architecture and returns the image.
func assembleFile(targetArch arch.Arch, f *ast.File) (*asm.Image, error) {
func assembleFile(targetArch arch.Arch, f *ast.File, goos string) (*asm.Image, error) {
switch targetArch {
case arch.AMD64:
return asm.AssembleFile(f)
return asm.AssembleFile(f, asm.WithGOOS(goos))
case arch.RISCV:
return asm.AssembleFileRISCV(f)
case arch.ARM64:
@@ -813,18 +836,18 @@ func assemblePath(path string, forced arch.Arch, dirs includeDirs) (*asm.Image,
if err != nil {
return nil, err
}
f, errs := parser.ParseWithOptions(path, src, parser.Options{Expand: true, IncludeDirs: dirs})
target := forced
if target == arch.Unknown {
target = arch.FromFilename(path)
}
f, errs := parser.ParseWithOptions(path, src, parser.Options{Expand: true, IncludeDirs: dirs, Predefines: platformPredefinesFor(string(target), "")})
for _, e := range errs {
fmt.Fprintf(os.Stderr, "%s: %v\n", path, e)
}
if len(errs) > 0 {
return nil, fmt.Errorf("parse errors")
}
target := forced
if target == arch.Unknown {
target = arch.FromFilename(path)
}
return assembleFile(target, f)
return assembleFile(target, f, "")
}
// printByteDiff shows the first few byte differences between two code blocks.
@@ -920,7 +943,7 @@ func cmdVerifyNonJIT(path string, targetArch arch.Arch, groundTruth, profile boo
if len(errs) > 0 {
return 1
}
img, err := assembleFile(targetArch, f)
img, err := assembleFile(targetArch, f, "")
if err != nil {
fmt.Fprintf(os.Stderr, "gasm verify: %v\n", err)
return 1
+14
View File
@@ -119,6 +119,20 @@ identifier is a register or a label is an *architecture* question, so it is
left to `arch` and resolved in the lint/lsp layers. This keeps the parser
arch-agnostic and its output deterministic.
### Optional preprocessing
With `Options{Expand: true}` the parser runs a pre-parse pass
(`preproc.go`) that splices `#include` files (the source directory, then the
`-I` directories), expands object and parameterised `#define` macros,
applies `#undef` and the `#ifdef`/`#ifndef`/`#else`/`#endif` family, and
folds constant expressions left in operands. The go command's platform
macros (`GOARCH_<arch>`, `GOOS_<goos>`) arrive through `Options.Predefines`.
The assembly path (`asm`, `diff`, `audit`) expands; `lint`, `fmt` and the
language server read the raw file. The command layer adds the go_asm.h
generator (`asmhdr.go`): a file that includes go_asm.h gets the package's
defines type-checked out of its Go files for the target architecture and
GOOS, with no compiler in the loop.
### `arch`
Register files are generated programmatically (the regular `R8`-`R15`,
+4 -2
View File
@@ -365,7 +365,7 @@ add: 16 bytes, args=24, frame=0 NOSPLIT
## audit-instructions
```text
Usage: gasm audit-instructions [--corpus [dir]] [-I dir] [amd64|arm64|riscv64|loong64]
Usage: gasm audit-instructions [--corpus [dir]] [--list] [-I dir] [amd64|arm64|riscv64|loong64]
```
Compare the gasm encoder for the given architecture (default amd64) against the
@@ -403,7 +403,9 @@ architecture; a file without one is attempted for all four, exactly as a
The report gives the headline number (files
that assemble for every target architecture), the per-architecture pass rates
and the most common failure reasons with one representative file each, which
drive the encodability backlog by frequency rather than by table order. A run
drive the encodability backlog by frequency rather than by table order. With
`--list` the report additionally prints every failing file with its failure
reason, per architecture. A run
over GOROOT takes under a second.
```sh
+657
View File
@@ -0,0 +1,657 @@
# The GOOBJ object file format
This document is a complete specification of GOOBJ, the object file format
that the Go toolchain's assembler, compiler and linker exchange, written for
implementers of independent producers and consumers. It documents the format
as shipped by Go 1.27.1, identified by the magic string `"\x00go120ld"`.
No comparable document exists upstream. The format is defined only by the
source of the `cmd/internal/goobj` package inside the toolchain tree, it is an
internal interface with no stability promise, and it can change in any
release. This specification was therefore produced by reverse engineering
that source and by parsing real objects produced by `go tool asm` and
`go tool compile`, byte for byte, against the layout described here. Within
gasm-devkit it is kept honest by the differential tests in `asm/goobj_test.go`
and `asm/link_test.go`, which compare `gasm asm --format goobj` output against
the toolchain's own products and feed gasm objects to `go build`.
Every numeric value in this document, every block index, structure size, flag
bit, type code and relocation number, was read from the Go 1.27.1 source at
`/usr/local/go/src/cmd/internal/goobj`, `cmd/internal/obj` and
`cmd/internal/objabi`, and exercised against assembled objects.
## Containers
The unit this document specifies is the **object**: one package's worth of
symbols, relocations and data. An object is never consumed naked. Two
wrappers exist in practice, and the linker dispatches on the first bytes of
the file.
**The bare object**, written by `go tool asm`:
```text
"go object linux amd64 go1.27.1 GOAMD64=v1 X:regabiwrappers,...\n"
"!\n"
<GOOBJ blob>
```
The first line is the toolchain configuration string, produced by
`objabi.HeaderString`: `go object`, the GOOS, the GOARCH, the toolchain
version, an optional architecture qualifier such as `GOAMD64=v1`, and
`X:` followed by the enabled experiments, comma separated. The linker requires
this line to match its own configuration exactly and rejects the file
otherwise; the `-f` linker flag waives the check. Header lines may be
followed by export data delimited by `$$` markers; the header region always
ends at the first line consisting of exactly `!`, and the GOOBJ blob starts
immediately after that line.
**The package archive**, written by the compiler output pipeline and consumed
by `go build`: the classic `ar` format, magic `!<arch>\n`, with the export
data in a `__.PKGDEF` member and one or more objects as further members, each
carrying the bare-object structure above. `go tool pack` creates and
inspects such archives.
| Consumer | Role |
|---|---|
| `cmd/asm` | writes objects from `.s` files |
| `cmd/compile` | writes objects from Go source |
| `cmd/link` | reads objects and archives, produces executables |
| `cmd/nm`, `cmd/objdump` | read objects through `cmd/internal/objfile` |
## Conventions
- All integers are **little endian**.
- There is **no alignment or padding** anywhere in the file; structures follow
one another byte by byte.
- Every offset stored in the file is **relative to the first byte of the GOOBJ
blob**, not to the start of the container.
- The blob opens with a 96 byte header that carries the byte offset of every
block. A block's length is the difference between its own offset and the
next block's, so the offset array is the only index the format needs.
### Layout overview
```mermaid
flowchart TB
A[Container header line and ! terminator] --> B[File header, 96 bytes]
B --> C[String table, implicit region]
C --> D[Autolib]
D --> E[PkgIndex]
E --> F[Files]
F --> G[Symbol definition arrays: Symdef, Hashed64def, Hasheddef, Nonpkgdef, Nonpkgref]
G --> H[RefFlags]
H --> I[Hash64 and Hash]
I --> J[RelocIndex, AuxIndex, DataIndex]
J --> K[Relocs]
K --> L[Aux]
L --> M[Data]
M --> N[RefNames]
N --> O[BlkEnd marks the end of the blob]
```
## The file header
Exactly 96 bytes: 8 magic, 8 fingerprint, 4 flags, and 19 four byte block
offsets.
| Offset | Size | Field | Meaning |
|---|---|---|---|
| 0 | 8 | Magic | `"\x00go120ld"`. A reader rejects anything else. The digits are the format version and have moved before; a new toolchain release may move them again. |
| 8 | 8 | Fingerprint | Identifies the package build. The compiler writes a hash of the export data; the assembler leaves all zero. The linker compares this against the fingerprint recorded by importers. |
| 16 | 4 | Flags | Bit field, see below. |
| 20 | 76 | Offsets | 19 `uint32` entries, one per block index 0 to 18. |
Header flags:
| Bit | Value | Name | Meaning |
|---|---|---|---|
| 0 | 1 | ObjFlagShared | built with `-shared` |
| 1 | 2 | reserved | was `ObjFlagNeedNameExpansion`, now unused |
| 2 | 4 | ObjFlagFromAssembly | produced from assembly source; `go tool asm` and gasm set this |
| 3 | 8 | ObjFlagUnlinkable | package path is invalid, the linker refuses to link |
| 4 | 16 | ObjFlagStd | standard library package |
### Block indices
The offset array is indexed by these constants, in file order:
| Index | Constant | Contents |
|---|---|---|
| 0 | BlkAutolib | imported packages |
| 1 | BlkPkgIndex | referenced packages, indexed |
| 2 | BlkFile | source file names |
| 3 | BlkSymdef | symbol definitions, package scope |
| 4 | BlkHashed64def | short hashed definitions |
| 5 | BlkHasheddef | hashed definitions |
| 6 | BlkNonpkgdef | non-package definitions |
| 7 | BlkNonpkgref | non-package references |
| 8 | BlkRefFlags | flags of referenced symbols |
| 9 | BlkHash64 | 8 byte hashes for short hashed definitions |
| 10 | BlkHash | 16 byte hashes for hashed definitions |
| 11 | BlkRelocIndex | per symbol relocation start index |
| 12 | BlkAuxIndex | per symbol aux start index |
| 13 | BlkDataIndex | per symbol data offset |
| 14 | BlkReloc | relocations |
| 15 | BlkAux | aux symbol entries |
| 16 | BlkData | symbol payloads |
| 17 | BlkRefName | names of referenced symbols, for tools |
| 18 | BlkEnd | no contents; its offset is the end of the blob |
## The string table
There is no block index for strings. The table occupies the implicit region
between the end of the header (offset 96) and `Offsets[BlkAutolib]`, and every
string offset in the file points into that region. The writer de-duplicates:
each distinct string is stored once, in first-use order, and the empty string
is always the first entry, so its reference is length 0 and offset 96.
A **string reference** is 8 bytes: `uint32` length, then `uint32` absolute
offset of the bytes. The bytes are stored raw, with no terminator.
## Symbol references and the package index
A **symbol reference** (SymRef) is 8 bytes: two `uint32`, `PkgIdx` and
`SymIdx`. The pair `{0, 0}` means nil. `PkgIdx` says which array the symbol
lives in:
| Value | Constant | SymIdx indexes |
|---|---|---|
| 0 | PkgIdxInvalid | never valid in a written file |
| 1 and up, ascending | (imported packages) | the SymbolDefs array of the package named at PkgIndex entry `PkgIdx` |
| 0x7ffffffb | PkgIdxSelf | this object's Symdef array |
| 0x7ffffffc | PkgIdxBuiltin | the compiler's builtin table, see Builtins |
| 0x7ffffffd | PkgIdxHashed | this object's Hasheddef array |
| 0x7ffffffe | PkgIdxHashed64 | this object's Hashed64def array |
| 0x7fffffff | PkgIdxNone | NonPkgDefs, overflowing into NonPkgRefs |
Assignment rules, as the toolchain performs them:
- Every definition a package exports to the linker by index lands in Symdefs
with PkgIdxSelf. The compiler puts its functions and data here; the
assembler puts only its file-local static symbols here, everything else by
name, see below.
- External package references take indices 1, 2, 3, in order of first
reference during assembly; the package names go into PkgIndex at those
indices, entry 0 is the empty package and is never referenced.
- References to the compiler's builtin functions become PkgIdxBuiltin with
SymIdx set to the builtin's index.
- A symbol referenced **by name** rather than by index becomes PkgIdxNone and
its index counts through NonPkgDefs first, then continues into NonPkgRefs.
A producer must emit the definitions it made in NonPkgDefs and the pure
references in NonPkgRefs.
- The assembler's rule, from `cmd/internal/obj/sym.go`: every assembly symbol
is referenced by name, PkgIdxNone, **except** file-local static symbols,
whose names carry `<>` and which are referenced by index. The compiler also
forces references by name for symbols marked `//go:linkname` and for any
symbol with the DUPOK attribute, which the linker de-duplicates by name.
## Symbol definition entries
The five definition and reference arrays (block indices 3 to 7) share one
element layout, 21 bytes:
| Offset | Size | Field | Meaning |
|---|---|---|---|
| 0 | 8 | Name | string reference |
| 8 | 2 | ABI | see table below |
| 10 | 1 | Type | symbol kind, see the kind table |
| 11 | 1 | Flag | bit field, see below |
| 12 | 1 | Flag2 | second bit field, see below |
| 13 | 4 | Siz | payload size in bytes, `uint32` |
| 17 | 4 | Align | alignment the linker must honour, `uint32` |
The Name is a real string reference for hand-written symbols. The auxiliary
symbols the toolchain generates per function, the FuncInfo payload, the DWARF
entries, have empty names: length 0, and their identity is only via the Aux
entries that point at them by index.
### The ABI field
| Value | Meaning |
|---|---|
| 0 | ABI0, the stack based ABI, the ABI of every hand-written assembly function |
| 1 | ABIInternal, the register ABI of compiler-generated functions |
| 0xffff | static, a file-local symbol (`name<>(SB)`), `SymABIstatic` |
### The Flag byte
| Bit | Value | Name | Meaning |
|---|---|---|---|
| 0 | 1 | SymFlagDupok | duplicates allowed, the linker merges them |
| 1 | 2 | SymFlagLocal | file-local |
| 2 | 4 | SymFlagTypelink | belongs in the typelink table |
| 3 | 8 | SymFlagLeaf | leaf function |
| 4 | 16 | SymFlagNoSplit | no stack-split preamble |
| 5 | 32 | SymFlagReflectMethod | `//go:reflectmethod` reachability |
| 6 | 64 | SymFlagGoType | a Go type descriptor, `type:` name and SRODATA |
Note that NoSplit is not reserved for explicit `NOSPLIT` declarations. On
amd64 the assembler itself marks any function whose frame is below
`abi.StackSmall` and whose body calls nothing that needs stack as NoSplit and
omits the split check, so a `TEXT` without `NOSPLIT` can still carry the bit.
### The Flag2 byte
| Bit | Value | Name | Meaning |
|---|---|---|---|
| 0 | 1 | SymFlagUsedInIface | type or itab reachable through an interface |
| 1 | 2 | SymFlagItab | an itab, `go:itab.` name and SRODATA |
| 2 | 4 | SymFlagDict | a generic dictionary symbol |
| 3 | 8 | SymFlagPkgInit | package initialisation function |
| 4 | 16 | SymFlagLinkname | reachable through `//go:linkname`; the assembler also sets it on `main.main` |
| 5 | 32 | SymFlagLinknameStd | linkname into the standard library |
| 6 | 64 | SymFlagABIWrapper | ABI transition wrapper |
| 7 | 128 | SymFlagWasmExport | `//go:wasmexport` target |
### The Type byte: symbol kinds
Values of `objabi.SymKind`, in numeric order:
| Value | Name | Meaning |
|---|---|---|
| 0 | Sxxx | invalid zero value |
| 1 | STEXT | executable code |
| 2 | STEXTFIPS | executable code, FIPS section |
| 3 | SRODATA | read only data |
| 4 | SRODATAFIPS | read only data, FIPS section |
| 5 | SNOPTRDATA | data without pointers |
| 6 | SNOPTRDATAFIPS | data without pointers, FIPS section |
| 7 | SDATA | data, may contain pointers |
| 8 | SDATAFIPS | data, FIPS section |
| 9 | SBSS | zero initialised data |
| 10 | SNOPTRBSS | zero initialised data without pointers |
| 11 | STLSBSS | thread local zero initialised data |
| 12 | SDWARFCUINFO | DWARF compile unit information |
| 13 | SDWARFCONST | DWARF constants |
| 14 | SDWARFFCN | DWARF function entry |
| 15 | SDWARFABSFCN | DWARF absolute function entry |
| 16 | SDWARFTYPE | DWARF type information |
| 17 | SDWARFVAR | DWARF variable information |
| 18 | SDWARFRANGE | DWARF range lists |
| 19 | SDWARFLOC | DWARF location lists |
| 20 | SDWARFLINES | DWARF line programs |
| 21 | SDWARFADDR | DWARF address table |
| 22 | SLIBFUZZER_8BIT_COUNTER | libFuzzer coverage counter |
| 23 | SCOVERAGE_COUNTER | coverage counter |
| 24 | SCOVERAGE_AUXVAR | coverage auxiliary variable |
| 25 | SSEHUNWINDINFO | Windows SEH unwind information |
## Referenced symbol flags (RefFlags)
Element size 10 bytes, one per referenced external indexed symbol that
carries a non-zero Flag2:
| Offset | Size | Field |
|---|---|---|
| 0 | 8 | Sym, a SymRef into another package |
| 8 | 1 | Flag, always 0 in current writers |
| 9 | 1 | Flag2, only SymFlagUsedInIface is ever written |
The linker uses these to preserve reachability of interface conversions
across package boundaries. Entries with no flags are omitted entirely.
## Hashes
**Hash64**, block 9: one `uint64` per Hashed64def entry, in array order. Not
a hash at all: the writer copies the **first 8 bytes of the symbol's
payload**. Only symbols whose content-hash section byte is 0 may use the
short form.
**Hash**, block 10: 16 bytes per Hasheddef entry: the first 16 bytes of a
SHA-256 computation over a seed byte `0x01` followed by the hash input. The
input, from `cmd/internal/obj/objfile.go`:
1. the payload size, little endian `uint64`;
2. the section byte, one of `t` for STEXT, `f` for STEXTFIPS, `P` for pcdata,
`F` for the `go:func.*` and `go:funcrel.*` families, `T` for `type:`
symbols, otherwise 0;
3. for text symbols, the symbol name, which keeps distinct functions from
merging;
4. the payload with trailing zero bytes trimmed;
5. for each relocation: a 14 byte record, offset `uint32`, size `uint8`,
low type byte `uint8`, addend `int64`, followed by an encoding of the
target: tag byte 0 then the target's short hash, tag 1 then its full
hash, tag 2 then its expanded name, tag 3 then its builtin index, or,
for PkgIdxSelf and imported packages, no tag, then the package path
and the symbol index.
Two symbols with equal hashes are interchangeable at link time, which is what
makes content addressing work. A producer that computes these hashes wrongly
produces objects that link but de-duplicate wrongly; gasm verifies them by
byte comparison against `go tool asm`.
## The index arrays
Three arrays of `uint32`, one element per **defined** symbol plus one final
element, in the order Symdefs, Hashed64defs, Hasheddefs, NonPkgDefs. With N
defined symbols, each array holds N + 1 entries, and the entry at N is the
total.
- RelocIndex: entry i is where symbol i's relocations start in BlkReloc;
entry i + 1 minus entry i is its count.
- AuxIndex: the same construction over BlkAux.
- DataIndex: entry i is the byte offset of symbol i's payload within BlkData;
the count is the difference of neighbours.
The toolchain writes relocations grouped per symbol in definition order, and
sorts each symbol's relocations by their Off field first. A producer that
skips the sort produces objects the linker still accepts, but that no longer
compare byte-for-byte with the toolchain's output.
## Relocations
Element size 23 bytes:
| Offset | Size | Field | Meaning |
|---|---|---|---|
| 0 | 4 | Off | patch position, bytes from the start of the symbol's payload, `int32` |
| 4 | 1 | Siz | patch width in bytes |
| 5 | 2 | Type | relocation type, `uint16`, see the table |
| 7 | 8 | Add | addend, `int64` |
| 15 | 8 | Sym | target SymRef |
The computed value `payload[Off:Off+Siz] += address(Sym) + Add` in the
flavour the type prescribes is the linker's job; the object only records the
request. A size 0 relocation patches nothing and exists purely as a marker
for the linker's reachability analysis.
### Relocation types
Values of `objabi.RelocType`. The assembler and compiler emit the generic
ones plus their own architecture's family; the rest exist for other ports and
for the linker itself.
| Value | Name | Meaning |
|---|---|---|
| 1 | R_ADDR | absolute address |
| 2 | R_ADDRPOWER | ppc64: high adjusted plus low 16 bits across two D-form instructions |
| 3 | R_ADDRARM64 | arm64: adrp plus add pair |
| 4 | R_ADDRMIPS | mips: low 16 bits of an external address |
| 5 | R_ADDROFF | 32-bit offset from the section start to the symbol |
| 6 | R_SIZE | size of the referenced symbol |
| 7 | R_CALL | direct call, PC relative |
| 8 | R_CALLARM | arm: call with a shifted 24-bit field |
| 9 | R_CALLARM64 | arm64: BL |
| 10 | R_CALLIND | indirect call marker |
| 11 | R_CALLPOWER | ppc64: call |
| 12 | R_CALLMIPS | mips: non-PC-relative call target |
| 13 | R_CONST | constant value of the symbol |
| 14 | R_PCREL | PC relative displacement |
| 15 | R_TLS_LE | thread local, local exec offset |
| 16 | R_TLS_IE | thread local, initial exec GOT offset |
| 17 | R_GOTOFF | offset from the GOT base |
| 18 | R_PLT0 | PLT sequence, first instruction |
| 19 | R_PLT1 | PLT sequence, second instruction |
| 20 | R_PLT2 | PLT sequence, third instruction |
| 21 | R_USEFIELD | field reachability marker |
| 22 | R_USETYPE | type reachability marker, no bytes patched |
| 23 | R_USEIFACE | interface conversion marker, size 0 |
| 24 | R_USEIFACEMETHOD | interface method marker, size 0, addend is the method offset |
| 25 | R_USENAMEDMETHOD | keeps named methods alive |
| 26 | R_METHODOFF | like R_ADDROFF, the linker may zero it when the method is dead |
| 27 | R_KEEP | keeps the target alive if the source survives |
| 28 | R_POWER_TOC | ppc64: TOC relative |
| 29 | R_GOTPCREL | 32-bit PC relative GOT slot |
| 30 | R_JMPMIPS | mips: non-PC-relative jump target |
| 31 | R_DWARFSECREF | offset of the symbol from its section, DWARF use |
| 32 | R_ARM64_TLS_LE | arm64: MOV[NZ] immediate, TLS local exec |
| 33 | R_ARM64_TLS_IE | arm64: adrp plus ldr, TLS initial exec |
| 34 | R_ARM64_GOTPCREL | arm64: adrp plus ldr GOT slot |
| 35 | R_ARM64_GOT | arm64: GOT relative sequence |
| 36 | R_ARM64_PCREL | arm64: adrp plus add PC relative |
| 37 | R_ARM64_PCREL_LDST8 | arm64: adrp plus 8-bit load or store |
| 38 | R_ARM64_PCREL_LDST16 | arm64: adrp plus 16-bit load or store |
| 39 | R_ARM64_PCREL_LDST32 | arm64: adrp plus 32-bit load or store |
| 40 | R_ARM64_PCREL_LDST64 | arm64: adrp plus 64-bit load or store |
| 41 | R_ARM64_LDST8 | arm64: 12-bit load or store immediate, byte |
| 42 | R_ARM64_LDST16 | arm64: bits 11 to 1 of the address |
| 43 | R_ARM64_LDST32 | arm64: bits 11 to 2 |
| 44 | R_ARM64_LDST64 | arm64: bits 11 to 3 |
| 45 | R_ARM64_LDST128 | arm64: bits 11 to 4 |
| 46 | R_POWER_TLS_LE | ppc64: TLS local exec across two instructions |
| 47 | R_POWER_TLS_IE | ppc64: TLS initial exec via GOT |
| 48 | R_POWER_TLS | ppc64: marks the X-form instruction completing a TLS sequence |
| 49 | R_POWER_TLS_IE_PCREL34 | ppc64: prefixed TLS initial exec load |
| 50 | R_POWER_TLS_LE_TPREL34 | ppc64: prefixed TLS local exec |
| 51 | R_ADDRPOWER_DS | ppc64: DS-form second instruction, bits 15 to 2 |
| 52 | R_ADDRPOWER_GOT | ppc64: GOT entry relative to TOC |
| 53 | R_ADDRPOWER_GOT_PCREL34 | ppc64: PC relative GOT, prefixed |
| 54 | R_ADDRPOWER_PCREL | ppc64: PC relative across two D-form instructions |
| 55 | R_ADDRPOWER_TOCREL | ppc64: TOC relative across two D-form instructions |
| 56 | R_ADDRPOWER_TOCREL_DS | ppc64: TOC relative, DS form |
| 57 | R_ADDRPOWER_D34 | ppc64: prefixed absolute, 34 bits |
| 58 | R_ADDRPOWER_PCREL34 | ppc64: prefixed PC relative, 34 bits |
| 59 | R_RISCV_JAL | riscv64: 20-bit J-type offset |
| 60 | R_RISCV_JAL_TRAMP | riscv64: as R_RISCV_JAL, linker-generated trampolines only |
| 61 | R_RISCV_CALL | riscv64: AUIPC plus JALR pair |
| 62 | R_RISCV_PCREL_ITYPE | riscv64: AUIPC plus I-type pair |
| 63 | R_RISCV_PCREL_STYPE | riscv64: AUIPC plus S-type pair |
| 64 | R_RISCV_TLS_IE | riscv64: TLS initial exec, AUIPC plus I-type |
| 65 | R_RISCV_TLS_LE | riscv64: TLS local exec, LUI plus I-type |
| 66 | R_RISCV_GOT_HI20 | riscv64: high 20 bits of a GOT address |
| 67 | R_RISCV_GOT_PCREL_ITYPE | riscv64: GOT entry, AUIPC plus I-type |
| 68 | R_RISCV_PCREL_HI20 | riscv64: high 20 bits of a PC relative address |
| 69 | R_RISCV_PCREL_LO12_I | riscv64: low 12 bits, I-type |
| 70 | R_RISCV_PCREL_LO12_S | riscv64: low 12 bits, S-type |
| 71 | R_RISCV_BRANCH | riscv64: 12-bit branch offset |
| 72 | R_RISCV_ADD32 | riscv64: in-place addition, V + S + A |
| 73 | R_RISCV_SUB32 | riscv64: in-place subtraction, V - S - A |
| 74 | R_RISCV_RVC_BRANCH | riscv64: 8-bit compressed branch offset |
| 75 | R_RISCV_RVC_JUMP | riscv64: 11-bit compressed jump offset |
| 76 | R_PCRELDBL | s390x: PC relative, 2-byte aligned |
| 77 | R_LOONG64_ADDR_HI | loong64: bits 31 to 12 of an address |
| 78 | R_LOONG64_ADDR_LO | loong64: low 12 bits |
| 79 | R_LOONG64_ADDR64_HI | loong64: bits 63 to 52 |
| 80 | R_LOONG64_ADDR64_LO | loong64: bits 51 to 32 |
| 81 | R_LOONG64_ADDR_PCREL20_S2 | loong64: 22-bit aligned PC relative, PCADDI |
| 82 | R_LOONG64_TLS_LE_HI | loong64: TLS local exec, high bits |
| 83 | R_LOONG64_TLS_LE_LO | loong64: TLS local exec, low bits |
| 84 | R_CALLLOONG64 | loong64: 28-bit aligned BL |
| 85 | R_LOONG64_CALL36 | loong64: 38-bit aligned PCADDU18I plus JIRL |
| 86 | R_LOONG64_TLS_IE_HI | loong64: TLS initial exec via GOT, high |
| 87 | R_LOONG64_TLS_IE_LO | loong64: TLS initial exec via GOT, low |
| 88 | R_LOONG64_GOT_HI | loong64: GOT entry, high bits |
| 89 | R_LOONG64_GOT_LO | loong64: GOT entry, low bits |
| 90 | R_LOONG64_GOT64_HI | loong64: 64-bit GOT entry, high |
| 91 | R_LOONG64_GOT64_LO | loong64: 64-bit GOT entry, low |
| 92 | R_LOONG64_ADD64 | loong64: 64-bit in-place addition |
| 93 | R_LOONG64_SUB64 | loong64: 64-bit in-place subtraction |
| 94 | R_JMP16LOONG64 | loong64: 18-bit aligned conditional jump |
| 95 | R_JMP21LOONG64 | loong64: 23-bit aligned BEQZ or BNEZ |
| 96 | R_ADDRMIPSU | mips: sign-adjusted upper 16 bits |
| 97 | R_ADDRMIPSTLS | mips: TLS low 16 bits |
| 98 | R_ADDRCUOFF | pointer-sized offset from the DWARF compile unit start |
| 99 | R_WASMIMPORT | wasm: import module and name indices |
| 100 | R_XCOFFREF | aix: keeps the target alive, patches nothing |
| 101 | R_PEIMAGEOFF | windows: offset from the image base |
| 102 | R_INITORDER | orders inittask records, patches nothing |
| 103 | R_DWTXTADDR_U1 | writes a 1-byte ULEB .debug_addr index for the target function |
| 104 | R_DWTXTADDR_U2 | as above, 2 bytes |
| 105 | R_DWTXTADDR_U3 | as above, 3 bytes |
| 106 | R_DWTXTADDR_U4 | as above, 4 bytes; the assembler always picks this one |
| -32768 | R_WEAK | mask: the target need not be reachable, see below |
| -32767 | R_WEAKADDR | R_WEAK or R_ADDR |
| -32763 | R_WEAKADDROFF | R_WEAK or R_ADDROFF |
R_WEAK is bit 15 set on a negative `int16`: a weak relocation is the base
type's value with bit 15 set. The linker strips the bit before dispatch.
## Aux symbol entries
Element size 9 bytes: a `uint8` type then a SymRef. Aux entries attach
auxiliary symbols to a definition; the arrays run per symbol in the order
given by AuxIndex.
| Value | Name | Attaches |
|---|---|---|
| 0 | AuxGotype | the Go type of a data symbol |
| 1 | AuxFuncInfo | the FuncInfo payload of a text symbol |
| 2 | AuxFuncdata | one funcdata symbol; one entry per slot, nil slots carry the {0,0} reference |
| 3 | AuxDwarfInfo | DWARF debug info for the function |
| 4 | AuxDwarfLoc | DWARF location lists |
| 5 | AuxDwarfRanges | DWARF range lists |
| 6 | AuxDwarfLines | DWARF line program |
| 7 | AuxPcsp | pc-value table: SP adjustments |
| 8 | AuxPcfile | pc-value table: source file indices |
| 9 | AuxPcline | pc-value table: line numbers |
| 10 | AuxPcinline | pc-value table: inlining tree positions |
| 11 | AuxPcdata | one pc-value table per live variable slot |
| 12 | AuxWasmImport | wasm import description |
| 13 | AuxWasmType | wasm export type description |
| 14 | AuxSehUnwindInfo | Windows SEH unwind info |
The writer emits them in the order Gotype, FuncInfo, Funcdata entries,
DwarfInfo, DwarfLoc, DwarfRanges, DwarfLines, Pcsp, Pcfile, Pcline, Pcinline,
SehUnwindInfo, Pcdata entries, WasmImport, WasmType, and skips any whose
payload would be empty. A function assembled from `.s` source by Go 1.27.1
carries exactly: FuncInfo, the Funcdata slots including nils, DwarfInfo,
DwarfLines, Pcsp, Pcfile, Pcline and Pcinline; gasm's writer produces the
same set.
The aux targets are either PkgIdxSelf definitions, PkgIdxHashed pcdata
symbols, or, for the funcdata of assembly functions, PkgIdxNone references
carrying names such as `pkg.Fn.args_stackmap` and `pkg.Fn.arginfo0`, which
resolve to definitions in the package's compiled Go code when there is any.
## Symbol payloads (BlkData)
The payloads of all defined symbols, in definition order, concatenated with
no padding; DataIndex gives each symbol's slice. A text symbol's payload is
its machine code, with the stack-split preamble and any morestack block
already included. A data symbol's payload is the bytes laid down by its DATA
directives, zero filled to its declared size. If a symbol was created from an
embedded file, the file's bytes follow the payload and count towards its
DataIndex extent; assembly producers never write this extension.
### The FuncInfo payload
An SDATA symbol with no name, referenced by AuxFuncInfo. 28 bytes minimum,
little endian:
| Offset | Size | Field | Meaning |
|---|---|---|---|
| 0 | 4 | Args | argument area in bytes; 0x80000000 when the producer declared none |
| 4 | 4 | Locals | frame size in bytes |
| 8 | 1 | FuncID | runtime function classification, 0 means normal |
| 9 | 1 | FuncFlag | TopFrame = 1, SPWrite = 2, Asm = 4 |
| 10 | 2 | padding | zero, reserved to a 4 byte boundary |
| 12 | 4 | StartLine | source line of the TEXT declaration |
| 16 | 4 | NumFile | count of file indices that follow |
| 20 | 4 × NumFile | Files | indices into the Files block, ascending |
| then | 4 | NumInlTree | count of inlining tree nodes that follow |
| then | 24 × NumInlTree | InlTree | nodes, see below |
One InlTree node, 24 bytes: `int32` parent index, `uint32` file index,
`int32` line, `uint32` PkgIdx and `uint32` SymIdx of the inlined function, and
`int32` parent PC.
The assembler derives FuncID from the symbol name through
`objabi.GetFuncID`, so a runtime function with a name the runtime treats
specially gets that classification even when defined in assembly; an ordinary
name yields 0. FuncFlag carries the Asm bit, 4, for every assembly function.
### The pc-value tables
The AuxPcsp, AuxPcfile, AuxPcline, AuxPcinline and AuxPcdata payloads are
pc-value tables, each a sequence of value deltas and PC deltas:
- a signed value delta, zig-zag encoded, `binary.PutVarint` form;
- an unsigned PC delta in ULEB128 form, counted in instruction units, the
raw delta divided by the architecture's minimum instruction length;
- the table ends with a final PC delta to the end of the function followed by
a zero byte.
The first value applies from function entry. The encoding is the one
`cmd/internal/obj/pcln.go` calls funcpctab, and it is the same encoding the
final runtime pclntable carries.
### The DWARF payloads
AuxDwarfInfo, AuxDwarfLoc, AuxDwarfRanges and AuxDwarfLines reference SDWARF
symbols whose payloads are DWARF byte streams. The object format treats them
as opaque: the linker concatenates them into the final `.debug_*` sections
and resolves the relocations recorded inside them. The compiler produces
DWARF content per its own generation; gasm produces DWARF5 streams in
`asm/goobj_dwarf.go`.
## Builtins
Frequently referenced runtime functions are referenced by index rather than
by name: PkgIdxBuiltin with SymIdx set to the position in the generated table
`cmd/internal/goobj/builtinlist.go`, 299 entries in Go 1.27.1, names such as
`runtime.newobject` at index 0; 232 entries carry ABI 1 and the remaining 67
ABI 0. Builtin names never enter the string table. The mapping only applies
while the object is not linked against shared libraries, and a linkname'd
symbol never counts as a builtin even when its name matches.
## Fingerprints
The 8 byte fingerprint identifies one build of a package. The compiler fills
it with a hash of the package's export data; the assembler leaves it zero.
The linker checks a package's fingerprint against the fingerprints its
importers recorded in their Autolib entries and rejects a mismatched build,
which is how stale objects are caught.
## What a producer must do
The checklist a third-party writer must satisfy for `go build` to accept its
objects, in one place:
1. Write the container exactly: the `go object` line matching the target
toolchain's configuration string, the `!\n` terminator, then the blob.
2. Emit the 19 block offsets, in order, and make BlkEnd the blob length.
3. Deduplicate the string table, keep the empty string at offset 96, and
reference it everywhere a name appears.
4. Index relocations, aux entries and data per symbol with the N + 1 arrays,
definitions ordered Symdefs, Hashed64defs, Hasheddefs, NonPkgDefs.
5. Sort relocations by offset within each symbol.
6. Fill Siz with the true payload length, set Align for every
content-addressable symbol, and keep symbols under 2 GB.
7. Reference symbols by the package-index rules. An assembly producer
references everything outside the object by name, PkgIdxNone,
except its own file-local statics and the builtins; PkgIdxSelf is
reserved for definitions in this object. Assembly TEXT symbols
carry ABI 0.
8. Compute the content hashes exactly as the toolchain does, or emit no
hashed definitions at all.
## How gasm-devkit implements and verifies it
The writer lives in `asm/goobj.go`, which carries the shared container and the
amd64 relocation emission, with per-architecture relocation emitters in
`asm/goobjarm64.go`, `asm/goobjriscv.go` and `asm/goobjloong64.go`, symbol
resolution in `asm/goobj_resolve.go` and DWARF generation in
`asm/goobj_dwarf.go`. `gasm asm --format goobj -p pkg/path` writes objects
that `go build` consumes in place of the toolchain's own.
Verification is differential and continuous:
- `asm/goobj_test.go` compares gasm's GOOBJ output against `go tool asm`
output for the same source, byte for byte;
- `asm/link_test.go` builds real Go programs whose assembly comes from gasm
objects and runs them;
- `gasm verify` keeps the machine code itself identical to the toolchain's,
which is the precondition for the object comparison to be meaningful.
## Versioning and drift
The magic string carries the format generation, `go120ld` in Go 1.27.1. When
a toolchain release changes the format, it changes that string first, and the
linker refuses blobs whose magic it does not know. The watch points for a new
release are, in order: the magic, the block index list, the Aux type list,
the tail of the relocation table, the FuncInfo layout, and the builtin table
count. gasm's tests fail against any of these changes, which is the mechanism
that keeps this document and the writer current.
The authoritative sources, for the release this document covers:
- `cmd/internal/goobj/objfile.go`: the format, every structure in this
document;
- `cmd/internal/goobj/funcinfo.go`: FuncInfo and the inlining tree;
- `cmd/internal/goobj/builtinlist.go`: the builtin table;
- `cmd/internal/obj/objfile.go`: the writer, hash inputs and aux order;
- `cmd/internal/obj/sym.go`: package index assignment and the by-name rule;
- `cmd/internal/obj/pcln.go`: the pc-value encoding;
- `cmd/internal/objabi/reloctype.go`: relocation types;
- `cmd/internal/objabi/symkind.go`: symbol kinds;
- `cmd/link/internal/ld/lib.go`: container parsing and fingerprint checks.
+124
View File
@@ -0,0 +1,124 @@
# AMD64
Layer 1, target page. Verified against `go tool asm` of Go 1.27.1 and against
gasm's encoder, whose output is compared byte for byte with the toolchain's
and executed on real hardware (`gasm verify`). The complete mnemonic
inventory lives in the generated appendix
[INSTRUCTIONS-AMD64.md](INSTRUCTIONS-AMD64.md); this page is the grammar and
the conventions.
## Registers
| Group | Names | Notes |
|---|---|---|
| General purpose, 64-bit | `AX` `BX` `CX` `DX` `SI` `DI` `BP` `SP` `R8` to `R15` | bare names, no prefix |
| Sub-registers | `AL` `CL` `DL` `BL` `AH` family; `R8B` `R8W` `R8D` for the byte, word and double word of `R8` | width rides the mnemonic as well |
| Vector | `X0` to `X15` (128-bit), `Y0` to `Y15` (256-bit), `Z0` to `Z31` (512-bit) | SSE, AVX and AVX-512 |
| Mask | `K0` to `K7` | AVX-512 opmask |
| System | `TLS` | the thread pointer, see below |
Roles the calling convention fixes, which assembly must respect and can rely
on:
- `SP` is the hardware stack pointer; the virtual frame pointer of the
common language is the pseudo-register SP of OPERANDS.md, a different
spelling with a different meaning.
- `BP` is callee-save. The assembler inserts the save and restore whenever
the function has a non-zero frame, so using BP as a general register
interferes with sampling profilers that walk the frame chain.
- `R14` holds `g`, the goroutine pointer, in the register ABI; `RDX` holds
the closure context; `R12` and `R13` are the register ABI's scratch pair
and `R15` its GOT temporary; `X15` is the zeroing register the compiler
uses. An ABI0 assembly function called from Go sees none of these live
across the call, but runtime assembly reads them directly.
- The legacy spellings for the goroutine pointer are the macros of
`runtime/go_tls.h`: `get_tls(r)` expands to `MOVQ TLS, r` and `g(r)` to
`0(r)(TLS*1)`, the segment base riding the index field.
## Addressing
The common forms of OPERANDS.md, with the amd64 specifics:
```text
offset(base) MOVQ 16(BX), AX
offset(base)(index*scale) MOVL foo+32(SP)(R9*8), CX
scale is 1, 2, 4 or 8
name±offset(SB) MOVQ ·table(SB), CX
```
- Global references assemble as absolute addresses and produce R_ADDR
relocations; branch targets produce R_PCREL.
- Vector indexed memory, the VSIB form with an X, Y or Z register in the
index position, exists for the gather and scatter families.
- There are no segment overrides in source; the one segment-flavoured form
is the TLS base in the index field shown above.
## The frame and the split check
The assembler manages the frame, not the programmer:
- It inserts the `BP` save and restore for any non-zero frame.
- It inserts the stack-split check for any function that is not NoSplit:
the check compares SP against the guard, and on exhaustion calls
`runtime.morestack_noctxt`. Frames at or below 128 bytes, StackSmall, use
the small compare; frames at or below 4096 bytes, StackBig, use the
adjusted form; larger frames compare in two steps.
- On amd64 the assembler marks a function NoSplit itself when the frame is
under StackSmall and the body calls nothing that needs stack: such a
function carries the NoSplit flag in the object without the source ever
writing NOSPLIT.
Results and arguments are stack-only in ABI0: the caller's frame carries
them at FP offsets, per the Go prototype.
## Instructions
The inventory counts 1654 recognised mnemonics today, of which the encoder
emits 1113; both numbers are generated in the appendix, and the gap is the
encoder backlog that `gasm audit-instructions` measures. The families:
- **Integer base.** The ALU and move set with width suffixes, `MOVB`,
`MOVW`, `MOVL`, `MOVQ`; the extension moves `MOVBLZX`, `MOVWLSX`,
`MOVLQSX` and their siblings, which the compiler's output leans on;
`LEA`; `PUSH` and `POP`; the shifts and rotates; the bit operations `BT`
through `BTC`, `BSF`, `BSR`, `LZCNT`, `TZCNT`, `POPCNT`, `BSWAP`; the
string primitives `MOVS` and `STOS`.
- **Exchange and atomics.** `XCHG`, `CMPXCHG`, `XADD`; the extended-carry
pair `ADCX` and `ADOX`; `CRC32`.
- **Scalar floating point.** The SSE2 scalar moves and arithmetic
(`MOVSD`, `MOVSS`, `ADDSD`, and the `CVT` family). Floating-point
immediates are not encodable on this target, so the assembler
materialises them: the constant lands in a synthesised read-only pool,
and a positive zero collapses to `XORPS` of the register with itself,
exactly as the toolchain does.
- **Legacy SIMD, SSE.** The `MOVO`, `MOVOU`, `MOVAPS` family and the packed
integer and floating operations, shuffles, lane extracts and inserts and
the imm8-controlled forms.
- **VEX and EVEX.** The `V`-prefixed forms for 256 and 512-bit work,
opmask operations on `K0` to `K7`, gathers and scatters, and the
quad-register families 4FMAPS, 4FNMADD, 4VNNIW, VP4DPWSSD and VP4DPWSSDS,
whose register list rides the inverted V′VVV field. Mixing VEX and legacy
SSE in one loop pays the AVX-SSE transition penalty on every switch: keep
a loop in one dialect.
- **Cryptographic and counting extensions.** AES-NI, SHA-1 and SHA-256,
PCLMULQDQ, GFNI.
- **System.** `CPUID`, `RDTSC`, `SYSCALL`, the fences, `LDMXCSR` and
`STMXCSR`, the prefetch family.
- **Pseudo-operations.** `BYTE`, `WORD`, `LONG`, `QUAD` lay raw bytes or
words into the stream for encodings the assembler does not know; `ADJSP`
adjusts the stack pointer; `DUFFCOPY` and `DUFFZERO` and `GETCALLERPC`
are compiler-side names the table recognises but an encoder need not
emit.
A mnemonic the appendix lists with `gasm encodes: no` assembles nowhere:
gasm reports it as an explicit error, never as wrong bytes, and the
`unencodable-instruction` lint flags it at edit time.
## Relocations
The relocations an amd64 object carries, all specified in
[GOOBJ.md](../GOOBJ.md): `R_ADDR` for absolute globals, `R_PCREL` for
relative addresses, `R_CALL` for direct calls, `R_TLS_LE` and `R_TLS_IE` for
thread local access and `R_GOTPCREL` for GOT relative sequences, plus
`R_DWTXTADDR_U4` inside the DWARF records, which the assembler always
emits in the four-byte flavour.
+122
View File
@@ -0,0 +1,122 @@
# ARM64
Layer 1, target page. Verified against `go tool asm` of Go 1.27.1, against the
toolchain's own arm64 assembler manual (`cmd/internal/obj/arm64/doc.go`) and
against gasm's encoder, whose output is compared byte for byte with the
toolchain's. The complete mnemonic inventory lives in the generated appendix
[INSTRUCTIONS-ARM64.md](INSTRUCTIONS-ARM64.md).
## Registers
- General purpose: `R0` to `R30`, plus `ZR`, the zero register, and `RSP`,
the stack pointer. There is no R31: thirty-one names and ZR.
- Floating-point and SIMD share one file written `Vn`; where an instruction
is scalar floating point the operand may be written `Fn` (`F0` to `F31`).
- SVE register names (`Z0` to `Z31`, `P0` to `P15`) exist in the assembler's
tables.
- Roles the convention fixes: `RSP` is the stack pointer, `R29` the frame
pointer, `R30` the link register, `R26` the closure context and `R27` the
assembler's scratch register. The goroutine pointer lives in `R28` and is
written `g` in source, its fields as `g_m(g)`, `g_sched(g)`; `R18` is the
platform-reserved register and the Go toolchain never addresses it.
## Loads, stores and the width suffixes
The MOV series is the load and store interface, with the width in the
mnemonic rather than the register name:
| Mnemonic | Machine instruction |
|---|---|
| `MOVD` | ldr, str, stur, 64-bit |
| `MOVW` | ldrsw, str, stur, 32-bit sign extending |
| `MOVWU` | ldr, 32-bit zero extending |
| `MOVH` | ldrsh, strh, sturh |
| `MOVHU` | ldrh |
| `MOVB` | ldrsb, strb, sturb |
| `MOVBU` | ldrb |
Post-index and pre-index addressing take the `.P` and `.W` suffixes on the
mnemonic: `MOVD.P -8(R10), R8` is `ldr x8, [x10],#-8`, and `MOVB.W
16(R16), R10` is `ldrsb x10, [x16,#16]!`.
## Addressing
```text
imm(Rn|RSP) 28(R17)
(Rn|RSP) (R22)
(Rn)(Rm) (R27)(R23)
(Rn)(Rm<<scale) (R4)(R12<<2)
(Rn)(Rm.UXTW<<3) extended and shifted index
(Rt1, Rt2) register pair for LDP, STP and the exclusive pair forms
```
Branch targets are labels, `(R3)` for indirect, `name(SB)` for static.
## Operand order and the special forms
Most instructions appear in left-to-right assignment order: `ADD R11,
RSP, R25` computes into R25. The exceptions the toolchain's manual lists,
each with its own order:
- stores and `CBZ`, `CBNZ` keep the GNU order: `MOVD R29, 384(R19)`.
- The multiply-accumulate family `MADD`, `MSUB`, `SMADDL` and friends are
`<Rm>, <Ra>, <Rn>, <Rd>`.
- The scalar FMA family `FMADDD` and friends are `<Fm>, <Fa>, <Fn>, <Fd>`.
- The bitfield family `BFI`, `BFXIL`, `SBFIZ`, `SBFX`, `UBFIZ`, `UBFX` is
`$<lsb>, <Rn>, $<width>, <Rd>`.
- The conditional compare and select families carry the condition as the
**first** operand: `CSEL GT, R0, R19, R1`, `CCMP MI, R22, $12, $13`,
`FCCMPD AL, F8, F26, $0`.
- The exclusive stores are `<Rf>, (<Rn>), <Rs>` with the status register
last: `STLXR ZR, (R15), R16`.
- `TBZ` and `TBNZ` are `$<imm>, <Rt>, <label>`.
Shifted and extended register operands ride the register: `R19>>30`,
`R26->24` for arithmetic right shift, `@>` for rotate, and the extend forms
`R19.UXTB<<4`, `R14.SXTX` with extend operators UXTB, UXTH, UXTW, UXTX,
SXTB, SXTH, SXTW, SXTX.
## Conditions, branches and names
- Conditions ride the branch mnemonic: `B.EQ`, or the canonical
per-condition names such as `BEQ`. Both spellings exist; the canonical
names are what the generated inventory lists.
- `br` is `JMP` and `blr` is `CALL` in this dialect; indirect branches are
`JMP (R3)` and `CALL (R17)`.
- `NOP` is a zero-width pseudo-instruction; the hardware nop is `NOOP`,
an alias of `HINT $0`.
- `umov` is written as `VMOV`.
## Constants
- A 16-bit immediate optionally shifted: `MOVK $(10<<32), R20`, with
`MOVZ`, `MOVN` and their W variants; a zero shift is rejected by the
assembler.
- Large integer constants: `MOV` materialises any 64-bit constant, the
closest-instruction way.
- Vector constants: `VMOVS`, `VMOVD` and `VMOVQ`, the last taking two
64-bit halves for a 128-bit value:
`VMOVQ $0x1122334455667788, $0x99aabbccddeeff00, V2`.
## SIMD
Floating-point and SIMD instructions mostly carry a `V` prefix
(`VADD`, `VFMLA`), the cryptographic extensions (`AESD`, `SHA256H`) and the
scalar floating-point instructions being the exceptions. Operands carry an
arrangement suffix, `V5.H8`, and structure loads and stores use bracket
lists, `[V21.B16]`, with element selection as `V9.S[1]`.
## Alignment
`PCALIGN $n` pads to a power-of-two boundary between 8 and 2048 and also
raises the function's alignment to the coarsest boundary any of its PCALIGN
directives asks for. Functions default to 16-byte alignment on this target.
## Relocations
`R_ADDRARM64` for the adrp-plus-add pair, `R_ARM64_PCREL` and the
`R_ARM64_PCREL_LDST` family for PC relative addressing, `R_ARM64_LDST` for
the load and store immediates, `R_ARM64_GOTPCREL` and `R_ARM64_GOT` for the
GOT, `R_ARM64_TLS_LE` and `R_ARM64_TLS_IE` for thread local storage and
`R_CALLARM64` for direct calls, all specified in
[GOOBJ.md](../GOOBJ.md).
+148
View File
@@ -0,0 +1,148 @@
# Directives: TEXT, DATA, GLOBL and the annotations
Layer 1, the common language, with the flag vocabulary both layers share.
Verified against `go tool asm` of Go 1.27.1, against the shipped headers
`textflag.h` and `funcdata.h` in `$GOROOT/pkg/include`, and against gasm's
parser. Where gasm extends a directive, the extension says so and is marked.
Six directives exist. Three define things: TEXT, DATA, GLOBL. Three
annotate: FUNCDATA, PCDATA, PCALIGN.
## TEXT
```text
// func Add(a, b int64) int64
TEXT ·Add(SB), NOSPLIT, $0-24
...instructions...
RET
```
```text
TEXT symbol(SB), [flags,] $framesize[-argsize]
```
- The symbol is an `·Name(SB)` reference into the current package, or a
fully qualified name.
- The optional flag argument is a constant expression, normally an OR of the
names from `textflag.h`, the table below. Without `#include "textflag.h"`
the names are not macros and the assembler reports the misleading error
`illegal or missing addressing mode for symbol NOSPLIT`: include the
header first.
- `$framesize-argsize` is two constants, not a subtraction: the local frame
size in bytes, and the caller's argument area in bytes. The argument size
may be omitted entirely, `$16`, which marks the argument size unknown
(0x80000000 in the object, the value of `ArgsSizeUnknown` from
`funcdata.h`); a frame size may be negative only in the generated ABI
wrappers.
- A function whose last instruction is not a branch cannot fall through into
the next TEXT: the toolchain appends a jump to itself, so end functions
with `RET` deliberately.
- One TEXT per symbol; redeclaring is an error. The TEXT line also fixes the
function's source line for traceback: it is the line number that pcln
reports for the function's start.
The framesize and argsize fields do real work: the framesize drives the
stack-split preamble (RUNTIME.md carries the contract), and both travel into
the FuncInfo record of the object (GOOBJ.md carries its layout).
### The flag table
Values from `textflag.h`, in agreement with `cmd/internal/obj/textflag.go`:
| Name | Value | Applies to | Meaning |
|---|---|---|---|
| NOPROF | 1 | both | do not profile; deprecated |
| DUPOK | 2 | both | the linker may keep one of several duplicates |
| NOSPLIT | 4 | TEXT | no stack-split preamble |
| RODATA | 8 | data | put the data in a read-only section |
| NOPTR | 16 | data | the data contains no pointers |
| WRAPPER | 32 | TEXT | a wrapper; must not disable `recover` |
| NEEDCTXT | 64 | TEXT | a closure consuming the context register |
| TLSBSS | 256 | data | a thread local word in BSS |
| NOFRAME | 512 | TEXT | no frame setup; only valid with a frame size of 0 |
| REFLECTMETHOD | 1024 | TEXT | the function calls `reflect.Type.Method` or `MethodByName` |
| TOPFRAME | 2048 | TEXT | the outermost frame; unwinders stop here |
| ABIWRAPPER | 4096 | TEXT | an ABI transition wrapper |
Rules with teeth:
- `NOSPLIT` removes the split check, so the frame plus everything the
function calls must fit in the stack segment that remains. It exists to
protect the splitting code itself; reaching for it to save two instructions
is how stack overflows corrupt memory. On amd64 the assembler additionally
marks small leaf functions NoSplit itself and omits the check, so the
absence of the preamble is not proof the flag was written.
- A TEXT whose symbol is declared `ABIInternal` must carry NOSPLIT: the
assembler rejects it otherwise, because it cannot generate
the split path for a register-ABI function.
- `RODATA` implies NOPTR for the garbage collector.
## DATA
```text
DATA ·table+0(SB)/8, $0x0102030405060708
DATA ·msg+0(SB)/14, $"hello, world\n"
GLOBL ·msg(SB), RODATA, $14
```
```text
DATA symbol+offset(SB)/width, value
```
- `width` is exactly 1, 2, 4 or 8: the initialiser is written into the data
image at `symbol+offset` in that many bytes.
- The value is an integer or character constant of the width, or a string
literal whose byte length equals the width exactly; escapes count. Long
data is written as successive DATA lines at increasing offsets; bytes the
directives never name are zero.
- Every symbol initialised with DATA ends with a GLOBL line declaring its
total size, after all of its DATA lines.
A symbol containing pointers cannot be defined in assembly, because the
collector cannot see into it: define it in Go and refer to it by name. As a
rule, data that is not read-only belongs in Go.
Extension, gasm only: a DATA initialiser may name a symbol,
`DATA ·fn+0(SB)/8, $·handler(SB)`, which gasm lays down as an absolute
relocation on that field. The toolchain offers no ground truth for this
form; gasm's behaviour is verified by linking and execution.
## GLOBL
```text
GLOBL symbol(SB), [flags,] $size
```
Declares the symbol global with its total size in bytes. The useful flags
are RODATA, NOPTR, DUPOK and TLSBSS from the table above. Uninitialised
bytes are zero, which makes GLOBL with no DATA the language's BSS.
## FUNCDATA and PCDATA
```text
FUNCDATA $functypeid, symbol(SB)
PCDATA $pctypeid, $value
```
The compiler's annotations for the garbage collector and traceback, named by
the ids in `funcdata.h`: FUNCDATA 0 to 7 (args pointer maps, locals pointer
maps, stack objects, inline tree, open-coded defer info, argument info,
argument liveness, wrap info), PCDATA 0 to 4 (unsafe point, stack map index,
inline tree index, argument liveness index, panic bounds). Assembly code
normally reaches them only through the macro forms in `funcdata.h`, which
RUNTIME.md explains. Outside the macros, hand-written PCDATA is meaningless:
the values are pc-value tables the compiler builds from its own view of the
program.
## PCALIGN
```text
PCALIGN $32
```
Pads the code so that the next instruction lands on the given boundary,
which must be a power of two and at least the target's instruction
alignment. Supported on amd64, arm64, ppc64, loong64 and riscv64. The
padding instructions are the target's NOP encoding, so the bytes between
functions differ from what the instruction stream alone would produce, which
matters to anyone comparing encodings byte for byte.
File diff suppressed because it is too large Load Diff
+570
View File
@@ -0,0 +1,570 @@
# ARM64: instruction inventory
Generated by gasm-devkit's `_gen` from the Go toolchain's instruction table
(`cmd/internal/obj/arm64/anames.go`, go1.27.1); DO NOT EDIT. This page lists every mnemonic
`go tool asm` accepts on this target, which is the upper bound of the
language on it: a name absent here is not an instruction of the target,
and a name present here may still be one gasm's encoder cannot emit yet.
The inventory carries no per-mnemonic encoder column: on this target
encodability is decided per operand shape, and the live measured
coverage is reported by `gasm audit-instructions`.
| Mnemonic | Notes |
|---|---|
| `CALL` | |
| `DUFFCOPY` | |
| `DUFFZERO` | |
| `END` | |
| `FUNCDATA` | |
| `GETCALLERPC` | |
| `JMP` | |
| `NOP` | No operation |
| `PCALIGN` | |
| `PCALIGNMAX` | |
| `PCDATA` | |
| `RET` | Return |
| `TEXT` | |
| `UNDEF` | |
| `ADC` | ADC (64-bit) |
| `ADCS` | ADCS (64-bit) |
| `ADCSW` | ADCS (32-bit) |
| `ADCW` | ADC (32-bit) |
| `ADD` | ADD (64-bit) |
| `ADDS` | ADDS (64-bit) |
| `ADDSW` | ADDS (32-bit) |
| `ADDW` | ADD (32-bit) |
| `ADR` | Address of label/page |
| `ADRP` | Address of label/page |
| `AESD` | AES round |
| `AESE` | AES round |
| `AESIMC` | AES round |
| `AESMC` | AES round |
| `AND` | AND (64-bit) |
| `ANDS` | ANDS (64-bit) |
| `ANDSW` | ANDS (32-bit) |
| `ANDW` | AND (32-bit) |
| `ASR` | ASR shift |
| `ASRW` | ASR shift (32-bit) |
| `AT` | |
| `AUTIA1716` | |
| `AUTIASP` | |
| `AUTIB1716` | |
| `AUTIBSP` | |
| `BCC` | Conditional branch |
| `BCS` | Conditional branch |
| `BEQ` | Conditional branch |
| `BFI` | |
| `BFIW` | |
| `BFM` | |
| `BFMW` | |
| `BFXIL` | Bitfield extract |
| `BFXILW` | |
| `BGE` | Conditional branch |
| `BGT` | Conditional branch |
| `BHI` | Conditional branch |
| `BHS` | Conditional branch |
| `BIC` | BIC (64-bit) |
| `BICS` | BICS (64-bit) |
| `BICSW` | BICS (32-bit) |
| `BICW` | BIC (32-bit) |
| `BLE` | Conditional branch |
| `BLO` | Conditional branch |
| `BLS` | Conditional branch |
| `BLT` | Conditional branch |
| `BMI` | Conditional branch |
| `BNE` | Conditional branch |
| `BPL` | Conditional branch |
| `BRK` | Breakpoint |
| `BTI` | |
| `BVC` | Conditional branch |
| `BVS` | Conditional branch |
| `CASAD` | |
| `CASALB` | |
| `CASALD` | |
| `CASALH` | |
| `CASALW` | |
| `CASAW` | |
| `CASB` | |
| `CASD` | |
| `CASH` | |
| `CASLD` | |
| `CASLW` | |
| `CASPD` | |
| `CASPW` | |
| `CASW` | |
| `CBNZ` | Compare/test and branch |
| `CBNZW` | Compare/test and branch (32-bit) |
| `CBZ` | Compare/test and branch |
| `CBZW` | Compare/test and branch (32-bit) |
| `CCMN` | Conditional compare |
| `CCMNW` | Conditional compare |
| `CCMP` | Conditional compare |
| `CCMPW` | Conditional compare |
| `CINC` | Conditional select |
| `CINCW` | Conditional select (32-bit) |
| `CINV` | Conditional select |
| `CINVW` | Conditional select (32-bit) |
| `CLREX` | |
| `CLS` | Bit manipulation |
| `CLSW` | Bit manipulation |
| `CLZ` | Bit manipulation |
| `CLZW` | Bit manipulation |
| `CMN` | CMN (64-bit) |
| `CMNW` | CMN (32-bit) |
| `CMP` | CMP (64-bit) |
| `CMPW` | CMP (32-bit) |
| `CNEG` | Conditional select |
| `CNEGW` | Conditional select (32-bit) |
| `CRC32B` | |
| `CRC32CB` | |
| `CRC32CH` | |
| `CRC32CW` | |
| `CRC32CX` | |
| `CRC32H` | |
| `CRC32W` | |
| `CRC32X` | |
| `CSEL` | Conditional select |
| `CSELW` | Conditional select (32-bit) |
| `CSET` | Conditional select |
| `CSETM` | Conditional select |
| `CSETMW` | Conditional select (32-bit) |
| `CSETW` | Conditional select (32-bit) |
| `CSINC` | Conditional select |
| `CSINCW` | Conditional select (32-bit) |
| `CSINV` | Conditional select |
| `CSINVW` | Conditional select (32-bit) |
| `CSNEG` | Conditional select |
| `CSNEGW` | Conditional select (32-bit) |
| `DC` | Data cache maintenance |
| `DCPS1` | |
| `DCPS2` | |
| `DCPS3` | |
| `DMB` | Barrier |
| `DRPS` | |
| `DSB` | Barrier |
| `DWORD` | |
| `EON` | EON (64-bit) |
| `EONW` | EON (32-bit) |
| `EOR` | EOR (64-bit) |
| `EORW` | EOR (32-bit) |
| `ERET` | |
| `EXTR` | Bitfield extract |
| `EXTRW` | |
| `FABSD` | |
| `FABSS` | |
| `FADDD` | |
| `FADDS` | |
| `FCCMPD` | |
| `FCCMPED` | |
| `FCCMPES` | |
| `FCCMPS` | |
| `FCMPD` | |
| `FCMPED` | |
| `FCMPES` | |
| `FCMPS` | |
| `FCSELD` | |
| `FCSELS` | |
| `FCVTDH` | |
| `FCVTDS` | |
| `FCVTHD` | |
| `FCVTHS` | |
| `FCVTSD` | |
| `FCVTSH` | |
| `FCVTZSD` | |
| `FCVTZSDW` | |
| `FCVTZSS` | |
| `FCVTZSSW` | |
| `FCVTZUD` | |
| `FCVTZUDW` | |
| `FCVTZUS` | |
| `FCVTZUSW` | |
| `FDIVD` | |
| `FDIVS` | |
| `FLDPD` | Register-pair load or store |
| `FLDPQ` | |
| `FLDPS` | |
| `FMADDD` | |
| `FMADDS` | |
| `FMAXD` | |
| `FMAXNMD` | |
| `FMAXNMS` | |
| `FMAXS` | |
| `FMIND` | |
| `FMINNMD` | |
| `FMINNMS` | |
| `FMINS` | |
| `FMOVD` | Move / load / store |
| `FMOVQ` | |
| `FMOVS` | Move / load / store |
| `FMSUBD` | |
| `FMSUBS` | |
| `FMULD` | |
| `FMULS` | |
| `FNEGD` | |
| `FNEGS` | |
| `FNMADDD` | |
| `FNMADDS` | |
| `FNMSUBD` | |
| `FNMSUBS` | |
| `FNMULD` | |
| `FNMULS` | |
| `FRINTAD` | |
| `FRINTAS` | |
| `FRINTID` | |
| `FRINTIS` | |
| `FRINTMD` | |
| `FRINTMS` | |
| `FRINTND` | |
| `FRINTNS` | |
| `FRINTPD` | |
| `FRINTPS` | |
| `FRINTXD` | |
| `FRINTXS` | |
| `FRINTZD` | |
| `FRINTZS` | |
| `FSQRTD` | |
| `FSQRTS` | |
| `FSTPD` | Register-pair load or store |
| `FSTPQ` | |
| `FSTPS` | |
| `FSUBD` | |
| `FSUBS` | |
| `HINT` | |
| `HLT` | |
| `HVC` | Exception generation |
| `IC` | |
| `ISB` | Barrier |
| `LDADDAB` | |
| `LDADDAD` | |
| `LDADDAH` | |
| `LDADDALB` | |
| `LDADDALD` | |
| `LDADDALH` | |
| `LDADDALW` | |
| `LDADDAW` | |
| `LDADDB` | |
| `LDADDD` | |
| `LDADDH` | |
| `LDADDLB` | |
| `LDADDLD` | |
| `LDADDLH` | |
| `LDADDLW` | |
| `LDADDW` | |
| `LDAR` | Atomic memory operation |
| `LDARB` | Atomic memory operation |
| `LDARH` | Atomic memory operation |
| `LDARW` | Atomic memory operation |
| `LDAXP` | |
| `LDAXPW` | |
| `LDAXR` | Atomic memory operation |
| `LDAXRB` | Atomic memory operation |
| `LDAXRH` | Atomic memory operation |
| `LDAXRW` | Atomic memory operation |
| `LDCLRAB` | |
| `LDCLRAD` | |
| `LDCLRAH` | |
| `LDCLRALB` | |
| `LDCLRALD` | |
| `LDCLRALH` | |
| `LDCLRALW` | |
| `LDCLRAW` | |
| `LDCLRB` | |
| `LDCLRD` | |
| `LDCLRH` | |
| `LDCLRLB` | |
| `LDCLRLD` | |
| `LDCLRLH` | |
| `LDCLRLW` | |
| `LDCLRW` | |
| `LDEORAB` | |
| `LDEORAD` | |
| `LDEORAH` | |
| `LDEORALB` | |
| `LDEORALD` | |
| `LDEORALH` | |
| `LDEORALW` | |
| `LDEORAW` | |
| `LDEORB` | |
| `LDEORD` | |
| `LDEORH` | |
| `LDEORLB` | |
| `LDEORLD` | |
| `LDEORLH` | |
| `LDEORLW` | |
| `LDEORW` | |
| `LDORAB` | |
| `LDORAD` | |
| `LDORAH` | |
| `LDORALB` | |
| `LDORALD` | |
| `LDORALH` | |
| `LDORALW` | |
| `LDORAW` | |
| `LDORB` | |
| `LDORD` | |
| `LDORH` | |
| `LDORLB` | |
| `LDORLD` | |
| `LDORLH` | |
| `LDORLW` | |
| `LDORW` | |
| `LDP` | Register-pair load or store |
| `LDPSW` | |
| `LDPW` | Register-pair load or store |
| `LDXP` | |
| `LDXPW` | |
| `LDXR` | |
| `LDXRB` | |
| `LDXRH` | |
| `LDXRW` | |
| `LSL` | LSL shift |
| `LSLW` | LSL shift (32-bit) |
| `LSR` | LSR shift |
| `LSRW` | LSR shift (32-bit) |
| `MADD` | Multiply / multiply-accumulate |
| `MADDW` | |
| `MNEG` | Multiply / multiply-accumulate |
| `MNEGW` | |
| `MOVB` | Move / load / store |
| `MOVBU` | Move / load / store |
| `MOVD` | Move / load / store |
| `MOVH` | Move / load / store |
| `MOVHU` | Move / load / store |
| `MOVK` | Move wide constant |
| `MOVKW` | Move wide constant |
| `MOVN` | Move wide constant |
| `MOVNW` | Move wide constant |
| `MOVP` | |
| `MOVPD` | |
| `MOVPQ` | |
| `MOVPS` | |
| `MOVPSW` | |
| `MOVPW` | |
| `MOVW` | Move / load / store |
| `MOVWU` | Move / load / store |
| `MOVZ` | Move wide constant |
| `MOVZW` | Move wide constant |
| `MRS` | System register access |
| `MSR` | System register access |
| `MSUB` | Multiply / multiply-accumulate |
| `MSUBW` | |
| `MUL` | Multiply / multiply-accumulate |
| `MULW` | |
| `MVN` | MVN (64-bit) |
| `MVNW` | MVN (32-bit) |
| `NEG` | NEG (64-bit) |
| `NEGS` | |
| `NEGSW` | |
| `NEGW` | NEG (32-bit) |
| `NGC` | NGC (64-bit) |
| `NGCS` | |
| `NGCSW` | |
| `NGCW` | NGC (32-bit) |
| `NOOP` | |
| `ORN` | ORN (64-bit) |
| `ORNW` | ORN (32-bit) |
| `ORR` | ORR (64-bit) |
| `ORRW` | ORR (32-bit) |
| `PACIASP` | |
| `PACIBSP` | |
| `PRFM` | Memory prefetch |
| `PRFUM` | |
| `RBIT` | Bit manipulation |
| `RBITW` | Bit manipulation |
| `REM` | |
| `REMW` | |
| `REV` | Bit manipulation |
| `REV16` | Bit manipulation |
| `REV16W` | |
| `REV32` | Bit manipulation |
| `REVW` | Bit manipulation |
| `ROR` | ROR shift |
| `RORW` | ROR shift (32-bit) |
| `SBC` | SBC (64-bit) |
| `SBCS` | SBCS (64-bit) |
| `SBCSW` | SBCS (32-bit) |
| `SBCW` | SBC (32-bit) |
| `SBFIZ` | |
| `SBFIZW` | |
| `SBFM` | Bitfield extract |
| `SBFMW` | |
| `SBFX` | Bitfield extract |
| `SBFXW` | |
| `SCVTFD` | |
| `SCVTFS` | |
| `SCVTFWD` | |
| `SCVTFWS` | |
| `SDIV` | Divide |
| `SDIVW` | Divide |
| `SEV` | |
| `SEVL` | |
| `SHA1C` | SHA round |
| `SHA1H` | SHA round |
| `SHA1M` | SHA round |
| `SHA1P` | SHA round |
| `SHA1SU0` | SHA round |
| `SHA1SU1` | SHA round |
| `SHA256H` | SHA round |
| `SHA256H2` | SHA round |
| `SHA256SU0` | SHA round |
| `SHA256SU1` | SHA round |
| `SHA512H` | SHA round |
| `SHA512H2` | SHA round |
| `SHA512SU0` | SHA round |
| `SHA512SU1` | SHA round |
| `SMADDL` | Multiply / multiply-accumulate |
| `SMC` | Exception generation |
| `SMNEGL` | |
| `SMSUBL` | Multiply / multiply-accumulate |
| `SMULH` | Multiply / multiply-accumulate |
| `SMULL` | Multiply / multiply-accumulate |
| `STLR` | Atomic memory operation |
| `STLRB` | Atomic memory operation |
| `STLRH` | Atomic memory operation |
| `STLRW` | Atomic memory operation |
| `STLXP` | |
| `STLXPW` | |
| `STLXR` | |
| `STLXRB` | |
| `STLXRH` | |
| `STLXRW` | |
| `STP` | Register-pair load or store |
| `STPW` | Register-pair load or store |
| `STXP` | |
| `STXPW` | |
| `STXR` | Atomic memory operation |
| `STXRB` | Atomic memory operation |
| `STXRH` | Atomic memory operation |
| `STXRW` | Atomic memory operation |
| `SUB` | SUB (64-bit) |
| `SUBS` | SUBS (64-bit) |
| `SUBSW` | SUBS (32-bit) |
| `SUBW` | SUB (32-bit) |
| `SVC` | Exception generation |
| `SWPAB` | |
| `SWPAD` | |
| `SWPAH` | |
| `SWPALB` | |
| `SWPALD` | |
| `SWPALH` | |
| `SWPALW` | |
| `SWPAW` | |
| `SWPB` | |
| `SWPD` | |
| `SWPH` | |
| `SWPLB` | |
| `SWPLD` | |
| `SWPLH` | |
| `SWPLW` | |
| `SWPW` | |
| `SXTB` | |
| `SXTBW` | |
| `SXTH` | |
| `SXTHW` | |
| `SXTW` | |
| `SYS` | |
| `SYSL` | |
| `TBNZ` | Compare/test and branch |
| `TBZ` | Compare/test and branch |
| `TLBI` | |
| `TST` | TST (64-bit) |
| `TSTW` | TST (32-bit) |
| `UBFIZ` | |
| `UBFIZW` | |
| `UBFM` | Bitfield extract |
| `UBFMW` | |
| `UBFX` | Bitfield extract |
| `UBFXW` | |
| `UCVTFD` | |
| `UCVTFS` | |
| `UCVTFWD` | |
| `UCVTFWS` | |
| `UDIV` | Divide |
| `UDIVW` | Divide |
| `UMADDL` | Multiply / multiply-accumulate |
| `UMNEGL` | |
| `UMSUBL` | Multiply / multiply-accumulate |
| `UMULH` | Multiply / multiply-accumulate |
| `UMULL` | Multiply / multiply-accumulate |
| `UREM` | |
| `UREMW` | |
| `UXTB` | |
| `UXTBW` | |
| `UXTH` | |
| `UXTHW` | |
| `UXTW` | |
| `VADD` | NEON SIMD vector operation |
| `VADDP` | |
| `VADDV` | NEON SIMD vector operation |
| `VAND` | NEON SIMD vector operation |
| `VBCAX` | Three-way XOR / rotate crypto vector operation |
| `VBIF` | NEON SIMD vector operation |
| `VBIT` | |
| `VBSL` | NEON SIMD vector operation |
| `VCMEQ` | |
| `VCMTST` | |
| `VCNT` | NEON SIMD vector operation |
| `VDUP` | NEON SIMD vector operation |
| `VEOR` | NEON SIMD vector operation |
| `VEOR3` | Three-way XOR / rotate crypto vector operation |
| `VEXT` | NEON SIMD vector operation |
| `VFMLA` | NEON SIMD vector operation |
| `VFMLS` | NEON SIMD vector operation |
| `VLD1` | NEON SIMD vector operation |
| `VLD1R` | |
| `VLD2` | NEON SIMD vector operation |
| `VLD2R` | |
| `VLD3` | NEON SIMD vector operation |
| `VLD3R` | |
| `VLD4` | NEON SIMD vector operation |
| `VLD4R` | |
| `VMOV` | NEON SIMD vector operation |
| `VMOVD` | |
| `VMOVI` | NEON SIMD vector operation |
| `VMOVQ` | NEON SIMD vector operation |
| `VMOVS` | |
| `VORR` | NEON SIMD vector operation |
| `VPMULL` | |
| `VPMULL2` | |
| `VRAX1` | Three-way XOR / rotate crypto vector operation |
| `VRBIT` | |
| `VREV16` | NEON SIMD vector operation |
| `VREV32` | NEON SIMD vector operation |
| `VREV64` | NEON SIMD vector operation |
| `VSHL` | NEON SIMD vector operation |
| `VSLI` | |
| `VSRI` | |
| `VST1` | NEON SIMD vector operation |
| `VST2` | NEON SIMD vector operation |
| `VST3` | NEON SIMD vector operation |
| `VST4` | NEON SIMD vector operation |
| `VSUB` | NEON SIMD vector operation |
| `VTBL` | NEON SIMD vector operation |
| `VTBX` | NEON SIMD vector operation |
| `VTRN1` | NEON SIMD vector operation |
| `VTRN2` | NEON SIMD vector operation |
| `VUADDLV` | |
| `VUADDW` | |
| `VUADDW2` | |
| `VUMAX` | |
| `VUMIN` | |
| `VUSHLL` | |
| `VUSHLL2` | |
| `VUSHR` | NEON SIMD vector operation |
| `VUSRA` | |
| `VUXTL` | |
| `VUXTL2` | |
| `VUZP1` | NEON SIMD vector operation |
| `VUZP2` | NEON SIMD vector operation |
| `VXAR` | Three-way XOR / rotate crypto vector operation |
| `VZIP1` | NEON SIMD vector operation |
| `VZIP2` | NEON SIMD vector operation |
| `WFE` | |
| `WFI` | |
| `WORD` | |
| `YIELD` | |
| `B` | Unconditional branch |
| `BL` | Branch with link |
Recognised: 554 mnemonics.
+830
View File
@@ -0,0 +1,830 @@
# LoongArch 64: instruction inventory
Generated by gasm-devkit's `_gen` from the Go toolchain's instruction table
(`cmd/internal/obj/loong64/anames.go`, go1.27.1); DO NOT EDIT. This page lists every mnemonic
`go tool asm` accepts on this target, which is the upper bound of the
language on it: a name absent here is not an instruction of the target,
and a name present here may still be one gasm's encoder cannot emit yet.
The inventory carries no per-mnemonic encoder column: on this target
encodability is decided per operand shape, and the live measured
coverage is reported by `gasm audit-instructions`.
| Mnemonic | Notes |
|---|---|
| `CALL` | |
| `DUFFCOPY` | |
| `DUFFZERO` | |
| `END` | |
| `FUNCDATA` | |
| `GETCALLERPC` | |
| `JMP` | |
| `NOP` | No operation |
| `PCALIGN` | |
| `PCALIGNMAX` | |
| `PCDATA` | |
| `RET` | Return |
| `TEXT` | |
| `UNDEF` | |
| `ABSD` | |
| `ABSF` | |
| `ADD` | Integer add (word) |
| `ADDD` | Add doubleword |
| `ADDF` | |
| `ADDV` | |
| `ADDV16` | |
| `ADDVU` | |
| `ADDW` | Add word |
| `ALSLV` | |
| `ALSLW` | |
| `ALSLWU` | |
| `AMADDDBV` | |
| `AMADDDBW` | |
| `AMADDV` | |
| `AMADDW` | |
| `AMANDDBV` | |
| `AMANDDBW` | |
| `AMANDV` | |
| `AMANDW` | |
| `AMCASB` | |
| `AMCASDBB` | |
| `AMCASDBH` | |
| `AMCASDBV` | |
| `AMCASDBW` | |
| `AMCASH` | |
| `AMCASV` | |
| `AMCASW` | |
| `AMMAXDBV` | |
| `AMMAXDBVU` | |
| `AMMAXDBW` | |
| `AMMAXDBWU` | |
| `AMMAXV` | |
| `AMMAXVU` | |
| `AMMAXW` | |
| `AMMAXWU` | |
| `AMMINDBV` | |
| `AMMINDBVU` | |
| `AMMINDBW` | |
| `AMMINDBWU` | |
| `AMMINV` | |
| `AMMINVU` | |
| `AMMINW` | |
| `AMMINWU` | |
| `AMORDBV` | |
| `AMORDBW` | |
| `AMORV` | |
| `AMORW` | |
| `AMSWAPB` | |
| `AMSWAPDBB` | |
| `AMSWAPDBH` | |
| `AMSWAPDBV` | |
| `AMSWAPDBW` | |
| `AMSWAPH` | |
| `AMSWAPV` | |
| `AMSWAPW` | |
| `AMXORDBV` | |
| `AMXORDBW` | |
| `AMXORV` | |
| `AMXORW` | |
| `AND` | Bitwise AND |
| `ANDN` | |
| `BEQ` | Branch if equal |
| `BFPF` | |
| `BFPT` | |
| `BGE` | Branch if greater or equal |
| `BGEU` | Branch if greater or equal unsigned |
| `BGEZ` | |
| `BGTZ` | |
| `BITREV4B` | |
| `BITREV8B` | |
| `BITREVV` | |
| `BITREVW` | |
| `BLEZ` | |
| `BLT` | Branch if less than |
| `BLTU` | Branch if less than unsigned |
| `BLTZ` | |
| `BNE` | Branch if not equal |
| `BREAK` | Breakpoint |
| `BSTRINSV` | |
| `BSTRINSW` | |
| `BSTRPICKV` | |
| `BSTRPICKW` | |
| `CLOV` | |
| `CLOW` | |
| `CLZV` | |
| `CLZW` | |
| `CMPEQD` | |
| `CMPEQF` | |
| `CMPGED` | |
| `CMPGEF` | |
| `CMPGTD` | |
| `CMPGTF` | |
| `CPUCFG` | |
| `CRCCWBW` | |
| `CRCCWHW` | |
| `CRCCWVW` | |
| `CRCCWWW` | |
| `CRCWBW` | |
| `CRCWHW` | |
| `CRCWVW` | |
| `CRCWWW` | |
| `CTOV` | |
| `CTOW` | |
| `CTZV` | |
| `CTZW` | |
| `DBAR` | Barrier |
| `DIV` | Divide (word) |
| `DIVD` | Divide doubleword |
| `DIVF` | |
| `DIVU` | |
| `DIVV` | |
| `DIVVU` | |
| `DIVW` | Divide word |
| `DIVWU` | |
| `EXTWB` | |
| `EXTWH` | |
| `FCLASSD` | |
| `FCLASSF` | |
| `FCOPYSGD` | |
| `FCOPYSGF` | |
| `FFINTDV` | |
| `FFINTDW` | |
| `FFINTFV` | |
| `FFINTFW` | |
| `FLOGBD` | |
| `FLOGBF` | |
| `FMADDD` | |
| `FMADDF` | |
| `FMAXAD` | |
| `FMAXAF` | |
| `FMAXD` | |
| `FMAXF` | |
| `FMINAD` | |
| `FMINAF` | |
| `FMIND` | |
| `FMINF` | |
| `FMSUBD` | |
| `FMSUBF` | |
| `FNMADDD` | |
| `FNMADDF` | |
| `FNMSUBD` | |
| `FNMSUBF` | |
| `FSCALEBD` | |
| `FSCALEBF` | |
| `FSEL` | |
| `FTINTRMVD` | |
| `FTINTRMVF` | |
| `FTINTRMWD` | |
| `FTINTRMWF` | |
| `FTINTRNEVD` | |
| `FTINTRNEVF` | |
| `FTINTRNEWD` | |
| `FTINTRNEWF` | |
| `FTINTRPVD` | |
| `FTINTRPVF` | |
| `FTINTRPWD` | |
| `FTINTRPWF` | |
| `FTINTRZVD` | |
| `FTINTRZVF` | |
| `FTINTRZWD` | |
| `FTINTRZWF` | |
| `FTINTVD` | |
| `FTINTVF` | |
| `FTINTWD` | |
| `FTINTWF` | |
| `JIRL` | Jump indirect with link |
| `LL` | |
| `LLV` | |
| `LU12IW` | |
| `LU32ID` | |
| `LU52ID` | |
| `LUI` | |
| `MASKEQZ` | |
| `MASKNEZ` | |
| `MOVB` | |
| `MOVBU` | |
| `MOVD` | |
| `MOVDF` | |
| `MOVDV` | |
| `MOVDW` | |
| `MOVF` | |
| `MOVFD` | |
| `MOVFV` | |
| `MOVFW` | |
| `MOVH` | |
| `MOVHU` | |
| `MOVV` | |
| `MOVVD` | |
| `MOVVF` | |
| `MOVVP` | |
| `MOVW` | |
| `MOVWD` | |
| `MOVWF` | |
| `MOVWP` | |
| `MOVWU` | |
| `MUL` | Multiply (word) |
| `MULD` | Multiply doubleword |
| `MULF` | |
| `MULH` | |
| `MULHU` | |
| `MULHV` | |
| `MULHVU` | |
| `MULV` | |
| `MULVU` | |
| `MULW` | Multiply word |
| `MULWVW` | |
| `MULWVWU` | |
| `NEGD` | |
| `NEGF` | |
| `NEGV` | |
| `NEGW` | |
| `NOOP` | |
| `NOR` | Bitwise NOR |
| `OR` | Bitwise OR |
| `ORN` | |
| `PCADDU12I` | |
| `PCALAU12I` | |
| `PRELD` | |
| `PRELDX` | |
| `RDTIMED` | |
| `RDTIMEHW` | |
| `RDTIMELW` | |
| `REM` | |
| `REMU` | |
| `REMV` | |
| `REMVU` | |
| `REMW` | |
| `REMWU` | |
| `REVB2H` | |
| `REVB2W` | |
| `REVB4H` | |
| `REVBV` | |
| `REVH2W` | |
| `REVHV` | |
| `RFE` | |
| `ROTR` | Rotate right |
| `ROTRV` | |
| `SC` | |
| `SCV` | |
| `SGT` | |
| `SGTU` | |
| `SLL` | Shift left logical |
| `SLLV` | |
| `SQRTD` | |
| `SQRTF` | |
| `SRA` | Shift right arithmetic |
| `SRAV` | |
| `SRL` | Shift right logical |
| `SRLV` | |
| `SUB` | Subtract (word) |
| `SUBD` | Subtract doubleword |
| `SUBF` | |
| `SUBV` | |
| `SUBVU` | |
| `SUBW` | Subtract word |
| `SYSCALL` | System call |
| `TEQ` | |
| `TNE` | |
| `TRUNCDV` | |
| `TRUNCDW` | |
| `TRUNCFV` | |
| `TRUNCFW` | |
| `VADDB` | |
| `VADDBU` | |
| `VADDD` | |
| `VADDF` | |
| `VADDH` | |
| `VADDHU` | |
| `VADDQ` | |
| `VADDV` | |
| `VADDVU` | |
| `VADDW` | |
| `VADDWEVHB` | |
| `VADDWEVHBU` | |
| `VADDWEVQV` | |
| `VADDWEVQVU` | |
| `VADDWEVVW` | |
| `VADDWEVVWU` | |
| `VADDWEVWH` | |
| `VADDWEVWHU` | |
| `VADDWODHB` | |
| `VADDWODHBU` | |
| `VADDWODQV` | |
| `VADDWODQVU` | |
| `VADDWODVW` | |
| `VADDWODVWU` | |
| `VADDWODWH` | |
| `VADDWODWHU` | |
| `VADDWU` | |
| `VANDB` | |
| `VANDNV` | |
| `VANDV` | |
| `VBITCLRB` | |
| `VBITCLRH` | |
| `VBITCLRV` | |
| `VBITCLRW` | |
| `VBITREVB` | |
| `VBITREVH` | |
| `VBITREVV` | |
| `VBITREVW` | |
| `VBITSETB` | |
| `VBITSETH` | |
| `VBITSETV` | |
| `VBITSETW` | |
| `VDIVB` | |
| `VDIVBU` | |
| `VDIVD` | |
| `VDIVF` | |
| `VDIVH` | |
| `VDIVHU` | |
| `VDIVV` | |
| `VDIVVU` | |
| `VDIVW` | |
| `VDIVWU` | |
| `VEXTRINSB` | |
| `VEXTRINSH` | |
| `VEXTRINSV` | |
| `VEXTRINSW` | |
| `VFCLASSD` | |
| `VFCLASSF` | |
| `VFRECIPD` | |
| `VFRECIPF` | |
| `VFRINTD` | |
| `VFRINTF` | |
| `VFRINTRMD` | |
| `VFRINTRMF` | |
| `VFRINTRNED` | |
| `VFRINTRNEF` | |
| `VFRINTRPD` | |
| `VFRINTRPF` | |
| `VFRINTRZD` | |
| `VFRINTRZF` | |
| `VFRSQRTD` | |
| `VFRSQRTF` | |
| `VFSQRTD` | |
| `VFSQRTF` | |
| `VILVHB` | |
| `VILVHH` | |
| `VILVHV` | |
| `VILVHW` | |
| `VILVLB` | |
| `VILVLH` | |
| `VILVLV` | |
| `VILVLW` | |
| `VMADDB` | |
| `VMADDH` | |
| `VMADDV` | |
| `VMADDW` | |
| `VMADDWEVHB` | |
| `VMADDWEVHBU` | |
| `VMADDWEVHBUB` | |
| `VMADDWEVQV` | |
| `VMADDWEVQVU` | |
| `VMADDWEVQVUV` | |
| `VMADDWEVVW` | |
| `VMADDWEVVWU` | |
| `VMADDWEVVWUW` | |
| `VMADDWEVWH` | |
| `VMADDWEVWHU` | |
| `VMADDWEVWHUH` | |
| `VMADDWODHB` | |
| `VMADDWODHBU` | |
| `VMADDWODHBUB` | |
| `VMADDWODQV` | |
| `VMADDWODQVU` | |
| `VMADDWODQVUV` | |
| `VMADDWODVW` | |
| `VMADDWODVWU` | |
| `VMADDWODVWUW` | |
| `VMADDWODWH` | |
| `VMADDWODWHU` | |
| `VMADDWODWHUH` | |
| `VMODB` | |
| `VMODBU` | |
| `VMODH` | |
| `VMODHU` | |
| `VMODV` | |
| `VMODVU` | |
| `VMODW` | |
| `VMODWU` | |
| `VMOVQ` | |
| `VMSUBB` | |
| `VMSUBH` | |
| `VMSUBV` | |
| `VMSUBW` | |
| `VMUHB` | |
| `VMUHBU` | |
| `VMUHH` | |
| `VMUHHU` | |
| `VMUHV` | |
| `VMUHVU` | |
| `VMUHW` | |
| `VMUHWU` | |
| `VMULB` | |
| `VMULD` | |
| `VMULF` | |
| `VMULH` | |
| `VMULV` | |
| `VMULW` | |
| `VMULWEVHB` | |
| `VMULWEVHBU` | |
| `VMULWEVHBUB` | |
| `VMULWEVQV` | |
| `VMULWEVQVU` | |
| `VMULWEVQVUV` | |
| `VMULWEVVW` | |
| `VMULWEVVWU` | |
| `VMULWEVVWUW` | |
| `VMULWEVWH` | |
| `VMULWEVWHU` | |
| `VMULWEVWHUH` | |
| `VMULWODHB` | |
| `VMULWODHBU` | |
| `VMULWODHBUB` | |
| `VMULWODQV` | |
| `VMULWODQVU` | |
| `VMULWODQVUV` | |
| `VMULWODVW` | |
| `VMULWODVWU` | |
| `VMULWODVWUW` | |
| `VMULWODWH` | |
| `VMULWODWHU` | |
| `VMULWODWHUH` | |
| `VNEGB` | |
| `VNEGH` | |
| `VNEGV` | |
| `VNEGW` | |
| `VNORB` | |
| `VNORV` | |
| `VORB` | |
| `VORNV` | |
| `VORV` | |
| `VPCNTB` | |
| `VPCNTH` | |
| `VPCNTV` | |
| `VPCNTW` | |
| `VPERMIW` | |
| `VROTRB` | |
| `VROTRH` | |
| `VROTRV` | |
| `VROTRW` | |
| `VSADDB` | |
| `VSADDBU` | |
| `VSADDH` | |
| `VSADDHU` | |
| `VSADDV` | |
| `VSADDVU` | |
| `VSADDW` | |
| `VSADDWU` | |
| `VSEQB` | |
| `VSEQH` | |
| `VSEQV` | |
| `VSEQW` | |
| `VSETALLNEB` | |
| `VSETALLNEH` | |
| `VSETALLNEV` | |
| `VSETALLNEW` | |
| `VSETANYEQB` | |
| `VSETANYEQH` | |
| `VSETANYEQV` | |
| `VSETANYEQW` | |
| `VSETEQV` | |
| `VSETNEV` | |
| `VSHUF4IB` | |
| `VSHUF4IH` | |
| `VSHUF4IV` | |
| `VSHUF4IW` | |
| `VSHUFB` | |
| `VSHUFH` | |
| `VSHUFV` | |
| `VSHUFW` | |
| `VSLLB` | |
| `VSLLH` | |
| `VSLLV` | |
| `VSLLW` | |
| `VSLTB` | |
| `VSLTBU` | |
| `VSLTH` | |
| `VSLTHU` | |
| `VSLTV` | |
| `VSLTVU` | |
| `VSLTW` | |
| `VSLTWU` | |
| `VSRAB` | |
| `VSRAH` | |
| `VSRAV` | |
| `VSRAW` | |
| `VSRLB` | |
| `VSRLH` | |
| `VSRLV` | |
| `VSRLW` | |
| `VSSUBB` | |
| `VSSUBBU` | |
| `VSSUBH` | |
| `VSSUBHU` | |
| `VSSUBV` | |
| `VSSUBVU` | |
| `VSSUBW` | |
| `VSSUBWU` | |
| `VSUBB` | |
| `VSUBBU` | |
| `VSUBD` | |
| `VSUBF` | |
| `VSUBH` | |
| `VSUBHU` | |
| `VSUBQ` | |
| `VSUBV` | |
| `VSUBVU` | |
| `VSUBW` | |
| `VSUBWEVHB` | |
| `VSUBWEVHBU` | |
| `VSUBWEVQV` | |
| `VSUBWEVQVU` | |
| `VSUBWEVVW` | |
| `VSUBWEVVWU` | |
| `VSUBWEVWH` | |
| `VSUBWEVWHU` | |
| `VSUBWODHB` | |
| `VSUBWODHBU` | |
| `VSUBWODQV` | |
| `VSUBWODQVU` | |
| `VSUBWODVW` | |
| `VSUBWODVWU` | |
| `VSUBWODWH` | |
| `VSUBWODWHU` | |
| `VSUBWU` | |
| `VXORB` | |
| `VXORV` | |
| `WORD` | |
| `XOR` | Bitwise XOR |
| `XVADDB` | |
| `XVADDBU` | |
| `XVADDD` | |
| `XVADDF` | |
| `XVADDH` | |
| `XVADDHU` | |
| `XVADDQ` | |
| `XVADDV` | |
| `XVADDVU` | |
| `XVADDW` | |
| `XVADDWEVHB` | |
| `XVADDWEVHBU` | |
| `XVADDWEVQV` | |
| `XVADDWEVQVU` | |
| `XVADDWEVVW` | |
| `XVADDWEVVWU` | |
| `XVADDWEVWH` | |
| `XVADDWEVWHU` | |
| `XVADDWODHB` | |
| `XVADDWODHBU` | |
| `XVADDWODQV` | |
| `XVADDWODQVU` | |
| `XVADDWODVW` | |
| `XVADDWODVWU` | |
| `XVADDWODWH` | |
| `XVADDWODWHU` | |
| `XVADDWU` | |
| `XVANDB` | |
| `XVANDNV` | |
| `XVANDV` | |
| `XVBITCLRB` | |
| `XVBITCLRH` | |
| `XVBITCLRV` | |
| `XVBITCLRW` | |
| `XVBITREVB` | |
| `XVBITREVH` | |
| `XVBITREVV` | |
| `XVBITREVW` | |
| `XVBITSETB` | |
| `XVBITSETH` | |
| `XVBITSETV` | |
| `XVBITSETW` | |
| `XVDIVB` | |
| `XVDIVBU` | |
| `XVDIVD` | |
| `XVDIVF` | |
| `XVDIVH` | |
| `XVDIVHU` | |
| `XVDIVV` | |
| `XVDIVVU` | |
| `XVDIVW` | |
| `XVDIVWU` | |
| `XVEXTRINSB` | |
| `XVEXTRINSH` | |
| `XVEXTRINSV` | |
| `XVEXTRINSW` | |
| `XVFCLASSD` | |
| `XVFCLASSF` | |
| `XVFRECIPD` | |
| `XVFRECIPF` | |
| `XVFRINTD` | |
| `XVFRINTF` | |
| `XVFRINTRMD` | |
| `XVFRINTRMF` | |
| `XVFRINTRNED` | |
| `XVFRINTRNEF` | |
| `XVFRINTRPD` | |
| `XVFRINTRPF` | |
| `XVFRINTRZD` | |
| `XVFRINTRZF` | |
| `XVFRSQRTD` | |
| `XVFRSQRTF` | |
| `XVFSQRTD` | |
| `XVFSQRTF` | |
| `XVILVHB` | |
| `XVILVHH` | |
| `XVILVHV` | |
| `XVILVHW` | |
| `XVILVLB` | |
| `XVILVLH` | |
| `XVILVLV` | |
| `XVILVLW` | |
| `XVMADDB` | |
| `XVMADDH` | |
| `XVMADDV` | |
| `XVMADDW` | |
| `XVMADDWEVHB` | |
| `XVMADDWEVHBU` | |
| `XVMADDWEVHBUB` | |
| `XVMADDWEVQV` | |
| `XVMADDWEVQVU` | |
| `XVMADDWEVQVUV` | |
| `XVMADDWEVVW` | |
| `XVMADDWEVVWU` | |
| `XVMADDWEVVWUW` | |
| `XVMADDWEVWH` | |
| `XVMADDWEVWHU` | |
| `XVMADDWEVWHUH` | |
| `XVMADDWODHB` | |
| `XVMADDWODHBU` | |
| `XVMADDWODHBUB` | |
| `XVMADDWODQV` | |
| `XVMADDWODQVU` | |
| `XVMADDWODQVUV` | |
| `XVMADDWODVW` | |
| `XVMADDWODVWU` | |
| `XVMADDWODVWUW` | |
| `XVMADDWODWH` | |
| `XVMADDWODWHU` | |
| `XVMADDWODWHUH` | |
| `XVMODB` | |
| `XVMODBU` | |
| `XVMODH` | |
| `XVMODHU` | |
| `XVMODV` | |
| `XVMODVU` | |
| `XVMODW` | |
| `XVMODWU` | |
| `XVMOVQ` | |
| `XVMSUBB` | |
| `XVMSUBH` | |
| `XVMSUBV` | |
| `XVMSUBW` | |
| `XVMUHB` | |
| `XVMUHBU` | |
| `XVMUHH` | |
| `XVMUHHU` | |
| `XVMUHV` | |
| `XVMUHVU` | |
| `XVMUHW` | |
| `XVMUHWU` | |
| `XVMULB` | |
| `XVMULD` | |
| `XVMULF` | |
| `XVMULH` | |
| `XVMULV` | |
| `XVMULW` | |
| `XVMULWEVHB` | |
| `XVMULWEVHBU` | |
| `XVMULWEVHBUB` | |
| `XVMULWEVQV` | |
| `XVMULWEVQVU` | |
| `XVMULWEVQVUV` | |
| `XVMULWEVVW` | |
| `XVMULWEVVWU` | |
| `XVMULWEVVWUW` | |
| `XVMULWEVWH` | |
| `XVMULWEVWHU` | |
| `XVMULWEVWHUH` | |
| `XVMULWODHB` | |
| `XVMULWODHBU` | |
| `XVMULWODHBUB` | |
| `XVMULWODQV` | |
| `XVMULWODQVU` | |
| `XVMULWODQVUV` | |
| `XVMULWODVW` | |
| `XVMULWODVWU` | |
| `XVMULWODVWUW` | |
| `XVMULWODWH` | |
| `XVMULWODWHU` | |
| `XVMULWODWHUH` | |
| `XVNEGB` | |
| `XVNEGH` | |
| `XVNEGV` | |
| `XVNEGW` | |
| `XVNORB` | |
| `XVNORV` | |
| `XVORB` | |
| `XVORNV` | |
| `XVORV` | |
| `XVPCNTB` | |
| `XVPCNTH` | |
| `XVPCNTV` | |
| `XVPCNTW` | |
| `XVPERMIQ` | |
| `XVPERMIV` | |
| `XVPERMIW` | |
| `XVROTRB` | |
| `XVROTRH` | |
| `XVROTRV` | |
| `XVROTRW` | |
| `XVSADDB` | |
| `XVSADDBU` | |
| `XVSADDH` | |
| `XVSADDHU` | |
| `XVSADDV` | |
| `XVSADDVU` | |
| `XVSADDW` | |
| `XVSADDWU` | |
| `XVSEQB` | |
| `XVSEQH` | |
| `XVSEQV` | |
| `XVSEQW` | |
| `XVSETALLNEB` | |
| `XVSETALLNEH` | |
| `XVSETALLNEV` | |
| `XVSETALLNEW` | |
| `XVSETANYEQB` | |
| `XVSETANYEQH` | |
| `XVSETANYEQV` | |
| `XVSETANYEQW` | |
| `XVSETEQV` | |
| `XVSETNEV` | |
| `XVSHUF4IB` | |
| `XVSHUF4IH` | |
| `XVSHUF4IV` | |
| `XVSHUF4IW` | |
| `XVSHUFB` | |
| `XVSHUFH` | |
| `XVSHUFV` | |
| `XVSHUFW` | |
| `XVSLLB` | |
| `XVSLLH` | |
| `XVSLLV` | |
| `XVSLLW` | |
| `XVSLTB` | |
| `XVSLTBU` | |
| `XVSLTH` | |
| `XVSLTHU` | |
| `XVSLTV` | |
| `XVSLTVU` | |
| `XVSLTW` | |
| `XVSLTWU` | |
| `XVSRAB` | |
| `XVSRAH` | |
| `XVSRAV` | |
| `XVSRAW` | |
| `XVSRLB` | |
| `XVSRLH` | |
| `XVSRLV` | |
| `XVSRLW` | |
| `XVSSUBB` | |
| `XVSSUBBU` | |
| `XVSSUBH` | |
| `XVSSUBHU` | |
| `XVSSUBV` | |
| `XVSSUBVU` | |
| `XVSSUBW` | |
| `XVSSUBWU` | |
| `XVSUBB` | |
| `XVSUBBU` | |
| `XVSUBD` | |
| `XVSUBF` | |
| `XVSUBH` | |
| `XVSUBHU` | |
| `XVSUBQ` | |
| `XVSUBV` | |
| `XVSUBVU` | |
| `XVSUBW` | |
| `XVSUBWEVHB` | |
| `XVSUBWEVHBU` | |
| `XVSUBWEVQV` | |
| `XVSUBWEVQVU` | |
| `XVSUBWEVVW` | |
| `XVSUBWEVVWU` | |
| `XVSUBWEVWH` | |
| `XVSUBWEVWHU` | |
| `XVSUBWODHB` | |
| `XVSUBWODHBU` | |
| `XVSUBWODQV` | |
| `XVSUBWODQVU` | |
| `XVSUBWODVW` | |
| `XVSUBWODVWU` | |
| `XVSUBWODWH` | |
| `XVSUBWODWHU` | |
| `XVSUBWU` | |
| `XVXORB` | |
| `XVXORV` | |
| `JAL` | |
Recognised: 814 mnemonics.
+991
View File
@@ -0,0 +1,991 @@
# RISC-V 64: instruction inventory
Generated by gasm-devkit's `_gen` from the Go toolchain's instruction table
(`cmd/internal/obj/riscv/anames.go`, go1.27.1); DO NOT EDIT. This page lists every mnemonic
`go tool asm` accepts on this target, which is the upper bound of the
language on it: a name absent here is not an instruction of the target,
and a name present here may still be one gasm's encoder cannot emit yet.
The inventory carries no per-mnemonic encoder column: on this target
encodability is decided per operand shape, and the live measured
coverage is reported by `gasm audit-instructions`.
| Mnemonic | Notes |
|---|---|
| `CALL` | Call subroutine |
| `DUFFCOPY` | |
| `DUFFZERO` | |
| `END` | |
| `FUNCDATA` | |
| `GETCALLERPC` | |
| `JMP` | Unconditional jump |
| `NOP` | |
| `PCALIGN` | |
| `PCALIGNMAX` | |
| `PCDATA` | |
| `RET` | Return |
| `TEXT` | |
| `UNDEF` | |
| `ADD` | Integer add |
| `ADDI` | Add immediate |
| `ADDIW` | Add immediate (32-bit) |
| `ADDUW` | |
| `ADDW` | Add (32-bit) |
| `AMOADDD` | Atomic add doubleword |
| `AMOADDW` | Atomic add word |
| `AMOANDD` | |
| `AMOANDW` | |
| `AMOMAXD` | |
| `AMOMAXUD` | |
| `AMOMAXUW` | |
| `AMOMAXW` | |
| `AMOMIND` | |
| `AMOMINUD` | |
| `AMOMINUW` | |
| `AMOMINW` | |
| `AMOORD` | |
| `AMOORW` | |
| `AMOSWAPD` | Atomic swap doubleword |
| `AMOSWAPW` | Atomic swap word |
| `AMOXORD` | |
| `AMOXORW` | |
| `AND` | Bitwise AND |
| `ANDI` | AND immediate |
| `ANDN` | |
| `AUIPC` | Add upper immediate to PC |
| `BCLR` | |
| `BCLRI` | |
| `BEQ` | Branch if equal |
| `BEQZ` | |
| `BEXT` | |
| `BEXTI` | |
| `BGE` | Branch if greater or equal |
| `BGEU` | Branch if greater or equal unsigned |
| `BGEZ` | |
| `BGT` | |
| `BGTU` | |
| `BGTZ` | |
| `BINV` | |
| `BINVI` | |
| `BLE` | |
| `BLEU` | |
| `BLEZ` | |
| `BLT` | Branch if less than |
| `BLTU` | Branch if less than unsigned |
| `BLTZ` | |
| `BNE` | Branch if not equal |
| `BNEZ` | |
| `BSET` | |
| `BSETI` | |
| `CADD` | |
| `CADDI` | |
| `CADDI16SP` | |
| `CADDI4SPN` | |
| `CADDIW` | |
| `CADDW` | |
| `CAND` | |
| `CANDI` | |
| `CBEQZ` | |
| `CBNEZ` | |
| `CEBREAK` | |
| `CFLD` | |
| `CFLDSP` | |
| `CFSD` | |
| `CFSDSP` | |
| `CJ` | |
| `CJALR` | |
| `CJR` | |
| `CLD` | |
| `CLDSP` | |
| `CLI` | |
| `CLUI` | |
| `CLW` | |
| `CLWSP` | |
| `CLZ` | |
| `CLZW` | |
| `CMV` | |
| `CNOP` | |
| `COR` | |
| `CPOP` | |
| `CPOPW` | |
| `CSD` | |
| `CSDSP` | |
| `CSLLI` | |
| `CSRAI` | |
| `CSRLI` | |
| `CSRRC` | |
| `CSRRCI` | |
| `CSRRS` | |
| `CSRRSI` | |
| `CSRRW` | |
| `CSRRWI` | |
| `CSUB` | |
| `CSUBW` | |
| `CSW` | |
| `CSWSP` | |
| `CTZ` | |
| `CTZW` | |
| `CXOR` | |
| `CZEROEQZ` | |
| `CZERONEZ` | |
| `DIV` | Divide |
| `DIVU` | Divide unsigned |
| `DIVUW` | |
| `DIVW` | Divide (32-bit) |
| `DRET` | |
| `EBREAK` | Breakpoint |
| `ECALL` | Environment call |
| `FABSD` | |
| `FABSS` | |
| `FADDD` | FP add (double) |
| `FADDQ` | |
| `FADDS` | FP add (single) |
| `FCLASSD` | |
| `FCLASSQ` | |
| `FCLASSS` | |
| `FCVTDL` | |
| `FCVTDLU` | |
| `FCVTDQ` | |
| `FCVTDS` | |
| `FCVTDW` | |
| `FCVTDWU` | |
| `FCVTLD` | |
| `FCVTLQ` | |
| `FCVTLS` | |
| `FCVTLUD` | |
| `FCVTLUQ` | |
| `FCVTLUS` | |
| `FCVTQD` | |
| `FCVTQL` | |
| `FCVTQLU` | |
| `FCVTQS` | |
| `FCVTQW` | |
| `FCVTQWU` | |
| `FCVTSD` | |
| `FCVTSL` | |
| `FCVTSLU` | |
| `FCVTSQ` | |
| `FCVTSW` | |
| `FCVTSWU` | |
| `FCVTWD` | |
| `FCVTWQ` | |
| `FCVTWS` | |
| `FCVTWUD` | |
| `FCVTWUQ` | |
| `FCVTWUS` | |
| `FDIVD` | FP divide (double) |
| `FDIVQ` | |
| `FDIVS` | FP divide (single) |
| `FENCE` | Memory barrier |
| `FEQD` | |
| `FEQQ` | |
| `FEQS` | |
| `FLD` | FP load doubleword |
| `FLED` | |
| `FLEQ` | |
| `FLES` | |
| `FLQ` | |
| `FLTD` | |
| `FLTQ` | |
| `FLTS` | |
| `FLW` | FP load word |
| `FMADDD` | |
| `FMADDQ` | |
| `FMADDS` | |
| `FMAXD` | |
| `FMAXQ` | |
| `FMAXS` | |
| `FMIND` | |
| `FMINQ` | |
| `FMINS` | |
| `FMSUBD` | |
| `FMSUBQ` | |
| `FMSUBS` | |
| `FMULD` | FP multiply (double) |
| `FMULQ` | |
| `FMULS` | FP multiply (single) |
| `FMVDX` | |
| `FMVSX` | |
| `FMVWX` | |
| `FMVXD` | |
| `FMVXS` | |
| `FMVXW` | |
| `FNED` | |
| `FNEGD` | |
| `FNEGS` | |
| `FNES` | |
| `FNMADDD` | |
| `FNMADDQ` | |
| `FNMADDS` | |
| `FNMSUBD` | |
| `FNMSUBQ` | |
| `FNMSUBS` | |
| `FSD` | FP store doubleword |
| `FSGNJD` | |
| `FSGNJND` | |
| `FSGNJNQ` | |
| `FSGNJNS` | |
| `FSGNJQ` | |
| `FSGNJS` | |
| `FSGNJXD` | |
| `FSGNJXQ` | |
| `FSGNJXS` | |
| `FSQ` | |
| `FSQRTD` | |
| `FSQRTQ` | |
| `FSQRTS` | |
| `FSUBD` | FP subtract (double) |
| `FSUBQ` | |
| `FSUBS` | FP subtract (single) |
| `FSW` | FP store word |
| `JAL` | Jump and link |
| `JALR` | Jump and link register |
| `LB` | Load byte |
| `LBU` | Load byte unsigned |
| `LD` | Load doubleword |
| `LH` | Load halfword |
| `LHU` | Load halfword unsigned |
| `LRD` | Load-reserved doubleword |
| `LRW` | Load-reserved word |
| `LUI` | Load upper immediate |
| `LW` | Load word |
| `LWU` | Load word unsigned |
| `MAX` | |
| `MAXU` | |
| `MIN` | |
| `MINU` | |
| `MOV` | |
| `MOVB` | |
| `MOVBU` | |
| `MOVD` | |
| `MOVF` | |
| `MOVH` | |
| `MOVHU` | |
| `MOVW` | |
| `MOVWU` | |
| `MRET` | |
| `MUL` | Multiply |
| `MULH` | Multiply high |
| `MULHSU` | Multiply high signed/unsigned |
| `MULHU` | Multiply high unsigned |
| `MULW` | Multiply (32-bit) |
| `NEG` | |
| `NEGW` | |
| `NOT` | |
| `OR` | Bitwise OR |
| `ORCB` | |
| `ORI` | OR immediate |
| `ORN` | |
| `RDCYCLE` | |
| `RDINSTRET` | |
| `RDTIME` | |
| `REM` | Remainder |
| `REMU` | Remainder unsigned |
| `REMUW` | |
| `REMW` | |
| `REV8` | |
| `ROL` | |
| `ROLW` | |
| `ROR` | |
| `RORI` | |
| `RORIW` | |
| `RORW` | |
| `SB` | Store byte |
| `SBREAK` | |
| `SCALL` | |
| `SCD` | Store-conditional doubleword |
| `SCW` | Store-conditional word |
| `SD` | Store doubleword |
| `SEQZ` | |
| `SEXTB` | |
| `SEXTH` | |
| `SFENCEVMA` | |
| `SH` | Store halfword |
| `SH1ADD` | |
| `SH1ADDUW` | |
| `SH2ADD` | |
| `SH2ADDUW` | |
| `SH3ADD` | |
| `SH3ADDUW` | |
| `SLL` | Shift left logical |
| `SLLI` | Shift left logical immediate |
| `SLLIUW` | |
| `SLLIW` | |
| `SLLW` | |
| `SLT` | Set if less than |
| `SLTI` | Set if less than immediate |
| `SLTIU` | Set if less than unsigned immediate |
| `SLTU` | Set if less than unsigned |
| `SNEZ` | |
| `SRA` | Shift right arithmetic |
| `SRAI` | Shift right arithmetic immediate |
| `SRAIW` | |
| `SRAW` | |
| `SRET` | |
| `SRL` | Shift right logical |
| `SRLI` | Shift right logical immediate |
| `SRLIW` | |
| `SRLW` | |
| `SUB` | Integer subtract |
| `SUBW` | Subtract (32-bit) |
| `SW` | Store word |
| `VAADDUVV` | |
| `VAADDUVX` | |
| `VAADDVV` | |
| `VAADDVX` | |
| `VADCVIM` | |
| `VADCVVM` | |
| `VADCVXM` | |
| `VADDVI` | |
| `VADDVV` | |
| `VADDVX` | |
| `VANDVI` | |
| `VANDVV` | |
| `VANDVX` | |
| `VASUBUVV` | |
| `VASUBUVX` | |
| `VASUBVV` | |
| `VASUBVX` | |
| `VCOMPRESSVM` | |
| `VCPOPM` | |
| `VDIVUVV` | |
| `VDIVUVX` | |
| `VDIVVV` | |
| `VDIVVX` | |
| `VFABSV` | |
| `VFADDVF` | |
| `VFADDVV` | |
| `VFCLASSV` | |
| `VFCVTFXUV` | |
| `VFCVTFXV` | |
| `VFCVTRTZXFV` | |
| `VFCVTRTZXUFV` | |
| `VFCVTXFV` | |
| `VFCVTXUFV` | |
| `VFDIVVF` | |
| `VFDIVVV` | |
| `VFIRSTM` | |
| `VFMACCVF` | |
| `VFMACCVV` | |
| `VFMADDVF` | |
| `VFMADDVV` | |
| `VFMAXVF` | |
| `VFMAXVV` | |
| `VFMERGEVFM` | |
| `VFMINVF` | |
| `VFMINVV` | |
| `VFMSACVF` | |
| `VFMSACVV` | |
| `VFMSUBVF` | |
| `VFMSUBVV` | |
| `VFMULVF` | |
| `VFMULVV` | |
| `VFMVFS` | |
| `VFMVSF` | |
| `VFMVVF` | |
| `VFNCVTFFW` | |
| `VFNCVTFXUW` | |
| `VFNCVTFXW` | |
| `VFNCVTRODFFW` | |
| `VFNCVTRTZXFW` | |
| `VFNCVTRTZXUFW` | |
| `VFNCVTXFW` | |
| `VFNCVTXUFW` | |
| `VFNEGV` | |
| `VFNMACCVF` | |
| `VFNMACCVV` | |
| `VFNMADDVF` | |
| `VFNMADDVV` | |
| `VFNMSACVF` | |
| `VFNMSACVV` | |
| `VFNMSUBVF` | |
| `VFNMSUBVV` | |
| `VFRDIVVF` | |
| `VFREC7V` | |
| `VFREDMAXVS` | |
| `VFREDMINVS` | |
| `VFREDOSUMVS` | |
| `VFREDUSUMVS` | |
| `VFRSQRT7V` | |
| `VFRSUBVF` | |
| `VFSGNJNVF` | |
| `VFSGNJNVV` | |
| `VFSGNJVF` | |
| `VFSGNJVV` | |
| `VFSGNJXVF` | |
| `VFSGNJXVV` | |
| `VFSLIDE1DOWNVF` | |
| `VFSLIDE1UPVF` | |
| `VFSQRTV` | |
| `VFSUBVF` | |
| `VFSUBVV` | |
| `VFWADDVF` | |
| `VFWADDVV` | |
| `VFWADDWF` | |
| `VFWADDWV` | |
| `VFWCVTFFV` | |
| `VFWCVTFXUV` | |
| `VFWCVTFXV` | |
| `VFWCVTRTZXFV` | |
| `VFWCVTRTZXUFV` | |
| `VFWCVTXFV` | |
| `VFWCVTXUFV` | |
| `VFWMACCVF` | |
| `VFWMACCVV` | |
| `VFWMSACVF` | |
| `VFWMSACVV` | |
| `VFWMULVF` | |
| `VFWMULVV` | |
| `VFWNMACCVF` | |
| `VFWNMACCVV` | |
| `VFWNMSACVF` | |
| `VFWNMSACVV` | |
| `VFWREDOSUMVS` | |
| `VFWREDUSUMVS` | |
| `VFWSUBVF` | |
| `VFWSUBVV` | |
| `VFWSUBWF` | |
| `VFWSUBWV` | |
| `VIDV` | |
| `VIOTAM` | |
| `VL1RE16V` | |
| `VL1RE32V` | |
| `VL1RE64V` | |
| `VL1RE8V` | |
| `VL1RV` | |
| `VL2RE16V` | |
| `VL2RE32V` | |
| `VL2RE64V` | |
| `VL2RE8V` | |
| `VL2RV` | |
| `VL4RE16V` | |
| `VL4RE32V` | |
| `VL4RE64V` | |
| `VL4RE8V` | |
| `VL4RV` | |
| `VL8RE16V` | |
| `VL8RE32V` | |
| `VL8RE64V` | |
| `VL8RE8V` | |
| `VL8RV` | |
| `VLE16FFV` | |
| `VLE16V` | |
| `VLE32FFV` | |
| `VLE32V` | |
| `VLE64FFV` | |
| `VLE64V` | |
| `VLE8FFV` | |
| `VLE8V` | |
| `VLMV` | |
| `VLOXEI16V` | |
| `VLOXEI32V` | |
| `VLOXEI64V` | |
| `VLOXEI8V` | |
| `VLOXSEG2EI16V` | |
| `VLOXSEG2EI32V` | |
| `VLOXSEG2EI64V` | |
| `VLOXSEG2EI8V` | |
| `VLOXSEG3EI16V` | |
| `VLOXSEG3EI32V` | |
| `VLOXSEG3EI64V` | |
| `VLOXSEG3EI8V` | |
| `VLOXSEG4EI16V` | |
| `VLOXSEG4EI32V` | |
| `VLOXSEG4EI64V` | |
| `VLOXSEG4EI8V` | |
| `VLOXSEG5EI16V` | |
| `VLOXSEG5EI32V` | |
| `VLOXSEG5EI64V` | |
| `VLOXSEG5EI8V` | |
| `VLOXSEG6EI16V` | |
| `VLOXSEG6EI32V` | |
| `VLOXSEG6EI64V` | |
| `VLOXSEG6EI8V` | |
| `VLOXSEG7EI16V` | |
| `VLOXSEG7EI32V` | |
| `VLOXSEG7EI64V` | |
| `VLOXSEG7EI8V` | |
| `VLOXSEG8EI16V` | |
| `VLOXSEG8EI32V` | |
| `VLOXSEG8EI64V` | |
| `VLOXSEG8EI8V` | |
| `VLSE16V` | |
| `VLSE32V` | |
| `VLSE64V` | |
| `VLSE8V` | |
| `VLSEG2E16FFV` | |
| `VLSEG2E16V` | |
| `VLSEG2E32FFV` | |
| `VLSEG2E32V` | |
| `VLSEG2E64FFV` | |
| `VLSEG2E64V` | |
| `VLSEG2E8FFV` | |
| `VLSEG2E8V` | |
| `VLSEG3E16FFV` | |
| `VLSEG3E16V` | |
| `VLSEG3E32FFV` | |
| `VLSEG3E32V` | |
| `VLSEG3E64FFV` | |
| `VLSEG3E64V` | |
| `VLSEG3E8FFV` | |
| `VLSEG3E8V` | |
| `VLSEG4E16FFV` | |
| `VLSEG4E16V` | |
| `VLSEG4E32FFV` | |
| `VLSEG4E32V` | |
| `VLSEG4E64FFV` | |
| `VLSEG4E64V` | |
| `VLSEG4E8FFV` | |
| `VLSEG4E8V` | |
| `VLSEG5E16FFV` | |
| `VLSEG5E16V` | |
| `VLSEG5E32FFV` | |
| `VLSEG5E32V` | |
| `VLSEG5E64FFV` | |
| `VLSEG5E64V` | |
| `VLSEG5E8FFV` | |
| `VLSEG5E8V` | |
| `VLSEG6E16FFV` | |
| `VLSEG6E16V` | |
| `VLSEG6E32FFV` | |
| `VLSEG6E32V` | |
| `VLSEG6E64FFV` | |
| `VLSEG6E64V` | |
| `VLSEG6E8FFV` | |
| `VLSEG6E8V` | |
| `VLSEG7E16FFV` | |
| `VLSEG7E16V` | |
| `VLSEG7E32FFV` | |
| `VLSEG7E32V` | |
| `VLSEG7E64FFV` | |
| `VLSEG7E64V` | |
| `VLSEG7E8FFV` | |
| `VLSEG7E8V` | |
| `VLSEG8E16FFV` | |
| `VLSEG8E16V` | |
| `VLSEG8E32FFV` | |
| `VLSEG8E32V` | |
| `VLSEG8E64FFV` | |
| `VLSEG8E64V` | |
| `VLSEG8E8FFV` | |
| `VLSEG8E8V` | |
| `VLSSEG2E16V` | |
| `VLSSEG2E32V` | |
| `VLSSEG2E64V` | |
| `VLSSEG2E8V` | |
| `VLSSEG3E16V` | |
| `VLSSEG3E32V` | |
| `VLSSEG3E64V` | |
| `VLSSEG3E8V` | |
| `VLSSEG4E16V` | |
| `VLSSEG4E32V` | |
| `VLSSEG4E64V` | |
| `VLSSEG4E8V` | |
| `VLSSEG5E16V` | |
| `VLSSEG5E32V` | |
| `VLSSEG5E64V` | |
| `VLSSEG5E8V` | |
| `VLSSEG6E16V` | |
| `VLSSEG6E32V` | |
| `VLSSEG6E64V` | |
| `VLSSEG6E8V` | |
| `VLSSEG7E16V` | |
| `VLSSEG7E32V` | |
| `VLSSEG7E64V` | |
| `VLSSEG7E8V` | |
| `VLSSEG8E16V` | |
| `VLSSEG8E32V` | |
| `VLSSEG8E64V` | |
| `VLSSEG8E8V` | |
| `VLUXEI16V` | |
| `VLUXEI32V` | |
| `VLUXEI64V` | |
| `VLUXEI8V` | |
| `VLUXSEG2EI16V` | |
| `VLUXSEG2EI32V` | |
| `VLUXSEG2EI64V` | |
| `VLUXSEG2EI8V` | |
| `VLUXSEG3EI16V` | |
| `VLUXSEG3EI32V` | |
| `VLUXSEG3EI64V` | |
| `VLUXSEG3EI8V` | |
| `VLUXSEG4EI16V` | |
| `VLUXSEG4EI32V` | |
| `VLUXSEG4EI64V` | |
| `VLUXSEG4EI8V` | |
| `VLUXSEG5EI16V` | |
| `VLUXSEG5EI32V` | |
| `VLUXSEG5EI64V` | |
| `VLUXSEG5EI8V` | |
| `VLUXSEG6EI16V` | |
| `VLUXSEG6EI32V` | |
| `VLUXSEG6EI64V` | |
| `VLUXSEG6EI8V` | |
| `VLUXSEG7EI16V` | |
| `VLUXSEG7EI32V` | |
| `VLUXSEG7EI64V` | |
| `VLUXSEG7EI8V` | |
| `VLUXSEG8EI16V` | |
| `VLUXSEG8EI32V` | |
| `VLUXSEG8EI64V` | |
| `VLUXSEG8EI8V` | |
| `VMACCVV` | |
| `VMACCVX` | |
| `VMADCVI` | |
| `VMADCVIM` | |
| `VMADCVV` | |
| `VMADCVVM` | |
| `VMADCVX` | |
| `VMADCVXM` | |
| `VMADDVV` | |
| `VMADDVX` | |
| `VMANDMM` | |
| `VMANDNMM` | |
| `VMAXUVV` | |
| `VMAXUVX` | |
| `VMAXVV` | |
| `VMAXVX` | |
| `VMCLRM` | |
| `VMERGEVIM` | |
| `VMERGEVVM` | |
| `VMERGEVXM` | |
| `VMFEQVF` | |
| `VMFEQVV` | |
| `VMFGEVF` | |
| `VMFGEVV` | |
| `VMFGTVF` | |
| `VMFGTVV` | |
| `VMFLEVF` | |
| `VMFLEVV` | |
| `VMFLTVF` | |
| `VMFLTVV` | |
| `VMFNEVF` | |
| `VMFNEVV` | |
| `VMINUVV` | |
| `VMINUVX` | |
| `VMINVV` | |
| `VMINVX` | |
| `VMMVM` | |
| `VMNANDMM` | |
| `VMNORMM` | |
| `VMNOTM` | |
| `VMORMM` | |
| `VMORNMM` | |
| `VMSBCVV` | |
| `VMSBCVVM` | |
| `VMSBCVX` | |
| `VMSBCVXM` | |
| `VMSBFM` | |
| `VMSEQVI` | |
| `VMSEQVV` | |
| `VMSEQVX` | |
| `VMSETM` | |
| `VMSGEUVI` | |
| `VMSGEUVV` | |
| `VMSGEVI` | |
| `VMSGEVV` | |
| `VMSGTUVI` | |
| `VMSGTUVV` | |
| `VMSGTUVX` | |
| `VMSGTVI` | |
| `VMSGTVV` | |
| `VMSGTVX` | |
| `VMSIFM` | |
| `VMSLEUVI` | |
| `VMSLEUVV` | |
| `VMSLEUVX` | |
| `VMSLEVI` | |
| `VMSLEVV` | |
| `VMSLEVX` | |
| `VMSLTUVI` | |
| `VMSLTUVV` | |
| `VMSLTUVX` | |
| `VMSLTVI` | |
| `VMSLTVV` | |
| `VMSLTVX` | |
| `VMSNEVI` | |
| `VMSNEVV` | |
| `VMSNEVX` | |
| `VMSOFM` | |
| `VMULHSUVV` | |
| `VMULHSUVX` | |
| `VMULHUVV` | |
| `VMULHUVX` | |
| `VMULHVV` | |
| `VMULHVX` | |
| `VMULVV` | |
| `VMULVX` | |
| `VMV1RV` | |
| `VMV2RV` | |
| `VMV4RV` | |
| `VMV8RV` | |
| `VMVSX` | |
| `VMVVI` | |
| `VMVVV` | |
| `VMVVX` | |
| `VMVXS` | |
| `VMXNORMM` | |
| `VMXORMM` | |
| `VNCLIPUWI` | |
| `VNCLIPUWV` | |
| `VNCLIPUWX` | |
| `VNCLIPWI` | |
| `VNCLIPWV` | |
| `VNCLIPWX` | |
| `VNCVTXXW` | |
| `VNEGV` | |
| `VNMSACVV` | |
| `VNMSACVX` | |
| `VNMSUBVV` | |
| `VNMSUBVX` | |
| `VNOTV` | |
| `VNSRAWI` | |
| `VNSRAWV` | |
| `VNSRAWX` | |
| `VNSRLWI` | |
| `VNSRLWV` | |
| `VNSRLWX` | |
| `VORVI` | |
| `VORVV` | |
| `VORVX` | |
| `VREDANDVS` | |
| `VREDMAXUVS` | |
| `VREDMAXVS` | |
| `VREDMINUVS` | |
| `VREDMINVS` | |
| `VREDORVS` | |
| `VREDSUMVS` | |
| `VREDXORVS` | |
| `VREMUVV` | |
| `VREMUVX` | |
| `VREMVV` | |
| `VREMVX` | |
| `VRGATHEREI16VV` | |
| `VRGATHERVI` | |
| `VRGATHERVV` | |
| `VRGATHERVX` | |
| `VRSUBVI` | |
| `VRSUBVX` | |
| `VS1RV` | |
| `VS2RV` | |
| `VS4RV` | |
| `VS8RV` | |
| `VSADDUVI` | |
| `VSADDUVV` | |
| `VSADDUVX` | |
| `VSADDVI` | |
| `VSADDVV` | |
| `VSADDVX` | |
| `VSBCVVM` | |
| `VSBCVXM` | |
| `VSE16V` | |
| `VSE32V` | |
| `VSE64V` | |
| `VSE8V` | |
| `VSETIVLI` | |
| `VSETVL` | |
| `VSETVLI` | |
| `VSEXTVF2` | |
| `VSEXTVF4` | |
| `VSEXTVF8` | |
| `VSLIDE1DOWNVX` | |
| `VSLIDE1UPVX` | |
| `VSLIDEDOWNVI` | |
| `VSLIDEDOWNVX` | |
| `VSLIDEUPVI` | |
| `VSLIDEUPVX` | |
| `VSLLVI` | |
| `VSLLVV` | |
| `VSLLVX` | |
| `VSMULVV` | |
| `VSMULVX` | |
| `VSMV` | |
| `VSOXEI16V` | |
| `VSOXEI32V` | |
| `VSOXEI64V` | |
| `VSOXEI8V` | |
| `VSOXSEG2EI16V` | |
| `VSOXSEG2EI32V` | |
| `VSOXSEG2EI64V` | |
| `VSOXSEG2EI8V` | |
| `VSOXSEG3EI16V` | |
| `VSOXSEG3EI32V` | |
| `VSOXSEG3EI64V` | |
| `VSOXSEG3EI8V` | |
| `VSOXSEG4EI16V` | |
| `VSOXSEG4EI32V` | |
| `VSOXSEG4EI64V` | |
| `VSOXSEG4EI8V` | |
| `VSOXSEG5EI16V` | |
| `VSOXSEG5EI32V` | |
| `VSOXSEG5EI64V` | |
| `VSOXSEG5EI8V` | |
| `VSOXSEG6EI16V` | |
| `VSOXSEG6EI32V` | |
| `VSOXSEG6EI64V` | |
| `VSOXSEG6EI8V` | |
| `VSOXSEG7EI16V` | |
| `VSOXSEG7EI32V` | |
| `VSOXSEG7EI64V` | |
| `VSOXSEG7EI8V` | |
| `VSOXSEG8EI16V` | |
| `VSOXSEG8EI32V` | |
| `VSOXSEG8EI64V` | |
| `VSOXSEG8EI8V` | |
| `VSRAVI` | |
| `VSRAVV` | |
| `VSRAVX` | |
| `VSRLVI` | |
| `VSRLVV` | |
| `VSRLVX` | |
| `VSSE16V` | |
| `VSSE32V` | |
| `VSSE64V` | |
| `VSSE8V` | |
| `VSSEG2E16V` | |
| `VSSEG2E32V` | |
| `VSSEG2E64V` | |
| `VSSEG2E8V` | |
| `VSSEG3E16V` | |
| `VSSEG3E32V` | |
| `VSSEG3E64V` | |
| `VSSEG3E8V` | |
| `VSSEG4E16V` | |
| `VSSEG4E32V` | |
| `VSSEG4E64V` | |
| `VSSEG4E8V` | |
| `VSSEG5E16V` | |
| `VSSEG5E32V` | |
| `VSSEG5E64V` | |
| `VSSEG5E8V` | |
| `VSSEG6E16V` | |
| `VSSEG6E32V` | |
| `VSSEG6E64V` | |
| `VSSEG6E8V` | |
| `VSSEG7E16V` | |
| `VSSEG7E32V` | |
| `VSSEG7E64V` | |
| `VSSEG7E8V` | |
| `VSSEG8E16V` | |
| `VSSEG8E32V` | |
| `VSSEG8E64V` | |
| `VSSEG8E8V` | |
| `VSSRAVI` | |
| `VSSRAVV` | |
| `VSSRAVX` | |
| `VSSRLVI` | |
| `VSSRLVV` | |
| `VSSRLVX` | |
| `VSSSEG2E16V` | |
| `VSSSEG2E32V` | |
| `VSSSEG2E64V` | |
| `VSSSEG2E8V` | |
| `VSSSEG3E16V` | |
| `VSSSEG3E32V` | |
| `VSSSEG3E64V` | |
| `VSSSEG3E8V` | |
| `VSSSEG4E16V` | |
| `VSSSEG4E32V` | |
| `VSSSEG4E64V` | |
| `VSSSEG4E8V` | |
| `VSSSEG5E16V` | |
| `VSSSEG5E32V` | |
| `VSSSEG5E64V` | |
| `VSSSEG5E8V` | |
| `VSSSEG6E16V` | |
| `VSSSEG6E32V` | |
| `VSSSEG6E64V` | |
| `VSSSEG6E8V` | |
| `VSSSEG7E16V` | |
| `VSSSEG7E32V` | |
| `VSSSEG7E64V` | |
| `VSSSEG7E8V` | |
| `VSSSEG8E16V` | |
| `VSSSEG8E32V` | |
| `VSSSEG8E64V` | |
| `VSSSEG8E8V` | |
| `VSSUBUVV` | |
| `VSSUBUVX` | |
| `VSSUBVV` | |
| `VSSUBVX` | |
| `VSUBVV` | |
| `VSUBVX` | |
| `VSUXEI16V` | |
| `VSUXEI32V` | |
| `VSUXEI64V` | |
| `VSUXEI8V` | |
| `VSUXSEG2EI16V` | |
| `VSUXSEG2EI32V` | |
| `VSUXSEG2EI64V` | |
| `VSUXSEG2EI8V` | |
| `VSUXSEG3EI16V` | |
| `VSUXSEG3EI32V` | |
| `VSUXSEG3EI64V` | |
| `VSUXSEG3EI8V` | |
| `VSUXSEG4EI16V` | |
| `VSUXSEG4EI32V` | |
| `VSUXSEG4EI64V` | |
| `VSUXSEG4EI8V` | |
| `VSUXSEG5EI16V` | |
| `VSUXSEG5EI32V` | |
| `VSUXSEG5EI64V` | |
| `VSUXSEG5EI8V` | |
| `VSUXSEG6EI16V` | |
| `VSUXSEG6EI32V` | |
| `VSUXSEG6EI64V` | |
| `VSUXSEG6EI8V` | |
| `VSUXSEG7EI16V` | |
| `VSUXSEG7EI32V` | |
| `VSUXSEG7EI64V` | |
| `VSUXSEG7EI8V` | |
| `VSUXSEG8EI16V` | |
| `VSUXSEG8EI32V` | |
| `VSUXSEG8EI64V` | |
| `VSUXSEG8EI8V` | |
| `VWADDUVV` | |
| `VWADDUVX` | |
| `VWADDUWV` | |
| `VWADDUWX` | |
| `VWADDVV` | |
| `VWADDVX` | |
| `VWADDWV` | |
| `VWADDWX` | |
| `VWCVTUXXV` | |
| `VWCVTXXV` | |
| `VWMACCSUVV` | |
| `VWMACCSUVX` | |
| `VWMACCUSVX` | |
| `VWMACCUVV` | |
| `VWMACCUVX` | |
| `VWMACCVV` | |
| `VWMACCVX` | |
| `VWMULSUVV` | |
| `VWMULSUVX` | |
| `VWMULUVV` | |
| `VWMULUVX` | |
| `VWMULVV` | |
| `VWMULVX` | |
| `VWREDSUMUVS` | |
| `VWREDSUMVS` | |
| `VWSUBUVV` | |
| `VWSUBUVX` | |
| `VWSUBUWV` | |
| `VWSUBUWX` | |
| `VWSUBVV` | |
| `VWSUBVX` | |
| `VWSUBWV` | |
| `VWSUBWX` | |
| `VXORVI` | |
| `VXORVV` | |
| `VXORVX` | |
| `VZEXTVF2` | |
| `VZEXTVF4` | |
| `VZEXTVF8` | |
| `WFI` | |
| `WORD` | |
| `XNOR` | |
| `XOR` | Bitwise XOR |
| `XORI` | XOR immediate |
| `ZEXTH` | |
Recognised: 975 mnemonics.
+140
View File
@@ -0,0 +1,140 @@
# Language: lexicon, statements and expressions
Layer 1, the common language, the same on every target. Verified against
`go tool asm` of Go 1.27.1 and against gasm's parser, which is differentially
tested against the toolchain. The authoritative sources behind this page are
the assembler's lexer (`cmd/asm/internal/lex`), its parser
(`cmd/asm/internal/asm/parse.go`) and the toolchain's own test data.
## Source files and targets
An assembly source is a `.s` file. The Go build convention names a
target-specific file with the architecture suffix, `_amd64.s`, `_arm64.s`,
`_riscv64.s` or `_loong64.s`; files without a suffix are portable across
targets. The same assembler program assembles every target: `go tool asm`
picks the target from the `GOOS` and `GOARCH` environment variables, and gasm
from the file name suffix or the `--arch` flag.
## Character set and identifiers
Sources are ASCII text. An identifier is a sequence of ASCII letters, digits
and underscores, digits never first, with exactly two additions:
- U+00B7, the middle dot `·`, stands for the period in a symbol's
package-qualified name;
- U+2215, the division slash `∕`, stands for the slash in a package path.
The two substitutions exist because the parser treats a real period and a
real slash as punctuation. The syntax is otherwise uppercase throughout:
instructions, registers and directives are written in upper case. The one
inherited exception is the `g` register name on 32-bit ARM.
## Comments
Two comment forms, both Go's:
```text
// a line comment
/* a block comment */
```
A comment of the form `//go:build` or the legacy `+build` comment is not a
plain comment: the lexer reports it to the build system as a build
constraint.
## Statements
The grammar of one line, from the parser:
```text
{label:} WORD[.qualifier] [ arg {, arg} ] (';' | '\n')
```
- A **label** is an identifier followed by a colon. Labels are
function-local: two functions in one file may reuse the same name, and a
reference resolves within the function that contains it. A branch
instruction names its target with a bare label operand, and the assembler
resolves it PC-relative. The explicit forms `offset(PC)`, a constant
counting instructions from the branch, and `name(SB)`, a cross-function
static reference, appear as branch targets as well.
- **WORD** is the instruction or directive name, upper case. On the ARM
family the word may carry a dot qualifier selecting a condition or shift
mode, such as the condition suffixes on 32-bit ARM; the amd64, arm64,
riscv64 and loong64 assemblies carry no instruction qualifiers apart from
their own width suffixes, which are part of the mnemonic.
- **Arguments** are separated by commas, with no trailing comma.
- A statement ends at a newline or at a semicolon, so several statements fit
on one line separated by `;`. Blank lines are free.
The first word of a line is a directive if it is one of the directive names
(TEXT, DATA, GLOBL, FUNCDATA, PCDATA, PCALIGN) and an instruction otherwise.
Unknown instruction names are errors; the instruction set is the set the
toolchain itself defines per target, plus the common pseudo-instructions.
## Literals
| Form | Examples | Notes |
|---|---|---|
| Integer | `0`, `42`, `0x2a`, `0o52`, `0b101010`, `1_000` | decimal, hexadecimal, octal and binary forms with Go's digit separators |
| Character | `'a'`, `'\n'`, `'\x41'` | single quoted, Go escape rules |
| String | `"this program can only run\n"` | double quoted, Go escape rules; accepted where an operand takes raw bytes, in practice a DATA initialiser |
| Float | `1.5`, `1e9` | accepted by the lexer; only meaningful where the target's encoding takes a float operand |
## Expressions
Constant expressions may appear wherever a constant is expected: in
immediates after `$`, in memory offsets, in frame and data sizes. The
evaluator works on unsigned 64-bit values with Go's operator precedence, and
the parser states its grammar in exactly those terms:
```text
expr = term { '+' term | '-' term | '|' term | '^' term }
term = factor { '*' factor | '/' factor | '%' factor | '<<' factor | '>>' factor | '&' factor }
factor = const | '+' factor | '-' factor | '~' factor | '(' expr ')'
```
Two consequences are worth naming, because the arithmetic surprises people
who read it as C:
- Shifts bind at the multiplicative level, next to `*` and `&`, while `|`
and `^` bind at the additive level. `$x<<1|3` computes `(x<<1)|3`, which
differs from `x*2+3` whenever `x` is odd. Plan 9 arithmetic is Go
precedence applied to a byte-oriented language, not the C expression it
resembles.
- The evaluator is unsigned and guarded: division or modulo by zero is an
error, and so is dividing a value with the high bit set; shift counts must
be non-negative; and a right shift of a value with the high bit set is
rejected rather than sign-extended.
An address expression such as `(index*4)(base)` is evaluated at assembly
time only if every name in it is a constant; a name that resolves to a
symbol turns the expression into a relocation request, never into a folded
constant.
Named constants enter expressions through the preprocessor (`#define`,
`-D`) and, in Go-embedded packages, through the generated `go_asm.h`; see
PREPROCESSOR.md and RUNTIME.md.
## The common pseudo-instructions
A handful of instructions exist on every target, assembled by the assembler
itself rather than the encoder: `NOP`, which emits the target's no-operation
encoding, and the frame-management pseudo-instructions the compiler emits
(`FUNCDATA`, `PCDATA`) which DIRECTIVES.md specifies. Everything else is the
target's own instruction set, and the assembler knows only the instructions
the toolchain's compiler emits; a hand-written kernel wanting more lays the
encoding down with `BYTE` on amd64 or waits for the extended layer.
## Case study: three lines, decomposed
```text
B.EQ 1(PC) // arm64: condition qualifier on the mnemonic,
// target one instruction past the branch
JMP done // every target: bare label, function-local,
// resolved PC-relative
MOVQ $reader__size>>3, CX // amd64: expression over a go_asm.h constant
```
The first shows a qualifier and the explicit relative target form; the second
the ordinary label reference; the third an expression over a generated
constant. Labels are reusable between functions without conflict.
+94
View File
@@ -0,0 +1,94 @@
# LoongArch 64
Layer 1, target page. Verified against `go tool asm` of Go 1.27.1, against
the toolchain's own loong64 assembler manual (`cmd/internal/obj/loong64/doc.go`)
and against gasm's encoder, whose output is compared byte for byte with the
toolchain's. The complete mnemonic inventory lives in the generated appendix
[INSTRUCTIONS-LOONG64.md](INSTRUCTIONS-LOONG64.md).
## Registers
- General purpose `R0` to `R31`, floating point `F0` to `F31`, LSX vectors
`V0` to `V31` and LASX vectors `X0` to `X31`.
- Fixed roles from the toolchain's table: `R0` is the constant zero, `R1`
the return address, `R3` the stack pointer, `R22` the goroutine pointer,
`R29` the closure context and `R30` the assembler's temporary. `R12`,
`R13`, `R14`, `R15` and `R20` serve the PLT and trampoline sequences:
usable in assembly, but saved before any call.
## Widths ride the mnemonic
| Suffix | Width |
|---|---|
| `B`, `BU` | 8-bit, 8-bit unsigned |
| `H`, `HU` | 16-bit, 16-bit unsigned |
| `W`, `WU` | 32-bit, 32-bit unsigned |
| `V` | 64-bit |
| `F`, `D` | 32-bit and 64-bit float |
| `V` prefix (LSX) | 128-bit vector |
| `XV` prefix (LASX) | 256-bit vector |
The MOV series is the load and store interface: `MOVB (R2), R3` loads a
byte, `MOVV (R2), R3` a double word, `VMOVQ (R2), V1` a 128-bit vector and
`XVMOVQ (R2), X1` a 256-bit one.
## Operand order
Most instructions appear in left-to-right assignment order: `ADDV R11, R12,
R13` is `add.d R13, R12, R11`, and the two-operand form
`OR R5, R6` assigns into R6. Exceptions:
- Jump and branch instructions keep the GNU order: `BEQ R0, R4, label1`.
- The bitfield family is `BSTRINSW`, `BSTRINSV`, `BSTRPICKW`, `BSTRPICKV`
`$<msb>, <Rj>, $<lsb>, <Rd>`.
## Addressing
- Plain: `offset(Rbase)`.
- Base plus offset **register**, no scale: `(R4)(R5)`, as in
`MOVB (R4)(R5), R6`, the `ldx` family.
- The pointer loads and stores `MOVWP` and `MOVVP` take a source-level
16-bit offset that the encoder halves into the 14-bit field, writing
`MOVWP 8(R4), R5` as `ldptr.w r5, r4, $2`.
## Vector element syntax
The `VMOVQ` and `XVMOVQ` transfer family covers register-to-vector moves
with arrangement and index suffixes: `VMOVQ Rj, Vd.B[index]` inserts a
general register into one lane, `VMOVQ Vj.B[index], Rd` extracts one,
`VMOVQ Rj, Vd.B16` broadcasts across all sixteen, and `VMOVQ Vj.B[index],
Vd.B16` replicates one lane. The broadcast-from-memory form takes the true
byte offset at source level, which the encoder rescales per arrangement.
The permute and extract families take their 8-bit control word first:
`VPERMIW ui8, Vj, Vd`, `VEXTRINSB ui8, Vj, Vd`.
## Alignment
`PCALIGN $n` pads with NOOP to a power-of-two boundary between 8 and 2048,
and this target additionally auto-aligns loop heads to 16 bytes.
## Atomics, barriers and prefetch
- The `AM` atomic family comes in plain and `_DB` flavours; the `_DB`
forms, such as `AMSWAPDBW`, complete the atomic sequence and act as a
full data barrier. Within the AM family the destination and base
registers may not coincide and the destination may not equal the operand
register: one is an exception, the other silently unspecified.
- `DBAR` carries the graded hint encoding documented for LA664 and later,
with hint 0x700 as the read-after-read lightweight barrier; older cores
treat every hint as the full barrier.
- `PRELD offset(Rbase), $hint` prefetches with the documented hints (0
load to L1, 2 load to L3, 8 store to L1); `PRELDX` adds the encoded
block descriptor.
- `ALSL`-family shift-and-add writes the desired shift amount in source and
encodes one less: `ALSLV $4, R4, R5, R6` shifts by 4.
- `ADDV16 si16<<16, Rj, Rd` is the high-immediate add paired with the
pointer loads for GOT relative access.
## Relocations
`R_CALLLOONG64` for the 28-bit BL, `R_LOONG64_CALL36` for the
PCADDU18I-plus-JIRL pair, the `R_LOONG64_ADDR`, `ADDR64`, `TLS_LE`, `TLS_IE`,
`GOT` and `GOT64` high and low pairs, the aligned conditional jump forms
`R_JMP16LOONG64` and `R_JMP21LOONG64`, and `R_LOONG64_ADD64` and `SUB64`
for in-place arithmetic, all specified in [GOOBJ.md](../GOOBJ.md).
+114
View File
@@ -0,0 +1,114 @@
# Operands: grammar, pseudo-registers, addressing and symbols
Layer 1, the common language. Verified against `go tool asm` of Go 1.27.1 and
against gasm's parser. The operand grammar is the part of the language that
varies most between targets, so this page fixes the common grammar and the
pseudo-registers; the per architecture pages carry the register names and the
addressing quirks each target adds.
## The four operand kinds
Every operand is one of four kinds:
```text
R1 register
$4 immediate
label branch target or symbol
-8(BX)(DI*4) memory
```
**Operands go source first, destination last**: `MOVQ x+0(FP), AX` loads the
argument into AX. This is the opposite of Intel order and the same order as
AT&T, with the sigils removed: registers are bare names, immediates take
`$`, memory is `offset(base)`.
## Registers
A register operand is its bare name, with no prefix: `AX`, `X15`, `R14` on
amd64; `R0` to `R30`, `ZR`, `V0` to `V31` on arm64; `X0` to `X31`, `F0` to
`F31`, `V0` on riscv64; `R0` to `R31`, `F0` to `F31`, `V0` on loong64.
Sub-register and width selection rides the mnemonic, not the operand: the
amd64 family spells `MOVB`, `MOVW`, `MOVL`, `MOVQ`, and the arm64 family
suffices `B`, `H`, `S`, `D`, `Q` on the shared forms. Each architecture page
lists its registers and the reserved ones.
## Immediates
`$` introduces a constant: `$42`, `$-1`, `$0x2a`, `$'A'`, `$bufSize`. The
`$` applies to the whole constant expression that follows, so
`$(4*8+reader__size)` is one immediate. Without the `$`, a number in operand
position is an address, not a value; the classic error `ADDQ 1, AX` asks the
assembler for the byte at address 1.
The one place a `$` number is not an immediate is the frame and argument
size field of TEXT, `$16-24`, which is two separate constants and not a
subtraction; DIRECTIVES.md specifies it.
## Memory
```text
offset(base)
offset(base)(index*scale)
```
Both parts are optional where the target allows them: `(BX)` is the memory
at BX, `foo+16(SB)` is a global, and on amd64 `foo+32(SP)(R9*8)` adds a
scaled index. `offset` is a constant expression, optionally carrying a
symbol name. The extensions beyond `offset(base)` are where the targets
diverge, and each belongs to its architecture page: amd64 carries the
`index*scale` form with scale 1, 2, 4 or 8 and its own rules on which
registers may index; loong64 writes base plus index as `(R4)(R5)`; the ARM
family attaches shift amounts to the index register in its own spelling.
The address arithmetic is on **byte addresses**: the offset is added to the
base as it stands, whatever the operand width of the instruction. Loading
the third 8-byte word of an array at BX is `16(BX)`, not `2(BX)`.
## The four pseudo-registers
Four names denote locations no target register holds, and they mean the same
on every architecture:
- **FP**, the frame pointer: the arguments and results of the current
function, at positive offsets, in the order the Go prototype declares
them. Every FP reference must carry a name: `x+0(FP)`, and an unnamed
`0(FP)` is rejected. Results follow arguments; an unnamed result is called
`ret`.
- **SP**, the virtual stack pointer: the high end of the function's local
frame, so locals live at negative offsets, `x-8(SP)`. A reference without
a name and without a plus, `-8(SP)`, addresses the **hardware** stack
pointer instead: the two spellings are one character apart and mean
different registers. That is the sharpest edge in the language and the
source of the deepest bugs.
- **SB**, the static base: the origin of memory, used for globals and
cross-package symbols, always with a name: `foo(SB)`, `foo+4(SB)`.
- **PC**, the program counter: branch targets, and the explicit relative
form `1(PC)`.
## Symbol names
A symbol's full name is the package path, a period, and the base name. In
source, the period is written U+00B7 (`·`) and a slash in the path U+2215
(`∕`), because the parser treats the ASCII forms as punctuation. Inside the
package's own file, `·Name` is enough and is the preferred spelling, since
it survives a rename of the import path.
| Spelling | Meaning |
|---|---|
| `·Name(SB)` | this package's Name |
| `runtime·morestack(SB)` | another package's morestack |
| `sourcedock.dev∕petrbalvin∕pkg·Name(SB)` | fully qualified |
| `msg<>(SB)` | file-local, the static of this language; `<>` also makes the ABI field static in the object |
| `Name<ABIInternal>(SB)` | ABI-qualified reference, the ABI in angle brackets after the name |
The object file these symbols produce, with the index rules that decide what
is referenced by name and what by index, is specified in
[GOOBJ.md](../GOOBJ.md).
## What vet adds in Go
Inside a Go package, `go vet`'s asmdecl analyzer checks every FP offset and
name against the Go prototype, and checks the declared argument area against
the frame. That layer, the prototype requirement and `go_asm.h`, belongs to
RUNTIME.md; the grammar above is the whole of what the assembler itself
requires.
+79
View File
@@ -0,0 +1,79 @@
# Preprocessing: include, define and selection
Layer 1, the common language. Verified against the preprocessor inside
`go tool asm` of Go 1.27.1 (`cmd/asm/internal/lex`), whose directives are
`#define`, `#undef`, `#include`, `#ifdef`, `#ifndef`, `#else`, `#endif` and
`#line`, and against gasm's implementation, which is differentially tested
against the toolchain's.
Input runs through a simplified C preprocessor before the parser sees it.
The set is deliberately small: there is no `#if` with constant expressions
and no token pasting with `##`. `#line` is honoured, so it changes the
positions the assembler reports and records.
## #include
```text
#include "textflag.h"
#include "go_asm.h"
#include "defs_linux_amd64.h"
```
The search path, in order: the directory of the including file, then the
directories given by repeatable `-I` flags. The assembler seeds no default
of its own: a bare `go tool asm` invocation finds none of the standard
headers, and it is the `go` build system that passes `$GOROOT/pkg/include`
among the `-I` directories when it drives the build. That directory ships
`textflag.h`, `funcdata.h` and the per architecture register headers.
Includes nest; a file included twice through different paths is processed
twice, which is why headers guard their defines.
## #define and #undef
```text
#define bufSize 1024
#define MOVD(d, s) MOVQ s, d
#undef bufSize
```
- An object macro replaces its name with its token sequence at the point of
use.
- A parameterised macro takes its arguments in parentheses and substitutes
them into the body. Macro parameters compose with the rest of the
language: an argument used with an element suffix, as in `A.S4` on the
vector forms, substitutes correctly.
- Redefinition is an error; `#undef` first, or pick a new name.
- The `-D name[=value]` flag predefines an object macro from the command
line, repeatable, exactly as `#define` would; a `-D` without a value
defines the name as `1`.
- Expansion happens when the name is used, so a macro may expand to
instructions, operands or fragments of either, and a macro body may use
macros defined before it.
`textflag.h` and `funcdata.h` are themselves ordinary `#define` files: the
flag names and the runtime macros are preprocessor definitions, not language
keywords. That is why a missing include produces a parser error at the first
use of `NOSPLIT` rather than a complaint about the name.
## #ifdef, #ifndef, #else, #endif
```text
#ifdef GOOS_windows
#define SYSCALL_INT 0x2b
#endif
```
Selection is by defined-name only: `#ifdef`, `#ifndef`, `#else`, `#endif`,
nesting freely. There is no `#if defined(x) && y`, because the preprocessor
evaluates no expressions; reach that with a build-tag Go file generating a
header, which is exactly how the runtime's own `go_asm.h` and defs headers
are produced.
## What preprocessing does not cover
The preprocessor is textual and runs first, so it knows nothing of assembly
semantics: it does not check that a macro expansion is a legal instruction,
and it does not participate in the constant expression evaluator, which runs
later, in the parser. A constant folded with `#define` and a constant folded
in an operand expression end at the same value through different doors;
GOOBJ.md records both in the object identically.
+56
View File
@@ -0,0 +1,56 @@
# The Plan 9 assembly language
This directory is the reference for the Plan 9 assembly language as the Go
toolchain and gasm accept it, written to be complete enough to implement
against. It exists because no such reference exists upstream: Go documents
the language on a single page, and the rest of the knowledge lives in the
toolchain's source and in the practice of reading it.
Every page carries the same conformance statement: which layer of the system
it describes, which toolchain release it was verified against, and how the
claims were checked. Pages in this directory are verified against Go 1.27.1
and against gasm's own differential test suite, which compares gasm's
behaviour with `go tool asm` byte for byte and output for output.
## The three layers
The reference deliberately separates three layers, because their rules have
different owners and different lifetimes:
1. **The common language** (LANGUAGE, OPERANDS, DIRECTIVES,
PREPROCESSOR): the syntax, operands, directives and preprocessing, the
same on every target and meaningful without a Go runtime.
2. **The Go-embedded layer** (RUNTIME): everything that exists only because
the code runs inside a Go program: the ABI0 contract, generated wrappers,
`go_asm.h`, the garbage collector annotations and `go vet` checks.
3. **The standalone layer** (STANDALONE, planned with the standalone
compilation phase): using the language outside Go, through gasm's ELF
output and the extended instruction set, where the toolchain offers no
ground truth and execution testing is the only verification.
A rule stated in layer 1 holds on every target. A rule stated in layer 2
says which part of the Go machinery imposes it. Nothing in layer 3 changes
layers 1 or 2; it extends them.
## Pages
| Page | Layer | Contents |
|---|---|---|
| [LANGUAGE.md](LANGUAGE.md) | 1 | lexicon, statement structure, labels, literals, expressions |
| [OPERANDS.md](OPERANDS.md) | 1 | operand grammar, pseudo-registers, addressing modes, symbol naming |
| [DIRECTIVES.md](DIRECTIVES.md) | 1 | TEXT, DATA, GLOBL, FUNCDATA, PCDATA, PCALIGN and the function flags |
| [PREPROCESSOR.md](PREPROCESSOR.md) | 1 | `#include`, `#define`, `#ifdef` and friends, `-D`, `-I` |
| [RUNTIME.md](RUNTIME.md) | 2 | ABI0, prototypes, `go_asm.h`, `funcdata.h`, `go vet` |
| [AMD64.md](AMD64.md) | 1 | registers, addressing, the frame and split check, families, relocations |
| [ARM64.md](ARM64.md) | 1 | registers, the MOV load and store series, special operand orders, SIMD |
| [RISCV64.md](RISCV64.md) | 1 | registers and their constrained names, per class operand order, profiles, vector extension |
| [LOONG64.md](LOONG64.md) | 1 | registers, width suffixes, vector element syntax, atomics and barriers |
| INSTRUCTIONS-AMD64.md and the other three | 1 | generated per architecture inventory of every accepted mnemonic |
| STANDALONE.md | 3 | the language outside Go |
## Status
The common-language core, the Go-embedded layer, all four per-architecture
pages and the generated instruction appendices are written and verified.
STANDALONE.md lands with the standalone compilation phase. The object format
these pages feed is specified in [GOOBJ.md](../GOOBJ.md).
+104
View File
@@ -0,0 +1,104 @@
# RISC-V 64
Layer 1, target page. Verified against `go tool asm` of Go 1.27.1, against
the toolchain's own riscv64 assembler manual (`cmd/internal/obj/riscv/doc.go`)
and against gasm's encoder, whose output is compared byte for byte with the
toolchain's. The complete mnemonic inventory lives in the generated appendix
[INSTRUCTIONS-RISCV64.md](INSTRUCTIONS-RISCV64.md).
## Registers
- Integer: `X0` to `X31`. `X0` is hardwired zero. Three names the toolchain
constrains: `X4` must be written through its ABI name `TP`; `X27`, the
goroutine pointer, must be written `g` and may not be written `S11`; in
shared builds `X3` is off limits and must be written `GP`.
- The other integer registers may be written `Xn` or by their ABI names
(`A0`, `T0`, `S1`, and so on).
- Floating point: `F0` to `F31`. Vector: `V0` to `V31`.
- `X26` is the closure pointer and `X31` is the assembler's own scratch
register: its value may be clobbered by instruction sequences the
assembler inserts, so hand-written code must not rely on it.
- There is no reserved frame pointer register on this target.
## Operand order
The ordering differs from the ISA manual, and per instruction class:
- **R-type** is reversed: `ADD X10, X11, X12` is `add x12, x11, x10`.
- **I-type arithmetic** keeps that shape with the immediate first:
`ADDI $1, X11, X12`.
- **Loads and stores** are source first, like every Plan 9 dialect:
`MOV 16(X2), X10` loads and `MOV X10, (X2)` stores. The MOV series hides
the width; `MOVB` through `MOVD` spell it out.
- **Branches** keep the ISA order: `BLT X12, X23, loop1`, which jumps when
X12 < X23, the reverse of the SLT operand order.
- **FMA** is rotated one place left so the destination comes last:
`FMADDS F1, F2, F3, F4`.
- **AMO** is likewise rotated: `AMOSWAPW X5, (X6), X7`.
- **Ternary abbreviation** is supported and encouraged: `ADD X10, X12` means
`ADD X10, X12, X12`.
Where an R-type instruction has an I-type sibling, the assembler picks the
immediate form from the operand: `AND $3, X12, X13` assembles as `ANDI`.
## Names, suffixes and rounding
Dots are removed and suffixes are upper-cased: the ISA's `fmv.w.x` is
`FMVWX`. Floating-point rounding modes become suffixes, `FCVTLUS.RNE F0,
X5`, with RTZ assumed when the suffix is omitted; the toolchain never sets
the FCSR.
## Constants
- `MOV` materialises any 64-bit integer constant, synthesising it from a
few arithmetic instructions where possible and otherwise loading it from
a literal pool in the binary.
- A 32-bit constant is accepted by `ADDI`, `ANDI`, `ORI` and `XORI`, and
the assembler synthesises values that exceed the 12-bit encoding window.
- `MOVF` and `MOVD` materialise floating-point constants, encoding them as
`FLW` and `FLD` from a pool location unless the constant is exactly 0.0.
## Extensions and profiles
The default target profile is rva20u64, selected or raised with the
GORISCV64 environment variable. A short list of instructions outside the
default profile is synthesised by the assembler when the profile does not
provide them, so they are safe without guards: `ANDN`, `MAX`, `MAXU`, `MIN`,
`MINU`, `MOVB`, `MOVH`, `MOVHU`, `MOVWU`, `ORN`, `ROL`, `ROLW`, `ROR`,
`RORI`, `RORIW`, `RORW`, `XNOR`. The header `asm_riscv64.h` defines the
`hasZba`, `hasZbb`, `hasZbs` and `hasV` macros for guarding everything else.
## Fences and atomics
`FENCE` takes predecessor and successor sets in that order, uppercase
letters, `FENCE R, RW`; a bare `FENCE` is a full fence, as is
`FENCE IORW, IORW`. `FENCE.TSO` exists. The ordering bits of `LR`, `SC`
and the AMO instructions are not specifiable in source: the assembler sets
acquire and release on the AMO instructions, acquire on `LR` and release on
`SC`, always.
## Compressed instructions
The assembler converts 32-bit instructions to their compressed encodings
automatically; the conversion is a property of the emitted machine code, not
of the source, and register choice influences how much compresses.
Hand-writing compressed instructions in source is accepted but discouraged.
The debug flag `compressinstructions=0` turns the automatic conversion off.
## Vector extension
`VSETVLI` writes its vtype components in uppercase with the destination
last: `VSETVLI X10, E8, M1, TU, MU, X12`. Vector loads and stores are
source first like the scalar ones, with an optional stride or index register
second and the mask register, when present, always penultimate:
`VLE8V (X10), V3`, `VLE8V (X10), V0, V3` for the masked form. Vector
arithmetic reverses its operands, `VADDVV V1, V2, V3`, with the mask again
penultimate.
## Relocations
`R_RISCV_JAL`, `R_RISCV_CALL`, the `R_RISCV_PCREL_ITYPE` and `STYPE` pairs,
`R_RISCV_BRANCH`, the compressed branch and jump forms, the TLS and GOT
families and `R_RISCV_ADD32` and `SUB32`, all specified in
[GOOBJ.md](../GOOBJ.md). The assembler always emits the four-byte
`R_DWTXTADDR_U4` flavour inside its DWARF records.
+120
View File
@@ -0,0 +1,120 @@
# The Go-embedded layer: ABI0, prototypes and the runtime contract
Layer 2: everything that exists only because the assembly runs inside a Go
program. Without a Go runtime this page does not apply; the language of
OPERANDS.md and DIRECTIVES.md still does. Verified against Go 1.27.1, against
the shipped `funcdata.h` header, and against the object files the toolchain
produces, which were parsed and checked field by field while writing
[GOOBJ.md](../GOOBJ.md).
## Hand-written assembly is ABI0
Go functions compiled from source use ABIInternal, the register-based
calling convention, which the toolchain documents as unstable and free to
change between releases. A `.s` function is written against ABI0, the stack
based convention: arguments and results live in the caller's frame at
positive FP offsets, byte-addressed, in declaration order, with no registers
assigned at all. The toolchain generates the wrapper that translates between
the two; a caller in Go calling an assembly function goes through it, and it
is marked `ABIWRAPPER` in the object. Hand-writing a bridge is never needed
and never correct.
## Every assembly function carries a Go prototype
```go
package add
func Add(x, y int64) int64
```
The body-less declaration is not optional, and not only for the linker: it
is what tells the garbage collector which arguments and results hold
pointers, and what `go vet` checks the assembly against. Even a function
nothing in Go calls gets one. Consequences:
- The FP operand names and offsets are checked by vet's asmdecl analyzer
against the prototype: `x+0(FP)` must name an argument that exists, at the
offset the prototype says. A file that assembles and links can still fail
vet.
- The declared argument area in `$framesize-argsize` is checked against the
prototype's size. An omitted argsize marks the argument size unknown
(0x80000000 in the object, the value of `ArgsSizeUnknown` from
`funcdata.h`), which is the normal spelling for functions with no Go
callers.
- `//go:noescape` on the declaration tells the compiler that a pointer
argument does not escape, for assembly that keeps the pointer beyond the
call.
## The frame, the stack and the collector
The runtime owns the stack and the pointer map, and assembly must hold up
its end of four rules:
1. **Arguments are initialised on entry; results are not.** A function whose
results hold live pointers across a call must zero them and then execute
`GO_RESULTS_INITIALIZED`. Designing functions that return no pointers
avoids the problem.
2. **A frame with calls and no local pointers says so** with
`NO_LOCAL_POINTERS`. A frame with local pointers that the runtime cannot
see is not allowed at all: assembly cannot describe a pointer-containing
local, so it must not have one. Data symbols containing pointers are the
same: define them in Go.
3. **The stack may move.** Stack growth copies the frame, so no pointer into
the frame may be held across a call, and the raw hardware SP register may
not be cached across a call either.
4. **The split check is not optional by default.** Without NOSPLIT, the
assembler inserts the stack-growth preamble, including the morestack
block for framed functions; NOSPLIT is a contract that the frame and
everything below it fit in the remaining stack segment. On amd64 the
assembler also marks small leaf functions NoSplit itself and skips the
preamble, so silence is not a promise.
The simplest safe shape is a leaf function with no local frame and no calls:
it needs no annotation beyond the prototype.
## go_asm.h: Go constants and layout in assembly
A package with `.s` files gets a generated header. Include it and use the
generated names instead of hard-coding layouts, which lie silently when the
Go side changes:
| Go declaration | Assembly name |
|---|---|
| `const bufSize = 1024` | `const_bufSize` |
| field `r` of `type reader struct` | `reader_r` |
| size of `type reader struct` | `reader__size` |
The constants arrive as macros, usable as immediates and offsets, computed
from the Go declarations. An ambiguous name, such as a struct that really
has a `_size` field, fails the generation with a redefinition error.
## funcdata.h: the runtime macros
`$GOROOT/pkg/include/funcdata.h` defines the PCDATA and FUNCDATA ids and the
three macros assembly normally uses instead:
| Macro | Expands to | Meaning |
|---|---|---|
| `GO_ARGS` | `FUNCDATA $FUNCDATA_ArgsPointerMaps, go_args_stackmap(SB)` | the Go prototype defines the argument pointer map |
| `GO_RESULTS_INITIALIZED` | `PCDATA $PCDATA_StackMapIndex, $1` | results are initialised; treat them as live from here |
| `NO_LOCAL_POINTERS` | `FUNCDATA $FUNCDATA_LocalsPointerMaps, no_pointers_stackmap(SB)` | the frame holds no pointers |
`GO_ARGS` is inserted implicitly by the assembler for any function whose
package-qualified name belongs to the current package, which is why most
assembly never writes it. `NOSPLIT` leaf functions that call nothing need
none of the three.
The underlying ids, for reading toolchain output rather than for writing
source: FUNCDATA 0 to 7 are args pointer maps, locals pointer maps, stack
objects, inline tree, open-coded defer info, argument info, argument
liveness and wrap info; PCDATA 0 to 4 are unsafe point, stack map index,
inline tree index, argument liveness index and panic bounds.
## What the runtime does with all of this
The object file records the annotations as aux symbols and FuncInfo records;
GOOBJ.md specifies the encoding. The linker assembles them into the runtime's
pclntable, which traceback and the collector consume. An assembly function
that misdeclares its frame is not a compile error and usually not a link
error: it is a wrong collector decision or a wrong traceback at runtime,
which is why the annotations are a contract and not documentation.
+8 -2
View File
@@ -1,8 +1,8 @@
.TH GASM-AUDIT-INSTRUCTIONS 1 "2026-09-19" "gasm" "User Commands"
.TH GASM-AUDIT-INSTRUCTIONS 1 "2026-09-21" "gasm" "User Commands"
.SH NAME
gasm-audit-instructions \- diff the encoder against the Go toolchain, or measure a corpus
.SH SYNOPSIS
.B gasm audit\-instructions [\-\-corpus [\fIdir\fR]] [\-I dir] [amd64|arm64|riscv64|loong64]
.B gasm audit\-instructions [\-\-corpus [\fIdir\fR]] [\-\-list] [\-I dir] [amd64|arm64|riscv64|loong64]
.SH DESCRIPTION
Compare the gasm encoder for the given architecture (default amd64)
against
@@ -39,6 +39,12 @@ second.
Assemble a corpus of .s files and report pass rates and failure
reasons.
.TP
.B \-\-list
With
.BR \-\-corpus ,
print every failing file with its failure reason, per architecture,
instead of one representative file per reason.
.TP
.B \-I \fIdir\fR
Directory to search for #include files; may be repeated, searched in
order after the source directory. A corpus run whose files include
+31
View File
@@ -489,6 +489,17 @@ func parseImmediate(g []token.Token) ast.Immediate {
} else if g[i].Kind == token.Plus {
i++
}
// A constant expression after the sign: $-(R - 8), $+(32-shift). The
// toolchain folds the negated value in place (the cgo ABI macros write
// ADJSP $-(REGS_HOST_TO_ABI0_STACK - 8)), so the sign applies to the
// folded value exactly as it does to a bare literal.
if i < len(g) && (g[i].Kind == token.LParen || g[i].Kind == token.Tilde) {
if v, rest, ok := foldExpr(g[i:]); ok && len(rest) == 0 {
imm.Val = v
imm.HasVal = true
return imm
}
}
if i < len(g) && g[i].Kind == token.Number {
text := g[i].Text
if v, ok := tryInt(text); ok {
@@ -673,6 +684,26 @@ func parseAddress(g []token.Token) ast.Address {
if i > 0 && i < len(g) {
addr.Shift = joinRaw(g[i:])
}
// A lone (possibly signed) number is an absolute address: MOVL $0xf1,
// 0xf1 stores through the bare displacement with no base at all. In
// operand position a number without $ is an address, never a value.
if addr.Sym == nil && addr.Base == "" && addr.Index == "" && !addr.HasOff {
neg := false
j := 0
if j < len(g) && (g[j].Kind == token.Minus || g[j].Kind == token.Plus) {
neg = g[j].Kind == token.Minus
j++
}
if j == len(g)-1 && g[j].Kind == token.Number {
v := parseInt(g[j].Text)
if neg {
v = -v
}
addr.Offset = v
addr.HasOff = true
return addr
}
}
return addr
}
+9
View File
@@ -34,6 +34,12 @@ type Options struct {
// Expand enables macro expansion, include splicing and the
// statement-separator reading of ';' that the expanded bodies rely on.
Expand bool
// Predefines names the macros defined before the file is read. The
// go command drives go tool asm with -D GOOS_<goos> -D GOARCH_<arch>,
// and GOROOT's own headers (go_tls.h, asm_riscv64.h) select their
// platform blocks with #ifdef on exactly those names, so an assembler
// without them cannot see the platform definitions at all.
Predefines map[string]string
}
// ParseWithOptions parses src like Parse, optionally preprocessing it first.
@@ -44,6 +50,9 @@ func ParseWithOptions(path, src string, opts Options) (*ast.File, []error) {
var errs []error
if opts.Expand {
pp := &preproc{opts: opts, macros: map[string]*macroDef{}}
for name, value := range opts.Predefines {
pp.macros[name] = &macroDef{name: name, body: lexer.Tokenize(value)}
}
lines = pp.fileLines(path, tokens, token.Position{})
errs = pp.errs
} else {
+32
View File
@@ -0,0 +1,32 @@
// The runtime bookkeeping statements: FUNCDATA and PCDATA contribute no
// text bytes on any architecture, and amd64 now matches. They sit between
// real instructions here, with plain, static and offset symbol references
// on the FUNCDATA lines, so the byte counts prove the zero contribution.
#include "textflag.h"
// func bookkeep(x int64) int64
TEXT ·bookkeep(SB), NOSPLIT, $0-16
PCDATA $0, $-1
MOVQ x+0(FP), AX
PCDATA $1, $-2
FUNCDATA $0, args_stackmap(SB)
ADDQ $1, AX
FUNCDATA $5, arginfo0(SB)
PCDATA $1, $3
MOVQ AX, ret+8(FP)
FUNCDATA $1, externalfuncdata(SB)
PCDATA $0, $0
RET
// func bookkeepstatic() int64
TEXT ·bookkeepstatic(SB), NOSPLIT, $0-8
// A static symbol and a defined data symbol as the funcdata target.
// (A symbol+offset target the toolchain itself refuses.)
FUNCDATA $2, fdtable<>(SB)
FUNCDATA $3, undefsym(SB)
MOVQ $7, AX
MOVQ AX, ret+0(FP)
RET
GLOBL fdtable<>(SB), NOPTR, $16
+55
View File
@@ -0,0 +1,55 @@
// Floating-point immediates on the SSE scalar paths: the constant is
// rewritten into a read from a read-only pool symbol ($f64.<hex> or
// $f32.<hex>, the IEEE-754 bits in the name), RIP-relative with the
// displacement left to the relocation. A positive zero on the moves
// collapses to XORPS dst, dst; a negative zero keeps its sign bit and
// takes the pool. The parenthesised $(-1.0) spelling is the one
// math/floor_amd64.s uses. Every result is folded back so no
// instruction is dead.
#include "textflag.h"
// func floatimm(x float64) float64
TEXT ·floatimm(SB), NOSPLIT, $0-16
MOVQ x+0(FP), AX
MOVQ AX, X0
// The floor kernel's sign fold: the parenthesised negative spelling.
MOVSD $ (-1.0), X2
ANDPD X2, X0
// Positive and fractional constants on the scalar moves.
MOVSD $1.5, X3
MOVSD $0.5, X4
MOVSS $2.5, X5
MOVSS $-0.5, X6
// A positive zero collapses to XORPS; a negative zero does not.
MOVSD $0.0, X7
MOVSS $0.0, X8
MOVSD $-0.0, X9
// The scalar arithmetic reads the pool through r/m (hypot's shape).
ADDSD $1.0, X3
SUBSD $0.5, X4
MULSD $-2.5, X4
DIVSD $2.0, X3
ADDSS $0.25, X5
// Fold everything into one double.
ADDSD X5, X3
ADDSD X6, X3
ADDSD X7, X3
ADDSD X8, X3
ADDSD X9, X3
ADDSD X4, X3
ADDSD X0, X3
MOVSD X3, ret+8(FP)
RET
// func floatimmfloat32() float32
TEXT ·floatimmfloat32(SB), NOSPLIT, $0-4
// The single-width pool constants ride the F3 prefix.
MOVSS $1.0, X0
MOVSS $-1.0, X1
MOVSS $0.0, X2
ADDSS $0.5, X0
ADDSS X1, X0
ADDSS X2, X0
MOVSS X0, ret+0(FP)
RET
+52
View File
@@ -0,0 +1,52 @@
// Kernel: the operand forms the GOROOT campaign surfaced — numeric
// PC-relative jumps, symbol-immediate materialisation (the toolchain rewrites
// MOVQ $sym(SB) into a RIP-relative LEA) and the negated constant-expression
// ADJSP the cgo ABI macros write. Bytes are pinned against go tool asm by
// TestDifferentialKernels.
#include "textflag.h"
DATA sd<>(SB)/4, $7
GLOBL sd<>(SB), RODATA, $4
// func Jumps(flag int64) int64
TEXT ·Jumps(SB), NOSPLIT, $0-16
MOVQ flag+0(FP), AX
TESTQ AX, AX
JEQ 2(PC)
MOVQ $1, AX
JMP 3(PC)
MOVQ $2, AX
MOVQ AX, ret+0(FP)
RET
// func SymImm() int64
TEXT ·SymImm(SB), NOSPLIT, $0-16
MOVQ $sd<>(SB), AX
MOVQ $·SymImm(SB), CX
MOVQ AX, ret+0(FP)
RET
// func Frame()
TEXT ·Frame(SB), NOSPLIT, $0
PUSHFQ
CLD
ADJSP $(64 - 8)
ADJSP $-(64 - 8)
POPFQ
RET
// func Tls() int64
TEXT ·Tls(SB), NOSPLIT, $0-8
MOVQ TLS, BX
MOVQ 0(BX)(TLS*1), AX
MOVQ AX, ret+0(FP)
RET
// func Aligned() int64
TEXT ·Aligned(SB), NOSPLIT, $0-8
MOVQ $1, AX
PCALIGN $16
MOVQ $2, AX
PCALIGN $32
MOVQ AX, ret+0(FP)
RET
+44
View File
@@ -0,0 +1,44 @@
// The quad-register instructions: the 4FMAPS family (V4FMADDPS,
// V4FMADDSS, V4FNMADDPS, V4FNMADDSS) and the 4VNNIW pair (VP4DPWSSD,
// VP4DPWSSDS). The bracketed list's low register travels the inverted
// V'VVVV field, the memory source keeps r/m, the opmask rides aaa and the
// vector length follows the destination (512-bit for the ZMM forms,
// 128-bit for the scalar ones) while the disp8xN multiplier stays 16 for
// every member. Every result is folded back so no instruction is dead.
#include "textflag.h"
// func quadf4(src *[16]uint32, n int) float32
TEXT ·quadf4(SB), NOSPLIT, $0-20
MOVQ src+0(FP), SI
MOVQ n+8(FP), CX
// The packed 4-FMA form over four consecutive ZMM accumulators,
// masked with K2, K3 and unmasked alike; the displacements exercise
// the disp32 form and the disp8x16 compressed form.
V4FMADDPS 17(SI), [Z0-Z3], K2, Z0
V4FMADDPS 64(SI), [Z10-Z13], K2, Z1
V4FMADDPS (SI), [Z20-Z23], Z2
V4FNMADDPS 96(SI), [Z1-Z4], K3, Z5
// The scalar form reads XMM lists and takes the 128-bit length; the
// displacement compresses by 16.
V4FMADDSS 7(AX), [X0-X3], K5, X22
V4FMADDSS (DI), [X10-X13], K5, X23
V4FNMADDSS 16(SI), [X20-X23], K1, X24
// The 4-VNNI dot products, indexed source included.
VP4DPWSSD 15(DX)(BX*8), [Z2-Z5], K4, Z17
VP4DPWSSDS -7(DI)(R8*1), [Z4-Z7], K1, Z31
VP4DPWSSD (SI), [Z12-Z15], Z6
// Zeroing keeps the usual rule: a mask register must ride along.
V4FMADDPS.Z 128(SI), [Z24-Z27], K4, Z3
// Fold every accumulator into one scalar.
VPADDD Z0, Z1, Z9
VPADDD Z2, Z5, Z10
VPADDD Z9, Z17, Z11
VPADDD Z10, Z31, Z12
VPADDD Z11, Z12, Z13
VPADDD Z13, Z14, Z15
VADDSS X22, X23, X0
VADDSS X24, X0, X1
VADDSS X1, X2, X3
VMOVSS X3, ret+16(FP)
RET
+6
View File
@@ -124,10 +124,16 @@ func TestGroundTruthAMD64(t *testing.T) {
"../testdata/verify/avx_amd64.s",
"../testdata/verify/pfx_amd64.s",
"../testdata/verify/vsib_amd64.s",
"../testdata/verify/floatimm_amd64.s",
"../testdata/verify/bookkeep_amd64.s",
"../testdata/verify/quadreg_amd64.s",
"../testdata/verify/rawdata_amd64.s",
"../testdata/verify/avx512_amd64.s",
"../testdata/verify/pfx_amd64.s",
"../testdata/verify/vsib_amd64.s",
"../testdata/verify/floatimm_amd64.s",
"../testdata/verify/bookkeep_amd64.s",
"../testdata/verify/quadreg_amd64.s",
"../testdata/verify/rawdata_amd64.s",
"../testdata/verify/avx512_amd64.s",
"../testdata/verify/doubleshift_amd64.s",